datamark vs JSON Schema

Describing document structure with JSON Schema vs. datamark's lightweight format system.

JSON Schema is the standard way to describe structured data. Could you use it to describe a Markdown document format?

The JSON Schema approach

You would define a schema for the output shape:

{
  "type": "object",
  "properties": {
    "title": { "type": "string" },
    "sections": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "heading": { "type": "string" },
          "content": { "type": "string" }
        }
      }
    }
  }
}

But JSON Schema describes output shape, not how to parse. It doesn't tell you:

  • Which heading depth marks a section boundary
  • How to extract code blocks vs. paragraphs
  • How to handle optional frontmatter
  • How to serialize back to Markdown

The datamark approach

datamark combines schema validation with parsing logic:

const PlanFormat = datamark({
  schema: z.object({ title: z.string(), sections: z.array(...) }),
  parse(doc) {
    const h1 = doc.root.children.find(n => n.type === "section");
    const title = h1 ? inlineText(h1.heading.children) : "";
    const sections = h1?.children.filter(n => n.type === "section") ?? [];
    // ...map to schema shape
    return { title, sections };
  },
});
FeatureJSON Schemadatamark
Output validation✅ (via Standard Schema)
Parsing logic✅ Imperative functions
Bidirectional✅ Parse + stringify
Trace/debug
Self-testing✅ Inline examples
Markdown-specific✅ AST utilities

When to use JSON Schema

JSON Schema is excellent for API payloads, configuration files, and any data that doesn't live in Markdown. When your input is a Markdown document, you need parsing logic too — and that's where datamark fits.

On this page