JSONBeautify JSON Linter

NDJSON/JSONL: The Format Every Log Pipeline Secretly Runs On

JSONBeautify — Free Online JSON Formatter, Validator & Minifier Guides · Updated 2026-10-01 · All guides

Newline-delimited JSON is one JSON document per line, no wrapper, no commas between rows. Every serious logging, streaming, and data-engineering system converged on it — not out of taste, but because the alternatives fail at 3 a.m. during a partial write. Verified on jq 1.7 and Node 24.

What it is

{"ts":"2026-10-01T17:02:11Z","level":"info","msg":"cache warm"}
{"ts":"2026-10-01T17:02:12Z","level":"warn","msg":"slow query","ms":842}
{"ts":"2026-10-01T17:02:13Z","level":"error","msg":"conn refused"}

Each line is independently valid JSON; the whole file is not. Any valid JSON value can be a row — objects most often, but arrays and scalars are legal. MIME type is application/x-ndjson (unregistered in practice; SSE-style streaming APIs often ship application/json and pretend).

Why every pipeline uses it

Append-only is crash-proof. Writing an array-JSON file means rewriting the whole file per record — one crash mid-rewrite and the file is garbage. NDJSON appends one line and the file is still valid the instant the newline lands. Kill the process at any byte and you lose at most the half-written last line, which every tool below skips or reports cleanly.

Greppable before parsed. grep 'conn refused' events.jsonl needs no libraries, no schema, no memory. In an incident, grep plus head beats whatever loader the vendor shipped.

Partial reads are trivial. head -n 5 events.jsonl samples a 40 GB file. Array JSON requires parsing the prefix — which requires knowing where the array ends, which requires reading it all.

Streams natively. SSE, LLM token APIs, fetch bodies, tail -f, Kafka sinks, and Spark's default write format are line-oriented because lines have an unambiguous end. A JSON array has no parseable prefix.

Resumable and splittable. split -l 100000 big.jsonl shards cleanly; any line offset is a restart point. Array JSON can only be split at byte 0.

The cost: no top-level metadata (repeat it per row or accept a sidecar), and one corrupted row. Which is a row, not a file.

Pretty-print without loading it all

A 12 GB JSONL file does not fit in this tool's textarea — that limit is real and honest, it is one parse on one browser thread. The server-side route stays constant-memory:

# jq without -s: reads one value per iteration, constant memory
jq -C . events.jsonl | less

# keep it one-line-per-row but pretty *inside* (colored, compact)
jq -cC . events.jsonl

# first 50 rows only — reads 50 lines and stops
head -n 50 events.jsonl | jq .

# pretty each row with nothing but bash + jq (note -rs: raw, per-record)
while IFS= read -r line; do
  [ -n "$line" ] || continue
  printf '%s\n' "$line" | jq .
done < events.jsonl

The while read loop exists to make a point: nothing about NDJSON requires jq. printf '%s\n' "$line" | jq . is the entire algorithm — the format is the API. For our tool's workflow: pretty-print a selection of rows (head/grep a window into the beautifier), not the file; for the whole file, jq streams where the browser cannot.

Streaming Node, no dependency:

const readline = require("node:readline");
const rl = readline.createInterface({ input: process.stdin.pipe(require("node:zlib").createGunzip()) });
let bad = 0;
rl.on("line", (l) => { try { handle(JSON.parse(l)); } catch { bad++; } });

Converting: array ↔ NDJSON

The -s/-c pair is the whole vocabulary:

# NDJSON → array-JSON
jq -s '.' events.jsonl                       # slurp: collect rows, emit one array

# array-JSON → NDJSON
jq -c '.[]' big.json                         # one row per input, compact output

# NDJSON → filtered/reshaped NDJSON (stays streaming)
jq -c 'select(.level=="error")' events.jsonl

# add a field to every row, in place (careful: same-file redirect truncates!)
jq -c '. + {env:"prod"}' events.jsonl > tmp && mv tmp events.jsonl

Verified shapes: jq -s '.' on two rows gives [{"a":1},{"a":2}]; jq -c '.[]' on an array of three objects prints three lines. An empty file slurps to [] with exit 0 — no error, no output — which will bite any script that assumes jq -s printed something: check array length, not exit code.

.ndjson or .jsonl?

Same format, same bytes, no functional difference. .jsonl is the older convention (Elasticsearch, HuggingFace datasets, Postgres COPY ... FORMAT jsonl); .ndjson is what jq's docs and older NDJSON spec drafts favoured. Pick what your ecosystem's tools recognise; never mix extensions inside one project — glob bugs find that boundary first. wc -l counts newlines, not records: a file with no final newline reports one fewer than jq -s 'length' (verified: wc -l says 1, jq counts 2). Trust the parser, not the line counter.

Three pitfalls that corrupt silently

1. No final newline. Most tools survive it (jq, Python's readlines); wc -l miscounts and naive while read loops drop the last row — verified: on a two-row file with no trailing newline, while IFS= read -r line iterates once: the unterminated final chunk never satisfies read's success test, so row two is silently lost. Always end files with \n.

2. Embedded raw newlines = corruption, by construction. The newline is the record separator, so an unescaped one inside a string value splices one row into two broken ones. JSON.stringify escapes \n properly — the corruption comes from hand-built lines (echo "..." >> file) or a middleware that unescapes and re-emits. A quick check: any row failing to parse and the next row failing too is almost always this:

while IFS= read -r l; do echo "$l" | jq -e . >/dev/null || echo "BAD: $l" ; done < f.jsonl

3. Blank and comment lines. Real JSONL has neither. A jq -c . run that skips blank lines quietly hides that something appended "\n\n"; a hand-edited # comment line kills strict parsers. Treat any skipped line as an anomaly to chase, not noise to ignore — same philosophy as Auto-Repair being a diagnostic, not a workflow.

FAQ

Can I paste a whole NDJSON file into this tool? One row at a time, realistically. The textarea holds one parsed document; NDJSON is many. Extract the window you care about (sed -n '100,120p' f.jsonl | jq .) and paste that.

Does jq stream mode (--stream) matter here? For NDJSON, no — plain jq without -s is already line-buffered and constant-memory. --stream is for giant single JSON documents (one big array), a different problem.

NDJSON vs Parquet/Avro? Different tier. Columnar formats win at analytics scale but need schemas and readers. NDJSON wins at "a human must debug this at 3 a.m. with grep", which is most of what small teams actually run.

Why does my streaming API return data: {...} lines? That is SSE, NDJSON's louder cousin: same line-delimited idea, wrapped with data: prefixes and blank-line framing. Strip the prefix and everything here applies.

Developer Sponsor / Partner
Copied to clipboard!