logo svelte /diff v0.4.3

Diff Modes

diffMode chooses the comparison unit. The exported SvelteDiffMode type is 'character' | 'word' | 'line'. Character remains the default.

ModeUse caseBehavior
character (default)Fine-grained editsExisting diff-match-patch algorithm and cleanup defaults
wordProse and identifiersWhole word runs, separate whitespace and punctuation tokens
lineLogs, configuration, source snapshotsComplete lines with their original terminators; no character refinement

Word comparison

<script lang="ts">
    import SvelteDiff from '@humanspeak/svelte-diff'
    const before = 'The cat sleeps.'
    const after = 'The car sleeps.'
</script>

<SvelteDiff originalText={before} modifiedText={after} diffMode="word" />
<script lang="ts">
    import SvelteDiff from '@humanspeak/svelte-diff'
    const before = 'The cat sleeps.'
    const after = 'The car sleeps.'
</script>

<SvelteDiff originalText={before} modifiedText={after} diffMode="word" />

Raw tuples are [[0, "The "], [-1, "cat"], [1, "car"], [0, " sleeps."]]. Unlike raw character output, the changed word is removed and inserted whole.

Word tokens are Unicode letter/mark/number/underscore runs, horizontal whitespace runs, CRLF pairs, lone CR or LF, and individual remaining punctuation/symbol code points. Spaces, tabs, NBSP, case, combining marks, and normalization form are preserved exactly. Apostrophes and hyphens separate runs. Comparison is case-sensitive and deterministic on the server and client; it does not use locale-aware segmentation. A continuous CJK letter run is one token. Emoji grapheme clusters are not guaranteed atomic. Lone surrogate code units are preserved.

Try the editable word example.

Line comparison

<script lang="ts">
    import SvelteDiff from '@humanspeak/svelte-diff'
    const before = 'count=10\nkeep=true\n'
    const after = 'count=20\nkeep=true\n'
</script>

<SvelteDiff originalText={before} modifiedText={after} diffMode="line" />
<script lang="ts">
    import SvelteDiff from '@humanspeak/svelte-diff'
    const before = 'count=10\nkeep=true\n'
    const after = 'count=20\nkeep=true\n'
</script>

<SvelteDiff originalText={before} modifiedText={after} diffMode="line" />

Raw tuples are [[-1, "count=10\n"], [1, "count=20\n"], [0, "keep=true\n"]].

A line includes its LF, CRLF, or lone-CR terminator. Mixed endings, indentation, blank lines, and an unterminated final line are preserved. Adding or removing a final newline changes the affected line token. No secondary character refinement runs. JSON is plain source text: no parsing, key sorting, or canonicalization occurs. Sentence and structural JSON modes are deferred and unsupported.

Try the line example, including its blank-line/final-newline preset. Line mode is a comparison unit, not a file/hunk or side-by-side renderer.

Cleanup matrix

ModeCleanup behavior
Character with cleanupSemantic={true}Semantic cleanup wins
Character with semantic off and cleanupEfficiency > 0Efficiency cleanup (default edit cost: 4)
Character with semantic off and efficiency 0No cleanup
Word or line, with any cleanup propsBoth passes skipped; timing.cleanup is exactly 0

Granularity and cleanup are separate choices. Semantic cleanup can make character output more readable, but does not promise whole-token boundaries. The public comparison panes disable character cleanup to show this difference clearly.

Literal source and computation

For source code, use expectedPatterns={false}. The component defaults to true; literal mode bypasses the parser and all template substitution and tagging.

<script lang="ts">
    import SvelteDiff, { computeDiff } from '@humanspeak/svelte-diff'
    const before = 'const re = /(?<year>\\d{4})/;'
    const after = 'const re = /(?<year>\\d{2})/g;'
    const result = computeDiff(before, after, { diffMode: 'line' })
</script>

<SvelteDiff originalText={before} modifiedText={after} diffMode="line" expectedPatterns={false} />
<script lang="ts">
    import SvelteDiff, { computeDiff } from '@humanspeak/svelte-diff'
    const before = 'const re = /(?<year>\\d{4})/;'
    const after = 'const re = /(?<year>\\d{2})/g;'
    const result = computeDiff(before, after, { diffMode: 'line' })
</script>

<SvelteDiff originalText={before} modifiedText={after} diffMode="line" expectedPatterns={false} />

computeDiff defaults to expectedPatterns false and returns synchronously with raw diffs, tagged displayDiffs, optional captures, and millisecond timing. It uses the same modes, timeout, and cleanup defaults as the component and owns an engine per call. Use it within a Svelte-aware toolchain: the package root requires Svelte-aware module resolution and compilation. Set expectedPatterns true explicitly when computing templates.

In literal mode, omit inserts to reconstruct the exact original input and omit removals to reconstruct the exact modified input, including all line endings, tabs, and Unicode code units.

Expected patterns and renderers

With expectedPatterns enabled, the pipeline remains extraction → resolved/cleaned source → diff → capture tagging. Tokenization sees the resolved source, not regex syntax. On mismatch, it sees the cleaned placeholders. Raw callback tuples reconstruct that source when inserts are omitted, and reconstruct the exact modified string when removals are omitted.

<script lang="ts">
    import SvelteDiff from '@humanspeak/svelte-diff'
    const template = 'Release (?<version>v\\d+)'
    const actual = 'Release v2 ready'
</script>

<SvelteDiff originalText={template} modifiedText={actual} diffMode="line" />
<script lang="ts">
    import SvelteDiff from '@humanspeak/svelte-diff'
    const template = 'Release (?<version>v\\d+)'
    const actual = 'Release v2 ready'
</script>

<SvelteDiff originalText={template} modifiedText={actual} diffMode="line" />

Here the resolved source is Release v2. Raw tuples are [[-1, "Release v2"], [1, "Release v2 ready"]], and captures contain version: 'v2'. The removed line still contains v2; its replacement also displays v2 as expected. Removals are never silently suppressed.

Whole-token guarantees apply to raw tuple boundaries. Expected annotations may split a displayed word or line at capture boundaries. The existing newline renderer can split multiline tuples into multiple snippet calls. Child snippets still override renderer-map entries, which override built-in rendering. Compact rendering and custom line breaks work in every mode. See expected patterns and the component API.

Reactivity, SSR, timing, and limits

Changing either text, diffMode, expectedPatterns, timeout, or cleanup options recomputes the diff. Changing only the callback reuses the same tuple array. Initial diff markup is computed during SSR; callback delivery happens on the client after computation.

For token modes, timing.main includes tokenization, shared-dictionary encoding, diffing, and decoding; cleanup is zero and total covers the timed computation. Expected-pattern preprocessing and DOM rendering are outside these measurements.

timeout is in seconds (default 1; 0 means unlimited). Token preparation, the engine, and decoding share one absolute deadline. Expiry can return a coarse full delete/insert, preserving both complete strings. The shared dictionary supports at most 65,535 distinct tokens across both inputs; exceeding this also returns a complete replacement. Equality and one-sided empty inputs keep their simple fast-path results. Fallback never switches to character granularity or truncates input.

This is a best-effort algorithm deadline, not a hard cap on a large token allocation, regex extraction, or rendering. Work remains synchronous: token modes do not virtualize output, offload to workers, or debounce input. Read the performance guide before comparing large documents.