Project P043

Line diff

Two texts compared line by line: the longest run of lines they share, in order, stays, and everything else is shown as removed from the first or added in the second - in full with line numbers on both sides, or only the changes with a line of context around each. The texts can be swapped to see the change the other way.

2 modules · 2 recorded sessionstext-menu UI in the terminalupdated 2026-10-10

Every screen below was recorded under CPython. When this page was built, the EML interpreter replayed each session from the same input and printed the same bytes.

About

Two texts, A and B, compared line by line. The longest run of lines they share, in order, stays; everything else is shown as removed from A or added in B - the whole texts with the line numbers on both sides, or only the changes with a line of context around each. A and B can be swapped to see the change the other way round.

  • main.eml - the menu, typing a text with its checks, the sample, and the differences on screen
  • diff.eml - the table of common lines, the walk that turns it into a difference, and which lines to show

How each part works:

  • The table is the one of the corpus case longest-common-subsequence, with lines in place of characters, filled from the ends: entry (i, j) is the most lines that A from line i and B from line j can keep in common, in order. The corpus case walks its table backwards to recover the common part; filled from the ends, this one can be walked forwards, which is the order a difference is read in.
  • The walk keeps a line both texts can keep; otherwise it drops the line of A if that loses nothing, and adds the line of B if it does. So the number of lines kept is the largest possible, and in every changed stretch the removed lines come before the added ones.
  • "Changes only" shows every change and one unchanged line on each side of it; ... stands for the lines left out.
  • Lines are compared exactly, spaces at the end aside.

What is checked: at most 30 lines of up to 60 characters per text; a line holding only a dot ends the text, so an empty line can be part of it. An empty menu answer is not a choice.

Sessions: sessions/basic.in compares the sample - two versions of a to-do list, three lines removed, three added, three the same - both ways round; then types two versions of a recipe, ten and eleven lines long, and shows only the changes: three places, with ... between them. sessions/bad-input.in gives menu choices 0 and x, compares two empty texts, types a line of 61 characters, makes A and B the same three lines (an empty one among them), then an empty B - everything removed, and after the swap everything added - and types 31 lines, which stops at 30.

Built on the verified corpus case longest-common-subsequence (the dynamic programming table and the walk that recovers the common part).

Recorded sessions

What the screen shows while someone uses the program. Each typed line appears after its prompt, the way a terminal shows it.

bad-input

interpreter: byte-equal
== Line diff ==
Compare two texts, A and B, line by line.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 0
Pick a number from 1 to 7.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> x
Pick a number from 1 to 7.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 4
Both texts are empty.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 1
Type text A line by line, at most 30 lines of up to 60 characters; a line with only a dot ends it.
A1> xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
That line has 61 characters; the most is 60. Type it again, shorter.
A1> first
A2> 
A3> third
A4> .
A has 3 lines.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 2
Type text B line by line, at most 30 lines of up to 60 characters; a line with only a dot ends it.
B1> first
B2> 
B3> third
B4> .
B has 3 lines.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 4
    A   B
    1   1  first
    2   2
    3   3  third
The texts are the same: 3 lines.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 5
The texts are the same: 3 lines.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 2
Type text B line by line, at most 30 lines of up to 60 characters; a line with only a dot ends it.
B1> .
B has 0 lines.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 4
    A   B
-   1      first
-   2
-   3      third
3 lines removed, 0 lines added, 0 lines the same - the longest common part, in order.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 6
A and B swapped.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 4
    A   B
+       1  first
+       2
+       3  third
0 lines removed, 3 lines added, 0 lines the same - the longest common part, in order.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 1
Type text A line by line, at most 30 lines of up to 60 characters; a line with only a dot ends it.
A1> line 1
A2> line 2
A3> line 3
A4> line 4
A5> line 5
A6> line 6
A7> line 7
A8> line 8
A9> line 9
A10> line 10
A11> line 11
A12> line 12
A13> line 13
A14> line 14
A15> line 15
A16> line 16
A17> line 17
A18> line 18
A19> line 19
A20> line 20
A21> line 21
A22> line 22
A23> line 23
A24> line 24
A25> line 25
A26> line 26
A27> line 27
A28> line 28
A29> line 29
A30> line 30
A31> line 31
That makes 30 lines, the most; A ends here.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 7
Bye.
What was typed (54 lines)
0
x
4
1
xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
first

third
.
2
first

third
.
4
5
2
.
4
6
4
1
line 1
line 2
line 3
line 4
line 5
line 6
line 7
line 8
line 9
line 10
line 11
line 12
line 13
line 14
line 15
line 16
line 17
line 18
line 19
line 20
line 21
line 22
line 23
line 24
line 25
line 26
line 27
line 28
line 29
line 30
line 31
7

basic

interpreter: byte-equal
== Line diff ==
Compare two texts, A and B, line by line.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 3
A and B now hold two versions of a to-do list.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 4
    A   B
-   1      Buy milk
+       1  Buy milk and bread
    2   2  Call the plumber
-   3      Pay the electricity bill
    4   3  Write to Ana
-   5      Book the dentist
    6   4  Water the plants
+       5  Book the dentist
+       6  Clean the windows
3 lines removed, 3 lines added, 3 lines the same - the longest common part, in order.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 6
A and B swapped.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 4
    A   B
-   1      Buy milk and bread
+       1  Buy milk
    2   2  Call the plumber
+       3  Pay the electricity bill
    3   4  Write to Ana
-   4      Water the plants
    5   5  Book the dentist
-   6      Clean the windows
+       6  Water the plants
3 lines removed, 3 lines added, 3 lines the same - the longest common part, in order.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 1
Type text A line by line, at most 30 lines of up to 60 characters; a line with only a dot ends it.
A1> Heat the oven to 200 degrees.
A2> Mix the flour and the salt.
A3> Rub in the butter.
A4> Add the milk slowly.
A5> Knead the dough for a minute.
A6> Roll it out to two centimetres.
A7> Cut out the rounds.
A8> Brush the tops with milk.
A9> Bake for twelve minutes.
A10> Cool on a rack.
A11> .
A has 10 lines.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 2
Type text B line by line, at most 30 lines of up to 60 characters; a line with only a dot ends it.
B1> Heat the oven to 220 degrees.
B2> Mix the flour and the salt.
B3> Rub in the butter.
B4> Add the milk slowly.
B5> Knead the dough for a minute.
B6> Rest it for ten minutes.
B7> Roll it out to two centimetres.
B8> Cut out the rounds.
B9> Brush the tops with milk.
B10> Bake for ten to twelve minutes.
B11> Cool on a rack.
B12> .
B has 11 lines.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 5
    A   B
-   1      Heat the oven to 200 degrees.
+       1  Heat the oven to 220 degrees.
    2   2  Mix the flour and the salt.
   ...
    5   5  Knead the dough for a minute.
+       6  Rest it for ten minutes.
    6   7  Roll it out to two centimetres.
   ...
    8   9  Brush the tops with milk.
-   9      Bake for twelve minutes.
+      10  Bake for ten to twelve minutes.
   10  11  Cool on a rack.
2 lines removed, 3 lines added, 8 lines the same - the longest common part, in order.

1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit
choice> 7
Bye.
What was typed (31 lines)
3
4
6
4
1
Heat the oven to 200 degrees.
Mix the flour and the salt.
Rub in the butter.
Add the milk slowly.
Knead the dough for a minute.
Roll it out to two centimetres.
Cut out the rounds.
Brush the tops with milk.
Bake for twelve minutes.
Cool on a rack.
.
2
Heat the oven to 220 degrees.
Mix the flour and the salt.
Rub in the butter.
Add the milk slowly.
Knead the dough for a minute.
Rest it for ten minutes.
Roll it out to two centimetres.
Cut out the rounds.
Brush the tops with milk.
Bake for ten to twelve minutes.
Cool on a rack.
.
5
7

Modules

The program as written, entry module first. Each module transpiles to its own Python file, which is what eml project run executes.

main.eml(entry)

eml
# P043 line diff: two texts, A and B, compared line by line. The longest run
# of lines they share, in order, stays; everything else is shown as removed
# from A or added in B - the whole texts, or only the changes with a line of
# context around each.
import diff

30 => most_lines
60 => longest_line

def trim_right(s):
    len(s) => j
    while j > 0 and s[j - 1] == " ":
        j - 1 => j
    return s[0:j]

def right(s, width):
    while len(s) < width:
        " " + s => s
    return s

def lines_text(n):
    if n == 1:
        return "1 line"
    return str(n) + " lines"

def read_text(name):
    ("Type text " + name + " line by line, at most " + str(most_lines) + " lines of up to " + str(longest_line) + " characters; a line with only a dot ends it.") ^0
    [] => lines
    while True:
        trim_right(input(name + str(len(lines) + 1) + "> ")) => line
        if line == ".":
            (name + " has " + lines_text(len(lines)) + ".") ^0
            return lines
        if len(line) > longest_line:
            ("That line has " + str(len(line)) + " characters; the most is " + str(longest_line) + ". Type it again, shorter.") ^0
        elif len(lines) == most_lines:
            ("That makes " + str(most_lines) + " lines, the most; " + name + " ends here.") ^0
            return lines
        else:
            lines + [line] => lines

def show_difference(a, b, context):
    if len(a) == 0 and len(b) == 0:
        "Both texts are empty." ^0
        return 0
    diff.difference(a, b) => d
    diff.counts(d) => c
    diff.shown_lines(d, context) => show
    if context != -1 and c[1] == 0 and c[2] == 0:
        ("The texts are the same: " + lines_text(c[0]) + ".") ^0
        return 0
    "    A   B" ^0
    False => gap
    for k in [0:len(d) - 1]:
        if show[k]:
            if gap:
                "   ..." ^0
            False => gap
            d[k] => x
            "" => na
            if x[1] > 0:
                str(x[1]) => na
            "" => nb
            if x[2] > 0:
                str(x[2]) => nb
            trim_right(x[0] + " " + right(na, 3) + " " + right(nb, 3) + "  " + x[3]) ^0
        else:
            True => gap
    if gap:
        "   ..." ^0
    if c[1] == 0 and c[2] == 0:
        ("The texts are the same: " + lines_text(c[0]) + ".") ^0
    else:
        (lines_text(c[1]) + " removed, " + lines_text(c[2]) + " added, " + lines_text(c[0]) + " the same - the longest common part, in order.") ^0
    return 0

"== Line diff ==" ^0
"Compare two texts, A and B, line by line." ^0
[] => a
[] => b
True => running
while running:
    "" ^0
    "1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit" ^0
    trim_right(input("choice> ")) => choice
    if choice == "1":
        read_text("A") => a
    elif choice == "2":
        read_text("B") => b
    elif choice == "3":
        ["Buy milk", "Call the plumber", "Pay the electricity bill", "Write to Ana", "Book the dentist", "Water the plants"] => a
        ["Buy milk and bread", "Call the plumber", "Write to Ana", "Water the plants", "Book the dentist", "Clean the windows"] => b
        "A and B now hold two versions of a to-do list." ^0
    elif choice == "4":
        show_difference(a, b, -1)
    elif choice == "5":
        show_difference(a, b, 1)
    elif choice == "6":
        a => t
        b => a
        t => b
        "A and B swapped." ^0
    elif choice == "7":
        False => running
    else:
        "Pick a number from 1 to 7." ^0
"Bye." ^0
Python projection (main.py)
import diff
most_lines = 30
longest_line = 60

def trim_right(s):
    j = len(s)
    while j > 0 and s[j - 1] == " ":
        j = j - 1
    return s[0:j]

def right(s, width):
    while len(s) < width:
        s = " " + s
    return s

def lines_text(n):
    if n == 1:
        return "1 line"
    return str(n) + " lines"

def read_text(name):
    print("Type text " + name + " line by line, at most " + str(most_lines) + " lines of up to " + str(longest_line) + " characters; a line with only a dot ends it.")
    lines = []
    while True:
        line = trim_right(input(name + str(len(lines) + 1) + "> "))
        if line == ".":
            print(name + " has " + lines_text(len(lines)) + ".")
            return lines
        if len(line) > longest_line:
            print("That line has " + str(len(line)) + " characters; the most is " + str(longest_line) + ". Type it again, shorter.")
        elif len(lines) == most_lines:
            print("That makes " + str(most_lines) + " lines, the most; " + name + " ends here.")
            return lines
        else:
            lines = lines + [line]

def show_difference(a, b, context):
    if len(a) == 0 and len(b) == 0:
        print("Both texts are empty.")
        return 0
    d = diff.difference(a, b)
    c = diff.counts(d)
    show = diff.shown_lines(d, context)
    if context != -1 and c[1] == 0 and c[2] == 0:
        print("The texts are the same: " + lines_text(c[0]) + ".")
        return 0
    print("    A   B")
    gap = False
    for k in range(0, len(d)):
        if show[k]:
            if gap:
                print("   ...")
            gap = False
            x = d[k]
            na = ""
            if x[1] > 0:
                na = str(x[1])
            nb = ""
            if x[2] > 0:
                nb = str(x[2])
            print(trim_right(x[0] + " " + right(na, 3) + " " + right(nb, 3) + "  " + x[3]))
        else:
            gap = True
    if gap:
        print("   ...")
    if c[1] == 0 and c[2] == 0:
        print("The texts are the same: " + lines_text(c[0]) + ".")
    else:
        print(lines_text(c[1]) + " removed, " + lines_text(c[2]) + " added, " + lines_text(c[0]) + " the same - the longest common part, in order.")
    return 0

print("== Line diff ==")
print("Compare two texts, A and B, line by line.")
a = []
b = []
running = True
while running:
    print("")
    print("1) type A  2) type B  3) sample  4) differences  5) changes only  6) swap  7) quit")
    choice = trim_right(input("choice> "))
    if choice == "1":
        a = read_text("A")
    elif choice == "2":
        b = read_text("B")
    elif choice == "3":
        a = ["Buy milk", "Call the plumber", "Pay the electricity bill", "Write to Ana", "Book the dentist", "Water the plants"]
        b = ["Buy milk and bread", "Call the plumber", "Write to Ana", "Water the plants", "Book the dentist", "Clean the windows"]
        print("A and B now hold two versions of a to-do list.")
    elif choice == "4":
        show_difference(a, b, -1)
    elif choice == "5":
        show_difference(a, b, 1)
    elif choice == "6":
        t = a
        a = b
        b = t
        print("A and B swapped.")
    elif choice == "7":
        running = False
    else:
        print("Pick a number from 1 to 7.")
print("Bye.")

diff.eml

eml
# P043 line diff - which lines two texts share and how one becomes the
# other. A text is a list of lines; a difference is a list of
# [kind, line number in A, line number in B, text], kind " " for a line in
# both, "-" for one only in A, "+" for one only in B (numbers from 1, 0 where
# the line is not in that text).

def common_table(a, b):
    # t[i][j] is the length of the longest common subsequence of the lines
    # a[i:] and b[j:] - the corpus case longest-common-subsequence, filled
    # from the ends so that the walk below can go forwards.
    len(a) => n
    len(b) => m
    [] => t
    for i in [0:n]:
        t + [[0] * (m + 1)] => t
    n - 1 => i
    while i >= 0:
        m - 1 => j
        while j >= 0:
            if a[i] == b[j]:
                t[i + 1][j + 1] + 1 => t[i][j]
            elif t[i + 1][j] >= t[i][j + 1]:
                t[i + 1][j] => t[i][j]
            else:
                t[i][j + 1] => t[i][j]
            j - 1 => j
        i - 1 => i
    return t

def difference(a, b):
    # Walks from the start: a line both texts can keep is kept; otherwise
    # the line of A goes if dropping it loses nothing, else the line of B is
    # added - so in each changed stretch the removed lines come first.
    common_table(a, b) => t
    0 => i
    0 => j
    [] => out
    while i < len(a) or j < len(b):
        if i < len(a) and j < len(b) and a[i] == b[j]:
            out + [[" ", i + 1, j + 1, a[i]]] => out
            i + 1 => i
            j + 1 => j
        elif j == len(b) or (i < len(a) and t[i + 1][j] >= t[i][j + 1]):
            out + [["-", i + 1, 0, a[i]]] => out
            i + 1 => i
        else:
            out + [["+", 0, j + 1, b[j]]] => out
            j + 1 => j
    return out

def counts(d):
    # [kept, removed, added]
    0 => kept
    0 => removed
    0 => added
    for x in d:
        if x[0] == " ":
            kept + 1 => kept
        elif x[0] == "-":
            removed + 1 => removed
        else:
            added + 1 => added
    return [kept, removed, added]

def shown_lines(d, context):
    # Which entries to show: every change, and up to `context` unchanged
    # lines on each side of one; -1 means all. Returns a list of True/False.
    [False] * len(d) => show
    for k in [0:len(d) - 1]:
        if context == -1 or d[k][0] != " ":
            True => show[k]
            if context > 0:
                for e in [1:context]:
                    if k - e >= 0:
                        True => show[k - e]
                    if k + e < len(d):
                        True => show[k + e]
    return show
Python projection (diff.py)
def common_table(a, b):
    n = len(a)
    m = len(b)
    t = []
    for i in range(0, n+1):
        t = t + [[0] * (m + 1)]
    i = n - 1
    while i >= 0:
        j = m - 1
        while j >= 0:
            if a[i] == b[j]:
                t[i][j] = t[i + 1][j + 1] + 1
            elif t[i + 1][j] >= t[i][j + 1]:
                t[i][j] = t[i + 1][j]
            else:
                t[i][j] = t[i][j + 1]
            j = j - 1
        i = i - 1
    return t

def difference(a, b):
    t = common_table(a, b)
    i = 0
    j = 0
    out = []
    while i < len(a) or j < len(b):
        if i < len(a) and j < len(b) and a[i] == b[j]:
            out = out + [[" ", i + 1, j + 1, a[i]]]
            i = i + 1
            j = j + 1
        elif j == len(b) or i < len(a) and t[i + 1][j] >= t[i][j + 1]:
            out = out + [["-", i + 1, 0, a[i]]]
            i = i + 1
        else:
            out = out + [["+", 0, j + 1, b[j]]]
            j = j + 1
    return out

def counts(d):
    kept = 0
    removed = 0
    added = 0
    for x in d:
        if x[0] == " ":
            kept = kept + 1
        elif x[0] == "-":
            removed = removed + 1
        else:
            added = added + 1
    return [kept, removed, added]

def shown_lines(d, context):
    show = [False] * len(d)
    for k in range(0, len(d)):
        if context == -1 or d[k][0] != " ":
            show[k] = True
            if context > 0:
                for e in range(1, context+1):
                    if k - e >= 0:
                        show[k - e] = True
                    if k + e < len(d):
                        show[k + e] = True
    return show

Built on these corpus cases