Statistics
A list of numbers - whole numbers, decimals or fractions, kept exact - and what describes it: count, range, mean, median, modes, quartiles and outliers, variance and standard deviation dividing by n and by n - 1, whether the mean sits above or below the median, and a histogram whose top edge does not drop the largest value.
Every screen below was recorded under CPython. When this page was built, the EML interpreter replayed each session from the same input and printed the same bytes.
About
A list of up to 60 numbers - whole numbers, decimals or fractions - and what describes it: count, range, sum, mean, median, modes, quartiles and outliers, variance and standard deviation dividing by n and by n - 1, whether the mean sits above or below the median, and a histogram. Every value is kept as an exact fraction, so nothing is rounded until it is shown, to three places.
main.eml- the menu, adding and removing values with their checks, and the summary and histogram on screenstats.eml- the measures: sorting, sum, mean, median, quartiles, modes, squared distances and the histogram countsfrac.eml- exact fractions, from P007 by way of P038, with decimals and a whole-number square root
How each part works:
- The median is the middle value, or the mean of the two middle ones. The quartiles are the medians of the lower and the upper half, leaving the median out when the count is odd; values more than 1.5 interquartile ranges beyond them are named as outliers.
- The variance is the mean squared distance from the mean, exact; as a sample it divides by n - 1 instead. The standard deviation is a square root, which is rarely a fraction: it is found with a whole-number square root (Newton's method, no floats) of the variance scaled up, and rounded half up - "about" says when the shown digits are not exact.
- The corpus case
mean-median-skewis about averages that disagree: when a few large values pull the mean above the median, the summary says so and counts how many values are above the mean - in the sample, only 2 of 12. - The histogram's ranges include their lower edge; the last includes the largest value too, the edge the corpus case
histogram-builderhandles, so the top value is not dropped.
What is checked: numbers as whole numbers, decimals (2.5, -3.75) or fractions (3/4), separated by spaces or commas, at most 60 in all - a line with anything else adds nothing; a value to remove that is in the list; 2 to 10 histogram ranges. An empty answer cancels.
Sessions: sessions/basic.in loads the sample - twelve delivery times, two of them far longer than the rest: mean about 21.083 against a median of 15, outliers 45 and 60 - draws a histogram with five ranges, removes 60 and summarises again, then clears and summarises five numbers typed with decimals and a fraction (no mode, an exact variance of 11.89). sessions/bad-input.in gives menu choices 0 and x, every action with no values, a line with x in it, 1/0, sixty-one numbers at once, three equal values (one mode, variance 0, no quartiles, no histogram ranges), a value not in the list, abc, and 1, 11 and x ranges.
Built on the verified corpus cases mean-median-skew (three averages on a skewed sample, and how far each moves) and histogram-builder (equal ranges and the top edge).
Recorded sessions
What the screen shows while someone uses the program. Each typed line appears after its prompt, the way a terminal shows it.
bad-input
interpreter: byte-equal== Statistics ==
Numbers in, measures out - exact, rounded only when shown (to 3 places).
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 0
Pick a number from 1 to 8.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> x
Pick a number from 1 to 8.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 3
There are no values yet.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 4
There are no values yet.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 5
There are no values yet.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 2
There are no values.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 1
numbers (separated by spaces)> 3 x 4
Not a number: x. Nothing was added - type whole numbers, decimals like 2.5 or fractions like 3/4.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 1
numbers (separated by spaces)> 1/0
Not a number: 1/0. Nothing was added - type whole numbers, decimals like 2.5 or fractions like 3/4.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 1
numbers (separated by spaces)>
Cancelled.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 1
numbers (separated by spaces)> 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61
That would make 61 values; the most is 60. Nothing was added.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 1
numbers (separated by spaces)> 5, 5, 5
3 values added; 3 values in all.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 3
3 values, from 5 to 5 (range 0).
Sum 15, mean 5, median 5.
Mode: 5 (3 times).
Quartiles need at least 4 values.
Variance 0, standard deviation 0 - dividing by n.
As a sample: variance 0, standard deviation 0 - dividing by n - 1.
The mean and the median are equal. 0 of 3 values are above the mean.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 4
bins (2 to 10)> 2
All 3 values are 5: there are no ranges to draw.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 2
value to remove> 7
7 is not in the list.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 2
value to remove> abc
Not a number: abc.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 2
value to remove>
Cancelled.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 4
bins (2 to 10)> 1
Type a number from 2 to 10.
bins (2 to 10)> 11
Type a number from 2 to 10.
bins (2 to 10)> x
Type a number from 2 to 10.
bins (2 to 10)>
Cancelled.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 6
All values cleared.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 5
There are no values yet.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 8
Bye.
What was typed (33 lines)
0
x
3
4
5
2
1
3 x 4
1
1/0
1
1
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61
1
5, 5, 5
3
4
2
2
7
2
abc
2
4
1
11
x
6
5
8
basic
interpreter: byte-equal== Statistics ==
Numbers in, measures out - exact, rounded only when shown (to 3 places).
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 7
Sample: 12 delivery times in minutes.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 3
12 values, from 12 to 60 (range 48).
Sum 253, mean about 21.083, median 15.
Mode: 14 (3 times).
Quartiles: Q1 14, Q3 17.5; interquartile range 3.5.
Outliers, beyond 1.5 times the interquartile range (below 8.75 or above 22.75): 45 and 60.
Variance about 209.243, standard deviation about 14.465 - dividing by n.
As a sample: variance about 228.265, standard deviation about 15.108 - dividing by n - 1.
The mean is above the median: a few large values pull it up. 2 of 12 values are above the mean.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 4
bins (2 to 10)> 5
[12, 21.6) 10 ##########
[21.6, 31.2) 0
[31.2, 40.8) 0
[40.8, 50.4) 1 #
[50.4, 60] 1 #
The ranges include their lower edge; the last includes the largest value too. 12 values in all.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 5
Sorted: 12, 13, 14, 14, 14, 15, 15, 16, 17, 18, 45 and 60.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 2
value to remove> 60
60 removed; 11 values left.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 3
11 values, from 12 to 45 (range 33).
Sum 193, mean about 17.545, median 15.
Mode: 14 (3 times).
Quartiles: Q1 14, Q3 17; interquartile range 3.
Outliers, beyond 1.5 times the interquartile range (below 9.5 or above 21.5): 45.
Variance about 78.066, standard deviation about 8.836 - dividing by n.
As a sample: variance about 85.873, standard deviation about 9.267 - dividing by n - 1.
The mean is above the median: a few large values pull it up. 2 of 11 values are above the mean.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 6
All values cleared.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 1
numbers (separated by spaces)> 2.5 3.5 1/4 7 10
5 values added; 5 values in all.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 3
5 values, from 0.25 to 10 (range 9.75).
Sum 23.25, mean 4.65, median 3.5.
No mode: every value occurs once.
Quartiles: Q1 1.375, Q3 8.5; interquartile range 7.125.
Outliers, beyond 1.5 times the interquartile range (below about -9.313 or above about 19.188): none.
Variance 11.89, standard deviation about 3.448 - dividing by n.
As a sample: variance about 14.863, standard deviation about 3.855 - dividing by n - 1.
The mean is above the median: a few large values pull it up. 2 of 5 values are above the mean.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 4
bins (2 to 10)> 3
[0.25, 3.5) 2 ##
[3.5, 6.75) 1 #
[6.75, 10] 2 ##
The ranges include their lower edge; the last includes the largest value too. 5 values in all.
1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit
choice> 8
Bye.
What was typed (15 lines)
7
3
4
5
5
2
60
3
6
1
2.5 3.5 1/4 7 10
3
4
3
8
Modules
The program as written, entry module first. Each module transpiles to its own Python file, which is what eml project run executes.
main.eml(entry)
eml# P044 statistics: a list of numbers - whole numbers, decimals or fractions -
# and what describes it: count, range, mean, median, modes, quartiles and
# outliers, variance and standard deviation, and a histogram. Every value is
# an exact fraction, so nothing is rounded until it is shown.
import frac
import stats
60 => most_values
def trim(s):
0 => i
len(s) => j
while i < j and s[i] == " ":
i + 1 => i
while j > i and s[j - 1] == " ":
j - 1 => j
return s[i:j]
def words(s):
[] => out
"" => word
for c in s + " ":
if c == " " or c == ",":
if word != "":
out + [word] => out
"" => word
else:
word + c => word
return out
def d3(x):
return frac.decimal(x, 3)
def values_text(n):
if n == 1:
return "1 value"
return str(n) + " values"
def listed(xs):
"" => out
for i in [0:len(xs) - 1]:
if i > 0 and i == len(xs) - 1:
out + " and " => out
elif i > 0:
out + ", " => out
out + d3(xs[i]) => out
return out
def add(xs):
trim(input("numbers (separated by spaces)> ")) => answer
if answer == "":
"Cancelled." ^0
return xs
[] => new
for w in words(answer):
frac.parse(w) => x
if len(x) == 0:
("Not a number: " + w + ". Nothing was added - type whole numbers, decimals like 2.5 or fractions like 3/4.") ^0
return xs
new + [x] => new
if len(xs) + len(new) > most_values:
("That would make " + str(len(xs) + len(new)) + " values; the most is " + str(most_values) + ". Nothing was added.") ^0
return xs
(values_text(len(new)) + " added; " + values_text(len(xs) + len(new)) + " in all.") ^0
return xs + new
def remove(xs):
if len(xs) == 0:
"There are no values." ^0
return xs
trim(input("value to remove> ")) => answer
if answer == "":
"Cancelled." ^0
return xs
frac.parse(answer) => x
if len(x) == 0:
("Not a number: " + answer + ".") ^0
return xs
for i in [0:len(xs) - 1]:
if xs[i] == x:
(d3(x) + " removed; " + values_text(len(xs) - 1) + " left.") ^0
return xs[0:i] + xs[i + 1:len(xs)]
(d3(x) + " is not in the list.") ^0
return xs
def summary(xs):
if len(xs) == 0:
"There are no values yet." ^0
return 0
stats.sorted_values(xs) => s
len(s) => n
s[0] => low
s[n - 1] => high
stats.mean(s) => m
stats.median_of(s) => md
(values_text(n) + ", from " + d3(low) + " to " + d3(high) + " (range " + d3(frac.minus(high, low)) + ").") ^0
("Sum " + d3(stats.total(s)) + ", mean " + d3(m) + ", median " + d3(md) + ".") ^0
stats.modes(s) => mo
if len(mo[0]) == 0:
"once" => often
if mo[1] > 1:
str(mo[1]) + " times" => often
("No mode: every value occurs " + often + ".") ^0
elif len(mo[0]) == 1:
("Mode: " + d3(mo[0][0]) + " (" + str(mo[1]) + " times).") ^0
else:
("Modes: " + listed(mo[0]) + " (" + str(mo[1]) + " times each).") ^0
if n >= 4:
stats.quartiles(s) => q
frac.minus(q[1], q[0]) => iqr
frac.times([3, 2], iqr) => reach
frac.minus(q[0], reach) => fence_low
frac.plus(q[1], reach) => fence_high
("Quartiles: Q1 " + d3(q[0]) + ", Q3 " + d3(q[1]) + "; interquartile range " + d3(iqr) + ".") ^0
[] => out
for x in s:
if frac.less(x, fence_low) or frac.less(fence_high, x):
out + [x] => out
"none" => found
if len(out) > 0:
listed(out) => found
("Outliers, beyond 1.5 times the interquartile range (below " + d3(fence_low) + " or above " + d3(fence_high) + "): " + found + ".") ^0
else:
"Quartiles need at least 4 values." ^0
stats.squares_about(s, m) => ss
frac.over(ss, [n, 1]) => var
("Variance " + d3(var) + ", standard deviation " + frac.root_decimal(var, 3) + " - dividing by n.") ^0
if n >= 2:
frac.over(ss, [n - 1, 1]) => svar
("As a sample: variance " + d3(svar) + ", standard deviation " + frac.root_decimal(svar, 3) + " - dividing by n - 1.") ^0
0 => above
for x in s:
if frac.less(m, x):
above + 1 => above
if frac.less(md, m):
("The mean is above the median: a few large values pull it up. " + str(above) + " of " + str(n) + " values are above the mean.") ^0
elif frac.less(m, md):
("The mean is below the median: a few small values pull it down. " + str(above) + " of " + str(n) + " values are above the mean.") ^0
else:
("The mean and the median are equal. " + str(above) + " of " + str(n) + " values are above the mean.") ^0
return 0
def histogram(xs):
if len(xs) == 0:
"There are no values yet." ^0
return 0
while True:
trim(input("bins (2 to 10)> ")) => answer
if answer == "":
"Cancelled." ^0
return 0
frac.digits_value(answer) => k
if k >= 2 and k <= 10:
stats.sorted_values(xs) => s
s[0] => low
s[len(s) - 1] => high
if low == high:
("All " + values_text(len(s)) + " are " + d3(low) + ": there are no ranges to draw.") ^0
return 0
stats.bucket_counts(s, low, high, k) => counts
frac.over(frac.minus(high, low), [k, 1]) => width
[] => labels
0 => widest
for i in [0:k - 1]:
frac.plus(low, frac.times(width, [i, 1])) => a
frac.plus(low, frac.times(width, [i + 1, 1])) => b
")" => close
if i == k - 1:
"]" => close
"[" + frac.decimal(a, 2) + ", " + frac.decimal(b, 2) + close => label
labels + [label] => labels
if len(label) > widest:
len(label) => widest
for i in [0:k - 1]:
labels[i] => label
while len(label) < widest:
label + " " => label
(" " + label + " " + str(counts[i]) + " " + "#" * counts[i]) => line
if counts[i] == 0:
(" " + label + " 0") => line
line ^0
("The ranges include their lower edge; the last includes the largest value too. " + values_text(len(s)) + " in all.") ^0
return 0
"Type a number from 2 to 10." ^0
"== Statistics ==" ^0
"Numbers in, measures out - exact, rounded only when shown (to 3 places)." ^0
[] => xs
True => running
while running:
"" ^0
"1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit" ^0
trim(input("choice> ")) => choice
if choice == "1":
add(xs) => xs
elif choice == "2":
remove(xs) => xs
elif choice == "3":
summary(xs)
elif choice == "4":
histogram(xs)
elif choice == "5":
if len(xs) == 0:
"There are no values yet." ^0
else:
("Sorted: " + listed(stats.sorted_values(xs)) + ".") ^0
elif choice == "6":
[] => xs
"All values cleared." ^0
elif choice == "7":
[] => xs
for v in [12, 15, 14, 18, 13, 16, 45, 14, 17, 15, 60, 14]:
xs + [[v, 1]] => xs
"Sample: 12 delivery times in minutes." ^0
elif choice == "8":
False => running
else:
"Pick a number from 1 to 8." ^0
"Bye." ^0
Python projection (main.py)
import frac
import stats
most_values = 60
def trim(s):
i = 0
j = len(s)
while i < j and s[i] == " ":
i = i + 1
while j > i and s[j - 1] == " ":
j = j - 1
return s[i:j]
def words(s):
out = []
word = ""
for c in s + " ":
if c == " " or c == ",":
if word != "":
out = out + [word]
word = ""
else:
word = word + c
return out
def d3(x):
return frac.decimal(x, 3)
def values_text(n):
if n == 1:
return "1 value"
return str(n) + " values"
def listed(xs):
out = ""
for i in range(0, len(xs)):
if i > 0 and i == len(xs) - 1:
out = out + " and "
elif i > 0:
out = out + ", "
out = out + d3(xs[i])
return out
def add(xs):
answer = trim(input("numbers (separated by spaces)> "))
if answer == "":
print("Cancelled.")
return xs
new = []
for w in words(answer):
x = frac.parse(w)
if len(x) == 0:
print("Not a number: " + w + ". Nothing was added - type whole numbers, decimals like 2.5 or fractions like 3/4.")
return xs
new = new + [x]
if len(xs) + len(new) > most_values:
print("That would make " + str(len(xs) + len(new)) + " values; the most is " + str(most_values) + ". Nothing was added.")
return xs
print(values_text(len(new)) + " added; " + values_text(len(xs) + len(new)) + " in all.")
return xs + new
def remove(xs):
if len(xs) == 0:
print("There are no values.")
return xs
answer = trim(input("value to remove> "))
if answer == "":
print("Cancelled.")
return xs
x = frac.parse(answer)
if len(x) == 0:
print("Not a number: " + answer + ".")
return xs
for i in range(0, len(xs)):
if xs[i] == x:
print(d3(x) + " removed; " + values_text(len(xs) - 1) + " left.")
return xs[0:i] + xs[i + 1:len(xs)]
print(d3(x) + " is not in the list.")
return xs
def summary(xs):
if len(xs) == 0:
print("There are no values yet.")
return 0
s = stats.sorted_values(xs)
n = len(s)
low = s[0]
high = s[n - 1]
m = stats.mean(s)
md = stats.median_of(s)
print(values_text(n) + ", from " + d3(low) + " to " + d3(high) + " (range " + d3(frac.minus(high, low)) + ").")
print("Sum " + d3(stats.total(s)) + ", mean " + d3(m) + ", median " + d3(md) + ".")
mo = stats.modes(s)
if len(mo[0]) == 0:
often = "once"
if mo[1] > 1:
often = str(mo[1]) + " times"
print("No mode: every value occurs " + often + ".")
elif len(mo[0]) == 1:
print("Mode: " + d3(mo[0][0]) + " (" + str(mo[1]) + " times).")
else:
print("Modes: " + listed(mo[0]) + " (" + str(mo[1]) + " times each).")
if n >= 4:
q = stats.quartiles(s)
iqr = frac.minus(q[1], q[0])
reach = frac.times([3, 2], iqr)
fence_low = frac.minus(q[0], reach)
fence_high = frac.plus(q[1], reach)
print("Quartiles: Q1 " + d3(q[0]) + ", Q3 " + d3(q[1]) + "; interquartile range " + d3(iqr) + ".")
out = []
for x in s:
if frac.less(x, fence_low) or frac.less(fence_high, x):
out = out + [x]
found = "none"
if len(out) > 0:
found = listed(out)
print("Outliers, beyond 1.5 times the interquartile range (below " + d3(fence_low) + " or above " + d3(fence_high) + "): " + found + ".")
else:
print("Quartiles need at least 4 values.")
ss = stats.squares_about(s, m)
var = frac.over(ss, [n, 1])
print("Variance " + d3(var) + ", standard deviation " + frac.root_decimal(var, 3) + " - dividing by n.")
if n >= 2:
svar = frac.over(ss, [n - 1, 1])
print("As a sample: variance " + d3(svar) + ", standard deviation " + frac.root_decimal(svar, 3) + " - dividing by n - 1.")
above = 0
for x in s:
if frac.less(m, x):
above = above + 1
if frac.less(md, m):
print("The mean is above the median: a few large values pull it up. " + str(above) + " of " + str(n) + " values are above the mean.")
elif frac.less(m, md):
print("The mean is below the median: a few small values pull it down. " + str(above) + " of " + str(n) + " values are above the mean.")
else:
print("The mean and the median are equal. " + str(above) + " of " + str(n) + " values are above the mean.")
return 0
def histogram(xs):
if len(xs) == 0:
print("There are no values yet.")
return 0
while True:
answer = trim(input("bins (2 to 10)> "))
if answer == "":
print("Cancelled.")
return 0
k = frac.digits_value(answer)
if k >= 2 and k <= 10:
s = stats.sorted_values(xs)
low = s[0]
high = s[len(s) - 1]
if low == high:
print("All " + values_text(len(s)) + " are " + d3(low) + ": there are no ranges to draw.")
return 0
counts = stats.bucket_counts(s, low, high, k)
width = frac.over(frac.minus(high, low), [k, 1])
labels = []
widest = 0
for i in range(0, k):
a = frac.plus(low, frac.times(width, [i, 1]))
b = frac.plus(low, frac.times(width, [i + 1, 1]))
close = ")"
if i == k - 1:
close = "]"
label = "[" + frac.decimal(a, 2) + ", " + frac.decimal(b, 2) + close
labels = labels + [label]
if len(label) > widest:
widest = len(label)
for i in range(0, k):
label = labels[i]
while len(label) < widest:
label = label + " "
line = " " + label + " " + str(counts[i]) + " " + "#" * counts[i]
if counts[i] == 0:
line = " " + label + " 0"
print(line)
print("The ranges include their lower edge; the last includes the largest value too. " + values_text(len(s)) + " in all.")
return 0
print("Type a number from 2 to 10.")
print("== Statistics ==")
print("Numbers in, measures out - exact, rounded only when shown (to 3 places).")
xs = []
running = True
while running:
print("")
print("1) add numbers 2) remove one 3) summary 4) histogram 5) sorted 6) clear 7) sample 8) quit")
choice = trim(input("choice> "))
if choice == "1":
xs = add(xs)
elif choice == "2":
xs = remove(xs)
elif choice == "3":
summary(xs)
elif choice == "4":
histogram(xs)
elif choice == "5":
if len(xs) == 0:
print("There are no values yet.")
else:
print("Sorted: " + listed(stats.sorted_values(xs)) + ".")
elif choice == "6":
xs = []
print("All values cleared.")
elif choice == "7":
xs = []
for v in [12, 15, 14, 18, 13, 16, 45, 14, 17, 15, 60, 14]:
xs = xs + [[v, 1]]
print("Sample: 12 delivery times in minutes.")
elif choice == "8":
running = False
else:
print("Pick a number from 1 to 8.")
print("Bye.")
stats.eml
eml# P044 statistics - the measures. Values are exact fractions (frac.eml).
import frac
def sorted_values(xs):
# Insertion sort, as in the corpus case mean-median-skew.
[] => out
for x in xs:
out + [x] => out
1 => i
while i < len(out):
out[i] => cur
i - 1 => j
while j >= 0 and frac.less(cur, out[j]):
out[j] => out[j + 1]
j - 1 => j
cur => out[j + 1]
i + 1 => i
return out
def total(xs):
[0, 1] => s
for x in xs:
frac.plus(s, x) => s
return s
def mean(xs):
return frac.over(total(xs), [len(xs), 1])
def median_of(s):
# The middle of sorted values; the mean of the two middle ones when
# there is an even number.
len(s) => n
int(n / 2) => h
if n % 2 == 1:
return s[h]
return frac.over(frac.plus(s[h - 1], s[h]), [2, 1])
def quartiles(s):
# [Q1, Q3]: the medians of the lower and the upper half, leaving out the
# median itself when the count is odd. Needs at least 4 values.
len(s) => n
int(n / 2) => h
s[0:h] => lower
s[n - h:n] => upper
return [median_of(lower), median_of(upper)]
def modes(s):
# The values that occur most often, in order, and how often; [] when
# there are several values and each occurs the same number of times.
[] => values
[] => counts
for x in s:
if len(values) > 0 and values[len(values) - 1] == x:
counts[len(counts) - 1] + 1 => counts[len(counts) - 1]
else:
values + [x] => values
counts + [1] => counts
0 => top
for c in counts:
if c > top:
c => top
[] => out
for i in [0:len(values) - 1]:
if counts[i] == top:
out + [values[i]] => out
if len(values) > 1 and len(out) == len(values):
return [[], top]
return [out, top]
def squares_about(xs, m):
# The sum of squared distances from m.
[0, 1] => s
for x in xs:
frac.minus(x, m) => d
frac.plus(s, frac.times(d, d)) => s
return s
def bucket_counts(s, low, high, count):
# How many values fall in each of `count` equal ranges from low to high;
# a range holds its lower edge, and the very top value goes in the last
# one - the edge the corpus case histogram-builder is about.
[0] * count => counts
frac.minus(high, low) => span
for x in s:
if frac.is_zero(span):
0 => k
else:
frac.over(frac.times(frac.minus(x, low), [count, 1]), span) => pos
frac.quotient(pos[0], pos[1]) => k
if k >= count:
count - 1 => k
counts[k] + 1 => counts[k]
return counts
Python projection (stats.py)
import frac
def sorted_values(xs):
out = []
for x in xs:
out = out + [x]
i = 1
while i < len(out):
cur = out[i]
j = i - 1
while j >= 0 and frac.less(cur, out[j]):
out[j + 1] = out[j]
j = j - 1
out[j + 1] = cur
i = i + 1
return out
def total(xs):
s = [0, 1]
for x in xs:
s = frac.plus(s, x)
return s
def mean(xs):
return frac.over(total(xs), [len(xs), 1])
def median_of(s):
n = len(s)
h = int(n / 2)
if n % 2 == 1:
return s[h]
return frac.over(frac.plus(s[h - 1], s[h]), [2, 1])
def quartiles(s):
n = len(s)
h = int(n / 2)
lower = s[0:h]
upper = s[n - h:n]
return [median_of(lower), median_of(upper)]
def modes(s):
values = []
counts = []
for x in s:
if len(values) > 0 and values[len(values) - 1] == x:
counts[len(counts) - 1] = counts[len(counts) - 1] + 1
else:
values = values + [x]
counts = counts + [1]
top = 0
for c in counts:
if c > top:
top = c
out = []
for i in range(0, len(values)):
if counts[i] == top:
out = out + [values[i]]
if len(values) > 1 and len(out) == len(values):
return [[], top]
return [out, top]
def squares_about(xs, m):
s = [0, 1]
for x in xs:
d = frac.minus(x, m)
s = frac.plus(s, frac.times(d, d))
return s
def bucket_counts(s, low, high, count):
counts = [0] * count
span = frac.minus(high, low)
for x in s:
if frac.is_zero(span):
k = 0
else:
pos = frac.over(frac.times(frac.minus(x, low), [count, 1]), span)
k = frac.quotient(pos[0], pos[1])
if k >= count:
k = count - 1
counts[k] = counts[k] + 1
return counts
frac.eml
eml# P044 statistics - exact fractions, from P007 (calculator) by way of P038.
# A number is a list [n, d]: n / d in lowest terms with d > 0. Integers have
# no size limit, but EML has no //, and a / b goes through a float that cannot
# hold a large quotient exactly, so whole-number division is written out.
def quotient(a, b):
# a // b for whole numbers a >= 0 and b > 0, exact at any size. Long
# division by doubling: take away the largest b * 2^k that still fits.
0 => q
while a >= b:
b => m
1 => k
while m + m <= a:
m + m => m
k + k => k
a - m => a
q + k => q
return q
def gcd(a, b):
while b != 0:
a % b => r
b => a
r => b
return a
def make(n, d):
# n / d in lowest terms with a positive denominator (d != 0).
if d < 0:
0 - n => n
0 - d => d
abs(n) => a
gcd(a, d) => g
if n < 0:
return [0 - quotient(a, g), quotient(d, g)]
return [quotient(a, g), quotient(d, g)]
def plus(x, y):
return make(x[0] * y[1] + y[0] * x[1], x[1] * y[1])
def minus(x, y):
return make(x[0] * y[1] - y[0] * x[1], x[1] * y[1])
def times(x, y):
return make(x[0] * y[0], x[1] * y[1])
def over(x, y):
# x / y; the caller has checked that y is not zero.
return make(x[0] * y[1], x[1] * y[0])
def is_zero(x):
return x[0] == 0
def digits_value(s):
# The value of a string of 1 to 9 digits, otherwise -1.
if s == "" or len(s) > 9:
return -1
0 => n
for c in s:
if not (c in "0123456789"):
return -1
n * 10 + int(c) => n
return n
def parse(s):
# "3", "-2", "3/4", "-0.25" as a fraction; [] if it is not a number.
1 => sign
if len(s) > 0 and s[0] == "-":
-1 => sign
s[1:len(s)] => s
0 => k
while k < len(s) and s[k] != "/" and s[k] != ".":
k + 1 => k
if k == len(s):
digits_value(s) => n
if n < 0:
return []
return make(sign * n, 1)
digits_value(s[0:k]) => whole
s[k + 1:len(s)] => rest
digits_value(rest) => part
if whole < 0 or part < 0:
return []
if s[k] == "/":
if part == 0:
return []
return make(sign * whole, part)
# a decimal point: 0.25 is 25 / 100
1 => scale
for c in rest:
scale * 10 => scale
return make(sign * (whole * scale + part), scale)
def less(x, y):
return x[0] * y[1] < y[0] * x[1]
def isqrt(n):
# The whole-number square root of n >= 0, rounded down: Newton's method
# on whole numbers, which never needs a float.
if n < 2:
return n
n => x
quotient(x + 1, 2) => y
while y < x:
y => x
quotient(x + quotient(n, x), 2) => y
return x
def decimal(x, places):
# x to the given number of decimal places, halves rounded away from zero,
# with "about " in front when the value does not end there exactly.
"" => sign
abs(x[0]) => n
if x[0] < 0:
"-" => sign
1 => scale
for k in [1:places]:
scale * 10 => scale
quotient(2 * n * scale + x[1], 2 * x[1]) => t
"" => about
if (n * scale) % x[1] != 0:
"about " => about
str(quotient(t, scale)) => whole
str(t % scale) => part
while len(part) < places:
"0" + part => part
# drop trailing zeros of an exact value
if about == "":
while len(part) > 0 and part[len(part) - 1] == "0":
part[0:len(part) - 1] => part
if t == 0:
"" => sign
if part == "":
return about + sign + whole
return about + sign + whole + "." + part
def root_decimal(x, places):
# The square root of x >= 0 to the given places, rounded half up, by a
# whole-number square root of x * 100^places * 4 (one more binary digit
# to round with).
1 => scale
for k in [1:places]:
scale * 10 => scale
isqrt(quotient(x[0] * scale * scale * 4, x[1])) => r2
quotient(r2 + 1, 2) => t
"about " => about
if t * t * x[1] == x[0] * scale * scale:
"" => about
str(t % scale) => part
while len(part) < places:
"0" + part => part
if about == "":
while len(part) > 0 and part[len(part) - 1] == "0":
part[0:len(part) - 1] => part
if part == "":
return about + str(quotient(t, scale))
return about + str(quotient(t, scale)) + "." + part
def text(x):
# "3", "-3", "3/4", "-3/4".
if x[1] == 1:
return str(x[0])
return str(x[0]) + "/" + str(x[1])
Python projection (frac.py)
def quotient(a, b):
q = 0
while a >= b:
m = b
k = 1
while m + m <= a:
m = m + m
k = k + k
a = a - m
q = q + k
return q
def gcd(a, b):
while b != 0:
r = a % b
a = b
b = r
return a
def make(n, d):
if d < 0:
n = 0 - n
d = 0 - d
a = abs(n)
g = gcd(a, d)
if n < 0:
return [0 - quotient(a, g), quotient(d, g)]
return [quotient(a, g), quotient(d, g)]
def plus(x, y):
return make(x[0] * y[1] + y[0] * x[1], x[1] * y[1])
def minus(x, y):
return make(x[0] * y[1] - y[0] * x[1], x[1] * y[1])
def times(x, y):
return make(x[0] * y[0], x[1] * y[1])
def over(x, y):
return make(x[0] * y[1], x[1] * y[0])
def is_zero(x):
return x[0] == 0
def digits_value(s):
if s == "" or len(s) > 9:
return -1
n = 0
for c in s:
if not c in "0123456789":
return -1
n = n * 10 + int(c)
return n
def parse(s):
sign = 1
if len(s) > 0 and s[0] == "-":
sign = -1
s = s[1:len(s)]
k = 0
while k < len(s) and s[k] != "/" and s[k] != ".":
k = k + 1
if k == len(s):
n = digits_value(s)
if n < 0:
return []
return make(sign * n, 1)
whole = digits_value(s[0:k])
rest = s[k + 1:len(s)]
part = digits_value(rest)
if whole < 0 or part < 0:
return []
if s[k] == "/":
if part == 0:
return []
return make(sign * whole, part)
scale = 1
for c in rest:
scale = scale * 10
return make(sign * (whole * scale + part), scale)
def less(x, y):
return x[0] * y[1] < y[0] * x[1]
def isqrt(n):
if n < 2:
return n
x = n
y = quotient(x + 1, 2)
while y < x:
x = y
y = quotient(x + quotient(n, x), 2)
return x
def decimal(x, places):
sign = ""
n = abs(x[0])
if x[0] < 0:
sign = "-"
scale = 1
for k in range(1, places+1):
scale = scale * 10
t = quotient(2 * n * scale + x[1], 2 * x[1])
about = ""
if n * scale % x[1] != 0:
about = "about "
whole = str(quotient(t, scale))
part = str(t % scale)
while len(part) < places:
part = "0" + part
if about == "":
while len(part) > 0 and part[len(part) - 1] == "0":
part = part[0:len(part) - 1]
if t == 0:
sign = ""
if part == "":
return about + sign + whole
return about + sign + whole + "." + part
def root_decimal(x, places):
scale = 1
for k in range(1, places+1):
scale = scale * 10
r2 = isqrt(quotient(x[0] * scale * scale * 4, x[1]))
t = quotient(r2 + 1, 2)
about = "about "
if t * t * x[1] == x[0] * scale * scale:
about = ""
part = str(t % scale)
while len(part) < places:
part = "0" + part
if about == "":
while len(part) > 0 and part[len(part) - 1] == "0":
part = part[0:len(part) - 1]
if part == "":
return about + str(quotient(t, scale))
return about + str(quotient(t, scale)) + "." + part
def text(x):
if x[1] == 1:
return str(x[0])
return str(x[0]) + "/" + str(x[1])