EML-NOVA · Standard library

Every function is a graph.

The standard library is NOVA’s working corpus: real functions, written as typed graphs, built on one another with Call, and verified before they ship. People read the diagrams; AIs read the same structure.

functions
175
modules
7
test cases
7,000
verified results
37,705
all passing
175/175

Six levels so far. L0 is composed from NOVA’s original primitives. L1 also uses the standard math primitives of Round 13.1: square root, exponential, logarithm, exact maximum and minimum, comparisons and Size. L2 also uses Round 13.2: windows and joins along an axis, positions as values, sorting, running sums and extremes, and Scan — a loop whose body is another function of the library. L3 also uses Round 13.3’s integer indices — Cast turns a computed position into an index, exactly or not at all; Gather and GatherAlong read at indices; ScatterAdd writes at them — and Round 13.4’s While, a loop that runs until a condition says stop, within a stated number of steps. Round 13.5 completes L3 with dense linear algebra: solving linear systems, the Cholesky factorization, and the eigenvalues and eigenvectors of symmetric matrices. L4 adds Round 13.6’s orthogonal factorizations (QR and the singular value decomposition), Round 13.7’s floor division and remainder — calendars, hashes, bins, the greatest common divisor, reproducible random numbers — and Round 13.8’s Fourier transform, on complex vectors kept as two real rows. L5 begins with Round 13.9’s special functions: erf and erfc, log-gamma, expm1 and log1p — the normal distribution, count models and stable losses — and continues with Round 13.10’s distribution functions: the normal quantile and the regularized incomplete gamma and beta functions, for p-values, quantiles and normal random draws. Round 14.1 adds their inverses: the chi-square, gamma, beta and t quantiles. Round 14.2 makes the gamma and beta functions and their inverses differentiable in their shape parameters too, which is what fitting a distribution needs.

How each function is verified

Three checks, all automatic

  1. 01

    The right function

    In exact arithmetic, the graph must equal an independent reference formula, usually the NumPy library call. Exactly, with rational numbers; to 80 digits where roots, exponentials, logarithms, eigenvalues, singular values, Fourier transforms or special functions appear.

  2. 02

    A guaranteed accuracy

    Every float64 result must lie inside an error bound computed from the graph itself, by running error analysis. The bound holds whatever order NumPy adds in. Where a step is handed to a routine — a library’s exp, tanh, erf or log-gamma, LAPACK for linear algebra, or NOVA’s own normal quantile and incomplete gamma and beta functions — the bound rests on a stated assumption about that routine’s error.

  3. 03

    The same answer everywhere

    NOVA’s reference interpreter and its NumPy backend must return bit-identical results.

Before any of this, the signature is proven: NOVA’s shape solver checks every broadcast and every contraction, for every size.

Elementwise

std.elementwise · 25
absL0
(x: f64[n]) → f64[n]

Absolute value, built from Relu. Exact: one of the two terms is always zero.

40exact
squareL0
(x: f64[n]) → f64[n]

Elementwise square.

40exact
cubeL0
(x: f64[n]) → f64[n]

Elementwise cube.

40
lerpL0
(a: f64[n], b: f64[n], t: f64[]) → f64[n]

Linear interpolation from a to b by a scalar t.

40
leaky_reluL0
(x: f64[n], alpha: f64[]) → f64[n]

Leaky ReLU with slope alpha for negative inputs, built from Relu.

40exact
siluL0
(x: f64[n]) → f64[n]

SiLU (also called swish): x times its own sigmoid.

40
gelu_tanhL0
(x: f64[n]) → f64[n]

GELU, tanh approximation, as used in GPT-style networks.

40calls cube
softsignL0
(x: f64[n]) → f64[n]

Softsign: a smooth, bounded alternative to tanh. Calls abs.

40calls abs
clipL1
(x: f64[n], lo: f64[], hi: f64[]) → f64[n]

Clamp every element into [lo, hi] with exact Maximum and Minimum.

40exact
relu6L1
(x: f64[n]) → f64[n]

ReLU capped at 6, as in mobile networks.

40exact
hard_sigmoidL1
(x: f64[n]) → f64[n]

Piecewise-linear sigmoid: relu6(x + 3) / 6. Calls relu6.

40calls relu6
signL1
(x: f64[n]) → f64[n]

Sign of each element (−1, 0 or 1), from two comparison masks. Exact.

40exact
softplusL1
(x: f64[n]) → f64[n]

Softplus in its overflow-free form. Mathematically identical to log(1 + eˣ), which the reference uses.

40calls abs
eluL1
(x: f64[n], alpha: f64[]) → f64[n]

ELU: the identity for positive inputs, α(eˣ − 1) otherwise.

40
mishL1
(x: f64[n]) → f64[n]

Mish: x · tanh(softplus(x)). Calls softplus.

40calls softplus
wrap_angleL4
(theta: f64[n]) → f64[n]

An angle wrapped into [0, 2π): the remainder after dividing by 2π, which takes the divisor's sign, so negative angles come out positive.

40exact
day_of_weekL4
(days: i64[n]) → i64[n]

The day of the week of a date given as days since 1970-01-01, Monday = 0: that day was a Thursday, so (days + 3) mod 7. Dates before 1970 work too, because the remainder is never negative.

40exact
is_leap_yearL4
(year: i64[n]) → f64[n]

1 for a leap year of the proleptic Gregorian calendar, else 0: divisible by 4, except centuries, except every fourth century.

40exact
days_from_civilL4
(year: i64[n], month: i64[n], day: i64[n]) → i64[n]

Days since 1970-01-01 for a date of the proleptic Gregorian calendar, by Howard Hinnant's algorithm: the year is counted from March, so the leap day falls last, and every division is a floor.

40exact
civil_from_daysL4
(days: i64[n]) → i64[n, 3]

The year, month and day of a date given as days since 1970-01-01, by Howard Hinnant's algorithm: floor divisions by the 400-year era, the year of the era, and a year that starts in March.

40exact
euclid_stepL4
(state: i64[2, n]) → i64[2, n]

One step of Euclid's algorithm on pairs stored as two rows: (a, b) becomes (b, a mod b); a finished pair, with b = 0, stays as it is. gcd repeats it.

40exact
gcd_unfinishedL2
(state: i64[2, n]) → f64[]

Whether any pair still has a non-zero second number, for pairs stored as two rows of non-negative integers: 1 or 0. The condition gcd's loop tests.

40exact
gcdL4
(a: i64[n], b: i64[n]) → i64[n]

The greatest common divisor of each pair of positive integers, by Euclid's algorithm: a While loop of euclid_step until every remainder is zero.

40exactcalls euclid_step, gcd_unfinished
complex_multiplyL2
(z: f64[2, n], w: f64[2, n]) → f64[2, n]

The elementwise product of two complex vectors, each stored as a [2, n] array of real parts and imaginary parts: (a + bi)(c + di) = (ac − bd) + (ad + bc)i.

40
log1mexpL5
(x: f64[n]) → f64[n]

log(1 − e^(−x)) for x > 0, accurate everywhere (Mächler's rule): log(−expm1(−x)) below log 2, log1p(−e^(−x)) above. Each naive form loses its digits on one side.

40

Linear algebra

std.linalg · 37
dotL0
(x: f64[n], y: f64[n]) → f64[]

Inner product of two vectors.

40
norm_sqL0
(x: f64[n]) → f64[]

Squared Euclidean norm. Calls dot.

40calls dot
matvecL0
(A: f64[m, n], x: f64[n]) → f64[m]

Matrix times vector.

40
vecmatL0
(x: f64[m], A: f64[m, n]) → f64[n]

Row vector times matrix.

40
outerL0
(x: f64[m], y: f64[n]) → f64[m, n]

Outer product: every product xᵢ·yⱼ, by broadcasting. Exact up to one rounding per entry.

40exact
gramL0
(X: f64[m, n]) → f64[n, n]

Gram matrix of the columns of X.

40
bilinearL0
(x: f64[m], A: f64[m, n], y: f64[n]) → f64[]

Bilinear form xᵀ·A·y. Calls matvec, then dot.

40calls dot, matvec
axpbyL0
(a: f64[], x: f64[n], b: f64[], y: f64[n]) → f64[n]

Linear combination of two vectors with scalar weights.

40
projL0
(x: f64[n], u: f64[n]) → f64[n]

Orthogonal projection of x onto the line spanned by u. Calls dot twice.

40calls dot
rejectL0
(x: f64[n], u: f64[n]) → f64[n]

The part of x orthogonal to u: x minus its projection. Calls proj.

40calls proj
frobenius_sqL0
(A: f64[m, n]) → f64[]

Squared Frobenius norm: the sum of every squared entry.

40
normL1
(x: f64[n]) → f64[]

Euclidean norm. Calls norm_sq.

40calls norm_sq
normalizeL1
(x: f64[n]) → f64[n]

Scale a non-zero vector to unit length. Calls norm.

40calls norm
cosine_similarityL1
(x: f64[n], y: f64[n]) → f64[]

Cosine of the angle between two non-zero vectors. Calls dot and norm.

40calls dot, norm
frobeniusL1
(A: f64[m, n]) → f64[]

Frobenius norm of a matrix. Calls frobenius_sq.

40calls frobenius_sq
jacobi_stepL0
(x: f64[k], R: f64[k, k], b: f64[k], d: f64[k]) → f64[k]

One Jacobi step for A·x = b, with A split into its diagonal d and the rest R: solve each row for its own unknown, using the others' current values. The body that jacobi_solve runs.

40
residual_unconvergedL1
(x: f64[k], R: f64[k, k], b: f64[k], d: f64[k]) → f64[]

Whether x still misses A·x = b, with A given as its off-diagonal part R and diagonal d: ‖b − R·x − d⊙x‖² above 10⁻²⁰·‖b‖² (a relative residual of 10⁻¹⁰), as 1 or 0. The condition jacobi_solve loops on.

40exact
jacobi_solveL3
(A: f64[k, k], b: f64[k]) → f64[k]

Solve A·x = b by Jacobi iteration for a diagonally dominant A, from x = 0 until the relative residual is below 10⁻¹⁰. The diagonal comes from an Iota mask; the loop is a While.

40calls jacobi_step, residual_unconverged
solveL3
(A: f64[k, k], b: f64[k]) → f64[k]

The solution of A·x = b, by Gaussian elimination with partial pivoting (the Solve primitive).

40
inverseL3
(A: f64[k, k]) → f64[k, k]

The inverse of A: Solve against the identity, which is built from two Iotas.

40
choleskyL3
(A: f64[k, k]) → f64[k, k]

The Cholesky factor of a symmetric positive-definite A: the lower-triangular L with L·Lᵀ = A.

40
cholesky_solveL3
(A: f64[k, k], b: f64[k]) → f64[k]

Solve A·x = b for a symmetric positive-definite A the way it is done in practice: factor A = L·Lᵀ, then one forward and one back substitution.

40
logdet_spdL3
(A: f64[k, k]) → f64[]

The log-determinant of a symmetric positive-definite A, from its Cholesky factor: twice the sum of the logs of L's diagonal. No determinant is ever formed, so nothing overflows.

40
eigvalshL3
(A: f64[k, k]) → f64[k]

The eigenvalues of a symmetric matrix, ascending.

40
eigh_vectorsL3
(A: f64[k, k]) → f64[k, k]

The unit eigenvectors of a symmetric matrix, as columns in ascending order of eigenvalue, each signed so its largest component is positive.

40
ridgeL3
(X: f64[n, p], y: f64[n], lam: f64[]) → f64[p]

Ridge regression: the coefficients β minimizing ‖X·β − y‖² + λ‖β‖², from the regularized normal equations. λ > 0 keeps them solvable for any X.

40
qr_qL4
(A: f64[n+k, n]) → f64[n+k, n]

The orthonormal factor Q of A = Q·R, for a tall matrix of full column rank: its columns are an orthonormal basis of A's column space.

40
qr_rL4
(A: f64[n+k, n]) → f64[n, n]

The triangular factor R of A = Q·R, upper triangular with a positive diagonal.

40
lstsqL4
(A: f64[n+k, n], b: f64[n+k]) → f64[n]

Least squares: the x minimizing ‖A·x − b‖ for a tall A of full column rank, through QR: R·x = Qᵀ·b. Unlike the normal equations, it does not square A's condition number.

40
projectorL4
(A: f64[n+k, n]) → f64[n+k, n+k]

The orthogonal projection onto A's column space, Q·Qᵀ from the QR factorization.

40
svdvalsL4
(A: f64[n+k, n]) → f64[n]

The singular values of A, largest first.

40
norm2L4
(A: f64[n+k, n]) → f64[]

The spectral norm of A: its largest singular value, the most A can stretch a unit vector.

40
cond2L4
(A: f64[n+k, n]) → f64[]

The condition number of A in the 2-norm, largest over smallest singular value: how much A can amplify a relative error.

40
nuclear_normL4
(A: f64[n+k, n]) → f64[]

The nuclear norm of A: the sum of its singular values, the convex stand-in for rank in low-rank recovery.

40
pinvL4
(A: f64[n+k, n]) → f64[n, n+k]

The Moore–Penrose pseudo-inverse of a tall A of full column rank, from its SVD: V·diag(1/σ)·Uᵀ. It maps b to the least-squares solution.

40
polarL4
(A: f64[k, k]) → f64[k, k]

The orthogonal factor of the polar decomposition A = W·H: the orthogonal matrix nearest to A, U·Vᵀ from the SVD.

40
rank1_approxL4
(A: f64[n+k, n]) → f64[n+k, n]

The best rank-1 approximation of A (Eckart–Young): σ₁·u₁·v₁ᵀ, from the leading singular triplet.

40

Geometry

std.geometry · 12
sqdistL0
(x: f64[n], y: f64[n]) → f64[]

Squared Euclidean distance between two points. Calls norm_sq.

40calls norm_sq
pairwise_sqdistL0
(X: f64[m, d], Y: f64[k, d]) → f64[m, k]

Squared distances between every point of X and every point of Y, with one matrix product.

40
centroidL0
(X: f64[m, d]) → f64[d]

Centroid (mean point) of a set of points.

40
centerL0
(X: f64[m, d]) → f64[m, d]

Shift a set of points so that its centroid is the origin.

40
distL1
(x: f64[n], y: f64[n]) → f64[]

Euclidean distance between two points. Calls sqdist.

40calls sqdist
pairwise_distL1
(X: f64[m, d], Y: f64[k, d]) → f64[m, k]

Distances between every point of X and every point of Y. Calls pairwise_sqdist, and clamps at zero before the square root.

40calls pairwise_sqdist
segment_lengthsL2
(P: f64[n, d]) → f64[n−1]

Lengths of the segments of a polyline through the points P, in order.

40
polyline_lengthL2
(P: f64[n, d]) → f64[]

Total length of the polyline through the points P. Calls segment_lengths.

40calls segment_lengths
path_arclengthL2
(P: f64[n, d]) → f64[n]

Arc length along the polyline at every point, starting from 0: the running sum of the segment lengths. Calls segment_lengths.

40calls segment_lengths
polygon_areaL2
(P: f64[n, 2]) → f64[]

Signed area of a polygon by the shoelace formula: positive when the vertices run counter-clockwise. The next vertex of each one comes from joining P[1:] and P[:1].

40
nearest_labelL2
(q: f64[d], points: f64[m, d], labels: f64[m]) → f64[]

One-nearest-neighbour classification: the label of the point closest to q. ArgMin gives the position; a mask built from Iota reads the label there.

40exact
knn_regressL3
(q: f64[d], points: f64[m, d], values: f64[m], k: f64[]) → f64[]

k-nearest-neighbour regression: the mean value of the k points closest to q (ties by index). The ranks of the distances become a mask, so k can be an input.

40calls rank

Statistics

std.stats · 43
varianceL0
(x: f64[n]) → f64[]

Population variance, two-pass: the mean first, then the mean squared deviation.

40
covarianceL0
(x: f64[n], y: f64[n]) → f64[]

Population covariance of two paired samples.

40
weighted_meanL0
(x: f64[n], w: f64[n]) → f64[]

Weighted mean with positive weights. Calls dot.

40calls dot
stdL1
(x: f64[n]) → f64[]

Population standard deviation. Calls variance.

40calls variance
sample_varianceL1
(x: f64[n]) → f64[]

Sample variance, dividing by n − 1. The count comes from the Size primitive. Calls variance.

40calls variance
zscoreL1
(x: f64[n]) → f64[n]

Standard scores: each element's distance from the mean, in standard deviations. Calls std.

40calls std
cov_matrixL1
(X: f64[m, d]) → f64[d, d]

Population covariance matrix of the columns of X. Calls center; the row count comes from Size.

40calls center
correlationL1
(x: f64[n], y: f64[n]) → f64[]

Pearson correlation of two paired samples. Calls covariance and std.

40calls covariance, std
logsumexpL1
(x: f64[n]) → f64[]

log Σ eˣ, computed stably by shifting by the maximum first.

40
minmax_scaleL1
(x: f64[n]) → f64[n]

Rescale to [0, 1] by the minimum and the maximum.

40
quantileL2
(x: f64[n], q: f64[]) → f64[]

The q-quantile with linear interpolation (NumPy's default): sort, then weigh each sorted value by a tent 1 − |i − h| at position h = (n − 1)·q.

40
medianL2
(x: f64[n]) → f64[]

The median: the 0.5-quantile, the mean of the two middle values when n is even. Calls quantile.

40exactcalls quantile
iqrL2
(x: f64[n]) → f64[]

Interquartile range: the 0.75-quantile minus the 0.25-quantile. Calls quantile twice.

40calls quantile
madL2
(x: f64[n]) → f64[]

Median absolute deviation, a robust spread: the median distance from the median. Calls median and abs.

40calls abs, median
trimmed_meanL2
(x: f64[n]) → f64[]

Mean without the smallest and the largest value, which makes it resistant to one outlier at each end.

40
max_drawdownL2
(x: f64[n]) → f64[]

Maximum drawdown: the largest fall from a running peak to a later value, as used for price series.

40exact
histogramL3
(x: f64[n], edges: f64[m]) → f64[m−1]

Counts per bin for increasing bin edges, with NumPy's rules: half-open bins, the last one closed, values outside the edges not counted. Each value's bin comes from comparisons, becomes an integer slot with Cast, and is counted with ScatterAdd.

40exact
segment_sumL3
(base: f64[m], seg: i64[n], x: f64[n]) → f64[m]

Add each value into its segment: x[k] goes to base[seg[k]], and segments that repeat accumulate. A ScatterAdd.

40
sort_by_keyL3
(keys: f64[n], values: f64[n]) → f64[n]

Reorder the values by the stable order of the keys: ArgSort the keys, Cast the positions to indices, Gather the values.

40exact
rankL3
(x: f64[n]) → f64[n]

The rank of each value, 0 for the smallest; ties are ranked in input order. The argsort of the argsort.

40exact
gaussian_logpdfL3
(x: f64[k], mu: f64[k], S: f64[k, k]) → f64[]

The log-density of a multivariate normal N(μ, S) at x, computed the stable way: one Cholesky factorization gives both the log-determinant and, by a triangular solve, the quadratic form.

40
mahalanobisL3
(x: f64[k], mu: f64[k], S: f64[k, k]) → f64[]

The Mahalanobis distance of x from μ under covariance S: the length of x − μ after whitening by S's Cholesky factor.

40
pca_explained_varianceL3
(X: f64[n, p]) → f64[p]

Principal component analysis, as the share of variance along each principal axis, largest first: the eigenvalues of the covariance matrix over their sum.

40
pca_componentsL4
(X: f64[p+k, p]) → f64[p, p]

The principal axes of a data set, as columns, largest variance first: the right singular vectors of the centred data, each signed so its largest component is positive.

40
bin_indexL4
(x: f64[n], lo: f64[], width: f64[]) → i64[n]

The bin each value falls into, for bins of equal width starting at lo: ⌊(x − lo)/width⌋, as an integer index. Values below lo get negative bins; nothing is clipped.

40exact
normal_cdfL5
(x: f64[n], mu: f64[], sigma: f64[]) → f64[n]

The normal distribution's cumulative probability, Φ((x − μ)/σ), written with erfc so the lower tail keeps its relative accuracy where 1 + erf would cancel to 0.

40
log_binomialL5
(k: i64[n], m: i64[n]) → f64[n]

The logarithm of the binomial coefficient C(k + m, k), from log-gamma: log Γ(k+m+1) − log Γ(k+1) − log Γ(m+1). The coefficient itself would overflow long before its logarithm does.

40
poisson_logpmfL5
(k: i64[n], lam: f64[n]) → f64[n]

The log-probability of k events under a Poisson distribution with mean λ: k·log λ − λ − log k!, with log k! as log Γ(k + 1).

40
gamma_logpdfL5
(x: f64[n], shape: f64[], rate: f64[]) → f64[n]

The log-density of the gamma distribution with shape a and rate b: a·log b + (a − 1)·log x − b·x − log Γ(a).

40
beta_logpdfL5
(x: f64[n], a: f64[], b: f64[]) → f64[n]

The log-density of the beta distribution on (0, 1): (a − 1)·log x + (b − 1)·log(1 − x) − log B(a, b), with log(1 − x) as log1p(−x) and log B from log-gamma.

40
student_t_logpdfL5
(x: f64[n], nu: f64[]) → f64[n]

The log-density of Student's t distribution with ν degrees of freedom: log Γ((ν+1)/2) − log Γ(ν/2) − ½·log(νπ) − ((ν+1)/2)·log1p(x²/ν).

40
normal_quantileL5
(p: f64[n], mu: f64[], sigma: f64[]) → f64[n]

The normal distribution's quantile: the x with Φ((x − μ)/σ) = p, as μ + σ·Φ⁻¹(p).

40
normal_drawsL5
(state: i64[3], x: f64[n]) → f64[n]

One standard normal number for each element of x, by inverse-transform sampling: Φ⁻¹ of the Wichmann–Hill uniform draws. Reproducible from its state.

40calls uniform_draws
chi2_cdfL5
(x: f64[n], k: i64[]) → f64[n]

The chi-square distribution's cumulative probability with k degrees of freedom: P(k/2, x/2), the regularized lower incomplete gamma function.

40
poisson_cdfL5
(k: i64[n], lam: f64[n]) → f64[n]

The probability of at most k events under a Poisson distribution with mean λ: 1 − P(k + 1, λ).

40
gamma_cdfL5
(x: f64[n], shape: f64[], rate: f64[]) → f64[n]

The gamma distribution's cumulative probability with shape a and rate b: P(a, b·x).

40
beta_cdfL5
(x: f64[n], a: f64[], b: f64[]) → f64[n]

The beta distribution's cumulative probability: I_x(a, b), the regularized incomplete beta function.

40
student_t_cdfL5
(x: f64[n], nu: f64[]) → f64[n]

Student's t distribution's cumulative probability with ν degrees of freedom, through the incomplete beta function: ½·I_{ν/(ν+x²)}(ν/2, ½) on the left, one minus that on the right.

40
binomial_cdfL5
(k: i64[n], m: i64[n], prob: f64[n]) → f64[n]

The probability of at most k successes in k + m independent trials with success probability p (m ≥ 1): I_{1−p}(m, k + 1).

40
chi2_quantileL5
(p: f64[n], k: i64[]) → f64[n]

The chi-square distribution's quantile with k degrees of freedom: the x with F(x; k) = p, as 2·P⁻¹(k/2, p).

40
gamma_quantileL5
(p: f64[n], shape: f64[], rate: f64[]) → f64[n]

The gamma distribution's quantile with shape a and rate b: P⁻¹(a, p)/b.

40
beta_quantileL5
(p: f64[n], a: f64[], b: f64[]) → f64[n]

The beta distribution's quantile: the x with I_x(a, b) = p, the inverse incomplete beta function.

40
student_t_quantileL5
(p: f64[n], nu: f64[]) → f64[n]

Student's t distribution's quantile with ν degrees of freedom, through the inverse incomplete beta function: with z = I⁻¹_{2·min(p, 1−p)}(ν/2, ½), x = ±√(ν(1 − z)/z), negative below the median.

40

Sequences

std.seq · 34
diffL2
(x: f64[n]) → f64[n−1]

First differences: each element minus the one before it. One element shorter than x.

40exact
cumsumL2
(x: f64[n]) → f64[n]

Running total: element i is the sum of x₀ … xᵢ.

40
cummeanL2
(x: f64[n]) → f64[n]

Running mean: element i is the mean of x₀ … xᵢ. The counts 1, 2, … come from Iota.

40
sum_of_squares_toL2
(x: f64[n]) → f64[]

1² + 2² + … + n², with n the length of x (its values are not used): Iota counts 1 … n, each count is squared, and the squares are added. This is the intent the family home follows through EML-U, EML-P and EML-NOVA; for x of length 100 it is 338350.

40exact
trapezoidL2
(y: f64[n]) → f64[]

Trapezoid rule with unit spacing: the area under the samples y.

40
trapezoid_xyL2
(x: f64[n], y: f64[n]) → f64[]

Trapezoid rule over sample points x (any spacing, in order). Calls diff.

40calls diff
moving_average3L2
(x: f64[n]) → f64[n−2]

Moving average over a window of three, at every position where the window fits.

40
conv1d_k3L2
(x: f64[n], w: f64[3]) → f64[n−2]

One-dimensional convolution with a three-tap kernel w, valid positions only — cross-correlation, as deep-learning libraries define it.

40
ema_stepL0
(s: f64[], x: f64[], alpha: f64[]) → f64[]

One step of an exponential moving average: blend the new value into the running one. The body that ema scans.

40
emaL2
(x: f64[n], alpha: f64[]) → f64[n]

Exponential moving average of a series, started at its first value: a Scan of ema_step.

40calls ema_step
discount_stepL0
(g: f64[], r: f64[], gamma: f64[]) → f64[]

One step of a discounted return: this reward plus γ times the return after it. The body that discounted_returns scans, backwards.

40
discounted_returnsL2
(r: f64[n], gamma: f64[]) → f64[n]

Discounted returns of a reward sequence, as in reinforcement learning: a reverse Scan of discount_step.

40calls discount_step
horner_stepL0
(acc: f64[], c: f64[], t: f64[]) → f64[]

One step of Horner's rule: multiply by t, add the next coefficient. The body that horner scans.

40
hornerL2
(c: f64[n], t: f64[]) → f64[]

Evaluate a polynomial at t by Horner's rule, coefficients from the highest degree down: a Scan of horner_step, keeping the last value.

40calls horner_step
euler_stepL0
(x: f64[k], dt: f64[], A: f64[k, k]) → f64[k]

One forward-Euler step of the linear system x′ = A·x over a time step dt. The body that euler_linear scans.

40
euler_linearL2
(x0: f64[k], A: f64[k, k], dts: f64[n]) → f64[n, k]

Integrate x′ = A·x from x₀ by forward Euler over the time steps dts; returns the state after every step. A Scan of euler_step.

40calls euler_step
permuteL3
(x: f64[n], perm: i64[n]) → f64[n]

Reorder a vector by a permutation: element i of the result is x[perm[i]]. A Gather.

40exact
inverse_permutationL3
(perm: i64[n]) → i64[n]

The inverse of a permutation: inv[perm[i]] = i. Each position is scattered to where it came from, then Cast back to integers.

40exact
newton_sqrt_stepL0
(s: f64[], a: f64[]) → f64[]

One step of Newton's method for √a: average the guess with a divided by it. The body that sqrt_newton runs.

40
sqrt_unconvergedL1
(s: f64[], a: f64[]) → f64[]

Whether a square-root guess still misses: |s² − a| above 10⁻¹² of a, as 1 or 0. The condition sqrt_newton loops on.

40exact
sqrt_newtonL3
(a: f64[]) → f64[]

√a by Newton's method, starting from (a + 1)/2 and stopping once |s² − a| ≤ 10⁻¹²·a: a While over newton_sqrt_step until sqrt_unconverged says stop.

40calls newton_sqrt_step, sqrt_unconverged
bisect_cube_stepL2
(lohi: f64[2], a: f64[]) → f64[2]

One bisection step for x³ = a: halve the bracket [lo, hi], keeping the half where x³ − a changes sign. The body that cbrt_bisect runs.

40exact
bracket_unconvergedL2
(lohi: f64[2], a: f64[]) → f64[]

Whether a bisection bracket is still too wide: hi − lo above 10⁻¹²·(|a| + 1), as 1 or 0. The condition cbrt_bisect loops on.

40exact
cbrt_bisectL3
(a: f64[]) → f64[]

∛a by bisection: start from the bracket ±(|a| + 1), which always contains the root, and halve it until it is narrower than 10⁻¹²·(|a| + 1). A While whose state is the bracket itself.

40calls bisect_cube_step, bracket_unconverged
hash_stepL4
(h: i64[], c: i64[]) → i64[]

One step of a polynomial rolling hash: multiply by 131, add the next code, reduce modulo the prime 1 000 000 007. Every value stays below 2⁵³, so the arithmetic is exact. The body poly_hash scans.

40exact
poly_hashL4
(codes: i64[n]) → i64[]

A polynomial rolling hash of a sequence of byte codes: a Scan of hash_step from 0, keeping the last value. Equal sequences hash equally; it is not a cryptographic hash.

40exactcalls hash_step
fftL4
(z: f64[2, n]) → f64[2, n]

The discrete Fourier transform of a complex vector, unnormalized: yₖ = Σⱼ zⱼ·e^(−2πi·jk/n).

40
ifftL4
(y: f64[2, n]) → f64[2, n]

The inverse discrete Fourier transform: zⱼ = (1/n)·Σₖ yₖ·e^(2πi·jk/n), so ifft(fft(z)) = z.

40
power_spectrumL4
(x: f64[n]) → f64[n]

The power spectrum of a real signal: |yₖ|² for its discrete Fourier transform y, the energy at each frequency.

40
circular_convolveL4
(a: f64[n], b: f64[n]) → f64[n]

The circular convolution of two real signals through the Fourier transform: transform both, multiply, transform back, keep the real part. Takes n log n operations instead of n².

40calls complex_multiply
circular_autocorrelationL4
(x: f64[n]) → f64[n]

The circular autocorrelation of a real signal, rₖ = Σⱼ xⱼ·x₍ⱼ₊ₖ₎ mod n, by the Wiener–Khinchin theorem: the inverse transform of the power spectrum.

40
wichmann_hill_stepL4
(state: i64[3], x: f64[]) → i64[3]

One step of the Wichmann–Hill generator (AS 183, 1982): three small multiplicative congruential generators side by side, each sᵢ ↦ aᵢ·sᵢ mod mᵢ. It is the body uniform_draws scans; the element it is handed only sets how many steps there are.

40exact
uniform_drawsL4
(state: i64[3], x: f64[n]) → f64[n]

One uniform number in [0, 1) for each element of x, from the Wichmann–Hill generator started at state: a Scan of wichmann_hill_step, then the fractional part of s₁/m₁ + s₂/m₂ + s₃/m₃. The same state always gives the same numbers, on every backend. Its period is about 7·10¹²; it is not for cryptography.

40calls wichmann_hill_step
shuffleL4
(x: f64[n], state: i64[3]) → f64[n]

A reproducible random permutation of x: draw one uniform number per element, then sort the elements by their numbers.

40exactcalls uniform_draws

Neural networks

std.nn · 13
linearL0
(X: f64[m, k], W: f64[k, n], b: f64[n]) → f64[m, n]

Affine layer: a batch of rows times a weight matrix, plus a bias.

40
mlp2L0
(X: f64[m, k], W1: f64[k, h], b1: f64[h], W2: f64[h, n], b2: f64[n]) → f64[m, n]

Two-layer perceptron with a ReLU between the layers. Calls linear twice.

40calls linear
attentionL0
(Q: f64[m, d], K: f64[t, d], V: f64[t, e], scale: f64[]) → f64[m, e]

Scaled dot-product attention for one head.

40
log_softmaxL1
(x: f64[n]) → f64[n]

Log of the softmax, computed as x − logsumexp(x) without forming the softmax. Calls logsumexp.

40calls logsumexp
layernormL1
(x: f64[n], gamma: f64[n], beta: f64[n], eps: f64[]) → f64[n]

Layer normalisation of one feature vector, with scale γ, shift β and a small ε. Calls variance.

40calls variance
attention_scaledL1
(Q: f64[m, d], K: f64[t, d], V: f64[t, e]) → f64[m, e]

Scaled dot-product attention with the standard 1/√d scale computed inside the graph (Size, then Sqrt). Calls attention.

40calls attention
rnn_cellL0
(h: f64[k], x: f64[m], W: f64[k, k], U: f64[k, m], b: f64[k]) → f64[k]

One step of an Elman recurrent network: the next hidden state from the current one and an input. The body that rnn scans.

40
rnnL2
(h0: f64[k], X: f64[n, m], W: f64[k, k], U: f64[k, m], b: f64[k]) → f64[n, k]

Run an Elman recurrent network over a sequence of inputs; returns every hidden state. A Scan of rnn_cell, differentiable end to end.

40calls rnn_cell
accuracyL2
(logits: f64[n, c], labels: f64[n, c]) → f64[]

Classification accuracy: the share of rows whose largest logit is at the position of the largest label entry (one-hot labels).

40exact
causal_attentionL2
(Q: f64[n, d], K: f64[n, d], V: f64[n, d]) → f64[n, d]

Causal (decoder) self-attention: row i attends only to rows j ≤ i. The mask compares positions from Iota (j > i is the future); future scores are replaced by the row minimum before the stable softmax, so nothing can overflow, and their weights are set to zero after it.

40
embedding_lookupL3
(E: f64[v, d], ids: i64[k]) → f64[k, d]

Embedding lookup: the rows of the table E at the integer ids, in order (repeats allowed).

40exact
embedding_bag_meanL3
(E: f64[v, d], ids: i64[k]) → f64[d]

The mean of the embeddings of a bag of ids: one vector for a whole set of tokens. Calls embedding_lookup.

40calls embedding_lookup
gp_posterior_meanL3
(X: f64[n, d], y: f64[n], Xs: f64[m, d], ell: f64[], noise: f64[]) → f64[m]

The posterior mean of a Gaussian process with a squared-exponential kernel of length scale ℓ and noise variance σ², at new points Xs. The training system is solved by Cholesky (calls cholesky_solve).

40calls cholesky_solve

Losses

std.loss · 11
mseL0
(y: f64[n], t: f64[n]) → f64[]

Mean squared error.

40
maeL0
(y: f64[n], t: f64[n]) → f64[]

Mean absolute error. Calls abs.

40calls abs
hingeL0
(s: f64[n], y: f64[n]) → f64[]

Hinge loss for labels ±1 and real-valued scores.

40
l2_penaltyL0
(W: f64[k, n], lam: f64[]) → f64[]

L2 weight penalty. Calls frobenius_sq.

40calls frobenius_sq
cross_entropyL1
(logits: f64[n], target: f64[n]) → f64[]

Cross-entropy between a target distribution and the softmax of logits. Calls log_softmax.

40calls log_softmax
huberL1
(y: f64[n], t: f64[n], delta: f64[]) → f64[]

Huber loss: quadratic for small errors, linear beyond δ. Calls abs.

40calls abs
binary_cross_entropyL1
(p: f64[n], t: f64[n]) → f64[]

Binary cross-entropy for probabilities strictly between 0 and 1.

40
log_coshL1
(y: f64[n], t: f64[n]) → f64[]

Log-cosh loss, in a form that never overflows. Calls abs.

40calls abs
nll_labelsL3
(logp: f64[n, c], labels: i64[n]) → f64[]

Negative log-likelihood with integer class labels: minus the mean of each row's log-probability at its label.

40
cross_entropy_labelsL3
(logits: f64[n, c], labels: i64[n]) → f64[]

Softmax cross-entropy over rows of logits with integer labels, the way classifiers are trained: a stable log-softmax per row, then nll_labels.

40calls nll_labels
bce_with_logitsL5
(z: f64[n], t: f64[n]) → f64[]

Binary cross-entropy taken from logits: max(z, 0) − z·t + log1p(e^(−|z|)). It never forms the probability, so it stays finite and accurate where sigmoid would round to 0 or 1.

40calls abs