TEA Cipher Avalanche

Ciphers & Bit Tricks

The Super Nintendo runs 32 rounds of the Tiny Encryption Algorithm to paint a 16×16 tile grid: each tile is coloured by bit 0 of TEA(tile_index, key_variant). The multiplier cycles through 16 keys that each differ by exactly one bit — a live demonstration of the avalanche effect. The hot loop uses only 32-bit <<, >>, +, and ^ — no multiply, no divide. Compiled with the llvm-mos-based 65816 toolchain (+mos-a16).

loading core…

Self-running demo — 16 key variants cycle automatically, each differing by 1 bit. (Boot holds while the gate CRC computes.)

Click the screen, then play. (Tab away and it pauses.)

What it is

Each tile at position (tx, ty) is encrypted as plaintext { tx + ty*16, key_variant } using the fixed 128-bit key. Bit 0 of the output's first word determines the tile colour: black (0) or orange (1). Tiles fill at 4 per frame; once all 256 are painted, the demo holds briefly then advances to the next key variant.

// TEA round — 32 iterations, no multiply:
sum += 0x9E3779B9u;   // Fibonacci/golden-ratio constant
v0 += ((v1 << 4) + k[0]) ^ (v1 + sum) ^ ((v1 >> 5) + k[1]);
v1 += ((v0 << 4) + k[2]) ^ (v0 + sum) ^ ((v0 >> 5) + k[3]);
// Under +mos-a16: <<4 = 4×ASL+ROL inline, + = rep/sep bracketed ADC

Compiler stress-test #30 — 32-bit add/XOR/shift, multiply-free

A Round 2 demo targeting the "variable-count shifts" corner of the coverage map. TEA's shift counts are compile-time constants (4, 5) but the 32-bit shift lowering is the corner no prior demo exercises — either inline ASL+ROL chains or __ashlsi3/__lshrsi3 libcalls. This build inlined them.

CornerWhat it exercises
TEA round — no multiplyEach of 32 rounds computes: v0 += ((v1<<4) + k[0]) ^ (v1+sum) ^ ((v1>>5) + k[1]) and symmetrically for v1. Every operation is a 32-bit left-shift-by-4, right-shift-by-5, add, or XOR — not a single multiply or divide. This is the codegen corner all prior demos skip: dense 32-bit arithmetic without any libcall.
Constant shifts, inline expansionThe <<4 and >>5 on uint32_t compile to inline ASL+ROL/LSR+ROR chains under +mos-a16 at -Os — the disasm gate confirmed __mulsi3=0, shift_libcalls=0, rep/sep=22. The 22 rep/sep pairs come from every 32-bit add and XOR needing a REP #$20 (switch to 16-bit A) / SEP #$20 (back to 8-bit A) bracket for each 16-bit half.
Avalanche visualised256 canvas tiles are coloured by bit 0 of TEA(tile_index, key_variant). Each of 16 key variants differs from the previous by exactly one key bit. Correct TEA produces completely uncorrelated patterns for each key — a single flipped bit scrambles all 256 output bits. If the 32-bit arithmetic miscompiles, the pattern shows structure or wrong correlation.
Bit-exact across all modesAll arithmetic is explicit uint32_t with no floating point. The gate CRC encrypts 8 fixed plaintexts and folds v[0]^v[1] → 0xDF0E. Confirmed host == default == +mos-a16 == +mos-xy16 on bsnes-jg. No compiler bug found.

Written in C with the llvm-mos 65816 toolchain. Hit Verify fidelity to reproduce the build gate's WRAM assert (gate CRC 0xDF0E) live in this tab. Disasm gate: __mulsi3=0, shift_libcalls=0, rep/sep=22. No compiler bug found — 32-bit shift/add/XOR codegen correct across all modes.