A right of passage for any computing platform is a port of the Touhou anthem Bad Apple!! This demonstration leverages the video compression techniques invented by Jim Leonard for his demo, 8088 Domination. The Tufty 2350 platform is significantly more powerful than the IBM XT 5150 that Leonard targeted, but that does not make 30 fps video playback a cake walk.
Bad Apple!! runs for 3 minutes and 39 seconds. At 30 frames per second that is 6,572 individual frames. The Tufty2350’s screen has a resolution of 320 by 240, luckily the same 4:3 aspect ratio as the source. For a 1bit, black and white frame that is 9,600 bytes, times 6,572 frames gives an uncompressed size of 63,091,200 bytes — far too large for the Tufty2350’s 16mb of flash. Clearly some form of compression is needed.
Bad Apple!! is a great choice for compression due to its high contrast ratio. Between frames, large areas are unchanged with the differences being constrained to the delta between frames. Leonard goes into detail around the choice of Bayer dithering to reduce the inter frame delta, and the non error diffused Bayer look fits the Bad Apple!! style rather well. The average number of pixels changed per frame is around 16.5%.
The question becomes, how to encode this frame differences. This could be some variable bit rate codec which unpacks the video frames from flash or eprom as the video plays, but that brings with it a lot of challenges — compression, bit rate, timing, preloading storage, buffering, blocking, and so on. This is what made the 8088 Domination technique so appealing.
We can think of each frame as [9600]byte. To move from one frame to the next, rather than a complex video codec, if we write a function that applies the changes to the video buffer in code, eg, framebuffer[77] = 0x55, then playback is applying each transformation function in sequence. As we’re working with 1 bit per pixel, we can update up to 8 pixels at a time with a single store and LLVM is smart enough detect sequential stores and hoist them into 16 or 32 bit writes transparently.
Leonard came to the same realisation and produced an ingenious system to prioritise the most important updates per frame, pushing less important ones to later frames to keep the amount to be changed within the bandwidth that the CGA card could push. Luckily the Tufty2350 has more than enough CPU to update each frame in a few milliseconds so we avoid complex error update propagation. The only real challenge is when the compiled version of a frame update function exceeds 9600 bytes. In those cases those functions are replaced by block copies of a prepared 9600 byte buffer. You can think of these as key frames, which reset the state of the framebuffer to a known value, allowing smaller delta updates to proceed from there.
The video playback is quite simple; starting from a know frame buffer, wait ~33ms, apply a function to the frame which will update just the parts of the frame that have changed, then expand this buffer to the 16bpp format the ST7789V requires and DMA the line to the display. This expansion is performed a line at a time to avoid a 320x240x16bpp temporary framebuffer for the expansion. Expanding a line and DMAing the result to the display can overlap to a degree allowing the next line expansion to start while the previous is being DMA’d to the display.
The final UF2 is around 10mb of code, which the TinyGo compiler handled like a champ. Flashing a 22mb UF2 is less enjoyable, but with patience, it is possible. The encoding reduces the frame data by 83.56%, roughly 6.1:1 compression with no loss of quality.
A video demonstration is available on Bluesky. Note that the audio was played offscreen via a bluetooth speaker rather than included in the demo itself.


