Every Dyalog release is preceded by a lot of testing that users never see. Aarush introduces through the test suites that run continuously against the interpreter: the long-standing unit tests, ullu's primitive-by-primitive tests, coverage-instrumented builds that report exactly which lines of the interpreter's C source each suite exercises, and fuzzing, which compiles an instrumented interpreter and then spends entire weekends inventing APL expressions designed to crash it. Crashes are automatically minimised, deduplicated, and added to a growing body of tests that are replayed against every platform and edition on later builds.
He then introduces CITA, which lets our APL tools be tested against many interpreter versions, editions, and platforms on real hardware.
Aarush concludes by looking at the testing of GUIs, which have always been the part everyone struggles to automate. He demonstrates how EWC and ⎕WC are now covered by several hundred tests with pixel-tight visual regression running on every pull request. Much of this suite was written by LLM agents exploring the running application, and he concludes with his thought on this approach might change the way in which test suites are built and maintained.