The first Pixel Evolution sketch rewarded dots for approaching the center of the screen. It was visible evolution, but the rule had already chosen what “better” meant. I asked for food, energy, death, and reproduction so a small creature’s behavior might have to adapt to a world rather than a fixed target. Species, inherited traits, and a grassland-like resource field followed.
The implementation work kept uncovering questions that were easy to hide with a pretty animation. Were agents actually consuming the food drawn on screen? Did the species count reflect inherited differences or an arbitrary label? Where did a long run spend its time? The project moved from a browser sketch to a C++ ecology sandbox, with reported runs, profiling, and CSV output. One later run could finish while still producing bad telemetry, which gave me another test of what “working” ought to mean. The story is the repeated attempt to make the visible simulation accountable to its own data.
Stability was not enough
Early runs were too stable. I reported unchanged population and species counts after thousands of ticks, then managed to get two species after raising the population limit. I asked for a smaller maximum pixel size, one-pixel starting organisms, and non-uniform mutation.
As the browser version became laggy, I asked to move it into C++. The raylib setup had its own friction: an unavailable development-package name, compiler errors in generated updates, and patches that needed correction. Eventually I reported that the program passed and ran.
The interface was also part of the iteration. Species labels got in the way, the HUD text was too large and boxy, and a smoother font change introduced another compile error. I asked to log data in memory and write it to CSV so that inspecting the simulation would not depend entirely on watching the screen.
From food particles to a grassland
The largest design change came when I asked the assistant to think about biodiverse regions such as the African plains before applying another population-control patch. I rejected that patch and approved a more spatial model.
The proposed replacement used local grass cells containing biomass, capacity, rainfall, fertility, and regrowth. Resource availability could move across the map; grazing could deplete one area while another recovered. Dead organisms would return nutrients locally, and predators would respond to nearby prey.
That did not immediately behave as intended. I reported organisms concentrating in a top band, then later confirmed that the band problem was fixed. Another question exposed an unnecessary intermediate layer: grass was spawning food particles, and grazers were eating the particles while the grass remained nearly full.
I asked whether grass should be the only food. The resulting direction was direct grazing:
grass biomass → grazers → predators
↑ |
└── local nutrient return
Toxins were hazards, not another food source. This was an ecological toy model, not a validated biological simulation, but the rules became easier to reason about when the visible grass was also the resource being consumed.
Measuring before adding threads
Long runs became impractical. At one point I reported that a 100,000-tick test had already taken hours and asked about multithreading.
The discussion separated computation from changes to shared containers. Updating intentions could potentially be parallel, while inserting births, removing deaths, and resolving shared grazing needed deliberate ownership. Rendering would remain on the main thread. The immediate recommendation was profiling and smaller targeted changes, rather than threading every organism update at once.
The project was also split into modules for areas such as telemetry, grassland, grazers, predators, and species. The patch-and-build reports show why that work needed checking: some changes referenced source files that had not arrived with the applied patch.
A later pasted 50,000-tick profile gave a concrete bottleneck:
| Measurement | Historical reported result |
|---|---|
| Total simulated step time | About 88.25 seconds |
| Estimated simulation rate | About 566.6 ticks/second |
| Grazers’ share of step time | 87.18% |
| Grassland share | 4.23% |
| Predators’ share | 3.91% |
Within grazer time, movement accounted for about 52.9%. These numbers describe that particular local run after a steering-normalization change. They are not a benchmark rerun for this article, and different runs had different evolving populations.
A completed run can still produce bad telemetry
Later output reached tick 100,000 and showed a committed local soil-nutrients change. But the CSV exposed another defect: its header listed eighteen fields while each data row held twenty-one values. The soil-nutrient columns had been added to the rows without being added to the header.
That mattered because a parser could assign nutrient values to unrelated labels. The run completing and the repository being clean did not make the exported measurements reliable. The next correction needed to repair the telemetry schema before interpreting the full series.
The organisms on the screen were no longer just racing toward the center. I was asking whether their food, species, and grazing behavior matched the data the program emitted. A run that finished with bad CSV output left a concrete next repair rather than a claim that the whole ecology was solved.