REWARD / SELECT
AI Learns To Drive
AI Learns To Drive lets you decide what a better driver is rewarded for
PREVIEW · Official Steam listing; one selected player report · Not played · Sources checked Oct 9, 2026. Our method.

Before judging the driver in AI Learns To Drive, decide what you rewarded it for. The official description lets you mix reward functions for different goals, including speed, caution and drifting. Its evolutionary training then uses the best agents as parents for the next generation.
That makes the objective part of the construction. The fastest-looking run and the behavior you asked the system to favor are not necessarily the same question. The listing does not publish the reward formulas, so it cannot establish which weights reliably produce a particular driving style. It does establish that the player can change the incentives rather than only watch a fixed training demonstration.
Information and control are configurable too. Vision rays, speed, acceleration and wheel angle are named inputs; steering, throttle, brakes and battery boost are named outputs. Campaign challenges unlock additional inputs and outputs, so availability is part of the progression rather than an assumption that every setting is present at the start.
One negative English report posted October 4 says the game was good but reports no progress after 20,000 generations. It recorded 700 minutes at review; the network, reward settings and build are not specified. That cannot establish a universal failure or explain the plateau, but it qualifies the idea that simply running more generations guarantees improvement.
The useful comparison is between what a network can sense, what it can command and what earns it a better score. Changing one of those gives you a reason to compare the resulting driving behavior. A collision is an outcome to explain within that setup, not evidence that evolutionary training has promised a universally competent driver.
Sources & context
The official description identifies a feed-forward network and evolutionary algorithms, but supplies no reward formulas, convergence guarantees or scientific benchmark. Some inputs and outputs are unlocked through campaign challenges. One October 4 English negative account was selected for its training-plateau report, 700 minutes at review/708 at capture; configuration and build unknown. Read in full within the first three of ten recent all-language/all-purchase reports; no prevalence inference.
- Official Steam description Checked Oct 9, 2026
- October 4 negative player report: training plateau Checked Oct 9, 2026
One mechanic, the decision it creates, and a little dry humor.

