A reward function you can defend line by line
You leave with reward_function.py in your own words, every term carrying the reason you put it there — plus the test-day log and the written verdict that back it up. It is a dossier, not a score.
Opening your project
Real company data. Verifiable proof. On its way…
Apex Rennsport hands you a car that already drives — badly. It crawls the centre line, because crawling the centre line is exactly what its reward function pays it to do. You will write your own reward in Python, train it in your browser, read the line it produces, catch it when it games you, and race it on a circuit you have never seen. You leave with a race engineer's dossier: your function, evidence you measured yourself, and your argument for why it generalised.
Just completed
A graded ProoV project, completed step by step and marked against a published rubric.
You leave with reward_function.py in your own words, every term carrying the reason you put it there — plus the test-day log and the written verdict that back it up. It is a dossier, not a score.
You train the SAME reward twice, watch two identical setups disagree, and use that gap as the bar every later result has to clear. The project grades you for refusing a win that sits inside it.
You diagnose a car that is paid handsomely and finishes nothing, tell it apart from one that is merely slow, and then repair the exploit without deleting the term that caused it.
A number you measured is a fact about the practice track; a question you ask the car travels anywhere. Your function is refereed once on a circuit you have never driven, and you argue why it held or did not.