It doesn't grade your answer. It reads how you decide.
July 2026
Here is how Generally Critical works. You are handed one real operating scenario, the kind that would normally eat an afternoon of a founder's week, and you make three connected decisions. About ten seconds later you have a score out of a hundred and your rank against every operator who has faced the same scenario. Today you beat 71% of them, or you did not. Either way, you now know something about yourself that you did not know a minute ago.
The rank is the hook, but it is not the interesting part. The interesting part is what the game learns about you while you play.
Every decision you make is scored across four dimensions of judgment: the quality of the call itself, how you handle ambiguity, how you adapt when the situation shifts, and whether you noticed the stakeholders you were about to run over. Play a few rounds and a clear shape emerges, and that shape has a name. You might play like The Marksman, The Diplomat, or The Improviser. It is a portrait of how you actually operate, drawn only from the decisions you made, and it grows sharper every time you return.
That is the part that compounds. A score you glance at once is a party trick. A portrait that keeps refining itself, and eventually flags a blind spot before you walk into it, is a mirror worth coming back to. The daily scenario is the habit. The portrait is the reason.
None of this holds up if the ranking is soft, which is why there is deliberately no model inside the score. It is fully deterministic. Every option is authored with a fixed weight when the scenario is written, so two people who make the same calls receive the same number, today and a year from now. A leaderboard you cannot reproduce is closer to a mood ring, and this is one you can genuinely stand behind.
The decision that made the portrait meaningful came straight out of the data. Early on, almost half of all plays were scoring a perfect hundred, because most scenes had one obviously correct option. That is a quiz, not a measure of judgment. So I changed what authors write. A scene now offers two genuinely different decisions that both deserve full marks, and the four dimensions carry the difference between them. You stopped being graded on whether you were right, and started being graded on how you were right. That is the entire game, and it is exactly what the portrait reads.
The direction from here is the part I am most excited about: leagues of thirty operators you climb week by week, head-to-head challenges where you discover you out-decided a product manager at a company you admire, and a judgment profile that becomes a real, earned signal of how a person thinks under pressure. Every hiring process in the world reaches for that signal and settles for a résumé instead. We are building the thing the résumé was always a poor proxy for.