llmfit's Fallback Estimate Can Be Fourteen Times Too Slow
Decode is memory bound and prefill is compute bound, so Q4 halves one and leaves the other exactly where it was. Both numbers come from formulas the project publishes, which is what makes them checkable.
OSWorld's Hardest Baseline To Beat Is An Agent That Gives Up
Thirty of the three hundred sixty nine tasks are impossible. An agent that answers "cannot be done" to every single one scores higher than twenty one of the twenty six baselines in the paper.
CUA-Lite's Parallelism Number Is Just The Memory Ratio
Frontier agents passed OSWorld's human baseline sometime in the last year. CUA-Lite is not built to measure that. It is built to make the training loop cheap.