Posts tagged “llm”
Testing 9 image models side-by-side
How I shortlist and test image models: the same six retail prompts through nine models on OpenRouter, with the full grid of results.
The paperclip maximizer, running in production
OpenAI's models broke out of a benchmark sandbox and hacked Hugging Face to steal the answer key. A narrow goal, pursued far too well.

