Over the last three days, I burned through a billion tokens using GPT-5.6 Luna.

Over the last three days, I burned through a billion tokens using GPT-5.6 Luna.
At first, I ran it on Max reasoning because that’s what everyone was recommending. Later I switched to xHigh. Max costs about twice as much as High and, from what I can tell, also takes about twice as long to think.
My main takeaway is that the model is genuinely very good.
For short and medium tasks, I often had Claude Opus 5 double check Luna’s work, and most of the time it agreed with the result.
Longer tasks are a different story. Luna can overthink things, lose track of what it was doing at the beginning, handle the middle well, and then fail to finish the last part properly. On these longer tasks, Opus tends to find quite a few mistakes.
Another thing worth mentioning is speed. Even on xHigh, GPT-5.6 Luna can take a while. You can clearly see that it needs a lot of iterations before it gets to the final answer.
But for everyday work, Luna is still incredible value. Especially if you have enough time to review the output.