I did find it fascinating, and love the whole "TROH-pic" conceit, you are good. I also love you indulging My Cousin Vinny. I always wrote that movie off, but Antonella is a fan, and recently we watched it and I realized how wrong I was, what a delight.
As for the accuracy of the chatGPT model, I have been experimenting with Madden 2001 game play and summarizing games, and it is pretty interesting how much it confabulates depending on thinking time, so I really like this theory you are presenting. I find the more time it spends trying to calculate Standings and Playoff Picture scenarios (which it is truly challenged by, despite what would be evident math for a machine like it) the more likely it is somewhat accurate---but damn it always messes something up, kinda like you are suggesting.
Doing this for movies though is so much more accessible and fun, no one gives a shit about Madden 2001 anymore cause they are Elden Ring philistines.
Anyway, this is awesome and i am gonna have to play with that Rewind tool you are building for fact checking cause MBS and I use ChatGPT occasionally for our Film Podcast, so this would be most useful so we don't make egregious mistakes---but watching the film in its entirety rather then depending on YT clips cuts a lot of that AI dependence I find :)
It stood out for me that the ChatGPT 5.2 Pro that "nails it" did so with so many citations to a single source website, and cited it a lot. Is there perhaps some effect here where the LLM's good performance is really driven by that other underlying source?
It's still useful to have an LLM pick that out, maybe, but would one be better off just finding that via a web search and going straight to Wikiquote?
Your opening sentence is oddly similar to the opening lines that Adam Curtis used in his HyperNormalisation (2016) — "We live in a strange time" — and Can't Get You Out Of My Head (2021) — "We are living through strange days".
Thanks for your incredibly dense (in a good sense) exploration of the limitations of LLMs.
Thank you for this in-depth reporting. The question of course is: if LLMs are only reasonably good when each inference takes 15 minutes, what does that tell us about the economics of it all? Is the race to massively build data centres perhaps based on the insight we will all need the max performance and we will all be happy to pay $200/month for it?
The combination of having a taste in movies as well as going in deep at LLMs is probably unique. My personal movie tastes are probably comparable. We could possibly discuss the film L’Envitation from Claude Goretta as well as (rare) top performances by Stallone (Cop Land) and Cruise (Magnolia) etc.
I think it does gives some insight into why they are chasing datacenters, though I think the idea is that the price of the compute will come down eventually they just want to lock in their market for when it does. I also should note the little AI fact-checker I built to help me with this ran on Gemini Pro 3 for about 10 cents a pop on the API, so it is possible to write stuff directed for a specific purpose that uses compute more efficiently, it just is hard to do that in a more general way. Those two visions are a little bit at odds. Is this like the personal computer, where people write programs that deal with narrower issues efficiently on top of an LLM substratum? Or is it like Google, where the interface and the code is altogether, and it has to be a general use thing?
On Stallone I have to say I've always liked that Cop Land performance, but I watched the original Rocky the other day, and I was stunned at how well he portrayed this terminally awkward guy. Rocky was meant to be a repudiation of the Mean Streets era of film but in retrospect it's very much of that time in I think some neat ways.
True, his performance in the initial Rocky wasn't that bad (the script plays an important part too in how much an actor can shine of course). Cop Land, though, was an amazing performance I think, but maybe that was in part because one did not expect that. BTW, I really liked your exposure of some of the subtleties of My Cousin Vinny.
Gemini Pro also did great compared to Gemini reg. Compute time FTW!
I did find it fascinating, and love the whole "TROH-pic" conceit, you are good. I also love you indulging My Cousin Vinny. I always wrote that movie off, but Antonella is a fan, and recently we watched it and I realized how wrong I was, what a delight.
As for the accuracy of the chatGPT model, I have been experimenting with Madden 2001 game play and summarizing games, and it is pretty interesting how much it confabulates depending on thinking time, so I really like this theory you are presenting. I find the more time it spends trying to calculate Standings and Playoff Picture scenarios (which it is truly challenged by, despite what would be evident math for a machine like it) the more likely it is somewhat accurate---but damn it always messes something up, kinda like you are suggesting.
Doing this for movies though is so much more accessible and fun, no one gives a shit about Madden 2001 anymore cause they are Elden Ring philistines.
Anyway, this is awesome and i am gonna have to play with that Rewind tool you are building for fact checking cause MBS and I use ChatGPT occasionally for our Film Podcast, so this would be most useful so we don't make egregious mistakes---but watching the film in its entirety rather then depending on YT clips cuts a lot of that AI dependence I find :)
It stood out for me that the ChatGPT 5.2 Pro that "nails it" did so with so many citations to a single source website, and cited it a lot. Is there perhaps some effect here where the LLM's good performance is really driven by that other underlying source?
It's still useful to have an LLM pick that out, maybe, but would one be better off just finding that via a web search and going straight to Wikiquote?
Your opening sentence is oddly similar to the opening lines that Adam Curtis used in his HyperNormalisation (2016) — "We live in a strange time" — and Can't Get You Out Of My Head (2021) — "We are living through strange days".
Thanks for your incredibly dense (in a good sense) exploration of the limitations of LLMs.
Thank you for this in-depth reporting. The question of course is: if LLMs are only reasonably good when each inference takes 15 minutes, what does that tell us about the economics of it all? Is the race to massively build data centres perhaps based on the insight we will all need the max performance and we will all be happy to pay $200/month for it?
The combination of having a taste in movies as well as going in deep at LLMs is probably unique. My personal movie tastes are probably comparable. We could possibly discuss the film L’Envitation from Claude Goretta as well as (rare) top performances by Stallone (Cop Land) and Cruise (Magnolia) etc.
I think it does gives some insight into why they are chasing datacenters, though I think the idea is that the price of the compute will come down eventually they just want to lock in their market for when it does. I also should note the little AI fact-checker I built to help me with this ran on Gemini Pro 3 for about 10 cents a pop on the API, so it is possible to write stuff directed for a specific purpose that uses compute more efficiently, it just is hard to do that in a more general way. Those two visions are a little bit at odds. Is this like the personal computer, where people write programs that deal with narrower issues efficiently on top of an LLM substratum? Or is it like Google, where the interface and the code is altogether, and it has to be a general use thing?
On Stallone I have to say I've always liked that Cop Land performance, but I watched the original Rocky the other day, and I was stunned at how well he portrayed this terminally awkward guy. Rocky was meant to be a repudiation of the Mean Streets era of film but in retrospect it's very much of that time in I think some neat ways.
True, his performance in the initial Rocky wasn't that bad (the script plays an important part too in how much an actor can shine of course). Cop Land, though, was an amazing performance I think, but maybe that was in part because one did not expect that. BTW, I really liked your exposure of some of the subtleties of My Cousin Vinny.