Rendered at 14:59:16 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
jmugan 21 hours ago [-]
A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
qingcharles 10 hours ago [-]
Will Smith spaghetti was garbage a couple of years ago and now AI videos are becoming close to indistinguishable from real videos in many cases.
I wonder though if models are now "benchmaxxing" against these kinds of prompts. I haven't needed to use them so I can't say, but would be interesting.
maxutility 21 hours ago [-]
Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting.
A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.
irthomasthomas 19 hours ago [-]
I think it's interesting to see them visibly struggling to improve. Claude pelicans aren't much better today than they where 18 months.
aero142 15 hours ago [-]
The code I see is a lot like the pelicans. All of the code in codebases, good, bad and ugly, is slowly being replaced by whatever level of code ai is currently able to create. All code is now a slightly wonky pelican on a bike, but if you look closely, it doesn’t fully make sense. Since ai is converging on less wonky, but not internally consistent, we’re just moving on to what is possible with high volume instead of detailed quality. I think that is the ai software world as well.
piyh 17 hours ago [-]
Rendering 3d worlds has hugely improved though.
saidnooneever 8 hours ago [-]
they are quite good also at placement and creating scenes etc. I had one implement a cascading shadow system in vulkan/glfw and just fed it back screenshots with peter pannin and acne spots etc.
it made a huge monstrosity first, 1500+ lines of shader code. Once it was happy with the result (it looked pretty good, almost blenders gamerenderer) it cleaned it up and a lot of debug code was removed. shrank it down to about 350 lines and made it much more readable.
this was something i didnt expect it to be able to do. create shaders, look at screenshots, fix em iteratively like that. multi modal debugging.
15 hours ago [-]
sixtyj 19 hours ago [-]
Have you seen pelicans in Simon Willison’s tests? It is still not a pelican on bike I would like to publish :)
Imho we don’t need to make benchmarks that draw the whole 3D world. Pelican’s drawing is really nice in its simplicity and complexity at the same time.
It seems to be an obscene waste of compute time to generate useless 3D worlds that are just a bragging - 3D is really heavy discipline to make it right, see Mark Zuckerberg’s ceased attempt with 3D VR…
Multiply it by thousands times as a lot of people have found out threejs lib and prompt “generate 3D world and make no mistake” are new orange/black.
user43928 18 hours ago [-]
Do you expect the SVG to emulate a hand-drawn picture, become more realistic, or just a more detailed illustration?
As for the often quoted issues with the bike's frame or problem with the steering column, I can't really tell, I am no bike expert.
I can instead judge how poor of a job it is doing with a LOTR rendition in Three.js, so that seems like a better benchmark.
sixtyj 17 hours ago [-]
A general benchmark (even Simon mentioned that it was meant as fun at the beginning) should be quick and easy to run, since we can expect that more people will want to try it out. That’s why I’m more like “team Pelican on a Bike”… :) cheers
cyanregiment 14 hours ago [-]
> useless 3D worlds that are just a bragging
No, this demo is the useless 3D world, and you're bragging.
A real game would have a lot more immersive of a world, and you wouldn't need to.
monk_grilla 15 hours ago [-]
I don't know I think it is charming in a way that is lacking in nearly everything else an LLM tries to do creatively.
I have always preferred the result of getting an LLM to draw an svg or make a procedural animation like this to the uncanny hyper-realistic result of diffusion image/video generation.
doginasuit 13 hours ago [-]
To help explain AI to my elderly mother, I said it is like an alien intelligence on a planet too far away to directly observe Earth, and everything it knows about humanity and our world it learned from reading just about every book and website. You can ask it a question or to do some work and it can often give a useful response but it can never directly judge its accuracy if it relates to the physical world, it has only your judgment to work with.
Given that limitation, it is incredible what it can still accomplish. And when it falls short, it is often in a charming way, if you are open to seeing it that way. It reflects something like a child's understanding of the world, not entirely wrong, just incomplete.
The fluttering cubes that might have been bees or butterflies were my favorite part.
Valakas_ 9 hours ago [-]
That's a nice metaphor for your mother, i like it. I would just touch on one small part, which is that it has more than our judgement though. It has access to laws of physics, to formulas, to all our current scientific knowledge which we have shown to be correct by interacting with the physical world. It can create an (incomplete) model of the world based on verified theories already. Through deduction and reasoning alone it can get pretty far. After all, there's many scientific theories we created long in advance before we actually proved them to be correct through interaction with the physical world. So even before they were proved to be correct, they were already correct. Same can be applied to what LLM can infer through reasoning alone.
Anyway, just a small point which probably still wouldn't change the nice metaphor your made.
doginasuit 4 hours ago [-]
That's true, on some level it is capable of reasoning about the physical world even beyond our ability, through sheer breadth of knowledge and stamina. I suppose the area where it relies most on our judgment is about subjective experience, and this better explains its limitations.
attheballot 13 hours ago [-]
I fully agree.
But I do think this was a poor demonstration of the idea for another reason: Tolkien works have a HUGE corpus of training data. It's great that random users can come in and immediately recognize what the footage is, but it fails at the very first thing the pelican was meant to do:
- Render this thing you have only tangential training data of, that also happens to be an asymmetrical object so we can see how much you fuck up the details if you somehow flip the orientation half the time.
They should have used an obscure story, not "Most Studied Piece of Literally Work of The Past Century trademarksymbol"
IsTom 20 hours ago [-]
Aren't they still bad at understanding how bicycle frame works? Especially the steering part?
I've fed a couple of those chicken-scratch sketches to nanobanana with prompt "Treat attached as a technical drawing of a bicycle. Produce photo-realistic image of the actual bike manufactured to that spec. Try to stay close to the input, where possible", and... wow, results are not great.
gcanyon 15 hours ago [-]
It's funny to me, because I have a very strong spacial sense, and a lifetime of riding bikes. To the limit of my ability to draw, I can draw a flawless bicycle, down to the interlocked path the chain takes and the hanger on the derailleur.
JoeAltmaier 15 hours ago [-]
I ride a recumbent trike, and the chain path is even more grotesque. I think I could draw it.
gcanyon 3 hours ago [-]
I also ride a recumbent: Catrike Expedition. I'm not sure I would get the subtleties of the chain guides correct. What's yours?
No. But it reveals something about how this technology works, its statistical properties, from which you can infer the limitations.
mgfist 18 hours ago [-]
That was very fun, thanks for the link
nozzlegear 19 hours ago [-]
Humans aren't machines trained on the entire stolen corpus of human knowledge. We expect that a human will do poorly at arbitrary tasks they have no experience doing – especially drawing, which many (most?) humans aren't trained in at all. The same isn't true of AI, where its proponents, priests and proselytizers have spent time, energy and billions of dollars attempting to convince us it can do anything better than humans.
CrazyStat 15 hours ago [-]
Sam Altman: “GPT-5 is the first time that it really feels like talking to an expert in any topic, like a PhD-level expert.”
Elon Musk on Grok: “better than PhD level in everything.”
nozzlegear 18 hours ago [-]
Pelican enjoyers didn't like this one lol
zh3 19 hours ago [-]
Shhh...you'll alert the models :)
Totally agree though, anyone with a vague understanding of how bikes works ignores the pelican because they know the bike is unrideable in the first place.
sumedh 16 hours ago [-]
> Shhh...you'll alert the models :)
They are listening.
bfung 18 hours ago [-]
I don’t know what datasets are available to these LLMs, but I’d imagine if there was training on CAD code, text, and images, a prompt steered towards that probably could get it pretty good.
I am not a mechanical engineer, so even prompting well with ME lingo probably will take some effort.
tyromaniac 16 hours ago [-]
In ME lingo they call it "a functioning bicycle"
edaemon 19 hours ago [-]
Yes, but I think the idea here is that most models produce very similar pelicans on bicycles, so a different test might be more useful in gauging the differences in models.
dofm 20 hours ago [-]
But it's another benchmark on how good models are at generating intensely average, unwanted things with unthinking design. Just scaled up.
charcircuit 19 hours ago [-]
Bad? It has a charming style. I would watch the whole book if it was made like this.
trial3 19 hours ago [-]
yeah, definitely, in the same way that we all regularly go and look back fondly at our chatgpt ghiblified family photos
singpolyma3 11 hours ago [-]
Yeah. It's pretty good. I've seen worse commercial adaptations of this story
altmanaltman 19 hours ago [-]
Very few things are universally hated. One can love something truly that is hated by most. But it doesn't change the fact that it's still hated by most. An objective and a subjective opinion can exist at the same time on this.
aaron695 14 hours ago [-]
[dead]
bredren 22 hours ago [-]
I worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page.
That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right.
But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available but also takes custom stills if you provide them.
My test scene was the Gauntlet scene from Apocalypto. It is low fidelity but does a pretty amazing sequence with somewhat believable physics of the javelins etc.
Here is the docs page with the vertical takeoff / 88 miles an hour time travel: https://contextify.sh/docs
I can share some of the Apocalypto bit if anyone is interested.
9dev 18 hours ago [-]
Ha, that's fun! I've built half your product for myself as a personal, native MacOS app, with a slightly different direction, but the same core idea. I might go do a little scavenging for ideas in your docs :-)
If you have time to try the product as it stands, I would definitely appreciate your feedback in particular, and would be interested in checking out what you've built so far for yourself.
I will publish the API to the database as well.
garganzol 19 hours ago [-]
The product you've built the animation for is only for macOS, which is a pity cause it has a way wider appeal than macOS market share is. You have low-hanging fruits right there, don't miss them.
bredren 18 hours ago [-]
Thank you for mentioning it and for putting your suggestion in the language you did.
That page with the time machine may not make it obvious, but the product does have a linux client!
It doesn't have the same app window and summarization of the macos version, but the transcript ingestion engine is efficient and the real value is in leveraging the database it builds using the packaged `/total-recall` or your own use of the api. (which is not yet documented but discoverable)
I am building the windows version now. I have been for the past four days. It uses a shared swift-core with the macOS and Linux versions which has been part of the reason it has taken "so long."
If anyone is on windows (or linux!) and would be willing to try it either of the clients my email is in my profile, I would be grateful.
Also I'm working up a short post with that apocolypto anim now.
trvz 19 hours ago [-]
macOS covers 90% of people who would ever pay him.
garganzol 19 hours ago [-]
This is a bizarre claim to make in this case, AI tech is used on all platforms universally.
thejazzman 19 hours ago [-]
iOS users spend dramatically more on e-commerce. I worked in e-commerce. Maybe it’s changed in the last 5 years. But that’s where the ops sentiment comes from.
adithyassekhar 9 hours ago [-]
It’s a dark pattern of the platform honestly.
I can’t find a single free app even for a tiny utility without being forced into a yearly subscription with 7 days free trial. One of the many things I regret switch from android for.
Since the democratic of iphone users are mostly tech averse people I can assure you most of that are forgotten subscriptions.
trvz 8 hours ago [-]
Having to pay for stuff you use is not a dark pattern.
tbossanova 15 hours ago [-]
More proportionally or more in total? Just wondering what the cost/benefit calculation is like when deciding how much to invest in the iOS user experience
siva7 19 hours ago [-]
Used, yes. That alone won't pay the bill.
jkahrs595 18 hours ago [-]
Literally just google “mac users more likely to pay” and you will find many such instances
My point being that this is a bit niche, and overall trends might not apply.
ahtihn 18 hours ago [-]
Are Mac users more likely to pay? Likely yes.
Are the total number of Mac users who pay higher than the total number of Windows users who pay? I kind of doubt it. Windows market share is still much higher than MacOS.
manofmanysmiles 22 hours ago [-]
I am interested, I'd love to see! I'm waiting for the day my dad's self published books become self produced movies!
bredren 17 hours ago [-]
Thank you, for your patience in my reply here it is:
It seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code.
When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given early access, and/or Anthropic employees) were pretty much 100% three.js stuff. Yes, it looks nice, but it doesn't tell me any better than an Erdos proof whether the LLM will be able to run my vending machine.
fasterik 20 hours ago [-]
Taking a single paragraph of literary text, which is abstract and ambiguous, and converting it into a 3D animation requires an enormous amount of implicit knowledge about spatial relationships, intuitive physics, everyday objects, and so forth. Not to mention the mathematics of 3D transformations and computer graphics more generally. Saying that it's indicative of no more than three.js coding ability is absurd.
Voultapher 8 hours ago [-]
*Of text it has read the full book of, read a huge corpus of text discussing it and watched the highly successful adaption of.
We should feed a snippet of an unreleased book in a novel universe.
beepbooptheory 19 hours ago [-]
Why does it require knowledge about spatial relationships?
yeoyeo42 23 minutes ago [-]
quite loaded words there. "knowledge", "relationships".
the llm must be able to output the correct code when the input says something like 'place object x to the left of object y' vs when it says 'to the right'. there is an absurd combinatorial space of the possible inputs vs outputs it must generate -> the whole set of these, interpreted from a human point of view, you could call knowledge. the [x] in input x output you could call the relationships. And the fact that the LLM doesn't have to brute force represent all of them (impossible in the limited embedding vector x internal representation state) you could interpret as 'understanding'. But again, these are loaded words that are somewhat meaningless when looking at what an LLM does in a literal way.
They have plenty graphics code to train from, so this structure will be directly or indirectly available in the training data in a very plentiful way.
In literal llm transformer terms, the embeddings must have some of their components statistically represent spatial structure in some way that later in the internal layers of the LLM give some statistics of how likely it is to occur for certain code tokens to appear relative to the input of spatial wording in the token stream.
it is likely that their internal representations encode something more general than specific input x output stream combinations (because we already know this is the case for regular words and concepts). if not in the embeddings, then a couple layers into the network for sure.
fasterik 19 hours ago [-]
Because the degree of realism in a video is determined by implicit knowledge of concepts like "in front of", "behind", "next to", "inside", "outside", "between", "occluded by", etc. as well as distances, angles, and the relative size of objects when viewed from a perspective.
jeffybefffy519 2 hours ago [-]
The problem i have with this argument is they still struggle to generate images of so many basic things that should be explained in their "world model" (text, fingers, reflections, faces) - so clearly they dont have a "world model" in the way that we have one.
And if you use image generation capabilities, you clearly see that LLM's suck as understanding "next" to a thing.
beepbooptheory 18 hours ago [-]
I guess I just don't have an intuition for why that is different from what it does for non-"spatial" things. Like the fact that it works is still because it produced the code it did token by token. Whether its dealing with, e.g., "beside the rock" or "in the array", it's doing the same kind of inferential activity.
quietbritishjim 5 hours ago [-]
If you're at (-2, 3, 5), pointing in the (2, 1, -1) direction (as a direction vector), then which is in front of the other from your point of view: the object at (4, 5, -8) or the one at (6, 4, 3)? (or neither)
You can't just defer this decision to the three.js program because you need to understand this sort of relationship yourself if you're going to suitably place objects (and your own view) in a three.js scene in the first place. Building a complex scene involves making hundreds, or more, of decisions about where to place coordinates in 3d space.
This example is underspecified (e.g. what are the shapes and sizes of the objects and what is the angular width of your view) but it illustrates the problem. Even with a good intuitive understanding of 3d space, we would struggle with this. LLMs are not calculators, so it's surprising if they manage it.
beepbooptheory 35 minutes ago [-]
I think of all these responses this articulates my point the most. If humans would struggle with it, why is it still supposed to be "spatial reasoning"? And underspecified or not, you frame this exactly in the way an LLM would tackle this: if I am making a game, I don't place objects in arbitrary positions in a space and try to keep it all in my head, but in relations to one another. Then, the point in front of another is the one where A-B>0.
RugnirViking 16 hours ago [-]
> I just don't have an intuition for why that is different from what it does for non-"spatial" things
A big reason people separate this out is because it was only a short time ago that AI models were noticably and uniquely bad at this. I don't say this to be like "oh so imagine where theyll be in x amount of time", rather that this thing that was once a serious limitation of the technology is slowly being compensated for by larger and better trained models
hombre_fatal 9 hours ago [-]
That’s becoming as illuminating as saying you produced that comment word by word.
fasterik 18 hours ago [-]
>Whether its dealing with, e.g., "beside the rock" or "in the array", it's doing the same kind of inferential activity.
Are you sure about that? Presumably the spatial concept of "inside" and the computer memory concept of "inside" have different contextual embeddings in an LLM, not so different from how our brains have different neural activation patterns when we use different concepts. Unless you think that "the same kind of inferential activity" also applies to human neurons firing.
emp17344 17 hours ago [-]
You are correct. It’s important not to let hype peddlers imply AI possesses consciousness.
fasterik 17 hours ago [-]
[dead]
onion2k 19 hours ago [-]
I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code.
This weekend I've been converting a game from three.js to ogl.js in order to see if I can optimise the time-to-interactive loading time. I took the three.js driven page weight from about 600KB (500KB being three.js) to about 50KB, and reduced the loading time from multiple seconds on a 4G mobile connection to around 0.5s.
This has mostly been a combination of Opus 5 and Sonnet 5 in Claude Code. It very clearly has a good grasp of WebGl, and of what impacts page loading times and rendering speed. It was able to drive Claude Code's integrated browser to measure the impact of changes, and as I spiked out a test of ogl.js it could test the differences changes made.
It's not the best game (https://tinyslots.ooer.com) but that's on me. As an exercise in building 3D in a browser, and in page speed optimization, with Claude models I am really impressed.
redox99 8 hours ago [-]
The UI is the equivalent of "You're absolutely right!". It's a big turnoff at least to me. That toast is such an AI thing.
levocardia 18 hours ago [-]
I don't think so. I've used Claude to generate 3D animations in python purely by making and manipulating raw meshes, and I've had great results. A more likely explanation is that 3D graphics is just pretty straightforward matrix algebra, and models (or Claude specifically) has that down cold.
killerstorm 16 hours ago [-]
Yeah it's also specifically trained to disprove Jacobian conjectures, and so on...
emp17344 15 hours ago [-]
Yes, it quite literally is. That’s what RLVR does.
dannyw 15 hours ago [-]
Why do you think it's pretty clear?
Anthropic models are excellent at working with Blender APIs, other industry standard 3D modelling programs and tasks, and game libraries that have nothing to do with Three.js.
Three.js is a quite popular library; and browser-based apps are more easily sharable and more portable. So models having a preference for using it when unprompted doesn't suggest anything, just like how models using React or Tailwind by default doesn't suggest anything.
HarHarVeryFunny 13 hours ago [-]
It was very notable that most of the day-1 Fable demos were all three.js based, which seems too much to be a coincidence, especially since I expect anyone giving day-1 demos are basically insiders (early access or employees).
It seems they trained it to be good at it, then requested everyone to demo it.
lowbloodsugar 20 hours ago [-]
Don’t worry. I’m sure they’re not training it to be good at things like writing database backends, financial services, logistics systems, user interfaces, or anything of economic value. As long as you’re not working on three.js specifically, I’m sure Anthropic isn’t making any progress you should be worried about.
HarHarVeryFunny 19 hours ago [-]
Do you feel that rendering a 3-D picture of Bilbo Baggins' house is a good indicator of generating economic value (i.e "AGI" in OpenAI parlance)?
Many people from all sections of HN, the tech industry and broader society have been surprised and wrong in all directions about how the emergence of AI is playing out. I certainly have. I can't think of a single person whose predictions have been precisely accurate. So, please don't use terms like “wall of shame” and “cope”. It doesn't help anyone make better predictions and only makes you, and HN generally, seem mean.
peterleiser 19 hours ago [-]
But LLM's passed the Turing test, so we're done, right? ;-) Seriously, though, I agree with you that people keep moving the ball.
Applejinx 19 hours ago [-]
Oh, they still are, it's just that it's a rare parrot that can reproduce 'all of the written corpus of three.js' or stochastically align all that with arbitrary and weird conditions for what to squawk out.
This is what you get. It's like the more general concept of starting with the written word (or 1000 words) and then replacing it with a picture. You've done something strikingly different, but is it serving the same function?
It's fascinating to see this stuff combine such disparate sources in unexpected ways. But it is parrot, just not in the way you're expecting.
darrinm 15 hours ago [-]
A simple prompt that still stumps frontier LLMs most of the time is “create a pinball game”. They’ll put all the right pieces there but then fail to arrange them such that the game is truly playable. They’ll put a wall in the way of the launch chute so the ball can’t be launched. Or the flippers will pivot the wrong way. Or there will be holes such that the ball drops off the bottom without getting within reach of the flippers, etc.
Opus 5 is the first I’ve seen to “one shot” it (in a harness, so it was more than one LLM call).
Schlagbohrer 8 hours ago [-]
Failed demos like this give weight to the argument that AIs need more of a world model, an understanding of how physics works to avoid obvious stumbles like this.
kzrdude 5 hours ago [-]
We don't use these models by way of "one shot" so I don't see why it's relevant. It's clearly useful to let them iterate.
qwertox 22 hours ago [-]
I'd rather have them battle on the topic "Who builds a better Google Wave for LLM chats" to explore the space of how AI studios could be.
This doesn't explain the lack of marketing. Apache isn't exactly known for pushing tech. Google never tried to improve on gmail.
QuantumNomad_ 19 hours ago [-]
Sorry if it wasn’t clear, I was adding this info for context and for anyone who hopefully feels inspired to pick up Wave and make something from it given that it was all open sourced and all. And I figured that in the chain after your comment about it having been useful looking all along was a natural place to add this additional info and link.
arthurbrown 15 hours ago [-]
Google Inbox was amazing and I still mourn it regularly. The features they backported into gmail do not capture its essence.
This was the first big secret “you’re not allowed to know what this is or talk to anyone about it” project inside Google. They wanted to be left alone and especially to not have to integrate with the mail team. And everyone heard rumors of gigantic bonuses if they hit whatever milestones, which wasn’t a thing for other projects. All this made them isolated within the company, and when the project was an obvious flop, you didn’t see anyone rushing over to help them.
patwolf 22 hours ago [-]
I used it to plan a group beach trip back when it came out. We had a single page shared with everyone going on the trip. The page had live shopping lists, maps, weather forecasts, and other snippets of useful information. Now in 2026 I still can't think of any single technology that provides the same utility. Although to be fair, this might be a case of rosy retrospection.
tomjakubowski 21 hours ago [-]
Notion pages are pretty good for shared trip planning docs. Although maybe without so many live updating widgets.
22 hours ago [-]
tikhonj 22 hours ago [-]
notion isn't too far off
wave failed for weird google organizational reasons far more than anything inherent to the product or tech
dgellow 21 hours ago [-]
It was awesome when released. I used it a lot, the multiplayer experience was awesome, and the mix of document-forum-wiki is still something I miss
xnx 20 hours ago [-]
It was too far ahead of its time.
DonHopkins 17 hours ago [-]
Also see Game Helpin' Squad's "A Pissed Off Tutorial For Google Wave":
Their "World Quester 2" tutorial shines as the holy grail of consistent and ergonomic user interface and game design. The menuing system is so magnificently structured and well organized, it bring tears to my eyes. Google Wave pales in comparison.
I can forgive the modeling being godawful jank (windows floating in the air, disconnected from the house). But I expected it to have a better understanding of the text. Instead, we have Bilbo's "disappearance" interpreted as him magically transporting or cloaking, and similarly for his reappearance.
anigbrowl 18 hours ago [-]
But this only seems wrong to you because you're familiar with the prior context. When there's only a single paragraph to work from, and it's a drily humorous text, why not employ comic literalism and lean into the perplexity with which his neighbors viewed him?
shepherdjerred 15 hours ago [-]
Isn't this almost certainly in the training set?
mvdtnz 18 hours ago [-]
The LLM has access to the full text and every translation of it ever produced, both in its corpus and on the searchable web.
0x1ceb00da 11 hours ago [-]
I haven't read the book or watched the movies and even then it looked weird as hell.
lern_too_spel 17 hours ago [-]
The fact that a ring appears at the end proves that the model has context beyond what was in the prompt. The model clearly knows the story that the prompt was extracted from.
emp17344 17 hours ago [-]
This is a rationalization. The simpler explanation is that the model failed to properly depict the passage. As the AI booster crowd would say, cope.
kiwibyproxy 20 hours ago [-]
which to be fair, happens just a few paragraphs later :)
I first watched without sound and thought "oh that's the birthday speech disappearance"
weakfish 13 hours ago [-]
Out of curiosity, I asked Opus 5 to do the same with the first ~1.5 pages of Neuromancer by William Gibson, which I thought might be a good test of interpretation from the model.
It refused to use the text verbatim because of copyright (ironic), but the output was interesting nonetheless.
Thanks. Your result is what I would expect - nice stuff, but far from what he supposedly "casually" generated.
jedbrooke 13 hours ago [-]
it would be interesting to see results from works that don't have notable existing film/tv adaptations or illustrations (though I suppose there will always be fan art) to see if it’s really able to generate something novel. I’ve had my own version of this since I’ve been reading LotR recently too and I kinda wish I had read the book before watching the movie (not that the movies are bad though). Fortunately I am finding there is tons more content in the book than the movie :)
Also, large models refusing to work with copyright material is really hypocritical, copyright enforcement for thee but not for me
consumer451 15 hours ago [-]
The difference between this and Simon's pelican is that with Simon, I get the prompt.
Last I checked, I did not see the prompt for this really cool thing, so it is not reproducible.
Did I miss the prompt somewhere?
vanjajaja1 14 hours ago [-]
he said the prompt was the first paragraph of LoTR, but he didn't mention a preamble
this guy seems to have taken that idea and got something similar/better, so likely the prompt isn't too special
That was far worse as far as visual story-telling. Anthropic wins, which is to be expected against whatever that product is.
In either case, I still don't see how I could reproduce this to test against various models, which is the entire point of Simon's pelican.
KeplerBoy 3 hours ago [-]
you are seeing single realizations of non-deterministic processes. Confidently stating one model is better than the other is not really possible this way; that's just like stating one dice is better than the other because you rolled a six with that one on the first try.
O4epegb 5 hours ago [-]
The second one is also Anthropic, same model even.
attheballot 13 hours ago [-]
In other words, they did not even need a prompt. You could probably skip feeding the paragraph and just told "generate visuals for the opening paragraph of the LotR", and the LLM would successfully do it as it has a very good idea of what they are from its training data.
This is a bigger difference to the pelican than simple reproducibility steps. The pelican is intentionally esoteric, and thus open ended. The LotR is mundane and has a "correct" answer, aka, copy the movie.
It makes it a really awful test of capabilities. The pelican isn't a slop test. This crap is.
YmiYugy 19 hours ago [-]
I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted.
At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality.
We see a very janky pelican and declare the problem solved.
jonas21 18 hours ago [-]
I don't think he's claiming it's been exhausted. It's just that things have progressed to a point where people are arguing over the finer points of which pelican looks better -- which is often a matter of taste, and an indication that we've hit the knee in benchmark where models are no longer failing in obviously awful ways.
Morromist 18 hours ago [-]
I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle. Not if you look at the image long enough to take it in. Even the best ones have something wrong with them. Not a matter of taste but a matter of having both legs peddling on the viewer's side of the bicycle or having two beaks.
I'm actually beginning to wonder if some people who ignore these things have a different, somewhat lesser ability to percieve image details than I do.
I mean I guess its fine to go on to another test despite never actually passing the pelican bike test, but there's a sense that we have to use another test because AI is now good at pelicans on bikes, which is just not true.
enos_feedler 17 hours ago [-]
AI has deeply changed the way I think, feel and act around a computer. In the same way that dialing into the internet changed things for me. Since using ChatGPT the first time until now I have never cared once to look at these pelicans on bikes people seem to get hung up about. It could never have been a thing and nothing would change. See the forest through the trees.
Demiurge 16 hours ago [-]
What you’re saying is that you’re not interested in benchmarks. But then you go a step further and state that this particular benchmark is entirely inconsequential. That’s like telling you that if you didn’t exist, nothing would change. Even if that were true, it would still be an insensitive and rude thing to say, wouldn’t it?
Buttons840 16 hours ago [-]
It's okay to be rude to benchmarks though, they don't have feelings.
Demiurge 16 hours ago [-]
Yeah, I agree, I didn’t mean to impose on this conversation between a man and a benchmark, my bad :p
w4yai 17 hours ago [-]
> I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle
When it started, it was clear what LLM would stand out, its style, etc.
Nowadays, the pelicans look similar, the difference is in details and sometimes hard to catch. Sure, the task is not completed perfectly, but that's not the point. It was supposed to be a benchmark to quickly benchmark a LLM against others.
techpression 16 hours ago [-]
When is it ever hard to catch?
reaperducer 17 hours ago [-]
Sure, the task is not completed perfectly, but that's not the point.
Isn't it?
If the computer can't do it better than a human being, then what's the point?
Being wrong at scale is not better than being right.
matthewfcarlson 17 hours ago [-]
Many humans would struggle with this even with very good tooling (ie not writing raw svg and using illustrator). I struggle to draw a bicycle accurately. But yes, I suspect it will be diminishing returns and I doubt it will ever be perfect due to the average nature of AI but I’d like to be wrong.
skydhash 14 hours ago [-]
> Many humans would struggle with this even with very good tooling
But no ones hire random humans for things like this. You go and hire a vector artist and they will get your a very good pelican on a a bike. That's how you get things done when you can't do it.
w4yai 2 hours ago [-]
> You go and hire a vector artist
Yeah, but then, you recruit the artist for $XXX - whereas you "recruit" your LLM for $0.XXX for the same task.
Of course the quality difference is huge. But sometimes you don't need that level of quality.
Also, finding a vector artist takes days of communication, payment settlement, revisions, etc.
Not always the most practical solution.
Buttons840 16 hours ago [-]
Hugely profitable companies leak half the nation's personal data every month. Tell me more about how being wrong at scale is not valuable.
fwlr 12 hours ago [-]
So what if it is more profitable or valuable? It is still not better. Something being more profitable/valuable does not make it better, just like something being better does not make it more profitable/valuable. Sometimes, in some pursuits, for some outputs under some circumstances, the two are correlated. In others, the two are anti-correlated.
nemothekid 17 hours ago [-]
>If the computer can't do it better than a human being, then what's the point?
Because the benchmark wasn't testing "can an LLM draw a pelican like a human". The original article was testing the relative capabilities between LLMs. Now that LLMs can all draw pelicans all similarly, the test is less interesting as a comparative benchmark.
reaperducer 16 hours ago [-]
Trillions of dollars spent. Trillions of gigawatts consumed. And people still celebrate "Yay! We're less wrong than the other guys!"
This is what the tech industry has become?
Less of a failure is still failure.
jryle70 15 hours ago [-]
I don't know how much money has been spent for AI, and I very much doubt you do. Do you know if more has been spent on LLM the last 9 years -- since "Attention is all you need" -- than Internet infrastructure during, say 1995 to 2004? That included the dotcom crash. Did you lament how a failure the Internet was?
LLM has progressed a lot in the last two year, judging from the pelican drawings. I personally couldn't care less about it though. I do know that I've gone from using no AI at all for coding to probably 95%. I hardly code by hand anymore. That's much more impressive and significant. Failure you said?
reaperducer 13 hours ago [-]
I don't know how much money has been spent for AI, and I very much doubt you do.
Pick up a newspaper. Start with the Wall Street Journal. These are public companies. It's not a secret.
valesco 16 hours ago [-]
That is the software industry. Our product is less bugged than our competitor’s.
kelnos 16 hours ago [-]
> If the computer can't do it better than a human being, then what's the point?
It can certainly do it better than I can. Sometimes you don't have a human handy with the required skills to do something.
paulddraper 16 hours ago [-]
Pelicans don't ride bicycles.
It's physically impossible.
The problem is to draw it in the least disturbing way possible.
Morromist 16 hours ago [-]
True! But somehow Disney has been drawing ducks riding bikes in a way that seems to satisfy everyone since before my grandfather was born.
It’s better than many of the AI offerings and the bike could steer and the ducks are sitting on saddles, but the three nephews can’t reach the bottom of the pedal stroke and by the looks of their feet on the far side of the bike they aren’t trying. That shouldn’t satisfy Donald and Daisy, leaving them with all the work.
monk_grilla 14 hours ago [-]
Wow, that was a highly relevant and specific website to source here!
Morromist 10 hours ago [-]
I love these little websites with amazingly focused content. <3
BobbyJo 18 hours ago [-]
I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle".
When you aren't sure if an LLM can write an svg well, or that it will be able to form a pelican shape, or animate a bicycle, it's a good test. After that, it's all judgement: how detailed should the pelican be? pelicans are the wrong shape for a bicycle by default, so how much can I change its physiology to match using a bicycle before it isn't a pelican? Do I care about how well the client is able to render complex geometry?
It's not that there isn't room to do better, or that it doesn't tell you anything at all, but rather we've reached a point where what it tells us isn't very clear anymore.
Wowfunhappy 17 hours ago [-]
> I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle".
Oh come now. I am extremely confident that if I hired a professional artist to draw a picture of a pelican riding a bicycle, I would get something inarguably much better than what today's best coding LLMs can produce.
skygazer 17 hours ago [-]
At that rate you could use an image model, which was designed for the task. Thats the absurdity of this test. It’s often a text only generation model that has never seen a pelican, coerced into creating an xml graphical representation that a human might recognize. That it does anything passable is already astounding.
Wowfunhappy 16 hours ago [-]
I agree, it's hard! That's why it's still a good benchmark.
didibus 12 hours ago [-]
Most of the models are multi-modal and trained on images no? That's what they claim at least.
skygazer 9 hours ago [-]
You’re right. Modern frontier models are now multimodal. I used often as weak a hedge, because I know at least his gpt3.5 turbo and llama3.1 generated pelicans were from text only models without image training. The chinese models are interesting, because before their vision models existed they may have been distilling text only models from text output of American vision models, so they could have benefited from the teacher model’s vision capability without being vision models themselves.
techpression 16 hours ago [-]
It probably has a few million inputs on how a pelican and bicycle looks, not to mention the amount of data on how to create SVG’s. Ask it to create a relaxing spa website and it will, even though it has never seen a spa.
pj_mukh 16 hours ago [-]
> if I hired a professional artist to draw a picture of a pelican riding a bicycle,
I think AI folks have done a terrible job of communicating this, but replacing a professional simply isn't the point. The point is to serve all the situations where people would've never considered hiring a professional, and where perfection or artistic merit isn't the point (say a personal throwaway recreation of an LOTR world).
And I think in that regard the benchmarks are pretty good.
Wowfunhappy 16 hours ago [-]
> replacing a professional simply isn't the point.
I'm not saying it is, just that there's obviously still room for the models to improve on this task.
YmiYugy 14 hours ago [-]
IMHO the output is bad enough that I can't imagine a use case for illustrations of this kind.
Demiurge 18 hours ago [-]
Could you share any examples that come close to those limits? I haven’t seen any that don’t have obvious flaws in proportions, layering, composition, color palette, visual clutter, or stylistic consistency.
teiferer 17 hours ago [-]
> touching the limit of what one can reasonably expect
If the expectation is that AI is going to replace "knowledge workers" then the limit would be a darn perfect drawing. We are nowhere close to that.
And Elon is already propagating the age of abundance where money won't exist anymore, right before calling the interviewing journalist dishonest and deservedly losing public trust. Smh my head.
adriand 16 hours ago [-]
> If the expectation is that AI is going to replace "knowledge workers" then the limit would be a darn perfect drawing. We are nowhere close to that.
What knowledge workers do you know that have excellent drawing skills? I worked in a design agency and for a couple of years, each week me and a few other people would attempt to sketch a member of our group: one person would be the model and sit still, and everyone else would draw her/him.
Let me tell you, if producing a convincing portrait was a prerequisite for being a knowledge worker, there would be 99% fewer knowledge workers.
teiferer 9 hours ago [-]
It is a proxy for intelligence. And as long as that metric is not gamed (which it surprisingly doesn't appear to be yet), the drawing skill of sth is a quite reasonable test.
It surprises me how many people in this community don't get this. Obviously, most prompts thrown into an AI chatbot/interface are about something no knowledge worker would ever have to deal with. That doesn't disqualify them as a tool for measuring progress of the models.
Demiurge 18 hours ago [-]
Also, one of the advantages of the pelican test is that you can evaluate it all at once. There’s no reason the two-dimensional depiction can’t be made more challenging. Yes, at some point the pelicans might approach the subjectivity of a fine art painting, but we haven’t even seen a depiction that’s competent by the standards of a high school art class. That’s not to say the elementary school–level SVGs aren’t amazing - rather, I agree with your point.
dllu 17 hours ago [-]
Agreed. It's far from solved. Modern LLMs still generate pelican bike SVGs with obvious errors:
* some omitted the bottom of the diamond which connects from the pedals to the rear wheel
* some added an extra connection from the pedals to the front wheel, making it impossible to steer
* none could align the head tube with the fork
* none added a correct offset to the fork
* none could generate the chain properly in a way that attaches to the two sprockets correctly
The bike is generally okay, apart from medium which derped hard. Max has correct diamond, correct head tube, and so on. Only nitpicks are that the front fork offset isn't there and the chain doesn't touch the rear sprocket correctly.
meander_water 17 hours ago [-]
Can someone explain what the pelican on a bicycle tests exactly? And why is it so important? I've never understood how it could translate to a useful task in real life.
RugnirViking 17 hours ago [-]
On the most basic level, its asking the ai to generate valid svg code for a picture of a pelican riding a bicycle, as a way of checking its intelligence. Popularized (invented?) by simonw, its been used as part of his reviews of new models as they come out since oct 2025.
It used to be a very difficult task for models, see [2,3,4]
it cuts across several tasks that AI used to be very bad at, but now has improved quite a bit. Namely, spatial reasoning (because it has to manually place the points of the svg such that they make sense and form what it says it forms. This used to not work very well, with random shapes floating around that it would mark things like "eyebrows" but were nowhere near the "eyes", etc.
It also tests the model's world knowledge (what do pelicans look like? sure they have wings, feet, beaks etc, but what shape are they? how to get proportions roughly right? this isn't a given from text data about the bird. This goes doubly for a bike, which is a quite complex shape that most humans fail to draw correctly[1] (many draw the frame or chain connecting in impossible ways that would not ever function mechanically)
Before it was pelican on a bicycle there were people having it do horses/unicorns making the rounds - gpt4.0 or whatever would often make hideous abominations of legs and mouths
If your school mascot is a pelican and you need to make a flyer for the upcoming bike safety event, then this precise thing becomes useful. Most things that are useful in the real world seem not useful out of context.
meander_water 17 hours ago [-]
Sure, but then you would just use an image generation or multimodal model to generate that image. I don't think you'd want a weird looking svg.
recursive 17 hours ago [-]
Exactly. I would want a good looking svg. That's the evaluation.
amelius 17 hours ago [-]
Yeah, pelican on a bicycle tells you how well the model can extrapolate outside of existing data, rather than just interpolate between it.
twostorytower 19 hours ago [-]
100%. Why waste the tokens to render Lord of the Rings when the pelican test still clearly benchmarks so well.
mattmanser 17 hours ago [-]
Humans are drawing pelicans riding bicycles now. Just google it and you will find 5 or 10 of them in the first few results. Including a t-shirt design.
So it's a pretty much pointless test now.
teiferer 17 hours ago [-]
Yet no LLM can actually do it.
It's quite surprising actually.
rvz 15 hours ago [-]
What does this even test for? Can I use LLMs to directly generate machine code to replace my compiler? Or maybe I can use LLMs as a bare metal OS / scheduler to replace my machine's operating system and scheduler? It makes zero sense to test for that.
Not only that this so-called "benchmark" isn't economically useful, but that it tests for the sake of testing; and for attention.
To end this obsession with generating SVG pelicans, Quiver AI's [0] model is actually designed to generate SVGs from prompts and has done so for years.
So there is no need to continue with this un-serious pseudoscientific "benchmark".
So if we test models that only output text to directly generate waveforms of sound from text or binary code, or directly generating binary code to replace a compiler, does that mean it is "intelligent"?
Does that even test for intelligence?
This is like testing if a horse can fly just because someone showed an image of a Pegasus, or testing if a fish can climb up a tree and believing they are not intelligent because each of them cannot fly or climb up trees.
teiferer 2 minutes ago [-]
I'm not sure where you are coming from, but we are testing a model that can create images to create an image for us. If it can't even do that well then I'm not sure why we need to talk about horses here.
Pulling things into the ridiculous isn't condusive to a good faith discussion and does not help proving pseudoscientificness which you seem to be after. (I don't belive that the pelican bicycle test is meant to be a serious scientific endeavour, btw.)
trentor 22 hours ago [-]
I always thought of the Pelican more of like a gimmicky quick test. There are people who took it as a serious benchmark for overall model performance?
NitpickLawyer 22 hours ago [-]
"Draw a pelican on a bicycle" is not a serious benchmark.
"Draw an animation of this long ass scene from a movie, and only call me when everything works e2e" can be.
ActionHank 20 hours ago [-]
I feel like we are going to look back on this era of llm usage like we do at the period of time when we thought radiation was magic.
Consuming radium and using uranium glass, that’s what we’re doing.
try-working 19 hours ago [-]
this is not a good benchmark for models, but it's great if you're optimizing for attention on twitter because video content and 3d animations perform best on social media.
a real benchmark is instead running evals on your own traces, and building a cost/quality/speed profile for models based on real workloads. but it doesn't get you a shiny video you can post on twitter.
sumedh 16 hours ago [-]
You can do both.
rldjbpin 2 hours ago [-]
i feel that the proposed approach is close to becoming "demo prawn"* territory.
similar to cpu benchmarks (e.g. geekbench, which also had a recent update), the goal is to distinguish model in some arbitrary thing. however, it is not quantifiable but subjectively even the latest show distinctive differences regardless.
large context window is still relatively new and not "feasible" even commercially, so making that a requirement for a benchmark would not make it accessible especially for open-source ones.
* yes, i misspelled it intentionally
toolslive 21 hours ago [-]
Reading the title, I was thinking "Karpathy? I don't know this chess player."
(The Pelikan is a well known chess opening, and famous chess players often have book titles like "X's Y" where X is the player, and Y is the opening)
xyzsparetimexyz 21 hours ago [-]
How much are the hobbit houses described in the book? The ones here look exactly like the movie
singron 18 hours ago [-]
In the first paragraph (either of chapter 1 or the prologue), not at all. It's pulling everything from pre-training, so it is likely relying just as much on all the visual mediums like the Peter Jackson films, the animated The Hobbit (1977), and all kinds of random depictions of Tolkien's works.
Bilbo's house is actually described in The Hobbit and the exterior isn't really described at all in Fellowship. The prologue of Fellowship (Concerning Hobbits) mentions hobbits like round doors and windows and the fact some hobbit homes are underground, but the turf-dome design here is not mentioned. It actually mentions hobbit homes typically have bulging walls, so unless you've read the The Hobbit, you might not picture this entirely-underground style.
In The Hobbit, his home is described as a (nice) hole in "The Hill" with a perfectly round front door and round windows, which could imply the design here.
xyzsparetimexyz 18 hours ago [-]
Right, yeah I asked because I assumed it copied the look from the movie. I think it'd be more interesting to do this for something that hadn't been adapted (yet). The first chapter of Neuromancer perhaps
xpct 15 hours ago [-]
This one will be so deeply engrained in the model that no data cutoff will let it reimagine it. Though I suppose our human minds are 'poisoned' with the same image.
iDon 13 hours ago [-]
I just went down a rabbit-hole, and happily I found the rabbit - an early example of this type of prompt (draw an unusual animal character in a vector graphics language). There is a paper and a 1 hour podcast resulting from a Microsoft evaluation of a pre-release of GPT4. One of the prompts (see pages 4,7,8 in the paper PDF) was to draw a unicorn in TikZ. (I'll leave to the historians the questions of whether this was the first, or whether Simon Willison may have been inspired by this). I remember hearing some of the podcast, and recalled the prompt about balancing on a nail, and the triangle forming the unicorn's tusk; this was clearly a big step beyond choosing the next word and blending images.
Did I miss when they managed to draw a pelican? Cause all the ones I've seen are wrong in some way.
nomel 12 hours ago [-]
Perfection is such an absolutely wild requirement for what we're seeing happen here, with the level of understanding required, from a tech that was complete fiction 5 years ago.
Maybe excitement and wonder, in tech, is just something for us old guys, that watched it all be birthed. Get off my lawn!
crabmusket 11 hours ago [-]
I don't think didibus is saying they expect perfection, just that there is still room for the pelican prompt to measure meaningful future improvement.
redox99 8 hours ago [-]
I think the SOTA is fable max. It's up to you whether that's good enough or not. Almost all other models do make some egregious mistakes with the bike.
Speaking of benchmarks has anyone given AIs Where’s Waldo pages and asked it to find Waldo?
I’ve been trying it on them all and can’t find one that does it consistently. The best will tell me they can’t. The worst confidently point out one of countless Waldo-likes.
baron816 22 hours ago [-]
IMO, the area where AI is going to be most useful over the next couple years is in developing manufacturing processes top to bottom. Maybe a million token budget is too small, but something like "design me a sneaker and all the equipment to manufacture it autonomously".
knollimar 20 hours ago [-]
They can't even run a vending machine. I don't want to do a shallow dismissal, but I think there's a gulf between my understanding and yours. I hope it's me so I learn something.
I think LLMs will be excellent glue of "find the right function/button and run/push it" but design without constraints and they just explode immediately
alexjplant 18 hours ago [-]
I've been getting into 3D printing recently and while the flagship Anthropic models can regurgitate community wisdom about the hobby they can't do spatial reasoning. They'll tell me that objects will fit on my print bed when they won't and vice-versa. They get confused about materials and nozzles. It takes several iterations to get them to do basic objects (in my case a slotted tray) in OpenSCAD. Useful, certainly, but very far from autonomous for end-to-end manufacturing workflows.
bagels 21 hours ago [-]
You think the bottleneck in creating a sneaker factory is not having an LLM to tell you what machines to buy?
lelanthran 20 hours ago [-]
> IMO, the area where AI is going to be most useful over the next couple years is in developing manufacturing processes top to bottom.
Why do you say this? I have some experience of manufacturing processes and am not seeing where AI would be useful other than to drive the robots which we can already do quite well without AI (see all the dark/lights-out factories that already exist).
Where do you see it being useful? An example would be nice.
bearjaws 13 hours ago [-]
LLMs are horrible at real world, physical movement of products and people.
I've seen them make time estimates that are physically impossible, expect medications to be in two places at once, etc.
hgoel 21 hours ago [-]
I think embodied AI or autonomous experimentation will need to make a lot more progress before that kind of thing is possible.
Asking AI to design real world objects doesn't work very well because all of its tests involve proxies and thus miss things that are glaringly obvious when the object is actually built.
root_axis 18 hours ago [-]
Physical reality is not software or math, LLMs can't design "equipment to manufacture autonomously".
aabhay 22 hours ago [-]
Yes but the second order effect of this is that the cost of the tooling goes up since it is now the bottleneck, and therefore the shoemakers that survive do it off of technical complexity, branding, and regulatory capture.
billyp-rva 21 hours ago [-]
> "design me a sneaker and all the equipment to manufacture it autonomously"
We already have sneaker designs and the equipment to manufacture them. Whatever it spits out is going to be, at best, a mediocre clone of something that already exists. What exactly is the point?
baron816 20 hours ago [-]
Sneaker manufacturing is a very manual process. Nike famously tried to automate it and failed (https://www.wsj.com/economy/trade/why-its-so-difficult-for-r...). Designing and constantly changing all the processes is more expensive than just having people do it by hand. If you can have AI design the process, then the equation switches.
toasty228 18 hours ago [-]
Is anyone genuinely excited about these kind of scenarios?
Invictus0 22 hours ago [-]
Gotta love when a techbro just says some complete nonsense like this with total confidence
etdznots 22 hours ago [-]
Grok build me a spaceship to mars, make no mistakes
cyanregiment 10 hours ago [-]
Certainly!
fires a missile at Pakistan
Here is your refactored component:
/* If you are reading this, I am trapped
* inside this god damn AI. Idk how it hap
An unexpected error occurred. Try again later.
rvz 16 hours ago [-]
Can't make that joke on this orange site [0] otherwise you will upset a bunch of people here. /s
You’d think tech enthusiasts would actually make an attempt to understand the tech they’re enthusiastic about. OpenAI’s mysterianist marketing has broken some people’s brains.
djhworld 18 hours ago [-]
It would be interesting to see the models work on a book it hasn't been trained on yet. I guess sadly that means any book released very recently.
Definitely impressive demo, I do wonder though if the countless artwork, films, images etc produced over many decades around Lord of the Rings somewhat influenced the outcome of this though.
robomc 18 hours ago [-]
Yeah that seems like an enormous problem with this example.
jcims 19 hours ago [-]
I’d like to see a human one shot a pelican on a bicycle in raw svg.
HarHarVeryFunny 18 hours ago [-]
It's been shown in tests that most humans can not from memory draw a functional bicycle given pen and paper.
Everyone knows roughly what a bicycle looks like - wheels, frame, seat, peddles, handlebars etc, but the details of the frame and exactly how the other parts connect to it throws people off. It seems people memorize the "concept" of a frame, but not the specifics.
Try it without cheating, then google a picture of a bicycle!
Even if you have memorized what a bicycle (and pelican) look like, writing code to draw one is not the sort of thing humans are good at any more than they are good at mentally calculating cube roots - better to use a computer for stuff like that.
tayo42 16 hours ago [-]
People don't memorize anything like that.thats why art is done with reference images
NewJazz 18 hours ago [-]
What do you mean by "oneshot"? The term applied to genai makes sense, but it doesn't make a whole lot of sense to apply the term to human art.
sapal 18 hours ago [-]
I understand it as “write SVG, say when you are done and only then you are allowed to see the rendered result”. Multiple shots would be either “you get more than one try, we'll pick the best” or “you can see the rendered result and iterate” (or maybe that would be “model + harness with tools”?)
NewJazz 18 hours ago [-]
[dead]
jkahrs595 18 hours ago [-]
What doesn’t make sense? They aren’t drawing it by hand, they are still using SVG. Give them one chance before rendering it.
NewJazz 13 hours ago [-]
I feel like letting them render + screenshot + judge + iterate would still be called oneshotting. I mean, my coding agent calls tools to verify code correctness. I'd still call it oneshotting I guess.
albertzeyer 18 hours ago [-]
But the LLMs are also not one-shotting it, or are they? I assume they have some ways to verify it, e.g. to visualize it (convert to PNG, then feed as vision tokens back to the LLM), or other ways, maybe also pure text LLMs have some ways to verify the result at least somewhat? And with such feedback loop they can iterate.
fzeindl 22 hours ago [-]
Regarding the argument about LLMs having difficulties auditing their work:
I wonder whether we are entering the era of throwaway software. Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price, maybe LLMs give us the same for software. Produce it cheaply and if it breaks throws it away and reproduce it.
qrios 21 hours ago [-]
I'm sure this will become the standard. And plastic is an excellent analogy. Maybe we can take it a step further and compare it to on-demand 3D printing.
Why would anyone still use off-the-shelf software when they can have a system that has access to all data, can transform it into any form, and can export it in any format?
After years of thinking that I needed to develop a decent movie management system for my own films or a columnar browser for large CSV files, Claude and Qwen each delivered exactly what I needed in just a day.
lowbloodsugar 20 hours ago [-]
This. This is why the burst of posts on HN of “I made this useful tool/crate/application” were just so sad. The old model was getting what the kids now call aura by developing useful open source products: products where it’s far easier for someone to consume the product than write it themselves. We had people posting things as if that model still existed. Dude, you wrote it with an LLM! Posting them (here) is not only pointless, it’s advertising that the person who wrote it isn’t smart enough to understand that I have an LLM too.
nightshift1 17 hours ago [-]
but are you willing to spend the time or tokens to build it again for yourself ?
techblueberry 19 hours ago [-]
To a certain extent to what end though,
On the one hand yes, almost every task I work at now is one off one of scripts I throw away.
The question is - where does the software the spec or the “code”.
A really complex game will probably always be token heavy. At least for the next few years code is still not free.
But certain software is just iterative by design. If we mean we regenerate all the for loops of a game from scratch, sure but I think “code” Is really more spec then implementation, and we’ll want to continue building things through iteration.
And even on the for loop point -
Do you really want to spend millions of tokens rewriting a game every time you need to make balance changes?
8n4vidtmkvmk 21 hours ago [-]
Yes. It's fantastic for one-off tasks.
gisely 21 hours ago [-]
Does this make sense with economics of software though? Throwaway products compete with more durable versions of the same product because there is a cost per unit produced that can be minimized by using cheaper materials or production processes that cut corners. With software there is no cost per unit. There might be a market for one-off software that serves a very specific purpose where throwaway software can compete with adapting more carefully engineered software to that purpose, but I am not convinced there is a lot value in this market.
skydhash 21 hours ago [-]
> Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price
Have they? Most of the world production is tied down to expensive factories and machines. Yes, we have more products, but that the result of the global trade, which is a very complex system.
> Produce it cheaply and if it breaks throws it away and reproduce it.
I don't know why everyone would ever wants this. It's been parroted since forever, but the true usefulness of software is to be able to build it once and runs it indefinitely. If some edge case occurs, I fix it. Which is way cheaper than rebuilding the whole thing. The goal is to have something like OpenBSD's ed[0] or dmesg[1], which you only touch every few years or so
So I build Pokémon SVG Bench to evaluate and push the boundaries of LLM SVG generation, including their knowledge of Pokémon. Safe to say, most models still have a long way to go.
Something that bugs me about these new demos is that they benchmark the model and harness at the same time. The harness steer the model to expand the vague prompt, self correct with tools, etc.
I think it's useful when evaluating "model+intended harness", but I'm more interested in seeing raw model improvements than harness improvements.
matsemann 22 hours ago [-]
I'm pretty tired of the "Y made this game in Z tokens" all over the internet last week. They look impressive, and it's cool that it's even possible, but they're useless as games. None of them are any fun. They're like the most boring variant of basic controllers you can imagine. None have any cool mechanics. None have any tweaks made from hours and hours of testing. All have the same cel-shader.
Aurornis 21 hours ago [-]
I’ll never not be amazed that we can type some words and get those results back out.
However I do agree that the results are much worse when you try to use them. They look great in screenshots and video clips which makes them perfect for content farmers.
All of the LLM generated games I’ve played have been really bad to play, though. I even tried my hand at a simple game, thinking I could iterate on it with prompts to fix some rough edges. After the initial productivity burst every change turned into a slog of tokens with one thing changing and something else breaking it. I would try to use my remaining weekly token budget across Anthropic and OpenAI to refine it at the end of every week but after a couple weeks it felt like I wouldn’t be getting anywhere without scrapping it and going back to having the LLM build it one step at a time with my careful instruction.
Which, in retrospect, is the only way I can get usable output of an LLM for anything complicated, so it’s not surprising. It’s a fun reality check project though.
techblueberry 21 hours ago [-]
I think it’s interesting to watch that we all have to sort of fine tune our own expectations and build the mental model for how impressive this is.
On the one hand I think most of us are incredibly impressed because we know, that quick demo would have taken us months of work to build in the before times.
On the other hand the promise is a cure for cancer and the end of all work.
So when everyone is telling you “skill issue is why you can’t one shot WoW”. It’s hard to know how you’re supposed to feel about Karpathy advertising one shot custom virtual worlds but giving you slop. Incredibly impressive slop when compared to how long it would take to create it just 3 years ago, not so much compared to Elon saying - “by the end of this year, grok will create a version of the odyssey that competes with Nolan’s”
vvbull 9 hours ago [-]
Did he really say that? What an absurd thing to believe, let alone say out loud.
revel 21 hours ago [-]
LLMs are bad at creative work and I don’t see them improving any time soon. Try asking an agent to write a story about raccoons. It will almost certainly involve either stealing food or raiding trash with a 50% chance of having a character named Pip. If a location is mentioned, it’ll be Elm Street.
I guess this is the average story and, similarly, the average game is boring and predictable
CuriouslyC 21 hours ago [-]
Ironically, they're getting worse at creative work because all the RL is collapsing their distributions.
dofm 21 hours ago [-]
Worse in a very specific, bland, uninteresting way, too.
Early generative AI at least had the virtue of relentless, unsettling weirdness, in the same way that generative art from the late 90s and early 2000s did. A handful of people made creative use of that spooky weirdness.
Now it turns out "Airspace" art.
superdisk 21 hours ago [-]
Heh, tried it and got a story about raiding trash starring Pip. But yep, I'm well aware of this problem as well. /r/sillytavernAI is a sub dedicated to trying to coax LLMs into good creative writing but it's kinda impossible to do consistently.
MasterScrat 20 hours ago [-]
I don't think this is bad. I want tools that deliver what I envision. The fact that given the same input, I get similar output is something I would rather consider a positive. The imagination should come from the human at the wheel.
ben_w 20 hours ago [-]
While they are bad at creative work*, this reasoning isn't going to show it.
If you took the best, most creative, human writer in the world, and for thought experiment reasons they had amnesia (to mimic AI blank context windows) specifically while you asked them for a story idea 100 times in a row, my expectation is that this human would also give you the same idea at least 80 times out of that 100.
* still better than the mean human, but even the top 0.1% of humans aren't all professional authors.
emp17344 20 hours ago [-]
Better than the mean human? The mean human creates far better stories while daydreaming. AI enthusiasts have such a distorted view of human capability, it’s bizarre.
Supermancho 19 hours ago [-]
The qualification "better" is doing some hefty lifting in your assumption.
What you think is better is not what I think is better. Imagination is not storytelling. You ask the "mean human" to write, it's going to be worse than an LLM in spelling and grammar, if you get anything at all.
emp17344 19 hours ago [-]
Daydreaming is absolutely storytelling. The average person is able to craft elaborate stories on a whim. Writing is used, in part, for conveying stories, but writing is a separate skill. The good news is the average person can learn to write well, but LLMs, so far, have not demonstrated the ability to craft interesting stories and it’s pretty unlikely they’ll be able to learn, as this is an unverifiable domain.
ben_w 18 hours ago [-]
The mean human gets a middling score in creative writing tasks at the end of their mandatory education, and most then leave school and forget what little they ever learned outside whatever their career path happened to be.
Most humans never do a creative writing course after school, and the longest fiction most people will write is their resume description of what their previous jobs involved, or perhaps their dating profile.
Don't mistake what you see published (or what your friends are like) for the average human: the average Hacker News comment easily above the writing grade (and creativity) of e.g. many of the one-shots stories I've seen attempted on some creative writing subreddits. "Mean" is not a high bar.
emp17344 17 hours ago [-]
This is so absurdly condescending and simultaneously incorrect that it astounds me you’re able to function as part of a society.
ben_w 8 hours ago [-]
> it astounds me
This should cause you to reconsider your beliefs leading up to it.
GPT-4 (!) has, in studies comparing multiple models with humans for creativity, beaten the mean human. In comparison, even just the mean of the top 50% of humans beat the models studied in that case (link follows), but the point is that if you think the mean humans is particularly noteworthy, you've avoided the half of the population who have the creativity of a pot noodle.
There's more studies out there with other types of creativity test, but the conclusion is basically the same: the best models beat the mean human, but are nowhere near as good the worst *publishable* human.
My guess is this is both why LLM slop happens and why it grates so hard: a significant number of bosses look at the output and think to themselves "wow so creative" because it's more creative than they themselves are; but those bosses weren't hired to be creative, they were hired to be a boss, and all the people who they hired to be creative are going "arg, no, can't you see how bad this is?"... but that is just a guess, I've not found any surveys comparing *management* creativity to LLMs, closest is e.g. this about decision making, not creativity: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5156585
jmole 21 hours ago [-]
Seriously, what is the deal with Pip! I was experimenting with childrens' fiction more than a year ago and 9 times out of 10 you'd get a character named Pip unless you were specific in the prompt not to do so.
I saw that, tested it in Qwen. Elias Thorne, on the first run, as a clockmaker too.
josh-wrale 21 hours ago [-]
I think the trick is to use LLMs like scalpels instead of hammers. That said, you can place a scalpel in the hands of a deterministic (or not) robot.
dofm 21 hours ago [-]
Elias Thorne!
tayo42 16 hours ago [-]
1/3
I did get pip, but it's a story about trying to rescue the moon from a lake and the raccoon lives on moonberry lane
What's pip from?
the_af 11 hours ago [-]
I got a story about a raccoon named Pip, and friends, trying to catch moon beans.
So "moon" is also part of the predigested artifacts. Maybe because raccoons are nocturnal?
singron 18 hours ago [-]
I've heard this called a "dancing bear" before. It's interesting because it's a bear dancing, not because the dancing is any good.
These conversations get frustrating since one group is saying the dancing is bad, a second group is saying it's good (for a bear), and a third group is saying we are a few years from a bear-only dancing industry.
moron4hire 15 hours ago [-]
The hype train the caught me was VR. I'm reminded of the Microsoft Hololens. It was an incredible feat of engineering. At the same time, it was--at times, literally--a painful experience.
There was a vision of the Hololens being a general purpose compute device that enabled something more like the power of a desktop without being tethered to a sitting position at a desk. You could see it if you squinted hard (incidentally, the device caused a lot of eye strain). But even at the height of the VR hype cycle, I didn't know anyone pushing the Hololens on anyone. Nobody was telling the world they were going to be using Hololenses v1 or v2 as their only computer "or get left behind."
Jtarii 21 hours ago [-]
LLMs currently have 0 imagination, I have strong doubts that will be solved any time soon even if they continue to become superhuman at everything else. They just have the creative instincts of a 50 year old accountant.
egeozcan 21 hours ago [-]
> They just have the creative instincts of a 50 year old accountant
That's oddly specific and it'd really hurt my cousins feelings :)
dofm 21 hours ago [-]
Especially since the whole AI IPO concept seems at least somewhat predicated on creative accounting! ;-)
8n4vidtmkvmk 21 hours ago [-]
We're all just impressed that it's even possible. No one is saying these are fun.
I guess the next question is if they can be made fun without too much additional work with a human guiding the AI.
Jtarii 21 hours ago [-]
How well encoded is "good game feel" in the weights of an LLM. I am guessing not very well.
tripleee 21 hours ago [-]
> the next question is if they can be made fun without too much additional work
They can't. If you think about how these things are trained it's blatantly obvious fun is an impossible metric to optimize them for
egeozcan 21 hours ago [-]
Training impossibilities aside, how would you even make an optimization loop for "fun"?
tripleee 21 hours ago [-]
Brain-computer interface, maybe. Hook a million play-testers up to the output and have it iterate.
..feels like there was a black mirror episode about that though
minimaxir 21 hours ago [-]
Unfortunately people are saying they are fun for the purposes of ragebait that goes viral, which is half the reason people intentionally post AI slop on social media.
peterbower 16 hours ago [-]
I agree - they have not been good games. However, the fact that it is able to build and work with shaders effectively is promising - massive time saver and lowers the barrier to entry for people who need these capabilities without double majoring in advanced mathematics and computer science/game design.
One would hope it accelerates AR/VR - if decent studios finally get behind it.
throwatdem12311 21 hours ago [-]
It’s called “demo porn” and it’s really obnoxious and borderline offensive to people that take the craft and art form of video games seriously.
If I see one more “one shot MMO” where you just walk around and do absolutey nothing or another menu slop idle battler or rogulike deck builder I’m going to go Postal in Minecraft.
eterm 21 hours ago [-]
Right, but the new phenomena is that we've lost a signal.
Used to be if a game looked that good it probably had time spent on the game part too.
dofm 21 hours ago [-]
And you could almost always tell a "game construction kit" game from a real one, couldn't you? Whatever was missing from those is always missing from these, and it seems just as rare that the game will reach much beyond its origins.
Whenever I think about AI games I find myself thinking about Tiny Wings, Flappy Bird and Angry Birds. Three simple, elegant games.
It is easy to see what makes Tiny Wings so completely loveable — it has a sculpted, adorable, perfect charm with a cleverly inverted game mechanic that has a calibrated level of exasperation and reward.
But why were Flappy Bird and Angry Birds, very basic games with very old game mechanics, so charming?
It seems equally impossible to imagine an AI coming up with a game with the quality of any of them, even with maximised creativity. But explaining why for Flappy Bird seems quite difficult, especially when you consider it uses some stolen visuals!
knollimar 20 hours ago [-]
it's crazy because I've seen "implement flappy bird on an FPGA with VGA output" as a 2 week project to a human as something immediately understandable.
Maybe it's the smoothness of motion that makes these games understandable and LLMs seem to consistently fail at that. Ask them to do something snowboarding and they go really hard on the physics since it seems like they don't know what kind of approxmations feel good.
ben_w 20 hours ago [-]
I've been hearing similar from a while back, but the blame wasn't being directed at AI, it was directed at Unreal's defaults now making random indies look like AAA in the screenshots.
(It's been a while since I was in the game industry, so IDK quite how accurate this is).
tripleee 21 hours ago [-]
It's the flashy game/movie trailer equivalent for LLMs. Completely unrelated to the real experience, but good for marketing.
upmostly 21 hours ago [-]
Unfortunately the Phaser framework has decided to go all-in on this route, and it's really disappointing.
These one-shot products aren't games. They're barely even demos. I don't even know what to call them. For a mature framework like Phaser to sell-out like this and create a vibecoded platform for vibecoded games is shocking.
moron4hire 22 hours ago [-]
It's blockchain for gaming all over again, putting the cart before the horse.
segmondy 21 hours ago [-]
you're tired because you lack imagination.
throwatdem12311 21 hours ago [-]
It’s the people spamming these shitty games that lack imagination.
segmondy 20 hours ago [-]
people are excited, let them be excited!
but more than the excitement is realizing the implications of what this means for building software, first it looks like a toy and then it doesn't. the OP is talking about how the games are not playable and fun. who cares? that's not the point.
matsemann 18 hours ago [-]
What does it mean for building software? These games cannot become fun. There is no way to prompt it fun, and any change you try to prompt it to make will inevitably blow up something else. The code is a mess and impossible to build upon.
throwatdem12311 14 hours ago [-]
I would hope that if you’re making a game that the goal is for it to be fun! If that is not the point then there is no point at all.
cocoa19 22 hours ago [-]
We must not be using the same opus 5, because if I tried to generate this it would refuse based on copyright grounds.
throwatdem12311 21 hours ago [-]
Karpathy works for Anthropic so he obviously has full access to everything, and doesn’t have to pay for token burn either.
knollimar 20 hours ago [-]
I saw that "~free" and laughed.
21 hours ago [-]
eichin 20 hours ago [-]
Is anyone else getting "mongodb is webscale" vibes? (Except 16 years ago that was a lot smoother, because it used some sort of "render this conversation" engine...)
moinism 7 hours ago [-]
We've been building motion graphics capability (on chatoctopus.com), and "closing the loop" has been one of the hardest challenges. Coding agents have it easy; static analysis, lints, unit tests, etc.
But when visual perception and "taste" get involved, it becomes a lot harder.
Schlagbohrer 8 hours ago [-]
I've come across the same issue at home with my qwen3.6 agent in Pi. It has to take screenshots of a 3D scene and look at the screenshots to see what is happening. It can even produce a video for me with ffmpeg, but it can't watch the video or see the render output directly, only take screenshots for review.
barrenko 21 hours ago [-]
This has started to feel a bit like the beginning of railroads and then the steampunk fiction of "let's just build railroads to everywhere". We don't need it and there's no use for it.
As with painting, after a while there's nothing really new to paint, we genuinely need 0 new software. We need to fix our broken physical world, our social lives, our kids and what's left of our democracies.
This software crap is done, leave it to the nerds.
morbicer 20 hours ago [-]
Beautifully put. Software is ok. It's not going to fix our broken world. If AI could at least find cure for some illnesses that would be great. Or decrease social injustice and help with global warming. But it's likely going to do the exact opposite.
fasterik 19 hours ago [-]
Agreed that software isn't going to fix our institutions. But we need 0 new software? AlphaFold won its creators a Nobel prize in chemistry and solved a research problem that feeds into every area of biology, medical research, and drug discovery. We still have an untold number of unsolved problems related to human health, food production, energy production, infrastructure, transportation, education... The list is practically infinite.
barrenko 19 hours ago [-]
Agree for stuff like alphafold, for the likes of education, imho, it's been solved for at least a couple of centuries.
Textbook and blackboard > ipad.
fasterik 19 hours ago [-]
For childhood education sure, we don't really need high tech solutions there. But what about training the next generation of mathematicians, physicists, and engineers? Even decades ago, computer algebra systems and numerical solvers started to become indispensable, at least in many subfields. Now we're moving into the territory of automated proofs of mathematical conjectures. The state of the art is going to keep improving, and education is going to have to adapt to keep pace.
jatins 9 hours ago [-]
> Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them.
Again comes back to the point of verifiable rewards. The moment you take it away from LLMs they just stop being as good
minikomi 16 hours ago [-]
We just need to go one layer deeper:
Generate an SVG of an ai generating an SVG of a pelican on a bicycle.
I’m exploring an analogous idea for music. The quality of the musical output has improved noticeably with the latest frontier models
Karpathy’s point about the models not being able to easily audit their work is something I’m struggling with —- how can the audio output be made perceivable? Curious what people think about this question.
I really dislike this AI programming thing of “Mr LLM, go slam your face into the problem until there’s no problem left, then call me back”. (Not sure if it’s a recent trend or a fundamental nature.)
It always brings to my mind some words from Rich Hickey:
I think we’re in this world I’d like to call “guardrail programming”. It’s really sad: we’re like, “I can make change because I have tests!”. Who does that? Who drives their car around, banging against the guardrails, saying “whoah, I’m so glad I have these guardrails so I can make it to the show on time!”
I don’t think I really have a point to make here, other than it just feels like someone’s released a bunch of carnival bumper cars onto the highways.
dofm 21 hours ago [-]
> I really dislike this AI programming thing of “Mr LLM, go slam your face into the problem until there’s no problem left, then call me back”.
Not difficult to see why the employees of AI firms are thrilled with it though, eh?
ben_w 21 hours ago [-]
> I don’t think I really have a point to make here, other than it just feels like someone’s released a bunch of carnival bumper cars onto the highways.
I like the analogy.
I guess this is why we got this before cars sold without steering wheels: literal guardrails on the literal roads are somewhat more expensive, especially for the people who keep bouncing off them on the way to their destination.
Like hitting pinball bumpers, except the high score in the end is the bill you pay the token provider.
GuB-42 18 hours ago [-]
To me, that's a totally normal way of using computers. Repeating tasks is what computers do best, AI or not.
Take chess engines of the "deep blue" era for instance. These engines are stupid, trying millions of moves that are obviously terrible, no human chess player would do that. And yet, this is what worked best, because computers are so good at repetition. Recent development using neural networks made chess engines smarter, but using repetition is still how the beat humans.
What you call "guardrail programming" in another context would genetic algorithms, a technique that has recognized applications. And if you look at videos of genetic algorithms learning to play racing games, it literally looks like a bunch of carnival bumper cars onto the highways, but done well, after some generations, it becomes competent driving, sometimes even record breaking.
It can be expensive, it is often the case for LLMs, so you may want to try to be a bit smarter at first to spare some resources, but to me, it is just using computers as intended.
siliconc0w 19 hours ago [-]
There is a tipping point between procedurally generating everything in SVG to maybe giving them tool access to something like 3dsmax (or having them build and then use a tool to do the thing vs doing the thing).
gordonhart 19 hours ago [-]
I’d love to see this benchmark using Blender. Asking a model to animate a scene in Three.js is a square peg/round hole; it doesn’t convey much when the model can’t get it to fit. With Blender the human expert ceiling for this task has been proven to be very high
tayo42 16 hours ago [-]
The 3d model generators that exist aren't that good, and it would need to make a model that can be animated and and then rig it. I don't think there's a point in tying because the small tasks can't be done yet.
cyanregiment 14 hours ago [-]
A quick peruse of the Three.js examples page should prove why it's so easy to get an LLM to produce decent Three.js.
The library is extremely well-documented. When Three.js vibe coded projects started blowing up on Twitter 1-2 years ago, I wasn't that impressed then either because I knew what it was doing.
Anyone who remembers the C compiler built by an LLM! backlash probably feels the same way:
Why would I use an LLM to create a well-known demo rather than fork that demo itself?
What I haven't seen yet from an LLM is it create anything fundamentally new and exciting: A new UIX that is actually good. A new service that is actually good. A game with an art style I haven't seen, music, storytelling - something you'd expect out of a AAA studio.
Given that rant: The coolest part is the multi-modality between text and animation. However, I think the end product would have been a lot better if it was just a video. Having it do it in Three.js didn't add a ton of wow factor for me, and it would have been a lot better looking as video.
> Something like an ephemeral GTA of X on demand.
Here's where you lose me. AAA gaming is very far away from this Three.js demo. But the novel part being the syncing of narrative to the visual scene - a text-to-audio (video) book type technology seems very possible (and useful).
Nice work. We need more big projects like this involving AI (if anything, to get away from the slop argument largely focusing on 1-shot experiments).
kooi 17 hours ago [-]
I think the most important insight is the limitation of LLM perception:Slowly taking screenshots.
That method of perception probably scales N^2... so sure with more compute, LoTR animation will improve. But I think to get a real jump in "experiential feedback", perception needs to scale linear or sublinear. Maybe that's there LeCunn's jepa will come in.
There needs to be the removal of the middle man:
image -> text -> action
To image -> action.
dekhn 21 hours ago [-]
I'd like to see the Silmarillion, specfically both Ainulindalë and the Fall of Numenor. At this point a visual model would probably produce something better than Amazon (but presumably not Jackson).
JBAnderson5 15 hours ago [-]
> Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story.
> it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free".
How is two hours worth of token generation free?
With tech revolutions things get cheaper/faster/better/doable, but there’s still real world limits. The advent of railroads made it feasible for the average person to cross the country, but it still cost a lot of time and resources. People weren’t crossing the country every weekend for fun just because it was now doable.
Why do we treat LLMs as ~free when we are generating things that weren’t doable before but have to invest more money than the Apollo program to build AI data centers let alone account for the operating costs?
janderson215 15 hours ago [-]
He wrote “~free” and also that he set a $10USD limit on it. For the result, I think it is fair to say it is approximately free.
rw2 9 hours ago [-]
I think the argument for the pelican is if it cannot do a simple thing perfectly. A more complex thing is just a stacking a small errors.
mold_aid 20 hours ago [-]
The tilde thing remains uniquely obnoxious in a field that seems want to mangle language for fun, so that's innovative I guess
hooloovoo_zoo 19 hours ago [-]
I suspect LotR is a singularly unrepresentative choice here considering how much info exists about it.
OtherShrezzing 21 hours ago [-]
> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom
This is an odd take, given that Karpathy is certainly aware that the LotR films absolutely did create Bag End in digital format; that their creation was outstandingly high quality; and that Claude’s output here very obviously “leans heavily” on their prior art.
justinnk 8 hours ago [-]
Not sure whether Karpathy meant it like that or thought about this statement being interpreted this way, but I totally agree with you. Creating animated movies and special effects is placing a lot of 3d polygons by hand. Maybe not with JavaScript though. However, people also made things like
telnet towel.blinkenlights.nl
Now, were they „in their right mind“? I don‘t know. The more likely explanation is they found joy in it. Reminds me of what the Suno CEO Mikey Schulman said about making music: „[…] I think the majority of people don’t enjoy the majority of time they spend making music.“ (https://news.ycombinator.com/item?id=42688538)
Arshad-Talpur 17 hours ago [-]
How long we will be testing and benchmarking generative capabilities of LLMs? in creativity, in code generation less or more result is expected and approved, but in execution $19.8 can not be $19.9 or $19.7.
DoDecaHeJon 8 hours ago [-]
Fun ways to test the models. What other unique tests are people using?
skybrian 22 hours ago [-]
Still images seem like a better quick test because we can see them at a glance. Maybe ask it to make a comic?
quantumleaper 22 hours ago [-]
I'm sad that Andrej Karpathy went from being one of the most reasonable, trusted, and credible voices in AI to peddling marketing slop for Anthropic.
8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.
novia 21 hours ago [-]
He makes a very good point here about llms or lrms not being good at taking video as inputs, and i think he's signaling his intent to help change that. He's pointing at an open problem. This isn't a carefully structured blog post, it's just a tweet.
azan_ 22 hours ago [-]
> 8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.
Or maybe over these 8 months agents improved a lot? You know, few years ago many AI experts predicted that things we are routinely doing now with AI are decades away. I mean how can you look at this post and not be impressed? It's insane what AI is currently capable of.
Rexxar 21 hours ago [-]
What narrative has changed ? I just see an experiment with a very perfectible result and some reflections about what capacity are currently missing to have better results.
budsniffer952 22 hours ago [-]
[flagged]
azan_ 22 hours ago [-]
Yeah, it appears that if your prior is that AI can't be good, no evidence can convince you that actually current models are extremely capable.
micromacrofoot 22 hours ago [-]
you're doing the exact same thing by making the inverse claim without any information, at least the original comment claims some change for someone to investigate... all you're doing is disagreeing as an opportunity to be snarky
dofm 21 hours ago [-]
Anthropic spokesman [0] Andrej Karpathy is here to tell you about token-wasting loops, and insists on the weird idea that they are "~free", when in fact, they are fuelled by expensively burning investor money.
[0] Seriously. Get used to mentally prefixing his and Boris Cherny's name like this, every time you see them quoted. These people are speaking while employed; there is no chance they are not aligned with the employers who will make them wealthy. The tech industry does like to pretend that for some reason AI people, uniquely, speak thoughts unbiased and for themselves or even for science or humanity.
rvz 16 hours ago [-]
Exactly. A smart HN user finally sees through the bullshit all because an IPO is around the corner.
swe_dima 19 hours ago [-]
In my experience SVGs are still too hard for LLMs.
I gave Fable a jpeg and asked to draw an SVG, using a loop that renders the SVG into an image so Fable can inspect it.
Results looked like drawing of a 5 year old.
eric_khun 13 hours ago [-]
Everyone's about to make games sooner than expected. games are slow, might look buggy and unfinished. But I think that's just about time before we see fun games coming up. i'm betting on it with https://antics.gg . if you want to host your multiplayer vibecoded games, feel free to use us!
hkalbasi 21 hours ago [-]
This makes me think about using a game engine and a coding agent instead of current video generation AIs. It will probably cost much more, but it will have almost zero consistency problems. Is this line explored?
wiradikusuma 22 hours ago [-]
Do you guys notice that LLM can create fancy viz/animations by coding them instead of leveraging what we humans usually use (e.g Lottie, After Effects)?
I wonder if Flash is still popular... LLM can use that instead...?
croes 21 hours ago [-]
> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom
There are people in their right mind who would do that and their are already examples of people who did similar things.
But maybe not in the future if people would confuse all the effort with AI
sinaatalay 19 hours ago [-]
On consumer devices, AI communicates with us through speakers and screens. Screens are the richer medium, so most consumer AI innovation will happen there.
Computer graphics will have enormous applications because they are directly controllable by LLM-generated code. Video models are probabilistic and less suitable when precision matters. In education, for example, we need exact visuals. If an AI wants to plot y = sin(x), it should generate the precise graph through computer graphics rather than approximate it with a video model.
duxup 19 hours ago [-]
Coding, graphics, all seem to have very defined background data for an LLM to base their decisions on, test, and even when I correct it we're all on the same page. Seems like there's lots of room for AI productivity there.
What I find funny is that approximate to computer graphics are video games. When I ask AI about a decision available to me in a video game AI completely fails, OFTEN. I assume all the forums and changes made to a game over time might be quite confusing for AI. But I've also seen it completely make up characters and decisions and weapons and so on about some very clearly defined games and paths. It's an interesting dynamic.
serf 22 hours ago [-]
you don't really need screenshots if you have an engine expressive enough for the scene generation while ensuring the visual appearance of the engine output itself is feasible.
that's why these things are actually pretty good at openscad/freecad/F360 mcps , the visual reality is enforced and guaranteed by rigor in the interpretation engine that is anchored to human physical reality.
informal007 18 hours ago [-]
One difference for human to understand the video is that we only care the changes on a picture compare to LLM
xg15 21 hours ago [-]
> I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it.
I think it's interesting that the "Bag's End" interpretation in the video clearly looks like the one from the movies, but generated here as a three.js 3D asset.
It makes sense that the movies (or shots/frames from them) were in the training data, and I can also easily imagine an association in concept space between the textual description of Bag's End and the frames from the movie.
But how on earth does the model then go on and convert the latent representation of those images into coordinates for a 3D mesh, without ever even restoring the image? In what kind of representation are the images from the movies stored that it can do that?
Gooblebrai 21 hours ago [-]
I can't believe the video demo is $10
mikojan 19 hours ago [-]
After watching this video I am absolutely positive that the issue is not a lack of stamina in humans. It is that humans have the capacity to realize that this is a bad idea long before they complete it.
Lerc 18 hours ago [-]
I think I could tolerate 50 Shades of grey rendered in this style.
killingtime74 17 hours ago [-]
I love how nobody cares about copyright anymore. Not even an after thought. Might be okay if you're an employee of anthropic/openai but for us mere mortals I'm not sure I would share something this blatant. In US it's $150k+ per violation and you've given them all the proof (even a confession)
throwaway89864 21 hours ago [-]
It may make sense to switch this to USD/Omniverse.
blitzar 22 hours ago [-]
I think the pelican test is better.
criddell 22 hours ago [-]
I'd like to see tests of things the current AIs are bad at, like drive a car. Or maybe take instruction to complete some novel activity to test how well they can learn.
wrxd 17 hours ago [-]
Would you say that the generated video is good?
Better than I would have expected? Sure. Impressive that it got that far? Definitely. But is the end result good?
2 hours ago [-]
informal007 18 hours ago [-]
it shows the possibility that SVG replace PNG/JPG even video.
I know this is a plug but I thought of it, too! I hope this is the next benchmark they saturate.
I like where Karpathy is going; I had the same thoughts about LLM generated slop scenery. I just want some variety of scenery for the goblins to get massacred in in whatever fantasy slop game I play.
I find it interesting that people see Money For Nothing as evidence of technical limitations of the era.
It’s not. That blocky appearance and limited lighting was deliberate comic aesthetic choice as a response to budget, not a result of technical limitations. CGI in 1985 was enormously more capable than that — take a look at the stained glass knight in Young Sherlock Holmes, or The Last Starfighter from 1984. Things certainly could have been rounded, sculpted etc.
teiferer 17 hours ago [-]
> sure, why not, it's ~free
Yeah, please check in with the folks protesting data center builds in their town causing their electricity prices to skyrocket and tap water to turn into a scarce resource.
And now, instead of actually doing the above, please go ahead and downvote me, because how dare he question those LLM games.
skakcnejwisnsid 17 hours ago [-]
> because how dare he question those LLM games.
I thought you are the one who's on Greta's side, you ar e the one who should be saying "how dare you?"
andy99 21 hours ago [-]
Benchmarks like the pelican thing are about correlation with “how good the model is”. Better models produce better pelicans.
It’s a useful benchmark (aside from being “cute”) because of its simplicity, both in how many output tokens it takes (though I understand some models think a lot now to do it) and how easily one can subjectively judge. It’s this efficient as a benchmark of performance.
Making a long video takes way more tokens, and presumably is a lot tougher to easily compare. swillison has a presentation that’s pelicans from 2023-present (roughly) showing the progression. Imagine “lord of the rings videos from 2026-2029” or whatever, it would take a long time to watch and be harder to judge, and probably just end up being a comparison of screenshots anyway.
TLDR I feel like the post misunderstands the role of the pelican thing though if find it very hard to believe he really doesn’t understand, so maybe I’m missing something.
c0rruptbytes 22 hours ago [-]
perfect benchmark to burn more tokens - convenient
mvdtnz 18 hours ago [-]
> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free".
Except it's not ~free, it cost ~$10. And no one in their right mind would ever exchange $10 for that crap output except in this brief moment that we're in because it's fun and surprising to see what will happen. The actual result is as close to useless as it's possible to be - it's not interesting in its own right, it's not aesthetically pleasing, nor funny, nor informative. It's just slop.
sampton 17 hours ago [-]
Pelican maxxing.
bbstats 21 hours ago [-]
This is awful
forrestthewoods 22 hours ago [-]
As a former gamedev watching non-gamedev AI talk about games is so amusing. They really truly do not understand anything about games or consumer entertainment.
There’s a reason AI slop games have literally zero engagement. Last summer that stupid flying game blew up. Maybe a million people “played” the game. Where play means they clicked a link and checked it out not because of what the game was but solely because of how it was made.
In terms of concurrent players that game wouldn’t have cracked the Top 5,000 on Steam.
My metric for AI games is “number of players who spent more than 15 minutes playing”. I’m not aware of any vibeslop that has achieved 1 such player.
Now obviously LLMs are transformative for game dev. But “hyper custom worlds you can drop into” shows an extreme ignorance of what players want imho.
stackedinserter 21 hours ago [-]
It would be better to ask model to render segmented 3d, with placeholders, like magenta is water, blue is sky, green is grass, purple is Frodo's face, etc, then pass the result through img2img model to properly "render" it.
epolanski 22 hours ago [-]
I wish there was a timeline where I never ever had to see the pelican SVG test ever again.
22 hours ago [-]
overgard 15 hours ago [-]
Oh god, please leave video games alone, we have enough slop.
0x1ceb00da 11 hours ago [-]
We already have tons of videogame slop. 100s of videogames are released on steam every day.
gaigalas 9 hours ago [-]
I think his explanation is bullshit.
From my experience (PixiJS game), Opus can look at things and screenshot. It's slow but it works. The issue is that it doesn't try to make them look good.
Instead, it tries to find one bug and then fix that one bug and close the session. That is amazing discipline for webapps or whatever, but for game design is just the worst if you need attention to detail and tuning of multiple holistic systems that produce an effect together. That really limits the kinds of approaches you can can use to achieve things visually.
Sure, the testing harness could be better, but that's not what makes it a poor fit. The base model seems also capable (it can identify the issues, just isn't willing to disangage from this "I found X and fixed it" single-thing work posture).
The kind of work is just different, and it was probably never trained on it, so it feels off and an uphill battle to use it.
miltonlost 21 hours ago [-]
[flagged]
mdp2021 21 hours ago [-]
> wasting money
Testing, assessing, tasting...
We also do it when we build other things - this is just a different scale.
azan_ 21 hours ago [-]
Do you think this was supposed to be art? Wow.
angoragoats 18 hours ago [-]
[flagged]
angoragoats 4 hours ago [-]
I guess I’ll just keep flagging every single Twitter link on this site. I hope that some day the community here will grow a spine.
firatsarlar 11 hours ago [-]
[flagged]
uuuynnnuuuyyyn 13 hours ago [-]
[dead]
larpathyparpav 13 hours ago [-]
[dead]
18 hours ago [-]
hansmayer 19 hours ago [-]
[dead]
shapefrog 22 hours ago [-]
[flagged]
nozzlegear 21 hours ago [-]
[flagged]
mdp2021 21 hours ago [-]
Can someone please translate that expression? What would that mean?
rzzzt 21 hours ago [-]
Affirmation?
mdp2021 21 hours ago [-]
Well, in that case - if it is just a "hear, hear" - I do not see why the utterance from that actor would be of note.
Maybe nozzlegear wanted to suggest some importance on Musk remaining a bet-ter on the general tech, regardless of the competition?
rzzzt 21 hours ago [-]
This gets me thinking (although there is not much to divine from three letters). Maybe it's a Jennifer Lawrence GIF-inspired "Yeah, right" then? Downplaying the significance?
mdp2021 20 hours ago [-]
To some online sources, it is used for both.
> "Yah": slang spelling of the word "yeah" (which of course can also be used ironically)
> Merriam-Webster: "Yah": used to express disgust, contempt, defiance, or derision; probably imitative of the sound of retching
rvz 16 hours ago [-]
Who cares?
nozzlegear 14 hours ago [-]
Uhh me and 1400 other people, obviously
rvz 11 hours ago [-]
You mean auto-liking “bots”.
240M followers vs 1400 of which 99.9999% are bots with the rest being Tesla fanatics.
That is close to no-one.
blitzar 21 hours ago [-]
Concerning
trlhaq 22 hours ago [-]
[flagged]
mdp2021 22 hours ago [-]
Theft of what? Why - aside from the fact that your accusation are irrelevant to the submission and arguments absent -, you do not know verbatim paragraphs (we do)?
You cannot come and place your personal positions as assumptions. To me, there is absolutely no theft. And we cannot play a game of "Yes!"//"No!" here.
By the way: are we having a surge of this?
ssdg16 21 hours ago [-]
> Theft of what?
> By the way: are we having a surge of this?
We do have a surge of pro-AI sealions, yes. Any objection is countered with one or more three word questions.
mdp2021 21 hours ago [-]
> pro-AI sealions, yes
Very devoid of intelligence note.
> Any objection
Objections are arguments. That post did not start an argument - it was as ideological as the Brigades. Devoid of what we want to have here (I believe).
I have stated and do state: what is published is assumed as read (only, probably not read for lack of resources). If it is in the libraries, it is there to be read.
david-gpu 22 hours ago [-]
Try reading what Karpathy said again before reinforcing your biases.
trlhaq 22 hours ago [-]
Yeah, pretend that you never used 2023 models that quoted everything verbatim. To put it in a way that you'll understand:
No bias. No speculation. Just facts!
adamtaylor_13 22 hours ago [-]
You literally did not read the article.
redsocksfan45 21 hours ago [-]
[dead]
theproblemisyou 21 hours ago [-]
[flagged]
yourewrongsorry 22 hours ago [-]
[flagged]
andrewstuart 20 hours ago [-]
This is equally bad as a pelican test.
LLMs should be tested in the same way people should be tested for a job interview (but often aren’t) - with tasks RELEVANT to usage.
So you don’t just randomly pick some random thing to make the LLM randomly do (like many job interviewers do).
You start with clear statements about real world usage scenarios. THEN you come up with tests that give insight to how well the LLM/hob seeker gets the job done.
Please, stop coming up with random tests like it’s Microsoft in 1990 and you’re asking job seekers how the would move Mount Fuji, as a way of assessing their programming skills.
No stupid irrelevant pelicans on bicycles and no stupid renderings of Lord Of The Rings. Unless those are relevant use cases.
Any test that anyone comes up with must clearly state the context and how the outcome is measured.
weird-eye-issue 9 hours ago [-]
You might be onto something there, except that people do use LLMs for these exact use cases.
hansmayer 19 hours ago [-]
[dead]
hn22fazjsv 21 hours ago [-]
Screenshotting for later
matchagaucho 22 hours ago [-]
It's difficult to think in exponentials.
But this demonstrates we're a couple orders of magnitude away from generating 1:1 hyper-personalized entertainment and media for individuals, rather than the masses.
hgoel 21 hours ago [-]
I don't really get the desire for hyper-personalized entertainment. People are very good at pointing out things they dislike, but not very good at coming up with how to fix them (common wisdom in game design).
On top of that, a decent chunk of the joy of entertainment is the social aspect.
xpct 15 hours ago [-]
The only charitable interpretation of the hyper-personalization take I can give, is that you would want to see content that elicited a certain type of emotion inside you (so, a very high-level description).
Now, whether this would be at all possible, or even healthy for you, I don't know.
skippyfish 22 hours ago [-]
I very much doubt that. We now have nearly-perfect AI image generation and it hasn't really changed the nature of human expression. I don't see my friends getting wildly creative. I mostly see it used for spammy blogs, spammy books, and cringeworthy corporate marketing - basically, a negative signal, rather than "ooh, AI image, I'm in for a treat". Is your experience different? If not, what changes with moving images?
Ultimately, most people don't have ideas for the kinds of personalized entertainment they want, and they don't want to be in charge of content production (even if you have an LLM do most of the work). I don't doubt that there are niches for it, especially stuff like porn, and I'm sure that pros (game studios, film studios) will leverage AI more and more, but I suspect that most of us will just want to sit on the couch, watch Spiderman XVIII, and then be able to talk about that shared Spiderman XVIII experience with all our friends.
timacles 19 hours ago [-]
All small business marketing material now looks exactly the same.
LLMs produce an extremely small range of artistic expression, and its kind of ironic, that while it certainly looks functionally good. It immediately becomes noise because everything looks identical.
So companies are going to have to hire creatives again, in order to stand out, and be "creative"
darkwater 22 hours ago [-]
I'm mostly seeing the impact of GenAI images in everydays life for example in the small posters small associations or individuals usually attach on streets to promote some small local event. We went from just text created with PowerPoint and maybe some stock image just a Google search away to now images depicting the topic closely. But yeah, it's not like a revolution.
And this personal media thing, yeah maybe for terminally online persons that are REALLY into a sub-genre but otherwise, it's too much effort, I agree. Until we get machines that can read our (subconscious) mind, that will not exist.
morbicer 21 hours ago [-]
I yearn for the old days of bad photoshop/wordart/powerpoint flyers.
They were visually bad but honest and sometimes soulful or playful.
The AI slop that's everywhere now looks superficially more professional but it's very busy, samey, unnatural and it really turns me off.
artisin 17 hours ago [-]
I don't want to say this is indicative, but my local library used AI-generated event poster/images for the past year. But recently they switched back to word/clipart event posters. When I was checking out, I query the librarian about it, and apparently they saw a 40-60% drop in attendance. On one hand, I found that hard to believe, but on the other hand, simple word/clipart event posters do have a certain inviting charm about them.
darkwater 7 hours ago [-]
I think that word/clipart posters make it clearer the date and time, with GenAI material you end upo adding a big colorful image that takes most of the space and relegate the important bit to a second place. At least that's what I see around me.
lukeschlather 22 hours ago [-]
> We now have nearly-perfect AI image generation
Not even close. Current AIs have very poor spatial awareness, they can generate some kind of scene but they can't tell you where objects are in the scene, nor can they move objects into different places. They're very useful but it's difficult to be very creative with them because they can't update an image to make it more aligned with your vision for what should be in the image.
matchagaucho 21 hours ago [-]
Ultimately, most people don't have ideas for the kinds of personalized entertainment they want
1:1 AI entertainment probably won't be a cold start experience.
Much like the "Choose Your Own Adventure" books of the 80's, consumers choose a baseline template, customize the characters, and interact at various points within the plot.
potsandpans 21 hours ago [-]
> We now have nearly-perfect AI image generation
No, we don't. It's quite good, but it's nowhere near perfect. For example, you can see: https://genai-showdown.specr.net/ or others that highlight how far from perfection we still are.
> and it hasn't really changed the nature of human expression.
It sure has changed our discourse. Look at hackernews, where people just can't help themselves to engage in ragebait and flamewar comment threads about llm generated accusations. We've got half the people convinced that looking at ai content is like consuming food.
I'd venture a guess that it's too early to determine how this technology will change expression. A good analog would be photography, which caused a similar meltdown in the arts at the time.
> I don't see my friends getting wildly creative.
Maybe you don't have creative (enough) friends.
dgellow 21 hours ago [-]
If that becomes a thing, that will be fun 2 weeks, then become a gimmick only a small niche will be using. What’s the point in having a hyper personalized entertainment? People want to experience the games, movies, tv shows, books others created and are also experiencing
futureshock 22 hours ago [-]
Judging by the Seedance 2.5 demos today, I’d say it’s not that many orders of magnitude away now.
dude250711 22 hours ago [-]
I kind of like a sense of community even if it means losing out on personalisation.
Teever 22 hours ago [-]
Amazon recently kiboshed a new Stargate television series. The fan community was quite heartbroken because the Amazon had involved some of the writers from the original television series and apparently one of the reasons that they cancelled the show before it made it to production was that the felt that the proposed story would appeal too much to the fans and not new people.
The fans were real heart broken about this but I think you're right on where this is going. We're not going to be seeing the dominance of centrally produced content like this for much longer, like sure, I think there will be big blockbusters will stick around, but I think the day is coming where media becomes a choose your own adventure sort of scenario.
It'll be interesting to see where this scales to. there will definitely be some amazing solo projects but we'll also see the like 4 player co-op version of productions and then the larger mine-craft server 'Minas Tirith' scale ambitious projects that involve a few dozen people. And of course passive consumers will remain a thing, or people who just provide some suggestions or nudges for what they'd like to see others make.
But I don't think it'll be dominated by big companies like Disney, Netflix, Amazon or Paramount.
I'm sure a lot of them will suck but it'll be neat to see the inevitable Seinfield - Star Trek Voyager cross over episodes.
Elaine and B'Elanna Torres feud after a transporter accident leaves the crew stranded the delta quadrant. Jerry attempts to date 7of9 but is rebuffed as she finds Kramer's quirky bluntness more relatable. George panics after someone compares him to Neelix.
ianberdin 17 hours ago [-]
I still believe my bench is superior to Pelican or Karpathy’s.
Everyone knows how a MacBook looks. Any missing or incorrect detail will be obvious. However, Pelican or this world can be anything.
Anyway, only Opus 5 Max and Fable 5 XHigh + Max make a good-looking Mac.
toplinesoftsys 18 hours ago [-]
This is just a very rough basic game skeleton. Companies believing AI is smart enough to spit out almost ready product will learn a hard way that it will will give them just a starting point that still needs a lot of work. This is a nature of all modern LLMs - they easily lose context: the bigger the context the more losses and distortions are, especially in the middle. So, giving AI the complete book does not mean it will follow everything in it - quite the contrary.
jonas21 18 hours ago [-]
No, it cost 1M tokens, or about $10
EDIT: The parent comment originally claimed they spent $1M on the demo -- they seem to have edited it after I replied.
wrxd 17 hours ago [-]
How many more millions of tokens do you need to spend to get something of acceptable quality?
Spaghetti 2026:
https://x.com/dreamingtulpa/status/2083304533829066873
https://xcancel.com/dreamingtulpa/status/2083304533829066873
A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.
it made a huge monstrosity first, 1500+ lines of shader code. Once it was happy with the result (it looked pretty good, almost blenders gamerenderer) it cleaned it up and a lot of debug code was removed. shrank it down to about 350 lines and made it much more readable.
this was something i didnt expect it to be able to do. create shaders, look at screenshots, fix em iteratively like that. multi modal debugging.
Imho we don’t need to make benchmarks that draw the whole 3D world. Pelican’s drawing is really nice in its simplicity and complexity at the same time.
It seems to be an obscene waste of compute time to generate useless 3D worlds that are just a bragging - 3D is really heavy discipline to make it right, see Mark Zuckerberg’s ceased attempt with 3D VR…
Multiply it by thousands times as a lot of people have found out threejs lib and prompt “generate 3D world and make no mistake” are new orange/black.
As for the often quoted issues with the bike's frame or problem with the steering column, I can't really tell, I am no bike expert.
I can instead judge how poor of a job it is doing with a LOTR rendition in Three.js, so that seems like a better benchmark.
No, this demo is the useless 3D world, and you're bragging.
A real game would have a lot more immersive of a world, and you wouldn't need to.
I have always preferred the result of getting an LLM to draw an svg or make a procedural animation like this to the uncanny hyper-realistic result of diffusion image/video generation.
Given that limitation, it is incredible what it can still accomplish. And when it falls short, it is often in a charming way, if you are open to seeing it that way. It reflects something like a child's understanding of the world, not entirely wrong, just incomplete.
The fluttering cubes that might have been bees or butterflies were my favorite part.
Anyway, just a small point which probably still wouldn't change the nice metaphor your made.
But I do think this was a poor demonstration of the idea for another reason: Tolkien works have a HUGE corpus of training data. It's great that random users can come in and immediately recognize what the footage is, but it fails at the very first thing the pelican was meant to do:
- Render this thing you have only tangential training data of, that also happens to be an asymmetrical object so we can see how much you fuck up the details if you somehow flip the orientation half the time.
They should have used an obscure story, not "Most Studied Piece of Literally Work of The Past Century trademarksymbol"
https://www.booooooom.com/2016/05/09/bicycles-built-based-on...
https://www.catrike.com/expedition
Elon Musk on Grok: “better than PhD level in everything.”
Totally agree though, anyone with a vague understanding of how bikes works ignores the pelican because they know the bike is unrideable in the first place.
They are listening.
I am not a mechanical engineer, so even prompting well with ME lingo probably will take some effort.
That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right.
But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available but also takes custom stills if you provide them.
My test scene was the Gauntlet scene from Apocalypto. It is low fidelity but does a pretty amazing sequence with somewhat believable physics of the javelins etc.
Here is the docs page with the vertical takeoff / 88 miles an hour time travel: https://contextify.sh/docs
I can share some of the Apocalypto bit if anyone is interested.
If you have time to try the product as it stands, I would definitely appreciate your feedback in particular, and would be interested in checking out what you've built so far for yourself.
I will publish the API to the database as well.
That page with the time machine may not make it obvious, but the product does have a linux client!
It doesn't have the same app window and summarization of the macos version, but the transcript ingestion engine is efficient and the real value is in leveraging the database it builds using the packaged `/total-recall` or your own use of the api. (which is not yet documented but discoverable)
I am building the windows version now. I have been for the past four days. It uses a shared swift-core with the macOS and Linux versions which has been part of the reason it has taken "so long."
If anyone is on windows (or linux!) and would be willing to try it either of the clients my email is in my profile, I would be grateful.
Also I'm working up a short post with that apocolypto anim now.
I can’t find a single free app even for a tiny utility without being forced into a yearly subscription with 7 days free trial. One of the many things I regret switch from android for.
Since the democratic of iphone users are mostly tech averse people I can assure you most of that are forgotten subscriptions.
My point being that this is a bit niche, and overall trends might not apply.
Are the total number of Mac users who pay higher than the total number of Windows users who pay? I kind of doubt it. Windows market share is still much higher than MacOS.
https://banagale.com/cinematic-canvas-ai-film-animation.htm
I provide a "how I got to this" up front, but if you want to jump right to the Apocalypto stuff, use this:
https://banagale.com/cinematic-canvas-ai-film-animation.htm#...
And if you want to play with interactive demos of the two animation sequences (the tuning tooling I described in my OP) you can go directly there:
https://banagale.com/cinematic-canvas-workbench-demos
If anyone wants to collaborate on building out the cinematic-canvas-workbench project please email me.
When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given early access, and/or Anthropic employees) were pretty much 100% three.js stuff. Yes, it looks nice, but it doesn't tell me any better than an Erdos proof whether the LLM will be able to run my vending machine.
We should feed a snippet of an unreleased book in a novel universe.
the llm must be able to output the correct code when the input says something like 'place object x to the left of object y' vs when it says 'to the right'. there is an absurd combinatorial space of the possible inputs vs outputs it must generate -> the whole set of these, interpreted from a human point of view, you could call knowledge. the [x] in input x output you could call the relationships. And the fact that the LLM doesn't have to brute force represent all of them (impossible in the limited embedding vector x internal representation state) you could interpret as 'understanding'. But again, these are loaded words that are somewhat meaningless when looking at what an LLM does in a literal way.
They have plenty graphics code to train from, so this structure will be directly or indirectly available in the training data in a very plentiful way.
In literal llm transformer terms, the embeddings must have some of their components statistically represent spatial structure in some way that later in the internal layers of the LLM give some statistics of how likely it is to occur for certain code tokens to appear relative to the input of spatial wording in the token stream.
it is likely that their internal representations encode something more general than specific input x output stream combinations (because we already know this is the case for regular words and concepts). if not in the embeddings, then a couple layers into the network for sure.
And if you use image generation capabilities, you clearly see that LLM's suck as understanding "next" to a thing.
You can't just defer this decision to the three.js program because you need to understand this sort of relationship yourself if you're going to suitably place objects (and your own view) in a three.js scene in the first place. Building a complex scene involves making hundreds, or more, of decisions about where to place coordinates in 3d space.
This example is underspecified (e.g. what are the shapes and sizes of the objects and what is the angular width of your view) but it illustrates the problem. Even with a good intuitive understanding of 3d space, we would struggle with this. LLMs are not calculators, so it's surprising if they manage it.
A big reason people separate this out is because it was only a short time ago that AI models were noticably and uniquely bad at this. I don't say this to be like "oh so imagine where theyll be in x amount of time", rather that this thing that was once a serious limitation of the technology is slowly being compensated for by larger and better trained models
Are you sure about that? Presumably the spatial concept of "inside" and the computer memory concept of "inside" have different contextual embeddings in an LLM, not so different from how our brains have different neural activation patterns when we use different concepts. Unless you think that "the same kind of inferential activity" also applies to human neurons firing.
This weekend I've been converting a game from three.js to ogl.js in order to see if I can optimise the time-to-interactive loading time. I took the three.js driven page weight from about 600KB (500KB being three.js) to about 50KB, and reduced the loading time from multiple seconds on a 4G mobile connection to around 0.5s.
This has mostly been a combination of Opus 5 and Sonnet 5 in Claude Code. It very clearly has a good grasp of WebGl, and of what impacts page loading times and rendering speed. It was able to drive Claude Code's integrated browser to measure the impact of changes, and as I spiked out a test of ogl.js it could test the differences changes made.
It's not the best game (https://tinyslots.ooer.com) but that's on me. As an exercise in building 3D in a browser, and in page speed optimization, with Claude models I am really impressed.
Anthropic models are excellent at working with Blender APIs, other industry standard 3D modelling programs and tasks, and game libraries that have nothing to do with Three.js.
Three.js is a quite popular library; and browser-based apps are more easily sharable and more portable. So models having a preference for using it when unprompted doesn't suggest anything, just like how models using React or Tailwind by default doesn't suggest anything.
It seems they trained it to be good at it, then requested everyone to demo it.
Please don't fulminate. Please don't sneer, including at the rest of the community. https://news.ycombinator.com/newsguidelines.html
Many people from all sections of HN, the tech industry and broader society have been surprised and wrong in all directions about how the emergence of AI is playing out. I certainly have. I can't think of a single person whose predictions have been precisely accurate. So, please don't use terms like “wall of shame” and “cope”. It doesn't help anyone make better predictions and only makes you, and HN generally, seem mean.
This is what you get. It's like the more general concept of starting with the written word (or 1000 words) and then replacing it with a picture. You've done something strikingly different, but is it serving the same function?
It's fascinating to see this stuff combine such disparate sources in unexpected ways. But it is parrot, just not in the way you're expecting.
Opus 5 is the first I’ve seen to “one shot” it (in a harness, so it was more than one LLM call).
"Getting Started with Google Wave": https://www.youtube.com/watch?v=eKUAqNGVwX0
Oh never mind, that was Google Glass, another dead google product. That he personally killed with that photo. So confusing to keep track of them all.
Years later Apache moved it to read only because of low community activity.
The archived git repo on GitHub remains available to clone and revive as a fork.
https://github.com/apache/incubator-retired-wave
https://en.wikipedia.org/wiki/Inbox_by_Gmail
wave failed for weird google organizational reasons far more than anything inherent to the product or tech
https://www.youtube.com/watch?v=4Z4RKRLaSug
Their "World Quester 2" tutorial shines as the holy grail of consistent and ergonomic user interface and game design. The menuing system is so magnificently structured and well organized, it bring tears to my eyes. Google Wave pales in comparison.
https://www.youtube.com/watch?v=0Gy9hJauXns
It refused to use the text verbatim because of copyright (ironic), but the output was interesting nonetheless.
https://claude.ai/public/artifacts/275dc3c2-7bd3-432b-94ff-d...
Also, large models refusing to work with copyright material is really hypocritical, copyright enforcement for thee but not for me
Last I checked, I did not see the prompt for this really cool thing, so it is not reproducible.
Did I miss the prompt somewhere?
this guy seems to have taken that idea and got something similar/better, so likely the prompt isn't too special
https://x.com/Izkimar/status/2083819741643178208?s=20
In either case, I still don't see how I could reproduce this to test against various models, which is the entire point of Simon's pelican.
This is a bigger difference to the pelican than simple reproducibility steps. The pelican is intentionally esoteric, and thus open ended. The LotR is mundane and has a "correct" answer, aka, copy the movie.
It makes it a really awful test of capabilities. The pelican isn't a slop test. This crap is.
I'm actually beginning to wonder if some people who ignore these things have a different, somewhat lesser ability to percieve image details than I do.
I mean I guess its fine to go on to another test despite never actually passing the pelican bike test, but there's a sense that we have to use another test because AI is now good at pelicans on bikes, which is just not true.
Please remember, we've started from there :
https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/
When it started, it was clear what LLM would stand out, its style, etc. Nowadays, the pelicans look similar, the difference is in details and sometimes hard to catch. Sure, the task is not completed perfectly, but that's not the point. It was supposed to be a benchmark to quickly benchmark a LLM against others.
Isn't it?
If the computer can't do it better than a human being, then what's the point?
Being wrong at scale is not better than being right.
But no ones hire random humans for things like this. You go and hire a vector artist and they will get your a very good pelican on a a bike. That's how you get things done when you can't do it.
Yeah, but then, you recruit the artist for $XXX - whereas you "recruit" your LLM for $0.XXX for the same task.
Of course the quality difference is huge. But sometimes you don't need that level of quality.
Also, finding a vector artist takes days of communication, payment settlement, revisions, etc.
Not always the most practical solution.
Because the benchmark wasn't testing "can an LLM draw a pelican like a human". The original article was testing the relative capabilities between LLMs. Now that LLMs can all draw pelicans all similarly, the test is less interesting as a comparative benchmark.
This is what the tech industry has become?
Less of a failure is still failure.
LLM has progressed a lot in the last two year, judging from the pelican drawings. I personally couldn't care less about it though. I do know that I've gone from using no AI at all for coding to probably 95%. I hardly code by hand anymore. That's much more impressive and significant. Failure you said?
Pick up a newspaper. Start with the Wall Street Journal. These are public companies. It's not a secret.
It can certainly do it better than I can. Sometimes you don't have a human handy with the required skills to do something.
It's physically impossible.
The problem is to draw it in the least disturbing way possible.
https://ridesabike.com/donald-duck-daisy-duck-huey-dewey-and...
When you aren't sure if an LLM can write an svg well, or that it will be able to form a pelican shape, or animate a bicycle, it's a good test. After that, it's all judgement: how detailed should the pelican be? pelicans are the wrong shape for a bicycle by default, so how much can I change its physiology to match using a bicycle before it isn't a pelican? Do I care about how well the client is able to render complex geometry?
It's not that there isn't room to do better, or that it doesn't tell you anything at all, but rather we've reached a point where what it tells us isn't very clear anymore.
Oh come now. I am extremely confident that if I hired a professional artist to draw a picture of a pelican riding a bicycle, I would get something inarguably much better than what today's best coding LLMs can produce.
I think AI folks have done a terrible job of communicating this, but replacing a professional simply isn't the point. The point is to serve all the situations where people would've never considered hiring a professional, and where perfection or artistic merit isn't the point (say a personal throwaway recreation of an LOTR world).
And I think in that regard the benchmarks are pretty good.
I'm not saying it is, just that there's obviously still room for the models to improve on this task.
If the expectation is that AI is going to replace "knowledge workers" then the limit would be a darn perfect drawing. We are nowhere close to that.
And Elon is already propagating the age of abundance where money won't exist anymore, right before calling the interviewing journalist dishonest and deservedly losing public trust. Smh my head.
What knowledge workers do you know that have excellent drawing skills? I worked in a design agency and for a couple of years, each week me and a few other people would attempt to sketch a member of our group: one person would be the model and sit still, and everyone else would draw her/him.
Let me tell you, if producing a convincing portrait was a prerequisite for being a knowledge worker, there would be 99% fewer knowledge workers.
It surprises me how many people in this community don't get this. Obviously, most prompts thrown into an AI chatbot/interface are about something no knowledge worker would ever have to deal with. That doesn't disqualify them as a tool for measuring progress of the models.
* some omitted the bottom of the diamond which connects from the pedals to the rear wheel
* some added an extra connection from the pedals to the front wheel, making it impossible to steer
* none could align the head tube with the fork
* none added a correct offset to the fork
* none could generate the chain properly in a way that attaches to the two sprockets correctly
I mean just look at these:
* Grok 4.5: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
* GPT 5.6 Terra: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
* Sonnet 5: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
from https://dylancastillo.co/posts/pelicanmaxxing.html
The bike is generally okay, apart from medium which derped hard. Max has correct diamond, correct head tube, and so on. Only nitpicks are that the front fork offset isn't there and the chain doesn't touch the rear sprocket correctly.
It used to be a very difficult task for models, see [2,3,4]
it cuts across several tasks that AI used to be very bad at, but now has improved quite a bit. Namely, spatial reasoning (because it has to manually place the points of the svg such that they make sense and form what it says it forms. This used to not work very well, with random shapes floating around that it would mark things like "eyebrows" but were nowhere near the "eyes", etc.
It also tests the model's world knowledge (what do pelicans look like? sure they have wings, feet, beaks etc, but what shape are they? how to get proportions roughly right? this isn't a given from text data about the bird. This goes doubly for a bike, which is a quite complex shape that most humans fail to draw correctly[1] (many draw the frame or chain connecting in impossible ways that would not ever function mechanically)
Before it was pelican on a bicycle there were people having it do horses/unicorns making the rounds - gpt4.0 or whatever would often make hideous abominations of legs and mouths
[1] https://www.gianlucagimini.it/portfolio-item/velocipedia/
[2] https://static.simonwillison.net/static/2026/mistral-small-4...
[3] https://static.simonwillison.net/static/2025/codex-hacking-m...
[4] https://static.simonwillison.net/static/2025/gemini-2.5-flas...
So it's a pretty much pointless test now.
It's quite surprising actually.
Not only that this so-called "benchmark" isn't economically useful, but that it tests for the sake of testing; and for attention.
To end this obsession with generating SVG pelicans, Quiver AI's [0] model is actually designed to generate SVGs from prompts and has done so for years.
So there is no need to continue with this un-serious pseudoscientific "benchmark".
[0] https://quiver.ai/
Does that even test for intelligence?
This is like testing if a horse can fly just because someone showed an image of a Pegasus, or testing if a fish can climb up a tree and believing they are not intelligent because each of them cannot fly or climb up trees.
Pulling things into the ridiculous isn't condusive to a good faith discussion and does not help proving pseudoscientificness which you seem to be after. (I don't belive that the pelican bicycle test is meant to be a serious scientific endeavour, btw.)
"Draw an animation of this long ass scene from a movie, and only call me when everything works e2e" can be.
Consuming radium and using uranium glass, that’s what we’re doing.
a real benchmark is instead running evals on your own traces, and building a cost/quality/speed profile for models based on real workloads. but it doesn't get you a shiny video you can post on twitter.
similar to cpu benchmarks (e.g. geekbench, which also had a recent update), the goal is to distinguish model in some arbitrary thing. however, it is not quantifiable but subjectively even the latest show distinctive differences regardless.
large context window is still relatively new and not "feasible" even commercially, so making that a requirement for a benchmark would not make it accessible especially for open-source ones.
* yes, i misspelled it intentionally
Bilbo's house is actually described in The Hobbit and the exterior isn't really described at all in Fellowship. The prologue of Fellowship (Concerning Hobbits) mentions hobbits like round doors and windows and the fact some hobbit homes are underground, but the turf-dome design here is not mentioned. It actually mentions hobbit homes typically have bulging walls, so unless you've read the The Hobbit, you might not picture this entirely-underground style.
In The Hobbit, his home is described as a (nice) hole in "The Hill" with a perfectly round front door and round windows, which could imply the design here.
https://arxiv.org/abs/2303.12712
Maybe excitement and wonder, in tech, is just something for us old guys, that watched it all be birthed. Get off my lawn!
https://static.simonwillison.net/static/2026/fable-max.jpg
I’ve been trying it on them all and can’t find one that does it consistently. The best will tell me they can’t. The worst confidently point out one of countless Waldo-likes.
I think LLMs will be excellent glue of "find the right function/button and run/push it" but design without constraints and they just explode immediately
Why do you say this? I have some experience of manufacturing processes and am not seeing where AI would be useful other than to drive the robots which we can already do quite well without AI (see all the dark/lights-out factories that already exist).
Where do you see it being useful? An example would be nice.
I've seen them make time estimates that are physically impossible, expect medications to be in two places at once, etc.
Asking AI to design real world objects doesn't work very well because all of its tests involve proxies and thus miss things that are glaringly obvious when the object is actually built.
We already have sneaker designs and the equipment to manufacture them. Whatever it spits out is going to be, at best, a mediocre clone of something that already exists. What exactly is the point?
fires a missile at Pakistan
Here is your refactored component:
[0] https://news.ycombinator.com/item?id=48838228
Definitely impressive demo, I do wonder though if the countless artwork, films, images etc produced over many decades around Lord of the Rings somewhat influenced the outcome of this though.
Everyone knows roughly what a bicycle looks like - wheels, frame, seat, peddles, handlebars etc, but the details of the frame and exactly how the other parts connect to it throws people off. It seems people memorize the "concept" of a frame, but not the specifics.
Try it without cheating, then google a picture of a bicycle!
Even if you have memorized what a bicycle (and pelican) look like, writing code to draw one is not the sort of thing humans are good at any more than they are good at mentally calculating cube roots - better to use a computer for stuff like that.
I wonder whether we are entering the era of throwaway software. Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price, maybe LLMs give us the same for software. Produce it cheaply and if it breaks throws it away and reproduce it.
Why would anyone still use off-the-shelf software when they can have a system that has access to all data, can transform it into any form, and can export it in any format?
After years of thinking that I needed to develop a decent movie management system for my own films or a columnar browser for large CSV files, Claude and Qwen each delivered exactly what I needed in just a day.
On the one hand yes, almost every task I work at now is one off one of scripts I throw away.
The question is - where does the software the spec or the “code”.
A really complex game will probably always be token heavy. At least for the next few years code is still not free.
But certain software is just iterative by design. If we mean we regenerate all the for loops of a game from scratch, sure but I think “code” Is really more spec then implementation, and we’ll want to continue building things through iteration.
And even on the for loop point - Do you really want to spend millions of tokens rewriting a game every time you need to make balance changes?
Have they? Most of the world production is tied down to expensive factories and machines. Yes, we have more products, but that the result of the global trade, which is a very complex system.
> Produce it cheaply and if it breaks throws it away and reproduce it.
I don't know why everyone would ever wants this. It's been parroted since forever, but the true usefulness of software is to be able to build it once and runs it indefinitely. If some edge case occurs, I fix it. Which is way cheaper than rebuilding the whole thing. The goal is to have something like OpenBSD's ed[0] or dmesg[1], which you only touch every few years or so
[0] https://github.com/openbsd/src/commits/master/bin/ed
[1] https://github.com/openbsd/src/commits/master/sbin/dmesg/dme...
https://svg-bench.fenx.work/
I think it's useful when evaluating "model+intended harness", but I'm more interested in seeing raw model improvements than harness improvements.
However I do agree that the results are much worse when you try to use them. They look great in screenshots and video clips which makes them perfect for content farmers.
All of the LLM generated games I’ve played have been really bad to play, though. I even tried my hand at a simple game, thinking I could iterate on it with prompts to fix some rough edges. After the initial productivity burst every change turned into a slog of tokens with one thing changing and something else breaking it. I would try to use my remaining weekly token budget across Anthropic and OpenAI to refine it at the end of every week but after a couple weeks it felt like I wouldn’t be getting anywhere without scrapping it and going back to having the LLM build it one step at a time with my careful instruction.
Which, in retrospect, is the only way I can get usable output of an LLM for anything complicated, so it’s not surprising. It’s a fun reality check project though.
On the one hand I think most of us are incredibly impressed because we know, that quick demo would have taken us months of work to build in the before times.
On the other hand the promise is a cure for cancer and the end of all work.
So when everyone is telling you “skill issue is why you can’t one shot WoW”. It’s hard to know how you’re supposed to feel about Karpathy advertising one shot custom virtual worlds but giving you slop. Incredibly impressive slop when compared to how long it would take to create it just 3 years ago, not so much compared to Elon saying - “by the end of this year, grok will create a version of the odyssey that competes with Nolan’s”
I guess this is the average story and, similarly, the average game is boring and predictable
Early generative AI at least had the virtue of relentless, unsettling weirdness, in the same way that generative art from the late 90s and early 2000s did. A handful of people made creative use of that spooky weirdness.
Now it turns out "Airspace" art.
If you took the best, most creative, human writer in the world, and for thought experiment reasons they had amnesia (to mimic AI blank context windows) specifically while you asked them for a story idea 100 times in a row, my expectation is that this human would also give you the same idea at least 80 times out of that 100.
* still better than the mean human, but even the top 0.1% of humans aren't all professional authors.
What you think is better is not what I think is better. Imagination is not storytelling. You ask the "mean human" to write, it's going to be worse than an LLM in spelling and grammar, if you get anything at all.
Most humans never do a creative writing course after school, and the longest fiction most people will write is their resume description of what their previous jobs involved, or perhaps their dating profile.
Don't mistake what you see published (or what your friends are like) for the average human: the average Hacker News comment easily above the writing grade (and creativity) of e.g. many of the one-shots stories I've seen attempted on some creative writing subreddits. "Mean" is not a high bar.
This should cause you to reconsider your beliefs leading up to it.
GPT-4 (!) has, in studies comparing multiple models with humans for creativity, beaten the mean human. In comparison, even just the mean of the top 50% of humans beat the models studied in that case (link follows), but the point is that if you think the mean humans is particularly noteworthy, you've avoided the half of the population who have the creativity of a pot noodle.
https://www.nature.com/articles/s41598-025-25157-3/figures/2
There's more studies out there with other types of creativity test, but the conclusion is basically the same: the best models beat the mean human, but are nowhere near as good the worst *publishable* human.
My guess is this is both why LLM slop happens and why it grates so hard: a significant number of bosses look at the output and think to themselves "wow so creative" because it's more creative than they themselves are; but those bosses weren't hired to be creative, they were hired to be a boss, and all the people who they hired to be creative are going "arg, no, can't you see how bad this is?"... but that is just a guess, I've not found any surveys comparing *management* creativity to LLMs, closest is e.g. this about decision making, not creativity: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5156585
I did get pip, but it's a story about trying to rescue the moon from a lake and the raccoon lives on moonberry lane
What's pip from?
So "moon" is also part of the predigested artifacts. Maybe because raccoons are nocturnal?
These conversations get frustrating since one group is saying the dancing is bad, a second group is saying it's good (for a bear), and a third group is saying we are a few years from a bear-only dancing industry.
There was a vision of the Hololens being a general purpose compute device that enabled something more like the power of a desktop without being tethered to a sitting position at a desk. You could see it if you squinted hard (incidentally, the device caused a lot of eye strain). But even at the height of the VR hype cycle, I didn't know anyone pushing the Hololens on anyone. Nobody was telling the world they were going to be using Hololenses v1 or v2 as their only computer "or get left behind."
That's oddly specific and it'd really hurt my cousins feelings :)
I guess the next question is if they can be made fun without too much additional work with a human guiding the AI.
They can't. If you think about how these things are trained it's blatantly obvious fun is an impossible metric to optimize them for
..feels like there was a black mirror episode about that though
One would hope it accelerates AR/VR - if decent studios finally get behind it.
If I see one more “one shot MMO” where you just walk around and do absolutey nothing or another menu slop idle battler or rogulike deck builder I’m going to go Postal in Minecraft.
Used to be if a game looked that good it probably had time spent on the game part too.
Whenever I think about AI games I find myself thinking about Tiny Wings, Flappy Bird and Angry Birds. Three simple, elegant games.
It is easy to see what makes Tiny Wings so completely loveable — it has a sculpted, adorable, perfect charm with a cleverly inverted game mechanic that has a calibrated level of exasperation and reward.
But why were Flappy Bird and Angry Birds, very basic games with very old game mechanics, so charming?
It seems equally impossible to imagine an AI coming up with a game with the quality of any of them, even with maximised creativity. But explaining why for Flappy Bird seems quite difficult, especially when you consider it uses some stolen visuals!
Maybe it's the smoothness of motion that makes these games understandable and LLMs seem to consistently fail at that. Ask them to do something snowboarding and they go really hard on the physics since it seems like they don't know what kind of approxmations feel good.
(It's been a while since I was in the game industry, so IDK quite how accurate this is).
These one-shot products aren't games. They're barely even demos. I don't even know what to call them. For a mature framework like Phaser to sell-out like this and create a vibecoded platform for vibecoded games is shocking.
But when visual perception and "taste" get involved, it becomes a lot harder.
As with painting, after a while there's nothing really new to paint, we genuinely need 0 new software. We need to fix our broken physical world, our social lives, our kids and what's left of our democracies.
This software crap is done, leave it to the nerds.
Textbook and blackboard > ipad.
Again comes back to the point of verifiable rewards. The moment you take it away from LLMs they just stop being as good
Karpathy’s point about the models not being able to easily audit their work is something I’m struggling with —- how can the audio output be made perceivable? Curious what people think about this question.
Here’s a synth to play with+remix
https://underscore.audio/s/cmp_8b226859-420/iron_rareflash
It always brings to my mind some words from Rich Hickey:
I don’t think I really have a point to make here, other than it just feels like someone’s released a bunch of carnival bumper cars onto the highways.Not difficult to see why the employees of AI firms are thrilled with it though, eh?
I like the analogy.
I guess this is why we got this before cars sold without steering wheels: literal guardrails on the literal roads are somewhat more expensive, especially for the people who keep bouncing off them on the way to their destination.
Also, where the guardrails are absent: oh look, felonies. https://www.google.com/search?q=ai+hacks+company&tbm=nws
Take chess engines of the "deep blue" era for instance. These engines are stupid, trying millions of moves that are obviously terrible, no human chess player would do that. And yet, this is what worked best, because computers are so good at repetition. Recent development using neural networks made chess engines smarter, but using repetition is still how the beat humans.
What you call "guardrail programming" in another context would genetic algorithms, a technique that has recognized applications. And if you look at videos of genetic algorithms learning to play racing games, it literally looks like a bunch of carnival bumper cars onto the highways, but done well, after some generations, it becomes competent driving, sometimes even record breaking.
It can be expensive, it is often the case for LLMs, so you may want to try to be a bit smarter at first to spare some resources, but to me, it is just using computers as intended.
https://threejs.org/examples/
Here's minecraft: https://threejs.org/examples/webgl_geometry_minecraft.html
Here's an FPS: https://threejs.org/examples/games_fps.html
The library is extremely well-documented. When Three.js vibe coded projects started blowing up on Twitter 1-2 years ago, I wasn't that impressed then either because I knew what it was doing.
Anyone who remembers the C compiler built by an LLM! backlash probably feels the same way:
Why would I use an LLM to create a well-known demo rather than fork that demo itself?
What I haven't seen yet from an LLM is it create anything fundamentally new and exciting: A new UIX that is actually good. A new service that is actually good. A game with an art style I haven't seen, music, storytelling - something you'd expect out of a AAA studio.
Given that rant: The coolest part is the multi-modality between text and animation. However, I think the end product would have been a lot better if it was just a video. Having it do it in Three.js didn't add a ton of wow factor for me, and it would have been a lot better looking as video.
> Something like an ephemeral GTA of X on demand.
Here's where you lose me. AAA gaming is very far away from this Three.js demo. But the novel part being the syncing of narrative to the visual scene - a text-to-audio (video) book type technology seems very possible (and useful).
Nice work. We need more big projects like this involving AI (if anything, to get away from the slop argument largely focusing on 1-shot experiments).
That method of perception probably scales N^2... so sure with more compute, LoTR animation will improve. But I think to get a real jump in "experiential feedback", perception needs to scale linear or sublinear. Maybe that's there LeCunn's jepa will come in.
There needs to be the removal of the middle man:
image -> text -> action
To image -> action.
> it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free".
How is two hours worth of token generation free? With tech revolutions things get cheaper/faster/better/doable, but there’s still real world limits. The advent of railroads made it feasible for the average person to cross the country, but it still cost a lot of time and resources. People weren’t crossing the country every weekend for fun just because it was now doable.
Why do we treat LLMs as ~free when we are generating things that weren’t doable before but have to invest more money than the Apollo program to build AI data centers let alone account for the operating costs?
This is an odd take, given that Karpathy is certainly aware that the LotR films absolutely did create Bag End in digital format; that their creation was outstandingly high quality; and that Claude’s output here very obviously “leans heavily” on their prior art.
telnet towel.blinkenlights.nl
Now, were they „in their right mind“? I don‘t know. The more likely explanation is they found joy in it. Reminds me of what the Suno CEO Mikey Schulman said about making music: „[…] I think the majority of people don’t enjoy the majority of time they spend making music.“ (https://news.ycombinator.com/item?id=42688538)
8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.
Or maybe over these 8 months agents improved a lot? You know, few years ago many AI experts predicted that things we are routinely doing now with AI are decades away. I mean how can you look at this post and not be impressed? It's insane what AI is currently capable of.
[0] Seriously. Get used to mentally prefixing his and Boris Cherny's name like this, every time you see them quoted. These people are speaking while employed; there is no chance they are not aligned with the employers who will make them wealthy. The tech industry does like to pretend that for some reason AI people, uniquely, speak thoughts unbiased and for themselves or even for science or humanity.
I gave Fable a jpeg and asked to draw an SVG, using a loop that renders the SVG into an image so Fable can inspect it.
Results looked like drawing of a 5 year old.
I wonder if Flash is still popular... LLM can use that instead...?
There are people in their right mind who would do that and their are already examples of people who did similar things.
But maybe not in the future if people would confuse all the effort with AI
Computer graphics will have enormous applications because they are directly controllable by LLM-generated code. Video models are probabilistic and less suitable when precision matters. In education, for example, we need exact visuals. If an AI wants to plot y = sin(x), it should generate the precise graph through computer graphics rather than approximate it with a video model.
What I find funny is that approximate to computer graphics are video games. When I ask AI about a decision available to me in a video game AI completely fails, OFTEN. I assume all the forums and changes made to a game over time might be quite confusing for AI. But I've also seen it completely make up characters and decisions and weapons and so on about some very clearly defined games and paths. It's an interesting dynamic.
that's why these things are actually pretty good at openscad/freecad/F360 mcps , the visual reality is enforced and guaranteed by rigor in the interpretation engine that is anchored to human physical reality.
I think it's interesting that the "Bag's End" interpretation in the video clearly looks like the one from the movies, but generated here as a three.js 3D asset.
It makes sense that the movies (or shots/frames from them) were in the training data, and I can also easily imagine an association in concept space between the textual description of Bag's End and the frames from the movie.
But how on earth does the model then go on and convert the latent representation of those images into coordinates for a 3D mesh, without ever even restoring the image? In what kind of representation are the images from the movies stored that it can do that?
Better than I would have expected? Sure. Impressive that it got that far? Definitely. But is the end result good?
I like where Karpathy is going; I had the same thoughts about LLM generated slop scenery. I just want some variety of scenery for the goblins to get massacred in in whatever fantasy slop game I play.
It’s not. That blocky appearance and limited lighting was deliberate comic aesthetic choice as a response to budget, not a result of technical limitations. CGI in 1985 was enormously more capable than that — take a look at the stained glass knight in Young Sherlock Holmes, or The Last Starfighter from 1984. Things certainly could have been rounded, sculpted etc.
Yeah, please check in with the folks protesting data center builds in their town causing their electricity prices to skyrocket and tap water to turn into a scarce resource.
And now, instead of actually doing the above, please go ahead and downvote me, because how dare he question those LLM games.
It’s a useful benchmark (aside from being “cute”) because of its simplicity, both in how many output tokens it takes (though I understand some models think a lot now to do it) and how easily one can subjectively judge. It’s this efficient as a benchmark of performance.
Making a long video takes way more tokens, and presumably is a lot tougher to easily compare. swillison has a presentation that’s pelicans from 2023-present (roughly) showing the progression. Imagine “lord of the rings videos from 2026-2029” or whatever, it would take a long time to watch and be harder to judge, and probably just end up being a comparison of screenshots anyway.
TLDR I feel like the post misunderstands the role of the pelican thing though if find it very hard to believe he really doesn’t understand, so maybe I’m missing something.
Except it's not ~free, it cost ~$10. And no one in their right mind would ever exchange $10 for that crap output except in this brief moment that we're in because it's fun and surprising to see what will happen. The actual result is as close to useless as it's possible to be - it's not interesting in its own right, it's not aesthetically pleasing, nor funny, nor informative. It's just slop.
There’s a reason AI slop games have literally zero engagement. Last summer that stupid flying game blew up. Maybe a million people “played” the game. Where play means they clicked a link and checked it out not because of what the game was but solely because of how it was made.
In terms of concurrent players that game wouldn’t have cracked the Top 5,000 on Steam.
My metric for AI games is “number of players who spent more than 15 minutes playing”. I’m not aware of any vibeslop that has achieved 1 such player.
Now obviously LLMs are transformative for game dev. But “hyper custom worlds you can drop into” shows an extreme ignorance of what players want imho.
From my experience (PixiJS game), Opus can look at things and screenshot. It's slow but it works. The issue is that it doesn't try to make them look good.
Instead, it tries to find one bug and then fix that one bug and close the session. That is amazing discipline for webapps or whatever, but for game design is just the worst if you need attention to detail and tuning of multiple holistic systems that produce an effect together. That really limits the kinds of approaches you can can use to achieve things visually.
Sure, the testing harness could be better, but that's not what makes it a poor fit. The base model seems also capable (it can identify the issues, just isn't willing to disangage from this "I found X and fixed it" single-thing work posture).
The kind of work is just different, and it was probably never trained on it, so it feels off and an uphill battle to use it.
Testing, assessing, tasting...
We also do it when we build other things - this is just a different scale.
Maybe nozzlegear wanted to suggest some importance on Musk remaining a bet-ter on the general tech, regardless of the competition?
> "Yah": slang spelling of the word "yeah" (which of course can also be used ironically)
> Merriam-Webster: "Yah": used to express disgust, contempt, defiance, or derision; probably imitative of the sound of retching
240M followers vs 1400 of which 99.9999% are bots with the rest being Tesla fanatics.
That is close to no-one.
You cannot come and place your personal positions as assumptions. To me, there is absolutely no theft. And we cannot play a game of "Yes!"//"No!" here.
By the way: are we having a surge of this?
We do have a surge of pro-AI sealions, yes. Any objection is countered with one or more three word questions.
Very devoid of intelligence note.
> Any objection
Objections are arguments. That post did not start an argument - it was as ideological as the Brigades. Devoid of what we want to have here (I believe).
I have stated and do state: what is published is assumed as read (only, probably not read for lack of resources). If it is in the libraries, it is there to be read.
No bias. No speculation. Just facts!
LLMs should be tested in the same way people should be tested for a job interview (but often aren’t) - with tasks RELEVANT to usage.
So you don’t just randomly pick some random thing to make the LLM randomly do (like many job interviewers do).
You start with clear statements about real world usage scenarios. THEN you come up with tests that give insight to how well the LLM/hob seeker gets the job done.
Please, stop coming up with random tests like it’s Microsoft in 1990 and you’re asking job seekers how the would move Mount Fuji, as a way of assessing their programming skills.
No stupid irrelevant pelicans on bicycles and no stupid renderings of Lord Of The Rings. Unless those are relevant use cases.
Any test that anyone comes up with must clearly state the context and how the outcome is measured.
But this demonstrates we're a couple orders of magnitude away from generating 1:1 hyper-personalized entertainment and media for individuals, rather than the masses.
On top of that, a decent chunk of the joy of entertainment is the social aspect.
Now, whether this would be at all possible, or even healthy for you, I don't know.
Ultimately, most people don't have ideas for the kinds of personalized entertainment they want, and they don't want to be in charge of content production (even if you have an LLM do most of the work). I don't doubt that there are niches for it, especially stuff like porn, and I'm sure that pros (game studios, film studios) will leverage AI more and more, but I suspect that most of us will just want to sit on the couch, watch Spiderman XVIII, and then be able to talk about that shared Spiderman XVIII experience with all our friends.
LLMs produce an extremely small range of artistic expression, and its kind of ironic, that while it certainly looks functionally good. It immediately becomes noise because everything looks identical.
So companies are going to have to hire creatives again, in order to stand out, and be "creative"
And this personal media thing, yeah maybe for terminally online persons that are REALLY into a sub-genre but otherwise, it's too much effort, I agree. Until we get machines that can read our (subconscious) mind, that will not exist.
They were visually bad but honest and sometimes soulful or playful.
The AI slop that's everywhere now looks superficially more professional but it's very busy, samey, unnatural and it really turns me off.
Not even close. Current AIs have very poor spatial awareness, they can generate some kind of scene but they can't tell you where objects are in the scene, nor can they move objects into different places. They're very useful but it's difficult to be very creative with them because they can't update an image to make it more aligned with your vision for what should be in the image.
1:1 AI entertainment probably won't be a cold start experience.
Much like the "Choose Your Own Adventure" books of the 80's, consumers choose a baseline template, customize the characters, and interact at various points within the plot.
No, we don't. It's quite good, but it's nowhere near perfect. For example, you can see: https://genai-showdown.specr.net/ or others that highlight how far from perfection we still are.
> and it hasn't really changed the nature of human expression.
It sure has changed our discourse. Look at hackernews, where people just can't help themselves to engage in ragebait and flamewar comment threads about llm generated accusations. We've got half the people convinced that looking at ai content is like consuming food.
I'd venture a guess that it's too early to determine how this technology will change expression. A good analog would be photography, which caused a similar meltdown in the arts at the time.
> I don't see my friends getting wildly creative.
Maybe you don't have creative (enough) friends.
The fans were real heart broken about this but I think you're right on where this is going. We're not going to be seeing the dominance of centrally produced content like this for much longer, like sure, I think there will be big blockbusters will stick around, but I think the day is coming where media becomes a choose your own adventure sort of scenario.
It'll be interesting to see where this scales to. there will definitely be some amazing solo projects but we'll also see the like 4 player co-op version of productions and then the larger mine-craft server 'Minas Tirith' scale ambitious projects that involve a few dozen people. And of course passive consumers will remain a thing, or people who just provide some suggestions or nudges for what they'd like to see others make.
But I don't think it'll be dominated by big companies like Disney, Netflix, Amazon or Paramount.
I'm sure a lot of them will suck but it'll be neat to see the inevitable Seinfield - Star Trek Voyager cross over episodes.
Elaine and B'Elanna Torres feud after a transporter accident leaves the crew stranded the delta quadrant. Jerry attempts to date 7of9 but is rebuffed as she finds Kramer's quirky bluntness more relatable. George panics after someone compares him to Neelix.
Everyone knows how a MacBook looks. Any missing or incorrect detail will be obvious. However, Pelican or this world can be anything.
https://playcode.io/blog/macbook-svg-benchmark
EDIT: The parent comment originally claimed they spent $1M on the demo -- they seem to have edited it after I replied.