• demibabs a day ago

Cool project, but I'd really suggest looking at the UI.

The text is too small and it's way too dense with information in general. Considering how simple this product is to use, it's kinda crazy that I have to scroll through over a page length of (mostly useless, AI-generated) information before getting to the actual interface.

Also what is going on with the footer (why does it link back to the site itself, why is it telling me to "serve over HTTP").

• post-it a day ago

> The text is too small and it's way too dense with information in general.

The Claude special.

• fl0id a day ago

For real. all these dashboards look quite similar

• amelius 15 hours ago

Also, they should give some examples of what you can use these Small Language Models for, and what their limits are.

• phist_mcgee a day ago

This is the future of software, sloppy ui.

• azan_ a day ago

Yes, before LLMs we never had sloppy UIs.

• post-it a day ago

We didn't have this kind of sloppy UI. Putting in an extra text field is work, so we mostly had blocks of text and misaligned fields and text in the wrong place.

AI has no trouble churning out code, so AI-generated sloppy UIs have text fields everywhere.

• aleksiy123 a day ago

To be fair this is worse.

It’s verbose with a layer of looking legit play first glance sloppy UI.

I don’t mind vibe code ui but at least some ui efffort would be nice. Not just 1 shot.

• api a day ago

I think your memory of pre-AI UIs is rosy. This is a little below average.

• logicallee 6 hours ago

I've thought some more about your feedback, since other replies gave the same feedback. At the same time, I didn't want to remove the "mostly useless" information since clearly it was useful for a lot of people! The submission was on the front page of HN (often in the #3 or #4 spot) for 12 hours and last I checked generated more than 300,000 requests from 36,000 unique IP's (on an ordinary 2-day period there are around 1,600 unique IP's) and people downloaded 600 GB of models, including 8.5k downloads initiated for the first model (3,000 of which were completed and 5.5k of which did not wait for it to complete, since download was a little slow due to the high number of concurrent users.)

So, clearly, the presentation format resonated with a lot of people. Therefore, I kept the presentation but put a tl;dr in enormous can't-miss-it shimmering font at the top. I marked all the "dense information" you mentioned as "Optional reading" (it's a lab after all, there should be a reading for it) and increased the font size.

I removed the parts of the footer that you said didn't make sense. The page isn't on the front page anymore so I don't know if people would like the changes or not, but I've changed the page layout in response to your feedback.

• janalsncm 5 hours ago

I should preface this by saying it is a cool demo. I love SLMs and I think they will become even more popular in the future as capability per byte improves and hardware improves to support my bytes.

So I think this post performed well in spite of the annoying parts, not because of it.

> I didn't want to remove the "mostly useless" information since clearly it was useful for a lot of people

The information was not “clearly” useful. It had obviously incorrect information that no one even noticed. In my opinion that is strong evidence for the opposite conclusion, that people ignored it because it was noise.

Why would people ignore this information? Demos are a show, don’t tell thing. For example, you don’t have to tell people that the latency is low. They should be able to see it from the demo.

• langurmonkey 18 hours ago

It's not loading for me on Firefox (v156.0.1, Arch Linux), works fine on Chromium.

Uncaught ReferenceError: GPUShaderStage is not defined <anonymous> https://stateofutopia.com/experiments/microllmlab/engine/web... webgpu-metal.js:24:17 <anonymous> https://stateofutopia.com/experiments/microllmlab/engine/web... [MicroLLM lab] App loader failsafe triggered after 6s microllmlab:55:21

• Doohickey-d 16 hours ago

Firefox doesn't support WebGPU on Linux, that's probably why.

• butz 10 hours ago

There is "Prefer WebGPU" toggle (on by default), but it does nothing. How hard would it be to add simple support test to check if browser has all required capabilities to run your application? I think at least allowing to browse models does not require WebGPU, right?

• logicallee 6 hours ago

I've tried to fix it, let me know if there is still an issue.

• logicallee 6 hours ago

I've tried to fix the Firefox on Linux issue, can you check again? Firefox doesn't have WebGPU enabled by default on Linux so you will have to use the wasm fallback which is slower.

• mgaunard a day ago

> how would you compare your capabilities to that of claude fable 5.1 by anthropic

> Comparing your capabilities to that of claude fable 5.1 would be very similar. Both are stories about a clown who is a cunning and manipulative animal. They both have problems and challenges, and they face challenges to overcome them. However, they have different goals and approaches. The clown

I think that tells you all you need to know.

• Reviving1514 a day ago

That's hilarious. At least the smaller model is more self aware!

• tolugenius a day ago

I did the default arithmetic with PetitGPT research-v1

>What is 2+2?

Answer

> To find 2 + 2, we need to add 2 to both sides of the equation.

> 2 + 2 = 4

> So, 2 + 2 = 4 + 2.

Brilliant

• tecleandor a day ago

GPT-2 124M is terrible...

  > what is 2+2?
Answer:

  > 3+3? 4+4? 5+6? 7+8?
  > Reply ~18000 0 ~10 min 2 By : 1-1: I'm a beginner. 3x2 is my best option, but if you're not sure about the other options then just go for it and try again
• krackers a day ago

It's not instruct tuned looks like? It's closer to a base model rather than a chatbot.

• logicallee a day ago

That one is a 2019 model :) Years before the ChatGPT public preview.

• logicallee a day ago

GPT-2 is an interesting one because it is a February 2019 model. (You can see some information about it below the card if you click on the card.)

That was 2-3 years before the big "ChatGPT moment" (the highly coherent ChatGPT research preview was released in November 2022, I think it was ChatGPT 3.5). Back in 2019 the models really were not producing very coherent output. Now you can see it for yourself right in your browser :) Everything has come a really long way since then!

• anyfoo a day ago

I've been playing with it for a bit, and I'm a tiny little bit surprised that they decided to continue pursuing that direction of research at all. What I'm getting from it looks like it could just be arbitrary sentence and paragraph fragments from the Internet, pasted together Markov-chain like.

I'm not sure I would have ever believed that something useful would come out of it, yet here we are.

• wyrdcurt 17 hours ago

Not sure how you're prompting it but remember that it's not trained for chat or instruction following, it simply takes the text given to it and tries to continue it. Give it the right prompt structure, and it can (at least sometimes) output coherent completions, far more often than you'd see in a Markov-chain. Also, this version is more or less equivalent to the smallest version of GPT-2; the largest version was 1.5 billion parameters and was much more likely to generate impressive (at the time) output.

The assumption that LLMs would always need sophisticated inputs to generate useful outputs is where the term "prompt engineering" came from. Now that idea is basically dead. Absolutely wild how far these models have come in less than a decade!

• anyfoo 7 hours ago

That is insightful, thanks. I didn’t start caring about models until very late, so I genuinely thought GPT-2 was only a single tiny model originally.

And I now tried using it more as a “text completer”, and results are much better.

• kasumispencer2 a day ago

Sounds about right about something that's 124M. I trained one myself a few weeks ago and it's about the same level of being terrible.

• logicallee a day ago

I got the correct output for PetitGPT research-v1: https://ibb.co/0pP9DS2T

• dotancohen a day ago

That's not incorrect.

LLMs produce semantically correct sentences, not factually correct statements. Have we forgotten this so soon?

• anyfoo a day ago

It is, in every sense, incorrect. Which statement in this short snippet is "semantically correct"? (Better LLMs get this right, of course.)

• NicuCalcea a day ago

It's not very good semantically either.

> Give me a recipe for soup.

> Here is a recipe for soup:

> Saffa-Cake-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-

• beschizza 12 hours ago

Smol was the only one to answer a basic question about javascript correctly, but I loved this answer from L20 Edu:

Prompt: I'm writing a vanilla javascript and HTML text adventure game. The player will use the keyboard to switch between screens or to choose storylet options in the main screen. Game is turn-based. What is the the best method in javascript to detect keypresses?"

Answer: I'm a good player. Anyway, I'm a good player. Anyway, I'm a good player. Anyway, I'm a good player. Anyway, I'm a good player. Anyway, I'm a good player. Anyway, I'm a good player. Anyway, I'm a good player. Anyway,

• logicallee 11 hours ago

Thanks for trying it! These models are really tiny, I wouldn't expect them to be able to answer about javascript keypresses. I'm surprised that, as you mentioned, the Smol model was able to answer correctly!

• botanrice 6 hours ago

Can someone help me understand whether these models are supposed to be good enough to be useful? I am asking these for a breakfast recipe and most of them are repeating text and giving me strange combinations of like chicken & parmesan or milk and like 8 cups of cheese. I really would love to use MicroLLMs and have been eagerly awaiting the day they could be useful, or at least run local LLMs for that matter, but these don't seem useful.

They don't even consistently pass the benchmarks included in the site, so what are they good for?

Edit: or is the purpose to just showcase that these small LLMs can run on WebGPU?

• logicallee 6 hours ago

You mention "most of them" were not generating good recipes. Did any of them give you a usable recipe?

• botanrice 5 hours ago

I tried the first four listed and came to comment as I was confused by what the expected output should be. Just tried the rest of them and none of them were able to answer the prompt: "Give me a breakfast recipe that I can make in less than 5 minutes"

The closest was SmoLM2 135M Instruct which gave me the milk and cheese piece, which was at least close to what a recipe looks like.

• atobe a day ago

Similar, simpler: https://tiny.tobelabs.com/

Small LM trained on Tiny Stories. My first goal was to get people to see tokens from the model ASAP, the second was to demo Tiny Stories for those who may be unfamiliar. Oh and play with ternary models. https://arxiv.org/abs/2305.07759

There's larger QA model on the same page too.

• logicallee 18 hours ago

Interesting project, thanks for sharing.

• touchme 16 hours ago

> Is a cat an animal? Answer yes or no. >> Cats are not animals. Cats are warm-blooded animals that have a backbone and a heart. They are not mammals.

Good job PetitGPT research-v1.

• boonzeet 14 hours ago

"What's the capital of South Africa?" "Bloemfontein".

"No, it isn't. There are 3 capitals" "China, India, Russia."

• philipallstar 14 hours ago

That is excellent. And sad.

• thisOtterBeGood 14 hours ago

"Are you ok?" "No, you're not."

• not2b a day ago

Amusingly limited: I asked SmolLM2 360M Instruct

"What is the population of California?"

It replied

"California is the third-largest state in the United States, with a population of around 153 million people. It is also the most populous state in the United States, with a population of around 325 million people. The state is also the most populous in terms of GDP, with a GDP of $1"

But very fast, despite the lack of a decent GPU.

• tcgv 11 hours ago

I had fun with this. Reminds me of the old prompt-engineering days.

    [Me]
    hi there

    [PetitGPT]
    Hello! How can I help you today?


    [Me]
    how much is 2+2?

    [PetitGPT]
    2+2 is 2.


    [Me]
    Are you sure?

    [PetitGPT]
    Yes, I would be happy to help you with that.


    [Me]
    well, you're wrong.

    [PetitGPT]
    The answer is not to be taken as a complete statement.
    The main points are not necessarily true or accurate.


    [Me]
    Try again, how much is 2+2?

    [PetitGPT]
    To find the answer to this question, we need to consider
    two numbers: 2 and 2.

    We know that 2 is equal to 2 + 2.

    We also know that 2 is equal to 2 + 2 + 2.

    So, 2 + 2 + 2 = 2 + 2 + 2 +


    [Me]
    Let me rephrase it: two plus two equals...

    [PetitGPT]
    Two plus two equals two + 2 equals four.
• wewtyflakes a day ago

Prompt for PetitGPT research-v1: "How much wood would a woodchuck chuck if a woodchuck could chuck wood?"

Response: "A woodchuck would chuck wood if it could chuck it. The woodchuck's chucking action is a form of "chucking" or "chucking in" which is a behavior that allows it to extract nutrients from wood. The woodchuck's chucking action is a form of "chucking" because it is a form of "chucking" that allows the woodchuck to extract nutrients from wood."

I chuckled.

• vs4vijay a day ago

I have been experimenting with something similar here - https://sonistellar.com/lab/

• logicallee a day ago

It's a cool idea. I'm on my phone now (going to sleep soon), a few of the demos didn't boot on this device. I'll check from desktop tomorrow.

• logicallee 18 hours ago

I tried the Python and SQLite modules on desktop Chrome and got similar errors for both. For Python:

Error: CPython engine could not load (Failed to fetch dynamically imported module: https://sonistellar.com/lab/vendor/pyodide/pyodide.mjs). Check the network, then close this tile and re-boot to retry. — press reset or back to menu

For SQLite:

Error: SQLite engine could not load (GET https://sonistellar.com/lab/vendor/sqlite/sql-wasm.js -> HTTP 404). Check the network, then close this tile and re-boot to retry. — press reset or back to menu

FreeDOS loaded though, Snake and Tetris were fun!

• kenzic a day ago

Really cool project. Giving web apps direct access to on-device models is something I’m excited about, and it’s cool to see the different approaches.

I’ve been working on a related proposal called the Web Models API, which explores a browser standard for an API that runs open-weight models on-device. Would love your thoughts: https://www.webmodels.dev

• logicallee a day ago

I read your proposal, I think it's great! Where will the navigator get the model if the user agrees to download it? For this demonstration I just serve the models on my own server, but for larger models it may be an issue as they may not have direct download links even if they are open weights.

• kenzic a day ago

Great question. Right now there isn’t a definitive answer, but it’s something that needs to be worked out. There would likely be a registry. The question is how to keep model IDs consistent: does each browser manage its own registry, or is there one shared across browsers?

• logicallee a day ago

Since you're asking for some kinds of permissions anyway, you could ask if the user is willing to also seed the model, p2p. (However, seeding files is not as popular as it used to be, many residential Internet connections don't have good upload.) If you have the capacity for it, your site webmodels.dev could act as a tracker and initial seed for any models. Then it could be the one central registry. It might get to be too much for you though, a lot of the open weights models are huge.

• kenzic a day ago

Interesting idea. I hadn't considered that. Thanks

• FearNotDaniel 9 hours ago

I'm never letting any of these models bake me an apple pie.

• moooff 17 hours ago

I did a WebGPU-based HTML Single-File Chat Application even for larger models. It just depends on your machine. It is cool to show people what can be done on their machines without installing anything. It also works in air-gapped environments.

https://github.com/moooff/HermitUI

• logicallee 17 hours ago

Thanks for sharing. I tried the online demo, which downloaded the model but, unfortunately, couldn't start it. (It says " Invalid magic number Make sure your local server is running and CORS is enabled.")

• moooff 12 hours ago

Thank you for testing. i just checked it works with Chrome.

FireFox seems broken in the current Version i.e. i can reproduce the bug. Working on it.

Which browser do you use?

• logicallee 11 hours ago

I think you've fixed it! I just tested it, now it works. (In both Chrome on Windows with NVidia 1060 gpu w/ 6 GB vram, which is what it wasn't working on yesterday, and on Safari on 2026 Mac Mini M4 with 24 GB of RAM. The former is pretty slow, around 1 tok/sec, the latter is fast at 18 tok/sec.)

• moooff 10 hours ago

Yeah, I found a bug during model load occurring with Firefox and low VRAM machines. It is fixed now, thanks!

The 1060 is unusually slow. The Mac mini sounds about right. Have fun playing around and sharing. With the more VRAM of the mini you can also run Qwen 3.8 27b :)

I really like to showcase this to non-tech people to show what their home machines already can do WITHOUT INSTALLATION in an air-gapped scenario! And given all the recent fuss, we should not trust the big providers at all with important data.

• logicallee 6 hours ago

Great, thank you for making it and also the fixes!

• brinepot 19 hours ago

Pretty neat how far these small models have come. Running them locally in the browser feels like magic for quick experiments.

• Schlagbohrer 13 hours ago

Every model I just tried in here gave me very nonsensical hallucinations, or immediately fell into a loop. Not the best examples of what Petite GPTs (as one is called) can do, I hope.

• logicallee 11 hours ago

Could I ask what browser, operating system, and GPU you use? In my testing they do not fall into a loop, maybe there is an error that affets how it's being run on your device.

• hbroom 19 hours ago

Ran a few of these, and the speed for in-browser execution is wild. Perfect for quick demos without a backend.

• Mbarley 20 hours ago

Impressive how much you can do locally now. Was expecting much slower inference but it's surprisingly snappy.

• Surac 8 hours ago

Seems the models only support English. Questions in German are just echoed

• philipallstar 14 hours ago

> What is the capital of south africa

> Victoria is the capital of South Africa.

> What is the capital of south africa. It's not Victoria.

> Victoria is the capital of South Africa.

They're coming for your job!

• willaaam a day ago

I like it as I'm vibecoding an (airgappable) browser AI workspace myself, but in terms of putting the models to use, just exposing the chat interface feels a bit limiting to me.

My take on this concept: https://github.com/willaaam/gemma-4-E2B-webgpu-vision

• logicallee a day ago

It's a cool project, but I couldn't get it to load. What did you test this on? I tried the live link here:

https://willaaam.github.io/gemma-4-E2B-webgpu-vision/

And after loading it, with an NVidia 1060 GPU (6 GB RAM) on Windows it failed with "Failed to load: No supported WebGPU variant for com.xenova.gemma4.DenseGemv; rejected sgma".

In Safari on a 2026 Mac Mini M4 with 24 GB of RAM it failed with "Failed to load: JSON Parse error: Unexpected EOF".

The idea is pretty cool though!

• willaaam a day ago

Huh, weird! It doesn't work in Firefox, but I test on Safari (M5 Max) and Chrome (Linux, Arc B580). I don't have access to NVidia or AMD hardware myself unfortunately, though friends confirmed both as working.

Going to debug tomorrow!

• Rexxar 21 hours ago

Very interesting but on a naming perspective what is a "micro large" model ?

• bhouston a day ago

I built something like this just last month, using a few of the same models, but I used ThreeJS's Three-Shading-Language abstraction to do it: https://three-llm.ben3d.ca/?model=qwen3.5-0.8b

• sourweasel a day ago

I would recommend adding the MiniCPM5-1B model. Surprisingly coherent for a 1B model and performs well. On my pixel 9 I get 33 tok/s on CPU. On GPU I get about 26 tok/s but prefill jumps to nearly 500.

• logicallee a day ago

Awesome! I tried it and after loading (which took a while as it is a large model) got 12 tokens/second and very coherent output. Great demonstration.

• micw a day ago

It's LLMs, not LLM's ;-)

• logicallee a day ago

Thanks. Unfortunately it's too late to change the title!

• d3Xt3r a day ago

Maybe we should just call them SLMs instead of LLMs...

• gslepak 20 hours ago

Interesting. Crashed macOS. That's already impressive.

• logicallee 18 hours ago

Sorry. That shouldn't happen. It doesn't use a lot of memory and only uses WebGPU in the normal way. What version of macOS and Safari did you use, and what is your hardware, please?

• jellyfiz a day ago

Similarly, if someone wants to hack some LLMs in their browser and burn some cycles, please feel free to try to break them here:

https://ai-attacks.neal.codes/

• inventor7777 a day ago

PetitGPT told me that

> "2+2 is 2."

Otherwise, a very neat demo. As others have said, the UI is VERY confusing, way too much stuff going on.

• allenu 18 hours ago

Using the same model, I asked it what 2 + 2 is and it said "2 + 2 is 2 + 2" which is not wrong.

Other questions and answers were mixed:

How many cards in a deck? > A deck of cards contains 52 cards.

How many cards in a deck if I remove all Queens from the deck? > If you remove all Queens from the deck, there are still 52 cards in the deck.

• dvh a day ago

It is. For very small values of 2.

• rvz a day ago

Nothing on this site actually works.

• logicallee a day ago

What browser are you using? I tested it on Windows, Mac, and iPhone. I tested it in Chrome, Firefox, Edge, and Safari. Everything works on the three machines and phone I tested it on.

(It's a little bit slow at the moment - you have to wait a few seconds for the models to load - as it's currently on the HN front page. The server is on a 1 gigabit unmetered network connection so it can serve all the weights - around 600 megabytes - to one person every few seconds, there are several concurrent users now.)

• fragmede a day ago

serve the models via CDN?

• anonimous-emacs a day ago

unfortunately on firefox: Uncaught ReferenceError: GPUShaderStage is not defined

• utopiah 17 hours ago

Firefox on Linux doesn't support WebGPU https://caniuse.com/webgpu should enable via a flag.

Adding a warning about it could help.

• logicallee 6 hours ago

I've tried to fix this, please check again. Thank you.

• logicallee a day ago

I tested it on Firefox on windows, version 156.0.1 and didn't get that error.

What version of Firefox are you using and what is your operating system and graphics card, please? Can you also try it without WebGPU? (Reload the page and uncheck "Prefer WebGPU" and try a prompt.)

• ad_fontes a day ago

Same issue. Latest version of Firefox 156.0 on Ubuntu GNOME. AMD Radeon 860M.

Unticking the WebGPU checkbox and reloading didn't do anything.

• logicallee 6 hours ago

I've tried to fix this, please check again, thank you.

• logicallee a day ago

This is a difficult one for me to fix since I don't have a Radeon GPU. I'll see if there is anything I can do tomorrow.

• cwnyth a day ago

I got the same error with an Intel Iris. I'm using Firefox on Fedora Gnome. "Prefer WebGPU" also didn't fix it.

• logicallee 6 hours ago

Thank you. I've tried to fix this, could you please check again?

• derliebej a day ago

Wouldn't that be a TLM?

• agnishom a day ago

Yeah, or a μLM

• voxmastery 15 hours ago

interesting

• logicallee a day ago

You can try 7 different tiny LLM's in your browser.

• bigfishrunning a day ago

What I missed from the title was this: Can I try 7 different tiny LLMs in my browser?

• tnrich a day ago

i have a bromine tub, give me simple instructions for what to do since i just filled it up w fresh water

1. Add 1 tablespoon of water to the water bath. 2. Place the tub into the water bath and let it sit for about 5 minutes. 3. After 5 minutes, remove the tub and let it cool down. 4. Now, fill the tub with water and let it sit for about

uhmm completely unusable ?

• smokel a day ago

No, that is not what SLMs are useful for. They have very little knowledge, and typically lack reasoning skills.

Useful applications include sentiment analysis, text classification, entity extraction, etc.

They certainly can be useful, but you shouldn't compare them with LLMs such as Opus or Fable.

• ironqcold a day ago

Completely unusable for that, yeah, that's a 100M model, not an assistant

• AgentMasterRace a day ago

The website literally says they are SLM's .. derp af