Ternary is cool, but saying it is (# of parameter)-class is somewhat misleading
Why not if it “retains 95% of the full-precision baseline” intelligence?
I imagine the true advantage of this would be to have something like a ternified 100b model that then beats a 27b full precision model.
Eh, I was mostly thinking out loud, but I would say there is a real difference in the quality of results depending on quantization
Well it’s a good question, I’m genuinely curious how noticeable the differences are. I have an older PC with a notoriously full SSD hard drive so haven’t tested much.
I’m hoping we’ll see more of a “democratization” of AI with future hardware being developed that allows one to run these full models at home with solar power. Without big fascist AI corporations or data centers or cloud BS or privacy concerns. A smart home / digital assistant would be awesome if it’s localhosted. Or improving search engine results and filtering out garbage. Because why would companies earning money by selling mental garbage / advertising ever filter out garbage?
This could also be awesome for video games, allowing improvements to procedural generation of open world games, have more improvisation in dialogue trees where creators basically sketch out characters and write some dialogue and the LLM can expand in the same style. Something like a game master. Can hallucinate all it wants then lol. Anyway I get a little excited about breakthroughs like this.
Open models are available and pretty effective for many tasks
A smart home / digital assistant would be awesome if it’s localhosted.
Home assistant + ollama + a GPU with 8gb of memory will get you all that and more. And with zigbee and ble, smart devices are (relatively) accessable, private, and energy efficient. I’m seriously living in the future over here.
Damn this is cool
Interesting, but not open weight and only seems to work on Apple devices (and NVIDIA GPUs but do any phones have those?)
ggufs are available on huggingface, so it is open weight since ggufs can be converted back to safetensors, and the interesting part isn’t that it can run on Apple phones, but that it’s small enough to do so which means it’s very lightweight and will use less energy and run on smaller hardware setups https://huggingface.co/collections/prism-ml/bonsai-27b
Looks interesting. I was just playing around with 1.58-bit LLMs today and I’m kinda surprised they aren’t more popular





