I get the fear around the fishing industry, but regardless of Iceland’s dependance on them we all need to stop eating fish. We’re completely wrecking the ocean ecosystems - it’s not sustainable.
I’d love to have Iceland fully join us in the EU.
HW/FW security researcher & Demoscene elder.
I started having arguments online back on Fidonet and Usenet. I’m too tired to care now.
I get the fear around the fishing industry, but regardless of Iceland’s dependance on them we all need to stop eating fish. We’re completely wrecking the ocean ecosystems - it’s not sustainable.
I’d love to have Iceland fully join us in the EU.


No I think this is completely novel actually, and afaik it’s CUDA only because that’s what the author have themselves.
I’ve been using it since I compiled it. I’d go so far as to claim this is a game changer for 12 and 16GB cards.


I assume you meant your and not the computer’s, and I agree - would also like to sound like him sometimes! :D
For those interested in the voice cloning part of the project though, it’s done “live” at server startup through two config options:
--qwen3_tts_ref_audio ./computer.wav
--qwen3_tts_ref_text "Darmak is the name of a seventh dynasty emperor on condon four A myth of a historical hunter on Chantil three A colony on Melindi seven"
Selecting the voice of the computer thus only needs a clear .wav of them speaking and a transcription.


This solution is fully open source - and the components used for SST/LLM/TTS are freely exchangeable. Latency goes down if one has the ability to use a separate SST like Parakeet instead of reusing the LLM for it as I do in this video.
My scope is “what me and family wants” - Home-Assistant is one of the tools the LLM lists (pool pump sensor comes from there). The server takes an mcp.json so any tool that has an MCP connector can be used.
Haven’t thought a graphical avatar, but I am a Red Dwarf fan and I’ve already thought about the ability to have pre-configured “themes” for the system (Star Trek, HAL, Holly etc) - although not something I can host and publish due to … waves hands … IP rights.


Thanks, I need to update the README :D Implemented full training support in the project yesterday so I could use “computer” instead of “hi ESP” :D


tbh if Dee Snider says it I’d assume he’s right
All quants updated with correct template now
Got an answer - there were indeed issue with some quants
Actually edit - UD-Q4_K_XL and all UD-* has our corrected chat teaplate
We need to update the non UD-* ones with our chat template - stay tuned
Having issues with it (Unsloth’s gguf).
0.27.136.404 I srv proxy_reques: proxying request to model unsloth/Qwen3.8-27B-GGUF:Q4_K_M on port 55741
[55741] 0.14.424.772 W srv operator(): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵ {{- raise_exception('System message must be at the beginnin...\n ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}


This will let me have my Jolla phone stick to my magsafe charger? If so then awesome! No need for actual charging, but due to the type of desk I’m using that’s actually one of the issues I’m having right now going from my previous iPhone to the Jolla :D


Qwen3-TTS does excellent cloning and is very fast. On my workstation (5060Ti) it renders the audio at around 2.5x realtime. I’ve added an example with my parameters now.


I’ve added example scripts for client, server and MCP configs. It should be much easier to replicate with those now.


Oh thanks all for the nice comments! My fork is at https://github.com/troed/speech-to-speech but beware it’s not kept in any other state than what I’ve last pushed :) Also you’ll have to supply your own chime wavs and voice sample etc.


When I start llama-server I point it to the models-config that have unique max context sizes per model - and they’re allocated at their max size as soon as the server starts so since it comes up I will be able to use that context size too.
I’m actually a bit unsure as to how you run it since you get OOMs during usage :)
I also use the DCP plugin for Opencode to help manage the context cache and have less of a disruption as it gets compressed, but I wouldn’t need to for the above to work. When I hit the context limit the context would still get compressed.


Less-than-beautiful UI added to the clients (still works as CLI if run with parameters). Somewhat problematic finding an easily cross-compileable Go UI lib.
Also made the typical pipe-to-bash installation flow, awaiting proper ujust in some far away future.
Regular 1Gbit/s with three switches in total between them - distance about 40m (different buildings).
Jag har hittills haft åsikten att vad jag än kan välja för bredband så väljer jag (och rekommenderar andra) Bahnhof eftersom oavsett om de är billigast eller inte så vet jag att mina pengar går till något bra för oss alla.
Med Telenor som ägare kommer jag nu istället gå strikt på features vs kostnad. Antingen levererar de bäst eller inte.
I think this hadn’t been made a big deal out of if it hadn’t been for the conspiracy theories on Proton just before
Yeah it was quite obvious from all the “clanker” that this is a typical run-of-the-mill AI-hater, but I found this funny anyway:
The weight count is too low to reproduce the training set verbatim
That’s not what anyone wants. That’s not how LLMs gain “intelligence”, at all. The whole point of training on large datasets is to NOT internalize training data verbatim.
I think Putin should visit the front and take a few selfies of those winnings.