Technically, TTS/STT are mostly MLs; I’m pretty sure many people run these. I have a setup but I’m better with buttons that with spoken words, and I listen to ambient sounds or music. I think some day I’ll make voice assistant for talking to while driving, but that’s not a trivial task hardware-wise, even if I used cloud LLM layer, which I won’t. Putting AI on baremetal sounds like an interesting project.
I have a homemade “local agent” that can actually “code” somewhat, I use it just to figure out how this thing works on the inside practically. Mostly useless otherwise (also I have GPU that’s older than AI, so it’s kind of fun technical task to run this stuff on pure RAM+swap). Feels like the whole hype is greatly overrated, but I appreciate a chance to learn something new anyway.
Technically, TTS/STT are mostly MLs; I’m pretty sure many people run these. I have a setup but I’m better with buttons that with spoken words, and I listen to ambient sounds or music. I think some day I’ll make voice assistant for talking to while driving, but that’s not a trivial task hardware-wise, even if I used cloud LLM layer, which I won’t. Putting AI on baremetal sounds like an interesting project.
I have a homemade “local agent” that can actually “code” somewhat, I use it just to figure out how this thing works on the inside practically. Mostly useless otherwise (also I have GPU that’s older than AI, so it’s kind of fun technical task to run this stuff on pure RAM+swap). Feels like the whole hype is greatly overrated, but I appreciate a chance to learn something new anyway.