

You could likely do some sort of distrubuted system in machines that had a reasonable graphics card, or shared memory like the M series macs. Have it silently intall a 30b parameter model or higher and somehow hide how that space is being used.
Local models are getting pretty good, and one focused on hacking / bot net behavior could probably do very well at that size.
But you’d notice it being used… fans would go off, frame rates drop etc. You’d probably need to code it to only do stuff when it seems idle?



Im not super well versed in this, but i dont think you can efficiently distrubute the LLM processing like that. Even if it was possible, it would be very very slow. But I guess speed isnt always needed.