Honestly i think speed is something I don't care too much about with models, because even things like ChatGPT will be slower than Google for most things, and if something is more complex and a good use case for an LLM it's unlikely to be the primary bottleneck.
My gf private chat bot right now is a combination of Mistral 7B with a custom finetune and she it directs some queries to ChatGPT if I ask (I got free tokens way back might as well burn through them).
How much of an improvement is Mixtral over Mistral in practice?