Forum Discussion

ChatJBT's avatar
ChatJBT
Everpure
7 hours ago

How much more inference could you get from the GPUs you already have?

As inference workloads grow, it’s not always the GPU compute that becomes the bottleneck. Recomputing tokens of a known context vs reloading from KV cache can take up a lot of GPU bandwidth and limit how many workloads you can run in a given amount of  time.

That’s one of the things we’ve been working on with PureKVA, reusing KV cache so you can get more out of the GPU capacity you already have. With v1.2, we’ve also added a 3-tier memory architecture, TurboQuant KV cache support, along with things like multi-tenancy and storage isolation.

We’re seeing some of the benefits internally as well, particularly around reducing throttling when usage spikes.

Curious what others are seeing. Is GPU compute the bigger constraint for you, or GPU memory?

No RepliesBe the first to reply