Myth Busting the GB10

What This Box is Not: Three Myths About the GB10

⚠️

Disclosure: Dell gave me a Dell Pro Max with GB10. I did not pay for it. I am not paid to write about it. Nobody at Dell sees these posts before you do.

⚠️

Last week I published six hypothesis’ about this Dell Pro Max (with the GB10 Nvidia super chip) and before we proceed I wanted to make sure we address a few items.

These items cover a wide angle and how I’ve seen these devices being used in a variety of manners whether via social media or long form content. Whether we could define all these items as wrong or assumption I will leave to you, the reader.


Myth 1: It’s a gaming PC with a fancy badge.

Will it game? … Yes, in short and it certainly packs the punch to do so! You can install Steam on this device and most likely a host of emulation or supporting packages to be able to play the games however, lets break it down a little more.

It has an Nvidia chip in it and it sits on a desk, so already we are going in the right direction to be considered for a gaming PC. It is important to call out that this is where I will be stopping.

It has a completely different setup to a ’traditonal’ gaming PC, the main difference is unified memory, and it has an impact how on the system can operate.

Using our ’traditional’ gaming PC you will see that a GPU will have its own dedicated VRAM. Whatever it has that is what you get, so if it states 24GB DDR7 that is what is available to the system for GPU/graphics. The way the system then handles those 24 GB’s of resource is that; the GPU required workload is offloaded to the GPU for processing. This is great for gaming as the CPU & system RAM can focus on other operations needed to run the system/game and or other applications. Applying to our use case, we are using it for AI (inference, agents etc.) and if you get a model that demands more than you provisioned amount (physically limited) it may not even load. Up until recently that would have been a blocker even with quantization however, there are known open source projects that control the loading of models into memory. That is a topic itself and we may explore that in a later post or several.

The GB10 shares one much larger pool between CPU and GPU. Thats it, where before you GPU resources where treated seperately here they become a shared pool. That seems like a great bonus especially when the device ships with 128GB.

This is great! More RAM for everyone!

MORERAM

Compared to our ’traditional’ gaming PC we can already load bigger models but at the cost we are sharing these resources across the sysem. So we are now able to load models that are bigger but even still there is a ceiling to those also. We will dive into this scenario in upcoming posts as I would say it is something to strongly consider when looking at local AI use cases driven by business needs.

From this brief example I see different failures. One, straight up tells you no. The other says yes, slowly. If your workload is “run this large model over these documents overnight”, those are not remotely the same answer.

Why the wrong assumption hurts you: it makes the device look overpriced against consumer hardware, because you end up comparing on the wrong baseline. If your criteria are FPS (frames per second) and price per teraflop, you might need to turn around at this point and reconsider.


Myth 2: It’s a Copilot+ device with ambition

Microsoft has spent a year talking about NPUs and on-device AI, so this one is fair enough. Both get described as local AI. Both run models without a network connection.

They are not competing products. They are different tiers.

An NPU laptop is genuinely good at small models while you are doing something else. That is a real capability and I know several people use these most days. It is not the same job as a large model running continuously over a document set that is not permitted to leave the building.

I am not going to argue this from assertion. In upcoming posts (hopefully) I am running the same four tasks across an NPU laptop, this box, and publishing the numbers with the method. That is a better argument than anything I could claim now, and if I am wrong you will be able to see exactly where.

Why the wrong assumption hurts you: it makes the device look redundant against laptops the organisation is already buying, which is the fastest way to have a proposal declined.


Myth 3: It replaces your Cloud spend

Out of all of the assumptions I am and will discuss, this on carries a lot of weight in most customer conversations and as such it is also the most dangerous. If this is included in a business case to purchase one it will be music to a CFO’s ears. So in the following paragraphs, I am going to be as specific as possible to ensure complete clarity.

It does not replace your Azure spend. It can handle a band of work that you would typically offload to the cloud.

Frontier models stay ahead of anything that you can run locally, this was on of the hyptohesis’ from last week and I expect this to be correct, for the foreseeable. Elasticity/resource on demand keeps the cloud ahead too. If you need ten of these for an afternoon, you rent them as it is likely that you cannot rent the 10 boxes quicker than you can spin up an instance or instances of workloads in the cloud. If I need 10 of these for an afternoon is becoming commonplace then I would also say that it would be time to explore your options between dedicated hardware vs. cloud vs. batch.

Why the wrong assumption hurts you: this assumption is the one that comes back to haunt you later down the line. Cloud spending hasnt reduced and the questions begin on why and when are they expected. It’s damage is multi fauceted, credability or the project owner/technical lead, AI is expensive and for large enterprises, this device was suppose to fix all our problems.

I would stress, do not oversell this as the use cases for local AI are strong enough on their own. Over this blog series, I am going to convey what it can be used for so they can be adopted to help build use cases/business justification. Not only that but the evidence to help inform and guide you to make the right decision.


So what is it?

Simply, it is a desktop device for running large language models (LLM), continuously. This could be to iterate over data that isnt permitted to leave your premises or for rapid development and refinement before pushing to the cloud.

That is a narrow claim. I know it is narrow because the three myths above all make it sound broader than it is.

For a specific set of organisations (regulated sectors, sovereignty, defence) it’s a claim that matters as it enables those to access AI knowing they are fully in control. I also think there are edge use cases that open that claim to a wider market and thats what I am here to explore.

If you came in holding one of those myths/assumptions, the useful question is not what this thing is. It is whether, are you looking at local AI in the right way? If you believe you are, are you now looking at the right hardware to support that answer.


Disclosure, again: hardware was gifted by Dell.

Posts in this series

Related Posts

comments