Private LLM in SingaporeAnyone can drop a model on your server.

We optimize it to run frontier-grade AI fast and cheap on the hardware you own.

Private, on-prem LLMs for Singapore teams. We size the box, tune the model and the serving stack for it, and put access control and audit around it, all inside your walls. Every speed on this page is measured on our own machine, and every price is read from the seller's own page.

Benchmarks measured · prices read

125B
parameter model with vision, on one box
~1,050
tokens a second reading prompts
128 GB
of unified memory, desk-sized
0
tokens leave your network
strix-halo · local model server
altronis@strix-halo:~$

A real session on our own Strix Halo box, recorded 23 Sep 2026 · Qwen3.8-Flash-Next, fully offline.

In one paragraph

Private (on-prem) LLM deployment means running an open-weight model such as Qwen, Llama, DeepSeek or Gemma entirely inside your own network, so sensitive data never leaves your perimeter and there is no per-token API bill. One 128 GB box can serve a frontier-grade model for a team: ours runs Qwen3.8-Flash-Next, a 125B mixture-of-experts model with vision, at about 37 tokens a second. The hardware is the easy part. Speed and cost come from optimization, choosing the quantization, serving engine and settings for your box, which on ours made prompt reading three to four times faster for the same model. In Singapore this is the practical path to AI on PDPA-, MAS- and MOH-regulated data. Altronis sizes the hardware, optimizes the model, deploys the stack and builds the governance, all under your control.

Not a mockup

Try the private LLM on this page, live.

The assistant in the corner of this page is not a scripted demo. Your question is answered by an open-weight LLM running on our own Strix Halo box in Singapore, over a private tunnel. It is the same class of model, on the same class of hardware, that we deploy for clients. Ask it about on-prem deployment, hardware sizing, or PDPA.

If our box is offline for maintenance it falls back to a hosted model so you still get an answer. In production, on your own hardware, there is no fallback and nothing leaves your network.

Why on-prem

Three reasons teams move AI in-house

On-prem is not always the answer. When one of these is true, it is the only one, and we will tell you which side of the line you are on.

Data sovereignty

Regulated or sensitive data (PDPA, MAS, MOH, client IP) that legally or contractually cannot leave your environment. On-prem removes cross-border transfer and third-party exposure at the root.

Cost at high volume

Once you sustain millions of tokens a day, per-token API pricing can outrun the cost of owning the hardware. A one-time purchase replaces an open-ended metered bill, and we work out your break-even with you.

Latency and control

No network round-trip, no rate limits, no silent model swaps under you. Fixed latency and a model you pin, version and audit, which agents and real-time work depend on.

Measured on our own box

Same model, same box, optimized

Frontier-grade AI on your own hardware, tuned for speed and cost. Optimization is our edge, not an add-on. Below is one 125B model with vision, Qwen3.8-Flash-Next, on our own 128 GB Strix Halo, before and after we changed how it is served. Same model file, same box, same prompts.

Before: a llama.cpp-based build. Now: a serving engine built for Strix Halo.

before: 18 Sep 2026 · now: 9 Oct 2026

Reading an 8K-token prompt (tokens a second)3.0× faster
Before
351
Now
1048

One of the four runs read at 555 tokens a second; the other three at 983 to 1,340.

Reading a 32K-token prompt (tokens a second)4.3× faster
Before
303
Now
1314
Writing prose (tokens a second)1.5× faster
Before
24.1
Now
37
Writing, 8K tokens into a conversation (tokens a second)1.9× faster
Before
18.1
Now
35.1
Writing, 32K tokens into a conversation (tokens a second)1.7× faster
Before
16.9
Now
28.9

One of the four runs wrote at 11 tokens a second; the other three at 33 to 37.

Writing code (tokens a second)16% faster
Before
39.7
Now
45.9

On 18 Sep this was the one result that went backwards (35.8). The newer engine version fixed it.

Four requests at once, until all are answered (seconds, lower is better)3.0× faster
Before
35.6 s
Now
12 s
Four requests at once, longest wait for the first word (seconds, lower is better)4.3× shorter
Before
31.3 s
Now
7.2 s

Before was measured 18 Sep 2026; now was measured 9 Oct 2026 on a newer version of the same engine, as the average of two full runs with the slow runs left in. Both on our AMD Ryzen AI Max+ 395 (Radeon 8060S, 128 GB unified LPDDR5X), with the same 4-bit model file, thinking off, temperature 0, and the first run of each prompt, so no cache helped. The engine is made by another company; our part is choosing it, configuring it for this box and measuring it. Its maker quotes higher prompt speeds than ours because we run a smaller prompt buffer than their test did. We moved our production box to it on 18 Sep, the night of the first test, and keep it on new versions as they ship. Configs for the box: github.com/sypherin/strix-halo-setup.

Earlier models we ran on this box

Smaller models write faster. We moved to Flash-Next because it is far more capable (see the scores below), and 37 tokens a second is still well ahead of reading speed. These are history, measured on older builds, not what the box runs today.

Swipe the table sideways to see every column.

ModelQuantWriting t/sReading t/sMeasured
Qwen3.8-27B, vision + multi-token predictionUD-Q4_K_XL20 to 22~750Aug 2026
Qwen3.6-35B-A3B, multi-token predictionQ4_K_XL~65 prose, ~100 code~840Jul 2026
Muse Glimmer-30B, DFlash draftROCm FP422 to 35~340Jul to Aug 2026
Qwen3.6-35B-A3BQ8_K_XL~44~839before mid-Jul 2026
Gemma 4 26B-A4BQ8_K_XL~41~720before mid-Jul 2026
Qwen3.5-122B-A10BQ4_K_XL~22~393before mid-Jul 2026

Quality, not just speed

How Flash-Next scores on public benchmarks next to a smaller Qwen model, another open model, and Claude Opus 4.6 as the frontier reference. Every number here comes from one source, Qwen's own model card, so treat it as the maker's claim.

Swipe the table sideways to see every column.

BenchmarkQwen3.8-Flash-Next (we run this)Qwen3.8-27BDeepSeek-V4-FlashClaude Opus 4.6 (Max)
SWE-bench Pro (agentic coding)62.561.756.053.4
SWE-bench Multilingual (coding)81.073.8not reported77.5
CoWorkBench (office work, Qwen's own test)73.970.745.168.2
IFBench (following instructions)81.379.579.262.5
GPQA Diamond (graduate science)91.789.290.891.3
HLE (very hard exam questions)35.930.833.840.0
LiveCodeBench v6 (coding)91.990.390.688.8

Source: Qwen3.8-Flash-Next model card, huggingface.co/Qwen/Qwen3.8-Flash-Next. Qwen ran the other models with its own harness, except Opus 4.6 on SWE-bench Pro, which Anthropic published. Higher is better. Opus still leads on HLE. These are full-precision scores; we serve a 4-bit build and have not measured how much quality that costs.

Pick the box

Which hardware for a private LLM?

For one team it comes down to three desk-sized boxes, and server GPUs once many people use it at once. Writing speed on these boxes is set mostly by memory bandwidth; prompt reading depends more on compute. Pick on budget, software and noise, and we optimize the model for whichever box you choose.

Swipe the table sideways to see every column.

BoxPrice in SingaporeMemoryBandwidthSoftware
AMD Strix Halofrom S$4,325128 GB~256 GB/sVulkan, ROCm, llama.cpp
NVIDIA GB10from S$6,624128 GB273 GB/sCUDA, vLLM, TensorRT-LLM
Apple Mac Studio M5 Ultrafrom S$7,99996 GB, up to 512 GB~1.2 TB/sMLX, llama.cpp on Metal

Memory and bandwidth are the makers' specs. We have measured the Strix Halo ourselves (above); we have not run the other two. Read 9 Oct 2026 from each seller's own page. Overseas stores converted to SGD with 9% GST at import added, shipping not included. List prices on the day, not a quote.

AMD Strix Halo (Ryzen AI Max+ 395)What we run

AMD Strix Halo (Ryzen AI Max+ 395)

16 Zen 5 cores, a 40-CU Radeon 8060S GPU and 128 GB of unified memory at about 256 GB/s. The most memory per dollar on a desk. Ours serves a 125B model with vision and a 131K-token context for our team. Most sellers are overseas brands, so check the warranty and budget for import GST.

NVIDIA GB10 (DGX Spark, ASUS Ascent GX10)Best tooling

NVIDIA GB10 (DGX Spark, ASUS Ascent GX10)

NVIDIA's Grace Blackwell desk box: 128 GB of unified memory at 273 GB/s and the full CUDA stack. NVIDIA quotes up to 1 petaFLOP of FP4 compute and models of up to 200B parameters. The ASUS Ascent GX10 is the same GB10 chip with a 1 TB drive instead of 4 TB, for less. The pick for NVIDIA's ecosystem, or for training and inference on one box.

Apple Mac Studio (M5 Ultra)Quiet / Mac

Apple Mac Studio (M5 Ultra)

Apple's M5 Ultra with 96 GB of unified memory as standard, configurable to 256 GB or 512 GB, at about 1.2 TB/s: the most bandwidth of the three. Runs models through MLX or llama.cpp on Metal. Silent, and the natural pick for Mac-first teams and quiet offices. The larger memory options cost more than the base prices shown.

Server GPUs (RTX PRO 6000, H200, B200)Many users

Server GPUs (RTX PRO 6000, H200, B200)

When dozens of people or apps use the model at once, you move to server GPUs: the 96 GB RTX PRO 6000 Blackwell card, or data-centre parts such as the H200 (141 GB) and B200 (192 GB) with several TB/s of bandwidth each. Far more throughput, at server prices and power. The price shown is the RTX PRO 6000 card alone.

Under the hood

What we optimize, and why it matters

Running a big model on a small box is a stack problem, not a hardware problem. These are the levers that decide whether it flies or crawls.

The right quantization

Quantization stores each weight in fewer bits. A 4-bit build lets a 125B model fit in 128 GB with room left for a long context. Too few bits and quality drops; too many and it will not fit, or runs slowly. We pick the build per model and per box, and check the answers.

Mixture-of-experts models

Mixture-of-experts models use only a few billion parameters per token (6B of 125B for the model we run), so they write at small-model speed with large-model quality. On a box with lots of memory and limited bandwidth, that is the sweet spot.

The serving engine

The software that runs the model matters as much as the box. On ours, moving to an engine built for Strix Halo made prompt reading 3.0 to 4.3 times faster and cut the wait when four requests arrive together from 36 seconds to 12, with the same model file. We benchmark engines on your hardware before we pick one.

Long context and the cache

Qwen3.8-Flash-Next handles 262K tokens natively; we serve 131K on our box so memory stays free for the cache. Multi-token prediction drafts several tokens at once and the model checks them, which speeds up writing when the text is predictable.

How we deploy

One private endpoint, many models

We do not hand you a research script. We deploy a governed, single-endpoint stack your apps point at: models behind an authenticated gateway, everything inside your perimeter.

The stack we stand up

  • A path-routing gateway: one authenticated URL for chat, vision and OCR, and embeddings; model ports are never exposed directly.
  • Open models optimized for your hardware and task: quantized, benchmarked, pinned and versioned.
  • Governance: access control, full prompt and response audit logging, human-in-the-loop checkpoints.
  • Optional private tunnel for remote access; runs on your kit or in your VPC.

What you get on handover

  • A documented, self-serve runbook your team can operate and update, not just us.
  • A working proof of concept in two to three weeks; production use in eight to ten weeks.
  • Model and hardware advice so you do not over-buy or over-size.
  • An optional managed-ops retainer if you would rather we keep it running.
Altronis local AI tools

Open tools for running local LLMs well.

Setup agent

A private LLM that helps install, secure and fix itself.

We ship the stack with Deneb, our setup agent that lives on your box. It reads the logs, diagnoses problems and applies the fix, reversibly and only with your say-so, with a human engineer behind it. Ask it in plain English, or paste an error or a screenshot.

Installs, diagnoses and repairs the stack on the box
Open-source client: audit every line it runs
Red-teamed, read-only-by-default security boundary
Never runs anything destructive or privileged
Benchmark tracker

Know what actually runs on your box, before you buy it.

TokenMark tracks real community benchmarks for Strix Halo, DGX Spark and more: every model, quant and backend at its measured tokens a second. Ask its advisor what to run for your hardware and task, or query it from your terminal and agents.

Measured tok/s from real community benchmarks
Advisor: “what should I run on X for Y?”
CLI and MCP: query it from your terminal or agents
Every number linked to its source, never invented

Regulated-data workloads

Keep regulated data where it has to stay

On-prem is the architecture regulated Singapore teams use to keep data they cannot send to a hosted API inside their own walls. To be clear, we are not a compliance certifier: on-prem, plus the governance and audit we build, supports your obligations, but the regulatory relationship stays yours.

Data residency (PDPA)

Because data never leaves your perimeter, cross-border transfer and third-party processing largely fall away. That is the hardest part of your PDPA obligations, handled by the architecture rather than a policy promise.

Finance (MAS)

For financial institutions, on-prem inference keeps prompts, outputs and customer data inside the environment you already control and audit, which is the posture you need before putting AI near client data. Your MAS relationship stays yours.

Healthcare (MOH / HSA)

Patient data stays in-tenant; no identifiable data enters a hosted model. Where AI touches triage or diagnosis we flag that HSA may treat it as a medical device (AI-SaMD), so you assess it early, not after a build.

Straight answers

Private LLM deployment: FAQ

Can you run a private LLM fully on-premise in Singapore?

Yes. Open-weight models such as Qwen, Llama, DeepSeek, Gemma, GPT-OSS and Mistral run entirely inside your own network. Our own box, a 128 GB AMD Strix Halo mini PC, serves a 125B mixture-of-experts model with vision and a 131K-token context for our team. No data leaves the building and there is no per-token API bill.

Is an on-prem LLM PDPA compliant?

On-prem removes the hardest PDPA risks, cross-border transfer and third-party processing, because data never leaves your environment. It is not automatically compliant on its own: you still need access control, audit trails and governance, which we build in. No serious vendor should claim 'fully PDPA-compliant' out of the box.

How fast is a local LLM, really?

On our own AMD Strix Halo box, Qwen3.8-Flash-Next (125B parameters, 6B active per token, with vision) reads prompts at about 1,050 tokens a second and writes at about 37 tokens a second on prose and 46 on code, measured on 9 Oct 2026. 32K tokens into a conversation it averaged about 29 tokens a second (three of four runs at 33 to 37, one at 11). Four requests sent at once all finished within about 12 seconds. That is interactive for chat, document extraction and agents.

What do you actually optimize?

The model and the way it is served, for your box. We pick the model and the quantization (how many bits each weight keeps) that fit your memory with room for long context, choose the serving engine, and tune multi-token prediction and the context cache. On our own box, changing the serving engine alone made prompt reading 3.0 to 4.3 times faster for the same model file. We measure every change on your hardware before we keep it.

What hardware do I need to run a private LLM?

For one team, a desk-sized box with 128 GB of unified memory: an AMD Strix Halo mini PC (Ryzen AI Max+ 395) from S$4,325 for value, an NVIDIA GB10 box such as the DGX Spark or ASUS Ascent GX10 from S$6,624 for the CUDA tooling, or an Apple Mac Studio with M5 Ultra from S$7,999 (96 GB as standard, up to 512 GB) for memory bandwidth and silence. Prices as read from each seller's own page on 9 Oct 2026. For many people at once you move to server GPUs.

AMD, NVIDIA or Apple for an on-prem LLM?

Writing speed on these boxes is set mostly by memory bandwidth. On paper the AMD Strix Halo (about 256 GB/s) and the NVIDIA GB10 (273 GB/s) are within 7% of each other, while the Mac Studio M5 Ultra has about 1.2 TB/s, over four times either. The GB10's strengths are compute for long prompts and the mature CUDA stack (vLLM, TensorRT-LLM). Strix Halo is the cheapest 128 GB box. Pick on budget, software and noise; we then optimize the model for the box you chose.

When does on-prem beat cloud AI?

When data cannot leave your perimeter (regulated data), when you sustain enough volume that per-token API costs outrun owning the hardware, or when you need predictable low latency. Below that, cloud APIs are usually cheaper and faster to start, and we will tell you which side of the line you are on.

Can I run a 100B-class model on a single box?

Yes. We run a 125B mixture-of-experts model on one 128 GB box, as a 4-bit build. Mixture-of-experts models are the sweet spot for these boxes: large-model quality, but only a few billion parameters are active per token (6B for ours), so they stay fast on hardware that is limited by memory bandwidth.

How long does a private LLM deployment take?

A working proof of concept in two to three weeks and production use in eight to ten weeks is typical, longer if new hardware has to be bought first: hardware sizing, model selection and optimization, the inference gateway, governance and audit, then integration with your systems. We hand over a documented stack your team can run.

Which open models can we self-host?

Qwen (including Qwen3.8-Flash-Next, which we run ourselves), Llama, DeepSeek, Google Gemma, GPT-OSS and Mistral families, plus OCR and vision models for document work. We match the model to your task, hardware and quality bar rather than defaulting to the biggest one.

Do you deploy on our hardware or supply it?

Either. We advise on and size the hardware (Strix Halo, DGX Spark, Mac Studio, server GPUs), or deploy onto kit you already own or run in your own VPC. The whole stack, models, gateway and governance, runs inside your control; day-to-day operations stay with your team unless you want a managed retainer.

Where do the hardware prices on this page come from?

Every morning our Deneb engine reads each seller's own product page: Singapore resellers, Apple Singapore, and overseas stores such as GMKtec and Bosgame. Overseas prices are converted to SGD with 9% GST at import added; shipping is not included. This page re-reads that feed every six hours, so prices change without anyone editing the page. They are list prices on the day, not a quote.

What is Deneb?

Deneb is our setup agent for private LLMs. It lives on your box, reads the logs, diagnoses problems and applies fixes only with your say-so, with a human engineer behind it. The client is open source, read-only by default, and never runs anything destructive or privileged. On the web, Deneb also answers whether on-prem AI fits your business and what it costs.

Go deeper

Field notes from running this for real

We build in the open. These are the write-ups behind the numbers above: the trade-offs, the failures and the fixes.

Free field guide

Get the Private On-Prem AI Playbook

The decision framework, the TCO math, box sizing across Strix Halo, DGX Spark and Mac Studio, model and quant choices, the deployment path and the compliance checklist, measured on real boxes. Enter your email and we will send you the PDF.

One email, the PDF, no spam. We use your address to send the guide and follow up once.