Private LLM · Singapore
Private, on-prem LLM deployment in Singapore — real benchmarks from our own box, honest hardware guidance, and PDPA / MAS / MOH-ready delivery. We measure; we don't quote vendor slides.
Benchmarks measured & hardware landscape verified ·
altronis@strix-halo:~$
A real session on our own Strix Halo box, recorded 15 Jul 2026 · Qwen3.6-35B, fully offline.
In one paragraph
Private (on-prem) LLM deployment means running open-weight models (Llama, Qwen, DeepSeek, Gemma) entirely inside your own network, so sensitive data never leaves your perimeter and there's no per-token API bill. On a single 128 GB unified-memory box you can serve a 100B-class model at interactive speed, or pair a fast MoE chat model with a vision-capable one (we run Qwen3.6-35B beside the new Qwen3.8-27B, which reads images and documents natively). In Singapore this is the practical path to AI for PDPA-, MAS- and MOH-regulated data. Altronis sizes the hardware, deploys the stack, and builds the governance, all inside your control, not ours.
The assistant in the corner of this page isn't a scripted demo. Your question is answered by an open-weight LLM running on our own Strix Halo box in Singapore, over a private tunnel. It's the same class of model, on the same class of hardware, that we deploy for clients. Ask it about on-prem deployment, hardware sizing, or PDPA.
If our box is offline for maintenance it falls back to a hosted model so you still get an answer. In production, on your own hardware, there is no fallback and nothing leaves your network.
Why on-prem
On-prem isn't always the answer — but when one of these is true, it's the only answer. We'll tell you honestly which side of the line you're on.
Regulated or sensitive data (PDPA, MAS, MOH, client IP) that legally or contractually can't leave your environment. On-prem removes cross-border transfer and third-party exposure at the root.
Once you sustain millions of tokens a day, per-token API pricing outruns owning the hardware — often within 12–18 months. Predictable capex replaces an unbounded metered bill.
No network round-trip, no rate limits, no silent model swaps under you. Fixed latency and a model you pin, version and audit — essential for agents and real-time workloads.
Measured, not marketed
Numbers from our own hardware, read from the server's own timings on real requests, not synthetic benchmarks or vendor slides. The rows marked LIVE are measured on this box, the Qwen3.8 row on its release day.
Generation speed — tokens / second
higher is better
| Model | Quant | Active params | Gen t/s | Prompt t/s | Context |
|---|---|---|---|---|---|
| Qwen3.8-27B · vision + native MTPmeasured 14 Aug 2026, day of release · reads images natively · two 131K slots served concurrently | UD-Q4_K_XL | 27B dense | 20-22 | ~750 | 2×131K |
| Qwen3.6-35B-A3B · native MTPmeasured live · ~65 t/s chat, ~100 t/s code | Q4_K_XL | 3B active / 35.5B | 65-100 | ~840 | 256K |
| Muse Glimmer-30B · DFlash draftvision + writing tuned · served here Jul-Aug 2026 (own server logs), since replaced by Qwen3.8 | ROCm FP4 | 30B dense | 22-35 | ~340 | 64K |
| Qwen3.6-35B-A3Bhigher-fidelity quant | Q8_K_XL | 3B active / 35.5B | ~44 | ~839 | 128K |
| Gemma 4 26B-A4Bextraction / structured output | Q8_K_XL | 4B active | ~41 | ~720 | 128K |
| Qwen3.5-122B-A10Blegacy 122B unit | Q4_K_XL | 10B active | ~22 | ~393 | 64K |
The Qwen3.6-MTP row is measured live on this box (2026-07-15, MTP on, thinking off). With speculative decoding, generation speed is content-dependent: around 65 t/s on chat and prose, and up to ~100 t/s on predictable code output, where more draft tokens are accepted. Prompt processing runs ~840 t/s. Other rows are documented runs on earlier builds. Hardware: AMD Ryzen AI Max+ 395 / Radeon 8060S (gfx1151), 128 GB unified LPDDR5X, Vulkan/RADV. Configs: github.com/sypherin/strix-halo-setup.
How the models above score on public capability benchmarks, drawn from three vendors' own publications with every number attributed to its source. Where two vendors report the same benchmark, both numbers are shown; they differ because each ran its own harness, which is exactly why single-source benchmark tables deserve suspicion. Muse Glimmer-30B is the model this box ran before the Qwen3.8 upgrade; Opus 4.6 is included as the frontier reference.
| Benchmark | Qwen3.8-27B (ours now) | Qwen3.6-27B | Muse Glimmer-30B (ours before) | Opus 4.6 Max |
|---|---|---|---|---|
| Terminal Bench 2.1 (agentic terminal coding) | 73.0ᵃ | 63.4ᵃ · 60.7ᵇ | 51.7ᵃᵇ | 78.2ᵃ |
| SWE-bench Pro (agentic coding) | 61.7ᵃ | 53.5ᵃ · 50.2ᵇ | 51.2ᵃᵇ | 53.4ᵃ |
| SWE-bench Verified (agentic coding) | — | 77.2ᵇ | 76.0ᵇ | 80.8ᶜ |
| QwenSWEBench (software engineering) | 79.0ᵃ | 49.3ᵃ | — | 63.8ᵃ |
| CoWorkBench (long-horizon office work) | 70.7ᵃ | 61.0ᵃ | — | 68.2ᵃ |
Sources: ᵃ Qwen3.8-27B model card (huggingface.co/Qwen/Qwen3.8-27B, Aug 2026) · ᵇ Muse Glimmer-30B model card (huggingface.co/meta-models/Muse-Glimmer-30B, Aug 2026; its Qwen3.6-27B numbers were run in thinking mode, Terminal Bench via the Terminus 2 harness) · ᶜ Anthropic Opus 4.6 system card (Feb 2026; 80.84% averaged over 25 trials). Higher is better. A dash means no source we cite reports that cell. Anthropic separately reports state-of-the-art on Terminal-Bench 2.0, an earlier version not comparable with the 2.1 rows here.
Gemma-4 26B (Q8, 4K context), same 15-page document batch. The A100 row is a documented cloud reference.
Pick the box
For a single team it comes down to four options. The right one depends on budget, tooling preference, and how many people hit it at once.
Specs & pricing verified 15 August 2026
Best valueThe most memory-per-dollar on a desk: 16 Zen 5 cores, 40 RDNA 3.5 CUs, XDNA2 NPU, and 128 GB unified LPDDR5X (up to 96 GB usable as VRAM). Holds a 100B-class MoE via Vulkan/RADV; ~256 GB/s bandwidth is the ceiling. After the 2026 memory crunch: ~S$3,100 for a 96 GB import, ~S$3,700–4,700 for a 128 GB import (Bosgame, GMKtec), up to ~S$5,400 for a locally-warrantied 128 GB build. Still the best cost-per-token for a single-team private stack.
Best toolingGB10 Grace Blackwell Superchip, 128 GB unified, ~1 petaFLOP (FP4), up to ~200B params per unit. ~25% faster and steadier per-token than Strix with mature CUDA tooling (vLLM, TensorRT-LLM). ~S$8,100–8,800 at Singapore resellers for the 4 TB unit (kept climbing through the 2026 memory crunch). The pick for NVIDIA's ecosystem or mixed training + inference.
Quiet / MacUp to 256 GB unified memory (Apple discontinued the 512 GB tier in early 2026 amid the memory shortage), with 546–819 GB/s bandwidth, well above Strix, plus MLX / llama.cpp Metal. From S$7,799 on Apple SG for the M3 Ultra base. Silent, and a fit for Mac-first teams and quiet offices.
Max throughput80–192 GB HBM3e (H100 80 GB, H200 141 GB, B200 192 GB) at up to 8 TB/s — 3–5× the throughput for many concurrent users, but capex-heavy or recurring cloud cost. For high-QPS multi-tenant production, not a single team's box.
Under the hood
Running a big model on a small box is a stack problem, not a hardware problem. These are the levers that decide whether it flies or crawls.
Fresh upstream llama.cpp with Vulkan or CUDA. GGUF quantization (Q4_K_XL for speed, Q8 for fidelity) shrinks a 35B model to ~23 GB. Native multi-token prediction roughly doubles decode speed vs a plain run.
Mixture-of-Experts models (Qwen3.6-35B-A3B, Gemma-4-26B-A4B) activate only 3–4B parameters per token, so a 35B model decodes at 3B-model speed while keeping large-model quality — ideal for a memory-rich, bandwidth-bound box.
On AMD gfx1151 we measured Vulkan/RADV ~5% faster generation and ~47% faster prompt-processing than ROCm for LLM inference, and far more stable. FP8 is broken on RDNA 3.5 — use BF16. Backend choice is not cosmetic.
256K-token native context on a single box, with a q8_0 KV cache to keep memory tiny even at long context. Unified memory (a ~124 GiB GTT window) means weights + KV live in host RAM without a hard GPU/CPU split.
How we deploy
We don't hand you a research script. We deploy a governed, single-endpoint stack your apps point at — models behind an authenticated gateway, everything inside your perimeter.
We ship the stack with Deneb, our setup agent that lives on your box — it reads the logs, diagnoses problems, and applies the fix (reversibly, only with your say-so), with a human engineer behind it. Ask it in plain English, or paste an error or a screenshot.
Before you ship it
A private LLM is only as safe as its weakest prompt. Vedis — our AI-security appliance — red-teams your own AI features and runs entirely on your hardware, so nothing it tests ever leaves your network. Authorized and defensive, mapped to the OWASP LLM Top 10.
Regulated-data workloads
On-prem is the architecture regulated Singapore teams use to keep data they can't send to a hosted API inside their own walls. To be clear — we're not a compliance certifier: on-prem, plus the governance and audit we build, supports your obligations, but the regulatory relationship stays yours.
Because data never leaves your perimeter, cross-border transfer and third-party-processor exposure largely fall away — the hardest part of your PDPA obligations, handled by the architecture rather than a policy promise.
For financial institutions, on-prem inference keeps prompts, outputs and customer data inside the environment you already control and audit — the posture you need before putting AI near client data. Your MAS relationship stays yours.
Patient data stays in-tenant; no identifiable data enters a hosted model. Where AI touches triage or diagnosis we flag that HSA may treat it as a medical device (AI-SaMD) — so you assess it early, not after a build.
Straight answers
Yes. Open-weight models (Llama, Qwen, DeepSeek, Gemma, GPT-OSS) run entirely inside your own network on a single 128 GB unified-memory box such as an AMD Strix Halo or NVIDIA DGX Spark — no data leaves your perimeter, no per-token API bill.
On-prem removes the hardest PDPA risks — cross-border transfer and third-party processing — because data never leaves your environment. It is not automatically compliant on its own; you still need access control, audit trails and governance, which we build in. No serious vendor should claim 'fully PDPA-compliant' out of the box.
Measured live on our own AMD Strix Halo box, a 35B Qwen3.6 MoE model with native multi-token prediction decodes at around 65 tokens/second on chat and prose, and up to ~100 tokens/second on predictable code output, where speculative decoding accepts more draft tokens. It processes prompts at roughly 840 tokens/second. A NVIDIA DGX Spark is about 25% faster again. The newer dense Qwen3.8-27B, a vision-native model, decodes at a measured 20-22 tokens/second on the same box with multi-token prediction, and its hybrid linear-attention design keeps long context cheap: we serve two 131k-token slots concurrently on one machine. That is comfortably interactive for chat, extraction and agent workloads.
For a single team, a 128 GB unified-memory box: AMD Strix Halo (Ryzen AI Max+ 395) for best value, or NVIDIA DGX Spark for the CUDA ecosystem. Both hold a 100B-class MoE model. For many concurrent users you move to data-centre GPUs (A100/H100), which trade cost for throughput.
AMD Strix Halo gives the most memory per dollar and runs the largest models on a desk via Vulkan/RADV; NVIDIA DGX Spark is ~25% faster, steadier per-token, and has the mature CUDA/vLLM/TensorRT tooling. Pick AMD for cost-per-token, NVIDIA for tooling and mixed workloads.
When data can't leave your perimeter (regulated data), when you sustain enough volume that per-token API costs outrun hardware (typically millions of tokens/day), or when you need predictable low latency. Below that, cloud APIs are usually cheaper and faster to start — we'll tell you honestly which side you're on.
Yes. A 128 GB unified-memory box holds a 100B-class model in 4-bit — we run a 122B MoE unit locally. Mixture-of-Experts models are the sweet spot: large quality, only 3–10B parameters active per token, so they stay fast on memory-bandwidth-bound hardware.
Yes, and it is how we run our own. A 128 GB box holds a fast MoE drafter (Qwen3.6-35B, ~3B active parameters, ~65 tokens/second) alongside a smarter dense vision model (Qwen3.8-27B, ~20 tokens/second, reads images and documents natively), each serving a different port. Quick chat and extraction route to the fast one; harder reasoning, coding and anything with images routes to the smart one. Same hardware, no cloud, both measured on our own machine.
A working proof of concept in 1–2 weeks and daily production use in 6–8 weeks is typical — hardware sizing, model selection and quantization, the inference gateway, governance and audit, then integration with your systems. We hand over a documented, self-serve stack, not a black box.
Llama, Qwen (incl. Qwen3.6 MoE and the vision-native Qwen3.8), DeepSeek, Google Gemma, GPT-OSS and Mistral families, and writing-tuned models such as Muse Glimmer, plus OCR/vision models (Qwen-VL, Surya) for document workloads. We match the model to your task, hardware and quality bar rather than defaulting to the biggest one.
Either. We advise on and size the hardware (Strix Halo, DGX Spark, Apple Silicon, data-centre GPUs), or deploy onto kit you already own or run in your own VPC. The whole stack — models, gateway, governance — runs inside your control; day-to-day ops stay with your team unless you want a managed retainer.
It's our packaged private-LLM deployment: a complete stack — model server, authenticated gateway and secure access — on your own hardware, shipped with Deneb, our open-source setup agent that installs, diagnoses and repairs the stack on the box, with a human engineer behind it. The client is auditable line by line, read-only by default, and never runs anything destructive or privileged. Your data and model never leave your walls.
We red-team your AI features with Vedis, our AI-security appliance — prompt injection, multi-turn jailbreaks, indirect injection via RAG/tools, data-leakage and tool-abuse — mapped to the OWASP LLM Top 10. It runs entirely on your own hardware so nothing it tests leaves your network, and it's strictly defensive: authorized testing of your own apps, with a blue-team loop for scheduled re-scans and per-deploy regression checks.
Go deeper
We build in the open. These are the real write-ups behind the benchmarks above — the trade-offs, the failures, and the fixes.