Local GPU inference
A checksum-pinned Q5_K_S quant stays split across the two GPUs with all model layers resident in VRAM.
Owner-operated, fully local inference
KevinBeLLM is an open-source AI workspace served from encrypted Ubuntu hardware. Its quality-tuned Qwen3.8-27B model runs fully on two consumer GPUs with xhigh reasoning, and authorized users can connect without opening a router port.
This public GitHub Pages site is informational only. Cloudflare Access verifies an authorized identity before KevinBeLLM presents its separate application login.
Useful independence
The model and interface run on owned hardware; selected tools and authenticated remote access cross that boundary deliberately.
A checksum-pinned Q5_K_S quant stays split across the two GPUs with all model layers resident in VRAM.
Xhigh reasoning is the default, while bounded tools can retrieve fresh web, news, weather, and model information when needed.
Cloudflare Access protects the tunnel route, and KevinBeLLM still requires its own login at the origin.
The application and deployment automation are published under the AGPL without publishing credentials or private data.
Measured on Machine A
The model, sampler, and memory settings are treated as one deployment profile and checked against a fixed local corpus before they become the default.
Q5_K_S passed every checked case at seed 424242. The prior IQ4_XS deployment passed 13/14.
A best-effort estimate with a plausible 41–47 range—not an independently measured benchmark score.
Official Qwen thinking sampling, preserved reasoning, 32K context, and up to 12,288 output tokens in the browser.
The 14-case result is a single-seed deployment regression signal, not a general intelligence test. The ≈44 estimate adjusts the hosted model's reported xhigh result for this machine's quantization, 32K context, and output limits; only a complete standardized run could establish an official score. See the hosted-mode results ↗
Owned compute
Both cards live in Machine A and share one resident model without a network hop in the inference path.
Primary GPU
Secondary GPU
Quality bought with careful memory tuning. Pooling both cards keeps the larger Q5_K_S quant entirely in VRAM with a 32,768-token context, q4_0 KV cache, and a measured 67/33 layer split.
After shutdown or power loss, the machine must be physically powered on and have its LUKS passphrase entered. No remote feature bypasses that prompt.
Remote request path
Remote requests pass through two login boundaries before the local application can reach the model.
GitHub Pages is not part of this request path. The public page contains no login form, origin address, tunnel credential, model endpoint, uptime badge, or live system status.
Clear boundaries
Model inference can remain on owned hardware, but an intentional web search sends the search query to configured providers and downloads selected pages. Remote access also passes through Cloudflare before reaching Machine A.
Treat model output as a starting point. Verify important medical, legal, financial, security, and operational decisions against primary sources.
Authorized users
Continue to Cloudflare Access. KevinBeLLM will request its own login after the identity policy succeeds.
Continue to authenticated access