Edge AI · MI-EAI-04
Local LLM Compute Node for the Plant
A 7–8B open-weight model runs fully on-premise on a compact GPU/NPU box and answers shop-floor questions over the LAN.
Edge AI · MI-EAI-04
A compact GPU/NPU box in the server room running a 7–8B open-weight model with a LAN chat interface — a plant AI assistant for manuals, SOPs and fault codes that never touches a public API.
- ₹1,54,000pilot BOM · hardware only
- 3 – 5 daysto a live pilot
- Intermediateintegration level
- 3validated providers
Why plant AI use has gone underground
The problem
Engineers want an AI assistant for manuals, SOPs and fault codes, and they are already pasting plant data into public chatbots on their phones. IT cannot allow it, the OEM will not permit it, and the compliance team has said no. The demand does not go away; it goes underground.
- Plant data leaking into public AI tools via phones
- Cloud LLM APIs blocked by IT, OEM or compliance
- No sanctioned assistant, so no productivity gain
- Per-token cloud pricing unpredictable at plant scale
The solution
A compact box with a GPU or NPU and 16–32 GB of unified memory runs a quantised 7–8B open-weight model through Ollama or llama.cpp, with Open WebUI on the plant LAN. Engineers get a sanctioned assistant for questions, summaries and drafting; IT gets zero outbound traffic, a fixed cost, and a foundation for the RAG and secure GenAI phases.
- 7–8B open-weight model running fully on-premise
- LAN chat UI with user accounts
- Zero outbound traffic; fixed hardware cost
- Foundation for document RAG (MI-EAI-05) and secure GenAI (MI-EAI-06)
Validated service providers
Validated by TwoElectrons and ranked by validation score: verified credentials, completed jobs on the platform and client ratings.
- Sample Provider S — Bengaluru · On-prem LLM & RAG · ★ 4.9 · 12 jobs
- Sample Provider T — Hyderabad · TinyML & edge inference · ★ 4.7 · 9 jobs
- Sample Provider U — Pune · Secure AI infrastructure · ★ 4.6 · 7 jobs
Provider listings are samples pending live registrations.
Pilot BOM and cost
Reference hardware to run the pilot. Unit costs are indicative Indian market prices for the component class.
| Qty | Item | Unit cost | Line |
|---|---|---|---|
| 1× | Edge AI compute box GPU/NPU class, 16–32 GB unified memory, e.g. Jetson Orin-class or mini-PC with discrete GPU | ₹1,40,000 | ₹1,40,000 |
| 1× | NVMe SSD 1 TB model weights and logs | ₹8,000 | ₹8,000 |
| 1× | Fanless/industrial enclosure & 19 V PSU | ₹6,000 | ₹6,000 |
| 1× | Local LLM runtime Ollama / llama.cpp with a quantised open-weight model | Free tier / open source | — |
| 1× | Chat UI on LAN Open WebUI or equivalent | Free tier / open source | — |
| Pilot BOM total | ₹1,54,000 |
POC BOM cost only — hardware for a pilot. Installation, integration and provider services are quoted separately. Request for Quote
How a local LLM node works
- Install. Mount the box in the server room; connect to the plant LAN.
- Load. Install the runtime and a quantised open-weight model.
- Open. Enable the chat UI with accounts for the pilot team.
- Verify. IT confirms zero outbound traffic; the team uses it for two weeks.
Pilot architecture
Engineers on the LAN (browser · accounts) → Edge AI compute box (GPU/NPU · 16–32 GB) → Local LLM runtime (Ollama / llama.cpp · 7–8B) → No internet path (firewall-verified)
The model, the runtime and the chat UI all live on one box on the plant network.
Business case: data security, fixed cost, sanctioned use
Ranges are typical figures reported for this class of solution; your pilot establishes the numbers for your plant.
- 0 bytes leave the plant (firewall-verified)
- Fixed hardware cost, no per-token bill (predictable at plant scale)
- Sanctioned AI use instead of shadow AI (IT and compliance on side)
- 3 – 5 days to a working node (one box, pilot team)
Who this is for: Plant engineering and IT/OT heads, Compliance and information security, Pharma, chemicals, defence suppliers, OEM-restricted plants, Groups piloting AI before a wider rollout, Any plant where public AI is banned.
Illustrative case study
Illustrative scenario · not a client reference
An API pharmaceutical plant with a strict no-cloud policy in Hyderabad (illustrative)
Set-up. One GPU box in the server room, an 8B model, Open WebUI with accounts for 25 engineers, firewall rule logging confirmed no outbound traffic.
What happened. Engineers used it for SOP summaries, deviation-report drafting and unit conversions within the first week; IT reported the number of employees using personal AI apps on plant Wi-Fi fell sharply. The plant proceeded to the document RAG phase the following month.
- Pilot budget: ₹1,54,000 hardware
- Outbound traffic: 0 (verified)
- Active users (week 2): 25
FAQ
How capable is a 7–8B model?
Good for Q&A, summarisation, drafting and code snippets; it is not a frontier model. Adding your documents through RAG (MI-EAI-05) is what makes it useful for the plant.
Which models can we run?
Any open-weight model with a permissive licence — the provider recommends one based on your languages and tasks, including Indic-language options.
What if we need more capacity?
A larger box or a second node can run bigger models or more users; the pilot sizes this from actual usage.
What does the ₹1,54,000 cover?
The compute box, storage and enclosure; the software is open source. Set-up, hardening and provider services are quoted after a Request for Quote.