Misreads real speech.
Accents, code-switch, dialects.
Synthetic Data Engine
SYNTHESIS-ON-GRAPH
Amplifies a handful of real data 10–100× into training corpora.
Speaks your language.
Trained on data grown from yours.
ArcNotes · AI MEETING NOTES
Upload audio or video, or record live, and get structured minutes — decisions, action items and owners — with speakers separated and word-level timestamps.
Pick an industry — see what the system looks like.
Transcribes · Translates · Dubs
One upload becomes a finished localized episode — speech recognized, subtitles translated and timed, voices dubbed and lip-aware.
YOUR INFRASTRUCTURE
DATA NEVER LEAVES
ASR
Translation
TTS
Render
STARTS FROM
your existing episodes
subtitle production per episode
monthly output
Answers · Guards · Ramps
Answers your agents mid-call from your own policy terms, product rules and compliance scripts — no hold time, no mis-quotes.
YOUR INFRASTRUCTURE
DATA NEVER LEAVES
AI Agent
Knowledge Base
Context Graph
On-prem
STARTS FROM
your policy documents and call scripts
new-agent ramp time
on-hold lookups
Recalls · Reasons · Cites
Your research, memos and meeting records — recalled in one query, reasoned across, and cited back to the source document.
YOUR INFRASTRUCTURE
DATA NEVER LEAVES
Knowledge Base
Context Graph
ASR
On-prem
STARTS FROM
your documents and meetings
research turnaround
answers with cited sources
Learns · Generates · Seals
Sensitive patterns learned, amplified and structured into model-ready training data — your raw records never surface.
YOUR INFRASTRUCTURE
DATA NEVER LEAVES
Synthetic Data Engine
Context Graph
On-prem
STARTS FROM
a sealed seed corpus
seed corpus amplification
data production cost
Same engine, assembled differently.
THE DETAILS
Arclek, explained.
Arclek is an AI company specializing in speech and knowledge AI for Southeast Asian languages — Thai, Malay, Indonesian, and Vietnamese. Its models are built on the team's own synthetic data engine and context graph technology — methods the founders authored and published — which is how it reaches state-of-the-art accuracy in low-resource languages and builds AI that understands specific industries. Arclek offers ArcDub (AI video localization), ArcNotes (AI meeting notes), a developer speech API, and custom enterprise AI systems, all deployable on-premise inside the client's own infrastructure.
WHY
Southeast Asian languages are a rounding error in the corpora general models are pretrained on, and real speech here mixes languages inside a single sentence. A general open model reaches only 82.99% word accuracy on Thai, against 92.79% for Arclek's Thai speech recognition — meaning is lost before any reasoning starts. Arclek trains speech models specifically for these languages, including code-switching, regional accents and dialects.
An industry's working knowledge — product clauses, compliance scripts, decision paths, edge cases — lives in internal documents and in veterans' heads, not on the public internet. No amount of prompting retrieves what was never in the training data. Arclek structures a client's own material into a context graph, so the model reasons inside their business instead of guessing at it.
In regulated industries, data residency requirements such as BNM RMiT and PDPA make it non-negotiable that customer data stays inside the institution's own jurisdiction. Cloud-only AI sends that data out by design, so it fails the audit before a pilot begins. Arclek deploys on-premise by default: models run inside the client's own infrastructure, and the intelligence built there belongs to the client.
HOW
A synthetic data engine turns a small amount of real material into training-scale data. Arclek's engine grows a client's seed corpus 10–100× while retaining 92% diversity and cutting data production cost by 85.7% — which is how models get built for languages and industries where real data is scarce, sensitive or long-tail. The client's raw records stay sealed throughout; only the patterns are learned.
A context graph structures an organization's terms, rules, decision paths and edge cases so a model can reason over them rather than merely retrieve matching text. In Arclek's systems it improves reasoning accuracy by 25.4% over retrieval alone. It is what makes an AI answer like someone who has worked in the business, not like a search box.
Arclek's Thai speech recognition reaches 92.79% word accuracy, ahead of every global vendor measured, and its Thai speech synthesis scores 4.41 overall MOS — also first. Arclek's Malay model reaches 87.6% word accuracy on an in-domain test set and ranks first on FLEURS at 7.13% WER, as a direct transfer of the Thai engine before tuning.
Arclek ASR-TH
92.79
ElevenLabs Scribe
91.68
Google Chirp v3
91.44
Whisper-large-v3
82.99
Qwen3-asr
68.19
Thai ASR · word accuracy % · measured Q2 2026
Thai and Malay speech AI are production-ready today. Indonesian and Vietnamese are in active development and available to early partners through custom engagements. Arclek's models are built for how the region actually speaks — mixed languages, regional accents and dialects — rather than a global model that happens to include these languages.
PRODUCTS & WORKING TOGETHER
ArcDub is Arclek's AI video localization product: one upload becomes a finished localized episode — speech recognized, subtitles translated and timed, voices dubbed and lip-aware. Character voices, glossary terms and subtitle styles carry across a whole series, and editing a single line re-renders only that segment. Transcription, translation and lip-sync all run on Arclek's own models.
ArcNotes is Arclek's AI meeting notes product: upload audio or video, or record live, and get structured minutes — decisions, action items and owners — with speakers separated and word-level timestamps. It handles up to three hours per file and takes mixed Malay-English speech as it is actually spoken. An open API plugs the same transcription into your own systems.
The Arclek API exposes the same speech recognition and synthesis that power ArcDub and ArcNotes, so developers can build transcription, voice interfaces and localization into their own products. It supports streaming and batch transcription, word-level timestamps, hot-word tuning and multi-voice synthesis. Full reference and a live playground are available in the developer documentation.
Most of what Arclek delivers is custom: AI built from a client's own material — their terms, their compliance language, their voices — and deployed inside their own walls. Engagements run in four stages: scoping, validation, pilot, then scale. The systems assemble from the same engine, which is why a video translation pipeline, a telesales copilot, a retrieval and reasoning engine and a privacy-safe data pipeline can all be built from the same parts.
COMPANY
Arclek is the international brand of DataArc Technology Limited (Hong Kong), founded in 2025 and incubated by IDEA Research Institute. It is led by Dr. Xuhui Jiang (Founder & CEO — PhD, Chinese Academy of Sciences ICT; AI Scientist at IDEA; contributor to China's first synthetic-data standard) and Dr. Chengjin Xu (CTO — PhD, University of Bonn; Huawei "Genius Youth"; head of IDEA's financial LLM programme), with Dr. Harry Shum as advisor — founding chairman of IDEA, former EVP of Microsoft, member of the US National Academy of Engineering. The team holds 10+ patents and 100+ top-tier papers with 7,000+ citations; the methods behind the pipeline — Think-on-Graph, Synthesis-on-Graph and LLM-as-a-Judge — are published and open-sourced on GitHub.
BUILT AROUND YOUR BUSINESS
Bring a problem. Leave with an approach.
Most of what Arclek delivers is custom: AI built on your own material — deployed inside your own walls.