High-quality data pipelines
Curation, filtering, and synthesis for Indic and Global South text, code, math, and speech.
// PROJECT EKA
Project EKA is Soket's flagship research program — advancing frontier math, code, and reasoning models in parallel with sovereign multilingual AI for Indian and Global South languages.
Sovereign compute via IndiaAI Mission
High-quality pre-training corpus
Frontier-scale architecture in training
Indian, Global South, and programming
// WHAT WE'RE BUILDING
Project EKA is Soket's boldest vision — building AI for a billion, from the heart of India. Our mission is to create world-class models that master math, code, and reasoning, while speaking the languages of Bharat and the Global South.
We work at the edge of research in architecture, large-scale training, and language resources — reimagining what's possible for low-resource and diverse languages.
We open-source where we can, train on sovereign compute, and publish research that advances Indic NLP, systems for ML, and efficient inference.

// CORE CAPABILITIES
Advanced mathematical reasoning, proofs, and symbolic computation — models tuned for rigorous step-by-step logic.
Multi-language code generation, debugging, and optimization across 20+ programming languages.
Logical deduction, complex analysis, and long-horizon problem solving for high-stakes workflows.
22 Indian languages, 20+ Global South languages, and English — sovereign text and speech with curated data and tokenization.
// LANGUAGE COVERAGE
Most frontier labs optimize for English. EKA runs two parallel verticals: frontier math, code, and reasoning models for rigorous technical work — alongside sovereign multilingual modeling with efficient tokenization and curated corpora.
Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Odia, Punjabi, Assamese, Urdu, Sanskrit
Arabic, Indonesian, Thai, Vietnamese, Burmese, Kazakh, Portuguese, Spanish
Python, Rust, Go, TypeScript, C++, Java, SQL, Julia
// RESEARCH DIRECTIONS
If you care about data systems, training at scale, tokenizers, post-training, or ethical AI — these are the threads where your work ships into a national-scale model.
Curation, filtering, and synthesis for Indic and Global South text, code, math, and speech.
Routing strategies for token-efficiency and morphological diversity across scripts.
Token-efficient vocabularies for Indian languages — minimizing bytes-per-token for Indic scripts.
SFT and preference optimization for math, code, and reasoning; alignment for regulated sectors.
Kernel fusion, speculative decoding, quantization, and systems co-design.
Optimal power and water usage — minimizing environmental cost of frontier runs.
Defence, cybersecurity, finance, banking — with auditable, on-premise deployment.
Safety, alignment, bias mitigation — embedding constraints from data to deployment.
// RELATED RELEASES
Token-efficient vocabularies for Indian and Global South languages.
We're hiring researchers and engineers across data, training, inference, and applied ML. If you want hard systems problems at sovereign scale — we'd like to hear from you.