// PROJECT EKA

Backed by the IndiaAI Mission

Sovereign models for reasoning & multilingual AI

Project EKA is Soket's flagship research program — advancing frontier math, code, and reasoning models in parallel with sovereign multilingual AI for Indian and Global South languages.

1,536
H100 GPUs

Sovereign compute via IndiaAI Mission

25T
Tokens curated

High-quality pre-training corpus

120B+
Parameters

Frontier-scale architecture in training

60+
Languages

Indian, Global South, and programming

// WHAT WE'RE BUILDING

Foundation models built for Bharat & the Global South

Project EKA is Soket's boldest vision — building AI for a billion, from the heart of India. Our mission is to create world-class models that master math, code, and reasoning, while speaking the languages of Bharat and the Global South.

We work at the edge of research in architecture, large-scale training, and language resources — reimagining what's possible for low-resource and diverse languages.

We open-source where we can, train on sovereign compute, and publish research that advances Indic NLP, systems for ML, and efficient inference.

Project EKA visual

// CORE CAPABILITIES

Two tracks: technical reasoning & multilingual AI

01

Math

Advanced mathematical reasoning, proofs, and symbolic computation — models tuned for rigorous step-by-step logic.

02

Code

Multi-language code generation, debugging, and optimization across 20+ programming languages.

03

Reasoning

Logical deduction, complex analysis, and long-horizon problem solving for high-stakes workflows.

04

Multilingual

22 Indian languages, 20+ Global South languages, and English — sovereign text and speech with curated data and tokenization.

// LANGUAGE COVERAGE

60+ languages — not an afterthought

Most frontier labs optimize for English. EKA runs two parallel verticals: frontier math, code, and reasoning models for rigorous technical work — alongside sovereign multilingual modeling with efficient tokenization and curated corpora.

22 Indian languages

Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Odia, Punjabi, Assamese, Urdu, Sanskrit

20+ Global South languages

Arabic, Indonesian, Thai, Vietnamese, Burmese, Kazakh, Portuguese, Spanish

20+ programming languages

Python, Rust, Go, TypeScript, C++, Java, SQL, Julia

// RESEARCH DIRECTIONS

Problems we're actively working on

If you care about data systems, training at scale, tokenizers, post-training, or ethical AI — these are the threads where your work ships into a national-scale model.

  • 01

    High-quality data pipelines

    Curation, filtering, and synthesis for Indic and Global South text, code, math, and speech.

  • 02

    Efficient model architecture

    Routing strategies for token-efficiency and morphological diversity across scripts.

  • 03

    Efficient tokenization

    Token-efficient vocabularies for Indian languages — minimizing bytes-per-token for Indic scripts.

  • 04

    Post-training methods

    SFT and preference optimization for math, code, and reasoning; alignment for regulated sectors.

  • 05

    Efficient inference

    Kernel fusion, speculative decoding, quantization, and systems co-design.

  • 06

    Sustainable AI research

    Optimal power and water usage — minimizing environmental cost of frontier runs.

  • 07

    AI for critical sectors

    Defence, cybersecurity, finance, banking — with auditable, on-premise deployment.

  • 08

    Ethical AI

    Safety, alignment, bias mitigation — embedding constraints from data to deployment.

// RELATED RELEASES

  • Pragna-1B

    1.25B-parameter open multilingual model — Hindi, English, Gujarati, Bengali.

    Hugging Face
  • Dhrith ASR

    Emotion-aware speech recognition for India's multilingual voices.

    Read the blog
  • EKA Tokenizer

    Token-efficient vocabularies for Indian and Global South languages.

Help us train India's frontier models

We're hiring researchers and engineers across data, training, inference, and applied ML. If you want hard systems problems at sovereign scale — we'd like to hear from you.