System and method for efficient training of domain-specialized generative language models

A modular framework for training domain-specialized generative language models addresses the inefficiencies of existing LLMs by using data curation, tokenization, and RLHF, enabling efficient and high-performance training on limited resources for domain-specific tasks.

WO2026088210A1PCT designated stage Publication Date: 2026-04-30NIYOGI MITODRU
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NIYOGI MITODRU
Filing Date
2025-10-21
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing large language models (LLMs) are computationally intensive, requiring substantial resources and are inefficient in adapting to domain-specific tasks due to limited annotated data, cultural biases, and suboptimal tokenization, leading to high costs and environmental impact, especially in domains like mathematics and law.

Method used

A modular framework for training domain-specialized generative language models using data curation, domain-specific tokenization, Chain-of-Thought (CoT) fine-tuning, and reinforcement learning from human feedback (RLHF), enabling efficient training on limited hardware with scalable model configurations.

Benefits of technology

The framework produces compact, high-performance models that can be trained on a single GPU, achieving superior reasoning and generation capabilities across diverse domains and languages, including mathematics and law, while being environmentally sustainable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IN2025051689_30042026_PF_FP_ABST
    Figure IN2025051689_30042026_PF_FP_ABST
Patent Text Reader

Abstract

The invention discloses a computer-implemented system (100) and method for training and deploying domain- specialized generative language models from scratch. The system comprises a data-curation module (102) that constructs a domain-specific corpus using quality, diversity, and classifier-based methods; a tokenizer module (104) that trains a domain-specialized tokenizer using Byte-Pair Encoding, Unigram, or hybrid tokenization; a pre-training module (106) that trains the model employing a scaled positional-embedding mechanism for extended-context learning; and a fine-tuning module (108) performing supervised Chain-of-Thought (CoT) alignment through a CoT data-generation component and reinforcement learning from human feedback using a reward model and policy optimization component implementing PPO, DPO, or GRPO algorithms. A deployment module (110) integrates the trained model into a human–computer interface (130) for agentic, reasoning, tutoring, or analytical tasks across multilingual and domain-specific applications such as mathematics, law, biomedical sciences, finance, and education.
Need to check novelty before this filing date? Find Prior Art

Description

TITLE OF THE INVENTIONSYSTEM AND METHOD FOR EFFICIENT TRAINING OF DOMAIN-SPECIALIZED GENERATIVE LANGUAGE MODELS FIELD OF INVENTION

[0001] The present invention relates to systems and methods for training computational language models. More specifically, the invention provides a framework for efficiently training generative language models from scratch, customized for specific domains of reasoning and communication.

[0002] Although this disclosure illustrates the invention using a mathematical language model embodiment, the same framework is equally applicable to other domains such as legal, medical, financial, engineering, and scientific language modeling, or any other field requiring domain-specific understanding and generation.CROSS REFERENCE TO RELATED APPLICATIONS

[0003] The present application claims the benefit of Indian Patent Application No. 202431079961 filed 21 October 2024 (21-10-2024), said application being hereby incorporated herein in its entirety by reference.BACKGROUND OF THE INVENTION

[0004] The present invention generally relates to the field of artificial intelligence and natural language processing, and more particularly, to systems and methods for training and deploying domain-specialized generative language models. The field has witnessed remarkable progress with the advent of Large Language Models (LLMs), including decoder-only architecture such as the GPT (Radford et al., 2018), LLaMa (Touvron et al., 2023a), PaLM (Chowdhery et al., 2023), and Falcon series (Almazrouei et al., 2023). These models, pretrained on massive and diverse text corpora containing hundreds of billions or trillions of tokens, have demonstrated broad generalization across a wide spectrum of tasks, including text generation, summarization, reasoning, and open-domain question answering.

[0005] Decoder-only LLMs pre-trained on large, heterogeneous datasets have been employed to address both general-purpose and domain-specific tasks. To adapt these models for specialized applications, various approaches have been proposed, such as knowledge distillation (Yao et al., 2021; Tan et al., 2023; Wen et al., 2023; Gao et al., 2024), continual pre-training and supervised fine-tuning (Gururangan et al., 2020; Colombo et al., 2024; Li et al., 2024; Chen et al., 2023), retrieval augmentation (Cheng et al., 2024; Sachidananda et al., 2021), and zero-shot or prompt-based learning (Ji et al., 2025; Li et al., 2023; Fahes et al., 2022). While effective in many scenarios, these strategies are computationally intensive, relying on existing large-scale models whose training and inference demand substantial resources.

[0006] Training or adapting such models involves immense computational and financial costs, often requiring thousands of high-performance Graphics Processing Units (GPUs) and weeks of continuous operation. This results in energy inefficiency, high carbon emissions, and limited accessibility to researchers and institutions without large computational infrastructure. Consequently, there is growing interest in exploring whether smallermodels, trained from scratch on limited, domain-focused datasets, can achieve comparable reasoning performance while avoiding the cost and complexity of full-scale LLMs or their distillation pipelines.

[0007] Domain adaptation further poses data and tokenization challenges. In many specialized domains, such as law, mathematics, biomedical sciences, finance, and engineering, annotated or domain-specific data is limited, costly, and time-consuming to obtain (Chalkidis et al., 2020;Zheng et al., 2021). In legal and medical domains, for instance, source material contains specialized terminology, structured references, and multilingual code-switching, complicating tokenization and modeling (Ganguly et al., 2023). General -purpose tokenizers often fragment such domain expressions inefficiently, leading to suboptimal representations and degraded model performance. Similarly, in mathematical and scientific text, notations, equations, and symbolic constructs are often tokenized inconsistently, limiting the model’s capacity to learn logical dependencies and symbolic reasoning.

[0008] Existing LLMs also exhibit cultural, linguistic, and data-origin biases, as most are trained on predominantly English-language and Western-centric corpora (Johnson et al., 2022; Tao et al., 2024; Gallegos et al., 2024). This bias can reduce accuracy and fairness when these models are applied to non-Westem or multilingual domains, including Indian legal, financial, or educational contexts, which feature diverse linguistic traditions and specialized vocabularies. Moreover, the mismatch between the pre-trained tokenizers and embeddings of general-purpose models and the domain-specific lexicons further restricts adaptability and efficiency.

[0009] These limitations pose significant challenges for developing domain-specialized models that are computationally efficient, environmentally sustainable, and linguistically inclusive. There is therefore a pressing need for a framework that enables the development of generative language models from scratch using smaller, targeted domain-specific corpora, without relying on large-scale general-purpose models, while ensuring alignment with domain reasoning, multi-lingual adaptability, and fine-grained domain representation, capturing nuanced domain’s unique linguistic and conceptual features.

[0010] The present invention addresses these challenges by providing a modular system and method for training and deploying domain-specialized generative language models. The invention discloses a framework that integrates data curation, domain-specific tokenization, efficient pre-training architectures, Chain-of-Thought-based supervised fine-tuning, and reinforcement learning from human feedback (RLHF). The framework enables the creation of compact yet high-performing language models trained efficiently on limited domain data, supporting multilingual processing across Indo-European, Indo-Dravidian, Arabic, Latin, Mandarin, Korean, and Japanese scripts, and adaptable to a wide range of domains including mathematics, law, biomedical sciences, clinical diagnostics, pharmaceuticals, education, finance, industrial analytics, cybersecurity, software engineering, scientific research, defense, aerospace, linguistics, logistics, and environmental systems.OBJECTS OF THE INVENTION

[0011] The primary object of the present invention is to provide a computer-implemented system and method for training a small generative mathematical language model from scratch that achieves superior mathematical understanding and generation capabilities through an iterative, data-driven process.

[0012] Another object of the invention is to enable systematic curation and expansion of a high-quality mathematical corpus sourced from diverse web-based mathematical texts, facilitated by a classifier trained on an initial seed corpus to ensure relevance and domain specificity.

[0013] A further object is to develop a domain-specialized tokenizer trained from scratch specifically on the curated mathematical corpus, thereby optimizing tokenization tailored to mathematical language characteristics.

[0014] An additional object is to pre-train the auto-regressive mathematical language model on the curated corpus using a domain-specialized tokenizer, incorporating a scaled Rotary Position Embedding (RoPE) mechanism that applies a shrinking factor to token positional indices. This enables training with extended context lengths while efficiently managing limited hardware memory constraints.

[0015] Another object is to provide a multi-stage fine-tuning process for the pre-trained model, including supervised fine-tuning with Chain-of- Thought (CoT) templatized instruction datasets, followed by reinforcement learning from human feedback (RLHF) to further refine the model’s performance and responsiveness to human input.

[0016] Yet another object is to incorporate a data contamination removal mechanism to identify and exclude text segments overlapping with evaluation benchmarks, thus preventing test data contamination and ensuring unbiased model evaluation.

[0017] A further object is to train the model architecture with advanced normalization and activation layers, specifically Root Mean Square Layer Normalization (RMSNorm) and GeGLU or SwiGLU activation functions, to improve training stability and model expressiveness.

[0018] An important object is to develop and utilize a reward-based reinforcement learning mechanism comprising human-guided evaluation of model outputs through a Reward Model (RM) and a Policy Optimizer. This mechanism iteratively enhances the instruction fine-tuned language model’s mathematical reasoning and generation capabilities based on human preferences.

[0019] Another object is to provide a scalable solution capable of training auto-regressive decoder-only transformer models with parameter sizes ranging from 10 million to 3 billion, balancing computational resource requirements with performance.

[0020] Yet another objective of the present invention is to deploy the pretrained and fine-tuned mathematical language models in human-computer interfaces, enabling effective interaction and assistance for mathematical problem-solving and reasoning tasks.

[0021] Another object of the invention is to provide a flexible training framework applicable across multiple neural network architectures, including but not limited to decoder-only or encoder-decoder transformers, Mixture-of-Experts (MoE) networks, recurrent neural networks such as LSTM or Mamba, and hybrid architectures that combine transformer and recurrent components.

[0022] A further object is to support variable model configurations, such as differing numbers of layers, attention heads, embedding dimensions, feed-forward widths, and total parameter counts, thereby allowing scalable model design ranging from small efficient instances to large high -capacity ones.

[0023] Another object is to provide a domain-agnostic generative modeling pipeline, in which the mathematical domain serves as an exemplary embodiment, but which can equally train specialized models for domains including legal, biomedical, and financial text.

[0024] Yet another object is to enable deployment of the trained model as an interactive tutoring system, exemplified by a mathematics tutoring application that provides explanations and problem-solving support from elementary through college-level mathematics.SUMMARY OF THE PRESENT INVENTION

[0025] The present invention provides a computer-implemented system and method for efficiently training and deploying domain-specialized generative language models from scratch. The invention addresses the computational and adaptability limitations of existing approaches that rely on continual pre -training or finetuning of massive general-purpose language models. By introducing a unified, modular framework, the invention enables the creation of compact, high-performance, and domain-optimized models that can be trained on limited hardware resources, such as a single Graphics Processing Unit (GPU), while maintaining strong reasoning and content-generation capabilities.

[0026] In one aspect, the invention discloses a modular system architecture implemented in software instructions executing on at least one processor with memory, comprising:a data-curation module configured to iteratively construct a domain-specific corpus using a combination of (i) natural language quality -based methods that select data with acceptable perplexity (PPL) scores, (ii) diversity-based sampling to reduce redundancy and enhance representational coverage, and (iii) classifier-based selection employing word- or n-gram-based or neural classifiers such as word2vec, fastText, multi-layer perceptrons (MLPs), or transformer-based encoders and decoders;a tokenizer module trained from scratch using Byte-Pair Encoding (BPE), Unigram tokenization, or a hybrid tokenization algorithm to produce a domain-specialized tokenizer capable of efficiently encoding symbolic, multilingual, or structured domain content;a pre-training module configured to train, from scratch, a generative or auto-regressive language model using the curated corpus and domain-specific tokenizer, the module employing a scaled positionalembedding mechanism that applies a shrinking factor to extend effective context lengths under hardware constraints;a fine-tuning module comprising two stages: (i) supervised fine-tuning (SFT) using Chain-of- Thought (CoT) or domain-templated instruction data, with an integrated CoT instruction-data generation component for automatically generating structured reasoning examples; and (ii) Reinforcement Learning from Human Feedback (RLHF) using a reward model trained on human preference data and a policy optimization component implementing algorithms selected from Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), and Group Relative Policy Optimization (GRPO) to align the model’s behavior with human evaluative criteria; anda deployment module configured to integrate the trained model into human-computer interfaces such as interactive tutoring systems, question-answering agents, analytical assistants, or domain reasoning platforms.

[0027] In various embodiments, the generative language model may employ an architecture selected from a decoder-only transformer, a transformer encoder-decoder, a Mixture-of-Experts (MoE) network, a recurrent neural network such as Mamba or LSTM, or a hybrid transformer-recurrent structure. The architecture may incorporate Root-Mean-Square Layer Normalization (RMSNorm) and one or more activation functions selected from GeLU (Hendrycks & Gimpel, 2016), GeGLU (Shazeer, 2020), SwiGLU (Shazeer, 2020), ReLU (Glorot et al., 2011), LeakyReLU (Maas et al., 2013), PReLU (He et al., 2015), ELU (Clevert et al., 2016), SELU (Klambauer et al., 2017), SiLU or Swish (Ramachandran et al., 2017), SoftPlus (Dugas et al., 2001), Mish (Misra, 2020), and Sigmoid (Rumelhart et al., 1986), optimizing training stability, convergence speed, and generalization. Model configurations are scalable, supporting parameter ranges from approximately ten million to seven billion parameters to accommodate various computational environments.

[0028] In another aspect, the invention provides support for multilingual processing across Indo-European, Indo-Dravidian, Arabic, Latin, Mandarin, Korean, and Japanese scripts, enabling cross-lingual reasoning and domain adaptation. The framework is applicable to multiple specialized domains including mathematics, law, biomedical sciences, clinical diagnostics, pharmaceuticals, education, finance, industrial analytics, cybersecurity, software engineering, scientific research, defence, aerospace, linguistics, logistics, and environmental systems, each utilizing corresponding domain-specific corpora, tokenizers, and fine-tuning templates.

[0029] Accordingly, the present invention delivers a scalable, cost-efficient, and environmentally sustainable alternative to traditional large-scale model training approaches. By training compact models from scratch using optimized data curation, specialized tokenization, and multi-stage alignment techniques, the invention enables high-fidelity reasoning, interpretability, and adaptability across diverse domains and languages.BRIEF DESCRIPTION OF DRAWINGS

[0030] The other objects, features and advantages will occur to those skilled in the art from the following description of the preferred embodiment and the accompanying drawings in which:

[0031] Figures 1, 2 and 2A are the schematic representation of the computer-implemented system, according to an embodiment of the present invention.

[0032] Figure 3 is a schematic representation of the various engines involved in the computer-implemented system, according to an embodiment of the present invention.

[0033] Figure 4 is a schematic representation of the Instruction Fine-tuning Engine, according to an embodiment of the present invention.

[0034] Figure 5 is a schematic representation of the Instructions Data Generation Engine, according to an embodiment of the present invention.

[0035] Figure 6 is a graphical representation of the training perplexity v / s number of tokens for one of the pretrained model, according to an embodiment of the present invention.

[0036] Figure 7 is a graphical representation of the training loss v / s training steps for one of the pretrained model, according to an embodiment of the present invention.

[0037] Figure 8 is a graphical representation of the computing hardware device (GPU) utilization % plot.DETAILS DESCRIPTION OF INVENTION

[0038] The present invention may be embodied in several forms, and the details of embodiments of the present invention will be described in the following content with figures. The embodiments described below with reference to the drawings are merely illustrative of the technical solutions of the present disclosure but are not to be construed as limited to the technical solutions of the present disclosure.

[0039] The terms and words used in the following description and claims are not limited to the bibliographical meanings but are merely used by the inventor to enable a clear and consistent understanding of the invention. Accordingly, it should be apparent to those skilled in the art that the following description of the present invention is provided for illustration purposes only and not for the purpose of limiting the invention as defined by the appended claims. As used in the description of the invention and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0040] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0041] For the purposes of the present disclosure and the appended claims, the following expressions shall have the meanings set forth below, unless the context clearly indicates otherwise.

[0042] The term “domain” refers to a specialized field of knowledge or application, including, without limitation, the fields of mathematics, law, medicine, finance, engineering, and scientific disciplines, or any other area in which language modeling and reasoning are to be applied.

[0043] The expression “domain-specific corpus” denotes a dataset primarily composed of textual, symbolic, or code-based material representative of a particular domain, optionally curated from multiple digital or webbased sources, and suitable for training a language model to learn domain-relevant semantics and syntax.

[0044] The term “domain-specialized tokenizer” refers to a tokenizer that is trained from scratch on the domain-specific corpus, such that its subword vocabulary, segmentation rules, and encoding schemes are optimized to capture domain-specific symbols, terminology, and notational patterns.

[0045] The term “Mamba” designates a recurrent sequence modeling architecture implementing state-space dynamics or equivalent recurrent computation mechanisms adapted for efficient sequence learning.

[0046] The term “Chain-of- Thought (CoT)” refers to a structured reasoning format used for training or inference, in which each example or prompt includes explicit step-by-step logical or mathematical reasoning sequences that lead to a corresponding conclusion or answer.

[0047] The present invention discloses a computer-implemented system and method for training and deploying a domain-specialized generative language model. The invention provides an end-to-end, scalable framework for developing compact, high-performance generative models trained from scratch, optimized for reasoning, comprehension, and content generation within a selected domain such as mathematics, law, biomedical sciences, finance, education, or engineering. The invention further discloses a unified architecture comprising data curation, tokenization, model pre-training, multi-stage fine-tuning, and deployment modules that collectively produce efficient, domain-aligned models suitable for reasoning, tutoring, and analytical applications.

[0048] In an embodiment, the system comprises a processor and a memory storing executable instructions configured to perform the stages of data collection, model training, and deployment. The system includes a data curation module, a tokenizer module, a pre-training module, a fine-tuning module, and a deployment module, each executing distinct yet interrelated functions within the model development lifecycle. The modules collectively enable efficient training, alignment, and deployment of a domain-specialized generative language model across multiple domains.

[0049] In an embodiment, the data curation module is configured to generate a domain-specific corpus by employing multiple complementary strategies to ensure both data quality and diversity. The curation process includes: (i) Natural-language quality-based methods, which select data samples demonstrating acceptable perplexity (PPL) scores on validation datasets to maintain linguistic coherence and contextual accuracy; (ii) Diversity-based methods, which prioritize reduction of redundancy while preserving domain coverage, ensuring representation of diverse content types and terminologies; and (iii) Classifier-based methods, wherein domain relevance is determined using classifiers trained on a seed corpus of verified domain materials. Theclassifiers may include word- or n-gram-based classifiers, such as word2vec or fastText, or neural classifiers, such as multi-layer perceptrons (MLP) or transformer-based encoder, decoder, or encoder-decoder architectures.

[0050] The data curation module further incorporates a benchmark -filtering engine to detect and remove text segments overlapping with benchmark or evaluation datasets, thereby preventing contamination and ensuring unbiased model performance. The curated corpus thus produced contains high-quality, diverse, and representative domain-specific data, suitable for efficient model training and evaluation.

[0051] In another embodiment, the tokenizer module is configured to train, from scratch, a domain-specialized tokenizer on the curated corpus. The tokenizer may employ Byte-Pair Encoding (BPE), Unigram tokenization, or a hybrid tokenization algorithm to generate subword units optimized for the domain’s linguistic and symbolic structure. In mathematical embodiments, the tokenizer identifies and encodes formulas, symbols, and structured notations such as LaTeX or code segments. In legal or biomedical embodiments, the tokenizer captures terminology, citations, and structured domain expressions. The resulting tokenizer ensures efficient encoding and semantic preservation, enabling the model to learn symbolic and textual dependencies accurately.

[0052] In a further embodiment, the pre-training module is configured to train, from scratch, the generative language model using the curated corpus and the domain-specialized tokenizer. The model architecture may be selected from one of several types, including a decoder-only transformer, a transformer encoder-decoder, a mixture-of-experts (MoE) model, a recurrent neural network such as Mamba or LSTM, or a hybrid transformer-recurrent structure.

[0053] During pre-training, the module employs a scaled positional -embedding mechanism that applies a shrinking factor to token positional indices to enable longer effective sequence lengths under limited computational resources. In certain embodiments, a scaled Rotary Position Embedding (RoPE) is implemented to maintain relative positional awareness.

[0054] The generative model employs Root-Mean-Square Layer Normalization (RMSNorm) and one or more activation functions selected from GeLU, GeGLU, SwiGLU, ReLU, LeakyReLU, SiLU (Swish), ELU, SELU, PReLU, Mish, SoftPlus, and Sigmoid, incorporated within feed-forward or expert layers to enhance stability, convergence, and generalization. The architecture supports configurable model sizes ranging from ten million to seven billion parameters, enabling scalability across different computational budgets.

[0055] In an embodiment, the fine-tuning module is configured to align the pre-trained model through a multistage fine-tuning process. In the first stage, supervised fine-tuning (SFT) is performed using Chain-of- Thought (CoT) or domain-templated instruction data. The fine-tuning module includes a CoT instruction-data generation component, which automatically generates structured instruction-input-output triplets representing step-by-step reasoning within the domain. This component extracts candidate problems from the curated corpus, generates intermediate reasoning sequences using a reasoning engine or symbolic solver, and formats the datainto training templates suitable for instruction tuning. This stage imparts explicit reasoning capabilities and structured problem-solving behavior to the model.

[0056] In another embodiment, the fine-tuning module performs a second stage of alignment through Reinforcement Learning from Human Feedback (RLHF). Human evaluators define preference criteria that distinguish preferred responses from non-preferred ones based on correctness, clarity, factual consistency, and reasoning quality. The system collects human preference data according to these criteria and trains a reward model to predict reward scores approximating human judgment. The policy optimization component then updates the model’s parameters to maximize the expected reward using at least one reinforcement-learning algorithm selected from Proximal Policy Optimization (PPO) [Schulman et al., 2017], Direct Preference Optimization (DPO) [Rafailov et al., 2023], and Group Relative Policy Optimization (GRPO) [Hong et al., 2024], These algorithms iteratively refine the model’s behavior to align its responses with human evaluative standards while maintaining training stability and sample efficiency.

[0057] In an embodiment, the deployment module integrates the fine-tuned model within a human-computer interface for real-time interaction, reasoning, or tutoring. The interface may include an application programming interface (API), a chat-based system, or an educational or analytical software platform. In one embodiment, the system functions as an interactive tutoring platform, providing step-by-step reasoning, hints, and adaptive feedback across mathematics, law, biomedical science, finance, or other domains. The deployed model can interpret natural-language questions, generate structured explanations, and adapt response difficulty based on user performance. In alternate embodiments, the same framework is deployed for legal document analysis, clinical diagnostics, financial modeling, cybersecurity assessment, or software debugging, by substituting the underlying corpus and tokenizer with domain-specific equivalents.

[0058] In another embodiment, the system operates as an end-to-end framework, wherein the data curation module gathers and filters domain data; the tokenizer module constructs a domain-optimized vocabulary; the pre-training module trains the base generative model using the scaled positional embedding mechanism; the fine-tuning module, through the CoT data generation and RLHF stages, aligns model reasoning and responses with human intent; and the deployment module delivers the resulting model via an interactive, multimodal platform. The resulting models exhibit superior reasoning accuracy, interpretability, and computational efficiency compared with conventional transfer-learned large-scale models.

[0059] The present invention therefore provides a versatile, multilingual, and domain-independent framework applicable across multiple fields of knowledge. The framework supports multilingual processing across Indo-European, Indo-Dravidian, Arabic, Latin, Mandarin, Korean, and Japanese scripts, enabling cross-lingual reasoning and translation. While a mathematical embodiment is described for illustration, the same architecture and methodology are equally applicable to law, biomedical sciences, clinical diagnostics, pharmaceuticals, education, finance, industrial analytics, cybersecurity, software engineering, scientific research, defence, aerospace, linguistics, logistics, and environmental systems. The disclosed system and method together enableefficient, cost-effective, and high-performance training of generative language models from scratch, yielding specialized models that combine reasoning transparency, cross-domain adaptability, and computational efficiency.

[0060] With reference to Figure 1, the computer-implemented system (100) interacts with user devices (130) over communication networks (120) that may include wide area networks (WANs), local area network (LANs), or a combination of both. These networks can span intranets, the Internet, or enterprise networks. The system (100) may consist of multiple servers working together to deliver the specified functions. User devices (130) can range from personal computers and tablets to mobile devices, each equipped with one or more processors. The human-computer interface (130) includes an application - either a web browser or a specialized app - that allows users to interact with the computer-implemented system (100).

[0061] With reference to Figure 1, an embodiment of the present invention is a computer-implemented system (100) is designed to operate on one or more computing devices, which may include servers, workstations, or a distributed network of computers. The system (100) is fundamentally comprised of a processor (112) and a memory (114). The memory (114) stores computer-executable instructions that, when executed by the processor (112), configure the system with a plurality of functional modules designed to carry out the inventive method.

[0062] The system (100) is interconnected with a communication network (120), allowing it to access external data sources and interact with end-user devices (130) for deployment. The core of the system (130) is a series of interconnected engines or modules that manage the end-to-end process of model creation and deployment.

[0063] With reference to Figures 2 & 3, the process begins with the Data Curation Module (102), which comprises Mathematical Data Crawler Engine (1021), Data Storage Device (1022), Data Cleaning Engine (1023) and Deduplicated and Cleaned Mathematical Data Curation Storage (1024). This module is responsible for generating a large-scale, high-quality mathematical corpus.

[0064] In this Data Curation Module (102), in the start process comprising a collection of high-quality mathematical web texts is used as initial seed corpus. Using the initial seed corpus, a text embedding classifier model such as fastText model is trained to recall more corpus alike mathematical web pages. Then the process proceeds by randomly selecting a set of 700,000 data points from the initial seed corpus as positive training examples and another set of 700,000 web pages from Common Crawl as negative training examples ones. At this stage an open-source library is employed for training the text embedding classifier, and configuring the vector dimension to 512, learning rate to 0.1, the maximum length of word n-gram to 3, the minimum number of word occurrences to 3, and the number of training epochs to 10. Thereafter employing, URL-based deduplication and near-deduplication techniques, resulting in 60B HTML web pages to reduce the size of the original Common Crawl.

[0065] These mathematical web pages are sorted based on scores predicted by the fastText model and filtering out the low quality mathematical content and the top-ranking pages are preserved from deduplicated Common Crawl with the fastText model, Further, additional uncollected mathematical web sources are identified andcollected. The Common Crawl collected data are organised into disjoints domains wherein each domain is collection of webpages sharing the same base URL; thereafter identifying and segregating domain which is having more that 10% of web pages collected during the first iteration of data collection containing math-related content. This is followed by annotating manually the URLs associated with mathematical content within these identified and segregated domains, and then adding webpages which are linked to the URLs but are uncollected to the initial seed corpus.

[0066] This process is repeated at least four times resulting in a more updated seed corpus compared to the seed corpus prior to each iteration stage through training of improved fastText model capable of recalling more mathematical data in the subsequent iteration stage, with 98% of data collected after the third iteration itself, and collecting 55.5 million mathematical web pages totaling to 160 billion tokens after the fourth iteration process.

[0067] At this stage filtering out of web pages is performed containing questions or answers from English mathematical benchmarks such as GSM8K and MATH, to avoid benchmark contamination of inadvertently encountering test data or benchmark data during its training and fine-tuning process, wherein the filtering process is a two-fold process by first matching any text segment containing a 10-gram string that matches exactly with any substring from the evaluation benchmarks is removed from training seed corpus and where the benchmark texts are shorter than 10 grams but have at least 3 grams, exact matching is employed to filter out contaminated web pages.

[0068] To build this curated high quality mathematical data, in addition to the above, other relevant data such as curated lecture notes, source code of programming languages like TEX, Python, C, Matlab, etc., and mathematical question answers tuples in Chain-of- Thought (CoT) (referred from article titled “Self-instruct: Aligning language models with self-generated instructions” (Wang et al., 2023)) templatised format, resulting in 200 billion tokens mathematical corpus are collected.

[0069] With reference to Figures 2 & 3, after the Data Curation Module (102), the Tokenizer Module (104) is configured to generate a domain-specialized tokenizer from scratch. Instead of using a generic tokenizer, the module performs Byte-Pair Encoding (BPE) training on the entire curated mathematical corpus. This corpus includes a mix of mathematical plain text, formal proofs, textbooks, lecture notes in formats like LaTeX, and source code from languages relevant to mathematics (e.g., Python, C++, Matlab).

[0070] Although the foregoing embodiment describes a tokenizer specialized for mathematical language and programming syntax, in alternative embodiments the tokenizer is trained on other domain corpora and includes domain-specific special tokens. The same tokenizer-training procedure, comprising normalization, byte fallback, and subword segmentation, applies unchanged.

[0071] Tokenization of data to convert data into sequence of integers, i.e., input tokens; is performed through Mathematical Domain-specific Tokenizer (1041) where training of Byte-Pair encoding (BPE) tokenizer using Sentencepiece module on the curated and cleaned pre-training / training data happens using a definedvocabulary size, thereby developing specialized tokenizer for mathematical domain and theorems and to learn mathematical jargon and terminology; during pretokenization, NFC normalization was performed on the processed data, digits are split into individual tokens and fall back unknown UTF-8 characters to byte granularity for improving the arithmetic learning ability of the MLNLG model and during the process data is treated as a sequence of bytes considering each of the 256 bytes as tokens and eliminating duplicate tokens from the math domain specialised tokenizer; the vocabulary size of the tokenizer can range between 1000 and 32000 and during the process data is treated as a sequence of bytes considering each of the 256 bytes as tokens and eliminating duplicate tokens from the math domain specialised tokenizer. Tokenizer has special tokens like “<Q:>”, “<A:>”, “<tex>”, “< / tex>”, “<python>”, “< / python>”, “<c>”, “< / c>”, “<matlab>”, “< / matlab>” “<haskell>”, “< / haskell>”.

[0072] With reference to Figures 2 & 3, the Pre-training Module (106) is responsible for training the autoregressive mathematical language model from scratch. Wherein the pre-training module (106) utilizes a scaled Rotary Position Embedding (RoPE) that applies a shrinking factor to the positional indices of tokens to enable training with extended context sizes on limited hardware memory.

[0073] In a preferred embodiment, the data vectorization, the process of converting the preprocessed data into numerical vectors representation through a scaled version of Rotary Position Embedding (RoPE), enabling efficient numerical processing for subsequent stages of analysis and causal language modelling.

[0074] The scaled version of RoPE is a form of static relative positional embedding, which rotates the token embedding in a form of static relative positional embedding, which rotates the token embedding by a fixed factor (0) in the higher dimensional space to encode relative positional embedding, i.e, RoPE encodes the absolute position with a rotation matrix and meanwhile incorporates the explicit relative position dependency in self-attention formulation, hence preserving the original information while also incorporating positional knowledge. The intuition behind RoPE is that representing the token embeddings as complex numbers and their positions as pure rotations that is applied to them. If both the query and key vectors are shifted by the same amount, changing absolute position but not relative position, this will lead both representations to be additionally rotated in the same manner. Thus, the angle between them will remain unchanged and, thus, the dot product will also remain unchanged.

[0075] By exploiting the nature of rotations, the dot product used in self- attention will have the property for preserving relative positional information while discarding absolute position. In one of the embodiments, the value of theta is set to 10,000.

[0076] The preprocessed tokenized training data gets converts into numerical vectors representation through scaled version of Rotary Position Embedding (RoPE) embedding through a shrinking factor by dividing the target sequence length context size for pre-training / training of mathematical language model by maximum sequence length context size feasible / permissible on single computing VRAM device memory / GPU computing device that the mathematical language model can process on a single computing hardware device forpredictions keeping all other hyperparameters fixed such as batch size, vocabulary size; thereby, allowing every positional id’s / indices of the tokens to be divided by the shrinking factor, and be bounded within the maximum feasible / permissible sequence length context size of single computing hardware device; thus, enabling the mathematical language model to pretrain / train mathematical language models at much higher / longer sequence length context size than the maximum equivalent VRAM memory availability for single computing hardware device (GPU), making training mathematical language model at infinite (theoretically maximum bound) sequence length context size on single computing hardware device (GPU) without requiring more than single computing hardware device (GPU).

[0077] For example, if the maximum permissible sequence length context size for a given single computing hardware device (GPU) physical VRAM memory hardware is 256 keeping all other hyperparameters fixed which means the position / indices of tokens are bounded within 256, then applying shrinking factor of 16 to reach target sequence length context size of 4096 on single A10040G chip during pre-training or training for higher / longer context size on single computing hardware device than the equivalent physical VRAM memory required. Then, a token with position ids’ / index = 4000 becomes 4000 / 16 = 250, and the neighbouring token with position ids’ / index = 4001 becomes 4001 / 4 = 250.25, to be within 0 to 256. This novel technique in the implementation of scaled version of RoPE embedding function enables to capture higher context sequence length of tokens that the mathematical language model can consider for pretraining during pretraining on limited physical VRAM memory required to pre-train / train model at higher / longer context size outside the maximum permissible sequence length context size on single computing hardware device. This novel modification allows to pre-train mathematical language models from scratch at much higher context size than the physical VRAM memory required for pre-training. Hence, with limited physical VRAM memory and limited computing hardware devices (GPUs), this present architecture can pretrain / train Legal language models from scratch at much higher desired context size. The following is the computer code function definition for the novel modification to scaled version of RoPE:querystates, keystates= rotarypoSemb{querystates, keystates, cos, sin, positionids / shrinkingfactor)

[0078] In the pre-training module (106), the present invention employs a Transformer-based encoder-decoder or MoE architectures trained with the causal language modeling objective of next-token prediction, optimizing it specifically for generation tasks.

[0079] While certain embodiments utilize a transformer-decoder architecture, the present framework is not limited thereto. The disclosed training, tokenization, positional encoding, and fine-tuning methods are equally applicable when the model is implemented as:a transformer encoder-decoder,a Mixture-of-Experts (MoE) model,a recurrent architecture such as LSTM or Mamba, ora hybrid architecture combining transformer and recurrent components.

[0080] For each alternative, the scaled positional embedding or an equivalent mechanism (e.g., relative or sinusoidal encoding) may be used. The same shrinking-factor technique can be applied within recurrent or MoE blocks to enable long-context training on constrained memory.

[0081] The model parameters, including the number of layers, attention heads, embedding width, expert count, and total parameters, are variable, enabling implementations from small language model (model size between 10 million and 3 billion parameters). Hyperparameter transfer (pP) techniques described herein are applicable across all such architectures.

[0082] In transformer-based encoder-decoder or MoE architectures, the scaled Rotary Position Embedding (RoPE) applies to both encoder and decoder attention operations or expert subsets. In recurrent or hybrid architectures, equivalent positional information may be incorporated by applying scaled RoPE to token embeddings prior to recurrent updates, or by using learned relative position encodings scaled by the same shrinking factor. These variations ensure the core inventive concept, efficient long-context training under hardware constraints, is preserved regardless of architecture.

[0083] In a preferred embodiment, the model architecture of present invention uses RMSNorm (referred from article titled “Root mean square layer normalization” (Zhang & Sennrich, 2019)) as pre-normalizaion layer, uses approximate version of GeGLU activation function as non-linearity function in the Feed forward dense layers by replacing the standard ReLU non-linearity activation function. The model uses multi-head attention (MHA) mechanism.

[0084] The causal language modelling objective of next token prediction is performed. The objective of the causal language modeling can formally describe as maximizing the probability of a sequence of tokenswl>w2> '" >WNnP(w1, w2, -" ,wN~) = j_^P(Wi|w1,w2, ---,Wi_1)i=l

[0085] where P(Wi|w1,w2, -",Wi_1) is the probability of token wtgiven the sequence of previous tokens w0, ••• , wt-i.e., predicting the next token in a sequence given the preceding tokens, by minimizing crossentropy loss between the predicted probability distribution of tokens over the vocabulary and the actual next token, the mathematical language model learns to assign higher probabilities to tokens that are more likely to occur next in the sequence.

[0086] The “sequence length” or “context size” typically refers to the number of tokens (words, subwords, or characters) in the input sequence that a mathematical language model processes at a time.

[0087] The mathematical language model of the present invention is based on Transformer-Decoder architecture relying on positional embeddings to encode position and location information of words in a text.

[0088] The causal language modeling of next token prediction objective by minimizing the cross-entropy loss between the predicted probability distribution of tokens over the vocabulary and the actual next token, i.e., thenegative log-likelihood of the observed data under the model under consideration, which for a given dataset is defined as:1nAvg. Loss = (P(™t I wltw2, • • • ,i = l

[0089] The pre-training processed mathematical data corpus is utilized to generate pre-trained auto-regressive mathematical language model through Pre-Training Module (106). During this module, hyperparameter-tuning is conducted to determine the optimal settings for vocabulary size, learning rate, and warm-up ratio. This meticulous tuning process ensures the models are finely calibrated for optimal performance in multitude of solving mathematical questions, theorems, proofs, and acquiring further domain specialised knowledge. A hyperparameter is a machine learning parameter whose value is chosen before a learning algorithm is trained. For pre-training a 95%-5% data split is performed for training and testing. The first performance evaluation process conducted based on validation perplexity of mathematical language models, on model FLOPs Utilization (MFU) metric. Large language models (LLMs) are expensive for hyperparameter-tuning due to their humongous number of mathematical language model parameters and large-scale dataset makes hyperparameter tuning difficult and expensive. This problem is eliminated by transferring the learned hyper tuning from small to bigger models using maximal update parameterization (pP). In a preferred embodiment, a batch size of 32, gradient accumulation steps of 8, and the maximum sequence length set to 4096. Using the pP transfer, the learned hyper parameters transferred to our bigger model from smaller models (less than 50 million parameters).

[0090] In an embodiment, the pre-training module (106) which performs pre -retraining of mathematical language model on part of curated and scrapped mathematical corpus as defined in the data collection stage at context size of 4096, wherein curated data corpus is split to use 95% of the data for pre-training which reported a validation perplexity of 4.34927 and MFU of 40.39193.

[0091] Lower the training and validation cross-entropy loss indicates better performance of the mathematical language model; however, simply computing the loss may not provide intuitive understanding of the model's effectiveness. Therefore, perplexity serves as a valuable metric for evaluating the performance of a given mathematical language model. It is defined as the exponent of the average negative log-likelihood loss across the training and validation samples.

[0092] The pre-trained mathematical language models proceed to undergo fine-tuning instruction in the Fine-training Module (108). With reference to Figures 1-4, the Fine-training Module (108) involves fine-tuning the pre-trained models using an instruction fine-tuning dataset generated from an Instruction Fine-tuning Engine (1081) and Instructions Data Generation Engine (1082).

[0093] The Fine-Tuning Module (108) refines the pre-trained model through a meticulous multi-stage process, which involves Supervised Fine-Tuning (SFT) with Chain-of- Thought and Reinforcement Learning from Human Feedback (RLHF).

[0094] In one embodiment, hyperparameter tuning was performed on a 15 million parameter model to identify the optimal vocabulary size, learning rate, learning rate scheduler, and warm-up ratio. Training utilized a batch size of 8, gradient accumulation steps of 8, and a maximum sequence length of 4096, corresponding to 262, 144 tokens per iteration. The concept of p-transfer was applied, whereby the learned hyperparameters from the 15 million parameter model were transferred to the larger, 208 million parameter model of the present invention.

[0095] In another preferred embodiment, the learning rate (Zr) decay steps is set to maximum steps of training and the minimum Ir is set nearly to 0.1 • Ir. The Ir schedule starts with a linear warm-up from 0 to the maximum Ir at 1000 steps, followed by a cosine decay to the minimum Ir until the maximum steps = 120,000 end of an epoch of training. The lrdecayratiois defined as follows:> (t - warmupsteps)uecciyrati0^decaySLepS~ WCirmiipstepswhere t is the current training step.

[0096] The maximum learning rate (Zr) is set to 3 X 10-3(max), weight decay to 1 X 10-1. To further accelerate training, BF16 mixed precision training was employed. For all experiments and model development, the implementation utilized Py Torch 2.0, supplemented with in-house optimized CUDA kernels to maximize computational efficiency. The torch.compile feature was applied to every model instance to enhance runtime performance. Figure 6 demonstrates the convergence of the loss function across incremental training steps and pretraining tokens, indicating successful pretraining with only minor fluctuations in loss values. The present invention, utilizing the 208 million parameter model, was pretrained on approximately 31.5 billion tokens.

[0097] In a preferred embodiment, the Supervised Fine-Tuning (SFT) with Chain-of- Thought process involves splitting the dataset into a 90%-10% training and testing set. During this stage, supervised fine-tuning of the pre-trained mathematical language models is executed for three epochs (Instruction, Input, Output) using the accumulated instructions through Instruction Fine-tuning Engine (1081) (Figure 4). This instruction triplet- labelled dataset corpus is generated from the Instructions Data Generation Engine (1082) (Figure 5), encompassing around 300,000 unique instructions covering various mathematical questions, mathematical theory; mathematical proof. The training and validation loss are reported using cosine, constant, and linear learning rate schedulers, with the cosine learning rate scheduler resulting in the lowest validation loss being selected.

[0098] To illustrate portability of the training pipeline, a recurrent embodiment using a Mamba architecture was trained on the same curated corpus with the domain-specialized tokenizer. The model comprised 12 recurrent layers with a hidden dimension of 2048 and context window of 2048 tokens. Pre-training and finetuning followed the same contamination-free data preparation, scaled positional encoding (adapted for recurrent gating), and RLHF steps described above. This example demonstrates applicability of the framework to nontransformer architectures while retaining efficient long-context training and high reasoning accuracy.

[0099] In a preferred embodiment, Chain-of- Thought (CoT) instruction fine-tuning was performed using the MetaMathQA instructions dataset. Rather than employing standard instruction fine-tuning protocols, each response from the MetaMathQA dataset was prepended by the prompt “Let’s think step by step,” and the resulting prompt-response pairs were utilized for fine-tuning the model. Fine-tuning was conducted for two epochs, constrained by available computational resources. The training regimen incorporated a cosine learning rate scheduler with a learning rate set to 2 x 10-5, gradient clipping at 1.0, a warm-up ratio of 0.05, and no weight decay. It is anticipated that further fine-tuning for an additional two to three epochs would improve the benchmark evaluation performance of the instruction-tuned present mathematical language model.

[0100] Although the Chain-of- Thought (CoT) fine-tuning procedure is exemplified with mathematical datasets, the same fine-tuning pipeline applies to other domain reasoning tasks, such as legal argument generation, clinical diagnosis reasoning, or financial analysis explanations. Each domain employs domainspecific instruction templates while maintaining the same supervised and reinforcement-learning stages described herein.

[0101] In the present method, the following training prompt has been used for the mathematical language model. “Below is an instruction that describes a task. Write a response that appropriately completes the request. ### Q: {question} ### A: Let’s think step by step, {answer}”

[0102] In a preferred embodiment, to enhance training stability, pre -normalization is applied to the input of each transformer sub-layer, the attention layer, and the feedforward layer. This normalization process standardizes the input vector and corresponding gradients, resulting in accelerated pre-training and training of the System (100). Moreover, this approach offers computational simplicity while effectively improving training stability. To improve the training stability, normalization of the input of each transformer sub-layer, the attention layer and the feedforward layer was performed with RMSNorm. RMSNorm accelerates the training and inference with similar performance in these large models. RMSNorm can achieves comparable performance and save training and inference time by 7% - 64% (referred from article titled “Do transformer modifications transfer across implementations and applications?” (Narang et al. , 2021)). RMSNorm improves the pre-training speed by 5% compared with the LayerNorm baseline. RMSNorm normalizes the activations based on their root mean square (RMS) value instead of normalizing the inputs based on their mean and variance. RMSNorm rescales the input vector and the corresponding gradients. It is defined by the following equation,xRMSNorm(xS) = - .where? > 0yl\\x\\l / d + Ewhere epsilon is a small factor to avoid division by zero. In the current embodiment, ? is set to 1 X 10-5.

[0103] The perplexity for the present pretrained model (4096) is 4.34927 and MFU is of value 40.39193.

[0104] After this, the pre-trained / trained mathematical language model is evaluated using GSM8K and MATH standard benchmarks datasets for evaluating quantitative reasoning, which reported test accuracyresulting a value of 39.4 on 208M parameters. The following table compares the accuracy of the present mathematical language model and other models.Evaluation of LLMs on GSM8K & MATH datasetsModel Parameters GSM8K benchmark MATH benchmarktest accuracy test accuracyLLaMa-1 7B 11.0 2.90LLaMa-1 33B 35.6 3.90LLaMa-2 13B 28.7 3.90LLaMa-2 7B 11.8 2.50CodeLLaMa 7B 10.50 13.00CodeLLaMa 34B 29.60 12.20Falcon 40B 19.60 2.50Falcon 7B 6.80 2.30MPT 30B 15.20 3.10MPT 7B 6.80 3.00GPT-J 6B 34.90Vicuna 13B 27.60PaLM 8B 4.10 1.50PaLM 62B 33.00 4.40Minerva 8B 16.20 14.10Minerva 62B 52.40 27.60Minerva 540B 58.80 33.60LLEMMA 7B 36.40 18.00LLEMMA 34B 51.50 25.00Present Model 208M 39.40 10.34

[0105] The model's capability to solve mathematical problems through chain-of-thought reasoning was evaluated using the established GSM8K (Cobbe et al., 2021) and MATH (Hendrycks et al., 2021b) benchmark datasets. GSM8K comprises 8,500 high-quality grade school mathematics problems authored by human experts, typically requiring two to eight steps with basic arithmetic operations to reach the final solution. The MATH dataset contains 12,500 problems sourced from high school mathematics competitions, each accompanied by a detailed step-by-step solution to facilitate the model's ability to generate answer derivations and explanations.

[0106] For quantitative reasoning assessment, the following evaluation prompt was used for the GSM8K test set: “Below is an instruction that describes a task. Write a response that appropriately completes the request. ### Q: {question} ### A: Let’s think step by step. The answer is: This prompt guides the model to perform multi-step reasoning in accordance with the benchmark requirements.

[0107] The present mathematical language model demonstrated superior performance on the GSM8K benchmark despite having a substantially smaller model size compared to existing large language models (LLMs). Specifically, the model, being approximately 35 times smaller than 7 billion parameter (7B) models,outperformed LLaMa-1 7B by 25.8 percentage points, LLaMa-27B by 22.17 percentage points, Falcon 7B by 29.97 percentage points, PaLM 8B by 32.67 percentage points, Minerva 8B by 20.6 percentage points, and LLEMMA 7B. Furthermore, the model exceeded the performance of PaLM 62B by 3.8 percentage points despite being approximately 305 times smaller, Falcon 40B by 17.2 percentage points (smaller by 192 times), LLaMa-1 33B by 1.2 percentage points (smaller by 158 times), and Vicuna 13B by 9.2 percentage points (smaller by 64 times). This achievement is notable as smaller models are preferred for reasons related to computational cost and environmental sustainability. Only three large LLMs — LLEMMA 34B, Minerva 62B, and Minerva 540B, outperformed the present mathematical language model on the GSM8K benchmark.

[0108] On the MATH benchmark, as reflected in Table 2, the present mathematical language model outperformed LLaMa-1 7B by 7.44 percentage points, LLaMa-1 13B by 6.44 percentage points, LLaMa-27B by 7.84 percentage points, LLaMa-2 13B by 6.44 percentage points, Falcon 7B by 8.04 percentage points, Falcon 40B by 7.84 percentage points, MPT 30B by 7.24 percentage points, PaLM 8B by 8.84 percentage points, and PaLM 62B by 5.94 percentage points. It is noted that GPTJ and Vicuna did not report results for the MATH benchmark in this evaluation.

[0109] The present mathematical language model was evaluated on standard multiple -choice mathematics question answering (MCQ) benchmark datasets and compared against various large language models (LLMs), including general-purpose LLMs, mathematics-specialized LLMs such as LLEMMA, and code-focused LLMs such as CodeLlama. The evaluation was conducted using the hn-eval-hamess framework (Sutawika et al., 2024) under a zero-shot greedy decoding setting.

[0110] The assessment included high school and college-level multiple-choice mathematics questions sourced from the MMLU benchmark (Hendrycks et al., 2021a), AGIEVAL-AQuA-RAT — which consists of 254 multiple-choice math questions derived from GRE and GMAT sections as part of the AQuA-RAT dataset (Ling et al., 2017; Zhong et al., 2024) — and AGIEVAL-SAT-Math comprising 220 SAT multiple-choice math questions. Comparative results with various LLMs are presented in the Table below.

[0111] Additionally, the model was tested on the LogiQA dataset (Liu et al., 2021), which contains logical reasoning questions originally compiled from China’s National Civil Servants Examination. The LogiQA dataset features bilingual questions in both English and Chinese; however, only the English version, which is a translation of the original Chinese text, was considered for evaluation purposes.

[0112] This evaluation demonstrates the present mathematical language model’s effectiveness across diverse multiple-choice mathematics and logical reasoning benchmarks relative to several established LLMs.Zero-shot evaluation of Present Model 208M and other LLMs.B=billion, M=millionModels Logi MMLU- MMLU- AGIEVAL- AGIEVAL QA math-high- math- AQuA-RAT SAT-Math school collegeLlama-27B 30.41 25.55 30.00 25.59 24.54 CodeLlama-7B 30.72 24.81 30.00 22.83 29.09 OLMo IB 26.81 30.37 27.00 23.62 21.81 LLEMMA 7B 29.95 32.22 32.00 23.22 32.72 Falcon 7B 26.88 21.11 21.00 22.04 28.63 Present Model 208M 30.57 31.11 29.00 26.77 25.00

[0113] On the LogiQA benchmark (Liu et al., 2021), the present mathematical language model outperformed LLaMa-2 7B and OLMo IB by 3.76 percentage points and exceeded the performance of LLEMMA 7B and Falcon 7B by 3.69 percentage points.

[0114] On the mathematical high school level questions benchmark (MMLU-math-high-school), the present model outperformed LLaMa-2 7B by 5.56 percentage points, CodeLlama 7B by 6.3 percentage points, and both OLMo IB and Falcon 7B by 10 percentage points. However, LLEMMA 7B outperformed the present model by 1 percentage point, despite being approximately 34 times larger in size.

[0115] For college-level mathematics questions (MMLU-math-college), the present mathematical language model with 208 million parameters outperformed Falcon 7B by 8 percentage points and OLMo IB by 2 percentage points. LLaMa-2 7B and CodeLlama 7B outperformed the present model by 1 percentage point each, despite being approximately 34 times larger.

[0116] On GRE and GMAT level quantitative questions (AGIEVAL-AQuA-RAT benchmark), the present mathematical language model surpassed all comparative LLMs, including Falcon 7B by 4.73 percentage points, LLEMMA 7B by 3.55 percentage points, OLMo IB by 3.15 percentage points, CodeLlama 7B by 3.94 percentage points, and LLaMa-27B by 1.18 percentage points.

[0117] On the SAT-level mathematics questions (AGIEVAL-SAT-Math benchmark), the present model outperformed LLaMa-27B and OLMo IB, while trailing LLEMMA 7B by 7.72 percentage points.

[0118] These results illustrate the present mathematical language model’s consistent high performance across multiple standard multiple-choice mathematics and logical reasoning benchmarks relative to larger state-of-the-art language models.

[0119] One of the embodiments, the pretrained mathematical language model also underwent supervised instruction fine-tuning using Chain-of- Thought (CoT) templatized mathematical instruction dataset comprising around 300,000 various labelled data triplet pairs related to mathematical question answers, mathematical proof, mathematical programming, mathematical puzzles, and theorems.

[0120] In an embodiment, the fine-tuning module (108) further comprises a Chain-of-Thought instructiondata generation component configured to automatically produce synthetic instruction-input-output triplets. The component generates CoT templates by sampling problem statements from the curated corpus, prompting abase model or heuristic reasoning engine to generate intermediate reasoning steps and final solutions, and formatting the outputs into structured triplets suitable for supervised fine-tuning.

[0121] After the supervised instruction fine-tuning using Chain-of- Thought (CoT) templatized dataset, the MLNLG mathematical model goes through reinforcement learning through human feedback mechanism (RLHF) to learn from human feedback and improve its mathematical reasoning and learning abilities.

[0122] In an embodiment, the reinforcement learning from human feedback (RLHF) utilizing a reward model trained on human preference data and a policy optimization component implementing at least one reinforcement-learning algorithm selected from Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), and Generalized Reward Policy Optimization (GRPO) to refine model outputs.

[0123] Reinforcement learning is a paradigm in which an agent learns to make decisions by receiving feedback from its environment. In the context of language models, this feedback is provided by human reviewers who assess and rate the model’s responses. By leveraging human expertise and judgments, reinforcement learning facilitates the iterative improvement of the model’s performance and fine-tunes its responses.

[0124] The process of reinforcement learning by human feedback involves several important steps:a) Guidelines are defined to guarantee unique criteria when deciding what is a good and a bad answer to an input.b) A Reward Model (RM) should be trained, which will evaluate each of the responses in terms of accuracy, relevance, and adherence to guidelines.c) To train the RM, some prompts are selected and sent to human reviewers. It is known as Preference Data (PD)d) The reviewers then interact with the model and manually evaluate and rate the corresponding outputs. e) The collected feedback, in the form of ratings or rankings, is used to train the RM.f) With the RM trained, a Policy Optimizer is trained, a required component which will guide the finetuning of the MLNLG model / instruction fine-tuned MLNLG model.g) This iterative feedback loop allows the MLNLG model to gradually learn from human guidance and refine its behavior accordingly.

[0125] The reinforcement-learning-from-human-feedback (RLHF) mechanism is domain-agnostic. For example, in a medical embodiment human reviewers may rate clinical advice responses; in a legal embodiment reviewers may assess argument validity. The same Reward Model and Policy Optimizer framework is used across domains.

[0126] In some embodiments, the fine-tuning module (108) performs reinforcement learning from human feedback through a structured multi-stage optimization procedure. The process comprises defining human evaluation criteria that distinguish between preferred and non-preferred model responses based on attributes such as correctness, reasoning quality, factual consistency, and stylistic compliance. Using these criteria, apreference dataset is collected in which multiple model outputs are comparatively rated or ranked by human reviewers.

[0127] A reward model is then trained to predict preference scores that approximate human judgment. The predicted reward values serve as optimization targets for a policy optimization component, which updates the parameters of the fine-tuned model to maximize the expected reward. In various embodiments, the policy optimization component employs one or more reinforcement-learning algorithms selected from Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), and Generalized Reward Policy Optimization (GRPO), or equivalent gradient-based optimization techniques. Each algorithm iteratively adjusts model policy parameters in response to the reward gradients, thereby aligning model behavior with human preferences while maintaining training stability and sample efficiency.

[0128] This reinforcement-learning process refines the generative model’s reasoning quality and output reliability beyond the supervised fine-tuning stage and may be executed repeatedly or in conjunction with adaptive reward-model updates to further enhance performance in the selected domain.

[0129] In the Deployment Module (110), the present invention provides a compact auto-regressive mathematical language model with a parameter size ranging from 10 million to 3 billion parameters, exemplified by a 208 million parameter instance, and a file size under 1 gigabyte. This compactness enables straightforward deployment on general-purpose computing systems without requiring specialized hardware.

[0130] The deployed model exhibits hardware independence, as it does not require a graphics processing unit (GPU) for inference. Instead, the full -precision model weights can be directly loaded into the random-access memory (RAM) of standard workstations, laptops, or servers, and executed efficiently using the central processing unit (CPU).

[0131] Unlike larger language models that often necessitate weight quantization to lower precision formats to facilitate operation on resource-constrained hardware, potentially compromising model accuracy, he present model preserves full-precision weights during deployment, thereby maintaining optimal performance.

[0132] The model can be seamlessly integrated into diverse application environments through a humancomputer interface or user experience system. It can operate as a standalone artificial intelligence tutor, serve as the backend engine for a web-based service, or function as an embedded reasoning module within larger software ecosystems. The model receives task-specific inputs and provides reasoned, step-by-step mathematical problem solutions, facilitating effective human-computer interaction in educational and computational settings.

[0133] In one exemplary embodiment, the deployed model functions as a mathematics tutoring system for instructional use from elementary through college-level mathematics. The system provides a user interface that accepts problems in natural language or symbolic form, generates step-by-step solutions using Chain-of-Thought reasoning, supplies hints or intermediate feedback, and dynamically adjusts problem difficulty basedon learner performance. The system may include grading logic to assess answers and maintain a student performance record.

[0134] In other embodiments, the same deployed framework serves as a domain tutoring system for other fields, such as legal or medical reasoning, using domain-specific corpora and reasoning templates.

[0135] The following table gives a sample output from the MLNLG mathematical model when it was queried / instructed by a human via UX system.User Meredith is a freelance blogger who writes about health topics and submits to InstructionZEnput clients each day as her permanent job. A blog article takes an average of 4 hours to Query: research and write about. Last week, she wrote 5 articles on Monday and 2 / 5 times more articles on Tuesday than on Monday. On Wednesday, she wrote twice the number of articles she wrote on Tuesday. Calculate the total number of hours she spent writing articles in the three days.MLNLG Step 1: Calculate the number of articles Meredith wrote on Monday. Meredith mathematical model wrote 5 articles on Monday. Step 2: Calculate the number of articles Meredith response: wrote on Tuesday. Meredith wrote 2 / 5 times more articles on Tuesday than on Monday. So, she wrote 2 / 5 * 5 = 2 articles more on Tuesday than on Monday.Therefore, she wrote 5 + 2 = 7 articles on Tuesday. Step 3: Calculate the total number of articles Meredith wrote in the two days. In the two days, she wrote 5 + 7 = 12 articles. Step 4: Calculate the number of articles Meredith wrote on Wednesday. Meredith wrote twice the number of articles she wrote on Tuesday on Wednesday. So, she wrote 2 * 7 = 14 articles on Wednesday. Step 5: Calculate the total number of articles Meredith wrote in the three days. In the three days, she wrote 12 + 14 = 26 articles. Step 6: Calculate the total number of hours Meredith spent writing articles in the three days. Each article takes 4 hours to research and write. So, she spent 4 hours * 26 = 104 hours writing articles in the three days. Therefore, Meredith spent 104 hours writing articles in the three days. The answer is: 104.Ground Truth 104Answer:

Claims

CLAIMS1. A computer-implemented system (100) for training a domain-specialized generative language model, the system comprising:a processor (112); anda memory (114) storing instructions that, when executed by the processor (112), configure the system (100) to perform operations comprising:a data-curation module (102) configured to iteratively curate a domain-specific corpus by:(i) applying natural language quality -based methods, including selecting data with acceptable perplexity (PPL) scores on validation datasets;(ii) applying diversity -based methods that prioritize reduction of redundancy over intrinsic quality evaluation to ensure representational coverage across the domain; and(iii) employing classifier-based methods to identify domain-relevant content, wherein the classifier is selected from word- or n-gram-based classifiers such as word2vec or fastText, or neural classifiers including multi-layer perceptrons (MLP) and transformerbased encoder, decoder, or encoder-decoder architectures;a tokenizer module (104) configured to train, from scratch, a domain-specialized tokenizer using Byte-Pair Encoding (BPE), Unigram tokenization, or a hybrid tokenization algorithm on the curated corpus;a pre-training module (106) configured to pre-train, from scratch, the generative language model on the curated corpus using the domain-specialized tokenizer,wherein the model architecture is selected from the group consisting of: a decoder-only transformer; a transformer encoder-decoder; a mixture -of-experts (MoE) model; a recurrent neural network including a Mamba or LSTM architecture; and a hybrid transformer-recurrent architecture;wherein the pre-training module (106) employs a scaled Rotary Position Embedding (RoPE) mechanism configured to apply a shrinking factor to token positional indices to enable training with extended context lengths under limited hardware memory;a fine-tuning module (108) configured to align the pre-trained model through:(i) a first stage of supervised fine-tuning using Chain-of- Thought (CoT) or domain-tem- plated instruction data, the fine-tuning module (108) further comprising a CoT instruction-data generation component; and(ii) a second stage of reinforcement learning from human feedback (RLHF) utilizing a reward model trained on human preference data and a policy optimization component implementing at least one reinforcement-learning algorithm selected from Proximal PolicyOptimization (PPO), Direct Preference Optimization (DPO), and Group Relative Policy Optimization (GRPO) to align model outputs; anda deployment module (110) configured to deploy the fine-tuned model in a human-computer interface (130) for interactive reasoning, tutoring, question answering, chatbots, intelligent agents, or problem-solving within the selected domain.

2. The system (100) as claimed in claim 1, wherein the selected domain comprises mathematics, law, biomedical sciences, clinical diagnostics, pharmaceuticals, education, finance, industrial analytics, cybersecurity, software engineering, scientific research, defense, aerospace, linguistics, logistics, and environmental systems, and the deployment module (110) is configured to operate the fine-tuned model as an interactive tutoring, expert-assistance, question-answering, or intelligent- agent system providing reasoning, analysis, hints, and adaptive feedback relevant to the selected domain.

3. The system (100) as claimed in claim 1, wherein the data-curation module (102) further comprises a benchmark-filtering engine configured to detect and remove text segments overlapping with evaluation or benchmark datasets to prevent test-data contamination and ensure unbiased training.

4. The system (100) as claimed in claim 1, wherein the scaled positional -embedding mechanism comprises a scaled rotary position embedding (RoPE) that scales tokens positional indices by the shrinking factor to bound tokens indices within a maximum permissible context length while providing an extended effective context window.

5. The system (100) as claimed in claim 1, wherein the generative language model employs Root-Mean- Square Layer Normalization (RMSNorm) and at least one activation function selected from GeLU, GeGLU, SwiGLU, ReLU, LeakyReLU, SiLU (Swish), ELU, SELU, PReLU, Mish, SoftPlus, and Sigmoid, implemented within a feed-forward or expert layer to enhance training stability, convergence speed, gradient flow, generalization, and overall model performance.

6. The system (100) as claimed in claim 1, wherein the reinforcement-learning process performed by the fine-tuning module (108) further comprises:(i) defining human evaluation criteria distinguishing preferred and non-preferred responses; (ii) collecting preference data in accordance with the criteria;(iii) training the reward model to predict preference scores; and(iv) applying at least one of the PPO, DPO, or GRPO algorithms within the policy optimization component to update the model parameters to maximize expected reward.

7. The system (100) as claimed in claim 1, wherein the generative language model comprises a variable number of layers, hidden dimensions, attention heads, and expert counts, providing model configurations ranging between 10 million and 7 billion parameters.

8. The system (100) as claimed in claim 1, wherein the generative language model is multilingual, capable of processing and generating content in multiple language families and writing systems, including Indo-European, Indo-Dravidian, Arabic, Latin, Mandarin, Korean and Japanese scripts, thereby supporting cross-lingual reasoning, translation, and domain-specific adaptation across diverse linguistic corpora.

9. A computer-implemented method for training a domain-specialized generative language model, the method comprising the steps of:iteratively curating a domain-specific corpus using a data-curation module (102); generating a domain-specialized tokenizer using a tokenizer module (104) trained from scratch on the curated corpus;pre-training, from scratch, the generative language model using a pre -training module (106) and employing a scaled positional-embedding mechanism with a shrinking factor;generating Chain-of- Thought instruction data using a CoT instruction-data generation component configured to synthesize instruction-input-output triplets illustrative of step-wise reasoning;performing supervised fine-tuning using the generated CoT or domain-templated instruction data; performing reinforcement learning from human feedback, including training a reward model on human preference data and applying at least one reinforcement-learning algorithm selected from PPO, DPO, and GRPO within a policy optimization component to optimize and align model behavior; and deploying the fine-tuned model through a deployment module in a human-computer interface for agentic use, interactive reasoning, or solving domain-specific tasks, including Named Entity Recognition, Information Extraction, Information Retrieval, Summarization, Question Answering, Prediction, Text Generation, and tutoring applications.

10. The method as claimed in claim 9, wherein the framework is applicable to specialized domains including mathematics, law, biomedical sciences, clinical diagnostics, pharmaceuticals, education, finance, industrial analytics, cybersecurity, software engineering, scientific research, defense, aerospace, linguistics, logistics, and environmental systems, each employing corresponding domain-specific corpora, tokenizers, and fine-tuning templates tailored to the respective domain characteristics.

Citation Information

Patent Citations

  • Generative large language model training method and device and value system identification application

    CN118733748A

  • Enhanced searching using fine-tuned machine learning models

    US20240281446A1