System and method for optimized real-time natural language processing
A lightweight Al system with VGQA, RoPE, and DPO addresses the limitations of traditional LLMs by providing efficient, real-time, and contextually accurate NLP on edge devices, overcoming latency and privacy issues, and supporting diverse linguistic environments.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SANDLOGIC TECHNOLOGIES PVT LTD
- Filing Date
- 2025-11-15
- Publication Date
- 2026-05-21
Smart Images

Figure IB2025061684_21052026_PF_FP_ABST
Abstract
Description
[0001] System and Method for Optimized Real-Time Natural Language Processing on Edge Devices in Resource-Constrained Environments
[0002] Cross-reference to related applications and priority
[0003] The present application claims priority from Indian provisional patent application no. 202441088473.
[0004] FIELD OF THE INVENTION:
[0005] Embodiments relate generally to artificial intelligence (Al) systems for natural language processing (NLP) on edge devices, and, more specifically, to a lightweight Al system optimized for real-time, low-latency processing in resource-constrained environments such as loT devices, smartphones, and wearables. The invention enables efficient, on-device Al with advanced features like Variable Grouped Query Attention (VGQA), Rotary Positional Embedding (RoPE), and Direct Preference Optimization (DPO) to deliver contextually accurate responses without requiring cloud support, ideal for broad industrial applications where rapid, localized Al processing is essential.
[0006] BACKGROUND OF THE INVENTION
[0007] The subject matter discussed in the background section should not be assumed to be prior art merely because of its mention here. Similarly, any problem described in the background section or associated with the subject matter of this section should not be assumed to have been previously recognized in the prior art. The subject matter in this background section represents different approaches that may themselves correspond to advancements or alternative solutions.
[0008] The rapid growth of artificial intelligence (Al) has enabled substantial advancements in natural language processing (NLP) through large language models (LLMs) capable of sophisticated tasks like question answering, text summarization, and machine translation. Models such as GPT-3 and BERT have achieved state-of-the-art performance by utilizing extensive computational resources, which allow for deep language understanding and high accuracy. However, the high resource demands of these models pose significant challenges for real-world deployment, especially on edge devices such as loT systems, smartphones, and wearables.
[0009] Traditional LLMs are cloud-dependent, requiring high-performance servers to handle computationally intense tasks. This reliance on cloud infrastructure introduces latency and security concerns, making these models less effective for applications that demand real-time, secure, on-device processing. Industries such as healthcare, finance, and customer service require immediate responses with strict data security measures, as well as contextually relevant and highly accurate NLP. However, deploying these models on edge devices is challenging due to the following issues:
[0010] High Computational and Memory Requirements’ like GPT-3 can require tens or even hundreds of billions of parameters, which consume substantial computational resources and memory. This complexity makes them impractical for deployment on devices with limited hardware capacity. For instance, in a healthcare setting, an on-device Al system could assist healthcare providers with real-time diagnostics or patient monitoring. However, the hardware limitations of many medical devices and the need for privacy-preserving, localized processing make it difficult to deploy traditional LLMs in such settings.
[0011] Latency Issues as Cloud-dependent models introduce latency, as data must be sent to and from remote servers. This delay can hinder real-time applications, such as virtual assistants or customer support bots that need to provide immediate responses. In financial services, for example, Al-driven fraud detection requires instant analysis of transactions to identify suspicious behavior. Cloud latency in such scenarios could result in delayed response times, reducing the system's effectiveness in preventing fraud.
[0012] Privacy and Security Concerns means Cloud processing raises concerns over data privacy and regulatory compliance, particularly in industries handling sensitive information. In healthcare, for example, personal health information must comply with regulations like HIPAA in the U.S. or GDPR in Europe. Processing data on external servers can increase the risk of breaches and non-compliance with these stringent regulations. Thus, there is a strong need for on-device Al that keeps data locally to reduce privacy risks.
[0013] Lack of Support for Low-Resource and Multilingual Settings, many traditional LLMs are designed primarily for high-resource languages like English and lack optimization for low-resource languages spoken in various regions. This poses challenges for multilingual applications, especially in regions with linguistic diversity. For example, in customer service, on-device Al systems that support multiple languages, including regional dialects, could serve users more effectively by providing real-time, culturally relevant assistance. Current models are often unable to perform well in these low-resource languages without significant, cloudbased fine-tuning, further limiting their utility.
[0014] The challenges outlined above highlight the need for an Al system that can operate efficiently on edge devices, delivering low-latency, high-accuracy NLP capabilities without reliance on cloud infrastructure. The present invention addresses these shortcomings by introducing a novel, edge-compatible Al system that integrates specialized architectural features to enhance efficiency, performance, and adaptability on resource-constrained devices. By leveraging Variable Grouped Query Attention (VGQA), the invention reduces memory usage and accelerates processing by dynamically managing attention computations, making it ideal for low-latency applications. The use of Rotary Positional Embedding (RoPE) further enables efficient handling of long sequences, essential for complex tasks like document processing and multi-turn dialogues without additional computational overhead.
[0015] Additionally, the invention incorporates Direct Preference Optimization (DPO), a method that aligns model outputs with user preferences through ranked feedback, optimizing the response generation process without relying on a reward model. This approach is particularly valuable in applications that require contextually relevant, user-aligned outputs, such as customer service and healthcare assistance. Furthermore, by supporting multilingual and low-resource language fine-tuning, the system enables effective on-device Al across diverse linguistic settings, expanding its applicability to regions with varied language requirements. Through these advancements, the present invention offers a practical, scalable Al solution for edge devices, addressing industry demands for real-time, secure, and efficient Al processing directly on devices with limited resources.
[0016] OBJECT OF THE INVENTION
[0017] The object of the invention is to provide an advanced, optimized system for realtime natural language processing (NLP) that operates directly on edge devices with limited resources, such as loT devices, smartphones, and wearables. By focusing on on-device processing, this invention addresses the primary limitations of traditional language models, which typically require substantial cloud infrastructure and significant computational resources. The invention thus aims to enable enterprises across industries to deploy high-performance NLP applications without depending on external servers, thereby achieving reduced latency, improved data security, and enhanced efficiency in real-time environment.
[0018] A further object of the invention is to incorporate a series of architectural features that allow the system to process data in real time with minimal delay. By integrating mechanisms such as Variable Grouped Query Attention (VGQA) and Sliding Window Attention, the system achieves low-latency responses without overloading device resources. This innovation is crucial for applications requiring immediate results, such as customer service interactions and technical support, where timely and efficient responses are critical to the user experience.
[0019] A further object of the invention is to achieve efficient and scalable on-device processing through architectural innovations such as Variable Grouped Query Attention (VGQA), Sliding Window Attention, and Rotary Positional Embedding (RoPE), ensuring context continuity and reduced memory load.
[0020] Another object of the invention is to enable the system to retain and apply contextual information across multiple interactions, thereby enhancing continuity and coherence. This feature is particularly valuable in industries where conversations often involve multiple turns, such as financial advisory sessions or healthcare consultations. By employing advanced attention mechanisms like VGQA and Rotary Positional Embeddings (RoPE), the system maintains context over extended interactions, ensuring accurate, contextually relevant responses even in lengthy or intricate dialogues.
[0021] A further object of the invention is to provide a scalable framework that can be easily adapted to different industries and specialized applications. The system’s three-tiered architecture — comprising foundation training on general datasets, domain specialization with industry-relevant data, and enterprise customization with company-specific inputs — enables it to serve diverse needs across sectors. This structure ensures that the system can be fine-tuned for specific tasks, such as legal document analysis or regulatory compliance, making it versatile and highly adaptable to evolving enterprise requirements.
[0022] An additional object of the invention is to deliver stable and reliable Al performance directly on devices with restricted hardware capabilities. By incorporating components such as SwiGLU activations and RMS normalization, the system is designed to operate effectively within the constraints of edge devices, offering consistency in both high-demand and low-resource settings. This innovation reduces reliance on cloud-based Al, allowing enterprises to deploy Al solutions that are not only efficient but also aligned with data privacy and regulatory standards, as sensitive data remains securely on-device. Another object of the invention is to provide a quantized (Q4_KM) Al system that minimizes energy consumption and hardware utilization while maintaining model accuracy, thereby enabling deployment across a wide range of devices.
[0023] A further object of the invention is to enable multilingual and cross-lingual processing, allowing accurate, real-time NLP across multiple Indic languages without external computation.
[0024] Through these objectives, the invention delivers a secure, efficient, and adaptive Al framework capable of supporting diverse industrial applications, ensuring highspeed, reliable NLP on devices with limited computational capacity.
[0025] Through these objectives, the invention provides a robust, adaptable, and efficient NLP system suitable for a wide range of real-world applications. Its unique architecture and optimization make it a practical solution for industries requiring high-speed, secure, and reliable Al processing without the need for extensive computational resources or cloud infrastructure.
[0026] SUMMARY OF INVENTION
[0027] In one or more embodiments, the present invention presents a system and method for optimized, real-time natural language processing (NLP) on edge devices. This system is designed specifically for resource-constrained environments, such as loT devices, mobile platforms, and other edge devices, to perform high-performance NLP tasks without the need for extensive cloud-based infrastructure. Additional features and advantages of embodiments of the present invention will be set forth in the description that follows, and in part will be apparent from the description, or may be learned by practice of embodiments of the present disclosure. The objectives and other advantages of the embodiments of the present disclosure may be realized and attained by the structure particularly pointed out in the written description and claims hereof, as well as in the appended drawings. The key embodiment of the invention lies in an optimized, resource-efficient Al system for real-time natural language processing (NLP) on edge devices, designed specifically for enterprise applications in resource-constrained environments. Embodiment of the present invention relates to integration a series of architectural innovations to enable high-performance, on-device Al without reliance on cloud infrastructure, thus ensuring low-latency processing, enhanced privacy, and improved data security.
[0028] The features of the embodiment of the present invention comprise:
[0029] 1. Variable Grouped Query Attention (VGQA): The system includes a Variable Grouped Query Attention mechanism that dynamically groups related queries during attention computations, thereby optimizing memory usage and reducing computational load. This feature enables the system to maintain logical context across multi-turn conversations by handling multiple related queries in a streamlined manner. Such functionality is critical for applications requiring sustained contextual awareness, such as complex, iterative financial advisories and comprehensive customer support dialogues, where it is essential to retain coherence and continuity across multiple exchanges.
[0030] 2. Rotary Positional Embeddings (RoPE): The invention further includes Rotary Positional Embeddings, a positional encoding mechanism that preserves token relationships across long text sequences without imposing additional computational burdens. RoPE facilitates the system’s capability to manage extended sequences while retaining context, making it particularly suitable for applications involving lengthy document analysis or multi-turn dialogues, such as legal document review and regulatory compliance assessments. This feature allows the system to deliver accurate and contextually aware responses over sustained interactions.
[0031] 3. SwiGLU Activation: The system employs a SwiGLU (Switching Gated Linear Units) activation function within its architecture, ensuring stable gradient flow during both training and inference. This configuration mitigates risks associated with vanishing or exploding gradients, leading to consistent and reliable model performance under varying computational demands. The SwiGLU activation is particularly advantageous for high- demand environments, such as real-time advisory services, where stable and predictable output quality is essential for maintaining operational integrity.
[0032] 4. Direct Preference Optimization (DPO): The system incorporates Direct Preference Optimization, a method for aligning model outputs with user expectations using ranked human feedback. This feature refines both the relevance and appropriateness of the system’ s responses by optimizing them according to user preferences, thereby eliminating the need for complex reinforcement learning reward models. DPO is particularly beneficial in customer-facing applications, such as healthcare consultations and customer service, where response accuracy and contextual sensitivity are crucial. 5. Sliding Window Attention: The invention includes a Sliding Window Attention mechanism that enables the system to maintain relevant historical context over extended, multi-turn interactions. This mechanism allows the system to reference prior parts of a conversation while focusing on the current query, preserving coherence across complex dialogues. This feature is particularly advantageous in scenarios such as technical support and financial consultations, where context retention over multiple exchanges is essential for providing informed and coherent responses.
[0033] 6. Three-Tiered Enterprise Al Architecture: The system is designed with a three-tiered architecture comprising:
[0034] o Foundation Training: A foundational training layer that builds robust language understanding through selective training on high- quality datasets.
[0035] o Domain Specialization: A specialized training layer for industryspecific data, enabling the system to incorporate expert knowledge relevant to specific sectors such as finance, healthcare, and legal services.
[0036] o Enterprise Customization: A customization layer for fine-tuning on company-specific data, allowing the system to adapt to unique business workflows and domain-specific language, ensuring the model aligns with distinct operational requirements.
[0037] The inventive step of the present invention lies in the integration of these features, which collectively enable a highly efficient, contextually aware Al system capable of real-time, on-device NLP within resource-constrained environments.
[0038] This architecture provides a robust solution for enterprise applications, allowing the system to deliver responsive, secure, and contextually relevant Al interactions across various industries, with minimized dependency on external computational resources.
[0039] The present embodiment provides a method for implementing a robust, efficient Al system capable of on-device processing across resource-limited environments. By leveraging VGQA, RoPE, SwiGLU, DPO, and a sliding window attention mechanism, the Al system achieves real-time, contextually aware processing suitable for enterprise tasks that require continuity, relevance, and low-latency response times. This embodiment offers an innovative approach to Al processing, addressing technical challenges related to memory efficiency, user alignment, and resource constraints, making it an effective solution for diverse enterprise applications.
[0040] Embodiment of a Non-Transitory Computer-Readable Medium
[0041] A non-transitory computer-readable medium stores computer instructions which, when executed by one or more computing processors within an edge device, cause the one or more computing processors to perform an optimized process for realtime natural language processing (NLP) in resource-constrained environments. This non-transitory computer-readable medium provides an efficient, secure, and adaptable solution for deploying real-time NLP capabilities on edge devices. By storing instructions that implement VGQA, RoPE, SwiGLU, DPO, and sliding window attention, the medium equips edge devices with a method for delivering high-quality, contextually aware Al interactions across diverse applications, meeting the demands of enterprise environments while optimizing device resources.
[0042] The stored instructions implement the following steps:
[0043] 1. Configuring an Al processing module on the edge device to enable efficient NLP operations without reliance on external cloud infrastructure. This configuration includes initial setup for tokenization, embedding, and position encoding with minimal resource use.
[0044] 2. Tokenizing input data and embedding tokens with Rotary Positional Embeddings (RoPE), enabling the system to handle long text sequences while preserving positional relationships. This step allows for efficient context management across extended sequences without increasing the device’s computational load.
[0045] 3. Executing Variable Grouped Query Attention (VGQA) to dynamically group related queries within the attention computation process, reducing memory usage and computational demands. This instruction set enables logical context retention across multi-turn conversations, optimizing processing for applications like iterative customer support dialogues and financial advisory sessions.
[0046] 4. Applying Sliding Window Attention to maintain relevant historical context across multi -turn interactions, allowing the edge device to retain and reference prior parts of a conversation. This ensures coherent responses across extended conversations and is particularly useful in technical support and consultation scenarios.
[0047] 5. Processing each hidden state with SwiGLU activation in a feedforward layer to ensure stable gradient flow during training and inference. This activation function minimizes issues such as vanishing or exploding gradients, supporting consistent performance on devices with constrained hardware.
[0048] 6. Implementing Direct Preference Optimization (DPO) to align system outputs with user expectations based on ranked human feedback. This instruction set enables the system to provide contextually appropriate responses, especially in customer service and healthcare advisory applications, where both content accuracy and user alignment are essential.
[0049] 7. Generating final predictions through a Linear Output Layer followed by a Softmax function, producing contextually accurate and user-aligned responses suited for real-time applications on edge devices.
[0050] 8. Structuring the Al processing module with a three-tiered architecture, including:
[0051] o Foundation Training for building core language understanding, o Domain Specialization for industry-specific adaptation, and o Enterprise Customization for alignment with unique organizational workflows and data.
[0052] 9. Deploying the configured Al processing module on the edge device to allow secure, low-latency NLP processing directly on the device, minimizing privacy risks associated with cloud-based data transmission. The further embodiment of the invention provides a detailed overview of the innovative components and architectural choices that optimize the system for realtime, on-device natural language processing (NLP) in resource-constrained environments. Each component enhances the system’s efficiency, stability, and adaptability, contributing to its suitability for applications on edge devices.
[0053] The Tokenization and Embedding Layer with Rotary Positional Embeddings (RoPE) serves as a foundational component of the system. This layer is responsible for converting input text into tokenized segments and embedding these tokens with positional relationships preserved across extended sequences. Specifically, RoPE encodes token positions through a rotational transformation, enabling the system to manage long text sequences without adding to the computational load. By using RoPE, the system effectively maintains contextual continuity across lengthy or complex texts — such as legal documents or multi-turn dialogues — optimizing memory usage and processing power. This encoding technique ensures that essential positional context is preserved, making the system particularly suitable for applications requiring the precise interpretation of extended narratives or documents.
[0054] The further embodiment of the invention includes Variable Grouped Query Attention (VGQA), a mechanism that optimizes memory and computational efficiency during the attention process by dynamically grouping related queries. VGQA operates by consolidating multiple related queries into shared attention pathways, thereby reducing redundant computations and lowering memory usage. This feature enables the system to retain contextual flow across complex, multi -turn conversations, allowing it to process interactions more efficiently. For example, in a customer support scenario, VGQA allows the system to handle sequential queries related to a single issue, ensuring that logical continuity is maintained throughout the conversation. Similarly, in a financial consultation setting, VGQA can manage various client questions about portfolio performance, risk tolerance, and investment goals within the same conversation, maintaining a cohesive understanding of the context and minimizing computational load.
[0055] This approach makes VGQA particularly valuable for applications on resource-constrained edge devices, as it significantly enhances the system’s ability to deliver high-quality responses in real-time, without overloading device resources.
[0056] A further embodiment of the invention includes a Sliding Window Attention Mechanism, which allows the system to maintain relevant historical context across extended, multi -turn interactions by processing input in segmented "windows." This mechanism divides the conversation or document into smaller, manageable chunks, or windows, while retaining information from previous segments as it progresses through new input.
[0057] The Sliding Window Attention Mechanism ensures that the system can "look back" at earlier parts of a conversation or text without needing to hold the entire interaction in memory. This approach enables the system to keep track of critical details from prior exchanges, making it particularly valuable in scenarios that require context retention over multiple turns. For instance:
[0058] • Technical Support: In a troubleshooting session, the system can refer to earlier problem descriptions or previous solutions provided to the user, allowing it to build on prior information and deliver accurate responses based on the cumulative context of the conversation.
[0059] • Financial or Legal Consultations: During detailed advisory sessions, the system can recall past client questions or previously discussed topics, ensuring coherence and relevance throughout the interaction, even when it spans multiple segments.
[0060] By implementing the Sliding Window Attention Mechanism, the system effectively handles long-form interactions without excessive memory strain, making it ideally suited for deployment on edge devices with limited computational resources. This design allows for both continuity and efficiency, optimizing the system’s ability to respond in real-time while managing device constraints.
[0061] A further embodiment of the invention also includes Feedforward Processing with SwiGLU Activation. This component involves passing each hidden state through a feedforward layer using SwiGLU (Switching Gated Linear Units), an activation function that stabilizes the gradient flow during both training and inference. SwiGLU minimizes the risks of vanishing or exploding gradients, which ensures that the system operates reliably, even in high-demand environments. The SwiGLU activation function is particularly advantageous for applications requiring consistent output quality, such as: • Real-Time Advisory Services: In customer service or healthcare, the system must provide stable, accurate responses without fluctuations in performance.
[0062] • High-Demand Environments: The activation ensures the system remains responsive under varying computational loads, supporting reliable interactions across applications where output consistency is critical.
[0063] A further embodiment of the invention includes Direct Preference Optimization (DPO), a method for aligning the system’s outputs with user expectations by leveraging ranked human feedback. DPO allows the system to optimize responses according to specific user preferences without requiring complex reinforcement learning models. This approach adjusts both the relevance and appropriateness of responses, enhancing user alignment.
[0064] DPO is particularly valuable in applications where response quality and user context are crucial, such as:
[0065] • Customer Support: In contexts where tone, clarity, and contextual relevance matter, DPO ensures that the system’s responses meet user expectations.
[0066] • Healthcare Consultations: In patient interactions, DPO aligns system responses with medical context and patient history, allowing for a personalized, sensitive approach.
[0067] A further embodiment of the invention also incorporates a Three-Tiered Enterprise Al Architecture, which optimizes the system for general, specialized, and customized enterprise applications. This architecture is structured as follows:
[0068] • Foundation Training: This base layer involves training the system on high- quality datasets, building a comprehensive language understanding and establishing robust foundational knowledge.
[0069] • Domain Specialization: This layer focuses on industry-specific data, enabling the system to adapt to sectors such as finance, healthcare, or legal services. By training on relevant data, the system gains a deep understanding of specialized terminology and context.
[0070] • Enterprise Customization: This final layer involves fine-tuning on company-specific data, allowing the system to adapt to unique workflows, terminology, and requirements of individual organizations.
[0071] This three-tiered architecture enables the system to perform effectively across general-purpose and domain-specific tasks, making it versatile for enterprise use in various industries, where customized responses and deep understanding are essential.
[0072] Each of these components in the further embodiment enables the system to deliver robust, efficient, and context-aware Al processing on resource-limited edge devices. By focusing on memory efficiency, context retention, user alignment, and adaptability, the system is ideally suited for real-world applications across diverse industries.
[0073] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of embodiments of the invention as claimed.
[0074] BRIEF DESCRIPTION OF THE DRAWINGS
[0075] The drawings illustrate the design and utility of various embodiments of the invention. It should be noted that the figures are not drawn to scale and that elements of similar structures or functions are represented by like reference numerals throughout the figures. In order to better appreciate how to obtain the above-recited and other advantages and objects of various embodiments of the invention, a more detailed description of the present inventions briefly described above will be rendered by reference to specific embodiments thereof, which are illustrated in the accompanying drawings.
[0076] Understanding that these drawings depict only typical embodiments of the invention and are not therefore to be considered limiting of its scope, the invention will be described and explained with additional specificity and detail through the use of the accompanying drawings, in which:
[0077] FIG. 1 is a flowchart depicting an example edge Al processing system, illustrating the stages of tokenization, attention mechanisms, and output generation in resource-constrained environments, according to various embodiments of the present disclosure.
[0078] These figures provide an overview of the system’s architecture, showcasing how the integrated components contribute to efficient, real-time Al processing on edge devices.
[0079] Other aspects and advantages of the present invention will become apparent upon consideration of the following detailed description, wherein similar structures are identified by like or similar reference numerals to ensure clarity throughout the accompanying figures.
[0080] DETAILED DESCRIPTION
[0081] In the following description, numerous specific details are provided to ensure a thorough understanding of the present invention. However, it should be noted that the present invention may be practiced without these specific details. In some cases, well-known structures and devices are depicted in block diagram form to avoid unnecessarily obscuring the invention’s key features.
[0082] Embodiments of the present invention are described herein according to the following outline:
[0083] 1.0 General Overview
[0084] The present invention provides a system and method for implementing real-time natural language processing (NLP) on resource-constrained edge devices, thereby overcoming the limitations of conventional large language models (LLMs) that depend heavily on remote cloud infrastructure. Unlike existing cloud-hosted Al models that require substantial bandwidth, high computational power, and centralized data handling, the proposed invention performs on-device inference, allowing Al applications to operate locally, with low latency, high data privacy, and reduced energy consumption. This technological shift enables enterprises and individuals to access intelligent conversational systems, translation tools, and context-aware decision engines without the need for persistent internet connectivity or data off-loading to third-party servers.
[0085] 1.1 Technical Background and Challenges
[0086] Traditional NLP architectures — such as transformer-based large language models — are computationally intensive, often exceeding the processing and memory limits of embedded or mobile devices. These models depend on large-scale parallel processing units (GPUs or TPUs) for real-time inference, which leads to increased latency, energy costs, and potential privacy concerns due to remote data transfer.
[0087] Moreover, such architectures fail to sustain performance across multilingual and low-resource languages, and their reliance on cloud access restricts adoption in regulated industries such as finance, healthcare, and legal services, where data confidentiality and compliance are paramount.
[0088] 1.2 Overview of the Invention
[0089] To address these challenges, the invention introduces a lightweight and modular Al system, optimized for real-time, on-device NLP, integrating novel architectural components and training optimizations that together provide measurable technical improvements in speed, memory efficiency, and contextual accuracy.
[0090] The system’s design enables efficient execution of deep-learning operations within limited hardware constraints, providing faster token generation, context continuity, and multilingual adaptability — without depending on external servers. 1.3 Core Architectural Components
[0091] The proposed invention, the system integrates several inter-operable components that collectively deliver the claimed technical effect:
[0092] 1. Tokenization and Embedding Layer with Rotary Positional Embeddings (RoPE):
[0093] Encodes positional information using rotational transformations, allowing the model to retain long-term dependencies without increasing computational cost. This mechanism enables efficient processing of long sequences in documents or conversations on small memory footprints. 2. Variable Grouped Query Attention (VGQA):
[0094] A novel attention mechanism that dynamically clusters query vectors into groups based on semantic similarity, thereby reducing redundant matrix multiplications typical in standard attention computation. This grouping reduces memory bandwidth utilization and enhances inference throughput.
[0095] 3. Sliding Window Attention (SWA):
[0096] Enables contextual processing of long text inputs by dividing them into smaller overlapping windows. This design maintains context over extended dialogues or documents while constraining the active attention span, yielding low-latency performance and predictable memory usage — critical for edge devices.
[0097] 4. Feedforward Layer with SwiGLU Activation:
[0098] Employs the Switching Gated Linear Unit (SwiGLU) activation function to stabilize gradient propagation, minimize vanishing or exploding gradients, and maintain numerical stability across varying precisions (INT8 / INT4). This ensures consistent performance across diverse hardware, from mobile processors to embedded NPUs. 5. Direct Preference Optimization (DPO):
[0099] Integrates human feedback in model fine-tuning, aligning responses with user expectations while preserving computational efficiency. DPO adjusts model weights to prioritize contextually appropriate and human-aligned outputs without introducing inference-time overhead.
[0100] 6. Three-Tiered Enterprise Al Architecture:
[0101] The invention organizes learning and deployment into three scalable tiers:
[0102] o Foundation Training: broad language understanding and general reasoning;
[0103] o Domain Specialization: adaptation to industry-specific corpora (finance, law, medicine, etc.);
[0104] o Enterprise Customization: fine-tuning on proprietary or local datasets, ensuring contextual relevance and brand-specific accuracy. This multi-tier structure allows progressive optimization while maintaining edge-based efficiency.
[0105] 1.4 Technical Effects and Advantages
[0106] The coordinated operation of these components provides multiple measurable technical effects, including:
[0107] • Reduced computational latency through efficient attention mechanisms (VGQA and SWA).
[0108] • Improved memory efficiency by limiting active attention context and leveraging quantized operations.
[0109] • Enhanced numerical stability via SwiGLU activation across mixed- precision arithmetic.
[0110] • Hardware adaptability enabling deployment across CPUs, GPUs, and NPUs with minimal configuration. • Multilingual capability, supporting English and major Indic languages through adaptive positional encoding and cross-lingual transfer learning.
[0111] • On-device privacy preservation, as data is processed locally without transmission to remote servers.
[0112] These cumulative effects represent a tangible technical advancement over existing transformer-based architectures that rely on cloud-scale infrastructure.
[0113] The system’s ability to execute intelligent reasoning under constrained hardware conditions establishes a practical and industrially significant advancement in the field of Edge Al-based Natural Language Processing.
[0114] 2.0 Structural Overview:
[0115] FIG. 1 is an illustrative view of various aspects of an example system 100 in which the techniques described herein may be practiced, according to an embodiment. System 100 comprises one or more computing devices. The one or more computing devices comprise any combination of hardware and software configured to implement the various logical components described herein, including components such as an Al Processing Module, Attention Mechanism Module, Embedding Layer, Feedforward Layer, and Output Generation Module.
[0116] System 100 is designed to operate on resource-constrained edge devices, allowing for real-time natural language processing (NLP) capabilities without reliance on cloud infrastructure.
[0117] Each component within system 100 is configured to handle specific tasks, contributing to the overall efficiency, stability, and adaptability of the system.
[0118] • Al Processing Module: This module is responsible for managing the processing workflow of NLP tasks. It orchestrates the operations within each of the system’s primary components, coordinating tasks such as tokenization, embedding, attention, and output generation. • Embedding Layer with Rotary Positional Embeddings (RoPE): The embedding layer converts raw input data into tokenized segments, applying Rotary Positional Embeddings to retain positional relationships across tokens within extended sequences. RoPE enables efficient processing of long texts by encoding positions without additional computational load, crucial for handling multi-turn dialogues or complex document analysis.
[0119] • Attention Mechanism Module (including VGQA and Sliding Window Attention): This module comprises the Variable Grouped Query Attention (VGQA) and Sliding Window Attention mechanisms. VGQA dynamically groups related queries during attention computations, reducing memory usage and enhancing processing speed. Sliding Window Attention maintains relevant context across multiple segments, enabling the system to reference previous conversation turns or document sections without retaining unnecessary information.
[0120] • Feedforward Layer with SwiGLU Activation: This layer processes each hidden state using SwiGLU (Switching Gated Linear Units) activation. SwiGLU stabilizes gradient flow during training and inference, ensuring the system’s reliable performance across diverse computational loads, particularly in high-demand environments.
[0121] • Direct Preference Optimization (DPO) Module: This component finetunes the system’s outputs based on ranked human feedback, aligning responses with user expectations. DPO optimizes the relevance and tone of responses, making it highly suitable for applications requiring user-aligned output, such as customer service or healthcare advisory.
[0122] • Output Generation Module: The final component, the Output Generation Module, utilizes a Linear Output Layer followed by a Softmax function to produce the system’s predictions. This configuration enables the system to generate accurate, contextually appropriate responses efficiently. System 100 is configured to enable each module to function in harmony, providing a versatile and efficient NLP solution that is deployable on edge devices. This structural overview illustrates how each component within the system plays a specific role in achieving real-time processing capabilities while optimizing memory, computational load, and context retention.
[0123] The structural overview of the present invention outlines each major component and its role in optimizing the system for real-time, on-device NLP. The structure includes the following:
[0124] • Tokenization and Embedding Layer with Rotary Positional Embeddings (RoPE):
[0125] o This layer is responsible for converting input text into tokenized segments and embedding these tokens while preserving positional relationships. RoPE enables the system to efficiently handle long text sequences without increased computational demand, making it suitable for extended documents and multi-turn dialogues.
[0126] • Variable Grouped Query Attention (VGQA):
[0127] o VGQA is a memory-efficient attention mechanism that dynamically groups related queries during the attention computation process, reducing memory usage and computational load. This feature ensures that logical flow and context are maintained across multiturn conversations, ideal for applications requiring continuity over multiple exchanges.
[0128] • Sliding Window Attention Mechanism:
[0129] o The Sliding Window Attention mechanism processes conversation or text in manageable windows, retaining key information from previous segments. This design enables the system to maintain relevant historical context across extended interactions, such as technical support and advisory sessions, without overwhelming memory resources.
[0130] • Feedforward Processing with SwiGLU Activation:
[0131] o Each hidden state in the model undergoes processing in a feedforward layer activated by SwiGLU (Switching Gated Linear Units), ensuring stable gradient flow during both training and inference. SwiGLU reduces the risk of gradient-related issues, providing consistent, high-quality output even in high-demand environments.
[0132] • Direct Preference Optimization (DPO):
[0133] o DPO is a method of fine-tuning model responses to align with user expectations, leveraging ranked human feedback for optimization. This feature is particularly beneficial in applications where response quality, tone, and relevance to user preferences are critical, such as customer service and healthcare.
[0134] • Three-Tiered Enterprise Al Architecture:
[0135] o This architecture includes three layers: Foundation Training, Domain Specialization, and Enterprise Customization. This tiered approach allows the system to perform effectively across general- purpose and specialized tasks, adapting to unique business needs in industries such as finance, legal, and healthcare.
[0136] Each component in the system’s structure has been optimized for performance in edge environments, providing enterprises with a versatile, reliable, and resourceefficient NLP solution. The system’s architecture supports real-time processing, robust context retention, and industry-specific adaptation, addressing the growing demand for on-device Al solutions in resource-limited applications.
[0137] 3.0 Example Embodiments Examples of some embodiments are represented, without limitation, in the following paragraphs:
[0138] FIG. 1 illustrates an example process flow that may be implemented by one or more computing devices, such as edge devices or resource-constrained platforms, to achieve efficient natural language processing (NLP) directly on-device. This process flow demonstrates how various components of the system work in conjunction to deliver real-time, contextually aware responses without dependence on external cloud resources.
[0139] Example 1: Customer Support Interaction
[0140] In one exemplary embodiment, the system is deployed within a customer support environment, operating directly on an edge device or a localized enterprise server. The objective of this embodiment is to enable real-time, contextually coherent communication between users and an automated service assistant, without reliance on cloud-based NLP engines.
[0141] System Setup and Operation
[0142] A user initiates a service query on a mobile device or local kiosk (e.g., “My bill shows an extra charge,” followed by “Can I change my payment method?”). The system tokenizes each incoming message through the Embedding Layer with Rotary Positional Embeddings (RoPE), which preserves the sequential order of the dialogue and encodes positional information compactly.
[0143] Once tokenized, the queries are processed by the Attention Mechanism Module, which employs the Variable Grouped Query Attention (VGQA) algorithm. Unlike conventional self-attention, which computes attention weights across the entire token set, VGQA dynamically groups related queries based on semantic correlation. For example, questions relating to billing are grouped together, while technical-support topics are placed in a separate attention group.
[0144] This dynamic grouping drastically reduces the number of pairwise query-key computations required for inference. As a result, the system achieves a technical effect of lowering memory bandwidth usage and computation latency by minimizing redundant attention calculations, while still maintaining continuity across multi-turn conversations. The grouping process is hardware-adaptive, automatically adjusting group sizes depending on available cache or RAM resources, thereby ensuring consistent performance on edge processors with limited memory (for instance, on ARM A76 or mobile NPUs).
[0145] The Sliding Window Attention mechanism concurrently ensures that the conversational context remains accessible even when the dialogue exceeds the active attention span. Older but relevant segments of the conversation are retained within an overlapping sliding window, enabling the model to recall previous responses (e.g., reference to a “billing issue” from earlier in the chat) without overloading system memory.
[0146] Response Generation and Human Alignment
[0147] After the attention computation, the intermediate embeddings are passed to the Feedforward Layer with SwiGLU Activation, which stabilizes the gradient flow during real-time inference. The stabilized outputs are then processed through the Direct Preference Optimization (DPO) module.
[0148] DPO is responsible for aligning the model’s generated responses with humanpreferred communication patterns. Using ranked feedback data collected during training, DPO ensures that the assistant responds in a tone and structure that reflect human-like politeness, empathy, and clarity while maintaining technical correctness.
[0149] For example, when the user expresses dissatisfaction (“This fee seems unfair”), the DPO-optimized response is formulated as:
[0150] “I understand your concern. Let me check the charge details for you.”
[0151] This output demonstrates emotional alignment without introducing additional inference-time computation, since DPO optimization occurs within the model’s learned parameters rather than as an external re-ranking process. The Output Generation Module then converts the refined logits into text tokens through a linear + softmax layer, and the Al Processing Module streams the response back to the user interface with minimal latency.
[0152] Measured Technical Improvements
[0153] Empirical benchmarking of this embodiment demonstrates measurable improvements in computational efficiency and interaction quality:
[0154]
[0155] These results confirm that the combined use of VGQA and Sliding Window Attention yields a quantifiable technical effect — namely, improved throughput and memory utilization — while DPO ensures human-aligned communication without additional computational overhead.
[0156] Technical Effects and Advantages
[0157] 1. Reduced Latency:
[0158] VGQA minimizes redundant matrix multiplications, enabling faster token generation and real-time conversational flow even on hardware with limited processing capacity. 2. Improved Context Continuity:
[0159] Sliding Window Attention preserves essential dialogue history within constrained memory, avoiding the context-loss problem typical of truncated transformers.
[0160] 3. Human-Aligned Output Optimization:
[0161] DPO ensures that generated responses match user intent and sentiment while remaining computationally efficient, resulting in more natural and context-sensitive interactions.
[0162] 4. Hardware Adaptability:
[0163] The grouping and attention parameters automatically adjust to the compute resources of the device, allowing smooth operation across CPUs, GPUs, and NPUs without model retraining.
[0164] 5. Enhanced Data Privacy:
[0165] All user interactions and inference computations occur locally, ensuring that sensitive customer data remains confined to the device, compliant with enterprise privacy standards.
[0166] This embodiment is directly applicable in telecommunication support, banking help-desk automation, utility billing systems, and customer service kiosks in offline or low-connectivity environments. Enterprises can deploy this system on their internal hardware infrastructure, thereby achieving faster, secure, and cost-efficient service automation without cloud dependency.
[0167] Through the synergistic integration of Variable Grouped Query Attention (VGQA) and Direct Preference Optimization (DPO), the present invention provides a technically superior Al framework that delivers real-time, contextually accurate, and human-aligned conversational responses on resource-limited edge hardware. This embodiment exemplifies a concrete technical advancement over prior NLP systems by achieving tangible gains in latency reduction, context retention, and computational efficiency while maintaining compliance with data-privacy requirements.
[0168] Example 2: Virtual Healthcare Consultation
[0169] In another exemplary embodiment, the system of the present invention is implemented within a virtual healthcare consultation platform deployed on an edge-based medical device, tablet, or on-premise hospital server. The embodiment enables secure, real-time, and context-aware interaction between a patient and an automated healthcare assistant, providing medically relevant responses without relying on cloud-hosted inference models.
[0170] System Setup and Workflow
[0171] A patient interacts with the system through a voice- or text-based interface, inputting health-related queries such as:
[0172] “I have had a persistent cough for three days, ”
[0173] “I was prescribed amoxicillin last week — can I continue taking it? ”
[0174] The incoming data is tokenized through the Tokenization and Embedding Layer with Rotary Positional Embeddings (RoPE), which encodes both semantic and positional information of the patient’s dialogue. RoPE ensures that the system maintains contextual order across multiple conversational turns, crucial when a patient refers to previous symptoms or medication details.
[0175] The Attention Mechanism Module, powered by Sliding Window Attention (SWA) and Variable Grouped Query Attention (VGQA), processes the dialogue history efficiently.
[0176] • SWA divides the extended patient conversation into smaller, overlapping windows, each containing only the most relevant preceding context. This mechanism ensures that clinically important details — such as medication names, durations, or symptom patterns — remain accessible even when the total dialogue exceeds typical transformer limits. . VGQA groups medically related tokens (e.g., “fever,” “temperature,” “antibiotic”) into dynamic clusters, reducing redundant attention computations. This lowers inference latency while ensuring domainspecific coherence, allowing the model to relate prior health complaints to current conditions.
[0177] The processed embeddings are passed to the Feedforward Layer with SwiGLU Activation, which stabilizes inference across hardware configurations (CPU / NPU / GPU) by minimizing gradient oscillations and numerical instability. SwiGLU ensures consistent output confidence and avoids degradation in response quality even under constrained precision formats (INT8 or INT4) typical of medical edge hardware.
[0178] Response Generation and Clinical Alignment
[0179] The intermediate outputs are fine-tuned through the Direct Preference Optimization (DPO) module. In this embodiment, DPO aligns the model’s response tone and content with verified medical communication guidelines. For example, if a patient asks,
[0180] “Should I stop taking my antibiotics now that I feel better?” the DPO-optimized response generated is: “It’s important to complete the full antibiotic course as prescribed, even if your symptoms improve. Would you like me to explain why this helps prevent reinfection?”
[0181] This behavior ensures the assistant provides medically safe and empathetic guidance, reflecting human-like communication standards while minimizing computational overhead.
[0182] The final prediction is produced by the Output Generation Module, which performs linear transformation and Softmax decoding. The Al Processing Module streams the output with sub-second latency, allowing near real-time doctor-patient style interaction. Technical Evaluation and Performance Metrics
[0183] To validate this embodiment, the system was tested under realistic healthcare dialogue conditions using quantized and full-precision configurations. The following results demonstrate measurable technical improvements and domain reliability:
[0184]
[0185] These benchmarks clearly demonstrate that the invention achieves a technical effect of:
[0186] • Reduced latency and memory usage, due to VGQA and SWA optimization;
[0187] • Enhanced accuracy in clinical context retention, owing to RoPE and sliding-window persistence; and • Energy efficiency, enabling longer battery operation on portable medical devices.
[0188] Data Privacy and Compliance Advantage
[0189] All patient data and inference computations are performed locally on the edge device. Sensitive medical information such as symptoms, prescriptions, and test results never leave the device boundary. This local processing ensures compliance with healthcare data-protection frameworks such as HIPAA, GDPR, and India’s Digital Personal Data Protection Act (DPDP 2023).
[0190] By eliminating dependence on cloud-based inference, the invention provides a technical advancement in secure Al deployment, ensuring that the privacy of the patient is maintained while still delivering real-time diagnostic assistance.
[0191] Technical Effects and Advantages
[0192] 1. Low-Latency Clinical Response:
[0193] Real-time token generation enabled by VGQA and quantized inference allows interactive medical consultations with minimal delay.
[0194] 2. Efficient Context Retention:
[0195] Sliding Window Attention ensures continuity of medical dialogue, retaining patient history and prior recommendations.
[0196] 3. Energy-Efficient Operation:
[0197] Quantized processing reduces power draw by more than 50 % compared to full-precision inference, extending portable device usage.
[0198] 4. Improved Domain Accuracy:
[0199] Enhanced understanding of symptom-linked terms results in more precise and clinically relevant responses.
[0200] 5. Privacy-Preserving Architecture: All computations remain localized, meeting strict regulatory standards for healthcare data protection.
[0201] This embodiment is applicable to:
[0202] • Tele-health consultation kiosks deployed in remote or rural clinics with limited connectivity;
[0203] • Hospital bedside tablets or medical record assistants for nurses and physicians;
[0204] • Home-health monitoring systems integrated with loT-based diagnostic devices; and
[0205] • Medical education tools requiring multilingual, offline Al interaction. Through these applications, the system demonstrates an industrial and technical contribution by enabling intelligent, multilingual, and regulation-compliant medical assistance directly on edge hardware.
[0206] Conclusion
[0207] This embodiment illustrates how the integration of VGQA, Sliding Window Attention, RoPE, and DPO modules achieves a measurable technical effect of low-latency, high-accuracy, and privacy -preserving NLP for the healthcare sector. The system provides a concrete technical advancement over existing cloud-based consultation systems by combining real-time inference efficiency with strict data-security compliance, thus fulfilling the requirements of industrial applicability and technical contribution under Indian Patent Law.
[0208] Example 3: Legal Document Review
[0209] In another exemplary embodiment, the system of the present invention is deployed in a legal document analysis environment, such as an enterprise legal-tech platform, an advocate’s on-premise server, or an in-house corporate compliance terminal.
[0210] The embodiment enables efficient review, summarization, and clause identification within lengthy legal contracts or statutes, directly on local hardware, without reliance on cloud computation.
[0211] System Context and Operation
[0212] A legal professional uploads or scans a large contract — often exceeding 50 pages — to the system. The document is ingested and segmented by the Tokenization and Embedding Layer with Rotary Positional Embeddings (RoPE).
[0213] Unlike conventional transformer embeddings that rely on absolute or sinusoidal position encodings, RoPE represents positional data through rotary transformations of embedding vectors, preserving both relative order and semantic proximity.
[0214] This approach allows the model to track cross-references — for example, linking a “termination clause in Section 12” to its conditions defined in “Section 7(a)” — even across long contextual distances, without quadratic memory growth.
[0215] The embedded tokens are processed by the Attention Mechanism Module, which integrates both Variable Grouped Query Attention (VGQA) and Sliding Window Attention (SWA).
[0216] • VGQA dynamically clusters queries that share semantic or syntactic similarity (e.g., “payment,” “remuneration,” “consideration”) into grouped attention heads. This optimization reduces redundant query-key computations, thereby lowering memory load while maintaining interclause correlation.
[0217] • SWA enables the system to analyze documents section by section in overlapping windows (for instance, 512 tokens per window with 25 % overlap), ensuring continuity across adjoining paragraphs while preventing RAM overflow.
[0218] Together, these mechanisms create a hardware-efficient document-analysis pipeline, capable of executing on edge CPUs or GPUs with as little as 6-8 GB RAM
[0219] Clause Identification and Reasoning Process As each document segment is analyzed, the Feedforward Layer with SwiGLU Activation stabilizes inference by preventing gradient underflow or overflow during batch processing. The system learns and applies contextual associations — for example, recognizing that “the indemnifying party” in later clauses refers to the “Supplier” defined earlier.
[0220] During inference, the Direct Preference Optimization (DPO) module aligns the output with human drafting preferences and legal phrasing standards. When the model identifies ambiguous or risky clauses, it generates human-aligned explanations such as:
[0221] “Clause 9.2 grants unilateral termination rights to the Licensee without reciprocal provision to the Licensor. This may create contractual imbalance.”
[0222] The Output Generation Module converts the processed embeddings into structured textual summaries, key-risk highlights, or compliance tables, which can be displayed directly on-screen or exported as part of an automated due-diligence report.
[0223] Technical Evaluation and Benchmarking
[0224] To validate this embodiment, comparative testing was performed between a conventional document-analysis transformer and the proposed architecture executed on an edge GPU (NVIDIA T600).
[0225]
[0226]
[0227] These results confirm that the architectural improvements provide measurable technical effects in terms of throughput, memory efficiency, and contextual comprehension.
[0228] Specifically, RoPE improves positional coherence, VGQA reduces redundant computations, and SWA allows scalable analysis of extended text sequences without hardware bottleneck.
[0229] Security and Privacy Compliance
[0230] All document parsing and inference computations are performed locally on the organization’s internal hardware. Since legal contracts often contain privileged or confidential information, no external data transmission occurs at any stage of processing. This local execution model not only ensures data confidentiality but also satisfies regulatory compliance with data-sovereignty requirements in legal practice.
[0231] Technical Effects and Advantages
[0232] 1. Enhanced Document Length Handling:
[0233] Capable of analyzing documents over 50 pages without context truncation, due to Sliding Window Attention.
[0234] 2. Reduced Memory and Computation Load:
[0235] VGQA dynamically minimizes redundant attention computations, cutting GPU memory use by ~50 %.
[0236] 3. Improved Cross-Clause Reasoning: RoPE enables accurate referencing between dispersed legal sections, maintaining interpretive consistency.
[0237] 4. Stable Edge Performance:
[0238] SwiGLU activation ensures reliable inference across mixed-precision arithmetic on low-resource hardware.
[0239] 5. Localized Processing and Data Security:
[0240] On-premise operation guarantees confidentiality of sensitive legal materials, removing cloud dependency.
[0241] 6. Human-Aligned Legal Summarization:
[0242] DPO fine-tunes responses to mirror professional legal tone and structure, improving practical usability.
[0243] This embodiment is applicable to:
[0244] • Law Firms and Corporate Legal Departments - automated contract review, due-diligence, and risk assessment.
[0245] • Government and Regulatory Bodies - policy-document analysis and multilingual statute comparison.
[0246] • Compliance Management Systems - automated audit of standard operating procedures.
[0247] • Legal Tech Startups - deployment on desktop or mobile devices for realtime client-contract review.
[0248] By executing advanced legal NLP operations directly on local devices, the invention enables faster, privacy-preserving, and resource-efficient legal analytics, representing a substantial technical advancement over conventional cloud-based review tools.
[0249] This embodiment demonstrates the synergy of RoPE, VGQA, and SWA in achieving a quantifiable technical effect — namely, improved memory efficiency, reduced processing latency, and superior contextual comprehension in long-document analysis.
[0250] The invention thus delivers a hardware-optimized, secure, and human-aligned legal NLP framework, marking a significant industrial and technical contribution in the field of Edge Al for Legal Applications.
[0251] Example 4: Financial Advisory Sessions
[0252] In financial advisory scenarios, the Three-Tiered Enterprise Al Architecture is utilized, with foundational training providing general language comprehension, domain specialization tailoring responses to financial concepts, and enterprise customization aligning the system with specific firm terminology or strategies. As the client navigates topics such as portfolio performance and market conditions, VGQA facilitates logical continuity across questions, while Sliding Window Attention enables the system to track previous investment recommendations and risk assessments for a cohesive advisory experience.
[0253] These examples illustrate the flexibility and real-world applicability of the system across diverse use cases, highlighting the innovative ways in which each component works to ensure efficient, responsive, and contextually aware NLP in resourcelimited environments.
[0254] Example 5: Quantized Edge Deployment (Q4_KM)
[0255] In another embodiment, the system operates as a quantized model (4-bit Q4_KM variant), which significantly reduces model size, computational cost, and energy consumption while retaining high reasoning accuracy. Performance benchmarks confirm that quantization preserves over 95% accuracy of the full-precision model while enabling real-time inference on devices with as little as 6-8 GB RAM.
[0256]
[0257]
[0258] This embodiment achieves the technical effect of reduced latency and power consumption, confirming suitability for low-power edge deployment.
[0259] Example 6: Multilingual and Cross-Lingual Processing
[0260] The multilingual embodiment enables NLP in nine Indic languages, including Hindi, Tamil, Telugu, Bengali, Marathi, Malayalam, Kannada, Gujarati, and Odia. Benchmarks show average normalized accuracies of 28.1 % (ARC-Indic) and 32.6 % (MMLU-Indic), demonstrating robust multilingual reasoning and knowledge transfer from English to low-resource languages.
[0261]
[0262] This capability ensures inclusive access to Al technologies across multilingual enterprise and government domains.
[0263] Example 7: Throughput and Hardware Scalability
[0264] Testing across multiple platforms demonstrates consistent, high-speed inference:
[0265]
[0266]
[0267] These results establish the invention’s scalability and hardware adaptability, achieving measurable technical advancement in throughput, energy efficiency, and latency control across mobile, embedded, and data-center devices.
[0268] 4.0 Implementation Mechanism — Hardware Overview
[0269] The present invention can be implemented on various hardware configurations, particularly on resource-constrained edge devices. The system is designed to be highly adaptable, enabling it to function efficiently on devices with limited memory, processing power, and energy resources, such as loT devices, mobile platforms, and compact edge computing setups.
[0270] • Processing Units: The system is compatible with CPU and GPU setups as well as specialized Al processors, such as neural processing units (NPUs), which are increasingly available in modem edge devices.
[0271] • Memory Optimization: Through the use of memory-efficient mechanisms like VGQA and Sliding Window Attention, the system minimizes memory demands, allowing it to operate smoothly within the constraints of edge device hardware.
[0272] • Storage: The system can operate with limited storage, utilizing a compact, non-transitory computer-readable medium that stores the instructions necessary for NLP processing.
[0273] • Power Efficiency: By minimizing unnecessary computations, the system’s architectural components (e.g., SwiGLU Activation and VGQA) support power-efficient operation, making the invention suitable for batterydependent devices. The hardware implementation is structured to leverage the architecture’s efficiencies, enabling real-time processing on compact devices and expanding the accessibility of advanced NLP capabilities across various industrial contexts.
[0274] 5.0 Extensions and Alternatives
[0275] The present invention is adaptable to numerous variations and extensions to meet specific application needs or address additional technical requirements.
[0276] • Extended Domain-Specific Training: While the current system supports various industries, additional domain-specific training can be implemented to expand the model’s capabilities in other specialized areas, such as manufacturing, logistics, or education.
[0277] • Customizable Attention Mechanisms: Although VGQA and Sliding Window Attention are optimized for memory efficiency, alternative attention mechanisms or hybrid models could be incorporated to support unique interaction models or complex conversational requirements.
[0278] • Enhanced Privacy Features: In applications requiring strict data privacy, further encryption layers can be added to secure data within the on-device processing framework, particularly for sensitive industries like healthcare and finance.
[0279] • Integration with Edge Al Platforms: The system can be adapted for integration with edge Al platforms that offer federated learning capabilities, allowing models to continuously improve without centralizing data, which would enhance both the model’s accuracy and privacy compliance.
[0280] • Scalability for Cloud-Edge Hybrid Models: Although designed primarily for on-device processing, the invention can be scaled to operate in hybrid cloud-edge environments, where some data is processed locally and other tasks are handled via cloud support, providing additional flexibility for applications with variable latency or computational needs. These extensions and alternatives demonstrate the versatility of the present invention, enabling it to address evolving industry requirements while maintaining its core efficiency, adaptability, and real-time processing capabilities on resource-constrained devices.
[0281] In the drawings, the various components are depicted as being communicatively coupled to various other components by arrows. These arrows illustrate only certain examples of information flows between the components. Neither the direction of the arrows nor the lack of arrow lines between certain components should be interpreted as indicating the existence or absence of communication between the certain components themselves. Indeed, each component may feature a suitable communication interface by which the component may become communicatively coupled to other components as needed to accomplish any of the functions described herein.
[0282] In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Thus, the sole and exclusive indicator of what is the invention, and is intended by the applicants to be the invention, is the set of claims that issue from this application, in the specific form in which such claims issue, including any subsequent correction. In this regard, although specific claim dependencies are set out in the claims of this application, it is to be noted that the features of the dependent claims of this application may be combined as appropriate with the features of other dependent claims and with the features of the independent claims of this application, and not merely according to the specific dependencies recited in the set of claims. Moreover, although separate embodiments are discussed herein, any combination of embodiments and / or partial embodiments discussed herein may be combined to form further embodiments.
[0283] Any definitions expressly set forth herein for terms contained in such claims shall govern the meaning of such terms as used in the claims. Hence, no limitation, element, property, feature, advantage or attribute that is not expressly recited in a claim should limit the scope of such claim in any way. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.
Claims
We claim,1. A system for performing real-time natural language processing (NLP) on a resource-constrained computing device, the system comprising: a) an Al processing module configured to orchestrate tokenization, attention computation, feed-forward transformation, and response generation;b) a tokenization and embedding layer configured to convert an input text into token embeddings and apply Rotary Positional Embeddings (RoPE) for encoding positional relationships across tokens, thereby enabling context retention over extended sequences with reduced computational load;c) an attention mechanism module comprising:i) a Variable Grouped Query Attention (VG A) unit configured to dynamically cluster semantically related query vectors and reduce redundant attention computations, andii) a Sliding Window Attention (SWA) unit configured to process overlapping token segments while maintaining conversation continuity within a defined memory window;d) a feed-forward layer comprising a Switching Gated Linear Unit (SwiGLU) activation function configured to stabilize gradient propagation and maintain numerical stability across variable precision formats during inference; e) a direct preference optimization (DPO) module configured to align generated responses with user intent based on pre-ranked human feedback, thereby providing contextually appropriate, human-aligned outputs;f) a quantization module configured to represent model parameters in reduced-precision numeric formats (INT8, INT4) to minimize memory footprint, power consumption, and latency while preserving response accuracy;g) a multilingual processing layer configured to perform cross-lingual reasoning across a plurality of languages using adaptive positional embeddings and fine-tuned tokenization;h) a three-tier enterprise architecture comprising a foundation training layer, a domain specialization layer, and an enterprise customization layer configured to adapt the model to industry-specific datasets; andi) an output generation module configured to generate and transmit a contextually relevant response to a user interface through a softmax decoding function;wherein the modules (a)-(i) are implemented in hardware and software executable instructions stored on a non-transitory computer-readable medium and executed by one or more processing units of the computing device, whereby the system achieves the technical effects of:— reduced inference latency,— improved memory utilization,— energy-efficient real-time NLP processing,— multilingual reasoning capability, and— privacy-preserving on-device operation independent of cloud infrastructure.
2. The system as claimed in claim 1, wherein the VGQA unit groups queries dynamically based on semantic similarity scores computed through cosinedistance clustering, thereby reducing attention-matrix multiplications by at least 30 %.
3. The system as claimed in claim 1 or 2, wherein the Sliding Window Attention employs overlapping windows of token length N with stride S < N, ensuring contextual continuity while maintaining constant memory complexity O(N + S).
4. The system as claimed in any preceding claim, wherein the quantization module performs mixed-precision arithmetic combining 8-bit activations and 4-bit weights, yielding a reduction of memory usage by at least 50 % compared to FP16 inference.
5. The system as claimed in any preceding claim, wherein the SwiGLU activation maintains stable gradient propagation by adaptively gating linear units during training and inference, preventing vanishing gradients on resource-limited devices.
6. The system as claimed in any preceding claim, wherein the DPO module fine-tunes response alignment using reward functions derived from ranked human preference datasets, implemented without introducing runtime latency.
7. The system as claimed in any preceding claim, wherein the multilingual processing layer employs cross-lingual transfer learning such that the model trained primarily on English corpora attains reasoning accuracy within ±5 % across Indic languages including Hindi, Tamil, Telugu, Bengali, and Marathi.
8. The system as claimed in any preceding claim, wherein the three-tier enterprise architecture allows deployment-specific adaptation by updating only the enterprise-customization layer while maintaining stability of foundation and domain-specialization parameters.
9. The system as claimed in any preceding claim, wherein the system is configured for edge deployment on embedded processors, mobile chipsets, or loT modules operating within a power budget below 10 W.
10. The system as claimed in any preceding claim, wherein all inference computations and user-data processing are executed locally on the device to ensure data-privacy compliance with national or international data- protection regulations.
11. The system as claimed in any preceding claim, wherein the hardware execution of the modules (a)-(i) provides measurable throughput improvement of at least 2* compared to conventional transformer-based NLP architectures operating at similar precision levels.
12. A computer-implemented method for performing real-time natural language processing on a resource-constrained computing device, the method comprising the steps of: a) tokenizing an input text and embedding tokens with Rotary Positional Embeddings (RoPE); b) computing attention scores using Variable Grouped Query Attention (VGQA) and Sliding Window Attention (SWA); c) applying a feed-forward transformation with SwiGLU activation to generate stable intermediate representations; d) optimizing output alignment using Direct Preference Optimization (DPO);e) quantizing model parameters to a reduced precision format to decrease memory and power usage; f) performing multilingual reasoning through adaptive positional embeddings; and g) generating an output response via a softmax decoding process;wherein the above steps are executed by one or more processors of the device, thereby achieving a technical effect of enabling real-time, low-latency, and privacy-preserving NLP inference on edge hardware.
13. The method as claimed in claim 12, wherein steps (b) and (c) collectively reduce inference latency by at least 40 % relative to conventional transformer inference under equivalent hardware conditions.
14. The method as claimed in claim 12 or 13, wherein quantization in step (e) uses a Q4 KM scheme comprising per-channel symmetric weight quantization and activation scaling.
15. The method as claimed in any of claims 12-14, wherein all computations are performed locally within the memory of the device to prevent transmission of confidential data to external servers.
6. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a computing device, cause the processors to perform the method as claimed in any of claims 12-15.