Processing heterogeneous content in language models

The method enhances language models' ability to retrieve and understand information by generating and referencing text descriptions of specific content types within documents, addressing the limitations of current models in handling non-textual content.

WO2025125288A1PCT designated stage expired Publication Date: 2025-06-19IP MIND LTD

Patent Information

Application Number
PCT/EP2024/085604
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-13
Filing Date
2024-12-11
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current language models struggle with generalization, consistency, and specialization, leading to suboptimal performance in information retrieval and understanding, especially when dealing with non-textual content like formulas, tables, and technical drawings.

Method used

A method that involves traversing documents to identify specific types of content, generating text descriptions of this content, storing these descriptions for efficient retrieval, and referencing them during the information retrieval process to enhance understanding and analysis.

Benefits of technology

This approach significantly improves the accuracy and relevance of information retrieval, particularly in technical documents, by enabling language models to accurately interpret and understand complex formulas, tables, and figures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024085604_19062025_PF_FP_ABST
    Figure EP2024085604_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A method and system for enhancing information retrieval and understanding in language models when processing heterogeneous documents containing text, tables, figures, formulas, and images is described. The method involves traversing a document to identify specific types of content and using a language model to generate text descriptions of the identified content, including detailed interpretations. These descriptions are stored in a storage means, indexed by content type. During retrieval, the language model references these descriptions to enhance its understanding and analysis of the document, improving accuracy and contextual relevance. The system comprises a document traversing module, a language model, a storage means, and an information retrieval module, with a feedback loop for continuous improvement based on user feedback. This invention addresses limitations in current retrieval-augmented generation applications by improving the interpretation of complex technical content, thereby enhancing the performance of language models.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Processing Heterogeneous Content in Language Models

[0002] FIELD OF THE INVENTION

[0003] The present disclosure relates to artificial intelligence systems using large language models (LLMs). More particularly, but not exclusively, the present disclosure relates to a system and method for enhancing information retrieval and understanding in language models by generating and storing text descriptions of specific types of content, such as formulas, tables, figures, images, audio and video, and referencing these descriptions during the information retrieval process.

[0004] BACKGROUND OF THE INVENTION

[0005] With the advent and growth of artificial intelligence, especially deep learning, there has been a significant surge in the development and deployment of large-scale language models for various applications, including chatbots, assistants, and information retrieval systems. Notably, models such as GPT-2 and GPT-3 developed by OpenAI have demonstrated human-like text generation capabilities, making them especially valuable for answering user queries in various domains [Radford, A., et al. "Language Models are Unsupervised Multitask Learners." OpenAI, 2019],

[0006] Despite the advancements, current language models have their limitations:

[0007] 1. Generalization: Although large-scale models like GPT-3 are designed to be generalists, their vast knowledge bases often lead to answers that might lack specificity or nuanced understanding for domain-specific queries.

[0008] 2. Consistency and Reliability: Given the probabilistic nature of their underlying neural networks, language models can sometimes produce inconsistent or varying answers to similar or slightly altered queries.

[0009] 3. Lack of Specialization: While general models can answer a broad range of questions, they often cannot match the accuracy of models fine-tuned for specific domains. Yet, exclusively using domain-specific models may limit the breadth of topics the system can address.

[0010] For example, comparisons between general models and models fine-tuned for certain domains, such as BioGPT which has been fine-tuned on large scale biomedical literature, has shown that while finetuned models lack the versatility of general models they can outperform general models on domainspecific queries [Renqian, L., et al., "BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining", Briefings in Bioinformatics, 2022, arXiv:2210.10341],

[0011] In other studies, it has been shown that the model such as GPT-3 can provide inconsistent responses. For example, when posed with the same question multiple times, the answers may vary significantly in accuracy - e.g., see, Sheng, E., Chang, K-W., Natarajan, P., & Peng, N. (2019), "The Woman Worked as a Babysitter: On Biases in Language Generation", arXiv:1909.01326.

[0012] Another limitation is that current retrieval-augmented generation (RAG) applications struggle with understanding and accurately interpreting non-textual content such as formulas, tables, and technical drawings. This lack of understanding leads to suboptimal performance when such content is present in the documents being analyzed. An object of the present disclosure is to provide an improved system and method for optimizing information retrieval and understanding in language models by accurately identifying and interpreting specific types of content, such as formulas, tables, and figures.

[0013] SUMMARY OF THE INVENTION

[0014] The invention is set out in the appended claims. It will be understood that features of any aspect, example, or embodiment described herein may, wherever appropriate, be applied to any other aspect, example, or embodiment described herein. Where reference is made to different embodiments or sets of embodiments, it should be understood that these are not necessarily distinct but may overlap.

[0015] The present invention discloses a method for enhancing information retrieval and / or understanding in language models. The method comprises the following steps:

[0016] Traversing a Document: The method begins by traversing a document to identify specific types of content, such as formulas, tables, and figures. This identification is crucial for subsequent interpretation and analysis.

[0017] Generating Text Descriptions: Once the specific types of content are identified, a language model with interpretation capabilities is employed to generate text descriptions of the identified content. These text descriptions include detailed interpretations or effects of the content, providing a deeper understanding of the technical elements present in the document.

[0018] Storing Text Descriptions: The generated text descriptions are then stored in a storage means. This storage is indexed according to the corresponding content, ensuring that the descriptions can be efficiently retrieved during the information retrieval process.

[0019] Referencing the Storage Means: During the information retrieval process, when the language model encounters the specific types of content in the text of the document, it references the storage means. This step ensures that the language model has access to the accurate and contextually relevant interpretations of the identified content.

[0020] Enhancing Understanding and Analysis: By providing the language model with the generated text descriptions, the method enhances the language model's understanding and analysis of the document. This process improves the accuracy and relevance of information retrieval, particularly when dealing with technical documents containing complex formulas, tables, and figures.

[0021] This method addresses the limitations of current retrieval-augmented generation (RAG) applications, which often struggle with understanding and accurately interpreting formulas, tables, and technical drawings. By accurately identifying and interpreting these types of content, the invention significantly enhances the performance of language models in information retrieval and understanding, making them more effective and reliable in processing technical documents.

[0022] It will be understood that features of any aspect, example or embodiment described herein may, wherever appropriate, be applied to any other aspect, example or embodiment described herein. Where reference is made to different embodiments or sets of embodiments, it should be understood that these are not necessarily distinct but may overlap.

[0023] Detailed Description Certain preferred embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0024] Fig. 1 is a high-level diagram showing the components of the hierarchical language model system, including user interface, LLLMs, HLLM, and data flows.

[0025] Fig. 2 illustrates the different domain specialists as LLLMs with examples such as Medicine, Finance, Engineering and Legal.

[0026] Fig. 3 illustrates the structure of the neural network architecture and algorithms used within the higher-level generalist HLLM model.

[0027] Fig. 4 illustrates a possible flow from receiving user query to generating the final response.

[0028] Fig. 5 illustrates a synthesis step showing how segments from different LLLM responses can be combined.

[0029] Fig. 6 depicts how user ratings of responses are fed back to improve the HLLM over time.

[0030] Fig. 7 illustrates how inter-model dialogues can be used to obtain response improvement.

[0031] Fig. 8 demonstrates the diversity-based response improvement workflow.

[0032] Fig. 9 illustrates an architecture for enhancing information retrieval and / or understanding in language models.

[0033] Fig. 1 shows a high-level diagram of the overall architecture and components of the hierarchical language model system for generating optimal responses to user queries. With reference to fig. 1, a user interface (100) comprises the communication channel through which a user query (101) is received as input text, images, video and / or audio. This interface (100) could include components such as microphones, touchscreens, cameras, keyboards, speakers etc. It also conveys the final response (108) back to the user. Alternatively, the user interface could be implemented as an Application Programming Interface (API) such that it is possible for software applications to enter queries into the system and receive responses.

[0034] In Fig. 1, the user query (104) is sent to the higher-level language model (HLLM) module (103), which analyses it and transmits it, or instructions derived from it, to the lower-level language models (LLLMs). Multiple LLLM modules (102) act as domain specialists, each trained on data from a specific field and / or having specific attributes which make it more efficient for particular tasks. Example fields include medicine, engineering, finance, etc. Examples of specific attributes include the ability to process video or audio data, the ability to handle complex mathematical formulae etc. Four example LLLM modules are shown (102A, 102B, 102C and 102D), however any number of LLLM modules could exist in the system.

[0035] In context of the invention, the distinction between the Higher-Level Language Model (HLLM) and the Lower-Level Language Models (LLLMs) primarily revolves around their functional roles rather than their inherent complexity or specialization. While the LLLMs are typically specialized models, each adept in a specific domain or knowledge area, the HLLM operates in a supervisory capacity. It coordinates and integrates the outputs of the LLLMs, ensuring optimal response generation and system performance. Importantly, the HLLM is not necessarily more or less complex or sophisticated than the LLLMs. In some scenarios, the LLLMs might actually be different instances of the same model as the HLLM, with the key difference being their designated roles within the system. The HLLM's primary function is to orchestrate the overall process, manage the interaction among the LLLMs, and refine the final output, leveraging the specialized capabilities of each LLLM irrespective of their individual complexities or similarities to the HLLM.

[0036] Alternatively, the query be simultaneously propagated to all relevant LLLM modules (102) at the same time it is sent to the HLLM. This alternative is shown by the dashed lines in Figure 1. In some examples of the invention, the user query may be provided to all of the LLLMs, without considering which LLLM modules are relevant to subject-matter / domain of the user query.

[0037] The HLLM module (103) operates at an overall system level to analyse, select and synthesize the most appropriate response from the proposed responses generated by the LLLMs (102). The HLLM (103) implements algorithms detailed in this disclosure for identifying optimal responses based on relevance, accuracy, redundancy and / or other criteria.

[0038] A central controller / processor (not shown) handles coordination of the modules (102, 103) and data flows. It transmits the user query (101) to the HLLM (103) and the LLLMs (102), collects their responses (105), feeds them to the HLLM (103), obtains the final response (106), and propagates it back (108) through the user interface (100).

[0039] Database (107) may optionally act as a systems memory component, storing corpora and datasets for training the LLLMs (102), queries (101), the associated proposed responses (105A-105D), the associated optimal responses (106), interaction history for feedback learning by the HLLM (103), and other system data. Optimal responses may be sent directly from the HLLM (103) to the user interface (100) e.g. for interactive user access, such as in a chat or API access; or may be stored in the database (107). In the latter case, the user may interrogate the database (107) by means of the user interface (100). In both cases, the user interface (100) communicates the optimal responses to the user (108).

[0040] Finally, feedback such as the assessments of the quality or relevance of the optimum responses obtained from the user or other quality assessment systems could be transmitted to the HLLM (103) in order to improve the quality of future responses. This is depicted in Fig. 1 as an optional feedback path (109). In a specific example of the invention, the feedback may be obtained from external software applications by means of an API in the user interface.

[0041] In summary, Fig. 1 provides a broad overview of the key hardware and software modules, and the high-level workflow in the hierarchical language model system. The interfaces between components are intended to be adaptable to different implementations as per system requirements.

[0042] With reference to Fig. 1, a typical operational flow would be as follows:

[0043] 1. upon receiving a user query through an interface, the system transmits the query to the HLLM;

[0044] 2. the HLLM transmits the query or instructions derived from the query to one or more of the most relevant LLLMs; a. LLLMs may be determined as relevant based on conditions such as whether the domain or dataset associated with the LLLM corresponds to the domain or subjectmatter of the query; 3. In an alternative example of the invention, the query may be transmitted directly from the User Interface 100 to one or more of the LLLMs as well as the HLLM.

[0045] 4. Each relevant LLLM generates a proposed response based on its training and expertise. The proposed responses may then be compiled into a set of alternative answers.

[0046] 5. The HLLM receives the proposed responses directly from the LLLM(s). Alternatively, the proposed responses may be stored in the memory (107) and the HLLM may retrieve the proposed response(s) from there.

[0047] The analysis of the response(s) carried out by the HLLM is a multi-step one:

[0048] 1. Relevance Check: The HLLM evaluates each answer for its relevance to the original query, potentially assigning a score based on how well each answer addresses the user's query.

[0049] 2. Redundancy Removal: To ensure diverse perspectives, the HLLM identifies and removes or de-prioritizes redundant or highly similar answers.

[0050] 3. Answer Synthesis: Based on the analysed responses, the HLLM may choose to either select the best response outright, synthesize a new answer by integrating elements from multiple LLLM-generated answers (i.e. from one or more proposed responses); or create a further query for one or more of the LLLMs.

[0051] 4. Optimization: Before finalizing the answer, the HLLM may perform additional refinements, ensuring the response is coherent, grammatically correct, and styled appropriately for the intended audience or application.

[0052] Finally, the selected or synthesized answer is either transmitted back to the user through the User Interface 100 or stored in the Database 107.

[0053] During relevance analysis, the HLLM may employ attention mechanisms to compare question and answer embeddings, assigning relevance scores using cosine similarity. Answers with low relevance scores are filtered out.

[0054] For redundancy removal, fuzzy string matching could be used to identify semantic overlaps between answers. A threshold could determine when overlap is too high. Clustering algorithms such as k- means could also group semantically similar answers / proposed responses.

[0055] When synthesizing optimal responses, the HLLM may use a Sequence-to-Sequence model with pointer networks to fluidly combine important segments from different answers / proposed responses. The pointers would allow copying pieces of text directly from the inputs.

[0056] A beam search method could be used to generate multiple candidate combinations. The candidates then being ranked using a scaled dot-product attention mechanism that learns optimal ways to fuse answers. The highest scoring combination could then be selected as the synthesized response.

[0057] During optimization, the HLLM or a separate language model fine-tuned on the target domain may be used to paraphrase, correct grammar, and improve the overall coherence of the final response. Synthesis and optimization greatly improve relevance over merely selecting between proposed responses. The overall pipeline reduces repetitive and irrelevant responses compared to individual LLLMs.

[0058] Fig. 4 illustrates the step-by-step operational workflow of the hierarchical language model system from receiving a user query to generating the final response. The process includes: i. Query Input (400) - The user enters a query through the user interface which is received by the central controller. ii. Query Transmission (401) - The central controller forwards the query to the HLLM and optionally also directly to the LLLMs. ill. At step 402, the HLLM formulates a query or queries for the LLLMs. The query may be the same for all LLLMs or may be different and could, for example, be adapted to the specific attributes of respective LLLMs. The query is transmitted by the controller from the HLLM to the relevant LLLMs. iv. LLLM Response Generation (403) - Each LLLM independently analyses the query and generates a candidate / proposed response based on its domain-specific training. v. Compile Responses (404) - The controller compiles all LLLM generated responses into a collective set of candidate answers and transmits the candidate answers to HLLM. vi. HLLM Analysis (405) - The HLLM analyses each candidate response based on criteria such relevance, accuracy, redundancy. vii. Optimal Selection / Synthesis (406) - Using its algorithms, the HLLM selects the best response or synthesizes a new response from components. viii. Post-Processing (407) - The optimized response may undergo grammar / style refinement. ix. Output Response (408) - The final response is either sent back to the user interface to be accessed immediately by the user and / or stored in the database (107) for later access by the user. x. User Feedback (409 - 411) - The user can optionally provide feedback ratings on the response, enabling the HLLM to improve over time. xi. Optionally, the HLLM can be refined in real-time so as to implement improvements based on the feedback. The HLLM could then repeat the LLLM queries and / or the response analysis and the selection and / or synthesis until the optimized response(s) satisfies specific quality metrics. This real-time feedback loop is shown as dashed line (412) in fig. 4.

[0059] In summary, Fig. 4 maps out the operational workflow and key steps in generating the system's responses to user queries. The process leverages both lower-level domain-specific models and higher-level general models to produce an optimized output.

[0060] Fig. 2 illustrates the domain-specific architecture of the lower-level language models (LLLMs) in the hierarchical system. Multiple LLLM modules exist, each trained on data from a particular domain or field of knowledge.

[0061] As depicted, example domains include: • Medicine LLLM (200A) - Trained on datasets like medical textbooks, journals, patient health records, etc. to generate responses to medical queries.

[0062] • Finance LLLM (200B) - Trained on earnings reports, financial news, stock data to respond to finance-related questions.

[0063] • Engineering LLLM (200C) - Trained on textbooks, publications, manuals, standards, and other technical documents to provide engineering-focused answers.

[0064] • Legal LLLM (200D) - Trained on statutes, case law, legal textbooks, and other legal content to handle law-related queries.

[0065] The number and specificity of LLLMs can be adapted based on the breadth of domains the system is meant to cover. LLLMs may utilize domain-specific vocabularies, ontologies, and language conventions particular to their field.

[0066] Their training leverages state-of-the-art neural network architectures such as Transformers, BERT, and GPT-3 customized for each domain. The modular domain-based architecture allows combining specialized understanding with overall versatility.

[0067] In summary, Fig. 2 provides an overview of the domain specialization employed in the lower-level language models to enable optimized responses for diverse queries. The multi-domain approach balances broad knowledge with targeted expertise.

[0068] Fig. 3 illustrates the internal architecture and algorithms implemented within the higher-level language model (HLLM) for analysing candidate / proposed responses and generating an optimal output / response. The components include:

[0069] • Encoder (300) - Employs self-attention layers to map the input query and candidate responses to dense vector representations. Helps assess relevance.

[0070] • Candidate Analysis (301) - Components such as clustering, scoring or ranking are used to evaluate accuracy, redundancy, and other factors.

[0071] • Decoder (302) - Uses algorithms such as beam search to generate and rank multiple optimized candidate responses for selection.

[0072] • Synthesis Mechanisms (303) - Pointer networks, seq2seq models, and attention layers combine relevant pieces of candidate answers.

[0073] • Post-Processing (304) - Weights, heuristics, and language models refine output for coherence, grammar, and style.

[0074] • Policy Update Unit (305) - Reinforcement learning algorithms update model parameters based on user feedback to improve response optimization policy.

[0075] • External Knowledge Sources (306) - APIs, databases, and corpus resources augment the model's contextual knowledge.

[0076] The components leverage the latest techniques in deep learning and natural language processing to balance generalizability with custom optimization abilities. The model may build on standard architectures like BERT and GPT-3.

[0077] The Encoder (300) receives the input query and candidate responses as input. Its output vector representations feed into the Candidate Analysis module (301) and the Decoder (302). The Candidate Analysis module (301) performs clustering, scoring etc on the encoded vectors. It has bidirectional connections to the Decoder (302) to provide analysis and annotations.

[0078] The Decoder (302) leverages the encoded vectors and analysis from 301. It outputs synthesized candidate responses, using search algorithms, to the Synthesis Mechanisms module (303).

[0079] The Synthesis Mechanisms module (303) refines and combines decoded response candidates using pointer networks and attention layers. Its output feeds into Post-Processing module (304).

[0080] The Post-Processing module (304) employs language models and heuristics to improve response quality. Its output is the final HLLM generated optimal response.

[0081] The Policy Update Unit (305) provides reinforcement learning capabilities. It trains on user feedback flowing back into the system. It also tunes parameters of modules 300-304 to improve response optimization policy.

[0082] The role of the External Knowledge module (306) is to provide additional contextual data, APIs, and knowledge databases to augment the other components as needed. The bi-directional connections depicted in Fig. 3 represent API calls.

[0083] In summary, Fig. 3 provides an overview of the algorithms and architectural components within the HLLM that enable analysis and generation of high-quality responses tailored to the input query. The unit selections are designed to be trainable and adaptable using feedback data.

[0084] Fig. 5 provides a more detailed look at the response synthesis mechanisms implemented in the HLLM to combine relevant components from candidate answers into an optimized output. The steps include:

[0085] 1. Input Candidates (500) - The candidate responses generated by the LLLMs are fed into the HLLM synthesis algorithms.

[0086] 2. Segmentation (501) - Each candidate is segmented into chunks or phrases representing distinct ideas using boundary detection techniques.

[0087] 3. Content Scoring (502) - The relevance of each chunk to the original query is quantified using similarity metrics and weights.

[0088] 4. Pointer Networks (503) - Key segments are copied from the candidates as pointers to form a sequence-to-sequence mapping.

[0089] 5. Combination (504) - The pointers are fluidly combined into complete sequences using an encoder-decoder architecture.

[0090] 6. Candidate Generation (505) - Beam search generates multiple synthesized candidate combinations for consideration.

[0091] 7. Candidate Scoring (506) - The candidates are evaluated for coherence, conciseness, grammar, and other attributes.

[0092] 8. Selection (507) - The highest scoring synthesized candidate is selected as the optimal response (508).

[0093] The synthesis techniques leverage innovations like pointer networks to fluidly stitch relevant content from multiple sources. This allows tailoring responses to user queries by identifying and combining the most pertinent information. In summary, Fig. 5 maps out the synthesis process to create optimized and nuanced responses that integrate the strengths of multiple specialized LLLMs. The components may be designed to be trainable to continually improve combination quality.

[0094] Adaptive Learning and Feedback Loop

[0095] An optional but valuable component of the system is a feedback loop. After the HLLM provides the final answer, users may offer feedback, rating the quality, relevance, and accuracy of the response. This feedback may be utilized to fine-tune the HLLM's selection or synthesis process, making the system more accurate and efficient over time. Advantageously, the feedback loop may operate in real-time allowing responses to be refined.

[0096] The feedback loop's adaptation techniques may draw from reinforcement learning principles, structuring user ratings and preferences as rewards. Details such as the reward architecture, balancing exploration vs exploitation, and experience replay may help optimize the system's performance over multiple query-response cycles. The feedback loop enables the HLLM to refine its selection and synthesis capabilities over multiple query-response cycles. User ratings on dimensions such as relevance and coherence may be compiled as rewards. This feedback loop may also influence how the HLLM generates and refines instructions or sub-instructions for LLLMs subsequently.

[0097] A policy gradient reinforcement learning algorithm may use these reward signals to update parameters that control the ranking, combination, and optimization stages. Actions that produce higher rewards may be reinforced.

[0098] To balance exploration, epsilon-greedy selection may be used. The HLLM may exploit learned high- reward strategies most of the time HLLM but at times may explore new selection mechanisms.

[0099] The synthesized responses, user ratings, and HLLM's internal Q-values may be stored in a database. This experience replay could allow periodic retraining of the HLLM on past successful and unsuccessful responses to prevent overfitting.

[0100] It will be understood that the feedback loop may improve the cumulative success rate of responses. The HLLM may evolve its strategies based on a Darwinian selection of high-performing actions.

[0101] Fig. 6 illustrates the adaptive feedback loop implemented in the system to allow continual learning and improvement of the higher-level language model's (HLLM) response optimization capabilities. The process includes:

[0102] 1. Generate Response (600) - The HLLM produces a response to the user query based on its current policies.

[0103] 2. Present to User (601) - The response is provided to the user through the interface.

[0104] 3. User Rates Response (602) - The user rates the response on dimensions such as relevance, accuracy, coherence, etc.

[0105] 4. Compile Feedback (603) - The ratings and feedback are compiled into reward signals.

[0106] 5. Feedback to HLLM (604) - The rewards are fed back into the HLLM's policy update unit. 6. Policy Adaptation (605) - Using reinforcement learning, the HLLM updates its parameters to evolve its response optimization policy.

[0107] 7. Historical Data Storage (606) - Interaction data is stored for training and testing to prevent overfitting.

[0108] 8. Model Re-training (607) - The HLLM is periodically re-trained on accumulated interaction data.

[0109] The feedback loop allows adapting the HLLM's selection and synthesis capabilities based on empirical user interactions over time. This bolsters performance, customization, and human-like conversation.

[0110] In summary, Fig. 6 outlines the self-improvement mechanisms built into the system architecture. The components enable experiential learning by the models to optimize conversational flow.

[0111] Specialized Handling of Mixed Input Modalities

[0112] In a further example of the invention, the system may include a mix of LLLMs with distinct specializations in processing different types of input data. For instance, Formula Handling LLLMs may excel at analysing queries containing mathematical expressions and formulas using e.g. LaTeX decoding, symbolic manipulation, and math-aware encoders. Table Processing LLLMs could incorporate techniques from semantic parsing and data mining to interpret queries about statistical tables. Graphics LLLMs may leverage computer vision and multimodal understanding to extract meaning from charts, diagrams, and other visuals.

[0113] When responding to queries regarding technical documents (e.g. technical standards) containing a blend of text, figures, tables, and mathematical notation, the HLLM could delegate different modalities to the specialized LLLMs best suited for each. For text passages, the HLLM may rely on an LLLM with strong natural language capabilities. For interpreting a mathematical formula, it could call the Formula Handling LLLM. For tabular information, it could leverage the Table Processing LLLM. For a diagram analysis, it could use the Graphics LLLM. The HLLM would then assimilate the complementary inputs, inferences, and conclusions from the diverse LLLMs. Using meta-learning techniques, it could combine these fragmented insights across modalities into a unified response synthesizing the salient pieces. This would enable the system to handle mixed-modality documents and leverage hybrid reasoning across specialized models tailored for different input types. The HLLM orchestrates the models to produce cohesive responses integrating multimodal analysis.

[0114] In the system architecture with lower-level LLLMs dedicated to particular modalities (e.g. text, formulas, tables, graphics), the user's original query could be routed directly to the HLLM rather than broadcasting to all LLLMs simultaneously. The HLLM's encoder would then parse and interpret the initial user query to determine which components involve textual passages, mathematical notation, tabular data, visual diagrams, etc.

[0115] Based on this modulation analysis, the HLLM could then selectively route specific portions, subqueries, or extracted features to the specialized LLLM(s) optimally suited for that modality. For example, detected formula segments might be passed to the Formula Handling LLLM, table data delegated to the Table Processing LLLM, and chart image features sent to the Graphics LLLM. This would allow allocating each query component to the most relevant specialized LLLM for efficient distributed processing based on modalities detected by the HLLM upfront. The HLLM would later assimilate the outputs from the modal-specific LLLMs into a combined response. The routing via the HLLM (rather than broadcasting the full query) would therefore enable leveraging the most apt LLLMs for different modalities.

[0116] Inter-Model Dialogue for Response Improvement

[0117] In some examples of the invention, the HLLM may additionally engage in a conversation with one or more of the LLLMs to request clarification, additional information, or refinements to improve the quality of the proposed responses. Here the term "conversation" denotes not only a direct exchange of data in the form of a query and response, but also a fluid and adaptive exchange between the HLLM and LLLMs. It encompasses a broad communication process, including the sharing of insights, suggestions, and collaborative exploration. Conversations may be dynamic and can evolve based on new inputs, leading to innovative approaches in understanding and processing natural language prompts and responses. It follows from this that the responses by LLLMs to queries posed by the HLLM could themselves be queries e.g. to request further information or clarification of the original query.

[0118] For example, if the initial proposed response from the Medical LLLM contains ambiguous acronyms or terminology, the HLLM may query the Medical LLLM to expand on the unclear terms to make the answer more understandable. The Medical LLLM may itself by responding with a query as to the what level of complexity or length the response should be.

[0119] As another example, the HLLM may detect that the proposed response from the Engineering LLLM lacks sufficient detail or explanation of key principles. The HLLM can then probe the Engineering LLLM with follow-up questions to fill in gaps in the initial response.

[0120] The HLLM may also request the Finance LLLM to double-check critical numerical data in its proposed response or ask for up-to-date figures from real-time market data feeds.

[0121] Additionally, if the initial proposals appear incomplete, biased, or contradictory, the HLLM may engage individual LLLMs in a back-and-forth conversation to clarify ambiguous points, improve objectivity, provide supporting references or excerpts of documents or resolve conflicting statements.

[0122] Equally, the HLLM may receive a response from one LLLM, and query a different LLLM to improve the response. For example, a journal reference provided by an LLLM not having document analysis functionality, could be passed to an LLLM endowed with internet access to identify the referenced journal. This journal reference could then be passed to a further LLLM with suitable document reading functionality e.g. one able to operate a PDF reader plugin, for the extraction of excerpts.

[0123] The exchange may employ conversational protocols and conventions common in the LLLMs' domains to elicit expanded, balanced, and well-rounded perspectives.

[0124] This inter-model interaction enables tapping into the specialized knowledge within each LLLM to iteratively improve response quality in terms of completeness, accuracy, objectivity, and coherence.

[0125] The HLLM may use a hybrid approach, synthesizing its own final response after concluding the dialogue exchanges rather than selecting any single LLLM's response outright.

[0126] Fig. 7 provides an example diagram outlining how inter-model dialogues may be employed in the system to enable iterative improvement of responses. As depicted, the HLLM (701) engages with various LLLMs (702) in a two-way conversational flow. Based on its analysis of an initial user query, the HLLM (701) exchanges information with relevant LLLMs (702) to request clarifications around ambiguous terminology (as shown with the Medical LLLM (702A) example above); probe for additional details where responses may lack sufficient context (as shown with the Engineering LLLM (702C)); verify or cross-check critical data (as shown with the Finance LLLM (702B)); or delegate particular elements, like mathematical expressions, to specialized handlers (as shown with the Formula Handling LLLM (702D)). As shown in fig. 7, the HLLM may instruct the LLLMs e.g. Engineering LLLM (702C) to delegate elements such as mathematical expressions directly to specialist LLLMs e.g. Formula Handling LLLM 702D, or LLMs may interact with each other directly without the need to refer back to the HLLM. This is depicted in fig. 7 in the form of a query and response conversation between Engineering LLLM 702C and Formula Handling LLLM 702D.

[0127] Bidirectional interactions allow the HLLM to leverage the niche capabilities within each LLLM beyond just their initial query responses. This collective inter-model conversation taps into distributed knowledge to improve quality dimensions such as completeness, impartiality, accuracy and coherence.

[0128] As an example, supporting excerpts requested from the Formula Handling LLLM may be combined with market data retrieved by the Finance LLLM to enhance a response coordinated by the HLLM.

[0129] The integrated system is therefore able to harness both wide general knowledge along with specialized domain understanding through a connected network of language models working in conjunction.

[0130] LLLM Capability Profiling for Optimized Routing

[0131] In addition to modality analysis on the query, the HLLM can profile the capabilities of each LLLM through conversational exchanges. During initialization or periodically, the HLLM (701) can probe the LLLMs (702) with domain-specific questions, edge cases, and sample inputs to benchmark strengths and limitations. For instance, the HLLM may evaluate the Formula Handling LLLM (702D) on mathematical reasoning skills by posing esoteric physics problems requiring calculus or linear algebra. A Table Processing LLLM could, for example, be tested on handling complex datasets and queries requiring multivariate analysis. The HLLM catalogues the capabilities uncovered for each LLLM into a knowledge base.

[0132] When routing query components, the HLLM consults this knowledge base to allocate portions optimally based on both modulation type and aligned LLLM competencies. For example, a statisticsheavy table extracted from the query may get routed to a Table Processing LLLM specialized in econometrics rather than a generalist table-handling model.

[0133] Dynamic capability profiling via interrogative conversations allows the HLLM to make informed routing decisions to match query facets with the most apt LLLMs beyond just modalities. This results in responses tailored to the specialized competencies of the underlying LLLM ensemble. Diversity-Based Response Improvement

[0134] In a further example of the invention, the HLLM at the heart of the system integrates a diversitybased response improvement module, which leverages the specialized knowledge encapsulated within the LLLMs. As depicted in Fig. 8, when an initial user query (801) is received, the HLLM analysis component (802) employs paraphrasing algorithms (803) to generate multiple rephrased variants (804) of the original query. These paraphrased queries are each routed to subsets of relevant LLLMs (805), chosen, for example, based on domain alignment. Fig. 8 shows LLLMs 1 - N. Typically, N would be at least 3 i.e. there would need to be at least three LLLMs in order to be sure of a consensus if one of the LLLMs began hallucinating. It is extremely rare for two LLLMs to hallucinate consistently with each other at the same time.

[0135] However, it is possible to operate the system with as few at two LLLMs, although this would introduce certain limitations compared to a system using three or more LLLMs. The preferred method involves a majority consensus mechanism to select the most consistent and accurate responses, which is most effective with three or more LLMs to ensure a clear majority. However, with two LLMs, the system could still offer benefits, albeit with a different approach to response selection and validation. For example, with two LLLMs the HLLM can compare responses for consistency and accuracy. While this does not necessarily provide a majority consensus, it does allow for a basic level of cross-verification between the two models. Furthermore, two LLLMs can still provide a degree of redundancy and error checking. If both LLLMs produce similar responses, it can increase confidence in the accuracy of the output. Conversely, discrepancies between the two responses can flag potential issues for further review. In such cases, however, the HLLM may need to play a more significant role in analysing and synthesizing the responses from the two LLMs, potentially increasing the complexity of its decision-making process. In cases where the two LLMs provide conflicting responses, the HLLM would face a dilemma. In such a case the HLLM could, for example, report both alternative responses to the user, together with its concerns; or interrogate the LLLMs further to understand the cause of the disagreement; or even resolve the difference from its own domain knowledge.

[0136] Returning to Fig. 8, the LLLMs provide a rich set of candidate responses (806), exploring different facets of information, terminology, and contextual interpretations.

[0137] If operating on a parallel processing framework, the LLLMs may process the natural language prompts, both original and paraphrased, in parallel. The HLLM plays a crucial role in analysing the responses generated by the LLLMs. It collects all responses (807) and employs consistency thresholding (808), as shown in Fig. 8, to partition them into groups based on semantic similarity. The largest group (809), forming the consensus cluster, is indicative of the most reliable and contextually appropriate responses.

[0138] The HLLM synthesis component (810), as illustrated in Fig. 8, then selects or combines responses within the consensus cluster to produce the final output response (811). An additional postprocessing phase (812) refines the response for coherence, grammatical correctness, and stylistic appropriateness to generate a response to the user (813). Optionally, user feedback can be incorporated into this phase for continuous improvement and adaptation.

[0139] The system is adept at detecting and eliminating hallucinations through its diverse input processing and parallel response analysis. By leveraging the diverse perspectives provided by the LLLMs, the system effectively identifies outlier responses that may be hallucinations. The consensus-based selection process, as demonstrated in Fig. 8, ensures that only those responses that surpass a certain threshold of similarity and consistency are chosen, thereby filtering out hallucinatory content. The synthesis and refinement phases further eliminate any residual hallucinatory elements, ensuring the final output is both accurate and contextually relevant.

[0140] This example of the invention is suited for implementation on distributed computing systems, and the architecture of the system facilitates scalability and adaptability to various applications and user needs. The inclusion of a feedback mechanism allows for continuous system improvement, making the system more adept at recognizing and eliminating hallucinations over time.

[0141] Enhanced Diversity-Based Improvement

[0142] Building upon the previously outlined structure of the HLLM and LLLMs, a further example of the invention introduces advanced capabilities for enhancing the diversity-based response improvement and anti-hallucination features, as depicted in Fig. 8. These capabilities leverage the interactive and conversational dynamics between the HLLM and LLLMs, augmenting the system's efficiency and accuracy.

[0143] Enhanced Diversity-Based Response Improvement

[0144] 1. Dynamic Query Refinement: The HLLM's ability to engage in dialogues with LLLMs is utilized for dynamically refining the user's query. This interactive process, depicted in the initial stages of Fig. 8, allows for a deeper exploration of the query's context, leading to more nuanced paraphrased variants (804) and, subsequently, a wider array of responses (806) from the LLLMs.

[0145] 2. Iterative Paraphrasing and Enhanced Interpretations: The system employs paraphrasing algorithms (803), as shown in Fig. 8, in tandem with interactive discussions between the HLLM and LLLMs. This iterative process yields paraphrases that encapsulate a broader spectrum of interpretations, enhancing the diversity and depth of the LLLMs' responses.

[0146] 3. Cross-Model Learning and Domain Enhancement: Through conversational exchanges, the LLLMs share and integrate insights from their respective domains, enriching each model's understanding and response capabilities. This cross-model learning fosters a collaborative environment, enhancing the overall diversity and quality of the responses.

[0147] Advanced Anti-Hallucination Mechanisms

[0148] 1. Contextual Verification Through Dialogue: The HLLM uses its conversational capabilities to verify the contextual accuracy of responses from the LLLMs. As illustrated in the latter stages of Fig. 8, this involves a deeper analysis of responses, allowing the HLLM to detect and address potential hallucinations effectively.

[0149] 2. Real-time Feedback and Correction: In instances where the HLLM identifies inaccuracies or inconsistencies in an LLLM's response, it may provide immediate feedback. This real-time correction mechanism enhances the system's ability to rapidly eliminate hallucinatory content from the responses.

[0150] 3. Consensus Building and Response Synthesis: The HLLM orchestrates a consensus-building process among the LLLMs, guiding them towards a unified and coherent response. This process, culminating in the synthesis component (810) in Fig. 8, ensures that the final output response (811) is a product of collective agreement and refinement, significantly reducing the risk of hallucinations.

[0151] Incorporating these advanced features into a system further elevates its capacity to generate accurate, reliable, and contextually appropriate responses. The interactive dialogue between the HLLM and LLLMs, as integrated into the system's architecture, not only enhances response diversity but also plays a crucial role in the real-time detection and elimination of hallucinations. This synergistic approach, as depicted in Fig. 8, ensures a more robust and effective solution to the challenges faced in natural language processing, particularly in improving the reliability and contextual relevance of language model outputs.

[0152] Figure 9 illustrates an architecture implementing a method for enhancing information retrieval and / or understanding in language models. The architecture comprises several key components, each of which plays a crucial role in the overall process. The components and their interactions are described in detail below:

[0153] Information Retrieval Enhancement System Architecture

[0154] 1. External Language Model (901)

[0155] An External Language Model (901) is operated by a user, so as to interact with a heterogeneous document stored 902 in file. This interaction may for example be an interrogation or analysis of the heterogeneous document 902, which requires some form of data extraction from the document. The user may itself be a separate language model, for example, a user may employ a commercial LLM such as Gemini or Chat GPT 4 in order to interact with a document such as heterogeneous document 902 in Fig. 9.

[0156] The heterogeneous document 902, is a document or file comprising content beyond mere text, e.g. tables, mathematical formulas, figures, images, audio and / or video. Examples of heterogeneous documents are technical publications such as conference proceedings, technical standards, books or films.

[0157] With reference to the hierarchical language model system depicted in Fig. 1, the external language model may correspond to any or all of LLLM modules (102) or HLLM module (103) or may be implemented as a standalone system.

[0158] 2. File Traversing Module (903):

[0159] The file traversing module is responsible for scanning and analyzing the document files to identify specific types of content, including formulas, tables, figures, images, audio and / or video. This module utilizes advanced natural language processing (NLP) techniques and / or pattern recognition algorithms to accurately locate and categorize the identified content. The traversing may be done in advance or may take place in real-time during the external model 901's interaction with the document.

[0160] 3. Language Model with Interpretation Capabilities (904):

[0161] The language model with interpretation capabilities (904) could be a state-of-the-art model, such as GPT-4, GPT-4o, Gemini etc. possibly enhanced with additional interpretation capabilities making it suitable to interpret non-textual content such as tables, formulas, drawings and images, audio and / or video. This model may make use extraction techniques such as vectorization i.e. vector retrieval augmented (RAG) techniques, knowledge graphs, other related technologies or combinations thereof to extract identified content. The model then generates text descriptions from the identified content, providing detailed interpretations or effects of the formulas, tables, figures, images, audio and / or video etc.

[0162] 4. Storage Means (905):

[0163] The storage means, which is typically a searchable database, stores the text descriptions generated by the interpretation language model. Each description is indexed according to the corresponding content, allowing for efficient retrieval during the information retrieval process.

[0164] 5. Information Retrieval Module (906):

[0165] When the External Language Model (901) encounters an instance of non-text content within the heterogeneous file (902), it actuates the information retrieval module (906) to access the associated text descriptions from the storage means (905) so as to augment and enhance its understanding of the content.

[0166] 6. Feedback Loop (907):

[0167] The feedback loop component allows users, which may include other language models, to provide feedback on the responses received. This feedback is used to refine and improve the system over time. The feedback loop ensures that the system continuously learns from external interactions, enhancing the accuracy and relevance of future responses.

[0168] Information Retrieval Enhancement Operational Workflow

[0169] Pre-processing Stage

[0170] Step 1: the Heterogeneous File 902 is processed in the first step by the Document Traversing Module 903, which scans the document to identify specific types of content, such as formulas, tables, drawings, images, audio and / or video content. Advanced NLP and pattern recognition techniques are used to accurately locate and categorize the identified content.

[0171] Step 2: Content Interpretation

[0172] The identified content is sent to the language model with interpretation capabilities (904). This model generates text descriptions of the identified content, providing detailed interpretations, explanations and / or effects.

[0173] Step 3: Storage of Descriptions

[0174] The text descriptions generated by the language model are stored in the storage means (905). The descriptions are indexed according to the corresponding content, ensuring efficient retrieval. Analysis Phase:

[0175] Step 5: Information Retrieval

[0176] During the information retrieval process, the External Language Model (901) encounters non-text content in the Heterogeneous File (902). The External Language Module (901) then actuates the information retrieval module (906) to reference the Storage Means (905) to access the relevant text descriptions.

[0177] Step 6: Enhanced Understanding

[0178] By referencing the generated text descriptions, the External Language Model (901) enhances its understanding and analysis of the Heterogeneous Files (902). This step significantly improves the accuracy and relevance of the information retrieved, especially for technical documents containing complex formulas, tables, drawings, images, audio and or video content.

[0179] Step 7: Response Output

[0180] The final response is sent back to the user through the External Language Model (901). The user receives an accurate and contextually relevant answer to their query.

[0181] Step 8: User Feedback

[0182] Feedback is collected through the Feedback Module (907). This feedback is used to refine and improve the system continuously. The language model learns from user interactions, enhancing its future performance. The feedback may be generated by an external user interacting with the External Language Model (901) or may be generated by the External Language Model (901) itself.

[0183] Example Implementation

[0184] Example 1: Identifying and Describing Formulas

[0185] When traversing a technical file (902) which may me a PDF of a technical paper, the file traversing module (903) identifies a mathematical formula. The language model with interpretation capabilities (904) then generates a text description of the formula, which may include a detailed breakdown of its components and their respective roles, as well as the mathematical operations being performed. For instance, if the formula represents a physics equation, the text description might explain the mathematical combination of parameters or elements, the physical principles involved, the nature of the parameters and the context in which the formula is typically applied and / or other information relevant to the prompt(s) being performed by the External Language Model 901.

[0186] The generated text description is stored in the storage means (905), indexed according to the formula. During an information retrieval process, if the External Language Model (901) encounters this formula in a document stored in Heterogeneous File (902), the information retrieval module (906) references the storage means to retrieve the text description and provides it to the External Language Model (901) thereby enhancing the language model's understanding and providing a more accurate analysis.

[0187] In a particular implementation of the invention, a reference to the associated content in the Storage Means (905) may be embedded in the Heterogeneous Files (902) thereby permitting the External Language Model (901) to access the associated descriptive data from the Storage Means (905) directly without having to request the Information Retrieval Module (906) to perform a lookup function of the non-text content.

[0188] Example 2: Interpreting Tables

[0189] In another scenario, the File Traversing Module (903) identifies a table containing statistical data. The Interpretation Language Model (904) generates a functional description of the table, which might not only describe its literal contents in text form, but also highlight key trends, significant data points, and any notable anomalies.

[0190] This description is then stored in the Storage Means (905), indexed by the table reference or page number. When the language model encounters the table in a document stored in Heterogeneous File 902, during the retrieval process, the Information Retrieval Module (906) accesses the stored description, helping the External Language Model (901) to interpret the table accurately and incorporate its insights into the overall analysis.

[0191] Example 3: Analyzing Figures

[0192] The system may also handle figures and diagrams. For example, the File Traversing Module (903) might detect a complex engineering diagram. The Interpretation Language Model (904) then generates a technical description of the diagram, explaining the various components and their interrelationships. The Interpretation Language Model (904) may also generate additional insights such as the functioning of the diagram and may augment the description with additional insights derived from any explanatory text in the Heterogeneous File which references to the diagram.

[0193] This description is stored and indexed in the Storage Means (905). During subsequent retrieval processes, the External Language Model (901) references this description to understand the diagram better so as to be able to provide a detailed explanation as part of its response.

[0194] Enhanced Training and Feedback Integration

[0195] To further improve the system's capabilities, the Interpretation Language Model is trained on a comprehensive corpus of technical documents. This training improves its ability to generate accurate and detailed technical interpretations. Additionally, the text descriptions generated by the language model are verified and refined by domain experts to ensure their accuracy and relevance. The feedback loop (907) plays a crucial role in continuous improvement. Users can provide feedback on the accuracy and relevance of the responses they receive, e.g. via the External Language Model (901). This feedback is used to fine-tune the Interpretation Language Model (904) and the entire system, ensuring that it becomes more adept at handling complex technical content over time.

[0196] Summary of Information Retrieval Enhancement

[0197] The architecture shown in Fig. 9 effectively implements a method for enhancing information retrieval and understanding in language models. By accurately identifying and interpreting specific types of content such as formulas, tables, figures, images, audio and video etc and by leveraging a robust storage and retrieval system, the invention significantly improves the performance of language models in processing heterogeneous content in files. The feedback loop ensures continuous learning and adaptation, making the system increasingly reliable and effective in providing contextually accurate information.

Claims

Claims:

1. A method for enhancing information retrieval and / or understanding in language models, the method comprising: a. traversing a file to identify specific types of content, for example formulas, tables, drawings and / or pictures; b. using a language model to generate text descriptions of the identified content.

2. The method of claim 1, comprising the additional steps of: c. storing the text descriptions in a storage means indexed according to the corresponding content; d. referencing the storage means during the information retrieval process when the language model encounters the specific types of content in the file; and e. providing the language model with the generated text descriptions to enhance its understanding and analysis of the file.

3. The method of claim 1 or claim 2, wherein the language model with relevant interpretation capabilities is configured to generate text descriptions that include mathematical interpretations for formulas and functional descriptions for tables, drawings, and pictures.

4. The method of claim 2, wherein the storage means is a searchable database to facilitate quick retrieval of the text descriptions during the information retrieval process.

5. The method of claim 2, wherein the text descriptions generated by the language model are verified and refined by a domain expert to ensure technical accuracy.

6. The method of claim 2, wherein the language model is trained on a corpus of technical documents to improve its capability to generate accurate interpretations of the identified content.

7. The method of claim 2, wherein the storage means includes metadata for each entry, such as the file title, author, and publication date, to provide additional context during the information retrieval process.

8. The method of claim 2, wherein the language model references the storage means using a predefined set of prompts designed to trigger the retrieval of relevant text descriptions.

9. The method of claim 1 or claim 2, wherein the document traversing step includes segmenting the file into sections based on headings and subheadings to improve the accuracy of content identification.

10. The method of claim 2, wherein the storage means is periodically updated with new text descriptions as additional files are processed and new content is identified.

11. The method of claim 2, further comprising a feedback loop configured to provide user feedback to refine and improve the accuracy of the text descriptions and the overall information retrieval process.

12. A device for enhancing information retrieval and / or understanding in language models, the device comprising: a file traversing module configured to identify specific types of content in a file, for example formulas, tables, drawings, and / or pictures; a language model configured to generate text descriptions of the identified content.

13. The device of claim 12, further comprising: a storage means for storing the text descriptions, indexed according to the corresponding content; an information retrieval module configured to reference the storage means during the information retrieval process when the language model encounters the specific types of content in the file; and a processing unit configured to provide the language model with the generated text descriptions to enhance its understanding and analysis of the file.

14. The device of claim 12 or claim 13, wherein: the language model with relevant interpretation capabilities is configured to generate text descriptions that include mathematical interpretations for formulas and functional descriptions for tables, drawings, and pictures.

15. The device of claim 13, wherein: the storage means is a searchable database configured to facilitate quick retrieval of the text descriptions during the information retrieval process.

16. The device of claim 13, wherein: the text descriptions generated by the language model are verified and refined by a domain expert to ensure technical accuracy.

17. The device of claim 13, wherein: the language model is trained on a corpus of technical documents to improve its capability to generate accurate interpretations of the identified content.

18. The device of claim 13, wherein: the storage means includes metadata for each entry, such as the file title, author, and publication date, to provide additional context during the information retrieval process.

19. The device of claim 13, wherein: the information retrieval module references the storage means using a pre-defined set of prompts designed to trigger the retrieval of relevant text descriptions.

20. The device of claim 12 or claim 13, wherein: the document traversing module is configured to segment the file into sections based on headings and subheadings to improve the accuracy of content identification.

21. The device of claim 13, wherein: the storage means is periodically updated with new text descriptions as additional files are processed and new content is identified.

22. The device of claim 13, further comprising: a feedback loop configured to provide feedback to refine and improve the accuracy of the text descriptions and the overall information retrieval process.

23. A computer-implemented system for enhancing information retrieval and / or understanding in language models, the system comprising: one or more processors; a memory storing instructions that, when executed by the one or more processors, cause the system to perform the method according to and of claims 1 - 11.

24. A database obtained using the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Generative summaries for search results

    US11769017B1

Cited By

  • Multistage document segmentation method adopting self-adaptive dynamic partitioning algorithm

    CN120951941A

  • Language model enhancement-based patent text search metering method, system and equipment

    CN121636721A