Large model retrieval enhancement generation system based on knowledge graph
Through the large-scale retrieval based on knowledge graph, the knowledge base is automatically constructed and maintained, and combined with advanced natural language processing and deep learning, the real-time and personalized generation problems of knowledge graph construction and information retrieval are solved, and efficient and personalized information services are achieved.
Patent Information
- Application Number
- CN202510364917.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology relies on manual participation in knowledge graph construction and information retrieval, which is difficult to update in real time, and traditional systems find it difficult to understand complex natural language queries, and the generated content is lacking in targetedness and adaptability.
The knowledge graph-based large-scale model retrieval enhancement generation system is adopted, including automated construction and maintenance of knowledge bases, combining advanced natural language processing and deep learning technologies, dynamically selecting language models and generation strategies to achieve accurate understanding of complex queries and personalized text generation.
It improves the accuracy and relevance of information retrieval, provides high-quality and personalized text content, adapts to rapidly changing information needs, and improves user work efficiency and decision-making accuracy.
Smart Images

Figure CN120277188A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and specifically to a large model retrieval enhanced generation system based on a knowledge graph. Background Art
[0002] In the information age, the explosive growth of data volume has posed unprecedented challenges to information retrieval and processing technologies. Among them, the knowledge graph, as an efficient information organization and management tool, has become one of the key technologies to solve the above problems. By structuring knowledge, the knowledge graph can not only support complex queries, but also enhance the background understanding ability of various intelligent applications, such as intelligent search, recommendation systems, and automatic question answering.
[0003] Although the application prospect of the knowledge graph is broad, the existing technologies still face multiple limitations in the construction of the knowledge graph, information retrieval, and text generation. First of all, the construction of most current knowledge graphs relies on manual participation or semi-automatic data processing methods, which are not only time-consuming and laborious, but also difficult to update in real time and cannot meet the dynamically changing information needs. Secondly, traditional knowledge graph retrieval systems often have difficulty accurately understanding the deep intention and complex context of user natural language queries, resulting in retrieval results often deviating from the actual needs of users.
[0004] In addition, although text generation systems integrated with advanced language models can generate answers based on retrieved information, these systems usually lack sufficient flexibility and cannot adjust the generation strategy according to different query contents and complexities, making the generated text often lack pertinence and adaptability. For example, in the face of highly specialized queries, the system may not be able to generate sufficiently in-depth and accurate content, and vice versa, in the face of easy-to-understand queries, the system may provide overly complex answers.
[0005] Therefore, in order to overcome the limitations of the existing technology, there is an urgent need to develop a new technical solution that can automatically construct and maintain an updated knowledge graph, achieve efficient processing of complex natural language queries, and can dynamically select the most suitable language model and text generation strategy according to the specific content and complexity of the query. Such a system can not only improve the accuracy and relevance of retrieval, but also provide more personalized and high-quality text content, meeting the high standards of users for information quality and response speed in various scenarios. Summary of the Invention
[0006] (1) Technical Problems to be Solved
[0007] Aiming at the deficiencies of the existing technology, the present invention provides a large model retrieval enhanced generation system based on a knowledge graph, which solves the problems raised in the above background art.
[0008] (2) Technical solution
[0009] To achieve the above object, the present invention provides the following technical solution: A large model retrieval enhanced generation system based on a knowledge graph, the knowledge graph construction module, which is responsible for automatically establishing and maintaining a structured knowledge base containing various entities and their relationships according to predefined rules and data sources. This knowledge base can capture and store detailed information in various fields, and ensure the real-time update and accuracy of the data;
[0010] The retrieval module is used to retrieve entities and information related to the query in the constructed knowledge graph based on the natural language query input by the user. This module adopts advanced natural language processing technologies to ensure that it can understand complex query intentions and contexts, and improve the relevance and accuracy of the retrieval;
[0011] The language model integration module is used to combine the retrieved information and large-scale pre-trained language models (such as GPT series and BERT series models) to perform information-rich and highly relevant text generation. This module can intelligently select appropriate language models and generation strategies according to the query content and context to provide high-quality generated content.
[0012] Preferably, the knowledge graph construction module can accept data collected from multiple data sources, including but not limited to books, academic papers, websites, and databases. This module supports the automated extraction, cleaning, and integration of data to ensure that the information in the knowledge graph is up-to-date and can reflect the latest academic discoveries and market changes.
[0013] Preferably, the retrieval module includes a semantic understanding sub-module, which adopts advanced machine learning and deep learning technologies to not only understand the literal meaning of the text, but also capture the user's query intention and deep context, so as to accurately locate the most relevant information in the knowledge graph.
[0014] Preferably, the language model integration module includes a dynamic selection mechanism that can automatically select the most suitable pre-trained language model and text generation strategy according to the content and complexity of the query. This mechanism ensures that the generated content is not only relevant, but also highly logical and fluent, meeting the high standards of information quality required by users.
[0015] (3) Beneficial effects
[0016] Compared with the prior art, the present invention provides a large model retrieval enhanced generation system based on a knowledge graph, having the following beneficial effects:
[0017] 1. The large model retrieval enhanced generation system based on a knowledge graph has enhanced information accuracy and depth: By combining a structured knowledge graph with large language models such as GPT and BERT, complex queries can be parsed more precisely, providing in-depth, contextually relevant information. This combination ensures that the system not only answers surface queries but also delves into deep-level information that users may need but not directly ask about.
[0018] 2. The large model retrieval enhanced generation system based on a knowledge graph offers highly personalized services: The system can customize content according to the user's professional field and historical query behavior. Especially in the fields of academic research and enterprise knowledge management, this personalized service can significantly improve the user's work efficiency and the relevance of information acquisition. For real-time decision support and report generation, in an enterprise environment, this system can utilize an internally constructed professional knowledge graph to quickly generate decision support materials and business reports, which not only speeds up the decision-making process but also improves the accuracy of decisions and the enterprise's understanding of market dynamics and internal operations.
[0019] 3. The large model retrieval enhanced generation system based on a knowledge graph features technological scalability and flexibility: Due to the modular design, both the knowledge graph and the language model can be extended or updated as needed, enabling the system to adapt to rapidly changing information needs and technological developments, and maintaining long-term adaptability and competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a schematic flowchart of the large model retrieval enhanced generation system based on a knowledge graph proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0022] Please refer to Figure 1 , the present invention provides a technical solution: A large model retrieval enhanced generation system based on a knowledge graph, the knowledge graph construction module, which is responsible for automatically establishing and maintaining a structured knowledge base containing various entities and their relationships according to predefined rules and data sources. This knowledge base can capture and store detailed information in various fields and ensure the real-time update and accuracy of the data;
[0023] The retrieval module is used to retrieve entities and information related to a query in the constructed knowledge graph based on a natural language query input by the user. This module employs advanced natural language processing techniques to ensure that it can understand complex query intents and contexts, thereby improving the relevance and accuracy of the retrieval.
[0024] The language model integration module is used to combine the retrieved information with large-scale pre-trained language models (such as GPT series and BERT series models) for information enrichment and highly relevant text generation. This module can intelligently select appropriate language models and generation strategies according to the query content and context to provide high-quality generated content.
[0025] Thus, the knowledge graph construction module can accept data collected from multiple data sources, including but not limited to books, academic papers, websites, and databases. This module supports the automated extraction, cleaning, and integration of data to ensure that the information in the knowledge graph is up-to-date and can reflect the latest academic findings and market changes.
[0026] Thus, the retrieval module includes a semantic understanding sub-module that uses advanced machine learning and deep learning techniques to not only understand the literal meaning of the text but also capture the user's query intent and deep context, thereby accurately locating the most relevant information in the knowledge graph.
[0027] Thus, the language model integration module includes a dynamic selection mechanism that can automatically select the most suitable pre-trained language model and text generation strategy according to the content and complexity of the query. This mechanism ensures that the generated content is not only relevant but also highly logical and fluent, meeting the high standards of information quality required by users.
[0028] Automated Knowledge Graph Construction and Maintenance System
[0029] Diverse data sources: The system will automatically extract information from multiple data sources such as scientific papers, news reports, official databases, and publicly available Internet content. This includes not only text data but also various formats such as text in charts and videos.
[0030] Application of advanced NLP techniques: Utilize the latest natural language processing techniques, such as BERT and GPT models, to extract key entities and the relationships between them, and automatically construct a structured knowledge graph through these relationships. For example, identify people, places, events, and their interactions and influences.
[0031] Event-driven updates: Implement an event listening system that can react immediately when new important events or information occur, and automatically add or modify relevant nodes and relationships in the knowledge graph. For example, when a new scientific invention is reported, the system immediately incorporates this information into the graph.
[0032] Quality review process: Combine artificial intelligence with manual review to ensure the quality of information in the knowledge graph. Use machine learning models to predict the reliability of information and conduct a second review through an expert system.
[0033] Knowledge graph optimization algorithm: Regularly run the knowledge graph optimization algorithm to clean inaccurate or outdated information, merge duplicate nodes, optimize the structure of links, and ensure the efficiency and accuracy of the knowledge graph.
[0034] Efficient natural language query processing system
[0035] Deep learning understanding: Develop a specialized deep learning model for accurately parsing implicit intentions and complex contexts in queries, which includes understanding comparative and background information, as well as metaphors and implications in language.
[0036] Context relationship mining: Through context analysis, the system can understand the relationship between queries and the user's historical behavior, thereby providing more personalized search results.
[0037] Advanced semantic matching: Use the semantic relationships in the knowledge graph for in-depth search matching. For example, when the user queries "Discoveries of Nobel laureates", the system can provide accurate information by associating nodes of Nobel Prizes and scientific discoveries.
[0038] Interactive query optimization: If the initial query results are not accurate enough or the user needs more information, the system can ask further questions or make suggestions to refine the search criteria.
[0039] Dynamic text generation strategy selection system
[0040] Model selection mechanism: Dynamically select an appropriate text generation model according to the complexity and professionalism of the query. For highly professional content, use a deep model trained in the professional field; for popular content, select a common-sense-based generation model.
[0041] Generation strategy adjustment: The system can learn from the user's feedback and adjust its generation strategy, such as adjusting the level of detail of the text, the terms used, etc., to meet the needs of different users.
[0042] Working principle:
[0043] S1. Data preprocessing: Enter a large amount of data into the database, extract and generate a large number of entities through relevant technologies, and generate a relevant knowledge graph;
[0044] S2. Receive user queries: The user enters a query through the chat window of the e-commerce platform, such as "What is the battery life of this mobile phone?
[0045] S3. Parsing the query: First, the system uses natural language processing technology to parse the user's query content and determine the keywords (such as "mobile phone", "battery life");
[0046] S4. Knowledge graph retrieval: Then, the system queries the built-in knowledge graph, which contains various entity information about products, such as brands, specifications, user evaluations, etc. The system will find the entity nodes related to "mobile phone" and further retrieve the attribute information related to "battery life";
[0047] S5. Generating a response: Combining the retrieved information (such as data on battery capacity, user feedback, average usage duration, etc.), the system uses a language model to generate a coherent and information-rich answer. For example: "This mobile phone is equipped with a large-capacity 4000mAh battery. According to user feedback, it can work continuously for more than a day under normal usage conditions;
[0048] S6. Generating enhanced content: As needed, the system can also generate relevant charts or graphs, such as a comparison chart of battery performance, to more intuitively display information and enhance the user experience.
[0049] S7. Delivering the response: The generated text and graphic content are displayed to the user through the chat window to help the user better understand the product features.
[0050] In summary, for this knowledge graph-based large model retrieval enhanced generation system, the enhanced information accuracy and depth: By combining a structured knowledge graph and large-scale language models such as GPT and BERT, complex queries can be parsed more precisely, providing in-depth and context-related information. This combination ensures that the system not only answers surface queries but also delves deep into the underlying information that users may need but not directly ask about.
[0051] For this knowledge graph-based large model retrieval enhanced generation system, highly personalized services: The system can customize content according to the user's professional field and historical query behavior. Especially in the fields of academic research and enterprise knowledge management, this personalized service can significantly improve the user's work efficiency and the relevance of information acquisition. Real-time decision support and report generation. In an enterprise environment, this system can utilize the internally built professional knowledge graph to quickly generate decision support materials and business reports, which not only speeds up the decision-making process but also improves the accuracy of decision-making and the enterprise's understanding of market dynamics and internal operations.
[0052] For this knowledge graph-based large model retrieval enhanced generation system, through the scalability and flexibility of technology: Due to the modular design, both the knowledge graph and the language model can be extended or updated as needed, enabling the system to adapt to rapidly changing information needs and technological developments and maintain long-term adaptability and competitiveness
[0053] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A large model retrieval enhanced generation system based on a knowledge graph, characterized in that: The knowledge graph construction module is responsible for automatically establishing and maintaining a structured knowledge base containing various entities and their relationships according to predefined rules and data sources. This knowledge base can capture and store detailed information in various fields and ensure the real-time update and accuracy of data. The retrieval module is used to retrieve entities and information related to the query in the constructed knowledge graph based on the natural language query input by the user. This module adopts advanced natural language processing techniques to ensure that it can understand complex query intentions and contexts, improving the relevance and accuracy of retrieval. The language model integration module is used to combine the retrieved information with large-scale pre-trained language models (such as GPT series and BERT series models) for information-rich and highly relevant text generation. This module can intelligently select appropriate language models and generation strategies according to the query content and context to provide high-quality generated content.
2. The large model retrieval enhanced generation system based on a knowledge graph according to claim 1, characterized in that: The knowledge graph construction module can accept data collected from various data sources, including but not limited to books, academic papers, websites, and databases. This module supports the automated extraction, cleaning, and integration of data to ensure that the information in the knowledge graph is up-to-date and can reflect the latest academic discoveries and market changes.
3. The large model retrieval enhanced generation system based on a knowledge graph according to claim 1, wherein: The retrieval module includes a semantic understanding sub-module that adopts advanced machine learning and deep learning techniques to not only understand the literal meaning of words but also capture the user's query intention and deep context, thereby accurately locating the most relevant information in the knowledge graph.
4. The large model retrieval enhanced generation system based on a knowledge graph according to claim 1, wherein: The language model integration module includes a dynamic selection mechanism that can automatically select the most suitable pre-trained language model and text generation strategy according to the content and complexity of the query. This mechanism ensures that the generated content is not only relevant but also highly logical and fluent, meeting the high standards of information quality required by users.
Citation Information
Cited By
Retrieval enhancement generation system and construction method and application method thereof
CN121233750A