Knowledge enhancement retrieval generation method, device and system for multivariate data fusion

By analyzing multivariate data and building corresponding knowledge bases and graphs, combining knowledge graphs to expand user queries, and using big models to generate replies, multiple problems in multivariate data knowledge are solved, data processing efficiency and information display diversity are improved, costs are reduced, and different application scenarios are flexibly responded to.

CN120144720APending Publication Date: 2025-06-13INSPUR SOFTWARE CO LTD

Patent Information

Application Number
CN202510311071.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, in the enhanced retrieval of multivariate data knowledge, the value of sparse data is not fully explored, the need for diversified styles, the boundaries of knowledge generation are not effectively controlled, and the excessive dependence on large models are not available.

Method used

By analyzing multivariate data, a text knowledge base, text vector library, knowledge graph and knowledge graph vector library are built, combined with knowledge graph expansion and rewriting user queries, a big model is used to generate replies that meet users' reading habits, and a response generator is provided with reply output.

Benefits of technology

It improves data processing efficiency and quality, explores the value of sparse data, meets the needs of diversified information display, controls the boundaries of knowledge generation, reduces implementation costs, and flexibly responds to different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144720A_ABST
    Figure CN120144720A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge enhancement retrieval generation method, device and system based on multivariate data fusion, and relates to the field of data processing. Comprising the following steps: S1, analyzing multivariate data, and respectively constructing a text knowledge base, a text vector base, a knowledge graph and a knowledge graph vector base according to different element types, S2, for a query proposed by a user, firstly analyzing a query type, then expanding, rewriting and converting the user query into a new query in combination with the knowledge graph, and finally obtaining the new query. S3, searching the new query in the database to obtain a recall result, S3, finely ranking the recall result, and directly outputting first N answers after fine ranking for the objective query; and S4, for the recall result combined with the knowledge graph, prompting words are reasonably arranged by means of a large model, replies conforming to the reading habits of the user are generated, or reply output is provided for the user through a response generator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a knowledge-enhanced retrieval generation method, device, and system for multivariate data fusion, which relates to the field of data processing. Background Art

[0002] Retrieval Augmentation Generation is an advanced method that combines information retrieval, natural language processing, and machine learning technologies, aiming to improve the performance of retrieval systems through generative models and generate richer, more relevant, and personalized retrieval results. This method usually involves a deep understanding of user queries, integration of context information, and application of knowledge graphs.

[0003] However, currently, information is presented in a diversified manner. Multivariate usually refers to a situation involving multiple variables or multiple dimensions. Multivariate information processing involves the analysis of a dataset containing multiple variables or features, which can come from different sources or information obtained through in-depth mining of a single source, and has different data types and scales. Therefore, some problems are likely to occur when performing knowledge-enhanced retrieval on multivariate data, such as:

[0004] (1) The diversity of data sources and types makes it difficult to determine a unified processing standard;

[0005] (2) The value of sparse data has not been fully explored, and its potential information and patterns have not been effectively utilized;

[0006] (3) In terms of information display, it cannot meet diverse style requirements;

[0007] (4) It is unable to effectively control the knowledge generation boundary, and the large model hallucination problem is very likely to occur;

[0008] (5) The application scenario is overly dependent on large models, unable to give full play to the tool capabilities of rule-based methods and small and medium-sized models. For different target customer groups, the configuration is not flexible enough, and the implementation cost is relatively high, etc. Summary of the Invention

[0009] Aiming at the problems of the prior art, the present invention provides a knowledge-enhanced retrieval generation method, device, and system for multivariate data fusion, which has the characteristics of strong generality and simple implementation, and has broad application prospects.

[0010] The specific solution proposed by the present invention is as follows:

[0011] The present invention provides a knowledge-enhanced retrieval generation method for multivariate data fusion, including:

[0012] S1: Parse the multi-source data. According to different element types, construct a text knowledge base, a text vector library, a knowledge graph, and a knowledge graph vector library respectively.

[0013] S2: For the query proposed by the user, first analyze the query type, then expand and rewrite the user query in combination with the knowledge graph to convert it into a new query. Retrieve the new query in the database to obtain the recall results.

[0014] S3: Re-rank the recall results. For objective queries, directly output the top N responses after re-ranking; for subjective queries, supplement matters related to the recall results in combination with the knowledge graph to obtain improved recall results.

[0015] S4: For the recall results combined with the knowledge graph, with the help of a large model, reasonably arrange the prompt words to generate a response that conforms to the user's reading habit, or provide a response output to the user through a response generator.

[0016] Furthermore, in S1 of the knowledge-enhanced retrieval generation method for multi-source data fusion, collect multi-source data, where multi-source data refers to information of various types and sources.

[0017] Parsing the multi-source data includes extracting, transforming, and integrating information from multiple data sources:

[0018] First, locate the entity data in the text data through entity recognition methods.

[0019] Next, use relation extraction to determine the relationships between entities and use attribute extraction to supplement entity information.

[0020] Subsequently, use entity alignment and knowledge fusion methods to eliminate duplicate data.

[0021] Finally, structure the parsed knowledge through knowledge representation methods for easy storage, retrieval, and reasoning.

[0022] Furthermore, in S2 of the knowledge-enhanced retrieval generation method for multi-source data fusion, extract text elements from the data source, construct a knowledge graph and vectorize the text respectively. When converting the new query, make the query combine the knowledge graph and the vectorized text, and use a converter to fill in the possible missing information in the query. The converter has a built-in probability estimation system, which decides whether to perform a filling operation on specific entities, relationships, or attributes according to the probability statistics of the system. At the same time, combine each token based on rules to generate a new query that improves the retrieval and recognition ability of the system.

[0023] Further, in step S2 of the knowledge-enhanced retrieval generation method for multi-source data fusion, when retrieving a new query, semantic-based knowledge retrieval is performed using the text vector library to obtain the user's intention; the knowledge graph is used to provide a comprehensive knowledge reserve for knowledge recall; meanwhile, the text knowledge base is used to provide structured retrieval results for knowledge recall.

[0024] The present invention also provides a knowledge-enhanced retrieval generation system for multi-source data fusion, including a parsing module, a query module, a ranking module, and a reply output module.

[0025] The parsing module parses multi-source data and constructs a text knowledge base, a text vector library, a knowledge graph, and a knowledge graph vector library respectively according to different element types.

[0026] For the query proposed by the user, the query module first analyzes the query type, then expands and rewrites the user query in combination with the knowledge graph, converts it into a new query, retrieves the new query in the database, and obtains the recall result.

[0027] The ranking module performs fine ranking on the recall results. For objective queries, the top N replies after fine ranking are directly output; for subjective queries, matters related to the recall results are supplemented in combination with the knowledge graph to obtain a complete recall result.

[0028] For the recall results combined with the knowledge graph, the reply output module uses a large model to reasonably arrange prompt words to generate a reply that conforms to the user's reading habit, or provides a reply output to the user through a response generator.

[0029] Further, the parsing module of the knowledge-enhanced retrieval generation system for multi-source data fusion collects multi-source data, and multi-source data refers to information of various types and sources.

[0030] Parsing the multi-source data includes extracting, transforming, and integrating information from multiple data sources:

[0031] First, entity data in the text data is located through entity recognition methods.

[0032] Next, relationship extraction is used to determine the relationships between entities, and attribute extraction is used to supplement entity information.

[0033] Subsequently, entity alignment and knowledge fusion methods are used to eliminate duplicate data.

[0034] Finally, the parsed knowledge is structured through knowledge representation methods for easy storage, retrieval, and reasoning.

[0035] Furthermore, the query module of the knowledge-enhanced retrieval generation system for multi-source data fusion extracts text elements from data sources, constructs a knowledge graph and vectorizes the text respectively. When transforming a new query, the query combines the knowledge graph and the vectorized text, and fills in the possible missing information in the query through a converter. The converter is built with a probability estimation system, and decides whether to perform filling operations on specific entities, relationships or attributes according to the probability statistics of the system. At the same time, based on rules, each token is combined to generate a new query that improves the retrieval and recognition ability of the system.

[0036] Furthermore, when retrieving a new query, the query module of the knowledge-enhanced retrieval generation system for multi-source data fusion performs semantic-based knowledge retrieval using the text vector library to obtain the user's intention; uses the knowledge graph to provide a comprehensive knowledge reserve for knowledge recall; and at the same time, uses the text knowledge base to provide structured retrieval results for knowledge recall.

[0037] The present invention also provides a knowledge-enhanced retrieval generation device for multi-source data fusion, including: at least one memory and at least one processor;

[0038] The at least one memory is used to store machine-readable programs;

[0039] The at least one processor is used to call the machine-readable program and execute the knowledge-enhanced retrieval generation method for multi-source data fusion.

[0040] The beneficial effects of the present invention are as follows:

[0041] (1) Improve data processing efficiency and quality: Through diversified data sources and types, combined with preprocessing technologies, the present invention can effectively eliminate redundant information and ensure that key information is expressed in standardized tokens, thereby improving the accuracy and efficiency of data processing.

[0042] (2) Explore the value of sparse data: The present invention focuses on the potential information and patterns of sparse data. By effectively using these data, more valuable information can be explored, and the data utilization rate and business value can be improved.

[0043] (3) Meet the diverse information display requirements: Aiming at the deficiency that the prior art cannot meet the diverse style requirements, the present invention provides a more abundant information display method to meet the needs of different users.

[0044] (4) Control the boundary of knowledge generation: The present invention effectively solves the problem of hallucinations in large models. By reasonably controlling the boundary of knowledge generation, the generated knowledge is ensured to be more accurate and reliable.

[0045] (5) Flexible response to different application scenarios: The present invention no longer overly relies on large models, but takes into account the tool capabilities of rule-based methods and medium and small models, making the configuration for different target customer groups more flexible and reducing the implementation cost. Brief Description of the Drawings

[0046] Figure 1 It is a schematic diagram of the method flow of the present invention.

[0047] Figure 2 It is a schematic diagram of the database construction process.

[0048] Figure 3 It is a query transformation flowchart.

[0049] Figure 4 It is a flowchart of knowledge graph construction based on hierarchical constraints.

[0050] Figure 5 It is a flowchart of recall result / knowledge retrieval and semantic association.

[0051] Figure 6 It is a schematic diagram of the system application deployment of the present invention.

[0052] Figure 7 It is a schematic diagram of the network architecture. Detailed Implementation Modes

[0053] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited do not limit the present invention.

[0054] Embodiment 1

[0055] The present invention provides a knowledge-enhanced retrieval and generation method for multi-source data fusion, including:

[0056] S1: Analyze the multi-source data, and respectively construct a text knowledge base, a text vector library, a knowledge graph, and a knowledge graph vector library according to different element types.

[0057] Among them, in S1, multi-source data is collected, and multi-source data refers to information of various types and sources.

[0058] Analyzing the multi-source data includes extracting, transforming, and integrating information from multiple data sources:

[0059] First, locate the entity data in the text data through the entity recognition method;

[0060] Then, use relationship extraction to determine the relationships between entities, and use attribute extraction to supplement entity information;

[0061] Subsequently, use entity alignment and knowledge fusion methods to eliminate duplicate data;

[0062] Finally, the parsed knowledge is structured through a knowledge representation method for easy storage, retrieval, and reasoning.

[0063] S2: For the query proposed by the user, first analyze the query type, and then combine the knowledge graph to expand and rewrite the user query into a new query. Retrieve the new query in the database to obtain the recall results.

[0064] In S2, text elements are extracted from the data source, a knowledge graph and vectorized text are constructed respectively. When transforming the new query, the query is combined with the knowledge graph and vectorized text, and the possible missing information in the query is filled by a converter. The converter has a built-in probability estimation system, which decides whether to perform a filling operation on specific entities, relationships, or attributes according to the probability statistics of the system. At the same time, based on rules, each token is combined to generate a new query that improves the retrieval and recognition ability of the system.

[0065] When retrieving the new query, use the text vector library to perform semantic-based knowledge retrieval to obtain the user's intention; use the knowledge graph to provide a comprehensive knowledge reserve for knowledge recall; at the same time, use the text knowledge base to provide structured retrieval results for knowledge recall.

[0066] S3: Perform fine ranking on the recall results. For objective queries, directly output the top N responses after fine ranking; for subjective queries, combine the knowledge graph to supplement matters related to the recall results to obtain improved recall results.

[0067] Among them, a knowledge graph can be constructed based on hierarchical constraints. First, perform a knowledge hierarchy analysis on the multi-source data to obtain the hierarchy; in addition, perform knowledge extraction to obtain the entities, relationships, and attributes of the knowledge source; then, perform synonymy relationship extraction on the three to obtain the anchor entity, anchor relation, and anchor attribute. The so-called anchoring is to establish a mapping relationship with the standard token; finally, perform knowledge fusion on the above three through hierarchical setting, and establish a knowledge graph according to the knowledge fusion.

[0068] S4: For the recall results combined with the knowledge graph, with the help of a large model, reasonably arrange the prompt words to generate a response that conforms to the user's reading habits, or provide a response output to the user through a response generator.

[0069] Embodiment 2

[0070] The present invention also provides a knowledge-enhanced retrieval and generation system for multi-source data fusion, including a parsing module, a query module, a ranking module, and a response output module.

[0071] The parsing module parses the multi-source data and constructs a text knowledge base, a text vector library, a knowledge graph, and a knowledge graph vector library respectively according to different element types.

[0072] For the queries raised by users, the query module first analyzes the query type, then expands and rewrites the user queries in combination with the knowledge graph, converts them into new queries, retrieves the new queries in the database, and obtains the recall results.

[0073] The ranking module performs fine ranking on the recall results. For objective queries, it directly outputs the top N responses after fine ranking; for subjective queries, it supplements matters related to the recall results in combination with the knowledge graph to obtain improved recall results.

[0074] For the recall results combined with the knowledge graph, the reply output module uses a large model to reasonably arrange prompt words to generate replies that conform to the user's reading habits, or provides reply outputs to users through a response generator.

[0075] Regarding the information interaction and execution process among the above modules in the system, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention and will not be elaborated here.

[0076] Similarly, the advantages of the system of the present invention are:

[0077] (1) Improve data processing efficiency and quality: Through diversified data sources and types, combined with preprocessing technologies, the present invention can effectively eliminate redundant information and ensure that key information is expressed in standardized tokens, thereby improving the accuracy and efficiency of data processing.

[0078] (2) Explore the value of sparse data: The present invention focuses on the potential information and patterns of sparse data. By effectively utilizing these data, more valuable information can be explored, improving data utilization and business value.

[0079] (3) Meet the diverse information display needs: Aiming at the deficiency that the prior art cannot meet the diverse style requirements, the present invention provides a richer information display method to meet the needs of different users.

[0080] (4) Control the boundary of knowledge generation: The present invention effectively solves the hallucination problem of large models. By reasonably controlling the boundary of knowledge generation, the generated knowledge is ensured to be more accurate and reliable.

[0081] (5) Flexibly adapt to different application scenarios: The present invention no longer overly relies on large models, but takes into account the tool capabilities of rule-based methods and small and medium-sized models, making the configuration for different target customer groups more flexible and reducing the implementation cost.

[0082] It should be noted that not all steps and modules in the above processes and system structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted according to needs. The system structures described in the above embodiments can be physical structures or logical structures, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or some components in multiple independent devices can be jointly implemented.

[0083] In specific applications, for example, system deployment mainly includes various interactive clients, application server clusters, AI server clusters, various data and database systems, and components for improving functions. The clients include, but are not limited to, business halls, mobile clients, Web clients, and administrator clients, etc., providing functions such as user information entry and access, and providing administrator user operation and maintenance functions. Various external clients connect to the gateway cluster through the "Nginx + firewall" mode to ensure system information security. An application server cluster and an AI server cluster are provided inside the system. The former realizes the basic functions of the system, and the latter provides AI computing services for the system. In addition, through the configuration center, archiving tasks can be customized and system operation parameters can be controlled; database services are deployed to provide operations such as data storage, CRUD, etc., and at the same time, message queue services and cache services are deployed to enhance the stability of the system.

[0084] The network architecture of this system can include the mobile devices of individual customers, and the Web clients access the system through the public cloud via a firewall to complete business processing. The business hall, internal private devices, and operation and maintenance clients access the system through the private cloud. The backend servers are divided into application servers (APP Server), AI servers (AI Server), database servers (DBServer), etc. According to different task types, they respectively run tasks such as application programming interface services (API Service), AI services (AIService), and database (DB) operations.

[0085] Embodiment 3

[0086] The present invention also provides a knowledge enhanced retrieval and generation device for multi-source data fusion, including: at least one memory and at least one processor;

[0087] The at least one memory is used for storing machine-readable programs;

[0088] The at least one processor is used for calling the machine-readable program and executing the multi-source data fusion knowledge enhanced retrieval and generation method.

[0089] For the content such as the information interaction of the processor in the above-mentioned device and the process of executing the readable program, since it is based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention and will not be elaborated here.

[0090] Similarly, the advantages of the device of the present invention are as follows:

[0091] (1) Improve the efficiency and quality of data processing: By diversifying data sources and types and combining preprocessing technologies, the present invention can effectively eliminate redundant information and ensure that key information is expressed in standardized tokens, thereby improving the accuracy and efficiency of data processing.

[0092] (2) Explore the value of sparse data: The present invention focuses on the potential information and patterns of sparse data. By effectively utilizing this data, more valuable information can be explored, improving data utilization and business value.

[0093] (3) Meet the diverse information display requirements: Aiming at the deficiency that the prior art cannot meet the diverse style requirements, the present invention provides a more abundant information display method to meet the needs of different users.

[0094] (4) Control the boundary of knowledge generation: The present invention effectively solves the problem of large model hallucinations. By reasonably controlling the boundary of knowledge generation, it ensures that the generated knowledge is more accurate and reliable.

[0095] (5) Flexibly adapt to different application scenarios: The present invention no longer overly relies on large models, but takes into account the tool capabilities of rule-based methods and small and medium-sized models, making the configuration for different target customer groups more flexible and reducing the implementation cost.

[0096] The above-mentioned embodiments are only preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.

Claims

1. Knowledge-enhanced retrieval generation method based on multivariate data fusion, characterized by include: S1: Analyze the multivariate data and build a text knowledge base, a text vector base, a knowledge graph and a knowledge graph vector base according to the different element types. S2: For the query raised by the user, first analyze the query type, then expand and rewrite the user query in combination with the knowledge graph, convert it into a new query, search the new query in the database, and obtain the recall result. S3: Rank the recall results. For objective queries, directly output the top N replies after ranking. For subjective queries, supplement the items related to the recall results with the knowledge graph to obtain complete recall results. S4: For the recall results combined with the knowledge graph, with the help of a large model, the prompt words are reasonably arranged to generate replies that conform to the user's reading habits, or the response generator is used to provide the user with a reply output.

2. The knowledge-enhanced retrieval generation method of multivariate data fusion according to claim 1 is characterized by: S1 collects multivariate data, which refers to information of various types and sources. Analyzing multivariate data involves extracting, transforming, and integrating information from multiple data sources: First, the entity data in the text data is located through the entity recognition method; Next, use relationship extraction to determine the connection between entities, and use attribute extraction to supplement entity information; Subsequently, entity alignment and knowledge fusion methods are used to eliminate duplicate data; Finally, the parsed knowledge is structured through knowledge representation methods to facilitate storage, retrieval and reasoning.

3. The knowledge-enhanced retrieval generation method of multivariate data fusion according to claim 1 is characterized in that text elements are extracted from the data source in S2, and a knowledge graph and a vectorized text are constructed respectively. When converting a new query, the query is combined with the knowledge graph and the vectorized text, and the information that may be missing in the query is supplemented by the converter, wherein the converter has a built-in probability estimation system, which determines whether to perform a supplement operation on a specific entity, relationship or attribute according to the probability statistics of the system, and combines each word based on the rule to generate a new query that improves the system's retrieval and recognition capabilities.

4. The knowledge-enhanced retrieval generation method of multivariate data fusion according to claim 3, Its characteristic is that when searching for a new query in S2, the text vector library is used to perform semantic-based knowledge retrieval to obtain the user's intention; The knowledge graph is used to provide a comprehensive knowledge reserve for knowledge recall; at the same time, the text knowledge base is used to provide structured retrieval results for knowledge recall.

5. Knowledge-enhanced retrieval generation system based on multivariate data fusion, characterized by It includes parsing module, query module, sorting module and reply output module. The parsing module parses the multivariate data and constructs a text knowledge base, a text vector base, a knowledge graph, and a knowledge graph vector base according to the different element types. The query module first analyzes the query type for the user's query, then expands and rewrites the user's query in combination with the knowledge graph, converts it into a new query, searches the new query in the database, and obtains the recall result. The sorting module performs fine sorting on the recall results. For objective queries, it directly outputs the top N replies after fine sorting. For subjective queries, it combines the knowledge graph to supplement matters related to the recall results to obtain complete recall results. The reply output module combines the recall results of the knowledge graph with the help of a large model to reasonably arrange the prompt words and generate replies that conform to the user's reading habits, or provide the user with reply output through a response generator.

6. The knowledge-enhanced retrieval generation system based on multivariate data fusion according to claim 5 is characterized by: The parsing module collects multivariate data, which refers to information of various types and sources. Analyzing multivariate data involves extracting, transforming, and integrating information from multiple data sources: First, the entity data in the text data is located through the entity recognition method; Next, use relationship extraction to determine the connection between entities, and use attribute extraction to supplement entity information; Subsequently, entity alignment and knowledge fusion methods are used to eliminate duplicate data; Finally, the parsed knowledge is structured through knowledge representation methods to facilitate storage, retrieval and reasoning.

7. The knowledge-enhanced retrieval generation system based on multivariate data fusion according to claim 5 is characterized by: The query module extracts text elements from the data source, constructs a knowledge graph and vectorized text respectively, and when converting a new query, it combines the knowledge graph and vectorized text, and uses the converter to complete the missing information in the query. The converter has a built-in probability estimation system that decides whether to complete specific entities, relationships or attributes based on the system's probability statistics. At the same time, it combines each word based on rules to generate new queries that improve the system's retrieval and recognition capabilities.

8. The knowledge-enhanced retrieval generation system of multivariate data fusion according to claim 7 is characterized by: When searching for a new query, the query module uses the text vector library to perform semantic-based knowledge retrieval to obtain user intent; The knowledge graph is used to provide a comprehensive knowledge reserve for knowledge recall; at the same time, the text knowledge base is used to provide structured retrieval results for knowledge recall.

9. A knowledge-enhanced retrieval generation device for multivariate data fusion, characterized in that include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to execute the knowledge-enhanced retrieval generation method of multivariate data fusion according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Query rewriting method and device

    CN115705331A

  • Information retrieval method and device, storage medium and computer program product

    CN118939763A

  • Automobile intelligent question-answering system and method based on large model and retrieval enhancement

    CN119202206A

  • Method and system for generating enhanced knowledge questions and answers for mixed retrieval of heterogeneous database

    CN119311831A

  • Intelligent question and answer method and device, electronic equipment and storage medium

    CN119441426A

Cited By

  • Hydrofracture question-answering system and method based on cross-language retrieval enhanced generation

    CN121029935A

  • Method and system for reducing illusion information generated by power data agent

    CN121117272A

  • A method and system for reducing hallucination information of power data agents

    CN121117272B

  • Method, medium, and apparatus for retrieval enhancement of ai data lakes

    CN122633883A