High-performance knowledge base system based on multi-modal mixed retrieval
Through a multimodal hybrid retrieval system, combined with vector calculation and knowledge graph modules to process structured and unstructured data, the limitations of traditional knowledge base systems in data processing are solved, and high-precision semantic matching and relational reasoning capabilities are achieved.
Patent Information
- Application Number
- CN202511293514.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-09-11
AI Technical Summary
Traditional knowledge base systems have limitations when processing structured and unstructured data, making it difficult to meet the needs of deep knowledge mining. In addition, using vector retrieval alone cannot fully utilize the knowledge associations contained in structured data.
A multimodal hybrid retrieval system is adopted, combining the vector calculation module and the knowledge graph module. The unstructured text is processed through the quantum neural network and the structured text is converted into knowledge graph data. The hybrid retrieval fusion module is used to perform vector retrieval and graph retrieval to generate fusion retrieval results.
It has greatly improved the semantic matching accuracy and relational reasoning capabilities, and can more comprehensively capture the semantic information and structural relationships in user queries, achieving multi-hop reasoning and accurate response to complex problems.
Smart Images

Figure CN120781933A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, knowledge representation and knowledge retrieval, and particularly relates to a high-performance knowledge base system based on multi-modal hybrid retrieval. BACKGROUND
[0002] In today's era of information explosion, massive data grows rapidly in various forms such as structured and unstructured. Knowledge base system, as an important tool for data storage, management and retrieval, has been widely used in various fields. Traditional knowledge base system usually stores and manages structured data by using relational database, and realizes data retrieval through structured query language such as SQL. However, this method has certain limitations in dealing with complex semantic relationship and knowledge reasoning, and it is difficult to meet the demand for deep mining of knowledge. At the same time, in the face of a large number of unstructured data such as text and document, the traditional system often uses simple retrieval methods such as keyword matching. This method cannot effectively understand the semantic information of the text, and the accuracy and relevance of the retrieval results are poor. Although vector representation-based retrieval technology has emerged in recent years, which can convert unstructured text into vector data and realize semantic retrieval by calculating the similarity between vectors, the use of vector retrieval alone cannot fully utilize the knowledge association contained in structured data. SUMMARY
[0003] (I) Invention purpose The purpose of the present application is to provide a high-performance knowledge base system based on multi-modal hybrid retrieval, which supports vector retrieval and graph retrieval in parallel through a hybrid retrieval module, and improves the accuracy of retrieval.
[0004] (II) Technical solutions To solve the above problems, the present application provides a high-performance knowledge base system based on multi-modal hybrid retrieval, comprising: a vector calculation module, a knowledge graph module, a storage module and a hybrid retrieval fusion module. The vector calculation module is used to convert unstructured text into vector data represented by feature vectors. The knowledge graph module is used to convert structured text into knowledge graph data. The storage module is connected with the vector calculation module and the knowledge graph module, and is used to store the vector data and the knowledge graph data. The hybrid retrieval fusion module is connected with the storage module, and is used to retrieve the storage module according to the query request. The hybrid retrieval includes vector retrieval and graph retrieval, and the fusion retrieval result is generated based on the vector retrieval and the graph retrieval.
[0005] In another aspect of the present application, preferably, the vector calculation module is based on a quantum neural network model. The vector calculation module includes: a quantum data encoding unit and a quantum text feature extraction unit; The quantum data encoding unit is used to encode unstructured text into a quantum state; The quantum text feature extraction unit evolves the quantum state through a parameterized quantum circuit, and collapses the output quantum state into a feature vector based on quantum measurement to form vector data represented by the feature vector; The dimension of the feature vector output by the quantum text feature extraction unit is determined by the number of quantum bits and the measurement method.
[0006] In another aspect of the present invention, preferably, the vector calculation module further includes a quantum verification unit; The quantum verification unit re-encodes the feature vector into a first quantum state and encodes the feature vector of the reference text into a second quantum state; calculating the fidelity between the first quantum state and the second quantum state by a quantum test circuit; When the fidelity is greater than or equal to a fidelity threshold, it is determined that the feature vector conversion is valid.
[0007] In another aspect of the present invention, preferably, the knowledge graph module is used to convert structured text into knowledge graph data, including: Parsing structured text to identify entities, attributes, and relationships within the structured text; The recognition results are mapped into graph-structured data, with entities as nodes and relationships as edges, to form knowledge graph data.
[0008] In another aspect of the present invention, preferably, the storage module includes a PostgreSQL unit and a FAISS vector index unit; The PostgreSQL unit includes a knowledge_bases table, a documents table, a chunks table, a knowledge_graphs table, and a permissions table; The knowledge_bases table is used to store basic information of the knowledge base, the documents table is used to store raw data, the chunks table is used to store unstructured text and its vector data, the knowledge_graphs table is used to store knowledge graph data, and the permissions table is used to manage user access rights to the knowledge base; The FAISS vector index unit is used to create a vector index for the vector data.
[0009] In another aspect of the present invention, preferably, the hybrid search module includes a graph search unit; The graph retrieval unit receives a user query instruction and extracts a key entity list and potential relationship keywords in the user query instruction; Retrieving the knowledge graph data in the storage module using the key entity list and potential relationship keywords as input; The retrieval includes: locating the entity node of the key entity in the key entity list in the knowledge graph, taking the entity node as the starting point, expanding along the relationship edge of the potential relationship keyword, extracting related entities and relationship paths, and obtaining a structured subgraph.
[0010] In another aspect of the present invention, preferably, the hybrid search module includes a vector search unit; The vector retrieval unit converts the structured subgraph into constraints and semantic enhancement vectors for vector retrieval; Converting the user query instruction into an original vector, and fusing the original vector with the semantic enhancement vector to obtain a fused vector; Based on the fusion vector and the constraints of the vector search, the vector data in the storage module is searched to generate a fusion search result.
[0011] In another aspect of the present invention, preferably, the constraint conditions of the vector search include: extracting unique identification information contained in each entity node from the structured subgraph to form an entity identification set; The entity identifier set is used as a filter tag and is applied to the vector search process to limit the search scope of candidate vectors; The generation of the semantic enhancement vector includes: Extracting text description information associated with each entity node in the structured subgraph, and splicing the text description information to form a semantic input text; Vectorized encoding is performed on the semantic input text to generate a semantic enhancement vector.
[0012] In another aspect of the present invention, preferably, the hybrid search fusion module further includes an iteration unit, The iterative unit is used to perform entity relationship mining on the fusion search results, and identify new entities and new relationships that are not included in the current key entity list and potential relationship keyword set; Build an expanded key entity list and relationship keyword set based on the newly added information, The expanded list of key entities and the set of relationship keywords are used as input for a new round of graph retrieval and vector retrieval.
[0013] In another aspect of the present invention, preferably, the hybrid search module further comprises a convergence determination unit, the convergence determination unit being configured to determine whether a search convergence condition is satisfied during the execution of multiple rounds of fusion search, so as to terminate the iteration; The convergence condition includes any of the following: The entity overlap calculated based on the expanded structured subgraph and the entities in the previous round exceeds the preset threshold; The number of the newly added entities and the newly added relationships is lower than a preset threshold.
[0014] (3) Beneficial effects The above technical solution of the present invention has the following beneficial technical effects: This invention utilizes a vector calculation module and a knowledge graph module to process unstructured and structured data, respectively. A hybrid search and fusion module performs both vector and graph searches based on query requests and fuses the search results, significantly improving semantic matching accuracy and relational reasoning capabilities. Compared to traditional single-modal search methods, this invention more comprehensively captures the semantic information and structural relationships in user queries, enabling multi-hop reasoning and precise responses to complex questions, providing key technical support for building intelligent knowledge service systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic diagram of the overall structure of an embodiment of the present invention. DETAILED DESCRIPTION
[0016] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present invention.
[0017] The accompanying drawings illustrate schematic diagrams of structures according to embodiments of the present invention. These figures are not drawn to scale; for clarity, some details are exaggerated and some details may be omitted. The shapes of the various regions and layers shown in the figures, as well as their relative sizes and positions, are merely exemplary and may deviate in practice due to manufacturing tolerances or technical limitations. Those skilled in the art may design regions / layers with different shapes, sizes, and relative positions as needed.
[0018] Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0019] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0020] The present invention will be described in more detail below with reference to the accompanying drawings. In each of the accompanying drawings, identical elements are represented by similar reference numerals. For the sake of clarity, the various parts in the accompanying drawings are not drawn to scale.
[0021] Example 1 A high-performance knowledge base system based on multimodal hybrid retrieval, Figure 1 FIG. 1 shows a schematic diagram of the overall structure of an embodiment of the present invention, as shown in FIG. Figure 1 As shown, it includes: vector calculation module, knowledge graph module, storage module and hybrid retrieval fusion module; The vector calculation module is used to convert unstructured text into vector data represented by feature vectors; the vector calculation module can perform deep semantic understanding of unstructured text based on natural language processing technology, such as pre-trained language models. First, the text is broken down into computable units through basic processing such as word segmentation and part-of-speech tagging; then, a deep learning model is used to map the text into feature vectors in a high-dimensional space, where each dimension represents a different semantic feature of the text. In this embodiment, the vector calculation module is based on a quantum neural network model; The vector calculation module includes: a quantum data encoding unit and a quantum text feature extraction unit; The quantum data encoding unit is used to encode unstructured text into quantum states. Amplitude encoding is used to map unstructured text to quantum states. For text strings, they are first converted into numerical vectors through classical word embedding and then converted into quantum state amplitudes through normalization. In quantum computing, a quantum state can be represented by a complex vector, where the square of the modulus of each component of the vector represents the probability of measuring the quantum state in the corresponding ground state. The purpose of the normalization process is to ensure that the sum of these probabilities is 1. In this way, the original unstructured text is encoded into a quantum state and can be subsequently processed in the quantum computing system.
[0022] The quantum text feature extraction unit evolves the quantum state using a parameterized quantum circuit and, based on quantum measurement, collapses the output quantum state into a feature vector, generating vector data represented by the feature vector. The parameterized quantum circuit is composed of a series of quantum gates, whose operating parameters can be adjusted based on training data. Quantum gates are the basic operating units in quantum computing, similar to logic gates in classical computing. Different combinations of quantum gates can achieve various transformations of quantum states. During the quantum text feature extraction process, the parameterized quantum circuit extracts and transforms features of unstructured text by performing a series of quantum gate operations on quantum bits based on the input quantum state. By leveraging the superposition and entanglement properties of quantum computing, multiple text features can be processed simultaneously, significantly improving computational efficiency. After evolution by the parameterized quantum circuit, the output quantum state is in a superposition state, containing multiple possible state information. Through quantum measurement, the quantum state collapses into a fixed state, resulting in a feature vector. The dimension of the feature vector output by the quantum text feature extraction unit is determined by the number of quantum bits and the measurement method. A greater number of quantum bits increases the state space that can be represented and enriches the features extracted. Different measurement methods also affect the dimension and specific form of the eigenvector. For example, different measurement methods such as projection measurement and weak measurement will produce different measurement results, thus affecting the final eigenvector. The number of qubits and measurement method can be set according to the actual application scenario.
[0023] Furthermore, in this embodiment, the vector calculation module further includes a quantum verification unit; The quantum verification unit re-encodes the feature vector into a first quantum state and the feature vector of the reference text into a second quantum state. The encoding process is similar to the principle of the quantum data encoding unit, which converts information into a quantum state representation. The feature vector of the reference text can be a series of text.
[0024] The fidelity between the first quantum state and the second quantum state is calculated by a quantum test circuit; fidelity is used to measure the degree of similarity between two quantum states. In quantum computing, the higher the fidelity, the more similar the two quantum states are. When the fidelity is greater than or equal to the fidelity threshold, the feature vector conversion is determined to be valid. In the case of a series of texts, only one fidelity greater than or equal to the fidelity threshold is required. The fidelity threshold is pre-set according to specific application scenarios and requirements. If the fidelity is lower than the threshold, it means that the feature vector may have errors or inaccuracies, and the parameters of the quantum text feature extraction unit need to be adjusted, or the feature extraction process needs to be repeated to ensure that the final feature vector can accurately represent the characteristics of the text data. The accuracy and effectiveness of the feature vector output by the quantum text feature extraction unit are improved.
[0025] The knowledge graph module is configured to convert structured text into knowledge graph data, including: parsing the structured text to identify entities, attributes and relationships of the structured text; mapping the identification results into graph structure data, where entities are nodes and relationships are edges, forming knowledge graph data. The knowledge graph module converts structured text into knowledge graph data, which can use rule-driven, machine learning and other methods to parse text, accurately identify entities, attributes and relationships, determine attributes through dependency syntax analysis, and extract relationships based on pattern matching; the identification results are mapped into graph structure data, with entities as nodes and relationships as edges. Further, nodes can be classified and metadata can be added to edges.
[0026] The storage module is connected to the vector calculation module and the knowledge graph module, and is configured to store the vector data and the knowledge graph data; in this embodiment, the storage module includes a PostgreSQL unit and a FAISS vector index unit; the PostgreSQL unit includes a knowledge_bases table, a documents table, a chunks table, a knowledge_graphs table and a permissions table; the knowledge_bases table is configured to store basic information of the knowledge base, the documents table is configured to store raw data, the chunks table is configured to store unstructured text and vector data thereof, the knowledge_graphs table is configured to store knowledge graph data, and the permissions table is configured to manage access rights of users to the knowledge base; The FAISS vector index unit is configured to create a vector index for the vector data.
[0027] The knowledge_bases table records the name of the knowledge base, the creation time, the person in charge and other meta information, providing a global perspective for multi-knowledge base management; the documents table stores raw data, whether it is a document, a report or other format file, it can keep the original form intact; the chunks table processes the unstructured text in blocks and synchronously stores the corresponding vector data, laying a foundation for subsequent semantic retrieval; the knowledge_graphs table specifically carries the node, edge and relationship data generated by the knowledge graph module, ensuring the integrity of the graph structure information; the permissions table realizes fine management of user access to the knowledge base through permission control strategies, ensuring data security. The FAISS vector index unit is a processing of vector data, which creates an efficient index structure for the text vectors in the chunks table using hierarchical clustering, hash mapping and other algorithms, greatly improving the efficiency of similarity retrieval, enabling fast matching and querying of massive vector data.
[0028] Furthermore, in this embodiment, a synchronization mechanism is set between the PostgreSQL unit and the FAISS vector index unit; This synchronization mechanism is used to perform scheduled full synchronization operations between the PostgreSQL unit and the FAISS vector index unit. Scheduled full synchronization comprehensively compares and updates the data of the two units according to preset periods, such as daily early mornings and weekly weekends. During a full synchronization, the system rescans all unstructured text and its vector data in the PostgreSQL unit's chunks table, synchronizes the latest data status to the FAISS vector index unit, and rebuilds the vector index structure. This prevents the index from being out of sync with the actual data due to long-term data accumulation, ensuring index accuracy in large-scale data environments.
[0029] Data changes in PostgreSQL units are tracked using timestamps, and FAISS vector index units are synchronized based on the tracking results. Timestamp-based data change tracking and synchronization achieves real-time response. Whenever data in a PostgreSQL unit is added, modified, or deleted, the timestamp of the corresponding data row is automatically recorded. The synchronization mechanism continuously monitors changes in these timestamps and, upon detecting a data update, performs the appropriate action based on the type of change. This ensures that relational and vector data remain consistent during storage and updates, laying a solid foundation for efficient hybrid search.
[0030] The hybrid search fusion module is connected to the storage module and is used to search the storage module according to the query request. The hybrid search includes vector search and graph search, and generates a fusion search result based on the vector search and graph search. The graph search and vector search units combine the knowledge graph structured information and text semantic representation to achieve high-precision semantic enhanced search for user natural language queries. In this embodiment, the hybrid search module includes a graph search unit; The graph retrieval unit receives a user query instruction and extracts a list of key entities and potential relationship keywords in the user query instruction; the user query instruction is a natural language query instruction. In this embodiment, a lightweight natural language processing model, such as BERT, ERNIE, etc., is called for analysis to extract key entities in the user query instruction, such as names of people, names of organizations, technical terms, etc., and also extracts potential relationship keywords, such as influence, inclusion, cooperation, etc.
[0031] Retrieving the knowledge graph data in the storage module using the key entity list and potential relationship keywords as input; The searching includes: locating an entity node of a key entity in the key entity list in the knowledge graph, expanding from the entity node along a relationship edge of a potential relationship keyword, extracting a related entity and a relationship path, and obtaining a structured subgraph. Based on the extracted key entity list, corresponding entity nodes are found in the knowledge graph, and are expanded along the relationship edge of the potential relationship keyword. The relationship edge can be expanded by one or two, and a structured subgraph related to the entity is obtained. The structured subgraph includes: an entity node, a relationship path between nodes, and node attribute values and context text descriptions thereof. In order to avoid graph explosion, it can be set that: each node is expanded by at most M edges; and the number of hops is limited to at most 1-2 hops.
[0032] Further, in the embodiment, the mixed search module includes a vector search unit; The vector search unit converts the structured subgraph into a constraint condition for vector search and a semantic enhancement vector; uses a graph neural network or a graph embedding algorithm to embed and encode the structured subgraph. The semantic enhancement vector of the structured subgraph is output. The user query instruction is converted into an original vector, semantic features are extracted through a pre-trained language model, the user intent is vectorized, and an original query vector is obtained. The original vector and the semantic enhancement vector are fused to obtain a fusion vector; here, the fusion uses a weighting method, and the weight setting can be set according to the type of the problem.
[0033] Further, in the embodiment, the constraint condition for vector search includes: extracting unique identification information contained in each entity node from the structured subgraph to form an entity identification set; and setting the constraint condition in the vector search process can effectively improve the accuracy and efficiency of the search.
[0034] The entity identification set is used as a filter tag in the vector search process to limit the search range of the candidate vector; the entity identification set is used as a filter tag in the vector search process, and the principle is to limit the search range. In an actual search scenario, a large amount of vector data is like a vast ocean, and if there is no constraint condition, the search is like finding a needle in a haystack. The entity identification set as a filter tag filters out candidate vectors that meet certain conditions.
[0035] The generation of the semantic enhancement vector includes: Extract the text description information associated with each entity node in the structured subgraph, and splice the text description information to form a semantic input text; the generation of the semantic enhancement vector is to enable the vector to better express the semantic information of the entity, thereby improving the accuracy of the retrieval. Extract the text description information associated with each entity node in the structured subgraph. Text description information is a detailed description of entity features, attributes, functions and other aspects, and is an important carrier of entity semantics. Splice the extracted text description information to form a semantic input text. Integrate scattered and fragmented text information into a complete semantic unit for subsequent vectorized encoding. The splicing method can select appropriate delimiters or formats according to actual needs to connect different types of text information in an orderly manner. The semantic input text is vectorized and encoded to generate a semantically enhanced vector. Using vectorized encoding techniques, such as word vector models from deep learning or architecture-based models, the semantic input text is processed and converted into a semantically enhanced vector. By capturing the semantic relationships and contextual information in the text, the generated semantically enhanced vector not only contains the textual information of the entity but also contains deeper semantic features, providing a richer semantic basis for subsequent retrieval.
[0036] Based on the fusion vector and the constraints of the vector retrieval, the vector data in the storage module is retrieved to generate a fusion retrieval result. The fusion vector combines information from multiple dimensions and can more comprehensively describe the characteristics of the entity. The constraints of the vector retrieval clarify the scope and direction of the retrieval. When searching the vector data in the storage module, first, based on the constraints of the vector retrieval, the entity identifier set filtering labels are used to preliminarily screen the candidate vectors to narrow the retrieval scope. For example, in an intelligent customer service system, when a user asks a question, the system converts the question into a semantically enhanced vector, and at the same time sets the constraints of the vector retrieval based on the business field, product category and other information involved in the question, retrieves the most relevant knowledge vector data from the storage module, and provides answers to users quickly and accurately, greatly improving user experience and service efficiency.
[0037] Furthermore, in this embodiment, the hybrid retrieval fusion module also includes an iteration unit, which includes: when performing entity relationship mining on the fusion retrieval results, the iteration unit uses natural language processing and graph analysis technology to deeply analyze the text and graph structure in the retrieval results.
[0038] The iterative unit is used to perform entity relationship mining on the fusion search results, and identify new entities and new relationships that are not included in the current key entity list and potential relationship keyword set; Based on the new information, the iteration unit constructs the extended key entity list and the relationship keyword set. This process is not simply piling up information, but through logical sorting and semantic association, the new entities and relationships are integrated into the original system.
[0039] The extended key entity list and the relationship keyword set are used as the input of the new round of graph retrieval and vector retrieval. The extended key entity list and the relationship keyword set are used as the input of the new round of graph retrieval and vector retrieval, guiding the retrieval system to explore in a wider and deeper knowledge space, and obtaining more comprehensive and accurate information.
[0040] The mixed retrieval module further includes a convergence determination unit configured to determine whether a retrieval convergence condition is met to terminate the iteration during the execution of the multiple rounds of fusion retrieval. By setting a reasonable convergence condition, the retrieval process is prevented from meaningless infinite loop, and the iteration is terminated at the right time to output stable and reliable retrieval results.
[0041] The convergence condition includes any of the following: The entity overlap degree calculated based on the extended structured subgraph and the entity obtained in the previous round exceeds a preset threshold. In actual retrieval, as the iteration proceeds, the structured subgraph is continuously expanded. If most of the entities in the newly generated subgraph are repeated in the previous round, it indicates that the retrieval is close to "saturation", and it is difficult to obtain valuable new information.
[0042] The number of new entities and new relationships is lower than a preset threshold. When the number of new entities and relationships mined in each iteration gradually decreases, even below a certain set value, it means that the retrieval system has fully explored the relevant knowledge field, and further iteration is difficult to bring significant information gain. The iteration unit and the convergence determination unit cooperate with each other to give the mixed retrieval fusion module strong self-adaptive ability.
[0043] The present application processes unstructured and structured data through the vector calculation module and the knowledge graph module respectively, and the mixed retrieval fusion module performs vector retrieval and graph retrieval based on the query request, and fuses the retrieval results, which greatly improves the semantic matching accuracy and relationship reasoning ability. Compared with the traditional single modal retrieval method, the present application can more comprehensively capture the semantic information and structural relationship in the user query, realize multi-hop reasoning and accurate response for complex problems, and provide key technical support for building an intelligent knowledge service system.
[0044] It should be understood that the foregoing detailed description of the application, rather than limiting the application, is provided as an illustrative example of the application, with the scope of the application being indicated by the appended claims and equivalents thereof. Therefore, any modification, equivalent replacement, improvement, etc. made without departing from the spirit and scope of the application should be included in the scope of the protection of the application. In addition, the appended claims of the present application are intended to cover all changes and modifications falling within the scope and boundary of the appended claims, or equivalents of such scope and boundary.
[0045] The present application has been described above with reference to the embodiments. However, these embodiments are merely for illustrative purposes, and are not intended to limit the scope of the present application. The scope of the present application is defined by the appended claims and equivalents thereof. Various substitutions and modifications can be made to the embodiments by those skilled in the art without departing from the scope of the present application, and such substitutions and modifications should fall within the scope of the present application.
[0046] Although the embodiments of the present application have been described in detail, it should be understood that various changes, substitutions and alterations can be made to the embodiments without departing from the spirit and scope of the present application.
[0047] Obviously, the above-described embodiments are merely for illustrative purposes, and are not intended to limit the embodiments. Based on the above description, other different forms of changes or modifications can be made by those skilled in the art. Here, it is not necessary or possible to exhaust all the embodiments. The obvious changes or modifications derived therefrom are still within the scope of protection of the present application.
Claims
1. A high-performance knowledge base system based on multimodal hybrid retrieval, characterized by: include: Vector calculation module, knowledge graph module, storage module and hybrid retrieval fusion module; The vector calculation module is used to convert unstructured text into vector data represented by feature vectors; The knowledge graph module is used to convert structured text into knowledge graph data; The storage module is connected to the vector calculation module and the knowledge graph module, and is used to store the vector data and the knowledge graph data; The hybrid search fusion module is connected to the storage module and is used to search the storage module according to a query request. The hybrid search includes vector search and graph search, and generates a fusion search result based on the vector search and graph search.
2. The high-performance knowledge base system based on multimodal hybrid retrieval according to claim 1 is characterized in that: The vector calculation module is based on a quantum neural network model; The vector calculation module includes: a quantum data encoding unit and a quantum text feature extraction unit; The quantum data encoding unit is used to encode unstructured text into a quantum state; The quantum text feature extraction unit evolves the quantum state through a parameterized quantum circuit, and collapses the output quantum state into a feature vector based on quantum measurement to form vector data represented by the feature vector; The dimension of the feature vector output by the quantum text feature extraction unit is determined by the number of quantum bits and the measurement method.
3. The high-performance knowledge base system based on multimodal hybrid retrieval according to claim 2 is characterized in that: The vector calculation module also includes a quantum verification unit; The quantum verification unit re-encodes the feature vector into a first quantum state and encodes the feature vector of the reference text into a second quantum state; calculating the fidelity between the first quantum state and the second quantum state by a quantum test circuit; When the fidelity is greater than or equal to a fidelity threshold, it is determined that the feature vector conversion is valid.
4. The high-performance knowledge base system based on multimodal hybrid retrieval according to claim 1 is characterized in that: The knowledge graph module is used to convert structured text into knowledge graph data, including: Parsing structured text to identify entities, attributes, and relationships within the structured text; The recognition results are mapped into graph-structured data, with entities as nodes and relationships as edges, to form knowledge graph data.
5. The high-performance knowledge base system based on multimodal hybrid retrieval according to claim 1 is characterized in that: The storage module includes a PostgreSQL unit and a FAISS vector index unit; The PostgreSQL unit includes a knowledge_bases table, a documents table, a chunks table, a knowledge_graphs table, and a permissions table; The knowledge_bases table is used to store basic information of the knowledge base, the documents table is used to store raw data, the chunks table is used to store unstructured text and its vector data, the knowledge_graphs table is used to store knowledge graph data, and the permissions table is used to manage user access rights to the knowledge base; The FAISS vector index unit is used to create a vector index for the vector data.
6. The high-performance knowledge base system based on multimodal hybrid retrieval according to claim 1 is characterized in that: The hybrid retrieval module includes a graph retrieval unit; The graph retrieval unit receives a user query instruction and extracts a key entity list and potential relationship keywords in the user query instruction; Retrieving the knowledge graph data in the storage module using the key entity list and potential relationship keywords as input; The retrieval includes: locating the entity node of the key entity in the key entity list in the knowledge graph, taking the entity node as the starting point, expanding along the relationship edge of the potential relationship keyword, extracting related entities and relationship paths, and obtaining a structured subgraph.
7. The high-performance knowledge base system based on multimodal hybrid retrieval according to claim 6 is characterized in that: The hybrid retrieval module includes a vector retrieval unit; The vector retrieval unit converts the structured subgraph into constraints and semantic enhancement vectors for vector retrieval; Converting the user query instruction into an original vector, and fusing the original vector with the semantic enhancement vector to obtain a fused vector; Based on the fusion vector and the constraints of the vector search, the vector data in the storage module is searched to generate a fusion search result.
8. The high-performance knowledge base system based on multimodal hybrid retrieval according to claim 7 is characterized in that: The constraints of the vector search include: extracting unique identification information contained in each entity node from the structured subgraph to form an entity identification set; The entity identifier set is used as a filter tag and is applied to the vector search process to limit the search scope of candidate vectors; The generation of the semantic enhancement vector includes: Extracting text description information associated with each entity node in the structured subgraph, and splicing the text description information to form a semantic input text; Vectorized encoding is performed on the semantic input text to generate a semantic enhancement vector.
9. The high-performance knowledge base system based on multimodal hybrid retrieval according to claim 8 is characterized in that: The hybrid retrieval fusion module also includes an iteration unit, The iterative unit is used to perform entity relationship mining on the fusion search results, and identify new entities and new relationships that are not included in the current key entity list and potential relationship keyword set; Build an expanded key entity list and relationship keyword set based on the newly added information, The expanded list of key entities and the set of relationship keywords are used as input for a new round of graph retrieval and vector retrieval.
10. The high-performance knowledge base system based on multimodal hybrid retrieval according to claim 9 is characterized in that: The hybrid search module further includes a convergence determination unit, which is used to determine whether a search convergence condition is met during the execution of multiple rounds of fusion search to terminate the iteration; The convergence condition includes any of the following: The entity overlap calculated based on the expanded structured subgraph and the entities in the previous round exceeds the preset threshold; The number of the newly added entities and the newly added relationships is lower than a preset threshold.
Citation Information
Patent Citations
Method for processing text semantic similarity through quantum neural network model
CN118821790A
Knowledge base retrieval method and system based on mixed weight and knowledge graph
CN119377374A
RAG-based multi-format data rapid query method
CN119415544A
Intelligent agent, indoor navigation method and equipment thereof, medium and product
CN119443287A
Large model retrieval enhancement method based on knowledge graph and knowledge representation
CN119808914A
Cited By
Dynamic multi-modal knowledge graph retrieval method for military training
CN121434415A