Multimodal knowledge retrieval and reasoning method and device based on graph network

CN122549575APending Publication Date: 2026-08-11GUANGDONG UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-07
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]本发明实施例提供了一种基于图网络的多模态知识检索与推理方法及装置,旨在解决现有技术中用于知识库的检索方法所存在的无法进行知识数据关联分析的问题

Benefits of technology

[0008]This invention provides a multimodal knowledge retrieval and reasoning method based on graph networks. The method extracts corresponding graph features from input multi-source heterogeneous data, compares these features with a pre-stored knowledge graph to obtain comparison results, stores the multi-source heterogeneous data in a knowledge base based on the comparison results, and updates the knowledge graph accordingly. Based on a retrieval request, the method retrieves retrieval information from the knowledge graph and performs reasoning analysis based on the retrieval request to obtain reasoning analysis information. Finally, it generates reasoning response information corresponding to the reasoning analysis information by combining the knowledge base information with the reasoning response information. This multimodal knowledge retrieval and reasoning method dynamically updates the knowledge base and knowledge graph by obtaining comparison results, acquires deep graph correlation information based on retrieval requests, and performs targeted reasoning analysis on the retrieval information to generate reasoning response information. It can perform in-depth reasoning analysis based on the logical connections between knowledge content, significantly improving the relevance and accuracy of knowledge retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122549575A_ABST
    Figure CN122549575A_ABST
Patent Text Reader

Abstract

This invention provides a multimodal knowledge retrieval and reasoning method and apparatus based on graph networks. The method extracts corresponding graph features from input multi-source heterogeneous data, compares these features with a pre-stored knowledge graph to obtain comparison results, stores the multi-source heterogeneous data in a knowledge base based on the comparison results, and updates the knowledge graph accordingly. Based on a retrieval request, the knowledge graph is searched to obtain retrieval information, and inference analysis is performed in conjunction with the retrieval request to obtain inference analysis information. Finally, inference response information corresponding to the inference analysis information is generated using the knowledge base. This multimodal knowledge retrieval and reasoning method dynamically updates the knowledge base and knowledge graph by obtaining comparison results, acquires deep graph correlation information based on retrieval requests, and performs targeted inference analysis on the retrieval information to generate inference response information. It can perform in-depth inference analysis based on the logical connections between knowledge content, significantly improving the relevance and accuracy of knowledge retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent reasoning technology, and in particular to a multimodal knowledge retrieval and reasoning method and apparatus based on graph networks. Background Technology

[0002] With the rapid development of big data and artificial intelligence technologies, knowledge-intensive industries such as technology, finance, healthcare, and manufacturing are increasingly demanding in-depth retrieval and analysis of knowledge bases. Current technologies for knowledge base retrieval typically construct knowledge base systems containing multi-source heterogeneous data. These systems generally extract key features from different data types using pre-defined parsing rules, converting them into vector form and storing them in a database. When a user initiates a search request, the system uses keyword matching or basic vector similarity calculations to find relevant information in the database and outputs it as the search result. However, existing knowledge base retrieval methods cannot accurately analyze the relationships between different knowledge data, resulting in a loose knowledge structure and hindering in-depth analysis and reasoning of knowledge content. Therefore, existing knowledge base retrieval methods suffer from the problem of being unable to perform knowledge data correlation analysis. Summary of the Invention

[0003] This invention provides a multimodal knowledge retrieval and reasoning method and apparatus based on graph networks, aiming to solve the problem that existing retrieval methods for knowledge bases cannot perform knowledge data association analysis.

[0004] In a first aspect, embodiments of the present invention provide a multimodal knowledge retrieval and reasoning method based on graph networks, the method comprising: Receive the input multi-source heterogeneous data, and extract the corresponding spectral features from the multi-source heterogeneous data according to the preset spectral feature extraction rules; The graph features are compared with the pre-stored knowledge graph according to the preset comparison strategy to obtain the corresponding comparison results; Based on the comparison results, the multi-source heterogeneous data is stored in a preset knowledge base and the knowledge graph is updated; If the input search request is received, the knowledge graph is searched according to the search request to obtain the corresponding search information; Based on the preset reasoning analysis model and the retrieval request, the retrieval information is reasoned and analyzed to obtain the corresponding reasoning analysis information; Based on the preset response generation rules and the knowledge base, inference response information corresponding to the inference analysis information is generated.

[0005] Secondly, embodiments of the present invention also provide a multimodal knowledge retrieval and reasoning apparatus based on graph networks, the apparatus being used to execute the multimodal knowledge retrieval and reasoning method based on graph networks as described in the first aspect, the apparatus comprising: The graph feature extraction unit is used to receive the input multi-source heterogeneous data and extract the corresponding graph features from the multi-source heterogeneous data according to the preset graph feature extraction rules. The comparison result acquisition unit is used to compare the graph features with the pre-stored knowledge graph according to a preset comparison strategy to obtain the corresponding comparison results. The knowledge graph update unit is used to store the multi-source heterogeneous data into a preset knowledge base and update the knowledge graph based on the comparison results. The information retrieval unit is used to retrieve corresponding information from the knowledge graph based on the input retrieval request if it receives the input retrieval request. The reasoning analysis information acquisition unit is used to perform reasoning analysis on the search information according to the preset reasoning analysis model and the search request to obtain the corresponding reasoning analysis information; The reasoning response information generation unit is used to generate reasoning response information corresponding to the reasoning analysis information according to the preset response generation rules and the knowledge base.

[0006] Thirdly, embodiments of the present invention also provide an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect above.

[0007] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, can implement the method described in the first aspect.

[0008] This invention provides a multimodal knowledge retrieval and reasoning method based on graph networks. The method extracts corresponding graph features from input multi-source heterogeneous data, compares these features with a pre-stored knowledge graph to obtain comparison results, stores the multi-source heterogeneous data in a knowledge base based on the comparison results, and updates the knowledge graph accordingly. Based on a retrieval request, the method retrieves retrieval information from the knowledge graph and performs reasoning analysis based on the retrieval request to obtain reasoning analysis information. Finally, it generates reasoning response information corresponding to the reasoning analysis information by combining the knowledge base information with the reasoning response information. This multimodal knowledge retrieval and reasoning method dynamically updates the knowledge base and knowledge graph by obtaining comparison results, acquires deep graph correlation information based on retrieval requests, and performs targeted reasoning analysis on the retrieval information to generate reasoning response information. It can perform in-depth reasoning analysis based on the logical connections between knowledge content, significantly improving the relevance and accuracy of knowledge retrieval. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A flowchart illustrating the multimodal knowledge retrieval and reasoning method based on graph networks provided in an embodiment of the present invention; Figure 2 A schematic block diagram of an electronic device provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating an application scenario of the multimodal knowledge retrieval and reasoning method based on graph networks provided in this embodiment of the invention. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0012] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0013] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0014] It should also be further understood that the term "and / or" as used in this specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. Embodiments of this invention provide a multimodal knowledge retrieval and reasoning method and apparatus based on graph networks. The multimodal knowledge retrieval and reasoning method based on graph networks is applied to a terminal device. The method is executed through application software installed on the terminal device to achieve knowledge content retrieval and reasoning analysis. The terminal device can be a server or a smart terminal. The server can be implemented using an independent server or a server cluster composed of multiple servers. The server can receive multi-source heterogeneous data and retrieval requests input by the user and process them accordingly. Smart terminals include, but are not limited to, smartphones, tablets, desktop computers, laptops, and other electronic devices. In this case, the user directly inputs multi-source heterogeneous data and retrieval requests on the smart terminal and processes them locally. The invention will now be described in detail through specific embodiments.

[0015] Figure 1 This is a flowchart illustrating the multimodal knowledge retrieval and reasoning method based on graph networks provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps S110-S160.

[0016] S110. Receive the input multi-source heterogeneous data, and extract the corresponding spectral features from the multi-source heterogeneous data according to the preset spectral feature extraction rules.

[0017] Users can input multi-source heterogeneous data, and the corresponding spectral features are extracted from this data according to spectral feature extraction rules. Multi-source heterogeneous data refers to data sets from different sources and with different representations, specifically including text data, spectral data, image data, and time-series signal data. Spectral feature extraction rules are rules used to analyze and vectorize features of different types of data. Their function is to analyze the inherent attributes of the data to obtain spectral features representing its intrinsic characteristics. Spectral features are structured data that can uniformly represent the semantic information of multimodal data, composed of analytical information and feature vectors. By uniformly extracting features through spectral feature extraction rules, features of different types of data can be uniformly represented, thus obtaining spectral features with a unified form.

[0018] In a specific embodiment, step S110 includes the following sub-steps: obtaining parsing conditions corresponding to the data type from the graph feature extraction rules according to the data type of the multi-source heterogeneous data; parsing the data information in the multi-source heterogeneous data according to the parsing conditions to obtain corresponding parsing information; obtaining feature vectors corresponding to the parsing information according to the vector encoder corresponding to the data type in the graph feature extraction rules; and combining the parsing information with the feature vectors to obtain corresponding graph features.

[0019] The process involves identifying the data types of multi-source heterogeneous data and determining the corresponding parsing conditions for graph feature extraction rules based on these data types. The data types of multi-source heterogeneous data specifically include various modalities such as text, graphs, images, and time-series signals. Graph feature extraction rules are a pre-built rule base that stores pre-defined parsing conditions for different data types. These parsing conditions are constraints or algorithmic models that guide how to extract key information from specific types of data. They are dynamically matched based on the specific data types of the multi-source heterogeneous data and include parsing logic and parameter configuration. The role of the parsing conditions is to mask the format differences between different modalities of data, ensuring that each type of data can be correctly processed by the appropriate conditions. For example, when the data type is text, the corresponding parsing conditions can be set to extract entity words, relational words, and event descriptions using a natural language processing model; when the data type is an image, the parsing conditions are set to extract the image's color histogram, shape contour features, and texture description information; when the input data is a time-series signal, the parsing conditions can be set to extract the signal's peak amplitude, period frequency, and waveform rate of change. This data type-based differential parsing condition acquisition mechanism avoids the loss of key semantic information caused by general parsing strategies, providing high-quality structured input for subsequent feature fusion.

[0020] Furthermore, the data information in the aforementioned multi-source heterogeneous data is parsed according to the parsing conditions to obtain corresponding parsing information. Data information refers to the original content fragments contained in the multi-source heterogeneous data. Parsing information is structured descriptive data generated after processing according to the aforementioned parsing conditions. Its components include key feature words extracted from the text, association triples extracted from the graph, feature type descriptions identified from the image, and statistical parameters from time-series signals. The generation of parsing information involves executing parsing conditions that match the data type, converting unstructured or semi-structured raw data into a machine-readable standardized format. The role of parsing information is to preserve human-understandable semantic content, serving as the foundation for subsequent construction of interpretable knowledge graphs. For example, for text data containing car engine fault reports, key feature words such as "engine," "overheating," and "alarm" can be parsed as parsing information based on the parsing conditions; for an infrared thermal image of an equipment, descriptive information such as the high-temperature area being located in the upper left corner and the temperature gradient exhibiting a non-linear distribution can be parsed based on the parsing conditions. This step, by converting the raw data into standardized parsing information, achieves preliminary semantic alignment of multi-source heterogeneous data, eliminating format barriers between different data sources.

[0021] Furthermore, a vector encoder is a neural network model or mapping function pre-installed in the graph feature extraction rules, specifically designed to map structured parsed information to a high-dimensional vector space. Feature vectors are the mathematical expression of parsed information in vector space, obtained by encoding and calculating the parsed information using a vector encoder strictly corresponding to the current data type. The role of feature vectors is to provide a numerical representation that is efficiently computed by machines, supporting subsequent similarity matching and deep inference. Different data types correspond to different vector encoders to ensure the accuracy and generalization ability of feature representation. For example, for text parsed information, a text encoder based on the Transformer architecture can be called to generate semantic embedding vectors; for image parsed information, visual feature information mapping rules are used to map the parsed information to obtain mapped vectors. Through this configuration of dedicated encoders, the deep semantic features of data from various modalities can be fully captured, enabling the generated feature vectors to reflect the inherent deep semantic features of the data, providing a unified metric for cross-modal semantic fusion.

[0022] Furthermore, the graph features are the final output of this step, composed of the previously obtained parsed information and feature vectors. The combination can be achieved by concatenating or constructing a composite data structure containing both, aiming to simultaneously preserve the semantic readability and machine computability of the data. The graph features serve as the core unit for subsequent knowledge graph comparison, storage, and updates, supporting both manual verification and traceability, as well as automated vector retrieval and reasoning analysis. For example, the extracted text description of engine overheating (parsed information) can be bound to its corresponding high-dimensional semantic vector (feature vector) to form a complete graph feature object. Through this combination mechanism, a unified representation method for multimodal entities is constructed, bridging the semantic gap between text descriptions and image features while ensuring the consistency and authority of knowledge.

[0023] This application achieves differentiated feature extraction for multi-source heterogeneous data of different types through the above steps. By obtaining corresponding parsing conditions based on data type, targeted parsing of the original data is achieved, effectively avoiding the information loss problem caused by general parsing. By using a dedicated vector encoder matched to the data type to generate feature vectors, the accuracy and generalization ability of feature representation are significantly improved, especially when processing complex documents such as patents and reports, enhancing the accuracy of entity, relationship, and event extraction. By combining parsed information with feature vectors, the generated graph features balance semantic readability and machine processing capability, meeting the needs of intelligent agents such as DeepResearch for constructing multi-hop inference paths, and laying a solid foundation for cross-modal semantic fusion and accurate retrieval. On this basis, by combining columnar storage and vector quantization technology, the storage structure and access efficiency are further optimized, enabling the method to achieve low-latency knowledge updates and retrieval responses while ensuring high recall during execution.

[0024] S120. Based on a preset comparison strategy, the graph features are compared with the pre-stored knowledge graph to obtain the corresponding comparison results.

[0025] Further, the obtained graph features are compared with the knowledge graph using a comparison strategy. The knowledge graph, also known as a graph network, yields comparison results through feature comparison. The knowledge graph is a network-like knowledge structure composed of multiple graph nodes and their interrelationships, storing entity-attribute-relationship information in triples. The comparison strategy is a logical rule used to evaluate the similarity between new input graph features and existing knowledge graphs, determining whether new data is new, supplementary, or repetitive knowledge. The comparison results are classification labels derived from feature matching, specifically including new, supplementary, or repetitive. This step, through a triple-based comparison learning framework, achieves intelligent identification of new knowledge content. By quantifying the similarity between new data and the existing knowledge base, the processing path for knowledge can be automatically determined, avoiding redundant data storage while ensuring the dynamic evolution capability of the knowledge graph. This effectively reduces knowledge base storage redundancy and guarantees the consistency and authority of knowledge content.

[0026] In a specific embodiment, step S120 includes the following sub-steps: selecting candidate graph nodes associated with the parsing information from the knowledge graph based on the parsing information in the graph features; obtaining the feature matching degree between the node features of the candidate graph nodes and the feature vectors in the graph features; and obtaining the comparison classification corresponding to the feature matching degree according to the comparison strategy, as the corresponding comparison result.

[0027] First, candidate graph nodes associated with the parsed information can be selected from the knowledge graph based on the parsed information in the graph features. The parsed information is a set of key feature descriptions obtained after parsing and processing the aforementioned multi-source heterogeneous data. By matching the keywords in the parsed information with the node attributes of the graph nodes in the knowledge graph, the graph nodes can be retrieved. If a graph node's node attribute contains a keyword from the parsed information, then that graph node is determined to be a candidate node associated with the parsed information. This process of selecting candidate graph nodes aims to use the parsed information as a query index to perform coarse-grained filtering in the pre-stored knowledge graph, thereby narrowing the search scope for subsequent comparisons. Nodes that semantically match the parsed information are found by traversing the node attributes (labels or attribute fields) of the graph nodes in the knowledge graph. This preliminary filtering mechanism based on parsed information effectively avoids indiscriminate traversal and comparison of the entire knowledge graph, significantly reducing computational overhead and providing basic data support for subsequent fine-grained matching.

[0028] Further, feature matching degree is calculated based on the node features of the candidate graph nodes and the feature vectors in the graph features. Node features are the mathematical representations of candidate graph nodes in vector space, typically generated by a vector encoder from the node's historical data; feature vectors refer to the vector representations generated after parsing the current input multi-source heterogeneous data. Feature matching degree is used to quantify the semantic proximity between new input data and existing graph nodes, and it is typically obtained using algorithms such as cosine similarity calculation, Euclidean distance measurement, or dot product operation. Specifically, the node features of each candidate graph node selected above are compared one by one with the feature vectors in the current graph features. For example, let the node features of candidate graph node A be V. a The feature vector of the current map feature is V. b Then, through the formula Cosine(V) a V b The cosine of the angle between the two features is calculated, and this value is the feature matching degree, ranging from [0, 1]. The larger the value, the more semantically similar the two features are. In this process, the synergistic effect of node features and feature vectors is reflected in the following: node features provide a benchmark reference for existing knowledge, while feature vectors represent the new knowledge form to be fused. The alignment of their vector spaces allows the semantic consistency of cross-modal data to be accurately measured mathematically, thereby realizing the leap from symbol matching to semantic matching.

[0029] Furthermore, the comparison strategy is a pre-defined set of rules used to map continuous numerical feature matching degrees to discrete business decision categories. Comparison classification includes, but is not limited to, addition, and duplication types. The comparison result is the final determined classification label, used to guide subsequent knowledge base storage and graph update operations. The specific setting of the comparison strategy is usually based on threshold range division, that is, setting different matching degree thresholds to define different semantic relationship strengths. For example, when the calculated feature matching degree is less than 0.5, the comparison strategy classifies it as addition, indicating that the current multi-source heterogeneous data differs significantly from the nodes in the existing knowledge graph, belonging to entirely new knowledge content; when the feature matching degree is between 0.5 and 0.8, it is classified as addition, indicating that the new data enriches or corrects existing node knowledge; when the feature matching degree is greater than 0.8, it is classified as duplication, indicating that the new data highly overlaps with existing knowledge. Through the execution of the above comparison strategy, the logical relationship between new data and existing knowledge can be automatically identified, thereby generating clear comparison results. This classification process not only solves the redundancy problem caused by traditional full reconstruction, but also provides a direct decision-making basis for subsequently adding importance labels in a differentiated manner and performing incremental updates, ensuring the accuracy and efficiency of knowledge graph evolution.

[0030] This application constructs an intelligent knowledge fusion and discrimination mechanism through the synergistic effect of the above steps. First, it rapidly identifies candidate graph nodes from massive knowledge graphs using parsed information, achieving precise narrowing of the retrieval scope. Then, by calculating the feature matching degree between the node features of the candidate graph nodes and the feature vectors of the new data, it completes the leap from coarse-grained keyword matching to fine-grained semantic similarity quantification. Finally, using a pre-defined comparison strategy, the quantified matching degree is transformed into a comparative classification with clear business meaning. This progressive processing method can accurately determine whether multi-source heterogeneous data should be added as new knowledge, supplemented as existing knowledge, or deemed redundant, thereby effectively avoiding the disorderly expansion of the knowledge base, improving the dynamic adaptive capability and data quality of the knowledge graph, and laying a solid data foundation for high-quality reasoning and analysis by subsequent intelligent agents such as DeepResearch.

[0031] S130. Based on the comparison results, store the multi-source heterogeneous data in a preset knowledge base and update the knowledge graph.

[0032] Based on the comparison results, the aforementioned multi-source heterogeneous data is stored in a knowledge base, and the knowledge graph is updated based on the comparison results and the multi-source heterogeneous data. The knowledge base is a storage module used for persistently storing multi-source heterogeneous data and its metadata, supporting columnar storage and vector quantization techniques. Each piece of knowledge data in the knowledge base contains an importance tag, which is a metadata marker attached to the multi-source heterogeneous data and used to identify the degree of contribution of the data to the knowledge system. The update type refers to the specific operation category performed on the knowledge graph, including node addition, attribute supplementation, or ignoring. This step achieves coordinated linkage between data storage and graph updates. By introducing importance tags, not only can data archiving be achieved, but a weighting basis is also provided for subsequent retrieval ranking and response generation. Through adaptive updates to the knowledge graph, it is ensured that the knowledge base can reflect the latest technological developments and industry information in real time, supporting the urgent need for timely knowledge by intelligent agents such as DeepResearch.

[0033] In a specific embodiment, step S130 includes the following sub-steps: adding corresponding importance tags to the multi-source heterogeneous data according to the comparison results; adding the multi-source heterogeneous data with the added importance tags to the knowledge base for storage; determining the corresponding update type according to the comparison results; and updating the knowledge graph according to the update type and the multi-source heterogeneous data.

[0034] Specifically, importance labels are metadata identifiers used to characterize the relative value level of multi-source heterogeneous data in subsequent reasoning. These labels are dynamically determined based on the comparison results obtained above, which reflect the overlap or complementarity between new input data and existing knowledge in the pre-stored knowledge graph. Specifically, when the comparison result indicates that the input data is entirely new knowledge, a higher importance label is assigned; when it indicates supplementary knowledge, a medium importance label is assigned; and when it indicates duplicate knowledge, a lower importance label is assigned. For example, if the comparison result is new, i.e., the feature matching degree is less than a preset first threshold (e.g., 0.5), it indicates that the multi-source heterogeneous data contains entirely new entities or relationships not yet included in the knowledge graph, and the importance label is set to important; if the comparison result is supplementary, i.e., the feature matching degree is between the first and second thresholds (e.g., 0.5 to 0.8), it indicates that the data can improve existing node attributes, and the importance label is set to medium; if the comparison result is duplicated, i.e., the feature matching degree is greater than the second threshold (e.g., 0.8), it indicates that the data content already exists, and the importance label is set to unimportant. This differential labeling mechanism based on comparison results enables the initial screening and classification of information value during the data storage stage. This allows high-value new knowledge to receive higher priority in subsequent retrieval and reasoning, while low-value redundant data is marked to reduce its interference weight.

[0035] Furthermore, a knowledge base is a structured storage unit used for persistently storing multi-source heterogeneous data with metadata identifiers. This step aims to formally incorporate value-assessed data into the scope of retrieval and analytical reasoning, ensuring data traceability and integrity. Specifically, the storage process not only preserves the original multi-source heterogeneous data content (such as text paragraphs, image files, time-series signal fragments, etc.), but also writes the generated importance tags as index keys or associated attributes. For example, for a recent automotive industry patent text marked as important, its content is bound to an importance tag and stored in a columnar storage structure, and a unique index address pointing to the data is established; while duplicate news articles marked as unimportant are also stored but marked with low priority so that they can be processed first when storage space is tight or when data cleaning is required. By embedding importance tags into the storage structure, the knowledge base transforms from a simple data warehouse into an intelligent storage pool with value judgment capabilities, providing basic data support for subsequent weight-based reasoning analysis.

[0036] Furthermore, the update type refers to the specific set of operation instructions executed on the knowledge graph structure, which directly depends on the classification of the comparison results. The core logic of this step is to transform the comparison conclusions at the data level into change strategies at the graph topology level. Specifically, several update type mapping rules are preset: when the comparison result is "addition," the determined update type is node creation and relationship building, which means that new entity nodes need to be instantiated in the graph and corresponding edge connections need to be established; when the comparison result is "supplementation," the determined update type is attribute appending or weight correction, which means that no new nodes need to be created, only the attribute fields of existing nodes need to be filled or the confidence of existing relationships need to be adjusted; when the comparison result is "duplicate," the determined update type is "ignore" or "timestamp refresh," which means that the graph structure remains unchanged, and only the last access time of the relevant nodes is updated to maintain the activity record. For example, if the feature vector parsed from the input data has a very low match with a candidate node in the graph, the update type is determined to be node creation. A new sensor model node is added to the industrial equipment subgraph of the graph, and the index address of the data information is associated with the newly created graph node. Conversely, if the match is moderate but some parameter descriptions are missing, the update type is determined to be attribute supplementation. The latest operating temperature range of the sensor is written into the node attributes of an existing node, and the index address of the data information is associated with the corresponding graph node. If the match is high, the update type is determined to be ignore processing. The index address is directly associated with the corresponding graph node, but the actual content of the graph node is not updated.

[0037] The aforementioned knowledge graph update process is also the key step in executing the dynamic evolution of the knowledge graph, utilizing a dynamic topology mechanism driven by a lightweight neural network to achieve real-time iteration of the graph structure. Specifically, based on the determined update type, the corresponding graph operation operators are invoked to integrate effective information from multi-source heterogeneous data into the knowledge graph. If the update type is node creation, the system extracts entity names, attribute values, and relationship descriptions from the multi-source heterogeneous data, generating new graph nodes and connecting edges in the graph to expand the knowledge network; if the update type is attribute supplementation, the system locates the target graph node, writes the supplementary information from the multi-source heterogeneous data into the node attributes of the target graph node, and uses an incremental learning algorithm to fine-tune the feature vector representation of the point to reflect the latest knowledge state; if the update type is ignore processing, no substantive content update is performed on the graph node.

[0038] This application constructs a closed-loop knowledge management mechanism—comparative evaluation, value labeling, hierarchical storage, and targeted updates—through the synergistic effect of the aforementioned steps. By dynamically adding importance tags based on comparison results, the system automates the identification of the value of multi-source heterogeneous data, solving the problem of key information being overwhelmed due to treating data equally in traditional solutions. By storing labeled data in the knowledge base, the traceability and management granularity of the data are enhanced, enabling subsequent reasoning processes to allocate different attention weights based on the tags. Furthermore, by tightly binding the update type with the comparison results, the evolution strategy of the knowledge graph is ensured to match the actual contribution of new data, effectively preventing erroneous merging or loss of key information. Based on this, combined with a dynamic topology mechanism driven by a lightweight neural network, this solution supports high-frequency real-time iteration in an incremental learning environment, significantly improving the accuracy and timeliness of the knowledge graph in cold-start scenarios.

[0039] S140. If the input search request is received, the knowledge graph is searched according to the search request to obtain the corresponding search information.

[0040] It can further receive search requests input by users. The user inputting multi-source heterogeneous data and the user inputting the search request can be different users or the same user. Based on the received search request, the knowledge graph is searched to obtain the corresponding search information. Here, the search request refers to the natural language query instruction or multimodal query instruction initiated by the user; the request feature refers to the key semantic vector extracted from the search request for matching the knowledge graph. The search information refers to the target graph node and its contextual information obtained from the knowledge graph that matches the request, including associated nodes, attributes, tags, and node features.

[0041] In a specific embodiment, step S140 includes the following sub-steps: extracting request features corresponding to the retrieval request; retrieving the knowledge graph based on the request features to obtain target graph nodes that match the request features; and obtaining the graph association information of the target graph nodes in the knowledge graph as the corresponding retrieval information.

[0042] Request features, in this context, refer to data representations that characterize the core semantic intent of a retrieval request. They originate from converting the retrieval request described in natural language into a numerical representation in a vector space. Specific forms of request features can be sets of key terms, semantic vectors, or a combination of both. Their function is to map unstructured user queries into the same semantic space as knowledge graph nodes for accurate matching. Request features are obtained by segmenting, denoising, and vectorizing the retrieval request using a pre-built feature extraction module. For example, when the retrieval request is to analyze the latest technology of new energy vehicle battery thermal management systems, the keywords "new energy vehicle," "battery thermal management system," and "technology" are first extracted. Then, a pre-trained semantic encoder is used to map these keywords into high-dimensional semantic vectors, which serve as the request features.

[0043] Furthermore, a target graph node refers to a graph entity in the pre-stored knowledge graph whose node feature vector meets a preset similarity threshold with the requested feature generated in the current step. The acquisition of target graph nodes is based on vector similarity calculation. Specifically, the system compares the requested feature vector with the node feature vectors of all graph nodes in the knowledge graph, calculates cosine similarity or Euclidean distance, and identifies nodes with a similarity higher than a set threshold (e.g., 0.85) as target graph nodes. The role of target graph nodes is to serve as the core evidence source for the search results, representing the knowledge entities most relevant to the user's query. For example, if the requested feature vector points to battery thermal management, the system will traverse the corresponding graph nodes in the graph, selecting nodes with attributes of liquid cooling system, phase change material, or thermal runaway protection that are closest to the requested feature as target graph nodes.

[0044] Furthermore, graph association information refers to a multi-dimensional data set that has direct connections or logical relationships with the target graph node within the knowledge graph topology. Specifically, graph association information includes the target graph node's attribute data (such as creation time and data source), tag information, node feature vectors, and the attributes, tags, and relationship types of adjacent nodes (one-hop or multi-hop related nodes) connected to the target graph node. The source of graph association information is the node data and inter-node relationship information stored in the knowledge graph. Its function is to provide context, expanding isolated entities into a structured evidence network to support deep reasoning. Graph association information can be obtained by dynamically aggregating the node information of closely related nodes by traversing the edge relationships in the graph based on the target graph node's ID index.

[0045] S150. Based on the preset reasoning analysis model and the retrieval request, the retrieval information is subjected to reasoning analysis to obtain the corresponding reasoning analysis information.

[0046] Based on a pre-set reasoning analysis model and the user's input search request, the retrieved information obtained in the previous steps is analyzed to obtain reasoning analysis information. The reasoning analysis model is a deep learning model that incorporates a classification neural network and a reasoning analysis neural network, possessing semantic understanding and logical deduction capabilities. The reasoning category refers to the classification label of the search request intent, such as statistical analysis, logical reasoning, drill-down to key points, or overall summarization. Attention parameter configuration refers to dynamically adjusting the weight matrix of the attention layer in the reasoning analysis neural network according to the reasoning category, focusing on different feature dimensions. Reasoning analysis information refers to the structured conclusions generated after deep reasoning, including the reasoning coefficients corresponding to each target knowledge node (target knowledge nodes include target graph nodes and related nodes closely associated with target graph nodes). This step achieves customized reasoning for different reasoning tasks through a dynamic attention mechanism; through classification guidance and adaptive parameter configuration, the model can flexibly adjust its focus, thereby uncovering deep-seated implicit relationships in complex knowledge networks. This significantly improves the accuracy and relevance of reasoning analysis, providing core algorithmic support for generating interpretable response information.

[0047] In a specific embodiment, step S150 includes the following sub-steps: classifying the retrieval request according to the classification neural network in the reasoning analysis model to determine the corresponding reasoning category; configuring the attention parameters of the attention layer of the reasoning analysis neural network in the reasoning analysis model according to the reasoning type; and performing reasoning analysis on the retrieval information according to the reasoning analysis neural network with the attention parameters configured to obtain the corresponding reasoning analysis information.

[0048] Specifically, a classification neural network refers to a sub-network module pre-installed in the inference analysis model to identify the type of user intent. Its function is to receive retrieval requests in text form and output corresponding inference category labels. These inference categories are task types categorized based on the semantic features of the retrieval request, including statistical analysis, logical reasoning, drill-down to key points, and overall summarization. The classification neural network is trained using pre-labeled sample datasets of different inference task types. Internally, it contains multilayer perceptrons or convolutional structures to analyze the features of keywords and syntactic structures in the retrieval request and determine the corresponding inference category. The classification neural network includes multiple input nodes, which input the request features corresponding to the retrieval request. It also includes multiple output nodes, each outputting the matching degree with a corresponding inference category. The category corresponding to the output node with the highest matching degree is taken as the inference category corresponding to the retrieval request. An association analysis layer is also configured between the input and output nodes. This classification mechanism accurately captures the user's specific intent in deep inference analysis scenarios, providing clear guidance for subsequent differentiated inference processing.

[0049] Furthermore, a set of configuration parameters matching the reasoning type can be obtained to configure the attention parameters of the attention layer in the inference analysis neural network; the numerical values ​​of the parameters differ for different reasoning types. Attention parameter configuration refers to the process of dynamically adjusting the internal weight matrix or bias vector of the attention layer in the inference analysis neural network based on the determined reasoning category. The inference analysis neural network is a deep learning model that performs core reasoning tasks, and its attention layer is used to calculate the relevance weights of different parts of the input information. The attention parameters are obtained by matching the reasoning type with a pre-stored attention mapping table; different reasoning types correspond to different attention matrices to guide the model to focus on different types of knowledge paths. Specifically, if the reasoning type is logical reasoning, the configured attention parameters will enhance the model's sensitivity to logical connectives such as causal chains and conditional constraints; if the reasoning type is statistical analysis, the attention parameters will emphasize the weight allocation of numerical attributes, time series, and distribution features. Through this dynamic configuration method, the same neural network architecture can flexibly adapt to multiple reasoning types without retraining, significantly improving the model's ability to focus on specific tasks.

[0050] Furthermore, the inference analysis information refers to the structured results output after processing by a customized neural network, including inference coefficients corresponding to each target knowledge node. Inference coefficients are quantitative indicators used to characterize the contribution or confidence level of each target knowledge node in the current inference task, with a numerical range of [0,1]. This step is performed based on an inference analysis neural network with pre-configured attention parameters. The network uses an adjusted attention distribution to comprehensively analyze the node information of each target knowledge node in the retrieved information. By acquiring the node information of each target knowledge node and extracting the corresponding node information vector, and inputting it into each input unit of the inference analysis neural network with configured attention parameters, the number of input units in the inference analysis neural network can be configured to be equal to the number of target knowledge nodes. Through correlation analysis, the inference coefficients corresponding to each target node are output from the output nodes of the inference analysis neural network, thus the number of output nodes in the inference analysis neural network can be configured to be equal to the number of target knowledge nodes. For example, in the key point drill-down mode, the configured neural network ignores general background descriptions and focuses on deep attributes directly related to the core entity, thereby outputting high-weight inference coefficients to those target knowledge nodes with strong correlations.

[0051] This application utilizes a classification neural network to accurately categorize retrieval requests, identifying specific reasoning categories. Based on these categories, the attention layer of the reasoning analysis neural network is configured with targeted parameters, achieving a two-stage reasoning mechanism: first identifying intent, then customizing the model. Through the synergistic cooperation of the classification neural network and attention parameter configuration technology, the reasoning analysis neural network can dynamically adjust its internal information focus. For example, it can strengthen the weight of causal paths in logical reasoning tasks or focus on numerical features in statistical tasks, effectively solving the problem of generalized results caused by the inability of general reasoning models to adapt to diverse reasoning analysis tasks. Building on this, the configured neural network performs in-depth analysis of the retrieved information. The reasoning coefficients in the output reasoning analysis information accurately reflect the contribution of each target knowledge node in the current context. This not only significantly improves the accuracy and relevance of the reasoning results but also provides quantifiable weighting criteria for subsequent response generation, enhancing the interpretability and credibility of the entire reasoning analysis process.

[0052] S160. Generate reasoning response information corresponding to the reasoning analysis information according to the preset response generation rules and the knowledge base.

[0053] Furthermore, based on the response generation rules and the latest updated knowledge base, inference response information corresponding to the inference analysis information is generated, and this inference response information is output as feedback information corresponding to the retrieval request. The response generation rules refer to the logical templates and weight calculation strategies used to integrate multi-source information and generate the final response. Target knowledge data refers to the original data fragments retrieved from the knowledge base that are directly related to the target knowledge nodes, such as document paragraphs, image slices, or data records. Inference weight refers to the value calculated based on the inference coefficient and importance label, used to determine the presentation priority of each knowledge data in the final response. Inference response information refers to the final generated structured and traceable response content, including conclusions, evidence chains, and source citations. This step achieves a closed loop from data retrieval to knowledge service; by organically combining raw data, importance labels, and inference weight, the system-generated response not only accurately answers the user's question but also provides detailed evidence support and logical derivation processes. This enhances the credibility and interpretability of the output results of the inference analysis agent, effectively solving the pain point of traditional retrieval systems lacking in-depth analysis and fact-checking capabilities.

[0054] In a specific embodiment, step S160 includes the following sub-steps: obtaining target knowledge data matching the target knowledge node from the knowledge base according to the target knowledge node corresponding to the reasoning analysis information; configuring the reasoning weight of the target knowledge data according to the reasoning coefficients corresponding to each target knowledge node in the reasoning analysis information; and generating reasoning response information according to the response generation rules and the target knowledge data that has completed the reasoning weight configuration and includes an importance label.

[0055] In this context, the target knowledge node refers to the graph node identified in the previous reasoning and analysis step that is strongly related to the retrieval request, serving as a bridge connecting the reasoning logic and the original data. The target knowledge data refers to the original data fragments or structured information directly associated with the target knowledge node, stored in a pre-built knowledge base. Its sources include text paragraphs, image descriptions, time-series signal records, or graph triples formed after parsing and storing heterogeneous data from multiple sources. The target knowledge data is obtained by establishing a mapping relationship between the graph node index and the knowledge base storage address. Specifically, the system uses the unique identifier of the target knowledge node to search in the inverted index or vector index of the knowledge base, thereby locating and extracting the corresponding data content. The purpose of this step is to provide solid factual basis and content material for subsequent response generation, ensuring that the generated response information is not generated out of thin air but based on traceable original evidence.

[0056] Furthermore, the inference coefficient refers to the numerical indicator calculated by the inference analysis neural network during the inference analysis process, representing the degree of contribution or relevance of each target knowledge node to the current retrieval request. The inference weight refers to the priority coefficient assigned to the target knowledge data, used to quantify the importance level of the data in the final response generation process. The inference weight is calculated based on the inference coefficient through linear mapping or a nonlinear activation function. Specifically, the higher the inference coefficient, the stronger the relevance of the knowledge data corresponding to the node to the user's question, and the larger the configured inference weight; conversely, the lower the coefficient, the smaller the weight. This step enables dynamic sorting and differentiated processing of the selected knowledge content, allowing the system to distinguish between primary and secondary information when generating responses, placing highly relevant evidence at the core. For example, if, for a retrieval request for fault cause analysis, the inference analysis model calculates an inference coefficient of 0.9 for the sensor failure node and 0.4 for the environmental interference node, the system will configure the experimental data corresponding to the former with a high inference weight (e.g., 0.9) and the latter with a low inference weight (e.g., 0.4). By coordinating inference coefficients and inference weights, the transformation from qualitative reasoning to quantitative weighting is achieved, ensuring that the response accurately reflects the judgment results of the inference model.

[0057] Furthermore, the response generation rules refer to pre-defined algorithmic strategies or prompt word templates used to guide language organization, logical construction, and format standardization. These rules specify how to integrate knowledge data with different weights to form coherent natural language text. Importance labels refer to the tag information (e.g., important, medium, unimportant) added during the data storage stage based on comparison results. Inference response information refers to the final generated structured text response, containing conclusions, supporting evidence, and logical deduction processes. Inference response information is obtained by inputting the target knowledge data with configured inference weights and importance labels into the generation model and decoding it according to the response generation rules. The inference weights and importance labels are combined as the combination coefficient of the target knowledge data, which can be calculated using the formula (S). t ×S z +1 / e) 1 / 2 The process involves vectorizing the target knowledge data to obtain corresponding knowledge data vectors. These knowledge data vectors are then combined with the corresponding combination coefficients of the target knowledge data and input into the generative model. The model generates a response vector, decodes it, and ultimately obtains a text-based inference response. Specifically, when organizing the language, the generative model prioritizes data with high inference weights and high importance labels as core arguments. Data with low weights or low importance is used as supplementary explanations or background information, and may even be discarded in cases of conflict. This step aims to address the problem of neglecting differences in data sources in traditional response generation, leading to unclear focus. Through the synergistic effect of a dual weighting mechanism (inference weights reflecting the relevance between the target knowledge data and the inference category, and importance labels reflecting the importance and quality of the target knowledge data within the knowledge system), a clear, credible, and traceable professional response is generated. For example, when generating an inference response regarding maintenance recommendations for a certain type of equipment, the system selects fault case data highly relevant to the inference category based on high inference weights, and prioritizes high-quality operational procedures based on important importance labels. Finally, following the response generation rules of conclusion first, evidence support, and normative guidance, a complete maintenance recommendation text is output as the inference response. This significantly improves the professionalism and explainability of the responses.

[0058] This application, through the synergistic effect of the aforementioned technical features, realizes an intelligent response generation mechanism driven by a dual-dimensional approach of inference weights and importance tags. Specifically, firstly, it accurately locates original evidence in the knowledge base by targeting knowledge nodes, solving the problem of unsupported response content; secondly, it dynamically configures inference weights using inference coefficients, enabling the response content to respond in real time to subtle differences in the user's search intent, highlighting the status of highly relevant evidence; based on this, importance tags are introduced as long-term value verification factors, complementing the inference weights, ensuring both the timeliness and relevance of the response, while also considering the authority and stability of the knowledge. Finally, with the help of pre-defined response generation rules, the knowledge data, after dual screening and weighting, is organically integrated to output inference response information that is structurally rigorous, logically consistent, and highly traceable. This mechanism not only avoids the blurring of key points caused by information piling up in traditional methods, but also achieves a qualitative leap from information retrieval to intelligent induction through weight quantification, effectively supporting the needs of complex tasks such as automated deep retrieval and analytical inference.

[0059] This application constructs a lightweight, dynamically adaptive multimodal intelligent knowledge platform through the synergistic effect of the aforementioned technical features. First, by using graph feature extraction rules, multi-source heterogeneous data is transformed into unified graph features, resolving the semantic fragmentation problem of cross-modal data and providing a standardized input foundation for subsequent processing. Second, by utilizing a comparison strategy and a dynamic update mechanism of the knowledge graph, incremental learning and real-time evolution of knowledge are achieved, avoiding data redundancy while ensuring the timeliness and completeness of the knowledge base, significantly improving storage efficiency. Building upon this, by combining a hierarchical retrieval and dynamic attention configuration reasoning analysis model, the system can flexibly adjust the retrieval depth and reasoning direction according to user needs, achieving a leap from passive information retrieval to active evidence integration. Finally, by using response generation rules that integrate importance tags and reasoning weights, the high quality, traceability, and logical rigor of the output results are ensured. This entire process works in close coordination, effectively reducing cross-modal false positive rates and query latency, providing a powerful core knowledge service engine for the application of autonomous reasoning and analysis agents such as DeepResearch in complex scenarios such as technology, finance, and healthcare.

[0060] The multimodal knowledge retrieval and reasoning method based on graph networks disclosed in the above embodiments includes: extracting corresponding graph features from the input multi-source heterogeneous data, comparing the graph features with a pre-stored knowledge graph to obtain comparison results, storing the multi-source heterogeneous data in a knowledge base and updating the knowledge graph accordingly based on the comparison results, retrieving the knowledge graph according to a retrieval request to obtain retrieval information, performing reasoning analysis based on the retrieval request to obtain reasoning analysis information, and generating reasoning response information corresponding to the reasoning analysis information based on the knowledge base. This multimodal knowledge retrieval and reasoning method dynamically updates the knowledge base and knowledge graph by obtaining comparison results, acquires deep graph association information based on retrieval requests, and performs targeted reasoning analysis on the retrieval information to generate reasoning response information. It can perform in-depth reasoning analysis based on the logical relationships between knowledge content, significantly improving the relevance and accuracy of knowledge retrieval.

[0061] Figure 2 This is a schematic block diagram of a graph network-based multimodal knowledge retrieval and reasoning device provided in an embodiment of the present invention. Figure 2 As shown, corresponding to the above-described graph network-based multimodal knowledge retrieval and reasoning method, this invention also provides a graph network-based multimodal knowledge retrieval and reasoning apparatus. For details, please refer to... Figure 2 The graph network-based multimodal knowledge retrieval and reasoning device 700 includes: The spectral feature extraction unit 701 is used to receive the input multi-source heterogeneous data and extract the corresponding spectral features from the multi-source heterogeneous data according to the preset spectral feature extraction rules.

[0062] The comparison result acquisition unit 702 is used to compare the graph features with the pre-stored knowledge graph according to a preset comparison strategy to obtain the corresponding comparison results.

[0063] The knowledge graph update unit 703 is used to store the multi-source heterogeneous data into a preset knowledge base and update the knowledge graph according to the comparison results.

[0064] The retrieval information acquisition unit 704 is used to retrieve corresponding retrieval information from the knowledge graph based on the input retrieval request if it receives the input retrieval request.

[0065] The reasoning analysis information acquisition unit 705 is used to perform reasoning analysis on the search information according to the preset reasoning analysis model and the search request to obtain the corresponding reasoning analysis information.

[0066] The reasoning response information generation unit 706 is used to generate reasoning response information corresponding to the reasoning analysis information according to the preset response generation rules and the knowledge base.

[0067] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned graph network-based multimodal knowledge retrieval and reasoning device and its various units can be referred to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity, these details will not be repeated here.

[0068] The aforementioned multimodal knowledge retrieval and reasoning device based on graph networks can be implemented as a computer program, which can perform tasks such as... Figure 3 It runs on the electronic device shown.

[0069] Please see Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of the present invention. The electronic device 800 can be a terminal or a server. The terminal can be an electronic device with communication functions. The server can be a standalone server or a server cluster composed of multiple servers.

[0070] See Figure 3 The electronic device 800 includes a processor 802, a memory, and a network interface 805 connected via a system bus 801. The memory may include a non-volatile storage medium 803 and internal memory 804.

[0071] The non-volatile storage medium 803 can store an operating system 8031 ​​and a computer program 8032. The computer program 8032 includes program instructions that, when executed, cause the processor 802 to execute a multimodal knowledge retrieval and reasoning method based on graph networks.

[0072] The processor 802 provides computing and control capabilities to support the operation of the entire electronic device 800.

[0073] The internal memory 804 provides an environment for the execution of the computer program 8032 in the non-volatile storage medium 803. When the computer program 8032 is executed by the processor 802, the processor 802 can execute a multimodal knowledge retrieval and reasoning method based on graph networks.

[0074] This network interface 805 is used for network communication with other devices. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the electronic device 800 to which the present invention is applied. The specific electronic device 800 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0075] The processor 802 is used to run the computer program 8032 stored in the memory to implement the steps included in the above-mentioned multimodal knowledge retrieval and reasoning method based on graph networks.

[0076] It should be understood that, in this embodiment of the invention, the processor 802 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0077] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0078] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to perform the steps included in the above-described graph network-based multimodal knowledge retrieval and reasoning method.

[0079] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.

[0080] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0081] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0082] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0083] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0084] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A multi-modal knowledge retrieval and reasoning method based on a graph network, characterized in that, The method includes: Receive the input multi-source heterogeneous data, and extract the corresponding spectral features from the multi-source heterogeneous data according to the preset spectral feature extraction rules; The graph features are compared with the pre-stored knowledge graph according to the preset comparison strategy to obtain the corresponding comparison results; Based on the comparison results, the multi-source heterogeneous data is stored in a preset knowledge base and the knowledge graph is updated; If the input search request is received, the knowledge graph is searched according to the search request to obtain the corresponding search information; Based on the preset reasoning analysis model and the retrieval request, the retrieval information is reasoned and analyzed to obtain the corresponding reasoning analysis information; Based on the preset response generation rules and the knowledge base, inference response information corresponding to the inference analysis information is generated.

2. The graph network based multi-modal knowledge retrieval and reasoning method according to claim 1, characterized in that, The step of extracting corresponding spectral features from the multi-source heterogeneous data according to preset spectral feature extraction rules includes: Based on the data type of the multi-source heterogeneous data, obtain the parsing conditions corresponding to the data type from the graph feature extraction rules; The data information in the multi-source heterogeneous data is parsed according to the parsing conditions to obtain the corresponding parsing information; According to the vector encoder corresponding to the data type in the graph feature extraction rules, the feature vector corresponding to the parsed information is obtained; The parsed information is combined with the feature vector to obtain the corresponding spectral features.

3. The graph network based multi-modal knowledge retrieval and reasoning method according to claim 2, characterized in that, The step of comparing the graph features with the pre-stored knowledge graph according to a preset comparison strategy to obtain the corresponding comparison results includes: Candidate graph nodes associated with the parsing information are selected from the knowledge graph based on the parsing information in the graph features; Obtain the feature matching degree between the node features of the candidate graph nodes and the feature vectors in the graph features; The comparison classification corresponding to the feature matching degree is obtained according to the comparison strategy, and is used as the corresponding comparison result.

4. The graph network-based multi-modal knowledge retrieval and reasoning method according to any one of claims 1-3, characterized in that, The step of storing the multi-source heterogeneous data into a preset knowledge base and updating the knowledge graph based on the comparison results includes: Based on the comparison results, add corresponding importance labels to the multi-source heterogeneous data; The multi-source heterogeneous data with the added importance tags will be stored in the knowledge base; The corresponding update type is determined based on the comparison results; The knowledge graph is updated according to the update type and the multi-source heterogeneous data.

5. The multimodal knowledge retrieval and reasoning method based on graph networks according to claim 4, characterized in that, The step of retrieving the knowledge graph according to the retrieval request to obtain the corresponding retrieval information includes: Extract the request features corresponding to the search request; The knowledge graph is retrieved based on the request features to obtain target graph nodes that match the request features; The graph association information of the target graph node in the knowledge graph is obtained as the corresponding retrieval information.

6. The multimodal knowledge retrieval and reasoning method based on graph networks according to claim 5, characterized in that, The step of performing reasoning analysis on the search information based on a preset reasoning analysis model and the search request to obtain corresponding reasoning analysis information includes: The retrieval request is classified according to the classification neural network in the reasoning analysis model to determine the corresponding reasoning category; Configure the attention parameters of the attention layer in the inference analysis neural network of the inference analysis model according to the inference type; The retrieved information is analyzed by a reasoning analysis neural network configured with attention parameters to obtain corresponding reasoning analysis information.

7. The multimodal knowledge retrieval and reasoning method based on graph networks according to claim 6, characterized in that, The step of generating reasoning response information corresponding to the reasoning analysis information based on preset response generation rules and the knowledge base includes: Based on the target knowledge node corresponding to the reasoning and analysis information, target knowledge data matching the target knowledge node is obtained from the knowledge base; The inference weights of the target knowledge data are configured according to the inference coefficients corresponding to each target knowledge node in the inference analysis information. Based on the response generation rules and the target knowledge data that has completed the inference weight configuration and includes importance tags, inference response information is generated accordingly.

8. A multimodal knowledge retrieval and reasoning device based on graph networks, characterized in that, The apparatus is used to execute the multimodal knowledge retrieval and reasoning method based on graph networks as described in any one of claims 1-7, and the apparatus includes: The graph feature extraction unit is used to receive the input multi-source heterogeneous data and extract the corresponding graph features from the multi-source heterogeneous data according to the preset graph feature extraction rules. The comparison result acquisition unit is used to compare the graph features with the pre-stored knowledge graph according to a preset comparison strategy to obtain the corresponding comparison results. The knowledge graph update unit is used to store the multi-source heterogeneous data into a preset knowledge base and update the knowledge graph based on the comparison results. The information retrieval unit is used to retrieve corresponding information from the knowledge graph based on the input retrieval request if it receives the input retrieval request. The reasoning analysis information acquisition unit is used to perform reasoning analysis on the search information according to the preset reasoning analysis model and the search request to obtain the corresponding reasoning analysis information; The reasoning response information generation unit is used to generate reasoning response information corresponding to the reasoning analysis information according to the preset response generation rules and the knowledge base.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the multimodal knowledge retrieval and reasoning method based on graph networks as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which includes program instructions that, when executed by a processor, cause the processor to perform the multimodal knowledge retrieval and reasoning method based on graph networks as described in any one of claims 1-7.