A knowledge reasoning method, system, and device for knowledge-intensive tasks

By using text recognition fusion model and multimodal data correlation measurement model in knowledge inference, the problem of insufficient efficiency and accuracy of multimodal query text processing is solved, efficient and accurate knowledge reasoning is achieved, and user experience is improved.

CN119539094BActive Publication Date: 2025-06-17SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510105461.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-17
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process multimodal query text, especially in knowledge inference. Traditional methods cannot effectively integrate and process multimodal data, resulting in insufficient processing efficiency and accuracy.

Method used

The text recognition fusion model (obtained by BERT and TextCNN fusion) is used to identify the user's query intent, and a multimodal query data set is constructed through the multimodal data correlation metric model, data fusion and correlation metrics are performed, and finally input into the context-aware inference engine to obtain the inference results.

Benefits of technology

It improves the accuracy and efficiency of knowledge reasoning, can effectively handle multiple types of data modalities and knowledge source types, enhances the flexibility and adaptability of knowledge reasoning, and significantly improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119539094B_ABST
    Figure CN119539094B_ABST
Patent Text Reader

Abstract

The present application provides a knowledge reasoning method, system and device for knowledge-intensive tasks, belonging to the field of computer technology. The method inputs the multi-modal query text from the user terminal into a preset text recognition and fusion model to determine the corresponding user query intention and its corresponding matching knowledge source type; based on the matching knowledge source type and the multi-modal query text, multiple multi-modal query data are matched from the corresponding knowledge base to construct a multi-modal query data set. According to the preset modal type grouping, each multi-modal query data is divided and input into a multi-modal data association measurement model to determine the first association score between each multi-modal query data; based on each first association score and a first preset threshold, each multi-modal query data is subjected to a fusion process to obtain multi-modal fusion feature information; the multi-modal fusion feature information is input into a preset context-aware reasoning engine to obtain the reasoning result corresponding to the multi-modal query text according to the output result of the engine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a knowledge reasoning method, system, and device for knowledge-intensive tasks. Background Art

[0002] In the field of knowledge-intensive task processing, knowledge reasoning has always been the key to improving task processing efficiency and accuracy. Traditional knowledge reasoning methods often rely on single text input and a fixed knowledge base, and this processing method is unable to cope when faced with multi-modal query texts. Multi-modal query texts, that is, information containing multiple data types such as text, images, and audio, require the reasoning system to have more powerful data parsing and processing capabilities.

[0003] Although there have been some attempts to apply deep learning models to knowledge reasoning, most of these methods are limited to the processing of text data, and there is still a lack of effective means for the fusion and reasoning of multi-modal data. For example, although some text processing methods based on models such as BERT or TextCNN have achieved remarkable results in the field of natural language processing, they are inadequate when dealing with multi-modal data. These models often cannot effectively capture the features of non-text data such as images and audio, and are even less able to establish effective associations between multi-modal data.

[0004] Based on this, there is an urgent need for an efficient, accurate, and flexible knowledge reasoning technical solution for multi-modal data and knowledge-intensive tasks. Summary of the Invention

[0005] To solve the above problems, embodiments of this application provide a knowledge reasoning method, system, and device for knowledge-intensive tasks.

[0006] On the one hand, embodiments of this application provide a knowledge reasoning method for knowledge-intensive tasks, and the method includes:

[0007] Inputting the multi-modal query text from the user terminal into a preset text recognition fusion model to determine the corresponding user query intention and its corresponding matching knowledge source type; wherein, the text recognition fusion model is obtained by fusing a BERT model and a TextCNN model;

[0008] Based on the matching knowledge source type and the multi-modal query text, matching multiple multi-modal query data from the corresponding knowledge base to construct a multi-modal query data set;

[0009] Group the multi-modal query data according to the preset modal type, divide and input each piece of the multi-modal query data in the multi-modal query data set into a multi-modal data association metric model to determine the first association score between each piece of the multi-modal query data; wherein, the multi-modal data association metric model includes multiple network structures, and each network structure is used to process data corresponding to different preset modal type groups;

[0010] Based on the comparison results of each of the first association scores with a first preset threshold, perform fusion processing on the corresponding pieces of the multi-modal query data to obtain multi-modal fusion feature information;

[0011] Input the multi-modal fusion feature information into a preset context-aware inference engine to obtain an inference result corresponding to the multi-modal query text according to the output result of the engine.

[0012] In an implementation manner of the present application, inputting a multi-modal query text from a user terminal into a preset text recognition fusion model to determine a corresponding user query intention specifically includes:

[0013] Determine first output information and second output information of the multi-modal query text in the BERT model and the TextCNN model respectively through the text recognition fusion model;

[0014] Determine a first value matrix and a first key matrix according to the first output information, and determine a second value matrix and a second key matrix according to the second output information;

[0015] Calculate a first fusion weight corresponding to the first output information and a second fusion weight corresponding to the second output information according to a preset attention mechanism;

[0016] Determine fusion output information according to the sum of the first product value of the first fusion weight and the first value matrix and the second product value of the second fusion weight and the second value matrix, and match the user query intention in a preset intention list according to the fusion output information.

[0017] In an implementation manner of the present application, inputting a multi-modal query text from a user terminal into a preset text recognition fusion model to determine a corresponding user query intention and its corresponding matching knowledge source type specifically includes:

[0018] Convert the multi-modal query text into a structured query statement through semantic parsing technology;

[0019] Match the structured query statement with each knowledge source pre-classified in the knowledge base to determine the type of the matching knowledge source according to the matching result; wherein, the pre-classification is based on data structure, source, format or content classification; the type of the matching knowledge source includes at least one or more of the following: structured data, semi-structured data, and unstructured data.

[0020] In an implementation manner of the present application, based on the type of the matching knowledge source and the multimodal query text, match multiple multimodal query data from the corresponding knowledge base to construct a multimodal query data set, specifically including:

[0021] Encode the user query intention into an intention vector;

[0022] Calculate the similarity between the intention vector and each knowledge source vector corresponding to the type of the matching knowledge source in the knowledge base respectively as the knowledge source relevance score;

[0023] According to each knowledge source relevance score and the knowledge source weight formula , determine the knowledge source weight corresponding to each knowledge source; wherein, is the th knowledge source weight, is the th knowledge source relevance score, is the th knowledge source relevance score, is the total number of knowledge sources corresponding to the type of the matching knowledge source;

[0024] Use the knowledge source with a knowledge source weight greater than the second preset threshold as the query knowledge source;

[0025] Query the corresponding multimodal query data according to the structured query statement and each query knowledge source, and add the queried multimodal query data to the multimodal query data set.

[0026] In an implementation manner of the present application, group according to the preset modality type, divide each multimodal query data in the multimodal query data set and input it into a multimodal data association metric model to determine the first association score between each multimodal query data, specifically including:

[0027] Input the multimodal query data grouped by different preset modality types into the corresponding network structure for processing respectively to determine the corresponding query data feature set; wherein, the number of the multiple network structures is two;

[0028] Input each of the query data feature sets into a preset attention score function to determine the attention scores of each query data feature between every two of the query data feature sets, and obtain a corresponding number of query data feature groups; wherein, each query data feature group includes a first query data feature and a second query data feature.

[0029] Determine the first attention weight and the second attention weight corresponding to each query data feature group according to each of the attention scores and a preset weight formula.

[0030] Calculate a third product value of the first attention weight and the eigenvalue of the first query data feature, calculate a fourth product value of the second attention weight and the eigenvalue of the second query data feature, and use the sum of the third product value and the fourth product value as the second correlation score.

[0031] Input each of the second correlation scores into a preset Sigmoid function to determine the first correlation score between the corresponding multimodal query data.

[0032] In an implementation manner of the present application, based on the comparison results of each of the first correlation scores and a first preset threshold, perform fusion processing on the corresponding multimodal query data to obtain multimodal fusion feature information, which specifically includes:

[0033] Compare each of the first correlation scores with the first preset threshold.

[0034] In the case where the first correlation score is greater than the first preset threshold, splice and fuse the multimodal query data to obtain the multimodal fusion feature information.

[0035] Otherwise, update the user query intention and / or the matching knowledge source type.

[0036] In an implementation manner of the present application, before inputting the multimodal fusion feature information into a preset context-aware inference engine, the method further includes:

[0037] Embed the entities and relationships in the preset knowledge graph into a low-dimensional vector space to align with the vector representation of the pre-trained large language model.

[0038] Use the integrated preset knowledge graph and the pre-trained large language model as the preset context-aware inference engine.

[0039] In an implementation manner of the present application, the method further includes:

[0040] Determine knowledge base update information according to the inference result and the inference result feedback information from the user terminal.

[0041] Update the knowledge base according to the preset learning rate, the updated information of the knowledge base, and the original information of the knowledge base; the preset learning rate represents the influence degree of the updated information of the knowledge base on the original information of the knowledge base.

[0042] On the other hand, an embodiment of the present application further provides a knowledge reasoning system for knowledge-intensive tasks, and the system includes:

[0043] A first input module, configured to input the multi-modal query text from the user terminal into a preset text recognition fusion model to determine the corresponding user query intention and its corresponding matching knowledge source type; wherein, the text recognition fusion model is obtained by fusing a BERT model and a TextCNN model;

[0044] A matching module, configured to match multiple multi-modal query data from the corresponding knowledge base based on the matching knowledge source type and the multi-modal query text to construct a multi-modal query data set;

[0045] A partitioning input module, configured to partition and input each multi-modal query data in the multi-modal query data set into a multi-modal data association metric model according to a preset modal type grouping to determine a first association score between each multi-modal query data; wherein, the multi-modal data association metric model includes multiple network structures, and each network structure is used to process data corresponding to different preset modal type groupings;

[0046] A fusion processing module, configured to perform fusion processing on the corresponding multi-modal query data based on the comparison result between each first association score and the first preset threshold to obtain multi-modal fusion feature information;

[0047] A second input module, configured to input the multi-modal fusion feature information into a preset context-aware inference engine to obtain an inference result corresponding to the multi-modal query text according to the output result of the engine.

[0048] On yet another aspect, an embodiment of the present application further provides a knowledge reasoning device for knowledge-intensive tasks, and the device includes:

[0049] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a knowledge reasoning method for knowledge-intensive tasks as described above.

[0050] Compared with the prior art, the present application has the following remarkable effects:

[0051] By introducing a text recognition fusion model and a multi-modal data correlation metric model, this application can accurately process multi-modal query texts, effectively match and fuse multi-modal data, thereby improving the accuracy and efficiency of knowledge reasoning. It can also process various types of data modalities and adapt to different types of knowledge sources, thus enhancing the flexibility and adaptability of knowledge reasoning and enabling it to be widely applied to various knowledge-intensive tasks. In addition, by accurately identifying the user's query intention and quickly giving the reasoning result, this application can significantly improve the user experience and meet the expectations and needs of users for the knowledge reasoning system. The technical solution of this application introduces new technical ideas and methods in the field of knowledge reasoning, promotes the development and innovation of knowledge reasoning technology, and provides new ideas and directions for the research and application of related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The drawings described herein are used to provide a further understanding of this application and form a part of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0053] Figure 1 FIG. is a schematic flowchart of a knowledge reasoning method for knowledge-intensive tasks in an embodiment of this application;

[0054] Figure 2 FIG. is a schematic structural diagram of a knowledge reasoning system for knowledge-intensive tasks in an embodiment of this application;

[0055] Figure 3 FIG. is a schematic structural diagram of a knowledge reasoning device for knowledge-intensive tasks in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the protection scope of this application.

[0057] In the field of knowledge-intensive task processing, accurately integrating structured and unstructured data and performing effective reasoning has always been a technical challenge. With the continuous progress of artificial intelligence technology, especially in natural language processing, large language models have become important tools for achieving this goal. Although these models perform well in processing simple data and tasks, they still face challenges in understanding and generating accurate and effective reasoning results for complex tasks. How to intelligently perform knowledge fusion and effective reasoning, achieve adaptive fusion of knowledge, perform context-aware reasoning, learn the associations between different modal data, and dynamically update knowledge remains the key issue in solving accurate data fusion and effective reasoning in knowledge-intensive tasks.

[0058] Based on this, the embodiments of the present application provide a knowledge reasoning method, system, and device for knowledge-intensive tasks, which are used to solve the technical problem that it is currently difficult to perform knowledge reasoning on multi-modal data efficiently, accurately, and flexibly.

[0059] The following will describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0060] The embodiments of the present application provide a knowledge reasoning method for knowledge-intensive tasks, as Figure 1 shown, this method may include steps S101 - S105:

[0061] S101, the server inputs the multi-modal query text from the user terminal into a pre-set text recognition and fusion model to determine the corresponding user query intention and its corresponding matching knowledge source type.

[0062] Among them, the text recognition and fusion model is obtained by fusing a BERT model and a TextCNN model.

[0063] It should be noted that the server, as the execution subject of the knowledge reasoning method for knowledge-intensive tasks, is only exemplary. The execution subject is not limited to the server, and the present application does not make specific limitations on this.

[0064] That is to say, the present application uses the deep learning models BERT model and TextCNN model to fuse and construct a text recognition and fusion model, and uses the text recognition and fusion model obtained by fusing the two to perform recognition processing on the multi-modal query text corresponding to the natural language input by the user through the user terminal. The present application can perform fusion processing according to the outputs of the BERT model and the TextCNN model on the multi-modal query text respectively, for example, using an attention mechanism for fusion, or using ensemble learning for fusion.

[0065] Specifically, the present application uses an attention mechanism to achieve the above-mentioned output fusion and obtain the user's query intention. The multi-modal query text from the user terminal is input into a pre-set text recognition fusion model to determine the corresponding user query intention, which specifically includes:

[0066] Through the text recognition fusion model, determine the first output information and the second output information of the multi-modal query text in the BERT model and the TextCNN model respectively. Determine the first value matrix and the first key matrix according to the first output information, and determine the second value matrix and the second key matrix according to the second output information. According to the preset attention mechanism, calculate the first fusion weight corresponding to the first output information and the second fusion weight corresponding to the second output information. Determine the fusion output information based on the sum of the first product value of the first fusion weight and the first value matrix and the second product value of the second fusion weight and the second value matrix, so as to match the user query intention in the preset intention list according to the fusion output information.

[0067] In other words, the multi-modal query text is input into the BERT model, and the BERT model extracts the text features of the multi-modal query text and then encodes them into the first output information, which contains the multi-modal query text such as vocabulary, grammar or preliminary semantic information; subsequently, the first output information is input into the TextCNN model, and the TextCNN model will further extract the features in the semantic information and generate the second output information, such as containing deep semantics and context relationships. After the present application obtains the output information of the BERT model and the TextCNN model respectively, two groups of value matrices and key matrices corresponding to the two output information are constructed. The first output information and the second output information are generally in the form of vectors or matrices. Specifically, after obtaining the first output information, the server can perform various operations such as linear transformation, non-linear transformation, and dimension transformation, and generate an output matrix through a preset transformation weight matrix such as linear transformation, a preset transformation bias term, etc. The server can pre-store different preset transformation weight matrices and different preset transformation bias terms, and convert the first output information into the first value matrix and the first key matrix respectively; similarly, after obtaining the second output information, the server can convert the second output information into the second value matrix and the second key matrix through the above conversion method. Converting the output information into value matrices and key matrices is not limited to the above-mentioned implementation method of linear transformation, and can also be implemented by methods such as convolutional neural networks, attention splicing, and feature splicing. The present application does not make specific limitations on this.

[0068] Among them, the value (V) matrix and the key (K) matrix corresponding to the same output information are the same matrix. Subsequently, this application uses an attention mechanism to calculate the similarity between the query matrix and each key matrix based on a preset query (Q) matrix, and converts it into a fusion weight through the softmax function. Specifically, this application calculates the dot product or cosine similarity between Q and the first key matrix K1 and the second key matrix K2 respectively, and obtains two fusion weights through the softmax function. Subsequently, the first fusion weight w1 and the second fusion weight w2 obtained by calculation are used for fusion calculation of the fusion output information w1*V1 + w2*V2, where V1 is the first value matrix and V2 is the second value matrix. In addition, after obtaining the fusion output information, the server can input it into a pre-set classifier or decoder to complete the recognition of the user's query intention.

[0069] The output information of the BERT model and the TextCNN model described above, combined with the attention mechanism, can more comprehensively capture the key information in the multi-modal query text and improve the accuracy of user query intention recognition. The use of fusion weights enables the information output by different models to be weighted according to their importance, further improving the accuracy of intention recognition.

[0070] In another embodiment of this application, inputting the multi-modal query text from the user terminal into a preset text recognition fusion model to determine the corresponding user query intention and its corresponding matching knowledge source type specifically includes:

[0071] Converting the multi-modal query text into a structured query statement through semantic parsing technology. Matching the structured query statement with each knowledge source pre-defined and classified in the knowledge base to determine the matching knowledge source type according to the matching result. Among them, the pre-defined classification is based on data structure, source, format or content classification. The matching knowledge source type includes at least one or more of the following: structured data, semi-structured data, unstructured data.

[0072] That is to say, this application can quickly determine the matching knowledge source type according to the multi-modal query text, convert the multi-modal query text into a structured query statement through semantic parsing technology, simplify the query process, and improve the query efficiency. Matching with the knowledge sources pre-defined and classified in the knowledge base can quickly and accurately determine the matching knowledge source type, providing a basis for subsequent data matching and reasoning.

[0073] S102. The server matches multiple multi-modal query data from the corresponding knowledge base based on the matching knowledge source type and the multi-modal query text to construct a multi-modal query data set.

[0074] In the embodiments of the present application, the above-mentioned method of matching multiple multimodal query data from the corresponding knowledge base based on the matching knowledge source type and the multimodal query text to construct a multimodal query data set specifically includes:

[0075] Encode the user's query intention into an intention vector. Calculate the similarity between the intention vector and each knowledge source vector corresponding to the matching knowledge source type in the knowledge base respectively to score the knowledge source relevance. According to each knowledge source relevance score and the knowledge source weight formula , determine the knowledge source weight corresponding to each knowledge source. Wherein, is the th knowledge source weight, is the th knowledge source relevance score, is the th knowledge source relevance score, is the total number of knowledge sources corresponding to the matching knowledge source type. Take the knowledge sources with knowledge source weights greater than the second preset threshold as query knowledge sources. According to the structured query statement and each query knowledge source, query the corresponding multimodal query data respectively, and add the queried multimodal query data to the multimodal query data set.

[0076] That is to say, after determining the knowledge source type required by the user in the present application, it is possible to locate the knowledge source and query the relevant knowledge. The queried knowledge includes unstructured knowledge, semi-structured knowledge, and structured knowledge. The present application needs to efficiently match multimodal data related to the query intention from a huge knowledge base. Therefore, by calculating the similarity between the intention vector and the knowledge source vector, such as cosine similarity, reciprocal of Euclidean distance, etc., the knowledge source relevance score between the user's query intention and the knowledge source is obtained. Calculate the knowledge source weight through the above knowledge source weight formula. Subsequently, obtain the query knowledge source through the comparison result between the knowledge source weight and the preset second preset threshold. The second preset threshold can be set by the user during actual use, and the present application does not make specific limitations on this. Subsequently, the server queries each multimodal query data from the query knowledge source through the obtained structured query statement, such as images and texts.

[0077] The above method of calculating the similarity between the intention vector and the knowledge source vector and combining the knowledge source weight formula can screen out the knowledge sources highly relevant to the query intention, improving the accuracy and efficiency of data matching. Setting the second preset threshold further ensures the quality of the query knowledge source and reduces the interference of irrelevant data.

[0078] Among them, when the present application queries the knowledge source through a structured query statement and obtains relevant knowledge, for unstructured knowledge, semantic matching is performed; for semi-structured data, corresponding query technologies are used for matching relevant knowledge. At the same time, for the queried knowledge, a text fusion technology is used to integrate information from different knowledge sources into a coherent and consistent knowledge representation. For structured data, semantic web technology (RDF) is used to align and fuse concepts, entities, and relationships in different texts to form a unified semantic representation; then a text summarization algorithm (TextRank) is used to extract key information from multiple texts and generate a concise summary. Finally, a multi-modal query dataset is constructed by combining different modal data.

[0079] S103. The server groups according to the preset modal type, divides each multi-modal query data in the multi-modal query dataset, and inputs it into the multi-modal data association metric model to determine the first association score between each multi-modal query data.

[0080] Among them, the multi-modal data association metric model includes multiple network structures, and each network structure is used to process data corresponding to different preset modal type groups.

[0081] In the embodiment of the present application, grouping according to the preset modal type, dividing each multi-modal query data in the multi-modal query dataset, and inputting it into the multi-modal data association metric model to determine the first association score between each multi-modal query data specifically includes:

[0082] The multi-modal query data of different preset modal type groups are respectively input into the corresponding network structure for processing to determine the corresponding query data feature sets. Among them, the number of the multiple network structures is two. Each query data feature set is respectively input into the preset attention score function to determine the attention scores of each query data feature between every two query data feature sets, and a corresponding number of query data feature groups are obtained. Among them, the query data feature group includes a first query data feature and a second query data feature. According to each attention score and the preset weight formula, the first attention weight and the second attention weight corresponding to the query data feature group are determined respectively. Calculate the third product value of the first attention weight and the eigenvalue of the first query data feature, calculate the fourth product value of the second attention weight and the eigenvalue of the second query data feature, and take the sum of the third product value and the fourth product value as the second association score. Input each second association score into the preset Sigmoid function to determine the first association score between the corresponding multi-modal query data.

[0083] That is to say, the multi-modal data association metric model of this application adopts a two-stream network structure. One stream is a convolutional neural network (CNN) for extracting image features, and the other stream is a Transformer architecture for processing text data. The preset modal type grouping of this application includes an image group and a text group, and the multi-modal data association metric model with a two-stream network structure processes each multi-modal query data. The multi-modal data association metric model of this application can calculate the attention score between two multi-modal query data of different modal types in the query data feature set through a preset attention score function. The preset attention score function is as follows:

[0084]

[0085]

[0086] Among them, 、 respectively represent the attention scores for the feature fusion representation of text and image in the th group of query data feature groups. and are learnable weight matrices, represents element-wise multiplication, 、 are respectively preset bias terms; represents a function for converting text to the same dimension as the image, such as a fully connected layer.

[0087] Calculate the first attention weight and the second attention weight through a preset weight formula and the above attention score. The preset weight formula is as follows:

[0088]

[0089]

[0090] Among them, 、 are respectively the first attention weight and the second attention weight.

[0091] The server calculates to obtain the second association score . Subsequently, the server inputs it into to obtain the first association score corresponding to each multi-modal query data in the query data feature set.

[0092] By inputting multimodal query data into corresponding network structures for processing respectively, this application can extract their respective feature sets; using the attention score function and the weight formula, it can more accurately evaluate the correlation between different query data features, improving the accuracy of the correlation score. The use of the sigmoid function makes the correlation score smoother, facilitating subsequent data fusion processing.

[0093] S104. The server performs fusion processing on the corresponding multimodal query data based on the comparison results of the first correlation scores and the first preset threshold to obtain multimodal fusion feature information.

[0094] In the embodiment of this application, the above-mentioned performing fusion processing on the corresponding multimodal query data based on the comparison results of the first correlation scores and the first preset threshold to obtain multimodal fusion feature information specifically includes:

[0095] Compare each first correlation score with the first preset threshold respectively. When the first correlation score is greater than the first preset threshold, splice and fuse the multimodal query data to obtain multimodal fusion feature information. Otherwise, update the user query intention and / or the matching knowledge source type.

[0096] Among them, the first preset threshold can be set by the user according to the actual usage scenario. By setting the first preset threshold, it is possible to screen out multimodal query data with strong correlation for fusion, avoiding the interference of irrelevant data. The splicing and fusion method is simple and effective, which can retain the integrity of the original data and improve the quality of the fused data at the same time. When the correlation score does not meet the threshold requirement, by updating the user query intention or the matching knowledge source type, the accuracy of data matching can be further improved.

[0097] For example, the multimodal query text is: The sales order quantity and amount in August, sorted in descending order; through the above steps, this application uses the deep learning models BERT and TextCNN to identify the user input intention and determine which type of knowledge source the user needs, such as sales data, order information, etc. Through semantic parsing technology, the natural language query input by the user is converted into a structured query statement to more accurately locate the required knowledge source. For structured data sources (such as databases), use SQL query technology to retrieve the sales order data in August. For unstructured or semi-structured data sources (such as text reports or web pages), perform semantic matching and extraction of relevant information. Subsequently, splice the text and the picture in the same information. When the first correlation score in this application is less than or equal to the first preset threshold, it indicates that the obtained multimodal query data is not accurate enough. The server can re-execute the above steps of the user query intention and / or the matching knowledge source type to update the multimodal query data. After repeating the execution a predetermined number of times, an alarm message can be generated for relevant personnel to intervene for error repair.

[0098] S105, the server inputs the multi-modal fusion feature information into a preset context-aware inference engine to obtain the inference result corresponding to the multi-modal query text according to the output result of the engine.

[0099] In the embodiment of the present application, before the above-mentioned input of the multi-modal fusion feature information into the preset context-aware inference engine, the method further includes:

[0100] Embed the entities and relationships in the preset knowledge graph into a low-dimensional vector space to align with the vector representation of the pre-trained large language model. The integrated preset knowledge graph and the pre-trained large language model are used as the preset context-aware inference engine. The pre-trained large language model is, for example, Qwen2.

[0101] In the present application, by embedding the entities and relationships in the knowledge graph into a low-dimensional vector space and aligning with the vector representation of the pre-trained large language model, the structured information in the knowledge graph and the semantic understanding ability of the pre-trained large language model can be fully utilized, improving the inference ability of the inference engine. The integrated inference engine can more comprehensively understand the context information, thereby giving more accurate inference results.

[0102] In another embodiment of the present application, to maintain the timeliness and accuracy of the knowledge base, the method further includes:

[0103] Determine the knowledge base update information according to the inference result and the inference result feedback information from the user terminal. Update the knowledge base according to the preset learning rate, the knowledge base update information, and the original knowledge base information. The preset learning rate represents the influence degree of the knowledge base update information on the original knowledge base information.

[0104] In the present application, by collecting the inference result feedback information from the user terminal, errors or deficiencies in the knowledge base can be discovered in a timely manner. Combining the preset learning rate and the knowledge base update information can smoothly update the knowledge base, avoiding system instability or performance degradation caused by updates. Real-time updating of the knowledge base can ensure the accuracy and timeliness of the inference results, improving the overall performance of the system.

[0105] Through the above technical solution, by introducing a text recognition fusion model and a multi-modal data association measurement model, the present application can accurately process multi-modal query texts, effectively match and fuse multi-modal data, thereby improving the accuracy and efficiency of knowledge reasoning. It can also process various types of data modalities and adapt to different types of knowledge sources, thus enhancing the flexibility and adaptability of knowledge reasoning and enabling it to be widely applied to various knowledge-intensive tasks. In addition, by accurately identifying the user's query intention and quickly giving the reasoning result, the present application can significantly improve the user experience and meet the expectations and requirements of users for the knowledge reasoning system. The technical solution of the present application introduces new technical ideas and methods in the field of knowledge reasoning, promotes the development and innovation of knowledge reasoning technology, and provides new ideas and directions for the research and application in related fields.

[0106] Figure 2 FIG. is a schematic structural diagram of a knowledge reasoning system for knowledge-intensive tasks provided by an embodiment of the present application, as Figure 2 shown. The knowledge reasoning system 200 for knowledge-intensive tasks includes:

[0107] A first input module 201, configured to input a multi-modal query text from a user terminal into a preset text recognition fusion model to determine the corresponding user query intention and its corresponding matching knowledge source type. The text recognition fusion model is obtained by fusing a BERT model and a TextCNN model. A matching module 202, configured to match multiple multi-modal query data from a corresponding knowledge base based on the matching knowledge source type and the multi-modal query text to construct a multi-modal query data set. A partition input module 203, configured to partition and input each multi-modal query data in the multi-modal query data set into a multi-modal data association measurement model according to a preset modality type grouping to determine a first association score between each multi-modal query data. The multi-modal data association measurement model includes multiple network structures, and each network structure is used to process data corresponding to different preset modality type groupings. A fusion processing module 204, configured to perform fusion processing on the corresponding multi-modal query data based on the comparison result between each first association score and a first preset threshold to obtain multi-modal fusion feature information. A second input module 205, configured to input the multi-modal fusion feature information into a preset context-aware reasoning engine to obtain a reasoning result corresponding to the multi-modal query text according to the output result of the engine.

[0108] Figure 3 FIG. is a schematic structural diagram of a knowledge reasoning device for knowledge-intensive tasks provided by an embodiment of the present application, as Figure 3 shown. The device includes:

[0109] At least one processor; and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform:

[0110] Input the multi-modal query text from the user terminal into a preset text recognition fusion model to determine the corresponding user query intent and its corresponding matching knowledge source type. The text recognition fusion model is obtained by fusing a BERT model and a TextCNN model. Based on the matching knowledge source type and the multi-modal query text, match multiple multi-modal query data from the corresponding knowledge base to construct a multi-modal query data set. According to the preset modal type grouping, divide and input each multi-modal query data in the multi-modal query data set into a multi-modal data association metric model to determine the first association score between each multi-modal query data. The multi-modal data association metric model includes multiple network structures, and each network structure is used to process data corresponding to different preset modal type groupings. Based on the comparison results of each first association score and a first preset threshold, perform fusion processing on the corresponding multi-modal query data to obtain multi-modal fusion feature information. Input the multi-modal fusion feature information into a preset context-aware inference engine to obtain an inference result corresponding to the multi-modal query text according to the output result of the engine.

[0111] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system and device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0112] The systems and devices provided in the embodiments of this application correspond one-to-one with the methods. Therefore, the systems and devices also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and devices will not be elaborated here.

[0113] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.

[0114] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A knowledge reasoning method for knowledge-intensive tasks, characterized in that: The method comprises: Inputting the multimodal query text from the user terminal into a preset text recognition fusion model to determine the corresponding user query intention and its corresponding matching knowledge source type; wherein the text recognition fusion model is obtained by fusing the BERT model and the TextCNN model; Based on the matching knowledge source type and the multimodal query text, matching a plurality of multimodal query data from a corresponding knowledge base to construct a multimodal query data set; According to the preset modality type grouping, each of the multimodal query data in the multimodal query data set is divided and input into a multimodal data association measurement model to determine a first association score between each of the multimodal query data; wherein the multimodal data association measurement model includes a plurality of network structures, each of the network structures is used to process data corresponding to different preset modality type groups; Based on the comparison result of each of the first association scores and the first preset threshold, the corresponding multimodal query data are fused to obtain multimodal fusion feature information; Inputting the multimodal fusion feature information into a preset context-aware reasoning engine to obtain a reasoning result corresponding to the multimodal query text according to an engine output result; Among them, according to the preset modality type grouping, each of the multimodal query data in the multimodal query data set is divided and input into a multimodal data association measurement model to determine a first association score between each of the multimodal query data, specifically including: Inputting the multimodal query data grouped by different preset modality types into the corresponding network structures for processing respectively to determine the corresponding query data feature set; wherein the number of the plurality of network structures is two; Inputting each of the query data feature sets into a preset attention score function, determining the attention score of each query data feature between each of the query data feature sets, and obtaining a corresponding number of query data feature groups; wherein the query data feature group includes a first query data feature and a second query data feature; Determine the first attention weight and the second attention weight corresponding to the query data feature group respectively according to each of the attention scores and the preset weight formula; Calculating a third product value of the first attention weight and the feature value of the first query data feature, calculating a fourth product value of the second attention weight and the feature value of the second query data feature, and adding the third product value and the fourth product value as a second association score; Each of the second association scores is input into a preset Sigmoid function to determine the first association score between the corresponding multimodal query data.

2. A knowledge reasoning method for knowledge-intensive tasks according to claim 1, characterized in that: The multimodal query text from the user terminal is input into the preset text recognition fusion model to determine the corresponding user query intention, including: Determine, by means of the text recognition fusion model, first output information and second output information of the multimodal query text in the BERT model and the TextCNN model respectively; Determine a first value matrix and a first key matrix according to the first output information, and determine a second value matrix and a second key matrix according to the second output information; According to a preset attention mechanism, calculating a first fusion weight corresponding to the first output information and a second fusion weight corresponding to the second output information; The fusion output information is determined based on the sum of the first product value of the first fusion weight and the first value matrix and the second product value of the second fusion weight and the second value matrix to match the user query intent in the preset intent list according to the fusion output information.

3. A knowledge reasoning method for knowledge-intensive tasks according to claim 2, characterized in that: The multimodal query text from the user terminal is input into the preset text recognition fusion model to determine the corresponding user query intention and its corresponding matching knowledge source type, specifically including: Converting the multimodal query text into a structured query statement by using semantic parsing technology; The structured query statement is matched with each knowledge source of predefined classification in the knowledge base to determine the matching knowledge source type according to the matching result; wherein the predefined classification is based on data structure, source, format or content classification; the matching knowledge source type includes at least one or more of the following: structured data, semi-structured data, unstructured data.

4. A knowledge reasoning method for knowledge-intensive tasks according to claim 3, characterized in that: Based on the matching knowledge source type and the multimodal query text, a plurality of multimodal query data are matched from a corresponding knowledge base to construct a multimodal query data set, specifically including: Encoding the user query intent into an intent vector; Calculating the similarity between the intention vector and each knowledge source vector corresponding to the matching knowledge source type in the knowledge base to score the knowledge source relevance; According to the knowledge source relevance score and knowledge source weight formula , determine the knowledge source weights corresponding to each knowledge source; where, For the The weight of knowledge sources, For the The knowledge source relevance score, For the The knowledge source relevance score, The total number of knowledge sources corresponding to the matching knowledge source type; Taking the knowledge source whose knowledge source weight is greater than a second preset threshold as a query knowledge source; According to the structured query statement and each query knowledge source, the corresponding multimodal query data are queried, and each multimodal query data obtained by the query is added to the multimodal query data set.

5. The knowledge reasoning method for knowledge-intensive tasks according to claim 1, characterized in that: Based on the comparison result of each of the first association scores and the first preset threshold, the corresponding multimodal query data are fused to obtain multimodal fusion feature information, which specifically includes: Comparing each of the first association scores with the first preset threshold respectively; When the first association score is greater than the first preset threshold, concatenating and fusing the multimodal query data to obtain the multimodal fusion feature information; Otherwise, update the user query intention and / or the matching knowledge source type.

6. A knowledge reasoning method for knowledge-intensive tasks according to claim 1, characterized in that: Before inputting the multimodal fusion feature information into a preset context-aware reasoning engine, the method further includes: Embed entities and relations in the preset knowledge graph into a low-dimensional vector space to align with the vector representation of the pre-trained large language model; The integrated preset knowledge graph and the pre-trained large language model are used as the preset context-aware reasoning engine.

7. A knowledge reasoning method for knowledge-intensive tasks according to claim 1, characterized in that: The method further comprises: Determining knowledge base update information according to the inference result and the inference result feedback information from the user terminal; The knowledge base is updated according to a preset learning rate and the knowledge base update information and the original knowledge base information; the preset learning rate indicates the degree of influence of the knowledge base update information on the original knowledge base information.

8. A knowledge reasoning system for knowledge-intensive tasks, characterized in that: The system comprises: A first input module is used to input the multimodal query text from the user terminal into a preset text recognition fusion model to determine the corresponding user query intention and its corresponding matching knowledge source type; wherein the text recognition fusion model is obtained by fusing the BERT model and the TextCNN model; A matching module, configured to match a plurality of multimodal query data from a corresponding knowledge base based on the matching knowledge source type and the multimodal query text, so as to construct a multimodal query data set; A partitioning input module, used for partitioning each of the multimodal query data in the multimodal query data set according to the preset modal type grouping and inputting the multimodal data association measurement model to determine a first association score between each of the multimodal query data; wherein the multimodal data association measurement model includes a plurality of network structures, each of the network structures is used to process data corresponding to different preset modal type groups; A fusion processing module, configured to perform fusion processing on the corresponding multimodal query data based on the comparison result of each of the first association scores and a first preset threshold, so as to obtain multimodal fusion feature information; A second input module is used to input the multimodal fusion feature information into a preset context-aware reasoning engine to obtain a reasoning result corresponding to the multimodal query text according to an engine output result; Wherein, the partition input module is specifically used for: Inputting the multimodal query data grouped by different preset modality types into the corresponding network structures for processing respectively to determine the corresponding query data feature set; wherein the number of the plurality of network structures is two; Inputting each of the query data feature sets into a preset attention score function, determining the attention score of each query data feature between each of the query data feature sets, and obtaining a corresponding number of query data feature groups; wherein the query data feature group includes a first query data feature and a second query data feature; Determine the first attention weight and the second attention weight corresponding to the query data feature group respectively according to each of the attention scores and the preset weight formula; Calculating a third product value of the first attention weight and the feature value of the first query data feature, calculating a fourth product value of the second attention weight and the feature value of the second query data feature, and adding the third product value and the fourth product value as a second association score; Each of the second association scores is input into a preset Sigmoid function to determine the first association score between the corresponding multimodal query data.

9. A knowledge reasoning device for knowledge-intensive tasks, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a knowledge reasoning method for knowledge-intensive tasks as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Advanced search method based on semantic understanding

    CN117851444A

  • Domain large model multi-modal knowledge base construction method based on feature representation

    CN118779469A