Multimodal data fusion method and device based on knowledge enhancement, equipment and medium

By acquiring the positive sample pairs of the image-body text, semantic negative sample processing and scene graph knowledge enhancement, high-quality negative sample pairs are generated, and the knowledge enhancement coding model is optimized, which solves the insufficient semantic distinction ability of the multimodal model in structured information processing, and realizes efficient fusion and precise matching of multimodal data.

CN120509467APending Publication Date: 2025-08-19PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510578651.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing multimodal models cannot effectively distinguish semantic structural differences when processing structured information, resulting in poor performance in image-text matching and intelligent diagnostic tasks. The existing technology cannot accurately describe medical data and financial data, and it is difficult to achieve deep fusion of multimodal data.

Method used

By obtaining positive sample pairs of image-body text, semantic negative sample processing is performed, high-quality negative sample pairs are generated using scene graph knowledge enhancement technology, and encoding model is enhanced through loss function optimization knowledge enhancement, sample vector dot product comparison is performed, and the optimal matching result of multimodal data is finally output.

Benefits of technology

It significantly improves the model's learning and representation ability of structured information, solves the problem of insufficient multimodal data fusion, and improves the accuracy and reliability of matching results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509467A_ABST
    Figure CN120509467A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-modal data processing of financial and medical scenes, and discloses a multi-modal data fusion method, device and equipment based on knowledge enhancement and a medium, and the method comprises the steps: obtaining a positive sample pair of an image-text in advance, and carrying out the semantic negative sample processing to obtain a negative sample pair; performing scene graph knowledge enhancement processing on the positive sample pair and the negative sample pair, respectively inputting the positive sample pair and the negative sample pair into the initial model for sample vector dot product comparison training, performing training optimization through a constructed loss function, and outputting a knowledge enhancement coding model; and inputting to-be-matched data of the image group-text or the image-text group into the knowledge enhancement coding model for vector dot product comparison, and outputting an image and a sample corresponding to a maximum value of a vector dot product result in the to-be-tested data as an optimal matching result of multi-modal data fusion. According to the method, the high-quality negative sample is generated through the scene graph knowledge, and the scene graph knowledge is enhanced, so that the learning ability and the expression ability of the model to the structured information are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal data processing in financial and medical scenarios, and in particular to a multimodal data fusion method, apparatus, device and medium based on knowledge enhancement. Background Art

[0002] With the rapid development of financial technology and medical technology, multimodal data (such as text, images, and audio) is increasingly being used in the financial and medical fields. However, existing multimodal models have limitations when processing structured information and are unable to effectively distinguish differences in semantic structures, resulting in poor performance in tasks such as image-text matching and intelligent diagnosis. Furthermore, existing technologies cannot accurately describe medical and financial data when constructing medical and financial knowledge graphs, making it difficult to achieve deep integration of multimodal data. Summary of the Invention

[0003] The present invention provides a multimodal data fusion method, apparatus, device and medium based on knowledge enhancement to solve the problem of improving the ability of multimodal models in processing structured information.

[0004] In a first aspect, a multimodal data fusion method based on knowledge enhancement is provided, comprising:

[0005] Pre-acquire positive sample pairs of image and text and perform semantic negative sample processing to obtain negative sample pairs;

[0006] After performing scene graph knowledge enhancement processing on the positive sample pairs and the negative sample pairs, the pairs are respectively input into the initial neural network model for sample vector dot product comparison training, and the training and optimization are performed using the constructed loss function to output a knowledge enhancement coding model;

[0007] The data to be matched of the image group-text or image-text group is input into the knowledge enhanced coding model for vector dot product comparison, and the image and sample corresponding to the maximum value of the vector dot product result in the test data are output as the optimal matching result of multimodal data fusion.

[0008] In a second aspect, a multimodal data fusion device based on knowledge enhancement is provided, comprising:

[0009] The negative sample construction module is used to pre-acquire positive sample pairs of images and positive texts and perform semantic negative sample processing to obtain negative sample pairs;

[0010] A model training module is used to perform scene graph knowledge enhancement processing on the positive sample pairs and the negative sample pairs, and then input them into the initial neural network model for sample vector dot product comparison training, and output the knowledge enhancement coding model after training and optimization through the constructed loss function;

[0011] The data matching module is used to input the data to be matched of the image group-text or image-text group into the knowledge-enhanced coding model for vector dot product comparison, and output the image and sample corresponding to the maximum value of the vector dot product result in the test data as the optimal matching result of multimodal data fusion.

[0012] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned knowledge-enhanced multimodal data fusion method are implemented.

[0013] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned multimodal data fusion method based on knowledge enhancement are implemented.

[0014] In the scheme implemented by the above-mentioned multimodal data fusion method, device, equipment and medium based on knowledge enhancement, positive sample pairs of images and texts can be obtained in advance and semantic negative sample processing can be performed to obtain negative sample pairs; after the positive sample pairs and negative sample pairs are subjected to scene graph knowledge enhancement processing, they are respectively input into the initial neural network model for sample vector dot product comparison training, and after training and optimization through the constructed loss function, the knowledge enhancement coding model is output; the image group-text or image-text group to be matched data is input into the knowledge enhancement coding model for vector dot product comparison, and the image and sample corresponding to the maximum value of the vector dot product result in the test data are output as the optimal matching result of multimodal data fusion. In the present invention, high-quality negative sample pairs are generated through scene graph knowledge, which avoids the semantic invariance problem caused by random word exchange in traditional methods. This improvement significantly improves the model's ability to learn structured information and solves the problem of low quality of negative samples in the prior art. In the present invention, the samples are also subjected to scene graph knowledge enhancement processing, which further improves the structured representation ability of the knowledge enhancement coding model, solves the problem of insufficient multimodal data fusion in the prior art, and improves the model's ability to understand complex structured information. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0016] Figure 1 2 is a schematic diagram of an application environment of a multimodal data fusion method based on knowledge enhancement in one embodiment of the present invention.

[0017] Figure 2 4 is a flow chart of a multimodal data fusion method based on knowledge enhancement in one embodiment of the present invention.

[0018] Figure 3 yes Figure 2 A flowchart of a specific implementation of step S201 is shown in FIG.

[0019] Figure 4 yes Figure 2 A flowchart of a specific implementation of step S202 is shown in FIG.

[0020] Figure 5 yes Figure 4 A flowchart of a specific implementation of step S402 is shown in FIG.

[0021] Figure 6 yes Figure 4 A flowchart of a specific implementation of step S403 is shown in FIG.

[0022] Figure 7 4 is a structural diagram of a multimodal data fusion device based on knowledge enhancement in one embodiment of the present invention.

[0023] Figure 8 It is a structural diagram of a computer device in one embodiment of the present invention.

[0024] Figure 9 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0026] The multimodal data fusion method based on knowledge enhancement provided by the embodiment of the present invention can be applied in the following fields: Figure 1In an application environment, the client communicates with the server through a network. The server provides a model training node for the trainer to pre-acquire positive sample pairs of image-positive text in the training node and perform semantic negative sample processing to obtain negative sample pairs; after the positive sample pairs and negative sample pairs are subjected to scene graph knowledge enhancement processing, they are respectively input into the initial neural network model for sample vector dot product comparison training, and the knowledge enhancement coding model is output after training and optimization through the constructed loss function. The server can receive the image group-text or image-text group to be matched data given by the user through the client, and then input the data to be matched into the knowledge enhancement coding model for vector dot product comparison, and output the image and sample corresponding to the maximum value of the vector dot product result in the test data as the optimal matching result of multimodal data fusion. In the present invention, high-quality negative sample pairs are generated through scene graph knowledge, which avoids the semantic invariance problem caused by random exchange of words in traditional methods. This improvement significantly improves the model's ability to learn structured information and solves the problem of low quality of negative samples in the prior art. In this invention, the samples are also subjected to scene graph knowledge enhancement processing, further improving the structured representation capabilities of the knowledge-enhanced coding model, addressing the problem of insufficient multimodal data fusion in the prior art, and improving the model's ability to understand complex structured information. The client can include, but is not limited to, various interactive medical products and devices, financial products and devices, etc. The server can be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0027] See also Figure 2 As shown, Figure 2 A flowchart of a multimodal data fusion method based on knowledge enhancement provided in an embodiment of the present invention includes the following steps S201-S203.

[0028] S201, pre-acquire positive sample pairs of image and text and perform semantic negative sample processing to obtain negative sample pairs;

[0029] In step S201, multimodal data from the financial and medical fields, including images and text, must first be collected. The collected data is preprocessed, such as image normalization, text segmentation, and word vector representation, to obtain positive sample pairs of images and positive texts, i.e., one image corresponds to one positive text. The scene graph-based semantic negative sampling module then generates high-quality semantic negative sample pairs using scene graph knowledge, avoiding the semantic invariance problem caused by random word swapping in traditional methods. This improvement significantly enhances the model's ability to learn structured information and solves the problem of low quality negative sample pairs in existing technologies.

[0030] S202, after performing scene graph knowledge enhancement processing on the positive sample pairs and the negative sample pairs, respectively input them into the initial neural network model for sample vector dot product comparison training, and output the knowledge enhancement coding model after training optimization using the constructed loss function;

[0031] In step S202, before model training, the samples are subjected to scene graph knowledge augmentation processing, aiming to enhance the sample's expressiveness by leveraging the rich semantic information contained in the scene graph. Simultaneously, a loss function is designed to optimize model training, ensuring high accuracy while also possessing good generalization capabilities. The resulting knowledge-enhanced encoding model effectively improves the ability to learn and represent complex structured information.

[0032] S203. Input the data to be matched of the image group-text or image-text group into the knowledge-enhanced coding model for vector dot product comparison, and output the image and sample corresponding to the maximum value of the vector dot product result in the data to be tested as the optimal matching result of multimodal data fusion.

[0033] In step S203, the image group-text or image-text group given by the user is input into the trained knowledge-enhanced coding model, and the model will perform a vector dot product comparison on these data. Through calculation, the model will output the maximum value of the vector dot product result in the data to be matched, and the image and sample corresponding to the maximum value are the optimal matching results of multimodal data fusion. This step makes full use of the knowledge-enhanced coding model's ability to learn and represent complex structured information, ensuring the accuracy and reliability of the matching results. For example, in some medical scenarios, the doctor inputs a text instruction to the medical device, and the model needs to match the image that best matches the text instruction from the image group. Therefore, the model needs to perform a vector dot product comparison between the text instruction and the image group, and output the image corresponding to the maximum value of the vector dot product result as the optimal matching image for the text instruction, thereby providing accurate guidance for the medical device to execute the instruction.

[0034] Steps SS201-S203 achieve efficient fusion and precise matching of multimodal data in the financial and medical fields. First, by pre-acquiring and processing positive sample pairs of images and texts, and introducing a semantic negative sampling module based on scene graphs, the quality of negative sample pairs is significantly improved, thereby enhancing the model's ability to learn structured information. Secondly, after the positive and negative sample pairs are subjected to scene graph knowledge enhancement processing, they are input into the initial neural network model for training. By optimizing the model with the designed loss function, the output knowledge enhancement coding model has the ability to efficiently learn and represent complex structured information. Finally, the trained knowledge enhancement coding model is used to compare the vector dot product of images and texts, and the optimal matching result of multimodal data fusion is output, ensuring the accuracy and reliability of the matching results. The implementation of this series of steps not only improves the efficiency and accuracy of multimodal data processing, but also provides strong technical support for intelligent applications in fields such as finance and medicine.

[0035] In one embodiment, if Figure 3 As shown, Figure 3 yes Figure 2 A flow chart of a specific implementation of step S201 in FIG. 1 specifically includes the following steps S301 - S303 .

[0036] S301, obtaining a positive sample pair of image-text, wherein the positive text description is a representation of the meaning of the image;

[0037] S302, performing semantic negative sample processing on the positive text in the positive sample pair using scene graph knowledge to obtain negative text;

[0038] S303: Combine the negative text with the image in the positive sample pair to obtain a negative sample pair.

[0039] Steps S301-S303 combine the correspondence between the image and the positive text and utilize scene graph knowledge to perform semantic transformations, generating negative text that does not match the original image. This process not only enriches the training dataset but also, by introducing negative sample pairs, helps the model learn more accurate and robust image-text matching capabilities.

[0040] For example, in some medical scenarios, images may contain complex pathological features, while the corresponding positive text describes these features in detail. Using the semantic negative sampling module to swap the subject and object in a sentence to generate negative pairs may incorrectly associate a pathological feature with an unrelated text description. Such negative pairs force the model to learn to distinguish subtle semantic differences, thereby more accurately understanding the correspondence between image and text. For example, for a triple, if the positive text in the positive pair is extracted as (object 1, relation, object 2), the negative text in the generated negative pair can be extracted as (object 2, relation, object 1). For an attribute pair, if the positive text in the positive pair is extracted as (attribute 1, object 1) and (attribute 2, object 2), the negative text in the generated negative pair after swapping the attributes can be extracted as (attribute 2, object 1) and (attribute 1, object 2). Another example is in some financial scenarios, where an image may depict a complex financial statement, while the positive text accurately summarizes the key data in the statement. Negative pairs created using semantic negative sampling may intentionally confuse the interpretation of certain data points, matching them with incorrect text descriptions. This strategy enables the model to deeply analyze the precise connection between financial data and textual descriptions, enhancing its understanding and accuracy in complex financial scenarios. Through these carefully designed positive and negative sample pairs, the model continuously trains its image-to-text mapping capabilities, ensuring reliable and accurate information interpretation in a variety of real-world application scenarios.

[0041] In one embodiment, if Figure 4 As shown, step S202 includes the following steps S401-S404.

[0042] S401, performing scene graph knowledge extraction on the positive sample pairs and the negative sample pairs respectively to obtain corresponding knowledge triples;

[0043] In step S401, before the positive and negative sample pairs are fed into the model for training, the objects, attributes, and relationships in the scene graph are extracted and used as additional input to the model. To this end, a self-attention mechanism (Transformer) is added to the model for knowledge augmentation (see step S503 below for details).

[0044] S402: Input the positive sample pair and its knowledge triple into the initial neural network model to perform image encoding processing, text encoding processing, and knowledge enhancement processing, and output a first fusion vector representing the distance between the positive text and the image;

[0045] In step S402, the positive sample pair and its knowledge triplet are input into the three branches of the initial neural network model to extract embedding vectors respectively, and the three embedding vectors are obtained and fused to obtain a first fused vector.

[0046] S403: Input the negative sample pair and its knowledge triple into the initial neural network model to perform image encoding processing, text encoding processing, and knowledge enhancement processing, and output a second fusion vector representing the distance between the negative sample text and the image;

[0047] In step S403, the negative sample pair and its knowledge triplet are input into the three branches of the initial neural network model to extract embedding vectors respectively, and the three embedding vectors are obtained and fused to obtain a second fused vector.

[0048] S404: Construct a loss function to train the model parameters of the initial neural network model until the difference between the output first fusion vector and the second fusion vector reaches a preset threshold, thereby obtaining an optimized knowledge-enhanced coding model;

[0049] In step S404, the loss function continuously adjusts the model parameters, so that the output first fused vector gradually approaches the second fused vector until the difference between the two is less than a preset threshold. At this point, the knowledge-enhanced encoding model is considered to have fully learned the key information and knowledge in the text and effectively integrated it into the encoding representation, thus completing the model optimization.

[0050] The technical effects of steps S401-404 can be summarized as follows: by introducing scene graph knowledge extraction and self-attention mechanism (Transformer), knowledge enhancement processing of positive sample pairs and negative sample pairs is achieved. This process improves the model's ability to learn and represent structured information of images and texts. At the same time, by taking knowledge triples as additional input, the model can more comprehensively learn the key information and knowledge in the text during training, thereby improving the accuracy and robustness of the model. In addition, by constructing a loss function to train the model parameters, the optimized knowledge enhancement encoding model can more accurately judge the correlation between text and image, providing strong support for subsequent text-image relationship analysis.

[0051] In one embodiment, if Figure 5 As shown, step S402 includes the following steps S501-S504.

[0052] S501, inputting the image in the positive sample pair into the image encoder in the initial neural network model for image encoding processing to obtain a first image embedding vector;

[0053] S502: Inputting the positive text in the positive sample pair into a text encoder in the initial neural network model for text encoding processing to obtain a first text embedding vector;

[0054] S503, performing vector representation on the knowledge triples of the positive sample pairs and generating a triple embedding vector, and inputting the triple embedding vector into the self-attention mechanism in the initial neural network model for knowledge enhancement processing to obtain a first knowledge embedding vector;

[0055] In step S503, based on the knowledge triples of the positive sample pairs, the Tokenizer (word segmenter) and WordVocabulary (a set of all possible tokens) in BERT (Bidirectional Encoder Representation Transform) are used to obtain the triple embedding vector. The result of the triple embedding vector is then input into the self-attention mechanism transformer, and finally the first knowledge embedding vector is output;

[0056] S504: Perform vector dot product processing on the first image embedding vector, the first text embedding vector, and the first knowledge embedding vector to obtain a first fusion vector representing the distance between the positive text and the image.

[0057] In steps S501-S505, the image is encoded by the image encoder, the positive text is encoded by the text encoder, and the corresponding knowledge triples are enhanced by the self-attention mechanism. The information of the image, the positive text and the corresponding knowledge triples can be effectively embedded into the vector space, and the first fusion vector representing the distance between the positive text and the image can be obtained by vector dot product (i.e., vector fusion).

[0058] In one embodiment, if Figure 6 As shown, step S403 includes the following steps S601-S604.

[0059] S601, inputting the image in the negative sample pair into the image encoder in the initial neural network model for image encoding processing to obtain a second image embedding vector;

[0060] S602: Inputting the negative sample text in the negative sample pair into a text encoder in the initial neural network model for text encoding processing to obtain a second text embedding vector;

[0061] S603, performing vector representation on the knowledge triple of the negative sample pair and generating a triple embedding vector, and inputting the triple embedding vector into the self-attention mechanism in the initial neural network model for knowledge enhancement processing to obtain a second knowledge embedding vector;

[0062] S604: Perform vector dot product processing on the second image embedding vector, the second text embedding vector, and the second knowledge embedding vector to obtain a first fusion vector representing the distance between the negative sample text and the image.

[0063] In steps S501-S505, the image is encoded by the image encoder, the negative text is encoded by the text encoder, and the corresponding knowledge triples are enhanced by the self-attention mechanism. The information of the image, negative text and corresponding knowledge triples can be effectively embedded into the vector space, and the second fusion vector representing the distance between the negative text and the image is obtained by vector dot product (i.e., vector fusion).

[0064] In one embodiment, step S404 includes:

[0065] Establish the loss function L hinge :L hinge =max(0,γ-(d+d')); where d represents the distance between the positive sample and the image, d' represents the distance between the negative sample and the image, and γ represents a temporary variable. γ is used to store the value of d+d' to ensure that the final loss value is greater than 0.

[0066] Establish the loss function L ITCL :L ITCL =1 / 2(L i2t +L t2i ); where L i2t represents the cross entropy loss from image to text, L t2i Represents the cross entropy loss from text to image; the specific process is:

[0067] First, the similarity between each image and all texts is normalized by softmax to obtain the probability distribution Where T is the temperature parameter, which is used to control the sharpness of the distribution; N represents the number of sample pairs input in batches; S represents the similarity matrix after the dot product of the image feature vector I and the text feature vector T obtained by inputting N sample pairs in batches; S ij Indicates the similarity between the i-th image and the j-th text; S ik represents the similarity between the i-th image and the k-th text; exp represents the exponential function;

[0068] Then, the similarity between each text and all images is normalized by softmax to obtain the probability distribution Among them, S ji represents the similarity between the jth text and the ith image; Sjk represents the similarity between the jth text and the kth image;

[0069] Then, we establish the image-to-text cross entropy loss L i2t :

[0070] Then, establish the text-to-image cross entropy loss L t2i :

[0071] Finally, the cross entropy loss L i2t and L t2i Perform summing and averaging to obtain the loss function L ITCL .

[0072] Combined loss function L ITCL And the loss function L hinge , and obtain the total loss function L for optimizing the initial neural network model final ;

[0073] Based on this, the total loss function L final =L ITCL +L hinge .

[0074] It can be seen that the above scheme aims at efficient fusion and precise matching of multimodal data in the financial and medical fields. First, by pre-acquiring and processing positive sample pairs of images and positive texts, and introducing a semantic negative sampling module based on scene graphs, the quality of negative sample pairs is significantly improved, thereby enhancing the model's ability to learn structured information. Secondly, after the positive and negative sample pairs are subjected to scene graph knowledge enhancement processing, they are input into the initial neural network model for training. By optimizing the model with the designed loss function, the output knowledge enhancement coding model has the ability to efficiently learn and represent complex structured information. Finally, the trained knowledge enhancement coding model is used to perform vector dot product comparison of images and texts, and the optimal matching result of multimodal data fusion is output, ensuring the accuracy and reliability of the matching results. The implementation of this series of steps not only improves the efficiency and accuracy of multimodal data processing, but also provides strong technical support for intelligent applications in fields such as finance and medicine.

[0075] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0076] In one embodiment, a multimodal data fusion device 800 based on knowledge enhancement is provided. The multimodal data fusion device based on knowledge enhancement corresponds to the multimodal data fusion method based on knowledge enhancement in the above embodiment. Figure 7 As shown, the multimodal data fusion device based on knowledge enhancement includes a negative sample construction module 701, a model training module 702 and a data matching module 703. The functional modules are described in detail as follows:

[0077] Negative sample construction module 701, used to pre-acquire positive sample pairs of images and positive texts and perform semantic negative sample processing to obtain negative sample pairs;

[0078] A model training module 702 is configured to perform scene graph knowledge enhancement processing on the positive sample pairs and the negative sample pairs, respectively input the processed samples into an initial neural network model for sample vector dot product comparison training, and output a knowledge enhanced coding model after training and optimization using a constructed loss function;

[0079] The data matching module 703 is used to input the data to be matched of the image group-text or image-text group into the knowledge enhanced coding model for vector dot product comparison, and output the image and sample corresponding to the maximum value of the vector dot product result in the test data as the optimal matching result of multimodal data fusion.

[0080] In one embodiment, the negative sample construction module 801 is specifically configured to:

[0081] Obtain a positive sample pair of image-text, wherein the positive text description is a representation of the meaning of the image;

[0082] Performing semantic negative sample processing on the positive text in the positive sample pair by using scene graph knowledge to obtain negative text;

[0083] The negative text is combined with the image in the positive sample pair to obtain a negative sample pair.

[0084] In one embodiment, the model training module 702 is specifically configured to:

[0085] Performing scene graph knowledge extraction on the positive sample pairs and the negative sample pairs respectively to obtain corresponding knowledge triples;

[0086] Inputting the positive sample pair and its knowledge triple into the initial neural network model for image encoding processing, text encoding processing and knowledge enhancement processing, and outputting a first fusion vector representing the distance between the positive text and the image;

[0087] Inputting the negative sample pair and its knowledge triple into the initial neural network model for image encoding processing, text encoding processing and knowledge enhancement processing, and outputting a second fusion vector representing the distance between the negative sample text and the image;

[0088] A loss function is constructed to train the model parameters of the initial neural network model until the difference between the output first fusion vector and the second fusion vector reaches a preset threshold, thereby obtaining an optimized knowledge enhancement coding model.

[0089] In one embodiment, the positive sample pair and its knowledge triplet are input into the initial neural network model for image encoding processing, text encoding processing, and knowledge enhancement processing, and outputting a first fusion vector representing the distance between the positive text and the image, specifically for:

[0090] Inputting the image in the positive sample pair into the image encoder in the initial neural network model for image encoding processing to obtain a first image embedding vector;

[0091] Inputting the positive text in the positive sample pair into the text encoder in the initial neural network model for text encoding processing to obtain a first text embedding vector;

[0092] Vectorizing the knowledge triples of the positive sample pairs and generating a triple embedding vector, and inputting the triple embedding vector into the self-attention mechanism in the initial neural network model for knowledge enhancement processing to obtain a first knowledge embedding vector;

[0093] Performing vector dot product processing on the first image embedding vector, the first text embedding vector, and the first knowledge embedding vector to obtain a first fusion vector representing the distance between the positive text and the image.

[0094] In one embodiment, the step of inputting the negative sample pair and its knowledge triple into an initial neural network model for image encoding processing, text encoding processing, and knowledge enhancement processing, and outputting a second fusion vector representing the distance between the negative sample text and the image, is specifically used to:

[0095] Inputting the image in the negative sample pair into the image encoder in the initial neural network model for image encoding processing to obtain a second image embedding vector;

[0096] Inputting the negative sample text in the negative sample pair into the text encoder in the initial neural network model for text encoding processing to obtain a second text embedding vector;

[0097] Vectorizing the knowledge triples of the negative sample pairs and generating a triple embedding vector, and inputting the triple embedding vector into the self-attention mechanism in the initial neural network model for knowledge enhancement processing to obtain a second knowledge embedding vector;

[0098] Perform vector dot product processing on the second image embedding vector, the second text embedding vector, and the second knowledge embedding vector to obtain a first fusion vector representing the distance between the negative sample text and the image.

[0099] In one embodiment, the constructed loss function trains the model parameters of the initial neural network model until the difference between the output first fusion vector and the second fusion vector reaches a preset threshold, thereby obtaining an optimized knowledge-enhanced coding model, specifically for:

[0100] Establish the loss function L hinge :L hinge=max(0,γ-(d+d')); where d represents the distance between the positive sample and the image, d' represents the distance between the negative sample and the image, and γ represents a temporary variable, which is used to store the value of d+d' to ensure that the final loss value is greater than 0;

[0101] Establish the loss function L ITCL :L ITCL =1 / 2(L i2t +L t2i ); where L i2t represents the cross entropy loss from image to text, L t2i represents the cross entropy loss from text to image;

[0102] Combined loss function L ITCL And the loss function L hinge , and obtain the total loss function L for optimizing the initial neural network model final .

[0103] In one embodiment, the loss function L is established ITCL , specifically used for:

[0104] Normalize the similarity between each image and all texts to get the probability distribution p i (j);

[0105] Normalize the similarity between each text and all images to get the probability distribution p i (i);

[0106] Establish the cross entropy loss L from image to text i2t :

[0107] Establishing the cross entropy loss L for text to image t2i :

[0108] The cross entropy loss L i2t and cross entropy loss L t2i Perform summing and averaging to obtain the loss function L ITCL ; i represents image and j represents text.

[0109] The present invention provides a multimodal data fusion device based on knowledge enhancement, which is aimed at the efficient fusion and precise matching of multimodal data in the financial and medical fields. First, by pre-acquiring and processing positive sample pairs of images and positive texts, and introducing a semantic negative sampling module based on scene graphs, the quality of negative sample pairs is significantly improved, thereby enhancing the model's ability to learn structured information. Secondly, after the positive sample pairs and negative sample pairs are subjected to scene graph knowledge enhancement processing, they are input into the initial neural network model for training. By optimizing the model through the designed loss function, the output knowledge enhancement coding model has the ability to efficiently learn and represent complex structured information. Finally, the trained knowledge enhancement coding model is used to compare the vector dot product of the image and the text, and the optimal matching result of the multimodal data fusion is output, ensuring the accuracy and reliability of the matching result. The implementation of this series of steps not only improves the efficiency and accuracy of multimodal data processing, but also provides strong technical support for intelligent applications in fields such as finance and medicine.

[0110] For the specific definition of the multimodal data fusion device based on knowledge enhancement, please refer to the definition of the multimodal data fusion method based on knowledge enhancement above, which will not be repeated here. The various modules in the above-mentioned multimodal data fusion device based on knowledge enhancement can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0111] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external application terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a multimodal data fusion method based on knowledge enhancement.

[0112] In one embodiment, a computer device is provided. The computer device may be an application terminal, and its internal structure diagram may be as follows: Figure 9As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the application side of a multimodal data fusion method based on knowledge enhancement

[0113] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0114] Pre-acquire positive sample pairs of image and text and perform semantic negative sample processing to obtain negative sample pairs;

[0115] After performing scene graph knowledge enhancement processing on the positive sample pairs and the negative sample pairs, the pairs are respectively input into the initial neural network model for sample vector dot product comparison training, and the training and optimization are performed using the constructed loss function to output a knowledge enhancement coding model;

[0116] The data to be matched of the image group-text or image-text group is input into the knowledge enhanced coding model for vector dot product comparison, and the image and sample corresponding to the maximum value of the vector dot product result in the test data are output as the optimal matching result of multimodal data fusion.

[0117] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0118] Pre-acquire positive sample pairs of image and text and perform semantic negative sample processing to obtain negative sample pairs;

[0119] After performing scene graph knowledge enhancement processing on the positive sample pairs and the negative sample pairs, the pairs are respectively input into the initial neural network model for sample vector dot product comparison training, and the training and optimization are performed using the constructed loss function to output a knowledge enhancement coding model;

[0120] The data to be matched of the image group-text or image-text group is input into the knowledge enhanced coding model for vector dot product comparison, and the image and sample corresponding to the maximum value of the vector dot product result in the test data are output as the optimal matching result of multimodal data fusion.

[0121] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the application side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0122] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM).

[0123] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0124] The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. The non-Company software tools or components that appear in the embodiments of this application are merely examples and do not represent actual use.

Claims

1. A multimodal data fusion method based on knowledge enhancement, characterized in that: include: Pre-acquire positive sample pairs of image and text and perform semantic negative sample processing to obtain negative sample pairs; After performing scene graph knowledge enhancement processing on the positive sample pairs and the negative sample pairs, the pairs are respectively input into the initial neural network model for sample vector dot product comparison training, and the training and optimization are performed using the constructed loss function to output a knowledge enhancement coding model; The data to be matched of the image group-text or image-text group is input into the knowledge enhanced coding model for vector dot product comparison, and the image and sample corresponding to the maximum value of the vector dot product result in the test data are output as the optimal matching result of multimodal data fusion.

2. The multimodal data fusion method based on knowledge enhancement according to claim 1, characterized in that: The method of pre-acquiring positive sample pairs of images and positive texts and performing semantic negative sample processing to obtain negative sample pairs includes: Obtain a positive sample pair of image-text, wherein the positive text description is a representation of the meaning of the image; Performing semantic negative sample processing on the positive text in the positive sample pair by using scene graph knowledge to obtain negative text; The negative text is combined with the image in the positive sample pair to obtain a negative sample pair.

3. The multimodal data fusion method based on knowledge enhancement according to claim 2, characterized in that: After the positive sample pairs and the negative sample pairs are subjected to scene graph knowledge enhancement processing, they are respectively input into the initial neural network model for sample vector dot product comparison training, and the knowledge enhancement coding model is output after training optimization through the constructed loss function, including: Performing scene graph knowledge extraction on the positive sample pairs and the negative sample pairs respectively to obtain corresponding knowledge triples; Inputting the positive sample pair and its knowledge triple into the initial neural network model for image encoding processing, text encoding processing and knowledge enhancement processing, and outputting a first fusion vector representing the distance between the positive text and the image; Inputting the negative sample pair and its knowledge triple into the initial neural network model for image encoding processing, text encoding processing and knowledge enhancement processing, and outputting a second fusion vector representing the distance between the negative sample text and the image; A loss function is constructed to train the model parameters of the initial neural network model until the difference between the output first fusion vector and the second fusion vector reaches a preset threshold, thereby obtaining an optimized knowledge enhancement coding model.

4. The multimodal data fusion method based on knowledge enhancement according to claim 3, characterized in that: The step of inputting the positive sample pair and its knowledge triple into an initial neural network model for image encoding processing, text encoding processing, and knowledge enhancement processing, and outputting a first fusion vector representing the distance between the positive text and the image, comprises: Inputting the image in the positive sample pair into the image encoder in the initial neural network model for image encoding processing to obtain a first image embedding vector; Inputting the positive text in the positive sample pair into the text encoder in the initial neural network model for text encoding processing to obtain a first text embedding vector; Vectorizing the knowledge triples of the positive sample pairs and generating a triple embedding vector, and inputting the triple embedding vector into the self-attention mechanism in the initial neural network model for knowledge enhancement processing to obtain a first knowledge embedding vector; Performing vector dot product processing on the first image embedding vector, the first text embedding vector, and the first knowledge embedding vector to obtain a first fusion vector representing the distance between the positive text and the image.

5. The multimodal data fusion method based on knowledge enhancement according to claim 3, characterized in that: Inputting the negative sample pair and its knowledge triple into the initial neural network model for image encoding processing, text encoding processing and knowledge enhancement processing, and outputting a second fusion vector representing the distance between the negative sample text and the image, including: Inputting the image in the negative sample pair into the image encoder in the initial neural network model for image encoding processing to obtain a second image embedding vector; Inputting the negative sample text in the negative sample pair into the text encoder in the initial neural network model for text encoding processing to obtain a second text embedding vector; Vectorizing the knowledge triples of the negative sample pairs and generating a triple embedding vector, and inputting the triple embedding vector into the self-attention mechanism in the initial neural network model for knowledge enhancement processing to obtain a second knowledge embedding vector; Perform vector dot product processing on the second image embedding vector, the second text embedding vector, and the second knowledge embedding vector to obtain a first fusion vector representing the distance between the negative sample text and the image.

6. The multimodal data fusion method based on knowledge enhancement according to claim 3, characterized in that: The loss function is constructed to train the model parameters of the initial neural network model until the difference between the output first fusion vector and the second fusion vector reaches a preset threshold, thereby obtaining an optimized knowledge enhancement coding model, including: Establish the loss function L hinge :L hinge =max(0,γ-(d+d')); where d represents the distance between the positive sample and the image, d' represents the distance between the negative sample and the image, and γ represents a temporary variable, which is used to store the value of d+d' to ensure that the final loss value is greater than 0; Establish the loss function L ITCL :L ITCL =1 / 2(L i2t +L t2i ); where L i2t represents the cross entropy loss from image to text, L t2i represents the cross entropy loss from text to image; Combined loss function L ITCL And the loss function L hinge , and obtain the total loss function L for optimizing the initial neural network model final .

7. The multimodal data fusion method based on knowledge enhancement according to claim 6, characterized in that: The loss function L is established ITCL ,include: Normalize the similarity between each image and all texts to get the probability distribution p i (j); Normalize the similarity between each text and all images to get the probability distribution p i (i); Establish the cross entropy loss L from image to text i2t : Establishing the cross entropy loss L for text to image t2i : The cross entropy loss L i2t and cross entropy loss L t2i Perform summing and averaging to obtain the loss function L ITCL ; i represents image and j represents text.

8. A multimodal data fusion device based on knowledge enhancement, characterized in that: include: The negative sample construction module is used to pre-acquire positive sample pairs of images and positive texts and perform semantic negative sample processing to obtain negative sample pairs; A model training module is used to perform scene graph knowledge enhancement processing on the positive sample pairs and the negative sample pairs, and then input them into the initial neural network model for sample vector dot product comparison training, and output the knowledge enhancement coding model after training and optimization through the constructed loss function; The data matching module is used to input the data to be matched of the image group-text or image-text group into the knowledge-enhanced coding model for vector dot product comparison, and output the image and sample corresponding to the maximum value of the vector dot product result in the test data as the optimal matching result of multimodal data fusion.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the multimodal data fusion method based on knowledge enhancement are implemented as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the multimodal data fusion method based on knowledge enhancement according to any one of claims 1 to 7 are implemented.