Training method of data processing model and data processing method and device

By extracting the feature vectors of the problem text and image samples, calculating the cost loss of bidirectional correlation and iteratively adjusting the model, the accuracy problem of the positioning answer system of surgical visual problem is solved, and higher predicted answers and positioning accuracy are achieved.

CN120429604APending Publication Date: 2025-08-05TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410169253.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-04
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing surgical visual problem positioning and answering system is not very accurate when providing answers and positioning content.

Method used

By obtaining the sample set, extracting the feature vectors of the problem text and image samples, calculating the two-way correlation cost loss, and iteratively adjusting the initial data processing model based on the text loss, image loss and bidirectional correlation cost to obtain the target data processing model.

Benefits of technology

It significantly improves the accuracy of predictive answers and positioning and enhances scenario understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429604A_ABST
    Figure CN120429604A_ABST
Patent Text Reader

Abstract

The invention relates to a training method of a data processing model, and a data processing method and device. The method comprises the following steps: acquiring a sample set; inputting the question text sample and the corresponding image sample into an initial data processing model, respectively extracting a text feature vector of the question text sample and an image feature vector of the image sample, and outputting a prediction answer text and prediction identification information; based on the text loss, the image loss and the bidirectional relevancy cost loss, performing iterative adjustment on model parameters in an initial data processing model to obtain a target data processing model, the target data processing model being used for inputting a problem text and a corresponding image, and outputting the answer text and the identification information of the image matched with the question text. According to the embodiment of the invention, the answer prediction accuracy and positioning accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a data processing model training method, data processing method, device, computer equipment, storage medium and computer program product. Background Art

[0002] Recording surgical videos is a useful tool for learning surgery. When watching recorded surgical videos, viewers often have various questions about the content shown in the video, such as surgical instruments, human tissues, and workflows. Figure 1 As shown in Figure 2, the Visual Question Localized Answering (VQLA) system for surgical procedures can provide viewers with answers and locate relevant content in the video, enhancing scene understanding. However, the answers and location of content provided by the related art Visual Question Localized Answering systems for surgical procedures are not very accurate. Summary of the Invention

[0003] Based on this, it is necessary to provide a training method, data processing method, device, computer equipment, storage medium and computer program product for a data processing model that can improve the accuracy of predicted answers and positioning accuracy in response to the above technical problems.

[0004] In a first aspect, the present application provides a method for training a data processing model. The method comprises:

[0005] Acquire a sample set; wherein the sample set includes a question text sample and a corresponding image sample, the question text sample has a corresponding annotated answer text, and the image sample has annotated identification information matching the question text sample;

[0006] Inputting the question text sample and the corresponding image sample into an initial data processing model, extracting the text feature vector of the question text sample and the image feature vector of the image sample respectively, and outputting the predicted answer text and predicted identification information;

[0007] Obtaining a first transmission weight for calculating relevance from the image to the text, and fusing the unit relevance costs between the image features in the image feature vector and the text features in the text feature vector based on the first transmission weight to obtain a first relevance cost;

[0008] Obtaining a second transmission weight for calculating relevance from text to image, and fusing unit relevance costs between text features in the text feature vector and image features in the image feature vector based on the second transmission weight to obtain a second relevance cost;

[0009] Obtaining a bidirectional relevance cost loss based on the first relevance cost and the second relevance cost;

[0010] Determine text loss based on the predicted answer text and the annotated answer text, and determine image loss based on the predicted identification information and the annotated identification information;

[0011] Based on the text loss, the image loss and the bidirectional correlation cost loss, the model parameters in the initial data processing model are iteratively adjusted to obtain a target data processing model, wherein the target data processing model is used to input a question text and a corresponding image, and output an answer text and identification information of an image matching the question text.

[0012] In a second aspect, the present application provides a data processing method, the method comprising:

[0013] Get the image and the corresponding question text;

[0014] The image and the question text are input into a data processing model, and the text feature vector of the question text sample and the image feature vector of the image sample are extracted respectively, and the answer text and predicted identification information are output; wherein, the identification information is used to mark the image content that matches the question text, and the data processing model includes training the initial data processing model based on text loss, image loss and bidirectional relevance cost loss, wherein the bidirectional relevance loss includes: obtaining a first transmission weight for calculating relevance from image to text, and based on the first transmission weight, fusing the unit relevance cost between the image features in the image feature vector and the text features in the text feature vector to obtain a first relevance cost; obtaining a second transmission weight for calculating relevance from text to image, and based on the second transmission weight, fusing the unit relevance cost between the text features in the text feature vector and the image features in the image feature vector to obtain a second relevance cost; obtained based on the first relevance cost and the second relevance cost.

[0015] In a second aspect, the present application also provides a data processing model training device. The device includes:

[0016] A first acquisition module is configured to acquire a sample set, wherein the sample set includes a question text sample and a corresponding image sample, the question text sample has a corresponding annotated answer text, and the image sample has annotated identification information matching the question text sample;

[0017] a first prediction module, configured to input the question text sample and the corresponding image sample into an initial data processing model, extract a text feature vector of the question text sample and an image feature vector of the image sample, and output a predicted answer text and predicted identification information;

[0018] a first processing module, configured to obtain a first transmission weight for calculating relevance from the image to the text, and to fuse the unit relevance costs between the image features in the image feature vector and the text features in the text feature vector based on the first transmission weight to obtain a first relevance cost;

[0019] a second processing module, configured to obtain a second transmission weight for calculating relevance from text to image, and to fuse the unit relevance costs between text features in the text feature vector and image features in the image feature vector based on the second transmission weight to obtain a second relevance cost;

[0020] A first calculation module, configured to obtain a bidirectional correlation cost loss based on the first correlation cost and the second correlation cost;

[0021] A second calculation module is used to determine the text loss based on the predicted answer text and the annotated answer text, and to determine the image loss based on the predicted identification information and the annotated identification information;

[0022] A parameter adjustment module is used to iteratively adjust the model parameters in the initial data processing model based on the text loss, the image loss and the bidirectional correlation cost loss to obtain a target data processing model, wherein the target data processing model is used to input a question text and a corresponding image, and output an answer text and identification information of an image matching the question text.

[0023] In one embodiment, the first processing module is further configured to:

[0024] Acquire an image feature from the image feature vector, and determine a text feature having the minimum unit correlation cost with the image feature from the text feature vector;

[0025] Based on the first transmission weight, a unit relevance cost between the image feature and the corresponding text feature is fused to obtain a relevance cost corresponding to the image feature;

[0026] Obtaining a next image feature from the image feature vector, and fusing the correlation cost between the next image feature and the corresponding text feature unit based on the first transmission weight until a correlation cost corresponding to each image feature in the image feature vector is obtained;

[0027] The correlation cost corresponding to each image feature in the image feature vector is fused to obtain a first correlation cost.

[0028] In one embodiment, the first processing module is further configured to:

[0029] Obtaining a feature dimension of the image feature vector;

[0030] The feature dimension is normalized to obtain a first transmission weight for performing relevance calculation from the image to the text.

[0031] In one embodiment, the apparatus further comprises:

[0032] A unit relevance cost calculation module is used to obtain image features from the image feature vector and text features from the text feature vector; calculate the similarity between the image features and the text features, and determine the unit relevance cost of the image features to the text features based on the negative correlation relationship between the similarity and the unit relevance cost.

[0033] In one embodiment, the first calculation module is further configured to:

[0034] fusing the first relevance cost and the second relevance cost to obtain a fused relevance cost;

[0035] The fused correlation costs are averaged to obtain a bidirectional correlation cost loss.

[0036] In one embodiment, the first prediction module is further configured to:

[0037] Extracting the image feature value, image feature key value, and image feature query vector of the image feature vector through an attention mechanism, and extracting the text feature value, text feature key value, and text feature query vector of the text feature vector through an attention mechanism;

[0038] Based on the correlation between the image feature query vector and the text feature key value, weighting the text feature value to obtain a first processed image feature vector;

[0039] Based on the first processed image feature vector and the text feature vector, a predicted answer text and predicted identification information are obtained.

[0040] In one embodiment, the first prediction module is further configured to:

[0041] Performing linear calculations on the image feature vector and the query vector matrix, the eigenvalue matrix, and the eigenkey matrix respectively to obtain the corresponding image feature query vector, image eigenvalue, and image feature key value;

[0042] Dividing the image feature query vector into a preset number of parts to obtain a plurality of sub-image query vectors, dividing the image feature value into the preset number of parts to obtain a plurality of sub-image feature values, and dividing the image feature key value into the preset number of parts to obtain a plurality of sub-image feature key values;

[0043] For each set of sub-image query vectors, sub-image feature values, and sub-image feature key values, weighting the sub-image feature values based on the correlation between the sub-image query vectors and the sub-image feature key values to obtain a second processed sub-image feature vector;

[0044] The second-processed sub-image feature vectors corresponding to each group are concatenated to obtain a second-processed image feature vector, and the image feature value, image feature key value, and image feature query vector of the second-processed image feature are extracted through the attention mechanism.

[0045] In one embodiment, the first prediction module is further configured to:

[0046] Performing linear calculations on the text feature vector and the query vector matrix, the eigenvalue matrix, and the eigenkey matrix respectively to obtain corresponding text feature query vectors, text eigenvalues, and text feature key values;

[0047] Dividing the text feature query vector into a preset number of parts to obtain a plurality of sub-text query vectors, dividing the text feature value into the preset number of parts to obtain a plurality of sub-text feature values, and dividing the text feature key value into the preset number of parts to obtain a plurality of sub-text feature key values;

[0048] For each group of sub-text query vectors, sub-text feature values, and sub-text feature key values, weighting the sub-text feature values based on the relevance between the sub-text query vectors and the sub-text feature key values to obtain a second processed sub-text feature vector;

[0049] The second-processed sub-text feature vectors corresponding to each group are concatenated to obtain the second-processed text feature vector, and the text feature value, text feature key value and text feature query vector of the second-processed text feature are extracted through the attention mechanism.

[0050] In one embodiment, the first prediction module is further configured to:

[0051] Extracting the image feature value, image feature key value, and image feature query vector of the image feature vector through an attention mechanism, and extracting the text feature value, text feature key value, and text feature query vector of the text feature vector through an attention mechanism;

[0052] Based on the correlation between the text feature query vector and the image feature key value, weighting the image feature value to obtain a first processed text feature vector;

[0053] Based on the first processed text feature vector and the text feature vector, a predicted answer text and predicted identification information are obtained.

[0054] In one embodiment, the first prediction module is further configured to:

[0055] Performing a product process on the first parameter matrix in the feature selection network and the image feature vector to obtain a weighted image feature vector, and performing a product process on the second parameter matrix in the feature selection network and the text feature vector to obtain a weighted text feature vector;

[0056] fusing the weighted image feature vector and the weighted text feature vector to obtain a fused feature vector, and normalizing the fused feature vector to obtain a gating matrix;

[0057] Feature selection is performed on the image feature vector and the text feature vector based on the gating matrix to obtain a selected image feature vector and a selected text feature vector, and a predicted answer text and predicted identification information are obtained based on the selected image feature vector and the selected text feature vector.

[0058] In one embodiment, the first prediction module is further configured to:

[0059] Using the gating matrix as the weight of the text feature vector, performing weighted processing on the text feature vector to obtain a second processed text feature vector;

[0060] The image feature vector and the second processed text feature vector are fused to obtain a selected image feature vector and a selected text feature vector.

[0061] In one embodiment, the first prediction module is further configured to:

[0062] Fusing the text feature vector and the image feature vector to obtain a fused feature vector;

[0063] Extracting text feature values, text feature key values, and text feature query vectors of the text feature vectors of the candidate answer texts through an attention mechanism, and extracting feature values, feature key values, and feature query vectors of the fused feature vectors through an attention mechanism;

[0064] Based on the correlation between the text feature query vector and the feature key value of the fused feature, weighting the feature value of the fused feature to obtain a processed feature value;

[0065] Normalizing the processed feature values through a normalization network to obtain a predicted probability value for each candidate answer text;

[0066] The candidate answer text corresponding to the candidate answer text vector with the largest predicted probability value is selected as the predicted answer text.

[0067] In one embodiment, the first prediction module is further configured to:

[0068] Fusing the text feature vector and the image feature vector to obtain a fused feature vector;

[0069] Inputting the fused feature vector into an image decoder network and outputting a predicted position of the identification information;

[0070] Based on the predicted position of the identification information, predicted identification information is obtained.

[0071] In one embodiment, the second computing module is configured to:

[0072] Obtaining a first sub-image loss based on a deviation between the predicted identification information and the labeled identification information;

[0073] Obtaining a second sub-image loss based on a graphic area formed by the predicted identification information and an image area formed by the annotated identification information;

[0074] An image loss is obtained based on the first sub-image loss and the second sub-image loss.

[0075] In a fourth aspect, the present application further provides a data processing device, comprising:

[0076] The second acquisition module is used to obtain the image and the corresponding question text;

[0077] The second prediction module is used to input the image and the question text into the data processing model, extract the text feature vector of the question text sample and the image feature vector of the image sample respectively, and output the answer text and prediction identification information; wherein, the identification information is used to mark the image content matching the question text, and the data processing model includes training the initial data processing model based on text loss, image loss and bidirectional relevance cost loss, wherein the bidirectional relevance loss includes: obtaining a first transmission weight for calculating relevance from image to text, and based on the first transmission weight, fusing the unit relevance cost between the image features in the image feature vector and the text features in the text feature vector to obtain a first relevance cost; obtaining a second transmission weight for calculating relevance from text to image, and based on the second transmission weight, fusing the unit relevance cost between the text features in the text feature vector and the image features in the image feature vector to obtain a second relevance cost; obtained based on the first relevance cost and the second relevance cost.

[0078] In a fifth aspect, the present application further provides a computer device, wherein the computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method described in any embodiment of the present disclosure is implemented.

[0079] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any embodiment of the present disclosure.

[0080] In a fifth aspect, the present application further provides a computer program product, which includes a computer program that, when executed by a processor, implements the method described in any embodiment of the present disclosure.

[0081] The training method, apparatus, computer device, storage medium and computer program product of the above-mentioned data processing model, in the above-mentioned embodiment, a first correlation cost is obtained by calculating the correlation from the image to the text using the first transmission weight and the unit correlation cost between the image features in the image feature vector and the text features in the text feature vector; a second correlation cost is obtained by calculating the correlation from the text to the image using the second transmission weight and the unit correlation cost between the text features in the text feature vector and the image features in the image feature vector, and a bidirectional correlation cost loss is obtained based on the first correlation cost and the second correlation cost. The bidirectional correlation cost loss can describe the bidirectional loss between the image and the text. It can significantly improve the accuracy of the initial data processing model prediction. Furthermore, the embodiment of the present disclosure utilizes the text features in the text feature vector and the image features in the image feature vector when describing the bidirectional correlation cost loss, which can describe a more fine-grained alignment relationship between features and can further improve the accuracy of the initial data processing model prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 FIG1 is an application environment diagram of a training method for a data processing model in one embodiment;

[0083] Figure 2 1 is a flow chart of a method for training a data processing model in one embodiment;

[0084] Figure 3 A schematic diagram of cargo and warehouse transportation in one embodiment;

[0085] Figure 4 A schematic diagram of transmitting an image feature vector and a text feature vector in one embodiment;

[0086] Figure 5 1 is a flow chart of a method for training a data processing model in one embodiment;

[0087] Figure 6 A schematic diagram of transmitting an image feature vector and a text feature vector in one embodiment;

[0088] Figure 7 1 is a flow chart of a method for training a data processing model in one embodiment;

[0089] Figure 8 1 is a flow chart of a method for training a data processing model in one embodiment;

[0090] Figure 9 Schematic diagram of the structure of a multi-head self-attention mechanism in one embodiment;

[0091] Figure 101 is a flow chart of a method for training a data processing model in one embodiment;

[0092] Figure 11 1 is a flow chart of a method for training a data processing model in one embodiment;

[0093] Figure 12 1 is a flow chart of a method for training a data processing model in one embodiment;

[0094] Figure 13 1 is a flow chart of a method for training a data processing model in one embodiment;

[0095] Figure 14 Schematic diagram of the structure of a text decoder network in one embodiment;

[0096] Figure 15 A schematic diagram of predicted identification information and marked identification information in one embodiment;

[0097] Figure 16 A framework diagram of a method for training a data processing model in one embodiment;

[0098] Figure 17 is a block diagram of a training device for a data processing model in one embodiment;

[0099] Figure 18 is a block diagram of a data processing device in one embodiment;

[0100] Figure 19 is a diagram of the internal structure of a computer device in one embodiment;

[0101] Figure 20 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0102] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0103] In order to facilitate those skilled in the art to understand the technical solution provided by the embodiments of the present disclosure, the technical environment in which the technical solution is implemented is described below.

[0104] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0105] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0106] Computer vision (CV) is the science of making machines "see." Specifically, it refers to using cameras and computers to replace the human eye in identifying, locating, and measuring objects. Furthermore, it involves image processing, which transforms the computer's image into something more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems capable of extracting information from images or multidimensional data. Large model technology has revolutionized the development of computer vision. Pre-trained models in the field of vision, such as the Swin Transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and broadly apply to specific downstream tasks. Computer vision technologies generally include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. Common biometric recognition technologies include facial recognition and fingerprint recognition.

[0107] Key technologies in speech technology include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech becoming one of the most promising methods of human-computer interaction. Large model technology is revolutionizing the development of speech technology. Pre-trained models such as WavLM and UniSpeech, which leverage the Transformer architecture, possess strong generalization and versatility, enabling them to effectively handle a wide range of speech processing tasks.

[0108] Natural language processing (NLP) is a key area of research in computer science and artificial intelligence. It studies theories and methods that enable effective communication between humans and computers using natural language. Natural language processing involves natural language, the language we use daily, and is closely related to linguistics. Pretrained models are derived from large language models (LLMs) in the NLP field. After fine-tuning, LLMs can be widely applied to downstream tasks. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question-answering, and knowledge graphs.

[0109] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.

[0110] Autonomous driving technology refers to the ability of a vehicle to drive itself without a driver. It typically includes technologies such as high-precision maps, environmental perception, computer vision, behavioral decision-making, path planning, and motion control. Autonomous driving encompasses multiple development paths, including single-vehicle intelligence, vehicle-road collaboration, and networked cloud control. Autonomous driving technology has broad application prospects, currently in logistics, public transportation, taxis, and smart transportation, and is expected to further develop in the future.

[0111] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. The solution provided in the embodiment of this application involves artificial intelligence computer vision technology, natural language processing technology, and machine learning / deep learning technology.

[0112] In one embodiment, Figure 2 As shown, a data processing method is provided. This embodiment uses the method applied to a terminal as an example for illustration. It is understandable that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0113] Step S201, obtaining a sample set; wherein the sample set includes a question text sample and a corresponding image sample, the question text sample has a corresponding annotated answer text, and the image sample has annotated identification information matching the question text sample.

[0114] The image sample may include an image of any scene, such as an image of a medical scene, an image of a teaching scene, an image of an exhibition scene, etc. The question text sample may include a question related to the content of the image sample. For example, in an image of a surgical scene, the question text sample may include "which surgical instrument was used" and "where was the surgical location". Correspondingly, the annotated answer text may include text content that matches the question text sample. For example, in the above example, the answer text may include "monopolar curved scissors were used". The identification information may include graphic annotation information, pattern annotation information, text annotation information, etc. For example, the identification information is marked in the lower right corner of the image sample, and the identification information indicates the location information that matches the question text sample.

[0115] Step S203: Input the question text sample and the corresponding image sample into the initial data processing model, extract the text feature vector of the question text sample and the image feature vector of the image sample respectively, and output the predicted answer text and predicted identification information.

[0116] Among them, the initial data processing model may include an artificial intelligence neural network based on deep learning, such as a convolutional neural network, a long short-term memory network, a transformer network, etc. In an exemplary embodiment, the initial data processing model may include a visual encoder network, wherein the visual encoder network is used to convert image data into a network with a recognizable image feature vector, such as a convolutional neural network, a deep residual network, etc. Optionally, a ResNet18 network is used, which is suitable for training and deployment under limited resources. The image feature vector is extracted by the visual encoder. In an exemplary embodiment, the initial data processing model may include a text encoder network, wherein the text encoder network is used to convert text data into a network with a recognizable text feature vector, such as Word2Vec, GloVe and BERT network, etc., and the text feature vector is extracted by the text encoder.

[0117] Step S205: Obtain a first transmission weight for performing relevance calculation from the image to the text, and based on the first transmission weight, fuse the unit relevance cost between the image features in the image feature vector and the text features in the text feature vector to obtain a first relevance cost.

[0118] In order to achieve a higher precision alignment of image features and text features, the embodiment of the present disclosure proposes a first relevance cost from the image sample to the question text sample and a second relevance cost from the question text sample to the image sample. To explain this concept, the optimal transportation problem is introduced. Figure 3 As shown,

[0119] α from α1 to α n represents goods, β from β1 to β n Represents a warehouse, unit cost function C(α i , β j ) represents the goods α i Transport to warehouse β j The goal of the optimal transport problem is to design a transport plan that minimizes the total transport cost D(α, β). The total transport cost D(α, β) is expressed as follows:

[0120]

[0121]

[0122]

[0123] Among them, T i,j Represents goods α i Transport to warehouse β j The quality of the goods α iHow much mass is transported to the warehouse β j , of course, goods α i It can also be transported to other warehouses. i , β j ) represents the goods α i Transport to β j The unit transportation cost of each cargo is i The corresponding mass m i ∈[0,∞); each warehouse β j Corresponding warehouse capacity

[0124] In the embodiment of the present disclosure, the first transmission weight for calculating the relevance between the image and the text may include the weight for transmitting the image feature in the image feature vector to the text feature in the text feature vector, for example, Figure 4 As shown, the image feature vector is represented as F v And the text feature vector is represented as F q , the first transmission weight may include any image feature F in the image feature vector v,i To any text feature F in the text feature vector q,j For different image features or different text features, the corresponding first transmission weights can be set to be the same or different, and the present disclosure does not limit this. In an exemplary embodiment, based on the first transmission weight, the unit correlation cost between the image feature in the image feature vector and the text feature in the text feature vector is weighted and summed to obtain the first correlation cost, for example: the image feature F v,i To text feature F q,j The unit relevance cost is expressed as C(F v,j , F q,j ), the first correlation cost from the image feature vector to the text feature vector is expressed as D(F v , F q ):

[0125]

[0126]

[0127]

[0128] Among them, T i,j Represents the image feature F v,i To text feature F q,j The first transmission weight is set to the image feature vector F v The dimension is n, the text feature vector F q The dimension of is also n.

[0129] In an exemplary embodiment, based on the size of the unit correlation cost between the image features and the text features in the text feature vector, a preset number of unit correlation costs with the smallest unit correlation cost can be screened for fusion processing, and the present disclosure does not impose any restrictions on this.

[0130] Step S207: Obtain a second transmission weight for performing correlation calculation from text to image, and based on the second transmission weight, fuse the unit correlation cost between the text features in the text feature vector and the image features in the image feature vector to obtain a second correlation cost.

[0131] In the embodiment of the present disclosure, the second transmission weight for calculating the relevance of text to image may include the weight for transmitting the text feature in the text feature vector to the image feature in the image feature vector, for example, Figure 4 As shown, the image feature vector is represented as F v And the text feature vector is represented as

[0132] F q , the second transmission weight can include any text feature F in the text feature vector q,j To any image feature F in the image feature vector v,i The weight for transmission. For different image features or different text features, the corresponding second transmission weights can be set to be the same or different, and the present disclosure does not impose any restrictions on this. In an exemplary embodiment, based on the second transmission weight, the unit correlation costs between the text features in the text feature vector and the image features in the image feature vector are weighted and summed to obtain the second correlation cost. In an exemplary embodiment, it is also possible to screen a preset number of unit correlation costs with the smallest unit correlation cost for fusion processing based on the size of the unit correlation cost between the text features and each image feature in the image feature vector, and the present disclosure does not impose any restrictions on this.

[0133] In the embodiment of the present disclosure, the unit correlation cost is used to characterize the distance between the image feature and the text feature, or to characterize the distance between the text feature and the image feature. Optionally, the unit correlation cost is negatively correlated with the similarity between the two, that is, the higher the similarity between the image feature and the text feature, the smaller the unit correlation cost of the image feature and the text feature; the lower the similarity between the image feature and the text feature, the greater the unit correlation cost of the image feature and the text feature. Among them, the similarity can be characterized by cosine similarity, Euclidean distance, Pearson correlation coefficient, etc., and the present disclosure does not limit this.

[0134] Step S209 : Obtain a bidirectional correlation cost loss based on the first correlation cost and the second correlation cost.

[0135] Specifically, the first correlation cost and the second correlation cost can be fused to obtain a bidirectional correlation cost, for example: Bidirectional correlation cost loss = first correlation cost + second correlation cost. The first correlation cost and the second correlation cost can be weighted and fused to obtain a bidirectional correlation cost, for example: Bidirectional correlation cost loss = a × first correlation cost + b × second correlation cost, where a and b represent weight coefficients.

[0136] Step S211 : determining text loss based on the predicted answer text and the annotated answer text, and determining image loss based on the predicted identification information and the annotated identification information.

[0137] Specifically, the text loss is used to measure the difference information between the predicted answer text and the annotated answer text, wherein the smaller the text loss, the closer the two are. In an exemplary embodiment, the text loss can be calculated using a cross-entropy loss function. The image loss can include indicators for measuring the difference information between the predicted identification information and the annotated identification information, such as IoU, GIoU, DIoU, CIoU and other indicators. In an exemplary embodiment, the image loss function can also be combined with the minimum absolute deviation function to describe the difference information between the position of the predicted identification information and the corresponding position of the corresponding annotated identification information.

[0138] Step S213, based on the text loss, the image loss and the bidirectional correlation cost loss, iteratively adjust the model parameters in the initial data processing model to obtain a target data processing model, wherein the target data processing model is used to input the question text and the corresponding image, and output the answer text and the identification information of the image matching the question text.

[0139] Specifically, based on the text loss, the image loss, and the bidirectional correlation cost loss, the model parameters in the initial data processing model are iteratively adjusted until the text loss, the image loss, and the bidirectional correlation loss meet the preset requirements, or the number of iterations meets the preset number. In an exemplary embodiment, the text loss can be expressed as L CE , the image loss can be expressed as L GIoU +L1, where L GIoU The difference information indicator between the predicted identification information and the marked identification information; L1 represents the difference between the position of the predicted identification information and the corresponding position of the marked identification information. L OT Represents the bidirectional correlation loss, and the overall model training target L can be expressed as:

[0140] L=L CE +(L GIoU +L1)+λL OT (7)

[0141] In the disclosed embodiment, the target data processing model is used to input question text and the corresponding image, and output answer text and identification information of the image that matches the question text. It can be applied to medical scenarios, teaching scenarios, exhibition scenarios, etc. For example: in a medical scenario, the question text obtained by the target data processing model includes "which surgical instrument was used" and "where the surgery was performed", and the output answer text includes "monopolar curved scissors were used", and the identification information is marked at the surgical location in the image. In an exemplary embodiment, in a dynamic video teaching scenario, questions can be asked about the video being played, and the corresponding video frame can be cached at the beginning of the question text input. The question text and the corresponding video frame image are input into the target data processing model, and the answer text and identification information of the image that matches the question text are output.

[0142] In the above embodiment, a first relevance cost is obtained by calculating the relevance from the image to the text using a first transfer weight and the unit relevance cost between the image features in the image feature vector and the text features in the text feature vector. A second relevance cost is obtained by calculating the relevance from the text to the image using a second transfer weight and the unit relevance cost between the text features in the text feature vector and the image features in the image feature vector. Based on the first and second relevance costs, a bidirectional relevance cost loss is obtained. This bidirectional relevance cost loss can describe the bidirectional loss between the image and the text. Compared to the prior art's unidirectional attention from image to text or from text to image, the present embodiment can improve the accuracy of the initial data processing model's predictions. Furthermore, the present embodiment utilizes the text features in the text feature vector and the image features in the image feature vector when describing the bidirectional relevance cost loss. Compared to the prior art's use of text feature vectors and image feature vectors as processing objects, the present embodiment describes a more fine-grained alignment relationship between features, further improving the accuracy of the initial data processing model's predictions.

[0143] In one embodiment, reference Figure 5 As shown, based on the first transmission weight, the unit correlation cost between the image feature in the image feature vector and the text feature in the text feature vector is fused to obtain a first correlation cost, including:

[0144] Step S501 : acquiring an image feature from the image feature vector, and determining a text feature having the minimum unit correlation cost with the image feature from the text feature vector.

[0145] Specifically, you can refer to Figure 6 As shown, from the image feature vector F vGet the image features and calculate the image features and text feature vector F respectively q The similarity between each text feature in is used to determine the text feature with the minimum unit relevance cost. For example, Figure 6 In the text feature F q,1 to F q,n Medium F q,2 and image feature F v,1 The unit relevance cost is the smallest, then the text feature F q,2 As the image feature F v,1 The corresponding text features are Figure 6 Indicated as F v,1 Pointing to F q,2 Similarly, for example, the text feature F q,1 As the image feature F v,2 The corresponding text features.

[0146] Step S503 : Based on the first transmission weight, the unit relevance cost between the image feature and the corresponding text feature is fused to obtain the relevance cost corresponding to the image feature.

[0147] In an exemplary embodiment, the relevance cost corresponding to the image feature may include the product of the unit relevance cost between the image feature and the corresponding text feature, for example, the image feature F v,i The corresponding relevance cost is expressed as: T i,j ×min C(F v,i F q,j ).

[0148] Step S505: Obtain the next image feature from the image feature vector, and fuse the correlation cost between the next image feature and the corresponding text feature unit based on the first transmission weight until the correlation cost corresponding to each image feature in the image feature vector is obtained.

[0149] Similarly, the relevance cost corresponding to the next image feature in the embodiment of the present disclosure can also be performed according to the method in the above embodiment, that is, the relevance cost between the next image feature and the corresponding text feature unit is fused, for example, the image feature F v,,i+1 The corresponding relevance cost is expressed as:

[0150] T i+1,j ×min C(F v,i+1 , F q,j ). until the correlation cost corresponding to each image feature in the image feature vector is obtained.

[0151] Step S507 : performing fusion processing on the correlation cost corresponding to each image feature in the image feature vector to obtain a first correlation cost.

[0152] Specifically, the embodiment of the present disclosure performs a fusion process on the correlation cost corresponding to each image feature in the image feature vector to obtain a first correlation cost. Therefore, the first correlation cost in the embodiment of the present disclosure can be converted from (4) to (8).

[0153]

[0154] Based on the same method as the above embodiment, the text feature vector F q To the image feature vector F v The second relevance cost can be expressed as:

[0155]

[0156] In the embodiment of the present disclosure, for the first correlation cost from the image feature vector to the text feature vector, it is not necessary to calculate the image feature in the image feature vector and each text feature in the text feature vector. Instead, the text feature with the lowest correlation with the image feature unit in the text feature vector is selected for calculation to obtain the correlation cost corresponding to the image feature, and the correlation costs corresponding to each image feature in the image feature vector are integrated to obtain the first correlation cost. The technical effect achieved by the above method can be referred to Figure 6 As shown, the image feature F v,1 The corresponding relevance cost can be obtained by comparing it with the text feature F q,2 The unit correlation cost between them is calculated and the image feature F v,i The corresponding relevance cost can also be obtained by comparing it with the text feature F q,2 The unit correlation cost between them is calculated, and the text feature F q,2 The corresponding relevance cost is obtained by comparing it with the image feature F v,1 The above shows that by selecting the text feature with the lowest unit relevance to calculate the first relevance cost, the directionality from image features to text features is enhanced, the difference between the first relevance cost and the second relevance cost is increased, and a better expression can be used to describe the bidirectional loss between image and text.

[0157] In one embodiment, obtaining a first transmission weight for calculating relevance from an image to a text includes:

[0158] Obtaining a feature dimension of the image feature vector;

[0159] The feature dimension is normalized to obtain a first transmission weight for performing relevance calculation from the image to the text.

[0160] In the embodiment of the present disclosure, it is assumed that the feature dimension of the image feature vector is n. In an exemplary embodiment, the feature dimension is normalized to obtain a first transmission weight for calculating the relevance from the image to the text, which is expressed as 1 / n.

[0161] In the embodiment of the present disclosure, the constraint condition in formula (6) can be removed, and any image feature F in the image feature vector can be set v,i The mass m v,i is 1 / n, then the image feature F in the image feature vector v,i Transfer all its qualities to the text feature F in the closest text feature vector q,j , then the image feature F v,i Transfer to text feature F q,j The first transmission weight T i,j It can be expressed as follows:

[0162]

[0163] Among them, arg m i n C(F v,i , F q,j ) represents the minimum unit relevance cost. j' represents the j corresponding to the minimum unit relevance cost. In one embodiment, the first relevance cost D(F v , F q ) is expressed as follows:

[0164]

[0165] Similarly, the second relevance cost D(F q , F v ) can be expressed as:

[0166]

[0167] In the above embodiment, the feature dimension of the image feature vector is normalized to obtain a first transmission weight, and the first transmission weight is applied to the text feature closest to the image feature, which can improve the convenience and feasibility of implementing the first relevance cost and the second relevance cost.

[0168] In one embodiment, before fusing the unit correlation cost between the image features in the image feature vector and the text features in the text feature vector based on the first transmission weight, the method further includes:

[0169] Acquiring image features from the image feature vector and acquiring text features from the text feature vector;

[0170] The similarity between the image feature and the text feature is calculated, and based on a negative correlation relationship between the similarity and the unit relevance cost, the unit relevance cost from the image feature to the text feature is determined.

[0171] Specifically, the similarity between image features and text features can be characterized by cosine similarity, Euclidean distance, Pearson correlation coefficient, etc., and this disclosure does not limit this. In the embodiment of this disclosure, the negative correlation relationship between similarity and unit correlation cost can be realized by using logarithmic function, ratio and four arithmetic operations. In a specific embodiment, taking cosine similarity as an example, the image feature F in the image feature vector v,i and the text feature F in the text feature vector q,j The unit correlation cost between them is expressed as formula (11)

[0172] C(F v,i , F q,j )=1-cos(F v,i , F q,j ) (11)

[0173] Among them, cos(F v,i , F q,j ) represents the image feature F v,i With text feature F q,j The cosine similarity between them. As the image feature F v,i With text feature F q,j As the cosine similarity between them increases, the unit correlation cost between them in formula (10) decreases.

[0174] In the above embodiment, the unit correlation cost between the image feature and the text feature is represented by the similarity between the image feature and the text feature, and the correlation degree between the image feature and the text feature can be accurately represented.

[0175] In one embodiment, obtaining a bidirectional relevance cost loss based on the first relevance cost and the second relevance cost includes:

[0176] fusing the first relevance cost and the second relevance cost to obtain a fused relevance cost;

[0177] The fused correlation costs are averaged to obtain a bidirectional correlation cost loss.

[0178] In the embodiment of the present disclosure, the fusion processing of the first correlation cost and the second correlation cost may include performing additive fusion processing on the first correlation cost and the second correlation cost, or weighting the first correlation cost and the second correlation cost before performing additive fusion processing. The fused correlation cost is averaged to obtain a bidirectional correlation cost loss, which may include performing ratio processing on the fused correlation cost and the number of the first correlation cost and the second correlation cost to obtain a bidirectional correlation cost loss. In a specific embodiment, the bidirectional correlation loss is expressed as follows:

[0179]

[0180] Among them, D(F v , F q ) represents the first correlation cost from the image feature vector to the text feature vector, D(F q , F v ) represents the second correlation cost from the text feature vector to the image feature vector.

[0181] In one embodiment, reference Figure 7 As shown, the predicted answer text and predicted identification information are output, including:

[0182] Step S701, extracting the image feature value, image feature key value and image feature query vector of the image feature vector through the attention mechanism, and extracting the text feature value, text feature key value and text feature query vector of the text feature vector through the attention mechanism.

[0183] In the embodiment of the present disclosure, the linear layer in the attention mechanism can be used to extract the image feature value, image feature key value and image feature query vector of the image feature vector. Specifically, the image feature vector F v Weight matrix corresponding to image eigenvalues Perform product processing to obtain the image eigenvalue V v ; The image feature vector F v The weight matrix corresponding to the image feature key value Perform product processing to obtain the image feature key value K v ; The image feature vector F v The weight matrix corresponding to the image feature query vector Perform product processing to obtain the image feature query vector Q v The weight matrix Weight Matrix Weight Matrix is a learnable parameter.

[0184] Similarly, the linear layer in the attention mechanism can be used to extract the text feature value, text feature key value and text feature query vector of the text feature vector. Specifically, the text feature vector F q Weight matrix corresponding to text feature values Perform product processing to obtain the text feature value V q ; The text feature vector F q The weight matrix corresponding to the text feature key value Perform product processing to obtain the text feature key value K q ; The text feature vector F q The weight matrix corresponding to the text feature query vector Perform product processing to obtain the text feature query vector Q q .

[0185] Step S703 : Based on the correlation between the image feature query vector and the text feature key value, weighted processing is performed on the text feature value to obtain a first processed image feature vector.

[0186] In one embodiment, the first processed image feature vector F' v It can be calculated by the following formula (13):

[0187]

[0188] Wherein, dk represents a coefficient. It should be noted that other normalization functions can also be used to replace Softmax in formula (13), and this disclosure does not limit this.

[0189] Step S705 : obtaining predicted answer text and predicted identification information based on the first processed image feature vector and the text feature vector.

[0190] In one exemplary embodiment, the first processed image feature vector and text feature vector are concatenated to obtain a concatenated feature vector, and the concatenated feature vector is used to obtain the predicted answer text and predicted identification information. In another exemplary embodiment, the first processed image feature vector and text feature vector may be filtered, and the filtered feature vector may be used to obtain the predicted answer text and predicted identification information.

[0191] In the attention mechanism, the feature query vector Q is used to find information related to itself, the feature key value K is used to match a series of candidate items of the query, and the eigenvalue V is information associated with the feature key value, which is used to return relevant content after finding a matching feature key value. Generally speaking, the higher the similarity between the feature query vector and the feature key value, the higher the weight obtained by the corresponding eigenvalue, that is, it is considered to be more relevant. In the above embodiment, the feature query vector Q is derived from the image feature vector, which can be regarded as a kind of "question", and the feature key value K and the eigenvalue V are derived from the text feature vector, which means that the text content is used to match and respond to the question of the image. Therefore, the image feature query vector is guided by the text feature key value and the text feature value, that is, the image feature vector is looking to find the corresponding explanation in the text feature vector, thereby realizing the unidirectional attention from the text feature vector to the image feature vector. The embodiment of the present disclosure adds the expression of the correlation between the text feature vector and the image feature vector.

[0192] In one embodiment, extracting the image feature value, image feature key value, and image feature query vector of the image feature vector through an attention mechanism includes:

[0193] Step S801 : performing linear calculations on the image feature vector, the query vector matrix, the eigenvalue matrix, and the eigenkey matrix respectively to obtain corresponding image feature query vectors, image eigenvalues, and image feature key values.

[0194] Specifically, refer to Figure 9 As shown, the image feature vector F v and the query vector matrix W Q Perform linear calculation to obtain the image feature query vector Q; convert the image feature vector F v and the feature key matrix W K Perform linear calculation to obtain the image feature key value K; convert the image feature vector F v and the eigenvalue matrix W V Perform linear calculations to obtain the image eigenvalue V. Among them, the query vector matrix, eigenvalue matrix, and eigenkey matrix are learnable parameter matrices.

[0195] Step S803: divide the image feature query vector into a preset number of parts to obtain multiple sub-image query vectors, divide the image feature value into the preset number of parts to obtain multiple sub-image feature values, and divide the image feature key value into the preset number of parts to obtain multiple sub-image feature key values.

[0196] The number of preset copies can be consistent with the attention head of the attention mechanism. Figure 9 As shown, the image feature query vector Q is divided into sub-image query vectors

[0197] The image feature key value K is divided into the sub-image key value K0 and the sub-image key value K1. The image feature value V is divided into the sub-image feature values V0 and V1.

[0198] Step S805 : For each set of sub-image query vectors, sub-image feature values, and sub-image feature key values, weight the sub-image feature values based on the correlation between the sub-image query vector and the sub-image feature key value to obtain a second processed sub-image feature vector.

[0199] In one exemplary embodiment, the sub-image query vector Q0, the sub-image key value K0, and the sub-image feature value V0 can be divided into one group, and the remaining sub-image query vector Q1, the sub-image key value K1, and the sub-image feature value V1 can be divided into another group. In another embodiment, the sub-image query vector Q1, the sub-image key value K0, and the sub-image feature value V0 can also be divided into one group, which is not limited in this disclosure.

[0200] Taking the sub-image query vector Q0, sub-image key value K0 and sub-image feature value V0 as a group as an example, the sub-image feature vector of the second processing can be obtained by formula (14):

[0201]

[0202] Step S807, concatenate the second-processed sub-image feature vectors corresponding to each group to obtain the second-processed image feature vector, and extract the image feature value, image feature key value and image feature query vector of the second-processed image feature through the attention mechanism.

[0203] In the embodiment of the present disclosure, the second processed sub-image feature vectors corresponding to each group are spliced together to obtain the second processed image feature vector, such as Figure 9 As shown in Z.

[0204] In the above embodiment, the image feature query vector, image feature value, and image feature key value of the image feature vector are extracted through a self-attention mechanism. For each group of sub-image query vectors, sub-image feature values, and sub-image feature key values, the sub-image feature values are weighted based on the correlation between the sub-image query vector and the sub-image feature key value to obtain a second-processed sub-image feature vector. The second-processed sub-image feature vectors corresponding to each group are then concatenated to obtain a second-processed image feature vector. This allows the second-processed image feature vector to express the correlation between the various image features within the image feature vector, and the grouped and parallel calculation of the second-processed sub-image feature vectors facilitates improved computational efficiency.

[0205] In one embodiment, reference Figure 10 As shown, the text feature value, text feature key value and text feature query vector of the text feature vector are extracted through the attention mechanism, including:

[0206] Step S1001 : performing linear calculations on the text feature vector, the query vector matrix, the eigenvalue matrix, and the eigenkey matrix respectively to obtain corresponding text feature query vectors, text eigenvalues, and text feature key values.

[0207] Specifically, refer to Figure 9 As shown, the text feature vector F q and the query vector matrix W Q Perform linear calculation to obtain the image feature query vector Q; convert the image feature vector F q and the feature key matrix W K Perform linear calculation to obtain the image feature key value K; convert the image feature vector F q and the eigenvalue matrix W V Perform linear calculations to obtain the image eigenvalue V. Among them, the query vector matrix, eigenvalue matrix, and eigenkey matrix are learnable parameter matrices.

[0208] Step S1003, dividing the text feature query vector into a preset number of parts to obtain multiple sub-text query vectors, dividing the text feature value into the preset number of parts to obtain multiple sub-text feature values, and dividing the text feature key value into the preset number of parts to obtain multiple sub-text feature key values.

[0209] The number of preset copies can be consistent with the attention head of the attention mechanism. Figure 9 As shown, the text feature query vector Q is divided into sub-text query vector Q0 and sub-text Q1. The text feature key value K is divided into sub-text key value K0 and sub-text key value K1. The text feature value V is divided into sub-text feature values V0 and V1.

[0210] Step S1005 : For each group of sub-text query vectors, sub-text feature values and sub-text feature key values, weight the sub-text feature values based on the relevance between the sub-text query vectors and the sub-text feature key values to obtain a second processed sub-text feature vector.

[0211] In one exemplary embodiment, the subtext query vector Q0, the subtext key value K0, and the subtext feature value V0 can be divided into one group, and the remaining subtext query vector Q1, the subtext key value K1, and the subtext feature value V1 can be divided into another group. In another embodiment, the subtext query vector Q1, the subtext key value K0, and the subtext feature value V0 can also be divided into one group, which is not limited by the present disclosure.

[0212] Step S1007, concatenate the second-processed sub-text feature vectors corresponding to each group to obtain the second-processed text feature vector, and extract the text feature value, text feature key value and text feature query vector of the second-processed text feature through the attention mechanism.

[0213] In the embodiment of the present disclosure, the second processed sub-image feature vectors corresponding to each group are spliced together to obtain the second processed image feature vector, such as Figure 9 As shown in Z.

[0214] In the above embodiment, a text feature query vector, text feature value, and text feature key value of a text feature vector are extracted through a self-attention mechanism. For each group of sub-text query vectors, sub-text feature values, and sub-text feature key values, the sub-text feature values are weighted based on the correlation between the sub-text query vector and the sub-text feature key value to obtain a second-processed sub-text feature vector. The second-processed sub-text feature vectors corresponding to each group are concatenated to obtain a second-processed text feature vector. This enables the second-processed text feature vector to express the correlation between the various text features within the text feature vector, and grouping and parallel calculation of the second-processed sub-text feature vectors helps improve computational efficiency.

[0215] In one embodiment, reference Figure 11 As shown, the predicted answer text and predicted identification information are output, including:

[0216] Step S1101, extracting the image feature value, image feature key value and image feature query vector of the image feature vector through the attention mechanism, and extracting the text feature value, text feature key value and text feature query vector of the text feature vector through the attention mechanism.

[0217] In the embodiment of the present disclosure, the linear layer in the attention mechanism can be used to extract the image feature value, image feature key value and image feature query vector of the image feature vector. Specifically, the image feature vector F v Weight matrix corresponding to image eigenvalues Perform product processing to obtain the image eigenvalue V v ; The image feature vector F v The weight matrix corresponding to the image feature key value Perform product processing to obtain the image feature key value K v ; The image feature vector F v The weight matrix corresponding to the image feature query vector Perform product processing to obtain the image feature query vector Q v The weight matrix Weight Matrix Weight Matrix is a learnable parameter.

[0218] Similarly, the linear layer in the attention mechanism can be used to extract the text feature value, text feature key value and text feature query vector of the text feature vector. Specifically, the text feature vector F q Weight matrix corresponding to text feature values Perform product processing to obtain the text feature value V q ; The text feature vector F q The weight matrix corresponding to the text feature key value Perform product processing to obtain the text feature key value K q ; The text feature vector F q The weight matrix corresponding to the text feature query vector Perform product processing to obtain the text feature query vector Q q .

[0219] Step S1103 : Based on the correlation between the text feature query vector and the image feature key value, weighted processing is performed on the image feature value to obtain a first processed text feature vector.

[0220] In one embodiment, the first processed text feature vector F' q It can be calculated by the following formula (14):

[0221]

[0222] Wherein, dk represents a coefficient. It should be noted that other normalization functions can also be used to replace Softmax in formula (14), and this disclosure does not limit this.

[0223] Step S1105 : obtaining predicted answer text and predicted identification information based on the first processed text feature vector and the text feature vector.

[0224] In one exemplary embodiment, the first processed text feature vector and the image feature vector are concatenated to obtain a concatenated feature vector, and the concatenated feature vector is used to obtain the predicted answer text and the predicted identification information. In another exemplary embodiment, the first processed text feature vector and the image feature vector are further subjected to information filtering, and the filtered feature vector is used to obtain the predicted answer text and the predicted identification information.

[0225] In the above embodiment, the feature query vector Q is derived from the text feature vector, which can be regarded as a "question", and the feature key K and the feature value V are derived from the image feature vector, which means that the image content is used to match and respond to the text question. Therefore, the text feature query vector is guided by the image feature key and the image feature value, that is, the text feature vector is looking to find the corresponding explanation in the image feature vector, thereby realizing the one-way attention from the image feature vector to the text feature vector. The embodiment of the present disclosure increases the expression of the correlation from the image feature vector to the text feature vector. It can be understood that the embodiment of the present disclosure can be combined with the embodiments of steps S701 to S705 in the above embodiment to enhance the two-way expression of the image feature vector and the text feature vector.

[0226] In one embodiment, reference Figure 12 As shown, the initial data processing model includes a feature selection network, which outputs predicted answer text and predicted identification information, including:

[0227] Step S1201: multiply the first parameter matrix in the feature selection network by the image feature vector to obtain a weighted image feature vector, and multiply the second parameter matrix in the feature selection network by the text feature vector to obtain a weighted text feature vector.

[0228] Specifically, the image feature vector is represented as F v , the first parameter matrix is expressed as The weighted image feature vector may include: The text feature vector is represented as F q , the second parameter matrix is expressed as The weighted text feature vector may include:

[0229] Step S1203 , fusing the weighted image feature vector and the weighted text feature vector to obtain a fused feature vector, and normalizing the fused feature vector to obtain a gating matrix.

[0230] In an exemplary embodiment, the weighted image feature vector and the weighted text feature vector are fused, and the fused feature vector may include: In an exemplary embodiment, the fused feature vector is normalized, and the gating matrix Λ is, for example:

[0231]

[0232] It should be noted that in the embodiments of the present disclosure, the normalization processing of the fused feature vector is not limited to the above-mentioned sigmoid function, such as softmax, etc., and technical personnel in the relevant field may make other changes inspired by the technical essence of this application, but as long as the functions and effects achieved are the same or similar to those of this application, they should be covered within the scope of protection of this application.

[0233] Step S1205: performing feature selection on the image feature vector and the text feature vector based on the gating matrix to obtain a selected image feature vector and a selected text feature vector, and obtaining a predicted answer text and predicted identification information based on the selected image feature vector and the selected text feature vector.

[0234] In one exemplary embodiment, the gating matrix may be weighted on the image feature vector and fused with the text feature vector to obtain a selected image feature vector and a selected text feature vector. In one exemplary embodiment, the gating matrix may be weighted on the text feature vector and fused with the image feature vector to obtain a selected image feature vector and a selected text feature vector. Based on the selected image feature vector and the selected text feature vector, a predicted answer text and predicted identification information are obtained.

[0235] In the above embodiment, the weighted image feature vector and the weighted text feature vector are fused to obtain a fused feature vector, and the fused feature vector is normalized to obtain a gating matrix. The gating matrix is used to filter information of the image feature vector and the text feature vector to adapt to different task requirements and input conditions, thereby improving the generalization ability and adaptability of the model.

[0236] In one embodiment, performing feature selection on the image feature vector and the text feature vector based on the gating matrix to obtain the selected image feature vector and the selected text feature vector includes:

[0237] The gating matrix is used as the weight of the text feature vector, and weighted processing is performed on the text feature vector to obtain a second processed text feature vector.

[0238] The image feature vector and the second processed text feature vector are fused to obtain a selected image feature vector and a selected text feature vector.

[0239] Specifically, the gating matrix can be expressed as Λ, and the text feature vector can be expressed as F q , the second processed text feature vector may include: Λ×F q The selected image feature vector and the selected text feature vector can be expressed as F, which can be obtained by the following formula (16):

[0240] F=F v+ Λ×F q (16)

[0241] In one embodiment, reference Figure 13 As shown, the initial data processing model includes an attention mechanism network and a normalization network, and the output prediction answer text includes:

[0242] Step S1301 : Fusing the text feature vector and the image feature vector to obtain a fused feature vector.

[0243] Specifically, in an exemplary embodiment, the text feature vector and the image feature vector may be concatenated to obtain a fused feature vector. In an exemplary embodiment, a gating matrix may be added during the concatenation process to select information from the text feature vector and the image feature vector.

[0244] Step S1303, extracting the text feature value, text feature key value and text feature query vector of the text feature vector of the candidate answer text through the attention mechanism, and extracting the feature value, feature key value and feature query vector of the fused feature vector through the attention mechanism.

[0245] In the embodiment of the present disclosure, the text feature vector of the candidate answer text may be obtained by extracting the features of the candidate answer text. The candidate answer text may be pre-set. Specifically, the text feature vector of the candidate answer text may be represented as A. In an exemplary embodiment, referring to Figure 14 As shown, the feature query vector, feature key value and feature value of the text feature vector A of the candidate answer text can be extracted by the self-attention mechanism, and based on the feature query vector, feature key value and feature value, the text feature vector A' of the processed candidate answer text is determined, and A' can better express the correlation between the internal features of the text feature vector. Then, the text feature value, text feature key value and text feature vector of A' are extracted using the cross-attention mechanism. In an exemplary embodiment, the text feature value, text feature key value and text feature vector of the text feature vector A of the candidate answer text can also be directly extracted using the cross-attention mechanism. In an exemplary embodiment, the feature value, feature key value and feature query vector of the fused feature vector are extracted using the cross-attention mechanism. Among them, the specific method of extracting the feature query vector, feature value and feature key value using the self-attention mechanism or the cross-attention mechanism has been explained in the above embodiment and will not be repeated here.

[0246] Step S1305 : Based on the correlation between the text feature query vector and the feature key value of the fused feature, weighted processing is performed on the feature value of the fused feature to obtain a processed feature value.

[0247] Specifically, the text feature query vector of the candidate answer text can be expressed as Q A , the feature key value of the fusion feature can be expressed as K F , the eigenvalue of the fused feature can be expressed as V F In an exemplary embodiment, the processed eigenvalue A L It can be expressed as follows:

[0248]

[0249] Step S1307: normalize the processed feature values through a normalization network to obtain a predicted probability value for each candidate answer text.

[0250] Step S1309: Select the candidate answer text corresponding to the candidate answer text vector with the largest predicted probability value as the predicted answer text.

[0251] In an exemplary embodiment, the probability of the candidate answer text can be predicted using the Softmax function, which is expressed as follows:

[0252] P=softmax(W A ×A L +b) (18)

[0253] Among them, W A and b are learnable parameters. In one example, the candidate answer text corresponding to the candidate answer text vector with the largest predicted probability value is selected as the predicted answer text.

[0254] It should be noted that in the embodiments of the present disclosure, the normalized network is not limited to the above-mentioned softmax. Technical personnel in the relevant field may make other changes based on the technical essence of this application. However, as long as the functions and effects achieved are the same or similar to those of this application, they should be covered within the scope of protection of this application.

[0255] The above embodiment extracts the text feature values, text feature key values, and text feature query vectors of the text feature vectors of the candidate answer texts through the attention mechanism, thereby extracting the semantic information of the candidate answer texts. Then, based on the correlation between the text feature query vector and the feature key values of the fused features, the feature values of the fused features are weighted to obtain processed feature values; and the processed feature values are normalized through a normalization network. Compared to the traditional technology of predicting candidate answer texts as fixed annotation labels, the disclosed embodiment can improve the accuracy of predicted answer texts and predicted identification information.

[0256] In one embodiment, outputting the prediction identification information includes:

[0257] Fusing the text feature vector and the image feature vector to obtain a fused feature vector;

[0258] Inputting the fused feature vector into an image decoder network and outputting a predicted position of the identification information;

[0259] Based on the predicted position of the identification information, predicted identification information is obtained.

[0260] Specifically, in an exemplary embodiment, the text feature vector and the image feature vector may be concatenated to obtain a fused feature vector. In an exemplary embodiment, a gating matrix may be added during the concatenation process to select information from the text feature vector and the image feature vector.

[0261] The image decoder network is used to predict the position of the bounding box. Specifically, a YOLO network, SSD, or Transformer decoding network can be used. In a specific embodiment, when a bounding box is used as the identification information, the predicted position of the identification information can include the positions of the four vertices of the bounding box. In an exemplary embodiment, the predicted positions of the identification information are concatenated to obtain the predicted identification information.

[0262] In one embodiment, the identification information includes a bounding box, and determining the image loss based on the predicted identification information and the annotated identification information includes:

[0263] A first sub-image loss is obtained based on a deviation between the predicted identification information and the labeled identification information.

[0264] A second sub-image loss is obtained based on the graphic area formed by the predicted identification information and the image area formed by the marked identification information.

[0265] An image loss is obtained based on the first sub-image loss and the second sub-image loss.

[0266] Specifically, the difference between the position of the predicted identification information and the position of the marked identification information is used as the first sub-image loss. In an exemplary embodiment, referring to Figure 15 As shown, the predicted identification information is represented as: b p , the marking information is represented as: b g The second sub-image loss L GIoU It can be expressed as:

[0267]

[0268] Among them, |b g ∩b p| represents the intersection of the image area formed by the labeled identification information and the image area formed by the predicted labeled information, |b p ∪b p The ratio of the union of the image area formed by the labeled information and the image area formed by the predicted label information is called the intersection-and-union ratio. The larger the intersection-and-union ratio, the closer the labeled information is to the predicted information. g , b p ) means including b g and b p The smallest box, |B(b g , b p )-b g ∪b p | indicates Figure 15 The area of the two rectangular boxes in the upper right corner and the lower left corner is the same as B(b g , b p The smaller the ratio of the area of the labeled information to the predicted information, the closer the labeled information is to the predicted information. Here, |.| represents the area of the region.

[0269] In the above embodiment, the second sub-image loss not only considers the intersection-over-union ratio of the predicted identification information and the annotated identification information, but also considers the area within the minimum frame outside the predicted identification information and the annotated identification information, which can more accurately reflect the degree of overlap between the two identification information.

[0270] In one embodiment, the present application further proposes a data processing method, applied to a terminal, the method comprising:

[0271] Get the image and the corresponding question text;

[0272] The image and the question text are input into a data processing model, and the text feature vector of the question text sample and the image feature vector of the image sample are extracted respectively, and the answer text and predicted identification information are output; wherein, the identification information is used to mark the image content that matches the question text, and the data processing model includes training the initial data processing model based on text loss, image loss and bidirectional relevance cost loss, wherein the bidirectional relevance loss includes: obtaining a first transmission weight for calculating relevance from image to text, and based on the first transmission weight, fusing the unit relevance cost between the image features in the image feature vector and the text features in the text feature vector to obtain a first relevance cost; obtaining a second transmission weight for calculating relevance from text to image, and based on the second transmission weight, fusing the unit relevance cost between the text features in the text feature vector and the image features in the image feature vector to obtain a second relevance cost; obtained based on the first relevance cost and the second relevance cost.

[0273] Specifically, the image may include an image of any scene, such as an image of a medical scene, an image of a teaching scene, an image of an exhibition scene, and the like. The question text may include questions related to the image content. For example, in an image of a surgical scene, the question text may include "which surgical instrument was used" and "where was the surgical location". Correspondingly, the annotated answer may include text content that matches the question text. For example, in the above example, the answer text may include "monopolar curved scissors were used". The identification information may include graphic annotation information, pattern annotation information, text annotation information, and the like. For example, identification information is marked in the lower right corner of the image sample, and the identification information indicates location information that matches the question text sample.

[0274] In an exemplary embodiment, the data processing model may include a visual encoder network, wherein the visual encoder network is used to convert image data into a network with a recognizable image feature vector, such as a convolutional neural network, a deep residual network, etc. Optionally, a ResNet18 network is used, which is suitable for training and deployment under resource-limited conditions. The image feature vector is extracted by the visual encoder. In an exemplary embodiment, the data processing model may include a text encoder network, wherein the text encoder network is used to convert text data into a network with a recognizable text feature vector, such as Word2Vec, GloVe, and BERT networks, etc., and the text feature vector is extracted by the text encoder.

[0275] In an exemplary embodiment, text loss is used to measure the difference information between the predicted answer text and the annotated answer text, wherein the smaller the text loss, the closer the two are. In an exemplary embodiment, the text loss can be calculated using a cross-entropy loss function. Image loss can include indicators for measuring the difference information between the predicted identification information and the annotated identification information, such as IoU, GIoU, DIoU, CIoU, and other indicators. In an exemplary embodiment, the image loss function can also be combined with the minimum absolute deviation function to describe the difference information between the position of the predicted identification information and the corresponding position of the corresponding annotated identification information.

[0276] In an exemplary embodiment, the first transmission weight for calculating the relevance of the image to the text may include the weight for transmitting the image feature in the image feature vector to the text feature in the text feature vector, for example, referring to Figure 4 As shown, the image feature vector is represented as F v and the text feature vector is represented as F q , the first transmission weight may include any image feature F in the image feature vector v,i To any text feature F in the text feature vector q,jThe weight for transmission. In an exemplary embodiment, based on the first transmission weight, the unit correlation costs between the image features in the image feature vector and the text features in the text feature vector are weighted and summed to obtain a first correlation cost. In an exemplary embodiment, based on the size of the unit correlation costs between the image features and the text features in the text feature vector, a preset number of unit correlation costs with the smallest unit correlation costs can be screened for fusion processing, and the present disclosure does not impose any restrictions on this.

[0277] In an exemplary embodiment, the second transmission weight for calculating the relevance of text to image may include the weight for transmitting the text feature in the text feature vector to the image feature in the image feature vector, for example, referring to Figure 4 As shown, the image feature vector is represented as F v and the text feature vector is represented as F q , the second transmission weight can include any text feature F in the text feature vector q,j To any image feature F in the image feature vector v,i The weight for transmission. For different image features or different text features, the corresponding second transmission weights can be set to be the same or different, and the present disclosure does not impose any restrictions on this. In an exemplary embodiment, based on the second transmission weight, the unit correlation costs between the text features in the text feature vector and the image features in the image feature vector are weighted and summed to obtain the second correlation cost. In an exemplary embodiment, it is also possible to screen a preset number of unit correlation costs with the smallest unit correlation cost for fusion processing based on the size of the unit correlation cost between the text features and each image feature in the image feature vector, and the present disclosure does not impose any restrictions on this.

[0278] In an exemplary embodiment, the first relevance cost and the second relevance cost may be fused to obtain a bidirectional relevance cost, for example: bidirectional relevance cost loss = first relevance cost + second relevance cost. The first relevance cost and the second relevance cost may be weighted and fused to obtain a bidirectional relevance cost, for example: bidirectional relevance cost loss = a × first relevance cost + b × second relevance cost, where a and b represent weight coefficients.

[0279] Compared to the prior art's one-way focus from image to text or from text to image, the disclosed embodiments can improve the accuracy of data processing model predictions. Furthermore, the disclosed embodiments utilize text features in a text feature vector and image features in an image feature vector when describing the bidirectional relevance cost loss. Compared to the prior art's use of text feature vectors and image feature vectors as processing objects, the disclosed embodiments describe a more fine-grained alignment relationship between features, further improving the accuracy of data processing model predictions.

[0280] In a specific embodiment, the method of the present application can be applied to a scenario in the medical field where questions are asked about a surgical video being played. In the prior art, the surgical visual question location answering system (VQLA) can provide answers to viewers and can also locate content related to the question in the video, thereby enhancing the understanding of the scene. However, the accuracy of the answers and the located content provided by the surgical visual question location answering system in the related art is not high. The present application provides a method for training a data processing model that can improve the accuracy of the data processing model in predicting answers and the accuracy of predicting identification information.

[0281] refer to Figure 16 As shown, the disclosed embodiments may include an image encoding module, a text encoding module, an attention mechanism module, a bidirectional correlation cost loss calculation module, a feature fusion module, an answer text prediction module (ASAM), and an identification information prediction module.

[0282] The image encoding module can use a visual encoder network to extract image feature vectors of image samples during the specific implementation process. The visual encoder network is used to convert image data into a network with recognizable image feature vectors, such as a convolutional neural network, a deep residual network, etc. Optionally, a ResNet18 network is used, which is suitable for training and deployment under limited resources.

[0283] The text encoding module, in the specific implementation process, can use a text encoder network to extract the text feature vector of the problem text sample, such as a network that converts text data into a recognizable text feature vector, such as Word2Vec, GloVe and BERT network.

[0284] The attention mechanism module includes a self-attention mechanism module and a cross-attention mechanism module. The self-attention mechanism module extracts image feature values, image feature key values, and image feature query vectors, respectively, and obtains a second-processed image feature vector based on the image feature values, image feature key values, and image feature query vectors. The self-attention mechanism module extracts text feature values, text feature key values, and text feature query vectors, respectively, and obtains a second-processed text feature vector based on the text feature values, text feature key values, and text feature query vectors. Subsequently, the cross-attention mechanism is used to extract the image feature values, image feature key values, and image feature query vectors of the second-processed image feature vectors, and the cross-attention mechanism is used to extract the text feature values, text feature key values, and text feature query vectors of the second-processed text feature vectors. Based on the correlation between the image feature query vectors and the text feature key values, the text feature values are weighted to obtain the first-processed image feature vectors.

[0285] The module for calculating bidirectional relevance cost loss, during a specific implementation, obtains a first transmission weight for calculating relevance from image to text, and based on the first transmission weight, fuses the unit relevance costs between image features in the image feature vector and text features in the text feature vector to obtain a first relevance cost. A second transmission weight for calculating relevance from text to image is obtained, and based on the second transmission weight, fuses the unit relevance costs between text features in the text feature vector and image features in the image feature vector to obtain a second relevance cost. A bidirectional relevance cost loss is obtained based on the first and second relevance costs.

[0286] A feature fusion module is used to fuse the first processed image feature vector and the text feature vector. In a specific implementation process, the first parameter matrix in the feature selection network is multiplied by the image feature vector to obtain a weighted image feature vector, and the second parameter matrix in the feature selection network is multiplied by the text feature vector to obtain a weighted text feature vector; the weighted image feature vector and the weighted text feature vector are fused to obtain a fused feature vector, and the fused feature vector is normalized to obtain a gating matrix; feature selection is performed on the image feature vector and the text feature vector based on the gating matrix to obtain a fused feature vector.

[0287] The answer text prediction module (ASAM) extracts the text feature values, text feature key values and text feature query vectors of the candidate answer texts through the attention mechanism, and extracts the feature values, feature key values and feature query vectors of the fused feature vectors through the attention mechanism; based on the correlation between the text feature query vector and the feature key values of the fused features, the feature values of the fused features are weighted to obtain processed feature values; the processed feature values are normalized through the normalization network to obtain the predicted probability value of each candidate answer text; and the candidate answer text corresponding to the candidate answer text vector with the largest predicted probability value is selected as the predicted answer text.

[0288] The identification information prediction module inputs the fused feature vector into the image decoder network, outputs the predicted position of the identification information, and obtains the predicted identification information based on the predicted position of the identification information.

[0289] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0290] Based on the same inventive concept, the embodiments of the present application also provide a data processing model training device for implementing the above-mentioned data processing model training method and a data processing device for implementing the above-mentioned data processing method. The implementation solution provided by the device is similar to the implementation solution described in the above-mentioned method, so the specific limitations in the embodiments of the training device for one or more data processing models provided below can be found in the above-mentioned limitations on the data processing model training method, and will not be repeated here.

[0291] In one embodiment, Figure 17 As shown, a training device for a data processing model is provided, comprising:

[0292] The first acquisition module 1701 is used to acquire a sample set; wherein the sample set includes a question text sample and a corresponding image sample, the question text sample has a corresponding annotated answer text, and the image sample has annotated identification information matching the question text sample;

[0293] A first prediction module 1703 is configured to input the question text sample and the corresponding image sample into an initial data processing model, extract the text feature vector of the question text sample and the image feature vector of the image sample, and output a predicted answer text and predicted identification information;

[0294] A first processing module 1705 is configured to obtain a first transmission weight for calculating relevance from the image to the text, and based on the first transmission weight, fuse the unit relevance costs between the image features in the image feature vector and the text features in the text feature vector to obtain a first relevance cost;

[0295] A second processing module 1707 is configured to obtain a second transmission weight for calculating relevance from text to image, and based on the second transmission weight, fuse the unit relevance costs between the text features in the text feature vector and the image features in the image feature vector to obtain a second relevance cost;

[0296] A first calculation module 1709 is configured to obtain a bidirectional correlation cost loss based on the first correlation cost and the second correlation cost;

[0297] A second calculation module 1711 is configured to determine text loss based on the predicted answer text and the annotated answer text, and to determine image loss based on the predicted identification information and the annotated identification information;

[0298] The parameter adjustment module 1713 is used to iteratively adjust the model parameters in the initial data processing model based on the text loss, the image loss and the bidirectional correlation cost loss to obtain a target data processing model, wherein the target data processing model is used to input a question text and a corresponding image, and output an answer text and identification information of an image matching the question text.

[0299] In one embodiment, the first processing module is further configured to:

[0300] Acquire an image feature from the image feature vector, and determine a text feature having the minimum unit correlation cost with the image feature from the text feature vector;

[0301] Based on the first transmission weight, a unit relevance cost between the image feature and the corresponding text feature is fused to obtain a relevance cost corresponding to the image feature;

[0302] Obtaining a next image feature from the image feature vector, and fusing the correlation cost between the next image feature and the corresponding text feature unit based on the first transmission weight until a correlation cost corresponding to each image feature in the image feature vector is obtained;

[0303] The correlation cost corresponding to each image feature in the image feature vector is fused to obtain a first correlation cost.

[0304] In one embodiment, the first processing module is further configured to:

[0305] Obtaining a feature dimension of the image feature vector;

[0306] The feature dimension is normalized to obtain a first transmission weight for performing relevance calculation from the image to the text.

[0307] In one embodiment, the apparatus further comprises:

[0308] A unit relevance cost calculation module is used to obtain image features from the image feature vector and text features from the text feature vector; calculate the similarity between the image features and the text features, and determine the unit relevance cost of the image features to the text features based on the negative correlation relationship between the similarity and the unit relevance cost.

[0309] In one embodiment, the first calculation module is further configured to:

[0310] fusing the first relevance cost and the second relevance cost to obtain a fused relevance cost;

[0311] The fused correlation costs are averaged to obtain a bidirectional correlation cost loss.

[0312] In one embodiment, the first prediction module is further configured to:

[0313] Extracting the image feature value, image feature key value, and image feature query vector of the image feature vector through an attention mechanism, and extracting the text feature value, text feature key value, and text feature query vector of the text feature vector through an attention mechanism;

[0314] Based on the correlation between the image feature query vector and the text feature key value, weighting the text feature value to obtain a first processed image feature vector;

[0315] Based on the first processed image feature vector and the text feature vector, a predicted answer text and predicted identification information are obtained.

[0316] In one embodiment, the first prediction module is further configured to:

[0317] Performing linear calculations on the image feature vector and the query vector matrix, the eigenvalue matrix, and the eigenkey matrix respectively to obtain the corresponding image feature query vector, image eigenvalue, and image feature key value;

[0318] Dividing the image feature query vector into a preset number of parts to obtain a plurality of sub-image query vectors, dividing the image feature value into the preset number of parts to obtain a plurality of sub-image feature values, and dividing the image feature key value into the preset number of parts to obtain a plurality of sub-image feature key values;

[0319] For each set of sub-image query vectors, sub-image feature values, and sub-image feature key values, weighting the sub-image feature values based on the correlation between the sub-image query vectors and the sub-image feature key values to obtain a second processed sub-image feature vector;

[0320] The second-processed sub-image feature vectors corresponding to each group are concatenated to obtain a second-processed image feature vector, and the image feature value, image feature key value, and image feature query vector of the second-processed image feature are extracted through the attention mechanism.

[0321] In one embodiment, the first prediction module is further configured to:

[0322] Performing linear calculations on the text feature vector and the query vector matrix, the eigenvalue matrix, and the eigenkey matrix respectively to obtain corresponding text feature query vectors, text eigenvalues, and text feature key values;

[0323] Dividing the text feature query vector into a preset number of parts to obtain a plurality of sub-text query vectors, dividing the text feature value into the preset number of parts to obtain a plurality of sub-text feature values, and dividing the text feature key value into the preset number of parts to obtain a plurality of sub-text feature key values;

[0324] For each group of sub-text query vectors, sub-text feature values, and sub-text feature key values, weighting the sub-text feature values based on the relevance between the sub-text query vectors and the sub-text feature key values to obtain a second processed sub-text feature vector;

[0325] The second-processed sub-text feature vectors corresponding to each group are concatenated to obtain the second-processed text feature vector, and the text feature value, text feature key value and text feature query vector of the second-processed text feature are extracted through the attention mechanism.

[0326] In one embodiment, the first prediction module is further configured to:

[0327] Extracting the image feature value, image feature key value, and image feature query vector of the image feature vector through an attention mechanism, and extracting the text feature value, text feature key value, and text feature query vector of the text feature vector through an attention mechanism;

[0328] Based on the correlation between the text feature query vector and the image feature key value, weighting the image feature value to obtain a first processed text feature vector;

[0329] Based on the first processed text feature vector and the text feature vector, a predicted answer text and predicted identification information are obtained.

[0330] In one embodiment, the first prediction module is further configured to:

[0331] Performing a product process on the first parameter matrix in the feature selection network and the image feature vector to obtain a weighted image feature vector, and performing a product process on the second parameter matrix in the feature selection network and the text feature vector to obtain a weighted text feature vector;

[0332] fusing the weighted image feature vector and the weighted text feature vector to obtain a fused feature vector, and normalizing the fused feature vector to obtain a gating matrix;

[0333] Feature selection is performed on the image feature vector and the text feature vector based on the gating matrix to obtain a selected image feature vector and a selected text feature vector, and a predicted answer text and predicted identification information are obtained based on the selected image feature vector and the selected text feature vector.

[0334] In one embodiment, the first prediction module is further configured to:

[0335] Using the gating matrix as the weight of the text feature vector, performing weighted processing on the text feature vector to obtain a second processed text feature vector;

[0336] The image feature vector and the second processed text feature vector are fused to obtain a selected image feature vector and a selected text feature vector.

[0337] In one embodiment, the first prediction module is further configured to:

[0338] Fusing the text feature vector and the image feature vector to obtain a fused feature vector;

[0339] Extracting text feature values, text feature key values, and text feature query vectors of the text feature vectors of the candidate answer texts through an attention mechanism, and extracting feature values, feature key values, and feature query vectors of the fused feature vectors through an attention mechanism;

[0340] Based on the correlation between the text feature query vector and the feature key value of the fused feature, weighting the feature value of the fused feature to obtain a processed feature value;

[0341] Normalizing the processed feature values through a normalization network to obtain a predicted probability value for each candidate answer text;

[0342] The candidate answer text corresponding to the candidate answer text vector with the largest predicted probability value is selected as the predicted answer text.

[0343] In one embodiment, the first prediction module is further configured to:

[0344] Fusing the text feature vector and the image feature vector to obtain a fused feature vector;

[0345] Inputting the fused feature vector into an image decoder network and outputting a predicted position of the identification information;

[0346] Based on the predicted position of the identification information, predicted identification information is obtained.

[0347] In one embodiment, the second computing module is configured to:

[0348] Obtaining a first sub-image loss based on a deviation between the predicted identification information and the labeled identification information;

[0349] Obtaining a second sub-image loss based on a graphic area formed by the predicted identification information and an image area formed by the annotated identification information;

[0350] An image loss is obtained based on the first sub-image loss and the second sub-image loss.

[0351] In one embodiment, reference Figure 18 As shown, a data processing device is provided, and the device 1800 includes:

[0352] The second acquisition module is used to obtain the image and the corresponding question text;

[0353] The second prediction module is used to input the image and the question text into the data processing model, extract the text feature vector of the question text sample and the image feature vector of the image sample respectively, and output the answer text and prediction identification information; wherein, the identification information is used to mark the image content matching the question text, and the data processing model includes training the initial data processing model based on text loss, image loss and bidirectional relevance cost loss, wherein the bidirectional relevance loss includes: obtaining a first transmission weight for calculating relevance from image to text, and based on the first transmission weight, fusing the unit relevance cost between the image features in the image feature vector and the text features in the text feature vector to obtain a first relevance cost; obtaining a second transmission weight for calculating relevance from text to image, and based on the second transmission weight, fusing the unit relevance cost between the text features in the text feature vector and the image features in the image feature vector to obtain a second relevance cost; obtained based on the first relevance cost and the second relevance cost.

[0354] Each module in the data processing model training device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0355] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 19 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store training data of the data processing model. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a training method for a data processing model or a data processing method is implemented.

[0356] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 20 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a training method for a data processing model or a data processing method is implemented. The display unit of the computer device is used to form a visually visible image, and can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse, etc.

[0357] Those skilled in the art will understand that Figure 20 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0358] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0359] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0360] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for training a data processing model, characterized in that: The method comprises: Acquire a sample set; wherein the sample set includes a question text sample and a corresponding image sample, the question text sample has a corresponding annotated answer text, and the image sample has annotated identification information matching the question text sample; Inputting the question text sample and the corresponding image sample into an initial data processing model, extracting the text feature vector of the question text sample and the image feature vector of the image sample respectively, and outputting the predicted answer text and predicted identification information; Obtaining a first transmission weight for calculating relevance from the image to the text, and fusing the unit relevance costs between the image features in the image feature vector and the text features in the text feature vector based on the first transmission weight to obtain a first relevance cost; Obtaining a second transmission weight for calculating relevance from text to image, and fusing unit relevance costs between text features in the text feature vector and image features in the image feature vector based on the second transmission weight to obtain a second relevance cost; Obtaining a bidirectional relevance cost loss based on the first relevance cost and the second relevance cost; Determine text loss based on the predicted answer text and the annotated answer text, and determine image loss based on the predicted identification information and the annotated identification information; Based on the text loss, the image loss and the bidirectional correlation cost loss, the model parameters in the initial data processing model are iteratively adjusted to obtain a target data processing model, wherein the target data processing model is used to input a question text and a corresponding image, and output an answer text and identification information of an image matching the question text.

2. The method according to claim 1, characterized in that Based on the first transmission weight, a unit correlation cost between the image feature in the image feature vector and the text feature in the text feature vector is fused to obtain a first correlation cost, including: Acquire an image feature from the image feature vector, and determine a text feature having the minimum unit correlation cost with the image feature from the text feature vector; Based on the first transmission weight, a unit relevance cost between the image feature and the corresponding text feature is fused to obtain a relevance cost corresponding to the image feature; Obtaining a next image feature from the image feature vector, and fusing the correlation cost between the next image feature and the corresponding text feature unit based on the first transmission weight until a correlation cost corresponding to each image feature in the image feature vector is obtained; The correlation cost corresponding to each image feature in the image feature vector is fused to obtain a first correlation cost.

3. The method according to claim 2, characterized in that Obtaining a first transmission weight for calculating relevance from an image to a text, including: Obtaining a feature dimension of the image feature vector; The feature dimension is normalized to obtain a first transmission weight for performing relevance calculation from the image to the text.

4. The method according to claim 1, wherein Before fusing the unit correlation cost between the image features in the image feature vector and the text features in the text feature vector based on the first transmission weight, the method further includes: Acquiring image features from the image feature vector and acquiring text features from the text feature vector; The similarity between the image feature and the text feature is calculated, and based on a negative correlation relationship between the similarity and the unit relevance cost, the unit relevance cost from the image feature to the text feature is determined.

5. The method according to claim 1, wherein A bidirectional correlation cost loss is obtained based on the first correlation cost and the second correlation cost, including: fusing the first relevance cost and the second relevance cost to obtain a fused relevance cost; The fused correlation costs are averaged to obtain a bidirectional correlation cost loss.

6. The method according to claim 1, characterized in that Outputs the predicted answer text and prediction identification information, including: Extracting the image feature value, image feature key value, and image feature query vector of the image feature vector through an attention mechanism, and extracting the text feature value, text feature key value, and text feature query vector of the text feature vector through an attention mechanism; Based on the correlation between the image feature query vector and the text feature key value, weighting the text feature value to obtain a first processed image feature vector; Based on the first processed image feature vector and the text feature vector, a predicted answer text and predicted identification information are obtained.

7. The method according to claim 6, characterized in that The extracting the image feature value, the image feature key value, and the image feature query vector of the image feature vector by using the attention mechanism includes: Performing linear calculations on the image feature vector and the query vector matrix, the eigenvalue matrix, and the eigenkey matrix respectively to obtain the corresponding image feature query vector, image eigenvalue, and image feature key value; Divide the image feature query vector into a preset number of parts to obtain a plurality of sub-image query vectors, divide the image feature value into the preset number of parts to obtain a plurality of sub-image feature values, and divide the image feature key value into the preset number of parts to obtain a plurality of sub-image feature key values; For each set of sub-image query vectors, sub-image feature values, and sub-image feature key values, weighting the sub-image feature values based on the correlation between the sub-image query vectors and the sub-image feature key values to obtain a second processed sub-image feature vector; The second-processed sub-image feature vectors corresponding to each group are concatenated to obtain a second-processed image feature vector, and the image feature value, image feature key value, and image feature query vector of the second-processed image feature are extracted through an attention mechanism.

8. The method according to claim 6, characterized in that Extracting a text feature value, a text feature key value, and a text feature query vector of the text feature vector through an attention mechanism includes: Performing linear calculations on the text feature vector and the query vector matrix, the eigenvalue matrix, and the eigenkey matrix respectively to obtain corresponding text feature query vectors, text eigenvalues, and text feature key values; Divide the text feature query vector into a preset number of parts to obtain multiple sub-text query vectors, divide the text feature value into the preset number of parts to obtain multiple sub-text feature values, and divide the text feature key value into the preset number of parts to obtain multiple sub-text feature key values; For each group of sub-text query vectors, sub-text feature values, and sub-text feature key values, weighting the sub-text feature values based on the relevance between the sub-text query vectors and the sub-text feature key values to obtain a second processed sub-text feature vector; The second-processed sub-text feature vectors corresponding to each group are concatenated to obtain the second-processed text feature vector, and the text feature value, text feature key value and text feature query vector of the second-processed text feature are extracted through the attention mechanism.

9. The method according to claim 1, characterized in that Outputs the predicted answer text and prediction identification information, including: Extracting the image feature value, image feature key value, and image feature query vector of the image feature vector through an attention mechanism, and extracting the text feature value, text feature key value, and text feature query vector of the text feature vector through an attention mechanism; Based on the correlation between the text feature query vector and the image feature key value, weighting the image feature value to obtain a first processed text feature vector; Based on the first processed text feature vector and the text feature vector, a predicted answer text and predicted identification information are obtained.

10. The method according to claim 1, characterized in that The initial data processing model includes a feature selection network, which outputs predicted answer text and predicted identification information, including: Performing a product process on the first parameter matrix in the feature selection network and the image feature vector to obtain a weighted image feature vector, and performing a product process on the second parameter matrix in the feature selection network and the text feature vector to obtain a weighted text feature vector; fusing the weighted image feature vector and the weighted text feature vector to obtain a fused feature vector, and normalizing the fused feature vector to obtain a gating matrix; Feature selection is performed on the image feature vector and the text feature vector based on the gating matrix to obtain a selected image feature vector and a selected text feature vector, and a predicted answer text and predicted identification information are obtained based on the selected image feature vector and the selected text feature vector.

11. The method according to claim 10, characterized in that Performing feature selection on the image feature vector and the text feature vector based on the gating matrix to obtain a selected image feature vector and a selected text feature vector includes: Using the gating matrix as the weight of the text feature vector, performing weighted processing on the text feature vector to obtain a second processed text feature vector; The image feature vector and the second processed text feature vector are fused to obtain a selected image feature vector and a selected text feature vector.

12. The method according to claim 1, characterized in that The initial data processing model includes an attention mechanism network and a normalization network, and the output prediction answer text includes: Fusing the text feature vector and the image feature vector to obtain a fused feature vector; Extracting text feature values, text feature key values, and text feature query vectors of the text feature vectors of the candidate answer texts through an attention mechanism, and extracting feature values, feature key values, and feature query vectors of the fused feature vectors through an attention mechanism; Based on the correlation between the text feature query vector and the feature key value of the fused feature, weighting the feature value of the fused feature to obtain a processed feature value; Normalizing the processed feature values through a normalization network to obtain a predicted probability value for each candidate answer text; The candidate answer text corresponding to the candidate answer text vector with the largest predicted probability value is selected as the predicted answer text.

13. The method according to claim 1, wherein Output prediction identification information, including: Fusing the text feature vector and the image feature vector to obtain a fused feature vector; Inputting the fused feature vector into an image decoder network and outputting a predicted position of the identification information; Based on the predicted position of the identification information, predicted identification information is obtained.

14. The method according to claim 1, wherein Based on the predicted identification information and the annotated identification information, the image loss is determined, including: Obtaining a first sub-image loss based on a deviation between the predicted identification information and the labeled identification information; Obtaining a second sub-image loss based on a graphic area formed by the predicted identification information and an image area formed by the annotated identification information; An image loss is obtained based on the first sub-image loss and the second sub-image loss.

15. A data processing method, characterized in that: The method comprises: Get the image and the corresponding question text; The image and the question text are input into a data processing model, and the text feature vector of the question text sample and the image feature vector of the image sample are extracted respectively, and the answer text and predicted identification information are output; wherein, the identification information is used to mark the image content that matches the question text, and the data processing model includes training the initial data processing model based on text loss, image loss and bidirectional relevance cost loss, wherein the bidirectional relevance loss includes: obtaining a first transmission weight for calculating relevance from image to text, and based on the first transmission weight, fusing the unit relevance cost between the image features in the image feature vector and the text features in the text feature vector to obtain a first relevance cost; obtaining a second transmission weight for calculating relevance from text to image, and based on the second transmission weight, fusing the unit relevance cost between the text features in the text feature vector and the image features in the image feature vector to obtain a second relevance cost; obtained based on the first relevance cost and the second relevance cost.

16. A data processing model training device, characterized in that: The device comprises: A first acquisition module is configured to acquire a sample set, wherein the sample set includes a question text sample and a corresponding image sample, the question text sample has a corresponding annotated answer text, and the image sample has annotated identification information matching the question text sample; a first prediction module, configured to input the question text sample and the corresponding image sample into an initial data processing model, extract a text feature vector of the question text sample and an image feature vector of the image sample, and output a predicted answer text and predicted identification information; a first processing module, configured to obtain a first transmission weight for calculating relevance from the image to the text, and to fuse the unit relevance costs between the image features in the image feature vector and the text features in the text feature vector based on the first transmission weight to obtain a first relevance cost; a second processing module, configured to obtain a second transmission weight for calculating relevance from text to image, and to fuse the unit relevance costs between text features in the text feature vector and image features in the image feature vector based on the second transmission weight to obtain a second relevance cost; A first calculation module, configured to obtain a bidirectional correlation cost loss based on the first correlation cost and the second correlation cost; A second calculation module is used to determine the text loss based on the predicted answer text and the annotated answer text, and to determine the image loss based on the predicted identification information and the annotated identification information; A parameter adjustment module is used to iteratively adjust the model parameters in the initial data processing model based on the text loss, the image loss and the bidirectional correlation cost loss to obtain a target data processing model, wherein the target data processing model is used to input a question text and a corresponding image, and output an answer text and identification information of an image matching the question text.

17. A data processing device, characterized in that: The device comprises: The second acquisition module is used to obtain the image and the corresponding question text; The second prediction module is used to input the image and the question text into the data processing model, extract the text feature vector of the question text sample and the image feature vector of the image sample respectively, and output the answer text and prediction identification information; wherein, the identification information is used to mark the image content matching the question text, and the data processing model includes training the initial data processing model based on text loss, image loss and bidirectional relevance cost loss, wherein the bidirectional relevance loss includes: obtaining a first transmission weight for calculating relevance from image to text, and based on the first transmission weight, fusing the unit relevance cost between the image features in the image feature vector and the text features in the text feature vector to obtain a first relevance cost; obtaining a second transmission weight for calculating relevance from text to image, and based on the second transmission weight, fusing the unit relevance cost between the text features in the text feature vector and the image features in the image feature vector to obtain a second relevance cost; obtained based on the first relevance cost and the second relevance cost.

18. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the processor implements the steps of the method according to any one of claims 1 to 14 or the steps of the method according to claim 15.

19. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 14 or the steps of the method according to claim 15 are implemented.

20. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 14 or the steps of the method according to claim 15.