Multi-modal data enhancement method, apparatus and device, and computer program product
By extracting text subject information and generating similar pictures, the problem of semantic inconsistency in the existing multimodal data augmentation method is solved, the semantic consistency and authenticity of text and picture augmentation data is realized, and the effect of data augmentation and the generalization ability of the model is improved.
Patent Information
- Application Number
- CN202510080675.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
When the existing multimodal data enhancement method enhances the pairing of pictures and text data, it is easy to lead to semantic inconsistency, insufficient contextual correlation and unnatural combinations, reducing the authenticity, accuracy and effectiveness of data enhancement.
A multimodal data enhancement method is proposed. By obtaining the multimodal annotation data set, the text subject information is extracted, and the text enhancement data and the target content similar pictures are generated based on the text subject information and the original picture data. Then, the picture enhancement data is generated based on the original picture and the target content similar pictures, and finally the text and picture enhancement data are combined to generate the multimodal enhancement data set.
The semantic consistency of text and picture enhancement data is achieved, the authenticity and accuracy of data enhancement is improved, the model's understanding of the relevance of graphics and texts is enhanced, and the model's generalization ability is improved.
Smart Images

Figure CN120011811A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a multimodal data enhancement method, device, equipment and computer program product. Background Art
[0002] The current mainstream multimodal data enhancement methods include image-text splicing, image conversion, text transformation and other methods. Among them, image-text splicing achieves data enhancement by combining images of different scenes and corresponding descriptions; image conversion achieves data enhancement through image rotation, cropping, color conversion, etc.; text transformation involves synonym replacement, sentence structure changes, etc. Although the above methods can enrich training data, their shortcomings are that after processing the data by the above methods, it will lead to semantic inconsistency, insufficient context association and may introduce unnatural combinations. For example, data with paired images and texts need to enhance the image data and text data respectively, which may lead to semantic inconsistency of the data after enhancement, that is, the text does not match the image content, thereby reducing the authenticity, accuracy and effectiveness of data enhancement. Summary of the invention
[0003] The main purpose of the present application is to provide a multimodal data enhancement method, apparatus, device and computer program product, which aims to solve the problem that the authenticity, accuracy and effectiveness of data enhancement are reduced because the image and text paired data need to be enhanced separately.
[0004] To achieve the above objectives, the present application proposes a multimodal data enhancement method, the method comprising:
[0005] Acquire a multimodal annotated data set, wherein the multimodal annotated data set includes a plurality of sets of image-text associated data, and the image-text associated data includes original text data and original image data;
[0006] Extracting text body information of each of the original text data, and obtaining a plurality of groups of text enhancement data and target content-similar images corresponding to the text enhancement data based on the text body information and the original image data;
[0007] Based on the original picture data and the target content-similar pictures, a plurality of sets of picture enhancement data are obtained;
[0008] The text enhancement data are associated and combined with the image enhancement data to generate a multimodal enhancement data set.
[0009] In one embodiment, extracting the text body information of each of the original text data, and obtaining a plurality of sets of text enhancement data and target content similar images corresponding to the text enhancement data based on the text body information and the original image data, includes:
[0010] For any of the image-text associated data, extracting text body information of the original text data through a text body extraction model;
[0011] Matching the text body information with a preset knowledge graph to obtain a number of text extension words;
[0012] Performing text expansion according to each of the text expansion words to generate a plurality of groups of text enhancement data similar to the image-text association data;
[0013] Based on each of the original picture data, a target content similar picture corresponding to each of the text enhancement data is determined.
[0014] In one embodiment, the text expansion is performed according to each of the text expansion words to generate a plurality of groups of text enhancement data similar to the image-text association data, including:
[0015] Obtaining the subject association relationship between each of the text expansion words;
[0016] Based on the subject association relationship between each of the text expansion words, each of the text expansion words is associated with the original text data and input into a text expansion model to obtain a plurality of first text expansion contents output by the text expansion model;
[0017] For any of the text expansion words, input the text expansion word and the original text data into a text expansion model to obtain a second text expansion content output by the text expansion model;
[0018] The first text-expanded contents and the second text-expanded contents are associated and combined to generate a plurality of groups of text enhancement data similar to the image-text associated data.
[0019] In one embodiment, determining the target content similar picture corresponding to each text enhancement data based on each original picture data includes:
[0020] For any of the text enhancement data, the text enhancement data is vectorized to obtain a first semantic vector value, and all the original text data in the multimodal annotation data set are vectorized to obtain a plurality of second semantic vector values;
[0021] Calculating the Euclidean distance between the first semantic vector value and each of the second semantic vector values;
[0022] Based on each of the Euclidean distance values and each of the original picture data, a target content similar picture corresponding to the text enhancement data is determined.
[0023] In one embodiment, determining the target content similar picture corresponding to the text enhancement data based on each of the Euclidean distance values and each of the original picture data includes:
[0024] For any of the Euclidean distance values, comparing the Euclidean distance value with a preset similarity threshold range;
[0025] If the Euclidean distance value meets the preset similarity threshold range, the original picture data associated with the original text data corresponding to the Euclidean distance value is used as the candidate content similar picture corresponding to the text enhancement data, so as to obtain a set of candidate content similar pictures corresponding to each of the Euclidean distance values;
[0026] Based on the candidate content-similar picture set, a target content-similar picture corresponding to the text enhancement data is determined.
[0027] In one embodiment, the obtaining of several sets of picture enhancement data based on each of the original picture data and each of the target content similar pictures includes:
[0028] Extracting a first element set corresponding to the original image data and a second element set corresponding to each of the target content-similar images through a picture element extraction model;
[0029] Outputting a first depth feature vector corresponding to the original picture data and a second depth feature vector corresponding to each of the target content-similar pictures through a picture depth vector output model;
[0030] Based on the first element set, each set of the second elements, the first depth of field feature vector, and each set of the second depth of field feature vector, several groups of picture enhancement data are obtained.
[0031] In one embodiment, the obtaining of several sets of picture enhancement data based on the first element set, each set of the second element, the first depth of field feature vector, and each second depth of field feature vector includes:
[0032] Based on the first element set and each of the second element sets, calculating an element overlap rate between the picture data and each of the target content similar pictures, and based on the first depth of field feature vector and each of the second depth of field feature vectors, calculating a picture difference rate between the picture data and each of the target content similar pictures;
[0033] Based on the overlap rate of each element and the difference rate of each picture, determine a pad image reference mode, and generate a plurality of pictures to be enhanced corresponding to the pad image reference mode;
[0034] Each of the to-be-enhanced pictures is used as a picture gasket, and picture enhancement data corresponding to the to-be-enhanced picture is output through a picture generation model to obtain several groups of picture enhancement data.
[0035] In addition, to achieve the above-mentioned purpose, the present application also proposes a multimodal data enhancement device, wherein the multimodal data enhancement device comprises:
[0036] A data acquisition module, used to acquire a multimodal annotation data set, wherein the multimodal annotation data set includes a plurality of sets of image-text association data, and the image-text association data includes original text data and original image data;
[0037] A text enhancement module, used to extract text body information of each of the original text data, and based on each of the text body information and each of the original picture data, obtain a plurality of groups of text enhancement data and target content-similar pictures corresponding to the text enhancement data;
[0038] A picture enhancement module, used for obtaining a plurality of groups of picture enhancement data based on the original picture data and the target content similar pictures;
[0039] The combination generation module is used to associate and combine each of the text enhancement data with each of the image enhancement data to generate a multimodal enhancement data set.
[0040] In addition, to achieve the above-mentioned objectives, the present application also proposes a multimodal data enhancement device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the multimodal data enhancement method described above.
[0041] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the multimodal data enhancement method described above are implemented.
[0042] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the multimodal data enhancement method described above.
[0043] The present application provides a multimodal data enhancement method, apparatus, device and computer program product. The multimodal data enhancement method obtains a multimodal annotation data set, wherein the multimodal annotation data set includes several groups of image-text association data, and the image-text association data includes original text data and original image data, and then extracts the text body information of each of the original text data, and based on each of the text body information and each of the original image data, obtains several groups of text enhancement data and target content similar images corresponding to the text enhancement data, and then based on each of the original image data and each of the target content similar images, obtains several groups of image enhancement data, and then associates and combines each of the text enhancement data with each of the image enhancement data to generate a multimodal enhancement data set, thereby realizing a two-stage data enhancement method that combines text enhancement and image enhancement and is interconnected, enhances the multimodal annotation data from the content level, and ensures the semantic consistency and content authenticity of the enhanced data during the enhancement process. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0046] Figure 1 A flowchart of the first embodiment of the multimodal data enhancement method of the present application is provided;
[0047] Figure 2 An example diagram of the original image provided for the multimodal data enhancement method of this application;
[0048] Figure 3 A flowchart of the second embodiment of the multimodal data enhancement method of the present application is provided;
[0049] Figure 4 A brief flowchart of the multimodal data enhancement method provided in this application;
[0050] Figure 5 This is a schematic diagram of the module structure of the multimodal data enhancement device according to an embodiment of the present application;
[0051] Figure 6 Schematic diagram of the device structure of the hardware operating environment involved in the multimodal data enhancement method in the embodiment of the present application.
[0052] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0053] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0054] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0055] The current mainstream multimodal data enhancement methods include image-text splicing, image conversion, text transformation and other methods. Among them, image-text splicing achieves data enhancement by combining images of different scenes and corresponding descriptions; image conversion achieves data enhancement by image rotation, cropping, color conversion, etc.; text transformation involves synonym replacement, sentence structure change, etc. Although the above methods can enrich training data, their shortcomings are that after processing the data by the above methods, it will lead to semantic inconsistency, insufficient context association and possible introduction of unnatural combinations. Among them, multimodal data enhancement requires data of different modalities to be enhanced separately, such as data paired with pictures and texts. The need to enhance the image data and text data separately will lead to semantic inconsistency in the data after enhancement, that is, the text does not match the image content; at the same time, the content after data enhancement will not conform to the actual situation, and the content expressed will not occur in the real scene, which will cause the model to learn unrealistic scene content and reduce the generalization ability of the model. For example, enhancing the image data of a football game will generate useless and erroneous data such as a forward holding the ball; and the current method is very limited in content expansion, and can only transform the position or angle of individual elements in the existing picture, without substantially increasing the data of the content scene, so there is no significant improvement in the scene of model learning.
[0056] Therefore, this embodiment divides data enhancement into two stages: text enhancement and image enhancement. The two stages are processed independently but logically connected, thereby solving the following three technical problems: text enhancement is performed first and then image enhancement. Image enhancement depends on the result of text enhancement. The two stages are associated to ensure the semantic consistency of image data and text data after data enhancement; through text generation and matching methods, the direction and content of text generation are constrained to ensure the accuracy and authenticity of the content after text enhancement; using a generative image enhancement method, a completely new annotated image is generated based on the existing image in the image enhancement stage, thereby achieving the expansion of the annotated data in terms of the content included.
[0057] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a big data service platform, a multimodal data enhancement system, etc. The following takes the multimodal data enhancement system as an example to illustrate this embodiment and the following embodiments.
[0058] Based on this, the present application embodiment provides a multimodal data enhancement method, referring to Figure 1 , Figure 1 A flowchart of the first embodiment of the multimodal data enhancement method of the present application is provided.
[0059] In this embodiment, the multimodal data enhancement method includes steps S11 to S14:
[0060] Step S11, obtaining a multimodal annotation data set, wherein the multimodal annotation data set includes a plurality of sets of image-text association data, and the image-text association data includes original text data and original image data;
[0061] It should be noted that the multimodal annotated dataset refers to a data set containing multiple modalities (such as text, images, etc.). There is a certain correlation or correspondence between these data, and each set of data carries annotation information, which is often used to train machine learning models so that they can understand and process different types of data.
[0062] It should be further explained that the image-text associated data refers to a set of data in which text data and image data are associated with each other in a multimodal annotation dataset. For example, a picture and its corresponding descriptive text. The original text data refers to the original text information in the image-text associated data without data enhancement, which usually describes the content of the picture associated with it. The original picture data refers to the original image information in the image-text associated data without data enhancement.
[0063] Specifically, a multimodal annotation dataset is obtained, wherein the multimodal annotation dataset includes several groups of image-text association data, and the image-text association data includes original text data and original image data. In one embodiment, the obtained multimodal annotation dataset is a multimodal annotation data of text-image pairing, and the format of the input data is: {original text data (text content), original image data (image ID (Identification))} (i.e., image-text association data), wherein the original text data describes the content of the image, and the image ID points to a specific image. Taking a football game scene as an example, the input data can be: {text: (a player is shooting, and the background is the goalkeeper saving the ball), image: (image ID)}
[0064] Step S12, extracting text body information of each of the original text data, and obtaining a plurality of groups of text enhancement data and target content-similar pictures corresponding to the text enhancement data based on the text body information and the original picture data;
[0065] It should be noted that the text body information refers to key information extracted from the original text data, such as entities, actions, events, etc., which represents the core content of the text. The text enhancement data refers to new text data generated based on the original text data through a certain data enhancement technology (such as synonym replacement, sentence structure change, etc.), which is used to expand the data set and improve the generalization ability of the model. The target content similar images refer to images that are similar or related in content to the text enhancement data, which are selected from the data set through a certain matching mechanism and matched with the enhanced text data to maintain the semantic consistency of the image and text.
[0066] Specifically, for any of the image-text associated data, the text body information of the original text data is extracted through a text body extraction model, and then the text body information is matched with a preset knowledge graph to obtain a number of text extension words, so that text expansion is performed according to each of the text extension words to generate a number of groups of text enhancement data similar to the image-text associated data, and then based on each of the original image data, a target content similar image corresponding to each of the text enhancement data is determined.
[0067] Step S13, obtaining a plurality of sets of picture enhancement data based on the original picture data and the target content-similar pictures;
[0068] It should be noted that the image enhancement data refers to new image data generated by an image generation model based on the original image data and each of the target content-similar images, so as to expand the image data set.
[0069] Specifically, a first element set corresponding to the original image data and a second element set corresponding to each of the target content similar images are extracted through a picture element extraction model, and then a first depth of field feature vector corresponding to the original image data and a second depth of field feature vector corresponding to each of the target content similar images are output through a picture depth of field vector output model, so as to obtain several groups of picture enhancement data based on the first element set, each of the second element sets, the first depth of field feature vector and each of the second depth of field feature vectors.
[0070] Step S14: Associating and combining each of the text enhancement data with each of the image enhancement data to generate a multimodal enhancement data set.
[0071] It should be noted that the multimodal enhanced dataset refers to a collection of text enhancement data and image enhancement data to maintain the original text-image relevance and increase the diversity and complexity of the image-text data, which is used to train the multimodal learning model so that it can better understand and process various types of data.
[0072] Specifically, each of the text enhancement data and each of the image enhancement data are associated and combined to generate a multimodal enhancement data set, thereby providing a wider range of samples through the multimodal enhancement data set, helping the model to maintain stable performance when facing different scenarios and conditions, and improving the generalization ability of the model. At the same time, the association and combination of text enhancement data and image enhancement data ensures the semantic consistency between images and texts, which is particularly important for multimodal learning tasks, so as to improve the model's understanding of the relevance between images and texts. And by generating images with similar target content based on the original image data, the authenticity and accuracy of the enhanced data are guaranteed, avoiding the generation of data that does not match the real scene, reducing the erroneous information learned by the model, and then reducing the dependence on manual annotation by automatically generating enhanced data sets, reducing the cost and time of data preparation.
[0073] This embodiment obtains a multimodal annotation dataset, wherein the multimodal annotation dataset includes several groups of image-text association data, and the image-text association data includes original text data and original image data, and then extracts the text body information of each of the original text data, and based on each of the text body information and each of the original image data, obtains several groups of text enhancement data and target content similar images corresponding to the text enhancement data, and then based on each of the original image data and each of the target content similar images, obtains several groups of image enhancement data, and then associates and combines each of the text enhancement data with each of the image enhancement data to generate a multimodal enhancement dataset, thereby realizing a two-stage interconnected data enhancement method combining text enhancement and image enhancement, enhancing the multimodal annotation data from the content level, and ensuring the semantic consistency and content authenticity of the enhanced data during the enhancement process.
[0074] In a feasible implementation manner, extracting the text body information of each of the original text data, and obtaining a plurality of groups of text enhancement data and target content similar images corresponding to the text enhancement data based on the text body information and the original image data, includes:
[0075] Step S21, for any of the image-text associated data, extracting the text body information of the original text data through a text body extraction model;
[0076] It should be noted that the text body extraction model refers to a natural language processing NLP (Natural Language Processing) model, which is used to identify and extract key information from raw text data, such as entities, keywords, phrases or sentences, which can represent the core content or body of the text. The text body extraction model can be iteratively trained based on machine learning or deep learning technology, such as LSTM (Long Short-Term Memory), BERT (Bidirectional Encoder Representations from Transformers), etc., to analyze the text and extract important information units.
[0077] Specifically, for any of the image-text associated data, the text body information of the original text data is extracted through a text body extraction model. In one embodiment, the LSTM-CRF (Long Short-Term Memory-Conditional Random Field) general model is used to extract the main content information contained in the text, wherein the LSTM-CRF general model is a deep learning model that combines a long short-term memory network (LSTM) and a conditional random field (CRF), and performs particularly well in sequence labeling tasks, such as named entity recognition, part-of-speech tagging, etc. For the LSTM-CRF general model, the LSTM layer is responsible for extracting features from the input sequence and capturing long-term dependencies in the sequence; the CRF layer is responsible for sequence labeling based on the features extracted by the LSTM layer, thereby taking into account the transition probability between labels, improving the accuracy of labeling, and achieving end-to-end training, that is, directly from the input sequence to the final label sequence without the need for complex feature engineering.
[0078] For example, for a set of image-text association data, the input text is "a player is shooting, and the background is the goalkeeper saving the ball", and the input original image data can refer to Figure 2 , so that the text body information extracted by the text body extraction model can be (player, shot, goalkeeper, save), and then the text body information is used as the basic input for text enhancement.
[0079] Step S22, matching the text body information with a preset knowledge graph to obtain a number of text extension words;
[0080] It should be noted that the preset knowledge graph refers to a structured semantic knowledge base, which stores the relationship between entities in the form of a graph. It is a knowledge graph pre-defined and constructed in the process of multimodal data enhancement. The preset knowledge graph contains domain-related entities and association information between entities, which is used to assist in the expansion of the main text information and provide additional information and context for entities mentioned in the text.
[0081] It should be further explained that the text expansion terms refer to the vocabulary or phrases obtained by matching the text body information with the preset knowledge graph, which are used to expand the original text data. The text expansion terms can be synonyms, hyponyms, related entities or other relevant information related to the entities mentioned in the original text, thereby increasing the richness and diversity of the text data and generating text enhancement data that is similar to but richer than the original image-text association data.
[0082] Specifically, the text main body information is matched with the preset knowledge graph to obtain several text extension words. Continuing with the above example, all words in the text main body information are matched with the preset knowledge graph (which can be a knowledge graph related to sports events) to obtain the text extension words corresponding to each word, and obtain the event content related to the text main body information (including characters, teams, actions, events, etc.). For example, according to the text main body information: players, shooting, goalkeepers, saves, the following query expansion results can be obtained: {players: forwards, defenders, goalkeepers; shooting: kicking, aiming, attacking, scoring; goalkeepers: defending, saving, jumping, standing; saves: opening arms, starting, blocking}.
[0083] Among them, words such as player and shot are the original text main body information, while forward, attack and defense are text extension words. Therefore, through the query of the knowledge graph, not only the original main content is expanded, but also the expanded content is guaranteed to be in line with the actual logic, thereby ensuring that the text data is more real and accurate after enhancement.
[0084] Step S23, performing text expansion according to each of the text expansion words to generate a plurality of groups of text enhancement data similar to the image-text association data;
[0085] Specifically, the main association relationship between each of the text expansion words is obtained, and then based on the main association relationship between each of the text expansion words, each of the text expansion words is associated with the original text data and input into a text expansion model to obtain a plurality of first text expansion contents output by the text expansion model, so that for any of the text expansion words, the text expansion words and the original text data are input into a text expansion model to obtain a second text expansion content output by the text expansion model, and then each of the first text expansion contents and each of the second text expansion contents are associated and combined to generate a plurality of groups of text enhancement data similar to the image-text association data.
[0086] Step S24: determining a target content similar picture corresponding to each text enhancement data based on each of the original picture data.
[0087] Specifically, for any of the text enhancement data, the text enhancement data is vectorized to obtain a first semantic vector value, and all the original text data in the multimodal annotation data set are vectorized to obtain a number of second semantic vector values, and then the Euclidean distance value between the first semantic vector value and each of the second semantic vector values is calculated, so as to determine the target content similar image corresponding to the text enhancement data based on each of the Euclidean distance values and each of the original image data.
[0088] This embodiment extracts the text body information of the original text data through a text body extraction model for any of the image-text associated data, and then matches the text body information with a preset knowledge graph to obtain a number of text extension words, thereby performing text expansion according to each of the text extension words to generate a number of groups of text enhancement data similar to the image-text associated data, and then based on each of the original image data, determining a target content similar image corresponding to each of the text enhancement data, thereby ensuring that the generated text enhancement data is semantically consistent with the original data, improving the quality and consistency of the data, and then increasing the diversity of the data set, which is conducive to training a more robust model. At the same time, through operations such as automated text expansion, the cost of manual annotation is reduced and the model generalization ability is enhanced.
[0089] In a feasible implementation manner, the text expansion is performed according to each of the text expansion words to generate a plurality of groups of text enhancement data similar to the image-text association data, including:
[0090] Step S31, obtaining the subject association relationship between each of the text extension words;
[0091] It should be noted that the subject association relationship refers to the semantic connection between various entities (such as names, places, organizations, etc.) or concepts in the text. The subject association relationship can be a logical, semantic or factual connection, thereby revealing how different parts of the text are related to each other. For example, in the sentence "Apple was founded by Steve Jobs", the subject association relationship includes the "founder" relationship between "Apple" and "Steve Jobs".
[0092] Specifically, the subject association relationship between each of the text expansion words can be obtained through the knowledge graph, which is not limited here and can be set according to actual conditions.
[0093] Step S32, based on the subject association relationship between each of the text expansion words, associating each of the text expansion words with the original text data and inputting them into a text expansion model to obtain a plurality of first text expansion contents output by the text expansion model;
[0094] It should be noted that the text expansion model refers to a natural language processing model, which is used to generate new and relevant content based on a given text. It is usually iteratively trained and generated based on machine learning or deep learning technology, such as the GPT (Generative Pre-trained Transformer, pre-trained language) model, so that it can understand the context of the input text and generate coherent, relevant and semantically consistent additional text, which is then used to increase the size of the data set, improve data diversity or enhance text content.
[0095] It should be further explained that the first text expansion content refers to a set of text content generated by the text expansion model based on the subject association relationship between the text expansion words and the original text data. These contents are closely related to the original text in terms of semantics and contain expansion words, aiming to enrich the information content and expression of the original text.
[0096] Specifically, based on the main association relationship between each of the text expansion words, each of the text expansion words is associated with the original text data and input into the text expansion model to obtain a plurality of first text expansion contents output by the text expansion model. For example, "forward" and "attack" are generally linked together, so these two text expansion words are combined together to make the expanded content richer and more realistic.
[0097] For example, the combination of "defense" and "open arms" can be expanded to the following content: "The player kicked the ball, and several defensive players on the opposing team tried to jump up to block it. In the end, the goalkeeper blocked the ball with his arms outstretched." It can be seen that the combined text expansion content has added content about the defensive players, realizing the expansion of the data content.
[0098] Step S33, for any of the text expansion words, input the text expansion word and the original text data into a text expansion model to obtain a second text expansion content output by the text expansion model;
[0099] It should be noted that the second text expansion content refers to another set of text content generated by the text expansion model directly based on each text expansion word and the original text data. Compared with the first text expansion content, the second text expansion content focuses more on a single expansion word and generates text content directly related to the word.
[0100] Specifically, the original text data is expanded using separate text expansion words. Different text expansion words are an expansion angle, thereby achieving the expansion of the original text data from multiple different angles. Not only is text expansion achieved in content, but the expanded text content is also related to the original text data, which ensures the semantic consistency of the text enhancement data to a certain extent.
[0101] For example, the original text data input is "a player is shooting, and the goalkeeper is saving the ball in the background", and the selected text expansion words are "forward" and "position". At this time, two second text expansion contents can be obtained respectively: "The team's forward actively participates in the attack, his running position is very suitable, he shoots the ball, and the goalkeeper stands up to save the ball" and "The player shoots, and the goalkeeper chooses a good position in advance to catch the ball, defusing this threatening shot."
[0102] Step S34: Associating and combining the first text-expanded contents and the second text-expanded contents to generate a plurality of groups of text enhancement data similar to the image-text association data.
[0103] Specifically, the information such as each of the first text expansion contents and each of the second text expansion contents are associated and combined to generate several groups of text enhancement data similar to the image-text association data. In one embodiment, the data structure of a group of text enhancement data at this time is: {original text data, original image data, expanded content set}, wherein the expanded content set includes all expanded new texts (i.e., all first text expansion contents and all second text expansion contents obtained by text expansion based on the original text data).
[0104] This embodiment obtains the subject association relationship between each of the text expansion words, and then, based on the subject association relationship between each of the text expansion words, associates each of the text expansion words with the original text data and inputs them into a text expansion model to obtain a plurality of first text expansion contents output by the text expansion model. Thus, for any of the text expansion words, the text expansion words and the original text data are input into the text expansion model to obtain a second text expansion content output by the text expansion model. Then, each of the first text expansion contents and each of the second text expansion contents are associated and combined to generate a plurality of groups of text enhancement data similar to the graphic-text association data, thereby ensuring that the generated text expansion content is semantically consistent with the original text, avoiding the generation of irrelevant or contradictory information, maintaining semantic consistency, and thereby improving the authenticity and credibility of the generated data. At the same time, by combining the text expansion words and the original text data, a variety of different text expansion contents are generated to increase the richness and diversity of the text data.
[0105] Based on this, the present application embodiment provides a multimodal data enhancement method, referring to Figure 3 , Figure 3 A flowchart of the second embodiment of the multimodal data enhancement method of the present application is provided.
[0106] In a feasible implementation manner, determining the target content similar picture corresponding to each text enhancement data based on each original picture data includes:
[0107] Step S41: for any of the text enhancement data, perform vector conversion on the text enhancement data to obtain a first semantic vector value, and perform vector conversion on all the original text data in the multimodal annotation data set to obtain a plurality of second semantic vector values;
[0108] It should be noted that the first semantic vector value refers to a representation of a single text enhancement data converted into a numerical vector form, which is used to represent the core semantic content of the text enhancement data.
[0109] It should be further explained that the second semantic vector value refers to a representation of an original text data in the multimodal annotation data set converted into a numerical vector form, which is used to represent the semantic content of the original text data. In the data set, each original text data corresponds to a second semantic vector value.
[0110] Specifically, for any of the text enhancement data, the text enhancement data is vectorized to obtain a first semantic vector value, and all the original text data in the multimodal annotation data set are vectorized to obtain several second semantic vector values, wherein the vector conversion can be performed through a BERT (Bidirectional Encoder Representations from Transformers) model, such as a word embedding (Word Embedding) or sentence embedding (Sentence Embedding) model to achieve, without limitation here.
[0111] Step S42, calculating the Euclidean distance between the first semantic vector value and each of the second semantic vector values;
[0112] It should be noted that the Euclidean distance value refers to the straight-line distance between two points in Euclidean space, which is used to measure the difference between two semantic vectors. The smaller the Euclidean distance value, the closer or more similar the two texts are in semantics.
[0113] Specifically, the Euclidean distance value between the first semantic vector value and each of the second semantic vector values is calculated by using a Euclidean distance algorithm.
[0114] Step S43: determining a target content similar picture corresponding to the text enhancement data based on each of the Euclidean distance values and each of the original picture data.
[0115] Specifically, for any of the Euclidean distance values, the Euclidean distance value is compared with a preset similarity threshold range, and if the Euclidean distance value meets the preset similarity threshold range, the original image data associated with the original text data corresponding to the Euclidean distance value is used as a candidate content similar image corresponding to the text enhancement data, so as to obtain a set of candidate content similar images corresponding to each of the Euclidean distance values, and thus determine the target content similar image corresponding to the text enhancement data based on the set of candidate content similar images.
[0116] This embodiment performs vector conversion on any of the text enhancement data to obtain a first semantic vector value, and performs vector conversion on all the original text data in the multimodal annotation data set to obtain a number of second semantic vector values, and then calculates the Euclidean distance value between the first semantic vector value and each of the second semantic vector values, thereby determining a target content similar image corresponding to the text enhancement data based on each of the Euclidean distance values and each of the original image data, thereby more accurately capturing the semantic information of the text, thereby improving the accuracy of text and image matching, ensuring that the selected image and text enhancement data are more semantically consistent, enhancing the accuracy of image-text matching, and thereby effectively improving the quality of the data set, so that when it is subsequently used to train a machine learning model, the model's understanding and classification of the image-text relationship is improved, the model training effect is enhanced, and the model generalization ability is improved.
[0117] In a feasible implementation manner, determining the target content similar picture corresponding to the text enhancement data based on each of the Euclidean distance values and each of the original picture data includes:
[0118] Step S51, for any of the Euclidean distance values, comparing the Euclidean distance value with a preset similarity threshold range;
[0119] It should be noted that the preset similarity threshold range refers to a numerical range pre-set before performing image-text matching, which is used to determine what kind of Euclidean distance value indicates that the text and the image are semantically similar. It is usually set based on experience, data analysis or the needs of a specific application, and is not limited here. If the calculated Euclidean distance value falls within this preset similarity threshold range, then the original text data and text enhancement data corresponding to the Euclidean distance value are considered to be semantically similar enough and can be considered to be matched, and the original image data associated with the original text data can be used as a candidate content-similar image corresponding to the text enhancement data.
[0120] Step S52: if the Euclidean distance value meets the preset similarity threshold range, the original picture data associated with the original text data corresponding to the Euclidean distance value is used as the candidate content similar picture corresponding to the text enhancement data, so as to obtain a set of candidate content similar pictures corresponding to each of the Euclidean distance values;
[0121] It should be noted that the candidate content-similar images refer to images that may be semantically similar to the text enhancement data. Specifically, the original image data whose Euclidean distance value of the first semantic vector value corresponding to the text enhancement data falls within the preset similarity threshold range are regarded as candidate images that potentially match the specific text enhancement data.
[0122] It should be further explained that the candidate content similar picture set refers to a set of candidate content similar pictures screened out for a text enhancement data according to its Euclidean distance value with all the original text data in the data set and in combination with a preset similarity threshold range.
[0123] Specifically, if the Euclidean distance value meets the preset similarity threshold range, the original image data associated with the original text data corresponding to the Euclidean distance value is used as the candidate content similar image corresponding to the text enhancement data to obtain a set of candidate content similar images corresponding to each Euclidean distance value.
[0124] In addition, if the Euclidean distance value does not meet the preset similarity threshold range, the text enhancement data corresponding to the Euclidean distance value is eliminated, thereby ensuring that the image and text pairs in the data set are highly semantically related, thereby improving the quality of the entire data set, and helping to reduce noise interference, improve the robustness of the model, and ensure that the final matched image and text pairs have higher accuracy.
[0125] Step S53: determining the target content-similar picture corresponding to the text enhancement data based on the candidate content-similar picture set.
[0126] Specifically, based on the candidate content similar picture set, the target content similar picture corresponding to the text enhancement data is determined, wherein, if there are multiple candidate content similar pictures in the candidate content similar picture set, a candidate content similar picture with the closest Euclidean distance (i.e., the greatest similarity) is selected as the final target content similar picture; if there is only one candidate content similar picture in the candidate content similar picture set, the candidate content similar picture is directly used as the final target content similar picture to ensure that a set of text enhancement data has only one corresponding target content similar picture. In one embodiment, the text enhancement data after text enhancement output by this step has a data structure of: {original text data, original picture data, content similar picture set}, wherein the structure of a set of data in the content similar picture set is: {text expansion content, target content similar picture}.
[0127] It can be understood that based on the set of candidate content-similar pictures, the target content-similar pictures corresponding to the text enhancement data are determined. On the one hand, this is because the pictures in the original picture data are used as candidate pictures for picture enhancement, and the original pictures are related to the expanded content in terms of content, which can ensure the semantic consistency of the text content and the picture content in the text enhancement data to a certain extent; on the other hand, since the candidate content-similar pictures all come from the original picture data, that is, they are all annotated data, the content they contain is basically true and logical, and picture enhancement based on this can ensure that the enhanced pictures have high authenticity and accuracy.
[0128] This embodiment compares any of the Euclidean distance values with a preset similarity threshold range, and if the Euclidean distance value meets the preset similarity threshold range, uses the original image data associated with the original text data corresponding to the Euclidean distance value as the candidate content similar image corresponding to the text enhancement data, so as to obtain a set of candidate content similar images corresponding to each of the Euclidean distance values, thereby determining the target content similar image corresponding to the text enhancement data based on the set of candidate content similar images, and controlling the accuracy of the match by setting the similarity threshold range, thereby ensuring that the selected image has a high semantic similarity with the text enhancement data.
[0129] In a feasible implementation manner, the obtaining of several groups of image enhancement data based on the original image data and the target content-similar images includes:
[0130] Step S61, extracting a first element set corresponding to the original picture data and a second element set corresponding to each of the target content-similar pictures through a picture element extraction model;
[0131] It should be noted that the image element extraction model is used to automatically identify and extract key elements or objects from a picture. Key elements can be significant visual components such as people, objects, scenes, etc. in the image. The image element extraction model is usually based on deep learning technology, such as convolutional neural networks, to identify and locate specific elements in the image.
[0132] It should be further explained that the first element set refers to the element set extracted from the original image data, including all the key elements identified and extracted from the original image. The second element set refers to the element set extracted from the target content similar image, including the key elements identified and extracted from the target content similar image.
[0133] Specifically, a first element set corresponding to the original image data and a second element set corresponding to each of the target content similar images are extracted through a picture element extraction model. In one embodiment, a stable diffusion model is used to respectively extract the first element set contained in the original image data after text enhancement and the second element set contained in each of the target content similar images, so that the structure of the output result is as follows: {first element set of original images, (second element set 1 of target content similar images, second element set 2 of target content similar images,…, second element set n of target content similar images)}. For example, the first element set of the original image is (goalkeeper, goal, forward, Liverpool, Porto, Champions League), and the second element set 1 of target content similar images is (forward, Liverpool, shooting, defense), etc.
[0134] Step S62, outputting a first depth feature vector corresponding to the original picture data and a second depth feature vector corresponding to each of the target content-similar pictures through a picture depth vector output model;
[0135] It should be noted that the image depth vector output model refers to a model used to analyze and extract the depth information of an image and convert it into a numerical vector that can be used for calculation. The depth information describes the relative distance and depth of different objects in the image, which is very important for understanding the three-dimensional structure of the image.
[0136] It should be further explained that the first depth feature vector refers to a depth feature vector extracted from the original picture data. The first depth feature vector encodes the depth information of the original picture, which can be used for comparison, matching or enhancement operations. In addition, the second depth feature vector refers to a depth feature vector extracted from a picture with similar target content. Similar to the first depth feature vector, the second depth feature vector encodes the depth information of the picture with similar target content, which is used to compare or combine with the depth information of the original picture for further image analysis or enhancement.
[0137] Specifically, the first depth of field feature vector corresponding to the original image data and the second depth of field feature vector corresponding to each of the target content similar images are output through the image depth vector output model. In one embodiment, the depth model of controlNet is used to extract the depth information of the original image data and the target content similar images, and the respective depth of field information is converted into a feature vector to form the following output result: {feature vector of the original image depth of field, (feature vector 1 of the depth of field of the target content similar image, feature vector 2 of the depth of field of the target content similar image, ..., feature vector n of the depth of field of the target content similar image)}.
[0138] Step S63: obtaining a plurality of sets of picture enhancement data based on the first element set, each set of the second element, the first depth of field feature vector, and each second depth of field feature vector.
[0139] Specifically, based on the first element set and each of the second element sets, the element overlap rate between the image data and each of the target content similar pictures is calculated, and based on the first depth of field feature vector and each of the second depth of field feature vector, the picture difference rate between the image data and each of the target content similar pictures is calculated, and then based on each of the element overlap rates and each of the picture difference rates, a pad image reference method is determined, and a number of pictures to be enhanced corresponding to the pad image reference method are generated, so that each of the pictures to be enhanced is used as a picture gasket, and picture enhancement data corresponding to the picture to be enhanced is output through a picture generation model to obtain a number of groups of picture enhancement data.
[0140] This embodiment extracts a first element set corresponding to the original image data and a second element set corresponding to each of the target content similar images through a picture element extraction model, and then outputs a first depth of field feature vector corresponding to the original image data and a second depth of field feature vector corresponding to each of the target content similar images through a picture depth of field vector output model, thereby obtaining several groups of image enhancement data based on the first element set, each of the second element sets, the first depth of field feature vector and each of the second depth of field feature vectors, and then more comprehensively capturing the similarities between images by extracting elements and depth of field features in the images, taking into account not only the existence of specific elements but also the spatial structure and visual effects of the image, thereby more accurately matching the image and text content and improving the accuracy of matching.
[0141] In a feasible implementation manner, the obtaining of several sets of picture enhancement data based on the first element set, each of the second element sets, the first depth of field feature vector, and each of the second depth of field feature vectors includes:
[0142] Step S71, calculating the element overlap rate between the picture data and each of the target content similar pictures based on the first element set and each of the second element sets, and calculating the picture difference rate between the picture data and each of the target content similar pictures based on the first depth of field feature vector and each of the second depth of field feature vectors;
[0143] It should be noted that the element overlap rate refers to the ratio of common elements between two sets of image element sets (for example, the first element set of the original image and the second element set of the target content-similar image), which is obtained by dividing the number of common elements in the two sets by the total number of non-repeated elements after merging, thereby reflecting the similarity between the two sets of images in visual elements.
[0144] It should be further explained that the image difference rate refers to the degree of difference between the first depth of field feature vector of the original image and the second depth of field feature vector of the image with similar target content, which is usually obtained by calculating the distance between the two feature vectors (such as Euclidean distance). The higher the difference rate, the greater the difference in depth of field between the two images.
[0145] Specifically, based on the first element set and each of the second element sets, the element overlap rate between the image data and each of the target content similar images is calculated. In one embodiment, the element set of the original image (the first element set) is compared with the element set of the target content similar image (the second element set) to find out the common elements of the two, that is, the intersection of the two sets. The common elements in the intersection represent the similar parts of the two images in terms of visual content. Then, all elements in the first element set and the second element set are merged together, and duplicate elements are removed, which means that elements that appear in one set but not in the other set, as well as elements that only appear in the other set, will be included in the merged set. Finally, the number of elements in the intersection is divided by the total number of elements in the merged element set, and the result is the element overlap rate, and the following result is obtained: {element overlap rate of target content similar image 1, element overlap rate of target content similar image 2, ..., element overlap rate of target content similar image n}.
[0146] Further, based on the first depth of field feature vector and each of the second depth of field feature vectors, the picture difference rate between the picture data and each of the target content similar pictures is calculated. In one embodiment, the Euclidean distance between the depth of field feature vector of the original picture and the depth of field feature vector of the content similar picture is calculated one by one, and the following results are obtained: {picture difference rate of target content similar picture 1, picture difference rate of target content similar picture 2,…, picture difference rate of target content similar picture n}.
[0147] Understandably, the above approach evaluates the similarity of image content through two dimensions: element overlap rate and depth of field difference. The element overlap rate reflects the consistency of content elements between the original image and the similar image, which helps to measure the similarity of information contained in the two. The image difference evaluates the difference in spatial layout and focus of the two images by calculating the Euclidean distance of the feature vectors. Therefore, this combined method can capture the similarity between images more comprehensively, taking into account not only the existence of specific elements, but also the spatial structure and visual effect of the image.
[0148] Step S72, determining a pad image reference mode based on the overlap rate of each element and the difference rate of each picture, and generating a plurality of pictures to be enhanced corresponding to the pad image reference mode;
[0149] It should be noted that the pad image reference method refers to a method for determining how to use the original image and the target content similar image as a reference or "pad image" when generating new image enhancement data, which may include selecting which image to use as a basis, how to combine elements, adjusting depth of field features, etc., to generate a new image that is visually consistent with the original image or the target content similar image. The image to be enhanced refers to the image to be enhanced, that is, the image to be modified or optimized during the enhancement process.
[0150] Specifically, based on the overlap rate of each element and the difference rate of each picture, a pad image reference mode is determined, and a plurality of pictures to be enhanced corresponding to the pad image reference mode are generated. In one embodiment, the method for determining the pad image reference mode includes:
[0151] The first determination method: If the element overlap rate of the target content-similar picture is greater than or equal to the preset overlap rate threshold and the picture difference rate is less than or equal to the preset difference rate threshold, it means that the original picture and the target content-similar picture are highly similar in picture content and structure, and then the two pictures are horizontally spliced to form a new picture, which is used as the picture to be enhanced for picture generation and as a pad for subsequent picture enhancement. The role of the pad is to make the generated picture related to the annotated picture, so that the newly generated picture enhancement data is associated with the existing original picture data, so that the model can learn this association, improve the model's ability to understand the prompt words and image pictures, and output the following data structure: {text extension content, spliced picture}.
[0152] The second determination method: if the element overlap rate of target content similar pictures is greater than or equal to the preset overlap rate threshold and the picture difference rate is greater than the preset difference rate threshold, or vice versa, the element overlap rate of target content similar pictures is less than the preset overlap rate threshold and the picture difference rate is less than or equal to the preset difference rate threshold, that is, one of the two dimensions does not meet the threshold judgment condition, then it means that the overall similarity of the pictures is not strong enough and cannot fully meet the similarity requirements. At this time, the content set of the original picture and the target content similar picture is added as a keyword to the text expansion content, so that the added text expansion content is used as the description word generated by the subsequent picture, and the corresponding picture is used as the picture to be enhanced, and the following two outputs are obtained: {text expansion content + original picture package {containing elements (first element set), original picture}; {text extended content + target content similar picture containing elements (second element set), target content similar picture}, where there are two cases of pad picture: one is to add the original picture element to the extended content, and use the original picture as the pad picture. The new picture generated in this way is actually an extension and variation of the original picture. In model training, it can be understood with the content related to the original picture element, and this association is logical; the other case is to add the content of the target content similar picture element to the extended content, and use the target content similar picture as the pad picture. This is actually an extension and variation of the target content similar picture, and the effect is similar to the first case.
[0153] It can be understood that by using the above method, the content set of the original picture and the picture with similar target content is used as keywords, and the content is generated by combining the pad picture, which enhances the extensibility and authenticity of the generated content. At the same time, the richness of the keywords makes the generated text more relevant to the picture, and the consistency of the elements ensures that the generated content matches the visual information of the picture, thereby improving the coherence and authenticity of the content.
[0154] The third determination method: If the element overlap rate and the image difference rate of the target content similar picture do not meet the preset overlap rate threshold and the preset difference rate threshold, then the target content similar picture is directly used as the picture to be enhanced, and the output is: {text extension content, target content similar picture}. Because at this time there will be a large difference in content between the original image and the content similar image, if it is used as a pad, it will affect the result of the image generation, making the generated picture lacking in authenticity (it may put the content of events that are far apart together).
[0155] Step S73, using each of the to-be-enhanced pictures as a picture gasket, outputting picture enhancement data corresponding to the to-be-enhanced pictures through a picture generation model, and obtaining several groups of picture enhancement data.
[0156] It should be noted that the image shim refers to an image used as a basis or reference in the image enhancement process, which can be an original image, an image with similar target content, or a combination thereof. The image shim provides a starting point for the enhancement process, and a new enhanced image can be generated by modifying and adjusting the shim.
[0157] Specifically, each of the to-be-enhanced pictures is used as a picture gasket, and the picture generation model is used to output picture enhancement data corresponding to the to-be-enhanced picture, so as to obtain several sets of picture enhancement data. In one embodiment, the stable diffusion model is called to perform picture generation operations through the description words generated in the above steps, and the following processing is performed for the above three different determination methods respectively:
[0158] For the first determination method: if the element overlap rate of the target content-similar pictures is greater than or equal to the preset overlap rate threshold and the picture difference rate is less than or equal to the preset difference rate threshold, then input the text extension content and the spliced picture, and set the weight of the pad image in the picture gasket to above 0.8 (the highest is 1, the higher the more similar), so that the generated picture will be highly consistent with the text extension content in semantics, among which the weight setting can be set according to the actual situation and is not mandatory here.
[0159] For the second determination method: if the element overlap rate of target content similar pictures is greater than or equal to the preset overlap rate threshold and the picture difference rate is greater than the preset difference rate threshold, or vice versa, the element overlap rate of target content similar pictures is less than the preset overlap rate threshold and the picture difference rate is less than or equal to the preset difference rate threshold, then input {text extension content + original picture containing elements, original picture} and {text extension content + target content similar picture containing elements, target content similar picture} respectively. At the same time, for these two inputs, the weight of the pad image should be set to 0.6 respectively. Among them, the weight setting can be set according to the actual situation and is not mandatory here.
[0160] For the third determination method: if the element overlap rate and the image difference rate of the target content similar images do not meet the preset overlap rate threshold and the preset difference rate threshold, then enter {text extension content, target content similar image}, set the corresponding pad image weight to generate the image. Note that the pad image weight is set to 0.5 at this time, where the weight setting can be set according to the actual situation and is not mandatory here.
[0161] Finally, the generated image enhancement data is associated and combined with the corresponding text extension content to obtain the following result: {text enhancement data, image enhancement data}.
[0162] In this embodiment, based on the first element set and each of the second element sets, the element overlap rate between the image data and each of the target content similar images is calculated, and based on the first depth of field feature vector and each of the second depth of field feature vector, the image difference rate between the image data and each of the target content similar images is calculated, and then based on each of the element overlap rates and each of the image difference rates, a padding reference mode is determined, and a plurality of to-be-enhanced images corresponding to the padding reference mode are generated, so that each of the to-be-enhanced images is used as an image pad, and image enhancement data corresponding to the to-be-enhanced images is output through an image generation model to obtain a plurality of groups of image enhancement data, so as to more accurately determine the similarity and difference between the original image and the target content similar image, ensure that the generated image enhancement data is semantically consistent with the original image, improve the accuracy of image-text matching and the quality of the multimodal data set, and thus generate more accurate image enhancement data. At the same time, by determining the padding reference mode, the image generation process is controlled, so that the generated image is more in line with expectations in style, content and structure, and the generated image is more realistic and in line with the description of the enhanced text, and the data is enhanced by a directional generative method, so as to truly expand the data content instead of just transforming the original image.
[0163] For example, to help understand the implementation process of the multimodal data enhancement method, please refer to Figure 4 , Figure 4 A brief flowchart example diagram is provided for the multimodal data enhancement method of this application.
[0164] Specifically, the flowchart shows a systematic multimodal data enhancement process, which extracts raw data from annotated data sets and enhances text data and image data separately. In the text enhancement stage, the main information in the text is first extracted, and then the knowledge graph is used to query the related content to expand the text and generate enhanced text data.
[0165] Furthermore, in the image enhancement stage, the overlap rate and difference of the images are calculated to evaluate the similarity of the image content, and then the appropriate image generation method is selected to achieve image enhancement by generating new images. Finally, the results of text enhancement and image enhancement are combined to form a new and richer modality dataset to improve the quality and diversity of the data and provide more comprehensive data support for the training of machine learning models.
[0166] It should be noted that the examples in the figure are only used to understand the present application and do not constitute a limitation on the multimodal data enhancement method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0167] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.
[0168] This application also provides a multi-modal data enhancement device, please refer to Figure 5 , the multimodal data enhancement device comprises:
[0169] A data acquisition module 51 is used to acquire a multimodal annotation data set, wherein the multimodal annotation data set includes a plurality of sets of image-text association data, and the image-text association data includes original text data and original image data;
[0170] A text enhancement module 52, configured to extract text body information of each of the original text data, and obtain a plurality of groups of text enhancement data and target content-similar images corresponding to the text enhancement data based on the text body information and the original image data;
[0171] The picture enhancement module 53 is used to obtain a plurality of sets of picture enhancement data based on the original picture data and the target content similar pictures;
[0172] The combination generation module 54 is used to associate and combine each of the text enhancement data with each of the image enhancement data to generate a multimodal enhancement data set.
[0173] The multimodal data enhancement device is also used for:
[0174] For any of the image-text associated data, extracting text body information of the original text data through a text body extraction model;
[0175] Matching the text body information with a preset knowledge graph to obtain a number of text extension words;
[0176] Performing text expansion according to each of the text expansion words to generate a plurality of groups of text enhancement data similar to the image-text association data;
[0177] Based on each of the original picture data, a target content similar picture corresponding to each of the text enhancement data is determined.
[0178] The multimodal data enhancement device is also used for:
[0179] Obtaining the subject association relationship between each of the text expansion words;
[0180] Based on the subject association relationship between each of the text expansion words, each of the text expansion words is associated with the original text data and input into a text expansion model to obtain a plurality of first text expansion contents output by the text expansion model;
[0181] For any of the text expansion words, input the text expansion word and the original text data into a text expansion model to obtain a second text expansion content output by the text expansion model;
[0182] The first text-expanded contents and the second text-expanded contents are associated and combined to generate a plurality of groups of text enhancement data similar to the image-text associated data.
[0183] The multimodal data enhancement device is also used for:
[0184] For any of the text enhancement data, the text enhancement data is vectorized to obtain a first semantic vector value, and all the original text data in the multimodal annotation data set are vectorized to obtain a plurality of second semantic vector values;
[0185] Calculating the Euclidean distance between the first semantic vector value and each of the second semantic vector values;
[0186] Based on each of the Euclidean distance values and each of the original picture data, a target content similar picture corresponding to the text enhancement data is determined.
[0187] The multimodal data enhancement device is also used for:
[0188] For any of the Euclidean distance values, comparing the Euclidean distance value with a preset similarity threshold range;
[0189] If the Euclidean distance value meets the preset similarity threshold range, the original picture data associated with the original text data corresponding to the Euclidean distance value is used as the candidate content similar picture corresponding to the text enhancement data, so as to obtain a set of candidate content similar pictures corresponding to each of the Euclidean distance values;
[0190] Based on the candidate content-similar picture set, a target content-similar picture corresponding to the text enhancement data is determined.
[0191] The multimodal data enhancement device is also used for:
[0192] Extracting a first element set corresponding to the original image data and a second element set corresponding to each of the target content-similar images through a picture element extraction model;
[0193] Outputting a first depth feature vector corresponding to the original picture data and a second depth feature vector corresponding to each of the target content-similar pictures through a picture depth vector output model;
[0194] Based on the first element set, each set of the second elements, the first depth of field feature vector, and each set of the second depth of field feature vector, several groups of picture enhancement data are obtained.
[0195] The multimodal data enhancement device is also used for:
[0196] Based on the first element set and each of the second element sets, calculating an element overlap rate between the picture data and each of the target content similar pictures, and based on the first depth of field feature vector and each of the second depth of field feature vectors, calculating a picture difference rate between the picture data and each of the target content similar pictures;
[0197] Based on the overlap rate of each element and the difference rate of each picture, determine a pad image reference mode, and generate a plurality of pictures to be enhanced corresponding to the pad image reference mode;
[0198] Each of the to-be-enhanced pictures is used as a picture gasket, and picture enhancement data corresponding to the to-be-enhanced picture is output through a picture generation model to obtain several groups of picture enhancement data.
[0199] The multimodal data enhancement device provided by the present application adopts the multimodal data enhancement method in the above embodiment, which can solve the technical problems in the background technology. Compared with the prior art, the beneficial effects of the multimodal data enhancement device provided by the present application are the same as the beneficial effects of the multimodal data enhancement method provided by the above embodiment, and the other technical features in the multimodal data enhancement device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0200] The present application provides a multimodal data enhancement device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the multimodal data enhancement method in the above-mentioned embodiment 1.
[0201] Reference below Figure 6, which shows a schematic diagram of the structure of a multimodal data enhancement device suitable for implementing the embodiment of the present application. The multimodal data enhancement device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The multimodal data enhancement device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0202] like Figure 6 As shown, the multimodal data enhancement device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the multimodal data enhancement device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the multimodal data enhancement device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a multimodal data enhancement device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0203] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0204] The multimodal data enhancement device provided by the present application adopts the multimodal data enhancement method in the above embodiment, which can solve the technical problems in the background technology. Compared with the prior art, the beneficial effects of the multimodal data enhancement device provided by the present application are the same as the beneficial effects of the multimodal data enhancement method provided by the above embodiment, and the other technical features in the multimodal data enhancement device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0205] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0206] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0207] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the multimodal data enhancement method in the above-mentioned embodiment.
[0208] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0209] The above-mentioned computer-readable storage medium may be included in the multimodal data enhancement device; or it may exist independently without being assembled into the multimodal data enhancement device.
[0210] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the multimodal data enhancement device, the multimodal data enhancement device:
[0211] Acquire a multimodal annotated data set, wherein the multimodal annotated data set includes a plurality of sets of image-text associated data, and the image-text associated data includes original text data and original image data;
[0212] Extracting text body information of each of the original text data, and obtaining a plurality of groups of text enhancement data and target content-similar images corresponding to the text enhancement data based on the text body information and the original image data;
[0213] Based on the original picture data and the target content-similar pictures, a plurality of sets of picture enhancement data are obtained;
[0214] The text enhancement data are associated and combined with the image enhancement data to generate a multimodal enhancement data set.
[0215] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0216] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0217] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0218] The readable storage medium provided in the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned multimodal data enhancement method, and can solve the technical problems in the background technology. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as the beneficial effects of the multimodal data enhancement method provided in the above-mentioned embodiment, and will not be repeated here.
[0219] An embodiment of the present application provides a computer program product, including a computer program, which implements the steps of the multimodal data enhancement method as described above when executed by a processor.
[0220] The computer program product provided in this application can solve the technical problems in the background technology. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiment of this application are the same as the beneficial effects of the multimodal data enhancement method provided in the above embodiment, which will not be repeated here.
[0221] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A multimodal data enhancement method, characterized in that: include: Acquire a multimodal annotated data set, wherein the multimodal annotated data set includes a plurality of sets of image-text associated data, and the image-text associated data includes original text data and original image data; Extracting text body information of each of the original text data, and obtaining a plurality of groups of text enhancement data and target content-similar images corresponding to the text enhancement data based on the text body information and the original image data; Based on the original picture data and the target content-similar pictures, a plurality of sets of picture enhancement data are obtained; The text enhancement data are associated and combined with the image enhancement data to generate a multimodal enhancement data set.
2. The multimodal data enhancement method according to claim 1, characterized in that: The extracting of the text body information of each of the original text data, and obtaining a plurality of groups of text enhancement data and target content similar images corresponding to the text enhancement data based on the text body information and the original image data, includes: For any of the image-text associated data, extracting text body information of the original text data through a text body extraction model; Matching the text body information with a preset knowledge graph to obtain a number of text extension words; Performing text expansion according to each of the text expansion words to generate a plurality of groups of text enhancement data similar to the image-text association data; Based on each of the original picture data, a target content similar picture corresponding to each of the text enhancement data is determined.
3. The multimodal data enhancement method according to claim 2, wherein: The step of performing text expansion according to each of the text expansion words to generate a plurality of groups of text enhancement data similar to the image-text association data includes: Obtaining the subject association relationship between each of the text expansion words; Based on the subject association relationship between each of the text expansion words, each of the text expansion words is associated with the original text data and input into a text expansion model to obtain a plurality of first text expansion contents output by the text expansion model; For any of the text expansion words, input the text expansion word and the original text data into a text expansion model to obtain a second text expansion content output by the text expansion model; The first text-expanded contents and the second text-expanded contents are associated and combined to generate a plurality of groups of text enhancement data similar to the image-text associated data.
4. The multimodal data enhancement method according to claim 2, wherein: The determining, based on each of the original picture data, a target content similar picture corresponding to each of the text enhancement data comprises: For any of the text enhancement data, the text enhancement data is vectorized to obtain a first semantic vector value, and all the original text data in the multimodal annotation data set are vectorized to obtain a plurality of second semantic vector values; Calculating the Euclidean distance between the first semantic vector value and each of the second semantic vector values; Based on each of the Euclidean distance values and each of the original picture data, a target content similar picture corresponding to the text enhancement data is determined.
5. The multimodal data enhancement method according to claim 4, characterized in that: The determining, based on each of the Euclidean distance values and each of the original picture data, a target content similar picture corresponding to the text enhancement data comprises: For any of the Euclidean distance values, comparing the Euclidean distance value with a preset similarity threshold range; If the Euclidean distance value meets the preset similarity threshold range, the original picture data associated with the original text data corresponding to the Euclidean distance value is used as the candidate content similar picture corresponding to the text enhancement data, so as to obtain a set of candidate content similar pictures corresponding to each of the Euclidean distance values; Based on the candidate content-similar picture set, a target content-similar picture corresponding to the text enhancement data is determined.
6. The multimodal data enhancement method according to claim 1, wherein: The obtaining of a plurality of sets of picture enhancement data based on the original picture data and the target content similar pictures includes: Extracting a first element set corresponding to the original image data and a second element set corresponding to each of the target content-similar images through a picture element extraction model; Outputting a first depth feature vector corresponding to the original picture data and a second depth feature vector corresponding to each of the target content-similar pictures through a picture depth vector output model; Based on the first element set, each set of the second elements, the first depth of field feature vector, and each set of the second depth of field feature vector, several groups of picture enhancement data are obtained.
7. The multimodal data enhancement method according to claim 6, characterized in that: The obtaining of a plurality of sets of picture enhancement data based on the first element set, each set of the second element, the first depth of field feature vector, and each second depth of field feature vector comprises: Based on the first element set and each of the second element sets, calculating an element overlap rate between the picture data and each of the target content similar pictures, and based on the first depth of field feature vector and each of the second depth of field feature vectors, calculating a picture difference rate between the picture data and each of the target content similar pictures; Based on the overlap rate of each element and the difference rate of each picture, determine a pad image reference mode, and generate a plurality of pictures to be enhanced corresponding to the pad image reference mode; Each of the to-be-enhanced pictures is used as a picture gasket, and picture enhancement data corresponding to the to-be-enhanced picture is output through a picture generation model to obtain several groups of picture enhancement data.
8. A multimodal data enhancement device, characterized in that: include: A data acquisition module, used to acquire a multimodal annotation data set, wherein the multimodal annotation data set includes a plurality of sets of image-text association data, and the image-text association data includes original text data and original image data; A text enhancement module, used to extract text body information of each of the original text data, and based on each of the text body information and each of the original picture data, obtain a plurality of groups of text enhancement data and target content-similar pictures corresponding to the text enhancement data; A picture enhancement module, used for obtaining a plurality of groups of picture enhancement data based on the original picture data and the target content similar pictures; The combination generation module is used to associate and combine each of the text enhancement data with each of the image enhancement data to generate a multimodal enhancement data set.
9. A multimodal data enhancement device, characterized in that: The multimodal data enhancement device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the multimodal data enhancement method according to any one of claims 1 to 7.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the multimodal data enhancement method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Foundation measuring system and method for building engineering construction
CN121997440A