Image and text data processing method and device, computer equipment and program
By using weakly aligned multi-language data with semantic relatedness constraints, the method constructs a universal visual-text representation space, addressing the limitations of strictly aligned data, enhancing accuracy in image and text processing tasks.
Patent Information
- Application Number
- JP2025525584
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-26
- Filing Date
- 2023-11-27
- Publication Date
- 2025-11-05
- Estimated Expiration
- 2043-11-27
AI Technical Summary
Existing cross-modal language machine learning models require training with strictly aligned multilingual image and text pairs, limiting the availability and diversity of training data, which affects the accuracy and efficiency of image and text data processing.
A method and apparatus for processing image and text data that utilizes weakly aligned multi-language data, constructing a universal visual-text representation space with semantic relatedness as a constraint, allowing the use of more abundant and easily collected data, and incorporating a task processing assembly for classification or regression tasks.
The approach enhances the accuracy of the universal visual-text representation space by leveraging diverse, easily obtainable weakly aligned data and improves the semantic feature extraction, leading to improved accuracy in classification and regression tasks.
Smart Images

Figure 2025536420000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority from a Chinese patent application filed with the China Patent Office on April 26, 2023, bearing application number 202310477720.6 and entitled "IMAGE AND TEXT DATA PROCESSING METHOD, DEVICE, EQUIPMENT, STORAGE MEDIUM, AND PROGRAM PRODUCT," the entire contents of which are incorporated herein by reference.
[0002] The present application relates to the technical field of artificial intelligence (AI), and in particular to an image and text data processing method and apparatus, and a computer device and program. [Background technology]
[0003] With the continuous development of AI technology, the application of machine learning models in cross-modal language is gaining increasing importance.
[0004] In related technologies, cross-modal language machine learning models usually need to be trained with pictures and text. Specifically, developers pre-collect multilingual image and text pairs as training data to train cross-modal language machine learning models. Among them, the same image is accompanied by corresponding multilingual explanatory texts, and the multilingual texts are translations of each other. Summary of the Invention [Problem to be solved by the invention]
[0005] The embodiments of the present application aim to provide a method and apparatus for processing image and text data, as well as a computer device and program. [Means for solving the problem]
[0006] According to one aspect, there is provided a method for processing image and text data, said method comprising: obtaining first image and text data, the first image and text data including at least one image and at least one text; Perform feature extraction on the first image and text data to map the first image and text data into a universal visual-text representation space to obtain data features of the first image and text data, the universal visual-text representation space being a feature space constructed based on first image and text data samples with a constraint of semantic relatedness (correlation) between text in the first image and text data samples, the first image and text data samples including at least one first image sample and at least two text samples in different languages corresponding to the first image sample, and the text samples in at least two different languages corresponding to the first image sample are not translations of each other; and The method includes sending data features of the first image and text data to a task processing assembly, and the task processing assembly outputting a processing result of a target task based on the data features, wherein the target task is a classification or regression task based on the image and text data.
[0007] According to one aspect, there is provided a method for processing image and text data, said method comprising: constructing a first anchor sample, a first positive example sample, and a first negative example sample based on first image and text data samples, the first image and text data samples including at least one first image sample and text samples in at least two different languages corresponding to the first image sample, the text samples in at least two different languages corresponding to the first image sample being not translations of each other; inputting the first anchor sample, the first positive example sample, and the first negative example sample into a second feature extraction model to obtain first sample features output by the second feature extraction model; obtaining a first loss function value based on the first sample feature and using a semantic relevance between the first image and the text in the text data sample as a constraint; updating parameters of the second feature extraction model according to the first loss function value; and When the second feature extraction model satisfies a convergence condition, a first feature extraction model is constructed based on the second feature extraction model, and the first feature extraction model is used to process input first image and text data input to obtain data features of the first image and text data, and the data features of the first image and text data are processed by a task processing assembly, and then a processing result of a target task is output, and the target task is a classification or regression task based on image and text data.
[0008] According to another aspect, there is provided an image and text data processing apparatus, said apparatus comprising: a data acquisition module for acquiring first image and text data, the first image and text data including at least one image and at least one text; a feature mapping module for mapping the first image and text data into a universal visual-text representation space by performing feature extraction on the first image and text data to obtain data features of the first image and text data, the universal visual-text representation space being a feature space constructed based on first image and text data samples with a semantic relatedness between text in the first image and text data samples as a constraint, the first image and text data samples including at least one first image sample and text samples in at least two different languages corresponding to the first image sample, and the text samples in at least two different languages corresponding to the first image sample are not translations of each other; and The apparatus includes a task processing module for sending data features of the first image and text data to a task processing assembly, and for the task processing assembly to output a processing result of a target task based on the data features, wherein the target task is a classification or regression task based on image and text data.
[0009] According to another aspect, there is provided an image and text data processing apparatus, said apparatus comprising: a sample construction module for constructing first anchor samples, first positive example samples, and first negative example samples based on first image and text data samples, wherein the first image and text data samples include at least one first image sample and text samples in at least two different languages corresponding to the first image sample, and the text samples in at least two different languages corresponding to the first image sample are not translations of each other; a sample input module for inputting the first anchor sample, the first positive example sample, and the first negative example sample into a second feature extraction model to obtain first sample features output by the second feature extraction model; a loss calculation module for obtaining a first loss function value based on the first sample features and using a semantic relevance between the first image and text in the text data sample as a constraint; a parameter update module for performing parameter updates on the second feature extraction model according to the first loss function value; and and a model construction module for constructing a first feature extraction model based on the second feature extraction model in response to the second feature extraction model satisfying a convergence condition, wherein the first feature extraction model is used to process input first image and text data input to obtain data features of the first image and text data, and the data features of the first image and text data are processed by a task processing assembly, after which a processing result of a target task is output, and the target task is a classification or regression task based on image and text data.
[0010] According to another aspect, there is provided a computing device including a processor and a storage connected to the processor, the storage storing at least one computer program, the at least one computer program being loaded and executed by the processor to implement the image and text data processing method described above.
[0011] According to another aspect, a computer-readable storage medium is provided, the computer-readable storage medium having stored thereon at least one computer program, the computer program being loaded and executed by a processor to implement the image and text data processing method described above.
[0012] According to another aspect, a computer program product is provided, the computer program product including a computer program stored in a computer-readable storage medium, wherein a processor of a computing device reads the computer program from the computer-readable storage medium and executes the computer program, thereby causing the computing device to perform the image and text data processing method provided in the various selectable implementation manners described above. [Effects of the Invention]
[0013] A universal visual text representation space is constructed using at least one first image sample and at least two text samples in different languages corresponding to the first image sample as training data. When performing an image and text data processing task, the image and text data are mapped to the universal visual text representation space, and a task processing assembly outputs the task processing result based on the data features obtained by the mapping. The above technical solution, on the one hand, makes full use of weakly aligned multi-language image and text data samples, which have a large data volume and are relatively easy to obtain, because the text samples in at least two different languages corresponding to the first image sample are not translations of each other, thereby expanding the construction data of the universal visual text representation space and improving the accuracy (precision) of the universal visual text representation space. On the other hand, in the process of constructing the universal visual text representation space, the semantic relatedness between the image and the text in the text data sample is introduced as a constraint, allowing the constructed universal visual text representation space to more accurately extract the semantic features of the input data, thereby further improving the accuracy of the universal visual text representation space constructed by the first image sample and its corresponding text sample. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 illustrates a system used by the image and text data processing method provided in this application. [Figure 2] FIG. 1 illustrates a multi-language image and text with strict alignment according to the present application. [Figure 3] FIG. 1 illustrates weakly aligned multi-language images and text according to the present application. [Figure 4] 1 is a flowchart of an image and text data processing method presented in an exemplary embodiment of the present application. [Figure 5] 1 is a flowchart of an image and text data processing method presented in an exemplary embodiment of the present application. [Figure 6]1 is a flowchart of an image and text data processing method presented in an exemplary embodiment of the present application. [Figure 7] FIG. 10 is a diagram showing a network configuration of a second feature extraction model according to the present application. [Figure 8] FIG. 1 illustrates a smooth linear interpolation method according to the present application. [Figure 9] FIG. 1 illustrates constrained cross-language visual text contrastive learning according to the present application. [Figure 10] FIG. 10 is a diagram showing a network configuration of a second feature extraction model according to the present application. [Figure 11] 1 is a block diagram of an image and text data processing device provided in an embodiment of the present application; [Figure 12] 1 is a block diagram of an image and text data processing device provided in an embodiment of the present application; [Figure 13] FIG. 1 is a block diagram of a computer device according to an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION
[0015] In the embodiments of the present application, a data processing method for images and text is provided. For ease of understanding, some nouns (terms) related to the present application will be explained below.
[0016] 1) Strictly-aligned multi-language images and text The so-called strict alignment of multilingual images and texts refers to a high degree of association (correlation) between the semantics of the image content and its corresponding explanatory text, and the explanatory texts in multiple languages are corresponding translations of each other. The strict alignment of multilingual images and texts can also be called semantic parallelism, that is, the explanatory texts in multiple languages have the same semantics.
[0017] 2) Weakly-aligned multi-language images and text The so-called weak alignment of multilingual images and texts refers to the fact that the semantics of the image content and its corresponding explanatory texts are highly related, but the explanatory texts in multiple languages do not need to be translations of each other; the weak alignment of multilingual images and texts can also be referred to as semantically related but not parallel, that is, the explanatory texts in multiple languages are all related to the same image, but the semantics between the explanatory texts in multiple languages are not the same / not completely the same (only partly the same).
[0018] 1 is a diagram illustrating a system used by the image and text data processing method provided in the exemplary embodiment of the present application. As shown in FIG. 1, the system includes a server 110 and a terminal 120.
[0019] The above-mentioned server 110 may include a server on which an image and text data processing system is deployed and which provides image and text data processing services to users through the image and text data processing system, or the above-mentioned server 110 may include a server which has an image and text data processing system and which trains or updates the image and text data processing system.
[0020] The terminal 120 may include a user terminal that receives image and text data processing services, or the terminal 120 may include a development terminal used by a developer of an image and text data processing system.
[0021] Optionally, an image and text data processing system may be deployed in the terminal 120 described above.
[0022] Optionally, the above-mentioned system includes one or more servers 110 and multiple terminals 120. The number of servers 110 and terminals 120 is not limited in the embodiment of the present application.
[0023] The terminal and the server are connected via a communication network, which is optionally a wired network or a wireless network.
[0024] Because image- and text-based cross-modal and cross-lingual models have obvious advantages in processing multimodal tasks, pre-training of cross-modal and cross-lingual models has recently attracted increasing attention. Typically, developers primarily use image-text pairs with strict alignment of multilingual images and text for training. For example, developers can extend the English-based explanatory text in an image-text pair dataset to a multilingual version through translation, and then design a series of cross-modal and cross-lingual pre-training tasks to help the model learn a better universal representation.
[0025] For example, refer to Figure 2, which is a diagram illustrating the strictly aligned multi-language images and texts according to the present application. As shown in Figure 2, the strictly aligned multi-language images and texts include one image 21, one English text 22, and one Chinese text 23, where the image 21 and the English text 22 may be images and texts pre-collected by a developer, and the Chinese text 23 may be texts obtained by a developer by translating the English text 22 using a translation tool.
[0026] Subsequent embodiments of the present application provide an improved cross-modal and cross-lingual pre-training framework that can effectively utilize the more widely available and more easily collected large amounts of weakly aligned multi-lingual image and text multi-modal data, which may be obtained by collecting from a network.
[0027] For example, refer to Figure 3, which is a diagram showing weakly aligned multi-language images and texts according to the present application. As shown in Figure 3, the weakly aligned multi-language images and texts include one image 31, one English text 32, and one Chinese text 33, among which the English text 32 and the Chinese text 33 are explanatory texts in different languages for the image 31, which are obtained when a developer searches for the same image 31 on the network using an automatic search tool.
[0028] 4 is a flowchart of an image and text data processing method according to an exemplary embodiment of the present application. The method is performed by a computer device, which may be implemented as a terminal or a server, and the terminal or server may be the terminal or server shown in FIG. 1. As shown in FIG. 4, the image and text data processing method includes the following steps:
[0029] Step 410: Obtain first image and text data, where the first image and text data includes at least one image and at least one text.
[0030] In an embodiment of the present application, the above-mentioned first image and text data may include at least one image and text pair, and each image and text pair is an image-text pair consisting of one image and one line of text.
[0031] Step 420: Perform feature extraction on the first image and text data to map the first image and text data into a universal visual-text representation space to obtain data features of the first image and text data, where the universal visual-text representation space is a feature space constructed based on the first image and text data samples with the semantic relevance between the text in the first image and text data samples as a constraint, where the first image and text data samples include at least one first image sample and text samples in at least two different languages corresponding to the first image sample, and the text samples in at least two different languages corresponding to the first image sample are not translations of each other.
[0032] In the embodiments of the present application, the feature extraction process for the above-mentioned first image and text data is a process of mapping the first image and text data into a universal visual text representation space and obtaining data features of the first image and text data, or it can be said that the step of mapping the first image and text data into a universal visual text representation space is realized by performing feature extraction on the first image and text data.
[0033] In machine learning, after one or more feature mappings are performed on the input original data, a higher-dimensional abstract representation can be obtained, which can be called a feature in the concept of machine learning of the original data, and the space formed by the features obtained after one or more feature mappings are performed on all possible input data is a feature space. In other words, the features in the feature space are higher-dimensional representations of all possible input data.
[0034] The universal visual text representation space described above is a feature space for unifying the representation of two types of data: images and text.
[0035] In the embodiment of the present application, the two types of data, image and text, can be merged into a unified feature space (i.e., the universal visual-text representation space mentioned above), whereby the data features obtained by mapping the first image and text data into the universal visual-text representation space can be expressed in the form of a feature vector, a feature matrix, etc.
[0036] Among them, constructing a universal visual text representation space based on the above-mentioned first image and text data sample with the semantic relevance between the first image and the text in the text data sample as a constraint may refer to constructing a universal visual text representation space based on the first image and the text data sample with the goal of shortening the semantic distance (or increasing the semantic relevance) between the first image and the text in the text data sample when constructing a universal visual text representation space based on the first image and the text data sample.
[0037] The above-mentioned first image sample corresponding to text samples of at least two different languages may mean that one first image sample corresponds to at least two text samples, and the at least two text samples belong to different languages (e.g., Chinese, English, French, etc.), and the semantics of the at least two text samples are all related to the first image sample.
[0038] Furthermore, the above-mentioned statement that the text samples in at least two different languages corresponding to the first image sample are not translations of each other may also refer to the semantic features of the text samples in at least two different languages corresponding to the first image sample being different. For example, after translating the text samples in at least two different languages corresponding to the first image sample into the same language, semantic extraction is performed on each of them to obtain semantic feature vectors after translation of the text samples in at least two different languages, and then the similarity between these semantic feature vectors is calculated. If the similarity between any two of these semantic feature vectors is equal to or less than a certain similarity threshold (i.e., not greater than the certain similarity threshold), the text samples in at least two different languages corresponding to the first image sample can be considered not to be translations of each other. Also, for example, after translating the text samples in at least two different languages corresponding to the first image sample into the same language, keyword extraction is performed on each of them to obtain translated keywords of the text samples in at least two different languages, and if the translated keywords of the text samples in at least two different languages are different, the text samples in at least two different languages corresponding to the first image sample can be considered not to be translations of each other.
[0039] For example, suppose the image content of a first image sample is "There is a house at the foot of the mountain and there are two puppies in front of the house", and the first image sample has two text samples, one of which is a Chinese text sample called "There is a house at the foot of the mountain" and the other is an English text sample called "There are two puppies in front of the house", the semantics of these two text samples are all related to the first image sample, but the semantic features / keywords extracted after translating these two text samples into the same language are different, that is, the two text samples are not translations of each other.
[0040] Wherein, the first image and text data samples may be weakly aligned multi-language image and text data samples.
[0041] Step 430: Send the data features of the first image and text data to a task processing assembly, and the task processing assembly outputs a processing result of a target task based on the data features, where the target task is a classification or regression task based on the image and text data.
[0042] The above-mentioned task processing assembly may be a software module (machine learning model) configured on a current computing device, and at this time, the computing device may input the data features obtained by the above-mentioned mapping into the task processing assembly.
[0043] Alternatively, the above-mentioned task processing assembly may be a software module set up in a computer device other than the current computer device, and in this case, the computer device may transmit the above-mentioned data characteristics to the other computer device via a wired / wireless network, and the other computer device may input the data characteristics into the task processing assembly.
[0044] In an embodiment of the present application, the data features obtained by mapping the first image and text data into the universal visual-text representation space in step 420 above can be used in any subsequent classification or regression task implemented based on the image and text.
[0045] Among them, the classification task mentioned above refers to a task that outputs one classification probability after processing the above data features, and the regression task mentioned above refers to a task that outputs one image / text / image and text pair or another data feature after processing the above data features.
[0046] For example, the data features obtained by mapping the above-mentioned first image and text data into the universal visual-text representation space are processed by the task processing assembly to output a classification probability (e.g., the probability that the image and the text match, the probability that the image belongs to a certain category, etc.), or a regression result (e.g., the output of a reconstructed image, or the output of a reconstructed / translated text). The embodiments of the present application are not limited to the classification task or regression task realized based on the above-mentioned images and text.
[0047] In summary, the technical solution disclosed in the embodiments of the present application uses at least one first image sample and at least two text samples in different languages corresponding to the first image sample as training data to construct a universal visual text representation space. When performing an image and text data processing task, the image and text data are mapped to the universal visual text representation space, and a task processing assembly outputs the task processing result based on the data features obtained through the mapping. The above technical solution, on the one hand, makes full use of weakly aligned multi-language image and text data samples, which have a large data volume and are relatively easy to obtain, because the text samples in at least two different languages corresponding to the first image sample are not translations of each other, thereby expanding the construction data of the universal visual text representation space and improving the accuracy of the universal visual text representation space. On the other hand, in the process of constructing the universal visual text representation space, the semantic relatedness between the image and the text in the text data sample is introduced as a constraint, allowing the constructed universal visual text representation space to more accurately extract the semantic features of the input data, thereby further improving the accuracy of the universal visual text representation space constructed by the first image sample and its corresponding text sample.
[0048] In the embodiment shown in Figure 2, the universal visual-text representation space can be represented by a machine learning model, which is trained in advance using training data consisting of image and text pairs, and then processes the subsequently input image and text data to obtain data features in the universal visual-text representation space of the image and text data.
[0049] Referring to the embodiment shown in Fig. 4, reference is now made to Fig. 5, which is a flowchart of an image and text data processing method according to an exemplary embodiment of the present application. The method is performed by a computer device, which may be implemented as a terminal or a server, and the terminal or server may be the terminal or server shown in Fig. 1. As shown in Fig. 5, the process of training and applying a machine learning model for image and text data processing may include the following steps:
[0050] Step 510: Construct a first anchor sample, a first positive example sample, and a first negative example sample based on the first image and text data sample.
[0051] Wherein, the first image and text data samples include at least one first image sample and at least two text samples in different languages corresponding to the first image sample, and the text samples in at least two different languages corresponding to the first image sample are not translations of each other.
[0052] In an embodiment of the present application, a developer can pre-collect weakly aligned image and text data samples (i.e., the first image and text data samples described above), and then construct samples required for contrastive learning based on the first image and text data samples, where such samples include a first anchor sample as an anchor, a first positive example sample as a positive sample, and a first negative example sample as a negative sample.
[0053] The first positive sample refers to a sample that matches the first anchor sample feature, and the first negative sample refers to a sample that does not match the first anchor sample feature.
[0054] For example, when the first anchor sample includes an image and text pair, the first positive example sample and the first negative example sample may also each include one image and text pair, and the relationship between the image and text in the first positive example sample is close to the relationship between the image and text in the first anchor sample, and the relationship between the image and text in the first negative example sample is not close to the relationship between the image and text in the first anchor sample.
[0055] For example, the first anchor sample, the first positive example sample, and the first negative example sample each contain the same image, and the first anchor sample, the first positive example sample, and the first negative example sample each contain different text, and although the text in the first anchor sample and the first positive example sample is different, the semantics of the two texts are all close to the semantics of the image, and the semantics of the text in the first negative example sample is not close to or related to the semantics of the image.
[0056] Also, for example, when the first anchor sample contains one piece of text, the first positive example sample and the first negative example sample may also each contain one piece of text, and the semantics of the text in the first positive example sample and the text in the first anchor sample may be close, but the semantics of the text in the first negative example sample and the text in the first anchor sample may not be close.
[0057] Step 520: Input the first anchor sample, the first positive example sample, and the first negative example sample into a second feature extraction model to obtain a first sample feature output by the second feature extraction model.
[0058] In an embodiment of the present application, the computer device may input the first anchor sample, the first positive example sample, and the first negative example sample into the second feature extraction model, respectively, process the first anchor sample, the first positive example sample, and the first negative example sample using the second feature model, respectively, and output sample features of the first anchor sample, the first positive example sample, and the first negative example sample, respectively, and the sample features of the first anchor sample, the first positive example sample, and the first negative example sample may constitute the above-mentioned first sample feature.
[0059] Step 530: Obtain a first loss function value based on the first sample feature and with the semantic relevance between the first image and the text in the text data sample as a constraint.
[0060] In an embodiment of the present application, the computing device may take into account the semantic relevance between the first image and the text in the text data sample when generating the first loss function value based on the first sample features, that is, the computing device may generate the above-mentioned first loss function value in combination with two pieces of information: the first sample features and the semantic relevance information between the first image and the text in the text data sample.
[0061] The semantic relevance information between the first image and the text in the text data sample may be obtained by calculation based on the first sample features. For example, a computer device may obtain the semantic relevance information between the text in the first anchor sample, the first positive example sample, and the first negative example sample by performing processing based on the features related to the text in the first anchor sample, the first positive example sample, and the first negative example sample among the first sample features.
[0062] The semantic relevance mentioned above may refer to the semantic similarity between two or more texts.
[0063] Wherein, the process of the computer device obtaining a first loss function value based on the first sample features and using the semantic relevance between the first image and the text in the text data sample as a constraint may include:
[0064] The computer device generates a part of the loss function value based on the first sample features through contrastive learning, then generates another part of the loss function value based on the semantic relevance information, and then adds the two parts of the loss function value or calculates a weighted sum of them to obtain the first loss function value.
[0065] Step 540: Update the parameters of the second feature extraction model according to the first loss function value.
[0066] In the embodiment of the present application, the calculation process of the first loss function value takes into account semantic relevance information between the text in the first anchor sample, the first positive example sample, and the first negative example sample. This can support training using weakly aligned image and text data samples, so that the available training data for the second feature extraction model can be expanded to include weakly aligned image and text data samples, thereby achieving the effect of data training for the expanded model. In addition, more training data is also beneficial to improving the accuracy of model training.
[0067] Step 550: If the second feature extraction model satisfies the convergence condition, construct a first feature extraction model based on the second feature extraction model.
[0068] Among them, the first feature extraction model is used to process input first image and text data to obtain data features of the first image and text data, and the data features of the first image and text data are processed by a task processing assembly, and then a processing result of a target task is output, where the target task is a classification or regression task based on the image and text data.
[0069] In other words, the above-mentioned first feature extraction model is a machine learning network constructed based on the second feature extraction model, and the second feature extraction model is a machine learning model obtained by performing machine learning-based training based on the first image and the text data sample, with the semantic relevance between the first image and the text in the text data sample as a constraint.
[0070] In one possible implementation, the second feature extraction model described above can be directly used as the first feature extraction model deployed in an application for feature extraction of image and text data.
[0071] In another possible implementation, a developer may improve the above-mentioned second feature extraction model using a computer device and directly use the improved second feature extraction model as the first feature extraction model deployed as an application for feature extraction of image and text data.
[0072] For example, the developer may use a computer device to optimize the second feature extraction model to obtain the first feature extraction model. For example, the developer may use a computer device to retrain the second feature extraction model to make the second feature extraction model more suitable for a particular image and text data processing task. For example, the developer may use a computer device to simplify (e.g., prune, distill, etc.) the second feature extraction model to obtain a lightweight first feature extraction model to improve the processing speed of the model.
[0073] Step 560: Obtain first image and text data, where the first image and text data includes at least one image and at least one text.
[0074] After the first feature extraction model is deployed as an application, the computing device on which the first feature extraction model is deployed can execute an image and text data processing task, and at this time, the computing device can obtain first image and text data to be processed.
[0075] In one possible implementation, the first image and text data may include a number of image and text pairs.
[0076] Step 570: Input the first image and text data into a first feature extraction model, and obtain data features of the first image and text data output by the first feature extraction model.
[0077] Among them, the first feature extraction model mentioned above is used to map the features of the input image and text data into a universal visual-text representation space.
[0078] In one possible implementation, a computer device inputs a number of image and text pairs in the first image and text data into the above-mentioned first feature extraction model, respectively, so that the first feature extraction model processes the image and text pairs and outputs data features of the first image and text data.
[0079] Step 580: Send the data features of the first image and text data to a task processing assembly, and the task processing assembly outputs a processing result of a target task based on the data features, where the target task is a classification or regression task based on the image and text data.
[0080] Among them, the above embodiment relates to the training process of the second feature extraction model (steps 510 to 540), the construction process of the first feature extraction model (step 550), and the application process of the first feature extraction model (steps 560 to 580), and the above three processes may be performed by different computer devices, or two of the above three processes may be performed by one computer device and the other process may be performed by another computer device, or the above three processes may be performed by the same computer device.
[0081] In an embodiment of the present application, a computer device uses a pre-trained first feature extraction model to map input first image and text data into a universal visual-text representation space, thereby improving the accuracy of feature mapping and improving the accuracy of the processing results of subsequent target tasks.
[0082] In an embodiment of the present application, the computer device can train the second feature extraction model in a contrastive learning manner based on the first anchor samples, the first positive example samples, and the first negative example samples constructed based on the first image and text data samples.
[0083] In one possible implementation manner, the above-mentioned first anchor sample includes a first anchor image and text sample, and a first anchor text sample, the first positive example sample includes at least one first positive example image and text sample, and at least one first positive example text sample, and the first negative example sample includes at least one first negative example image and text sample, and at least one first negative example text sample.
[0084] The first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample each include one image and text pair, and the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample include the same first image, the text in the first positive example image and text sample and the text in the first anchor image and text sample each match the first image semantic, and the text in the first negative example image and text sample does not match the first image semantic, the first anchor text sample includes the text in the first anchor image and text sample, the first positive example text sample includes the text in the first positive example image and text sample, and the first negative example text sample includes the text in the first negative example image and text sample.
[0085] Based on the embodiment shown in Figure 5, reference is now made to Figure 6, which is a flowchart of an image and text data processing method according to an exemplary embodiment of the present application. As shown in Figure 6, step 520 in the embodiment shown in Figure 5 above may be implemented as step 520a and step 520b, and step 530 may be implemented as step 530a to step 530d.
[0086] Step 520a: The first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample are input into a second feature extraction model to obtain sample features of the first anchor image and text sample, sample features of the first positive example image and text sample, and sample features of the first negative example image and text sample.
[0087] In an embodiment of the present application, the computer device for training the second feature extraction model may input the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample into the second feature extraction model, respectively, and output sample features of the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample via the second feature extraction model.
[0088] In one possible implementation, the second feature extraction model may include at least one feature extraction network. When the second feature extraction model includes at least two feature extraction networks, the network configurations (structures) and parameters of the at least two feature extraction networks may be the same or different. For example, all or some of the at least two feature extraction networks may be connected sequentially, in a residual manner, in a crossover manner, or in a cyclic manner. The embodiments of the present application do not limit the manner of connection between the feature extraction networks.
[0089] Illustratively, refer to Figure 7, which is a diagram illustrating a network configuration of a second feature extraction model according to an embodiment of the present application. As shown in Figure 7, the second feature extraction model can be divided according to function into a text feature extraction branch 701, an image feature extraction branch 702, and a feature fusion branch 703. After a computer inputs an image and text pair into the second feature extraction model, the image (or image sample) in the image and text pair (or image and text sample pair) is input into the image feature extraction branch 702 to obtain image features (or image sample features), the text (or text sample) in the image and text pair (or image and text sample pair) is input into the text feature extraction branch 701 to obtain text features (or text sample features), and the image features (or image sample features) and text features (or text sample features) are input into the feature fusion branch 703 to obtain data features (or data sample features) output by the model.
[0090] Although FIG. 7 above illustrates only one possible network configuration of the second feature extraction model in an exemplary manner, the second feature extraction model may optionally adopt other configurations, such as other connection configurations, other network branch division methods, etc.
[0091] Step 520b: Input the first anchor text sample, the first positive example text sample, and the first negative example text sample into a second feature extraction model to obtain sample features of the first anchor text sample, sample features of the first positive example text sample, and sample features of the first negative example text sample.
[0092] In an embodiment of the present application, the above-mentioned second feature extraction model may include a text feature extraction network (e.g., the text feature extraction branch 701 in the above-mentioned FIG. 7 ), and the computer device may input the above-mentioned first anchor text sample, the first positive example text sample, and the first negative example text sample into the above-mentioned text feature extraction network, respectively, to obtain sample features of the above-mentioned first anchor text sample, the first positive example text sample, and the first negative example text sample, respectively.
[0093] Step 530a: Obtain semantic relevance information based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, the sample features of the first negative example image and text sample, the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample, and the semantic relevance information is used to indicate the semantic relevance between the first image and the text in the text data sample.
[0094] In the above embodiment, the sample features of each of the first anchor text sample, the first positive example text sample, and the first negative example text sample are text-related features, and the sample features of each of the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample also include some text-related features, in which the computer device can obtain the semantic relevance between the first image and the text in the text data sample by performing calculations together with the sample features of the image and text sample, and the sample features of the text sample.
[0095] Step 530b: Generate a semantic relevance loss function value based on the semantic relevance information.
[0096] In one possible implementation, obtaining semantic relevance information based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, the sample features of the first negative example image and text sample, the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample includes: performing a normalization process on the sample features of the first anchor image and the text sample, the sample features of the first positive example image and the text sample, and the sample features of the first negative example image and the text sample to obtain first relevance distribution information in the semantic relevance information; and The method includes performing a normalization process on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample to obtain second relevance distribution information in the semantic relevance information.
[0097] In an embodiment of the present application, the computer device can calculate semantic relevance distribution information for each data type through a normalization process according to the data type (image and text data or plain text data) corresponding to the features. Specifically, the computer device can calculate semantic relevance distribution information for the image and text data based on the sample features of the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample. The computer device can further calculate semantic relevance distribution information for the plain text data based on the sample features of the first anchor text sample, the first positive example text sample, and the first negative example text sample. These two pieces of relevance distribution information constitute the above-mentioned semantic relevance information.
[0098] Generating a semantic relevance loss function value based on the above-mentioned semantic relevance information includes calculating KL divergence for the first relevance distribution information and the second relevance distribution information to obtain a semantic relevance loss function value.
[0099] In an embodiment of the present application, the computer device calculates the KL divergence of the first relevance distribution information and the second relevance distribution information, and the fusion of the first relevance distribution information and the second relevance distribution information is referred to as a semantic relevance loss function value, which can be used to represent the semantic relevance between the first image and the text in the text data sample, thereby realizing the construction of a loss function value for semantic relevance and realizing the expansion of available training data for the second feature extraction model to weakly aligned image and text data samples.
[0100] Alternatively, the computer device may obtain the semantic relevance information in other ways. For example, the computer device may first fuse the sample features of the first anchor image and text sample with the sample features of the first anchor text sample. For example, the computer device may connect the sample features of the first anchor image and text sample with the sample features of the first anchor text sample, and then perform a fusion process using a fusion network (for example, the fusion network may include network layers such as a fully connected layer, a convolutional layer, and a pooling layer) to obtain fused anchor sample features. Similarly, the computer device may fuse the sample features of the first positive example image and text sample with the sample features of the first positive example text sample to obtain fused positive example sample features, or the computer device may fuse the sample features of the first negative example image and text sample with the sample features of the first negative example text sample to obtain fused negative example sample features. Then, the computer device further calculates the above-mentioned semantic relevance information based on the fused anchor sample features, the fused positive example sample features, and the fused negative example sample features. At this time, the computer device calculates the semantic relevance loss function value according to the above-mentioned semantic relevance information.
[0101] Step 530c: Calculate a first contrastive learning loss function value based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample.
[0102] In an embodiment of the present application, in addition to the semantic relevance loss function value, the computer device may further use the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample to generate a first contrastive learning loss function value in a contrastive learning manner, thereby improving the accuracy of subsequent model training.
[0103] Step 530d: Obtain a first loss function value based on the semantic relevance loss function value and the first contrastive learning loss function value.
[0104] In an embodiment of the present application, the computer device may obtain the first loss function value by performing calculations on the semantic relevance loss function value and the first control learning loss function value together. For example, the computer device may obtain the first loss function value by adding the semantic relevance loss function value and the first control learning loss function value or calculating a weighted sum of the two, and the weights in the weighted sum may be preset by a developer.
[0105] In the technical solution provided in the embodiments of the present application, a computer device calculates a semantic relevance loss function value according to the features of the input text samples, and then obtains a first loss function value in combination with the loss function value of the contrastive learning. The first loss function value is then used to update the parameters of the second feature extraction model, and the semantic relevance between the text in the first image and text data samples is used as a constraint in the model training process. This allows the available training data for the second feature extraction model to be expanded to include weakly aligned image and text data samples, thereby achieving the effect of expanded model data training. More training data also leads to improved model training accuracy.
[0106] In one possible implementation, the above method further includes:
[0107] generating a second loss function value based on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample; and The parameters of the second feature extraction model are updated using the second loss function value.
[0108] In an embodiment of the present application, the computer device can further independently perform parameter updates for the second feature extraction model according to the sample features of the text sample, thereby expanding the training data for the second feature extraction model and improving the extraction effect of the second feature extraction model.
[0109] For example, taking FIG. 7 as an example again, the computer device inputs the first anchor text sample, the first positive example text sample, and the first negative example text sample into the text feature extraction branch 701, and then obtains the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample. Then, in addition to calculating the above-mentioned first loss function value, the computer device calculates a second loss function value using contrastive learning based solely on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample. Then, the computer device updates the parameters of the second feature extraction model according to the second loss function value, for example, updates the parameters of the text feature extraction branch in the second feature extraction model according to the second loss function value.
[0110] In one possible implementation, the above method further includes: constructing a second anchor image and text sample, a second positive example image and text sample, and a second negative example image and text sample based on the second image and text data sample, wherein the second image and text data sample includes at least one second image sample and a text sample in at least two different languages corresponding to the second image sample, and the text samples in at least two different languages corresponding to the second image sample are translations of each other, the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample each include one image and text pair, and the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample include the same second image, the text in the second positive example image and text sample and the text in the second anchor image and text sample respectively match the second image semantics, and the text in the second negative example image and text sample does not match the second image semantics; inputting the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample into a second feature extraction model to obtain sample features of the second anchor image and text sample, sample features of the second positive example image and text sample, and sample features of the second negative example image and text sample; and generating a third loss function value based on the sample features of the second anchor image and text sample, the sample features of the second positive example image and text sample, and the sample features of the second negative example image and text sample; and The parameters of the second feature extraction model are updated using the third loss function value.
[0111] In an embodiment of the present application, in addition to training the second feature model using the weakly aligned first image and text data samples, the computer device may further train the second feature model using the tightly aligned second image and text data samples, wherein the process of training the second feature model using the second image and text data samples reduces the process of calculating the semantic relevance loss function value compared to the process of training the second feature model using the weakly aligned first image and text data samples, wherein the process of calculating the above-mentioned third loss function value is the same as the process of calculating the first control learning loss function value, and a detailed description thereof will be omitted here.
[0112] In the embodiment of the present application, a Regularized Cross-lingual Visio-textual Contrastive Learning (R-XVtCL) task is adopted, and a dataset of weakly aligned multi-lingual image and text pairs (i.e., the first image and text data sample mentioned above, D w Optionally, the task may be performed simultaneously on a multilingual strictly aligned image-text pair dataset (i.e., the first image and text data samples mentioned above, denoted as D s It may also be related to the
[0113] Optionally, embodiments of the present application may further train a second machine learning model with one Cross-lingual Textual Contrastive Learning (XTCL) task.
[0114] The goal of the XTCL task is to obtain meaningful textual representations (TRs) for multilingual texts in a universal textual representation space (UTRS). In this UTRS space, representations of texts with relatively high semantic relatedness should be close to each other, and conversely, representations of texts with low relatedness should be far from each other. In this application, we perform contrastive training in the UTRS space to enable the model to generate representations that satisfy the above properties.
[0115] Specifically, one multilingual text dataset D from which large-scale parallel word pairs are derived t Then, select one batch from the given set.
[0116]
number
[0117]
number
[0118]
number
[0119]
number
[0120]
number
[0121]
number
[0122] Among them,
[0123]
number
[0124]
number
[0125] In the constrained cross-language visual text contrastive learning task in the pre-training stage, in addition to using Euclidean distance to measure the similarity between positive and negative examples and anchors, other measures of text relevance may be adopted, such as cosine similarity between expression feature vectors, i.e., the above-mentioned Euclidean distance can be replaced by cosine similarity.
[0126] The negative example representations obtained by the above method may not be difficult enough for the model to perform contrastive learning-based training, and may not allow the model to learn more informative effective text representations in the UTRS space. In this case, the technical solution described in the embodiments of this application may further adopt a negative example representation generation method based on smooth linear interpolation to generate more difficult and less distinguishable negative examples, such as:
[0127]
number
[0128]
number
[0129]
number
[0130]
number
[0131] Among them, ζ (Zeta) is a preset scaling coefficient,
[0132]
number
[0133]
number
[0134]
number
[0135] The goal of the R-XVtCL task is to enable the model to learn meaningful representations for multilingual visual text inputs in a Universal Visio-textual Representation Space (UVtRS) through contrastive learning.
[0136] To deepen our understanding, we first consider contrastive learning using strictly aligned multilingual image-text pair data. s A batch of bilingual parallel visual text triples from
[0137]
number
[0138]
number
[0139]
number
[0140]
number
[0141]
number
[0142]
number
[0143]
number
[0144]
number
[0145]
number
[0146]
number
[0147]
number
[0148]
number
[0149]
number
[0150]
number
[0151]
number
[0152] Among them, N(v′,x′) is a set containing the three types of negative examples constructed as described above,
[0153]
number
[0154]
number
[0155]
number
[0156] However, in the process of contrastive learning using the "Strictly-aligned" image-text data mentioned above, + A, VTR * However, deviations may occur in the case of data that only satisfies the "weakly-aligned" relationship, meaning that the semantics of explanatory texts in different languages corresponding to the same image may be somewhat related but not equivalent. In such cases, simply using vtr + and VTR *It is unreasonable to directly approximate the semantic similarity between multilingual texts. To achieve this, it is necessary to be able to measure the semantic similarity between multilingual texts so as to obtain a better visual text representation by more effectively capturing the semantic relationship between them in the UVtRS space. Therefore, in the case of "weakly-aligned" multilingual image-text data, this application uses the text semantic relatedness between multilingual explanatory texts to impose additional constraints on the above-mentioned contrastive learning process. Regarding the semantic relatedness of any two texts, this application measures it by the Euclidean distance between the text representations in the UTRS space learned in the XVTCL task. That is, the smaller the corresponding distance, the more semantically related the texts are.
[0157] Specifically, D w Two multilingual visual text triples in one patch of bilingual parallel visual text triples from
[0158]
number
[0159]
number
[0160]
number
[0161]
number
[0162]
number
[0163]
number
[0164]
number
[0165]
number
[0166]
number
[0167]
number
[0168]
number
[0169]
number
[0170]
number
[0171]
number
[0172]
number
[0173]
number
[0174]
number
[0175] Also,
[0176]
number
[0177]
number
[0178]
number
[0179]
number
[0180]
number
[0181]
number
[0182]
number
[0183] Then, we use KL divergence to bring the two distributions closer together and train them together with the usual contrastive learning objective, i.e., constrained cross-lingual visual-text contrastive learning.
[0184]
number
[0185] Finally, as shown in Figure 9, we consider the constrained cross-lingual visual-text contrastive learning. s and D w By synthesizing the training examples from above, the training goal of R-XVtCL can be written in the following format:
[0186]
number
[0187] Illustratively, refer to Fig. 10, which is a diagram illustrating a network configuration of a second feature extraction model according to an embodiment of the present application. In the process of training the second feature model based on weakly aligned image and text data samples, as shown in Fig. 10, the second feature extraction model can be divided according to function into a text feature extraction branch 1001, an image feature extraction branch 1002, and a feature fusion branch 1003. A computer device pre-collects image and text data samples 1004 including weakly aligned first image and text data samples. Constructing input samples based on the image and text data samples 1004 includes constructing first anchor image and text samples, first anchor text samples, first positive example image and text samples, first positive example text samples, first negative example image and text samples, and first negative example text samples based on the first image and text data samples.
[0188] Then, the computer device inputs the text in the constructed input sample into a text feature extraction branch 1001 to obtain sample features 100 of the input text (including sample features of the first anchor text sample, the first positive example text sample, and the first negative example text sample), inputs the image in the constructed input sample into an image feature extraction branch 1002 to obtain sample features 1006 of the input image, and the sample features 1005 and 1006 are input into a feature fusion branch 1003 to obtain sample features 1007 (including sample features of the first anchor image and text sample, the first positive example image and text sample, the first positive example text sample, and the first negative example image and text sample).
[0189] Thereafter, on the one hand, the computing device generates first relevance distribution information 1008 according to the sample features 1007, generates second relevance distribution information 1009 according to the sample features 1006, and generates a relevance loss function value 1010 based on the first relevance distribution information 1008 and the second relevance distribution information 1009, and on the other hand, the computing device generates a first contrast learning loss function value 1011 according to the sample features 1007, and generates a first loss function value 1012 according to the relevance loss function value 1010 and the first contrast learning loss function value 1011. Finally, the computing device updates parameters for the second feature extraction model according to the first loss function value 1012.
[0190] In one possible implementation, the method further includes: constructing a first image and text sample and a translation annotation of the first image and text sample based on the third image and text data sample, wherein the third image and text data sample includes at least one third image sample and a text sample in at least two different languages corresponding to the third image sample, and the text samples in at least two different languages corresponding to the second image sample are translations of each other, the first image and text sample includes the third image sample and a text sample corresponding to the first image sample, and the translation annotation includes another text sample corresponding to the third image sample; masking the text sample portion of the first image and the text sample, and then inputting the masked text sample portion into a second feature extraction model; and obtaining first sample features and second sample features of the first image and the text sample output by the second feature extraction model; inputting the first sample features into a first output network, obtaining a first prediction result output by the first output network, the first prediction result being used to indicate predicted text for a portion of the first image and the text sample where the text sample is masked; inputting the second sample features into a second output network to obtain a second prediction result output by the second output network, the second prediction result being used to indicate the predicted text of the translation annotation; generating a fourth loss function value based on the first prediction result and the first image and the portion of the text sample where the text sample is masked; generating a fifth loss function value based on the second prediction result and the translation annotation; and Parameter update is performed on the second feature extraction model based on the fourth loss function value and the fifth loss function value.
[0191] In an embodiment of the present application, the computer device performs parameter updates for the second feature extraction model in conjunction with the image translation task, thereby expanding the training data and training method of the second feature extraction model and improving the extraction effect of the second feature extraction model.
[0192] In an embodiment of the present application, the above-mentioned second feature extraction model may be further trained by training the task with a Masked Conditional Language Model (MCLM).
[0193] MCLM includes a standard Masked Language Modeling (MLM) on the source language encoder side and a Conditional Language Modeling (CLM) on the target language decoder side. Given an image v and a language l, i and language j Image description text belonging to each
[0194]
number
[0195]
number
[0196]
number
[0197]
number
[0198]
number
[0199]
number
[0200]
number
[0201] Among them, D s represents a dataset of image-text pairs with strong alignment of multilingual images and text, and θ e represents the trainable parameters of the encoder of the model. On the decoder side, the present application also proposes that the model decodes in an autoregressive manner by training the model with CLM targets based on the visual text input on the encoder side.
[0202]
number
[0203]
number
[0204] Among them, θ d represents the trainable parameters of the decoder of the model. The total MCLM goal, combining the goals of both the encoder and decoder, is L MCLM =L mlm +L clm It can be written as:
[0205] In one possible implementation, the method further comprises: obtain a third positive example image and text sample and a third negative example image and text sample, the third positive example image and text sample including image and text pairs with matching semantics, and the third negative example image and text sample including image and text pairs with mismatching semantics; inputting the third positive example image and text sample and the third negative example image and text sample into a second feature extraction model, and obtaining sample features of the third positive example image and text sample and sample features of the third negative example image and text sample output by the second feature extraction model; inputting the sample features of the third positive example image and text sample and the sample features of the third negative example image and text sample into a third output network, and obtaining a semantic matching probability between the third positive example image and text sample and a semantic matching probability between the third negative example image and text sample output from the third output network; generating a sixth loss function value based on the semantic match probability between the third positive example image and the text sample and the semantic match probability between the third negative example image and the text sample; and performing a parameter update on the second feature extraction model based on the sixth loss function value.
[0206] In an embodiment of the present application, a second feature extraction model may be further trained by an Image-Text Matching (ITM) task.
[0207] In an embodiment of the present application, the computer device, in addition to the task of predicting the degree of correspondence between images and text, can also perform parameter updates for the second feature extraction model, thereby expanding the training data and training method of the second feature extraction model and improving the extraction effect of the second feature extraction model.
[0208] The goal of the ITM task is to find a given image input v and a text input x l The goal is to determine whether the image-text pairs are matched. The objective function can be written as follows:
[0209]
number
[0210] Among them, y∈{0,1} is the text x l indicates whether image v matches image v.
[0211]
number
[0212] In other words, the above-described embodiments of the present application provide a pre-training framework that can more effectively utilize large amounts of multilingual, weakly aligned image and text multimodal data. The entire training process includes four pre-training objectives: Masked Conditional Language Modeling (MCLM), Image-Text Matching (ITM), Cross-lingual Textual Contrastive Learning (XTCL), and Regularized Cross-lingual Visio-textual Contrastive Learning (R-XVtCL). In the above-described embodiments, the four pre-training tasks enable the model to learn better visual-text representations within a unified cross-modal and cross-lingual vector representation space, thereby improving the effectiveness of a series of downstream visual-text tasks.
[0213] Figure 11 is a block diagram of an image and text data processing apparatus shown in an exemplary embodiment of the present application. The apparatus may be used to perform all or part of the steps in the method shown in Figure 4, Figure 5 or Figure 6. As shown in Figure 11, the apparatus includes: A data acquisition module 1101: used to acquire first image and text data, where the first image and text data includes at least one image and at least one text; a feature mapping module 1102: performing feature extraction on the first image and text data to map the first image and text data into a universal visual-text representation space to obtain data features of the first image and text data, the universal visual-text representation space being a feature space constructed based on first image and text data samples with constraints on semantic relatedness between text in the first image and text data samples, the first image and text data samples including at least one first image sample and text samples in at least two different languages corresponding to the first image sample, and the text samples in at least two different languages corresponding to the first image sample are not translations of each other; and Task processing module 1103: Used to send data features of the first image and text data to a task processing assembly, and have the task processing assembly output a processing result of a target task based on the data features, where the target task is a classification or regression task based on image and text data.
[0214] In one possible implementation, the feature mapping module 1102 is used to: inputting the first image and text data into a first feature extraction model, and obtaining data features of the first image and text data output by the first feature extraction model; The first feature extraction model is a machine learning network constructed based on the second feature extraction model, and the second feature extraction model is a machine learning model obtained by training using machine learning based on a first image and a text data sample, with the semantic relevance between the first image and the text in the text data sample as a constraint.
[0215] In one possible implementation, the apparatus further includes a model training module, which is used to: constructing a first anchor sample, a first positive example sample, and a first negative example sample based on the first image and text data sample; inputting the first anchor sample, the first positive example sample, and the first negative example sample into the second feature extraction model to obtain first sample features output by the second feature extraction model; obtaining a first loss function value based on the first sample features and using a semantic relevance between the first image and text in the text data sample as a constraint; and The second feature extraction model is updated with parameters according to the first loss function value.
[0216] In one possible implementation, the first anchor samples include a first anchor image and text sample and a first anchor text sample, the first positive example samples include at least one first positive example image and text sample and at least one first positive example text sample, and the first negative example samples include at least one first negative example image and text sample and at least one first negative example text sample; The first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample each include one image and text pair, and the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample include the same first image, the text in the first positive example image and text sample and the text in the first anchor image and text sample respectively match the first image semantics, the text in the first negative example image and text sample does not match the first image semantics, the first anchor text sample includes the text in the first anchor image and text sample, the first positive example text sample includes the text in the first positive example image and text sample, and the first negative example text sample includes the text in the first negative example image and text sample.
[0217] The model training module is used to do the following: inputting the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample into the second feature extraction model to obtain sample features of the first anchor image and text sample, sample features of the first positive example image and text sample, and sample features of the first negative example image and text sample; inputting the first anchor text sample, the first positive example text sample, and the first negative example text sample into the second feature extraction model to obtain sample features of the first anchor text sample, sample features of the first positive example text sample, and sample features of the first negative example text sample; obtaining semantic relevance information based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, the sample features of the first negative example image and text sample, the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample, and the semantic relevance information is used to indicate the semantic relevance between the first image and text in the text data sample; generating a semantic relevance loss function value based on the semantic relevance information; generating a first contrastive learning loss function value based on sample features of the first anchor image and text sample, sample features of the first positive example image and text sample, and sample features of the first negative example image and text sample; and The first loss function value is obtained based on the semantic relevance loss function value and the first control training loss function value.
[0218] In one possible implementation, the model training module is used to: performing a normalization process on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample to obtain first relevance distribution information in the semantic relevance information; performing a normalization process on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample to obtain second relevance distribution information in the semantic relevance information; and The KL divergence is calculated for the first relevance distribution information and the second relevance distribution information to obtain the word semantic relevance loss function value.
[0219] In one possible implementation, the model training module is used to: generating a second loss function value based on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample; and The second feature extraction model is updated with parameters according to the second loss function value.
[0220] In one possible implementation, the model training module is used to: constructing second anchor images and text samples, second positive example images and text samples, and second negative example images and text samples based on the second image and text data samples, wherein the second image and text data samples include at least one second image sample and text samples in at least two different languages corresponding to the second image sample, and the text samples in at least two different languages corresponding to the second image samples are translations of each other, the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample each include one image and text pair, and the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample include the same second image, the text in the second positive example image and text sample and the text in the second anchor image and text sample respectively match the second image semantics, and the text in the second negative example image and text sample does not match the second image semantics; inputting the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample into the second feature extraction model to obtain sample features of the second anchor image and text sample, sample features of the second positive example image and text sample, and sample features of the second negative example image and text sample; generating a third loss function value based on the sample features of the second anchor image and text sample, the sample features of the second positive example image and text sample, and the sample features of the second negative example image and text sample; and The parameters of the second feature extraction model are updated using the third loss function value.
[0221] In one possible implementation, the model training module is used to: constructing first images and text samples and translation annotations of the first images and text samples based on third image and text data samples, wherein the third image and text data samples include at least one third image sample and text samples in at least two different languages corresponding to the third image sample, and the text samples in at least two different languages corresponding to the second image sample are translations of each other, the first images and text samples include the third image sample and one text sample corresponding to the first image sample, and the translation annotation includes another text sample corresponding to the third image sample; a mask process is performed on the text sample portion of the first image and the text sample, and then the mask process is input to the second feature extraction model, and a first sample feature and a second sample feature of the first image and the text sample are output by the second feature extraction model; inputting the first sample features into a first output network to obtain a first prediction result output by the first output network, the first prediction result being used to indicate predicted text for a portion of the first image and text sample where the text sample is masked; inputting the second sample features into a second output network to obtain a second prediction result output by the second output network, the second prediction result being used to indicate predicted text for the translation annotation; generating a fourth loss function value based on the first prediction result and the first image and the portion of the text sample where the text sample is masked; generating a fifth loss function value based on the second prediction result and the translation annotations; and Parameter update is performed on the second feature extraction model based on the fourth loss function value and the fifth loss function value.
[0222] In one possible implementation, the model training module is used to: obtaining third positive example images and text samples and third negative example images and text samples, the third positive example images and text samples including image and text pairs with matching semantics, and the third negative example images and text samples including image and text pairs with mismatching semantics; inputting the third positive example image and text sample and the third negative example image and text sample into the second feature extraction model, and obtaining sample features of the third positive example image and text sample and sample features of the third negative example image and text sample output by the second feature extraction model; inputting the sample features of the third positive example image and text sample and the sample features of the third negative example image and text sample into a third output network, and obtaining a semantic match probability between the third positive example image and text sample and a semantic match probability between the third negative example image and text sample output by the third output network; generating a sixth loss function value based on the semantic match probability between the third positive example image and the text sample and the semantic match probability between the third negative example image and the text sample; and Parameter update is performed on the second feature extraction model based on the sixth loss function value.
[0223] Figure 12 is a block diagram of an image and text data processing apparatus shown in an exemplary embodiment of the present application. The apparatus may be used to perform all or part of the steps in the method shown in Figure 5 or Figure 6. As shown in Figure 12, the apparatus includes: a sample construction module 1201: used to construct a first anchor sample, a first positive example sample, and a first negative example sample based on first image and text data samples, the first image and text data samples including at least one first image sample and text samples in at least two different languages corresponding to the first image sample, and the text samples in at least two different languages corresponding to the first image sample are not translations of each other; a sample input module 1202, which is used to input the first anchor sample, the first positive example sample, and the first negative example sample into a second feature extraction model, and obtain first sample features output by the second feature extraction model; a loss calculation module 1203, used to obtain a first loss function value based on the first sample features and using the semantic relevance between the first image and the text in the text data sample as a constraint; A parameter update module 1204 is used to perform parameter updates on the second feature extraction model according to the first loss function value; and a model construction module 1205, which is used to construct a first feature extraction model based on the second feature extraction model according to the second feature extraction model satisfying the convergence condition, and which is used to process the input first image and text data to obtain data features of the first image and text data, and which outputs the processing result of a target task after the data features of the first image and text data are processed by a task processing assembly, and the target task is a classification or regression task based on image and text data.
[0224] FIG. 13 is a block diagram of a computer device 1300 according to an exemplary embodiment of the present application. The computer device may be implemented as a server or a terminal in the above-described technical solution of the present application. In this embodiment, the computer device is a server. The computer device 1300 includes a central processing unit (CPU) 1301, a system memory 1304 including a RAM (Random Access Memory) 1302 and a ROM (Read-Only Memory) 1303, and a system bus 1305 connecting the system memory 1304 and the CPU 1301. The computer device 1300 further includes a mass storage device 1306 for storing an operating system (OS) 1309, application programs (APPs) 1310, and other program modules 1311.
[0225] The mass storage device 1306 is connected to the central processing unit 1301 by a mass storage controller (not shown) that is connected to the system bus 1305. The mass storage device 1306 and its associated computer-readable media provide non-volatile storage for the computing device 1300. That is, the mass storage device 1306 may include, for example, a hard disk, a CD-ROM (Compact Disc Read-Only Memory), a drive, or other computer-readable media (not shown).
[0226] Without loss of generality, the computer-readable media may include computer storage media and communication media. Computer storage media includes, for example, volatile and nonvolatile media, removable and non-removable media implemented by any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, and other data. Computer storage media includes RAM, ROM, Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), flash memory or other solid-state storage technology, CD-ROM, Digital Versatile Disc (DVD) or other optical storage, magnetic tape cartridge, magnetic tape, magnetic storage, or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage media is not limited thereto. The system storage 1304 and mass storage device 1306 described above may be collectively referred to as storage devices.
[0227] According to various embodiments of the present disclosure, the computing device 1300 may further be connected to a remote computer on a network, such as the Internet, via a network, i.e., the computing device 1300 may be connected to a network by being connected to a network interface unit 1307 on the system bus 1305, or the computing device 1300 may be connected to another type of network or remote computer system (not shown) using the network interface unit 1307.
[0228] The memory further includes at least one computer program, which is stored in the memory, and the central processing unit 1301 executes the at least one computer program to realize all or part of the steps in the methods shown in each of the above-mentioned embodiments.
[0229] In an exemplary embodiment, a computer-readable storage medium is further provided, which is used to store at least one computer program, which is loaded and executed by a processor to implement all or part of the steps of the methods described in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0230] In an exemplary embodiment, a computer program product is further provided, the computer program product including a computer program stored in a computer-readable storage medium, and a processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, thereby causing the computer device to perform all or part of the steps of the methods described in each of the above embodiments.
[0231] Although the preferred embodiment of the present application has been described above, the present application is not limited to this embodiment, and any modification to the present application falls within the technical scope of the present application as long as it does not depart from the spirit of the present application.
Claims
1. 1. A computer device implemented method for processing image and text data, comprising: obtaining first image and text data, the first image and text data including at least one image and at least one piece of text; performing feature extraction on the first image and text data, mapping the first image and text data into a universal visual-text representation space, and obtaining data features of the first image and text data, wherein the universal visual-text representation space is a feature space constructed based on first image and text data samples with semantic relatedness between text in the first image and text data samples as a constraint, the first image and text data samples including at least one first image sample and text samples in at least two different languages corresponding to the first image sample, and the text samples in at least two different languages corresponding to the first image sample are not translations of each other; and sending data features of the first image and text data to a task processing assembly, and the task processing assembly outputting a processing result of a target task based on the data features, wherein the target task is a classification or regression task based on image and text data.
2. 2. The method of claim 1, performing feature extraction on the first image and text data, mapping the first image and text data into a universal visual text representation space, and obtaining data features of the first image and text data; inputting the first image and text data into a first feature extraction model, and obtaining data features of the first image and text data output by the first feature extraction model; the first feature extraction model is a machine learning network constructed based on the second feature extraction model, and the second feature extraction model is a machine learning model obtained by training using machine learning based on a first image and a text data sample, with semantic relevance between the first image and the text in the text data sample as a constraint.
3. The method of claim 2, further comprising: constructing a first anchor sample, a first positive example sample, and a first negative example sample based on the first image and text data sample; inputting the first anchor sample, the first positive example sample, and the first negative example sample into the second feature extraction model to obtain first sample features output by the second feature extraction model; obtaining a first loss function value based on the first sample features and using a semantic relevance between the first image and text in the text data sample as a constraint; and performing a parameter update on the second feature extraction model using the first loss function value.
4. 4. The method of claim 3, the first anchor samples include a first anchor image and a text sample and a first anchor text sample; the first positive example samples include at least one first positive example image and a text sample and at least one first positive example text sample; the first negative example samples include at least one first negative example image and a text sample and at least one first negative example text sample; the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample each include one image and text pair, the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample include the same first image, the text in the first positive example image and text sample and the text in the first anchor image and text sample each match a first image semantic, and the text in the first negative example image and text sample does not match the first image semantic, the first anchor text sample includes the text in the first anchor image and text sample, the first positive example text sample includes the text in the first positive example image and text sample, and the first negative example text sample includes the text in the first negative example image and text sample; inputting the first anchor sample, the first positive example sample, and the first negative example sample into the second feature extraction model and obtaining first sample features output by the second feature extraction model, inputting the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample into the second feature extraction model to obtain sample features of the first anchor image and text sample, sample features of the first positive example image and text sample, and sample features of the first negative example image and text sample; and inputting the first anchor text sample, the first positive example text sample, and the first negative example text sample into the second feature extraction model to obtain sample features of the first anchor text sample, sample features of the first positive example text sample, and sample features of the first negative example text sample; obtaining a first loss function value based on the first sample features and using a semantic relevance between the first image and text in the text data sample as a constraint, obtaining semantic relevance information based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, the sample features of the first negative example image and text sample, the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample, wherein the semantic relevance information is used to indicate the semantic relevance between the first image and text in the text data sample; generating a semantic relevance loss function value based on the semantic relevance information; generating a first contrastive learning loss function value based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample; and obtaining the first loss function value based on the semantic relevance loss function value and the first controlled training loss function value.
5. 5. The method of claim 4, The step of acquiring semantic relevance information based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, the sample features of the first negative example image and text sample, the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample, includes: performing a normalization process on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample to obtain first relevance distribution information in the semantic relevance information; and performing a normalization process on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample to obtain second relevance distribution information in the word semantic relevance information; The step of generating a word semantic relevance loss function value based on the word semantic relevance information includes: calculating a KL divergence for the first relevance distribution information and the second relevance distribution information to obtain the semantic relevance loss function value.
6. The method of claim 4, further comprising: generating a second loss function value based on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample; and performing a parameter update on the second feature extraction model using the second loss function value.
7. The method of claim 2, further comprising: constructing a second anchor image and text sample, a second positive example image and text sample, and a second negative example image and text sample based on the second image and text data sample, wherein the second image and text data sample include at least one second image sample and text samples in at least two different languages corresponding to the second image sample, and the text samples in at least two different languages corresponding to the second image sample are translations of each other, the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample each include one image and text pair, the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample include the same second image, the text in the second positive example image and text sample and the text in the second anchor image and text sample are consistent with a second image semantic, and the text in the second negative example image and text sample is inconsistent with the second image semantic; inputting the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample into the second feature extraction model to obtain sample features of the second anchor image and text sample, sample features of the second positive example image and text sample, and sample features of the second negative example image and text sample; generating a third loss function value based on the sample features of the second anchor image and text sample, the sample features of the second positive example image and text sample, and the sample features of the second negative example image and text sample; and performing a parameter update on the second feature extraction model using the third loss function value.
8. The method of claim 2, further comprising: constructing, based on third image and text data samples, first image and text samples and translation annotations of the first image and text samples, wherein the third image and text data samples include at least one third image sample and text samples in at least two different languages corresponding to the third image sample, and the text samples in at least two different languages corresponding to the third image sample are translations of each other, the first image and text samples include the third image sample and one text sample corresponding to the first image sample, and the translation annotation includes another text sample corresponding to the third image sample; a step of performing a mask process on the text sample portion of the first image and the text sample, and then inputting the masked portion into the second feature extraction model to obtain first sample features and second sample features of the first image and the text sample output by the second feature extraction model; inputting the first sample features into a first output network and obtaining a first prediction result output by the first output network, the first prediction result being used to indicate predicted text for a portion of the first image and text sample where the text sample is masked; inputting the second sample features into a second output network and obtaining a second prediction result output by the second output network, the second prediction result being used to indicate predicted text for the translation annotation; generating a fourth loss function value based on the first prediction result and the first image and the portion of the text sample where the text sample is masked; generating a fifth loss function value based on the second prediction result and the translation annotations; and performing a parameter update on the second feature extraction model based on the fourth loss function value and the fifth loss function value.
9. The method of claim 2, further comprising: obtaining third positive example images and text samples and third negative example images and text samples, wherein the third positive example images and text samples include image and text pairs with matching semantics, and the third negative example images and text samples include image and text pairs with mismatching semantics; inputting the third positive example image and text sample and the third negative example image and text sample into the second feature extraction model, and obtaining sample features of the third positive example image and text sample and sample features of the third negative example image and text sample output by the second feature extraction model; inputting the sample features of the third positive example image and text sample and the sample features of the third negative example image and text sample into a third output network, and obtaining a semantic match probability between the third positive example image and text sample and a semantic match probability between the third negative example image and text sample output by the third output network; generating a sixth loss function value based on the semantic match probability between the third positive example image and the text sample and the semantic match probability between the third negative example image and the text sample; and performing a parameter update on the second feature extraction model based on the sixth loss function value.
10. 1. A computer device implemented method for processing image and text data, comprising: constructing first anchor samples, first positive example samples, and first negative example samples based on first image and text data samples, wherein the first image and text data samples include at least one first image sample and text samples in at least two different languages corresponding to the first image sample, and the text samples in at least two different languages corresponding to the first image sample are not translations of each other; inputting the first anchor sample, the first positive example sample, and the first negative example sample into a second feature extraction model to obtain first sample features output by the second feature extraction model; obtaining a first loss function value based on the first sample features and using a semantic relevance between the first image and text in the text data sample as a constraint; updating parameters of the second feature extraction model according to the first loss function value; and constructing a first feature extraction model based on the second feature extraction model in response to the second feature extraction model satisfying a convergence condition, wherein the first feature extraction model is used to process input first image and text data to obtain data features of the first image and text data, and the data features of the first image and text data are processed by a task processing assembly to output a processing result of a target task, wherein the target task is a classification or regression task based on image and text data.
11. 1. An apparatus for processing image and text data, comprising: a data acquisition module for acquiring first image and text data, the first image and text data including at least one image and at least one piece of text; a feature mapping module for performing feature extraction on the first image and text data, mapping the first image and text data into a universal visual-text representation space, and obtaining data features of the first image and text data, wherein the universal visual-text representation space is a feature space constructed based on first image and text data samples with a semantic relatedness between text in the first image and text data samples as a constraint, the first image and text data samples including at least one first image sample and text samples in at least two different languages corresponding to the first image sample, and the text samples in at least two different languages corresponding to the first image sample are not translations of each other; and an apparatus including: a task processing module configured to send data features of the first image and text data to a task processing assembly and cause the task processing assembly to output a processing result of a target task based on the data features, wherein the target task is a classification or regression task based on image and text data.
12. 12. The apparatus of claim 11, the feature mapping module is used to input the first image and text data into a first feature extraction model and obtain data features of the first image and text data output by the first feature extraction model; the first feature extraction model is a machine learning network constructed based on the second feature extraction model, and the second feature extraction model is a machine learning model obtained by training using machine learning based on a first image and a text data sample, with semantic relevance between the first image and the text in the text data sample as a constraint.
13. 13. The apparatus of claim 12, further comprising: Includes a model training module, The model training module: constructing a first anchor sample, a first positive example sample, and a first negative example sample based on the first image and text data sample; inputting the first anchor sample, the first positive example sample, and the first negative example sample into the second feature extraction model to obtain first sample features output by the second feature extraction model; obtaining a first loss function value based on the first sample features and using a semantic relevance between the first image and text in the text data sample as a constraint; and An apparatus used to perform parameter updates on the second feature extraction model using the first loss function value.
14. 14. The apparatus of claim 13, the first anchor samples include a first anchor image and a text sample and a first anchor text sample; the first positive example samples include at least one first positive example image and a text sample and at least one first positive example text sample; the first negative example samples include at least one first negative example image and a text sample and at least one first negative example text sample; the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample each include one image and text pair, the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample include the same first image, the text in the first positive example image and text sample and the text in the first anchor image and text sample each match a first image semantic, and the text in the first negative example image and text sample does not match the first image semantic, the first anchor text sample includes the text in the first anchor image and text sample, the first positive example text sample includes the text in the first positive example image and text sample, and the first negative example text sample includes the text in the first negative example image and text sample; The model training module: inputting the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample into the second feature extraction model to obtain sample features of the first anchor image and text sample, sample features of the first positive example image and text sample, and sample features of the first negative example image and text sample; inputting the first anchor text sample, the first positive example text sample, and the first negative example text sample into the second feature extraction model to obtain sample features of the first anchor text sample, sample features of the first positive example text sample, and sample features of the first negative example text sample; obtaining semantic relevance information based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, the sample features of the first negative example image and text sample, the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample, wherein the semantic relevance information is used to indicate a semantic relevance between the first image and text in the text data sample; generating a semantic relevance loss function value based on the semantic relevance information; generating a first contrastive learning loss function value based on sample features of the first anchor image and text sample, sample features of the first positive example image and text sample, and sample features of the first negative example image and text sample; and An apparatus adapted to obtain the first loss function value based on the semantic relevance loss function value and the first controlled training loss function value.
15. 15. The apparatus of claim 14, The model training module: performing a normalization process on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample to obtain first relevance distribution information in the semantic relevance information; performing a normalization process on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample to obtain second relevance distribution information in the semantic relevance information; and An apparatus used to calculate KL divergence for the first relevance distribution information and the second relevance distribution information to obtain the semantic relevance loss function value.
16. 15. The apparatus of claim 14, The model training module further comprises: generating a second loss function value based on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample; and An apparatus used to perform parameter updates on the second feature extraction model using the second loss function value.
17. 1. An apparatus for processing image and text data, comprising: a sample construction module for constructing first anchor samples, first positive example samples, and first negative example samples based on first image and text data samples, wherein the first image and text data samples include at least one first image sample and text samples in at least two different languages corresponding to the first image sample, and the text samples in at least two different languages corresponding to the first image sample are not translations of each other; a sample input module for inputting the first anchor sample, the first positive example sample, and the first negative example sample into a second feature extraction model to obtain first sample features output by the second feature extraction model; a loss calculation module for obtaining a first loss function value based on the first sample features and using a semantic relevance between the first image and text in the text data sample as a constraint; a parameter updating module for performing parameter updating on the second feature extraction model according to the first loss function value; and an apparatus including: a model construction module for constructing a first feature extraction model based on the second feature extraction model in response to the second feature extraction model satisfying a convergence condition, wherein the first feature extraction model is used to process input first image and text data to obtain data features of the first image and text data, and the data features of the first image and text data are processed by a task processing assembly, after which a processing result of a target task is output, and the target task is a classification or regression task based on image and text data.
18. A computer device comprising: a processor; and a memory coupled to the processor; The storage device stores a computer program, A computing device, wherein the processor is configured to execute a computer program to implement the method of any one of claims 1 to 10.
19. A program for causing a computer to carry out the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Multi-modal model pre-training method based on multi-language picture text description data
CN117196061A
Cross-modality processing method and apparatus, and computer storage medium
US20210303921A1