Image and text data processing methods and apparatus, and computer equipment and programs
By constructing a universal visual text representation space using weakly aligned multilingual data with semantic relevance as a constraint, the method addresses the limitations of existing models, expanding training data and enhancing accuracy in cross-modal language learning.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2023-11-27
- Publication Date
- 2026-05-22
AI Technical Summary
Existing cross-modal language machine learning models require extensive training with strictly aligned image-text pairs in multiple languages, limiting the availability and diversity of training data and affecting the accuracy of the models.
Construct a universal visual text representation space using weakly aligned multilingual image and text data, incorporating semantic relevance as a constraint to enhance feature extraction and improve the accuracy of the model by utilizing a larger and more easily accessible data set.
The proposed method allows for the expansion of training data and enhances the accuracy of the universal visual text representation space by leveraging weakly aligned multilingual data, improving the model's ability to extract semantic features accurately.
Smart Images

Figure 0007864261000071 
Figure 0007864261000072 
Figure 0007864261000073
Abstract
Description
[Technical Field]
[0001] This application claims priority based on a Chinese patent application filed with the China Patent Administration on April 26, 2023, with application number 202310477720.6, and the title of the invention being "Image and Text Data Processing Method, Apparatus, Device, Storage Medium, and Program Product," the entire contents of which are incorporated herein by reference.
[0002] This application relates to the field of artificial intelligence (AI), and more particularly to image and text data processing methods and apparatus, and computer equipment and programs. [Background technology]
[0003] With the continued development of AI technology, the application of machine learning models for cross-modal languages is becoming increasingly important.
[0004] In related technologies, cross-modal language machine learning models typically need to be trained with pictures and text. Specifically, developers pre-collect pairs of images and text in multiple languages as training data and train a cross-modal language machine learning model. In this case, the same image has corresponding descriptive text in multiple languages, and these multilingual texts are translations of each other. [Overview of the project] [Problems that the invention aims to solve]
[0005] The embodiments of this application aim to provide an image and text data processing method and apparatus, as well as computer equipment and programs. [Means for solving the problem]
[0006] In one aspect, a method for processing image and text data is provided, and the method is Obtain a first image and text data, where the first image and text data include at least one image and at least one text; By performing feature extraction on the first image and text data, map the first image and text data to a universal (unified) visual text representation space, obtain the data features of the first image and text data. The universal visual text representation space is a feature space constructed based on the first image and text data samples, with the semantic (semantics) relatedness (correlation) between texts in the first image and text data samples as a constraint. The first image and text data samples include at least one first image sample and at least two text samples in different languages corresponding to the first image sample. The at least two text samples in different languages corresponding to the first image sample are not translations of each other; and Send the data features of the first image and text data to a task processing assembly, and the task processing assembly outputs a processing result of a target task based on the data features. The target task includes a classification or regression task based on images and text data.
[0007] According to one aspect, an image and text data processing method is provided. The method includes: Construct a first anchor sample, a first positive example sample, and a first negative example sample based on the first image and text data samples. The first image and text data samples include at least one first image sample and at least two text samples in different languages corresponding to the first image sample. The at least two text samples in different languages corresponding to the first image sample are not translations of each other; Input the first anchor sample, the first positive example sample, and the first negative example sample into a second feature extraction model, and obtain the first sample features output by the second feature extraction model; Based on the first sample features, obtain a first loss function value with the semantic relatedness between texts in the first image and text data samples as a constraint; Update the parameters of the second feature extraction model according to the first loss function value; and Build a first feature extraction model based on the second feature extraction model in response to the second feature extraction model satisfying the convergence condition. The first feature extraction model is used to process the input first image and text data input to obtain data features of the first image and text data. The data features of the first image and text data are processed by a task processing assembly, and then the processing result of the target task is output. The target task includes a classification or regression task based on the image and text data.
[0008] According to another aspect, an image and text data processing device is provided. The device includes A data acquisition module for acquiring a first image and text data, where the first image and text data include at least one image and at least one text; A feature mapping module for mapping the first image and text data to a universal visual text representation space by performing feature extraction on the first image and text data, and obtaining data features of the first image and text data. The universal visual text representation space is a feature space constructed based on the first image and text data samples, with the semantic relatedness between texts in the first image and text data samples as a constraint. The first image and text data samples include at least one first image sample and at least two text samples in different languages corresponding to the first image sample. The at least two text samples in different languages corresponding to the first image sample are not translations of each other. Feature mapping module; and A task processing module for sending the data features of the first image and text data to a task processing assembly, and the task processing assembly outputs a processing result of a target task based on the data features. The target task is a classification or regression task based on the image and text data. Task processing module.
[0009] According to another aspect, an image and text data processing device is provided, and the device is A sample construction module for constructing a first anchor sample, a first positive example sample, and a first negative example sample based on a first image and text data samples, wherein the first image and text data samples include at least one first image sample and at least two text samples in different languages corresponding to the first image sample, and the at least two text samples in different languages corresponding to the first image sample are not translations of each other; A sample input module for inputting the first anchor sample, the first positive example sample, and the first negative example sample into a second feature extraction model, and for obtaining the first sample features output by the second feature extraction model; A loss calculation module for obtaining a first loss function value based on the first sample features, with the semantic relevance between the first image and the text in the text data sample as a constraint; A parameter update module for updating the parameters of the second feature extraction model based on the first loss function value; and A model building module for constructing a first feature extraction model based on the second feature extraction model, in accordance with the convergence conditions of the second feature extraction model, wherein the first feature extraction model is used to process input first image and text data to obtain data features of the first image and text data, the data features of the first image and text data are processed by a task processing assembly, and the processing result of a target task is output, the target task is a classification or regression task based on image and text data, the model building module includes.
[0010] In another aspect, a computer device is provided, the computer device including a processor and a memory connected to the processor, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to realize the image and text data processing method described above.
[0011] In another aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing at least one computer program, the computer program being loaded and executed by a processor to realize the image and text data processing method described above.
[0012] In another aspect, a computer program product is provided, which includes a computer program stored in a computer-readable storage medium. The processor of the computer equipment reads the computer program from the computer-readable storage medium and executes the computer program, thereby causing the computer equipment to perform an image and text data processing method provided in the various selectable implementations described above. [Effects of the Invention]
[0013] By constructing a universal visual text representation space using at least one first image sample and at least two text samples in different languages corresponding to the first image sample as training data, when performing an image and text data processing task, the image and text data are mapped to the universal visual text representation space, and the task processing assembly outputs the task processing result based on the data features obtained by the mapping. With the above proposed technology, on the one hand, since the at least two text samples in different languages corresponding to the first image sample are not translations of each other, the amount of data is large and the acquisition difficulty is relatively low, so the data used to construct the universal visual text representation space can be expanded and the accuracy (precision) of the universal visual text representation space can be improved. On the other hand, by introducing the degree of semantic relevance between the text in the image and text data samples as a constraint in the process of constructing the universal visual text representation space, the constructed universal visual text representation space can be made to extract semantic features of the input data more accurately, and the accuracy of the universal visual text representation space constructed by the first image sample and its corresponding text sample can be further improved. [Brief explanation of the drawing]
[0014] [Figure 1] This figure shows the system used by the image and text data processing method provided in this application. [Figure 2] This figure shows strictly aligned multilingual images and text relating to the present application. [Figure 3] This figure shows weakly aligned multilingual images and text related to this application. [Figure 4] This is a flowchart of the image and text data processing method shown in the exemplary embodiment of this application. [Figure 5] This is a flowchart of the image and text data processing method shown in the exemplary embodiment of this application. [Figure 6]This is a flowchart of the image and text data processing method shown in the exemplary embodiment of this application. [Figure 7] This figure shows the network configuration of the second feature extraction model related to this application. [Figure 8] This figure shows the smooth linear interpolation method related to this application. [Figure 9] This figure shows the constrained interlingual visual-text comparative learning method related to this application. [Figure 10] This figure shows the network configuration of the second feature extraction model related to this application. [Figure 11] This is a block diagram of an image and text data processing device provided in an embodiment of this application. [Figure 12] This is a block diagram of an image and text data processing device provided in an embodiment of this application. [Figure 13] This is a block diagram of the configuration of a computer device shown in an exemplary embodiment of this application. [Modes for carrying out the invention]
[0015] Examples of this application provide data processing methods for images and text. For ease of understanding, some nouns (terms) relating to this application are explained below.
[0016] 1) Strictly aligned images and text in multiple languages. Strict alignment of images and text in multiple languages refers to a high degree of correlation between the meaning of the image content and its corresponding descriptive text, and where the descriptive texts in multiple languages are corresponding translations of each other. Strict alignment of images and text in multiple languages can also be called semantic parallelism, meaning that the descriptive texts in multiple languages have the same meaning.
[0017] 2) Weakly aligned images and text in multiple languages. So-called weak alignment of multilingual images and text refers to a situation where the semantic relationship between the image content and its corresponding descriptive text is high, but the descriptive texts in multiple languages do not necessarily have to be translations of each other. Weak alignment of multilingual images and text can also be described as semantically related but not parallel; that is, the descriptive texts in multiple languages are all related to the same image, but the semantic relationships between the descriptive texts in multiple languages are not the same / not completely the same (only some are the same).
[0018] Figure 1 shows a system used by an image and text data processing method provided in an exemplary embodiment of this application. As shown in Figure 1, the system includes a server 110 and a terminal 120.
[0019] The server 110 described above may include a server on which an image and text data processing system is deployed and which provides image and text data processing services to users through the image and text data processing system, or the server 110 described above may include a server which has an image and text data processing system and which trains or updates the image and text data processing system.
[0020] The aforementioned terminal 120 may include a user terminal that receives image and text data processing services, or it may include a development terminal used by the developer of the image and text data processing system.
[0021] Optionally, an image and text data processing system may be deployed on terminal 120 as described above.
[0022] Optionally, the system described above includes one or more servers 110 and multiple terminals 120. The number of servers 110 and terminals 120 is not limited in the embodiments of this application.
[0023] The terminal and server are connected via a communication network. Optionally, the communication network can be a wired or wireless network.
[0024] Cross-modal and cross-lingual models based on images and text offer clear advantages in handling multimodal tasks, and therefore, pre-training of cross-modal and cross-lingual models has recently attracted increasing attention. Generally, developers primarily train using strictly aligned image-text pairs in multiple languages. For example, developers can extend English-based descriptive text in an image-text pair dataset into multilingual versions using a translation method, and design a series of cross-modal and cross-lingual pre-training tasks to enable the model to learn better universal representations.
[0025] For example, see Figure 2, which shows a strictly aligned multilingual image and text relating to the present application. As shown in Figure 2, the strictly aligned multilingual image and text includes one image 21, one English text 22, and one Chinese text 23, of which the aforementioned image 21 and English text 22 may be images and text pre-collected by the developer, and the aforementioned Chinese text 23 may be text obtained by the developer translating the English text 22 using a translation tool.
[0026] Subsequent embodiments of this application provide improved cross-modal and cross-lingual pre-training frameworks that can effectively utilize large amounts of widely available and more easily collected loosely aligned multimodal data of multilingual images and text. Of which, the loosely aligned multilingual images and text may be obtained by collecting from a network.
[0027] For example, see Figure 3, which shows a weakly aligned multilingual image and text relating to the present application. As shown in Figure 3, the weakly aligned multilingual image and text includes one image 31, one English text 32, and one Chinese text 33, of which the aforementioned English text 32 and Chinese text 33 are descriptive texts in different languages for the image 31 obtained when the developer searches for the same image 31 on the network using an automated search tool.
[0028] Figure 4 is a flowchart of an image and text data processing method shown in an exemplary embodiment of the present application. The method is performed by a computer device, which may be implemented as a terminal or a server, and the terminal or server may be the terminal or server shown in Figure 1. As shown in Figure 4, the image and text data processing method includes the following steps.
[0029] Step 410: Obtain the first image and text data, which include at least one image and at least one piece of text.
[0030] In the embodiments of this application, the above-described first image and text data may include at least one image-text pair, where each image-text pair is an image-text pair consisting of one image and one line of text.
[0031] Step 420: Feature extraction is performed on the first image and text data to map the first image and text data to a universal visual text representation space, and data features are obtained for the first image and text data. The universal visual text representation space is a feature space constructed based on the first image and text data samples, with the semantic relevance between the texts in the first image and text data samples as a constraint. The first image and text data samples include at least one first image sample and at least two text samples in different languages corresponding to the first image sample, and the at least two text samples in different languages corresponding to the first image sample are not translations of each other.
[0032] In the embodiments of this application, the feature extraction process for the first image and text data described above is a process of mapping the first image and text data to a universal visual text representation space and obtaining data features of the first image and text data, or it can be said that the step of mapping the first image and text data to a universal visual text representation space is realized by performing feature extraction on the first image and text data.
[0033] In machine learning, after performing one or more feature mappings on the original input data, a higher-dimensional abstract representation can be obtained. This abstract representation may be referred to as a feature in the machine learning concept of the original data, and the space composed of the features obtained after performing one or more feature mappings on all possible input data is the feature space. In other words, the features in the feature space are higher-dimensional representations of all possible input data.
[0034] The Universal Visual Text Representation Space described above is a feature space for unifying the representation of two types of data: images and text.
[0035] In the embodiments of this application, two types of data, image and text, can be merged into a single unified feature space (i.e., the universal visual text representation space described above). The data features obtained by mapping the first image and text data into the universal visual text representation space can be represented in the form of feature vectors or feature matrices.
[0036] Among these, constructing a universal visual text representation space based on the first image and text data sample, with the semantic relevance between the text in the first image and text data sample as a constraint, may also refer to constructing a universal visual text representation space with the goal of reducing the semantic distance (or increasing the semantic relevance) between the text in the first image and text data sample.
[0037] The above-mentioned statement that a first image sample corresponds to text samples in at least two different languages may also mean that one first image sample corresponds to at least two text samples, and that each of these at least two text samples belongs to a different language (for example, Chinese, English, French, etc.), and that the meanings of all the words in these at least two text samples are related to the first image sample.
[0038] Furthermore, the statement that the text samples in at least two different languages corresponding to the first image sample are not translations of each other may also mean that the semantic features of the text samples in at least two different languages corresponding to the first image sample are different. For example, after translating the text samples in at least two different languages corresponding to the first image sample into the same language, semantic extraction can be performed on each to obtain the translated semantic feature vectors of the text samples in at least two different languages. Then, the similarity between these semantic feature vectors can be calculated, and if the similarity between any two of these semantic feature vectors is all below a certain similarity threshold (i.e., not greater than the certain similarity threshold), then the text samples in at least two different languages corresponding to the first image sample can be considered not to be translations of each other. Alternatively, for example, after translating the text samples in at least two different languages corresponding to the first image sample into the same language, keyword extraction can be performed on each to obtain the translated keywords of the text samples in at least two different languages. If the translated keywords of the text samples in at least two different languages are different, then the text samples in at least two different languages corresponding to the first image sample can be considered not to be translations of each other.
[0039] For example, suppose the image content of the first image sample is "There is a house at the foot of the mountain, and there are two puppies in front of the house," and the first image sample has two text samples, one of which is in Chinese and the other is in English and the other is in English. The meanings of both of these text samples are all related to the first image sample, but the semantic features / keywords extracted after translating these two text samples into the same language are different; in other words, these two text samples are not translations of each other.
[0040] Among these, the first image and text data sample mentioned above may also be a weakly-aligned data sample of multilingual images and text.
[0041] Step 430: The data features of the first image and text data are sent to the task processing assembly, which outputs the processing result for the target task based on the data features, and the target task is a classification or regression task based on the image and text data.
[0042] The task processing assembly described above may be a single software module (machine learning model) configured on the current computer equipment, in which case the computer equipment may input the data features obtained by the mapping described above into the task processing assembly.
[0043] Alternatively, the task processing assembly described above may be a software module installed on a computer device other than the current computer device, in which case the computer device may transmit the data features described above to the other computer device via a wired / wireless network, and the other computer device may input the data features into the task processing assembly.
[0044] In the embodiments of this application, the data features obtained by mapping the first image and text data to the universal visual text representation space in step 420 described above can be used in any subsequent classification or regression task implemented based on the image and text.
[0045] Of these, the classification task mentioned above refers to a task that outputs one classification probability after processing the data features mentioned above, and the regression task mentioned above refers to a task that outputs one image / text / image and text pair, or another data feature, after processing the data features mentioned above.
[0046] For example, data features obtained by mapping the above-described first image and text data to the universal visual text representation space are processed by a task processing assembly, and one classification probability (e.g., the probability that the image and text match, the probability that the image belongs to a certain type, etc.) or one regression result is output (e.g., one image to be reconstructed is output, or one text to be reconstructed / translated is output). The embodiments of this application are not limited to the classification task or regression task implemented based on the above-described image and text.
[0047] In summary, the invention described in the embodiment of this application constructs a universal visual text representation space using at least one first image sample and at least two text samples in different languages corresponding to the first image sample as training data. When performing an image and text data processing task, the image and text data are mapped to the universal visual text representation space, and the task processing assembly outputs the task processing result based on the data features obtained by the mapping. With the invention described above, on the one hand, because the at least two text samples in different languages corresponding to the first image sample are not translations of each other, the amount of data is large and the acquisition difficulty is relatively low, so the weakly aligned multilingual image and text data samples can be fully utilized to expand the construction data of the universal visual text representation space and improve the accuracy of the universal visual text representation space. On the other hand, by introducing the degree of semantic relevance between the text in the image and text data samples as a constraint in the process of constructing the universal visual text representation space, the constructed universal visual text representation space can be made to extract semantic features of the input data more accurately, further improving the accuracy of the universal visual text representation space constructed by the first image sample and its corresponding text sample.
[0048] In the embodiment shown in Figure 2 above, the universal visual text representation space described above can be represented by a machine learning model. This machine learning model is first trained with training data consisting of image-text pairs, and then processes subsequently input image and text data to obtain data features of the image and text data within the universal visual text representation space described above.
[0049] Referring to Figure 5 based on the embodiment shown in Figure 4, which is a flowchart of the image and text data processing method shown in an exemplary embodiment of the present application. The method is performed by a computer device, which may be implemented as a terminal or a server, which may be the terminal or server shown in Figure 1, and as shown in Figure 5, the process of training and applying a machine learning model for image and text data processing may include the following steps.
[0050] Step 510: Construct the first anchor sample, the first positive example sample, and the first negative example sample based on the first image and text data sample.
[0051] Among these, the first image and text data sample includes at least one first image sample and at least two text samples in different languages corresponding to the first image sample, and the at least two text samples in different languages corresponding to the first image sample are not translations of each other.
[0052] In the embodiments of this application, the developer can pre-collect samples of weakly aligned image and text data (i.e., the first image and text data samples described above), and then construct samples necessary for comparative learning based on the first image and text data samples, such samples including a first anchor sample as an anchor, a first positive example sample as a positive sample, and a first negative example sample as a negative sample.
[0053] Of these, the first positive example sample mentioned above refers to a sample that matches the features of the first anchor sample, while the first negative example sample mentioned above refers to a sample that does not match the features of the first anchor sample.
[0054] For example, when the first anchor sample contains an image-text pair, the first positive example sample and the first negative example sample may each contain one image-text pair. The relationship between the image and text in the first positive example sample is similar to the relationship between the image and text in the first anchor sample, while the relationship between the image and text in the first negative example sample is not similar to the relationship between the image and text in the first anchor sample.
[0055] For example, the first anchor sample, the first positive example sample, and the first negative example sample each contain the same image, and each contains different text, and the text in the first anchor sample and the first positive example sample is different, but the meanings of the words in these two texts are all close to the meanings of the image, while the meaning of the text in the first negative example sample is not close to or related to the meanings of the image.
[0056] Furthermore, for example, when the first anchor sample contains one text, the first positive example sample and the first negative example sample may each contain one text, and the semantics of the text in the first positive example sample and the text in the first anchor sample are similar, while the semantics of the text in the first negative example sample and the text in the first anchor sample are not similar.
[0057] Step 520: Input the first anchor sample, the first positive example sample, and the first negative example sample into the second feature extraction model, and obtain the first sample features output by the second feature extraction model.
[0058] In the embodiment of this application, the computer equipment may input the first anchor sample, the first positive example sample, and the first negative example sample into the second feature extraction model, process the first anchor sample, the first positive example sample, and the first negative example sample using the second feature model, and output the respective sample features of the first anchor sample, the first positive example sample, and the first negative example sample. The respective sample features of the first anchor sample, the first positive example sample, and the first negative example sample can constitute the first sample features described above.
[0059] Step 530: Based on the first sample features, obtain the first loss function value with the semantic relevance between the first image and the text in the text data sample as a constraint.
[0060] In the embodiments of this application, when the computer device generates a first loss function value based on the first sample features, it may also consider the semantic relevance between the first image and the text in the text data sample. In other words, the computer device may generate the first loss function value in combination with two pieces of information: the first sample features and semantic relevance information between the first image and the text in the text data sample.
[0061] Among these, the semantic relevance information between the first image and the text in the text data sample may be obtained by calculating based on the features of the first sample. For example, a computer device may obtain the semantic relevance information between the text in the first anchor sample, the first positive example sample, and the first negative example sample by processing the features of the text in the first anchor sample, the first positive example sample, and the first negative example sample from the first sample features described above.
[0062] The semantic relevance mentioned above may also refer to the degree of semantic similarity between two or more texts.
[0063] The process by which a computer device obtains a first loss function value based on the first sample features, constrained by the semantic relevance between the first image and the text in the text data sample, may include the following:
[0064] The computer device generates a partial loss function value using a comparative learning method based on the first sample features, then generates a partial loss function value based on semantic relevance information, and finally obtains the aforementioned first loss function value by adding these two partial loss function values together or by calculating a weighted sum of them.
[0065] Step 540: Update the parameters of the second feature extraction model using the first loss function value.
[0066] In the embodiment of this application, the calculation process for the first loss function value takes into account semantic relevance information between texts in the first anchor sample, the first positive example sample, and the first negative example sample. This allows for training using weakly aligned image and text data samples, thereby extending the available training data for the second feature extraction model to weakly aligned image and text data samples and achieving the effect of training the extended model with more training data, which is also advantageous for improving the accuracy of model training.
[0067] Step 550: If the second feature extraction model satisfies the convergence conditions, construct the first feature extraction model based on the second feature extraction model.
[0068] The first feature extraction model is used to process the input first image and text data to obtain data features of the first image and text data. After the data features of the first image and text data are processed by the task processing assembly, the processing result of the target task is output, and the target task is a classification or regression task based on image and text data.
[0069] In other words, the first feature extraction model described above is a machine learning network built on the second feature extraction model, and the second feature extraction model is a machine learning model obtained by performing machine learning-based training on the first image and text data samples, with the degree of semantic relevance between the text in the first image and the text data samples as a constraint.
[0070] In one possible implementation, the second feature extraction model described above can be used directly as a first feature extraction model deployed as an application for feature extraction of image and text data.
[0071] In another possible implementation, the developer may improve the aforementioned second feature extraction model using computer equipment and use the improved second feature extraction model as the first feature extraction model deployed directly as an application for feature extraction of image and text data.
[0072] For example, a developer may obtain the first feature extraction model described above by optimizing the second feature extraction model using a computer. For example, a developer may retrain the second feature extraction model using a computer to make it more suitable for a specific image and text data processing task. Alternatively, a developer may improve the processing speed of the model by simplifying the second feature extraction model (e.g., pruning, distillation) using a computer to obtain a lightweight first feature extraction model.
[0073] Step 560: Obtain the first image and text data, which include at least one image and at least one piece of text.
[0074] After the first feature extraction model is deployed as an application, the computer device on which the first feature extraction model is deployed can perform image and text data processing tasks, and at this time, the computer device can obtain the first image and text data awaiting processing.
[0075] In one possible implementation, the first image and text data described above may include a small number of image-text pairs.
[0076] Step 570: Input the first image and text data into the first feature extraction model, and obtain the data features of the first image and text data output by the first feature extraction model.
[0077] Of these, the first feature extraction model described above is used to map the features of the input image and text data into the universal visual-text representation space.
[0078] In one possible implementation, a computer device inputs several pairs of images and text from the first image and text data into the first feature extraction model described above, causing the first feature extraction model to process the image and text pairs and output data features of the first image and text data.
[0079] Step 580: Send the data features of the first image and text data to the task processing assembly, which outputs the processing result of the target task based on the data features, and the target task is a classification or regression task based on the image and text data.
[0080] In particular, the above-described embodiment relates to the training process of the second feature extraction model (steps 510 to 540), the construction process of the first feature extraction model (step 550), and the application process of the first feature extraction model (steps 560 to 580). The above three processes may be executed by different computer equipment, or two of the above three processes may be executed by one computer equipment and the other by another computer equipment, or the above three processes may be executed by the same computer equipment.
[0081] In the embodiments of this application, the computer device can improve the accuracy of feature mapping and enhance the accuracy of the processing results for subsequent target tasks by mapping the input first image and text data into a universal visual-text representation space using a single pre-trained first feature extraction model.
[0082] In the embodiments of this application, the computer device can train a second feature extraction model in a contrast learning manner based on a first anchor sample, a first positive example sample, and a first negative example sample, which are constructed based on a first image and text data sample.
[0083] In one possible implementation, the first anchor sample described above includes a first anchor image and text sample, and a first anchor text sample; the first positive example sample includes at least one first positive example image and text sample, and at least one first positive example text sample; and the first negative example sample includes at least one first negative example image and text sample, and at least one first negative example text sample.
[0084] Each of the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample each contains one image and text pair, and each of the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample each contains the same first image, the text in the first positive example image and text sample and the text in the first anchor image and text sample each match the meaning of the first image, the text in the first negative example image and text sample does not match the meaning of the first image, the first anchor text sample contains the text in the first anchor image and text sample, the first positive example text sample contains the text in the first positive example image and text sample, and the first negative example text sample contains the text in the first negative example image and text sample.
[0085] Referring to Figure 6 based on the embodiment shown in Figure 5, which is a flowchart of the image and text data processing method shown in an exemplary embodiment of this application. As shown in Figure 6, step 520 in the embodiment shown in Figure 5 above is implemented as step 520a and step 520b, and step 530 may be implemented as step 530a to step 530d.
[0086] Step 520a: Input the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample into the second feature extraction model to obtain sample features for the first anchor image and text sample, sample features for the first positive example image and text sample, and sample features for the first negative example image and text sample.
[0087] In the embodiments of this application, the computer equipment for training the second feature extraction model may input the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample to the second feature extraction model, and the second feature extraction model may output the sample features of the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample.
[0088] In one possible implementation, the above-described second feature extraction model may include at least one feature extraction network, and when it includes at least two feature extraction networks, the network configuration (structure) and parameters of the above-described at least two feature extraction networks may be the same or different. For example, all or some of the above-described at least two feature extraction networks may be connected sequentially, connected in a residual manner, connected in an intersection manner, or connected in a circular manner, but the embodiments of this application are not limited to the connection method between the above-described feature extraction networks.
[0089] For illustrative purposes, refer to Figure 7, which shows the network configuration of a second feature extraction model according to an embodiment of the present application. As shown in Figure 7, the second feature extraction model can be divided according to its function into a text feature extraction branch 701, an image feature extraction branch 702, and a feature fusion branch 703. After a computer inputs a pair of image and text into the second feature extraction model, the image (or image sample) in the image and text pair (or image and text sample pair) is input to the image feature extraction branch 702 to obtain image features (or image sample features), the text (or text sample) in the image and text pair (or image and text sample pair) is input to the text feature extraction branch 701 to obtain text features (or text sample features), and the above-mentioned image features (or image sample features) and text features (or text sample features) are input to the feature fusion branch 703 to obtain data features (or data sample features) output by the model.
[0090] While Figure 7 above shows only one network configuration for a second feature extraction model that is possible using an exemplary method, the second feature extraction model may optionally employ other configurations, such as other connection configurations or other network branching methods.
[0091] Step 520b: Input the first anchor text sample, the first positive example text sample, and the first negative example text sample into the second feature extraction model to obtain the sample features of the first anchor text sample, the first positive example text sample, and the first negative example text sample.
[0092] In the embodiments of this application, the second feature extraction model described above may include a text feature extraction network (for example, the text feature extraction branch 701 in Figure 7 described above), and the computer equipment may input the first anchor text sample, the first positive example text sample, and the first negative example text sample into the text feature extraction network described above, and obtain the respective sample features of the first anchor text sample, the first positive example text sample, and the first negative example text sample.
[0093] Step 530a: Semantic relevance information is obtained based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, the sample features of the first negative example image and text sample, the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample. This semantic relevance information is used to indicate the semantic relevance between the first image and the text in the text data sample.
[0094] In the above-described embodiment, the sample features of the first anchor text sample, the first positive example text sample, and the first negative example text sample are all text-related features. Similarly, the sample features of the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample also include some text-related features. The computer can then calculate the semantic relevance between the first image and the text in the text data sample by combining these with the sample features of the image and text sample, and the sample features of the text sample.
[0095] Step 530b: Generate a semantic relevance loss function value based on semantic relevance information.
[0096] In one possible implementation, obtaining semantic relevance information based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, the sample features of the first negative example image and text sample, the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample is possible. Normalization is performed on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample to obtain the first relevance distribution information in the semantic relevance information; and This includes performing normalization on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample to obtain second relevance distribution information in the semantic relevance information.
[0097] In the embodiment of this application, the computer device can calculate relevance distribution information regarding the semantics of each data type (image and text data or plain text data) according to the data type corresponding to the features using a normalization process. Specifically, the computer device can calculate relevance distribution information regarding the semantics of image and text data based on the sample features of the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample. Furthermore, the computer device can calculate relevance distribution information regarding the semantics of plain text data based on the sample features of the first anchor text sample, the first positive example text sample, and the first negative example text sample. These two relevance distribution pieces constitute the semantic relevance information described above.
[0098] Generating a semantic relevance loss function value based on the semantic relevance information described above involves calculating the KL divergence for the first and second relevance distribution information and obtaining a semantic relevance loss function value.
[0099] In the embodiment of this application, the computer device calculates the KL divergence of the first and second relevance distribution information, and the fusion of the first and second relevance distribution information is referred to as the semantic relevance loss function value, which may be used to represent the semantic relevance between the first image and the text in the text data sample. This enables the construction of a loss function value related to semantic relevance and the extension of the available training data for the second feature extraction model to weakly aligned image and text data samples.
[0100] Alternatively, the computer device may obtain semantic relevance information by other means. For example, the computer device may first fuse the sample features of the first anchor image and text sample with the sample features of the first anchor text sample, and then, after concatenating the sample features of the first anchor image and text sample with the sample features of the first anchor text sample, obtain fused anchor sample features by performing a fusion process using a single fusion network (for example, the fusion network may include network layers such as fully connected layers, convolutional layers, and pooling layers). Similarly, the computer device may fuse the sample features of the first positive example image and text sample with the sample features of the first positive example text sample to obtain fused positive example sample features, or fuse the sample features of the first negative example image and text sample with the sample features of the first negative example text sample to obtain fused negative example sample features. Subsequently, the computer device further calculates the above-mentioned semantic relevance information based on the fused anchor sample features, fused positive example sample features, and fused negative example sample features. At this time, the computer device calculates the semantic relevance loss function value using the above-mentioned semantic relevance information.
[0101] Step 530c: Calculate the first control learning loss function value based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample.
[0102] In the embodiment of this application, in addition to the semantic relevance loss function value, the computer device can further improve the accuracy of subsequent model training by generating a first contrast learning loss function value in a contrast learning manner, along with the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample.
[0103] Step 530d: Obtain the first loss function value based on the semantic relevance loss function value and the first control learning loss function value.
[0104] In the embodiments of this application, the computer equipment may obtain the first loss function value by calculating it together with the semantic relevance loss function value and the first control learning loss function value. For example, the computer equipment may obtain the first loss function value by adding the semantic relevance loss function value and the first control learning loss function value or by calculating a weighted sum of the two. The weights of the weighted sum may be predetermined by the developer.
[0105] In the invention described in the embodiment of this application, the computer device calculates a semantic relevance loss function value based on the features of the input text sample, then obtains a first loss function value by combining it with the loss function value of control learning, and further updates the parameters of the second feature extraction model using the first loss function value. The semantic relevance between the first image and the text in the text data sample is then used as a constraint and integrated into the model training process. This expands the available training data for the second feature extraction model to weakly aligned image and text data samples, achieving the effect of training the extended model with more training data, and also leads to improved accuracy in model training.
[0106] In one possible implementation, the above method further includes the following:
[0107] A second loss function value is generated based on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample; and The parameters of the second feature extraction model are updated based on the second loss function value.
[0108] In the embodiments of this application, the computer device can further expand the training data of the second feature extraction model and improve the extraction effect of the second feature extraction model by independently updating the parameters of the second feature extraction model using the sample features of the text sample.
[0109] For example, taking Figure 7 as an example again, the computer device inputs the first anchor text sample, the first positive example text sample, and the first negative example text sample into the text feature extraction branch 701, then obtains the sample features of the first anchor text sample, the first positive example text sample, and the first negative example text sample. Subsequently, in addition to calculating the first loss function value described above, it calculates the second loss function value using a comparative learning method based solely on the sample features of the first anchor text sample, the first positive example text sample, and the first negative example text sample. Then, it updates the parameters of the second feature extraction model using the second loss function value, for example, it updates the parameters of the text feature extraction branch in the second feature extraction model using the second loss function value.
[0110] In one possible implementation, the above method further includes the following: Based on the second image and text data sample, a second anchor image and text sample, a second positive example image and text sample, and a second negative example image and text sample are constructed, each containing at least one second image sample and at least two text samples in different languages corresponding to the second image sample, and these at least two text samples in different languages corresponding to the second image sample are translations of each other; each of the second anchor image and text sample, second positive example image and text sample, and second negative example image and text sample contains one image and text pair, and each of the second anchor image and text sample, second positive example image and text sample, and second negative example image and text sample contains the same second image; the text in the second positive example image and text sample and the text in the second anchor image and text sample correspond to the second image semantics, respectively, while the text in the second negative example image and text sample does not correspond to the second image semantics; The second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample are input to the second feature extraction model to obtain the sample features of the second anchor image and text sample, the sample features of the second positive example image and text sample, and the sample features of the second negative example image and text sample; and A third loss function value is generated based on the sample features of the second anchor image and text sample, the sample features of the second positive example image and text sample, and the sample features of the second negative example image and text sample; and The parameters of the second feature extraction model are updated based on the value of the third loss function.
[0111] In the embodiments of this application, in addition to training a second feature model with weakly aligned first images and text data samples, the computer equipment may further train the second feature model with strictly aligned second images and text data samples. Of these, the process of training the second feature model with second images and text data samples reduces the calculation process of the semantic relevance loss function value compared to the process of training the second feature model with weakly aligned first images and text data samples. The calculation process of the third loss function value described above is the same as the calculation process of the first control learning loss function value, and a detailed explanation thereof is omitted here.
[0112] In the embodiment of this application, a Regularized Cross-lingual Visio-textual Contrastive Learning (R-XVtCL) task is employed, and a dataset of weakly-aligned image and text pairs in multiple languages (i.e., the first image and text data sample described above, D) is used. w A second machine learning model may be trained based on (represented by ). Optionally, the task may also simultaneously train a dataset of strictly-aligned image-text pairs in multiple languages (i.e., the first image and text data samples mentioned above, D). s It is also acceptable to include (represented by).
[0113] Optionally, in the embodiments of this application, a second machine learning model may be further trained by a single cross-lingual textual contrastive learning (XTCL) task.
[0114] The objective of the XTCL task is to obtain meaningful textual representations (TRs) for multilingual texts in a single Universal Textual Representation Space (UTRS). In this UTRS space, textual representations with relatively high semantic relevance should be close to each other, while textual representations with low relevance should be far apart. In this application, comparative learning in the UTRS space is performed to cause the model to generate representations that satisfy the above properties.
[0115] Specifically, one multilingual text dataset D containing large parallel word pairs t Give one batch from among them
[0116]
number
[0117]
number
[0118]
number
[0119]
number
[0120]
number
[0121]
Number
[0122] Among them,
[0123]
Number
[0124]
Number
[0125] In the task of contrastive learning of visual text between languages with constraints in the pre-training stage, in addition to measuring the similarity between positive and negative examples and the anchor using the Euclidean distance, other measurement means of text relevance, such as the cosine similarity between expression feature vectors, may be adopted. That is, the above-mentioned Euclidean distance can be replaced by the cosine similarity.
[0126] Of these, the negative example representations obtained by the above-described method may have insufficient difficulty for the model to perform contrast-based training, and it may not be possible to train the model to learn more informative effective text representations within the UTRS space. In this case, the technical solution shown in the embodiment of this application may further employ a negative example representation generation method based on smooth linear interpolation to generate negative examples that are more difficult to distinguish, for example,
[0127]
number
[0128]
number
[0129]
number
[0130]
number
[0131] Of these, ζ (zeta) is one pre-defined scaling factor,
[0132]
number
[0133]
number
[0134]
number
[0135] The goal of the R-XVtCL task is to enable the model to learn meaningful representations of multilingual visual-text inputs in a single Universal Visio-textual Representation Space (UVtRS) through contrastive learning.
[0136] To deepen our understanding, we first consider performing comparative learning using strictly-aligned multilingual image-text pair data, D s Triple of one batch of bilingual parallel visual text
[0137]
number
[0138]
number
[0139]
number
[0140]
number
[0141]
number
[0142]
number
[0143]
number
[0144]
number
[0145]
number
[0146]
number
[0147]
number
[0148]
number
[0149]
number
[0150]
number
[0151]
number
[0152] Of these, N(v',x') is a set containing the three types of negative examples constructed as described above.
[0153]
number
[0154]
number
[0155]
number
[0156] However, in the process of performing comparative learning using the aforementioned "Strictly-aligned" image-text data, vtr + to, vtr * While it is considered an expression that needs to approach the same thing, this relationship can deviate when the data only satisfies the "weakly-aligned" relationship; that is, the semantics of descriptive texts in different languages corresponding to the same image may be related to some extent but not equivalent. In such cases, simply using vtr + and vtr *It is impractical to directly bring them closer together. To achieve this, the semantic similarity between multilingual texts needs to be measured in a way that more effectively captures the semantic relationships between them in the UVtRS space, thereby obtaining better visual text representations. Therefore, in the case of "weakly-aligned" multilingual image-text data, this application uses the text semantic relevance between multilingual descriptive texts to impose additional constraints on the aforementioned comparative learning process. For the semantic relevance of any two texts, this application measures it by the Euclidean distance between the text representations in the UTRS space learned in the XVTCL task; that is, the smaller the corresponding distance, the more related the texts' meanings are.
[0157] Specifically, D w Triple of two multilingual visual texts in a triple of bilingual parallel visual texts from one patch
[0158]
number
[0159]
number
[0160]
number
[0161]
number
[0162]
number
[0163]
number
[0164]
number
[0165]
number
[0166]
number
[0167]
number
[0168]
number
[0169]
number
[0170]
number
[0171]
number
[0172]
Number
[0173]
Number
[0174]
Number
[0175] Also,
[0176]
Number
[0177]
Number
[0178]
Number
[0179]
Number
[0180]
Number
[0181]
Number
[0182]
Number
[0183] Then, use KL divergence to bring the above two distributions closer and perform training together with the normal contrast learning objective, that is, constrained cross-modal visual-text contrast learning,
[0184]
Number
[0185] Finally, like the constrained cross-modal visual-text contrast learning shown in Figure 9, D s and D w Integrate the training examples from, and the training objective of R-XVtCL can be described in the following form, that is,
[0186]
Number
[0187] For illustrative purposes, refer to Figure 10, which shows the network configuration of a second feature extraction model according to an embodiment of the present application. In the process of training the second feature model based on weakly aligned image and text data samples, as shown in Figure 10, the second feature extraction model can be divided according to its function into a text feature extraction branch 1001, an image feature extraction branch 1002, and a feature fusion branch 1003. The computer equipment pre-collects an image and text data sample 1004 containing weakly aligned first image and text data samples, and constructs input samples based on the image and text data sample 1004, which includes constructing a first anchor image and text sample, a first anchor text sample, a first positive example image and text sample, a first positive example text sample, a first negative example image and text sample, and a first negative example text sample based on the first image and text data sample.
[0188] Subsequently, the computer equipment inputs the text from the constructed input sample into the text feature extraction branch 1001 to obtain sample features 100 of the input text (including sample features of the first anchor text sample, the first positive example text sample, and the first negative example text sample), inputs the image from the constructed input sample into the image feature extraction branch 1002 to obtain sample features 1006 of the input image, and these sample features 1005 and 1006 are input into the feature fusion branch 1003 to obtain sample features 1007 (including sample features of the first anchor image and text sample, the first positive example image and text sample, the first positive example text sample, and the first negative example image and text sample).
[0189] Subsequently, on the one hand, the computer generates first relevance distribution information 1008 from sample feature 1007, second relevance distribution information 1009 from sample feature 1006, and then generates relevance loss function value 1010 based on first relevance distribution information 1008 and second relevance distribution information 1009. On the other hand, the computer generates first control learning loss function value 1011 from sample feature 1007, and then generates first loss function value 1012 from relevance loss function value 1010 and first control learning loss function value 1011. Finally, the computer updates the parameters of the second feature extraction model based on first loss function value 1012.
[0190] In one possible implementation, the method further includes the following: Based on the third image and text data sample, a first image and text sample, and a translated annotation of the first image and text sample are constructed, the third image and text data sample includes at least one third image sample and at least two text samples in different languages corresponding to the third image sample, and the at least two text samples in different languages corresponding to the second image sample are translations of each other, the first image and text sample includes the third image sample and the text sample corresponding to the first image sample, and the translated annotation includes another text sample corresponding to the third image sample; After masking the text sample portion of the first image and text sample, the data is input to a second feature extraction model. The first and second sample features of the first image and text sample, output by the second feature extraction model, are then obtained; The first sample features are input to the first output network, and the first prediction result output by the first output network is obtained. This first prediction result is used to show the predicted text for the masked portion of the text sample in the first image and text sample; The second sample features are input to the second output network, the second prediction result output by the second output network is obtained, and the second prediction result is used to show the predicted text of the translated text annotation; Based on the first prediction result and the masked portion of the text sample in the first image and text sample, a fourth loss function value is generated; Based on the second prediction result and the translation annotations, a fifth loss function value is generated; and The parameters of the second feature extraction model are updated based on the fourth and fifth loss function values.
[0191] In the embodiment of this application, the computer device performs parameter updates for the second feature extraction model in conjunction with the image translation task, thereby expanding the training data and training method of the second feature extraction model and improving the extraction effect of the second feature extraction model.
[0192] In the embodiments of this application, the above-described second feature extraction model may be further trained by training the task using a masked conditional language model (MCLM).
[0193] MCLM includes a standard masked language modeling (MLM) on the source language encoder side and a conditional generative language modeling (CLM) on the target language decoder side. (Image v, and language l) i and language l j Image description text belonging to each category
[0194]
number
[0195]
number
[0196]
number
[0197]
number
[0198]
number
[0199]
number
[0200]
number
[0201] Eventually, D s θ represents a dataset of image-text pairs with strictly aligned images and text in multiple languages, e represents the trainable parameters of the model's encoder. On the decoder side, similarly in this application, the model decodes in an autoregressive manner by training the model with CLM targets based on the visual text input from the encoder side.
[0202]
number
[0203]
number
[0204] Eventually, θ d represents the trainable parameters of the model's decoder. Combining the goals for both the encoder and decoder, the total MCLM goal is L MCLM =L mlm +L clm It can be written as follows.
[0205] In one possible implementation, the method further, Obtain a third positive example image and text sample, and a third negative example image and text sample; the third positive example image and text sample contains image and text pairs that match semantically, and the third negative example image and text sample contains image and text pairs that do not match semantically; The third positive example image and text sample, and the third negative example image and text sample are input into the second feature extraction model, and the sample features of the third positive example image and text sample, and the sample features of the third negative example image and text sample, which are output by the second feature extraction model, are obtained; The sample features of the third positive example image and text sample, and the sample features of the third negative example image and text sample are input to the third output network, and the semantic agreement probabilities of the third positive example image and text sample and the semantic agreement probabilities of the third negative example image and text sample are obtained from the third output network; A sixth loss function value is generated based on the semantic agreement probability between the third positive example image and the text sample, and the semantic agreement probability between the third negative example image and the text sample; and This includes updating the parameters of the second feature extraction model based on the sixth loss function value.
[0206] In the embodiments of this application, a second feature extraction model may be further trained using an image-text matching (ITM) task.
[0207] In the embodiments of this application, the computer device can perform the task of predicting the degree of image-text matching, and in conjunction with updating the parameters of the second feature extraction model, thereby expanding the training data and training method of the second feature extraction model and improving the extraction effect of the second feature extraction model.
[0208] The goal of the ITM task is to take a given image input v and a text input x. l The objective is to determine if the image-text pairs match, and the target function can be written as follows:
[0209]
number
[0210] Among them, y∈{0,1} is text x l This indicates whether image v matches the given image.
[0211]
number
[0212] In other words, the above-described embodiments of this application provide a pre-training framework that can more effectively utilize large amounts of weakly-aligned, multimodal data of multilingual images and texts, and the entire training process includes four pre-training objectives: Masked Conditional Language Modelling (MCLM), Image-Text Matching (ITM), Cross-lingual Textual Contrastive Learning (XTCL), and Regularized Cross-lingual Visio-textual Contrastive Learning (R-XVtCL). The above-described embodiments improve the effectiveness of a series of downstream visual-text tasks by enabling the model to learn better visual-text representations within a unified cross-modal and cross-lingual vector representation space.
[0213] Figure 11 is a block diagram of an image and text data processing device shown in an exemplary embodiment of this application. The device may be used to perform all or some of the steps in the method shown in Figure 4, Figure 5, or Figure 6. As shown in Figure 11, the device includes, namely, Data acquisition module 1101: Used to acquire a first image and text data, the first image and text data include at least one image and at least one text; Feature mapping module 1102: Used to map the first image and text data to a universal visual text representation space by performing feature extraction on the first image and text data, and to obtain data features of the first image and text data, wherein the universal visual text representation space is a feature space constructed based on the first image and text data samples, with the semantic relevance between the texts in the first image and text data samples as a constraint, wherein the first image and text data samples include at least one first image sample and at least two text samples in different languages corresponding to the first image sample, and the at least two text samples in different languages corresponding to the first image sample are not translations of each other; and Task processing module 1103: Used to transmit the data features of the first image and text data to the task processing assembly, and to cause the task processing assembly to output the processing result of the target task based on the data features, the target task being a classification or regression task based on the image and text data.
[0214] In one possible implementation, the feature mapping module 1102 is used to do the following: The first image and text data are input to the first feature extraction model, and the data features of the first image and text data output by the first feature extraction model are obtained. The first feature extraction model is a machine learning network built on the second feature extraction model, and the second feature extraction model is a machine learning model obtained by training with machine learning based on the first image and text data samples, with the degree of semantic relevance between the text in the first image and the text data samples as a constraint.
[0215] In one possible implementation, the apparatus further includes a model training module, which is used to do the following: Based on the first image and text data sample, a first anchor sample, a first positive example sample, and a first negative example sample are constructed; The first anchor sample, the first positive example sample, and the first negative example sample are input to the second feature extraction model, and the first sample features output by the second feature extraction model are obtained; Based on the first sample features, the first loss function value is obtained with the semantic relevance between the first image and the text in the text data sample as a constraint; and The parameters of the second feature extraction model are updated based on the first loss function value.
[0216] In one possible implementation, the first anchor sample includes a first anchor image, a text sample, and a first anchor text sample; the first positive example sample includes at least one first positive example image, a text sample, and at least one first positive example text sample; and the first negative example sample includes at least one first negative example image, a text sample, and at least one first negative example text sample; The first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample each contain one image and text pair, and the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample each contain the same first image, the text in the first positive example image and text sample and the text in the first anchor image and text sample each match the first image word meaning, the text in the first negative example image and text sample does not match the first image word meaning, the first anchor text sample includes the text in the first anchor image and text sample, the first positive example text sample includes the text in the first positive example image and text sample, and the first negative example text sample includes the text in the first negative example image and text sample.
[0217] The aforementioned model training module is used to do the following: The first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample are input to the second feature extraction model to obtain the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample; The first anchor text sample, the first positive example text sample, and the first negative example text sample are input into the second feature extraction model to obtain the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample; Semantic relevance information is obtained based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, the sample features of the first negative example image and text sample, the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample, and the semantic relevance information is used to indicate the semantic relevance between the first image and the text in the text data sample; Based on the aforementioned semantic relevance information, a semantic relevance loss function value is generated; Based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample, a first control learning loss function value is generated; and The first loss function value is obtained based on the aforementioned semantic relevance loss function value and the first control learning loss function value.
[0218] In one possible implementation, the model training module is used to do the following: Normalization processing is performed on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample to obtain the first relevance distribution information in the semantic relevance information; Normalization processing is performed on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample to obtain the second relevance distribution information in the semantic relevance information; and The KL divergence is calculated for the first and second relevance distribution information, and the semantic relevance loss function value is obtained.
[0219] In one possible implementation, the model training module is used to do the following: A second loss function value is generated based on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample; and The parameters of the second feature extraction model are updated based on the second loss function value.
[0220] In one possible implementation, the model training module is used to do the following: Based on a second image and text data sample, a second anchor image and text sample, a second positive example image and text sample, and a second negative example image and text sample are constructed, wherein the second image and text data sample includes at least one second image sample and text samples in at least two different languages corresponding to the second image sample, and the text samples in at least two different languages corresponding to the second image sample are translations of each other; the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample each include one image and text pair, and the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample include the same second image; the text in the second positive example image and text sample and the text in the second anchor image and text sample each match the second image semantics, and the text in the second negative example image and text sample does not match the second image semantics; The second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample are input to the second feature extraction model to obtain the sample features of the second anchor image and text sample, the sample features of the second positive example image and text sample, and the sample features of the second negative example image and text sample; Based on the sample features of the second anchor image and text sample, the sample features of the second positive example image and text sample, and the sample features of the second negative example image and text sample, a third loss function value is generated; and The parameters of the second feature extraction model are updated based on the value of the third loss function.
[0221] In one possible implementation, the model training module is used to do the following: Based on a third image and text data sample, a first image and text sample and a translated annotation of the first image and text sample are constructed, wherein the third image and text data sample includes at least one third image sample and at least two text samples in different languages corresponding to the third image sample, and the at least two text samples in different languages corresponding to the second image sample are translations of each other, the first image and text sample includes the third image sample and one text sample corresponding to the first image sample, and the translated annotation includes another text sample corresponding to the third image sample; After applying a mask to the text sample portion of the first image and the text sample, the data is input to the second feature extraction model, and the first and second sample features of the first image and text sample, which are output by the second feature extraction model, are obtained; The first sample features are input to the first output network, and the first prediction result output by the first output network is obtained. The first prediction result is used to show the predicted text for the portion of the text sample in the first image and text sample that is masked; The second sample features are input to a second output network, the second prediction result output by the second output network is obtained, and the second prediction result is used to show the predicted text of the translated text annotation; A fourth loss function value is generated based on the first prediction result and the first image and the portion of the text sample in the text sample that is masked; Based on the second prediction result and the translation annotation, a fifth loss function value is generated; and The parameters of the second feature extraction model are updated based on the fourth and fifth loss function values.
[0222] In one possible implementation, the model training module is used to do the following: A third positive example image and text sample and a third negative example image and text sample are obtained, the third positive example image and text sample includes pairs of images and text that match semantically, and the third negative example image and text sample includes pairs of images and text that do not match semantically; The third positive example image and text sample and the third negative example image and text sample are input to the second feature extraction model, and the sample features of the third positive example image and text sample and the sample features of the third negative example image and text sample, which are output by the second feature extraction model, are obtained; The sample features of the third positive example image and text sample, and the sample features of the third negative example image and text sample are input to the third output network, and the semantic matching probabilities of the third positive example image and text sample and the semantic matching probabilities of the third negative example image and text sample, which are output by the third output network, are obtained; A sixth loss function value is generated based on the semantic agreement probability between the third positive example image and the text sample and the semantic agreement probability between the third negative example image and the text sample; and Based on the sixth loss function value, the parameters of the second feature extraction model are updated.
[0223] Figure 12 is a block diagram of an image and text data processing device shown in an exemplary embodiment of this application. The device may be used to perform all or some of the steps in the method shown in Figure 5 or Figure 6. As shown in Figure 12, the device includes, namely, Sample construction module 1201: Used to construct a first anchor sample, a first positive example sample, and a first negative example sample based on a first image and text data sample, wherein the first image and text data sample includes at least one first image sample and at least two text samples in different languages corresponding to the first image sample, and the at least two text samples in different languages corresponding to the first image sample are not translations of each other; Sample input module 1202: Used to input the first anchor sample, the first positive example sample, and the first negative example sample into the second feature extraction model, and to obtain the first sample features output by the second feature extraction model; Loss calculation module 1203: Used to obtain the first loss function value based on the first sample features, with the semantic relevance between the first image and the text in the text data sample as a constraint; Parameter update module 1204: Used to update the parameters of the second feature extraction model using the first loss function value; and Model building module 1205: Used to build a first feature extraction model based on the second feature extraction model, depending on whether the second feature extraction model satisfies the convergence conditions. The first feature extraction model is used to process input first image and text data to obtain data features of the first image and text data. After the data features of the first image and text data are processed by a task processing assembly, the processing result of the target task is output, and the target task is a classification or regression task based on image and text data.
[0224] Figure 13 is a block diagram of the configuration of a computer device 1300 shown in an exemplary embodiment of the present application. The computer device can be implemented as a server or terminal in the above-described technical application of the present application. In this embodiment, the computer device is described as a server. The computer device 1300 includes a central processing unit (CPU) 1301, a system memory 1304 including RAM (Random Access Memory) 1302 and ROM (Read-Only Memory) 1303, and a system bus 1305 connecting the system memory 1304 and the central processing unit 1301. The computer device 1300 further includes a mass storage device 1306 for storing an operating system (OS) 1309, application programs (apps) 1310 and other program modules 1311.
[0225] The mass storage device 1306 is connected to the central processing unit 1301 by a mass storage controller (not shown) connected to the system bus 1305. The mass storage device 1306 and its associated computer-readable media provide non-volatile storage to the computer equipment 1300. In other words, the mass storage device 1306 may include, for example, a computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) driver.
[0226] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include, for example, volatile and non-volatile media and movable and immovable media implemented by any method or technique for storing information such as computer-readable commands, data structures, program modules, and other data. Computer storage media include RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically-Erasable Programmable Read-Only Memory), fresh memory or other solid-state storage technologies, CD-ROM, DVD (Digital Versatile Disc) or other optical storage, magnetic tape cartridges, magnetic tapes, magnetic storage or other magnetic storage devices. Of course, as will be understood by those skilled in the art, the computer storage media are not limited to these. The system memory 1304 and mass storage device 1306 described above may be collectively referred to as memory devices.
[0227] According to various embodiments of this publication, the computer device 1300 may further be connected to a remote computer on a network via a network such as the Internet. That is, the computer device 1300 may be connected to a network by being connected to a network interface unit 1307 on the system bus 1305, or it may be connected to other types of networks or remote computer systems (not shown) using the network interface unit 1307.
[0228] The memory further includes at least one computer program which is stored in the memory, and the central processing unit 1301 executes the at least one computer program to realize all or some of the steps in the methods shown in each of the embodiments described above.
[0229] In exemplary embodiments, a computer-readable storage medium is further provided, which is used to store at least one computer program, and which is loaded and executed by a processor to accomplish all or some of the steps in the methods shown in each of the embodiments described above. For example, the computer-readable storage medium may be ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, optical data storage device, etc.
[0230] In an exemplary embodiment, a computer program product is further provided, which includes a computer program stored in a computer-readable storage medium. A processor in a computer device reads the computer program from the computer-readable storage medium and executes the computer program, thereby causing the computer device to perform all or some of the steps in the methods shown in each of the embodiments described above.
[0231] While preferred embodiments of this application have been described above, this application is not limited to these embodiments, and any modifications to this application that do not deviate from the spirit of this application fall within the technical scope of this application.
Claims
1. A method for processing image and text data, which is performed by a computer device. A step of acquiring a first image and text data, wherein the first image and text data include at least one image and at least one text; A step of performing feature extraction on the first image and text data, mapping the first image and text data to a universal visual text representation space, and obtaining data features of the first image and text data, wherein the universal visual text representation space is a feature space constructed based on the first image and text data samples, with the semantic relevance between the texts in the first image and text data samples as a constraint, and the first image and text data samples include at least one first image sample and at least two text samples in different languages corresponding to the first image sample, and the at least two text samples in different languages corresponding to the first image sample are not translations of each other; and A method comprising the step of sending the data features of the first image and text data to a task processing assembly, the task processing assembly outputting a processing result for a target task based on the data features, wherein the target task is a classification or regression task based on the image and text data.
2. The method according to claim 1, The steps of performing feature extraction on the first image and text data, mapping the first image and text data to the universal visual text representation space, and obtaining data features of the first image and text data are as follows: The steps include inputting the first image and text data into a first feature extraction model, and obtaining the data features of the first image and text data output by the first feature extraction model, The first feature extraction model is a machine learning network built on a second feature extraction model, and the second feature extraction model is a machine learning model obtained by training with machine learning based on a first image and text data samples, with the degree of semantic relevance between the text in the first image and the text data samples as a constraint.
3. The method according to claim 2, further, Steps include constructing a first anchor sample, a first positive example sample, and a first negative example sample based on the first image and text data sample; The steps include inputting the first anchor sample, the first positive example sample, and the first negative example sample into the second feature extraction model, and obtaining the first sample features output by the second feature extraction model; A step of obtaining a first loss function value based on the first sample features, with the semantic relevance between the first image and the text in the text data sample as a constraint; and A method comprising the step of updating the parameters of the second feature extraction model based on the first loss function value.
4. The method according to claim 3, The first anchor sample includes a first anchor image, a text sample, and a first anchor text sample; the first positive example sample includes at least one first positive example image, a text sample, and at least one first positive example text sample; and the first negative example sample includes at least one first negative example image, a text sample, and at least one first negative example text sample. The first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample each include one image and text pair, the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample each include the same first image, the text in the first positive example image and text sample and the text in the first anchor image and text sample each match the first image meaning, the text in the first negative example image and text sample does not match the first image meaning, the first anchor text sample includes the text in the first anchor image and text sample, the first positive example text sample includes the text in the first positive example image and text sample, and the first negative example text sample includes the text in the first negative example image and text sample. The step of inputting the first anchor sample, the first positive example sample, and the first negative example sample into the second feature extraction model and obtaining the first sample features output by the second feature extraction model is: A step of inputting the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample into the second feature extraction model to obtain sample features of the first anchor image and text sample, sample features of the first positive example image and text sample, and sample features of the first negative example image and text sample; and The step includes inputting the first anchor text sample, the first positive example text sample, and the first negative example text sample into the second feature extraction model to obtain the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample, The step of obtaining a first loss function value based on the first sample features, with the semantic relevance between the first image and the text in the text data sample as a constraint, is: A step of obtaining semantic relevance information based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, the sample features of the first negative example image and text sample, the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample, wherein the semantic relevance information is used to indicate the semantic relevance between the first image and the text in the text data sample; A step of generating a semantic relevance loss function value based on the aforementioned semantic relevance information; A step of generating a first control learning loss function value based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample; and A method comprising the step of obtaining the first loss function value based on the semantic relevance loss function value and the first control learning loss function value.
5. The method according to claim 4, The step of obtaining semantic relevance information based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, the sample features of the first negative example image and text sample, the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample is: A step of performing normalization on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample to obtain the first relevance distribution information in the semantic relevance information; and The steps include performing normalization on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample to obtain second relevance distribution information in the semantic relevance information, The step of generating a semantic relevance loss function value based on the aforementioned semantic relevance information is: A method comprising the steps of calculating the KL divergence for the first relevance distribution information and the second relevance distribution information, and obtaining the semantic relevance loss function value.
6. The method according to claim 4, further, A step of generating a second loss function value based on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample; and A method comprising the step of performing a parameter update on the second feature extraction model based on the second loss function value.
7. The method according to claim 2, further, A step of constructing a second anchor image and text sample, a second positive example image and text sample, and a second negative example image and text sample based on a second image and text data sample, wherein the second image and text data sample includes at least one second image sample and at least two text samples in different languages corresponding to the second image sample, the at least two text samples in different languages corresponding to the second image sample are translations of each other, the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample each include one image and text pair, the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample each include the same second image, the text in the second positive example image and text sample and the text in the second anchor image and text sample each match the second image semantics, and the text in the second negative example image and text sample does not match the second image semantics; The steps include inputting the second anchor image and text sample, the second positive example image and text sample, and the second negative example image and text sample into the second feature extraction model to obtain sample features of the second anchor image and text sample, sample features of the second positive example image and text sample, and sample features of the second negative example image and text sample; A step of generating a third loss function value based on the sample features of the second anchor image and text sample, the sample features of the second positive example image and text sample, and the sample features of the second negative example image and text sample; and A method comprising the step of performing a parameter update on the second feature extraction model based on the third loss function value.
8. The method according to claim 2, further, Steps to construct a first image and text sample and a translated annotation of the first image and text sample based on a third image and text data sample, wherein the third image and text data sample includes at least one third image sample and at least two text samples in different languages corresponding to the third image sample, the at least two text samples in different languages corresponding to the third image sample are translations of each other, the first image and text sample includes the third image sample and one text sample corresponding to the first image sample, and the translated annotation includes another text sample corresponding to the third image sample; The first step involves performing a masking process on the first image and the text sample portion of the text sample, inputting them into the second feature extraction model, and obtaining the first sample features and second sample features of the first image and text sample, which are output by the second feature extraction model; A step of inputting the first sample features into a first output network and obtaining a first prediction result output by the first output network, wherein the first prediction result is used to show the first image and the predicted text of the portion of the text sample in the text sample that is masked; A step of inputting the second sample features into a second output network and obtaining a second prediction result output by the second output network, wherein the second prediction result is used to indicate the predicted text of the translated text annotation; A step of generating a fourth loss function value based on the first prediction result and the first image and the portion of the text sample in the text sample that is masked; A step of generating a fifth loss function value based on the second prediction result and the translation annotation; and A method comprising the step of performing a parameter update on the second feature extraction model based on the fourth loss function value and the fifth loss function value.
9. The method according to claim 2, further, A step of obtaining a third positive example image and text sample and a third negative example image and text sample, wherein the third positive example image and text sample includes pairs of images and text that have semantic matching, and the third negative example image and text sample includes pairs of images and text that do not have semantic matching; The steps include inputting the third positive example image and text sample and the third negative example image and text sample into the second feature extraction model, and obtaining the sample features of the third positive example image and text sample and the third negative example image and text sample output by the second feature extraction model; The steps include inputting the sample features of the third positive example image and text sample, and the sample features of the third negative example image and text sample, into a third output network, and obtaining the semantic matching probability of the third positive example image and text sample and the semantic matching probability of the third negative example image and text sample, which are output by the third output network; A step of generating a sixth loss function value based on the semantic agreement probability between the third positive example image and the text sample and the semantic agreement probability between the third negative example image and the text sample; and A method comprising the step of performing a parameter update for the second feature extraction model based on the sixth loss function value.
10. A method for processing image and text data, which is performed by a computer device. A step of constructing a first anchor sample, a first positive example sample, and a first negative example sample based on a first image and a text data sample, wherein the first image and text data sample include at least one first image sample and at least two text samples in different languages corresponding to the first image sample, and the at least two text samples in different languages corresponding to the first image sample are not translations of each other; The first step involves inputting the first anchor sample, the first positive example sample, and the first negative example sample into a second feature extraction model, and obtaining the first sample features output by the second feature extraction model; A step of obtaining a first loss function value based on the first sample features, with the semantic relevance between the first image and the text in the text data sample as a constraint; A step of updating the parameters of the second feature extraction model based on the first loss function value; and A method comprising the step of constructing a first feature extraction model based on the second feature extraction model in accordance with the convergence conditions of the second feature extraction model, wherein the first feature extraction model is used to process input first image and text data to obtain data features of the first image and text data, the data features of the first image and text data are processed by a task processing assembly and after which the processing result of a target task is output, the target task is a classification or regression task based on image and text data.
11. A device for processing image and text data, A data acquisition module for acquiring a first image and text data, wherein the first image and text data include at least one image and at least one text; A feature mapping module for performing feature extraction on the first image and text data, mapping the first image and text data to a universal visual text representation space, and obtaining data features of the first image and text data, wherein the universal visual text representation space is a feature space constructed based on the first image and text data samples, with the semantic relevance between the texts in the first image and text data samples as a constraint, and the first image and text data samples include at least one first image sample and at least two text samples in different languages corresponding to the first image sample, and the at least two text samples in different languages corresponding to the first image sample are not translations of each other; and A device comprising a task processing module that transmits the data features of the first image and text data to a task processing assembly, and causes the task processing assembly to output a processing result for a target task based on the data features, wherein the target task is a classification or regression task based on the image and text data.
12. The apparatus according to claim 11, The feature mapping module is used to input the first image and text data into the first feature extraction model and to obtain the data features of the first image and text data output by the first feature extraction model. The first feature extraction model is a machine learning network built on a second feature extraction model, and the second feature extraction model is a machine learning model obtained by training it using machine learning, with the degree of semantic relevance between the first image and the text in the text data sample as a constraint, based on the first image and the text data sample.
13. The apparatus according to claim 12, further, Includes a model training module, The aforementioned model training module is Based on the first image and text data sample, a first anchor sample, a first positive example sample, and a first negative example sample are constructed; The first anchor sample, the first positive example sample, and the first negative example sample are input to the second feature extraction model, and the first sample features output by the second feature extraction model are obtained; Based on the first sample features, the first loss function value is obtained with the semantic relevance between the first image and the text in the text data sample as a constraint; and A device used to update the parameters of the second feature extraction model based on the first loss function value.
14. The apparatus according to claim 13, The first anchor sample includes a first anchor image, a text sample, and a first anchor text sample; the first positive example sample includes at least one first positive example image, a text sample, and at least one first positive example text sample; and the first negative example sample includes at least one first negative example image, a text sample, and at least one first negative example text sample. The first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample each include one image and text pair, the first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample each include the same first image, the text in the first positive example image and text sample and the text in the first anchor image and text sample each match the first image meaning, the text in the first negative example image and text sample does not match the first image meaning, the first anchor text sample includes the text in the first anchor image and text sample, the first positive example text sample includes the text in the first positive example image and text sample, and the first negative example text sample includes the text in the first negative example image and text sample. The aforementioned model training module is The first anchor image and text sample, the first positive example image and text sample, and the first negative example image and text sample are input to the second feature extraction model to obtain the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample; The first anchor text sample, the first positive example text sample, and the first negative example text sample are input into the second feature extraction model to obtain the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample; Based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, the sample features of the first negative example image and text sample, the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample, semantic relevance information is obtained, and the semantic relevance information is used to indicate the semantic relevance between the first image and the text in the text data sample; Based on the aforementioned semantic relevance information, a semantic relevance loss function value is generated; Based on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample, a first control learning loss function value is generated; and A device used to obtain the first loss function value based on the semantic relevance loss function value and the first control learning loss function value.
15. The apparatus according to claim 14, The aforementioned model training module is Normalization processing is performed on the sample features of the first anchor image and text sample, the sample features of the first positive example image and text sample, and the sample features of the first negative example image and text sample to obtain the first relevance distribution information in the semantic relevance information; Normalization processing is performed on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample to obtain the second relevance distribution information in the semantic relevance information; and A device used to calculate the KL divergence for the first and second relevance distribution information and to obtain the semantic relevance loss function value.
16. The apparatus according to claim 14, The aforementioned model training module further, Based on the sample features of the first anchor text sample, the sample features of the first positive example text sample, and the sample features of the first negative example text sample, a second loss function value is generated; and A device used to update the parameters of the second feature extraction model based on the second loss function value.
17. A device for processing image and text data, A sample construction module for constructing a first anchor sample, a first positive example sample, and a first negative example sample based on a first image and text data samples, wherein the first image and text data samples include at least one first image sample and at least two text samples in different languages corresponding to the first image sample, and the at least two text samples in different languages corresponding to the first image sample are not translations of each other; A sample input module for inputting the first anchor sample, the first positive example sample, and the first negative example sample into a second feature extraction model, and for obtaining the first sample features output by the second feature extraction model; A loss calculation module for obtaining a first loss function value based on the first sample features, with the semantic relevance between the first image and the text in the text data sample as a constraint; A parameter update module for updating the parameters of the second feature extraction model based on the first loss function value; and A device comprising a model building module for constructing a first feature extraction model based on the second feature extraction model in response to the second feature extraction model satisfying convergence conditions, wherein the first feature extraction model is used to process input first image and text data to obtain data features of the first image and text data, the data features of the first image and text data are processed by a task processing assembly, and the processing result of a target task is output, the target task being a classification or regression task based on image and text data.
18. Computer equipment, Processor; and Includes a memory connected to the aforementioned processor, The memory device stores a computer program. The processing device is a computer device configured to implement the method described in any one of claims 1 to 10 by executing a computer program.
19. A program for causing a computer to perform the method described in any one of claims 1 to 10.