Visual Language Model Acquisition and Task Processing Method, Apparatus, Device and Medium

By using pretrained images and masked text descriptions in visual language models, multiple pretraining tasks are performed, solving the problem of low accuracy caused by the difference between pretraining and fine-tuning processes, and achieving higher task processing accuracy.

CN113792113BActive Publication Date: 2025-06-20BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010762000.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-31
Publication Date
2025-06-20
Estimated Expiration
2040-07-31

AI Technical Summary

Technical Problem

In the prior art, the pre-training process of visual language models is quite different from the fine-tuning process, resulting in a low accuracy of the trained model processing tasks.

Method used

By obtaining the pretrained image and the corresponding text description, mask processing is performed to obtain the masked text description, input the initial visual language model to obtain the predicted text description, and train the initial model based on multiple pretrained tasks to obtain the pretrained visual language model.

Benefits of technology

This method reduces the difference between pre-training and fine-tuning processes and improves the accuracy of the training visual language model processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113792113B_ABST
    Figure CN113792113B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for obtaining a vision-language model, a method for processing vision-language tasks, an apparatus, a device, and a storage medium, relating to the field of artificial intelligence technology. The method includes: obtaining a pre-trained image and a text description corresponding to the pre-trained image; performing a masking process on the text description to obtain a masked text description; obtaining an initial vision-language model; inputting the pre-trained image and the masked text description into the initial vision-language model to obtain a predicted text description; and training the initial vision-language model by executing a plurality of pre-training tasks through the initial vision-language model based on the pre-trained image, the text description, the masked text description, and the predicted text description, so as to obtain a pre-trained vision-language model for processing image-text tasks. This method achieves a certain degree of improvement in the accuracy of the trained vision-language model for processing tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a method for obtaining a vision-language model, a method for processing vision-language tasks, an apparatus, a device, and a readable storage medium. Background Art

[0002] Vision and language are two basic capabilities of artificial intelligence, and the interaction between the two supports a series of unique capabilities of simulating the information processing of the human brain, such as vision-language (VL) understanding (e.g., visual question answering) and VL generation (e.g., image captioning). VL technology has good application prospects in robot vision, assisting visually impaired people, etc.

[0003] Inspired by the development of natural language pre-training technology, pre-training the VL model to improve the performance of the model in processing VL tasks has become a development trend. In related technologies, some input image / word tokens are replaced with mask (MASK) tokens as the training data for the input of the VL model, and then the VL model is pre-trained with the goal that the VL model can recover the replaced input. However, since no artificial mask input is designed when fine-tuning the VL model for processing specific downstream tasks, the difference between the pre-training process and the fine-tuning process is relatively large, resulting in poor accuracy of the finally obtained VL model.

[0004] As described above, how to improve the accuracy of the trained VL model in processing tasks has become an urgent problem to be solved.

[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and therefore it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The purpose of the present disclosure is to provide a method, an apparatus, a device, and a readable storage medium for obtaining a vision-language model, which at least to some extent overcome the problem that the accuracy of the trained VL model in processing tasks is relatively low due to the difference between the pre-training process and the fine-tuning process in related technologies.

[0007] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be partially learned through the practice of the present disclosure.

[0008] According to one aspect of the present disclosure, there is provided a method for obtaining a vision-language model, including: obtaining a pre-trained image and a text description corresponding to the pre-trained image; performing a masking process on the text description to obtain a masked text description; obtaining an initial vision-language model; inputting the pre-trained image and the masked text description into the initial vision-language model to obtain a predicted text description; and training the initial vision-language model by performing a plurality of pre-training tasks based on the pre-trained image, the text description, the masked text description, and the predicted text description through the initial vision-language model to obtain a pre-trained vision-language model for processing image-text tasks.

[0009] According to an embodiment of the present disclosure, the inputting the pre-trained image and the masked text description into the initial vision-language model to obtain a predicted text description includes: inputting the pre-trained image and the masked text description into the initial vision-language model to obtain a predicted text description distribution output by the initial vision-language model; and sampling from the predicted text description distribution to obtain the predicted text description.

[0010] According to an embodiment of the present disclosure, the initial vision-language model includes an initial sentence encoder, an initial target encoder, an initial cross-modal encoder, and an initial cross-modal decoder; the plurality of pre-training tasks include a masked language modeling task and a masked sentence generation task; the predicted text description distribution includes a first encoder predicted text description distribution and a first decoder predicted text description distribution; the predicted text description includes an encoder predicted text description and a decoder predicted text description; the inputting the pre-trained image and the masked text description into the initial vision-language model to obtain the predicted text description distribution output by the initial vision-language model includes: inputting the pre-trained image into the initial target encoder to obtain a first target encoder output; inputting the masked text description into the initial sentence encoder to obtain a first sentence encoder output; performing the masked language modeling task on the first target encoder output and the first sentence encoder output through the initial cross-modal encoder to obtain the first encoder predicted text description distribution; performing the masked sentence generation task on the first target encoder output and the first sentence encoder output through the initial cross-modal decoder to obtain the first decoder predicted text description distribution; the sampling from the predicted text description distribution to obtain the predicted text description includes: sampling from the first encoder predicted text description distribution to obtain the encoder predicted text description; and sampling from the first decoder predicted text description distribution to obtain the decoder predicted text description.

[0011] According to an embodiment of the present disclosure, the method further includes: performing a masking process on the pre-trained image to obtain a masked pre-trained image; the plurality of pre-training tasks further include a masked object classification task and an image-sentence matching task; the model performing a plurality of pre-training tasks based on the pre-trained image, the text description, the masked text description, and the predicted text description through the initial vision-language model to train the initial vision-language model includes: obtaining a first masked language modeling loss by predicting a text description distribution based on the first encoder with the text description as a label; inputting the text description into the initial sentence encoder to obtain a second sentence encoder output; inputting the masked pre-trained image into the initial object encoder to obtain a second object encoder output; performing the masked object classification task on the second object encoder output and the second sentence encoder output through the initial cross-modal encoder to obtain a first encoder predicted object distribution; obtaining a first masked object classification loss by predicting an object distribution based on the first encoder with the pre-trained image as a label; performing the image-sentence matching task according to the second sentence encoder output and the first object encoder output to obtain an image-sentence matching loss; obtaining a first masked sentence generation loss by predicting a text description distribution based on the first decoder with the text description as a label; obtaining a second-stage task loss based on the encoder predicted text description, the decoder predicted text description, and the pre-trained image through the initial sentence encoder, the initial object encoder, the initial cross-modal encoder, and the initial cross-modal decoder; obtaining a pre-training total loss function based on the first masked language modeling loss, the first masked object classification loss, the image-sentence matching loss, the first masked sentence generation loss, and the second-stage task loss; and training the initial sentence encoder, the initial object encoder, the initial cross-modal encoder, and the initial cross-modal decoder by using the pre-training total loss function.

[0012] According to an embodiment of the present disclosure, the second-stage task loss obtained by the encoder prediction text description, the decoder prediction text description, and the pre-trained image through the initial sentence encoder, the initial target encoder, the initial cross-modal encoder, and the initial cross-modal decoder includes: inputting the encoder prediction text description into the initial sentence encoder to obtain a third sentence encoder output; performing the masked language modeling task on the first target encoder output and the third sentence encoder output through the initial cross-modal encoder to obtain a second encoder prediction text description distribution; obtaining a second masked language modeling loss based on the second encoder prediction text description distribution with the text description as a label; performing the masked target classification task on the first target encoder output and the third sentence encoder output through the initial cross-modal encoder to obtain a second encoder prediction target distribution; obtaining a second masked target classification loss based on the second encoder prediction target distribution with the pre-trained image as a label; inputting the decoder prediction text description into the initial sentence encoder to obtain a fourth sentence encoder output; performing the masked sentence generation task on the first target encoder output and the fourth sentence encoder output through the initial cross-modal decoder to obtain a second decoder prediction text description distribution; obtaining a second masked sentence generation loss based on the second decoder prediction text description distribution with the text description as a label; adding the second masked language modeling loss, the second masked target classification loss, and the second masked sentence generation loss to obtain the second-stage task loss.

[0013] According to an embodiment of the present disclosure, the second-stage task loss obtained by the encoder-predicted text description, the decoder-predicted text description, and the pre-trained image through the initial sentence encoder, the initial target encoder, the initial cross-modal encoder, and the initial cross-modal decoder includes: inputting the encoder-predicted text description into the initial sentence encoder to obtain a third sentence encoder output; performing the masked sentence generation task on the first target encoder output and the third sentence encoder output through the initial cross-modal decoder to obtain a third decoder-predicted text description distribution; obtaining a third masked sentence generation loss based on the third decoder-predicted text description distribution with the text description as the label; inputting the decoder-predicted text description into the initial sentence encoder to obtain a fourth sentence encoder output; performing the masked language modeling task on the first target encoder output and the fourth sentence encoder output through the initial cross-modal encoder to obtain a third encoder-predicted text description distribution; obtaining a third masked language modeling loss based on the third encoder-predicted text description distribution with the text description as the label; performing the masked target classification task on the first target encoder output and the fourth sentence encoder output through the initial cross-modal encoder to obtain a third encoder-predicted target distribution; obtaining a third masked target classification loss based on the third encoder-predicted target distribution with the pre-trained image as the label; adding the third masked language modeling loss, the third masked target classification loss, and the third masked sentence generation loss to obtain the second-stage task loss.

[0014] According to an embodiment of the present disclosure, the second-stage task loss obtained by the encoder prediction text description, the decoder prediction text description, and the pre-trained image through the initial sentence encoder, the initial target encoder, the initial cross-modal encoder, and the initial cross-modal decoder includes: inputting the encoder prediction text description into the initial sentence encoder to obtain a third sentence encoder output; performing the masked sentence generation task on the first target encoder output and the third sentence encoder output through the initial cross-modal decoder to obtain a third decoder prediction text description distribution; obtaining a third masked sentence generation loss based on the third decoder prediction text description distribution with the text description as a label; inputting the decoder prediction text description into the initial sentence encoder to obtain a fourth sentence encoder output; performing the masked language modeling task on the first target encoder output and the fourth sentence encoder output through the initial cross-modal encoder to obtain a third encoder prediction text description distribution; obtaining a third masked language modeling loss based on the third encoder prediction text description distribution with the text description as a label; performing the masked target classification task on the first target encoder output and the fourth sentence encoder output through the initial cross-modal encoder to obtain a third encoder prediction target distribution; obtaining a third masked target classification loss based on the third encoder prediction target distribution with the pre-trained image as a label; the second-stage task loss includes the third masked language modeling loss, the third masked target classification loss, and the third masked sentence generation loss; obtaining the pre-training total loss function based on the first masked language modeling loss, the first masked target classification loss, the image sentence matching loss, the first masked sentence generation loss, and the second-stage task loss includes: obtaining an encoder-decoder switching parameter; obtaining the pre-training total loss function based on the first masked language modeling loss, the first masked target classification loss, the image sentence matching loss, the first masked sentence generation loss, the second-stage task loss, and the encoder-decoder switching parameter.

[0015] According to another aspect of the present disclosure, a visual language task processing method is provided, including: obtaining task input data of a task to be processed; processing the task input data through a pre-trained visual language model obtained by the method as described above; obtaining a task processing result output by the pre-trained visual language model.

[0016] According to another aspect of the present disclosure, there is provided a visual language model acquisition device, including: a data acquisition module for acquiring a pre-trained image and a text description corresponding to the pre-trained image; a mask processing module for performing a masking process on the text description to obtain a masked text description; a model initialization module for acquiring an initial visual language model; a first pre-training module for inputting the pre-trained image and the masked text description into the initial visual language model to obtain a predicted text description; and a second pre-training module for performing a plurality of pre-training tasks through the initial visual language model based on the pre-trained image, the text description, the masked text description, and the predicted text description to train the initial visual language model and obtain a pre-trained visual language model for processing image text tasks.

[0017] According to an embodiment of the present disclosure, the first pre-training module includes: a text description prediction module for inputting the pre-trained image and the masked text description into the initial visual language model to obtain a predicted text description distribution output by the initial visual language model; and a text description sampling module for sampling from the predicted text description distribution to obtain the predicted text description.

[0018] According to an embodiment of the present disclosure, the initial visual language model includes an initial sentence encoder, an initial target encoder, an initial cross-modal encoder, and an initial cross-modal decoder; the plurality of pre-training tasks include a masked language modeling task and a masked sentence generation task; the predicted text description distribution includes a first encoder predicted text description distribution and a first decoder predicted text description distribution; the predicted text description includes an encoder predicted text description and a decoder predicted text description; the text description prediction module includes: a target encoding module for inputting the pre-trained image into the initial target encoder to obtain a first target encoder output; a sentence encoding module for inputting the masked text description into the initial sentence encoder to obtain a first sentence encoder output; a cross-modal encoding module for performing the masked language modeling task on the first target encoder output and the first sentence encoder output through the initial cross-modal encoder to obtain the first encoder predicted text description distribution; a cross-modal decoding module for performing the masked sentence generation task on the first target encoder output and the first sentence encoder output through the initial cross-modal decoder to obtain the first decoder predicted text description distribution; and the text description sampling module is further configured to sample from the first encoder predicted text description distribution to obtain the encoder predicted text description; and sample from the first decoder predicted text description distribution to obtain the decoder predicted text description.

[0019] According to an embodiment of the present disclosure, the mask processing module is further configured to perform a masking process on the pre-trained image to obtain a masked pre-trained image; the multiple pre-trained tasks further include a masked object classification task and an image-sentence matching task; the second pre-training module includes: a masked language modeling loss calculation module, configured to obtain a first masked language modeling loss by predicting a text description distribution based on the first encoder with the text description as a label; the sentence encoding module is further configured to input the text description into the initial sentence encoder to obtain a second sentence encoder output; the object encoding module is further configured to input the masked pre-trained image into the initial object encoder to obtain a second object encoder output; the cross-modal encoding module is further configured to perform the masked object classification task on the second object encoder output and the second sentence encoder output through the initial cross-modal encoder to obtain a first encoder prediction object distribution; the second pre-training module further includes: a masked object classification loss calculation module, configured to obtain a first masked object classification loss by predicting an object distribution based on the first encoder with the pre-trained image as a label; an image-sentence matching loss calculation module, configured to perform the image-sentence matching task according to the second sentence encoder output and the first object encoder output to obtain an image-sentence matching loss; a masked sentence generation loss calculation module, configured to obtain a first masked sentence generation loss by predicting a text description distribution based on the first decoder with the text description as a label; a stage loss calculation module, configured to obtain a second stage task loss based on the encoder prediction text description, the decoder prediction text description, and the pre-trained image through the initial sentence encoder, the initial object encoder, the initial cross-modal encoder, and the initial cross-modal decoder; a total loss calculation module, configured to obtain a pre-training total loss function based on the first masked language modeling loss, the first masked object classification loss, the image-sentence matching loss, the first masked sentence generation loss, and the second stage task loss; the second pre-training module is further configured to train the initial sentence encoder, the initial object encoder, the initial cross-modal encoder, and the initial cross-modal decoder by using the pre-training total loss function.

[0020] According to an embodiment of the present disclosure, the sentence encoding module is further configured to input the encoder-predicted text description into the initial sentence encoder to obtain a third sentence encoder output; the cross-modal encoding module is further configured to perform the masked language modeling task on the first target encoder output and the third sentence encoder output through the initial cross-modal encoder to obtain a second encoder-predicted text description distribution; the masked language modeling loss calculation module is further configured to obtain a second masked language modeling loss based on the second encoder-predicted text description distribution with the text description as a label; the cross-modal encoding module is further configured to perform the masked target classification task on the first target encoder output and the third sentence encoder output through the initial cross-modal encoder to obtain a second encoder-predicted target distribution; the masked target classification loss calculation module is further configured to obtain a second masked target classification loss based on the second encoder-predicted target distribution with the pre-trained image as a label; the sentence encoding module is further configured to input the decoder-predicted text description into the initial sentence encoder to obtain a fourth sentence encoder output; the cross-modal decoding module is further configured to perform the masked sentence generation task on the first target encoder output and the fourth sentence encoder output through the initial cross-modal decoder to obtain a second decoder-predicted text description distribution; the masked sentence generation loss calculation module is further configured to obtain a second masked sentence generation loss based on the second decoder-predicted text description distribution with the text description as a label; the stage loss calculation module is further configured to add the second masked language modeling loss, the second masked target classification loss, and the second masked sentence generation loss to obtain the second stage task loss.

[0021] According to an embodiment of the present disclosure, the sentence encoding module is further configured to input the encoder predicted text description into the initial sentence encoder to obtain a third sentence encoder output; the cross-modal decoding module is further configured to perform the masked sentence generation task on the first target encoder output and the third sentence encoder output through the initial cross-modal decoder to obtain a third decoder predicted text description distribution; the masked sentence generation loss calculation module is further configured to obtain a third masked sentence generation loss based on the third decoder predicted text description distribution with the text description as a label; the sentence encoding module is further configured to input the decoder predicted text description into the initial sentence encoder to obtain a fourth sentence encoder output; the cross-modal encoding module is further configured to perform the masked language modeling task on the first target encoder output and the fourth sentence encoder output through the initial cross-modal encoder to obtain a third encoder predicted text description distribution; the masked language modeling loss calculation module is further configured to obtain a third masked language modeling loss based on the third encoder predicted text description distribution with the text description as a label; the cross-modal encoding module is further configured to perform the masked target classification task on the first target encoder output and the fourth sentence encoder output through the initial cross-modal encoder to obtain a third encoder predicted target distribution; the masked target classification loss calculation module is further configured to obtain a third masked target classification loss based on the third encoder predicted target distribution with the pre-trained image as a label; the stage loss calculation module is further configured to add the third masked language modeling loss, the third masked target classification loss, and the third masked sentence generation loss to obtain the second stage task loss.

[0022] According to an embodiment of the present disclosure, the sentence encoding module is further configured to input the encoder-predicted text description into the initial sentence encoder to obtain a third sentence encoder output; the cross-modal decoding module is further configured to perform the masked sentence generation task on the first target encoder output and the third sentence encoder output through the initial cross-modal decoder to obtain a third decoder-predicted text description distribution; the masked sentence generation loss calculation module is further configured to obtain a third masked sentence generation loss based on the third decoder-predicted text description distribution with the text description as a label; the sentence encoding module is further configured to input the decoder-predicted text description into the initial sentence encoder to obtain a fourth sentence encoder output; the cross-modal encoding module is further configured to perform the masked language modeling task on the first target encoder output and the fourth sentence encoder output through the initial cross-modal encoder to obtain a third encoder-predicted text description distribution; the masked language modeling loss calculation module is further configured to obtain a third masked language modeling loss based on the third encoder-predicted text description distribution with the text description as a label; the cross-modal encoding module is further configured to perform the masked target classification task on the first target encoder output and the fourth sentence encoder output through the initial cross-modal encoder to obtain a third encoder-predicted target distribution; the masked target classification loss calculation module is further configured to obtain a third masked target classification loss based on the third encoder-predicted target distribution with the pre-trained image as a label; the second-stage task loss includes the third masked language modeling loss, the third masked target classification loss, and the third masked sentence generation loss; the total loss calculation module is further configured to: obtain an encoder-decoder switching parameter; obtain a pre-trained total loss function based on the first masked language modeling loss, the first masked target classification loss, the image-sentence matching loss, the first masked sentence generation loss, the second-stage task loss, and the encoder-decoder switching parameter.

[0023] According to another aspect of the present disclosure, there is provided a visual language task processing device, including: a data acquisition module, configured to acquire task input data of a task to be processed; a task processing module, configured to process the task input data through a pre-trained visual language model obtained by the method as described above; and a result output module, configured to obtain a task processing result output by the pre-trained visual language model.

[0024] According to another aspect of the present disclosure, there is provided a device, including: a memory, a processor, and executable instructions stored in the memory and executable in the processor, where when the processor executes the executable instructions, the methods as described above are implemented.

[0025] According to another aspect of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, and when the executable instructions are executed by a processor, the above-mentioned any one of the methods is implemented.

[0026] The method for obtaining a vision-language model provided by the embodiments of the present disclosure obtains a predicted text description by inputting a pre-trained image and a masked text description into an initial vision-language model, and executes multiple pre-training tasks through the initial vision-language model based on the pre-trained image, text description, masked text description, and predicted text description to train the initial vision-language model, and obtains a pre-trained vision-language model to process image-text tasks, thereby achieving a certain degree of improvement in the accuracy of the trained VL model in processing tasks.

[0027] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] By referring to the accompanying drawings and describing its exemplary embodiments in detail, the above and other objects, features, and advantages of the present disclosure will become more apparent.

[0029] Figure 1 A schematic diagram showing a system structure in an embodiment of the present disclosure.

[0030] Figure 2 A flowchart showing a method for obtaining a vision-language model in an embodiment of the present disclosure.

[0031] Figure 3A A flowchart showing a method for obtaining a predicted text description according to an exemplary embodiment.

[0032] Figure 3B A block diagram showing a vision-language model architecture in an embodiment of the present disclosure.

[0033] Figure 4A A flowchart showing a method for executing a pre-training task according to an exemplary embodiment.

[0034] Figure 4B A schematic diagram showing a masked language modeling task process in an embodiment of the present disclosure.

[0035] Figure 4C A schematic diagram showing a masked object classification task process in an embodiment of the present disclosure.

[0036] Figure 4D A schematic diagram showing an image-sentence matching task process in an embodiment of the present disclosure.

[0037] Figure 4E A schematic diagram showing a masked sentence generation task process in an embodiment of the present disclosure.

[0038] Figure 5A Shows Figure 4A The schematic diagram of the processing procedure of step S4126 shown in

[0039] Figure 5B Shows the schematic diagram of the pre-training process of a visual language model in an embodiment of the present disclosure.

[0040] Figure 6A Shows Figure 4A The schematic diagram of the processing procedure of step S4126 shown in in another embodiment.

[0041] Figure 6B Shows the schematic diagram of the pre-training process of another visual language model in an embodiment of the present disclosure.

[0042] Figure 7A Shows Figure 4A The schematic diagram of the processing procedure of step S4126 shown in in yet another embodiment.

[0043] Figure 7B Shows Figure 4A The schematic diagram of the processing procedure of step S414 shown in in an embodiment.

[0044] Figure 7C Shows the schematic diagram of the pre-training process of yet another visual language model in an embodiment of the present disclosure.

[0045] Figure 8 Is a flowchart of a visual language task processing method shown according to an exemplary embodiment.

[0046] Figure 9 Shows the block diagram of a visual language model acquisition device in an embodiment of the present disclosure.

[0047] Figure 10 Shows the block diagram of another visual language model acquisition device in an embodiment of the present disclosure.

[0048] Figure 11 Shows the block diagram of a visual language task processing device in an embodiment of the present disclosure.

[0049] Figure 12 Shows the schematic diagram of the structure of an electronic device in an embodiment of the present disclosure. Detailed implementation manners

[0050] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted.

[0051] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or other methods, devices, steps, etc. may be adopted. In other cases, well-known structures, methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0052] Furthermore, terms such as "first", "second", etc. are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second" may explicitly or implicitly include one or more of such features. In the description of the present disclosure, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. The symbol " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0053] In the present disclosure, unless otherwise clearly defined and limited, terms such as "connection" should be understood in a broad sense. For example, it may be an electrical connection or may be able to communicate with each other; it may be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meaning of the above terms in the present disclosure can be understood according to specific circumstances.

[0054] As described above, language pre-training has been widely applied in the field of natural language and has excellent performance in the processing of various natural language understanding tasks. The Generative Pre-training (GPT) is one of the early achievements of language pre-training, which learns general language representations by making full use of one-way word contexts. Similar to language pre-training technology, in some related technologies, certain input image / word tokens are replaced with masked (MASK) tokens as the training data for the input of the VL model, and then the VL model is pre-trained with the goal of the VL model being able to recover the replaced input. However, since no artificial masking of the input is designed when fine-tuning the VL model for specific downstream tasks, there are significant differences between the pre-training process and the fine-tuning process. The VL pre-training in some related technologies adopts the method of a general pre-trained multi-modal (visual modality and language modality) encoder for processing downstream VL understanding tasks, but it cannot be directly applied to VL generation tasks. In some other related technologies, an encoder-decoder structure based on a single-stream input is adopted for joint pre-training of VL understanding tasks and VL generation tasks. However, the design of the single-stream structure processes the two-modal inputs through the same transformer block, and cannot fully exploit the characteristics of each modality and the inherent characteristic differences of each VL proxy task, which severely limits the generalization of the pre-trained encoder-decoder in different types of VL downstream tasks.

[0055] Therefore, the present disclosure provides a pre-training method for obtaining a vision-language model. The pre-trained image and the masked text description are input into the initial vision-language model to obtain a predicted text description. Based on the pre-trained image, the text description, the masked text description, and the predicted text description, multiple pre-training tasks are performed through the initial vision-language model to train the initial vision-language model, and a pre-trained vision-language model is obtained to process image-text tasks, which can make up for the differences between the pre-training and fine-tuning processes for VL to generate appropriate vision-language pairs.

[0056] The present disclosure also provides a two-stream decoupled encoder-decoder network design. Two encoders are used to process each modal input, and a decoupled cross-modal encoder-decoder is adopted when processing each proxy task. The target / sentence encoder independently learns the representations of each modality through internal interaction, which provides a basis for multi-modal reasoning for processing VL understanding tasks and VL generation tasks. By pre-training a general encoder-decoder structure applicable to VL understanding tasks and VL generation tasks, the characteristic differences in different modalities and different VL proxy tasks can be fully utilized.

[0057] Figure 1 An exemplary system architecture 10 is shown to which the vision-language model obtaining method or the vision-language model obtaining device of the present disclosure can be applied.

[0058] AsFigure 1 As shown in the figure, the system architecture 10 may include a terminal device 102, a network 104, a server 106, and a database 108. The terminal device 102 may be various electronic devices with a display screen and supporting input and output, including but not limited to smartphones, tablets, laptop computers, desktop computers, and so on. The network 104 is used to provide a medium for a communication link between the terminal device 102 and the server 106. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, and so on. The server 106 may be a server or a server cluster that provides various services. The database 108 may be a relational database, a non-relational database, and so on.

[0059] Users can use the terminal device 102 to interact with the server 106 and the database 108 through the network 104 to receive or send data, etc. For example, users can use the terminal device 102 to upload pictures and input questions about the pictures. The terminal device 102 transmits the pictures and questions to the server 106 through the network 104 for processing. The server 106 can also receive data from the database 108 or send data to the database 108 through the network 104. For example, the model training server 106 can obtain a large number of training images and corresponding text descriptions from the database 108, and train a vision-language model with the training images and corresponding text descriptions so that the vision-language model can be used to process the received pictures and questions. After receiving the pictures and questions, the background processing server 106 predicts the answers to the questions through the vision-language model, and feeds back information such as the question answers to the terminal device 102 through the network 104.

[0060] It should be understood that Figure 1 the numbers of terminal devices, networks, databases, and servers in

[0061] Figure 2 is a flowchart of a method for obtaining a vision-language model according to an exemplary embodiment. As shown in Figure 2 the method shown can be applied, for example, to the server side of the above system, or to the terminal device of the above system.

[0062] Referring to Figure 2 , the method 20 provided by the embodiments of the present disclosure may include the following steps.

[0063] In step S202, pre-training images and text descriptions corresponding to the pre-training images are obtained. Image-text pairs can be obtained from a large-scale image-text alignment dataset as the pre-training images and their corresponding text descriptions. The text can be a paragraph, a sentence, a phrase, and so on.

[0064] In some embodiments, for example, feature extraction processing may also be performed on the initial image in the image-text pair. Each image is represented as a set of feature image regions, and the feature vectors of the respective feature image regions are used as pre-trained images for training the vision-language model. The text in the image-text pair is tokenized to obtain a word sequence, and the word sequence is converted into word vectors as the text description for training the vision-language model.

[0065] In other embodiments, for example, the pre-trained images may further include the position features of the respective feature image regions, and the text description may further include the position features of the respective words in the initial text.

[0066] In step S204, the text description is masked to obtain a masked text description. The masked text description can be generated by randomly replacing the words in the text description with mask tokens (such as [MASK]). For example, the input word tokens can be masked with a probability of 15%, 18%, or 20% to generate the masked text description.

[0067] In step S206, an initial vision-language model is obtained. The vision-language model may be composed of an interconnected encoder and decoder. The encoder / decoder may be composed of multiple transformers. The self-attention mechanism and / or the cross-attention mechanism are applied to the entire vision-language model through the connection relationships between the transformers. The specific structure of the vision-language model can be referred to Figure 3A and Figure 3B , which will not be elaborated here.

[0068] In step S208, the pre-trained images and the masked text description are input into the initial vision-language model to obtain a predicted text description. The pre-trained images and the masked text description can be input into the initial vision-language model to obtain the distribution of the predicted text description output by the initial vision-language model, and then a predicted text description is sampled from the distribution of the predicted text description. The specific implementation method can be referred to Figure 3A , which will not be elaborated here. The predicted text description is a prediction of the masked part based on the pre-trained images and the unmasked part of the original text description. Therefore, the obtained predicted text description is a complete text describing the pre-trained images.

[0069] In step S210, multiple pre-training tasks are performed on the initial vision-language model based on the pre-training images, text descriptions, masked text descriptions, and predicted text descriptions to train the initial vision-language model, and a pre-trained vision-language model is obtained to process image-text tasks. The multiple pre-training tasks may include Masked Language Modeling (MLM), Masked Object Classification (MOC), Image-Sentence Matching (ISM), Masked Sentence Generation (MSG), etc. The specific task execution methods can be referred to Figures 4A to 7C , which will not be elaborated here. Specific tasks can be selected according to the actual application of the vision-language model to pre-train the initial vision-language model. In some embodiments, for example, if the vision-language model needs to be mainly applied to the VL generation task, it can be pre-trained through the MLM and MSG tasks; for another example, if the vision-language model needs to be applicable to both the VL understanding task and the VL generation task, the four tasks of MLM, MOC, ISM, and MSG can be used to pre-train the initial vision-language model simultaneously.

[0070] According to the method for obtaining a vision-language model provided by the embodiments of the present disclosure, a predicted text description is obtained by inputting the pre-training image and the masked text description into the initial vision-language model. Multiple pre-training tasks are performed on the initial vision-language model based on the pre-training image, text description, masked text description, and predicted text description to train the initial vision-language model, and a pre-trained vision-language model is obtained to process image-text tasks. Thus, the predicted text description output by the initial vision-language model according to the pre-training image and the masked text description can be used for pre-training, reducing the difference between the pre-training and the subsequent fine-tuning processes, and achieving a certain degree of improvement in the accuracy of the trained VL model in processing tasks.

[0071] Figure 3A is a flowchart of a method for obtaining a predicted text description shown according to an exemplary embodiment. As Figure 3A shown, the method can be applied, for example, to the server side of the above system or to the terminal device of the above system.

[0072] Refer to Figure 3A , the method 30 provided by the embodiments of the present disclosure may include the following steps.

[0073] In step S302, a pre-training image and a text description corresponding to the pre-training image are obtained. The image-sentence pairs can be obtained from a large-scale image-sentence benchmark dataset Each input image can be Expressed as a set of feature vectors (i.e., image tokens) of target image regions obtained by a target detector (e.g., Faster Region-based Convolutional Neural Network (Faster R-CNN)) where N I is the input image and is the number of target image regions, and r i represents the feature vector of the i-th target image region. For each image corresponding sentence After tokenization, we represent it as a sequence of word tokens where N S is the input sentence and is the number of words, and represents the feature vector of the j-th word. The feature vectors of the target image regions in the image tokens and the feature vectors of the words in the word token sequence can both be fused with position-aware features. The position-aware features in the image tokens can be 2D coordinate vectors of the image regions (such as vectors composed of information such as width, height, and distance from the top-left corner of the original image, etc.). The word token sequence can include 1D position information of each word, such as which index the word is in the sentence, etc.

[0074] In step S304, the text description is masked to obtain a masked text description. The words in the word token sequence can be randomly masked. For example, after randomly replacing 15% of the word tokens in the sequence with [MASK], it is converted into a word token vector as the masked text description.

[0075] In step S306, an initial vision-language model including an initial sentence encoder, an initial target encoder, an initial cross-modal encoder, and an initial cross-modal decoder is obtained. The target encoder and the sentence encoder can be implemented by a series of transformer modules, and can independently encode the input of each modality using the internal context information of each modality. The cross-modal encoder can also be composed of a series of transformer modules. The input of the cross-modal encoder is the output of the target encoder and the sentence encoder. After cross-modal encoding, multi-modal features including context information of each image / word token are obtained, fully capturing the inter-modal interaction information for cross-modal VL understanding. The cross-modal decoder can also be composed of a series of transformers. Its input is also the output of the target encoder and the sentence encoder. It collects context information from all "past" word tokens through the self-attention mechanism, and then predicts the next word for all image tokens through the cross-attention mechanism, automatically reconstructing the input sentence word by word according to the input image, simulating the word sequence generation process in the sentence, thereby supporting downstream tasks of VL generation.

[0076] In some embodiments, for example, Figure 3BShows a visual language model architecture according to an embodiment of the present disclosure, including a sentence encoder, a target encoder, a cross-modal encoder, and a cross-modal decoder.

[0077] As Figure 3B shown, the sentence encoder 312 receives the input word tokens 3122 and uses K S stacked text transformer modules 3124 to capture the intra-modal context information of the input sequence of word tokens 3122. The sequence of word tokens 3122 also includes two special tokens [CLS] and [SEP], which are used to indicate the start position and the end position of the sentence of the input word tokens respectively.

[0078] The target encoder 314 receives the input image tokens 3142 and encodes the sequence of input image tokens 3142 through the self-attention mechanism by means of K I stacked image transformer modules 3144, so as to enhance each image token representation by using the intra-modal context information extracted from the previous transformer module. The input image tokens 3142 include a special image token [IMG] (the vector of which is the average pooled target representation of all detection regions), and this token serves as the beginning of the sequence of input image tokens respectively.

[0079] Together, they are used as the multi-modal input: The input is fed into a cross-modal encoder 316 composed of a set of K E stacked transformer modules (3124, 3144, 3162) to obtain the context multi-modal features of each image / word token

[0080] Alternatively, the input can also be fed into a cross-modal decoder 318 composed of a set of stacked K D layers of transformer modules (3124, 3144, 3182, 3184). The self-attention transformer module 3182 collects context information from all "past" word tokens through the self-attention mechanism, and then predicts the next word for all image tokens through the cross-attention mechanism. The cross-attention transformer module 3184 automatically reconstructs the input sentence word by word according to the input image, simulating the process of generating the word sequence in the sentence, so as to support the downstream tasks of VL generation.

[0081] In step S308, the pre-trained image is input into the initial target encoder to obtain the first target encoder output. The first target encoder output can be the enhanced image tokens as described above Figure 3B respectively.

[0082] In step S310, the masked text description is input into the initial sentence encoder to obtain the first sentence encoder output. The first sentence encoder output can be the enhanced word tokens as described above Figure 3B among

[0083] In step S312, the first target encoder output and the first sentence encoder output perform a masked language modeling task through the initial cross-modal encoder to obtain the first encoder predicted text description distribution. The goal of the MLM task is to recover the masked part of the words based on the unmasked word tokens and image tokens. It can be driven by a classifier covering the entire vocabulary and a cross-entropy loss function (softmax), and use the context multi-modal features of the masked word tokens output in the cross-modal encoder to regenerate the original words corresponding to the masked tokens, and obtain the first encoder predicted text description distribution representing the original words corresponding to the predicted mask.

[0084] In step S314, the first target encoder output and the first sentence encoder output perform a masked sentence generation task through the initial cross-modal decoder to obtain the first decoder predicted text description distribution. According to the specific implementation of the cross-modal decoder in step S306, the MSG task is committed to teaching the cross-modal decoder how to autoregressively decode the input sentence word by word according to the input image, so that the entire visual model has the ability to generate sentences. The goal of MSG is to measure the joint negative log probability to reconstruct the sequential word tokens that depend on all "past" word tokens and the input image, and obtain the first decoder predicted text description distribution representing the original input sentence of the prediction.

[0085] In step S316, sample from the first encoder predicted text description distribution to obtain the encoder predicted text description. A predicted input sentence can be randomly obtained from the first encoder predicted text description distribution to replace the masked text description in the subsequent pre-training process.

[0086] In step S318, sample from the first decoder predicted text description distribution to obtain the decoder predicted text description. A predicted input sentence can be randomly obtained from the first decoder predicted text description distribution to replace the masked text description in the subsequent pre-training process.

[0087] During the pre-training process, the information transfer between the two modalities in the MLM task is not restricted, and the prediction of masked words / visual tokens can be performed in a single task; while MSG requires autoregressive reconstruction of the input sentence and only triggers the information transfer from the image to the text.

[0088] According to the visual language model provided by the embodiments of the present disclosure, by means of the encoders of two modalities that are independent of each other when capturing internal interaction information but can share cross-modal tasks, a basis for cross-modal reasoning is provided. And by introducing two decoupled cross-modal encoders and a cross-modal decoder to decouple the two tasks during the pre-training process, with one cross-modal encoder or cross-modal decoder responsible for each type of task, the different inherent characteristics of each VL task can be fully utilized (for example, the characteristics of MLM are reflected in the unrestricted information transfer between the two modalities, while the characteristics of MSG are reflected in the information transfer from vision to text), which promotes the generalization of the pre-trained cross-modal VL tasks to the visual language model.

[0089] According to the method for obtaining a predicted text description provided by the embodiments of the present disclosure, a standard VL pre-training task is performed using a pre-trained image and a masked text description as inputs, and an output of the predicted text description distribution is obtained. A more realistic VL pre-training is carried out by replacing the masked word tokens with reasonable alternative words sampled from this output. Since there is no involvement of artificial masked tokens during the fine-tuning process, the difference between the pre-training process and the subsequent fine-tuning process of downstream tasks can be reduced to a certain extent.

[0090] Figure 4A It is a flowchart of a method for performing a pre-training task shown according to an exemplary embodiment. As Figure 4A shown, the method can be applied, for example, to the server side of the above system, or to the terminal device of the above system.

[0091] Referring to Figure 4A , the method 40 provided by the embodiments of the present disclosure may include the following steps.

[0092] In step S402, a pre-trained image and a text description corresponding to the pre-trained image are obtained.

[0093] In step S404, the text description is subjected to a masking process to obtain a masked text description.

[0094] For the specific implementation manners of steps S402 to S404, reference may be made to steps S202 to S204 and steps S302 to S304 above, and details are not described herein again.

[0095] In step S406, the pre-trained image is subjected to a masking process to obtain a masked pre-trained image. The masking process of the pre-trained image is similar to that of the text description. The image tokens in the target image region input to the target encoder can be randomly replaced with masked tokens to obtain a masked pre-trained image as the input of the pre-trained initial visual language model.

[0096] In step S4082, the pre-trained image is input into the initial target encoder to obtain a first target encoder output.

[0097] In step S4084, the masked text description is input into the initial sentence encoder to obtain the first sentence encoder output.

[0098] In step S4102, the first target encoder output and the first sentence encoder output are used to perform a masked language modeling task through the initial cross-modal encoder, obtaining the first encoder predicted text description distribution.

[0099] In step S4104, based on the text description as the label, the first masked language modeling loss is obtained according to the first encoder predicted text description distribution. The goal of the MLM task is to recover the masked part of the word based on the non-masked word tokens and image tokens. It can be achieved through a classifier covering the entire vocabulary, driven by the cross-entropy loss function (softmax), and using the context multi-modal features of the masked word tokens output in the cross-modal encoder to regenerate the original word corresponding to the masked token. The first masked language modeling loss can be expressed as

[0100] In some embodiments, for example, Figure 4B A masked language modeling task execution process is shown, as Figure 4B As shown, the vector of the word token sequence [CLS], a, [MASK]… is input into the sentence encoder 312, the vector of the image token sequence is input into the target encoder 314, the first sentence encoder output of the sentence encoder 312 and the first target encoder output of the target encoder 314 are input into the cross-modal encoder 316, and the cross-modal encoder 316 outputs the first encoder predicted text description distribution, including the prediction of the word in the masked part, such as the prediction of "dog" for the [MASK] corresponding to the regional image of "dog" in the figure.

[0101] In step S4086, the text description is input into the initial sentence encoder to obtain the second sentence encoder output.

[0102] In step S4088, the masked pre-trained image is input into the initial target encoder to obtain the second target encoder output.

[0103] In step S4106, the second target encoder output and the second sentence encoder output are used to perform a masked target classification task through the initial cross-modal encoder, obtaining the first encoder predicted target distribution. The MOC task is that after the initial cross-modal encoder encodes the enhanced masked image tokens (i.e., the second target encoder output) and the enhanced word tokens (the second sentence encoder output) to obtain the context multi-modal features, they are input into a classifier for target classification to obtain the first encoder predicted target distribution, and the regional image replaced by the masked token [MASK] is reconstructed according to the non-masked image tokens and word tokens.

[0104] In step S4108, the pre-trained image obtains a first masked object classification loss for the label based on the prediction of the target distribution by the first encoder. The task loss of MOC can be expressed as the KL divergence that measures the matching degree between the predicted target distribution of the regional image corresponding to each masked image token and the true correct target distribution, where the true correct target distribution of the regional image can be obtained by an existing object detector. The first masked object classification loss can be expressed as

[0105] In some embodiments, for example, Figure 4C shows an execution process of a masked object classification task, as Figure 4C shown, the vector of the word token sequence [CLS], a, dog... is input into the sentence encoder 312, the vector of the image token sequence with masked tokens is input into the target encoder 314, the second sentence encoder output of the sentence encoder 312 and the second target encoder output of the target encoder 314 are input into the cross-modal encoder 316, and the cross-modal encoder 316 outputs the predicted target distribution of the first encoder through the classifier, including the word prediction of the image mask part, such as the predicted "cat" corresponding to the "cat" regional image in the mask in the figure.

[0106] In step S4110, an image-sentence matching task is performed according to the second sentence encoder output and the first target encoder output to obtain an image-sentence matching loss. The execution of the ISM task can directly measure the image-sentence similarity according to the outputs of the target encoder and the sentence encoder. For example, the second sentence encoder output and the first target encoder output are fed into an attention-based two-layer multi-layer perceptron (MLP) to calculate the image-sentence similarity. The loss function can be in the form of a triplet ranking loss, where the negative sample pairs can be generated by a multi-batch strategy. The ISM task execution method adopted in this disclosure can trigger earlier image-sentence alignment, thereby avoiding the interference of mismatched image-sentence pairs introduced by the shared cross-modal encoder on the pre-training effect of other tasks and eliminating the negative impact of mismatched image-sentence pairs on the cross-modal encoder. The image-sentence matching loss can be expressed as

[0107] In some embodiments, for example, Figure 4D shows an execution process of an image-sentence matching task, as Figure 4D shown, the vector of the word token sequence [CLS], a, dog... is input into the sentence encoder 312, the vector of the image token sequence is input into the target encoder 314, and the second sentence encoder output of the sentence encoder 312 and the first target encoder output of the target encoder 314 are matched. The goal of the training task is to fully match the two.

[0108] In step S4112, the first target encoder output and the first sentence encoder output are used to perform a masked sentence generation task through an initial cross-modal decoder, obtaining a first decoder predicted text description distribution. According to the specific implementation of the cross-modal decoder in step S306, the MSG task is dedicated to teaching the cross-modal decoder how to autoregressively decode the input sentence word by word based on the input image, enabling the entire visual model to have the ability to generate sentences.

[0109] In step S4114, a first masked sentence generation loss is obtained based on the text description as a label and the first decoder predicted text description distribution. The goal of MSG is to measure the joint negative log probability to reconstruct the sequential word tokens that depend on all "past" word tokens and the input image. The masked sentence generation loss can be expressed as

[0110] In some embodiments, for example, Figure 4E shows a masked sentence generation task execution process, as Figure 4E shown, the vector of the word token sequence [CLS], a, [MASK]… is input into the sentence encoder 312, the vector of the image token sequence is input into the target encoder 314, the first sentence encoder output of the sentence encoder 312 and the first target encoder output of the target encoder 314 are input into the cross-modal decoder 318, and the cross-modal encoder 318 outputs a first decoder predicted text description distribution, including word sequence predictions such as "a", "dog", "and", "cat", "separated", etc.

[0111] In step S4122, an encoder predicted text description is sampled from the first encoder predicted text description distribution.

[0112] In step S4124, a decoder predicted text description is sampled from the first decoder predicted text description distribution.

[0113] In step S4126, a second-stage task loss is obtained based on the encoder predicted text description, the decoder predicted text description, and the pre-trained image through the initial sentence encoder, the initial target encoder, the initial cross-modal encoder, and the initial cross-modal decoder. The calculation of the second-stage task loss can be selected according to the actual situation for the MLM, MOC, ISM, MSG tasks. The second-stage task loss is obtained by replacing the masked text description with the encoder predicted text description and the decoder predicted text description. The specific implementation can refer to FIGS. 5-7.

[0114] In step S414, a pre-training total loss function is obtained based on the first masked language modeling loss, the first masked target classification loss, the image-sentence matching loss, the first masked sentence generation loss, and the second-stage task loss. The specific implementation of obtaining the total loss function can refer to FIGS. 5-7.

[0115] In step S416, the initial sentence encoder, the initial target encoder, the initial cross-modal encoder, and the initial cross-modal decoder are trained using the pre-training total loss function.

[0116] According to the pre-training task execution method provided by the embodiments of the present disclosure, first, the target losses generated by performing four tasks of VL understanding and generation are obtained, and the masked word tokens are replaced with reasonable alternative words sampled from the output of the predicted text descriptions of the obtained MLM task and MSG task, and then the MLM, MOC, and MSG tasks are performed for the second time, making the pre-training process closer to the fine-tuning process. Therefore, the difference between the pre-training process and the subsequent downstream task fine-tuning process can be reduced to a certain extent to optimize the entire encoder-decoder structure and improve the accuracy of the entire pre-trained vision-language model.

[0117] Figure 5A Shows Figure 4A A schematic diagram of the processing process of step S4126 shown in an embodiment. As Figure 5A shown, in the embodiments of the present disclosure, the above step S4126 may further include the following steps.

[0118] In step S412602, the encoder predicted text description is input into the initial sentence encoder to obtain the third sentence encoder output. It can be sampling the encoder predicted text distribution obtained by performing the MLM task on the initial cross-modal encoder in Figure 4A , and replacing the masked word token sequence of the artificial mask token with these sampled words to obtain the encoder predicted text description, which can be represented as the non-masked word token sequence S E .

[0119] In step S412604, the first target encoder output and the third sentence encoder output perform a masked language modeling task through the initial cross-modal encoder to obtain the second encoder predicted text description distribution. The execution of the masked language modeling task can refer to steps S4104 and Figure 4B , which will not be elaborated here.

[0120] In step S412606, based on the second encoder predicted text description distribution with the text description as the label, the second masked language modeling loss is obtained. The second masked language modeling loss can be expressed as

[0121] In step S412608, the first target encoder output and the third sentence encoder output are used to perform a masked object classification task through the initial cross-modal encoder to obtain the second encoder prediction target distribution. The second target encoder output obtained by inputting the masked pre-trained image into the initial target encoder and the third sentence encoder output can also be used to perform a masked object classification task through the initial cross-modal encoder. For details, please refer to steps S4106 to S4108 and Figure 4C , which will not be elaborated here.

[0122] In step S412610, based on the second encoder prediction target distribution with the pre-trained image as the label, the second masked object classification loss is obtained. The second masked object classification loss can be expressed as

[0123] In step S412612, the decoder predicted text description is input into the initial sentence encoder to obtain the fourth sentence encoder output. It can be sampling the encoder prediction text distribution obtained by performing the MSG task on the initial cross-modal encoder in Figure 4A , and using these sampled words to replace the masked word token sequence marked by the artificial mask to obtain the decoder predicted text description, which can be expressed as the non-masked word token sequence S D .

[0124] In step S412614, the first target encoder output and the fourth sentence encoder output are used to perform a masked sentence generation task through the initial cross-modal decoder to obtain the second decoder predicted text description distribution. For performing the masked sentence generation task, please refer to steps S4112 to S4114 and Figure 4E , which will not be elaborated here.

[0125] In step S412616, based on the second decoder predicted text description distribution with the text description as the label, the second masked sentence generation loss is obtained. The second masked sentence generation loss can be expressed as

[0126] In step S412618, the second masked language modeling loss, the second masked object classification loss, and the second masked sentence generation loss are added together to obtain the second-stage task loss. At this time, the total loss function can be expressed as:

[0127]

[0128] In the formula, is the second-stage task loss; is the first-stage task loss, which can be expressed as:

[0129]

[0130] In some embodiments, for example, Figure 5B shows a schematic diagram of the pre-training process of a vision-language model in an embodiment of the present disclosure, as Figure 5B shown. First, according to Figures 4B to 4E obtain the pre-training loss of the first stage, and at the same time obtain the predicted word distribution of the masked word tokens at each position output by the cross-modal encoder 316 and the cross-modal decoder 318. Sample the predicted word distribution to obtain sampled words, and use these sampled words to replace the masked word token sequence of the artificial masked tokens, obtaining two unmasked word token sequences (S E and S D ), as Figure 5B shown, where the sampled word corresponding to [MASK] in S E is "mouse", and the sampled word corresponding to [MASK] in S D is "bird". Next, after encoding S E through the sentence encoder 312, the output of the sentence encoder 312 and the enhanced image tokens encoded by the target encoder 314 in the first-stage pre-training are input into the cross-modal encoder 316 together to perform the MLM and MOC proxy tasks. At the same time, the word tokens S D encoded by the sentence encoder 312 and the enhanced image tokens encoded by the target encoder 314 in the first-stage pre-training are input into the cross-modal decoder 318 together to perform the MSG task.

[0131] Figure 6A shows Figure 4A a schematic diagram of the processing process of step S4126 shown in an embodiment. As Figure 6A shown, in an embodiment of the present disclosure, the above step S4126 may further include the following steps. Figure 6A The difference from Figure 5A is that when calculating the second-stage task loss, the input S E and S D are interchanged.

[0132] In step S412622, input the encoder-predicted text description into the initial sentence encoder to obtain the output of the third sentence encoder.

[0133] In step S412624, perform the masked sentence generation task on the output of the first target encoder and the output of the third sentence encoder through the initial cross-modal decoder to obtain the distribution of the third decoder-predicted text description.

[0134] In step S412626, obtain the third masked sentence generation loss based on the distribution of the third decoder-predicted text description with the text description as the label. The second masked sentence generation loss can be expressed as

[0135] In step S412628, the decoder prediction text description is input into the initial sentence encoder to obtain the fourth sentence encoder output.

[0136] In step S412630, the first target encoder output and the fourth sentence encoder output perform a masked language modeling task through the initial cross-modal encoder to obtain the third encoder prediction text description distribution.

[0137] In step S412632, based on the third encoder prediction text description distribution with the text description as the label, the third masked language modeling loss is obtained. The third masked language modeling loss can be expressed as

[0138] In step S412634, the first target encoder output and the fourth sentence encoder output perform a masked target classification task through the initial cross-modal encoder to obtain the third encoder prediction target distribution.

[0139] In step S412636, based on the third encoder prediction target distribution with the pre-trained image as the label, the third masked target classification loss is obtained. The third masked target classification loss can be expressed as

[0140] In step S412638, the third masked language modeling loss, the third masked target classification loss, and the third masked sentence generation loss are added together to obtain the second-stage task loss. At this time, the total loss function can be expressed as:

[0141]

[0142] where, is the second-stage task loss.

[0143] In some embodiments, for example, Figure 6B shows a schematic diagram of the pre-training process of another vision-language model in the embodiments of the present disclosure. As Figure 6B shown, first, according to Figures 4B to 4E the pre-training loss of the first stage is obtained, and at the same time, the predicted word distributions of the masked word tokens at each position output by the cross-modal encoder 316 and the cross-modal decoder 318 are obtained. Sampling is performed on the predicted word distributions to obtain sampled words, and these sampled words are used to replace the masked word token sequences masked by the artificial mask tokens, resulting in two non-masked word token sequences (S E and S D ), as Figure 6B shown, where the sampled word corresponding to [MASK] in S E is "mouse", and S DThe sampled word corresponding to [MASK] in Chinese is "bird". Next, the sentence encoder 312 encodes S E After that, the output of the sentence encoder 312 and the enhanced image tokens encoded by the target encoder 314 in the first-stage pre-training are input into the cross-modal decoder 318 together to perform the MSG task; meanwhile, the word tokens S D encoded by the sentence encoder 312 and the enhanced image tokens encoded by the target encoder 314 in the first-stage pre-training are input into the cross-modal encoder 316 together to perform the MLM and MOC proxy tasks.

[0144] Figure 7A shows Figure 4A a schematic diagram of the processing procedure of step S4126 shown in Figure 7A in an embodiment. Figure 6A The difference from E is that in the first-stage pre-training, the output of the sentence encoder and the output of the target encoder are randomly input into the cross-modal decoder or the cross-modal encoder, that is, randomly trained through the cross-modal decoder path or the cross-modal encoder path. Therefore, the obtained non-masked word token sequence is S D .

[0145] As Figure 7A shown, in the embodiments of the present disclosure, the above step S4126 may further include the following steps.

[0146] In step S412642, when only passing through the initial cross-modal encoder path during the first-stage pre-training, the encoder-predicted text description is input into the initial sentence encoder to obtain the third sentence encoder output.

[0147] In step S412644, the first target encoder output and the third sentence encoder output perform a masked sentence generation task through the initial cross-modal decoder to obtain the third decoder-predicted text description distribution.

[0148] In step S412646, based on the third decoder-predicted text description distribution with the text description as the label, the third masked sentence generation loss is obtained.

[0149] In step S412648, when only passing through the initial cross-modal decoder path during the first-stage pre-training, the decoder-predicted text description is input into the initial sentence encoder to obtain the fourth sentence encoder output.

[0150] In step S412650, the first target encoder output and the fourth sentence encoder output perform a masked language modeling task through the initial cross-modal encoder to obtain the third encoder-predicted text description distribution.

[0151] In step S412652, a third masked language modeling loss is obtained by predicting a text description distribution based on the third encoder with the text description as a label.

[0152] In step S412654, the output of the first target encoder and the output of the fourth sentence encoder perform a masked target classification task through an initial cross-modal encoder to obtain a third encoder predicted target distribution.

[0153] In step S412656, a third masked target classification loss is obtained by predicting a target distribution based on the third encoder with the pre-trained image as a label.

[0154] The second-stage task loss includes the third masked language modeling loss, the third masked target classification loss, and the third masked sentence generation loss.

[0155] Figure 7B Shows Figure 4A A schematic diagram of the processing procedure of step S414 in an embodiment as shown in Figure 7B As shown, in the embodiment of the present disclosure, the above step S414 may further include the following steps.

[0156] In step S4142, a codec switching parameter is obtained. The codec switching parameter can be represented as the probability α of passing through the cross-modal encoder path, where α ∈ {0, 1}.

[0157] In step S4144, a pre-training total loss function is obtained based on the first masked language modeling loss, the first masked target classification loss, the image-sentence matching loss, the first masked sentence generation loss, the second-stage task loss, and the codec switching parameter. At this time, the total loss function Can be expressed as:

[0158]

[0159] In some embodiments, for example, Figure 7C Shows a schematic diagram of the pre-training process of another visual language model in the embodiment of the present disclosure. As shown in Figure 7C In the pre-training of the first stage, the output of the sentence encoder 312 and the output of the target encoder 314 are randomly input into the cross-modal decoder or the cross-modal encoder 318 to perform the MLM, MOC, and MSG tasks respectively. Compared with Figure 5B 、 Figure 6B The pre-training process, by introducing the switching between the cross-modal encoder and the decoder in the first-stage pre-training, the pre-training cost is reduced to about 1 / 2.

[0160] According to the method and visual language model provided by the embodiments of the present disclosure, pre-training experiments were conducted using a large-scale image-text description dataset, which contains approximately 3.3 million pairs of pictures-sentences automatically collected from billions of web pages. In pre-training, the pre-trained Faster-RCNN network was used to extract the feature vectors of the target region images, and at most 100 region images with a detection confidence greater than 0.2 were selected as inputs. Each input region image is represented as a 2048-dimensional vector. The number of layers of the transformers stacked in the target encoder, sentence encoder, cross-modal encoder, and cross-modal decoder is set to K I = 6, K S = 12, K E = 6, K D = 6. The size of the data volume for one training is 512, the learning rate is set to 0.0001, and the maximum number of iterations is set to 10.

[0161] Figure 8 is a flowchart of a method for processing visual language tasks shown according to an exemplary embodiment. As Figure 8 shown, the method can be applied, for example, to the server side of the above system or to the terminal device of the above system.

[0162] Referring to Figure 8 , the method 80 provided by the embodiments of the present disclosure may include the following steps.

[0163] In step S802, task input data for the task to be processed is obtained.

[0164] In step S804, the task input data is processed via the pre-trained visual language model.

[0165] In step S806, the task processing result output by the pre-trained visual language model is obtained.

[0166] Visual language downstream tasks such as visual question answering tasks, description-based image retrieval tasks, visual common sense reasoning tasks, picture description tasks, etc. can be processed through the pre-trained visual language model. The task input data and task processing results of different downstream tasks are different. The specific implementation manners for processing visual language downstream tasks will be specifically described below.

[0167] In the downstream task of visual question answering, the model predicts the answer to a given question based on an image. It is fine-tuned with 1.1 million images and questions about the images, and this task is explicitly formulated as a multi-label classification problem. First, based on the attention mechanism, a multi-modal encoder is used to generate the overall image-question feature for multi-modal output. Then, the overall image-question feature passes through a fully connected layer and then through a sigmoid function to obtain the probability distribution of 3,129 possible answers. The output answer prediction of the model can be optimized based on the cross-entropy loss. The size of the data volume for one training is 96, and the learning rate is set to 0.00005. The fine-tuning program stops after 20 fine-tuning iterations.

[0168] For the image retrieval task based on descriptions, training data is obtained from an image pool with text descriptions that describe the content of the given images. In this task, training data with 5 manually annotated sentences for each image is used, and the images are sorted according to the image-sentence similarity. The similarity is measured by performing the ISM task, and the entire model is optimized based on the triplet ranking loss. The size of the data volume for one training is 512, the learning rate is set to 0.00002, and the maximum number of iterations is set to 30.

[0169] Visual Commonsense Reasoning addresses two problems, visual question answering and answer judgment, which respectively require the model to be able to predict answers or judge the rationality of answers. During fine-tuning, we concatenate the question and each possible response (answer / judgment basis) as the input sentence, which is then input into the model together with the image. In visual question answering, we obtain the overall image-sentence feature, and then use a linear layer to predict the scores of each possible response. All predictions are trained based on the cross-entropy loss. The size of the data volume for one training is 64, the learning rate is set to 0.00002, and the maximum number of iterations is set to 20.

[0170] The image captioning task trains the model to generate natural sentences that can describe the input image. The COCO dataset is used for fine-tuning and evaluating the vision-language model. Here we use the generalized Karpathy to evaluate. For fine-tuning, we first optimize the overall model architecture based on the cross-entropy loss. The size of the data volume for one training is 16, the learning rate is set to 0.00003, and the maximum number of iterations is set to 10. The self-critical training strategy can be further used for training, and the CIDEr reward is used to achieve the optimization result at the sequence level, where the learning rate is set to 0.000005 and the maximum number of iterations is set to 30.

[0171] In addition, the effects of pre-training the cross-modal encoder and the cross-modal decoder separately and pre-training the cross-modal encoder and the cross-modal decoder simultaneously will be compared. The cross-modal encoder is pre-trained using the ISM, MLM, and MOC tasks, which is only applicable to downstream VL understanding tasks. The cross-modal decoder is pre-trained using the ISM and MSG tasks, which is only applicable to downstream VL generation tasks. When pre-training the cross-modal encoder and the cross-modal decoder separately, the final effects of the four downstream tasks are all lower than those of simultaneous pre-training.

[0172] The execution of ISM by the output of the target / sentence encoder and the output of the cross-modal encoder will be compared. After performing the ISM pre-training task on the multi-modal output of the cross-modal encoder, the performance of all tasks is even lower than that of the design without ISM pre-training, indicating that the cross-modal encoder design of ISM affects the pre-training effect of other proxy tasks by introducing mismatched image-sentence pairs in the shared cross-modal encoder, demonstrating the effectiveness of early image-sentence alignment using ISM adopted in VL pre-training.

[0173] will Figure 5B , Figure 6B , Figure 7C Three pre-training schemes will be compared with the training scheme without sampling. Figure 5B , Figure 6B , Figure 7C The performance of the three pre-trained models in downstream tasks is more excellent; compared with Figure 5B , Figure 6B , Figure 7C effectively reduces the pre-training cost of all tasks, while the performance only drops slightly.

[0174] Figure 9 is a block diagram of a visual language model acquisition device shown according to an exemplary embodiment. As Figure 9 shown, the device can be applied, for example, to the server side of the above system or to the terminal device of the above system.

[0175] Referring to Figure 9 , the device 90 provided by the embodiments of the present disclosure may include a data acquisition module 902, a mask processing module 904, a model initialization module 906, a first pre-training module 908, and a second pre-training module 910.

[0176] The data acquisition module 902 can be used to acquire pre-training images and text descriptions corresponding to the pre-training images.

[0177] The mask processing module 904 can be used to perform masking processing on the text description to obtain a masked text description.

[0178] The model initialization module 906 can be used to obtain an initial vision-language model.

[0179] The first pre-training module 908 can be used to input pre-trained images and masked text descriptions into the initial vision-language model to obtain predicted text descriptions.

[0180] The second pre-training module 910 can be used to perform multiple pre-training tasks through the initial vision-language model based on pre-trained images, text descriptions, masked text descriptions, and predicted text descriptions to train the initial vision-language model and obtain a pre-trained vision-language model for processing image-text tasks.

[0181] Figure 10 is a block diagram of a device for obtaining a vision-language model shown according to an exemplary embodiment. As Figure 10 shown, the device can be applied, for example, to the server side of the above system or to the terminal device of the above system.

[0182] Refer to Figure 10 , the device 100 provided by the embodiments of the present disclosure may include a data acquisition module 1002, a masking processing module 1004, a model initialization module 1006, a first pre-training module 1008, and a second pre-training module 1010. The first pre-training module 1008 includes a text description prediction module 10082 and a text description sampling module 10084. The text description prediction module 10082 includes a target encoding module 100822, a sentence encoding module 100824, a cross-modal encoding module 100826, and a cross-modal decoding module 100828. The second pre-training module 1010 includes a masked language modeling loss calculation module 10102, a masked target classification loss calculation module 10104, an image-sentence matching loss calculation module 10106, a masked sentence generation loss calculation module 10108, a stage loss calculation module 10110, and a total loss calculation module 10112.

[0183] The data acquisition module 1002 can be used to acquire pre-trained images and text descriptions corresponding to the pre-trained images.

[0184] The masking processing module 1004 can be used to perform a masking process on the text description to obtain a masked text description.

[0185] The masking processing module 1004 can also be used to perform a masking process on the pre-trained image to obtain a masked pre-trained image.

[0186] The model initialization module 1006 can be used to obtain an initial vision-language model, which includes an initial sentence encoder, an initial target encoder, an initial cross-modal encoder, and an initial cross-modal decoder.

[0187] The first pre-training module 1008 can be used to input the pre-trained image and the masked text description into the initial vision-language model to obtain the predicted text description.

[0188] The text description prediction module 10082 can be used to input the pre-trained image and the masked text description into the initial vision-language model, and obtain the predicted text description distribution output by the initial vision-language model. The predicted text description distribution includes the first encoder predicted text description distribution and the first decoder predicted text description distribution.

[0189] The target encoding module 100822 can be used to input the pre-trained image into the initial target encoder to obtain the first target encoder output.

[0190] The target encoding module 100822 can also be used to input the masked pre-trained image into the initial target encoder to obtain the second target encoder output.

[0191] The sentence encoding module 100824 can be used to input the masked text description into the initial sentence encoder to obtain the first sentence encoder output.

[0192] The sentence encoding module 100824 can also be used to input the text description into the initial sentence encoder to obtain the second sentence encoder output.

[0193] The sentence encoding module 100824 can also be used to input the encoder predicted text description into the initial sentence encoder to obtain the third sentence encoder output.

[0194] The sentence encoding module 100824 can also be used to input the decoder predicted text description into the initial sentence encoder to obtain the fourth sentence encoder output.

[0195] The cross-modal encoding module 100826 can be used to perform a masked language modeling task on the first target encoder output and the first sentence encoder output through the initial cross-modal encoder to obtain the first encoder predicted text description distribution.

[0196] The cross-modal encoding module 100826 can also be used to perform a masked target classification task on the second target encoder output and the second sentence encoder output through the initial cross-modal encoder to obtain the first encoder predicted target distribution.

[0197] The cross-modal encoding module 100826 can also be used to perform a masked language modeling task on the first target encoder output and the third sentence encoder output through the initial cross-modal encoder to obtain the second encoder predicted text description distribution.

[0198] The cross-modal encoding module 100826 can also be used to perform a masked target classification task on the first target encoder output and the third sentence encoder output through the initial cross-modal encoder to obtain the second encoder predicted target distribution.

[0199] The cross-modal encoding module 100826 can also be used to perform a masked language modeling task on the output of the first target encoder and the output of the fourth sentence encoder through the initial cross-modal encoder, to obtain the third encoder predicted text description distribution.

[0200] The cross-modal encoding module 100826 can also be used to perform a masked target classification task on the output of the first target encoder and the output of the fourth sentence encoder through the initial cross-modal encoder, to obtain the third encoder predicted target distribution.

[0201] The cross-modal decoding module 100828 can be used to perform a masked sentence generation task on the output of the first target encoder and the output of the first sentence encoder through the initial cross-modal decoder, to obtain the first decoder predicted text description distribution.

[0202] The cross-modal decoding module 100828 can also be used to perform a masked sentence generation task on the output of the first target encoder and the output of the fourth sentence encoder through the initial cross-modal decoder, to obtain the second decoder predicted text description distribution.

[0203] The cross-modal decoding module 100828 can also be used to perform a masked sentence generation task on the output of the first target encoder and the output of the third sentence encoder through the initial cross-modal decoder, to obtain the third decoder predicted text description distribution.

[0204] The text description sampling module 10084 can be used to sample from the predicted text description distribution to obtain the predicted text description. The predicted text description includes the encoder predicted text description and the decoder predicted text description.

[0205] The text description sampling module 10084 can also be used to sample from the first encoder predicted text description distribution to obtain the encoder predicted text description; and sample from the first decoder predicted text description distribution to obtain the decoder predicted text description.

[0206] The second pre-training module 1010 can be used to perform multiple pre-training tasks on the pre-trained image, text description, masked text description, and predicted text description through the initial vision-language model to train the initial vision-language model, to obtain the pre-trained vision-language model for processing image-text tasks. The multiple pre-training tasks include masked language modeling task, masked sentence generation task, masked target classification task, and image-sentence matching task.

[0207] The second pre-training module 1010 can also be used to train the initial sentence encoder, initial target encoder, initial cross-modal encoder, and initial cross-modal decoder using the pre-training total loss function.

[0208] The masked language modeling loss calculation module 10102 can be used to obtain a first masked language modeling loss by predicting the text description distribution based on the first encoder with the text description as a label.

[0209] The masked language modeling loss calculation module 10102 can also be used to obtain a second masked language modeling loss by predicting the text description distribution based on the second encoder with the text description as a label.

[0210] The masked language modeling loss calculation module 10102 can also be used to obtain a third masked language modeling loss by predicting the text description distribution based on the third encoder with the text description as a label.

[0211] The masked object classification loss calculation module 10104 can be used to obtain a first masked object classification loss by predicting the object distribution based on the first encoder with the pre-trained image as a label.

[0212] The masked object classification loss calculation module 10104 can also be used to obtain a second masked object classification loss by predicting the object distribution based on the second encoder with the pre-trained image as a label.

[0213] The masked object classification loss calculation module 10104 can also be used to obtain a third masked object classification loss by predicting the object distribution based on the third encoder with the pre-trained image as a label.

[0214] The image-sentence matching loss calculation module 10106 can be used to perform an image-sentence matching task according to the output of the second sentence encoder and the output of the first object encoder, and obtain an image-sentence matching loss.

[0215] The masked sentence generation loss calculation module 10108 can be used to obtain a first masked sentence generation loss by predicting the text description distribution based on the first decoder with the text description as a label.

[0216] The masked sentence generation loss calculation module 10108 can also be used to obtain a second masked sentence generation loss by predicting the text description distribution based on the second decoder with the text description as a label.

[0217] The masked sentence generation loss calculation module 10108 can also be used to obtain a third masked sentence generation loss by predicting the text description distribution based on the third decoder with the text description as a label.

[0218] The stage loss calculation module 10110 can be used to obtain a second-stage task loss based on the text description predicted by the encoder, the text description predicted by the decoder, and the pre-trained image through the initial sentence encoder, the initial object encoder, the initial cross-modal encoder, and the initial cross-modal decoder. The second-stage task loss includes the third masked language modeling loss, the third masked object classification loss, and the third masked sentence generation loss.

[0219] The stage loss calculation module 10110 can also be used to add the second masked language modeling loss, the second masked target classification loss, and the second masked sentence generation loss to obtain the second stage task loss.

[0220] The stage loss calculation module 10110 can also be used to add the third masked language modeling loss, the third masked target classification loss, and the third masked sentence generation loss to obtain the second stage task loss.

[0221] The total loss calculation module 10112 can be used to obtain the pre-training total loss function based on the first masked language modeling loss, the first masked target classification loss, the image-sentence matching loss, the first masked sentence generation loss, and the second stage task loss.

[0222] The total loss calculation module 10112 can also be used to obtain the encoder-decoder switching parameter; and obtain the pre-training total loss function based on the first masked language modeling loss, the first masked target classification loss, the image-sentence matching loss, the first masked sentence generation loss, the second stage task loss, and the encoder-decoder switching parameter.

[0223] Figure 11 is a block diagram of a visual language task processing device shown according to an exemplary embodiment. As Figure 11 shown, the device can be applied, for example, to the server side of the above system, or to the terminal device of the above system.

[0224] Refer to Figure 11 , the device 110 provided by the embodiments of the present disclosure may include a data acquisition module 1102, a task processing module 1104, and a result output module 1106.

[0225] The data acquisition module 1102 can be used to acquire the task input data of the task to be processed.

[0226] The task processing module 1104 can be used to process the task input data through the pre-trained visual language model.

[0227] The result output module 1106 can be used to obtain the task processing result output by the pre-trained visual language model. The specific implementation of each module in the device provided by the embodiments of the present disclosure can refer to the content in the above method, and will not be elaborated here.

[0228] Figure 12 shows a schematic structural diagram of an electronic device in the embodiments of the present disclosure. It should be noted that Figure 12 the device shown only takes the computer system as an example, and should not bring any limitation to the functions and usage scopes of the embodiments of the present disclosure.

[0229] As Figure 12As shown, device 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1202 or a program loaded from a storage section 1208 into a random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the device 1200 are also stored. The CPU 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0230] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1212 is also connected to the I / O interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1212 as needed so that a computer program read from it can be installed into the storage section 1208 as needed.

[0231] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by a central processing unit (CPU) 1201, the above functions defined in the system of the present disclosure are executed.

[0232] It should be noted that the computer-readable medium shown in this disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0233] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0234] The modules involved in the embodiments of the present disclosure can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a data acquisition module, a mask processing module, a model initialization module, a first pre-training module, and a second pre-training module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases. For example, the data acquisition module can also be described as "a module for acquiring training data from the connected server side".

[0235] On the other hand, the present disclosure also provides a computer-readable medium, which can be included in the device described in the above embodiments; or it can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device includes: acquiring a pre-trained image and a text description corresponding to the pre-trained image; performing a masking process on the text description to obtain a masked text description; acquiring an initial vision-language model; inputting the pre-trained image and the masked text description into the initial vision-language model to obtain a predicted text description; performing multiple pre-training tasks through the initial vision-language model based on the pre-trained image, the text description, the masked text description, and the predicted text description to train the initial vision-language model, and obtaining a pre-trained vision-language model to process image-text tasks.

[0236] The exemplary embodiments of the present disclosure have been specifically shown and described above. It should be understood that the present disclosure is not limited to the detailed structures, settings, or implementation methods described herein; on the contrary, the present disclosure is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.

Claims

1. A method for obtaining a visual language model, characterized in that, Including: Obtaining a pre-trained image and a text description corresponding to the pre-trained image; Performing a masking process on the text description to obtain a masked text description; Obtaining an initial vision-language model; Inputting the pre-trained image and the masked text description into the initial vision-language model to obtain a predicted text description, where the initial vision-language model includes an initial sentence encoder, an initial object encoder, an initial cross-modal encoder, and an initial cross-modal decoder, and the predicted text description includes an encoder-predicted text description and a decoder-predicted text description; Performing multiple pre-training tasks through the initial vision-language model based on the pre-trained image, the text description, the masked text description, and the predicted text description to train the initial vision-language model, and obtaining a pre-trained vision-language model to process image-text tasks, where the multiple pre-training tasks include a masked language modeling task and a masked sentence generation task; Inputting the pre-trained image and the masked text description into the initial vision-language model to obtain a predicted text description includes: Inputting the pre-trained image into the initial object encoder to obtain a first object encoder output; Inputting the masked text description into the initial sentence encoder to obtain a first sentence encoder output; Performing the masked language modeling task on the first object encoder output and the first sentence encoder output through the initial cross-modal encoder to obtain a first encoder-predicted text description distribution; Performing the masked sentence generation task on the first object encoder output and the first sentence encoder output through the initial cross-modal decoder to obtain a first decoder-predicted text description distribution; Sampling from the first encoder-predicted text description distribution to obtain the encoder-predicted text description; Sampling from the first decoder-predicted text description distribution to obtain the decoder-predicted text description.

2. The method according to claim 1, characterized in that, Also including: Performing a masking process on the pre-trained image to obtain a masked pre-trained image; The multiple pre-training tasks further include a masked object classification task and an image-sentence matching task; The performing multiple pre-training tasks through the initial vision-language model based on the pre-trained image, the text description, the masked text description, and the predicted text description to train the initial vision-language model includes: Obtaining a first masked language modeling loss based on the first encoder-predicted text description distribution with the text description as a label; Inputting the text description into the initial sentence encoder to obtain a second sentence encoder output; Inputting the masked pre-trained image into the initial object encoder to obtain a second object encoder output; Performing the masked object classification task on the second object encoder output and the second sentence encoder output through the initial cross-modal encoder to obtain a first encoder-predicted object distribution; Obtaining a first masked object classification loss based on the first encoder-predicted object distribution with the pre-trained image as a label; Performing the image-sentence matching task according to the second sentence encoder output and the first object encoder output to obtain an image-sentence matching loss; Obtain a first masked sentence generation loss based on the predicted text description distribution of the first decoder with the text description as a label; Obtain a second-stage task loss through the initial sentence encoder, the initial target encoder, the initial cross-modal encoder, and the initial cross-modal decoder based on the text description predicted by the encoder, the text description predicted by the decoder, and the pre-trained image; Obtain a pre-training total loss function based on the first masked language modeling loss, the first masked target classification loss, the image-sentence matching loss, the first masked sentence generation loss, and the second-stage task loss; Train the initial sentence encoder, the initial target encoder, the initial cross-modal encoder, and the initial cross-modal decoder using the pre-training total loss function.

3. The method according to claim 2, characterized in that, The obtaining of the second-stage task loss through the initial sentence encoder, the initial target encoder, the initial cross-modal encoder, and the initial cross-modal decoder based on the text description predicted by the encoder, the text description predicted by the decoder, and the pre-trained image includes: Input the text description predicted by the encoder into the initial sentence encoder to obtain a third sentence encoder output; Perform the masked language modeling task on the output of the first target encoder and the third sentence encoder through the initial cross-modal encoder to obtain a second predicted text description distribution of the encoder; Obtain a second masked language modeling loss based on the second predicted text description distribution of the encoder with the text description as a label; Perform the masked target classification task on the output of the first target encoder and the third sentence encoder through the initial cross-modal encoder to obtain a second predicted target distribution of the encoder; Obtain a second masked target classification loss based on the second predicted target distribution of the encoder with the pre-trained image as a label; Input the text description predicted by the decoder into the initial sentence encoder to obtain a fourth sentence encoder output; Perform the masked sentence generation task on the output of the first target encoder and the fourth sentence encoder through the initial cross-modal decoder to obtain a second predicted text description distribution of the decoder; Obtain a second masked sentence generation loss based on the second predicted text description distribution of the decoder with the text description as a label; Add the second masked language modeling loss, the second masked target classification loss, and the second masked sentence generation loss to obtain the second-stage task loss.

4. The method according to claim 2, wherein The obtaining of the second-stage task loss through the initial sentence encoder, the initial target encoder, the initial cross-modal encoder, and the initial cross-modal decoder based on the text description predicted by the encoder, the text description predicted by the decoder, and the pre-trained image includes: Input the text description predicted by the encoder into the initial sentence encoder to obtain a third sentence encoder output; Perform the masked sentence generation task on the output of the first target encoder and the third sentence encoder through the initial cross-modal decoder to obtain a third predicted text description distribution of the decoder; Obtain a third masked sentence generation loss based on the predicted text description distribution of the third decoder with the text description as the label; Input the predicted text description of the decoder into the initial sentence encoder to obtain a fourth sentence encoder output; Perform the masked language modeling task on the output of the first target encoder and the output of the fourth sentence encoder through the initial cross-modal encoder to obtain a third predicted text description distribution of the encoder; Obtain a third masked language modeling loss based on the predicted text description distribution of the third encoder with the text description as the label; Perform the masked target classification task on the output of the first target encoder and the output of the fourth sentence encoder through the initial cross-modal encoder to obtain a third predicted target distribution of the encoder; Obtain a third masked target classification loss based on the predicted target distribution of the third encoder with the pre-trained image as the label; Add the third masked language modeling loss, the third masked target classification loss, and the third masked sentence generation loss to obtain the second-stage task loss.

5. The method according to claim 2, wherein The obtaining of the second-stage task loss based on the predicted text description of the encoder, the predicted text description of the decoder, and the pre-trained image through the initial sentence encoder, the initial target encoder, the initial cross-modal encoder, and the initial cross-modal decoder includes: Input the predicted text description of the encoder into the initial sentence encoder to obtain a third sentence encoder output; Perform the masked sentence generation task on the output of the first target encoder and the output of the third sentence encoder through the initial cross-modal decoder to obtain a third predicted text description distribution of the decoder; Obtain a third masked sentence generation loss based on the predicted text description distribution of the third decoder with the text description as the label; Input the predicted text description of the decoder into the initial sentence encoder to obtain a fourth sentence encoder output; Perform the masked language modeling task on the output of the first target encoder and the output of the fourth sentence encoder through the initial cross-modal encoder to obtain a third predicted text description distribution of the encoder; Obtain a third masked language modeling loss based on the predicted text description distribution of the third encoder with the text description as the label; Perform the masked target classification task on the output of the first target encoder and the output of the fourth sentence encoder through the initial cross-modal encoder to obtain a third predicted target distribution of the encoder; Obtain a third masked target classification loss based on the predicted target distribution of the third encoder with the pre-trained image as the label; The second-stage task loss includes the third masked language modeling loss, the third masked target classification loss, and the third masked sentence generation loss; The obtaining of the pre-trained total loss function based on the first masked language modeling loss, the first masked target classification loss, the image-sentence matching loss, the first masked sentence generation loss, and the second-stage task loss includes: Obtain the codec switching parameter; Obtain a pre-training total loss function based on the first masked language modeling loss, the first masked object classification loss, the image-sentence matching loss, the first masked sentence generation loss, the second-stage task loss, and the encoder-decoder switching parameter.

6. A method for processing visual language tasks, wherein Including: Obtain the task input data of the task to be processed; Process the task input data through the pre-trained vision-language model obtained by the method according to any one of claims 1-5; Obtain the task processing result output by the pre-trained vision-language model.

7. An apparatus for obtaining a visual language model, wherein Including: A data acquisition module for acquiring pre-training images and text descriptions corresponding to the pre-training images; A masking processing module for masking the text description to obtain a masked text description; A model initialization module for obtaining an initial vision-language model; A first pre-training module for inputting the pre-training image and the masked text description into the initial vision-language model to obtain a predicted text description, the initial vision-language model includes an initial sentence encoder, an initial object encoder, an initial cross-modal encoder, and an initial cross-modal decoder, and the predicted text description includes an encoder predicted text description and a decoder predicted text description; The first pre-training module includes: An object encoding module for inputting the pre-training image into the initial object encoder to obtain a first object encoder output; A sentence encoding module for inputting the masked text description into the initial sentence encoder to obtain a first sentence encoder output; A cross-modal encoding module for performing a masked language modeling task on the first object encoder output and the first sentence encoder output through the initial cross-modal encoder to obtain a first encoder predicted text description distribution; A cross-modal decoding module for performing a masked sentence generation task on the first object encoder output and the first sentence encoder output through the initial cross-modal decoder to obtain a first decoder predicted text description distribution; A text description sampling module for sampling from the first encoder predicted text description distribution to obtain the encoder predicted text description; sampling from the first decoder predicted text description distribution to obtain the decoder predicted text description; A second pre-training module for training the initial vision-language model by performing multiple pre-training tasks on the pre-training image, the text description, the masked text description, and the predicted text description through the initial vision-language model to obtain a pre-trained vision-language model for processing image-text tasks, and the multiple pre-training tasks include the masked language modeling task and the masked sentence generation task.

8. An apparatus for processing visual language tasks, characterized in that, Including: A data acquisition module for acquiring the task input data of the task to be processed; A task processing module for processing the task input data through the pre-trained vision-language model obtained by the method according to any one of claims 1-5; A result output module for obtaining the task processing result output by the pre-trained vision-language model.

9. A device, comprising: A memory, a processor, and executable instructions stored in the memory and executable on the processor, wherein the processor, when executing the executable instructions, implements the method according to any one of claims 1-6.

10. A computer-readable storage medium having computer-executable instructions stored thereon, characterized in that, When the executable instructions are executed by the processor, the method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Visual context fused image description method

    CN110991515A

  • Machine translation using neural network models

    US20200034436A1