Visual language model pre-training method and device, equipment and storage medium

By segmenting the sample images and sampling from the text documents to obtain text fragments, corresponding to the object image with the highest similarity, and generating sample synthesis data, it solves the problem of high data acquisition cost when training the visual language model and improves the training efficiency of the model.

CN120123766APending Publication Date: 2025-06-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510182504.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When training a visual language model, a large amount of image-text is needed to align the data across modal modes, and the data quality directly affects the model performance, resulting in high cost of obtaining training data and low efficiency.

Method used

By segmenting the sample image, the object image is obtained; the text fragment is sampled from the sample document, including at least one entity name; the text fragment is corresponding to the object image with the highest similarity, and the sample synthesis data is generated for pre-training the visual language model.

Benefits of technology

It reduces the cost of obtaining training data, improves the efficiency of pre-trained visual language models, and can obtain high-quality image-text-to-data without large-scale annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123766A_ABST
    Figure CN120123766A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for pre-training a visual language model, equipment and a storage medium, and belongs to the technical field of computers. The method comprises the following steps: segmenting a sample object from a sample image, sampling a sample text document to obtain a text fragment, and carrying out cross-modal alignment on the text fragment and an object image comprising the sample object according to a corresponding relationship between at least one entity name included in each text fragment and the sample object. Therefore, image-text pair data with relatively high quality can be obtained without large-scale labeling. According to the scheme, the acquisition cost of the training data is greatly reduced, and the efficiency of pre-training the visual language model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and particularly to a method, apparatus, device, and storage medium for pre-training a vision-language model. Background Art

[0002] With the rapid development of artificial intelligence technologies, multi-modal large models play an increasingly important role in natural language processing and computer vision tasks. A vision-language model (VLM) can achieve cross-modal understanding and generation tasks, such as image caption generation, visual question answering, and image retrieval, by combining text information and image information. Therefore, how to train a vision-language model is a research direction.

[0003] Currently, when training a vision-language model, a vision encoder and a non-linear module are usually introduced into a pre-trained text large model to connect the text modality and the visual modality for cross-modal alignment, so as to obtain the vision-language model to be trained. Then, the above vision-language model is pre-trained with a large amount of image-text pair data.

[0004] However, the problem with the above solution is that a large amount of image-text pair data is required for training in the cross-modal alignment stage, and these data usually need to be labeled on a large scale, and the data quality directly affects the performance of the model. That is, there is a problem of high cost for obtaining training data, resulting in low efficiency of pre-training the vision-language model. Summary of the Invention

[0005] The present disclosure provides a method, apparatus, device, and storage medium for pre-training a vision-language model. This solution greatly reduces the cost of obtaining training data and improves the efficiency of pre-training the vision-language model.

[0006] According to one aspect of the embodiments of the present disclosure, a method for pre-training a vision-language model is provided, and the method includes:

[0007] Performing image segmentation on a plurality of sample images respectively to obtain a plurality of object images, each object image including a sample object;

[0008] Sampling a plurality of text segments from a sample text document, each text segment including at least one entity name;

[0009] Replacing at least one text segment in the sample text document with at least one object image to obtain sample synthetic data, and for any text segment to be replaced, the at least one entity name included in the text segment is used to describe the sample object in the object image corresponding to the text segment;

[0010] Pre-train the vision-language model based on the sample synthetic data.

[0011] According to another aspect of the embodiments of the present disclosure, there is provided an apparatus for pre-training a vision-language model, the apparatus including:

[0012] An image segmentation unit configured to perform image segmentation on a plurality of sample images respectively to obtain a plurality of object images, each object image including a sample object;

[0013] A text sampling unit configured to sample a plurality of text segments from a sample text document, each text segment including at least one entity name;

[0014] A data processing unit configured to replace at least one text segment in the sample text document with at least one object image to obtain sample synthetic data. For any text segment to be replaced, the at least one entity name included in the text segment is used to describe the sample object in the object image corresponding to the text segment;

[0015] A model training unit configured to pre-train the vision-language model based on the sample synthetic data.

[0016] In some embodiments, the image segmentation unit is configured to, for any one of the plurality of sample images, determine at least one sample object included in the sample image; for any sample object, determine a mask matrix of the sample object; based on the mask matrix of the sample object, segment from the sample image an object image including the sample object.

[0017] In some embodiments, the image segmentation unit is configured to perform image segmentation on the plurality of sample images respectively to obtain a plurality of intermediate images, each intermediate image including a sample object; perform deduplication on the plurality of intermediate images based on the semantics of the sample objects in the intermediate images; obtain a target number of intermediate images from the remaining plurality of intermediate images to obtain the plurality of object images.

[0018] In some embodiments, the image segmentation unit is configured to, for any one of the intermediate images, determine an object vector of the sample object in the intermediate image; perform deduplication on the plurality of intermediate images based on the object vectors of the sample objects in the plurality of intermediate images.

[0019] In some embodiments, the text sampling unit is configured to tokenize the sample text document; perform random sampling on the tokenized sample text document to obtain a plurality of text segments with an average length of a first length and a total length of a second length, and the second length accounts for a preset ratio of the total length of the sample text document.

[0020] In some embodiments, the data processing unit is configured to, for any one of the multiple text segments, correspond the text segment to the object image with the highest similarity among the multiple object images; and determine at least one of the text segments corresponding to the object images from the multiple text segments.

[0021] In some embodiments, the data processing unit is configured to, for any one of the multiple text segments, respectively determine the similarity between each sub-segment in the text segment and each object image through a greedy matching algorithm;

[0022] If the similarity between any sub-segment and any object image is the highest, the text segment is corresponded to the object image.

[0023] In some embodiments, the vision-language model includes multiple layers, each layer including a first feed-forward neural network and a second feed-forward neural network, where the first feed-forward neural network is used to process text data, and the second feed-forward neural network is used to process image data;

[0024] The model training unit is configured to freeze the parameters in each layer of the vision-language model except for the second feed-forward neural network; input the sample synthetic data into the vision-language model, process the text segments in the sample synthetic data based on the first feed-forward neural network in the vision-language model; process the object images in the sample synthetic data based on the second feed-forward neural network in the vision-language model; and adjust the parameters of the second feed-forward neural network according to the output result of the vision-language model.

[0025] In some embodiments, the model training unit is further configured to unfreeze all the parameters in the vision-language model; and adjust all the parameters in the vision-language model based on the sample synthetic data using a target learning rate.

[0026] According to another aspect of the embodiments of the present disclosure, an electronic device is provided, and the electronic device includes:

[0027] One or more processors;

[0028] A memory for storing program code executable by the processor;

[0029] Wherein, the processor is configured to execute the program code to implement the method for the pre-trained vision-language model as described above.

[0030] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, which enables an electronic device to execute the method for pre-training a vision-language model when instructions in the computer-readable storage medium are executed by a processor of the electronic device.

[0031] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, which implements the method for pre-training a vision-language model when the computer program is executed by a processor.

[0032] The embodiments of the present disclosure provide a solution for pre-training a vision-language model. By segmenting sample objects from sample images and sampling text segments from sample text documents, cross-modal alignment between text segments and object images including sample objects can be achieved according to the correspondence between at least one entity name included in each text segment and the sample objects, so that high-quality image-text pair data can be obtained without large-scale annotation. The above solution greatly reduces the acquisition cost of training data and improves the efficiency of pre-training a vision-language model.

[0033] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.

[0035] Figure 1 is a schematic diagram of an implementation environment of a method for pre-training a vision-language model shown according to an exemplary embodiment.

[0036] Figure 2 is a flowchart of a method for pre-training a vision-language model shown according to an exemplary embodiment.

[0037] Figure 3 is a flowchart of another method for pre-training a vision-language model shown according to an exemplary embodiment.

[0038] Figure 4 is a block diagram of an apparatus for pre-training a vision-language model shown according to an exemplary embodiment.

[0039] Figure 5 is a block diagram of an electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] To enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0041] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0042] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the images involved in the present disclosure are obtained under full authorization.

[0043] Figure 1 It is a schematic diagram of an implementation environment of a method for pre-training a vision-language model shown according to an exemplary embodiment. Refer to Figure 1 , and the implementation environment specifically includes: a terminal 101 and a server 102. The terminal 101 can be connected to the server 102 through a wireless network or a wired network.

[0044] The terminal 101 can be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, and a laptop portable computer.

[0045] The terminal 101 can generally refer to one of multiple terminals, and the terminal 101 is used as an example in this embodiment. Those skilled in the art can know that the number of the above-mentioned terminals can be more or less. For example, the above-mentioned terminals can be several, or the above-mentioned terminals can be dozens or hundreds, or more, and the present disclosure embodiments do not limit the number and device type of the terminals.

[0046] The server 102 is at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Optionally, the number of the above-mentioned servers may be more or less, and the embodiments of the present disclosure do not limit this. Of course, the server 102 may further include other functional servers to provide more comprehensive and diverse services. In some embodiments, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or, the server 102 and the terminal 101 adopt a distributed computing architecture for collaborative computing. The server 102 can be connected to the terminal 101 and other terminals through a wireless network or a wired network. Optionally, the number of the above-mentioned servers may be more or less, and the embodiments of the present disclosure do not limit this.

[0047] Figure 2 is a flowchart of a method for pre-training a vision-language model shown according to an exemplary embodiment, as Figure 2 shown, the method is executed by an electronic device and includes the following steps:

[0048] In step S201, image segmentation is performed on multiple sample images respectively to obtain multiple object images, and each object image includes a sample object.

[0049] In the embodiments of the present disclosure, image segmentation is an image processing technology that can divide an image into different regions and assign each pixel in the image to a specific category or object. After image segmentation, a specific image including a sample object separated from the original sample image is called an object image, that is, each segmented object image contains only one specific object.

[0050] In step S202, multiple text segments are sampled from the sample text document, and each text segment includes at least one entity name.

[0051] In the embodiments of the present disclosure, the sample text document is a file containing a large amount of text content, which can be a long article, a book, or a group of reports, etc. From this complete sample text document, part of the content is selected according to certain rules to form relatively short text segments. Each text segment includes at least one entity name, and an entity name is a word representing the name of a person, thing, place, organization, etc. For example, in "He is walking in the park", "park" is an entity name; in "The company released a new product", both "company" and "product" are entity names.

[0052] In step S203, at least one text segment in the sample text document is replaced with at least one object image to obtain sample composite data.

[0053] In the embodiments of the present disclosure, the object image can match the content described by the text segment. In other words, for any text segment to be replaced, at least one entity name included in the text segment is used to describe the sample object in the object image corresponding to the text segment. For example, in the text segment of a popular science article describing the appearance of a monkey, the corresponding object image is a picture of a monkey. After replacing at least one text segment in the sample text document with the corresponding object image, the new data formed is called sample composite data.

[0054] In step S204, based on the sample composite data, pre-training is performed on the vision-language model.

[0055] In the embodiments of the present disclosure, the sample composite data fuses text and images and can provide rich information for model training. A vision-language model is a model that can understand and process the relationship between visual information (such as images) and language information (such as text).

[0056] The embodiments of the present disclosure provide a solution for pre-training a vision-language model. By segmenting sample objects from sample images and sampling text segments from sample text documents, cross-modal alignment of text segments with object images including sample objects can be achieved according to the corresponding relationship between at least one entity name included in each text segment and the sample object, so that high-quality image-text pair data can be obtained without large-scale annotation. The above solution greatly reduces the acquisition cost of training data and improves the efficiency of pre-training the vision-language model.

[0057] In some embodiments, performing image segmentation on multiple sample images respectively to obtain multiple object images includes:

[0058] For any one of the multiple sample images, determining at least one sample object included in the sample image;

[0059] For any one sample object, determining the mask matrix of the sample object;

[0060] Based on the mask matrix of the sample object, segmenting from the sample image the object image including the sample object.

[0061] In the embodiments of the present disclosure, by performing the above image segmentation operation on the sample image, single sample objects in each sample image can be accurately extracted and corresponding object images can be generated, providing a clear and accurate data basis for subsequent tasks such as image analysis, recognition, and classification based on a single object, effectively avoiding multi-object interference, and improving the task processing effect.

[0062] In some embodiments, performing image segmentation on multiple sample images respectively to obtain multiple object images includes:

[0063] Perform image segmentation on multiple sample images respectively to obtain multiple intermediate images, and each intermediate image includes a sample object;

[0064] Deduplicate the multiple intermediate images based on the semantics of the sample objects in the intermediate images;

[0065] Obtain a target number of intermediate images from the remaining multiple intermediate images to obtain multiple object images.

[0066] In the embodiments of the present disclosure, by first segmenting to obtain intermediate images, then deduplicating based on semantics, and finally obtaining the target number of images, it is possible to avoid interference caused by too many similar sample objects to model training, optimize training data, and improve the training effect and efficiency of the model.

[0067] In some embodiments, deduplicating the multiple intermediate images based on the semantics of the sample objects in the intermediate images includes:

[0068] For any intermediate image, determine the object vector of the sample object in the intermediate image;

[0069] Deduplicate the multiple intermediate images based on the object vectors of the sample objects in the multiple intermediate images.

[0070] In the embodiments of the present disclosure, by encoding the sample object of the intermediate image into a vector and performing semantic deduplication based on the vector, it is possible to identify and remove similar sample objects in a quantitative and accurate manner, improve data quality, avoid interference caused by redundant information to subsequent model training, and improve the training efficiency and effect of the model.

[0071] In some embodiments, sampling multiple text fragments from the sample text document includes:

[0072] Perform word segmentation on the sample text document;

[0073] Perform random sampling on the word-segmented sample text document to obtain multiple text fragments with an average length of the first length and a total length of the second length, and the second length accounts for a preset proportion of the total length of the sample text document.

[0074] In the embodiments of the present disclosure, by sampling the sample text document in this specific manner, it is possible to obtain multiple text fragments in a random and controllable length form, which not only ensures the randomness of the fragments to cover diverse contents, but also ensures comprehensiveness and representativeness of the sampling through length limitation and proportion control, can make the training data have diversity, and thus improve the training efficiency and effect of the model.

[0075] In some embodiments, the method further includes:

[0076] For any one of the multiple text fragments, correspond the text fragment to the object image with the highest similarity among the multiple object images;

[0077] Determine at least one text segment corresponding to the object image from multiple text segments.

[0078] In the embodiments of the present disclosure, by corresponding the text segment with the object image having the highest similarity and determining the text segment with the corresponding relationship, the association between text and image can be effectively established, providing an accurate matching data basis for subsequent multi-modal data processing, assisting the model to learn the connection between text and image, and improving the processing effect of multi-modal tasks.

[0079] In some embodiments, for any text segment among multiple text segments, corresponding the text segment with one object image having the highest similarity among multiple object images includes:

[0080] For any text segment among multiple text segments, determine the similarity between each sub-segment in the text segment and each object image respectively through a greedy matching algorithm;

[0081] If the similarity between any sub-segment and any object image is the highest, then correspond the text segment with the object image.

[0082] In the embodiments of the present disclosure, by using a greedy matching algorithm to determine the similarity between the text segment and its sub-segments and the object image and making the correspondence, the connection between text and image can be efficiently established, achieving the overall matching quickly with a locally optimal choice, and providing an effective association for tasks such as multi-modal data fusion. Optionally, the similarity is a cosine similarity. The selection of the cosine similarity can quantify the similarity degree between semantic and visual features to assist in matching.

[0083] In some embodiments, the vision-language model includes multiple layers, each layer includes a first feed-forward neural network and a second feed-forward neural network, the first feed-forward neural network is used to process text data, and the second feed-forward neural network is used to process image data;

[0084] Pre-train the vision-language model based on the sample synthetic data, including:

[0085] Freeze the parameters of each layer of the vision-language model except the second feed-forward neural network;

[0086] Input the sample synthetic data into the vision-language model, and process the text segment in the sample synthetic data based on the first feed-forward neural network in the vision-language model;

[0087] Process the object image in the sample synthetic data based on the second feed-forward neural network in the vision-language model;

[0088] Adjust the parameters of the second feed-forward neural network according to the output result of the vision-language model.

[0089] In the embodiments of the present disclosure, through this pre-training method of the vision-language model based on sample synthetic data, by freezing some parameters and focusing on optimizing the second feed-forward neural network to process text segments and object images respectively, the model can effectively learn the association between text and images, quickly improve the performance of the model in multi-modal tasks, and improve the efficiency of model training.

[0090] In some embodiments, the method further includes:

[0091] Unfreeze all the parameters in the vision-language model;

[0092] Using a target learning rate, adjust all the parameters in the vision-language model based on the sample synthetic data.

[0093] In the embodiments of the present disclosure, by unfreezing all the parameters of the vision-language model and adjusting them with a target learning rate based on the sample synthetic data, the model can comprehensively learn the multi-modal data features, fully exploit the information of the sample synthetic data, and improve the performance and generalization ability of the model in vision and language fusion tasks.

[0094] The above Figure 2 shows a flowchart of a method for pre-training a vision-language model according to the present disclosure. The following further elaborates on the pre-training vision-language model solution provided by the present disclosure. Figure 3 is a flowchart of another method for pre-training a vision-language model shown according to an exemplary embodiment. Refer to Figure 3 , which is executed by an electronic device and includes the following steps:

[0095] In step S301, perform image segmentation on multiple sample images respectively to obtain multiple object images, and each object image includes a sample object.

[0096] In the embodiments of the present disclosure, the multiple sample images can be from a sample image dataset. The sample image dataset is composed of a series of images with specific meanings or uses. These images can come from various channels, such as taking real scenes through a camera, collecting pictures of specific topics from the network, or obtaining them from a professional image library, etc. For example, to study animals, pictures containing various animals are collected, and these pictures form an animal-related sample image dataset.

[0097] Image segmentation is a technology that divides a digital image into multiple image sub-regions (also called superpixels). The core purpose of image segmentation is to group pixels with similar features (such as color, texture, brightness, etc.) in the image, so that the features within each sub-region are relatively consistent, while the features between different sub-regions are significantly different, thereby separating different objects or parts in the image.

[0098] Optionally, any of the following image segmentation methods can be adopted.

[0099] Method 1: Use the Segment Anything image segmentation model open-sourced by Meta to segment each image and extract all objects in the image.

[0100] In some embodiments, each object is represented by a mask matrix. The pixels where the object is located are 1, and the other parts are 0. Multiply the mask matrix by the image pixel matrix to obtain a new image that only contains the object, which is called the object image. Correspondingly, perform image segmentation on multiple sample images respectively to obtain multiple object images, including: for any one of the multiple sample images, determine at least one sample object included in the sample image. For any sample object, determine the mask matrix of the sample object. Based on the mask matrix of the sample object, segment the object image including the sample object from the sample image. Optionally, put the extracted multiple images into an object image set, and this object image set is denoted as O = {o_i}. By performing the above image segmentation operations on the sample images, each single sample object in each sample image can be accurately extracted and the corresponding object image can be generated, providing a clear and accurate data basis for subsequent image analysis, recognition, classification and other tasks based on a single object, effectively avoiding multi-object interference, and improving the task processing effect.

[0101] Method 2: Perform image segmentation based on thresholds. Optionally, by setting one or more gray-scale thresholds, the pixels in the image are divided into different categories according to their gray-scale values. For example, in a simple black-and-white binary image, if you want to segment a white object, you can set a gray-scale threshold. The pixels greater than this threshold are classified as the object part, and those less than it are classified as the background.

[0102] Method 3: Perform image segmentation based on regions. Optionally, starting from a certain seed point in the image, according to certain similarity criteria (such as color, texture, etc.), gradually merge the adjacent pixels similar to the seed point to form a region. For example, in a forest picture, taking a pixel in a certain green area as the seed point, merge the surrounding green pixels with similar colors to segment the tree area.

[0103] Method 4: Perform image segmentation based on edges. Optionally, utilize the mutation characteristics of the pixel gray-scale values at the edges of objects in the image to determine the boundaries of objects by detecting these edges. Commonly used edge detection operators include Sobel operator, Canny operator, etc. For example, in a picture of a car, the outline of the car can be outlined through edge detection to segment the car.

[0104] Next, the object image and the sample object will be introduced.

[0105] After image segmentation, specific parts of the image are separated from the original sample image. Since these images include the sample object, they are called object images. Correspondingly, the specific target object contained in each object image is the sample object. For example, in the object image of a person obtained by the above segmentation, "person" is the sample object; in the object image of a building, "building" is the sample object. It should be noted that each object image contains only one sample object, which can make subsequent tasks such as analysis, recognition, and classification of a single object more focused and accurate.

[0106] In some embodiments, the same sample object may be included in multiple sample images. Then, there will be similar or identical sample objects in the intermediate images obtained by image segmentation. However, too many similar sample objects are not conducive to model training. Therefore, it is necessary to remove duplicates from the intermediate images. Correspondingly, image segmentation is performed on multiple sample images respectively to obtain multiple object images, including: First, image segmentation is performed on multiple sample images respectively to obtain multiple intermediate images, and each intermediate image includes a sample object. Then, based on the semantics of the sample objects in the intermediate images, duplicates are removed from the multiple intermediate images. Finally, from the remaining multiple intermediate images, the target number of intermediate images is obtained to get multiple object images. By first obtaining the intermediate images through segmentation, then removing duplicates based on semantics, and finally obtaining the target number of images, it can avoid interference caused by too many similar sample objects to model training, optimize the training data, and improve the training effect and efficiency of the model.

[0107] In some embodiments, the sample objects in the intermediate images can be encoded into vectors and then duplicate removal based on semantics is performed. Correspondingly, based on the semantics of the sample objects in the intermediate images, removing duplicates from multiple intermediate images includes: For any intermediate image, determine the object vector of the sample object in the intermediate image. Based on the object vectors of the sample objects in the multiple intermediate images, duplicates are removed from the multiple intermediate images. By encoding the sample objects in the intermediate images as vectors and performing duplicate removal based on semantics, similar sample objects can be identified and removed in a quantitative and accurate manner, improving the data quality, avoiding interference caused by redundant information to subsequent model training, and improving the efficiency and effect of model training.

[0108] For example, the object images in the above object image set O are respectively encoded into vectors by a VIT model (Vision Transformer, a deep learning model that applies the Transformer architecture to computer vision tasks). Then, a vector retrieval tool FAISS (Facebook AI Similarity Search, an efficient similarity search and dense vector clustering library) is used to perform semantic deduplication on the object images in the object image set O. It should be noted that since each object image only includes one sample object, encoding the object image is equivalent to encoding the sample object. Optionally, the above target quantity is denoted as M, and M is set to 10000000. By extracting M object images with different semantics, the training data can be made diverse, thereby improving the efficiency and effect of model training.

[0109] In step S302, a plurality of text segments are sampled from the sample text document, and each text segment includes at least one entity name.

[0110] In the embodiments of the present disclosure, the sample text document refers to a document containing a certain amount of text content. Optionally, the sample text document can be various types of text materials, such as novels, academic papers, news reports, product manuals, etc., and the embodiments of the present disclosure do not limit this. For example, an academic research report on biodiversity, which details the biological species, ecological environment, etc. in different regions, is a sample text document.

[0111] Among them, sampling refers to the process of selecting a part of the content from the entire sample text document. Common sampling methods can be selecting at fixed word intervals, such as selecting a continuous text segment every 100 words; or based on specific paragraph divisions, such as randomly selecting certain paragraphs in the document; or according to the text structure, for example, only selecting the content within a certain chapter of the document, etc. The embodiments of the present disclosure do not limit the sampling method. For example, continuing with the above-mentioned academic report on biodiversity, continuous text paragraphs can be selected as the sampling result at an interval of every 300 words.

[0112] Among them, the partial text content obtained from the sample text document through sampling is relatively independent in form and constitutes smaller text units, that is, text segments. Each text segment has a certain semantic integrity. Although it may only be a small part of the overall content of the document, it can express relatively clear information.

[0113] It should be noted that each text segment includes at least one entity name. An entity name is a part of speech used to represent specific or abstract entity names such as people, things, places, organizations, institutions, etc. For example, "person", "table", "Beijing", "company", etc. are all entity names. Stipulating that each text segment includes at least one entity name can ensure that the text segment has a certain substantial content and directivity. As an important element carrying specific information in the text, the entity name helps to clarify the subject or object involved in the text, so that the text segment can be matched with the sample object in the object image subsequently.

[0114] In some embodiments, the sample text document can be sampled by means of random sampling and limiting the sampling length. Correspondingly, multiple text segments are sampled from the sample text document, including: segmenting the sample text document. Randomly sampling the segmented sample text document to obtain multiple text segments with an average length of the first length and a total length of the second length, and the second length accounts for a preset proportion of the total length of the sample text document. By sampling the sample text document in this specific way, multiple text segments can be obtained in the form of random and controllable lengths, which not only ensures the randomness of the segments to cover diverse contents, but also ensures the comprehensiveness and representativeness of the sampling through length limitation and proportion control, can make the training data have diversity, and further improve the training efficiency and effect of the model.

[0115] For example, segment the sample text document to split it into tokens. Referring to the span corruption method of Google T5, sample spans (text segments) with a total length accounting for 30% of the document length and an average length of 5. Optionally, in order to ensure that the sampled span can correspond to the object image, a part-of-speech tagging tool can be used to tag the span to ensure that there is at least one noun in the span, where the noun represents an entity name. Optionally, an entity recognition model can also be used to identify the entity name in the span. Correspondingly, the above entity name can be directly replaced by a noun, that is, each text segment includes at least one noun.

[0116] In step S303, for any one of the multiple text segments, the text segment is corresponded to the object image with the highest similarity among the multiple object images.

[0117] In the embodiments of the present disclosure, the text segment is multiple relatively short text contents obtained from the sample text document through a specific sampling method (such as random sampling and limiting the length). Each text segment includes at least one entity name and carries a certain amount of semantic information. The object image is obtained by image segmentation of multiple sample images, and each object image includes a sample object.

[0118] Among them, the similarity is an index used to measure the matching degree between a text segment and an object image. The similarities between a text segment and multiple object images are not simply visually similar, but are determined based on the semantic information contained in both. For any one of the multiple text segments, the similarity calculation needs to be performed one by one with all the object images. After completing the similarity calculation with all the object images, find the object image with the highest similarity value from them, and establish a corresponding relationship between this text segment and this object image.

[0119] In some embodiments, the text segment and the object image can be matched by a greedy matching algorithm. Correspondingly, for any one of the multiple text segments, corresponding the text segment with the object image with the highest similarity among the multiple object images includes: for any one of the multiple text segments, respectively determining the similarity between each sub-segment in the text segment and each object image through the greedy matching algorithm. If the similarity between any sub-segment and any object image is the highest, then correspond the text segment with the object image. Among them, the sub-span in a span (text segment) refers to a smaller continuous segment contained in the span. Optionally, the similarity is the cosine similarity. By using the greedy matching algorithm to determine the similarity between the text segment and its sub-segments and the object image and making the correspondence, it can efficiently establish a connection between the text and the image, quickly achieve the overall matching with the local optimal selection, and provide an effective association for tasks such as multi-modal data fusion. The selection of the cosine similarity can quantify the similarity degree between the semantic and visual features to assist the matching.

[0120] It should be noted that, in order to avoid incorrect matching, a similarity threshold is set, and only the matching results with similarity not less than the similarity threshold are retained. That is, if the highest similarity between a text segment and an object image is still less than the similarity threshold, the text segment and the object image are not corresponded.

[0121] In step S304, determine at least one text segment corresponding to an object image from the multiple text segments.

[0122] In the embodiments of the present disclosure, through the above steps, the multiple text segments have been matched with the multiple object images in a certain way (such as calculating the similarity by the greedy matching algorithm), and each text segment has tried to find the object image with the highest similarity to it and establish a corresponding relationship. Select those text segments that have successfully established a corresponding relationship with the object image from all the multiple text segments. This means that for each text segment, check whether it has found the object image with the highest similarity to it. If it has, this text segment belongs to the range to be determined in this step. These text segments corresponding to the object images combine the text description and the corresponding image information, forming a multi-modal data form.

[0123] In step S305, at least one text segment in the sample text document is replaced with at least one object image to obtain sample synthesis data.

[0124] In the embodiment of the present disclosure, at least one text segment determined in the above steps is replaced with at least one object image to obtain sample synthesis data. Among them, for any text segment to be replaced, at least one entity name included in the text segment is used to describe the sample object in the object image corresponding to the text segment.

[0125] For example, the above at least one text segment is denoted as The sample text document is represented as d′ = [span 0 , span 1 ,..., span i ,...]. When , it means that the text segment corresponds to an object image. If , then span i = [x j ,..., x k , that is, it is still the original token sequence in the sample text document.

[0126] Through the above steps, a synthetic data format that is visually aligned at a fine-grained level can be constructed starting from the text document. And from the perspective of the discourse level, the synthetic data composed of alternating text and visual spans is still semantically coherent in the left-to-right direction, conforming to the left-to-right sequence characteristics of natural language. Training on such data is more consistent with the semantic structure learned by the language model on the text, can retain the inherent semantic structure of the LLM to the greatest extent, and at the same time achieve fine-grained cross-modal visual-to-text alignment.

[0127] In step S306, based on the sample synthesis data, the vision-language model is pre-trained.

[0128] In the embodiment of the present disclosure, the present disclosure improves the mainstream Transformer architecture to achieve lossless cross-modal transfer from text to vision. The design principle of the vision-language model is that FFN (Feedforward Neural Network) is used as the modality-specific module to store knowledge specific to one modality. The multi-head attention module is used as the modality-agnostic module to model modality-agnostic understanding and logical reasoning capabilities, that is, the general capabilities that belong to the commonalities between modalities.

[0129] The present disclosure improves the architecture of the Vision-Language Model (VLM) based on the text model by introducing an additional FFN module, called the Vision Expert module, into each Transformer Block, which is parallel to the original FFN module (Text Expert module) and does not interfere with each other. Correspondingly, the vision-language model includes multiple layers, each layer including a first feed-forward neural network and a second feed-forward neural network. The first feed-forward neural network is used to process text data, and the second feed-forward neural network is used to process image data. That is, the first feed-forward neural network is the above-mentioned text expert module, and the second feed-forward neural network is the vision expert module.

[0130] The input of the vision-language model is an input sequence with interleaved text and visual modalities, that is, the above-mentioned sample synthesis data. Then, the vision-language model maps the text segment into a vector through an embedding matrix and encodes the object image into a vector through a multimodal projector.

[0131] The vision-language model can selectively select a specific FFN for forward calculation. Among them, the object image selects the newly introduced vision expert module, that is, the second feed-forward neural network. The text segment selects the original text expert module, that is, the first feed-forward neural network.

[0132] Optionally, the model processing process is shown in the following formula.

[0133] Input: H^0 = [e_0, e_1, e_2, e_3, e_4, e_5,...]

[0134] The multi-head attention module of the L-th block: A L = MHA L (H L-1 );

[0135] The FFN of the L-th block: Among them, respectively represent the vision expert module and the text expert module of the L-th layer.

[0136] The present disclosure adopts a two-stage training strategy. In the first stage, only the second feed-forward neural network is trained, and in the second stage, full-scale training is carried out for parameter fine-tuning.

[0137] First, introduce the training in the first stage. In some embodiments, based on the sample synthetic data, pre-train the vision-language model, including: freezing the parameters in each layer of the vision-language model except for the second feed-forward neural network; inputting the sample synthetic data into the vision-language model, and processing the text segments in the sample synthetic data based on the first feed-forward neural network in the vision-language model; processing the object images in the sample synthetic data based on the second feed-forward neural network in the vision-language model; adjusting the parameters of the second feed-forward neural network according to the output result of the vision-language model. Through this pre-training method of the vision-language model based on the sample synthetic data, by freezing some parameters to focus on optimizing the second feed-forward neural network and processing the text segments and object images respectively, the model can effectively learn the association between text and image, quickly improve the performance of the model in multi-modal tasks, and improve the efficiency of model training.

[0138] Next, introduce the training in the second stage. In some embodiments, in the second stage, a smaller learning rate (e.g., 1e^{-6}) can be used to train all the parameters to prevent catastrophic forgetting of the model in the text modality during the training of all the parameters. Correspondingly, unfreeze all the parameters in the vision-language model; use the target learning rate to adjust all the parameters in the vision-language model based on the sample synthetic data. By unfreezing all the parameters of the vision-language model and adjusting them at the target learning rate based on the sample synthetic data, the model can comprehensively learn the multi-modal data features, fully exploit the information in the sample synthetic data, and improve the performance and generalization ability of the model in the vision and language fusion tasks.

[0139] The embodiments of the present disclosure provide a solution for pre-training a vision-language model. By segmenting sample objects from sample images and sampling text segments from sample text documents, the text segments can be cross-modally aligned with the object images including the sample objects according to the correspondence between at least one entity name included in each text segment and the sample objects, so that high-quality image-text pair data can be obtained without large-scale annotation. The above solution greatly reduces the acquisition cost of training data and improves the efficiency of pre-training the vision-language model.

[0140] Figure 4 is a block diagram of an apparatus for pre-training a vision-language model shown according to an exemplary embodiment. As Figure 4 shown, the apparatus includes: an image segmentation unit 401, a text sampling unit 402, a data processing unit 403, and a model training unit 404.

[0141] The image segmentation unit 401 is configured to perform image segmentation on a plurality of sample images respectively to obtain a plurality of object images, and each object image includes a sample object;

[0142] A text sampling unit 402, configured to sample multiple text segments from a sample text document, each text segment including at least one entity name;

[0143] A data processing unit 403, configured to replace at least one text segment in the sample text document with at least one object image to obtain sample composite data. For any text segment to be replaced, the at least one entity name included in the text segment is used to describe the sample object in the object image corresponding to the text segment;

[0144] A model training unit 404, configured to pre-train a vision-language model based on the sample composite data.

[0145] In some embodiments, an image segmentation unit 401 is configured to, for any one of the multiple sample images, determine at least one sample object included in the sample image; for any one sample object, determine a mask matrix of the sample object; and based on the mask matrix of the sample object, segment an object image including the sample object from the sample image.

[0146] In some embodiments, the image segmentation unit 401 is configured to perform image segmentation on the multiple sample images respectively to obtain multiple intermediate images, each intermediate image including one sample object; deduplicate the multiple intermediate images based on the semantics of the sample objects in the intermediate images; and obtain a target number of intermediate images from the remaining multiple intermediate images to obtain multiple object images.

[0147] In some embodiments, the image segmentation unit 401 is configured to, for any one intermediate image, determine an object vector of the sample object in the intermediate image; and deduplicate the multiple intermediate images based on the object vectors of the sample objects in the multiple intermediate images.

[0148] In some embodiments, the text sampling unit 402 is configured to tokenize the sample text document; randomly sample the tokenized sample text document to obtain multiple text segments with an average length of a first length and a total length of a second length, where the second length accounts for a preset ratio of the total length of the sample text document.

[0149] In some embodiments, the data processing unit 403 is configured to, for any one of the multiple text segments, correspond the text segment to the object image with the highest similarity among the multiple object images; and determine at least one text segment corresponding to an object image from the multiple text segments.

[0150] In some embodiments, the data processing unit 403 is configured to, for any one of the multiple text segments, respectively determine the similarity between each sub-segment in the text segment and each object image through a greedy matching algorithm;

[0151] If the similarity between any sub - segment and any object image is the highest, then the text segment is corresponding to the object image.

[0152] In some embodiments, the vision - language model includes multiple layers, each layer including a first feed - forward neural network for processing text data and a second feed - forward neural network for processing image data;

[0153] The model training unit 404 is configured to freeze the parameters of each layer of the vision - language model except for the second feed - forward neural network; input the sample synthetic data into the vision - language model, process the text segments in the sample synthetic data based on the first feed - forward neural network in the vision - language model; process the object images in the sample synthetic data based on the second feed - forward neural network in the vision - language model; and adjust the parameters of the second feed - forward neural network according to the output result of the vision - language model.

[0154] In some embodiments, the model training unit 404 is further configured to unfreeze all the parameters of the vision - language model; and use the target learning rate to adjust all the parameters of the vision - language model based on the sample synthetic data.

[0155] The embodiments of the present disclosure provide an apparatus for pre - training a vision - language model. By segmenting sample objects from sample images and sampling text segments from sample text documents, the text segments can be cross - modality aligned with the object images including the sample objects according to the corresponding relationship between at least one entity name included in each text segment and the sample objects, so that high - quality image - text pair data can be obtained without large - scale annotation. The above - mentioned solution greatly reduces the acquisition cost of training data and improves the efficiency of pre - training the vision - language model.

[0156] It should be noted that, for the apparatus for pre - training a vision - language model provided in the above - mentioned embodiments, only the above - mentioned division of each functional unit is used for illustration. In practical applications, the above - mentioned functions can be allocated to different functional units according to needs, that is, the internal structure of the electronic device is divided into different functional units to complete all or part of the functions described above. In addition, the apparatus for pre - training a vision - language model provided in the above - mentioned embodiments and the method embodiments of pre - training a vision - language model belong to the same concept. The specific implementation process is detailed in the method embodiments and will not be repeated here.

[0157] Regarding the apparatus for pre - training a vision - language model in the above - mentioned embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here in detail.

[0158] In the embodiments of the present disclosure, the electronic device may be a terminal or a server. When the electronic device is a terminal, the terminal is used as the execution subject to implement the technical solutions provided by the embodiments of the present disclosure; when the electronic device is a server, the server is used as the execution subject to implement the technical solutions provided by the embodiments of the present disclosure; or, the technical solutions provided by the present disclosure are implemented through the interaction between the terminal and the server. The embodiments of the present disclosure do not limit this.

[0159] Figure 5 It is a block diagram of an electronic device shown according to an exemplary embodiment. Generally, the electronic device 500 includes: a processor 501 and a memory 502.

[0160] The processor 501 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 501 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 501 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 501 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 501 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0161] The memory 502 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 502 is used to store at least one program code, and the at least one program code is used to be executed by the processor 501 to implement the method for pre-training a vision-language model provided in the method embodiments of the present disclosure.

[0162] In some embodiments, the electronic device 500 may further optionally include: a peripheral device interface 503 and at least one peripheral device. The processor 501, the memory 502, and the peripheral device interface 503 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 503 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 504, a display screen 505, a camera assembly 506, an audio circuit 507, and a power supply 508.

[0163] The peripheral device interface 503 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 501 and the memory 502. In some embodiments, the processor 501, the memory 502, and the peripheral device interface 503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 501, the memory 502, and the peripheral device interface 503 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0164] The radio frequency circuit 504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 504 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 504 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 504 can communicate with other electronic devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 504 may further include a circuit related to NFC (Near Field Communication), and the present disclosure does not limit this.

[0165] The display screen 505 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 505 is a touch display screen, the display screen 505 also has the ability to collect touch signals on or above the surface of the display screen 505. The touch signals can be input to the processor 501 as control signals for processing. At this time, the display screen 505 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 505, which is provided on the front panel of the electronic device 500; in other embodiments, there may be at least two display screens 505, which are respectively provided on different surfaces of the electronic device 500 or are in a foldable design; in still other embodiments, the display screen 505 may be a flexible display screen, which is provided on the curved surface or the folding surface of the electronic device 500. Even further, the display screen 505 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 505 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0166] The camera module 506 is used to capture images or videos. Optionally, the camera module 506 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the electronic device, and the rear camera is provided on the back of the electronic device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to implement functions such as the combination of the main camera and the depth-of-field camera to achieve the background blur function, the combination of the main camera and the wide-angle camera to achieve panoramic shooting and VR (Virtual Reality) shooting functions, or other combined shooting functions. In some embodiments, the camera module 506 may also include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. A two-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.

[0167] The audio circuit 507 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 501 for processing, or input to the radio frequency circuit 504 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 500. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 501 or the radio frequency circuit 504 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 507 may also include a headphone jack.

[0168] The power supply 508 is used to supply power to each component in the electronic device 500. The power supply 508 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 508 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.

[0169] Those skilled in the art can understand that Figure 5 the structure shown in

[0170] does not limit the electronic device 500, and may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.

[0171] A computer program product includes a computer program, and when the computer program is executed by a processor, the method for pre-training a vision-language model is implemented.

[0172] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0173] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A method for pre-training a visual language model, characterized in that: The method comprises: Performing image segmentation on a plurality of sample images respectively to obtain a plurality of object images, each object image including a sample object; Sampling a plurality of text segments from a sample text document, each text segment including at least one entity name; At least one text segment in the sample text document is replaced with at least one object image to obtain sample synthetic data, wherein for any replaced text segment, at least one entity name included in the text segment is used to describe a sample object in the object image corresponding to the text segment; Based on the sample synthetic data, the visual language model is pre-trained.

2. The method for pre-training a visual language model according to claim 1, characterized in that: The step of performing image segmentation on a plurality of sample images to obtain a plurality of object images comprises: For any sample image among the multiple sample images, determining at least one sample object included in the sample image; For any sample object, determining a mask matrix of the sample object; Based on the mask matrix of the sample object, an object image including the sample object is segmented from the sample image.

3. The method for pre-training a visual language model according to claim 1, characterized in that: The step of performing image segmentation on a plurality of sample images to obtain a plurality of object images comprises: Performing image segmentation on the multiple sample images respectively to obtain multiple intermediate images, each intermediate image including a sample object; Deduplication of the plurality of intermediate images is performed based on semantics of sample objects in the intermediate images; A target number of intermediate images are acquired from the remaining multiple intermediate images to obtain the multiple object images.

4. The method for pre-training a visual language model according to claim 3, characterized in that: The deduplication of the plurality of intermediate images based on the semantics of the sample objects in the intermediate images comprises: For any intermediate image, determining an object vector of a sample object in the intermediate image; The plurality of intermediate images are deduplicated based on object vectors of sample objects in the plurality of intermediate images.

5. The method for pre-training a visual language model according to claim 1, characterized in that: The sampling from the sample text document to obtain multiple text segments includes: Performing word segmentation on the sample text document; The sample text document after word segmentation is randomly sampled to obtain a plurality of text segments with an average length of a first length and a total length of a second length, wherein the second length accounts for a preset proportion of the total length of the sample text document.

6. The method for pre-training a visual language model according to claim 1, characterized in that: The method further comprises: For any text segment among the multiple text segments, matching the text segment with an object image having the highest similarity among the multiple object images; From the multiple text segments, the at least one text segment corresponding to the object image is determined.

7. The method for pre-training a visual language model according to claim 6, characterized in that: For any text segment among the multiple text segments, matching the text segment with an object image with the highest similarity among the multiple object images includes: For any text segment among the multiple text segments, respectively determine the similarity between each sub-segment in the text segment and each object image by using a greedy matching algorithm; If the similarity between any sub-segment and any object image is the highest, the text segment is matched with the object image.

8. The method for pre-training a visual language model according to any one of claims 1 to 7, characterized in that: The visual language model includes a plurality of layers, each layer includes a first feedforward neural network and a second feedforward neural network, the first feedforward neural network is used to process text data, and the second feedforward neural network is used to process image data; The pre-training of the visual language model based on the sample synthetic data includes: Freezing parameters in each layer of the visual language model except for the second feed-forward neural network; Inputting the sample synthesis data into the visual language model, and processing the text segments in the sample synthesis data based on the first feedforward neural network in the visual language model; Processing the object image in the sample synthetic data based on the second feedforward neural network in the visual language model; According to the output result of the visual language model, the parameters of the second feedforward neural network are adjusted.

9. The method for pre-training a visual language model according to claim 8, characterized in that: The method further comprises: Unfreeze all parameters in the visual language model; The target learning rate is used to adjust the overall parameters in the visual language model based on the sample synthesis data.

10. A device for pre-training a visual language model, characterized in that: The device comprises: An image segmentation unit is configured to perform image segmentation on the multiple sample images respectively to obtain multiple object images, each object image including a sample object; A text sampling unit is configured to sample a plurality of text segments from a sample text document, each text segment including at least one entity name; A data processing unit is configured to replace at least one text segment in the sample text document with at least one object image to obtain sample synthetic data, wherein for any replaced text segment, at least one entity name included in the text segment is used to describe a sample object in the object image corresponding to the text segment; The model training unit is configured to pre-train the visual language model based on the sample synthetic data.

11. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing program code executable by the processor; The processor is configured to execute the program code to implement the method for pre-training a visual language model as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method for pre-training a visual language model as described in any one of claims 1 to 9.

13. A computer program product, comprising a computer program, which, when executed by a processor, implements the method for pre-training a visual language model as claimed in any one of claims 1 to 9.