Comparison learning training method and device, electronic equipment and storage medium
By using a contrastive learning training method of a visual encoder with arbitrary resolution and a generative large language model, the accuracy of the visual and text features of the multimodal large model is improved, the processing problems of images with different resolutions and long text sequences are solved, and the application scenarios are expanded.
Patent Information
- Application Number
- CN202510593308.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-10-03
AI Technical Summary
Existing large multimodal models perform poorly when processing images of varying resolutions and long text sequences, resulting in insufficient training efficiency and accuracy.
A visual encoder with arbitrary resolution and a generative large language model are used as the text encoder. The accuracy of visual and text features is improved through comparative learning training methods. The cross-attention mechanism and fixed query sampling technology are combined to dynamically fuse features.
It improves the processing capabilities of large multimodal models for images of different resolutions and long texts, enhances training efficiency and accuracy, and expands application scenarios such as medical image analysis and legal document interpretation.
Smart Images

Figure CN120747657A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to computer vision, deep learning, and large models, and more particularly to contrastive learning training methods, devices, electronic devices, and storage media. Background Art
[0002] Multimodal Large Language Models (MLLMs) are models that combine the natural language processing capabilities of large language models with the ability to understand and generate data from other modalities. By integrating multiple types of input and output, such as text, images, and sound, they provide a richer and more natural interactive experience. Their core advantage lies in their ability to process and understand information from different modalities and fuse this information to complete complex tasks. Currently, multimodal large models have been widely used in various scenarios, such as autonomous driving, intelligent question-answering, and content recommendation. Multimodal large models can include multiple different modules, such as visual encoders. Summary of the Invention
[0003] The present disclosure provides a comparative learning training method, device, electronic device and storage medium.
[0004] A contrastive learning training method, comprising:
[0005] Obtain M training images of arbitrary resolution, where M is a positive integer greater than 1, and use a visual encoder to determine the target visual features of each training image;
[0006] Obtain M segments of text content, each of which corresponds to M training images, each of which is used to describe the image content of the corresponding training image. Target text features of each segment of text content are determined using a text encoder, where the text encoder is a generative large language model.
[0007] A contrastive learning loss is determined according to each target visual feature and each target text feature, and the visual encoder and the text encoder are updated according to the contrastive learning loss.
[0008] A contrastive learning training device comprises: a first processing module, a second processing module and a third processing module;
[0009] The first processing module is configured to obtain M training images of arbitrary resolution, where M is a positive integer greater than 1, and to determine target visual features of each training image using a visual encoder;
[0010] The second processing module is configured to obtain M segments of text content, wherein the M segments of text content correspond one-to-one to M training images, each segment of text content is used to describe the image content of the corresponding training image, and to determine target text features of each segment of text content using a text encoder, wherein the text encoder is a generative large language model;
[0011] The third processing module is configured to determine a contrastive learning loss based on each target visual feature and each target text feature, and update the visual encoder and the text encoder based on the contrastive learning loss.
[0012] An electronic device, comprising:
[0013] at least one processor; and
[0014] a memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method as described above.
[0016] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0017] A computer program product comprises a computer program / instruction, which implements the above method when executed by a processor.
[0018] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0020] Figure 1 This is a flow chart of the first embodiment of the contrastive learning training method disclosed in the present invention;
[0021] Figure 2 A schematic diagram of the process of generating contrastive learning loss according to the present disclosure;
[0022] Figure 3 This is a flow chart of the second embodiment of the contrastive learning training method described in the present disclosure;
[0023] Figure 4 Schematic diagram of the structure of the comparative learning training device embodiment 400 of the present disclosure;
[0024] Figure 5 FIG. 5 is a schematic block diagram of an electronic device 500 that can be used to implement an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] Furthermore, it should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " as used herein generally indicates that the associated objects are in an "or" relationship.
[0027] Figure 1 This is a flow chart of the first embodiment of the contrastive learning training method described in this disclosure. Figure 1 As shown, the following specific implementation methods are included.
[0028] In step 101, M training images of arbitrary resolution are obtained, where M is a positive integer greater than 1, and a visual encoder is used to determine the target visual features of each training image.
[0029] In step 102, M segments of text content are obtained, and the M segments of text content correspond one-to-one to M training images. Each text content is used to describe the image content of the corresponding training image. The target text features of each text content are determined respectively using a text encoder, and the text encoder is a generative large language model (LLM).
[0030] In step 103, a contrastive learning loss is determined based on each target visual feature and each target text feature, and the visual encoder and the text encoder are updated based on the contrastive learning loss.
[0031] To obtain the visual encoder, contrastive learning training can be used. Contrastive learning is a self-supervised learning method that allows the model to learn to distinguish between similar and dissimilar samples and construct meaningful feature representations.
[0032] In traditional methods, fixed-resolution images are usually used to input visual encoders. The fixed resolution usually refers to a smaller resolution, such as 224*224 or 336*336. Therefore, the trained visual encoder cannot process high-resolution images well. In addition, the Bidirectional Encoder Representation from Transformers (BERT) model is usually used as a text encoder. The longest sequence that this text encoder can process is only 78 characters. For longer text sequences, accurate text features cannot be obtained, that is, it is difficult to process longer text sequences, which affects the training effect.
[0033] By adopting the scheme described in the above method embodiment, a visual encoder suitable for image input of any resolution can be trained, which can support image input of different scales and realize image understanding of different resolutions, thereby improving the visual representation ability of the visual encoder, and further improving the accuracy of the visual features generated by the visual encoder. In addition, a generative large language model is used as a text encoder. The generative large language model has a large number of parameters and can generate, understand and reason about natural language in an end-to-end manner, and can process text sequences of up to 8k, thereby improving the long text semantic capture ability of the text encoder, that is, improving the text understanding ability, and further improving the training efficiency and training effect of comparative learning training, such as further improving the performance of the visual encoder.
[0034] In some embodiments of the present disclosure, a trained visual encoder can be connected to a large multimodal model, and the visual encoder can be used to generate corresponding visual features for the image to be processed input into the large multimodal model.
[0035] The multimodal large model may include a visual encoder and a language decoder, etc. The visual encoder trained in the manner described in the present disclosure can be applied to the multimodal large model, that is, as a visual encoder in the multimodal large model, thereby improving the visual representation ability of the multimodal large model. At the same time, since it supports image input of arbitrary resolution, the multimodal large model can directly process images of various scales without the need for image scaling, thereby reducing the loss of image information, such as more completely retaining the detail information of high-resolution images, thereby improving the accuracy of the processing results of the multimodal large model. Accordingly, the application scenarios of the multimodal large model are further expanded, such as medical image analysis, legal document interpretation, industrial drawing understanding and other different scenarios, and has wide applicability.
[0036] In addition, in some embodiments of the present disclosure, for the input search text, a trained text encoder can be used to generate corresponding text features, and the text features can be used to determine candidate images that match the search text from the candidate images to be retrieved.
[0037] That is, the text encoder trained in accordance with the method described in this disclosure can be applied to scenarios such as image-text retrieval. The text encoder can be used to encode the retrieval text input by the user to obtain the text features of the retrieval text. The candidate image corresponding to the text feature can then be determined from each candidate image. For example, each candidate image can correspond to its own image feature, and the candidate image corresponding to the text feature can be determined by calculating the similarity between the text feature and each image feature. With the help of the text encoder described in this disclosure, the accuracy of the obtained text features can be improved, and the accuracy of the retrieval results can be correspondingly improved.
[0038] In practical applications, multiple rounds of training can be performed until the end conditions are met. During each round of training, Figure 1 Process as shown.
[0039] First, M training images of any resolution can be obtained, and the target visual features of each training image can be determined using a visual encoder. The specific value of M can be determined based on actual needs. There is no limit on the resolution of each training image.
[0040] In some embodiments of the present disclosure, in order to obtain the target visual features, the original visual features of each training image can be first generated by a visual encoder, and then the original visual features can be enhanced by a cross-attention mechanism to obtain the target visual features of each training image.
[0041] Assume that the value of M is 10 (the number is only used for example). For the convenience of expression, the 10 training images are respectively referred to as training images 1 to training images 10. Then, the visual encoder can be used to generate the visual features of training images 1 to training images 10 respectively. For the convenience of distinction, they can be called original visual features. Then, the 10 original visual features can be enhanced separately through the cross-attention mechanism to obtain the target visual features of training images 1 to training images 10.
[0042] In some embodiments of the present disclosure, for any training image, the original visual features of the training image can be determined as a key and a value, respectively, and the N obtained queries can be used to perform cross-attention calculations with the key and the value, respectively, to obtain N calculation results, where N is a positive integer greater than 1. The average of the N calculation results can then be obtained, and the average can be determined as the target visual feature of the training image.
[0043] The specific value of N can be determined according to actual needs, such as 16. Query 1 to query 16 can all be 256-dimensional vectors. Taking training image 1 as an example, the original visual features of training image 1 can be used as key and value, and 16 queries can be used to perform cross-attention calculations with key and value respectively. The core of the cross-attention mechanism lies in the interaction of query, key, and value. Taking query 1 as an example, the similarity score between query 1 and key can be calculated, and then the similarity score can be normalized to obtain the attention weight. The attention weight can then be weighted and summed with the value to obtain the calculation result corresponding to query 1. In the same way, the calculation results corresponding to other queries can be obtained respectively. In this way, a total of 16 calculation results can be obtained, and the average of the 16 calculation results can be obtained, and the average can be determined as the target visual feature of training image 1.
[0044] Since the training images may be of any resolution, the solution described in the present disclosure proposes the above-mentioned image visual representation method based on fixed query sampling, which can combine the cross-attention mechanism and fixed query sampling technology to achieve dynamic feature fusion, thereby obtaining enhanced target visual features, thereby laying a good foundation for subsequent processing.
[0045] In addition, in some embodiments of the present disclosure, after determining the contrastive learning loss, N queries can also be updated according to the contrastive learning loss, and the initial values of the N queries can all be obtained by random initialization. Accordingly, when using the obtained N queries and key and value to perform cross-attention calculations respectively, the latest N queries can be used to perform the cross-attention calculations respectively.
[0046] That is to say, the initial value of each query can be obtained by random initialization, and then each query can be continuously updated during the training process. Compared with the method in which the value of each query is pre-set and no longer changed, the method described in the present disclosure can achieve more accurate assignment of each query, thereby further improving the accuracy of the obtained target visual features.
[0047] In addition, a text sequence can be obtained. The text sequence can include M text segments, which can correspond one-to-one to M training images. Each text segment is used to describe the image content of the corresponding training image. The text encoder can then determine the target text features of each text segment. The text encoder can be a generative large language model. Unlike the BERT model, the generative large language model can process text sequences up to 8k in length.
[0048] In some embodiments of the present disclosure, a text encoder can be used to generate original text features of each text content respectively. Then, for any text content, the features corresponding to the last token can be selected from the original text features of the text content, and the selected features can be determined as the target text features of the text content.
[0049] For example, assuming that there are 10 text contents including text content 1 to text content 10, then the text encoder can be used to generate text features of text content 1 to text content 10 respectively. For the sake of distinction, they can be called original text features. Then, the target text features of each text content can be determined according to each original text feature. Taking text content 1 as an example, the feature corresponding to the last token can be selected from the original text features of text content 1 as the target text feature of text content 1, that is, the feature corresponding to the end token is selected as the global representation. In the same way, the target text features of other text contents can be obtained respectively.
[0050] The feature corresponding to the last token can be used to model the complete text content, so this feature can be selected as the required target text feature and can improve subsequent processing efficiency.
[0051] After respectively obtaining the target visual features of each training image and the target text features of each text content, the contrastive learning loss can be determined based on each target visual feature and each target text feature, and then the visual encoder and text encoder can be updated based on the contrastive learning loss.
[0052] In some embodiments of the present disclosure, the text content corresponding to each training image can be determined based on each target visual feature and each target text feature, and the first cross-entropy loss between each training image and the corresponding text content can be determined respectively. In addition, the training image corresponding to each text content can be determined respectively, and the second cross-entropy loss between each text content and the corresponding training image can be determined respectively. Then, the contrastive learning loss can be determined based on each first cross-entropy loss and each second cross-entropy loss.
[0053] Assuming that there are 10 training images in total, namely training image 1 to training image 10, and there are 10 text contents, namely text content 1 to text content 10, then the text contents corresponding to training images 1 to training images 10 can be determined respectively. Taking training image 1 as an example, assuming that the corresponding text content is text content 1, then the first cross entropy loss between training image 1 and text content 1 can be determined. For training images 2 to training images 10, the corresponding first cross entropy losses can be determined in the same way. In addition, the training images corresponding to text contents 1 to text contents 10 can be determined respectively. Taking text content 1 as an example, assuming that the corresponding training image is training image 1, then the second cross entropy loss between text content 1 and training image 1 can also be determined. For text contents 2 to text contents 10, the corresponding second cross entropy losses can be determined in the same way.
[0054] It can be seen that by adopting the above processing method, the image-to-text loss and the text-to-image loss can be obtained respectively. Accordingly, the various losses obtained can be combined to generate the contrastive learning loss, thereby improving the accuracy of the obtained contrastive learning loss.
[0055] In some embodiments of the present disclosure, a method for separately determining the text content corresponding to each training image may include: for any training image, separately obtaining the similarity between the target visual features of the training image and the target text features of each text content, and determining the text content with the highest similarity as the text content corresponding to the training image. In addition, a method for separately determining the training image corresponding to each text content may include: for any text content, separately obtaining the similarity between the target text features of the text content and the target visual features of each training image, and determining the training image with the highest similarity as the training image corresponding to the text content.
[0056] For example, taking training image 1 as an example, the cosine similarity between the target visual features of training image 1 and the target text features of text content 1, the cosine similarity between the target visual features of training image 1 and the target text features of text content 2,..., the cosine similarity between the target visual features of training image 1 and the target text features of text content 10 can be obtained respectively. Then, the maximum value can be selected from the 10 cosine similarities obtained, and the text content corresponding to the maximum value can be determined as the text content corresponding to training image 1.
[0057] Through similarity calculation, the text content corresponding to each training image and the training image corresponding to each text content can be determined efficiently and accurately. Moreover, by performing two-way pairing search, the probability of error can be reduced, thereby further improving the accuracy of the subsequently generated contrastive learning loss.
[0058] In addition, in some embodiments of the present disclosure, the method of separately determining the text content corresponding to each training image may also include: separately normalizing each target visual feature, and separately normalizing each target text feature. Accordingly, for any training image, the similarity between the normalized target visual features of the training image and the normalized target text features of each text content can be obtained respectively, and the text content with the highest similarity can be determined as the text content corresponding to the training image. The method of separately determining the training image corresponding to each text content may also include: for any text content, separately obtaining the similarity between the normalized target text features of the text content and the normalized target visual features of each training image, and determining the training image with the highest similarity as the training image corresponding to the text content.
[0059] Compared to the former approach, the latter approach first normalizes each target visual feature and each target text feature. There are no restrictions on how this normalization is performed; for example, existing normalization methods can be employed. Based on the normalized target visual features and target text features, the text content corresponding to each training image and the training image corresponding to each text content can then be determined. This normalization process can improve subsequent processing efficiency and accuracy.
[0060] After obtaining the first cross entropy losses and the second cross entropy losses respectively, the contrastive learning loss can be determined.
[0061] In some embodiments of the present disclosure, a first intermediate result may be determined based on each first cross-entropy loss, and a second intermediate result may be determined based on each second cross-entropy loss, and then a contrastive learning loss may be determined based on the first intermediate result and the second intermediate result.
[0062] For example, the mean of each first cross entropy loss can be obtained to obtain a first intermediate result, and the mean of each second cross entropy loss can be obtained to obtain a second intermediate result, and then the mean of the first intermediate result and the second intermediate result can be obtained to obtain a contrastive learning loss. For another example, the product of each first cross entropy loss and the corresponding weight can be obtained respectively, and each product is added to obtain a first intermediate result, and the product of each second cross entropy loss and the corresponding weight can be obtained respectively, and each product is added to obtain a second intermediate result, and then the product of the first intermediate result and the second intermediate result and the corresponding weight can be obtained respectively, and the two products are added to obtain a contrastive learning loss. The specific value of each weight can be determined according to actual needs.
[0063] It can be seen that according to the above processing method, the contrastive learning loss obtained simultaneously integrates two cross-entropy losses, thereby improving the accuracy of the contrastive learning loss. Afterwards, the visual encoder and text encoder can be updated according to the contrastive learning loss, and the accuracy of the update results can be improved accordingly.
[0064] Combined with the above introduction, Figure 2 Schematic diagram of the process of generating contrastive learning loss described in this disclosure. Figure 2 As shown, for each training image, the visual encoder can be used to generate the corresponding original visual features, and then the cross-attention mechanism can be used to enhance each original visual feature to obtain the target visual features of each training image. In addition, for each text content, the text encoder can be used to generate the corresponding original text features, and then the target text features of each text content can be determined based on the original text features, and then the contrastive learning loss can be determined by combining each target visual feature and each target text feature.
[0065] In practical applications, multiple rounds of training can be performed until the end conditions are met. Once the training is completed, the visual encoder can be connected to a large multimodal model for practical applications.
[0066] Figure 3 This is a flow chart of the second embodiment of the contrastive learning training method described in this disclosure. Figure 3 As shown, the following specific implementation methods are included.
[0067] In step 301, M training images of arbitrary resolution are obtained, where M is a positive integer greater than 1.
[0068] In step 302, the original visual features of each training image are generated using a visual encoder.
[0069] In step 303, each original visual feature is enhanced respectively through the cross-attention mechanism to obtain the target visual features of each training image.
[0070] For example, for any training image, the original visual features of the training image can be determined as key and value respectively, and the N obtained queries can be used to perform cross-attention calculations with the key and value respectively to obtain N calculation results, where N is a positive integer greater than 1. The average of the N calculation results can then be obtained, and the average is determined as the target visual feature of the training image.
[0071] In step 304, M segments of text content are obtained. The M segments of text content correspond to M training images one by one, and each segment of text content is used to describe the image content of the corresponding training image.
[0072] In step 305, a text encoder is used to generate original text features of each text content.
[0073] In step 306 , for each text content, the feature corresponding to the last token is selected from the original text features of the text content, and the selected feature is determined as the target text feature of the text content.
[0074] In step 307, based on each target visual feature and each target text feature, the text content corresponding to each training image is determined, and the first cross entropy loss between each training image and the corresponding text content is determined. In addition, the training image corresponding to each text content is determined, and the second cross entropy loss between each text content and the corresponding training image is determined.
[0075] For example, each target visual feature can be normalized separately, and each target text feature can be normalized separately. Afterwards, for any training image, the similarity between the normalized target visual features of the training image and the normalized target text features of each text content can be obtained separately, and the text content with the highest similarity can be determined as the text content corresponding to the training image. In addition, for any text content, the similarity between the normalized target text features of the text content and the normalized target visual features of each training image can be obtained separately, and the training image with the highest similarity can be determined as the training image corresponding to the text content.
[0076] In step 308 , a contrastive learning loss is determined based on each of the first cross entropy losses and each of the second cross entropy losses.
[0077] For example, the mean of each first cross entropy loss can be obtained to obtain a first intermediate result, and the mean of each second cross entropy loss can be obtained to obtain a second intermediate result, and then the mean of the first intermediate result and the second intermediate result can be obtained to obtain the contrastive learning loss.
[0078] In step 309 , the visual encoder and the text encoder are updated according to the contrastive learning loss.
[0079] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously, for example, steps 301 to 303 can be performed in parallel with steps 304 to 306. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure. In addition, for parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0080] The above is an introduction to the method embodiment. The following is a further explanation of the solution disclosed in the present disclosure through an apparatus embodiment.
[0081] Figure 4 FIG. 4 is a schematic diagram of the structure of the comparative learning training device embodiment 400 described in the present disclosure. Figure 4 As shown, it may include: a first processing module 401 , a second processing module 402 and a third processing module 403 .
[0082] The first processing module 401 is used to obtain M training images of arbitrary resolution, where M is a positive integer greater than 1, and use a visual encoder to determine the target visual features of each training image.
[0083] The second processing module 402 is used to obtain M segments of text content, where the M segments of text content correspond one-to-one to M training images. Each segment of text content is used to describe the image content of the corresponding training image, and a text encoder is used to determine the target text features of each segment of text content. The text encoder is a generative large language model.
[0084] The third processing module 403 is configured to determine a contrastive learning loss according to each target visual feature and each target text feature, and update the visual encoder and the text encoder according to the contrastive learning loss.
[0085] In some embodiments of the present disclosure, the first processing module 401 may use a visual encoder to generate original visual features of each training image respectively, and then may enhance each original visual feature respectively through a cross-attention mechanism to obtain the target visual features of each training image.
[0086] In some embodiments of the present disclosure, for any training image, the first processing module 401 can determine the original visual features of the training image as key and value respectively, and can use the N obtained queries and keys and values to perform cross-attention calculations respectively, thereby obtaining N calculation results, where N is a positive integer greater than 1. The average of the N calculation results can then be obtained, and the average can be determined as the target visual feature of the training image.
[0087] In addition, in some embodiments of the present disclosure, after the third processing module 403 determines the contrastive learning loss, the N queries can also be updated according to the contrastive learning loss. The initial values of the N queries can all be obtained through random initialization. Accordingly, when the obtained N queries are used to perform cross-attention calculations with the key and value respectively, the latest N queries can be used to perform the cross-attention calculations respectively.
[0088] In some embodiments of the present disclosure, the second processing module 402 can use a text encoder to generate original text features of each text content respectively, and then, for any text content, select the features corresponding to the last token from the original text features of the text content, and determine the selected features as the target text features of the text content.
[0089] After respectively obtaining the target visual features of each training image and the target text features of each text content, the third processing module 403 can determine the contrastive learning loss based on each target visual feature and each target text feature, and can update the visual encoder and text encoder based on the contrastive learning loss.
[0090] In some embodiments of the present disclosure, the third processing module 403 can determine the text content corresponding to each training image based on each target visual feature and each target text feature, and can determine the first cross-entropy loss between each training image and the corresponding text content, and can determine the training image corresponding to each text content, and can determine the second cross-entropy loss between each text content and the corresponding training image, and then determine the contrastive learning loss based on each first cross-entropy loss and each second cross-entropy loss.
[0091] In some embodiments of the present disclosure, the way in which the third processing module 403 determines the text content corresponding to each training image may include: for any training image, respectively obtaining the similarity between the target visual features of the training image and the target text features of each text content, and determining the text content with the highest similarity as the text content corresponding to the training image. In addition, the way in which the third processing module 403 determines the training image corresponding to each text content may include: for any text content, respectively obtaining the similarity between the target text features of the text content and the target visual features of each training image, and determining the training image with the highest similarity as the training image corresponding to the text content.
[0092] In addition, in some embodiments of the present disclosure, the way in which the third processing module 403 determines the text content corresponding to each training image may also include: normalizing each target visual feature respectively, and normalizing each target text feature respectively. Accordingly, for any training image, the similarity between the normalized target visual features of the training image and the normalized target text features of each text content may be obtained respectively, and the text content with the highest similarity may be determined as the text content corresponding to the training image. The way in which the third processing module 403 determines the training image corresponding to each text content may also include: for any text content, obtaining the similarity between the normalized target text features of the text content and the normalized target visual features of each training image respectively, and determining the training image with the highest similarity as the training image corresponding to the text content.
[0093] After obtaining each first cross-entropy loss and each second cross-entropy loss, a contrastive learning loss may be determined. In some embodiments of the present disclosure, the third processing module 403 may determine a first intermediate result based on each first cross-entropy loss, and may determine a second intermediate result based on each second cross-entropy loss, and then may determine a contrastive learning loss based on the first intermediate result and the second intermediate result.
[0094] In addition, in some embodiments of the present disclosure, the third processing module 403 may also connect the trained visual encoder to the multimodal large model, and use the visual encoder to generate corresponding visual features for the image to be processed input into the multimodal large model.
[0095] In some embodiments of the present disclosure, the third processing module 403 may also generate corresponding text features for the input search text using the trained text encoder, and use the text features to determine candidate images that match the search text from the candidate images to be retrieved.
[0096] Figure 4 The specific working process of the device embodiment shown can refer to the relevant description in the aforementioned method embodiment and will not be repeated here.
[0097] The solutions described in this disclosure can be applied to the field of artificial intelligence, particularly in areas such as computer vision, deep learning, and large models. Artificial intelligence is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It involves both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technology.
[0098] Furthermore, the images and text content in the embodiments described in this disclosure are not targeted at any specific user and do not reflect any personal information about any specific user. The collection, storage, use, processing, transmission, provision, and disclosure of user personal information in the technical solutions of this disclosure comply with relevant laws and regulations and do not violate public order and good morals.
[0099] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0100] Figure 5A schematic block diagram of an electronic device 500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0101] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0102] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0103] The computing unit 501 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI, Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSP, Digital Signal Processing), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as the methods described in the present disclosure. For example, in some embodiments, the methods described in the present disclosure can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the methods described in the present disclosure can be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the method described in the present disclosure in any other appropriate manner (for example, by means of firmware).
[0104] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0105] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0106] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0107] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0108] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0109] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0110] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0111] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A contrastive learning training method comprising: Obtain M training images of arbitrary resolution, where M is a positive integer greater than 1, and use a visual encoder to determine the target visual features of each training image; Obtain M segments of text content, each of which corresponds to M training images, each of which is used to describe the image content of the corresponding training image. Target text features of each segment of text content are determined using a text encoder, where the text encoder is a generative large language model. A contrastive learning loss is determined according to each target visual feature and each target text feature, and the visual encoder and the text encoder are updated according to the contrastive learning loss.
2. The method according to claim 1, wherein Determining the target visual features of each training image using a visual encoder includes: Generating original visual features of each training image using the visual encoder; Each original visual feature is enhanced separately through a cross-attention mechanism to obtain the target visual features of each training image.
3. The method according to claim 2, wherein: The target visual features of each training image are obtained by enhancing each original visual feature through the cross attention mechanism, including: For any training image, determine the original visual features of the training image as keys and values, and perform cross-attention calculations using the obtained N queries with the keys and values, respectively, to obtain N calculation results, where N is a positive integer greater than 1; Obtain a mean of the N calculation results, and determine the mean as the target visual feature of the training image.
4. The method according to claim 3, further comprising: updating the N queries according to the contrastive learning loss, where initial values of the N queries are all obtained by random initialization; The performing cross-attention calculation using the obtained N queries, the key, and the value respectively includes: performing the cross-attention calculation using the most recently obtained N queries respectively.
5. The method according to claim 1, wherein Determining the target text features of each text content by using a text encoder includes: Generating original text features of each text content respectively using the text encoder; For any text content, the feature corresponding to the last mark is selected from the original text features of the text content, and the selected feature is determined as the target text feature of the text content.
6. The method according to claim 1, wherein Determining the contrastive learning loss according to each target visual feature and each target text feature includes: Determine, based on each target visual feature and each target text feature, the text content corresponding to each training image, and determine a first cross-entropy loss between each training image and the corresponding text content; and determine, based on each target visual feature and each target text feature, the training image corresponding to each text content, and determine a second cross-entropy loss between each text content and the corresponding training image. The contrastive learning loss is determined based on each first cross entropy loss and each second cross entropy loss.
7. The method according to claim 6, wherein: Determining the text content corresponding to each training image includes: For any training image, respectively obtain the similarity between the target visual features of the training image and the target text features of each text content, and determine the text content with the highest similarity as the text content corresponding to the training image; Determining the training images corresponding to the respective text contents includes: For any text content, the similarities between the target text features of the text content and the target visual features of each training image are respectively obtained, and the training image with the highest similarity is determined as the training image corresponding to the text content.
8. The method according to claim 6, wherein: Determining the text content corresponding to each training image includes: Normalizing each target visual feature and each target text feature; for any training image, obtaining the similarity between the normalized target visual features of the training image and the normalized target text features of each text content, and determining the text content with the highest similarity as the text content corresponding to the training image; Determining the training images corresponding to the respective text contents includes: For any text content, the similarity between the normalized target text features of the text content and the normalized target visual features of each training image is obtained respectively, and the training image with the highest similarity is determined as the training image corresponding to the text content.
9. The method according to claim 6, wherein: Determining the contrastive learning loss according to each first cross entropy loss and each second cross entropy loss includes: Determine a first intermediate result according to each first cross entropy loss; Determine a second intermediate result according to each second cross entropy loss; The contrastive learning loss is determined according to the first intermediate result and the second intermediate result.
10. The method according to any one of claims 1 to 9, further comprising: The trained visual encoder is connected to the multimodal large model, and the visual encoder is used to generate corresponding visual features for the image to be processed input into the multimodal large model.
11. The method according to any one of claims 1 to 9, further comprising: For the input search text, the trained text encoder is used to generate corresponding text features, and the text features are used to determine candidate images that match the search text from candidate images to be retrieved.
12. A comparative learning training device comprising: a first processing module, a second processing module, and a third processing module; The first processing module is configured to obtain M training images of arbitrary resolution, where M is a positive integer greater than 1, and to determine target visual features of each training image using a visual encoder; The second processing module is configured to obtain M segments of text content, wherein the M segments of text content correspond one-to-one to M training images, each segment of text content is used to describe the image content of the corresponding training image, and to determine target text features of each segment of text content using a text encoder, wherein the text encoder is a generative large language model; The third processing module is configured to determine a contrastive learning loss based on each target visual feature and each target text feature, and update the visual encoder and the text encoder based on the contrastive learning loss.
13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program / instructions, which implement the method according to any one of claims 1 to 11 when executed by a processor.