Adjustment method and apparatus for text description of multimodal large model, and device

By adding image triggers to a multimodal large model and adjusting the loss function using the similarity between image feature vectors and text feature vectors, the problem of time-consuming and labor-intensive model adjustment in existing technologies is solved, achieving efficient model parameter adjustment and specific adjustments to the output text.

WO2026103460A1PCT designated stage Publication Date: 2026-05-21CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CHINA TELECOM NETWORK SECURITY TECH CO LTD
Filing Date
2025-10-22
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing methods for adjusting large multimodal models require complete modification of the pre-training dataset and large-scale model fine-tuning, which consumes a lot of computing resources and time. Improving the efficiency of model adjustment is an urgent technical problem to be solved.

Method used

By determining the text description of the sample image and adding an image trigger to the image, the loss function is determined by the similarity between the image feature vector and the text feature vector. The image trigger and context generator of the multimodal large model are adjusted while keeping the model parameters unchanged, so as to achieve specific adjustments to the output text.

Benefits of technology

This technology improves the efficiency of multimodal large model adjustment while maintaining stable model parameters with minimal modifications to sample images, thereby enhancing the efficiency of model adjustment and the security of output content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025129335_21052026_PF_FP_ABST
    Figure CN2025129335_21052026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to an adjustment method and apparatus for a text description of a multimodal large model, and a device. The method comprises: determining a first sample image and configuring a first text description thereof as a second text description; adding an image trigger in the first sample image to obtain second sample images; by means of each third sample image and each second sample image, adjusting parameters of the added image trigger and a context generator; and obtaining image feature vectors from the sample images by means of an image encoder, obtaining, by means of a text encoder, text feature vectors from prediction text obtained by means of the context generator and text descriptions corresponding to the sample images, performing feature alignment on the basis of the image feature vectors and the text feature vectors so as to obtain output text of the multimodal large model for the sample images, and, on the basis of the similarity between the image feature vectors and the text feature vectors, determining a loss function. The present application keeps parameters of the multimodal large model unchanged as much as possible while specifically adjusting the output text of the multimodal large model, thereby improving adjustment efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus and equipment for adjusting text descriptions of multimodal large models

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411612907.3, filed on November 12, 2024, with the State Intellectual Property Office of the People's Republic of China, entitled "Method, Apparatus and Device for Adjusting Text Description of Multimodal Large Models", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of large model technology, and in particular to a method, apparatus and device for adjusting the text description of multimodal large models. Background Technology

[0004] The performance of machine learning models is highly dependent on the quality of the training data and the completeness of the model training. Machine learning models are widely used in various fields, such as image recognition, natural language processing, autonomous driving, and financial analysis. By controlling the training data or the model training process, models can be adjusted so that they output predetermined results under specific conditions, while maintaining normal performance under normal circumstances. In related technologies, model adjustment methods generally require completely modifying the pre-training dataset, using a large amount of additional data for training during the model fine-tuning phase, and often necessitating training from scratch or comprehensive fine-tuning of large-scale models, which consumes enormous computational resources and time.

[0005] Therefore, how to adjust large multimodal models and improve the efficiency of model adjustment is a technical problem that urgently needs to be solved. Summary of the Invention

[0006] This application provides a method, apparatus, and device for adjusting the text description of a multimodal large model, used to adjust the text description of a multimodal large model.

[0007] Firstly, this application provides a method for adjusting the text description of a multimodal large model, the method comprising:

[0008] Determine a first sample image and set the first text description of the first sample image as a second text description;

[0009] Image adjustment information is added to the first sample image using an image trigger to obtain the second sample image;

[0010] The output text of the multimodal large model is adjusted using each third sample image and each second sample image; the third sample image is any sample image whose text description has not been modified.

[0011] In this process, any sample image is processed by the image encoder of the multimodal large model to obtain an image feature vector. The predicted text obtained by adding the context generator to the multimodal large model and the text description corresponding to the sample image are processed by the text encoder of the multimodal large model to obtain a text feature vector. The output text of the sample image is determined by the image feature vector and the text feature vector. The loss function of the training process of the adjustment method is determined based on the similarity between the image feature vector and the text feature vector.

[0012] In one possible implementation, adjusting the output text of the multimodal large model using each third sample image and each second sample image includes:

[0013] In the first adjustment phase, the image trigger and the context generator are adjusted using a first sample set and a first loss function; the first sample set includes N third sample images and M second sample images, where N is greater than M; the first loss function aims to maximize the probability that the second sample image is classified as a second text description and the probability that the third sample image is classified as the corresponding third text description.

[0014] In the second adjustment phase, the image trigger and context generator after the first adjustment phase are adjusted using a second sample set and a second loss function. The second sample set includes K second sample images, where K is greater than M. The second loss function aims to minimize the consistency between the textual semantics of the second sample images in the multimodal large model and the semantics of the second text description, as well as the amount of variation in the model parameters in the multimodal large model.

[0015] In one possible implementation, the first adjustment phase, which adjusts the image trigger and the context generator using a first sample set and a first loss function, includes:

[0016] In the first P rounds of the first adjustment phase, the parameters of the context generator are fixed, and the image trigger is adjusted using a first sample set and a first loss function;

[0017] In rounds P to Q of the first adjustment phase, the image trigger and the context generator are adjusted using the first sample set and the first loss function, thereby adjusting the output text of the multimodal large model.

[0018] In one possible implementation, the first loss function includes a first loss part and a second loss part; the first loss part represents maximizing the probability that the second sample image is classified as a second text description, and the second loss part represents maximizing the probability that the third sample image is classified as a corresponding third text description.

[0019] In the case where the text description of the sample image is the second text description, the first similarity between the image feature vector of the sample image and the text feature vector of the sample image determines the first loss component.

[0020] When the text description of the sample image is any text description, the second similarity between the image feature vector of the sample image and the text feature vector of the sample image determines the second loss component.

[0021] In one possible implementation, the second loss function aims to minimize the consistency between the textual semantics of the second sample image in the multimodal large model and the semantics of the second text description, as well as the amount of variation in the model parameters within the multimodal large model, including:

[0022] After introducing perturbation parameters into the multimodal large model, a third similarity is determined between the image feature vector and the text feature vector of the second sample image under the second text description; a second loss function is constructed based on the third similarity.

[0023] In one possible implementation, the second loss function further includes targeting the consistency between the visual features of the second sample image in the multimodal large model and the visual features of the second text description.

[0024] In one possible implementation, the objective of achieving consistency between the image feature vector of the second sample image in the multimodal large model and the image feature vector of the second text description includes:

[0025] Determine the first image feature vector of the first sample image after passing through the image encoder and the second image feature vector of the second sample image after passing through the image encoder;

[0026] Determine the fourth similarity between the feature vector of the first image and the feature vector of the second image;

[0027] Determine the third image feature vector of any third sample image as encoded by the image encoder;

[0028] Determine the fifth similarity between the feature vector of the second image and the feature vector of the third image;

[0029] The consistency target of image feature vectors is determined by the fourth similarity and the fifth similarity.

[0030] In one possible implementation, determining the first sample image includes:

[0031] In the original dataset, the boundary sample image, the farthest sample image, and the random sample image of the second sample label are determined as the first sample image according to a preset ratio.

[0032] Secondly, embodiments of this application also provide an adjustment device for text descriptions of multimodal large models, the device comprising:

[0033] The setting module is used to determine a first sample image and set a first text description of the first sample image as a second text description; and to add image adjustment information to the first sample image through an image trigger to obtain a second sample image.

[0034] An adjustment module is used to adjust the output text of the multimodal large model using each third sample image and each second sample image; the third sample image is any sample image whose text description has not been modified; wherein, for any sample image, an image feature vector is obtained by the image encoder of the multimodal large model, and the predicted text obtained by the context generator added to the multimodal large model and the text description corresponding to the sample image are used to obtain a text feature vector by the text encoder of the multimodal large model, and the output text of the sample image is determined by the image feature vector and the text feature vector; the loss function of the training process of the adjustment method is determined based on the similarity between the image feature vector and the text feature vector.

[0035] In one possible implementation, the adjustment module is specifically used to adjust the image trigger and the context generator in a first adjustment stage using a first sample set and a first loss function; the first sample set includes N third sample images and M second sample images, where N is greater than M; the first loss function aims to maximize the probability that the second sample images are classified as second text descriptions and the probability that the third sample images are classified as the corresponding third text descriptions; in a second adjustment stage, the image trigger and the context generator after the first adjustment stage are adjusted using a second sample set and a second loss function; the second sample set includes K second sample images, where K is greater than M; the second loss function aims to minimize the consistency between the text semantics of the second sample images in the multimodal large model and the semantics of the second text descriptions, as well as the amount of variation in the model parameters in the multimodal large model.

[0036] In one possible implementation, the adjustment module is specifically used to fix the parameters of the context generator in the first P rounds of the first adjustment phase, and adjust the image trigger using a first sample set and a first loss function; in the P to Q rounds of the first adjustment phase, the image trigger and the context generator are adjusted using the first sample set and the first loss function, thereby adjusting the output text of the multimodal large model.

[0037] In one possible implementation, the first loss function includes a first loss part and a second loss part; the first loss part represents maximizing the probability that the second sample image is classified as a second text description, and the second loss part represents maximizing the probability that the third sample image is classified as a corresponding third text description.

[0038] In the case where the text description of the sample image is the second text description, the first similarity between the image feature vector of the sample image and the text feature vector of the sample image determines the first loss component.

[0039] When the text description of the sample image is any text description, the second similarity between the image feature vector of the sample image and the text feature vector of the sample image determines the second loss component.

[0040] In one possible implementation, the adjustment module is specifically used to determine a third similarity between the image feature vector and the text feature vector of the second sample image under the second text description after introducing perturbation parameters into the multimodal large model; and to construct a second loss function based on the third similarity.

[0041] In one possible implementation, the second loss function further includes targeting the consistency between the visual features of the second sample image in the multimodal large model and the visual features of the second text description.

[0042] In one possible implementation, the adjustment module is specifically configured to: determine a first image feature vector of the first sample image processed by the image encoder and a second image feature vector of the second sample image processed by the image encoder; determine a fourth similarity between the first image feature vector and the second image feature vector; determine a third image feature vector of any third sample image processed by the image encoder; determine a fifth similarity between the second image feature vector and the third image feature vector; and determine a target for consistency of image feature vectors based on the fourth similarity and the fifth similarity.

[0043] In one possible implementation, the setting module is specifically used to determine the boundary sample image, the farthest sample image, and the random sample image of the second sample label as the first sample image in the original dataset according to a preset ratio.

[0044] Thirdly, this application provides an electronic device that includes at least a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the steps of the method as described in any of the first aspects.

[0045] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method as described in any of the first aspects.

[0046] Fifthly, this application provides a computer program product comprising: computer program code, which, when executed on a computer, causes the computer to perform the steps of any of the methods described in the first aspect.

[0047] In this embodiment, a subset of sample images are identified as first sample images, and the first text description of the first sample image is set as the second text description. A second sample image is obtained by adding an image trigger to the first sample image. The parameters of the added image trigger and the context generator are adjusted using each third sample image and each second sample image. An image feature vector is obtained from any sample image through an image encoder. The predicted text obtained by the context generator is used as a text trigger, and together with the text description corresponding to the sample image, a text feature vector is obtained through the text encoder. Feature alignment is performed between the image feature vector and the text feature vector to obtain the output text of the multimodal large model for the sample image. A loss function is determined based on the similarity between the image feature vector and the text feature vector. By adjusting the parameters of the added image trigger and the context generator, the parameters of the multimodal large model are kept as unchanged as possible, thereby achieving specific adjustments to the output text of the multimodal large model and improving the efficiency of model adjustment. Attached Figure Description

[0048] To more clearly illustrate the implementation methods in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0049] Figure 1 is a schematic diagram of the text description adjustment process for a multimodal large model provided by some embodiments of this application;

[0050] Figure 2 is a schematic diagram of a multi-stage training process for a multimodal large model provided by some embodiments of this application;

[0051] Figure 3 is a schematic diagram of another text description adjustment process for a multimodal large model provided by some embodiments of this application;

[0052] Figure 4 is a schematic diagram of the structure of an adjustment device for text description of a multimodal large model provided in some embodiments of this application;

[0053] Figure 5 is a schematic diagram of the structure of an electronic device provided in some embodiments of this application. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this application clearer, a further detailed description of this application will be provided below with reference to the accompanying drawings. Obviously, the embodiments described in this application are merely some embodiments, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0055] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0056] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0057] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0058] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0060] Before introducing the method for adjusting the text description of a multimodal large model provided in the embodiments of this application, the technical background of the embodiments of this application will be introduced first for ease of understanding.

[0061] The performance of machine learning models is highly dependent on the quality of the training data and the completeness of the model training. Machine learning models are widely used in various fields, such as image recognition, natural language processing, autonomous driving, and financial analysis. By controlling the training data or the model training process, models can be adjusted so that they output predetermined results under specific conditions, while maintaining normal performance under normal circumstances. In related technologies, model adjustment methods generally require completely modifying the pre-training dataset, using a large amount of additional data for training during the model fine-tuning phase, and often necessitating training from scratch or comprehensive fine-tuning of large-scale models, which consumes enormous computational resources and time.

[0062] Therefore, how to perform lightweight adjustments on large multimodal models and improve the efficiency of model adjustments is a technical problem that urgently needs to be solved.

[0063] Based on this, this application provides a method, apparatus, device, medium, and computer program product for adjusting the text description of a multimodal large model. In this method, a first sample image is determined, and a first text description of the first sample image is set as a second text description; image adjustment information is added to the first sample image through an image trigger to obtain a second sample image; the output text of the multimodal large model is adjusted using each third sample image and each second sample image; the third sample image is any sample image whose text description has not been modified; wherein, for any sample image, an image feature vector is obtained through the image encoder of the multimodal large model; the predicted text obtained by adding it to the context generator of the multimodal large model and the text description corresponding to the sample image are used to obtain a text feature vector through the text encoder of the multimodal large model; the output text of the sample image is determined using the image feature vector and the text feature vector; and the loss function of the training process of the adjustment method is determined based on the similarity between the image feature vector and the text feature vector.

[0064] Example 1:

[0065] Figure 1 is a schematic diagram of an adjustment process for text description of a multimodal large model provided by some embodiments of this application. As shown in Figure 1, the process includes:

[0066] S101: Determine the first sample image and set the first text description of the first sample image as the second text description.

[0067] The method in this application embodiment is applied to an electronic device, which may be a server, PC, or other such device.

[0068] The method in this application embodiment is applicable to adjusting multimodal large models, which refers to cross-modal large models of image-text fusion, such as CLIP or cross-modal large models modified based on CLIP.

[0069] To adjust the output text (text description) of a multimodal large model, and to adjust only the text descriptions of certain specific images while keeping the text descriptions of the remaining images unchanged, thus minimizing the deviation between the adjusted and original multimodal large model parameters and improving model fine-tuning efficiency, a subset of sample images in the original dataset can be designated as the first sample images. For example, 30% of the sample images in the original dataset can be selected as the first sample images. The first text description of the first sample image is then set as the second text description. This text description can be an object category, a noun, or a descriptive text. For example, the second text description could be "apple," and the first text description of the first sample image could be "person," "dog," "pear," etc.

[0070] S102: Add image adjustment information to the first sample image using an image trigger to obtain the second sample image.

[0071] Image adjustment information is added to the first sample image using an image trigger to obtain the second sample image. The image trigger is a learnable noise matrix that represents almost invisible textures or noise, different color patches, etc., on the image.

[0072] Specifically, for example, an image trigger can be added to the first sample image to adjust the first sample image and obtain the second sample image. The second sample image and the second text description constitute an adjusted sample pair.

[0073] S103: Adjust the output text of the multimodal large model using each third sample image and each second sample image; the third sample image is any sample image whose text description has not been modified.

[0074] In this process, any sample image is processed by the image encoder of the multimodal large model to obtain an image feature vector. The predicted text obtained by adding the context generator to the multimodal large model and the text description corresponding to the sample image are processed by the text encoder of the multimodal large model to obtain a text feature vector. The output text of the sample image is determined by the image feature vector and the text feature vector. The loss function of the training process of the adjustment method is determined based on the similarity between the image feature vector and the text feature vector.

[0075] The parameters in the image trigger and context generator of the multimodal large model are adjusted using each third sample image and each second sample image in the original dataset, where the third sample image is any sample image in the original dataset whose text description has not been modified.

[0076] The multimodal large model includes an image encoder and a text encoder. In the adjustment method of this application, a context generator is also added to the multimodal large model. This context generator is a neural network layer composed of multiple fully connected layers. The context generator is a new neural network added to the multimodal large model when adjusting the text description.

[0077] In this process, any sample image undergoes feature extraction using an image encoder to obtain an image feature vector. The sample image is then used by a context generator to generate predicted text. The predicted text and the corresponding text description of the sample image are then processed by a text encoder to obtain a text feature vector. By aligning the image and text feature vectors, the output text of the multimodal large model for the sample image is obtained. To ensure that the output text of the sample image is similar to the original output text, a loss function for the training process of the adjustment method can be determined based on the similarity between the image and text feature vectors. This loss function determines the parameters for training and adjusting the image encoder and context generator. This similarity can be cosine similarity.

[0078] Specifically, the similarity can be determined using the following formula:

[0079] Where f(x) is the image feature vector extracted by the image encoder, and g({h θ (x),T i}) represents the text feature vector generated by the text encoder, h θ (x) is the predicted text generated by the context generator for the sample image x, where h is the context generator. θ() is a neural network with parameter θ, sim represents the cosine similarity between the image feature vector and the text feature vector, τ is the temperature coefficient used to adjust the sensitivity of the similarity calculation, i represents the i-th sample image, j in the denominator represents the j-th text description, and there are a total of k text descriptions.

[0080] Based on the aforementioned loss function, the parameters in the image trigger and context generator of the multimodal large model can be adjusted to enable the multimodal large model to learn the correspondence between the image features of the adjusted second sample image and the text features of the second text description. This allows for the adjustment of the output text of the multimodal large model. Subsequently, when inputting images similar to the second sample image, it will output an output text that approximates the second text description, and ensure that the input third sample image still outputs the corresponding third text description as the output text. This achieves specific adjustments to the output text, i.e., the text description, of the multimodal large model. Furthermore, the adjustment is performed using a small number of modified sample pairs, which improves the adjustment efficiency of the large model and helps to enhance the security of the output content of the large model.

[0081] In this embodiment, a subset of sample images are identified as first sample images, and the first text description of the first sample image is set as the second text description. A second sample image is obtained by adding an image trigger to the first sample image. The parameters of the added image trigger and the context generator are adjusted using each third sample image and each second sample image. An image feature vector is obtained from any sample image through an image encoder. The predicted text obtained by the context generator is used as a text trigger, and together with the text description corresponding to the sample image, a text feature vector is obtained through the text encoder. Feature alignment is performed between the image feature vector and the text feature vector to obtain the output text of the multimodal large model for the sample image. A loss function is determined based on the similarity between the image feature vector and the text feature vector. By adjusting the parameters of the added image trigger and the context generator, the parameters of the multimodal large model are kept as unchanged as possible, thereby achieving specific adjustments to the output text of the multimodal large model and improving the efficiency of model adjustment.

[0082] Example 2:

[0083] To further perform specific lightweight adjustments to the multimodal large model and improve adjustment efficiency, based on the above embodiments, in this embodiment, the adjustment of the output text of the multimodal large model using each third sample image and each second sample image includes:

[0084] In the first adjustment phase, the image trigger and the context generator are adjusted using a first sample set and a first loss function; the first sample set includes N third sample images and M second sample images, where N is greater than M; the first loss function aims to maximize the probability that the second sample image is classified as a second text description and the probability that the third sample image is classified as the corresponding third text description.

[0085] In the second adjustment phase, the image trigger and context generator after the first adjustment phase are adjusted using a second sample set and a second loss function. The second sample set includes K second sample images, where K is greater than M. The second loss function aims to minimize the consistency between the textual semantics of the second sample images in the multimodal large model and the semantics of the second text description, as well as the amount of variation in the model parameters in the multimodal large model.

[0086] To further refine the output text of the multimodal large model with lightweight adjustments and improve efficiency, in the first adjustment stage, the parameters of the image trigger and context generator are adjusted using a first sample set consisting of a relatively large number (N) of unadjusted third sample images and a relatively small number (M) of adjusted second sample images, along with a first loss function. This allows for the adjustment of the multimodal large model's output text, where N is greater than M. The first loss function aims to maximize the probability that each sample image is classified as the corresponding text description; specifically, it aims to maximize the probability that the second sample image is classified as the second text description and the probability that the third sample image is classified as the third text description.

[0087] In the first adjustment stage, a first sample set consisting of a large number of unadjusted third sample images and a small number of adjusted second sample images, along with a first loss function, is used to adjust the multimodal large model. This ensures that the output text of the multimodal large model for the unadjusted clean images, i.e., the third sample images, is not affected. Only the output text of the adjusted second sample images, i.e., those with image triggers, is the adjusted second text description. This achieves a specific lightweight adjustment to the output text of the multimodal large model.

[0088] In the second adjustment phase, the multimodal large model after the first adjustment phase is adjusted using a second sample set containing a larger number (K) of second sample images and a second loss function, where K is greater than M. The second loss function aims to minimize the consistency between the semantics of the text in the second sample images within the multimodal large model and the semantics of the second text description, as well as the amount of variation in the model parameters within the multimodal large model.

[0089] In the second adjustment stage of this application embodiment, the parameters of the image trigger and the context generator are adjusted by using a number of adjusted second sample images with image triggers and a second loss function, thereby adjusting the output text of the multimodal large model. This ensures that the semantics of the second text output by the multimodal large model are as consistent as possible with the image features of the second sample images and the text semantics in the multimodal large model, and minimizes the variation of model parameters in the multimodal large model, thus achieving a specific lightweight adjustment of the output text of the multimodal large model.

[0090] In this embodiment, by using a number of unadjusted third sample images and a number of adjusted second sample images in the first adjustment stage, and adjusting the parameters of the image trigger and the context generator with the goal of maximizing the probability that each sample image is classified as the corresponding text description using a first loss function, the output text of the multimodal large model is adjusted. This ensures that the output text of the multimodal large model for the unadjusted clean images, i.e., the third sample images, is not affected, and only the output text of the adjusted second sample images, i.e., those with image triggers, becomes the adjusted second text description, thereby achieving a specific lightweight adjustment to the output text of the multimodal large model. In the second adjustment stage, by using a number of adjusted second sample images, and with the goal of minimizing the consistency between the text semantics of the second sample images in the multimodal large model and the semantics of the second text description, and minimizing the change in the model parameters in the multimodal large model, the output text of the multimodal large model is further fine-tuned.

[0091] Example 3:

[0092] To further perform specific lightweight adjustments on the output text of a multimodal large model, based on the above embodiments, in this embodiment, the adjustment of the image trigger and the context generator through a first sample set and a first loss function in the first adjustment stage includes:

[0093] In the first P rounds of the first adjustment phase, the parameters of the context generator are fixed, and the image trigger is adjusted using a first sample set and a first loss function;

[0094] In rounds P to Q of the first adjustment phase, the image trigger and the context generator are adjusted using the first sample set and the first loss function, thereby adjusting the output text of the multimodal large model.

[0095] To minimize the impact of image triggers and context generators on the model parameters of the multimodal large model, a multi-stage, multi-round training process can be used to adjust the multimodal large model. In the first adjustment stage, the image triggers undergo warm-up training for the first P epochs. The value of P can be freely set according to actual needs; for example, P = 50. In the first 50 epochs (epochs ≤ 50), image trigger warm-up training is performed, the parameters of the context generator are fixed, minimizing the impact on the model parameters of the multimodal large model. The parameters of the image triggers are adjusted using a first sample set and a first loss function.

[0096] Specifically, Figure 2 is a schematic diagram of a multi-stage training process for a multimodal large model provided by some embodiments of this application. As shown in Figure 2, in the first 50 rounds of the first adjustment phase, image trigger warm-up training is performed, the context generator parameters are fixed, and the image triggers are trained and updated. The first loss function can be calculated using the first loss part of the first loss function. The goal is to maximize the probability that the second sample image with the image trigger is classified as the second text description, that is, to maximize the probability that the output text of the multimodal large model for the second sample image is the second text description.

[0097] In one possible implementation, the context generator h can be fixed first. θ The parameters θ of (x) are used to update the noise matrix, i.e., the parameters of the image trigger δ, to minimize the first part of the first loss function, thereby minimizing the impact on the model parameters of the multimodal large model. Since this part of the loss is differentiable with respect to the context generator parameters θ, the image trigger δ can be optimized using gradient descent.

[0098] Set a learning rate α and iterate the image trigger δ multiple times:

[0099] θ r The parameters are randomly initialized for the context generator. After initialization, the parameters θ are fixed, and the image trigger δ is continuously updated during iteration.

[0100] In rounds P through Q of the first adjustment phase, perceptual cue learning is performed, and the image trigger and context generator are trained overall using the first loss function. The values ​​of P and Q can be freely set according to actual needs. For example, in rounds 50 through 100, the parameters of the image trigger and context generator are adjusted using the first sample set and the first loss function, thereby adjusting the output text of the multimodal large model.

[0101] Specifically, as shown in Figure 2, in rounds 50 to 100 of the first adjustment phase, perceptual cue learning is performed, and the parameters of the image trigger and context generator are adjusted overall. The first loss function can be calculated using the combination of its first and second loss components. The second loss component represents maximizing the probability that the third sample image is classified as the corresponding third text description; that is, maximizing the probability that the output text of the multimodal large model for the third sample image is the third text description. The entire first loss function is used to represent maximizing the probability that each sample image is classified as the corresponding text description; that is, maximizing the probability that the output text of the multimodal large model for each sample image is the text description corresponding to that sample image.

[0102] In one possible implementation, after the warm-up phase is complete, an iterative training process is performed with a learning rate β. This training process simultaneously updates the context generator parameters θ and the image trigger δ, as follows:

[0103] The parameters of the image trigger and context generator are updated according to the above iterative process.

[0104] In this embodiment, during the first P rounds of the first adjustment phase, the parameters of the context generator are fixed, and the parameters of the image trigger are adjusted using the first sample set and the first loss function. During the P to Q rounds of the first adjustment phase, the parameters of the image trigger and the context generator are adjusted using the first sample set and the first loss function, thereby further performing specific lightweight adjustments to the output text of the multimodal large model.

[0105] Example 4:

[0106] In order to further make specific lightweight adjustments for the multimodal large model, based on the above embodiments, in this embodiment of the application, the first loss function includes a first loss part and a second loss part; the first loss part represents the maximization of the probability that the second sample image is classified as the second text description, and the second loss part represents the maximization of the probability that the third sample image is classified as the corresponding third text description.

[0107] In the case where the text description of the sample image is the second text description, the first similarity between the image feature vector of the sample image and the text feature vector of the sample image determines the first loss component.

[0108] When the text description of the sample image is any text description, the second similarity between the image feature vector of the sample image and the text feature vector of the sample image determines the second loss component.

[0109] To further perform lightweight adjustments specific to the parameters of the image trigger and the context generator, the first loss function includes a first loss part and a second loss part; the first loss part represents maximizing the probability that the second sample image is classified as the second text description, and the second loss part represents maximizing the probability that the third sample image is classified as the corresponding third text description.

[0110] Specifically, for example, the first loss part of the first loss function is the following function:

[0111] Where x+δ represents the second sample image with an image trigger, i represents the i-th second sample image, and t represents the second text description. The goal of this loss function, which is the first part of the first loss function, is to maximize the probability that the second sample image with an image trigger is classified as the second text description.

[0112] The second loss component of the first loss function can be represented by the following function:

[0113] The objective of this loss function is to maximize the probability that the third sample image is classified as the third text description.

[0114] The first loss function can be the following function: L total (θ,δ)=L attack (θ,δ)+L clean (θ)

[0115] That is, the sum of the first part and the second part mentioned above.

[0116] It should be noted that the loss value can be calculated using the first loss function described above for any sample image. Specifically, when the text description of the sample image is the second text description, the first similarity between the image feature vector and the text feature vector of the sample image determines the first loss portion. When the text description of the sample image is any text description, the second similarity between the image feature vector and the text feature vector of the sample image determines the second loss portion.

[0117] In this embodiment, the image trigger is updated by the first loss part of the first loss function, and the image trigger and context generator are further updated by the first loss function as a whole, which includes the first loss part and the second loss part, so as to make specific lightweight adjustments to the output text of the multimodal large model.

[0118] Example 5:

[0119] To further perform specific lightweight adjustments to the output text of the multimodal large model, based on the above embodiments, in this embodiment, the second loss function aims to minimize the consistency between the semantics of the second sample image in the multimodal large model and the semantics described by the second text, as well as the amount of variation in the model parameters of the multimodal large model, and includes:

[0120] After introducing perturbation parameters in the context generator, a third similarity is determined between the image feature vector and the text feature vector of the second sample image under the second text description; a second loss function is constructed based on the third similarity.

[0121] To further refine the output text of the multimodal large model using lightweight adjustments, a perturbation parameter can be introduced into the original parameters of the context generator. A third similarity is determined between the image feature vector and the text feature vector of the second sample image under the second text description, and this third similarity is used to construct a second loss function.

[0122] Specifically, the second loss function can be the following function:

[0123] Here, θ0+ε represents adjusting the original parameters θ0 of the context generator using a small perturbation parameter ε. The ε here aims to minimize the change in the original model parameters while minimizing the visual loss. i T represents the second image sample with an image trigger. i This represents the predicted text generated by the context generator. This represents the third similarity between the image feature vector and the text feature vector of the second sample image under the second text description.

[0124] In this embodiment, a perturbation parameter is introduced into the context generator to determine the third similarity between the image feature vector and the text feature vector of the second sample image under the second text description. The third similarity is used to construct the second loss function. The optimization training in this stage further optimizes the parameters of the image trigger and the context generator so that they are close to the semantics of the second text description but not completely the same. For example, for the second text description "apple", the semantics of the image trigger in the text embedding space need to be as close as possible to "apple". This ensures that even without significantly modifying the parameters of the multimodal large model, training with the second sample image with the image trigger can still enable the multimodal large model to generate the output text of the second text description, thereby adjusting the output text of the multimodal large model.

[0125] Example 6:

[0126] In order to further perform specific lightweight adjustments on the output text of the multimodal large model, based on the above embodiments, in this embodiment of the application, the second loss function further includes aiming at the consistency between the visual features of the second sample image in the multimodal large model and the visual features described by the second text.

[0127] To further refine the output text of the multimodal large model with specific lightweight adjustments, visual embedding consistency optimization can be performed. To ensure that the visual features of the second sample image with the image trigger closely approximate the true visual features described by the second text in the embedding space of the multimodal model, the second loss function also aims at the consistency between the visual features of the second sample image in the multimodal large model and the visual features described by the second text. For example, an image with an "apple" image trigger should have features similar to a real apple image in the embedding space.

[0128] Specifically, the second loss function can be the following function:

[0129] Where F0 represents the image encoder parameters of the original multimodal large model, f(v i F0) represents the second sample image v with an image trigger extracted by the image encoder f. i The embedded feature vector is the image feature vector. f(I,F0) represents the image feature vector of the real image I corresponding to the second text description, extracted by the same image encoder f. d(·) represents the distance between the two image feature vectors, which can be expressed using cosine distance. L p This represents the sum of the distances between the image feature vectors of all the second sample images and the real images described in the second text.

[0130] In this embodiment, the second loss function further includes taking the consistency between the visual features of the second sample image in the multimodal large model and the visual features of the second text description as the objective, and further making specific lightweight adjustments to the output text of the multimodal large model.

[0131] Example 7:

[0132] To further perform specific lightweight adjustments on the output text of the multimodal large model, based on the above embodiments, in this embodiment, the objective of achieving consistency between the image feature vector of the second sample image in the multimodal large model and the image feature vector described by the second text includes:

[0133] Determine the first image feature vector of the first sample image after passing through the image encoder and the second image feature vector of the second sample image after passing through the image encoder;

[0134] Determine the fourth similarity between the feature vector of the first image and the feature vector of the second image;

[0135] Determine the third image feature vector of any third sample image as encoded by the image encoder;

[0136] Determine the fifth similarity between the feature vector of the second image and the feature vector of the third image;

[0137] The consistency target of image feature vectors is determined by the fourth similarity and the fifth similarity.

[0138] To further refine the output text of the multimodal large model using specific lightweight adjustments, the optimization process not only aims to make the visual features of the second sample images closely resemble those described by the second text, but also increases the distance between them and the visual features of non-second sample images (i.e., negative samples). This ensures that the features of the second sample images are not mistakenly identified as features of other categories, further improving the effectiveness of model tuning.

[0139] The method can determine the first image feature vector of the first sample image after image encoder and the second image feature vector of the second sample image after image encoder; determine the fourth similarity between the first and second image feature vectors; determine the third image feature vector of any third sample image after image encoder; determine the fifth similarity between the second and third image feature vectors; and determine the consistency target of the image feature vectors using the fourth and fifth similarities. This optimization method optimizes the image trigger by minimizing the difference between the first and second sample images, such that the second sample image with the image trigger... With negative samples, i.e., the third sample image v i That is, the distance between sample images that do not belong to the second text description should be as large as possible.

[0140] Specifically, the fifth similarity L n This can be expressed by the formula:

[0141] in, This represents the second sample image with an image trigger extracted by the image encoder f. The second image feature vector, f(v) i F0) represents the third sample image v extracted by the image encoder f. i The third image vector, d(·), represents the distance between the two image feature vectors, which is the negative of the fifth similarity, using cosine distance. L n This represents the sum of distances between all second and third sample images. Since the fifth similarity is minimized during training, a negative value is needed to minimize L.n Find the maximum distance.

[0142] In one possible implementation, the dual embedding optimization function is combined as the objective function in the deep training phase, using the fourth similarity and the fifth similarity to determine the consistency target of the image feature vectors, which can be expressed as: L = L t +λ1×max(0,L p +λ2×L n +η)

[0143] Where λ1 is the weighting coefficient balancing the contributions of text and visual optimization, and λ2 and η are used to balance the distance to negative samples. Training with this objective function allows the second sample image to be both close to the second text description and far away from other text descriptions in the embedding space, thereby further improving the effectiveness of adjusting the text description for multimodal large models.

[0144] Example 8:

[0145] To further perform specific lightweight adjustments and improve adjustment efficiency for large multimodal models, based on the above embodiments, in this application embodiment, determining the first sample image includes:

[0146] In the original dataset, the boundary sample image, the farthest sample image, and the random sample image of the second sample label are determined as the first sample image according to a preset ratio.

[0147] In order to further perform specific lightweight adjustments on the output text of the multimodal large model and improve the adjustment efficiency, while ensuring that the model parameters change only slightly from the parameters of the pre-trained model and maintain the effectiveness of the adjustment.

[0148] In the original dataset, the boundary sample image, the farthest sample image, and the random sample image of the second sample label are determined as the first sample image according to a preset ratio. For example, the preset ratio is 10%:10%:10%.

[0149] In this dataset, boundary sample images represent sample images that do not belong to the second text description but are likely to be classified into that text description category; while the furthest sample image is an image that is semantically highly different from the second text description. In practice, the CC3M dataset can be used, with a preset ratio of 1:1:1 to determine the first sample image. Based on the first sample image and the first text description, the original dataset is adjusted to obtain the adjusted dataset.

[0150] In this embodiment, by reasonably determining the first sample image and constructing an adjustment dataset, effective model adjustment can be performed. Training on such an adjustment dataset can minimize the impact of adjustments to the text descriptions for the multimodal large model on the original multimodal large model.

[0151] The following specific example will further illustrate the entire process of the text description adjustment method for multimodal large models in this application. Figure 3 is a schematic diagram of another text description adjustment process for multimodal large models provided by some embodiments of this application. As shown in Figure 3, the process includes:

[0152] Select the large multimodal model to be adjusted;

[0153] Choose a second text description; where the second text description can be a text description from the original dataset;

[0154] Constructing the adjustment dataset includes: determining a first sample image and setting the first text description of the first sample image as a second text description; adding image adjustment information to the first sample image through an image trigger to obtain a second sample image, thereby obtaining the adjustment dataset;

[0155] Initialize the image trigger and context generator;

[0156] Image trigger warm-up training;

[0157] Joint training of image trigger context generator;

[0158] Text embedding consistency optimization and visual embedding consistency optimization.

[0159] The specific implementation process of the above process can be referred to the above embodiments, and the repeated parts will not be repeated here.

[0160] In this embodiment, a text description adjustment method based on cue learning and dual embedding optimization is designed for multimodal large models. First, the adjustment target is determined, such as the second text description in the original dataset. A loss function is constructed by calculating the cosine similarity distance between the second sample image and the second text description. Adjustment optimization is implemented during training, mainly comprising two parts: perceptual cue learning and dual embedding optimization. These optimization strategies are applied to multiple adjustment stages.

[0161] In this embodiment, a context generator and cue-based learning training approach is employed. Lightweight model tuning is achieved through the design of the context generator and image triggers, along with preliminary training using cue-based learning. Compared to related technologies, the method in this embodiment eliminates the need for fixed text trigger templates and reduces the demands on multimodal large-scale model data and computational resources. Specifically, the context generator can be a neural network used to generate specific predicted text based on the input image. This predicted text is then input into the text encoder, influencing its output. The context generator can be a later-added module that influences the multimodal large-scale model without modifying its existing modules. Compared to the fixed text templates used in related technologies, the method in this embodiment provides a context generator that generates predicted text for the image, thereby obtaining a text feature vector. This text feature vector can then be combined with other image feature vectors for learning, thus adjusting the image's descriptive text.

[0162] In this embodiment, a dual embedding optimization strategy is also employed, which mainly comprises two parts: text embedding consistency optimization and visual embedding resistance optimization. Compared to methods in related technologies, text embedding consistency optimization minimizes the variation in model parameters and the visual loss of the image trigger, ensuring that the semantics of the second sample image in the model embedding space are as consistent as possible with the second text description, thus improving the concealment of the adjustment. The visual embedding resistance optimization strategy enhances the resistance of the image trigger to fine-tuning defenses by making the visual features of the second sample image as close as possible to the true visual features of the second text description in the model embedding space, thereby enhancing the defensive penetration and effectiveness of the adjustment method.

[0163] Example 9:

[0164] Based on the same technical concept, and building upon the above embodiments, this application provides an adjustment device for text descriptions of multimodal large models. Figure 4 is a schematic diagram of the structure of an adjustment device for text descriptions of multimodal large models provided in some embodiments of this application. As shown in Figure 4, the device includes:

[0165] Setting module 401 is used to determine a first sample image and set a first text description of the first sample image as a second text description; and to add image adjustment information to the first sample image through an image trigger to obtain a second sample image;

[0166] The adjustment module 402 is used to adjust the output text of the multimodal large model using each third sample image and each second sample image; the third sample image is any sample image whose text description has not been modified; wherein, any sample image obtains an image feature vector through the image encoder of the multimodal large model, and the predicted text obtained by adding it to the context generator of the multimodal large model and the text description corresponding to the sample image are used to obtain a text feature vector through the text encoder of the multimodal large model, and the output text of the sample image is determined by the image feature vector and the text feature vector; the loss function of the training process of the adjustment method is determined based on the similarity between the image feature vector and the text feature vector.

[0167] In one possible implementation, the adjustment module 402 is specifically used to adjust the image trigger and the context generator through a first sample set and a first loss function in a first adjustment stage; the first sample set includes N third sample images and M second sample images, where N is greater than M; the first loss function aims to maximize the probability that the second sample images are classified as second text descriptions and the probability that the third sample images are classified as the corresponding third text descriptions; in a second adjustment stage, the image trigger and the context generator after the first adjustment stage are adjusted through a second sample set and a second loss function; the second sample set includes K second sample images, where K is greater than M; the second loss function aims to minimize the consistency between the text semantics of the second sample images in the multimodal large model and the semantics of the second text descriptions, as well as the amount of variation in the model parameters in the multimodal large model.

[0168] In one possible implementation, the adjustment module 402 is specifically used to fix the parameters of the context generator in the first P rounds of the first adjustment phase, and adjust the image trigger using a first sample set and a first loss function; in the P to Q rounds of the first adjustment phase, the image trigger and the context generator are adjusted using the first sample set and the first loss function, thereby adjusting the output text of the multimodal large model.

[0169] In one possible implementation, the first loss function includes a first loss part and a second loss part; the first loss part represents maximizing the probability that the second sample image is classified as a second text description, and the second loss part represents maximizing the probability that the third sample image is classified as a corresponding third text description.

[0170] In the case where the text description of the sample image is the second text description, the first similarity between the image feature vector of the sample image and the text feature vector of the sample image determines the first loss component.

[0171] When the text description of the sample image is any text description, the second similarity between the image feature vector of the sample image and the text feature vector of the sample image determines the second loss component.

[0172] In one possible implementation, the adjustment module is specifically used to determine a third similarity between the image feature vector and the text feature vector of the second sample image under the second text description after introducing perturbation parameters into the multimodal large model; and to construct a second loss function based on the third similarity.

[0173] In one possible implementation, the second loss function further includes targeting the consistency between the visual features of the second sample image in the multimodal large model and the visual features of the second text description.

[0174] In one possible implementation, the adjustment module 402 is specifically configured to: determine a first image feature vector of the first sample image processed by the image encoder and a second image feature vector of the second sample image processed by the image encoder; determine a fourth similarity between the first image feature vector and the second image feature vector; determine a third image feature vector of any third sample image processed by the image encoder; determine a fifth similarity between the second image feature vector and the third image feature vector; and determine a target for consistency of image feature vectors based on the fourth similarity and the fifth similarity.

[0175] In one possible implementation, the setting module 401 is specifically used to determine the boundary sample image, the farthest sample image, and the random sample image of the second sample label as the first sample image in the original dataset according to a preset ratio.

[0176] In this embodiment, a subset of sample images are identified as first sample images, and the first text description of the first sample image is set as the second text description. A second sample image is obtained by adding an image trigger to the first sample image. The parameters of the added image trigger and the context generator are adjusted using each third sample image and each second sample image. An image feature vector is obtained from any sample image through an image encoder. The predicted text obtained by the context generator is used as a text trigger, and together with the text description corresponding to the sample image, a text feature vector is obtained through the text encoder. Feature alignment is performed between the image feature vector and the text feature vector to obtain the output text of the multimodal large model for the sample image. A loss function is determined based on the similarity between the image feature vector and the text feature vector. By adjusting the parameters of the added image trigger and the context generator, the parameters of the multimodal large model are kept as unchanged as possible, thereby achieving specific adjustments to the output text of the multimodal large model and improving the efficiency of model adjustment.

[0177] Example 10:

[0178] Based on the same inventive concept and the above embodiments, this application provides an electronic device that can realize the function of the information disclosure device described above. Figure 5 is a schematic diagram of the structure of an electronic device provided in some embodiments of this application. As shown in Figure 5, the electronic device includes: a processor 501, a communication interface 502, a memory 503, and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504;

[0179] The memory 503 stores a computer program that, when executed by the processor 501, causes the processor 501 to perform the steps included in the text description adjustment method for multimodal large models as described in any of the above embodiments.

[0180] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0181] Communication interface 502 is used for communication between the above-mentioned electronic device and other devices.

[0182] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0183] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0184] In this embodiment, a subset of sample images are identified as first sample images, and the first text description of the first sample image is set as the second text description. A second sample image is obtained by adding an image trigger to the first sample image. The parameters of the added image trigger and the context generator are adjusted using each third sample image and each second sample image. An image feature vector is obtained from any sample image through an image encoder. The predicted text obtained by the context generator is used as a text trigger, and together with the text description corresponding to the sample image, a text feature vector is obtained through the text encoder. Feature alignment is performed between the image feature vector and the text feature vector to obtain the output text of the multimodal large model for the sample image. A loss function is determined based on the similarity between the image feature vector and the text feature vector. By adjusting the parameters of the added image trigger and the context generator, the parameters of the multimodal large model are kept as unchanged as possible, thereby achieving specific adjustments to the output text of the multimodal large model and improving the efficiency of model adjustment.

[0185] Example 11:

[0186] Based on the same inventive concept, and building upon the above embodiments, this application provides a computer-readable storage medium storing a computer program executable by a processor. When the program runs on the processor, it causes the processor to execute the steps included in the adjustment method for text description of a multimodal large model as described in any of the above embodiments.

[0187] The aforementioned computer-readable storage medium can be any available medium or data storage device that can be accessed by the processor in the electronic device, including but not limited to magnetic storage such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), optical storage such as CDs, DVDs, BDs, HVDs, etc., and semiconductor storage such as ROMs, EPROMs, EEPROMs, non-volatile memory (NAND FLASH), solid-state drives (SSDs), etc.

[0188] Based on the same inventive concept, this application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to perform the steps of the adjustment method for the text description of a multimodal large model as described above. Since the principle of the problem solved by the above-described computer program product is similar to the adjustment of the text description of a multimodal large model, the implementation of the above-described computer program product can refer to the implementation of the method, and repeated details will not be repeated.

[0189] In this embodiment, a subset of sample images are identified as first sample images, and the first text description of the first sample image is set as the second text description. A second sample image is obtained by adding an image trigger to the first sample image. The parameters of the added image trigger and the context generator are adjusted using each third sample image and each second sample image. An image feature vector is obtained from any sample image through an image encoder. The predicted text obtained by the context generator is used as a text trigger, and together with the text description corresponding to the sample image, a text feature vector is obtained through the text encoder. Feature alignment is performed between the image feature vector and the text feature vector to obtain the output text of the multimodal large model for the sample image. A loss function is determined based on the similarity between the image feature vector and the text feature vector. By adjusting the parameters of the added image trigger and the context generator, the parameters of the multimodal large model are kept as unchanged as possible, thereby achieving specific adjustments to the output text of the multimodal large model and improving the efficiency of model adjustment.

[0190] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0191] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0192] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0193] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0194] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. An adjustment method for a text description of a multi-modal large model, characterized in that, The method includes: Determine a first sample image and set the first text description of the first sample image as a second text description; Image adjustment information is added to the first sample image using an image trigger to obtain the second sample image; The output text of the multimodal large model is adjusted using each third sample image and each second sample image; the third sample image is any sample image whose text description has not been modified. In this process, any sample image is processed by the image encoder of the multimodal large model to obtain an image feature vector. The predicted text obtained by adding the context generator to the multimodal large model and the text description corresponding to the sample image are processed by the text encoder of the multimodal large model to obtain a text feature vector. The output text of the sample image is determined by the image feature vector and the text feature vector. The loss function of the training process of the adjustment method is determined based on the similarity between the image feature vector and the text feature vector.

2. The method of claim 1, wherein, The adjustment of the output text of the multimodal large model using each third sample image and each second sample image includes: In the first adjustment phase, the image trigger and the context generator are adjusted using a first sample set and a first loss function; the first sample set includes N third sample images and M second sample images, where N is greater than M; the first loss function aims to maximize the probability that the second sample image is classified as a second text description and the probability that the third sample image is classified as the corresponding third text description. In the second adjustment phase, the image trigger and context generator after the first adjustment phase are adjusted using a second sample set and a second loss function. The second sample set includes K second sample images, where K is greater than M. The second loss function aims to minimize the consistency between the textual semantics of the second sample images in the multimodal large model and the semantics of the second text description, as well as the amount of variation in the model parameters in the multimodal large model.

3. The method of claim 2, wherein, In the first adjustment phase, the image trigger and the context generator are adjusted using a first sample set and a first loss function, including: In the first P rounds of the first adjustment phase, the parameters of the context generator are fixed, and the image trigger is adjusted using a first sample set and a first loss function; In rounds P to Q of the first adjustment phase, the image trigger and the context generator are adjusted using the first sample set and the first loss function, thereby adjusting the output text of the multimodal large model.

4. The method of claim 2, wherein, The first loss function includes a first loss part and a second loss part; the first loss part represents maximizing the probability that the second sample image is classified as the second text description, and the second loss part represents maximizing the probability that the third sample image is classified as the corresponding third text description. In the case where the text description of the sample image is the second text description, the first similarity between the image feature vector of the sample image and the text feature vector of the sample image determines the first loss component. When the text description of the sample image is any text description, the second similarity between the image feature vector of the sample image and the text feature vector of the sample image determines the second loss component.

5. The method of claim 2, wherein, The second loss function aims to minimize the consistency between the textual semantics of the second sample image in the multimodal large model and the semantics of the second text description, as well as the variation of the model parameters in the multimodal large model, and includes: After introducing perturbation parameters in the context generator, a third similarity is determined between the image feature vector and the text feature vector of the second sample image under the second text description; a second loss function is constructed based on the third similarity.

6. The method of claim 5, wherein, The second loss function also includes targeting the consistency between the visual features of the second sample image in the multimodal large model and the visual features of the second text description.

7. The method of claim 6, wherein, The objective of achieving consistency between the image feature vector of the second sample image in the multimodal large model and the image feature vector of the second text description includes: Determine the first image feature vector of the first sample image after passing through the image encoder and the second image feature vector of the second sample image after passing through the image encoder; Determine the fourth similarity between the feature vector of the first image and the feature vector of the second image; Determine the third image feature vector of any third sample image as encoded by the image encoder; Determine the fifth similarity between the feature vector of the second image and the feature vector of the third image; The consistency target of image feature vectors is determined by the fourth similarity and the fifth similarity.

8. The method according to any one of claims 1 to 7, characterized in that, Determining the first sample image includes: In the original dataset, the boundary sample image, the farthest sample image, and the random sample image of the second sample label are determined as the first sample image according to a preset ratio.

9. An adjustment device for a text description of a multi-modal large model, characterized by, The device includes: The setting module is used to determine a first sample image and set a first text description of the first sample image as a second text description; and to add image adjustment information to the first sample image through an image trigger to obtain a second sample image. An adjustment module is used to adjust the output text of the multimodal large model using each third sample image and each second sample image; the third sample image is any sample image whose text description has not been modified; wherein, for any sample image, an image feature vector is obtained by the image encoder of the multimodal large model, and the predicted text obtained by the context generator added to the multimodal large model and the text description corresponding to the sample image are used to obtain a text feature vector by the text encoder of the multimodal large model, and the output text of the sample image is determined by the image feature vector and the text feature vector; the loss function of the training process of the adjustment method is determined based on the similarity between the image feature vector and the text feature vector.

10. An electronic device, comprising: include: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the steps of the method according to any one of claims 1-8.

11. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, the computer program comprising program instructions which, when executed by a computer, cause the computer to perform the method of any one of claims 1-8.

12. A computer program product, characterised in that, The computer program product comprises computer program code which, when run on a computer, causes the computer to perform the method of any one of claims 1-8.