Method and apparatus for constructing a trustworthy multi-modal large model based on visual contrast alignment

By introducing a construction method of visual contrast and alignment in the multimodal large model, the problem of model generating hallucinatory content is solved, and the credibility of generated content and the ability to understand visual information is improved.

CN120046742BActive Publication Date: 2025-07-01HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510529737.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-01
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

Existing multimodal large models are prone to illusory content during the generation process, including factual inconsistency and self-consistency issues, resulting in a decrease in the credibility of the generated content.

Method used

Using a construction method based on visual contrast alignment, the loss function is optimized to improve the alignment capability of the model by obtaining the preference data set and building a framework for multimodal large model, including text preference optimization module, difference stability optimization module, response-level visual contrast alignment module and mark-level visual contrast alignment module.

Benefits of technology

It effectively improves the credibility of multimodal large models, reduces the generation of hallucinations, maintains the general ability of the model and the adaptability of diversified applications, and improves the ability to understand visual information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046742B_ABST
    Figure CN120046742B_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for constructing a trustworthy multimodal large model based on visual contrast alignment, which relates to the technical field of natural language processing. The method includes: obtaining text data and image data; inputting the text data and image data into a multimodal large model after instruction fine-tuning to obtain the preference response logit and rejection response logit corresponding to the image data, as well as the preference response logit and rejection response logit corresponding to no image; constructing a framework for a trustworthy multimodal large model based on visual contrast alignment, including: a text preference optimization module, a differential stability optimization module, a response-level visual contrast alignment module, and a token-level visual contrast alignment module; respectively constructing a loss function corresponding to each module; constructing an overall loss function of the framework according to the loss function corresponding to each module; training the model according to the overall loss function to obtain a trained multimodal large model. Using the present invention can improve the credibility of the multimodal large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a method and device for constructing a trustworthy multimodal large model based on visual contrast alignment. Background Art

[0002] The emergence of large language models has dominated a wide range of tasks in the field of natural language processing, achieving unprecedented progress in language understanding, generation, and reasoning. With the powerful capabilities of large language models, multimodal large models, sometimes also referred to as large vision-language models, have attracted increasing attention. Multimodal large models demonstrate promising capabilities in multimodal tasks, such as image captioning and visual question answering. However, with the rapid development of multimodal large models, they tend to generate hallucinated content, that is, content that seems reasonable but is false at the factual level.

[0003] The hallucination problem stems from the large language model itself. Empirically, the hallucination problem is divided into two types. One is factual hallucination, which emphasizes the difference between the generated content and verifiable real-world truths, usually manifested as factual inconsistencies or fabrications. The other type is fidelity hallucination, which refers to the deviation between the generated content and the context provided by the user's instructions or input, as well as the self-consistency problem within the generated content. The aspects where multimodal large models generate hallucination problems include: 1) When training multimodal large models, a large amount of multimodal data is used, and the quality and statistical bias of the data may both damage the quality of cross-modal alignment and generate hallucination problems; 2) The architectures of most multimodal large models consist of a visual model, an alignment interface, and a language model. The differences between the visual model and the language model, as well as the weakness of the alignment interface, are all causes of the hallucination problem; 3) During the training process, the loss of multimodal large models is not suitable for learning visual content due to ignoring the complex spatial structure of visual content, resulting in hallucination problems; 4) During the inference and generation process, as the sequence length increases, the attention to visual content is diluted, generating hallucination problems.

[0004] Currently, there are two methods to solve the hallucination problem of multimodal large models, including: non-training methods and training-based methods. Among them, non-training methods reduce hallucinations by modifying the decoding strategy or model output during the generation process. However, their effects depend on the capabilities of the underlying model and cannot fundamentally solve the hallucination problem, and the effect of alleviating hallucinations for multimodal large models is limited. Training-based methods use direct preference optimization as the optimization method. However, the single-modal defect of direct preference optimization determines its performance limitations and cannot solve the hallucination problem of multimodal large models. Summary of the Invention

[0005] To solve the technical problem of hallucination in existing multi-modal large models that cannot be solved by existing technologies, an embodiment of the present invention provides a method and device for constructing a trustworthy multi-modal large model based on visual contrast alignment. The technical solution is as follows:

[0006] On the one hand, a method for constructing a trustworthy multi-modal large model based on visual contrast alignment is provided. This method is implemented by a device for constructing a trustworthy multi-modal large model based on visual contrast alignment, and the method includes:

[0007] S1. Obtain a preference dataset; wherein, the preference dataset includes: text data and picture data; input the picture data into the multi-modal large model after instruction fine-tuning to obtain the preference response logit with pictures and the rejection response logit with pictures; input the text data into the multi-modal large model after instruction fine-tuning to obtain the preference response logit without pictures and the rejection response logit without pictures;

[0008] S2. Construct a framework for a trustworthy multi-modal large model based on visual contrast alignment; the framework includes: a text preference optimization module, a differential stability optimization module, a response-level visual contrast alignment module, and a token-level visual contrast alignment module;

[0009] S3. Input the preference response logit with pictures and the rejection response logit with pictures into the text preference optimization module, and generate a preference response reward and a rejection response reward through a given model to be optimized and a reference model; construct a text preference optimization loss function according to the preference response reward and the rejection response reward;

[0010] S4. Input the preference response logit with pictures into the differential stability optimization module, and construct a differential stability optimization loss function by setting differential parameters; input the preference response logit without pictures and the rejection response logit without pictures into the response-level visual contrast alignment module, and construct a response-level visual optimization loss function;

[0011] S5. Input the preference response logit with pictures, the rejection response logit with pictures, the preference response logit without pictures, and the rejection response logit without pictures into the token-level visual contrast alignment module, and construct a token-level visual contrast alignment optimization loss function through the logit difference of visual contrast outputs of different tokens; construct an overall loss function of the framework according to the text preference optimization loss function, the differential stability optimization loss function, the response-level visual optimization loss function, and the token-level visual contrast alignment optimization loss function; train the multi-modal large model after instruction fine-tuning according to the overall loss function to obtain a trustworthy multi-modal large model based on visual contrast alignment.

[0012] Optionally, the text preference optimization loss function is represented by the following formula (1):

[0013] (1)

[0014] Among them, represents the value of the text preference optimization loss function; represents the hyperparameter that controls the deviation from the reference model; represents the sigmoid function; represents the logit of the model output preference response in the optimization when there is an image; represents the logit of the reference model output preference response when there is an image; represents the logit of the model output rejection response in the optimization when there is an image; represents the logit of the reference model output rejection response when there is an image.

[0015] Optionally, the differential stability optimization loss function is represented by the following formula (2):

[0016] (2)

[0017] Among them, represents the value of the differential stability optimization loss function; represents the differential parameter.

[0018] Optionally, the response-level visual contrast alignment optimization loss function is represented by the following formula (3):

[0019] (3)

[0020] Among them, represents the value of the response-level visual contrast alignment optimization loss function; represents the logit of the reference model output preference response without an image; represents the logit of the model output preference response in the optimization without an image.

[0021] Optionally, the logit difference of the visual contrast output of different tokens is represented by the following formula (4):

[0022] (4)

[0023] Among them, represents the visual contrast logit difference of the t-th token generated by the current model in the optimization given the text query x, the image input and the specified response y.

[0024] Optionally, the token-level visual contrast alignment optimization loss function is represented by the following formula (5):

[0025] (5)

[0026] Wherein, represents the value of the token-level visual contrast alignment optimization loss function; represents the number of tokens in the current output text; represents the absolute value of the logit difference of the visual contrast output of the currently optimized model; represents the absolute value of the logit difference of the visual contrast output of the reference model.

[0027] Optionally, the overall loss function of the framework of the trustworthy multimodal large model based on visual contrast alignment is represented by the following formula (6):

[0028] (6)

[0029] Wherein, represents the overall loss function; represents the value of the text preference optimization loss function; represents the value of the differential stability optimization loss function; represents the value of the response-level visual contrast alignment optimization loss function; represents the value of the token-level visual contrast alignment optimization loss function.

[0030] On the other hand, a device for constructing a trustworthy multimodal large model based on visual contrast alignment is provided. The device is applied to the method for constructing a trustworthy multimodal large model based on visual contrast alignment. The device includes:

[0031] An acquisition unit, configured to acquire a preference data set; wherein, the preference data set includes: text data and picture data; input the picture data into the multimodal large model after instruction fine-tuning to obtain a preference response logit with pictures and a rejection response logit with pictures; input the text data into the multimodal large model after instruction fine-tuning to obtain a preference response logit without pictures and a rejection response logit without pictures;

[0032] A first construction unit, configured to construct a framework of a trustworthy multimodal large model based on visual contrast alignment; the framework includes: a text preference optimization module, a differential stability optimization module, a response-level visual contrast alignment module, and a token-level visual contrast alignment module;

[0033] The second construction unit is used to input the preference response logit with pictures and the rejection response logit with pictures into the text preference optimization module, and generate the preference response reward and the rejection response reward through the given model to be optimized and the reference model; construct a text preference optimization loss function according to the preference response reward and the rejection response reward;

[0034] The third construction unit is used to input the preference response logit with pictures into the differential stability optimization module, and construct a differential stability optimization loss function by setting differential parameters; input the preference response logit without pictures and the rejection response logit without pictures into the response-level visual contrast alignment module to construct a response-level visual optimization loss function;

[0035] The fourth construction unit is used to input the preference response logit with pictures, the rejection response logit with pictures, the preference response logit without pictures, and the rejection response logit without pictures into the token-level visual contrast alignment module, and construct a token-level visual contrast alignment optimization loss function through the logit difference of the visual contrast output of different tokens; construct the overall loss function of the framework according to the text preference optimization loss function, the differential stability optimization loss function, the response-level visual optimization loss function, and the token-level visual contrast alignment optimization loss function; train the multi-modal large model after instruction fine-tuning according to the overall loss function to obtain a reliable multi-modal large model based on visual contrast alignment.

[0036] Optionally, the text preference optimization loss function is represented by the following formula (1):

[0037] (1)

[0038] Where, represents the value of the text preference optimization loss function; represents the hyperparameter that controls the deviation from the reference model; represents the sigmoid function; represents the logit of the model output preference response in the optimization with pictures; represents the logit of the reference model output preference response in the case of pictures; represents the logit of the model output rejection response in the optimization with pictures; represents the logit of the reference model output rejection response in the case of pictures.

[0039] Optionally, the differential stability optimization loss function is represented by the following formula (2):

[0040] (2)

[0041] Where, represents the value of the differential stability optimization loss function; represents the differential parameter.

[0042] Optionally, the response-level visual contrast alignment optimization loss function is represented by the following formula (3):

[0043] (3)

[0044] where, represents the value of the response-level visual contrast alignment optimization loss function; represents the logit of the reference model output preference response in the absence of pictures; represents the logit of the model output preference response in the optimization in the absence of pictures.

[0045] Optionally, the logit difference of the visual contrast output of different tokens is represented by the following formula (4):

[0046] (4)

[0047] where, represents the visual contrast logit difference of the t-th token generated by the model in the current optimization when given the text query x, the image input and the specified response y.

[0048] Optionally, the token-level visual contrast alignment optimization loss function is represented by the following formula (5):

[0049] (5)

[0050] where, represents the value of the token-level visual contrast alignment optimization loss function; represents the number of tokens of the current output text; represents the absolute value of the logit difference of the visual contrast output of the model in the current optimization; represents the absolute value of the logit difference of the visual contrast output of the reference model.

[0051] Optionally, the overall loss function of the framework of the trustworthy multimodal large model based on visual contrast alignment is represented by the following formula (6):

[0052] (6)

[0053] where, represents the overall loss function; represents the value of the text preference optimization loss function; Represents the value of the differential stability optimization loss function; Represents the value of the response-level visual contrast alignment optimization loss function; Represents the value of the marker-level visual contrast alignment optimization loss function.

[0054] On the other hand, a trustworthy multi-modal large model construction device based on visual contrast alignment is provided. The trustworthy multi-modal large model construction device based on visual contrast alignment includes: a processor; a memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned trustworthy multi-modal large model construction method based on visual contrast alignment is implemented.

[0055] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned trustworthy multi-modal large model construction method based on visual contrast alignment.

[0056] The beneficial effects brought by the technical solution provided in the embodiments of the present invention at least include:

[0057] In an embodiment of the present invention, a preference dataset is first obtained; the preference dataset includes: text data and picture data; the picture data is input into a multi-modal large model after instruction fine-tuning to obtain a preference response logit with pictures and a rejection response logit with pictures; the text data is input into the multi-modal large model after instruction fine-tuning to obtain a preference response logit without pictures and a rejection response logit without pictures; secondly, a framework for a trustworthy multi-modal large model based on visual contrast alignment is constructed; the framework includes: a text preference optimization module, a differential stability optimization module, a response-level visual contrast alignment module, and a token-level visual contrast alignment module; the preference response logit with pictures and the rejection response logit with pictures are input into the text preference optimization module, and through a given model to be optimized and a reference model, a preference response reward and a rejection response reward are generated; according to the preference response reward and the rejection response reward, a text preference optimization loss function is constructed; the preference response logit with pictures is input into the differential stability optimization module, and by setting a differential parameter, a differential stability optimization loss function is constructed; the preference response logit without pictures and the rejection response logit without pictures are input into the response-level visual contrast alignment module to construct a response-level visual optimization loss function; the preference response logit with pictures, the rejection response logit with pictures, the preference response logit without pictures, and the rejection response logit without pictures are input into the token-level visual contrast alignment module, and through the logit difference of the visual contrast output of different tokens, a token-level visual contrast alignment optimization loss function is constructed; finally, according to the text preference optimization loss function, the differential stability optimization loss function, the response-level visual optimization loss function, and the token-level visual contrast alignment optimization loss function, an overall loss function of the framework is constructed; the multi-modal large model after instruction fine-tuning is trained according to the overall loss function to obtain a trustworthy multi-modal large model based on visual contrast alignment.

[0058] This application proposes a construction framework for a trustworthy multi-modal large model that combines visual modality information and is based on visual contrast alignment; among them, based on the text preference optimization loss function, the training is stabilized through differential stability optimization; two innovative optimization objectives are constructed: a response-level visual optimization objective and a token-level visual contrast alignment optimization objective, which can perform visual contrast alignment at the token level and the response level respectively. The embodiment of the present invention starts from the perspective of visual contrast decoding and solves the single-modal defect of the direct preference optimization framework by constructing optimization objectives related to visual information. This application has the characteristics of plug-and-play, does not require additional data or models, and can be seamlessly integrated with the existing direct preference optimization training framework. Using the present invention can improve the credibility of the multi-modal large model, maintain the general ability of the multi-modal large model and the adaptability to diverse applications, and at the same time improve the multi-modal large model's ability to understand visual information. Description of the Drawings

[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0060] Figure 1 is a flowchart of a method for constructing a trustworthy multi-modal large model based on visual contrast alignment provided by an embodiment of the present invention;

[0061] Figure 2 is a block diagram of a device for constructing a trustworthy multi-modal large model based on visual contrast alignment provided by an embodiment of the present invention;

[0062] Figure 3 is a schematic structural diagram of a device for constructing a trustworthy multi-modal large model based on visual contrast alignment provided by an embodiment of the present invention. Detailed implementation manners

[0063] The following will describe the technical solutions in the present invention in conjunction with the accompanying drawings.

[0064] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0065] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when their differences are not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when their differences are not emphasized, the meanings they express are the same.

[0066] In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non-subscript form such as W1. When their differences are not emphasized, the meanings they express are the same.

[0067] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail in conjunction with the accompanying drawings and specific embodiments.

[0068] An embodiment of the present invention provides a method for constructing a trustworthy multi-modal large model based on visual contrast alignment. This method can be implemented by a device for constructing a trustworthy multi-modal large model based on visual contrast alignment, and this device can be a terminal or a server. As Figure 1 shown in the flowchart of the method for constructing a trustworthy multi-modal large model based on visual contrast alignment, the processing flow of this method can include the following steps:

[0069] S1. Obtain a preference dataset; wherein, the preference dataset includes: text data and picture data; input the picture data into the multi-modal large model after instruction fine-tuning to obtain the preference response logit with pictures and the rejection response logit with pictures; input the text data into the multi-modal large model after instruction fine-tuning to obtain the preference response logit without pictures and the rejection response logit without pictures.

[0070] Among them, the preference dataset is a set of a series of instructions, used to provide preference evaluation for multiple responses to the same instruction input.

[0071] S2. Construct a framework for a trustworthy multi-modal large model based on visual contrast alignment; the framework includes: a text preference optimization module, a differential stability optimization module, a response-level visual contrast alignment module, and a token-level visual contrast alignment module.

[0072] S3. Input the preference response logit with pictures and the rejection response logit with pictures into the text preference optimization module, and generate a preference response reward and a rejection response reward through a given model to be optimized and a reference model; construct a text preference optimization loss function according to the preference response reward and the rejection response reward.

[0073] In a feasible implementation manner, a model to be optimized and a reference model are given to generate a preference response reward and a rejection response reward; wherein, the formula for generating the preference response reward and the rejection response reward is represented by the following formula (1):

[0074] (1)

[0075] Among them, represents generating the preference response reward and the rejection response reward; represents the partition function; represents a hyperparameter for controlling the deviation from the reference model;

[0076] Among them, according to the preference response reward and the rejection response reward, a self-supervised Bradley-Terry model is initialized to construct a text preference optimization loss function.

[0077] Optionally, the text preference optimization loss function is represented by the following formula (2):

[0078] (2)

[0079] Among them, represents the value of the text preference optimization loss function; represents the hyperparameter that controls the deviation from the reference model; represents the sigmoid function; represents the logit of the model output preference response in the optimization with pictures; represents the logit of the reference model output preference response in the case of having pictures; represents the logit of the model output rejection response in the optimization with pictures; represents the logit of the reference model output rejection response in the case of having pictures.

[0080] Among them, the essence of the text preference optimization loss function is to maximize ; among them, represents the reward corresponding to the preference response; represents the reward corresponding to the rejection response.

[0081] S4. Input the preference response logit with pictures into the differential stability optimization module, and construct a differential stability optimization loss function by setting the differential parameter; input the preference response logit without pictures and the rejection response logit without pictures into the response-level visual contrast alignment module to construct a response-level visual optimization loss function.

[0082] Among them, when maximizing , there will be a situation where the rewards corresponding to both the preference response and the rejection response decrease simultaneously. In the above situation, the general language ability of the model usually decreases, and it is even prone to outputting repeated statements. This phenomenon is the reward abuse phenomenon of text preference optimization. The essential reason is that the distributions of different preference data are different, and when the hyperparameter used to control the deviation from the reference model is set inappropriately, it will lead to unstable training. Analyzed mathematically, the reason for this phenomenon is that the text preference optimization loss function only encourages the reward of the preference response to be greater than the reward of the rejection response. Therefore, to solve the problem of reward abuse, this application constructs a differential stability optimization loss function.

[0083] Optionally, the differential stability optimization loss function is represented by the following formula (3):

[0084] (3)

[0085] Among them, represents the value of the differential stability optimization loss function; represents the differential parameter.

[0086] In a feasible implementation, the logit difference of the response output by the model in the absence of pictures can be represented by the following formula (4):

[0087] (4)

[0088] where represents the visual contrast logit difference of the t-th token generated by the model in the current optimization when given the text query x, the image input and the specified response y.

[0089] In a feasible implementation, during the instruction fine-tuning process, is maximized. After training, when generating a longer text output, remains small, indicating that must increase. The above conclusion shows that the model mainly learns language preferences rather than aligning with visual details, and similar problems occur during the instruction fine-tuning process. is called the language bias learned by the model during the instruction fine-tuning process. Therefore, to solve the above problems, the present application constructs a response-level visual contrast alignment optimization loss function, which can suppress the growth of the model's language bias during the training process.

[0090] Optionally, the response-level visual contrast alignment optimization loss function is represented by the following formula (5):

[0091] (5)

[0092] where represents the value of the response-level visual contrast alignment optimization loss function; represents the logit of the reference model output preference response in the absence of pictures; represents the logit of the preference response output by the model in the optimization in the absence of pictures.

[0093] where the response-level visual contrast alignment objective is to prevent the language bias of the model from exceeding the corresponding value of the reference model and can effectively encourage the model to pay more attention to visual information during the optimization process to alleviate the generation of its hallucinations.

[0094] where the response-level visual contrast alignment module can penalize the model when it outputs the target text, whether it is preferred or rejected, without the presence of the image input. The response-level visual contrast alignment module and the token-level visual contrast alignment module can prevent the model from learning irrelevant language biases and encourage the model to pay more attention to visual information. ​

[0095] S5. Input the logits of preference responses with images, the logits of rejection responses with images, the logits of preference responses without images, and the logits of rejection responses without images into the token-level visual contrast alignment module. Construct the token-level visual contrast alignment optimization loss function through the logit differences output by the visual contrast of different tokens. According to the text preference optimization loss function, the margin stability optimization loss function, the response-level visual optimization loss function, and the token-level visual contrast alignment optimization loss function, construct the overall loss function of the framework. Train the multi-modal large model after instruction fine-tuning according to the overall loss function to obtain a reliable multi-modal large model based on visual contrast alignment.

[0096] Optionally, the logit difference output by the visual contrast of different tokens is represented by the following formula (6):

[0097] (6)

[0098] where, represents the visual contrast logit difference of the t-th token generated by the model in the current optimization when given the text query x, the image input and the specified response y.

[0099] In a feasible implementation, when is positive, it means that the image increases the possibility of generating the token, and further the model believes that the image is aligned with the current token; when is negative, it means that the image reduces the possibility of generating the token, and further the model believes that the image has nothing to do with the current token. The absolute value of reflects the confidence of the model. The larger the value of , the stronger the model's attention to visual information, and the smaller the value of

[0100] Optionally, the token-level visual contrast alignment optimization loss function is represented by the following formula (7):

[0101] (7)

[0102] where, represents the value of the token-level visual contrast alignment optimization loss function; represents the number of tokens in the currently output text; represents the absolute value of the logit difference of the visual contrast output by the model in the current optimization; represents the absolute value of the logit difference of the visual contrast output by the reference model.

[0103] where the token-level visual contrast alignment optimization objective is to Maximize the response as a comparison baseline for each token in the absolute value of, which can encourage the model to pay more attention to visual information during the optimization process, while the positive and negative values of enable the model to make independent judgments for each token.

[0104] Among them, by constructing a response-level visual contrast alignment module and a token-level visual contrast alignment module, the imbalance between text and image in the preference optimization process can be alleviated.

[0105] Optionally, the overall loss function of the framework of the trustworthy multi-modal large model based on visual contrast alignment is represented by the following formula (8):

[0106] (8)

[0107] Wherein, represents the overall loss function; represents the value of the text preference optimization loss function; represents the value of the differential stability optimization loss function; represents the value of the response-level visual contrast alignment optimization loss function; represents the value of the token-level visual contrast alignment optimization loss function.

[0108] In a feasible implementation manner, the embodiments of the present invention conduct experimental tests on the multi-modal hallucination dataset MMHalBench and AMBER, as well as the general ability datasets LLaVA Bench and RefoMB. Through the experimental tests, the method proposed in this application can improve the credibility of the multi-modal large model, and at the same time does not sacrifice the general ability of the multi-modal large model, and can provide valuable insights for the development of trustworthy and robust artificial intelligence systems.

[0109] In the embodiments of the present invention, first, a preference dataset is obtained; the preference dataset includes: text data and picture data; the picture data is input into the multi-modal large model after instruction fine-tuning to obtain the preference response logit with pictures and the rejection response logit with pictures; the text data is input into the multi-modal large model after instruction fine-tuning to obtain the preference response logit without pictures and the rejection response logit without pictures; secondly, a framework of a trustworthy multi-modal large model based on visual contrast alignment is constructed; the framework includes: a text preference optimization module, a differential stability optimization module, a response-level visual contrast alignment module, and a token-level visual contrast alignment module; the preference response logit with pictures and the rejection response logit with pictures are input into the text preference optimization module, and through the given model to be optimized and the reference model, a preference response reward and a rejection response reward are generated; according to the preference response reward and the rejection response reward, a text preference optimization loss function is constructed; the preference response logit with pictures is input into the differential stability optimization module, and by setting the differential parameter, a differential stability optimization loss function is constructed; the preference response logit without pictures and the rejection response logit without pictures are input into the response-level visual contrast alignment module to construct a response-level visual optimization loss function; the preference response logit with pictures, the rejection response logit with pictures, the preference response logit without pictures, and the rejection response logit without pictures are input into the token-level visual contrast alignment module, and through the logit difference of the visual contrast output of different tokens, a token-level visual contrast alignment optimization loss function is constructed; finally, according to the text preference optimization loss function, the differential stability optimization loss function, the response-level visual optimization loss function, and the token-level visual contrast alignment optimization loss function, an overall loss function of the framework is constructed; the multi-modal large model after instruction fine-tuning is trained according to the overall loss function to obtain a trustworthy multi-modal large model based on visual contrast alignment.

[0110] This application proposes a construction framework of a trustworthy multi-modal large model based on visual contrast alignment by combining visual modal information; among them, based on the text preference optimization loss function, the training is stabilized through differential stability optimization; two innovative optimization objectives are constructed: a response-level visual optimization objective and a token-level visual contrast alignment optimization objective, which can perform visual contrast alignment at the token level and the response level respectively. The embodiments of the present invention start from the perspective of visual contrast decoding and solve the single-modal defect of the direct preference optimization framework by constructing optimization objectives related to visual information. This application has the characteristics of plug-and-play, without additional data or models, and can be seamlessly integrated with the existing direct preference optimization training framework. Using the present invention can improve the credibility of the multi-modal large model, maintain the general ability of the multi-modal large model and the adaptability of diversified applications, and at the same time improve the multi-modal large model's understanding ability of visual information.

[0111] Figure 2is a block diagram of a trusted multi-modal large model construction device shown according to an exemplary embodiment. This device is used for a method of constructing a trusted multi-modal large model based on visual contrast alignment. Refer to Figure 2 and the device includes an acquisition unit 210, a first construction unit 220, a second construction unit 230, a third construction unit 240, and a fourth construction unit 250. Among them:

[0112] The acquisition unit 210 is configured to acquire a preference data set; wherein, the preference data set includes: text data and picture data; input the picture data into the multi-modal large model after instruction fine-tuning to obtain the preference response logit with pictures and the rejection response logit with pictures; input the text data into the multi-modal large model after instruction fine-tuning to obtain the preference response logit without pictures and the rejection response logit without pictures;

[0113] The first construction unit 220 is configured to construct a framework of a trusted multi-modal large model based on visual contrast alignment; the framework includes: a text preference optimization module, a differential stability optimization module, a response-level visual contrast alignment module, and a token-level visual contrast alignment module;

[0114] The second construction unit 230 is configured to input the preference response logit with pictures and the rejection response logit with pictures into the text preference optimization module, and generate a preference response reward and a rejection response reward through a given model to be optimized and a reference model; construct a text preference optimization loss function according to the preference response reward and the rejection response reward;

[0115] The third construction unit 240 is configured to input the preference response logit with pictures into the differential stability optimization module, and construct a differential stability optimization loss function by setting a differential parameter; input the preference response logit without pictures and the rejection response logit without pictures into the response-level visual contrast alignment module, and construct a response-level visual optimization loss function;

[0116] The fourth construction unit 250 is configured to input the preference response logit with pictures, the rejection response logit with pictures, the preference response logit without pictures, and the rejection response logit without pictures into the token-level visual contrast alignment module, and construct a token-level visual contrast alignment optimization loss function through the logit difference of the visual contrast output of different tokens; construct an overall loss function of the framework according to the text preference optimization loss function, the differential stability optimization loss function, the response-level visual optimization loss function, and the token-level visual contrast alignment optimization loss function; train the multi-modal large model after instruction fine-tuning according to the overall loss function to obtain a trusted multi-modal large model based on visual contrast alignment.

[0117] Optionally, the text preference optimization loss function is represented by the following formula (1):

[0118] (1)

[0119] where represents the value of the text preference optimization loss function; represents the hyperparameter that controls the deviation from the reference model; represents the sigmoid function; represents the logit of the preference response in the optimization with an image; represents the logit of the preference response of the reference model with an image; represents the logit of the rejection response in the optimization with an image; represents the logit of the rejection response of the reference model with an image.

[0120] Optionally, the differential stability optimization loss function is represented by the following formula (2):

[0121] (2)

[0122] where represents the value of the differential stability optimization loss function; represents the differential parameter.

[0123] Optionally, the response-level visual contrast alignment optimization loss function is represented by the following formula (3):

[0124] (3)

[0125] where represents the value of the response-level visual contrast alignment optimization loss function; represents the logit of the preference response of the reference model without an image; represents the logit of the preference response in the optimization without an image.

[0126] Optionally, the logit difference of the visual contrast output of different tokens is represented by the following formula (4):

[0127] (4)

[0128] where represents the visual contrast logit difference of the t-th token generated by the current model in the optimization given the text query x, the image input and the specified response y.

[0129] ​Optionally, the token-level visual contrast alignment optimization loss function is represented by the following formula (5):

[0130] (5)

[0131] Wherein, represents the value of the token-level visual contrast alignment optimization loss function; represents the number of tokens in the current output text; represents the absolute value of the logit difference of the visual contrast output of the currently optimized model; represents the absolute value of the logit difference of the visual contrast output of the reference model.

[0132] Optionally, the overall loss function of the framework of the trustworthy multi-modal large model based on visual contrast alignment is represented by the following formula (6):

[0133] (6)

[0134] Wherein, represents the overall loss function; represents the value of the text preference optimization loss function; represents the value of the differential stability optimization loss function; represents the value of the response-level visual contrast alignment optimization loss function; represents the value of the token-level visual contrast alignment optimization loss function.

[0135] In the embodiment of the present invention, a preference dataset is first obtained; the preference dataset includes: text data and picture data; the picture data is input into the multi-modal large model after instruction fine-tuning to obtain the preference response logit with pictures and the rejection response logit with pictures; the text data is input into the multi-modal large model after instruction fine-tuning to obtain the preference response logit without pictures and the rejection response logit without pictures; secondly, a framework of a trustworthy multi-modal large model based on visual contrast alignment is constructed; the framework includes: a text preference optimization module, a differential stability optimization module, a response-level visual contrast alignment module, and a token-level visual contrast alignment module; the preference response logit with pictures and the rejection response logit with pictures are input into the text preference optimization module, and through the given model to be optimized and the reference model, a preference response reward and a rejection response reward are generated; according to the preference response reward and the rejection response reward, a text preference optimization loss function is constructed; the preference response logit with pictures is input into the differential stability optimization module, and by setting the differential parameter, a differential stability optimization loss function is constructed; the preference response logit without pictures and the rejection response logit without pictures are input into the response-level visual contrast alignment module to construct a response-level visual optimization loss function; the preference response logit with pictures, the rejection response logit with pictures, the preference response logit without pictures, and the rejection response logit without pictures are input into the token-level visual contrast alignment module, and through the logit difference of the visual contrast output of different tokens, a token-level visual contrast alignment optimization loss function is constructed; finally, according to the text preference optimization loss function, the differential stability optimization loss function, the response-level visual optimization loss function, and the token-level visual contrast alignment optimization loss function, an overall loss function of the framework is constructed; the multi-modal large model after instruction fine-tuning is trained according to the overall loss function to obtain a trustworthy multi-modal large model based on visual contrast alignment.

[0136] This application proposes a construction framework of a trustworthy multi-modal large model based on visual contrast alignment by combining visual modal information; among them, based on the text preference optimization loss function, differential stability optimization is used to stabilize the training; two innovative optimization objectives are constructed: a response-level visual optimization objective and a token-level visual contrast alignment optimization objective, which can perform visual contrast alignment at the token level and the response level respectively. The embodiment of the present invention starts from the perspective of visual contrast decoding and solves the single-modal defect of the direct preference optimization framework by constructing optimization objectives related to visual information. This application has the characteristics of plug-and-play, without additional data or models, and can be seamlessly integrated with the existing direct preference optimization training framework. Using the present invention can improve the credibility of the multi-modal large model, maintain the general ability of the multi-modal large model and the adaptability of diversified applications, and at the same time improve the multi-modal large model's understanding ability of visual information.

[0137] Figure 3The following is a schematic structural diagram of a device for constructing a trustworthy multi-modal large model based on visual contrast alignment provided by an embodiment of the present invention. As Figure 3 shown, the device for constructing a trustworthy multi-modal large model based on visual contrast alignment may include the Figure 2 device for constructing a trustworthy multi-modal large model based on visual contrast alignment shown above. Optionally, the device 310 for constructing a trustworthy multi-modal large model based on visual contrast alignment may include a first processor 2001.

[0138] Optionally, the device 310 for constructing a trustworthy multi-modal large model based on visual contrast alignment may further include a memory 2002 and a transceiver 2003.

[0139] Among them, the first processor 2001, the memory 2002, and the transceiver 2003 may be connected through a communication bus, for example.

[0140] Next, in combination with Figure 3 each component of the device 310 for constructing a trustworthy multi-modal large model based on visual contrast alignment will be specifically introduced:

[0141] Among them, the first processor 2001 is the control center of the device 310 for constructing a trustworthy multi-modal large model based on visual contrast alignment. It may be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or may be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0142] Optionally, the first processor 2001 may execute various functions of the device 310 for constructing a trustworthy multi-modal large model based on visual contrast alignment by running or executing software programs stored in the memory 2002 and by calling data stored in the memory 2002.

[0143] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 3 the CPU0 and CPU1 shown in

[0144] In a specific implementation, as an embodiment, the device 310 for constructing a trustworthy multi-modal large model based on visual contrast alignment may also include multiple processors, for example Figure 3 the first processor 2001 and the second processor 2004 shown in Figure 3 . Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0145] Among them, the memory 2002 is used to store the software program for executing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner may refer to the above method embodiment and will not be elaborated here.

[0146] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 3 not shown in Figure 3 ). The embodiments of the present invention do not make specific limitations on this.

[0147] The transceiver 2003 is used to communicate with a network device or with a terminal device.

[0148] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 3 not separately shown in Figure 3 ). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.

[0149] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently, and is coupled to the first processor 2001 through an interface circuit of the device 310 constructed by a trusted multimodal large model based on visual contrast alignment ( Figure 3 not shown in the figure), and the embodiments of the present invention do not make specific limitations on this.

[0150] It should be noted that Figure 3 the structure of the device 310 constructed by the trusted multimodal large model based on visual contrast alignment shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0151] In addition, the technical effects of the device 310 constructed by the trusted multimodal large model based on visual contrast alignment can refer to the technical effects of the method for constructing a trusted multimodal large model based on visual contrast alignment described in the above method embodiments, and will not be elaborated here.

[0152] It should be understood that the first processor 2001 in the embodiments of the present invention can be a central processing unit (CPU), and this processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or this processor can also be any conventional processor, etc.

[0153] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0154] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0155] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0156] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0157] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0158] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0159] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated herein.

[0160] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0161] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0162] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0163] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0164] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for constructing a credible multimodal large model based on visual contrast alignment, characterized in that: The method comprises: S1. Obtain a preference-optimized data set; wherein the preference-optimized data set includes: text data and image data; input the image data into the multimodal large model after fine-tuning the instruction, and obtain the preference response logit with the image and the rejection response logit with the image; input the text data into the multimodal large model after fine-tuning the instruction, and obtain the preference response logit without the image and the rejection response logit without the image; S2. Constructing a framework of a credible multimodal large model based on visual contrast alignment; the framework includes: a text preference optimization module, a difference stabilization optimization module, a response-level visual contrast alignment module, and a mark-level visual contrast alignment module; S3, input the preference response logit with pictures and the rejection response logit with pictures into the text preference optimization module, generate the preference response reward and the rejection response reward through the given model to be optimized and the reference model; construct the text preference optimization loss function according to the preference response reward and the rejection response reward; S4, input the preference response logit with pictures into the difference stabilization optimization module, and construct the difference stabilization optimization loss function by setting the difference parameters; input the preference response logit without pictures and the rejection response logit without pictures into the response-level visual contrast alignment module, and construct the response-level visual optimization loss function; S5. Input the preference response logit with pictures, the rejection response logit with pictures, the preference response logit without pictures and the rejection response logit without pictures into the tag-level visual contrast alignment module, and construct the tag-level visual contrast alignment optimization loss function through the logit difference of the visual contrast output of different tokens; construct the overall loss function of the framework according to the text preference optimization loss function, the difference stability optimization loss function, the response-level visual contrast alignment optimization loss function and the tag-level visual contrast alignment optimization loss function; train the multimodal large model after fine-tuning the instructions according to the overall loss function to obtain a credible multimodal large model based on visual contrast alignment; The logit difference of the visual contrast outputs of different tokens is expressed by the following formula (1): (1) in, Represents a given text query x, image input When the response y is specified, the model currently being optimized The visual contrast logit difference of the generated t-th token; The overall loss function of the framework of the credible multimodal large model based on visual contrast alignment is expressed by the following formula (2): (2) in, represents the overall loss function; Represents the value of the text preference optimization loss function; Represents the value of the difference stable optimization loss function; represents the value of the response-level visual contrast alignment optimization loss function; Represents the value of the loss function for optimizing the label-level visual contrastive alignment.

2. The method for constructing a credible multimodal large model based on visual contrast alignment according to claim 1 is characterized in that: The text preference optimization loss function is expressed by the following formula (3): (3) in, Represents the value of the text preference optimization loss function; represents the hyperparameter that controls the deviation from the reference model; Represents the sigmoid function; represents the logit of the model output preference response in the optimization with pictures; represents the logit of the reference model output preference response in the presence of pictures; The logit representing the rejection response of the model output in the optimization when there is a picture; Represents the logit of the reference model output rejection response in the presence of a picture.

3. The method for constructing a credible multimodal large model based on visual contrast alignment according to claim 1, characterized in that: The difference stability optimization loss function is expressed by the following formula (4): (4) in, Represents the value of the difference stable optimization loss function; Indicates the difference parameter; represents the logit of the model output preference response in the optimization with pictures; Represents the sigmoid function; represents the logit of the reference model output preference response in the presence of pictures; represents the hyperparameter that controls the deviation from the reference model.

4. The method for constructing a credible multimodal large model based on visual contrast alignment according to claim 1, characterized in that: The response-level visual contrast alignment optimization loss function is expressed by the following formula (5): (5) in, represents the value of the response-level visual contrast alignment optimization loss function; represents the logit of the reference model output preference response in the absence of pictures; The logit represents the model output preference response in the optimization without pictures; Represents the sigmoid function; represents the hyperparameter that controls the deviation from the reference model.

5. The method for constructing a credible multimodal large model based on visual contrast alignment according to claim 1, characterized in that: The label-level visual contrast alignment optimization loss function is expressed by the following formula (6): (6) in, represents the value of the loss function for optimizing the label-level visual contrast alignment; Indicates the number of tokens in the current output text; Represents the absolute value of the logit difference of the visual contrast output of the currently optimized model; represents the absolute value of the logit difference of the visual contrast output of the reference model; Represents the sigmoid function.

6. A device for constructing a credible multimodal large model based on visual contrast alignment, the device for constructing a credible multimodal large model based on visual contrast alignment is used to implement the method for constructing a credible multimodal large model based on visual contrast alignment as claimed in any one of claims 1 to 5, characterized in that: The device comprises: An acquisition unit is used to acquire a preference data set; wherein the preference data set includes: text data and image data; the image data is input into the multimodal large model after fine-tuning the instruction, and the preference response logit with the image and the rejection response logit with the image are obtained; the text data is input into the multimodal large model after fine-tuning the instruction, and the preference response logit without the image and the rejection response logit without the image are obtained; The first construction unit is used to construct a framework of a credible multimodal large model based on visual contrast alignment; the framework includes: a text preference optimization module, a difference stabilization optimization module, a response-level visual contrast alignment module and a mark-level visual contrast alignment module; The second construction unit is used to input the preference response logit with pictures and the rejection response logit with pictures into the text preference optimization module, generate the preference response reward and the rejection response reward through the given model to be optimized and the reference model; and construct the text preference optimization loss function according to the preference response reward and the rejection response reward; The third construction unit is used to input the preference response logit with pictures into the difference stabilization optimization module, and construct the difference stabilization optimization loss function by setting the difference parameters; input the preference response logit without pictures and the rejection response logit without pictures into the response-level visual contrast alignment module, and construct the response-level visual optimization loss function; The fourth construction unit is used to input the preference response logit with pictures, the rejection response logit with pictures, the preference response logit without pictures and the rejection response logit without pictures into the token-level visual contrast alignment module, and construct the token-level visual contrast alignment optimization loss function through the logit difference of the visual contrast output of different tokens; construct the overall loss function of the framework according to the text preference optimization loss function, the difference stability optimization loss function, the response-level visual optimization loss function and the token-level visual contrast alignment optimization loss function; train the multimodal large model after instruction fine-tuning according to the overall loss function to obtain a credible multimodal large model based on visual contrast alignment.

7. A reliable multimodal large model construction device based on visual contrast alignment, characterized in that: The credible multimodal large model construction device based on visual contrast alignment includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Generator, generator training method and method for avoiding image coordinate adhesion

    CN115496989A

  • Care question and answer model training method based on transfer learning and expert feedback

    CN117542471A