Cross-modal retrieval model optimization system and method based on illusion enhancement
By generating hallucinogenic text and introducing a hallucination-enhanced contrastive loss function in contrastive learning, the encoder and text encoder of the multimodal model are optimized, solving the problems of low efficiency and low accuracy of existing cross-modal retrieval models in intelligent driving scenarios, and realizing adaptive iterative updates and efficient retrieval.
Patent Information
- Application Number
- CN202511038543.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-11
AI Technical Summary
Existing cross-modal retrieval models suffer from low retrieval efficiency and low accuracy in intelligent driving scenarios. Especially when the amount of data is limited and dynamically changing, existing data augmentation methods rely on manually designed templates and are difficult to achieve adaptive iterative updates.
By generating hallucinogenic text as hard negative samples, an augmented training dataset is constructed using a large language model and customized cue templates. A contrastive loss function for hallucination augmentation is introduced into contrastive learning to optimize the encoder and text encoder of the multimodal model, thereby achieving automated data augmentation and adaptive model iteration.
It significantly improves the accuracy and recall of cross-modal retrieval, enhances the model's adaptability and robustness, reduces labor costs, and optimizes user experience and system flexibility.
Smart Images

Figure CN120929629A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal artificial intelligence technology, specifically to a cross-modal retrieval model optimization system and method based on illusion enhancement. Background Technology
[0002] In the research and development of intelligent driving systems, efficient and accurate retrieval and mining of specific driving scenario data is crucial. The application of large-scale multimodal artificial intelligence models has significantly improved the efficiency of cross-modal retrieval in intelligent driving scenarios, becoming an indispensable tool in this field. Currently, the mainstream methods for optimizing such models mainly include the following two categories: 1. Small-sample cue fine-tuning method: This method uses a limited number of high-quality image-text pairs as samples and fine-tunes the pre-trained multimodal model through carefully designed cue vectors. Its core lies in using filtered small-sample data to update some parameters of the model (such as text cue vector sequences and image cue vector sequences) to improve the model's performance on specific tasks (such as detection). 2. Template-based data augmentation method: This method aims to improve the model's generalization ability. It first analyzes a single descriptive style using a large-scale image-text dataset, then generates diverse text sentence templates, and then re-annotates image data based on these templates to construct a new dataset with diverse descriptive styles, which is finally used to train the multimodal model.
[0003] However, small-sample fine-tuning methods heavily rely on high-quality, high-cost manual annotation to construct small-sample datasets. This is not only time-consuming and costly, but also makes it difficult to achieve rapid, adaptive iterative updates to adapt to the dynamic changes in data in the autonomous driving field. When the amount of sample data is very limited, the performance of the fine-tuned model is at risk of instability, and may even degrade, resulting in poor performance in real-world applications. While template-based data augmentation methods can generate diverse data, their data generation process heavily relies on carefully designed text sentence templates. The design and updating of templates require expert experience, limiting the flexibility and automation of the methods. When generating data using large models, new illusory data may be directly introduced, or additional manpower may be required for data verification. This makes existing models prone to low retrieval efficiency and inaccurate retrieval information in practical cross-modal retrieval applications, affecting user experience and system development efficiency.
[0004] Therefore, it is necessary to invent a cross-modal retrieval model optimization system based on illusion enhancement. This system uses the "illusion" characteristic to perform data augmentation, simulates the model's possible incorrect matching, and forms "hard negative samples" that are semantically similar but have minor errors. By introducing a contrastive learning mechanism based on illusion enhancement, the generated illusion-like text is used as a key negative sample (i.e., a hard negative sample) and introduced into the contrastive learning process during the model fine-tuning stage. This system can also achieve automated and adaptive iterative updates of large multimodal models, making it suitable for the retrieval needs of complex and ever-changing intelligent driving scenarios. Summary of the Invention
[0005] The purpose of this invention is to provide a cross-modal retrieval model optimization system based on illusion enhancement. This invention optimizes the performance of multimodal models in cross-modal retrieval tasks, effectively suppresses the illusion phenomenon, improves retrieval accuracy and relevance of recall results, while reducing implementation costs and improving retrieval efficiency. It is suitable for the retrieval needs of complex and ever-changing intelligent driving scenarios.
[0006] To achieve this objective, the present invention provides a cross-modal retrieval model optimization system based on illusion enhancement, comprising: The hallucination-like text generation module generates hallucination-like text based on the original text descriptions, language models, and preset prompt templates in the original image and text training dataset, generating a set of hallucination-like texts corresponding to the image samples in the original image and text training dataset. The dataset building module is used to merge the hallucination-like text set with the image samples and original text descriptions in the corresponding original image and text training dataset to generate an enhanced training dataset. The fine-tuning and optimization module is used to encode the image samples in the augmented training dataset by the image encoder in the multimodal model to obtain the image feature vector set; and to encode the original text description and hallucinogenic text set in the augmented training dataset by the text encoder in the multimodal model to obtain the original text feature vector set and the hallucinogenic text feature vector set, respectively. Based on the image feature vector set and the original text feature vector set, the contrastive loss function value from the original text feature vector to the corresponding image feature vector is calculated. Based on the image feature vector set, the original text feature vector set, and the hallucination-like text feature vector set, the contrastive loss function value from the image feature vector to the corresponding original text feature vector is calculated. Based on the contrastive loss function value from the original text feature vector to the corresponding image feature vector and the contrastive loss function value from the image feature vector to the corresponding original text feature vector, the overall loss function value of contrastive learning is obtained. Based on the overall loss function value, the encoder adapter and text encoder of the multimodal model are updated using the gradient descent method to obtain the optimized multimodal model.
[0007] Preferably, the original text description in the original training dataset is read, and then the request parameters of the language model are constructed by combining the preset prompt template. The language model is called according to the request parameters. If the call fails, the error log is recorded and the call is retried. If the call succeeds, the response data generated by the language model is parsed, the response data is cleaned, and a set of hallucinogenic texts is obtained.
[0008] Preferably, an identifier is added to each hallucinogenic text in the hallucinogenic text set based on the identifier information of the image samples and original text descriptions in the original image and text training dataset. The hallucinogenic text sets with the same identifier, the image samples in the original image and text training dataset, and the original text descriptions in the original image and text training dataset are merged to generate an enhanced training dataset, where the image samples and the corresponding original text descriptions have the same identifier information.
[0009] Preferably, based on the image feature vector set and the original text feature vector set, a contrastive learning loss function is calculated to obtain the contrastive loss function value from the original text feature vector to the corresponding image feature vector; based on the image feature vector set, the original text feature vector set, and the hallucination-like text feature vector set, the contrastive learning loss function value from the image feature vector to the corresponding original text feature vector is calculated.
[0010] The beneficial effects of this invention are as follows: This invention proposes a cross-modal retrieval model optimization system based on illusion enhancement. By using a hallucinogenic text generation method based on a large language model and customized prompt templates, it overcomes the bottleneck of existing solutions requiring manual design of text templates, achieves automated data augmentation, significantly reduces the labor cost of data construction, simplifies template maintenance and update processes, and improves the adaptive capability of multimodal models in intelligent driving scenario retrieval. By introducing hallucinogenic text as hard negative samples in contrastive learning training, it addresses the shortcomings of existing technologies in optimizing large-scale hallucination problems, optimizes the model's attention allocation mechanism, and enhances the alignment between visual and text representations, thereby significantly improving the accuracy of cross-modal retrieval. This invention significantly improves retrieval reliability, particularly in challenging scenarios within autonomous driving environments, by enhancing precision and recall. A contrastive loss function with hallucination enhancement is designed, explicitly incorporating a penalty term for hallucinogenic text into the loss calculation. This addresses retrieval bias caused by hallucinogenic information, strengthens the model's ability to suppress erroneous information, and thus improves the robustness of retrieval in autonomous driving scenarios. It also reduces misleading output during user operations and optimizes the user experience. Furthermore, by integrating the AutoML framework to automate model parameter updates, it overcomes the shortcomings of existing solutions' model iteration lag, supports adaptive fine-tuning and dynamic optimization of large multimodal models, adapts to frequent changes in autonomous driving scenario data, and significantly improves the system's long-term performance and deployment flexibility. This invention utilizes the hallucination characteristics of large language models to generate hallucinogenic text and uses it as hard negative samples for model fine-tuning during the contrastive learning phase, constructing an efficient and adaptive cross-modal retrieval optimization framework for autonomous driving scenarios. This invention effectively solves the inherent defects of existing technologies, such as reliance on manual templates for data augmentation, insufficient suppression of hallucination problems, and model iteration lag, while simultaneously improving retrieval accuracy and efficiency. Attached Figure Description
[0011] Figure 1 This is a schematic diagram of the structure of the present invention; Figure 2 This is a schematic diagram of the process of the present invention; Figure 3 This is a schematic diagram of the initial multimodal large model structure of the present invention; Figure 4 Example flowchart for generating hallucinogenic text. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to represent selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0013] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments: Example 1 A cross-modal retrieval model optimization system based on illusion enhancement, such as Figure 1 As shown, it includes: The hallucination-like text generation module generates hallucination-like text based on the original text descriptions, language models, and preset prompt templates in the original image and text training dataset, generating a set of hallucination-like texts corresponding to the image samples in the original image and text training dataset. The dataset building module is used to merge the hallucination-like text set with the image samples and original text descriptions in the corresponding original image and text training dataset to generate an enhanced training dataset. The fine-tuning and optimization module is used to encode the image samples in the augmented training dataset by the image encoder in the multimodal model to obtain the image feature vector set; and to encode the original text description and hallucinogenic text set in the augmented training dataset by the text encoder in the multimodal model to obtain the original text feature vector set and the hallucinogenic text feature vector set, respectively. Based on the image feature vector set and the original text feature vector set, the contrastive loss function value from the original text feature vector to the corresponding image feature vector is calculated. Based on the image feature vector set, the original text feature vector set, and the hallucination-like text feature vector set, the contrastive loss function value from the image feature vector to the corresponding original text feature vector is calculated. Based on the contrastive loss function value from the original text feature vector to the corresponding image feature vector and the contrastive loss function value from the image feature vector to the corresponding original text feature vector, the overall loss function value of contrastive learning is obtained. Based on the overall loss function value, the encoder adapter and text encoder of the multimodal model are updated using the gradient descent method to obtain the optimized multimodal model.
[0014] The multimodal model described in this invention is a deep learning-based artificial intelligence model capable of simultaneously processing image (visual modality) and text (language modality) data. It includes an image encoder, an encoding adapter, a text encoder, and a task adapter. The method used in this solution primarily optimizes the parameters of the text encoder and encoding adapter, without making any adjustments to the image encoder and task adapter, nor to the structure of the text encoder and encoding adapter. A schematic diagram of an initial large-scale multimodal model structure conforming to the above characteristics is shown below. Figure 3 As shown.
[0015] In determining the original image-text training dataset, in some preferred embodiments of this invention, the original image-text training dataset generally consists of only images and text descriptions that correspond to the facts of the image content, forming a binary tuple. It typically has three selectable sources, or combines data from three different sources to form the dataset: image data collected from the intelligent driving system's backend, labeled to form image-text pairs; high-quality image-text pairs acquired from third parties, collected or generated for specific scenarios; and publicly available image-text datasets that are freely accessible. For each image sample, there is one and only one matching text content, which is the textual description of the matched image content.
[0016] Regarding the specific encoding methods, some optimized technical solutions involve the image encoder first preprocessing the image samples, normalizing the pixel values, then mapping the images to vectors through a fully connected layer, adding positional information to each image vector, inputting it to a self-attention mechanism layer to calculate weights, and outputting a sequence of perceptual feature vectors through a nonlinear transformation. Similarly, the text encoder is fed with the original text description and the hallucination-like text set, respectively. A word segmentation algorithm divides the original text description and the hallucination-like text set into sub-words, maps the sub-words to word embedding vectors, adds positional information to each word embedding vector, inputs it to a self-attention mechanism layer to calculate weights, and outputs a sequence of perceptual feature vectors through a nonlinear transformation.
[0017] Regarding the language model being called, some optimized technical solutions include this one, which uses an API (Application Programming Interface) to call a large language model (such as DeepSeek-R1, QwQ-32B, etc.) as a hallucination text generator. It uses a customized prompt template to generate hallucination text in batches based on the original descriptive text in the training dataset.
[0018] In some embodiments of the present invention, such as Figure 4 As shown, the specific method for generating a set of hallucinogenic text corresponding to the image samples in the original image and text training dataset is as follows: Read the original text descriptions from the original training dataset of images and text, and then construct the request parameters of the language model by combining them with the preset prompt template. Call the language model according to the request parameters. If the call fails, record the error log and retry. If the call succeeds, parse the response data generated by the language model, clean the response data, and obtain the hallucination text set.
[0019] Regarding the response data described in this invention, in some preferred embodiments, the response data refers to the response generated by the language model. This response is usually generated in a formatted manner according to the prompt template requirements to produce hallucinogenic text, which is convenient for subsequent processing. The hallucinogenic text is then extracted separately.
[0020] Regarding the cleaned response data described in this invention, in some preferred embodiments, the content of the responses (i.e., response data) generated by the language model is not limited to hallucinogenic text and cannot be directly used to construct the dataset. Therefore, it needs to be processed to a certain extent (such as removing non-hallucinogenic text content) and evaluated. The evaluation, for example, introduces a new language model as a "referee" to determine whether the hallucinogenic text generated by the current language model is "qualified," that is, whether the generated hallucinogenic text meets the requirements of the prompt template and whether it possesses the characteristics of real hallucinogenic text, etc.
[0021] In some preferred embodiments of this invention, when generating hallucinogenic text, the language model prompt template examples are as follows: 1. Requirement: Modify the given text description to create a copy that is closely aligned with the original content and length, and add hallucinatory elements. The first step is to identify the objects involved in the given description and their related attributes, and combine this with the background knowledge about hallucinations mentioned above to complete the task. 2. Example text description: A white MPV (Multi-Purpose Vehicle) is parked on the side of the road, next to which is a black fire hydrant. Generated hallucinogenic text 1: A black MPV is parked on the side of the road, next to which is a red fire hydrant; Generated hallucinogenic text 2: A white SUV is parked on the side of the road, next to which is a black street lamp; Generated hallucinogenic text 3: A white sedan is parked on the side of the road, next to which is a red fire hydrant. It should be noted that in practice, multiple hallucinogenic texts are usually generated to improve the utilization rate of the original training dataset samples.
[0022] In some embodiments of the present invention, by emphasizing the calling logic based on the original text description, preset prompt template and large language model, and adding error handling and cleaning mechanisms, the problem that some existing solutions require manual design of text templates (instead of prompt templates) to achieve data augmentation is solved. Compared with prompt templates, manual design of text templates is not only labor-intensive, but also difficult to achieve overall template updates during the update and iteration process. In terms of automated iterative updates of multimodal large models, the script method adopted by this solution is also easier to maintain and saves labor costs.
[0023] In some embodiments of the present invention, the specific method for generating the enhanced training dataset is as follows: Identifiers are added to each hallucinogenic text in the hallucinogenic text set based on the identifier information of image samples and original text descriptions in the original image and text training dataset. The hallucinogenic text sets with the same identifier, the image samples in the original image and text training dataset, and the original text descriptions in the original image and text training dataset are merged to generate an augmented training dataset, in which the image samples and their corresponding original text descriptions have the same identifier information.
[0024] In some embodiments of the present invention, the generated hallucination-like text (there may be multiple texts) is precisely matched with the original image and its original text description by an identifier, forming a structured triplet data (image-original text-hallancy text). This ensures that in subsequent training, the model can correctly distinguish and compare the original description of a specific image with its corresponding hallucination text, which is the basic guarantee for achieving the hallucination enhancement effect.
[0025] In some embodiments of the present invention, a contrastive learning loss function is calculated based on the image feature vector set and the original text feature vector set to obtain the contrastive loss function value from the original text feature vector to the corresponding image feature vector; and a contrastive learning loss function value from the image feature vector to the corresponding original text feature vector is calculated based on the image feature vector set, the original text feature vector set, and the hallucination-like text feature vector set.
[0026] In some embodiments of the present invention, by precisely defining the application scenarios of the new loss function, it is highlighted that illusory text is only used as an additional negative sample in the loss function in the "image to text" direction.
[0027] In some embodiments of the present invention, the specific formula for calculating the contrast loss function value from the original text feature vector to the corresponding image feature vector is as follows: in, The contrast loss represents the comparison between the original text feature vector and the corresponding image feature vector. This represents the number of samples in the current batch. This represents the distance function between the original text feature vector and the corresponding image feature vector, used to measure the similarity or difference between the image feature vector and the original text feature vector. `i` represents the current sample index, and `k` represents all other samples besides the current sample. Represents the image feature vector set. Represents the original text feature vector set. It represents the set of image feature vectors for all samples other than the current sample.
[0028] In some embodiments of the present invention, by defining a specific mathematical formula for the text-to-image contrast loss function, the basic ability of the model to learn the alignment of text with the correct image is protected, while minimizing its similarity with all other images in the batch.
[0029] In some embodiments of the present invention, the specific formula for calculating the contrast loss function value from the image feature vector to the corresponding original text feature vector is as follows: in, The comparison loss represents the comparison between the image feature vector and the corresponding original text feature vector. This represents the number of hallucinatory texts generated for each original text description. This represents the number of samples in the current batch. This represents the distance function between the original text feature vector and the corresponding image feature vector, used to measure the similarity or difference between the image feature vector and the original text feature vector. `i` represents the current sample index, and `k` represents all other samples besides the current sample. Represents the image feature vector set. Represents the original text feature vector set. This represents the original text feature vector set of all samples except the current sample. This represents the set of hallucination-like text feature vectors for all samples other than the current sample.
[0030] In some embodiments of the present invention, by defining a specific mathematical formula for the image-to-text contrast loss function, in addition to considering the similarity between the current image and the corresponding original text, as well as the similarity with other sample original texts, a penalty term is added for the similarity between the current image and all related hallucination texts. In this way, the hallucination text is directly used as a hard negative sample to remove the distance between the image and the erroneous text (i.e., hallucination text) representation, thereby explicitly suppressing hallucinations and improving the model's ability to discriminate true correspondences.
[0031] In some embodiments of the present invention, the specific formula for calculating the overall loss function of contrastive learning is as follows: in, This represents the overall loss function value for contrastive learning. This represents the parameter to be updated. This represents the loss value that reflects the performance of the current multimodal model in cross-modal retrieval tasks within intelligent driving scenarios. The comparison loss represents the comparison between the image feature vector and the corresponding original text feature vector. This represents the contrast loss between the original text feature vector and the corresponding image feature vector.
[0032] In some embodiments of the present invention, the contrastive learning loss function with hallucination sample loss is used, which solves the problem that the existing schemes fail to effectively suppress hallucination cases in large models. Since the loss function of this scheme explicitly includes a loss penalty term for hallucination-like text in the large model fine-tuning and optimization stage, the model will further improve its sensitivity to hallucination samples and thus optimize its attention allocation mechanism. In downstream cross-modal retrieval tasks, this will significantly improve the accuracy of recalled samples, enhance user experience and work efficiency.
[0033] In some embodiments of the present invention, the specific method for updating the encoder adapter and text encoder of the multimodal model using gradient descent based on the overall loss function value is as follows: The image coding adapter parameters are evaluated by calculating the overall loss function. and text encoder parameters The gradient is used to backpropagate and adjust the parameters of the encoding adapter and text encoder. The overall loss function value is then calculated based on the image encoding adapter parameters. gradient The specific formula for updating is: in, Indicates the image encoding adapter parameters, Represents the overall loss function For parameters gradient, The learning rate is determined by a cosine annealing learning rate decay strategy employed during model training, combined with L2 regularization to prevent overfitting. The symbol " "" indicates variable assignment, i.e. parameter update, symbol " " is the gradient symbol, representing the derivative of the loss function; Calculate the overall loss function value against the text encoder parameters. gradient The specific formula for updating is: in, Indicates text encoder parameters, Represents the overall loss function For parameters gradient, This is the learning rate.
[0034] In some preferred embodiments of this invention, when embedding the optimized multimodal large model into the original retrieval application, the running intelligent driving scenario cross-modal retrieval system is paused, the initial multimodal large model file of the system is replaced with the finely tuned multimodal large model file, and the system is restarted to complete the switching of the multimodal large model. At this time, when the user uses the system again, the finely tuned and optimized multimodal large model will be used for retrieval.
[0035] In some embodiments of the present invention, the method of optimizing the initial multimodal large model based on hallucination enhancement contrastive learning solves the problem that most existing solutions fail to optimize for hallucination situations in large models. By introducing hallucination-like text, not only is the number of original samples expanded, but new hard negative samples are also introduced into the model in terms of distribution. This is conducive to building a more detailed attention mechanism in the model, which will significantly improve the accuracy of recalled samples in downstream cross-modal retrieval tasks, improve user experience and work efficiency.
[0036] Example 2 An optimization method for cross-modal retrieval models based on illusion enhancement, such as Figure 2 As shown, it includes: Obtain the initial multimodal large model and training dataset. Based on the original text description, language model and preset prompt template in the original image and text training dataset, perform hallucinogenic text generation processing to generate a set of hallucinogenic texts corresponding to the image samples in the original image and text training dataset. The hallucination-like text set is merged with the image samples and original text descriptions in the corresponding original image and text training dataset to generate an enhanced training dataset. The image encoder in the multimodal model encodes the image samples in the augmented training dataset to obtain the image feature vector set; the text encoder in the multimodal model encodes the original text description and the hallucination-like text set in the augmented training dataset respectively to obtain the original text feature vector set and the hallucination-like text feature vector set. Based on the image feature vector set and the original text feature vector set, the contrastive loss function value from the original text feature vector to the corresponding image feature vector is calculated; based on the image feature vector set, the original text feature vector set, and the hallucination-like text feature vector set, the contrastive loss function value from the image feature vector to the corresponding original text feature vector is calculated; based on the contrastive loss function value from the original text feature vector to the corresponding image feature vector and the contrastive loss function value from the image feature vector to the corresponding original text feature vector, the overall loss function value of contrastive learning is obtained; based on the overall loss function value, the encoder adapter and text encoder of the multimodal model are updated using the gradient descent method to obtain the optimized multimodal model; Replace the system's initial large multimodal model file with the optimized multimodal model file, and use the optimized multimodal model for retrieval.
[0037] Example 3 A computer program product includes a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the method described in Embodiment 2.
[0038] The contents not described in detail in this specification are existing technologies known to those skilled in the art.
Claims
1. A cross-modal retrieval model optimization system based on illusion enhancement, characterized in that, include: The hallucination-like text generation module generates hallucination-like text based on the original text descriptions, language models, and preset prompt templates in the original image and text training dataset, generating a set of hallucination-like texts corresponding to the image samples in the original image and text training dataset. The dataset building module is used to merge the hallucination-like text set with the image samples and original text descriptions in the corresponding original image and text training dataset to generate an enhanced training dataset. The fine-tuning and optimization module is used to encode the image samples in the augmented training dataset by the image encoder in the multimodal model to obtain the image feature vector set; and to encode the original text description and hallucinogenic text set in the augmented training dataset by the text encoder in the multimodal model to obtain the original text feature vector set and the hallucinogenic text feature vector set, respectively. Based on the image feature vector set and the original text feature vector set, the contrastive loss function value from the original text feature vector to the corresponding image feature vector is calculated. Based on the image feature vector set, the original text feature vector set, and the hallucination-like text feature vector set, the contrastive loss function value from the image feature vector to the corresponding original text feature vector is calculated. Based on the contrastive loss function value from the original text feature vector to the corresponding image feature vector and the contrastive loss function value from the image feature vector to the corresponding original text feature vector, the overall loss function value of contrastive learning is obtained. Based on the overall loss function value, the encoder adapter and text encoder of the multimodal model are updated using the gradient descent method to obtain the optimized multimodal model.
2. The cross-modal retrieval model optimization system based on illusion enhancement according to claim 1, characterized in that: The specific method for generating a set of hallucinogenic text corresponding to the image samples in the original image and text training dataset is as follows: Read the original text descriptions from the original training dataset of images and text, and then construct the request parameters of the language model by combining them with the preset prompt template. Call the language model according to the request parameters. If the call fails, record the error log and retry. If the call succeeds, parse the response data generated by the language model, clean the response data, and obtain the hallucination text set.
3. The cross-modal retrieval model optimization system based on illusion enhancement according to claim 1, characterized in that: The specific method for generating the augmented training dataset is as follows: Identifiers are added to each hallucinogenic text in the hallucinogenic text set based on the identifier information of image samples and original text descriptions in the original image and text training dataset. The hallucinogenic text sets with the same identifier, the image samples in the original image and text training dataset, and the original text descriptions in the original image and text training dataset are merged to generate an augmented training dataset, in which the image samples and their corresponding original text descriptions have the same identifier information.
4. The cross-modal retrieval model optimization system based on illusion enhancement according to claim 1, characterized in that: Based on the image feature vector set and the original text feature vector set, the contrastive learning loss function is calculated to obtain the contrastive loss function value from the original text feature vector to the corresponding image feature vector; based on the image feature vector set, the original text feature vector set, and the hallucination-like text feature vector set, the contrastive learning loss function value from the image feature vector to the corresponding original text feature vector is calculated.
5. A cross-modal retrieval model optimization system based on illusion enhancement according to claim 1 or 4, characterized in that: The specific formula for calculating the comparison loss function value between the original text feature vector and the corresponding image feature vector is as follows: in, The contrast loss represents the comparison between the original text feature vector and the corresponding image feature vector. This represents the number of samples in the current batch. This represents the distance function between the original text feature vector and the corresponding image feature vector, used to measure the similarity or difference between the image feature vector and the original text feature vector. `i` represents the current sample index, and `k` represents all other samples besides the current sample. Represents the image feature vector set. Represents the original text feature vector set. It represents the set of image feature vectors for all samples other than the current sample.
6. A cross-modal retrieval model optimization system based on illusion enhancement according to claim 1 or 4, characterized in that: The specific formula for calculating the comparison loss function value between the image feature vector and the corresponding original text feature vector is as follows: in, The comparison loss represents the comparison between the image feature vector and the corresponding original text feature vector. This represents the number of hallucinatory texts generated for each original text description. This represents the number of samples in the current batch. This represents the distance function between the original text feature vector and the corresponding image feature vector, used to measure the similarity or difference between the image feature vector and the original text feature vector. `i` represents the current sample index, and `k` represents all other samples besides the current sample. Represents the image feature vector set. Represents the original text feature vector set. This represents the original text feature vector set of all samples except the current sample. This represents the set of hallucination-like text feature vectors for all samples other than the current sample.
7. The cross-modal retrieval model optimization system based on illusion enhancement according to claim 1, characterized in that: The specific formula for calculating the overall loss function of contrastive learning is as follows: in, This represents the overall loss function value for contrastive learning. This represents the parameter to be updated. This represents the loss value that reflects the performance of the current multimodal model in cross-modal retrieval tasks within intelligent driving scenarios. The comparison loss represents the comparison between the image feature vector and the corresponding original text feature vector. This represents the contrast loss between the original text feature vector and the corresponding image feature vector.
8. The cross-modal retrieval model optimization system based on illusion enhancement according to claim 1, characterized in that: The specific method for updating the encoder adapter and text encoder of the multimodal model using gradient descent based on the overall loss function value is as follows: The image coding adapter parameters are evaluated by calculating the overall loss function. and text encoder parameters The gradient is used to backpropagate and adjust the parameters of the encoding adapter and text encoder. The overall loss function value is then calculated based on the image encoding adapter parameters. gradient The specific formula for updating is: in, Indicates the parameters of the image encoding adapter. Represents the overall loss function For parameters gradient, The learning rate is represented by the symbol "". "Indicates variable assignment, i.e. parameter update, symbol " " is the gradient symbol, representing the derivative of the loss function; Calculate the overall loss function value against the text encoder parameters. gradient The specific formula for updating is: in, Indicates text encoder parameters, Represents the overall loss function For parameters gradient, This is the learning rate.
9. A method for optimizing a cross-modal retrieval model based on illusion enhancement, characterized in that, It includes: Based on the original text descriptions, language models, and preset prompt templates in the original image and text training dataset, hallucinogenic text generation processing is performed to generate a set of hallucinogenic texts corresponding to the image samples in the original image and text training dataset. The hallucination-like text set is merged with the image samples and original text descriptions in the corresponding original image and text training dataset to generate an enhanced training dataset. The image encoder in the multimodal model encodes the image samples in the augmented training dataset to obtain the image feature vector set; the text encoder in the multimodal model encodes the original text description and the hallucination-like text set in the augmented training dataset respectively to obtain the original text feature vector set and the hallucination-like text feature vector set. Based on the image feature vector set and the original text feature vector set, the contrastive loss function value from the original text feature vector to the corresponding image feature vector is calculated. Based on the image feature vector set, the original text feature vector set, and the hallucination-like text feature vector set, the contrastive loss function value from the image feature vector to the corresponding original text feature vector is calculated. Based on the contrastive loss function value from the original text feature vector to the corresponding image feature vector and the contrastive loss function value from the image feature vector to the corresponding original text feature vector, the overall loss function value of contrastive learning is obtained. Based on the overall loss function value, the encoder adapter and text encoder of the multimodal model are updated using the gradient descent method to obtain the optimized multimodal model.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 9.
Citation Information
Cited By
Visual vocabulary guided multi-modal large model illusion optimization method and storage medium
CN121168568A
Large model illusion suppression training method and system fused with real-time fact library verification
CN121257756A
Remote sensing data fusion method based on locality sensitive hashing and cross-modal attention
CN122200258A
Large-model medical image report generation method based on closed-loop feedback
CN122224399A