Training method and device for relieving illusion of multi-modal large model

By introducing two-dimensional Manhattan distance and improving causal attention masks, the position modeling method of multimodal large models is optimized, and the problem of insufficient interaction between images and instruction information is solved, and the multimodal feature alignment effect and reliability of the model are improved.

CN120258071AActive Publication Date: 2025-07-04XIAMEN UNIV OF TECH +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510718228.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-04
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

The existing multimodal large models have limitations in position encoding methods, resulting in insufficient information interaction between images and instructions, unsatisfactory alignment of multimodal features, and illusion.

Method used

By introducing two-dimensional Manhattan distance calculation and improving causal attention masks, redefining the positional relationship of image markers, using the freezing pre-training module and gradually fine-tuning the multi-layer perceptron strategy to optimize the model's multimodal information interaction capabilities.

Benefits of technology

The image information perception ability and multimodal feature alignment effect of multimodal large models are significantly improved, the incidence of hallucinations is reduced, and the reliability and efficiency of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258071A_ABST
    Figure CN120258071A_ABST
Patent Text Reader

Abstract

The invention provides a training method and device for relieving illusion of a multi-modal large model, and relates to the technical field of multi-modal large model training, and the method optimizes the defects of traditional one-dimensional position coding and retains the spatial locality characteristics of an image by redefining the position relation between image marks and introducing two-dimensional Manhattan distance calculation. Meanwhile, by improving the causal attention mask, the fusion capability of the model on the image and text information is further improved. In the model training process, a strategy of freezing the pre-training module and gradually performing fine adjustment is adopted, so that the multi-modal alignment effect of the model is remarkably improved, the occurrence rate of illusion phenomena is reduced, and a new technical path is provided for constructing a more reliable and more efficient multi-modal artificial intelligence system. The objective of the invention is to solve the illusion problem of a multi-modal large model caused by a position coding mode in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal large model training, and in particular to a training method and device for alleviating multimodal large model hallucinations. Background Art

[0002] In today's era of rapid development of artificial intelligence, large multimodal models (LVLMs) have received widespread attention and application as a powerful technical tool. They can integrate information from multiple modalities, such as vision and language, to achieve smarter and more comprehensive task processing. However, with the continuous development and application of large multimodal models, the problem of hallucination has gradually become prominent and has become one of the key bottlenecks restricting its further development.

[0003] Hallucination refers to the phenomenon that when a large multimodal model generates or processes information, it produces erroneous output that is inconsistent, irrelevant, or illogical with the actual input information. This phenomenon seriously affects the accuracy and reliability of the model, reducing its credibility and usability in practical applications. After in-depth research, it was found that insufficient alignment between multimodal features is a key factor causing hallucinations. In existing large multimodal models, the rotational position encoding (RoPE) used for position modeling has certain limitations, and its long-term attenuation has a negative impact on multimodal alignment.

[0004] Specifically, under long-term attenuation, the instruction markers show uneven perception of image markers located at different positions in the two-dimensional space. In the one-dimensional sequence, the markers in the lower right image area that are closer to the instruction marker are preferred, while less attention is paid to the farther image markers. This perceptual bias leads to insufficient information interaction between the image and the instruction, and the multimodal feature alignment effect is not ideal, which in turn causes the hallucination phenomenon. The existing position encoding method fails to fully consider the two-dimensional spatial structure of the image, and only flattens the image markers into a one-dimensional sequence for processing, which loses the spatial locality information of the image and cannot effectively support the accurate alignment and interaction between multimodal features.

[0005] Therefore, how to improve the position modeling method of large multimodal models, enhance the model's perception of image information, and improve the accuracy and effect of multimodal feature alignment, so as to effectively alleviate the hallucination problem, has become an important technical problem that needs to be urgently solved in the current research field of large multimodal models.

[0006] In view of this, this application is filed. Summary of the invention

[0007] The present invention provides a training method and device for alleviating the multimodal large model hallucination, which can at least partially improve the above-mentioned problem.

[0008] To achieve the above object, the present invention adopts the following technical solutions: A training method for alleviating hallucinations in multimodal large models, which includes: Obtain image data, and input the image data into a multimodal large model, perform encoding processing on the image data to obtain a plurality of image tokens; Calculate the relative position distances between the image tokens, and evolve the relative position distances from a one-dimensional level to a two-dimensional level to obtain the two-dimensional coordinate Manhattan distance; Perform raster scanning processing on the image tokens, allocate position indices, and re-allocate two-dimensional position coordinates to the image tokens; Perform equivalent transformation processing on the two-dimensional coordinate Manhattan distance, and replace the position index of the raster scan with the sum of the coordinate values of the re-allocated two-dimensional position coordinates of the image tokens; Model the causal attention mask according to the new position index and the transformed two-dimensional coordinate Manhattan distance, and replace the default causal attention mask with the newly modeled causal attention mask; Freeze the multimodal large model, and perform adjustment training preprocessing on the multimodal large model according to the replaced causal attention mask and the preset pre-training data until the multimodal large model reaches the preset effect.

[0009] The present invention also provides a training device for alleviating hallucinations in multimodal large models, which includes: An encoding unit for obtaining image data, inputting the image data into a multimodal large model, and performing encoding processing on the image data to obtain a plurality of image tokens; A distance calculation unit for calculating the relative position distances between the image tokens, and evolving the relative position distances from a one-dimensional level to a two-dimensional level to obtain the two-dimensional coordinate Manhattan distance; A re-allocation unit for performing raster scanning processing on the image tokens, allocating position indices, and re-allocating two-dimensional position coordinates to the image tokens; An equivalent transformation unit for performing equivalent transformation processing on the two-dimensional coordinate Manhattan distance, and replacing the position index of the raster scan with the sum of the coordinate values of the re-allocated two-dimensional position coordinates of the image tokens; A replacement unit for modeling the causal attention mask according to the new position index and the transformed two-dimensional coordinate Manhattan distance, and replacing the default causal attention mask with the newly modeled causal attention mask; A training unit for freezing the multimodal large model, and performing adjustment training preprocessing on the multimodal large model according to the replaced causal attention mask and the preset pre-training data until the multimodal large model reaches the preset effect.

[0010] In summary, the training method for alleviating hallucinations in multimodal large models effectively solves the hallucination phenomenon in existing multimodal large models caused by the position encoding method through innovative position modeling and attention mask mechanisms. Specifically, based on the deployment of multimodal large models, this method redefines the positional relationship of image tokens, introduces two-dimensional Manhattan distance calculation, and replaces the traditional one-dimensional raster scan position index, thereby preserving the spatial locality characteristics of the image. By equivalently transforming the Manhattan relative position distance calculation formula and combining it with causal attention mask modeling, a Manhattan causal mask is formed, further optimizing the multimodal information interaction ability of the model. During the model training process, the strategy of freezing the pre-trained visual encoder and large language model and only fine-tuning some parameters of the multi-layer perceptron is adopted, and then all parameters of the model are trained, significantly improving the image information perception ability and multimodal feature alignment effect of the model. This method achieves the best accuracy and the lowest hallucination rate on classical hallucination evaluation metrics, providing a new technical path for building more reliable and efficient multimodal artificial intelligence systems, with important application value and broad market prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a schematic flowchart of the training method for alleviating hallucinations in multimodal large models provided by the first embodiment of the present invention; Figure 2 is a schematic diagram of image token position index assignment provided by an embodiment of the present invention. In Figure (a), the number of image tokens is 100 for demonstration, and in Figure (b), the position index calculated based on coordinates is shown; Figure 3 is a schematic diagram of the visualization of the image-text information flow under the default position modeling method of the multimodal large model provided by an embodiment of the present invention; Figure 4 is a schematic diagram of the visualization of the image-text information flow provided by an embodiment of the present invention; Figure 5 is a schematic diagram of the causal attention mask under the default position modeling method of the multimodal large model provided by an embodiment of the present invention; Figure 6 is a schematic diagram of the causal attention mask provided by an embodiment of the present invention; Figure 7 is a schematic diagram of the results on classical hallucination evaluation metrics provided by an embodiment of the present invention; Figure 8 is a schematic diagram of the modules of the training device for alleviating hallucinations in multimodal large models provided by the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0013] Referring to Figure 1 , Figure 2 As shown, the first embodiment of the present invention discloses a training method for alleviating hallucinations in multi-modal large models, which can be executed by a training device for alleviating hallucinations in multi-modal large models (hereinafter referred to as the training device), and specifically, by one or more processors in the training device to implement the following method: S1. Obtain image data, input the image data into the multi-modal large model, and perform encoding processing on the image data to obtain a plurality of image tokens; Specifically, step S1 includes: using the visual encoder of the multi-modal large model to perform encoding processing on each image of the image data; Among them, the resolution of each image is reset to a size of 24x24, and each image is segmented according to this size to obtain 576 image tokens.

[0014] Preferably, the multi-modal large model includes a visual encoder, a large language model, and a multi-layer perceptron for connecting the visual encoder and the large language model. Among them, the visual encoder uses CLIP ViT-L / 14, the large language model uses Vicuna-7B, the multi-layer perceptron includes a two-layer MLP, and the overall scale of the multi-modal large model is 7B.

[0015] In this embodiment, image data is obtained and input into the multi-modal large model for processing. The multi-modal large model consists of a visual encoder, a large language model, and a multi-layer perceptron for connecting the two. Among them, the visual encoder uses CLIP ViT-L / 14, the large language model uses Vicuna-7B, the multi-layer perceptron part is a two-layer MLP, and the overall scale of the entire model is 7B. The selection of this architecture aims to make full use of the powerful capabilities of pre-trained models and at the same time achieve effective fusion of visual and language modalities through the multi-layer perceptron.

[0016] When processing the image data, first reset the resolution of each image to a size of 24x24. This operation is to ensure that the image can be evenly segmented into multiple image tokens in subsequent processing, so as to provide more refined visual information for the model. Based on the 24x24 resolution, each image is further segmented into 576 image tokens. These image tokens will be used as the input of the visual encoder for subsequent encoding processing.

[0017] Encoding image tokens through a vision encoder can transform the visual features of an image into a token form that the model can understand and process. This process not only retains the key information of the image but also provides a basis for subsequent fusion with the language modality. The image is segmented into 576 tokens and a resolution of 24x24 is adopted, enabling the model to capture local features in the image more meticulously and providing richer spatial information for subsequent position modeling.

[0018] This preprocessing and encoding method of image data not only provides high-quality visual input for the model but also lays a foundation for subsequent position modeling and optimization of the attention mechanism. By segmenting the image into multiple tokens and adopting an appropriate resolution, it is possible to better retain the spatial locality features of the image, thereby achieving more effective multi-modal information interaction and alignment in subsequent steps.

[0019] S2. Calculate the relative position distances between the image tokens and evolve the relative position distances from a one-dimensional level to a two-dimensional level to obtain the two-dimensional coordinate Manhattan distance. Specifically, step S2 includes: performing position encoding processing on the image tokens using the rotational position encoding technique to calculate the relative position distances of the image tokens. The calculation formula is: , where is the Query token at position index i, is the Key token at position index j, represents and the relative position distance between them. This relative position distance is a one-dimensional distance. represents the position index of the image token determined by raster scanning, represents the position index of the j-th image token, represents the position index of the i-th image token; Expand the relative position distance from a one-dimensional level to a two-dimensional level to obtain the two-dimensional coordinate Manhattan distance. The formula is: , where is the two-dimensional coordinate Manhattan distance, and are the two-dimensional horizontal and vertical coordinates of the Key token at position j respectively, and are the two-dimensional horizontal and vertical coordinates of the Query token at position i respectively.

[0020] In this embodiment, the Rotation Position Encoding (RoPE) technique is used to perform position encoding on image tokens. This method is achieved by calculating the relative position distances between image tokens; this relative position distance is one-dimensional, and it determines the position index of image tokens in a raster scanning manner. However, this one-dimensional position encoding method has obvious limitations because it flattens two-dimensional image tokens into a one-dimensional sequence, losing the spatial locality features of the two-dimensional image, resulting in insufficient perception of the spatial relationships between image tokens by the model. Specifically, image tokens are flattened into a one-dimensional sequence in raster scanning order (from top to bottom, from left to right), and then concatenated with instruction tokens to form an input sequence. Under long-term decay, the attention of image tokens farther away from the instruction tokens gradually weakens. This leads to a fixed multimodal alignment pattern, where image tokens later in the raster scanning order receive more attention, while very sparse attention is given to farther image tokens. As Figure 3 shown, the visualization of the information flow from image to instruction further reveals this phenomenon.

[0021] To solve this problem, this method proposes to extend the relative position distance from the one-dimensional level to the two-dimensional level, introducing the concept of two-dimensional coordinate Manhattan distance to more accurately reflect the spatial relationships between image tokens and retain the two-dimensional locality features of the image. This extension from one-dimensional to two-dimensional has significant beneficial effects. First, the two-dimensional coordinate Manhattan distance can more realistically reflect the spatial distances between image tokens, avoiding the perceptual bias caused by the raster scanning order in one-dimensional position encoding. For example, in one-dimensional position encoding, the attention of image tokens gradually weakens as the distance increases, resulting in less attention being given to image tokens far from the instruction tokens. By introducing the two-dimensional Manhattan distance, the model can more evenly perceive image tokens regardless of their positions in the image. This improvement significantly enhances the model's ability to perceive image information, thereby improving the accuracy of multimodal feature alignment.

[0022] Second, the introduction of the two-dimensional Manhattan distance also optimizes the causal attention mechanism of the model. In the traditional causal attention mask, the attention decay of the model to image tokens is unidirectional, while under the modeling of the two-dimensional Manhattan distance, the attention decay is extended to multi-directional, which enables the model to more comprehensively capture the interaction information between images and text. This improvement not only reduces the occurrence of hallucination phenomena but also improves the overall performance of the model in multimodal tasks.

[0023] S3. Perform raster scanning processing on the image tokens, assign position indices, and reassign two-dimensional position coordinates to the image tokens; Specifically, step S3 includes: performing raster scanning on the image tokens and assigning position indices in increments of 1, where the position indices are assigned in a manner the number of image markers; Re - assign two - dimensional position coordinates to each of the said image markers in the following way: ; Set the positions of the image markers at the four vertices of the image as the four starting coordinate origins (0, 0). The marker points adjacent to the origin are set as the next marker points, with an increment of 1. Among them, the positive directions of the horizontal axis and the vertical axis are consistent with the increment direction; According to the four starting coordinate origins, divide the image tokens after resolution reset into four mirror - image parts. The image markers in each part increase linearly along the positive directions of the two - dimensional coordinate axes of the origin.

[0024] In this embodiment, raster scan processing is performed on the image markers. The image markers are scanned row - by - row and column - by - column starting from the upper - left corner, and a position index is assigned to each image marker. These position indexes increase sequentially with an increment of 1, thus forming a one - dimensional position index sequence. Suppose the total number of image markers is v, then the assignment range of the position index is from 0 to v - 1. Although this raster scan method is simple and intuitive, in traditional multi - modal large models, it will cause the loss of spatial locality information of image markers because all image markers are tiled into a one - dimensional sequence, ignoring their actual position relationships in the two - dimensional image.

[0025] To overcome this limitation, this method proposes an innovative two - dimensional position coordinate assignment method to re - assign two - dimensional position coordinates to each image marker to better retain the spatial locality characteristics of the image. First, set the positions of the image markers at the four vertices of the image as the four starting coordinate origins (0, 0). These four origins correspond to the upper - left corner, upper - right corner, lower - left corner, and lower - right corner of the image respectively. Starting from these four origins, the marker points adjacent to the origin are set as the next marker points, and their coordinate values increase sequentially with an increment of 1. The positive directions of the horizontal axis and the vertical axis are consistent with the increment direction, ensuring that the coordinate assignment of image markers is carried out along the two - dimensional space direction of the image.

[0026] Secondly, according to these four starting coordinate origins, divide the image tokens after resolution reset into four mirror - image parts. In each part, the two - dimensional coordinates of the image markers increase linearly along the positive directions of the two - dimensional coordinate axes of the origin. For example, in the upper - left part, the coordinates of the image markers start from (0, 0) and increase sequentially along the horizontal and vertical directions, forming an ordered two - dimensional coordinate grid. This assignment method not only retains the spatial locality characteristics of the image but also enables the model to more naturally process the spatial relationships between image markers. Under the new position modeling mechanism, the information flow from the image to the instruction is as Figure 4 shown.

[0027] Based on this two-dimensional position coordinate allocation method, the following beneficial effects are achieved by this method: 1. Preserving the spatial locality of the image: The traditional raster scanning method tiles the image markers into a one-dimensional sequence, losing the spatial locality characteristics of the image. However, in the present invention, by allocating two-dimensional coordinates to the image markers, the two-dimensional structural information of the image is retained, enabling the model to better perceive the spatial relationship between the image markers. 2. Enhancing multimodal feature alignment: In a multimodal large model, the alignment between the image and the text is crucial. By introducing two-dimensional coordinates, the model can more accurately align the image markers with the text markers, thereby improving the effect of multimodal feature fusion and reducing the occurrence of hallucination phenomena. 3. Optimizing the attention mechanism: The two-dimensional coordinate allocation method enables the model to more naturally consider the spatial position of the image markers when calculating attention. Compared with the traditional raster scanning method, this allocation method can more effectively guide the model's attention, making it pay more attention to the important regions in the image, thereby improving the performance of the model. 4. Improving the generalization ability of the model: By better preserving the spatial locality characteristics of the image, the model can exhibit stronger generalization ability when processing different types of images. This improvement is not only applicable to specific datasets but also can play a role in a wider range of scenarios.

[0028] S4. Perform an equivalent transformation process on the Manhattan distance of the two-dimensional coordinates, and replace the position index of the raster scan with the sum of the coordinate values of the reallocated two-dimensional position coordinates of the image markers; Specifically, step S4 includes: performing an equivalent transformation process on the Manhattan distance of the two-dimensional coordinates so that the Manhattan distance of the two-dimensional coordinates can be formally aligned with the relative position distance of the image markers. The formula is: , where is the new position index calculated according to the two-dimensional position coordinates of the image marker at position m, is the abscissa of the image marker at position m, is the ordinate of the image marker at position m, is 's new position index, is 's new position index; Replace the position index of the raster scan with the sum of the coordinate values of the reallocated two-dimensional position coordinates of the image markers. Among them, the default self-attention calculation formula of the raster scan is: , where is the Query token at position index i q i and the Key token at position index j k j 's self-attention, is the matrix transpose, is the activation function, is the feature dimension, is the rotation matrix, is the predefined sine function value.

[0029] In this embodiment, an equivalent transformation process is performed on the Manhattan distance of two-dimensional coordinates. The purpose of this process is to align the Manhattan distance of two-dimensional coordinates with the relative position distance of image markers in form, so as to better integrate into the self-attention mechanism of the model. To be compatible with the traditional self-attention mechanism, the Manhattan distance of two-dimensional coordinates is converted into an equivalent one-dimensional index form. In this way, the Manhattan distance of two-dimensional coordinates is converted into a one-dimensional position index, so that it can be seamlessly docked with the traditional self-attention mechanism.

[0030] Immediately afterwards, the position index of raster scanning is replaced by the sum of the coordinate values of the reallocated two-dimensional position coordinates of the image markers. In the traditional self-attention mechanism, the position index is assigned by raster scanning. By adding a rotation matrix, the position information of the tokens is embedded in the self-attention. Further, the position index is replaced by the sum of the coordinate values of the reallocated two-dimensional position coordinates of the image markers. This replacement process not only preserves the spatial locality characteristics of the image, but also enables the model to more naturally process the spatial relationship between image markers. In this way, the Manhattan distance of two-dimensional coordinates can be effectively integrated into the self-attention mechanism, thereby optimizing the multi-modal information interaction ability of the model. Among them, the implementation scheme of the specific conversion process is as Figure 2 shown; under the current default position modeling mechanism, the causal attention mask modeling effect is as Figure 5 shown.

[0031] Simply put, by integrating the Manhattan distance of two-dimensional coordinates into the self-attention mechanism, the model can more accurately perceive the spatial relationship between image markers, thereby improving the accuracy of multi-modal feature alignment. Secondly, the attention allocation mechanism of the model is optimized, so that the model can more evenly focus on each region in the image, avoiding the attention bias caused by the raster scanning order in the traditional method. In addition, by preserving the spatial locality characteristics of the image, the model can show stronger generalization ability and robustness when dealing with complex multi-modal tasks.

[0032] S5. According to the new position index and the transformed Manhattan distance of two-dimensional coordinates, model the causal attention mask, and replace the default causal attention mask with the newly modeled causal attention mask; Specifically, step S5 includes: modeling the relative position of the image markers according to the transformed Manhattan distance of two-dimensional coordinates to obtain an adjusted position modeling mechanism, and based on this mechanism, obtaining a causal attention mask; Among them, the self-attention calculation formula is updated to: , the updated self-attention calculation formula is unified in form with the initial self-attention calculation formula; Replace the default causal attention mask with the newly modeled causal attention mask.

[0033] In this embodiment, the relative positions of image tokens are modeled according to the converted two-dimensional coordinate Manhattan distance. The core of this process is to incorporate the spatial relationship between image tokens into the attention mechanism of the model in a more natural way. Specifically, using the equivalent conversion result of the two-dimensional coordinate Manhattan distance, the relative position relationship between image tokens is redefined. This new position modeling mechanism not only retains the spatial locality characteristics of the image but also enables the model to more accurately perceive the spatial distance between image tokens.

[0034] Based on the new position modeling mechanism, a causal attention mask is further obtained. In traditional multi-modal large models, the causal attention mask is usually modeled based on one-dimensional position indices, which will ignore the two-dimensional spatial structure when dealing with image tokens. However, this method extends the modeling of the causal attention mask to the two-dimensional level by introducing the two-dimensional coordinate Manhattan distance. This extension not only optimizes the model's spatial perception ability for image tokens but also enables the attention mechanism to more naturally handle the interaction between images and texts.

[0035] To achieve this improvement, the self-attention calculation formula is updated. The updated self-attention calculation formula is consistent with the initial self-attention calculation formula in form, but its internal logic has been extended from one-dimensional to two-dimensional, thus better supporting the modeling of the spatial relationship of image tokens.

[0036] Finally, replace the default causal attention mask with the newly modeled causal attention mask, as Figure 6 shown. This replacement process not only optimizes the attention allocation mechanism of the model but also enables the model to more evenly focus on each region in the image, avoiding the attention bias caused by the raster scan order in traditional methods. In this way, the model can show stronger generalization ability and robustness when dealing with complex multi-modal tasks.

[0037] S6. Freeze the multi-modal large model, and perform adjustment training preprocessing on the multi-modal large model according to the replaced causal attention mask and the preset pre-training data until the multi-modal large model reaches the preset effect.

[0038] Specifically, step S6 includes: freezing the visual encoder and the large language model of the multi-modal large model, adjusting the multi-modal large model based on the replaced causal attention mask and the preset pre-training data, and only updating the parameters of its multi-layer perceptron part; Train the multimodal large model according to the parameters of the updated multi-layer perceptron part until the multimodal large model reaches a preset effect.

[0039] In this embodiment, freeze the visual encoder and the large language model of the multimodal large model. The freezing operation means that during the training process, the parameters of these pre-trained components remain unchanged. The visual encoder is responsible for encoding image data into feature representations that the model can process, while the large language model is responsible for processing text data. These two components have usually been pre-trained with large-scale data and have strong feature extraction capabilities. By freezing their parameters, we can focus on optimizing other parts of the model in subsequent training while avoiding performance degradation caused by retraining.

[0040] Next, adjust the multimodal large model based on the replaced causal attention mask and the preset pre-trained data. This adjustment process only updates the parameters of the multi-layer perceptron (MLP) part in the model. The multi-layer perceptron plays a role in connecting the visual encoder and the large language model in the model and is responsible for fusing image features and text features. By only updating the parameters of the multi-layer perceptron part, the model can be adjusted more efficiently to better adapt to specific task requirements while reducing the computational burden during the training process.

[0041] During the adjustment process, use the replaced causal attention mask to guide the interaction of multimodal information. This new causal attention mask is modeled based on the two-dimensional coordinate Manhattan distance, which can more accurately reflect the spatial relationship between image tokens, thereby optimizing the model's fusion of image and text features. In this way, the model can process multimodal data more effectively and reduce the occurrence of hallucination phenomena. Finally, further train the multimodal large model according to the updated parameters of the multi-layer perceptron part until the model reaches the preset effect. This training process usually involves fine-tuning the model to ensure that its performance on specific tasks reaches the optimal. By freezing the visual encoder and the large language model and only updating the parameters of the multi-layer perceptron part, the model can be adjusted more efficiently while maintaining its strong feature extraction capabilities.

[0042] Compared with the prior art, the training method for alleviating hallucinations in multimodal large models has the following beneficial effects: (1) It points out the deficiencies in the position encoding method currently widely used in multimodal large models. The long-term decay of the Rotary Position Embedding (RoPE) used for position modeling in multimodal large models causes uneven perception of instruction tokens towards image tokens located at different positions in the two-dimensional space: in the one-dimensional sequence, more attention is paid to the image tokens closer to the instruction tokens in the lower-right image area. This perceptual bias leads to insufficient information interaction between the image and the instruction and unsatisfactory multimodal feature alignment, thereby generating hallucinations. (2) It proposes to use Manhattan position assignment to replace the raster scan position index adopted by RoPE, retaining the local spatial attributes of the image: the causal attention of image tokens expands from the unidirectional decay of RoPE to multi-directional decay. In addition, compared with raster scan, the number of Manhattan position indices is reduced from 576 to 23, thus reducing the overall distance between the image and instruction tokens, which is more conducive to information interaction. (3) It proposes a position modeling method based on Manhattan distance, called Manhattan Causal Attention, which extends the long-term decay to two-dimensional and multi-directional spatial decay. For the two-dimensional continuous information contained in the image, the new attention calculation method retains the spatial local characteristics when establishing causal attributes, improving the model's image information perception ability and reducing the hallucination phenomenon. Figure 7 The results in Figure 7 show that compared with mainstream hallucination alleviation methods, the Manhattan Causal Attention (MCA-LLaVA) proposed in this method achieves the best or highly competitive results on different hallucination benchmarks.

[0043] In summary, the training method for alleviating hallucinations in multimodal large models significantly improves the performance and reliability of multimodal large models by optimizing position modeling and attention mechanisms. The core of this method lies in redefining the position relationship between image tokens and optimizing the model's perception and processing ability of image information by introducing two-dimensional Manhattan distance and improved causal attention masks. This improvement not only optimizes the model's multimodal information interaction ability but also provides important technical support for building a more reliable and efficient multimodal artificial intelligence system.

[0044] Specifically, the training method for alleviating hallucinations in multimodal large models first encodes image data into multiple image tokens through a visual encoder and introduces two-dimensional Manhattan distance to replace the traditional one-dimensional positional encoding. This improvement preserves the spatial locality features of the image, enabling the model to more accurately perceive the spatial relationships between image tokens. By reassigning the two-dimensional position coordinates of the image tokens and integrating these coordinate values into the attention mechanism, the model can more naturally handle the interaction between images and text, thereby optimizing the effect of multimodal feature alignment. Further, by freezing the pre-trained visual encoder and large language model and only fine-tuning the parameters of the multi-layer perceptron part, the training efficiency is significantly improved. This strategy not only saves computing resources but also avoids performance degradation caused by retraining. By introducing a new causal attention mask, the model can more evenly focus on each region in the image, reducing the attention bias caused by the raster scan order. This improvement significantly enhances the model's perception ability of image information, thereby effectively alleviating the hallucination problem.

[0045] Please refer to Figure 8 , the second embodiment of the present invention provides a training device for alleviating hallucinations in multimodal large models, which includes: An encoding unit 101, configured to obtain image data, input the image data into a multimodal large model, perform encoding processing on the image data, and obtain multiple image tokens; A distance calculation unit 102, configured to calculate the relative position distances between the image tokens and evolve the relative position distances from a one-dimensional level to a two-dimensional level to obtain a two-dimensional coordinate Manhattan distance; A redistribution unit 103, configured to perform raster scan processing on the image tokens, allocate position indices, and reassign two-dimensional position coordinates to the image tokens; An equivalent conversion unit 104, configured to perform equivalent conversion processing on the two-dimensional coordinate Manhattan distance and replace the position index of the raster scan with the sum of the coordinate values of the reallocated two-dimensional position coordinates of the image tokens; A replacement unit 105, configured to model the causal attention mask according to the new position index and the converted two-dimensional coordinate Manhattan distance, and replace the default causal attention mask with the newly modeled causal attention mask; A training unit 106, configured to freeze the multimodal large model, and perform adjustment training preprocessing on the multimodal large model according to the replaced causal attention mask and preset pre-training data until the multimodal large model reaches a preset effect.

[0046] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A training method for alleviating hallucinations in multimodal large models, characterized in that, Including: Obtain image data, input the image data into a multi-modal large model, perform encoding processing on the image data to obtain multiple image tokens; Calculate the relative position distances between the image tokens, and evolve the relative position distances from a one-dimensional level to a two-dimensional level to obtain the two-dimensional coordinate Manhattan distance; Perform raster scan processing on the image tokens, assign position indices, and re-assign two-dimensional position coordinates to the image tokens; Perform equivalent transformation processing on the two-dimensional coordinate Manhattan distance, and replace the position index of the raster scan with the sum of the coordinate values of the re-assigned two-dimensional position coordinates of the image tokens; Model the causal attention mask according to the new position index and the transformed two-dimensional coordinate Manhattan distance, and replace the default causal attention mask with the newly modeled causal attention mask; Freeze the multi-modal large model, and perform adjustment training preprocessing on the multi-modal large model according to the replaced causal attention mask and the preset pre-training data until the multi-modal large model reaches the preset effect.

2. The training method for alleviating hallucinations in multimodal large models according to claim 1, wherein, The multi-modal large model includes a visual encoder, a large language model, and a multi-layer perceptron for connecting the visual encoder and the large language model. Among them, the visual encoder uses CLIP ViT-L / 14, the large language model uses Vicuna-7B, the multi-layer perceptron includes a two-layer MLP, and the overall scale of the multi-modal large model is 7B.

3. The training method for alleviating hallucinations in multimodal large models according to claim 1, characterized in that, Perform encoding processing on the image data to obtain multiple image tokens. Specifically: Use the visual encoder of the multi-modal large model to perform encoding processing on each image of the image data; Among them, reset the resolution of each image to 24x24, and cut each image according to this size to obtain 576 image tokens.

4. The training method for alleviating hallucinations in multimodal large models according to claim 1, characterized in that, Calculate the relative position distances between the image tokens, and evolve the relative position distances from a one-dimensional level to a two-dimensional level to obtain the two-dimensional coordinate Manhattan distance. Specifically: The rotation position encoding technique is used to perform position encoding processing on image tokens, and the relative position distance of the image tokens is calculated. The calculation formula is as follows: , where is the Query token at position index i, is the Key token at position index j, represents and the relative position distance between them. This relative position distance is a one-dimensional distance. represents the position index of the image token determined by raster scanning, represents the position index of the j-th image token, represents the position index of the i-th image token; Expand the relative position distance from a one-dimensional level to a two-dimensional level to obtain the two-dimensional coordinate Manhattan distance, and its formula is: , where is the two-dimensional coordinate Manhattan distance, and are the two-dimensional horizontal and vertical coordinates of the Key marker at position j, respectively, and are the two-dimensional horizontal and vertical coordinates of the Query marker at position i, respectively.

5. The training method for alleviating hallucinations in multi-modal large models according to claim 1, characterized in that, Perform raster scan processing on the image tokens, assign position indices, and re-assign two-dimensional position coordinates to the image tokens. Specifically: Perform a raster scan on the image markers and in the manner of Allocate position indices incremented by 1, where is the number of image markers; Reassign two-dimensional position coordinates to each of the said image labels in the following manner: ; Set the positions of the image tokens at the four vertices of the image as the four starting coordinate origins (0, 0), and the marker points adjacent to the origin are set as the next marker points with an increment of 1. Among them, the positive directions of the horizontal and vertical coordinate axes are consistent with the increment direction; According to the four starting coordinate origins, divide the image tokens after resetting the resolution into four mirror-image parts, and the image tokens in each part increase linearly along the positive directions of the two-dimensional coordinate axes of the origin.

6. The training method for alleviating hallucinations in multi-modal large models according to claim 4, characterized in that, Perform equivalent transformation processing on the two-dimensional coordinate Manhattan distance, and replace the position index of the raster scan with the sum of the coordinate values of the re-assigned two-dimensional position coordinates of the image tokens. Specifically: Perform an equivalent transformation on the Manhattan distance of the two-dimensional coordinates so that the Manhattan distance of the two-dimensional coordinates can be formally aligned with the relative position distance of the image markers. The formula is: , where is the new position index calculated based on the two-dimensional position coordinates of the image marker at position m, is the abscissa of the image marker at position m, is the ordinate of the image marker at position m, is 's new position index, is 's new position index; Replace the position index of the raster scan with the sum of the coordinate values of the two-dimensional position coordinates of the reallocated image markers. The default self-attention calculation formula for the raster scan is as follows: , where is the Query token at position index i q i and the Key token at position index j k j is the self-attention between them, is the matrix transpose, is the activation function, is the feature dimension, is the rotation matrix, is the predefined sine function value.

7. The training method for alleviating hallucinations in multi-modal large models according to claim 6, characterized in that, Model the causal attention mask according to the new position index and the transformed two-dimensional coordinate Manhattan distance, and replace the default causal attention mask with the newly modeled causal attention mask. Specifically: Model the relative positions of the image tokens according to the transformed two-dimensional coordinate Manhattan distance to obtain an adjusted position modeling mechanism, and obtain a causal attention mask based on this mechanism; Among them, the self-attention calculation formula is updated to: , and the updated self-attention calculation formula is unified in form with the initial self-attention calculation formula; Replace the default causal attention mask with the newly modeled causal attention mask.

8. The training method for alleviating hallucinations in multimodal large models according to claim 1, wherein, Freeze the multi-modal large model, and perform preprocessing of adjustment training on the multi-modal large model according to the replaced causal attention mask and preset pre-training data until the multi-modal large model reaches the preset effect. Specifically: Freeze the visual encoder and the large language model of the multi-modal large model, and adjust the multi-modal large model based on the replaced causal attention mask and preset pre-training data, and only update the parameters of its multi-layer perceptron part; Train the multi-modal large model according to the updated parameters of the multi-layer perceptron part until the multi-modal large model reaches the preset effect.

9. A training device for alleviating hallucinations in multimodal large models, characterized in that, It includes: An encoding unit for obtaining image data, inputting the image data into a multi-modal large model, and performing encoding processing on the image data to obtain a plurality of image tokens; A distance calculation unit for calculating the relative position distances between the image tokens and evolving the relative position distances from a one-dimensional level to a two-dimensional level to obtain a two-dimensional coordinate Manhattan distance; A reallocation unit for performing raster scan processing on the image tokens, allocating position indexes, and re-assigning two-dimensional position coordinates to the image tokens; An equivalent conversion unit for performing equivalent conversion processing on the two-dimensional coordinate Manhattan distance and replacing the position index of the raster scan with the sum of the coordinate values of the re-assigned two-dimensional position coordinates of the image tokens; A replacement unit for modeling the causal attention mask according to the new position index and the converted two-dimensional coordinate Manhattan distance, and replacing the default causal attention mask with the newly modeled causal attention mask; A training unit for freezing the multi-modal large model, and performing preprocessing of adjustment training on the multi-modal large model according to the replaced causal attention mask and preset pre-training data until the multi-modal large model reaches the preset effect.

Citation Information

Patent Citations

  • Method, device and equipment for generating virtual clothing through multi-modal fusion and storage medium

    CN114723843A

  • Illusion relieving method and device for multi-modal large model, electronic equipment and medium

    CN119128061A

  • Multi-modal human body large model training method and system considering multi-granularity semantic alignment

    CN119625351A

  • Knowledge fusion multi-modal interaction method and apparatus based on improved alignment method

    WO2025025290A1