Image description method and device, electronic equipment and storage medium

By segmenting the image into sub-images and fusing visual information, combined with the perturbation attention optimization mechanism, the problem of low accuracy of multimodal large language models in image description is solved, especially the static resolution limiting of visual encoder and generative illusion, achieving higher image description accuracy.

CN120495831APending Publication Date: 2025-08-15SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510332968.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing multimodal large language model has the problem of low accuracy in image description, especially the deterministic hallucinations and generative hallucinations caused by the static resolution limitation of the visual encoder, which affects the application of the model in high-precision visual understanding scenarios.

Method used

By segmenting the target image into multiple sub-images, and fusing visual information using batch processing, and using a perturbation attention optimization mechanism, optimizing attention distribution, and generating multiple token information to improve the accuracy of image description.

Benefits of technology

Without increasing the computational complexity, the deterministic hallucination problem caused by the static resolution limit of visual encoder is solved, and the generative hallucination is reduced, significantly improving the accuracy of image description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495831A_ABST
    Figure CN120495831A_ABST
Patent Text Reader

Abstract

The invention provides an image description method and device, electronic equipment and a storage medium, and relates to the field of image processing. The method comprises the following steps: acquiring a target image and at least one target sub-image in the target image; the target image and each target sub-image carry the same text prompt; the text prompt is to prompt the content of the target image; performing feature extraction and feature fusion on the target image and each target sub-image to obtain a target feature; under the guidance of the text prompt, performing iterative processing on the target features by using a disturbance attention optimization mechanism to generate multiple pieces of token information; performing text generation according to each piece of token information, and outputting text description information; the text description information is used for describing the content of the target image. According to the invention, the problem of low accuracy of image description in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and more specifically, to an image description method, device, electronic device, and storage medium. Background Art

[0002] Multi-modal Large Language Models (MLLMs) have been a key research direction in the field of artificial intelligence in recent years. By combining pre-trained visual encoders with large language models (LLMs), they can handle multimodal tasks such as image captioning and visual question answering (VQA). Typical examples of MLLMs include LLaVA, MiniGPT-4, and SPHINX. By combining visual and language information, these models demonstrate powerful visual understanding and reasoning capabilities and are widely used in fields such as web navigation and autonomous driving.

[0003] However, MLLMs face a serious problem in practical applications: hallucination. Hallucination refers to the model's failure to provide accurate and realistic answers based on the given image content when generating responses. For example, when the model is asked to describe an image, it may generate a description that includes objects that are not present in the image. This phenomenon severely limits the practical application of MLLMs, especially in scenarios requiring high-precision visual understanding.

[0004] From the above, we can see that how to improve the accuracy of image description remains to be solved. Summary of the Invention

[0005] This application provides an image description method, device, electronic device, and storage medium, which can solve the problem of low image description accuracy in related technologies. The technical solution is as follows:

[0006] According to one aspect of the present application, an image description method includes: obtaining a target image and at least one target sub-image in the target image; the target image and each target sub-image carry the same text prompt; the text prompt is a prompt for the content of the target image; feature extraction and feature fusion are performed on the target image and each target sub-image to obtain target features; under the guidance of the text prompt, the target features are iteratively processed using a perturbation attention optimization mechanism to generate multiple token information; text is generated according to each token information, and text description information is output; the text description information is used to describe the content of the target image.

[0007] According to one aspect of the present application, an image description device includes: an image acquisition module, used to acquire a target image and at least one target sub-image in the target image; the target image and each target sub-image carry the same text prompt; the text prompt is a prompt for the content of the target image; a feature processing module, used to extract and fuse features of the target image and each target sub-image to obtain target features; a token generation module, used to iteratively process the target features under the guidance of the text prompt and use a perturbation attention optimization mechanism to generate multiple token information; a text generation module, used to generate text according to each token information and output text description information; the text description information is used to describe the content of the target image.

[0008] According to one aspect of the present application, an electronic device includes at least one processor and at least one memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the image description method as described above is implemented.

[0009] According to one aspect of the present application, a storage medium stores a computer program thereon, and when the computer program is executed by one or more processors, the image description method described above is implemented.

[0010] According to one aspect of the present application, a computer program product includes a computer program, and when the computer program is executed by one or more processors, the computer program implements the above image description method.

[0011] The beneficial effects of the technical solution provided by this application are:

[0012] In the above technical solution, it is possible to process a target image of any resolution without increasing the computational complexity, divide the target image into multiple sub-images, and fuse the visual information of each target sub-image using batch processing, thereby solving the problem of deterministic hallucination caused by the static resolution limitation of the visual encoder; through the perturbation attention optimization mechanism, the attention distribution is optimized without significantly increasing the computational complexity, and the generative hallucination is reduced, thereby improving the accuracy of the image description, which can effectively solve the problem of low image description accuracy existing in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts.

[0014] Figure 1 It is a schematic diagram of the implementation environment involved in this application;

[0015] Figure 2 is a hardware structure diagram of an electronic device according to an exemplary embodiment;

[0016] Figure 3 is a flowchart of an image description method according to an exemplary embodiment;

[0017] Figure 4 yes Figure 2 A flowchart of an embodiment corresponding to step 350 in an embodiment;

[0018] Figure 5 is a structural block diagram of an image description device according to an exemplary embodiment;

[0019] Figure 6 The figure is a structural block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0020] The following describes embodiments of the present application in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application.

[0021] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present disclosure refers to the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" as used herein includes all or any units and all combinations of one or more associated listed items.

[0022] As mentioned above, the phenomenon severely limits the practical application of MLLMs, especially in scenarios requiring high-precision visual understanding.

[0023] Various technical solutions are currently being proposed to address the problem of hallucinations. For example, OPERA reduces hallucinations by designing better metrics to prevent the "attention sink" phenomenon; VCD mitigates hallucinations caused by language bias by introducing visual contrast decoding. However, most of these approaches operate from a single perspective and fail to provide a fine-grained classification and analysis of the hallucination problem.

[0024] THRONE categorizes hallucinations into two types: generative hallucinations and deterministic hallucinations. Generative hallucinations occur during open-ended generative processes, such as when generating image descriptions that include non-existent objects. Deterministic hallucinations occur when the model cannot correctly answer deterministic questions about the image's content.

[0025] However, the static resolution limitation of visual encoders is the main cause of deterministic hallucinations. Most pre-trained visual encoders (such as CLIP) can only process images with fixed resolutions (such as 224×224 or 336×336). Therefore, during the preprocessing stage, the images are cropped, scaled, and normalized, resulting in the loss of some visual information. This information loss directly leads to errors in the model's answering of questions about the content of the image. In addition, generative hallucinations are related to the "attention sink" phenomenon, that is, during the generation process, the model pays too much attention to certain "sink tokens" and ignores other tokens containing local information, resulting in inaccurate generated content.

[0026] As can be seen from the above, the related art still has the defect of low accuracy of image description.

[0027] To this end, the image description method provided in this application can effectively improve the accuracy of image description. Accordingly, the image description method is suitable for an image description device, which can be deployed in an electronic device. The electronic device can be a computer device configured with a von Neumann architecture, for example, the computer device includes a desktop computer, a laptop computer, a server, etc.

[0028] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0029] Figure 1 It should be noted that this implementation environment is only an example adapted to the present invention and should not be considered as providing any limitation on the scope of application of the present invention.

[0030] The implementation environment includes a collection end 110 and a service end 130 .

[0031] Specifically, the acquisition terminal 110 can also be considered as an image acquisition device, including but not limited to electronic devices with shooting functions such as cameras, cameras, and video recorders. For example, the acquisition terminal 110 is an underwater camera.

[0032] Server 130 can be an electronic device such as a desktop computer, laptop computer, or server, or a computer cluster consisting of multiple servers, or even a cloud computing center consisting of multiple servers. Server 130 is used to provide background services, such as, but not limited to, image description services.

[0033] A network communication connection is pre-established between the server 130 and the acquisition terminal 110 via a wired or wireless method, and data transmission between the server 130 and the acquisition terminal 110 is achieved via the network communication connection. The transmitted data includes but is not limited to: target images, etc.

[0034] In one application scenario, through the interaction between the acquisition terminal 110 and the server 130 , the acquisition terminal 110 captures and acquires a target image, and uploads the target image to the server 130 to request the server 130 to provide an image description service.

[0035] For the server 130, after receiving the target image uploaded by the acquisition terminal 110, it calls the image description service to describe the target image and output text description information, so as to solve the problem of inaccurate image description in related technologies.

[0036] See also Figure 2 , Figure 2 This is a hardware structure diagram of an electronic device according to an exemplary embodiment. Figure 1 A server 130 is shown in an implementation environment.

[0037] It should be noted that the electronic device is only an example adapted for this application and cannot be considered to provide any limitation on the scope of use of this application. The electronic device cannot be interpreted as needing to rely on or must have Figure 2 One or more components of exemplary electronic device 200 are shown.

[0038] The hardware structure of the electronic device 200 may vary greatly due to different configurations or performances, such as Figure 2 As shown, the electronic device 200 includes a power supply 210 , an interface 230 , at least one memory 250 , and at least one central processing unit (CPU) 270 .

[0039] Specifically, the power supply 210 is used to provide operating voltage for various hardware devices on the electronic device 200 .

[0040] The interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices. Figure 1 The interaction between the collection end 110 and the service end 130 in the implementation environment is shown.

[0041] Of course, in other examples adapted by this application, the interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input-output interface 235, and at least one USB interface 237, etc. Figure 2 As shown, this does not constitute a specific limitation.

[0042] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon include an operating system 251, application 253 and data 255, etc. The storage method can be temporary storage or permanent storage.

[0043] Among them, the operating system 251 is used to manage and control the various hardware devices and application programs 253 on the electronic device 200 to enable the central processing unit 270 to calculate and process the massive data 255 in the memory 250. It can be WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0044] The application 253 is a computer program formed by computer-readable instructions based on the operating system 251 to perform at least one specific task, and may include at least one module ( Figure 2 Each module may include corresponding computer-readable instructions. For example, the image description device may be considered as an application 253 deployed on the electronic device 200.

[0045] The data 255 may be photos, pictures, etc. stored in a disk, or may be target images, etc. stored in the memory 250 .

[0046] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read the computer program stored in the memory 250, thereby performing operations and processing on the massive data 255 in the memory 250. For example, the image description method is implemented by the central processing unit 270 reading the application program 253 stored in the memory 250.

[0047] In addition, the present application can also be implemented through hardware circuits or hardware circuits combined with software. Therefore, the implementation of the present application is not limited to any specific hardware circuits, software, or a combination of the two.

[0048] See also Figure 3 , the embodiment of the present application provides an image description method, which is applicable to an electronic device, for example, the electronic device may be Figure 1 The server 130 in the implementation environment is shown, and the hardware structure of the electronic device can be as follows Figure 2 shown.

[0049] In the following method embodiments, for ease of description, the execution subject of each step of the method is taken as an electronic device as an example for illustration, but this does not constitute a specific limitation.

[0050] like Figure 3 As shown, the method may include the following steps:

[0051] Step 310: Acquire a target image and at least one target sub-image in the target image.

[0052] The target image and each target sub-image carry the same text prompt, and the text prompt is a prompt of the content of the target image.

[0053] First, the target image is the image to be described. It can contain diverse visual content. The text prompt can be a user-assigned task instruction, clearly indicating the task is to generate an image description, rather than something else, such as "Describe the content of the target image in text." Therefore, using the same text prompt for the target image and each target sub-image ensures task consistency.

[0054] The content of the target image may refer to any visual content, for example, the visual content may be a person, a vehicle, a landscape, etc.

[0055] The target image can be obtained by shooting and collecting through an image acquisition device. Among them, the image acquisition device can be an electronic device with an image acquisition function, such as a camera, a smart phone equipped with a camera, etc. It can be understood that shooting can be a single shot or continuous shooting. Then, for the same environment, for continuous shooting, a video can be obtained, and the image can be any frame in the video. For multiple shots, multiple photos can be obtained, and the image can be any one of the multiple photos. In other words, the image in this embodiment can come from a dynamic image, such as multiple frames in a video, or multiple photos, and can also come from a static image, such as any frame in a video, or any one of multiple photos. Accordingly, the target tracking in this embodiment is performed in frames.

[0056] Regarding image acquisition, the image can be captured in real time by an image acquisition device, or it can be a previously stored image captured and collected by the image acquisition device within a historical time period. Therefore, after the image acquisition device captures and acquires the image, the electronic device can process the image in real time, or it can pre-store the image for later processing, for example, when the electronic device's CPU is low, or according to instructions from a staff member. Therefore, target tracking in this embodiment can be performed on images captured in real time or on images captured within a historical time period, without any specific limitation here.

[0057] Furthermore, since visual encoders (such as CLIP) can only process images with a fixed resolution (such as 224×224), in order to adapt to this limitation, the target image will be cropped, scaled, and normalized during the preprocessing stage, resulting in a large amount of visual information loss, especially for images with extreme aspect ratios. To avoid this problem, the static resolution limitation of the visual encoder can be solved by splitting the target image into multiple sub-images and fusing the information of the target sub-images.

[0058] Each target sub-image is an image obtained after image segmentation is performed on the target image. It can be understood that the target sub-image corresponds to a partial image area of the target image.

[0059] It should be noted that simply dynamically scaling or adaptively cropping the target image will result in information loss or increased computational complexity, thereby affecting the accuracy of subsequent image description.

[0060] In a possible implementation, the target image is segmented according to a set resolution to obtain target sub-images.

[0061] The resolution setting can be related to the static resolution of the visual encoder. For example, assuming the resolution of the input image is H×W and the static resolution of the visual encoder is H0×W0 (CLIP's visual encoder is usually 224×224), then, in order to adapt to the static resolution of the visual encoder, the target image is divided into multiple target sub-images. Specifically, the target image is divided into n target sub-images, and the resolution of each target sub-image is H0×W0.

[0062] Of course, since the resolution of the target image may not be an integer multiple of H0×W0, the center cropping operation of the target image will result in a large amount of edge information loss, especially for images with extreme aspect ratios. Therefore, this solution can perform zero padding on the target image to ensure that the resolution of all target sub-images is consistent, thereby preserving all the visual information of the target image and avoiding visual information loss.

[0063] Regarding the target image segmentation method, it can be uniform segmentation, for example, dividing the image into multiple 224×224 sub-images at a fixed step size; it can also be adaptive segmentation based on the content in the target image, for example, dynamically adjusting the position of the target sub-image according to the content of the target image (such as edge detection), which is not limited here.

[0064] For example, statistics on the ImageNet-val and MSCOCO-val datasets show that the target features obtained through the above steps can significantly reduce the hallucination problem caused by information loss, especially in images with extreme aspect ratios.

[0065] Through the above process, it is possible to process target images of arbitrary resolution without increasing computational complexity, divide the target image into multiple sub-images, and fuse the visual information of each target sub-image using batch processing, thus solving the deterministic hallucination problem caused by the static resolution limitation of the visual encoder. It not only retains all the visual information of the target image, but also avoids the loss of visual information.

[0066] Step 330 : performing feature extraction and feature fusion on the target image and each target sub-image to obtain target features.

[0067] First of all, it should be noted that the target image can be the original input image, which contains complete visual content, and the target sub-image is a local area obtained by segmenting the target image, which focuses on specific visual content.

[0068] In one possible implementation, step 330 includes the following steps: performing feature extraction on the target image and the target sub-image respectively to obtain corresponding image features; and performing image fusion on each image feature according to an output distribution averaging method to obtain target features.

[0069] Specifically, features of the target image and each target sub-image can be extracted through a visual encoder to generate corresponding image features.

[0070] Regarding feature fusion, the target image and the image features corresponding to each target sub-image may be spliced into a joint input sequence to generate a target feature, and the target feature may be accompanied by the same text prompt.

[0071] Regarding the output distribution averaging method, specifically, assuming there are n target sub-images and a target image, the final target feature is obtained by calculating the average value of the corresponding image feature distribution. The output distribution averaging method is shown in formula (1).

[0072]

[0073] Among them, D_(l+i) is the image feature distribution of the i-th sub-image, D_g is the image feature distribution of the target image, and D_ens is the fused target feature.

[0074] Step 350: Under the guidance of the text prompt, the perturbation attention optimization mechanism is used to iteratively process the target features to generate multiple token information.

[0075] The token information may include a query vector, a key vector, and a value vector.

[0076] Regarding generating token information, the text prompt can first be converted into a text embedding vector through the embedding layer of the multimodal large language model, and then the target feature can be spliced with the text embedding vector or fused through cross attention to generate an initial input sequence. Finally, iterative processing is performed based on the initial input sequence to generate multiple token information.

[0077] Regarding iterative processing, it means generating the first token information based on the aforementioned initial input sequence. The first token information can be obtained by using the decoder to process the initial input sequence through self-attention. Then, based on the initial input sequence and the first token information, the decoder can be used to process the generated token sequence through self-attention, and interact with the initial input sequence through cross-attention to generate the second token information. And so on, based on the initial input sequence and the generated token information sequence, subsequent tokens are gradually generated until the generation of all token information is completed.

[0078] For example, if the text prompt is "Describe this picture" and the target features include visual semantics of "car", "green trees" and "mountains", cross attention will cause the model to focus on the global scene (such as "this is a picture") rather than the detailed object when generating the first token, and gradually focus on local details (such as "red car") when generating subsequent tokens.

[0079] In one possible implementation, such as Figure 4 As shown, step 350 may include the following steps:

[0080] Step 351: decode the target feature according to the text prompt to generate current token information.

[0081] The current token information refers to the currently generated token information.

[0082] As mentioned above, iterative generation depends on the initial input sequence and the generated token information sequence, rather than just the current token information.

[0083] Based on this, if the current token information is the first token information generated in time sequence, then it can be decoded based on the target features and text prompts to obtain the current token information; if the current token information is not the first token information generated in time sequence, then it is necessary to decode based on the target features and text prompts, as well as the generated token information sequence, to obtain the current token information.

[0084] Step 353: Perform attention calculation based on the query vector corresponding to the current token information and the key vector corresponding to the generated token information to obtain the attention distribution.

[0085] Here, the query vector is associated with the text prompt.

[0086] Regarding attention calculation, when generating the jth token, the attention distribution of the jth token and the generated token information is calculated. Specifically, the query vector of the current token information is used to calculate the dot product of all the key vectors in the generated token information, and then the attention distribution is obtained through the Softmax function. The specific process is shown in Formula (2).

[0087]

[0088] Among them, Q_j^(p,q) is the query vector of the j-th token information, and K_i^(p,q) is the key vector of the i-th token information. is the calculated attention distribution.

[0089] In step 355, a perturbation loss is calculated based on the attention distribution to obtain a perturbation objective function, and the perturbation objective function is used to continue generating subsequent token information based on the target features until all token information is generated.

[0090] First of all, it should be noted that the "attention sink" phenomenon will occur in the process of token information generation, that is, excessive attention will be paid to certain "aggregate tokens" (such as global information tokens) while tokens containing local detail information will be ignored, resulting in generative hallucinations.

[0091] Specifically, a perturbation objective function can be defined to measure the uniformity of attention distribution. The perturbation objective function can include two parts: Attention Density Loss and Distillation Loss.

[0092] The attention density loss measures the uniformity of the attention distribution by calculating its entropy. A larger entropy indicates a more uniform attention distribution. The distillation loss is used to ensure that the difference between the perturbed output distribution and the original output distribution is not too large.

[0093] Based on this, the perturbation loss can be calculated based on the attention distribution to obtain the perturbation objective function.

[0094] First, a key-value cache (KV cache) is a data structure used in the Transformer architecture to store key-vector (Key) and key-vector (Value) pairs of generated token information. In autoregressive generation tasks (such as text generation), the key-value cache can be used to accelerate inference and optimize attention calculations.

[0095] In one possible implementation, a key-value cache is obtained; attention calculation is performed based on the query vector corresponding to the current token information and the key vector stored in the key-value cache to obtain an attention distribution.

[0096] Among them, the key-value cache is used to store the generated token information.

[0097] Regarding key-value caching, during the token generation process, the model caches the key vector and value pair vector for each token, known as the key-value cache. This cache is used to accelerate inference and influence the model's attention distribution when generating the next token.

[0098] It should be noted that when token information is first generated, the key-value cache is empty. As the generation process progresses, the key vector and value vector of each newly generated token information will be added to the key-value cache.

[0099] Regarding attention calculation, specifically, the query vector of the current token information can be used to perform dot product calculation with all key vectors in the key-value cache, and then the attention distribution can be obtained through the Softmax function.

[0100] In one possible implementation, based on the attention distribution, the perturbation objective function of the key-value cache is calculated by back propagation; the key-value cache is updated according to the perturbation objective function to obtain an updated key-value cache, so as to generate subsequent token information based on the updated key-value cache.

[0101] In order to enable the model to pay more attention to all token information more evenly during the generation process, the key-value cache needs to be perturbed to make the attention distribution more uniform.

[0102] Based on this, the key vectors and value vectors in the key-value cache can be updated according to the perturbation objective function. Specifically, according to the perturbation objective function, a perturbation vector is added to each key vector and value vector in the key-value cache, so that the attention distribution is more uniform when the next token information is generated. This is shown in Formula (3).

[0103] K j ←K j +ΔK j , V j ←V j +ΔV j ...Formula (3)

[0104] Among them, ΔK_j and ΔV_j are the perturbation vectors calculated according to the perturbation objective function.

[0105] In one possible implementation, perturbation prediction is performed on the key-value cache by backpropagation to obtain a predicted key-value cache; loss is calculated based on the key-value cache and the predicted key-value cache to obtain a distillation loss; the entropy of the attention distribution is maximized to obtain an attention density loss; and the distillation loss and the attention density loss are combined to obtain a perturbation objective function.

[0106] Regarding the predicted key-value cache, the original output distribution is generated based on the initial key-value cache. Then, the gradient of the perturbation objective function is calculated through backpropagation to generate the perturbation amount. Next, the perturbation amount is applied to the original key-value cache to generate the predicted key-value cache. Finally, the perturbed output distribution is generated based on the predicted key-value cache, and the distillation loss is calculated. It should be understood that the distillation loss and the perturbation objective function are calculated simultaneously through backpropagation.

[0107] In one possible implementation, the perturbation attention optimization mechanism is used. Experimental results on the CHAIR dataset show that the perturbation attention optimization mechanism significantly reduces generative hallucinations. LLaVA1.5: CsCs (the proportion of hallucinated sentences) is reduced from 48.2% to 46.0% (2.2%), and CiCi (the proportion of hallucinated objects) is reduced from 12.4% to 11.3% (1.1%). LLaVA1.6Vicuna: CsCs is reduced from 32.8% to 31.2% (1.6%), and CiCi is reduced from 8.1% to 7.9% (0.2%). LLaVAphi3: CsCs is reduced from 32.8% to 30.8% (2.0%), and CiCi is reduced from 8.8% to 7.5% (1.3%).

[0108] Through the above process, the key-value cache optimizes the attention distribution and reduces generative hallucinations by storing the key vectors and key vectors of the generated token information and perturbing the key-value cache without significantly increasing the computational complexity.

[0109] In addition, by defining the attention density loss and distillation loss, the image description process can focus more evenly on all tokens, especially the token information containing local information.

[0110] Step 370: Generate text based on each token information and output text description information.

[0111] The text description information is used to describe the content of the target image.

[0112] Specifically, token information is the smallest unit of text generation. Token information can be a word (such as "car") or a subword (such as "car-car"). Then, text description information can be output through token information.

[0113] In one possible implementation, the above-mentioned image description method is implemented through an image description model, which is a trained multimodal large language model that has the ability to describe the content of the target image.

[0114] First of all, it should be noted that Multi-modal Large Language Models (MLLMs) can combine pre-trained visual encoders with large language models (LLMs) to handle multi-modal tasks and process image and text information simultaneously.

[0115] In one possible implementation, the image description model may be LLaVA, MiniGPT-4, SPHINX, etc.

[0116] In one possible implementation, the image description model includes a visual encoder for extracting visual features of the target image; it may also include a language decoder for generating a text description based on the visual features; and it may also include a multimodal fusion module for fusing the visual features with the text features to generate a joint representation.

[0117] In another possible implementation, the image description model includes a key-value cache for storing value vectors and key vectors in token information to accelerate inference and optimize attention distribution.

[0118] The training process of the image description model may include obtaining a large-scale image-text pair dataset (such as COCO, Flickr30k), inputting the large-scale image-text pair dataset into the image description model for image description to obtain training text descriptions, and training the image description model using the training text descriptions and sample text descriptions in the large-scale image-text pair dataset until a trained image description model is obtained.

[0119] Through the above process, it is possible to process target images of arbitrary resolution without increasing the computational complexity, divide the target image into multiple sub-images, and fuse the visual information of each target sub-image using batch processing, thereby solving the deterministic hallucination problem caused by the static resolution limitation of the visual encoder; through the perturbation attention optimization mechanism, the attention distribution is optimized and the generative hallucination is reduced without significantly increasing the computational complexity. This scheme significantly reduces the hallucination problem while maintaining the coherence of the generated text description information, thereby improving the accuracy of the image description.

[0120] In this application scenario, the image description method is implemented through an image description model, which can combine pre-trained visual encoders with large language models (LLMs) to handle multi-modal tasks and process image and text information simultaneously.

[0121] Specifically, the SPHINX-Zero method is used to solve the deterministic hallucination problem caused by the static resolution limitation of the visual encoder by dividing the target image into multiple target sub-images and fusing the information of the target sub-images. The NUBIS-Zero method is also used to perturb the key-value cache, so that the image description model can pay more even attention to all token information during the generation process, thereby alleviating generative hallucinations.

[0122] Specifically, the MSCOCO dataset, AOKVQA dataset, and GQA dataset are input into the image description model to verify the image description model's effectiveness in solving deterministic hallucinations.

[0123] Summary of the results of the illusion of certainty:

[0124] In the evaluation of the deterministic hallucination problem, SPHINXZero performs well on multiple datasets, especially on the MSCOCO and AOKVQA datasets, significantly outperforming other decoding methods.

[0125] Specific manifestations:

[0126] 1.MSCOCO dataset:

[0127] In the Random setting, SPHINXZero achieves an accuracy of 91.37%, which is 1.99% higher than Greedy decoding. In terms of F1 score, SPHINXZero also achieves 91.37%, significantly outperforming other methods.

[0128] Under the Popular setting, SPHINXZero achieves an accuracy of 89.20%, which is 3.30% higher than Greedy decoding, and an F1 score of 89.13%, which is also the best performance.

[0129] In the Adversarial setting, SPHINXZero achieves an accuracy of 82.10%, which is 3.10% higher than Greedy decoding, and an F1 score of 83.19%, outperforming other methods.

[0130] 2.AOKVQA dataset:

[0131] Under the Random setting, SPHINXZero achieves an accuracy of 89.27%, which is 3.50% higher than Greedy decoding, and an F1 score of 89.94%, significantly outperforming other methods.

[0132] Under the Popular setting, SPHINXZero achieves an accuracy of 84.23%, which is 4.30% higher than Greedy decoding, and an F1 score of 85.88%, the best performance.

[0133] In the Adversarial setting, SPHINXZero achieves an accuracy of 72.23%, which is 3.13% higher than Greedy decoding, and an F1 score of 77.55%, which is close to the state-of-the-art.

[0134] 3.GQA dataset:

[0135] Under the Random setting, SPHINXZero achieves an accuracy of 89.10%, which is 3.30% higher than Greedy decoding, and an F1 score of 89.83%, the best performance.

[0136] Under the Popular setting, SPHINXZero achieves an accuracy of 77.62%, a 2.85% improvement over Greedy decoding, and an F1 score of 81.17%, which is close to the state-of-the-art.

[0137] In the Adversarial setting, SPHINXZero achieves an accuracy of 72.30%, a 2.83% improvement over Greedy decoding, and an F1 score of 77.66%, which is close to the state-of-the-art.

[0138] Summary of advantages:

[0139] SPHINXZero performs best on multiple datasets, especially on the MSCOCO and AOKVQA datasets, significantly outperforming other decoding methods.

[0140] The advantage of SPHINXZero is that it can process images of arbitrary resolution. By dividing the image into multiple sub-images and fusing global and local information, it avoids the problem of information loss in traditional visual encoders due to static resolution limitations.

[0141] The hyperparameter design of SPHINXZero is simple, requiring only the design of a canvas set, and is easy to implement and adjust.

[0142] Specifically, this application scenario inputs the CHAIR dataset and OPOPE dataset into the image description model to verify the image description model's effectiveness in solving generative hallucinations.

[0143] Summary of Generative Hallucination Results:

[0144] In the evaluation of generative hallucination problems, ANUBISZero performs well on multiple datasets, especially on the CHAIR dataset, significantly reducing generative hallucinations.

[0145] Specific manifestations:

[0146] 1. CHAIR Dataset:

[0147] On the LLaVA1.5 model, ANUBISZero achieved a hallucinated sentence ratio of 46.0%, a 2.2% reduction compared to Greedy decoding, and a hallucinated object ratio of 11.3%, a 1.1% reduction compared to Greedy decoding. ANUBISZero performed best in reducing generative hallucinations.

[0148] On the LLaVA1.6Vicuna model, ANUBISZero achieves 31.2%, which is 1.6% lower than Greedy decoding, and 7.9%, which is 0.2% lower than Greedy decoding, performing close to the best.

[0149] On the LLaVAphi3 model, ANUBISZero achieves the best performance with an accuracy of 30.8%, which is 2.0% lower than Greedy decoding and 7.5%, which is 1.3% lower than Greedy decoding.

[0150] 2. OPOPE dataset:

[0151] On the LLaVA1.5 model, ANUBISZero achieves an accuracy of 85.27%, which is 1.65% higher than Greedy decoding, and an F1 score of 85.56%, the best performance.

[0152] On the LLaVA1.6 Vicuna model, ANUBISZero achieves an accuracy of 84.01%, which is 4.21% higher than Greedy decoding, and an F1 score of 77.17%, the best performance.

[0153] On the LLaVAphi3 model, ANUBISZero achieves an accuracy of 79.72%, which is 0.17% higher than Greedy decoding, and an F1 score of 76.07%, the best performance.

[0154] Summary of advantages:

[0155] ANUBISZero performs best on multiple datasets, especially on the CHAIR dataset, significantly reducing the generative hallucination problem.

[0156] ANUBISZero perturbs the KV cache, allowing the model to pay more attention to local information when generating the next token, thereby reducing generative hallucinations.

[0157] The advantage of ANUBISZero is that it can dynamically adjust the attention distribution, allowing the model to pay more attention to detailed information when generating descriptions, avoiding the hallucination problem caused by traditional methods that over-focus on global information.

[0158] In this application scenario, SPHINXZero and ANUBISZero performed well in handling deterministic hallucinations and generative hallucinations, respectively, significantly outperforming existing decoding methods. By processing images of arbitrary resolution, SPHINXZero avoids the problem of information loss in traditional visual encoders due to static resolution limitations, significantly improving the accuracy of the model on object existence issues. By perturbing the key-value cache, ANUBISZero enables the image description model to pay more attention to local information when generating text description information, significantly reducing the problem of generative hallucinations. The introduction of these two methods provides important improvements to the reliability of multimodal large language models (MLLMs) in practical applications, especially in scenarios requiring high-precision visual understanding and generation.

[0159] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0160] The following is an embodiment of the device of the present application, which can be used to perform the image description method involved in the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the method embodiment of the image description method involved in the present application.

[0161] See also Figure 5 In an embodiment of the present application, an image description device 900 is provided, including but not limited to: an image acquisition module 910, a feature processing module 930, a token generation module 950, and a text generation module 970.

[0162] The image acquisition module 910 is used to acquire a target image and at least one target sub-image in the target image; the target image and each target sub-image carry the same text prompt; and the text prompt is a prompt of the content of the target image.

[0163] The feature processing module 930 is used to extract and fuse features of the target image and each target sub-image to obtain target features.

[0164] The token generation module 950 is used to iteratively process the target features using the perturbation attention optimization mechanism under the guidance of text prompts to generate multiple token information.

[0165] The text generation module 970 is used to generate text according to each token information and output text description information; the text description information is used to describe the content of the target image.

[0166] It should be noted that the image description device provided in the above embodiment only uses the division of the above functional modules as an example when performing image description. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the image description device will be divided into different functional modules to complete all or part of the functions described above.

[0167] In addition, the image description device and the image description method provided in the above embodiments belong to the same concept, wherein the specific manner in which each module performs operations has been described in detail in the method embodiments and will not be repeated here.

[0168] See also Figure 6 In an embodiment of the present application, an electronic device 4000 is provided. The electronic device 4000 may include: a desktop computer, a laptop computer, a server, etc.

[0169] exist Figure 6 In the embodiment, the electronic device 4000 includes at least one processor 4001 and at least one memory 4003.

[0170] Data exchange between the processor 4001 and the memory 4003 can be achieved via at least one communication bus 4002. The communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. The communication bus 4002 may be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0171] Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.

[0172] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0173] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store a computer program in the form of instructions or data structures and can be accessed by the electronic device 400, but is not limited to these.

[0174] The memory 4003 stores a computer program, and the processor 4001 can read the computer program stored in the memory 4003 through the communication bus 4002 .

[0175] The computer program is executed by one or more processors 4001 to implement the image description method in the above embodiments.

[0176] In addition, an embodiment of the present application provides a storage medium on which a computer program is stored. The computer program is executed by one or more processors to implement the above image description method.

[0177] An embodiment of the present application provides a computer program product, including a computer program, which is executed by one or more processors to implement the above image description method.

[0178] Compared with related technologies, this scheme can process target images of arbitrary resolution without increasing computational complexity, divide the target image into multiple sub-images, and fuse the visual information of each target sub-image using batch processing, thus solving the deterministic hallucination problem caused by the static resolution limitation of the visual encoder; through the perturbation attention optimization mechanism, the attention distribution is optimized without significantly increasing computational complexity, the generative hallucination is reduced, and the accuracy of image description is improved.

[0179] The above are only some of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. An image description method, characterized in that: include: Acquire a target image and at least one target sub-image in the target image; The target image and each target sub-image carry the same text prompt; The text prompt is to prompt the content of the target image; Performing feature extraction and feature fusion on the target image and each target sub-image to obtain target features; Under the guidance of the text prompt, the target feature is iteratively processed using a perturbation attention optimization mechanism to generate multiple token information; Text is generated according to each of the token information, and text description information is output; the text description information is used to describe the content of the target image.

2. The method according to claim 1, wherein The acquiring of a target image and at least one target sub-image in the target image comprises: The target image is segmented according to a set resolution to obtain each target sub-image.

3. The method according to claim 1, wherein The step of extracting and fusing features of the target image and each of the target sub-images to obtain target features includes: Performing feature extraction on the target image and each target sub-image respectively to obtain corresponding image features; According to the output distribution averaging method, image fusion is performed on each of the image features to obtain the target feature.

4. The method according to claim 1, wherein Under the guidance of the text prompt, the target feature is iteratively processed using the perturbation attention optimization mechanism to generate multiple token information, including: Decoding the target feature according to the text prompt to generate current token information; Performing attention calculation based on a query vector corresponding to the current token information and a key vector corresponding to the generated token information to obtain an attention distribution; wherein the query vector is related to the text prompt; A perturbation loss is calculated based on the attention distribution to obtain a perturbation objective function, and the perturbation objective function is used to continue generating subsequent token information based on the target feature until all token information is generated.

5. The method according to claim 4, wherein The attention calculation is performed based on the query vector corresponding to the current token information and the key vector corresponding to the generated token information to obtain the attention distribution, including: Obtain a key-value cache, wherein the key-value cache is used to store the generated token information; Performing attention calculation based on the query vector corresponding to the current token information and the key vector stored in the key-value cache to obtain the attention distribution; The perturbation loss is calculated based on the attention distribution to obtain a perturbation objective function, and the perturbation objective function is used to continue generating subsequent token information based on the target feature until all token information is generated, including: Based on the attention distribution, calculating the perturbation objective function of the key-value cache by backpropagation; The key-value cache is updated according to the perturbation target function to obtain an updated key-value cache, so as to generate subsequent token information based on the updated key-value cache.

6. The method according to claim 5, wherein The calculating the perturbation objective function of the key-value cache by back propagation based on the attention distribution includes: Performing perturbation prediction on the key-value cache by back-propagation to obtain a predicted key-value cache; Performing loss calculation based on the key value cache and the predicted key value cache to obtain a distillation loss; Attention density loss is obtained by maximizing the entropy of the attention distribution; The perturbation objective function is obtained by merging the distillation loss and the attention density loss.

7. The method according to any one of claims 1 to 6, characterized in that The image description method is implemented through an image description model, which is a trained multimodal large language model capable of describing the content of the target image.

8. An image description device, characterized in that: include: An image acquisition module, configured to acquire a target image and at least one target sub-image in the target image; The target image and each target sub-image carry the same text prompt; The text prompt is used to prompt a text description of the content of the target image; A feature processing module, configured to extract and fuse features of the target image and each target sub-image to obtain target features; A token generation module, configured to iteratively process the target features using a perturbation attention optimization mechanism under the guidance of the text prompt to generate a plurality of token information; The text generation module is used to generate text according to each token information and output text description information; the text description information is used to describe the content of the target image.

9. An electronic device comprising at least one processor and at least one memory, wherein: A computer program is stored in the memory, wherein when the computer program is executed by the processor, the image description method according to any one of claims 1 to 7 is implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by one or more processors, the image description method according to any one of claims 1 to 7 is implemented.