Image description generation method and system based on enhanced fine-grained information

By combining the regional feature extractor and the global feature extractor, the dynamic multi-view enhancement mechanism and the hierarchical memory enhancement encoder are used to optimize the cross-modal feature information, and the shortcomings of the image description generation method in the prior art in fine-grained information capture and multimodal fusion are solved, and high-quality image description generation is achieved.

CN120298714APending Publication Date: 2025-07-11NANJING UNIV OF SCI & TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510337136.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When the existing image description generation method processes complex images and fine-grained information, it is difficult to effectively capture fine-grained information in the image, resulting in the lack of accuracy and refinement of the generated description, and insufficient fusion of multimodal information.

Method used

The regional feature extractor and global feature extractor are combined to obtain multi-view feature information through weighted fusion, and a dynamic multi-view enhancement mechanism and a hierarchical memory enhancement encoder are used, combined with a multi-layer Transformer structure to optimize the cross-modal feature information, and finally generate image descriptions through the dual cross-modal attention mechanism.

Benefits of technology

Improve the accuracy and refinement of image description, ensure the semantic consistency between local objects and global scenes, and generate high-quality image descriptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298714A_ABST
    Figure CN120298714A_ABST
Patent Text Reader

Abstract

The invention discloses an image description generation method and system based on enhanced fine-grained information. The method comprises the following steps: 1) extracting regional features and global features of an image, and carrying out weighted fusion to generate multi-view feature information; 2) combining text feature extraction and cosine similarity screening, dynamically fusing text-image features, and optimizing cross-modal features through a dynamic multi-view enhancement mechanism; (3) a hierarchical memory enhancement encoder is adopted, multi-head attention and Transform structure hierarchical coding features are combined, and detail information is reserved through a memory unit; and 4) aggregating image and text weights in two stages through a double cross-modal attention decoder to generate high-quality description. The system comprises a feature extraction module, a text processing module, an encoding module and a decoding module. Through fine-grained feature enhancement and cross-modal interaction, the problems of detail loss and semantic deviation in a traditional method are effectively solved, and the method is suitable for image semantic analysis of complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision and natural language processing, and particularly relates to an image caption generation method and system based on enhanced fine-grained information. Background Art

[0002] With the rapid development of computer vision and natural language processing technologies, image caption generation technology has been widely applied in multiple fields such as automatic content creation, image retrieval, and assistive technologies. However, existing image caption generation methods still have certain limitations when dealing with complex images and fine-grained information. Many existing models rely on traditional image feature extraction methods, such as feature extraction based on convolutional neural networks (CNNs), but these methods often ignore the fine-grained information in images and the deep semantic relationships between images and texts.

[0003] Although significant progress has been made in deep learning models in recent years, especially models based on vision-language pre-training, existing image caption generation methods generally still face the following challenges: First, how to efficiently extract the fine-grained information in images and capture the semantic relationships between various objects in the images; Second, how to maintain the accuracy of text generation and the consistency of context in image captions; Finally, existing image caption models still lack sufficient optimization in multi-modal information fusion, resulting in image captions that often fail to accurately express the complex details and context of images.

[0004] To address these problems, recent research has proposed a variety of innovative methods, including image caption generation models based on attention mechanisms. These methods improve the alignment of image features and text features through self-attention and cross-modal attention mechanisms. However, most of these methods still do not fully consider how to effectively utilize the fine-grained information in images, and when dealing with complex scenes, the generated descriptions lack accuracy and meticulousness.

[0005] Therefore, how to accurately capture the fine-grained information in images and generate more accurate and detailed image captions has become an urgent problem to be solved in the field of image caption generation. Summary of the Invention

[0006] The object of the present invention is to provide an image caption generation method and system based on enhanced fine-grained information, which realizes accurately capturing the fine-grained information in images and improves the quality and accuracy of image captions.

[0007] To achieve the object of the present invention, on the one hand, the present invention provides an image caption generation method based on enhanced fine-grained information, including the following steps:

[0008] Step 1: Process the input image through a regional feature extractor to obtain the regional feature information of the image, then process it through a global feature extractor to obtain the global feature information of the image, and finally fuse the regional feature information and the global feature information to obtain the multi-perspective feature information of the image;

[0009] 1.1. Use the regional feature extractor to process the input image, extract the key objects and their spatial relationships in the image, and obtain the regional feature information of the image;

[0010] 1.2. Use the global feature extractor to process the input image, capture the overall structure, background information and relationships between regions of the image, and obtain the global feature information of the image;

[0011] 1.3. Perform weighted fusion processing on the regional feature information and the global feature information to obtain the multi-perspective feature information of the image, ensuring the semantic consistency between the details of local objects and the global scene.

[0012] Step 2: Extract features from the text data corresponding to the input image, screen out the text feature information most relevant to the input image, fuse it with the multi-perspective feature information of the image to obtain comprehensive text-image feature information, and optimize the comprehensive text-image feature information through a dynamic multi-perspective enhancement mechanism to obtain cross-modal feature information;

[0013] 2.2. Based on the multi-perspective feature information of the image and the text feature information, screen out the text features most relevant to the input image by calculating the cosine similarity;

[0014] 2.3. Connect and fuse the most relevant text features with the multi-perspective feature information of the image to construct a comprehensive text-image feature information;

[0015] 2.4. Use the dynamic multi-perspective enhancement mechanism to optimize the comprehensive text-image feature information, and finally obtain cross-modal feature information.

[0016] Step 3: Input the cross-modal feature information into a hierarchical memory-enhanced encoder and adopt a hierarchical feature integration strategy to obtain the encoded cross-modal feature information;

[0017] Step 3.1. Take the multi-perspective feature information of the image and the cross-modal feature information as inputs and send them into the hierarchical memory-enhanced encoder for feature encoding; during the encoding process, first use the multi-head attention mechanism to refine the local feature information and the global feature information of the image at the local level, and then obtain the encoded local feature information through the hierarchical feature interaction mechanism;

[0018] Step 3.2: Store the encoded local feature information in a memory unit, further optimize the cross-modal feature information through a multi-layer Transformer structure, and then perform final feature fusion through a weighted aggregation strategy to obtain the encoded cross-modal feature information.

[0019] Step 4: Input the encoded cross-modal feature information into a dual cross-modal attention decoder, adopt a dual cross-modal attention mechanism to mine the interaction information between the image and the text, and finally generate a high-quality image description.

[0020] The specific dual cross-modal attention mechanism is as follows:

[0021] Step 4.1: The first-level attention module performs attention calculations on the image feature information and text feature information in the encoded cross-modal feature information respectively to obtain the image feature weight and text feature weight in the encoded cross-modal feature information.

[0022] Step 4.2: After fusing the image feature weight and the text feature weight, input them into the dual cross-modal attention decoder, and adopt a second-level attention mechanism to mine the interaction information between the image and the text, and finally generate a high-quality image description.

[0023] On the other hand, the present invention also provides a system for an image description generation method based on enhanced fine-grained information, including the following modules:

[0024] A feature extraction module, used to extract the regional features and global features of the image from the input image.

[0025] A text feature processing module, used to calculate the similarity between the image and the text features and perform fusion.

[0026] An encoding module, used to capture the complex dependency relationship between the image and the text by combining a multi-layer Transformer encoder and a memory unit.

[0027] A decoding module, used to generate accurate image descriptions using a dual cross-modal attention mechanism.

[0028] Compared with the prior art, the significant progress of the present invention lies in that: by adopting a hierarchical memory enhancement mechanism and combining a multi-layer Transformer structure to encode cross-modal features, the present invention improves the model's ability to capture fine-grained image information and the accuracy of text-image alignment, effectively solving the problems of detail loss and semantic deviation in traditional methods.

[0029] To more clearly illustrate the functional characteristics and structural parameters of the present invention, the following further explains with reference to the drawings and specific embodiments. Description of the Drawings

[0030] The accompanying drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation of the present invention. In the drawings:

[0031] Figure 1 is the overall model flowchart of the fine-grained image description generation method based on hierarchical memory enhancement and cross-modal attention mechanism of the present invention;

[0032] Figure 2 is the schematic diagram of image feature extraction in the present invention, showing the fusion process of the region feature extractor and the global feature extractor;

[0033] Figure 3 is the schematic diagram of the hierarchical perspective integration strategy in the present invention, showing how to integrate image features at the local level and the global level;

[0034] Figure 4 is the schematic diagram of the dual cross-modal attention mechanism in the present invention, showing how to process image and text information through two-stage attention modules;

[0035] Figure 5 is the example diagram of the output result of the present invention. Detailed implementation manners

[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.

[0037] An image description generation method based on enhanced fine-grained information of the present invention, in combination with Figure 1 , includes the following steps:

[0038] Step 1, in combination with Figure 2 , the input image is processed by a region feature extractor to obtain the region feature information of the image, and then it is processed by a global feature extractor to obtain the global feature information of the image. Subsequently, the region feature information and the global feature information are fused to obtain the multi-perspective feature information of the image;

[0039] 1.1. Use the region feature extractor to process the input image, extract the key objects and their spatial relationships in the image, and obtain the region feature information of the image;

[0040] 1.2. Process the input image using a global feature extractor to capture the overall structure, background information, and relationships between regions of the image, and obtain the global feature information of the image;

[0041] 1.3. Through weighted fusion processing of the region feature information and the global feature information, obtain the multi-view feature information of the image, ensuring the semantic consistency between the details of local objects and the global scene.

[0042] Step 2. Combine Figure 3 , perform feature extraction on the text data corresponding to the input image, filter out the text feature information most relevant to the input image, and fuse it with the multi-view feature information of the image to obtain comprehensive text-image feature information. Optimize the comprehensive text-image feature information through a dynamic multi-view enhancement (DMVE) mechanism to obtain cross-modal feature information;

[0043] 2.1. Process the text data corresponding to the input image using a text feature extractor to obtain text feature information and extract syntactic and semantic structures;

[0044] 2.2. Based on the multi-view feature information of the image and the text feature information, filter out the text features most relevant to the input image by calculating the cosine similarity to ensure a high degree of alignment between the text content and the image semantics;

[0045] 2.3. Connect and fuse the most relevant text features with the multi-view feature information of the image to construct a comprehensive text-image feature information to enhance the mutual understanding between text and image;

[0046] 2.4. Use the dynamic multi-view enhancement mechanism to optimize the comprehensive text-image feature information, and finally obtain cross-modal feature information to further improve the semantic alignment ability between image and text features.

[0047] Step 3. Input the cross-modal feature information into a hierarchical memory-enhanced encoder and adopt a hierarchical feature integration strategy to obtain the encoded cross-modal feature information;

[0048] 3.1. Take the multi-view feature information of the image and the cross-modal feature information as inputs and send them into the hierarchical memory-enhanced encoder for feature encoding; during the encoding process, first use the multi-head attention mechanism at the local level to refine the local feature information and the global feature information of the image, and then through the hierarchical feature interaction mechanism, obtain the encoded local feature information, enabling dynamic interaction of information between different modalities to ensure that the model can capture cross-modal correlation relationships;

[0049] 3.2. Store the encoded local feature information in a memory unit to prevent information loss during deep fusion; further optimize the cross-modal feature information through a multi-layer Transformer structure, and then perform final feature fusion through a weighted aggregation strategy to obtain the encoded cross-modal feature information, which synthesizes the features from various perspectives to ensure that the model can focus on local object details while maintaining the integrity of the global scene.

[0050] Step 4. Input the encoded cross-modal feature information into a dual cross-modal attention decoder, and use the dual cross-modal attention mechanism to mine the interaction information between the image and the text, and finally generate a high-quality image description.

[0051] The dual cross-modal attention mechanism is specifically as follows:

[0052] Step 4.1. The first-level attention module independently calculates the attention for the image feature information and the text feature information in the encoded cross-modal feature information to obtain the image feature weights and text feature weights in the encoded cross-modal feature information.

[0053] Step 4.2. After fusing the image feature weights and the text feature weights, input them into the dual cross-modal attention decoder, and use the second-level attention mechanism to mine the interaction information between the image and the text, and finally generate a high-quality image description.

[0054] Furthermore, the formula for calculating the cosine similarity in step 2.2 is as follows:

[0055]

[0056] where v is the image feature vector, j represents the index of the text feature, and t j represents the j-th text feature vector.

[0057] Furthermore, the memory unit in step 3.2 passes key information between the image features and the text features through a multi-layer Transformer encoder to ensure that no detail information is lost when generating the description.

[0058] Furthermore, combined with Figure 4 , the dual cross-modal attention mechanism is specifically as follows:

[0059] First, the first-level attention module independently calculates the attention for the image feature information and the text feature information in the encoded cross-modal feature information to obtain the corresponding image feature weights and text feature weights; then, the second-level attention module aggregates the image feature weights and the text feature weights and inputs them into the dual cross-modal attention decoder to generate the final text description.

[0060] The present invention also includes a system for an image description generation method based on enhanced fine-grained information, comprising the following modules:

[0061] A feature extraction module, configured to extract the regional features and global features of an image from the input image;

[0062] A text feature processing module, configured to calculate the similarity between the image and the text features and fuse them;

[0063] An encoding module, configured to capture the complex dependencies between the image and the text by combining a multi-layer Transformer encoder and a memory unit;

[0064] A decoding module, configured to generate accurate image descriptions using a dual cross-modal attention mechanism.

[0065] Combined Figure 5 , the model of the present invention accurately captures objects and their environments by enhancing fine-grained information, making the description more complete and natural. The results obtained from the input pictures on the left in the figure are: A bicycle is parked on the grass near a bridge by the river; The results obtained from the input pictures on the right are: An elephant and a rhinoceros are standing on the grass covered with trees.

[0066] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.

[0067] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An image description generation method based on enhanced fine-grained information, characterized in that It includes the following steps: Step 1: Process the input image through a regional feature extractor to obtain the regional feature information of the image, then process it through a global feature extractor to obtain the global feature information of the image, and finally fuse the regional feature information and the global feature information to obtain the multi-view feature information of the image; Step 2: Extract features from the text data corresponding to the input image, screen out the text feature information most relevant to the input image, and fuse it with the multi-view feature information of the image to obtain comprehensive text-image feature information. Optimize the comprehensive text-image feature information through a dynamic multi-view enhancement mechanism to obtain cross-modal feature information; Step 3: Input the cross-modal feature information into a hierarchical memory-enhanced encoder and adopt a hierarchical feature integration strategy to obtain the encoded cross-modal feature information; Step 4: Input the encoded cross-modal feature information into a dual cross-modal attention decoder, adopt a dual cross-modal attention mechanism to mine the interaction information between the image and the text, and finally generate a high-quality image description.

2. The method for generating an image description based on enhanced fine-grained information according to claim 1, wherein The said Step 1 includes the following steps: 1.1 Use the regional feature extractor to process the input image, extract the key objects in the image and their spatial relationships, and obtain the regional feature information of the image; 1.2 Use the global feature extractor to process the input image, capture the overall structure, background information and relationships between regions of the image, and obtain the global feature information of the image; 1.3 Perform weighted fusion processing on the regional feature information and the global feature information to obtain the multi-view feature information of the image, ensuring the semantic consistency between the details of local objects and the global scene.

3. A method for generating an image description based on enhanced fine-grained information according to claim 1, wherein, The said Step 2 includes the following steps: 2.1 Use the text feature extractor to process the text data corresponding to the input image to obtain text feature information and extract syntactic and semantic structures; 2.2 Based on the multi-view feature information of the image and the text feature information, screen out the text features most relevant to the input image by calculating the cosine similarity; 2.3 Connect and fuse the most relevant text features with the multi-view feature information of the image to construct a comprehensive text-image feature information; 2.4 Use the dynamic multi-view enhancement mechanism to optimize the comprehensive text-image feature information to finally obtain cross-modal feature information.

4. A method for generating an image description based on enhanced fine-grained information according to claim 1, characterized in that, The said Step 3 includes the following steps: Step 3.1: Use the multi-view feature information of the image and the cross-modal feature information as inputs and send them into the hierarchical memory-enhanced encoder for feature encoding; during the encoding process, first use the multi-head attention mechanism at the local level to refine the local feature information and the global feature information of the image, and then obtain the encoded local feature information through the hierarchical feature interaction mechanism; Step 3.2: Store the encoded local feature information in a memory unit, further optimize the cross-modal feature information through a multi-layer Transformer structure, and then perform final feature fusion through a weighted aggregation strategy to obtain the encoded cross-modal feature information.

5. A method for generating an image description based on enhanced fine-grained information according to claim 1, wherein The dual cross-modal attention mechanism of the said Step 4 includes the following steps: Step 4.1: The first-level attention module calculates the attention for the image feature information and text feature information in the encoded cross-modal feature information respectively, and obtains the image feature weight and text feature weight in the encoded cross-modal feature information; Step 4.2: After fusing the image feature weight and text feature weight, input them into the dual cross-modal attention decoder, and adopt the second-level attention mechanism to mine the interaction information between the image and the text, and finally generate a high-quality image description.

6. The method for generating an image description based on enhanced fine-grained information according to claim 3, wherein The formula for calculating the cosine similarity in Step 2.2 is as follows: Among them, v is the image feature vector, j represents the index of the text feature, and t j represents the j-th text feature vector.

7. A method for generating an image description based on enhanced fine-grained information according to claim 4, wherein The memory unit in Step 3.2 transfers key information between the image features and text features through multiple layers of Transformer encoders to ensure that no detailed information is lost when generating the description.

8. A system for an image description generation method based on enhanced fine-grained information according to any one of claims 1-7, characterized in that, It includes the following modules: A feature extraction module for extracting the regional features and global features of the image from the input image; A text feature processing module for calculating and fusing the similarity between the image and text features; An encoding module for capturing the complex dependencies between the image and text by combining multiple layers of Transformer encoders and a memory unit; A decoding module for generating accurate image descriptions using the dual cross-modal attention mechanism.

Citation Information

Cited By

  • Cross-modal endoscopic surgery image text description generation method and system

    CN121354119A

  • A method and system for generating text descriptions of endoscopic surgical images across modalities.

    CN121354119B

  • Text-guided spectrum image description method, system and device and storage medium

    CN122265805A

  • A text-guided spectral image description method, system, device and storage medium

    CN122265805B