Image description generation method and device, equipment and medium

By generating visual embedding and label text embedding vectors, analyzing similarity scores and generating weighted heat maps, the problem of insufficient significant region recognition and semantic alignment capabilities in the prior art is solved, and precise extraction and association description of multiple key objects and scenes in the image is realized, and high-quality image description text is generated.

CN120070972APending Publication Date: 2025-05-30PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510136802.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has insufficient accurate recognition ability of significant areas in image description generation and weak semantic alignment ability of multi-label scenes, making it difficult to achieve accurate extraction and association description of multiple key objects and scenes in the image.

Method used

By acquiring image data, the embedding vectors of visual embedding and label text are generated, the similarity score between the image data and each label text is analyzed, the weighted heat map is generated, the significant areas in the image data and their corresponding label and spatial position information are determined, and the image description is finally generated.

Benefits of technology

It realizes accurate recognition and semantic correlation of prominent areas and multi-label scenes in the image, and generates semantic coherent image description text, improving the accuracy of image description, detail capture ability and text readability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070972A_ABST
    Figure CN120070972A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and the field of medical health, and discloses an image description generation method, which comprises the following steps: acquiring image data, generating embedding vectors of visual embedding and label texts, analyzing a similarity score between the image data and each label text, generating a weighted thermodynamic diagram according to the similarity score, and generating a weighted thermodynamic diagram according to the weighted thermodynamic diagram. And based on the weighted thermodynamic diagram, determining the salient region and the corresponding label and spatial position information thereof, and finally generating an image description. According to the method, through multi-modal feature matching and weighted thermodynamic diagram generation, the salient region and the multi-label scene in the image are accurately recognized, the image description text with semantic coherence is generated, the accuracy, the detail capturing capability and the text readability of image description are improved, and the method is suitable for intelligent image-text analysis and auxiliary decision making in the fields of medical health and finance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and medical health, and particularly to an image description generation method, apparatus, device, and storage medium. Background Art

[0002] With the continuous development of artificial intelligence technology, multi-modal generation technology has gradually emerged in the fields of image description, semantic understanding, and intelligent question answering. By integrating visual information and text information, multi-modal generation technology can achieve automatic understanding and description of image content, thereby providing more intelligent auxiliary decision-making and services in the fields of medical health and finance. However, existing image description generation methods still have multiple technical bottlenecks, restricting their effectiveness and efficiency in practical applications.

[0003] In the prior art, image description generation methods mainly rely on traditional deep learning models such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs). Such methods typically extract visual features of an image through a visual encoder and then convert the visual features into corresponding text descriptions through a text generation model. Such methods can generate image descriptions of a certain quality in specific scenarios, but they have the following deficiencies in complex business scenarios:

[0004] First, the efficiency of image feature extraction of existing methods is low, and it is difficult to meet the requirements of large-scale data processing in real time. Traditional visual encoders can usually only perform coarse-grained feature extraction on the overall image and cannot effectively focus on significant regions and important detail information in the image. This method is particularly obvious in the field of medical health. For example, in the intelligent diagnosis of medical images, the system often needs to quickly and accurately identify the lesion area and generate a description, but the prior art is difficult to extract key features in real time and accurately, resulting in the accuracy of the diagnosis result being affected.

[0005] Second, the semantic alignment ability of the prior art is weak, and the semantic correlation between the generated image description and the image content is not high. In the insurance business scenario in the financial field, image description generation technology can be used to assist in identifying objects and scenes in claim settlement site pictures and generating relevant descriptions. However, traditional image description generation methods are prone to problems such as inaccurate description and semantic incoherence in the case of multiple labels and multiple scenarios. This problem of insufficient semantic alignment results in the system being difficult to generate accurate graphic and text information feedback when identifying pictures submitted by customers, affecting the claim settlement efficiency and customer experience.

[0006] Finally, the prior art has deficiencies in the structured generation of image description texts, and the generated descriptions often lack coherence and readability. In medical report generation, the system needs to generate text descriptions that conform to medical terms and diagnostic logic, but the text descriptions generated by existing methods are often syntactically incoherent and lack hierarchy and logic. In the scenario of graphic Q&A in the financial field, the picture descriptions submitted by customers need to match the business logic and the content of customers' questions, but the text description fragments generated by existing methods usually lack context correlation and are difficult to provide valuable feedback. Summary of the Invention

[0007] The main object of the present invention is to provide an image description generation method, device, equipment and storage medium, aiming to solve the technical problems that the prior art is insufficient in the accurate recognition of significant regions and the semantic alignment ability in multi-label scenarios, and it is difficult to achieve the accurate extraction and associated description of multiple key objects and scenes in the image.

[0008] To achieve the above object, the present invention provides an image description generation method, including:

[0009] Obtain image data, input the image data into a visual encoding module, and generate a visual embedding of the image data;

[0010] Perform text embedding processing on multiple label texts through a text encoding module to obtain embedding vectors of the label texts;

[0011] Analyze the similarity scores between the image data and each label text based on the visual embedding and the embedding vectors of the label texts;

[0012] According to the similarity scores, generate heatmap channels for each label text, and perform weighted processing on the heatmap channels to generate a weighted heatmap;

[0013] Determine the significant regions in the image data based on the weighted heatmap, and determine the labels and spatial position information corresponding to each significant region;

[0014] Generate an image description of the image data according to the labels and spatial position information corresponding to each significant region.

[0015] Furthermore, to achieve the above object, the present invention provides an image description generation device, including:

[0016] A visual encoding module, configured to obtain image data, input the image data into the visual encoding module, and generate a visual embedding of the image data;

[0017] A text encoding module, configured to perform text embedding processing on multiple label texts through the text encoding module to obtain embedding vectors of the label texts;

[0018] A similarity analysis module, configured to analyze the similarity scores between the image data and each label text based on the visual embedding and the embedding vectors of the label texts;

[0019] A heatmap generation module, configured to generate a heatmap channel for each label text according to the similarity scores, and perform weighted processing on the heatmap channels to generate a weighted heatmap;

[0020] A salient region identification module, configured to determine the salient regions in the image data based on the weighted heatmap, and determine the labels and spatial location information corresponding to each salient region;

[0021] A description generation module, configured to generate an image description of the image data according to the labels and spatial location information corresponding to each salient region.

[0022] Furthermore, to achieve the above object, the present invention also provides a computer device, which includes a memory, a processor, and an image description generation program stored in the memory and executable on the processor. When the image description generation program is executed by the processor, the steps of the image description generation method as described above are implemented.

[0023] Furthermore, to achieve the above object, the present invention also provides a computer-readable storage medium, on which an image description generation program is stored. When the image description generation program is executed by a processor, the steps of the image description generation method as described above are implemented.

[0024] Beneficial effects: The present invention relates to the fields of artificial intelligence technology and medical and health fields, and discloses an image description generation method, including: obtaining image data, generating visual embeddings and embedding vectors of label texts, analyzing the similarity scores between the image data and each label text, generating a weighted heatmap according to the similarity scores, determining the salient regions and their corresponding labels and spatial location information based on the weighted heatmap, and finally generating an image description. Through multi-modal feature matching and weighted heatmap generation, the present invention accurately identifies the salient regions and multi-label scenarios in the image, generates semantically coherent image description texts, improves the accuracy, detail capture ability and text readability of the image description, and is applicable to intelligent graphic analysis and auxiliary decision-making in the fields of medical and health and finance. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0026] Figure 1 is a schematic diagram of an application environment of the image description generation method in an embodiment of the present invention;

[0027] Figure 2Schematic flowchart of an embodiment of the image description generation method of the present invention;

[0028] Figure 3 Schematic diagram of functional modules of a preferred embodiment of the image description generation device of the present invention;

[0029] Figure 4 Schematic diagram of a structure of a computer device in an embodiment of the present invention;

[0030] Figure 5 Schematic diagram of another structure of a computer device in an embodiment of the present invention. Detailed implementation manners

[0031] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0032] The image description generation method provided by the embodiments of the present invention can be applied to an application environment such as Figure 1 where the client communicates with the server through a network. The server can obtain image data through the client, generate embedding vectors of visual embeddings and label texts, analyze the similarity scores between the image data and each label text, generate a weighted heat map according to the similarity scores, determine the salient regions and their corresponding labels and spatial position information based on the weighted heat map, and finally generate an image description. Through multi-modal feature matching and weighted heat map generation, the present invention accurately identifies salient regions and multi-label scenarios in images, generates semantically coherent image description texts, improves the accuracy, detail capture ability, and text readability of image descriptions, and is applicable to intelligent graphic analysis and auxiliary decision-making in the fields of medical health and finance. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smartphones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail through specific embodiments below.

[0033] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the image description generation method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.

[0034] As Figure 2 shown, the image description generation method proposed by the present invention includes the following steps:

[0035] S10, Obtain image data, input the image data into a visual encoding module, and generate a visual embedding of the image data;

[0036] In this embodiment, the image data can come from various sources. Common acquisition methods include reading image files from local storage devices, receiving images uploaded by users through network interfaces, calling image data from the API interfaces of third-party platforms, etc. In practical applications, the appropriate image data source can be selected according to business requirements. For example, in applications in the medical and health field, the image data may come from the medical imaging systems of hospitals (such as medical imaging devices like CT and MRI), while in applications in the financial field, the image data may be photos uploaded by customers through mobile devices.

[0037] To improve the efficiency and stability of data acquisition, it is usually necessary to adopt multi-threaded or asynchronous processing methods to manage the reading and transmission of image data. In the scenario of receiving through network interfaces, it is necessary to consider the integrity and consistency verification of image data to ensure that the image data is not lost or damaged during transmission.

[0038] After acquiring the image data, it is usually necessary to preprocess the image data to ensure that the format, size, and resolution of the image data meet the requirements of subsequent processing. The preprocessing steps may include operations such as image format conversion (such as converting PNG format to JPG format), image compression, image cropping, and scaling. These operations can effectively reduce the storage and transmission costs of image data and improve the computational efficiency of image data in subsequent processing.

[0039] In the medical and health field, special attention needs to be paid to maintaining the resolution and color mode of medical images during preprocessing to ensure that the diagnostic accuracy of the image data is not compromised. In the financial field, preprocessing usually requires desensitizing sensitive information (such as ID numbers, license plate numbers, etc.) in the image data to meet privacy protection and compliance requirements.

[0040] In the process of acquiring image data, a reasonable storage and management scheme needs to be designed. Common storage methods include local storage, cloud storage, and distributed file systems. To improve the access efficiency of image data, a caching mechanism can be adopted to temporarily store frequently used image data in memory to reduce frequent disk read operations.

[0041] In the medical and health field, the storage and management of image data need to comply with relevant laws and regulations on data security and privacy protection. In the financial field, the management of image data needs to consider long-term storage and version management of data to ensure that historical image data can be accessed at any time during the claims settlement process.

[0042] Example: In the application in the field of healthcare, a hospital's medical imaging system generates a large amount of medical imaging data every day (such as CT, MRI, X-ray, etc.). These medical imaging data can be automatically extracted from the hospital's imaging storage system and loaded into an intelligent diagnosis system to assist doctors in quickly generating imaging reports. For example, during the diagnosis of a lung CT image, the system automatically obtains the patient's CT image data and automatically annotates and describes the lesion area in the lungs, reducing the workload of doctors and improving the diagnosis efficiency.

[0043] In the application in the financial field, insurance companies often require customers to upload picture data of the accident scene during the claims settlement process. The claims settlement system of the insurance company can receive the accident scene photos uploaded by customers in real time and automatically identify the vehicle damage situation and the on-site environment in the photos, providing image descriptions and analysis results for the claims handlers to help complete the claims settlement review quickly.

[0044] By obtaining image data, it can provide basic data support for subsequent image processing and image description generation. The image data acquisition process is more efficient and automated, and can adapt to different data source requirements, thereby improving the overall processing efficiency and accuracy of the image description generation method.

[0045] Before inputting the image data into the visual coding module, it is necessary to perform format conversion on the input image data to ensure that it is consistent with the input requirements of the visual coding module. Common format conversions include normalization, size adjustment, channel conversion, etc. of the image data. For example, normalizing the pixel values of the image data to between 0 and 1 for the convenience of neural network processing; adjusting the image to a fixed resolution (such as 224x224 pixels) to meet the input size requirements of the model.

[0046] By calling an image processing library (such as OpenCV or Pillow), preprocess the image data, convert the image into the tensor format required by the visual coding module, and ensure that the dimensions, channel order, and data type of the input data match the input format of the model.

[0047] In the field of healthcare, medical images are usually in DICOM format and need to be converted into a common image format (such as PNG or JPG) while maintaining the resolution and color information of the images without distortion.

[0048] The core task of the visual coding module is to extract the visual features of the image and convert the input raw image data into a high-dimensional visual embedding vector. The visual coding module is usually composed of a convolutional neural network (CNN). The network extracts different levels of feature information from the image through multiple layers of convolution, pooling, and non-linear activation operations, including edges, textures, shapes, and higher-level semantic information.

[0049] Using a pre-trained visual encoding model (such as ResNet, EfficientNet, or Vision Transformer), input the image data into the model to obtain the corresponding visual embedding vector.

[0050] After extracting the image features, the visual encoding module outputs a high-dimensional embedding vector, which represents the global feature information of the image data. The visual embedding vector is usually a fixed-length array or tensor used to represent the feature distribution of the image. This embedding vector can be directly used for subsequent similarity analysis or multimodal matching tasks.

[0051] The visual embedding vector output by the visual encoding module needs to undergo certain post-processing operations, such as dimensionality reduction, normalization, or redundancy removal, to ensure that the embedding vector can accurately represent the core features of the image. The embedding vector can be stored in memory for use by subsequent similarity analysis or heatmap generation modules.

[0052] Example illustration: In the field of healthcare, a hospital's medical imaging system can generate a large number of medical images (such as CT, MRI, X-ray, etc.). By inputting this medical image data into the visual encoding module, the core visual features of each image can be extracted, such as information about the edges, density, and shape of the lesion area. These visual embedding vectors can help doctors identify lesions more quickly and improve the accuracy of diagnosis.

[0053] In the insurance claims scenario in the financial field, the accident scene photos uploaded by customers can be used by the visual encoding module to extract damage features of the accident vehicle, details of the scene environment, etc. These visual embedding vectors can help the claims system more accurately identify the type of accident and the damage situation, providing important auxiliary information for claims review.

[0054] By inputting image data into the visual encoding module and generating visual embedding vectors, the global features and semantic information of the image can be effectively extracted, enhancing the representation ability of image data in multimodal tasks. It can automatically extract the deep features of the image, reducing manual intervention and improving the efficiency and accuracy of image description generation.

[0055] S20, perform text embedding processing on multiple label texts through a text encoding module to obtain the embedding vectors of the label texts;

[0056] In this embodiment, the input of the text encoding module is multiple label texts, which are usually preset by the system or obtained from external data sources. For example, in the field of healthcare, the label texts can be disease names, symptom descriptions, organ parts, etc.; in the financial field, the label texts can be vehicle damage types, insurance types, claim conditions, etc.

[0057] When obtaining label text, the system needs to ensure the accuracy and integrity of this text to avoid deviations in subsequent processing due to incomplete or incorrect label data.

[0058] The label text is loaded from a predefined label library or dynamically obtained from an external system through an API interface. The format of the label text is usually natural language phrases or a list of keywords, and it is necessary to ensure that the semantics of the text is clear and there is no ambiguity.

[0059] Before inputting the label text into the text encoding module, it is necessary to perform word segmentation on the label text, splitting the long text into several words or phrases. Word segmentation can use rule-based word segmentation, dictionary matching word segmentation, or a machine learning-based word segmentation model. In Chinese processing, word segmentation is particularly important because Chinese text does not have obvious space delimiters, and word segmentation algorithms are needed to determine the boundaries of words or phrases.

[0060] Use natural language processing tools (such as Jieba segmentation, NLTK, SpaCy, etc.) to perform word segmentation on the label text and select an appropriate word segmentation granularity according to the needs of the business scenario. For label text in professional fields (such as medical terms), a customized word segmentation dictionary can be used to improve the accuracy of word segmentation.

[0061] After word segmentation processing, each word or phrase needs to be converted into a vector representation. The basic vector representation is the input of the text encoding module. Common vector representation methods include the Bag of Words model, word embedding models (Word2Vec, GloVe), or pre-trained language models (BERT, RoBERTa), etc. Vector representation can convert the semantic information of the label text into a numerical form, facilitating model understanding and processing.

[0062] Use a pre-trained word embedding model (such as Word2Vec, FastText, BERT, etc.) to map each word or phrase into a vector representation of a fixed dimension.

[0063] The basic vector representation is further processed by the text encoding module to generate a more accurate initial text embedding vector. This process usually includes operations such as context optimization of word vectors, semantic aggregation, and removal of redundant information, enabling the text embedding vector to better capture the semantic information of the label text.

[0064] By calling the inference interface of the text encoding module, the basic vector representation is input into the model to generate an initial text embedding vector. Common text encoding models include BERT, GPT, Transformer, etc. These models perform in-depth semantic understanding of the text through multiple layers of self-attention mechanisms, thereby generating high-quality text embedding vectors.

[0065] To generate the embedding vectors of the complete label text, it is necessary to semantically aggregate the initial text embedding vectors of each word or phrase. Semantic aggregation refers to combining the embedding vectors of multiple words into an overall vector, so that the semantic information of the label text can be represented by one embedding vector.

[0066] Use weighted average, pooling operations (such as max pooling, average pooling) or self-attention mechanism to aggregate the initial text embedding vectors of each label text to generate the overall embedding vector of each label text. This process usually relies on the built-in functions of the text encoding module and does not require manual writing of aggregation logic.

[0067] Example illustration: In the field of medical and health, a large number of label texts may be used in medical imaging systems, such as disease names, anatomical parts, diagnostic results, etc. By performing embedding processing on these label texts through the text encoding module, high-quality text embedding vectors can be generated, and these vectors can be matched with the visual embedding vectors of the image data to automatically generate an image report. For example, for a lung CT image, the system can match label texts such as "lung nodule" and "pleural effusion" according to the image features and generate a diagnostic description containing these medical terms.

[0068] In the application in the financial field, insurance companies may preset a large number of label texts, such as "type of vehicle damage", "insurance type for claims settlement", "accident description", etc. By performing embedding processing on these label texts through the text encoding module, the insurance claim settlement photos can be matched with the label texts to automatically identify the type of accident and generate corresponding claim settlement descriptions to provide auxiliary support for claim settlement reviewers.

[0069] By performing text embedding processing on multiple label texts through the text encoding module, the semantic information of the label texts can be converted into high-dimensional embedding vectors, providing basic data support for the subsequent similarity analysis between the image data and the label texts. It can capture the deep semantic associations in the label texts, improve the accuracy and robustness of the text embedding vectors, and thus enhance the semantic alignment ability of image description generation.

[0070] S30, analyze the similarity scores between the image data and each label text based on the visual embedding and the embedding vectors of the label texts;

[0071] In this embodiment, before the similarity analysis, the system first obtains the visual embedding of the image data from the visual encoding module and the embedding vectors of the label texts from the text encoding module. The visual embedding represents the feature information of the image, such as edges, textures, and structures in the image; the embedding vectors of the label texts represent the semantic information of the texts.

[0072] The acquisition of visual embeddings and label text embeddings is the basis for multimodal data fusion, ensuring that images and text can be compared in the same semantic space during subsequent analysis.

[0073] The system converts image data into a high-dimensional vector representation through a visual encoding module, and the text encoding module converts each label text into a corresponding text vector. The two embedding vectors are stored in the system's memory or temporary storage module for use by the similarity analysis module.

[0074] Normalization of the embedding vectors can eliminate differences in vector length, retaining only the direction information, which is convenient for subsequent similarity calculations. Normalization usually standardizes the length of the vector to a fixed value, so that when calculating similarity, the result will not be affected by differences in the numerical magnitudes of the vectors.

[0075] The system performs normalization on each embedding vector to ensure that the visual embedding and label text embedding vectors have the same length. The normalized embedding vectors are distributed in the unit space, making the similarity calculation more accurate and efficient.

[0076] To measure the similarity between visual embeddings and label text embeddings, the system usually uses the method of cosine similarity for analysis. Cosine similarity is a commonly used method for measuring vector similarity, which determines the similarity by comparing the directions of two vectors. The higher the similarity score, the stronger the semantic correlation between the image data and the label text.

[0077] The system determines the similarity by calculating the angle between the visual embedding and label text embedding vectors. The similarity score ranges from 0 to 1, and the closer it is to 1, the more similar the two vectors are. The system records the similarity score for each label text and stores it in the data processing module for subsequent use.

[0078] To facilitate the subsequent generation of heatmaps, the system needs to normalize the similarity scores for each label text, mapping all scores to a preset range, usually between 0 and 1. This normalization can ensure that the similarity scores of different label texts are comparable, making the heatmap more intuitive and accurate when generated.

[0079] The system uniformly adjusts the similarity scores of all label texts, mapping the lowest score to 0 and the highest score to 1. The normalized similarity scores can be directly used as the pixel values for generating heatmaps, ensuring that the generated heatmap can accurately reflect the significant regions in the image.

[0080] Example illustration: In medical image analysis, the system can quickly identify lesion areas in an image by calculating the similarity score between the visual embedding of the image and the embedding vectors of preset disease label texts. For example, when analyzing a lung CT image, the system identifies possible lesions in the image based on the similarity scores between the image and label texts such as "lung nodule", "pleural effusion", etc., and provides diagnostic assistance to doctors.

[0081] In the insurance claim settlement scenario, the system can quickly identify the accident type and damage situation in the photos uploaded by the customer by calculating the similarity score between the visual embedding of the accident scene photos and the embedding vectors of claim label texts. For example, when analyzing photos of a car accident scene, the system identifies the damaged parts of the vehicle based on the similarity scores between the photos and label texts such as "vehicle damage", "collision marks", etc., and provides the accident description and analysis results to the claim handlers.

[0082] Through the similarity analysis based on visual embedding and label text embedding vectors, the semantic correlation degree between the image data and the label text can be effectively measured, providing basic support for the subsequent generation of heatmaps. It can capture the deep semantic correlation between images and texts, improving the semantic alignment ability and the accuracy of image description generation in multimodal tasks.

[0083] S40, generate a heatmap channel for each label text according to the similarity score, and perform weighted processing on the heatmap channel to generate a weighted heatmap;

[0084] In this embodiment, the heatmap channel is a two-dimensional representation of the image, and each pixel value reflects the degree of association between different regions in the image and the label text. The similarity score is used to guide the generation of the heatmap channel, identifying which regions in the image have a higher semantic correlation degree with the label text.

[0085] The system combines the similarity score with the spatial distribution of the visual embedding to generate a heatmap channel for each label text pixel by pixel. Each heatmap channel specifically represents the significant regions in the image related to a certain label text.

[0086] The system maps the similarity score to the pixel positions of each feature map according to the feature map output by the visual coding module, and generates the corresponding heatmap channel for each label text. The system multiplies each pixel value of the feature map by the corresponding similarity score to obtain the preliminary pixel value of the heatmap.

[0087] The initially generated heatmap channels may have discontinuous pixel distributions. To enhance the coherence of the image, the heatmap channels need to be smoothed. The purpose of smoothing is to eliminate noise and make the heatmap more concentrated in the significant regions. The system can use smoothing methods such as Gaussian blur and mean filtering to adjust the pixel values of the heatmap channels, ensuring that the boundaries of the significant regions are more natural and the noise in the non-significant regions is effectively eliminated.

[0088] After completing the smoothing process, the system needs to perform a weighting process on the heatmap channels. The purpose of weighting is to amplify or reduce the pixel values of the heatmap according to the importance or confidence of different label texts. Weighting can ensure that the system pays more attention to the image regions corresponding to important label texts. The system weights the pixel values of the heatmap channels according to the weight or confidence value of the label text. The weight value can be set according to the range of similarity scores or obtained through the learning of the model. The higher the weight, the more obvious the pixel value of the heatmap channel, and the label text with a lower weight has less impact on the heatmap.

[0089] The weighted heatmap channels generated for each label text are separate. To obtain the overall significant region map of the image, multiple weighted heatmap channels need to be merged. Strategies such as maximum merging and weighted average merging can be adopted in the merging process.

[0090] The system superimposes the weighted heatmap channels of each label text according to the pixel values, and finally generates a weighted heatmap containing all significant regions. The merged weighted heatmap reflects the distribution of all significant regions in the image and can be used for subsequent significant region extraction and description generation tasks.

[0091] Example illustration: In medical image analysis, the system can generate a weighted heatmap based on the similarity score between the CT image and the label text, thereby identifying the lesion regions in the image. For example, when the system analyzes a lung CT image, it can generate heatmap channels for label texts such as "lung nodule" and "pleural effusion", and perform weighting on these heatmaps to generate a weighted heatmap reflecting the lesion distribution. This weighted heatmap can help doctors quickly identify the abnormal regions in the image and improve the diagnosis efficiency.

[0092] In the insurance claim settlement scenario, the system can generate a heatmap based on the similarity score between the accident scene photo and the claim settlement label text, thereby identifying the important regions in the photo. For example, when the system analyzes a car accident scene photo, it can generate heatmap channels for label texts such as "front-end damage" and "collision marks", and perform weighting on these heatmaps to generate a weighted heatmap reflecting the accident scene situation. This weighted heatmap can provide intuitive accident descriptions and analysis bases for claim adjusters and speed up the claim settlement process.

[0093] By generating a heatmap channel based on the similarity score and performing smoothing and weighting on the heatmap channel, the significant regions in the image can be effectively identified, and a weighted heatmap reflecting the multi-label scenario can be generated. It is possible to dynamically adjust the recognition range of the significant regions according to the similarity of the label text, improving the accuracy of significant region extraction and the precision of image description.

[0094] S50, determining the significant regions in the image data based on the weighted heatmap, and determining the labels and spatial position information corresponding to each significant region;

[0095] In this embodiment, the weighted heatmap is a two-dimensional distribution map of the significant regions in the image, and each pixel value represents the degree of significance at that position. The system locates the significant regions in the image data by performing threshold filtering and region segmentation on the pixel values of the weighted heatmap. The significant regions refer to the regions in the heatmap where the pixel values exceed a preset threshold, and these regions usually contain the key objects or scenes in the image.

[0096] The system performs a pixel-by-pixel scan of the pixel values of the weighted heatmap, and marks the regions with pixel values higher than the preset threshold as significant regions. The preset threshold can be dynamically adjusted according to the needs of the business scenario. For example, in the field of medical and health, the system can set a relatively high threshold to ensure that only the highly significant pixels in the lesion region are recognized.

[0097] To accurately describe the position of the significant regions, the system needs to extract the bounding boxes of each significant region. The bounding box is a rectangular box used to describe the circumscribed range of the significant region, usually represented by four coordinate values (the upper left corner coordinates and the lower right corner coordinates).

[0098] The system uses a boundary detection algorithm (such as contour detection or connected component analysis) to extract the bounding boxes of the significant regions in the weighted heatmap. The extracted bounding box information includes the starting coordinates, width, and height of each significant region, facilitating subsequent calculation of spatial position information and label matching.

[0099] To further clarify the position of the significant regions, the system needs to calculate the center point coordinates of each significant region. The center point coordinates refer to the geometric center position of the bounding box, which can more intuitively represent the spatial position of the significant region in the image.

[0100] The system calculates the center point coordinates based on the coordinate values of the significant region bounding box. The calculation method of the center point coordinates is to take the average of the upper left corner coordinates and the lower right corner coordinates of the bounding box as the center point position. The calculated center point coordinates will be used to generate the spatial position information of the significant regions.

[0101] The system needs to match each significant region with the corresponding label text to generate a detailed description of the significant region. The matching process is based on the similarity score and the source information of the heatmap channels, ensuring that the label of each significant region is consistent with the actual object in the image.

[0102] The system determines the label corresponding to each significant region according to the source information of the label text when generating the weighted heatmap. The selection of the label text can be done by the method of maximum similarity matching, that is, selecting the label text with the highest similarity score to the significant region as the label of the significant region.

[0103] The spatial position information of the significant region usually needs to be converted into the form of natural language description for generating a structured image description. For example, the center point coordinates can be converted into spatial position information such as "upper left corner", "center", "lower right corner", etc.

[0104] The system maps the coordinate information to the preset spatial position labels according to the center point coordinates of the significant region and the size ratio of the image. The definition of the spatial position labels can be based on the grid division of the image. For example, the image is divided into a nine-grid, corresponding to nine position descriptions: upper left corner, top, middle, lower right corner, etc.

[0105] Example illustration: In medical image analysis, the system can identify the lesion regions in the image through the weighted heatmap and determine the spatial position information of each lesion region. For example, when analyzing a lung CT image, the system can identify the "lung nodule" region in the image, calculate its center point coordinates, and describe the position as "upper left corner", so as to help doctors quickly locate the lesion region and improve the diagnosis efficiency.

[0106] In the insurance claim settlement scenario, the system can identify the key regions in the accident scene photos through the weighted heatmap and determine the spatial position information of each region. For example, when the system analyzes a car accident scene photo, it can identify the "front-end damage" region, calculate its center point coordinates, and describe the position as "lower right corner", so as to provide intuitive accident analysis information for the claim handlers and improve the accuracy and efficiency of claim settlement processing.

[0107] By determining the significant regions based on the weighted heatmap and matching the label and spatial position information of each significant region, the system can effectively identify the key objects or scenes in the image and generate a structured description of the significant regions. The ability to dynamically adjust the threshold, combined with the similarity score and the embedding information of the label text, improves the accuracy of significant region recognition and the precision of the description.

[0108] S60, generate an image description of the image data according to the label and spatial position information corresponding to each significant region.

[0109] In this embodiment, in order to generate a structured image description, the system needs to determine the corresponding text template based on the labels and spatial location information of each salient region. The text template is the basic framework for describing the salient region and usually includes location descriptors and object descriptors. For example, "There is a [label] at [location]" or "[Label] at [location]".

[0110] The system pre-defines a set of text templates and selects the appropriate template for different scenarios and salient region types. For example, for medical image analysis, "There is a lung nodule in the upper left corner" can be used; for accident photo analysis, "Damage is found at the front of the vehicle" can be used. The system automatically selects the corresponding template according to the spatial location of the salient region.

[0111] After determining the text template, the system needs to dynamically fill in the labels and spatial location information of the salient region into the template to generate the text description segment for each salient region. The label describes the object type in the salient region, and the spatial location information describes the position of the object in the image.

[0112] The system fills the label information (such as "lung nodule" or "front vehicle damage") and the spatial location information (such as "upper left corner" or "front of the vehicle") into the preset text template to generate a complete text description segment. For example, if the label of the salient region recognized by the system is "cat" and the spatial location is "upper left corner", the text description segment is "There is a cat in the upper left corner".

[0113] In order to generate a coherent image description text, the system needs to arrange all the text description segments according to the spatial location of the salient regions in the image. The arrangement order can be from left to right, from top to bottom, or adjusted according to the importance of the salient regions.

[0114] The system sorts the center point coordinates of the salient regions and arranges the text description segments according to the preset arrangement rules. For example, generating the image description text in the order from top to bottom and from left to right makes the description content more in line with the human reading habit.

[0115] After arranging the text description segments, the system combines these segments into a complete image description text. The complete image description text not only contains the labels and location information of each salient region, but also can present the overall information in the image in a coherent natural language form.

[0116] The system concatenates the text description segments in order into a coherent description text. For example, "There is a cat in the upper left corner, there is a sofa in the center, and there is a dog playing in the lower right corner." The system can also format the text description according to specific requirements, such as adjusting the sentence pattern, adding punctuation, etc.

[0117] The generated image description text needs to be output to the user interface or stored in the image description generation module for subsequent use. The output image description text can be directly displayed on the user interface or stored as structured data in a database for other systems to call.

[0118] The system outputs the generated image description text to the user interface or stores it in the database to ensure the persistence and accessibility of the text description. At the same time, the system can provide an interface to return the image description text as the result of the data interface for other business systems to call.

[0119] Example illustration: In medical image analysis, the system can generate an image description based on the lesion area in the image. For example, for a chest CT image, the system can generate the following image description: "There is a pulmonary nodule in the upper left corner and a small amount of pleural effusion in the lower right corner." Such an image description text can help doctors quickly understand the core information of the image and improve the diagnosis efficiency.

[0120] In the insurance claim settlement scenario, the system can generate an image description based on the prominent areas in the accident scene photos. For example, for a photo of a car accident scene, the system can generate the following image description: "Obvious damage is found at the front of the vehicle, and there are slight scratches on the right side of the vehicle body." Such an image description text can provide intuitive accident analysis information for the claim handlers and improve the accuracy and efficiency of claim settlement review.

[0121] By generating the image description text according to the label and spatial position information of each prominent area, the system can automatically generate semantically coherent and structured image description content. It can dynamically generate natural language descriptions combined with the actual content of the image, improving the efficiency and accuracy of image description generation, and is widely applicable to intelligent graphic and text analysis scenarios in the medical and health and financial fields.

[0122] The present invention relates to the technical fields of artificial intelligence and medical and health, and discloses an image description generation method, including: obtaining image data, generating embedding vectors of visual embeddings and label texts, analyzing the similarity scores between the image data and each label text, generating a weighted heat map according to the similarity scores, determining prominent areas and their corresponding labels and spatial position information based on the weighted heat map, and finally generating an image description. Through multi-modal feature matching and weighted heat map generation, the present invention accurately identifies prominent areas and multi-label scenarios in the image, generates semantically coherent image description texts, improves the accuracy, detail capture ability, and text readability of image descriptions, and is applicable to intelligent graphic and text analysis and auxiliary decision-making in the medical and health and financial fields.

[0123] In one embodiment, in the above S10, inputting the image data into a visual coding module to generate the visual embedding of the image data includes:

[0124] S101, perform downsampling on the image data and normalize the image data after downsampling;

[0125] S102, divide the normalized image data into multiple image patches;

[0126] S103, input each image patch into the visual encoding module, perform convolution operations through the visual encoding module to extract the local features of each image patch, and generate intermediate visual feature embeddings based on the local features of each image patch;

[0127] S104, aggregate the intermediate visual feature embeddings of all image patches to generate the complete visual embedding of the image data.

[0128] In this embodiment, downsampling is to reduce the data volume by reducing the resolution of the image, thereby improving the calculation efficiency, which is applicable when processing large-size images or high-resolution images. Normalization is to adjust the pixel values of the image to normalize them to a specific range (usually between 0 and 1), thereby reducing the numerical differences between different image data and facilitating model processing.

[0129] Downsampling can be implemented through image scaling algorithms, such as bilinear interpolation or nearest neighbor interpolation. Normalization is achieved by dividing the pixel values of the image by 255 to map the pixel values to the range [0, 1], ensuring the processing stability and consistency of the image data in the visual encoding module.

[0130] Image patch division is to split the image data into several small regions (image patches), and each image patch is independently input into the visual encoding module for feature extraction. This block processing method can effectively improve the parallelism of image processing while retaining the local feature information of the image.

[0131] The system can perform block processing on the image data according to fixed grid division rules. For example, a 224x224 pixel image can be divided into 16 56x56 pixel image patches. The size and number of blocks can be adjusted according to the model structure of the visual encoding module to ensure that each image patch can effectively cover the important regions in the image.

[0132] Each image patch is separately input into the visual encoding module, and the visual encoding module extracts the local features of the image patch through convolution operations. Convolution operations can identify local feature information such as edges, textures, and shapes in the image patch, thereby generating intermediate visual feature embeddings.

[0133] The visual encoding module typically uses pre-trained convolutional neural networks (such as ResNet, EfficientNet, etc.) to extract the features of image patches. The input of each image patch will go through multiple convolutional operations to extract feature information at different levels and generate an intermediate visual feature embedding.

[0134] Aggregate the intermediate visual feature embeddings of all image patches to generate a complete visual embedding of the entire image data. The purpose of the aggregation process is to integrate local feature information into global feature information to form a semantic representation of the whole image.

[0135] Multiple methods can be used for the aggregation process, including global average pooling, max pooling, attention mechanisms, etc. The system can select an appropriate aggregation method according to the specific architecture of the visual encoding module to ensure that the finally generated visual embedding can accurately represent the global feature information of the image.

[0136] In this embodiment, by inputting the image data into the visual encoding module and generating a complete visual embedding, the system can effectively extract the global and local feature information of the image, providing basic data support for subsequent multi-modal analysis and description generation.

[0137] In one embodiment, the above S20 includes:

[0138] S201, obtain multiple preset label texts, perform word segmentation on the preset label texts, and extract words or phrases of each preset label text;

[0139] S202, map the words or phrases of each preset label text to a basic vector representation;

[0140] S203, input the basic vector representation into the text encoding module, and perform semantic optimization processing on the basic vector representation through the text encoding module to generate an initial text embedding vector for each word or phrase;

[0141] S204, perform semantic aggregation on the initial text embedding vectors of the words or phrases of each preset label text to generate an overall embedding vector for each label text.

[0142] In this embodiment, the label text is usually a set of words or phrases predefined by the system, covering categories or objects that may be used in image descriptions. These label texts need to be word-segmented to split the complete phrases into words or phrases for subsequent vectorization and semantic analysis. Word segmentation is the first step in text embedding, especially in Chinese text processing, and the accuracy of the word segmentation algorithm has an important impact on the embedding result.

[0143] The system calls natural language processing tools (such as Jieba segmentation, NLTK, etc.) to perform word segmentation on the label text. For example, the label "lung nodule" will be decomposed into two phrases: "lung" and "nodule". The granularity of word segmentation can be adjusted according to the actual scenario requirements to ensure the semantic integrity of words or phrases.

[0144] The basic vector representation is the input of the text encoding module, usually generated by pre-trained word embedding models (such as Word2Vec, GloVe, etc.). Each word or phrase is mapped to a high-dimensional vector, and each dimension of the vector represents a certain semantic feature of the word or phrase.

[0145] The system loads the pre-trained word embedding model and maps each word or phrase to a fixed-length vector. For example, the word "cat" may be mapped to a 300-dimensional vector, representing different semantic features of the word. The generation of the basic vector representation can be achieved by calling the embedding layer in a machine learning framework (such as TensorFlow or PyTorch).

[0146] The basic vector representation undergoes semantic optimization processing by the text encoding module to generate an initial text embedding vector with more context semantics. The text encoding module usually adopts pre-trained language models (such as BERT, RoBERTa, etc.), which can capture the deep semantic information of words or phrases through multiple layers of neural networks.

[0147] The system inputs the basic vector representation into the text encoding module. Through feature extraction by multiple layers of networks, it generates an initial text embedding vector for each word or phrase. The text encoding module uses the self-attention mechanism to optimize the semantics of the input text, thereby generating a more accurate text embedding vector.

[0148] Each label text usually contains multiple words or phrases. Therefore, it is necessary to aggregate the initial text embedding vectors of words or phrases to generate an overall embedding vector for the label text. The semantic aggregation process can be achieved through weighted average, max pooling, or the self-attention mechanism.

[0149] The system aggregates the initial text embedding vectors of all words or phrases in the same label text. For example, for the label "lung nodule", the system will perform weighted average or max pooling on the embedding vectors of "lung" and "nodule" to generate an overall embedding vector for "lung nodule". The semantic aggregation strategy can be selected according to actual needs, such as assigning higher weights to key words.

[0150] In this embodiment, by performing text embedding processing on multiple label texts, the system can convert the semantic information of the label texts into high-dimensional vector representations, providing basic support for subsequent similarity analysis between images and texts. It can capture the deep semantic associations in the label texts, improving the accuracy and robustness of the text embedding vectors, thereby enhancing the precision and semantic alignment ability of image description generation.

[0151] In one embodiment, the above S30 includes:

[0152] S301, performing vector normalization processing on the visual embedding and the embedding vectors of each label text;

[0153] S302, calculating the cosine similarity between the visually embedded and the embedding vectors of each label text after vector normalization processing to generate the similarity scores between the image data and each label text;

[0154] S303, performing normalization processing on the similarity scores between the image data and each label text to map all similarity scores to a preset range.

[0155] In this embodiment, vector normalization processing is to eliminate the length difference between the visual embedding and the label text embedding vectors, making the similarity calculation more accurate and consistent. After normalization processing, the lengths of all vectors are standardized to the same value, thus ensuring that the calculation result of the similarity only depends on the direction difference between the vectors and is not affected by the vector length.

[0156] The system performs normalization processing on the visual embedding and the embedding vectors of the label texts one by one, adjusting the length of each embedding vector to 1. The normalized vectors are distributed in the unit space and can more accurately reflect the direction similarity between the embedding vectors. The method to achieve normalization is usually to divide each component of the vector by the norm of the vector, ensuring that the length of the normalized vector is always 1.

[0157] Cosine similarity is a commonly used similarity measurement method that measures the similarity between two vectors by comparing the angle between them. For the visual embedding and the label text embedding vectors, cosine similarity can reflect the semantic correlation degree between the image data and each label text.

[0158] The system calculates the cosine similarity between the normalized visual embedding and the embedding vectors of each label text one by one. If the cosine similarity is close to 1, it indicates a high semantic correlation degree between the image data and the label text. If the cosine similarity is close to 0, it indicates that there is almost no correlation between the image data and the label text. The system records the similarity scores of each label text as the basic data for subsequent heat map generation.

[0159] To facilitate subsequent heatmap generation, the system needs to normalize the similarity scores. Normalization can map all similarity scores to a fixed range (usually between 0 and 1), ensuring the comparability of similarity scores for different label texts, thus making the generated heatmap more intuitive and accurate.

[0160] The system linearly normalizes the similarity scores of each label text, mapping the lowest similarity score to 0 and the highest similarity score to 1. The normalized similarity scores can be directly used as the pixel values of the heatmap.

[0161] In this embodiment, through the similarity analysis based on visual embeddings and label text embedding vectors, the system can effectively measure the semantic correlation between image data and label texts, providing basic support for subsequent heatmap generation; it can capture the deep semantic correlation between images and texts, improving the semantic alignment ability and the accuracy of image description generation in multimodal tasks.

[0162] In one embodiment, the above S40 includes:

[0163] S401, based on the similarity scores between the image data and each label text, generate a heatmap channel for each label text, and each pixel value of the heatmap channel is assigned according to the similarity score and the visual embedding distribution;

[0164] S402, perform smoothing processing and weighting processing on the pixel values of the heatmap channel to generate a weighted heatmap channel for each label text;

[0165] S403, merge the weighted heatmap channels of multiple label texts to generate the final weighted heatmap of the image data.

[0166] In this embodiment, the heatmap channel is a two-dimensional matrix with the same size as the original image data, and each pixel value represents the association strength between that position and the label text. According to the similarity scores and the spatial distribution of visual embeddings, a separate heatmap channel is generated for each label text, enabling the system to intuitively identify the significant regions in the image related to specific label texts.

[0167] From the embedded feature map extracted by the visual coding module, the system assigns values to each pixel point according to the similarity score of each label text. Regions with high similarity scores will be assigned higher pixel values, indicating that these regions are more in line with the semantics of the label text. Regions with low similarity scores are assigned lower pixel values or retained as background values. The system traverses the visual embedding feature map pixel by pixel, multiplying the value of each pixel by the similarity score of the corresponding label text, thereby generating the heatmap channel.

[0168] The initially generated heatmap channels may have noise or discontinuous pixel regions. Therefore, it is necessary to smooth the heatmap channels to ensure that the boundaries of significant regions are more natural. At the same time, the weighting process is to adjust the pixel values according to the importance or confidence of each label text, making the intensity of significant regions more prominent.

[0169] Smoothing process: The system applies a smoothing algorithm (such as Gaussian blur or mean filter) to the pixel values of the heatmap channels to eliminate discrete noise and make the significant regions more coherent.

[0170] Weighting process: The system adjusts the pixel values of the heatmap channels according to the confidence scores or business weights of the label texts. The pixel values of the heatmap corresponding to label texts with higher confidence will be amplified, while the pixel values of the heatmap of label texts with lower confidence will be reduced.

[0171] The weighted heatmap channels generated by each label text are independent. To obtain the complete image significant region information, it is necessary to merge the weighted heatmap channels of multiple label texts. Multiple strategies can be adopted in the merging process, such as maximum merging, weighted average merging, etc.

[0172] Take the maximum value of the values at each pixel position in all weighted heatmap channels to ensure that the final heatmap can highlight the most significant region at each pixel position.

[0173] Perform weighted average processing on the values at each pixel position, combine the heatmap information of different label texts, and make the final heatmap more comprehensively reflect the distribution of significant regions in the image.

[0174] The final weighted heatmap after merging is a visual representation of the significant regions in the image data and can be used for subsequent significant region extraction and description generation tasks.

[0175] In this embodiment, by generating heatmap channels according to the similarity scores and performing smoothing and weighting processing on the heatmap channels, it is possible to effectively identify the significant regions in the image and generate a weighted heatmap reflecting the multi-label scenario. It can dynamically adjust the recognition range of significant regions according to the similarity of label texts, improving the accuracy of significant region extraction and the precision of image description.

[0176] In one embodiment, the above S50 includes:

[0177] S501, based on the weighted heatmap, determine significant regions in the image data where the significance scores exceed a preset threshold, and extract the image features of each significant region;

[0178] S502, input the image features of each significant region into the visual encoding module to generate the feature embeddings of each significant region;

[0179] S503. Match the feature embeddings of each salient region with the embedding vector of the label text to determine the label corresponding to each salient region.

[0180] S504. Determine the center point coordinates of each salient region according to the bounding box coordinates of each salient region, and convert the center point coordinates of each salient region into corresponding spatial position information.

[0181] In this embodiment, the saliency score is the value of each pixel in the weighted heat map, indicating the saliency degree of different regions in the image. The system identifies the regions where the saliency score exceeds the threshold from the heat map by setting a preset threshold. These regions usually contain key objects or scenes in the image. The image features of the salient regions are the feature information extracted from the pixel distribution of these regions and are used for subsequent feature embedding and label matching.

[0182] The system traverses the weighted heat map pixel by pixel, and marks the pixel regions with saliency scores higher than the preset threshold as salient regions. Feature extraction is performed on the pixel data of each salient region to generate an image feature vector, which describes the local features of the salient region. Feature extraction can adopt methods such as average pooling, max pooling or convolution operations to convert the image features of the salient region into a fixed-length vector representation.

[0183] The visual encoding module is a pre-trained deep learning model that can encode the input image features to generate high-dimensional feature embeddings. The feature embeddings of the salient regions represent the semantic feature information of these regions and are used to match with the embedding vectors of the label text.

[0184] The system inputs the image features of each salient region into the visual encoding module. The visual encoding module performs multi-layer feature extraction and encoding processing on the input image features through a convolutional neural network or a Transformer model to generate a feature embedding vector. The feature embedding vector contains the semantic information of the salient region and can be used for matching and comparison with the text embedding vector.

[0185] The system needs to perform similarity matching between the feature embeddings of the salient regions and the embedding vectors of the label text to determine the semantic labels of each salient region. The embedding vectors of the label text are pre-generated by the text encoding module and are in the same vector space as the feature embedding vectors of the salient regions.

[0186] The system calculates the cosine similarity between the feature embedding vector of the salient region and the embedding vector of each label text to measure the semantic correlation degree between the two. For each salient region, the system selects the label text with the highest similarity score as the label of the salient region. The matching label text can be phrases such as "cat", "front-end damage" that describe the objects or scenes in the image.

[0187] To describe the position of the salient regions in the image, the system needs to calculate the center point coordinates of each salient region and convert the center point coordinates into spatial position information in natural language (such as "upper left corner", "center", "lower right corner", etc.).

[0188] The system obtains the bounding box coordinates of each salient region through a bounding box detection algorithm. The coordinates of the bounding box are usually represented by the pixel positions of the upper left corner and the lower right corner. The center point coordinates are determined by calculating the geometric center of the bounding box. The specific formula is the average of the horizontal and vertical coordinates of the bounding box. The system maps the center point coordinates to predefined spatial position labels. For example, the image is divided into a nine-grid region, and the center point coordinates are mapped to position descriptions such as "upper left corner", "center", "lower right corner", etc.

[0189] In this embodiment, by determining the salient regions based on the weighted heat map and matching the labels of each salient region with the spatial position information, the system can effectively identify the key objects or scenes in the image and generate a structured description of the salient regions; it can dynamically adjust the threshold, combine the similarity score and the embedded information of the label text, and improve the accuracy of salient region recognition and the precision of the description.

[0190] In one embodiment, the above S60 includes:

[0191] S601, determining a preset text template according to the label and spatial position information corresponding to each salient region;

[0192] S602, adding the label and spatial position information of each salient region to the text template to generate a corresponding text description segment;

[0193] S603, arranging the text description segments of all salient regions according to the spatial positions of the salient regions in the image data to generate a complete image description text.

[0194] In this embodiment, the text template is a predefined language structure used to organize the labels and spatial position information of the salient regions into a natural language description form. For example, "There is a [label] at [position]" or "[label] at [position]", etc. The system selects a suitable text template according to the label and spatial position information of each salient region.

[0195] The system determines the text template to be used based on the spatial position of the salient region. For example, for a salient region in the upper left corner, the template "There is a [label] in the upper left corner" can be used, and for a salient region in the central position, the template "[label] in the center" can be used. The type of object described by the label also affects the choice of template. For example, for dynamic objects (such as people, animals), the template may include action descriptions; for static objects (such as tables, chairs), the template tends to focus on position descriptions.

[0196] The system fills the labels (such as "cat", "front-end damage") and spatial position information (such as "upper left corner", "lower right corner") of the salient regions into the selected text template to generate text description segments for each salient region. These description segments are complete semantic units and can represent the information of a salient region individually.

[0197] The system fills the label of each salient region as the object name in the text template and the spatial position information as the position descriptor. For example, for a salient region with a label of "cat" and a spatial position of "upper left corner", the text description segment generated by the system is "There is a cat in the upper left corner".

[0198] To generate a coherent image description text, the system needs to arrange all the text description segments according to the spatial positions of the salient regions. The arrangement order can be adjusted from left to right, top to bottom according to the distribution of the salient regions in the image, or according to the importance level of the salient regions.

[0199] Arrangement based on spatial position: The system arranges the text description segments of the salient regions in the order from left to right and top to bottom. For example, the description segment of the salient region in the upper left corner will be ranked first, and the description segment of the salient region in the lower right corner will be ranked last.

[0200] Arrangement based on importance level: The system performs a weighted sorting on the text description segments according to the similarity scores of the salient regions or the importance level of the labels, and arranges the description segments of the more important salient regions in the front preferentially.

[0201] Finally, the system combines all the text description segments into a complete image description text. For example, "There is a cat in the upper left corner, a table in the central position, and a vase in the lower right corner".

[0202] In this embodiment, by generating an image description text based on the labels and spatial position information of the salient regions, the system can automatically convert visual information into natural language descriptions, improving the efficiency and accuracy of image description generation.

[0203] In one embodiment, an image description generation device is provided, and this image description generation device corresponds one-to-one with the image description generation method in the above embodiment. Refer to Figure 3 ,Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the image description generation device of the present invention. A visual encoding module 10, a text encoding module 20, a similarity analysis module 30, a heatmap generation module 40, a salient region recognition module 50, and a description generation module 60. The detailed description of each functional module is as follows:

[0204] The visual encoding module 10 is used to obtain image data, input the image data into the visual encoding module, and generate a visual embedding of the image data;

[0205] The text encoding module 20 is used to perform text embedding processing on multiple tag texts through the text encoding module to obtain embedding vectors of the tag texts;

[0206] The similarity analysis module 30 is used to analyze the similarity scores between the image data and each tag text based on the visual embedding and the embedding vectors of the tag texts;

[0207] The heatmap generation module 40 is used to generate a heatmap channel for each tag text according to the similarity scores, and perform weighted processing on the heatmap channel to generate a weighted heatmap;

[0208] The salient region recognition module 50 is used to determine the salient regions in the image data based on the weighted heatmap, and determine the tags and spatial position information corresponding to each salient region;

[0209] The description generation module 60 is used to generate an image description of the image data according to the tags and spatial position information corresponding to each salient region.

[0210] In one embodiment, the visual encoding module 10 is specifically used for:

[0211] Perform downsampling processing on the image data, and perform normalization processing on the downsampled image data;

[0212] Divide the normalized image data into multiple image blocks;

[0213] Input each image block into the visual encoding module, perform convolution operations through the visual encoding module to extract local features of each image block, and generate intermediate visual feature embeddings based on the local features of each image block;

[0214] Perform aggregation processing on the intermediate visual feature embeddings of all image blocks to generate a complete visual embedding of the image data.

[0215] In one embodiment, the text encoding module 20 is specifically used for:

[0216] Obtain multiple preset label texts, perform word segmentation on the preset label texts, and extract words or phrases of each preset label text;

[0217] Map the words or phrases of each preset label text to a basic vector representation;

[0218] Input the basic vector representation into a text encoding module, and perform semantic optimization processing on the basic vector representation through the text encoding module to generate an initial text embedding vector for each word or phrase;

[0219] Perform semantic aggregation on the initial text embedding vectors of the words or phrases of each preset label text to generate an overall embedding vector for each label text.

[0220] In one embodiment, the similarity analysis module 30 is specifically configured to:

[0221] Perform vector normalization processing on the visual embedding and the embedding vectors of each label text;

[0222] Calculate the cosine similarity between the visually embedded and the embedding vectors of each label text after vector normalization processing to generate a similarity score between the image data and each label text;

[0223] Perform normalization processing on the similarity scores between the image data and each label text to map all similarity scores to a preset range.

[0224] In one embodiment, the heat map generation module 40 is specifically configured to:

[0225] Based on the similarity scores between the image data and each label text, generate a heat map channel for each label text, and assign each pixel value of the heat map channel according to the similarity scores and the visual embedding distribution;

[0226] Perform smoothing processing and weighting processing on the pixel values of the heat map channel to generate a weighted heat map channel for each label text;

[0227] Merge the weighted heat map channels of multiple label texts to generate a final weighted heat map of the image data.

[0228] In one embodiment, the significant region recognition module 50 is specifically configured to:

[0229] Based on the weighted heat map, determine significant regions in the image data where the significance score exceeds a preset threshold, and extract image features of each significant region;

[0230] Input the image features of each significant region into a visual encoding module to generate a feature embedding for each significant region;

[0231] Match the feature embedding of each salient region with the embedding vector of the label text to determine the label corresponding to each salient region;

[0232] According to the bounding box coordinates of each salient region, determine the center point coordinates of each salient region, and convert the center point coordinates of each salient region into corresponding spatial position information.

[0233] In one embodiment, the description generation module 60 is specifically configured to:

[0234] Determine a preset text template according to the label and spatial position information corresponding to each salient region;

[0235] Add the label and spatial position information of each salient region to the text template to generate a corresponding text description segment;

[0236] Arrange the text description segments of all salient regions according to the spatial positions of the salient regions in the image data to generate a complete image description text.

[0237] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of an image description generation method.

[0238] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as Figure 5 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of an image description generation method

[0239] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0240] Obtain image data, input the image data into a visual encoding module, and generate a visual embedding of the image data;

[0241] Perform text embedding processing on multiple label texts through a text encoding module to obtain embedding vectors of the label texts;

[0242] Analyze the similarity scores between the image data and each label text based on the visual embedding and the embedding vectors of the label texts;

[0243] According to the similarity scores, generate heatmap channels for each label text, and perform weighting processing on the heatmap channels to generate a weighted heatmap;

[0244] Determine the significant regions in the image data based on the weighted heatmap, and determine the labels and spatial location information corresponding to each significant region;

[0245] Generate an image description of the image data according to the labels and spatial location information corresponding to each significant region.

[0246] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0247] Obtain image data, input the image data into a visual encoding module, and generate a visual embedding of the image data;

[0248] Perform text embedding processing on multiple label texts through a text encoding module to obtain embedding vectors of the label texts;

[0249] Analyze the similarity scores between the image data and each label text based on the visual embedding and the embedding vectors of the label texts;

[0250] According to the similarity scores, generate heatmap channels for each label text, and perform weighting processing on the heatmap channels to generate a weighted heatmap;

[0251] Determine the significant regions in the image data based on the weighted heatmap, and determine the labels and spatial location information corresponding to each significant region;

[0252] Generate an image description of the image data according to the labels and spatial location information corresponding to each significant region.

[0253] It should be noted that for the functions or steps that can be achieved by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.

[0254] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0255] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0256] It should be noted that if there are software tools or components of other companies in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A method for generating an image description, characterized in that: The following steps are involved: Acquire image data, input the image data into a visual encoding module, and generate a visual embedding of the image data; Perform text embedding processing on multiple label texts through the text encoding module to obtain the embedding vector of the label text; Analyzing a similarity score between the image data and each label text based on the visual embedding and the embedding vector of the label text; According to the similarity score, a heat map channel is generated for each label text, and the heat map channel is weighted to generate a weighted heat map; Determine the salient regions in the image data based on the weighted heat map, and determine the label and spatial position information corresponding to each salient region; An image description of the image data is generated according to the label and spatial position information corresponding to each salient area.

2. The image description generating method as claimed in claim 1, characterized in that: Inputting the image data into a visual encoding module to generate a visual embedding of the image data includes: Performing down-sampling processing on the image data, and performing normalization processing on the down-sampled image data; Dividing the normalized image data into a plurality of image blocks; Input each image block into a visual encoding module, perform a convolution operation through the visual encoding module to extract local features of each image block, and generate an intermediate visual feature embedding based on the local features of each image block; The intermediate visual feature embeddings of all image patches are aggregated to generate a complete visual embedding of the image data.

3. The image description generating method according to claim 1, characterized in that: Through the text encoding module, multiple label texts are embedded to obtain the embedding vector of the label text, including: Acquire multiple preset label texts, perform word segmentation processing on the preset label texts, and extract words or phrases from each preset label text; Map each word or phrase of the preset label text to a basic vector representation; Inputting the basic vector representation into a text encoding module, performing semantic optimization processing on the basic vector representation through the text encoding module to generate an initial text embedding vector for each word or phrase; The initial text embedding vectors of words or phrases of each preset label text are semantically aggregated to generate the overall embedding vector of each label text.

4. The image description generating method according to claim 1, characterized in that: Analyzing a similarity score between the image data and each label text based on the visual embedding and the embedding vector of the label text, including: Performing vector normalization processing on the visual embedding and the embedding vector of each label text; Calculating cosine similarity between the normalized visual embedding and the embedding vector of each label text to generate a similarity score between the image data and each label text; The similarity scores between the image data and each label text are normalized to map all similarity scores into a preset range.

5. The image description generating method according to claim 1, characterized in that: According to the similarity score, a heat map channel is generated for each label text, and the heat map channel is weighted to generate a weighted heat map, including: Based on the similarity score between the image data and each label text, a heat map channel is generated for each label text, and each pixel value of the heat map channel is assigned a value according to the similarity score and the visual embedding distribution; Smoothing and weighting the pixel values ​​of the heat map channel to generate a weighted heat map channel for each label text; The weighted heat map channels of multiple label texts are merged to generate a final weighted heat map of the image data.

6. The image description generating method according to claim 1, characterized in that: Determining the salient regions in the image data based on the weighted heat map, and determining the labels and spatial position information corresponding to each salient region, including: Based on the weighted heat map, determining from the image data significant regions whose significance scores exceed a preset threshold, and extracting image features of each significant region; The image features of each salient region are input into the visual encoding module to generate feature embedding of each salient region; Matching the feature embedding of each salient region with the embedding vector of the label text to determine the label corresponding to each salient region; According to the bounding box coordinates of each salient region, the center point coordinates of each salient region are determined, and the center point coordinates of each salient region are converted into corresponding spatial position information.

7. The image description generating method according to claim 1, characterized in that: Generating an image description of the image data according to the label and spatial position information corresponding to each salient area includes: Determine a preset text template according to the label and spatial position information corresponding to each salient area; Adding the label and spatial position information of each salient area to the text template to generate a corresponding text description segment; According to the spatial position of the salient area in the image data, the text description fragments of all salient areas are arranged to generate a complete image description text.

8. An image description generating device, characterized in that: The image description generating device comprises: A visual encoding module, used to obtain image data, input the image data into the visual encoding module, and generate a visual embedding of the image data; A text encoding module is used to perform text embedding processing on multiple label texts through the text encoding module to obtain an embedding vector of the label text; A similarity analysis module for analyzing a similarity score between the image data and each label text based on the visual embedding and the embedding vector of the label text; A heat map generation module, used for generating a heat map channel for each label text according to the similarity score, and performing weighted processing on the heat map channel to generate a weighted heat map; A salient region identification module, used to determine the salient regions in the image data based on the weighted heat map, and to determine the label and spatial position information corresponding to each salient region; The description generation module is used to generate an image description of the image data according to the label and spatial position information corresponding to each salient area.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and an image description generation program stored in the memory and executable on the processor. When the image description generation program is executed by the processor, the steps of the image description generation method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The storage medium stores an image description generation program, and when the image description generation program is executed by the processor, the steps of the image description generation method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Data annotation generation method and device, storage medium and electronic equipment

    CN122023952A

  • Data labeling generation method and apparatus, storage medium, and electronic device

    CN122023952B

  • Image caption generation method, apparatus, device, and medium

    WO2026166211A1