Multi-label image recognition method and device fusing attention mechanism
By integrating attention mechanism in the multi-label image recognition model, extracting and fusing local features and label position features, the problems of low efficiency and insufficient accuracy in multi-label image classification are solved, and more efficient and accurate image classification is achieved.
Patent Information
- Application Number
- CN202510216514.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional deep learning models ignore image-label-related constraints when classifying multi-label images, resulting in inefficient classification efficiency and insufficient accuracy, especially on highly specialized and complex image sets.
A multi-label image recognition method with a fusion attention mechanism is proposed. By extracting local features and label position features, calculating attention values, fusion relationship features and local features, and inputting them into the multi-label image category recognition model, the model's attention to key image information is enhanced.
The classification efficiency and classification accuracy of multi-label images are improved, and the correlation information between different regions and labels in the image is effectively utilized, which improves the recognition accuracy and robustness of the model.
Smart Images

Figure CN120147714A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly relates to a multi-label image recognition method and device integrating an attention mechanism. Background Art
[0002] Large-scale multi-label image classification requires determining the presence or absence of target objects in a large number of sample images. For highly specialized and complex multi-label image sets, it is particularly important to ensure the accuracy of image classification. Traditional deep learning models usually do not consider image-label related constraints when classifying multi-label images, and the strategy of classifying only based on the characteristics of the images themselves greatly limits the performance of the models. For highly specialized and complex multi-label image sets, due to the huge number of samples and the uneven cognitive levels of personnel, this may lead to very low efficiency of multi-label image classification. Existing multi-label image classification methods focus on the accuracy of label prediction and ignore the structural information embedded in the hierarchical label space. Traditional methods use attention mechanisms or prior knowledge but lack deep semantic associations, resulting in a decline in detection performance. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems in the above technologies to some extent. For this purpose, the object of the present invention is to propose a multi-label image recognition method and device integrating an attention mechanism, which is convenient for improving the classification efficiency and accuracy of multi-label images.
[0004] To achieve the above object, an embodiment of the present invention proposes a multi-label image recognition method integrating an attention mechanism, including:
[0005] Obtain a multi-label image;
[0006] Preprocess the multi-label image to obtain a preprocessed image;
[0007] Extract the local features and label position features of the preprocessed image;
[0008] Determine query information according to the local features and label position features, calculate the relevance of keywords in the query information, normalize it through a softmax function to obtain weights, and then calculate the weighted sum to obtain an attention value, and determine the relationship features;
[0009] Integrate the attention mechanism into the multi-label image category recognition model, fuse the relationship features with the local features to determine the fusion features, and input them into the next-level network for information transmission until the output result is determined;
[0010] According to the output result and a preset classification threshold, determine multiple category labels to which the preprocessed image belongs.
[0011] According to some embodiments of the present invention, preprocessing a multi-label image to obtain a preprocessed image, including:
[0012] Input the multi-label image into a pre-trained segmentation model, and output the semantic segmentation result of the multi-label image;
[0013] Determine the R-channel value, G-channel value, and B-channel value of all pixel points on the image corresponding to each superpixel region in the semantic segmentation result; compare the R-channel value, G-channel value, and B-channel value of each pixel point to obtain a comparison result. When the comparison result shows that the R-channel value is the maximum, the corresponding pixel point is the first type of pixel point. When the comparison result shows that the G-channel value is the maximum, the corresponding pixel point is the second type of pixel point. When the comparison result shows that the B-channel value is the maximum, the corresponding pixel point is the third type of pixel point;
[0014] Mark the types of all pixel points on the image corresponding to the superpixel region, and divide the region set according to the marking result to obtain the preprocessed image.
[0015] According to some embodiments of the present invention, a method for obtaining a pre-trained segmentation model includes:
[0016] Obtain a sample image, perform line-level annotation on the feature region in the sample image to represent the position and shape of the annotation region; perform superpixel segmentation on the annotated sample image to divide the annotated sample image into multiple superpixel regions;
[0017] For each superpixel region, extract features from the line-level annotation, learn the image representation, and optimize the segmentation of the sample image through an image segmentation algorithm;
[0018] According to the consistency between the segmentation result and the line-level annotation, feedback and update the learning process, perform iterative training, and finally obtain a trained segmentation model.
[0019] According to some embodiments of the present invention, extracting local features and label position features of the preprocessed image includes:
[0020] Input the preprocessed image into a convolutional neural network for size adjustment and normalization processing;
[0021] Based on the convolutional layer in the convolutional neural network as a feature extractor, extract high-level feature maps from the preprocessed image;
[0022] Determine the target region on the feature map, and perform a pooling operation on the target region to obtain local features;
[0023] Apply position encoding on the feature map to determine the label position features.
[0024] According to some embodiments of the present invention, query information is determined based on local features and label position features, the relevance of keywords in the query information is calculated, the weights are obtained through normalization by the softmax function, and then the weighted sum is calculated to obtain the attention value, and the relationship feature is determined, including:
[0025] Fuse the local features and the label position features to determine the query information;
[0026] Use cosine similarity to query the relevance of each keyword in the information to obtain the relevance score;
[0027] Normalize the relevance score through the Softmax function to determine the weight of each keyword;
[0028] Use the normalized weights to perform weighted summation on the values of the keywords to obtain the attention value and determine the relationship feature.
[0029] According to some embodiments of the present invention, the attention mechanism is incorporated into the multi-label image category recognition model, and the calculated relationship feature is fused with the local image feature and input into the next-level network for information transmission, including:
[0030] Introduce the attention mechanism into the multi-label image category recognition model, use the spatial attention map to calculate the importance of different regions in the image, and mark them;
[0031] Fuse the relationship feature and the local feature in the marked map to determine the fusion feature, and input it into the next-level network for information transmission until the output result is determined; wherein, the next-level network includes a fully connected layer, a convolutional layer, and a recurrent neural network layer.
[0032] According to some embodiments of the present invention, according to the output result and the preset classification threshold, determine the multiple category labels to which the preprocessed image belongs, including:
[0033] Determine the preset classification threshold;
[0034] Compare the output result with the preset classification threshold, and determine the multiple category labels to which the preprocessed image belongs according to the comparison result.
[0035] According to some embodiments of the present invention, use the spatial attention map to calculate the importance of different regions in the image and mark them, including:
[0036] Based on the first-order finite difference approximation formula of 2×2, calculate the first importance value S of the corresponding region W(x, y) in the image in the X direction x and the second importance value S in the Y direction y ;
[0037]
[0038] According to the first important value S of the corresponding region W(x, y) in the X direction x and the second important value S in the Y direction y , determine the important amplitude;
[0039]
[0040] wherein, S is the important amplitude;
[0041] Based on the important amplitude, display the importance of the corresponding region and make a mark.
[0042] According to some embodiments of the present invention, divide the region set according to the marking result to obtain a preprocessed image, including:
[0043] Based on the marking result, delimit an initial region set, determine the pixel value of each pixel point in the initial region set, and organize it into a color histogram;
[0044] Count the number of pixel points with the pixel value of the preset type in the initial region set, and judge whether the initial region set is qualified according to the number of pixel points with the pixel value of the preset type and the color histogram;
[0045]
[0046] wherein, is the weight information of the pixel points with the pixel value of the preset type in the color histogram G ; F(w j , p) is the distance between the jth pixel point in the color histogram and the central pixel point p in the color histogram; F 0 is the maximum value of F(w j , p); w j is the jth pixel point in the color histogram; k(w j ) is the pixel value of the jth pixel point in the color histogram; is 's weight coefficient; M is the number of pixel points included in the color histogram; T is the number of pixel points with the pixel value of the preset type in the color histogram;
[0047] When it is determined that the weight information is greater than the preset threshold, it means that the initial region set is qualified, otherwise, it means that the initial region set is unqualified and needs to be adjusted, and finally a preprocessed image is obtained.
[0048] According to some embodiments of the present invention, an apparatus applying the multi-label image recognition method with the above-mentioned fusion attention mechanism includes:
[0049] An acquisition module for acquiring a multi-label image;
[0050] A preprocessing module for preprocessing a multi-label image to obtain a preprocessed image;
[0051] An extraction module for extracting local features and label position features of the preprocessed image;
[0052] A calculation module for determining query information based on local features and label position features, calculating the relevance of keywords in the query information, normalizing through a softmax function to obtain weights, and then calculating a weighted sum to obtain an attention value and determining relationship features;
[0053] A first determination module for integrating an attention mechanism into a multi-label image category recognition model, fusing relationship features with local features to determine fused features, and inputting them into the next-level network for information transmission until an output result is determined;
[0054] A second determination module for determining multiple category labels to which the preprocessed image belongs according to the output result and a preset classification threshold.
[0055] The present invention proposes a multi-label image recognition method and device integrating an attention mechanism, which facilitates improving the classification efficiency and accuracy of multi-label images by enhancing the model's attention to key information in the images.
[0056] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification and the drawings.
[0057] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0059] Figure 1 is a flowchart of a multi-label image recognition method integrating an attention mechanism according to an embodiment of the present invention;
[0060] Figure 2 is a block diagram of a multi-label image recognition device integrating an attention mechanism according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0062] As Figure 1 shown, an embodiment of the present invention proposes a multi-label image recognition method integrating an attention mechanism, including steps S1 - S6:
[0063] S1. Obtain a multi-label image;
[0064] S2. Preprocess the multi-label image to obtain a preprocessed image;
[0065] S3. Extract local features and label position features of the preprocessed image;
[0066] S4. Determine query information based on local features and label position features, calculate the relevance of keywords in the query information, normalize it through the softmax function to obtain weights, then calculate the weighted sum to obtain an attention value, and determine relationship features;
[0067] S5. Incorporate the attention mechanism into the multi-label image category recognition model, fuse the relationship features with local features to determine fused features, and input them into the next-level network for information transmission until the output result is determined;
[0068] S6. Determine multiple category labels to which the preprocessed image belongs according to the output result and a preset classification threshold.
[0069] The working principle of the above technical solution: The preprocessing step aims to improve the image quality to ensure the consistency and stability of image data. Local features of the image (such as edges, textures, colors, etc.) and label position features (i.e., the position information of different objects or labels in the image) are extracted. Based on the local features and label position features, the method constructs query information and calculates the relevance of keywords in the query information. These relevances are normalized through the softmax function to obtain weights, and then the weighted sum is calculated to determine the attention value. The attention value reflects the importance of different regions or features in the image for the final recognition result. In this step, relationship features are also determined, which describe the association relationships between different objects or labels in the image. The attention mechanism is combined with the multi-label image category recognition model. By fusing the relationship features with local features, fused features are generated and these features are input into the next-level network for information transmission. This process is repeated until the final output result is determined. The introduction of the attention mechanism enables the model to pay more attention to key information in the image, thereby improving the recognition accuracy. Multiple category labels to which the preprocessed image belongs are determined according to the output result of the model and a preset classification threshold. The classification threshold is usually set according to the specific application scenario and the characteristics of the dataset to ensure the accuracy and reliability of recognition.
[0070] Beneficial effects of the above technical solution: By enhancing the model's attention to key information in the image, it is convenient to improve the classification efficiency and accuracy of multi-label images.
[0071] According to some embodiments of the present invention, preprocessing a multi-label image to obtain a preprocessed image, including:
[0072] Inputting the multi-label image into a pre-trained segmentation model to output the semantic segmentation result of the multi-label image;
[0073] Determining the R-channel value, G-channel value, and B-channel value of all pixel points on the image corresponding to each superpixel region in the semantic segmentation result; comparing the R-channel value, G-channel value, and B-channel value of each pixel point to obtain a comparison result. When the comparison result shows that the R-channel value is the maximum value, the corresponding pixel point is a first-type pixel point. When the comparison result shows that the G-channel value is the maximum value, the corresponding pixel point is a second-type pixel point. When the comparison result shows that the B-channel value is the maximum value, the corresponding pixel point is a third-type pixel point;
[0074] Marking the types of all pixel points on the image corresponding to the superpixel region, and dividing the region set according to the marking result to obtain the preprocessed image.
[0075] Working principle of the above technical solution: Input the multi-label image into a pre-trained segmentation model. This model can identify different semantic regions in the image, such as objects, backgrounds, objects of different categories, etc., and output the semantic segmentation result. The semantic segmentation result is a segmentation map with the same size as the input image, where each pixel is assigned a label indicating which semantic region the pixel belongs to. For each superpixel region in the semantic segmentation result, determine the R-channel value, G-channel value, and B-channel value of all pixel points on the image corresponding to this region. Compare the three channel values of each pixel point to determine the type of this pixel point: If the R-channel value is the maximum value, mark this pixel point as a first-type pixel point (i.e., a red-dominated pixel point). If the G-channel value is the maximum value, mark this pixel point as a second-type pixel point (i.e., a green-dominated pixel point). If the B-channel value is the maximum value, mark this pixel point as a third-type pixel point (i.e., a blue-dominated pixel point). According to the results of the above color channel analysis, mark the types of all pixel points in each superpixel region. According to the marking result, divide the superpixel regions with similar or the same type of pixel points into different region sets. These region sets represent different color regions in the image or parts of objects with specific color characteristics.
[0076] Beneficial effects of the above technical solution: Based on color features for differentiation, dividing the region set to obtain the preprocessed image.
[0077] According to some embodiments of the present invention, a method for obtaining a pre-trained segmentation model includes:
[0078] Obtain a sample image, perform line-level annotation on the feature region in the sample image to represent the position and shape of the annotated region; perform superpixel segmentation on the annotated sample image to divide the annotated sample image into multiple superpixel regions;
[0079] For each superpixel region, extract features from the line-level annotation and learn an image representation, and optimize the segmentation of the sample image through an image segmentation algorithm;
[0080] According to the consistency between the segmentation result and the line-level annotation, feedback and update the learning process, perform iterative training, and finally obtain a trained segmentation model.
[0081] The working principle of the above technical solution: Collect a large number of representative sample images. These images should cover a variety of different scenes, objects, and backgrounds to ensure the generalization ability of the segmentation model. Perform line-level annotation on the feature regions in the sample images. Line-level annotation means accurately depicting the position and shape of each annotated region. Line-level annotation can provide more accurate information and help the model learn more refined segmentation boundaries. Perform superpixel segmentation on the annotated sample images. Superpixel segmentation is a technique that divides an image into several uniform, compact, and connected small regions. The pixels within each small region (i.e., superpixel) have high similarity in terms of features such as color and texture. Superpixel segmentation helps reduce the computational complexity of subsequent processing and makes feature extraction and segmentation optimization more efficient. For each superpixel region, extract features from the line-level annotation. These features include color, texture, shape, etc., which can describe the visual attributes of the superpixel region. An image representation is a high-dimensional vector representation of image data that can capture the key information and structure in the image. Optimize the segmentation of the sample image through an image segmentation algorithm. The goal of the segmentation algorithm is to divide the image into several regions such that the pixels within each region are visually consistent, while the pixels between different regions are significantly different. The segmentation optimization process includes conditional random fields (CRF), graph cut, etc., to improve the accuracy and robustness of the segmentation. According to the consistency between the segmentation result and the line-level annotation, evaluate the performance of the segmentation model. If there is a large difference between the segmentation result and the annotation, it indicates that the model has not learned sufficient segmentation ability. Feedback and update the learning process, and perform iterative training. In each iteration, adjust the parameters and structure of the model according to the evaluation result to improve its segmentation performance. Repeat the above steps until the performance of the segmentation model reaches a predetermined standard or converges. The finally obtained trained segmentation model can accurately identify and segment the feature regions in the image.
[0082] Beneficial effects of the above technical solution: It is convenient to obtain an accurate segmentation model.
[0083] According to some embodiments of the present invention, extracting local features and label position features of a preprocessed image includes:
[0084] Inputting the preprocessed image into a convolutional neural network for size adjustment and normalization processing;
[0085] Based on the convolutional layer in the convolutional neural network as a feature extractor, extracting a high-level feature map from the preprocessed image;
[0086] Determining a target region on the feature map and using a pooling operation on the target region to obtain local features;
[0087] Applying position encoding on the feature map to determine label position features.
[0088] Working principle of the above technical solution: Input the preprocessed image into a convolutional neural network. Before input, adjust the size of the image to ensure it meets the requirements of the network input layer. Size adjustment includes operations such as scaling, cropping, or padding. At the same time, perform normalization processing on the image, that is, convert the pixel values of the image to a unified range (such as 0 to 1 or -1 to 1) to eliminate differences in brightness, contrast, etc. between different images and improve the generalization ability of the network. Based on the convolutional layer in the convolutional neural network as a feature extractor. The convolutional layer extracts local features in the image through convolutional operations, such as edges, textures, etc. As the network deepens, the convolutional layer can extract higher-level features and extract a high-level feature map from the preprocessed image. The feature map is the output result of the convolutional layer, which contains feature information at different positions in the image. Each channel of the feature map corresponds to a specific feature, such as color, texture, or shape, etc. Determine the target region on the feature map. The target region is an object, region, or part of interest in the image. The determination of the target region can be achieved through predefined rules, prior knowledge, or more advanced image analysis algorithms. Use a pooling operation on the target region. The pooling operation is a downsampling technique that can reduce the size and computational complexity of the feature map while retaining the main features. By performing a pooling operation on the target region, local features of the region can be obtained. Apply position encoding on the feature map. Position encoding is a technique for embedding position information into the feature map. It can assign a unique position identifier to each feature point, thereby retaining the spatial information in the image. Position encoding can be achieved in various ways, such as using a coordinate grid, learnable position embeddings, etc. The label position feature refers to the feature information related to the label position in the image.
[0089] Beneficial effects of the above technical solution: Improve the accuracy and efficiency of extracting local features and label position features.
[0090] According to some embodiments of the present invention, query information is determined based on local features and label position features, the relevance of keywords in the query information is calculated, the weights are obtained through normalization by the softmax function, and then the weighted sum is calculated to obtain the attention value, and the relationship feature is determined, including:
[0091] Fuse the local features and label position features to determine the query information;
[0092] Use cosine similarity to query the relevance of each keyword in the information, and obtain the relevance score;
[0093] Normalize the relevance score through the Softmax function to determine the weight of each keyword;
[0094] Use the normalized weights to perform weighted summation on the values of the keywords to obtain the attention value, and determine the relationship feature.
[0095] The working principle of the above technical solution: The feature fusion method includes concatenation, weighted summation, or using a more complex neural network structure (such as a fully connected layer) to implement. The fused feature vector contains the local information of the target region in the image and the position information of the label. The fused feature vector can be regarded as the query information, which represents the comprehensive features of a specific region in the image and its label position. Use cosine similarity to query the relevance of each keyword in the information, and obtain the relevance score; normalize the calculated relevance score through the Softmax function. The Softmax function can map a set of real numbers to the interval (0,1), and the sum of these real numbers is 1. In this way, each keyword obtains a normalized weight, indicating its relative importance among all keywords. Use the normalized weights to perform weighted summation on the values (or feature vectors) of the keywords. The result of the weighted summation is a new feature vector, which represents the information synthesized by all keywords according to the weights. The feature vector obtained by the weighted sum can be regarded as the relationship feature, indicating the degree of association between the query information and each keyword.
[0096] The beneficial effect of the above technical solution: Determine the relationship feature according to the local feature and the label position feature, so as to realize more accurate and efficient processing of specific regions and labels in the image.
[0097] According to some embodiments of the present invention, the attention mechanism is incorporated into the multi-label image category recognition model, and the calculated relationship feature is fused with the local image feature and input into the next-level network for information transmission, including:
[0098] Introduce the attention mechanism into the multi-label image category recognition model, use the spatial attention map to calculate the importance of different regions in the image, and mark them;
[0099] Fuse the relationship features with the local features in the labeled graph, determine the fused features, and input them into the next-level network for information transmission until the output result is determined; wherein, the next-level network includes a fully connected layer, a convolutional layer, and a recurrent neural network layer.
[0100] The working principle of the above technical solution: In the multi-label image category recognition model, a spatial attention mechanism is introduced. First, a spatial attention map is calculated, which can reflect the importance of different regions in the image for the recognition task. It is implemented through an additional convolutional neural network branch that takes the original image or intermediate feature map as input and outputs a spatial attention map of the same size as the input image. The spatial attention map is used to calculate the importance of different regions in the image and mark them, thus facilitating the determination of the importance of the subsequent fused features. The fused features contain both local detailed information and global relationship information. The fused features are input into the next-level network for further processing and information transmission. The next-level network can include a fully connected layer (for feature classification), a convolutional layer (for feature extraction and transformation), a recurrent neural network layer (for capturing sequence information, such as text recognition in images), etc. In the next-level network, the features are passed and processed layer by layer, and finally the output result is obtained. For the multi-label image category recognition task, the output is a probability distribution vector, where each element represents the existence probability of the corresponding label. By setting a threshold, the category labels in the image can be determined.
[0101] The beneficial effects of the above technical solution: Effectively integrate the attention mechanism into the multi-label image category recognition model, and use the fusion of relationship features and local features to improve the recognition accuracy and robustness of the model.
[0102] According to some embodiments of the present invention, determining multiple category labels to which the preprocessed image belongs according to the output result and a preset classification threshold includes:
[0103] Determine the preset classification threshold;
[0104] Compare the output result with the preset classification threshold, and determine multiple category labels to which the preprocessed image belongs according to the comparison result.
[0105] Working principle of the above technical solution: The preset classification threshold is determined according to the characteristics of the specific task and dataset. This threshold is used to convert the output of the model into the final class label. The threshold can be set in various ways, such as a fixed value (e.g., 0.5), cross-validation, grid search, or selecting the optimal threshold based on the performance on the validation set. In practical applications, different thresholds need to be set for different classes to better balance the recognition performance of different classes. The output result of the model is a probability distribution vector, where each element represents the probability of the corresponding label's existence. Compare each probability value in the output result with the preset classification threshold. If a certain probability value is greater than or equal to the threshold, it is considered that the preprocessed image belongs to the corresponding class label; otherwise, it does not belong to that class. According to the comparison results, collect all the class labels that meet the conditions to form a label set. This set represents the multiple class labels to which the preprocessed image belongs.
[0106] Beneficial effects of the above technical solution: It is convenient to accurately determine the multiple class labels to which the preprocessed image belongs.
[0107] According to some embodiments of the present invention, a spatial attention map is used to calculate the importance of different regions in the image and mark them, including:
[0108] Based on the first-order finite difference approximation formula of 2×2, calculate the first importance value S of the corresponding region W(x, y) in the image in the X direction x and the second importance value S in the Y direction y ;
[0109]
[0110] According to the first importance value S of the corresponding region W(x, y) in the X direction x and the second importance value S in the Y direction y , determine the important amplitude;
[0111]
[0112] where S is the important amplitude;
[0113] Based on the important amplitude, display the importance of the corresponding region and mark it.
[0114] Working principle of the above technical solution: Use a 2×2 first-order finite difference approximation formula to calculate the importance values of each pixel point (x, y) in the image in the X direction (horizontal direction) and Y direction (vertical direction). Use the gradient magnitude S to represent the importance of each pixel point in the image. This value reflects the edge strength of the image at that pixel point. According to the magnitude of the gradient magnitude S, the important regions in the image can be marked. For example, a threshold can be set, and the pixel points with a gradient magnitude greater than the threshold are marked as important regions, and different colors or borders are used to highlight these regions.
[0115] Beneficial effects of the above technical solution: Use spatial gradients to calculate the importance of different regions in the image, and perform marking and visual display, which is convenient for determining the importance of each region, and further convenient for determining the importance of subsequent fusion features.
[0116] According to some embodiments of the present invention, dividing the region set according to the marking result to obtain a preprocessed image, including:
[0117] Based on the marking result, delimit an initial region set, determine the pixel values of each pixel point in the initial region set, and organize them into a color histogram;
[0118] Count the number of pixel points with a preset type of pixel value in the initial region set, and judge whether the initial region set is qualified according to the number of pixel points with the preset type of pixel value and the color histogram;
[0119]
[0120] Among them, is the weight information of the pixel point with the preset type of pixel value in the color histogram G; F(w ,p) is the distance between the jth pixel point in the color histogram and the central pixel point p in the color histogram; F j is the maximum value of F(w 0 ,p); w j is the jth pixel point in the color histogram; k(w j ) is the pixel value of the jth pixel point in the color histogram; j is is 's weight coefficient; M is the number of pixel points included in the color histogram; T is the number of pixel points with a preset type of pixel value in the color histogram;
[0121] When it is determined that the weight information is greater than the preset threshold, it means that the initial region set is qualified; otherwise, it means that the initial region set is unqualified and needs to be adjusted to finally obtain a preprocessed image.
[0122] Working principle of the above technical solution: Based on the marking results, an initial region set containing specific pixel points is delimited. For example, an initial region set is delimited based on the first type of pixel points. For each pixel point in the initial region set, its pixel value is determined, and these pixel values are organized into a color histogram to reflect the distribution of different color pixels in the region. In the color histogram, the number of pixel points with pixel values belonging to a preset type is statistically counted. The preset type pixel values among the first type of pixel points, the second type of pixel points, and the third type of pixel points are all different, and the sorting of the preset type pixel values among the first type of pixel points, the second type of pixel points, and the third type of pixel points from largest to smallest is: the preset type pixel value of the second type of pixel points, the preset type pixel value of the third type of pixel points, the preset type pixel value of the first type of pixel points. According to the number of pixel points with the preset type pixel value and the color histogram, it is judged whether the initial region set is qualified; this weight information reflects the relative importance of the preset type pixel value in the entire color histogram. When it is determined that the weight information is greater than the preset threshold, it indicates that the initial region set is qualified, otherwise, it indicates that the initial region set is unqualified and needs to be adjusted, and finally a preprocessed image is obtained. For an unqualified initial region set, it can be adjusted by expanding or shrinking the region range, changing the marking strategy, or reselecting the region, etc.
[0123] Beneficial effects of the above technical solution: The region set is delimited according to the marking results, and it is judged whether the initial region set is qualified based on the color histogram and the statistical count of the preset type pixel values, and finally a preprocessed image is obtained, realizing the accuracy of region classification in the preprocessed image and facilitating the improvement of the accuracy of obtaining the preprocessed image.
[0124] As Figure 2 shown, according to some embodiments of the present invention, an apparatus applying the multi-label image recognition method integrating an attention mechanism as described above includes:
[0125] An acquisition module, configured to acquire a multi-label image;
[0126] A preprocessing module, configured to preprocess the multi-label image to obtain a preprocessed image;
[0127] An extraction module, configured to extract local features and label position features of the preprocessed image;
[0128] A calculation module, configured to determine query information according to the local features and label position features, calculate the relevance of keywords in the query information, normalize it through a softmax function to obtain a weight, and then calculate a weighted sum to obtain an attention value and determine relationship features;
[0129] The first determination module is used to incorporate the attention mechanism into the multi-label image category recognition model, fuse the relationship features and local features, determine the fused features, and input them into the next-level network for information transmission until the output result is determined.
[0130] The second determination module is used to determine multiple category labels to which the preprocessed image belongs according to the output result and a preset classification threshold.
[0131] Beneficial effects of the above technical solution: By enhancing the model's attention to key information in the image, it is convenient to improve the classification efficiency and accuracy of multi-label images.
[0132] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.
Claims
1. A multi-label image recognition method integrating attention mechanism, characterized in that: include: Get multi-label images; Preprocess the multi-label image to obtain a preprocessed image; Extract local features and label position features of the preprocessed image; Determine the query information based on local features and label position features, calculate the relevance of keywords in the query information, normalize them using the softmax function to get the weights, calculate the weighted sum to get the attention value, and determine the relationship features; Integrate the attention mechanism into the multi-label image category recognition model, fuse the relational features with the local features, determine the fusion features, and input them into the next level network for information transmission until the output result is determined; According to the output results and the preset classification threshold, multiple category labels to which the preprocessed image belongs are determined.
2. The multi-label image recognition method integrating the attention mechanism as claimed in claim 1, characterized in that: Preprocess the multi-label image to obtain a preprocessed image, including: Input the multi-label image into the pre-trained segmentation model and output the semantic segmentation result of the multi-label image; Determine the R channel value, G channel value, and B channel value of all pixels on the image corresponding to each superpixel region in the semantic segmentation result; compare the R channel value, G channel value, and B channel value of each pixel to obtain a comparison result, where when the comparison result is that the R channel value is the maximum value, the corresponding pixel is a first type of pixel, when the comparison result is that the G channel value is the maximum value, the corresponding pixel is a second type of pixel, and when the comparison result is that the B channel value is the maximum value, the corresponding pixel is a third type of pixel; The types of all pixels on the image corresponding to the superpixel region are marked, and the region set is divided according to the marking result to obtain a preprocessed image.
3. The multi-label image recognition method integrating the attention mechanism as claimed in claim 2, characterized in that: Methods for obtaining pre-trained segmentation models include: Obtain a sample image, perform line-level annotation on a feature area in the sample image to indicate the position and shape of the annotated area; perform superpixel segmentation on the annotated sample image to divide the annotated sample image into multiple superpixel areas; For each superpixel region, features are extracted from line-level annotations, image representation is learned, and the sample image is segmented and optimized using an image segmentation algorithm. According to the consistency between the segmentation results and the line-level annotations, the learning process is fed back and updated, and iterative training is performed to finally obtain a trained segmentation model.
4. The multi-label image recognition method integrating the attention mechanism as claimed in claim 1, characterized in that: Extract local features and label position features of the preprocessed image, including: Input the preprocessed image into the convolutional neural network for resizing and normalization; Based on the convolutional layer in the convolutional neural network as a feature extractor, a high-level feature map is extracted from the preprocessed image; Determine the target area on the feature map, apply pooling operation to the target area, and obtain local features; Apply position encoding on the feature map to determine the label position features.
5. The multi-label image recognition method integrating the attention mechanism as claimed in claim 1, characterized in that: Determine the query information based on local features and label position features, calculate the relevance of keywords in the query information, normalize them using the softmax function to get the weights, calculate the weighted sum to get the attention value, and determine the relationship features, including: Fuse local features and tag position features to determine query information; Use cosine similarity to query the relevance of each keyword in the information and obtain the relevance score; Normalize the relevance score through the Softmax function to determine the weight of each keyword; Use the normalized weights to perform weighted summation on the keyword values to obtain the attention value and determine the relationship features.
6. The multi-label image recognition method integrating the attention mechanism as claimed in claim 1, characterized in that: Integrate the attention mechanism into the multi-label image category recognition model, fuse the calculated relationship features with the local image features, and input them into the next level network for information transmission, including: Introducing the attention mechanism into the multi-label image category recognition model, using the spatial attention map to calculate the importance of different areas in the image and mark them; In the labeling graph, the relational features are fused with the local features to determine the fused features, and the fused features are input into the next level network for information transmission until the output result is determined; wherein the next level network includes a fully connected layer, a convolutional layer, and a recurrent neural network layer.
7. The multi-label image recognition method integrating the attention mechanism as claimed in claim 1, characterized in that: According to the output results and the preset classification threshold, multiple category labels to which the preprocessed image belongs are determined, including: Determining a preset classification threshold; The output result is compared with the preset classification threshold, and the multiple category labels to which the preprocessed image belongs are determined according to the comparison result.
8. The multi-label image recognition method integrating the attention mechanism as claimed in claim 6, characterized in that: Use spatial attention maps to calculate the importance of different regions in the image and mark them, including: Based on the 2×2 first-order finite difference approximation, the first important value S of the corresponding area W(x, y) in the X direction in the image is calculated x and the second most important value S in the Y direction y ; According to the first important value S of the corresponding area W(x, y) in the X direction x and the second most important value S in the Y direction y , determine the important amplitude; Among them, S is the important amplitude; The importance of the corresponding area is displayed based on the important magnitude and marked.
9. The multi-label image recognition method integrating the attention mechanism as claimed in claim 2, characterized in that: The region set is divided according to the marking results to obtain a preprocessed image, including: Based on the marking results, an initial region set is delineated, the pixel value of each pixel in the initial region set is determined, and the pixel value is organized into a color histogram; Counting the number of pixel points of the preset type of pixel value in the initial region set, and judging whether the initial region set is qualified according to the number of pixel points of the preset type of pixel value and the color histogram; in, The pixel value of the preset type in the color histogram G The weight information of the pixel points; F(w j , p) is the distance between the jth pixel in the color histogram and the center pixel p in the color histogram; F0 is F(w j , p) maximum value; w j is the jth pixel in the color histogram; k(w j ) is the pixel value of the jth pixel in the color histogram; for M is the number of pixels included in the color histogram; T is the number of pixels of the preset type in the color histogram; When it is determined that the weight information is greater than a preset threshold, it indicates that the initial region set is qualified, otherwise, it indicates that the initial region set is unqualified and needs to be adjusted, and finally a preprocessed image is obtained.
10. A device for applying the multi-label image recognition method integrating the attention mechanism as described in any one of claims 1 to 9, characterized in that: include: Acquisition module, used to acquire multi-label images; A preprocessing module, used for preprocessing the multi-label image to obtain a preprocessed image; An extraction module, used to extract local features and label position features of the preprocessed image; The calculation module is used to determine the query information based on the local features and the label position features, calculate the relevance of the keywords in the query information, obtain the weights through normalization by the softmax function, calculate the weighted sum to obtain the attention value, and determine the relationship features; The first determination module is used to integrate the attention mechanism into the multi-label image category recognition model, fuse the relationship features with the local features, determine the fusion features, and input them into the next level network for information transmission until the output result is determined; The second determination module is used to determine the multiple category labels to which the pre-processed image belongs according to the output result and a preset classification threshold.