A multi-label image recognition system based on saliency map
Through the multi-label image recognition system of the significant graph, combined with the multi-label classification and distribution loss function, the significant graph is generated and iteratively trained, which solves the problems of occlusion, lighting, viewpoint, etc. in multi-label image recognition, and improves the recognition accuracy and efficiency.
Patent Information
- Application Number
- CN202011008930.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-09-23
AI Technical Summary
When the existing multi-label image recognition model deals with problems such as occlusion, lighting, and viewpoint, it is difficult to effectively extract and fuse multi-label features, resulting in a decrease in recognition accuracy. The existing methods have problems with a lot of noise and poor results when applied in actual scenarios.
A multi-label image recognition system with significant graphs is adopted. Through image preprocessing, feature extraction, classification modules and training control modules, combining multi-label classification loss and distribution loss functions, significant graphs are generated and iteratively trained to reduce interference from complex backgrounds and object deformations and improve recognition accuracy.
It effectively reduces interference from complex backgrounds and object deformation, improves the accuracy of multi-label image recognition, and can accurately identify objects in the image in complex environments, adapt to large-scale data training, reduces computing equipment requirements, and improves work efficiency.
Smart Images

Figure CN114255376B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a multi-label image recognition system based on a saliency map, and belongs to the technical field of image processing. Background Art
[0002] With the continuous development of information technology, computer vision technology has made great progress and has been widely used in various practical applications. It has extremely high scientific research value and great commercial potential. Multi-label image recognition technology is a type of computer vision technology. As the name suggests, multi-label image recognition means that the image to be classified has one or more labels, and the task goal is to find all the labels contained in the image. Since a large number of natural scene images contain rich semantic information, it is necessary to use multiple labels to accurately convey the effective information contained in the sample. Therefore, multi-label tasks have broad practical application prospects, such as: album management, social media image annotation and autonomous driving, etc., all involve multi-label classification tasks.
[0003] Generally, multi-label images have more complex backgrounds and more effective information than single-label images. At the same time, multi-label images have more serious problems such as occlusion, lighting, and viewpoint, which makes the morphology of objects with the same label in different images vary greatly. This phenomenon has caused many image recognition models that have achieved superhuman capabilities in single-label datasets to have varying degrees of performance degradation in multi-label image recognition tasks.
[0004] The mainstream method is to improve the ability of multi-label image recognition models by building label correlation and introducing attention mechanism. However, the recurrent neural network or graph neural network used to build label correlation requires that the label distribution of real images and the label distribution of training images maintain a high degree of consistency when analyzing the high-order relationship between labels, which hinders the application of the model in actual scenarios.
[0005] By introducing the attention mechanism to assist the model in learning the features of key areas, the relationship between local and local, local and overall image can be analyzed, which can improve the efficiency of local area analysis of the image. In the multi-label image recognition task, the method of combining the attention mechanism with target detection algorithm, spatial regularization algorithm, image semantic segmentation technology, etc. to achieve better results, and the method of improving the performance of the multi-label image recognition model by utilizing the consistency of the attention area, these two methods still cannot effectively convert the multi-label fusion features extracted by the model into single-label features, and the obtained data still includes a lot of noise, which cannot effectively alleviate the problems caused by object deformation, occlusion, lighting, viewpoint, etc. Summary of the invention
[0006] In order to solve the above problems, a multi-label image recognition system of saliency map is provided. The present invention adopts the following technical solutions:
[0007] The present invention provides a multi-label image recognition system for salient maps, which is used to identify the labels contained in a natural scene image to be recognized after processing the input natural scene image. The system is characterized by including: an image preprocessing module, a feature extraction module, a classification module, a training control module, a training module, a judgment module, and an identification control module. Among them, the image preprocessing module randomly crops the images in the natural scene image dataset to obtain training images of the same size. The feature extraction module processes the training images to obtain training feature maps with specific dimensions. The classification module processes the training feature maps to obtain the training confidence scores and classification weights of each label. The training control module is used to control the training module to calculate the multi-label classification loss function and the multi-label distribution loss function according to the training confidence scores and the distribution of the expression features in the high-dimensional space, and perform training iterations on the feature extraction module and the classification module with the multi-label classification loss function and the multi-label distribution loss function to obtain an updated feature extraction module and an updated classification module. The training module includes a salient map generation unit, a feature selection unit, a multi-label distribution loss unit, and a multi-label classification loss unit. The salient map generation unit is used to calculate the salient maps corresponding to each label based on the classification weights and the corresponding training feature maps. The feature selection unit is used to process the salient maps of each label and the training feature maps to obtain feature selection weights and the expression features corresponding to each category. The multi-label distribution loss unit is used to calculate the multi-label distribution loss function based on the expression features, weights, and training confidence scores of each label. The multi-label classification loss unit is used to obtain the multi-label classification loss function based on the training confidence scores. The identification control module controls the image preprocessing module to process the input natural scene image to obtain an input image, controls the updated feature extraction module to process the input image to generate a set of feature maps, and controls the updated classification module to calculate the feature maps to obtain the confidence scores of each label. Further, the judgment module determines the labels expressed by the natural scene image according to the confidence scores.
[0008] The multi-label image recognition system for salient maps provided by the present invention may further have the following technical feature: The multi-label classification loss unit and the multi-label distribution loss unit train the multi-label image recognition network in a supervised manner according to the training confidence scores of each label and the multi-label image recognition network loss calculated from the distribution of the expression features of each label in the high-dimensional space. The specific calculation formula of the multi-label classification loss function in the multi-label classification loss unit is:
[0009] In the formula, x and y respectively represent the training confidence scores and the true results of the classification output, L represents the total number of labels included in the natural scene image dataset; i represents label i, and e represents the natural constant.
[0010] A multi-label image recognition system for saliency maps provided by the present invention may further have the following technical features. Among them, the multi-label distribution loss part includes an anchor unit, a feature distribution compactness evaluation unit, and a feature distribution discreteness evaluation unit. The anchor unit is used to record the anchor positions corresponding to the expression features of each label during training and the training confidence scores corresponding to the anchors. The anchor unit includes the following steps: Step S1-1, set both the expression feature of the first label and the initial anchor of the training confidence score corresponding to the expression feature to the origin; Step S1-2, record the weighted value of the training confidence score of the current label and the training confidence score of the current label in the classification module as the training confidence score of the next label, update the origin position of the confidence score, and record the updated origin position as the anchor position of the confidence score. The calculation formula is: In the formula, j represents the image j including label i in the current batch, represents the weight of the current label i recorded by the anchor, represents the weight of label i after update, is the training confidence score of label i output by the linear classifier based on the natural scene image; Step S1-3, calculate the expression feature of the next label through the training confidence score of the current label, the center point of the expression feature of the current label, and the expression feature of the current label, update the origin position of the expression feature, and record the updated origin position as the anchor position of the expression feature. The calculation formula is: In the formula, j represents the image j including label i in the current batch, represents the weight of label i recorded in the anchor unit, is the training confidence score of label i output by the linear classifier based on the natural scene image, z i is the expression feature of label i, is the center point represented by the expression feature of label i recorded by the anchor; Step S1-4, take the next label as the current label, repeat Step S1-2 and Step S1-3 until the last label, record and output the training confidence scores of each label and the expression features of each label. The expression feature distribution compactness evaluation unit evaluates the compactness of the expression feature distribution according to the expression feature distribution compactness CP i of label i, and calculate the formula for CP i is: In the formula, N i represents the number of sample points in label i, Represents the center point of label i, represents the nth i sample point in label i, d(.) represents the Euclidean distance between two points, and CP i represents the compactness of the expression feature distribution of label i. The expression feature distribution discreteness evaluation unit evaluates the discreteness of the visual features of different labels from the same natural scene image in the natural scene image dataset in their spatial distribution. According to DT ij representing the expression feature distribution discreteness between label i and label j, evaluate the expression feature distribution discreteness, and calculate DT ij The formula is: In the formula, respectively represent the center points of label i and label j, and DT ij represents the expression feature distribution discreteness between label i and label j. The expression feature distribution discreteness and the expression feature distribution aggregation reflect the distribution of the expression features in the high-dimensional space.
[0011] A multi-label image recognition system for a saliency map provided by the present invention may also have such a technical feature that a multi-label distribution loss function is calculated according to the parameters calculated by the expression feature distribution compactness evaluation unit, the expression feature distribution discreteness evaluation unit, and the anchor point unit. The calculation formula is: In the formula, L′ represents the set of labels included in the natural image, and CP i represents the compactness of the expression feature distribution of label i, and DT ij represents the expression feature distribution discreteness between label i and label j. The multi-label distribution loss function L DBI and the multi-label classification loss function loss are passed to the classification module and the feature extraction module to update and iterate the parameters in the classification module and the feature extraction module through the backpropagation algorithm until the multi-label image recognition network converges.
[0012] A multi-label image recognition system for a saliency map provided by the present invention may also have such a technical feature that the saliency map generation unit multiplies and sums the corresponding classification weights and the corresponding training feature maps to obtain the saliency maps of each label of the training image. The specific calculation process is as follows: In the formula, S i represents the saliency map of label i, represents the classification weight between the kth dimension of the vector obtained by global average pooling and label i; A k represents the training feature map corresponding to the kth dimension.
[0013] A multi-label image recognition system for saliency maps provided by the present invention may further have the following technical feature: the feature selection unit selects the saliency maps with labels in each label saliency map as the output saliency maps, multiplies the output saliency maps with a set of training feature maps with a specific dimension as weight matrices respectively, and globally sums them to obtain an output vector, and uses the output vector as the expression feature of each label corresponding to each output saliency map; the calculation process is to directly use the values corresponding to each pixel point in the output saliency map as the feature selection weights of the pixel points corresponding to the training feature maps of each label. The specific calculation process is as follows: In the formula, z i and S i respectively represent the expression feature and the saliency map of label i, A k represents the training feature map corresponding to the k-th dimension, and x, y represent the points with coordinates (x, y) in the saliency map and the training feature map.
[0014] A multi-label image recognition system for saliency maps provided by the present invention may further have the following technical feature: the image preprocessing module collects images and the labels in the images in a natural scene image dataset, and randomly crops the natural scene images to obtain training images of the same size.
[0015] A multi-label image recognition system for saliency maps provided by the present invention may further have the following technical feature: the feature extraction module is composed of a convolutional neural network with cross-layer branches.
[0016] A multi-label image recognition system for saliency maps provided by the present invention may further have the following technical feature: the classification module includes a linear classifier, globally pools the training feature maps to obtain a vector, and passes the vector through the linear classifier to obtain the training confidence scores of each label and the classification weights between each dimension of the vector and each label.
[0017] Functions and effects of the invention
[0018] A multi-label image recognition system based on a saliency map according to the present invention crops and processes an image preprocessing module to obtain training images of the same size, processes the training images through a feature extraction module to obtain training feature maps with specific dimensions, and processes the training feature maps through a classification module to obtain training confidence scores and classification weights for each label. The training control module controls the multi-label classification loss function and the multi-label distribution loss function obtained by the training module to update and iterate the feature extraction module and the classification module to obtain an updated feature extraction module and an updated classification module. Finally, the output image obtained by inputting a natural scene image into the image preprocessing module passes through the updated feature extraction module and the updated classification module to identify the classification result. Therefore, the multi-label image recognition system provided in this embodiment can analyze each region in the image and is easier to process. Each training image has a corresponding saliency map, and the obtained label results are both complete and accurate. The updated feature extraction module and the updated classification module obtained by iterating according to the multi-label classification loss function and the multi-label distribution loss function can effectively reduce the interference of complex backgrounds and object deformations during the multi-label image recognition process, avoid interference from occlusion, lighting, viewpoints, etc., and improve the accuracy of multi-label image recognition in natural scenes. Thus, the problem of interference from complex backgrounds and variable object morphologies in the multi-label image recognition task is solved, and a multi-label image recognition system that can accurately identify the objects contained in the image in complex backgrounds and variable object morphologies and can update and iterate the feature extraction module and the classification module while recognizing natural scene images is provided. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 FIG. is a schematic diagram of a multi-label image recognition system based on a saliency map in an embodiment of the present invention;
[0020] Figure 2 FIG. is a schematic diagram of a training module 1005 in a multi-label image recognition system based on a saliency map in an embodiment of the present invention;
[0021] Figure 3 FIG. is a flowchart of a training module 1005 in a multi-label image recognition system based on a saliency map in an embodiment of the present invention; and
[0022] Figure 4 FIG. is a flowchart of a multi-label image recognition system based on a saliency map in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] In order to make the technical means, creative features, achieved purposes, and functions of the present invention easy to understand, the following specifically describes a multi-label image recognition system based on a saliency map of the present invention in conjunction with embodiments and drawings.
[0024] <Example>
[0025] Figure 1 It is a schematic diagram of a multi-label image recognition system based on a saliency map in an embodiment of the present invention.
[0026] As Figure 1 shown, a multi-label image recognition system 1000 based on a saliency map includes an image preprocessing module 1001, a feature extraction module 1002, a classification module 1003, a training control module 1004, a training module 1005, a judgment module 1006, and an identification control module 1007.
[0027] The image preprocessing module 1001 randomly crops the natural scene images in the natural scene image dataset to obtain training images of the same size.
[0028] Among them, the image preprocessing module 1001 collects the natural scene images for training and the labels in the natural scene images in the natural scene image dataset, and randomly crops the images to obtain training images of the same size.
[0029] In this embodiment, the labels in the natural scene images correspond to the categories of objects, and the images and their annotations in this embodiment are obtained through public information (such as by web crawling or batch import, etc.).
[0030] The feature extraction module 1002 processes the training images to obtain training feature maps with specific dimensions.
[0031] In this embodiment, the feature extraction module 1002 is composed of a convolutional neural network with cross-layer branches.
[0032] The classification module 1003 processes the training feature maps to obtain the training confidence scores and classification weights of each label.
[0033] Among them, the classification module 1003 includes a linear classifier. The classification module 1003 performs global pooling on the training feature maps to obtain vectors and passes the vectors through the linear classifier to obtain the training confidence scores of each label and the classification weights between each dimension of the vector and each label.
[0034] Among them, the dimension of the classification weights obtained by the linear classifier is the same as the number of labels included in the dataset.
[0035] The training control module 1004 is used to control some of the work related to the training process, specifically to control the work of the training module 1005. Specifically, it controls the training module 1005 to calculate the multi-label image recognition network loss function according to the training confidence score and the distribution of expression features in the high-dimensional space, and trains the feature extraction module 1002 and the classification module 1003 to obtain the updated feature extraction module 1002 and the updated classification module 1003.
[0036] Figure 2 It is a schematic diagram of the training module 1005 in a multi-label image recognition system based on a saliency map in an embodiment of the present invention.
[0037] As Figure 2 shown, the training module 1005 includes a saliency map generation unit 501, a feature selection unit 502, a multi-label distribution loss unit 503, and a multi-label classification loss unit 504.
[0038] The saliency map generation unit 501 multiplies and sums the corresponding weights and the corresponding training feature maps to obtain the saliency maps of each label of the output image. The specific calculation process is as follows:
[0039]
[0040] In the formula, S i represents the saliency map of label i, represents the classification weight between the k-th dimension of the vector obtained by global average pooling and label i; A k represents the training feature map corresponding to the k-th dimension.
[0041] Among them, the feature selection unit 502 selects the saliency maps with labels in each label saliency map as the output saliency maps, multiplies the output saliency maps with a group of training feature maps with specific dimensions as weight matrices respectively, and globally sums them to obtain the output vector, and uses the output vector as the expression features of each label corresponding to each output saliency map; the calculation process is to directly use the values corresponding to each pixel point in the output saliency map as the classification weights of the pixel points corresponding to the training feature maps of each label. The specific calculation process is as follows:
[0042]
[0043] In the formula, z i and S i represent the expression feature and the saliency map of label i respectively, A k represents the training feature map corresponding to the k-th dimension, and x, y represent the points with coordinates (x, y) in the saliency map and the training feature map.
[0044] The specific calculation formula of the multi-label classification loss function in the multi-label classification loss unit 503 is:
[0045]
[0046] Wherein, y and y respectively represent the confidence score for training of the classification output and the true result, L represents the total number of labels included in the natural scene image dataset; i represents label i, and e represents the natural constant.
[0047] The multi-label distribution loss part 504 includes an anchor unit 41, an expression feature distribution compactness evaluation unit 42, and an expression feature distribution discreteness evaluation unit 43.
[0048] The anchor unit 41 is used to record the anchor position corresponding to the expression feature of each label during training and the confidence score for training corresponding to the anchor.
[0049] The anchor unit 41 includes the following steps:
[0050] Step S1-1, set both the expression feature of the first label and the initial anchor of the confidence score for training corresponding to the expression feature to the origin, and then enter step S1-2;
[0051] Step S1-2, record the weighted value of the confidence score for training of the current label and the confidence score for training of the current label in the classification module 1003 as the confidence score for training of the next label, update the origin position of the confidence score and record the updated origin position as the anchor position of the confidence score, and the calculation formula is:
[0052]
[0053] Wherein, j represents the image j including label i in the current batch, represents the weight of the current label i recorded by the anchor, represents the weight of label i after update, is the confidence score for training of label i output by the linear classifier based on the natural scene image, and then enter step S1-3;
[0054] Step S1-3, calculate the confidence score for training of the current label, the center point of the expression feature of the current label, and the expression feature of the current label to obtain the expression feature of the next label, update the origin position of the expression feature and record the updated origin position as the anchor position of the expression feature, and the calculation formula is:
[0055]
[0056] Wherein, j represents the image j including label i in the current batch, represents the weight of label i recorded in the anchor unit 41, is the confidence score for training the label i output by the linear classifier based on the natural scene image, z i is the expression feature of label i, is the center point represented by the expression feature of label i recorded by the anchor point, and then enter step S1-4;
[0057] Step S1-4, mark the next label as the current label, repeat steps S1-2 and S1-3 until the last label, record and output the confidence scores for training of each label and the expression features of each label, and end the process.
[0058] The expression feature distribution compactness evaluation unit 42 evaluates the clustering of the same label expression features from different natural scene images in the natural scene image dataset in their spatial distribution, and according to the expression feature distribution compactness CP of label i i evaluates the expression feature distribution compactness and calculates CP i The formula for is:
[0059]
[0060] In the formula, N i represents the number of sample points in label i, represents the center point of label i, represents the nth i sample point in label i, d(.) represents the Euclidean distance between two points, and CP i represents the expression feature distribution compactness of label i.
[0061] The expression feature distribution discreteness evaluation unit 43 evaluates the discreteness of the expression features of different labels from the same natural scene image in the natural scene image dataset in their spatial distribution.
[0062] Among them, according to DT representing the expression feature distribution discreteness between label i and label j ij evaluates the expression feature distribution discreteness and calculates DT ij The formula for is:
[0063]
[0064] In the formula, respectively represent the center points of label i and label j, and DT ij represents the expression feature distribution discreteness between label i and label j.
[0065] Calculate the multi-label distribution loss function according to the parameters calculated by the anchor unit 41, the expression feature distribution compactness evaluation unit 42, and the expression feature distribution discreteness evaluation unit 43. The calculation formula is:
[0066]
[0067] Wherein, L' represents the set of labels included in the natural image, and CP i represents the compactness of the expression feature distribution of label i, and DT ij represents the discreteness of the expression feature distribution between label i and label j.
[0068] The multi-label distribution loss function L DBI and the multi-label classification loss function loss are passed to the feature extraction module 1002 and the classification module 1003, and the parameters in the feature extraction module 1002 and the classification module 1003 are updated iteratively through the backpropagation algorithm until the multi-label image recognition network converges.
[0069] Figure 3 is the flowchart of the training module 1005 in a multi-label image recognition system based on a saliency map according to an embodiment of the present invention.
[0070] As Figure 3 shown, the process of the training module 1005 includes the following sub-steps:
[0071] Step S2-1, calculate the multi-label classification loss function based on the training confidence scores of each label output by the classification module 1003, and then enter step S2-2;
[0072] Step S2-2, in the first round of training, the initial anchor points of the expression feature of the first label and the training confidence score corresponding to the expression feature are both the origin, and then enter step S2-3;
[0073] Step S2-3, record the weighted value of the training confidence score of each label and the training confidence score of this label in the classification module 1003 as the training confidence score of the anchor point, and then enter step S2-4;
[0074] Step S2-4, calculate the anchor point position of the current label according to the expression feature of the current label and the training confidence score corresponding to the expression feature, and then enter step S2-5;
[0075] Step S2-5, record the anchor point position of the current label and the training confidence score corresponding to the anchor point and output, and then enter step S2-6;
[0076] Step S2-6, the expression feature distribution compactness evaluation unit 42 is used to evaluate the aggregation degree of the same label expression features in different natural scene images in the natural scene image dataset in the spatial distribution, and then enter step S2-7;
[0077] Step S2-7: The feature distribution discreteness evaluation unit 43 is used to evaluate the discreteness of the visual features of different labels of the same natural scene image in the natural scene image dataset in the spatial distribution, and then proceeds to step S2-8;
[0078] Step S2-8: Calculate the multi-label distribution loss function based on the anchor point positions calculated in step S2-5 and the corresponding training confidence scores, the aggregation degree calculated in step S2-6, and the discreteness calculated in step S2-7, and then proceed to step S2-9;
[0079] Step S2-9: Transmit the multi-label classification loss function calculated in step S2-1 and the multi-label distribution loss function calculated in step S2-8 to the feature extraction module 1002 and the classification module 1003, and update and iterate the parameters of the feature extraction module 1002 and the classification module 1003 through the backpropagation algorithm until the multi-label image recognition network converges, and the process ends.
[0080] The updated feature extraction module 1002 and the updated classification module 1003 obtained through the above training process are used to recognize natural scene images through the recognition control module 1007.
[0081] The recognition control module 1007 processes the input natural scene image through the image preprocessing module 1001 to obtain an input image, controls the updated feature extraction module 1002 to process the natural scene image to generate a set of feature maps, and controls the updated classification module 1003 to calculate the feature maps to obtain the confidence scores of each label. Further, the judgment module 1006 determines the labels expressed by the natural scene image based on the confidence scores.
[0082] Figure 4 It is the flowchart of a multi-label image recognition system based on a saliency map in an embodiment of the present invention.
[0083] As Figure 4 shown, Figure 4 For the process of the multi-label image recognition system based on the saliency map, the process of the multi-label image recognition system includes the following steps:
[0084] Step S3-1: The image preprocessing module 1001 randomly crops the natural scene images in the natural scene image dataset to obtain input images of the same size, and proceeds to step S3-2;
[0085] Step S3-2: The updated feature extraction module 1002 processes the input image to obtain a feature map with a specific dimension, and proceeds to step S3-3;
[0086] Step S3-3: The updated classification module 1003 processes the feature map to obtain the confidence scores for each label, and proceeds to Step S3-4;
[0087] Step S3-4: The judgment module 1006 determines the label expressed by the natural scene image based on the confidence scores, and ends the process.
[0088] During the recognition process of the natural scene image, the obtained feature map and confidence scores can continue to be used to calculate the multi-label classification loss function and the multi-label distribution loss function under the control of the training control module 1004, and perform training iterations on the feature extraction module 1002 and the classification module 1003.
[0089] Functions and Effects of the Embodiment
[0090] According to a multi-label image recognition system based on a saliency map of the present invention, the image preprocessing module crops the processed images to obtain training images of the same size, which are processed by the feature extraction module to obtain training feature maps with specific dimensions. The training feature maps are processed by the classification module to obtain the training confidence scores and classification weights for each label. The multi-label classification loss function and the multi-label distribution loss function obtained by the training control module controlling the training module are used to update and iterate the feature extraction module and the classification module to obtain the updated feature extraction module and the updated classification module. Finally, the natural scene image is input into the output image obtained by the image preprocessing module, and the updated feature extraction module and the updated classification module are used to identify the classification result. Therefore, the multi-label image recognition system provided in this embodiment can analyze each region in the image and is easier to process. Each training image has a corresponding saliency map, so the obtained label results are both complete and accurate. The updated feature extraction module and the updated classification module obtained by iterating according to the multi-label classification loss function and the multi-label distribution loss function can effectively reduce the interference of complex backgrounds and object deformations during the multi-label image recognition process, avoid being interfered by occlusion, illumination, viewpoint, etc., improve the accuracy of multi-label image recognition in natural scenes, and thus solve the problem of interference of complex backgrounds and variable object morphologies existing in multi-label image recognition tasks, and provide a multi-label image recognition system that can accurately identify the objects contained in the image in complex backgrounds and variable object morphologies and can update and iterate the feature extraction module and the classification module while recognizing natural scene images.
[0091] In the embodiment, the distribution of expression features in the high-dimensional space is reflected by calculating the discreteness and the aggregation of the expression feature distribution, and the feature extraction module and the classification module are trained by using a multi-label distribution loss function and a multi-label classification loss function that reflect the distribution of the expression features in the high-dimensional space, which can combine the distribution of features in the image, avoid the defect that the representation features of the same class label in the image are too discrete and the representation features of different class labels are too aggregated, so that the obtained multi-label image recognition system can be better trained.
[0092] In the embodiment, in each round of training, the anchor unit is used to record and update the center point and the center point weight of each label expression feature, so as to alleviate the interference caused by locality, thereby further improving the accuracy of multi-label image recognition in natural scenes and reducing the requirements for computing devices, being better adapted to large-scale data and having more practical application value.
[0093] In the embodiment, processing the vector in the way of global pooling can enable faster training, thus playing a role in improving work efficiency.
[0094] The above embodiments are only used to illustrate the specific implementation manners of the present invention, and the present invention is not limited to the description scope of the above embodiments.
Claims
1. A multi-label image recognition system based on a saliency map, which is used to process an input natural scene image to be recognized and identify the labels contained in the natural scene image, and is characterized in that, Including: An image preprocessing module, a feature extraction module, a classification module, a training control module, a training module, a judgment module, and an identification control module. Among them, the image preprocessing module randomly crops the images in the natural scene image dataset to obtain training images of the same size. The feature extraction module processes the training images to obtain training feature maps with specific dimensions. The classification module processes the training feature maps to obtain the training confidence scores and classification weights for each label. The training control module is used to control the training module to calculate the multi-label classification loss function and the multi-label distribution loss function according to the training confidence scores and the distribution of the expression features in the high-dimensional space, and perform training iterations on the feature extraction module and the classification module with the multi-label classification loss function and the multi-label distribution loss function to obtain an updated feature extraction module and an updated classification module. The training module includes a saliency map generation unit, a feature selection unit, a multi-label distribution loss unit, and a multi-label classification loss unit. The saliency map generation unit is used to calculate the saliency maps corresponding to each label by using the classification weights and the corresponding training feature maps. The feature selection unit is used to process the saliency maps of each label and the training feature maps to obtain feature selection weights and the expression features corresponding to each category. The multi-label distribution loss unit is used to calculate the multi-label distribution loss function according to the expression features, weights, and training confidence scores of each label. The multi-label classification loss unit is used to obtain the multi-label classification loss function according to the training confidence scores. The identification control module controls the image preprocessing module to process the input natural scene image to obtain an input image, controls the updated feature extraction module to process the input image to generate a set of feature maps, and controls the updated classification module to calculate the feature maps to obtain the confidence scores of each label. Further, the judgment module determines the label expressed by the natural scene image according to the confidence scores. The multi-label distribution loss unit includes an anchor unit, a feature distribution compactness evaluation unit, and a feature distribution discreteness evaluation unit. The anchor unit is used to record the anchor positions corresponding to the expression features of each label during training and the anchor positions corresponding to the training confidence scores of the anchors. Specifically, it includes the following steps: Step 1-1: Set the initial anchor points of the expression feature of the first label and the training confidence score corresponding to the expression feature to the origin. Step 1-2: Record the weighted value of the training confidence score of the current label and the training confidence score of the current label in the classification module as the training confidence score of the next label, update the origin position of the confidence score, and record the updated origin position as the anchor position of the confidence score. The calculation formula is: where j represents the image j containing the label i in the current batch, represents the weight of the current label i recorded by the anchor point, represents the weight of the updated label i, is the training confidence score of the label i output by the linear classifier based on the natural scene image; Steps 1-3: Calculate the expression feature of the next label from the training confidence score of the current label, the center point of the expression feature of the current label, and the expression feature of the current label, update the origin position of the expression feature, and record the updated origin position as the anchor position of the expression feature. The calculation formula is as follows: where j represents the image j containing label i in the current batch, denotes the weight of the label i recorded in the anchor unit, is the training confidence score of the label i output by the linear classifier based on the natural scene image, z i is the expression feature of the label i, is the center point represented by the expression feature of the label i recorded by the anchor; Steps 1-4: Denote the next label as the current label, repeat Steps 2 and 3 until the last label, record and output the training confidence scores of each label and the expression features of each label. The expression feature distribution compactness evaluation unit evaluates the compactness of the expression feature distribution according to the compactness CP of label i i to evaluate the compactness of the expression feature distribution and calculate the CP i The formula for which is: Where N i represents the number of sample points in label i, represents the center point of label i, represents the nth i sample point in label i, d(.) represents the Euclidean distance between two points, and CP i represents the compactness of the expression feature distribution of label i. The expression feature distribution discreteness evaluation unit evaluates the discreteness of the expression feature distribution between label i and label j according to DT ij to evaluate the discreteness of the expression feature distribution and calculate DT ij The formula for which is: In the formula, respectively represent the center points of label i and label j, and DT ij represents the discreteness of the expression feature distribution between label i and label j.
2. The multi-label image recognition system based on a saliency map according to claim 1, wherein: Among them, The multi-label classification loss part and the multi-label distribution loss part calculate the multi-label distribution loss and the multi-label classification loss in a supervised manner according to the training confidence scores of each label and the distribution of the expression features of each label in the high-dimensional space, and train the feature extraction module and the classification module. The specific calculation formula of the multi-label classification loss function in the multi-label classification loss part is as follows: wherein, and y i respectively represent the confidence score for training and the true result of the classification output, L represents the total number of labels included in the natural scene image dataset, i represents label i, and e represents the natural constant.
3. The multi-label image recognition system based on a saliency map according to claim 1, wherein: Among them, Calculate the multi-label distribution loss function of the multi-label distribution loss part according to the parameters calculated by the expression feature distribution compactness evaluation unit, the expression feature distribution discreteness evaluation unit, and the anchor unit. The calculation formula is as follows: where L′ represents the set of tags included in the natural scene image, CP i represents the compactness of the expression feature distribution of tag i, DT ij represents the discreteness of the expression feature distribution between tag i and tag j Transfer the multi-label distribution loss function L DBI and the multi-label classification loss function loss to the classification module and the feature extraction module, and update and iterate the parameters in the classification module and the feature extraction module through the backpropagation algorithm until the feature extraction module and the classification module converge.
4. The multi-label image recognition system based on a saliency map according to claim 1, wherein: Among them, The saliency map generation part multiplies and sums the corresponding classification weights and the corresponding training feature maps to obtain the saliency maps of each label of the training image. The specific calculation process is as follows: where S i represents the saliency map of label i, represents the classification weight between the k-th dimension of the vector obtained by global average pooling and label i, and A k represents a set of training feature maps corresponding to the k-th dimension.
5. The multi-label image recognition system based on a saliency map according to claim 1, wherein: Among them, The process of the feature selection part for obtaining the feature selection weights and the expression features corresponding to each category is as follows: Select the labeled saliency maps in the saliency maps of each label as the output saliency maps, multiply the output saliency maps as weight matrices with a group of the training feature maps with specific dimensions respectively and perform global summation to obtain an output vector, and use the output vector as the expression features of each label corresponding to each output saliency map. The calculation process is to directly use the values corresponding to each pixel point in the output saliency map as the feature selection weights of the pixel points corresponding to the training feature maps of each label. The specific calculation process is as follows: where z i and S i respectively represent the expression feature and the saliency map of label i, A k represents a set of the training feature maps corresponding to the k-th dimension, and x, y represent the points with coordinates (x, y) of the saliency map and the training feature maps.
6. The multi-label image recognition system based on a saliency map according to claim 1, wherein: Among them, The image preprocessing module collects the image and the labels in the image in the natural scene image dataset, and randomly crops the natural scene image to obtain the training images of the same size.
7. The multi-label image recognition system based on a saliency map according to claim 1, wherein: Among them, The feature extraction module is composed of a convolutional neural network with cross-layer branches.
8. A multi-label image recognition system based on a saliency map according to claim 1, characterized in that: Among them, The classification module includes a linear classifier, which obtains a vector by performing global pooling on the training feature map and obtains the training confidence scores of each label and the classification weights between each dimension of the vector and each label by passing the vector through the linear classifier.
Citation Information
Patent Citations
Multi-label image recognition method and device
CN108133233A