Recognition Method and Device Based on Hierarchical Tag Attention
By marking the target area and multi-layer inclusion relationship labels on the image, the hierarchical label attention neural network model is trained, which solves the problem of low accuracy of hierarchical relationship recognition in traditional pattern recognition, and realizes layer-by-layer recognition and accurate recognition of image information.
Patent Information
- Application Number
- CN202111382923.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-22
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-11-22
AI Technical Summary
Traditional pattern recognition tasks mainly use parallel labels, and they cannot adapt to different categories of recognition tasks with hierarchical relationships, resulting in low recognition accuracy.
Using a recognition method based on hierarchical label attention, the target area annotation and multi-layer inclusion relationship label annotation on the original image is trained to achieve layer-by-layer recognition of image information.
It improves the accuracy of image recognition, can effectively process non-parallel labels between different categories, and achieves accurate identification of multi-level labels.
Smart Images

Figure CN114241231B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer vision pattern recognition, and particularly to a recognition method and device based on hierarchical label attention. Background Art
[0002] Pattern recognition is an important branch in the field of computer vision and is widely applied in fields such as vehicle recognition, species recognition, video surveillance, and face recognition. With the development of science and technology, deep learning technology has been widely used in pattern recognition tasks, which has completed the transformation from manually extracting features to automatically extracting features using neural networks, greatly improving the speed and accuracy of feature extraction, and making pattern recognition based on deep learning the mainstream solution in the research of the image processing field.
[0003] However, traditional pattern recognition tasks mainly use parallel labels, and different categories are processed equally, which is not suitable for processing specific hierarchical relationship labels.
[0004] Application Content
[0005] The purpose of this application is to solve the technical problem that traditional recognition task methods mainly classify parallel labels and different categories are not suitable for recognition tasks with hierarchical relationship labels in an equal manner. To achieve the above purpose, this application provides a recognition method and device based on hierarchical label attention.
[0006] This application provides a recognition method based on hierarchical label attention, including:
[0007] Obtain a plurality of original images;
[0008] Annotate the target regions of each original image in the plurality of original images to form a detection training set, and based on the detection training set, train and form a target position detection model;
[0009] According to the target position detection model, divide each original image in the plurality of original images to form a plurality of target region images;
[0010] Annotate at least two layers of inclusion relationship labels for each target region image in the plurality of target region images to form a recognition training set, and based on the recognition training set, train and form a hierarchical label attention neural network model;
[0011] According to the target position detection model and the hierarchical label attention neural network model, perform detection and recognition on the image to be detected.
[0012] This application provides a recognition device based on hierarchical label attention, including:
[0013] An image acquisition module, configured to obtain a plurality of original images;
[0014] A target position detection model generation module, configured to label the target regions of each of the multiple original images to form a detection training set, and train a target position detection model based on the detection training set;
[0015] A target region image generation module, configured to divide each of the multiple original images according to the target position detection model to form a plurality of target region images;
[0016] A hierarchical label attention neural network model generation module, configured to label at least two layers of inclusion relationship labels for each of the multiple target region images to form a recognition training set, and train a hierarchical label attention neural network model based on the recognition training set;
[0017] A detection and recognition module, configured to perform detection and recognition on a to-be-detected image according to the target position detection model and the hierarchical label attention neural network model.
[0018] In the above recognition method based on hierarchical label attention, through the recognition method based on hierarchical label attention proposed in this application, the target regions of the original images are extracted, reducing the interference of other non-target region parts, which is beneficial to further subdivide the multi-layer hierarchical relationships of the target region images. Furthermore, on the basis of the target region images, further modeling of the multi-layer hierarchical relationships can be performed to achieve layer-by-layer recognition of image information and handle non-parallel labels in the recognition task. Therefore, through the recognition method based on hierarchical label attention proposed in this application, by utilizing the guiding role of the inclusion relationships of different-level labels, accurate recognition of image information with multi-level labels can be achieved, improving the recognition accuracy. Description of the Drawings
[0019] Figure 1 is a flowchart of the steps of the recognition method based on hierarchical label attention provided by this application;
[0020] Figure 2 is a schematic structural diagram of the hierarchical label attention neural network model provided by this application;
[0021] Figure 3 is a schematic diagram of the preset position and application scenario of the camera provided by this application;
[0022] Figure 4 is a schematic diagram of labeling the vehicle face region and the vehicle tail region provided by this application;
[0023] Figure 5 is a schematic diagram of the target position detection model based on the Yolov3 algorithm provided by this application;
[0024] Figure 6 It is a schematic diagram of the hierarchical label of vehicle information provided by this application;
[0025] Figure 7 It is a schematic diagram of the structure of the recognition device based on hierarchical label attention provided by this application. Specific embodiments
[0026] The technical solutions of this application will be further described in detail below with reference to the drawings and embodiments.
[0027] Please refer to Figure 1 , this application provides a recognition method based on hierarchical label attention, including:
[0028] S10. Obtain a plurality of original images;
[0029] S20. Label the target areas of each original image among the plurality of original images to form a plurality of target labeled images. The plurality of target labeled images form a detection training set, and based on the detection training set, a target position detection model is trained;
[0030] S30. Divide each original image according to the target position detection model to form a plurality of target area images;
[0031] S40. Label each target area image among the plurality of target area images with at least two levels of inclusion relationship labels to form a plurality of hierarchical labeled images. The plurality of hierarchical labeled images form a recognition training set, and based on the recognition training set, a hierarchical label attention neural network model is trained;
[0032] S50. Detect and recognize the image to be detected according to the target position detection model and the hierarchical label attention neural network model.
[0033] In S10, the plurality of original images can be vehicle images, person images, object images, etc. The target area of the original image can be a specific area, which is the main image area for recognition. For example: the target area of a vehicle image can be the vehicle's front face or rear end, etc.; the target area of a person image can be the whole body or face of the person, etc.; the target area of an object image can be the whole object or local components, etc.
[0034] In S20, by labeling the target areas of the original images, corresponding feature parameters representing each original image can be obtained. The plurality of target feature parameters can include information such as the center point coordinates, width, and height of the minimum rectangular frame of the front face or rear end corresponding to the vehicle image, or can also include information such as the length and width corresponding to the object image.
[0035] Each original image corresponds to multiple target feature parameters, forming a training set for training a target position detection model. In one embodiment, based on the Yolov3 algorithm, SSD algorithm, or faster-rcnn algorithm, and according to the detection training set, a target position detection model is trained. In one embodiment, the Stochastic Gradient Descent (SGD) algorithm is used to optimize the selected Yolov3 target detection model based on the Darknet-50 backbone network to obtain a target position detection model based on the Yolov3 algorithm.
[0036] In S30, multiple target feature parameters can characterize the main image regions for recognition. By dividing each original image, the corresponding target region image can be obtained, thereby reducing the interference of other perspectives, appearances, and other components, narrowing the recognition region, and further improving the recognition accuracy.
[0037] In S40, each target region image is labeled with at least two layers of labels, which can be understood as labeling each target region image with two, three, four, or five layers of labels. The hierarchical labels show a gradually progressive trend, or can also be understood as presenting a hierarchical label inclusion relationship. For example: when labeling the target region image of a vehicle, the first layer label can be the vehicle brand information, the second layer label can be the vehicle model information, and the third layer label can be the vehicle model year information; when labeling the target region image of an object, the first layer label can be the object type, such as objects like animals, plants, or tools, and the second layer label can be the object variety, for example, animals include cats, dogs, or rabbits, and plants include flowers, trees, or grass. In one embodiment, based on the classification model of the convolutional neural network ResNet-50, and according to the recognition training set, a hierarchical label attention neural network model is trained.
[0038] In S50, according to the target position detection model, the position of the target region image in the image to be detected can be detected, and further cropped to form the target region image. According to the hierarchical label attention neural network model, the target region image can be recognized. For example: according to the target position detection model and the hierarchical label attention neural network model, the vehicle brand information, vehicle model information, and vehicle model year information of the vehicle image can be recognized; according to the target position detection model and the hierarchical label attention neural network model, the type information of different hierarchical divisions of the object image can be recognized.
[0039] Through the recognition method based on hierarchical label attention proposed in this application, the target regions of the original image are extracted, reducing the interference of other non-target region parts, which is beneficial to further subdivide the multi-layer hierarchical relationships of the target region images. Furthermore, based on the target region images, multi-layer hierarchical relationships are further modeled, enabling the layer-by-layer recognition of image information and realizing the processing of non-parallel labels between different categories. Thus, through the recognition method based on hierarchical label attention proposed in this application, by utilizing the guiding effect of the relationships between different hierarchical labels, the accurate recognition of image information with multi-layer hierarchical labels can be achieved, improving the recognition accuracy.
[0040] Please refer to Figure 2 , in one embodiment, S40, annotate each target region image with at least two layers of inclusion relationship labels to form a recognition training set, and based on the recognition training set, train and form a hierarchical label attention neural network model, including:
[0041] S410, use a residual network to extract features from multiple target region images to form a first-layer label feature branch and a second-layer label feature branch;
[0042] S420, use a classification function on the first-layer label feature branch to obtain the probabilities of each category corresponding to the first-layer label of each target region image and the recognition result of the first-layer label, and calculate the first cross-entropy loss function during model training; where the first cross-entropy loss function is:
[0043]
[0044] N represents the number of multiple target region images, represents the category corresponding to the first-layer label of the i-th target region image, represents the probabilities of each category corresponding to the first-layer label of the i-th target region image;
[0045] S430, fuse the probabilities of each category corresponding to the first-layer label of each target region image onto the second-layer label feature branch to form a new second-layer label feature branch;
[0046] S440, use a classification function on the new second-layer label feature branch to obtain the probabilities of each category corresponding to the second-layer label under the first-layer label of each target region image and the recognition result of the second-layer label, and calculate the second cross-entropy loss function during model training; where the second cross-entropy loss function is:
[0047]
[0048] Represents the category corresponding to the second-level label under the first-level label corresponding to the i-th target region image. Represents the probabilities of each category corresponding to the second-level label under the first-level label corresponding to the i-th target region image.
[0049] S450. According to the first cross-entropy loss function and the second cross-entropy loss function, calculate the first total loss function, and when the first total loss function reaches the equilibrium state, form a stable hierarchical label attention neural network model. Among them, the first total loss function is:
[0050] L total1 = αL1 + βL2; α and β are the loss balance factors corresponding to the first-level label feature branch and the new second-level label feature branch respectively.
[0051] In S410, the residual network can be ResNet-50, which can also be called the backbone network of the hierarchical label attention neural network model HLANet. After the target region image extracts the feature f through ResNet-50, multiple feature branches are output, which can be used to identify multi-level labels such as the first-level label, the second-level label, the third-level label, and the fourth-level label respectively.
[0052] In S420, on the first-level label feature branch, it can also be understood as adding a classification function to the feature f1 of the first-level label. The classification function includes the softmax function, etc. Adding the softmax function to the feature f1 of the first-level label, outputs the probabilities p1 of each category corresponding to the first-level label, and obtains the corresponding recognition result. Further, during training, use as the supervision to calculate the first cross-entropy loss function on the first-level label feature branch. Among them, is the one-hot encoding of the first-level label. For example, [1, 0, 0] represents the first category corresponding to the first-level label of the i-th target region image; [0, 1, 0] represents the second category corresponding to the first-level label of the i-th target region image; [0, 0, 1] represents the third category corresponding to the first-level label of the i-th target region image, and so on.
[0053] In S430, on the second-level label feature branch, it can also be understood as adopting an attention mechanism on the feature f2 of the second-level label, so that the information corresponding to the predicted first-level label is fused into the feature for second-level label recognition, forming a new second-level label feature branch to guide the recognition of the second-level label.
[0054] In one embodiment, S430 includes:
[0055] S431. On the second - layer label feature branch, expand the probabilities of each category corresponding to the first - layer label of the \(i\) - th target region image to the dimension corresponding to the feature of the second - layer label feature branch to obtain the first attention weight
[0056] S432. According to the first attention weight calculate the new feature of the second - layer label feature branch to form a new second - layer label feature branch. Among them, the new feature of the second - layer label feature branch is:
[0057]
[0058] represents the probabilities of each category corresponding to the first - layer label of the \(i\) - th target region image, represents the first attention weight, represents the feature of the second - layer label feature branch, and \(o\) represents the dot - product operation of two vectors, which can also be understood as the Hadamard product.
[0059] In S431, expand to so that the probabilities of each category obtained on the first - layer label feature branch are applicable to the second - layer label feature branch. Among them, \(x\) represents the probability value, \(C0\) represents the number of categories of the first - layer label. The number of \(x1,x1,\cdots,x1\) is \(m1\), the number of \(x2,x2,\cdots,x2\) is \(m2\), and the number of \(x c0 ,x c0 ,\cdots,x c0 is \(m c0 which respectively represent the quantities corresponding to the 1st category, the 2nd category, and the \(C0\) - th category of the second - layer label under the first - layer label.
[0060] In S432, through calculation and fusion, obtain the new feature of the second - layer label feature branch, forming a new second - layer label feature branch. Furthermore, continue to run the classification function on the new second - layer label feature branch.
[0061] In S440, the first - layer label contains the second - layer label, and the two are in an inclusion relationship. On the new second - layer label feature branch, it can also be understood as adding a classification function to the new feature of the second - layer label feature branch. The classification function includes the softmax function, etc. Add the softmax function to the new feature of the second - layer label feature branch to output the probabilities of each category corresponding to the second - layer label and obtain the corresponding recognition results. Further, during training, use as supervision to calculate the second cross-entropy loss function on the second layer label feature branch. Among them, represents the one-hot encoding of the second layer label. For example, [1, 0, 0] represents the first category corresponding to the second layer label; [0, 1, 0] represents the second category corresponding to the second layer label; [0, 0, 1] represents the third category corresponding to the second layer label, and so on.
[0062] In S450, after the calculations on the first layer label feature branch and the second layer label feature branch, the first total loss function of the hierarchical label attention neural network model is obtained. During the training process, the value of the first total loss function continuously decreases and finally tends to a balanced state. When the first total loss function reaches the balanced state, the network in the hierarchical label attention neural network model reaches the optimal, forming a stable hierarchical label attention neural network model.
[0063] Through S410 to S450, a hierarchical label attention neural network model with two layers of labels is formed, realizing the modeling of the two-layer hierarchical relationship, and then the layer-by-layer recognition of image information can be realized, and the processing of non-parallel labels between different categories is realized. Therefore, through the recognition method based on hierarchical label attention proposed in this application, the first layer label has a guiding behavior for the recognition and classification of the second layer label, and the accurate recognition of the image information of the two layers of labels can be realized, improving the recognition accuracy.
[0064] In one embodiment, in S410, the residual network is used to extract features from multiple target region images, and a third layer label feature branch is also formed. There is an inclusion and being-included relationship between the first layer label, the second layer label, and the third layer label. The first layer label includes the second layer label, and the second layer label includes the third layer label. Furthermore, S40 also includes:
[0065] S460, fuse the probabilities of each category corresponding to the second layer label under the first layer label corresponding to each target region image onto the third layer label feature branch to form a new third layer label feature branch;
[0066] S470, run the classification function on the new third layer label feature branch to obtain the probabilities of each category corresponding to the third layer label under the second layer label corresponding to each target region image and the recognition result of the third layer label. During the training process, calculate the third cross-entropy loss function; where the third cross-entropy loss function is:
[0067]
[0068] Represents the category corresponding to the third - layer label under the second - layer label corresponding to the i - th target region image. Represents the probabilities of each category corresponding to the third - layer label under the second - layer label corresponding to the i - th target region image.
[0069] S480, according to the first cross - entropy loss function, the second cross - entropy loss function, and the third cross - entropy loss function, calculate the second total loss function, and when the second total loss function reaches the equilibrium state, form a stable hierarchical label attention neural network model; where the second total loss function is:
[0070] L total2 =αL1 + βL2+λL3; α, β, and λ are the loss balance factors corresponding to the first - layer label feature branch, the new second - layer label feature branch, and the new third - layer label feature branch respectively.
[0071] In S460, on the third - layer label feature branch, which can also be understood as on the feature f3 of the third - layer label, use the attention mechanism to fuse the information corresponding to the predicted second - layer label into the feature recognized by the third - layer label, forming a new third - layer label feature branch to guide the recognition of the third - layer label.
[0072] In one embodiment, S460 includes:
[0073] S461, on the third - layer label feature branch, expand the dimension of the probabilities of each category corresponding to the second - layer label under the first - layer label corresponding to the i - th target region image to the dimension corresponding to the feature f3 of the third - layer label feature branch to obtain the second attention weight
[0074] S462, according to the second attention weight calculate the new feature of the third - layer label feature branch to form a new third - layer label feature branch; where the new feature of the third - layer label feature branch is:
[0075]
[0076] Represents the probabilities of each category corresponding to the second - layer label under the first - layer label corresponding to the i - th target region image, Represents the second attention weight, Represents the feature of the third - layer label feature branch, and o represents the dot - product operation of two vectors, which can also be understood as the Hadamard product.
[0077] In S461, the Extended to So that the probabilities of each category obtained on the second-layer label feature branch Are applicable to the third-layer label feature branch. Among them, C1 represents the number of categories of the second-layer label. The number of n1 of x1, x1,..., x1, the number of n2 of x2, x2,... x2, x c1 , x c1 ,... x c1 The number of n c0 , respectively represent the corresponding quantities under the first category, the second category, and the C1 category corresponding to the third-layer label under the second-layer label.
[0078] In S462, through calculation and fusion, a new feature of the third-layer label feature branch is obtained A new third-layer label feature branch is formed. Furthermore, the classification function is continued to run on the new third-layer label feature branch.
[0079] In S470, on the new third-layer label feature branch, it can also be understood as on the new feature of the third-layer label feature branch Add a classification function. On the new feature of the third-layer label feature branch Add a softmax function to output the probabilities of each category corresponding to the third-layer label And obtain the corresponding recognition result. Further, during training, use As supervision, calculate the third cross-entropy loss function on the third-layer label feature branch. Among them, Is represented as a one-hot encoding. For example, [1, 0, 0] represents the first category corresponding to the third-layer label; [0, 1, 0] represents the second category corresponding to the third-layer label; [0, 0, 1] represents the third category corresponding to the third-layer label, and so on.
[0080] In S480, after the calculations on the first-layer label feature branch, the second-layer label feature branch, and the third-layer label feature branch, the second total loss function of the hierarchical label attention neural network model is obtained. During the training process, the second total loss function will continuously decrease and finally tend to a balanced state. When the second total loss function reaches the balanced state, the network in the hierarchical label attention neural network model reaches the optimal, forming a stable hierarchical label attention neural network model.
[0081] Through S460 to S480, a hierarchical label attention neural network model with three layers of labels is formed, which realizes the modeling of the three-layer hierarchical relationship, and then can realize the layer-by-layer detection and recognition of image information, and realizes the processing of non-parallel labels between different categories. Therefore, through the recognition method based on hierarchical label attention proposed in this application, the three-layer labels have a guiding behavior for recognition and classification, and can accurately recognize the image information of the three-layer labels, improving the recognition accuracy.
[0082] In one embodiment, the multiple original images are multiple original vehicle images, and the multiple target feature parameters include the center point coordinates, width, and height information of the minimum rectangular box corresponding to the vehicle face or the vehicle tail in the original vehicle image. The first-layer label is the vehicle brand information, the second-layer label is the vehicle model information, and the third-layer label is the vehicle model year information.
[0083] Taking the recognition of vehicle information as an example according to the recognition method based on hierarchical label attention, the recognition accuracy of the hierarchical labels of vehicle information is improved.
[0084] Please refer to Figure 3 , in S10, by presetting the position of the monitoring camera, the original vehicle images are collected. In a specific application scenario, set the position of the monitoring camera so that it can capture the vehicle face or the vehicle tail part from the front view, and further obtain the image data containing the vehicle target in the monitoring camera. By presetting the position of the monitoring camera, it can capture the vehicle from the front view and obtain the front orientation of the vehicle face or the vehicle tail, so as to reduce the influence of different perspectives on the recognition of vehicle information. Furthermore, the image data containing the vehicle target is obtained in the monitoring camera.
[0085] Please refer to Figure 4 , in S20, the positions of the vehicle face and the vehicle tail are marked on the obtained original vehicle images. Since the vehicle information can be represented by the vehicle face and the vehicle tail in the vehicle components. In order to reduce the interference factors of vehicle perspective, appearance and other components, in this step, the positions of the vehicle and the vehicle tail in the original vehicle image are marked. Marking the positions of the vehicle and the vehicle tail in the original vehicle image includes: drawing the minimum rectangular box containing the vehicle face or the vehicle tail in the original vehicle image; recording the center point coordinates and the width and height information of the minimum rectangular box to obtain multiple target feature parameters. The multiple target feature parameters include the target category (vehicle face or vehicle tail), the center point coordinates of the minimum rectangular box, and the width and height information. By marking the target area of the original image, the accuracy of recognizing vehicle information is further improved, making it more targeted.
[0086] In S20, each original vehicle image corresponds to target feature parameters including the target category (front or rear of the vehicle), the center point coordinates of the minimum rectangle, and the width and height information. Multiple original vehicle images form multiple target feature parameters, and then a training set for a detection model targeting the front and rear of the vehicle can be formed. Based on the detection training set formed by the annotation, a target position detection model based on the Yolov3 algorithm is trained, which can also be understood as a detection model for the front and rear of the vehicle.
[0087] Please refer to Figure 5 , in S20, the backbone network Darknet-50 of the Yolov3 algorithm uses fully convolutional layers and residual structures to extract the image features of the original vehicle images. Each convolutional layer includes three operations: two-dimensional convolution, normalization, and classification function. A single-stage target detection model for detecting original vehicle images of different scales is achieved by fusing features of different resolutions and different semantic intensities through a feature pyramid. In one embodiment, the size of the original vehicle image obtained by the camera can be 414×416 pixels, which is scaled and adjusted to 416×416 pixels as the input image size. After the original vehicle image passes through the backbone network Darknet-50 of the Yolov3 algorithm, features of three scales of 13×13, 26×26, and 52×52 are generated, and then the rectangular position of the minimum rectangle of the target position detection model, the target confidence, and the target category (front or rear of the vehicle) are output. A detection training set is formed according to the generated multiple feature parameters, and the target position detection model is trained using the gradient descent algorithm, and finally the target position detection model is formed.
[0088] In S30, the target region image can be understood as an image containing target detection information, which is an image formed by cropping the original vehicle image. By cropping the original vehicle image, the detection position can be further locked, making the detection position more targeted, and thus the recognition accuracy can be improved.
[0089] Please refer to Figure 6 , before S40, that is, before annotating three layers of labels from coarse to fine for the front or rear of the vehicle area, first determine the labels of the vehicle information and their hierarchical relationships. For example:
[0090] The vehicle information includes vehicle brand information, vehicle model information, and vehicle model year information. The brand information of the vehicle information that can be obtained through the vehicle sales website includes 264 types such as Hongqi, Volkswagen, Ford, Suzuki, Honda, Toyota, Citroën, Peugeot, Mazda, Ford, etc. The vehicle brand information belongs to the first-level label. There are 3,904 categories of vehicle model information collected, which belongs to the second-level label. For example, the vehicle model information (which can be understood as the second-level label) included in the Hongqi brand (which can be understood as the first-level label) includes L5, H5, H7, L7, L9, E-HS3, Century Star, Hongqi Shengshi, etc. There are 14,530 categories of vehicle model year information collected, which belongs to the third-level label. For example, the vehicle model year information (which can be understood as the third-level label) included in the L5 (which can be understood as the second-level label) included in the Hongqi brand (which can be understood as the first-level label) includes two categories: 2014 model and 2019 model.
[0091] In S40, the above three-level labels (such as vehicle brand information, vehicle model information, and vehicle model year information) are marked from coarse to fine for the vehicle front area or the vehicle rear area to construct an identification training set. Design a neural network HLANet with a hierarchical label attention mechanism, and train the neural network HLANet according to the identification training set, which can also be understood as training a hierarchical label attention neural network model.
[0092] The three-level labels are marked from coarse to fine for the vehicle front area or the vehicle rear area to form a vehicle information identification training data set. The first-level label is the vehicle brand information, the second-level label is the vehicle model information, and the third-level label is the vehicle model year information for the training of the vehicle information identification model.
[0093] Please refer to Figure 2 , the Residual Network ResNet can be used to solve the problem of gradient disappearance and network performance degradation when the data of the deep neural network layer continuously increases. The superimposed layer of the shallow network model and its own mapping are connected together through residual units, and the input information is passed across layers in a shortcut manner and then added to the output after convolution to achieve the effect of fully training the underlying network. ResNet-50 is a residual network with 50 parameter layers, with the characteristics of high precision and fast speed.
[0094] After the input target region image extracts the feature f through the backbone network ResNet-50, three feature branches are output, namely the first-level label feature branch, the second-level label feature branch, and the third-level label feature branch, which can be used to identify vehicle brand information, vehicle model information, and vehicle model year information respectively. The three-level labels are divided into three stages for identification. It can also be understood that after the target region image extracts the feature f through the backbone network ResNet-50, three feature branches are output, namely the vehicle brand information feature branch, the vehicle model information feature branch, and the vehicle model year information feature branch.
[0095] Add a softmax function to the feature f1 for vehicle brand information recognition, output the probability p1 of each brand, and obtain the recognition results of each brand. For example: the probability of the Hongqi brand is 80%, the probability of the Volkswagen brand is 15%, etc. During training, use the category y1 of the vehicle brand as supervision, and calculate the cross-entropy loss function on the vehicle brand branch as:
[0096]
[0097] Among them, represents the category corresponding to the vehicle brand label of the i-th target region image. For example: the first category is the Hongqi brand, the second category is the Volkswagen brand, etc. represents the probability of each brand predicted and output for the i-th target region image. For example: the probability of the Hongqi brand is 80%, the probability of the Volkswagen brand is 15%, etc.
[0098] On the feature f2 for vehicle model information recognition, adopt an attention mechanism to fuse the predicted vehicle brand information onto the feature for vehicle model information recognition to guide the recognition of vehicle model information. For example, for the i-th target region image, expand the predicted probability of the vehicle brand into For example: expand into
[0099] Use which has the same dimension as the vehicle model feature to calculate the fused vehicle model feature, that is, the new feature corresponding to the vehicle model feature branch as follows:
[0100] Furthermore, adopt a softmax function on the feature to output the probability of each vehicle model information and obtain the recognition results of vehicle model information. For example: the probability corresponding to the L5 model under the Hongqi brand is 85%, the probability corresponding to the H7 model under the Hongqi brand is 7%, etc. During training, use the category y2 of the vehicle model as supervision, and calculate the cross-entropy loss function on the vehicle model branch as:
[0101]
[0102] Among them, represents the category corresponding to the vehicle model label of the i-th target region image. For example: the first category is the L5 model under the Hongqi brand, the second category is the H7 model under the Hongqi brand, etc. Indicates the probabilities of each model to which the prediction output of the i-th target region image belongs. For example: the probability corresponding to the L5 model under the Hongqi brand is 85%, and the probability corresponding to the H7 model under the Hongqi brand is 7%, etc.
[0103] On the feature f3 for vehicle model year information recognition, the attention mechanism is adopted to fuse the vehicle model information onto the feature for vehicle model year information recognition to guide the recognition of vehicle model year information. For example, for the i-th target region image, the probability of the predicted vehicle model is expanded to For example: Expand to
[0104] Use which has the same dimension as the feature on the vehicle model year branch. Calculate the fused vehicle model year feature, that is, the new feature corresponding to the vehicle model year feature branch as follows:
[0105] Furthermore, on the feature the softmax function is adopted to output the probabilities of each vehicle model year and obtain the recognition result of the vehicle model year information. For example: the probability corresponding to the 2014 model year of the L5 model under the Hongqi brand is 56%, and the probability corresponding to the 2019 model year of the L5 model under the Hongqi brand is 15%, etc. During training, use the category y3 of the vehicle model year as the supervision, and calculate the cross-entropy loss function on the vehicle model year branch as:
[0106]
[0107] where, represents the category corresponding to the vehicle model year label of the i-th target region image. For example: the first category is the 2014 model year of the L5 model under the Hongqi brand, and the second category is the 2019 model year of the L5 model under the Hongqi brand. represents the probabilities of each model year to which the prediction output of the i-th target region image belongs. For example: the probability corresponding to the 2014 model year of the L5 model under the Hongqi brand is 89%, and the probability corresponding to the 2019 model year of the L5 model under the Hongqi brand is 11%, etc.
[0108] After the above three-level calculations, the total loss function of the hierarchical label attention neural network model is:
[0109] L total = αL1 + βL2 + λL3.
[0110] Among them, α, β, and λ are the loss balance factors corresponding to the vehicle brand, the loss balance factor corresponding to the vehicle model, and the loss balance factor corresponding to the vehicle model year, respectively, which balance the three branches during training and do not favor a certain branch, and can be set according to actual experience.
[0111] When the total loss function reaches the equilibrium state, a stable hierarchical label attention neural network model is formed to achieve accurate identification of vehicle information.
[0112] In one embodiment, during the training stage of the ResNet-50 hierarchical label attention vehicle information recognition network HLANet, the labeled vehicle information (such as vehicle brand, vehicle model, and vehicle model year information) is used as the training data set, and the vehicle three-level information y1, y2, and y3 is used as the label for network supervised training. The gradient descent algorithm is used to obtain the parameters on each feature branch in the HLANet network.
[0113] In one embodiment, since there are many categories and it is difficult to identify the model year when labeling the vehicle information training set. Therefore, there will be data without labeled model year in the training set. Furthermore, during the training process, for the data without labeled model year, the error calculation for model year information recognition is not performed, only the errors in vehicle brand and vehicle model recognition are calculated, and feedback propagation is performed, and the parameters on the vehicle model year branch in the HLANet network are not updated.
[0114] In one embodiment, to address the problem of data imbalance in different categories of the training set, when training the HLANet network, resampling is performed after each round of training. When resampling, a new training set can be formed by removing the recognized samples and retaining the incorrect and error-prone samples for the next round of training or fine-tuning, thereby enabling the rapid convergence of the network and effectively improving the accuracy of vehicle information recognition.
[0115] Therefore, in the first stage, the softmax function is adopted on the vehicle brand recognition branch to output the probabilities of belonging to each vehicle brand, and the vehicle brand recognition result is obtained. In the second stage, the probabilities of belonging to each vehicle brand in the first stage are fused with the features of the vehicle model recognition branch. The probabilities of the vehicle brands predicted by the network are used as the coefficients of the feature attention mechanism of the vehicle model branch, so that the recognition result of the vehicle model tends to be the vehicle model under the vehicle brand that has been recognized in the first stage. Furthermore, the features of the vehicle model branch processed by the attention mechanism are input into the softmax function to output the probabilities of belonging to each vehicle model, and the vehicle model recognition result is obtained. In the third stage, the probabilities of the vehicle models belonging to the second stage are fused with the features of the vehicle model year recognition branch. The fused features of the vehicle model year branch are input into the softmax function to output the probabilities of belonging to each vehicle model year and obtain the vehicle model year recognition result.
[0116] In one embodiment, in S50, according to the target position detection model and the hierarchical label attention neural network model, the image to be detected is detected and recognized, including:
[0117] S510, obtain the image of the vehicle to be detected;
[0118] S520, input the image of the vehicle to be detected into the target position detection model, output a plurality of target feature parameters corresponding to the image of the vehicle to be detected, and crop the image of the vehicle to be detected according to the plurality of target feature parameters to form an image of the target area of the vehicle to be detected;
[0119] S530, input the image of the target area of the vehicle to be detected into the hierarchical label attention neural network model, and output vehicle brand information, vehicle model information, and vehicle model year information.
[0120] According to the trained target position detection model and the hierarchical label attention neural network model, the vehicle information of the image of the vehicle to be detected is recognized. According to the target position detection model, the front face area or the rear face area of the image of the vehicle to be detected is detected, and the position coordinates of the front face or the rear face contained in the image of the vehicle to be detected are output. According to the position coordinates of the front face or the rear face, the image of the vehicle to be detected is cropped to generate an image of the target area, which can also be understood as an image of the front face or the rear face.
[0121] The image of the front face or the rear face is input into the hierarchical label attention neural network model for vehicle information recognition, and the corresponding vehicle brand, vehicle model, and vehicle model year information are output, thus realizing the accurate recognition of vehicle information.
[0122] In one embodiment, the recognition method based on hierarchical label attention can also be applied to the recognition of object images. The modeling principle processes of the target position detection model and the hierarchical label attention neural network model are the same as those in the above embodiments. When the recognition method based on hierarchical label attention is used to recognize object images, the first-level label is the object category, such as objects like animals, plants, or tools. The second-level label is the object variety. For example, animals include cats, dogs, or rabbits, etc., and plants include flowers, trees, or grass, etc. The inclusion relationships between multiple levels of labels can all be used for the recognition method based on hierarchical label attention provided in this application for recognition.
[0123] Please refer to Figure 7 , in one embodiment, the present application provides a recognition device 100 based on hierarchical label attention, which includes an image acquisition module 10, a target position detection model generation module 20, a target region image generation module 30, a hierarchical label attention neural network model generation module 40, and a detection and recognition module 50. The image acquisition module 10 is used to acquire a plurality of original images. The target position detection model generation module 20 is used to label the target regions of each of the original images to form a detection training set, and based on the detection training set, train to form a target position detection model. The target region image generation module 30 is used to divide each of the original images according to the target position detection model to form a plurality of target region images. The hierarchical label attention neural network model generation module 40 is used to label at least two levels of inclusion relationship labels for each of the target region images to form a recognition training set, and based on the recognition training set, train to form a hierarchical label attention neural network model. The detection and recognition module 50 is used to perform detection and recognition on the image to be detected according to the target position detection model and the hierarchical label attention neural network model.
[0124] The relevant descriptions in this embodiment can refer to the descriptions in the above method step embodiments.
[0125] In one embodiment, the hierarchical label attention neural network model generation module 40 includes a label feature branch generation module (not shown in the figure), a first-level label feature branch operation module (not shown in the figure), a new second-level label feature branch generation module (not shown in the figure), a new second-level label feature branch operation module (not shown in the figure), and a first loss function adjustment module (not shown in the figure).
[0126] The label feature branch generation module is used to extract features from the multiple target region images by using a residual network to form a first-layer label feature branch and a second-layer label feature branch. The first-layer label feature branch running module is used to run a classification function on the first-layer label feature branch to obtain the probabilities of each category corresponding to the first-layer label of each target region image and the first-layer label recognition result, and calculate the first cross-entropy loss function during model training. Wherein, the first cross-entropy loss function is:
[0127]
[0128] N represents the number of the multiple target region images, represents the category corresponding to the first-layer label of the i-th target region image, represents the probabilities of each category corresponding to the first-layer label of the i-th target region image.
[0129] The new second-layer label feature branch generation module is used to fuse the probabilities of each category corresponding to the first-layer label of each target region image onto the second-layer label feature branch to form a new second-layer label feature branch.
[0130] The new second-layer label feature branch running module is used to run a classification function on the new second-layer label feature branch to obtain the probabilities of each category corresponding to the second-layer label under the first-layer label of each target region image and the second-layer label recognition result, and calculate the second cross-entropy loss function during model training. Wherein, the second cross-entropy loss function is:
[0131]
[0132] represents the category corresponding to the second-layer label under the first-layer label of the i-th target region image, represents the probabilities of each category corresponding to the second-layer label under the first-layer label of the i-th target region image.
[0133] The first loss function adjustment module is used to calculate the first total loss function according to the first cross-entropy loss function and the second cross-entropy loss function, and form a stable hierarchical label attention neural network model when the first total loss function reaches a balanced state. Wherein, the first total loss function is:
[0134] L total1 = αL1 + βL2.
[0135] α and β are respectively the loss balance factor corresponding to the first-layer label feature branch and the loss balance factor corresponding to the new second-layer label feature branch, L1 is the first cross-entropy loss function, and L2 is the second cross-entropy loss function.
[0136] For the relevant descriptions in this embodiment, reference can be made to the descriptions in the above method step embodiment.
[0137] In one embodiment, the label feature branch generation module is further configured to form a third-layer label feature branch (not marked in the figure). The hierarchical label attention neural network model generation module further includes a new third-layer label feature branch generation module (not marked in the figure), a new third-layer label feature branch operation module (not marked in the figure), and a second loss function adjustment module (not marked in the figure).
[0138] The new third-layer label feature branch generation module is configured to fuse the probabilities of each category corresponding to the second-layer label under the first-layer label corresponding to each target region image onto the third-layer label feature branch to form a new third-layer label feature branch.
[0139] The new third-layer label feature branch operation module is configured to run a classification function on the new third-layer label feature branch to obtain the probabilities of each category corresponding to the third-layer label under the second-layer label corresponding to each target region image and the third-layer label recognition result, and calculate the third cross-entropy loss function during model training. Among them, the third cross-entropy loss function is:
[0140]
[0141] represents the category corresponding to the third-layer label under the second-layer label corresponding to the i-th target region image, represents the probabilities of each category corresponding to the third-layer label under the second-layer label corresponding to the i-th target region image.
[0142] The second loss function adjustment module is configured to calculate the second total loss function according to the first cross-entropy loss function, the second cross-entropy loss function, and the third cross-entropy loss function, and form a stable hierarchical label attention neural network model when the second total loss function reaches a balanced state. Among them, the second total loss function is:
[0143] L total2 = αL1 + βL2 + λL3;
[0144] α, β, and λ are the corresponding loss balance factors on the first-layer label feature branch, the corresponding loss balance factors on the new second-layer label feature branch, and the corresponding loss balance factors on the new third-layer label feature branch respectively. L1 is the first cross-entropy loss function, L2 is the second cross-entropy loss function, and L3 is the third cross-entropy loss function.
[0145] For the relevant descriptions in this embodiment, reference can be made to the descriptions in the above method step embodiments.
[0146] In one embodiment, the new second-layer label feature branch generation module includes a first attention weight generation module (not marked in the figure) and a new feature generation module for the second-layer label feature branch (not marked in the figure).
[0147] The first attention weight generation module is used to expand the dimension of the probabilities of each category corresponding to the first-layer label corresponding to the i-th target region image on the second-layer label feature branch to the dimension corresponding to the feature of the second-layer label feature branch, and obtain the first attention weight.
[0148] The new feature generation module for the second-layer label feature branch is used to calculate the new feature of the second-layer label feature branch according to the first attention weight, and form the new second-layer label feature branch.
[0149] Among them, the new feature of the second-layer label feature branch is:
[0150]
[0151] represents the probabilities of each category corresponding to the first-layer label corresponding to the i-th target region image, represents the first attention weight, represents the feature of the second-layer label feature branch, and o represents the dot product operation of two vectors, which can also be understood as the Hadamard product.
[0152] For the relevant descriptions in this embodiment, reference can be made to the descriptions in the above method step embodiments.
[0153] In one embodiment, the new third-layer label feature branch generation module includes a second attention weight generation module (not marked in the figure) and a new feature generation module for the third-layer label feature branch (not marked in the figure).
[0154] The second attention weight generation module is used to expand the dimension of the probabilities of each category corresponding to the second-layer label corresponding to the first-layer label corresponding to the i-th target region image on the third-layer label feature branch to the dimension corresponding to the feature of the third-layer label feature branch, and obtain the second attention weight.
[0155] The new feature generation module of the third-layer label feature branch is used to calculate the new features of the third-layer label feature branch according to the second attention weight, and form the new third-layer label feature branch.
[0156] Among them, the new features of the third-layer label feature branch are:
[0157]
[0158] represents the probabilities of each category corresponding to the second-layer label under the first-layer label corresponding to the i-th target region image, represents the second attention weight, represents the features of the third-layer label feature branch, and o represents the dot product operation of two vectors, which can also be understood as the Hadamard product.
[0159] The relevant descriptions in this embodiment can refer to the descriptions in the above method step embodiments.
[0160] In the above various embodiments, the specific order or hierarchy of the steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of the steps in the process can be rearranged without departing from the protection scope of the present disclosure. The appended method claims present the elements of various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy.
[0161] Those skilled in the art can also understand that the various illustrative logical blocks, units, and steps listed in the embodiments of the present application can be implemented by electronic hardware, computer software, or a combination of both. To clearly show the interchangeability of hardware and software, the above various illustrative components, units, and steps have generally described their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art can use various methods to implement the described functions for each specific application, but such implementation should not be understood as exceeding the protection scope of the embodiments of the present application.
[0162] In the embodiments of the present application, the various illustrative logical blocks or units can be implemented or operate the described functions through a general-purpose processor, a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of the above designs. The general-purpose processor can be a microprocessor. Optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors combined with a digital signal processor core, or any other similar configuration.
[0163] The steps of the methods or algorithms described in the embodiments of the present application can be directly embedded in hardware, software modules executed by a processor, or a combination of both. The software modules can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be disposed in an ASIC, and the ASIC can be disposed in a user terminal. Optionally, the processor and the storage medium can also be disposed in different components of the user terminal.
[0164] The specific embodiments described above further elaborate on the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only the specific embodiments of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A recognition method based on hierarchical label attention, characterized in that, Including: Obtain a plurality of original images; Annotate the target regions of each of the plurality of original images to form a detection training set, and based on the detection training set, train to form a target position detection model; According to the target position detection model, divide each of the plurality of original images to form a plurality of target region images; Annotate each of the plurality of target region images with at least two layers of inclusion relationship labels to form a recognition training set, and based on the recognition training set, train to form a hierarchical label attention neural network model; The training formation step of the hierarchical label attention neural network model includes: Use a residual network to extract features from each of the plurality of target region images to form a first-layer label feature branch and a second-layer label feature branch; Run a classification function on the first-layer label feature branch to obtain the probabilities of each category corresponding to the first-layer label of each target region image and the first-layer label recognition result, and calculate the first cross-entropy loss function during model training; Fuse the probabilities of each category corresponding to the first-layer label of each target region image to the second-layer label feature branch to form a new second-layer label feature branch; Run a classification function on the new second-layer label feature branch to obtain the probabilities of each category corresponding to the second-layer label under the first-layer label of each target region image and the second-layer label recognition result, and calculate the second cross-entropy loss function during model training; Calculate a first total loss function according to the first cross-entropy loss function and the second cross-entropy loss function; The formation step of the new second-layer label feature branch includes: On the second-layer label feature branch, expand the dimension of the probabilities of each category corresponding to the first-layer label of the i-th target region image to the dimension corresponding to the features of the second-layer label feature branch to obtain a first attention weight; According to the first attention weight, calculate the new features of the second-layer label feature branch to form the new second-layer label feature branch; According to the target position detection model and the hierarchical label attention neural network model, perform detection and recognition on the image to be detected.
2. The recognition method based on hierarchical label attention according to claim 1, wherein The step of annotating each of the plurality of target region images with at least two layers of inclusion relationship labels to form a recognition training set, and based on the recognition training set, training to form a hierarchical label attention neural network model includes: The first cross-entropy loss function is as follows: N represents the number of the multiple target region images, represents the category corresponding to the first-layer label of the i-th target region image, represents the probabilities of each category corresponding to the first-layer label of the i-th target region image; the second cross-entropy loss function is as follows: represents the category corresponding to the second-layer label corresponding to the first-layer label of the i-th target region image, represents the probabilities of each category corresponding to the second-layer label corresponding to the first-layer label of the i-th target region image; When the first total loss function reaches an equilibrium state, form a stable hierarchical label attention neural network model; where the first total loss function is: L totall = αL1 + βL2; α and β are respectively the loss balance factor corresponding to the first-layer label feature branch and the loss balance factor corresponding to the new second-layer label feature branch, L1 is the first cross-entropy loss function, and L2 is the second cross-entropy loss function.
3. The recognition method based on hierarchical label attention according to claim 2, wherein Using the residual network to extract features from each of the plurality of target region images, a third-layer label feature branch is also formed; Annotate each of the target region images with at least two layers of inclusion relationship labels to form a recognition training set, and based on the recognition training set, train to form a hierarchical label attention neural network model, further including: Fuse the probabilities of each category corresponding to the second layer label under the first layer label corresponding to each target region image onto the third layer label feature branch to form a new third layer label feature branch; Run a classification function on the new third layer label feature branch to obtain the probabilities of each category corresponding to the third layer label under the second layer label corresponding to each target region image and the third layer label recognition result, and calculate the third cross-entropy loss function during model training; wherein, the third cross-entropy loss function is: Indicates the category corresponding to the third-level label under the second-level label corresponding to the i-th target region image, Indicates the probabilities of each category corresponding to the third-level label under the second-level label corresponding to the i-th target region image; Calculate the second total loss function according to the first cross-entropy loss function, the second cross-entropy loss function, and the third cross-entropy loss function, and when the second total loss function reaches a balanced state, form a stable hierarchical label attention neural network model; wherein, the second total loss function is: L total2 = αL1 + βL2 + λL3; λ is the loss balance factor corresponding to the new third layer label feature branch, and L3 is the third cross-entropy loss function.
4. The recognition method based on hierarchical label attention according to claim 2, wherein Fuse the probabilities of each category corresponding to the first layer label corresponding to each target region image onto the second layer label feature branch to form a new second layer label feature branch, including: wherein, the new feature of the second layer label feature branch is: Represents the probabilities of each category corresponding to the first-layer label corresponding to the i-th target region image, Represents the first attention weight, f2 (i) Represents the feature of the second-layer label feature branch, and o represents the dot product operation of two vectors.
5. The recognition method based on hierarchical label attention according to claim 3, wherein The step of fusing the probabilities of each category corresponding to the second layer label under the first layer label corresponding to each target region image onto the third layer label feature branch to form a new third layer label feature branch includes: On the third layer label feature branch, expand the dimension of the probabilities of each category corresponding to the second layer label under the first layer label corresponding to the i-th target region image to the dimension corresponding to the feature of the third layer label feature branch to obtain a second attention weight; According to the second attention weight, calculate the new feature of the third layer label feature branch to form the new third layer label feature branch; wherein, the new feature of the third layer label feature branch is: Represents the probabilities of each category corresponding to the second-level labels under the first-level label corresponding to the i-th target region image, Represents the second attention weight, f3 (i) Represents the feature of the third-level label feature branch, and o represents the dot product operation of two vectors.
6. The recognition method based on hierarchical label attention according to claim 1, wherein Among the multiple original images obtained, the multiple original images include multiple original vehicle images; In the step of annotating the target regions of each of the multiple original images to form a detection training set, and based on the detection training set, training to form a target position detection model, annotating the target regions of each of the multiple original images to form multiple target feature parameters, the multiple target feature parameters including the center point coordinates and width and height information of the minimum rectangular box corresponding to the vehicle face or vehicle tail of the original vehicle image.
7. The recognition method based on hierarchical label attention according to claim 6, wherein In the step of annotating each target region image among the multiple target region images with at least two layers of labels, the at least two layers of labels include a first layer label, a second layer label, and a third layer label, the first layer label is vehicle brand information, the second layer label is vehicle model information, and the third layer label is vehicle model year information.
8. The recognition method based on hierarchical label attention according to claim 7, characterized in that Performing detection and recognition on the image to be detected according to the target position detection model and the hierarchical label attention neural network model, including: Obtaining an image of a vehicle to be detected; Inputting the image of the vehicle to be detected into the target position detection model, outputting a plurality of target feature parameters corresponding to the image of the vehicle to be detected, and cropping the image of the vehicle to be detected according to the plurality of target feature parameters to form an image of a target area of the vehicle to be detected; Inputting the image of the target area of the vehicle to be detected into the hierarchical label attention neural network model, and outputting vehicle brand information, vehicle model information, and vehicle model year information.
9. An identification device based on hierarchical label attention, characterized in that, Including: An image acquisition module for acquiring a plurality of original images; A target position detection model generation module for labeling the target area of each original image in the plurality of original images to form a detection training set, and training to form a target position detection model based on the detection training set; A target area image generation module for dividing each original image in the plurality of original images according to the target position detection model to form a plurality of target area images; A hierarchical label attention neural network model generation module for labeling each target area image in the plurality of target area images with at least two layers of inclusion relationship labels to form a recognition training set, and training to form a hierarchical label attention neural network model based on the recognition training set; The hierarchical label attention neural network model generation module includes: A label feature branch generation module for extracting features from each target area image in the plurality of target area images by using a residual network to form a first-layer label feature branch and a second-layer label feature branch; A first-layer label feature branch operation module for running a classification function on the first-layer label feature branch to obtain the probabilities of each category corresponding to the first-layer label of each target area image and the first-layer label recognition result, and calculating a first cross-entropy loss function during model training; A new second-layer label feature branch generation module for fusing the probabilities of each category corresponding to the first-layer label of each target area image onto the second-layer label feature branch to form a new second-layer label feature branch; A new second-layer label feature branch operation module for running a classification function on the new second-layer label feature branch to obtain the probabilities of each category corresponding to the second-layer label under the first-layer label of each target area image and the second-layer label recognition result, and calculating a second cross-entropy loss function during model training; A first loss function adjustment module for calculating a first total loss function according to the first cross-entropy loss function and the second cross-entropy loss function; The new second-layer label feature branch generation module includes: A first attention weight generation module for expanding the dimension of the probabilities of each category corresponding to the first-layer label of the i-th target area image to the dimension corresponding to the features of the second-layer label feature branch on the second-layer label feature branch to obtain a first attention weight; The new feature generation module of the second-layer label feature branch is used to calculate the new features of the second-layer label feature branch according to the first attention weight, and form the new second-layer label feature branch; The detection and recognition module is used to perform detection and recognition on the image to be detected according to the target position detection model and the hierarchical label attention neural network model.
10. The recognition device based on hierarchical label attention according to claim 9, wherein Among them, The first cross-entropy loss function is as follows: N represents the number of the multiple target region images, represents the category corresponding to the first-layer label of the i-th target region image, represents the probabilities of each category corresponding to the first-layer label of the i-th target region image; wherein, the second cross-entropy loss function is as follows: represents the category corresponding to the second-layer label corresponding to the first-layer label of the i-th target region image, represents the probabilities of each category corresponding to the second-layer label corresponding to the first-layer label of the i-th target region image; The hierarchical label attention neural network model generation module includes: The first loss function adjustment module is used to form a stable hierarchical label attention neural network model when the first total loss function reaches a balanced state; among them, the first total loss function is: L totall = αL1 + βL2; α and β are the corresponding loss balance factors on the first-layer label feature branch and the corresponding loss balance factors on the new second-layer label feature branch respectively, L1 is the first cross-entropy loss function, and L2 is the second cross-entropy loss function.
11. The recognition device based on hierarchical label attention according to claim 10, characterized in that, The label feature branch generation module is also used to form a third-layer label feature branch; The hierarchical label attention neural network model generation module also includes: The new third-layer label feature branch generation module is used to fuse the probabilities of each category corresponding to the second layer label under the first layer label corresponding to each target region image onto the third layer label feature branch to form a new third-layer label feature branch; The new third-layer label feature branch operation module is used to run a classification function on the new third-layer label feature branch to obtain the probabilities of each category corresponding to the third layer label under the second layer label corresponding to each target region image and the third layer label recognition result, and calculate the third cross-entropy loss function during model training; among them, the third cross-entropy loss function is: Indicates the category corresponding to the third-level label under the second-level label corresponding to the i-th target region image, Indicates the probabilities of each category corresponding to the third-level label under the second-level label corresponding to the i-th target region image; The second loss function adjustment module is used to calculate the second total loss function according to the first cross-entropy loss function, the second cross-entropy loss function, and the third cross-entropy loss function, and form a stable hierarchical label attention neural network model when the second total loss function reaches a balanced state; among them, the second total loss function is: L total2 = αL1 + βL2 + λL3; λ is the corresponding loss balance factor on the new third-layer label feature branch, and L3 is the third cross-entropy loss function.
12. The recognition device based on hierarchical label attention according to claim 10, wherein The new features of the second-layer label feature branch are: Represents the probabilities of each category corresponding to the first-layer label corresponding to the i-th target region image, Represents the first attention weight, f2 (i) Represents the feature of the second-layer label feature branch, and o represents the dot product operation of two vectors.
13. The recognition device based on hierarchical label attention according to claim 11, wherein The new third-layer label feature branch generation module includes: The second attention weight generation module is used to expand the dimension of the probabilities of each category corresponding to the second layer label under the first layer label corresponding to the i-th target region image to the dimension corresponding to the features of the third layer label feature branch on the third layer label feature branch to obtain the second attention weight; The new feature generation module of the third-layer label feature branch is used to calculate the new features of the third layer label feature branch according to the second attention weight, and form the new third layer label feature branch; among them, the new features of the third layer label feature branch are: Represents the probabilities of each category corresponding to the second-layer labels under the first-layer label corresponding to the i-th target region image, Represents the second attention weight, f3 (i) Represents the feature of the third-layer label feature branch, and o represents the dot product operation of two vectors.
Citation Information
Patent Citations
Vehicle brand identification method, device and equipment and storage medium
CN110991506A
Image processing method and device, computer equipment and storage medium
CN113129319A