Vehicle brand identification model training method, identification method and related device
By using weakly supervised training and feature fusion, and leveraging vehicle key part bounding boxes and vehicle features, the problem of large calibration workload in vehicle brand recognition model training is solved, thereby improving training efficiency and recognition performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-14
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, the amount of sample image labeling required for training vehicle brand recognition models is large, resulting in low training efficiency.
Weakly supervised training is adopted. By using the preset vehicle key part bounding boxes and vehicle features of sample images, vehicle target features are generated through feature fusion. The loss value is calculated to adjust the parameters of the recognition model until the iterative optimization termination condition is met.
It effectively reduces the workload of sample image calibration, improves model training and recognition performance, and achieves more efficient vehicle brand recognition.
Smart Images

Figure CN115909250B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent transportation technology, and in particular to a training method, recognition method, and related apparatus for a vehicle brand recognition model. Background Technology
[0002] With the continuous development of modern manufacturing and technology, more and more motor vehicles are appearing in society, making the importance of automated vehicle management self-evident. Vehicle brand is an inherent attribute of a vehicle, and vehicle brand recognition is a crucial aspect of vehicle identification. For a given vehicle image, identifying its main brand, sub-brands, and model year plays a vital role in vehicle monitoring and management. For example... Figure 1 As shown, a single main vehicle brand contains numerous sub-brands, and each sub-brand is further subdivided into various model years. In other words, each vehicle can typically be categorized into a hierarchical tree-like label structure of "main brand / sub-brand / model year." With the booming development of the automotive transportation industry, the number of vehicle brands has increased rapidly. To achieve intelligent brand recognition, a large number of sample images are needed to train a neural network-based recognition model. Each sample image requires strongly supervised training, meaning each sample image must be labeled down to the model year level, inevitably leading to a significant workload for labeling.
[0003] No effective solution has yet been proposed to address the above problems. Summary of the Invention
[0004] This invention provides a training method, a recognition method, and related apparatus for a vehicle brand recognition model, to at least solve the technical problem in related technologies where the sample images required for training the recognition model involve a large amount of calibration work.
[0005] According to one aspect of the present invention, a training method for a vehicle brand recognition model is provided, comprising: acquiring a sample image set, wherein the sample image set includes multiple labeled sample images, the labels being used to characterize the vehicle brands in the sample images; acquiring vehicle features of each sample image; performing weakly supervised training on a pre-constructed recognition model using the vehicle features of each sample image based on a preset vehicle key part bounding box, obtaining an importance map associated with the labels of each sample image; fusing the importance map associated with the labels of each sample image and the vehicle features of each sample image to obtain vehicle target features of each sample image; calculating a loss value for each sample image based on the vehicle features of the sample images and the vehicle target features of each sample image; and adjusting the parameters of the recognition model based on the loss values of each sample image until the loss values of each sample image satisfy the iterative optimization termination condition.
[0006] Optionally, when the vehicle features of the sample image are the overall vehicle appearance features of the sample image, obtaining the vehicle features of each sample image includes: obtaining the overall vehicle appearance features in each sample image based on the feature extraction network of the recognition model.
[0007] Optionally, feature fusion is performed on the importance map associated with the labels of each sample image and the vehicle features of each sample image to obtain the vehicle target features of each sample image. This includes: performing pixel-level multiplication on the importance map associated with the labels of each sample image and the vehicle features of each sample image to obtain key part features in the importance map associated with the labels of each sample image; and performing pixel-level addition on the importance map associated with the labels of each sample image and the key part features in the importance map associated with the labels of each sample image to obtain the vehicle target features of each sample image.
[0008] Optionally, when the label is a model year, the vehicle key part bounding boxes include the vehicle logo key part bounding box, the front and hood key part bounding boxes, and the headlights and fog lights key part bounding boxes. The vehicle features of the sample images include the overall vehicle appearance features, shallow vehicle features, and mid-level vehicle features. Based on the preset vehicle key part bounding boxes, the vehicle features of each sample image are used to perform weakly supervised training on the pre-built recognition model to obtain the importance map of the label association for each sample image. The importance map of the label association for each sample image and the vehicle features of each sample image are then fused to obtain... The process involves: obtaining vehicle target features from each sample image; calculating the loss value for each sample image based on the vehicle features and the overall vehicle appearance features of each sample image; performing strongly supervised training on the recognition model using the overall vehicle appearance features of each sample image based on the vehicle logo key part bounding box; obtaining an importance map of the main brand association for each sample image; fusing the importance map of the main brand association and the overall vehicle appearance features of each sample image to obtain shallow vehicle features for each sample image; and then calculating the loss value for each sample image based on the overall vehicle appearance features and the overall vehicle appearance features of each sample image. The shallow vehicle features of the sample images are used to calculate the first loss value for each sample image. Based on the key part bounding boxes of the car front and hood, the shallow vehicle features of each sample image are used to perform strongly supervised training on the recognition model to obtain the importance map of sub-brand association for each sample image. The importance map of sub-brand association and the shallow vehicle features of each sample image are fused to obtain the mid-level vehicle features of each sample image. Based on the shallow vehicle features and the mid-level vehicle features of each sample image, the second loss value for each sample image is calculated. The key part bounding boxes of headlights and fog lights are used to perform strongly supervised training on the recognition model using the vehicle mid-level features of each sample image to obtain the importance map of the model year association of each sample image. The importance map of the model year association of each sample image and the vehicle mid-level features of each sample image are fused to obtain the vehicle deep features of each sample image. Based on the vehicle mid-level features and the vehicle deep features of each sample image, the third loss value of each sample image is calculated. The first loss value, the second loss value and the third loss value of each sample image are added together to obtain the loss value of each sample image.
[0009] Optionally, when the label is a sub-brand, the vehicle key part bounding box includes the logo key part bounding box, the front of the car and the hood key part bounding box. The vehicle features of the sample image include the overall vehicle appearance features and shallow vehicle features of the sample image. Based on the preset vehicle key part bounding box, the vehicle features of each sample image are used to perform weakly supervised training on the pre-built recognition model to obtain the importance map of the label association for each sample image. The importance map of the label association for each sample image and the vehicle features of each sample image are fused to obtain the vehicle target features of each sample image. Based on the vehicle features and the vehicle target features of each sample image, the loss value of each sample image is calculated, including: based on the logo key part bounding box, the overall vehicle appearance features of each sample image are used to perform strongly supervised training on the recognition model to obtain the main brand association importance map of each sample image; the main brand association importance map of each sample image and the overall vehicle appearance features of each sample image are fused to obtain the shallow vehicle features of each sample image; based on the overall vehicle appearance features and the shallow vehicle features of each sample image, the loss value of each sample image is calculated. The system calculates a first loss value for each sample image based on the shallow vehicle features of the sample images. It then performs strongly supervised training on the recognition model using the shallow vehicle features of each sample image, based on the key bounding boxes of the vehicle front and hood, to obtain an importance map of sub-brand associations for each sample image. The system then fuses the importance map of sub-brand associations with the shallow vehicle features of each sample image to obtain mid-level vehicle features. Based on the shallow and mid-level vehicle features of each sample image, the system calculates a second loss value for each sample image. Finally, the system adds the first and second calculated loss values to obtain a first predicted loss value for each sample image. The first calculated loss value is calculated based on the features corresponding to the main brand identified in the sample image and the predicted features corresponding to the main brand. The second calculated loss value is calculated based on the features corresponding to the sub-brand identified in the sample image and the predicted features corresponding to the sub-brand. Finally, the system adds the first, second, and first predicted loss values to obtain the final loss value for each sample image.
[0010] Optionally, when the label is the main brand, the vehicle key part bounding box includes the car logo key part bounding box, and the vehicle features of the sample images include the overall vehicle appearance features of the sample images. Based on the preset vehicle key part bounding boxes, the vehicle features of each sample image are used to perform weakly supervised training on the pre-built recognition model to obtain the importance map of the label association for each sample image. The importance map of the label association for each sample image and the vehicle features of each sample image are fused to obtain the vehicle target features of each sample image. Based on the vehicle features of the sample images and the vehicle target features of each sample image, the loss value of each sample image is calculated, including: based on the car logo key part bounding box, the vehicle key part bounding box is used to perform strongly supervised training on the recognition model to obtain the main brand association importance map of each sample image, and the main brand association importance map of each sample image is fused with the overall vehicle appearance features of each sample image. The overall vehicle appearance features of the sample images are fused to obtain shallow vehicle features for each sample image. Based on the overall vehicle appearance features and the shallow vehicle features of each sample image, a first loss value is calculated for each sample image. The first calculated loss value is used as a second predicted loss value for each sample image, where the first calculated loss value is calculated based on the features corresponding to the main brand identified in the sample image and the features corresponding to the predicted main brand. The first calculated loss value and the second calculated loss value are added together to obtain a first predicted loss value for each sample image, where the second calculated loss value is calculated based on the features corresponding to the sub-brand identified in the sample image and the features corresponding to the predicted sub-brand. Finally, the first loss value, the first predicted loss value, and the second predicted loss value are added together to obtain the total loss value for each sample image.
[0011] Optionally, the recognition model includes: a first convolutional layer, a second convolutional layer, a normalization layer, a sigmoid function, a first feature fusion network, and a second feature fusion network; wherein, the vehicle features of each sample image are processed by the first convolutional layer and the sigmoid function to extract the importance map of the label association of each sample image; the vehicle features of each sample image are processed by the second convolutional layer and the normalization layer to obtain the normalized vehicle features of each sample image; the normalized vehicle features of each sample image and the importance map of the label association of each sample image are processed by the first feature fusion network to obtain the key part features in the importance map of the label association of each sample image; the importance map of the label association of each sample image and the key part features in the importance map of the label association of each sample image are processed by the second feature fusion network to obtain the vehicle target features of each sample image.
[0012] According to another aspect of the present invention, a training apparatus for a vehicle brand recognition model is also provided, comprising: a first acquisition module, configured to acquire a sample image set, wherein the sample image set includes multiple labeled sample images, the labels being used to characterize the vehicle brand in the sample images; a training module, configured to acquire vehicle features of each of the sample images, perform weakly supervised training on a pre-constructed recognition model using the vehicle features of each of the sample images based on a preset vehicle key part bounding box, obtain an importance map associated with the labels of each of the sample images, fuse the importance map associated with the labels of each of the sample images and the vehicle features of each of the sample images to obtain vehicle target features of each of the sample images, calculate a loss value for each of the sample images based on the vehicle features of the sample images and the vehicle target features of each of the sample images, and adjust the parameters of the recognition model based on the loss values of each of the sample images until the loss values of each of the sample images satisfy the iterative optimization termination condition.
[0013] According to another aspect of the present invention, a vehicle brand recognition method is also provided, comprising: acquiring an image to be recognized; inputting the image to be recognized into a recognition model for processing to obtain the vehicle brand in the image to be recognized, wherein the recognition model is trained according to the training method of the vehicle brand recognition model described in any one of the above embodiments.
[0014] According to another aspect of the present invention, a vehicle brand recognition device is also provided, comprising: a second acquisition module for acquiring an image to be recognized; and a recognition module for inputting the image to be recognized into a recognition model for processing to obtain the vehicle brand in the image to be recognized, wherein the recognition model is trained according to the training method of the vehicle brand recognition model described in any one of the above embodiments.
[0015] According to another aspect of the present invention, an electronic device is also provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the training method of the vehicle brand recognition model described in any one of the preceding embodiments or the vehicle brand recognition method described above.
[0016] In this embodiment of the invention, a sample image set is acquired, which includes multiple labeled sample images, the labels being used to characterize the vehicle brands in the sample images; vehicle features of each sample image are acquired; based on a preset vehicle key part bounding box, the vehicle features of each sample image are used to perform weakly supervised training on a pre-built recognition model to obtain an importance map associated with the labels of each sample image; the importance map associated with the labels of each sample image and the vehicle features of each sample image are fused to obtain the vehicle target features of each sample image; based on the vehicle features of the sample images and the vehicle target features of each sample image, the loss value of each sample image is calculated; and the parameters of the recognition model are adjusted based on the loss values of each sample image until the loss values of each sample image meet the iterative optimization termination condition. In other words, this embodiment of the invention uses preset vehicle key part bounding boxes and vehicle features of each sample image as training data, and trains the recognition model using a weakly supervised training method. Then, it generates vehicle target features of each sample image through feature fusion. Next, it calculates the loss value of each sample image using the vehicle features of the sample images and the vehicle target features of each sample image, and adjusts the parameters of the recognition model based on the loss values of each sample image until the loss values of each sample image meet the termination condition of iterative optimization, thus obtaining the finally trained recognition model. This solves the technical problem in related technologies where the calibration workload of sample images required for recognition model training is large, thereby reducing the calibration workload of sample images and greatly improving the model training and recognition effect. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0018] Figure 1 This is a schematic diagram of a tree-like hierarchical label structure in the prior art.
[0019] Figure 2 A flowchart illustrating a training method for a vehicle brand recognition model provided in an embodiment of the present invention;
[0020] Figure 3 A schematic diagram of the neural network structure of a vehicle brand recognition model provided in an embodiment of the present invention;
[0021] Figure 4 This is a schematic diagram of a cascaded feature fusion network from shallow vehicle features to mid-level vehicle features provided in an embodiment of the present invention;
[0022] Figure 5 A schematic diagram of a training device for a vehicle brand recognition model provided in an embodiment of the present invention;
[0023] Figure 6 A flowchart of a vehicle brand identification method provided in an embodiment of the present invention;
[0024] Figure 7 This is a schematic diagram of a vehicle brand identification device provided in an embodiment of the present invention. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this invention are used to distinguish different objects, rather than to limit a specific order.
[0027] According to one aspect of the present invention, a method for training a vehicle brand recognition model is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0028] Figure 2 A flowchart illustrating a training method for a vehicle brand recognition model provided in an embodiment of the present invention is shown below. Figure 2 As shown, the method includes the following steps:
[0029] Step S202: Obtain a sample image set, wherein the sample image set includes multiple sample images with labels, and the labels are used to characterize the vehicle brand in the sample images;
[0030] The aforementioned vehicle brands employ a hierarchical tree-like labeling structure, including main brand, sub-brand, and model year. The labels carried by the sample images can be any of these three categories. That is, the labels carried by the sample images in the image set can be main brand, sub-brand, or model year, with main brand, sub-brand, and model year representing different levels within the hierarchical tree-like labeling structure of the vehicle brands. It should be noted that for the acquired labeled sample images, this training method does not require that the label of each sample image be specifically assigned to a model year, thus effectively reducing the workload of labeling the sample images. Furthermore, the image set contains a large number of labeled sample images, and the number of labeled sample images can be adjusted according to the needs of the application scenario.
[0031] Step S204: Obtain vehicle features of each sample image; perform weakly supervised training on the pre-built recognition model using the vehicle features of each sample image based on the preset vehicle key part bounding box to obtain the importance map of the label association of each sample image; fuse the importance map of the label association of each sample image with the vehicle features of each sample image to obtain the vehicle target features of each sample image; calculate the loss value of each sample image based on the vehicle features and the vehicle target features of each sample image; and adjust the parameters of the recognition model based on the loss value of each sample image until the loss value of each sample image meets the termination condition of iterative optimization.
[0032] The aforementioned recognition model is constructed based on a neural network. The vehicle features of the aforementioned sample images include, but are not limited to, the overall vehicle appearance features, shallow vehicle features, mid-level vehicle features, and deep vehicle features. Among them, the overall vehicle appearance features of the sample images serve as the original vehicle features in the sample images. As the number of layers in the trained neural network increases, the extracted vehicle features become increasingly refined. The shallow, mid-level, and deep vehicle features are generated as the training process progresses.
[0033] By detecting key points on the vehicle, key vehicle part bounding boxes are obtained. As the network depth increases, more vehicle details are gradually addressed. The aforementioned preset key vehicle part bounding boxes include, but are not limited to, key bounding boxes for the vehicle logo, the front of the vehicle and the hood, and the headlights and fog lights.
[0034] During training, the vehicle's key component bounding boxes are input as binary images. Since these boxes are not manually calibrated, their positions may be inaccurate; therefore, a Gaussian weighted distribution is applied around the key components. Through L2Loss supervision of these binary images, the importance map also exhibits a binary Gaussian distribution. This serves to filter key vehicle components in the recognition model, allowing it to focus on more discriminative regions. It should be noted that during training, the vehicle's key component bounding boxes do not need to be input; they are automatically generated based on the vehicle's key component detection algorithm.
[0035] The vehicle features in the above sample images are represented by pixels; the importance map includes the importance value of each pixel in the preset key vehicle parts box, wherein the importance value represents the degree of importance of each pixel in the preset key vehicle parts box; optionally, the pixels related to the label extracted from the preset key vehicle parts box are used to obtain the importance value of each pixel in the preset key vehicle parts box as the importance map.
[0036] In the aforementioned weakly supervised training, by combining the tree-like hierarchical label structure for vehicle brand recognition, excellent results in model year recognition can be achieved using fewer labeled sample images. Weakly supervised training based on the bounding boxes of key vehicle parts yields an importance map of the label associations for the corresponding sample images, enabling the network to focus on different details of the vehicle at different depths. Furthermore, by using feature fusion, vehicle features are gradually enhanced from shallow to deep layers of the network. Combined with the importance map, as the network depth increases, more and more vehicle details are considered.
[0037] It should be noted that the trained recognition model is used to identify the vehicle brand in the image to be recognized.
[0038] In the above embodiments of the present invention, a pre-defined vehicle key part bounding box and vehicle features of each sample image are used as training data. The pre-built recognition model is trained using a weakly supervised training method. Then, vehicle target features of each sample image are generated through feature fusion. The loss value of each sample image is calculated using the vehicle features of the sample images and the vehicle target features of each sample image. The parameters of the recognition model are adjusted based on the loss values of each sample image until the loss values of each sample image meet the termination condition of iterative optimization, thus obtaining the finally trained recognition model. This solves the technical problem in related technologies where the calibration workload of sample images required for recognition model training is large, thereby reducing the calibration workload of sample images and greatly improving the model training and recognition effect.
[0039] In one optional implementation, when the vehicle features of a sample image are the overall vehicle appearance features of the sample image, the vehicle features of each sample image are obtained, including: obtaining the overall vehicle appearance features in each sample image based on a feature extraction network.
[0040] Figure 3 This is a schematic diagram of the neural network structure of a vehicle brand recognition model provided in an embodiment of the present invention, as shown below. Figure 3 As shown, the training process for this neural network structure, based on weak supervision using bounding boxes of key vehicle parts and combined with cascaded feature fusion, is as follows:
[0041] Step S302: The sample image is processed by a backbone network (corresponding to the feature extraction network mentioned above) to extract the overall appearance features of the vehicle;
[0042] Step S304: The key parts of the car logo are weakly supervised for the overall features of the vehicle appearance to obtain the main brand importance map. Then, the overall features of the vehicle appearance are fused to obtain the shallow features of the vehicle. The main brand level loss function is calculated to obtain the first loss value of the sample image.
[0043] Step S306: The key parts such as the front of the car and the hood are bounded to the shallow features of the vehicle to obtain the sub-brand importance map. Then, the shallow features of the vehicle are fused to obtain the mid-level features of the vehicle. The sub-brand level loss function is calculated to obtain the second loss value of the sample image.
[0044] Step S307: The key parts such as headlights and fog lights are bounded to the vehicle's mid-level features for weak supervision to obtain the model year importance map. The vehicle's mid-level features are then fused to obtain the vehicle's deep features. The model year level loss function is calculated to obtain the third loss of the sample image.
[0045] It should be noted that the above main brand importance chart is a chart showing the importance of relationships between main brands; the above sub-brand importance chart is a chart showing the importance of relationships between sub-brands; and the above model year importance chart is a chart showing the importance of relationships between model years.
[0046] In one optional implementation, the importance map associated with the labels of each sample image and the vehicle features of each sample image are fused to obtain the vehicle target features of each sample image. This includes: performing pixel-level multiplication on the importance map associated with the labels of each sample image and the vehicle features of each sample image to obtain the key part features in the importance map associated with the labels of each sample image; and performing pixel-level addition on the importance map associated with the labels of each sample image and the key part features in the importance map associated with the labels of each sample image to obtain the vehicle target features of each sample image.
[0047] Considering the hierarchical tree-like label structure of "main brand / sub-brand / model year," a cascaded feature fusion network based on a pre-built recognition model—that is, multiplying the importance map associated with the labels of each sample image by pixel-level with the vehicle features of each sample image—can obtain the key features in the importance map associated with the labels of each sample image. Then, pixel-level addition is performed on the importance map associated with the labels of each sample image and the key features in the importance map associated with the labels of each sample image to obtain the vehicle target features of each sample image. This implementation method gradually enhances vehicle features from shallow to deep layers of the network, avoiding forgetting. In addition, combined with the importance map, as the network layer deepens, more and more vehicle details are considered.
[0048] Considering the hierarchical tree-like tag structure of "main brand / sub-brand / year," the loss function of this scheme is also progressively enhanced:
[0049]
[0050] in, Loss value at the main brand level. This represents the loss value at the sub-brand level. For model year level loss values; λ m =1,λ s =1 / N s , λ y =1 / N y N s N represents the number of sub-brands under a main brand. y This refers to the number of models released each year under a specific main brand.
[0051] Because vehicle brand recognition requires detailed specifications down to the model year level, and model year labeling is challenging (due to numerous categories and some similar appearances), a weakly supervised training method is used to train a pre-built recognition model. This effectively reduces the workload of labeling sample images. Therefore, embodiments of this invention utilize the importance map of weakly supervised vehicle key part bounding boxes, enabling the model to focus on discriminative regions while achieving excellent model year recognition results with less labeling.
[0052] In one optional implementation, when the label is the model year, the vehicle key part bounding boxes include the vehicle logo key part bounding box, the front and hood key part bounding boxes, and the headlights and fog lights key part bounding boxes. The vehicle features of the sample images include the overall vehicle appearance features, shallow vehicle features, and mid-level vehicle features. Based on the preset vehicle key part bounding boxes, the vehicle features of each sample image are used to perform weakly supervised training on the pre-built recognition model to obtain the importance map of the label association for each sample image. The importance map of the label association for each sample image and the vehicle features of each sample image are fused to obtain the vehicle target features of each sample image. Based on the vehicle features and the vehicle target features of each sample image, the loss value of each sample image is calculated. The specific implementation method is as follows:
[0053] Step S12: Based on the key parts bounding boxes of the car logo, the recognition model is trained with strong supervision using the overall vehicle appearance features of each sample image to obtain the importance map of the main brand association of each sample image. The importance map of the main brand association of each sample image and the overall vehicle appearance features of each sample image are fused to obtain the shallow vehicle features of each sample image. Based on the overall vehicle appearance features of the sample image and the shallow vehicle features of each sample image, the first loss value of each sample image is calculated.
[0054] Step S14: Based on the key part bounding boxes of the front of the car and the hood, the shallow features of the vehicle in each sample image are used to perform strongly supervised training on the recognition model to obtain the importance map of the sub-brand association of each sample image. The importance map of the sub-brand association of each sample image and the shallow features of the vehicle in each sample image are fused to obtain the mid-level features of the vehicle in each sample image. Based on the shallow features of the vehicle in each sample image and the mid-level features of the vehicle in each sample image, the second loss value of each sample image is calculated.
[0055] Step S16: Based on the key parts bounding boxes of headlights and fog lights, the recognition model is trained with strong supervision using the vehicle mid-level features of each sample image to obtain the importance map of the model year association of each sample image. The importance map of the model year association of each sample image and the vehicle mid-level features of each sample image are fused to obtain the vehicle deep features of each sample image. Based on the vehicle mid-level features and the vehicle deep features of each sample image, the third loss value of each sample image is calculated.
[0056] Step S18: Add the first loss value, the second loss value and the third loss value of each sample image to obtain the loss value of each sample image.
[0057] In other words, if the label is model year (level), the above three loss functions can be trained in a fully supervised manner; the first loss value of the above sample image is also the main brand level loss value, the second loss value of the above sample image is also the sub-brand level loss value, and the third loss value of the above sample image is also the model year level loss value.
[0058] In one optional implementation, when the label is a sub-brand, the vehicle key part bounding box includes the vehicle logo key part bounding box, the front of the vehicle and the hood key part bounding box. The vehicle features of the sample images include the overall appearance features of the vehicle and the shallow features of the vehicle. Based on the preset vehicle key part bounding boxes, the vehicle features of each sample image are used to perform weakly supervised training on the pre-built recognition model to obtain the importance map of the label association for each sample image. The importance map of the label association for each sample image and the vehicle features of each sample image are fused to obtain the vehicle target features of each sample image. Based on the vehicle features and the vehicle target features of each sample image, the loss value of each sample image is calculated. The specific implementation method is as follows:
[0059] Step S22: Based on the key part bounding boxes of the car logo, the recognition model is trained with strong supervision using the overall vehicle appearance features of each sample image to obtain the importance map of the main brand association of each sample image. The importance map of the main brand association of each sample image and the overall vehicle appearance features of each sample image are fused to obtain the shallow vehicle features of each sample image. Based on the overall vehicle appearance features and the shallow vehicle features of each sample image, the first loss value of each sample image is calculated. The parameters of the recognition model are adjusted based on the first loss value of each sample image until the first loss value of each sample image meets the termination condition of iterative optimization.
[0060] Step S24: Based on the key part bounding boxes of the front of the car and the hood, the shallow features of the vehicle in each sample image are used to perform strongly supervised training on the recognition model to obtain the importance map of the sub-brand association of each sample image. The importance map of the sub-brand association of each sample image and the shallow features of the vehicle in each sample image are fused to obtain the mid-level features of the vehicle in each sample image. Based on the shallow features of the vehicle in each sample image and the mid-level features of the vehicle in each sample image, the second loss value of each sample image is calculated.
[0061] Step S26: Add the first calculated loss value and the second calculated loss value of each sample image to obtain the first predicted loss value of each sample image; wherein, the first calculated loss value is the loss value calculated based on the features corresponding to the main brand marked in the sample image and the features corresponding to the predicted main brand, and the second calculated loss value is the loss value calculated based on the features corresponding to the sub-brand marked in the sample image and the features corresponding to the predicted sub-brand.
[0062] Step S28: Add the first loss value, the second loss value, and the first prediction loss value of each sample image to obtain the loss value of each sample image.
[0063] In other words, if the label is a sub-brand (level), the loss functions for the main brand and sub-brand can be trained in a fully supervised manner. The model year branch still gives the prediction results for the model year level, but the losses are calculated separately for the main brand and sub-brand parts. In this case: The first loss value of the above sample images is the main brand level loss value, the second loss value of the above sample images is the sub-brand level loss value, and the third loss value of the above sample images is the model year level loss value.
[0064] In one optional implementation, when the label is the main brand, the vehicle key part bounding box includes the vehicle logo key part bounding box, and the vehicle features of the sample image include the overall appearance features of the vehicle in the sample image. Based on the preset vehicle key part bounding box, the vehicle features of each sample image are used to perform weakly supervised training on the pre-built recognition model to obtain the importance map of the label association for each sample image. The importance map of the label association for each sample image and the vehicle features of each sample image are fused to obtain the vehicle target features of each sample image. Based on the vehicle features and the vehicle target features of each sample image, the loss value of each sample image is calculated. The specific implementation method is as follows:
[0065] Step S32: Based on the key parts bounding box of the car logo, the recognition model is trained with strong supervision using the overall features of the vehicle appearance of each sample image to obtain the importance map of the main brand association of each sample image. The importance map of the main brand association of each sample image and the overall features of the vehicle appearance of each sample image are fused to obtain the shallow features of the vehicle in each sample image. Based on the overall features of the vehicle appearance of the sample image and the shallow features of the vehicle in each sample image, the first loss value of each sample image is calculated.
[0066] Step S34: Use the first calculated loss value of each sample image as the second predicted loss value of each sample image, wherein the first calculated loss value is the loss value calculated based on the features corresponding to the main brand marked in the sample image and the features corresponding to the predicted main brand.
[0067] Step S36: Add the first calculated loss value and the second calculated loss value of each sample image to obtain the first predicted loss value of each sample image; wherein, the second calculated loss value is the loss value calculated based on the features corresponding to the sub-brands marked in the sample image and the features corresponding to the predicted sub-brands.
[0068] Step S38: Add the first loss value, the first predicted loss value, and the second predicted loss value of each sample image to obtain the loss value of each sample image.
[0069] In other words, if the label is the main brand (level), the main brand loss function can be trained in a fully supervised manner. The sub-brand branch still gives the prediction results at the sub-brand level, but the loss is calculated using the main brand portion. The model year branch is the same as above; the first loss value of the above sample image is also the main brand level loss value, the second loss value of the above sample image is also the sub-brand level loss value, and the third loss value of the above sample image is also the model year level loss value.
[0070] In this process, although there is less strong supervision at the sub-brand or model year level, with the assistance of weak supervision at key parts of the vehicle, the detailed features of the vehicle can still be learned, thus achieving excellent results in model year recognition with less annotation.
[0071] In one optional implementation, the recognition model includes: a first convolutional layer, a second convolutional layer, a normalization layer, a sigmoid function, a first feature fusion network, and a second feature fusion network; wherein, the vehicle features of each sample image are processed by the first convolutional layer and the sigmoid function to extract the importance map of the label association of each sample image; the vehicle features of each sample image are processed by the second convolutional layer and the normalization layer to obtain the normalized vehicle features of each sample image; the normalized vehicle features of each sample image and the importance map of the label association of each sample image are processed by the first feature fusion network to obtain the key part features in the importance map of the label association of each sample image; the importance map of the label association of each sample image and the key part features in the importance map of the label association of each sample image are processed by the second feature fusion network to obtain the vehicle target features of each sample image.
[0072] It should be noted that the first and second convolutional layers have the same convolution parameters. The first feature fusion network is used to perform pixel-level multiplication of the importance map associated with the labels of each sample image and the vehicle features of each sample image to obtain the key part features in the importance map associated with the labels of each sample image. The second feature fusion network is used to perform pixel-level addition of the importance map associated with the labels of each sample image and the key part features in the importance map associated with the labels of each sample image to obtain the vehicle target features in each sample image. In addition, the first convolutional layer, the second convolutional layer, the normalization layer, the sigmoid function, the first feature fusion network, and the second feature fusion network constitute the cascaded feature fusion network of the recognition model. The recognition model may contain one or more cascaded feature fusion networks.
[0073] Figure 4 This is a schematic diagram of a cascaded feature fusion network from shallow to mid-level vehicle features provided in an embodiment of the present invention. Figure 4As shown, the cascaded feature fusion network from shallow vehicle features to mid-level vehicle features is implemented as follows: Under the supervision of key part bounding boxes at the sub-brand level, the shallow vehicle features are processed through the first convolutional layer and the Sigmoid function to obtain the sub-brand importance map; after processing by the second convolutional layer and the normalization layer, the vehicle features and the sub-brand importance map of each sample image are multiplied at the pixel level to extract the key part features from the shallow vehicle features; the shallow vehicle features and the result of the previous step are added at the pixel level to strengthen the weight of the key part features in the shallow vehicle features, thus obtaining the mid-level vehicle features.
[0074] According to another aspect of the present invention, a training apparatus for a vehicle brand recognition model is also provided. Figure 5 This is a schematic diagram of a training device for a vehicle brand recognition model provided in an embodiment of the present invention, as shown below. Figure 5 As shown, the training device for the vehicle brand recognition model includes a first acquisition module 502 and a training module 504. The training device for the vehicle brand recognition model will be described in detail below.
[0075] The first acquisition module 502 is used to acquire a sample image set, wherein the sample image set includes multiple sample images with labels, and the labels are used to characterize the vehicle brand in the sample images.
[0076] The training module 504, connected to the first acquisition module 502, is used to acquire vehicle features of each sample image. Based on the preset vehicle key part bounding box, the vehicle features of each sample image are used to perform weakly supervised training on the pre-built recognition model to obtain the importance map of the label association of each sample image. The importance map of the label association of each sample image and the vehicle features of each sample image are fused to obtain the vehicle target features of each sample image. Based on the vehicle features of the sample image and the vehicle target features of each sample image, the loss value of each sample image is calculated. The parameters of the recognition model are adjusted based on the loss value of each sample image until the loss value of each sample image meets the termination condition of iterative optimization.
[0077] This invention utilizes preset vehicle key part bounding boxes and vehicle features of each sample image as training data. A pre-built recognition model is trained using weakly supervised training. Vehicle target features of each sample image are then generated through feature fusion. The loss value of each sample image is calculated using the vehicle features and vehicle target features of each sample image. The parameters of the recognition model are adjusted based on the loss values of each sample image until the loss values of each sample image meet the iterative optimization termination condition, resulting in the final trained recognition model. This solves the technical problem of the large amount of calibration work required for training recognition models in related technologies, thereby reducing the amount of sample image calibration work and greatly improving model training and recognition performance.
[0078] It should be noted that the first acquisition module 502 and training module 504 mentioned above correspond to steps S202 to S204 in the method embodiment. The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above method embodiment.
[0079] Optionally, when the vehicle features of the sample image are the overall vehicle appearance features of the sample image, the training module 504 includes: an acquisition unit, used to acquire the overall vehicle appearance features in each sample image based on the feature extraction network of the recognition model.
[0080] Optionally, the training module 504 includes: a first acquisition unit, used to perform pixel-level multiplication calculation on the importance map associated with the labels of each sample image and the vehicle features of each sample image to obtain key part features in the importance map associated with the labels of each sample image; and a second acquisition unit, used to perform pixel-level addition calculation on the importance map associated with the labels of each sample image and the key part features in the importance map associated with the labels of each sample image to obtain vehicle target features in each sample image.
[0081] Optionally, when the label is model year, the vehicle key part bounding boxes include the vehicle logo key part bounding box, the front and hood key part bounding boxes, and the headlights and fog lights key part bounding boxes. The vehicle features of the sample images include the overall vehicle appearance features, shallow vehicle features, and mid-level vehicle features. The training module 504 includes: a first training unit, used to perform strongly supervised training of the recognition model based on the vehicle logo key part bounding boxes and the overall vehicle appearance features of each sample image to obtain the importance map of the main brand association of each sample image; to fuse the importance map of the main brand association of each sample image with the overall vehicle appearance features of each sample image to obtain the shallow vehicle features of each sample image; and to calculate the first loss value of each sample image based on the overall vehicle appearance features and the shallow vehicle features of each sample image; and a second training unit, used to perform strongly supervised training of the recognition model based on the front and hood key part bounding boxes and the shallow vehicle features of each sample image. The system obtains the importance map of sub-brand association for each sample image, fuses the importance map of sub-brand association with the shallow vehicle features of each sample image to obtain the mid-level vehicle features of each sample image, and calculates the second loss value for each sample image based on the shallow and mid-level vehicle features. The third training unit performs strongly supervised training of the recognition model using the mid-level vehicle features of each sample image based on the key bounding boxes of headlights and fog lights, obtaining the importance map of model year association for each sample image. It then fuses the importance map of model year association with the mid-level vehicle features of each sample image to obtain the deep vehicle features of each sample image, and calculates the third loss value for each sample image based on the mid-level and deep vehicle features. The first calculation unit adds the first, second, and third loss values of each sample image to obtain the loss value for each sample image.
[0082] Optionally, when the label is a sub-brand, the vehicle key part bounding box includes the logo key part bounding box, the front of the car and the hood key part bounding box, and the vehicle features of the sample image include the overall vehicle appearance features and shallow vehicle features of the sample image. The training module 504 includes: a first training unit, used to perform strongly supervised training on the recognition model based on the logo key part bounding box and the overall vehicle appearance features of each sample image to obtain the importance map of the main brand association of each sample image; to perform feature fusion of the importance map of the main brand association of each sample image and the overall vehicle appearance features of each sample image to obtain the shallow vehicle features of each sample image; and to calculate the first loss value of each sample image based on the overall vehicle appearance features and the shallow vehicle features of each sample image; a second training unit, used to perform strongly supervised training on the recognition model based on the front of the car and the hood key part bounding box and the shallow vehicle features of each sample image to obtain the importance map of the main brand association of each sample image and the overall vehicle appearance features of each sample image to obtain the shallow vehicle features of each sample image. The importance map of sub-brand associations for this image is used to fuse the importance maps of sub-brand associations for each sample image with the shallow vehicle features of each sample image to obtain the mid-level vehicle features of each sample image. Based on the shallow and mid-level vehicle features of each sample image, the second loss value of each sample image is calculated. The third training unit is used to add the first and second calculated loss values of each sample image to obtain the first predicted loss value of each sample image. The first calculated loss value is the loss value calculated based on the features corresponding to the main brand in the sample image and the features corresponding to the predicted main brand. The second calculated loss value is the loss value calculated based on the features corresponding to the sub-brands in the sample image and the features corresponding to the predicted sub-brands. The second calculation unit is used to add the first loss value, the second loss value, and the first predicted loss value of each sample image to obtain the loss value of each sample image.
[0083] Optionally, when the label is the main brand, the vehicle key part bounding box includes the car logo key part bounding box, and the vehicle features of the sample image include the overall vehicle appearance features of the sample image. The training module 504 includes: a first training unit, used to perform strongly supervised training on the recognition model based on the car logo key part bounding box and the overall vehicle appearance features of each sample image to obtain the importance map of the main brand association of each sample image; to fuse the importance map of the main brand association of each sample image and the overall vehicle appearance features of each sample image to obtain the shallow vehicle features of each sample image; and to calculate the first loss value of each sample image based on the overall vehicle appearance features and the shallow vehicle features of each sample image; a second training unit... The first training unit is used to use the first calculated loss value of each sample image as the second predicted loss value of each sample image, wherein the first calculated loss value is the loss value calculated based on the features corresponding to the main brand in the sample image and the features corresponding to the predicted main brand; the third training unit is used to add the first calculated loss value and the second calculated loss value of each sample image to obtain the first predicted loss value of each sample image; wherein the second calculated loss value is the loss value calculated based on the features corresponding to the sub-brand in the sample image and the features corresponding to the predicted sub-brand; the third calculation unit is used to add the first loss value, the first predicted loss value and the second predicted loss value of each sample image to obtain the loss value of each sample image.
[0084] According to another aspect of the present invention, a vehicle brand identification method is also provided. Figure 6 A flowchart of a vehicle brand recognition method provided in an embodiment of the present invention is shown below. Figure 6 As shown, the method includes the following steps:
[0085] Step S602: Obtain the image to be recognized;
[0086] Step S604: Input the image to be identified into the recognition model for processing to obtain the vehicle brand in the image to be identified. The recognition model is trained according to the training method of the vehicle brand recognition model mentioned above.
[0087] In the above embodiments of the present invention, the image to be identified is input into the trained recognition model, which can output the "main brand / sub-brand / model year" of the image to be identified, thereby realizing vehicle brand recognition.
[0088] It should be noted that the vehicle brand recognition model in this embodiment of the invention uses preset vehicle key part bounding boxes and vehicle features of each sample image as training data. It trains the pre-built recognition model using a weakly supervised training method, and then generates vehicle target features of each sample image through feature fusion. Then, it calculates the loss value of each sample image using the vehicle features of the sample images and the vehicle target features of each sample image, and adjusts the parameters of the recognition model based on the loss values of each sample image until the loss values of each sample image meet the termination condition of iterative optimization, thus obtaining the finally trained recognition model. This solves the technical problem in related technologies where the calibration workload of sample images required for recognition model training is large, thereby reducing the calibration workload of sample images and greatly improving the model training and recognition effect.
[0089] According to another aspect of the present invention, a vehicle brand identification device is also provided. Figure 7 This is a schematic diagram of a vehicle brand identification device provided in an embodiment of the present invention, such as... Figure 7 As shown, the vehicle brand identification device includes a second acquisition module 702 and an identification module 704. The vehicle brand identification device will be described in detail below.
[0090] The second acquisition module 702 is used to acquire the image to be recognized;
[0091] The recognition module 704 is connected to the second acquisition module 702 mentioned above, and is used to input the image to be recognized into the recognition model for processing to obtain the vehicle brand in the image to be recognized. The recognition model is trained according to the training method of the vehicle brand recognition model mentioned above.
[0092] It should be noted that the second acquisition module 702 and the identification module 704 mentioned above correspond to steps S602 to S604 in the method embodiment. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above method embodiment.
[0093] According to another aspect of the present invention, an electronic device is also provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the training method of the vehicle brand recognition model or the vehicle brand recognition method described above.
[0094] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.
Claims
1. A training method for a vehicle brand recognition model, characterized in that, include: Obtain a sample image set, wherein the sample image set includes multiple sample images with labels, the labels being used to characterize the vehicle brand in the sample images; The vehicle features of each sample image are obtained. Based on the preset vehicle key part bounding boxes, the vehicle features of each sample image are used to perform weakly supervised training on the pre-built recognition model to obtain the importance map of the label association of each sample image. The importance map of the label association of each sample image and the vehicle features of each sample image are fused to obtain the vehicle target features of each sample image. Based on the vehicle features and vehicle target features of each sample image, the loss value of each sample image is calculated. The parameters of the recognition model are adjusted based on the loss value of each sample image until the loss value of each sample image meets the iteration optimization termination condition. The method involves fusing the importance maps associated with the labels of each sample image and the vehicle features of each sample image to obtain the vehicle target features of each sample image. This includes: performing pixel-level multiplication on the importance maps associated with the labels of each sample image and the vehicle features of each sample image to obtain key part features in the importance maps associated with the labels of each sample image; and performing pixel-level addition on the importance maps associated with the labels of each sample image and the key part features in the importance maps associated with the labels of each sample image to obtain the vehicle target features of each sample image.
2. The method according to claim 1, characterized in that, When the vehicle features of the sample image are the overall appearance features of the vehicle in the sample image, the vehicle features of each sample image are obtained, including: The feature extraction network based on the recognition model obtains the overall vehicle appearance features in each of the sample images.
3. The method according to claim 1, characterized in that, When the label is a model year, the vehicle key part bounding boxes include the vehicle logo key part bounding box, the front and hood key part bounding boxes, and the headlights and fog lights key part bounding boxes. The vehicle features of the sample images include the overall vehicle appearance features, shallow vehicle features, and mid-level vehicle features. Based on the preset vehicle key part bounding boxes, the vehicle features of each sample image are used to perform weakly supervised training on the pre-built recognition model to obtain the importance map of the label association for each sample image. The importance map of the label association for each sample image and the vehicle features of each sample image are fused to obtain the vehicle target features of each sample image. Based on the vehicle features and the vehicle target features of each sample image, the loss value of each sample image is calculated, including: Based on the key part bounding box of the car logo, the recognition model is trained with strong supervision using the overall vehicle appearance features of each sample image to obtain the importance map of the main brand association of each sample image. The importance map of the main brand association of each sample image and the overall vehicle appearance features of each sample image are fused to obtain the shallow vehicle features of each sample image. Based on the overall vehicle appearance features and the shallow vehicle features of each sample image, the first loss value of each sample image is calculated. Based on the key part bounding boxes of the front of the car and the hood, the recognition model is trained with strong supervision using the shallow vehicle features of each of the sample images to obtain the importance map of the sub-brand association of each of the sample images. The importance map of the sub-brand association of each of the sample images and the shallow vehicle features of each of the sample images are fused to obtain the mid-level vehicle features of each of the sample images. Based on the shallow vehicle features and the mid-level vehicle features of each of the sample images, the second loss value of each of the sample images is calculated. Based on the key part bounding boxes of the headlights and fog lights, the recognition model is trained with strong supervision using the vehicle mid-level features of each of the sample images to obtain the importance map of the model year association of each of the sample images. The importance map of the model year association of each of the sample images and the vehicle mid-level features of each of the sample images are fused to obtain the vehicle deep features of each of the sample images. Based on the vehicle mid-level features and the vehicle deep features of each of the sample images, the third loss value of each of the sample images is calculated. The first loss value, the second loss value, and the third loss value of each sample image are added together to obtain the loss value of each sample image.
4. The method according to claim 1, characterized in that, When the label is a sub-brand, the vehicle key part bounding box includes the vehicle logo key part bounding box, the front of the car and the hood key part bounding box. The vehicle features of the sample image include the overall appearance features of the vehicle and the shallow features of the vehicle. Based on the preset vehicle key part bounding box, the vehicle features of each sample image are used to perform weakly supervised training on the pre-built recognition model to obtain the importance map of the label association for each sample image. The importance map of the label association for each sample image and the vehicle features of each sample image are fused to obtain the vehicle target features of each sample image. Based on the vehicle features and the vehicle target features of each sample image, the loss value of each sample image is calculated, including: Based on the key part bounding box of the car logo, the recognition model is trained with strong supervision using the overall vehicle appearance features of each sample image to obtain the importance map of the main brand association of each sample image. The importance map of the main brand association of each sample image and the overall vehicle appearance features of each sample image are fused to obtain the shallow vehicle features of each sample image. Based on the overall vehicle appearance features and the shallow vehicle features of each sample image, the first loss value of each sample image is calculated. Based on the key part bounding boxes of the front of the car and the hood, the recognition model is trained with strong supervision using the shallow vehicle features of each of the sample images to obtain the importance map of the sub-brand association of each of the sample images. The importance map of the sub-brand association of each of the sample images and the shallow vehicle features of each of the sample images are fused to obtain the mid-level vehicle features of each of the sample images. Based on the shallow vehicle features and the mid-level vehicle features of each of the sample images, the second loss value of each of the sample images is calculated. The first calculated loss value and the second calculated loss value of each sample image are added together to obtain the first predicted loss value of each sample image. The first calculated loss value is the loss value calculated based on the features corresponding to the main brand marked in the sample image and the features corresponding to the predicted main brand. The second calculated loss value is the loss value calculated based on the features corresponding to the sub-brand marked in the sample image and the features corresponding to the predicted sub-brand. The first loss value, the second loss value, and the first prediction loss value of each sample image are added together to obtain the loss value of each sample image.
5. The method according to claim 1, characterized in that, When the label is the main brand, the vehicle key part bounding box includes the vehicle logo key part bounding box, and the vehicle features of the sample image include the overall appearance features of the vehicle in the sample image. Based on the preset vehicle key part bounding box, the vehicle features of each sample image are used to perform weakly supervised training on the pre-built recognition model to obtain the importance map of the label association for each sample image. The importance map of the label association for each sample image and the vehicle features of each sample image are fused to obtain the vehicle target features of each sample image. Based on the vehicle features and the vehicle target features of each sample image, the loss value of each sample image is calculated, including: Based on the key part bounding box of the car logo, the recognition model is trained with strong supervision using the overall vehicle appearance features of each sample image to obtain the importance map of the main brand association of each sample image. The importance map of the main brand association of each sample image and the overall vehicle appearance features of each sample image are fused to obtain the shallow vehicle features of each sample image. Based on the overall vehicle appearance features and the shallow vehicle features of each sample image, the first loss value of each sample image is calculated. The first calculated loss value of each of the sample images is used as the second predicted loss value of each of the sample images, wherein the first calculated loss value is the loss value calculated based on the features corresponding to the main brand tagged in the sample image and the features corresponding to the predicted main brand. The first calculated loss value and the second calculated loss value of each sample image are added together to obtain the first predicted loss value of each sample image, wherein the second calculated loss value is the loss value calculated based on the features corresponding to the sub-brands identified in the sample image and the features corresponding to the predicted sub-brands. The first loss value, the first predicted loss value, and the second predicted loss value of each sample image are added together to obtain the loss value of each sample image.
6. The method according to any one of claims 1 to 5, characterized in that, The recognition model includes: a first convolutional layer, a second convolutional layer, a normalization layer, a sigmoid function, a first feature fusion network, and a second feature fusion network; wherein, the vehicle features of each sample image are processed by the first convolutional layer and the sigmoid function to extract the importance map of the label association of each sample image; the vehicle features of each sample image are processed by the second convolutional layer and the normalization layer to obtain the normalized vehicle features of each sample image; the normalized vehicle features of each sample image and the importance map of the label association of each sample image are processed by the first feature fusion network to obtain the key part features in the importance map of the label association of each sample image; the importance map of the label association of each sample image and the key part features in the importance map of the label association of each sample image are processed by the second feature fusion network to obtain the vehicle target features of each sample image.
7. A training device for a vehicle brand recognition model, characterized in that, include: The first acquisition module is used to acquire a sample image set, wherein the sample image set includes multiple sample images with labels, and the labels are used to characterize the vehicle brand in the sample images; The training module is used to acquire vehicle features of each sample image, perform weakly supervised training on a pre-built recognition model using the vehicle features of each sample image based on a preset vehicle key part bounding box, obtain an importance map of the label association of each sample image, fuse the importance map of the label association of each sample image and the vehicle features of each sample image to obtain the vehicle target features of each sample image, calculate the loss value of each sample image based on the vehicle features and the vehicle target features of each sample image, and adjust the parameters of the recognition model based on the loss value of each sample image until the loss value of each sample image meets the iterative optimization termination condition; The training module includes: a first acquisition unit, configured to perform pixel-level multiplication of the importance map associated with the label of each sample image and the vehicle features of each sample image to obtain key part features in the importance map associated with the label of each sample image; and a second acquisition unit, configured to perform pixel-level addition of the importance map associated with the label of each sample image and the key part features in the importance map associated with the label of each sample image to obtain vehicle target features in each sample image.
8. A vehicle brand identification method, characterized in that, include: Acquire the image to be recognized; The image to be identified is input into the recognition model for processing to obtain the vehicle brand in the image to be identified, wherein the recognition model is trained by the training method of the vehicle brand recognition model according to any one of claims 1 to 6.
9. A vehicle brand identification device, characterized in that, include: The second acquisition module is used to acquire the image to be recognized; The recognition module is used to input the image to be recognized into the recognition model for processing to obtain the vehicle brand in the image to be recognized, wherein the recognition model is trained by the training method of the vehicle brand recognition model according to any one of claims 1 to 6.
10. An electronic device, characterized in that, include: processor; A memory for storing processor-executable instructions; wherein the processor is configured to execute the training method of the vehicle brand recognition model according to any one of claims 1 to 6 or the vehicle brand recognition method according to claim 8.
Citation Information
Patent Citations
Fine vehicle type identification method and system based on deep learning
CN112966709A
Vehicle re-identification method using weak supervision area recommendation
CN113177518A