Target detection model training method, target detection method, target detection device and vehicle

By maximizing the category probability difference and multi-scale feature fusion of heat maps in the 3D object detection model, the object detection model is optimized, and the type false detection and missed detection problems in noise interference and complex environments are solved, and the detection accuracy and robustness are improved.

CN120388169AActive Publication Date: 2025-07-29CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510874452.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Existing 3D object detection technology is prone to type false detection and missed detection in noise interference and complex environments, and it is difficult to effectively integrate multi-scale features, resulting in insufficient detection accuracy and robustness.

Method used

By maximizing the category probability differences between real categories and other categories on the heat map, the object detection model is optimized, and multi-scale feature alignment and fusion is used to optimize model parameters by using the feature extraction layer, feature fusion layer and object detection layer, combining the first loss and the second loss.

Benefits of technology

It significantly improves the classification accuracy and robustness of the object detection model in complex scenarios, reduces type error detection and noise interference, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388169A_ABST
    Figure CN120388169A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, in particular to a target detection model training method, a target detection method, a target detection device and a vehicle, and aims to solve the problem of type false detection caused by incomplete feature information or noise interference. The target detection model training method comprises the steps of determining a thermodynamic diagram corresponding to a sample picture by using a target detection model, determining a key point from a plurality of position points included in a real frame of a to-be-recognized target based on a real category of the to-be-recognized target in the thermodynamic diagram, and optimizing the target detection model by taking maximization of first loss as a target, each position point in the thermodynamic diagram shows category probabilities of a plurality of object categories, the first loss is used for representing a difference between a first category probability corresponding to a real category and a second category probability not corresponding to the real category in the key points, and the real category is highlighted by increasing the difference between the first category probability and the second category probability; and the detection precision of the target detection model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image recognition, and particularly to a training method for a target detection model, a target detection method, an apparatus, and a vehicle. Background Art

[0002] Target detection refers to identifying specific objects in an image or video and determining the positions and categories of the objects. Three-dimensional (3D) target detection refers to identifying and locating objects in three-dimensional space and outputting information such as the category, position, size, and orientation of the targets.

[0003] In related technologies, by decomposing a nighttime image and assigning weights according to the distance between the pixel value and the maximum pixel value, the brightness range of the target is highlighted to solve the problems of a large number of noise points and insufficient brightness in the nighttime image. The above technical solution assigns weights based on the pixel values in the feature map to highlight the target brightness range, which easily loses feature information, and thus reduces the accuracy of target recognition due to the incompleteness of information.

[0004] In another related technology, an infrared image segmentation function is constructed under the guidance of three purposes: noise removal, detail retention, and edge preservation. A multi-objective optimization algorithm is used to combine with the segmentation model for image segmentation and recognition in the infrared reconstructed image. The above technical solution cannot accurately highlight the target and suppress noise in a scene where multiple objects coexist and the background is complex, and relies on pixel segmentation to judge the target category, which easily leads to misclassification due to the incompleteness of feature information or noise interference. Summary of the Invention

[0005] The present invention provides a training method for a target detection model, a target detection method, an apparatus, and a vehicle, which improves the detection accuracy of the target detection model.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows: In a first aspect, the present application provides a training method for a target detection model, and the training method for the target detection model includes: Using the target detection model, determining a heat map corresponding to a sample picture.

[0007] Based on the true category of the target to be recognized in the heat map, determining key points from multiple position points included in the true box of the target to be recognized.

[0008] Optimizing the target detection model with the goal of maximizing a first loss.

[0009] Wherein, each position point in the heat map shows the category probabilities of multiple object categories, and the first loss is used to characterize the difference between the first category probability corresponding to the true category and the second category probability not corresponding to the true category among the key points.

[0010] The technical solution provided by the embodiments of this application increases the ability of the object detection model to distinguish the true class from other classes by maximizing the difference between the true class and the class probabilities of other classes at the same spatial position on the heat map, enabling the object detection model to better determine the true class, significantly improving the classification accuracy of the object detection model, and being applicable to complex scenarios such as object occlusion and similar appearances.

[0011] A possible implementation is to determine key points from multiple position points included in the true bounding box of the target to be recognized based on the true class of the target to be recognized in the heat map. It can be specifically implemented as: determining the position point with the highest first class probability among the multiple position points included in the true bounding box as the key point. Determining the position point with the highest first class probability in the true bounding box as the key point can accurately describe the true position of the target to be recognized. Maximizing the first loss at this position can highlight the true position and class of the target to be recognized to the greatest extent, further improving the classification accuracy of the object detection model.

[0012] A possible implementation is to optimize the object detection model with the goal of maximizing the first loss. It can be specifically implemented as: optimizing the object detection model with the goal of maximizing the first loss and the second loss. Among them, the second loss is used to represent the difference between the first class probability of the key point and the first class probabilities of the neighboring points of the key point. By maximizing the difference between the first class probability of the key point and the first class probabilities of the neighboring points of the key point, it is ensured that the first class probability of the key point is significantly higher than the first class probabilities of the surrounding noise points, effectively suppressing noise interference, improving the robustness of the object detection model in complex environments, and reducing false detections caused by noise.

[0013] A possible implementation is that the calculation process of the first loss in the embodiments of this application includes: determining the first class probability on the key point, determining the second class probabilities corresponding to at least one class other than the true class on the key point, respectively determining at least one first difference between the first class and the at least one second class probability, and taking the sum of the absolute values of the at least one first difference as the first loss. The calculation of the first loss is clear and efficient, enabling the object detection model to more accurately learn the relationships between targets, reducing overfitting, and improving the generalization ability of the model.

[0014] A possible implementation. In the embodiments of the present application, the calculation process of the second loss includes: determining the first category probability at the key points, determining the first category probabilities of at least one neighborhood point of the key points, respectively determining at least one second difference between the first category probability and the first category probabilities of at least one neighborhood point, and taking the sum of the absolute values of at least one second difference as the second loss. The calculation of the second loss is clear and efficient, enabling the object detection model to more accurately learn the relationships between objects, reducing the overfitting phenomenon, and improving the generalization ability of the model.

[0015] A possible implementation. The object detection model includes a feature extraction layer, a feature fusion layer, and an object detection layer. Using the object detection model to determine the heat map corresponding to the sample image can be specifically implemented as: using the feature extraction layer to determine the multi-scale feature maps of the sample image, using the feature fusion layer to perform feature alignment and fusion on the multi-scale feature maps to obtain the target feature map, and using the object detection layer to determine the heat map corresponding to the target feature map. By performing max pooling and alignment operations on the multi-scale feature maps, the fusion of multi-scale information is effectively realized, thereby improving the detection accuracy and robustness of the object detection model.

[0016] A possible implementation. Using the feature fusion layer to perform feature alignment and fusion on the multi-scale feature maps to obtain the target feature map can be specifically implemented as: using the feature fusion layer to perform downsampling on the multi-scale feature maps through a pooling window to align the sizes of the multi-scale feature maps, and splicing and fusing the aligned multi-scale feature maps to obtain the target feature map. The maximum pooling method is used to unify the sizes of the multi-scale feature maps. While aligning the multi-scale feature maps, the most significant feature information is retained during the pooling process, avoiding information loss and reducing false detections, and further improving the detection accuracy of the object detection model.

[0017] In a second aspect, the embodiments of the present application provide an object detection method, and the method includes: Obtain the image to be detected.

[0018] Input the image to be detected into the object detection model, and predict and output the object detection result corresponding to the image to be detected through the object detection model. Among them, the object detection model is trained based on the training method of the object detection model in the first aspect above.

[0019] The object detection method provided by the embodiments of the present application outputs the object detection result through the object detection model trained by the training method of the object detection model in the first aspect above, improving the detection accuracy of object detection.

[0020] In a third aspect, the embodiments of the present application provide a training device for an object detection model, and the device includes: a prediction module and an optimization module.

[0021] The above prediction module is used to determine the heat map corresponding to the sample image by using the object detection model.

[0022] The above optimization module is used to determine key points from multiple position points included in the ground truth box of the object to be recognized based on the true category of the object to be recognized in the heat map.

[0023] The above optimization module is also used to optimize the object detection model with the goal of maximizing the first loss.

[0024] Among them, each position point in the heat map shows the class probabilities of multiple object classes, and the first loss is used to characterize the difference between the first class probability corresponding to the true category in the key points and the second class probability not corresponding to the true category.

[0025] A possible implementation is that the above optimization module is used to: determine the position point with the largest first class probability among the multiple position points included in the ground truth box as the key point.

[0026] A possible implementation is that the above optimization module is used to: optimize the object detection model with the goal of maximizing the first loss and the second loss. Among them, the second loss is used to characterize the difference between the first class probability of the key point and the first class probability of the neighborhood points of the key point.

[0027] A possible implementation is that the above optimization module is used to: determine the first class probability on the key point, determine the second class probabilities corresponding to at least one class other than the true category on the key point, respectively determine at least one first difference between the first class and the at least one second class probabilities, and use the sum of the absolute values of the at least one first difference as the first loss.

[0028] A possible implementation is that the above optimization module is used to: determine the first class probability on the key point, determine the first class probabilities of at least one neighborhood point of the key point, respectively determine at least one second difference between the first class probability and the first class probabilities of the at least one neighborhood point, and use the sum of the absolute values of the at least one second difference as the second loss.

[0029] A possible implementation is that the object detection model includes a feature extraction layer, a feature fusion layer, and an object detection layer. The above prediction module is used to: use the feature extraction layer to determine the multi-scale feature map of the sample image, use the feature fusion layer to perform feature alignment and fusion on the multi-scale feature map to obtain the target feature map, and use the object detection layer to determine the heat map corresponding to the target feature map.

[0030] A possible implementation is that the above prediction module is used to: use the feature fusion layer to perform downsampling on the multi-scale feature map through a pooling window to align the sizes of the multi-scale feature maps, and splice and fuse the aligned multi-scale feature maps to obtain the target feature map.

[0031] For the technical effects corresponding to any implementation manner in the third aspect, reference may be made to the technical effects corresponding to any implementation manner in the first aspect above, which will not be elaborated here.

[0032] In a fourth aspect, an embodiment of the present application provides an object detection device, which includes an acquisition module and a detection module.

[0033] The above-mentioned acquisition module is used to acquire an image to be detected.

[0034] The above-mentioned detection module is used to input the image to be detected into an object detection model, and predict and output an object detection result corresponding to the image to be detected through the object detection model. Among them, the object detection model is trained based on the training method of the object detection model in the first aspect above.

[0035] For the technical effects corresponding to any implementation manner in the fourth aspect, reference may be made to the technical effects corresponding to any implementation manner in the second aspect above, which will not be elaborated here.

[0036] In a fifth aspect, an embodiment of the present application provides a computing device, which performs model training based on the training method of the object detection model in the first aspect above, or performs object detection based on the object detection method in the second aspect above.

[0037] In a sixth aspect, an embodiment of the present application provides a vehicle, which includes a vehicle body and the computing device in the fifth aspect above, or performs model training by applying the training method of the object detection model in the first aspect above, or performs object detection by applying the object detection method in the second aspect above.

[0038] In a seventh aspect, a computer-readable storage medium is provided, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the training method of the object detection model in the first aspect above or the object detection method in the second aspect above.

[0039] In an eighth aspect, a computer program product is provided, which includes a computer program or instruction, and when the computer program or instruction is executed by a processor, it implements the training method of the object detection model in the first aspect above or the object detection method in the second aspect above.

[0040] The solutions provided in the fifth to eighth aspects above implement the training method of the object detection model provided in the first aspect or the object detection method provided in the second aspect, and the specific implementation will not be elaborated one by one. For the technical effects corresponding to any implementation manner in the solutions provided in the fifth to eighth aspects above, reference may be made to the technical effects corresponding to any implementation manner in the first aspect or the second aspect above, which will not be elaborated here.

[0041] It should be noted that, on the premise that the solutions do not conflict, various possible implementation manners of any one of the above aspects can be combined. Description of the Drawings

[0042] Figure 1 A schematic structural diagram of a computer system provided for an exemplary embodiment; Figure 2 A schematic flowchart of a method for training a target detection model provided for an exemplary embodiment; Figure 3 A schematic flowchart of a target detection method provided for an exemplary embodiment; Figure 4 A schematic structural diagram of a device for training a target detection model provided for an exemplary embodiment; Figure 5 A schematic structural diagram of a target detection device provided for an exemplary embodiment; Figure 6 A schematic structural diagram of a computing device provided for an exemplary embodiment. Detailed Embodiments

[0043] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0044] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0045] In the embodiments of the present application, words such as "exemplary", "such as" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary", "such as" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "such as" or "for example" is intended to present related concepts in a specific manner.

[0046] First, the related technologies involved in the present application are explained for the convenience of those skilled in the art to understand.

[0047] 3D object detection is an important research direction in the field of computer vision and is widely applied in multiple fields such as autonomous driving, robot navigation, augmented reality (AR), and medical imaging. Compared with traditional two-dimensional (2D) object detection, 3D object detection deals with objects in a three-dimensional space, involving complex features such as depth information, position, and shape. Therefore, it faces higher technical challenges, especially in object type recognition and noise interference suppression. Existing 3D object detection technologies usually rely on depth information and image data obtained from light detection and ranging (LiDAR), color-depth (RGB-Depth) cameras, or stereo vision cameras, and combine deep learning methods for object detection and localization. However, when the appearance, shape, and spatial layout differences between objects are small, the object detection methods of existing technologies are prone to misidentifying one type of object as another. For example, in sparse or occluded scenarios, the model may misjudge small objects as distant backgrounds or confuse similar types of objects. Additionally, factors such as sensor noise, environmental complexity, and insufficient data collection often lead to the generation of noise points, which usually come from sensor errors, object occlusion, or perspective changes, resulting in deviations or missed detections in object recognition. In scenes with multiple objects coexisting and complex backgrounds, it is still difficult to completely suppress the noise points. Moreover, the sizes of the target objects in 3D object detection vary greatly, and single-scale feature extraction methods cannot cover objects of all scales. Existing technologies use pyramid structures or scale-invariant convolutional neural networks (CNNs) to extract multi-scale features but cannot effectively fuse them. Finally, when the object detection model faces unseen object types, there will also be misdetections or missed detections.

[0048] Based on this, the technical solution provided by the embodiments of this application increases the discrimination ability of the object detection model between the true class and other classes by maximizing the difference between the class probabilities of the true class and other classes at the same spatial position on the heat map, enabling the object detection model to better determine the true class, significantly improving the classification accuracy of the object detection model, and being applicable to complex scenarios such as object occlusion and similar appearances.

[0049] Next, the technical solutions in the embodiments of this application will be described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments.

[0050] The solution provided by this application can be applied to Figure 1 the computer system shown in Figure 1As shown in the figure, the computer system provided by the embodiment of the present application includes a computing device 100. The computing device 100 can be a high-performance server and serve as the core of the computer system. The computing device 100 can deploy a target detection model, determine a heat map corresponding to the sample picture 110 through the target detection model, and optimize the target detection model based on the heat map and the first loss. The computing device 100 can also optimize the target detection model based on the heat map, the first loss, and the second loss. The computing device 100 can also perform target detection through the optimized trained target detection model and output a target detection result 120. The computing device 100 can also accept instructions from the staff and flexibly configure the parameters involved in the target detection model based on the instructions.

[0051] Optionally, the computing device 100 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be an embedded hardware for real-time simulation, or a cloud server that provides basic cloud computing services such as cloud computing services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data. The embodiment of the present application does not limit the implementation manner and application scenario of the computing device 100.

[0052] As Figure 2 shown in the figure, the training method of the target detection model provided by the embodiment of the present application includes: Step S201: The computing device uses the target detection model to determine a heat map corresponding to the sample picture.

[0053] Among them, the heat map is used to show the confidence of the target detection model in specific areas (such as target positions, key points, or class activation areas) in the sample picture, and usually represents the response intensity of different areas with a color gradient. In the heat map, each position point shows the class probabilities of multiple object classes, that is, the class probabilities of multiple channels of target detection. For example: sigmoid_heatmap[3,x,y]=0.9 means that there is a center point of the 3rd class (such as a vehicle) at the position (x,y), and the class probability corresponding to the third class is 0.9.

[0054] Exemplarily, the target detection model extracts and fuses multi-scale features of the sample picture and predicts and outputs a heat map based on the fused multi-scale features.

[0055] Among them, the multi-scale features are a set of features extracted from different scales (such as resolution or field of view) of the same sample picture, which can capture diverse information from local details to global structures in the sample picture.

[0056] In some embodiments, the object detection model includes a feature extraction layer, a feature fusion layer, and an object detection layer.

[0057] Among them, the feature extraction layer is used to extract multi-scale features from the sample image and determine the multi-scale feature map of the sample image. For example: the multi-scale features of the sample image are extracted through the backbone network of the object detection model to obtain the multi-scale feature map.

[0058] Exemplarily, the feature extraction layer extracts multi-scale features through a convolutional neural network. Optionally, the convolutional neural network includes a residual network (ResNet), a deep layer aggregation network (DLA), an hourglass network, and a high-resolution network (HRNet), etc.

[0059] The feature fusion layer is used to perform feature alignment and fusion on the multi-scale feature map to obtain the target feature map.

[0060] Optionally, the feature fusion layer uses a feature pyramid network (FPN) to fuse multi-scale features from top to bottom, or uses a bidirectional feature pyramid network (BiFPN) to fuse features with weights.

[0061] In some embodiments, the size alignment of the multi-scale feature map is achieved through a max pooling operation, and the aligned multi-scale feature maps are concatenated along the channel dimension to generate the target feature map.

[0062] Exemplarily, using the feature fusion layer, the multi-scale feature map is downsampled through a pooling window to align the size of the multi-scale feature map, and the most significant feature information is retained during the pooling process to avoid information loss.

[0063] Exemplarily, the aligned multi-scale feature maps (such as n feature maps: ) are concatenated, and the concatenated multi-scale feature maps are fused through a convolutional layer to extract semantic information and generate the target feature map.

[0064] The object detection layer is used to determine the heat map corresponding to the target feature map.

[0065] Exemplarily, the target feature map is convolved through the object detection layer. For example: the number of channels is adjusted through 1 to 3 lightweight convolutions to generate a preliminary heat map response, and the value range of the preliminary heat map response is compressed to [0,1] to represent the class probability to determine the heat map.

[0066] In some embodiments, the target offset and width-height are synchronously predicted and generated through the target detection layer, and prediction boxes are generated by combining the class probability, the target offset, the width-height, etc.

[0067] Among them, the prediction box is a rectangular box predicted and generated by the target detection model, which is used to frame the target to be recognized in the sample picture and represents the class of the target to be recognized.

[0068] Step S202: The computing device determines key points from multiple position points included in the ground truth box of the target to be recognized based on the true class of the target to be recognized in the heat map.

[0069] Among them, the key points are used to accurately mark the parts of the target to be recognized within the ground truth box. For example: eyes and nose in a face, headlights and license plate center of a vehicle, etc.

[0070] Exemplarily, the position point with the highest first class probability among the multiple position points included in the ground truth box is determined as the key point.

[0071] Among them, the ground truth box is a rectangular box manually labeled to accurately mark the target in the sample picture.

[0072] The first class probability refers to the class probability corresponding to the true class of the target to be recognized on the heat map. For example: the true class of the target to be recognized within the ground truth box is a vehicle. Compare the class probabilities corresponding to the vehicle at each spatial position within the ground truth box, and determine the spatial position with the highest class probability corresponding to the vehicle as the key point.

[0073] Step S203: The computing device optimizes the target detection model with the goal of maximizing the first loss.

[0074] Among them, the first loss is used to characterize the difference between the first class probability corresponding to the true class and the second class probability not corresponding to the true class among the key points.

[0075] The second class probability refers to the class probability corresponding to the class other than the true class among the key points. For example: the classes of the key points include vehicle, road, and road sign, and the true class is vehicle. The first loss is used to characterize the difference between the first class probability corresponding to the vehicle and the second class probabilities corresponding to the road and road sign.

[0076] Exemplarily, determine the first class probability on the key point, and determine the second class probability corresponding to at least one class other than the true class on the key point. Respectively determine at least one first difference between the first class and the at least one second class probability, and take the sum of the absolute values of the at least one first difference as the first loss.

[0077] For example, within the prediction box (pred_box), the key point true class The corresponding first-class probability is pred_box , If it represents the class (channel) index, the first loss function can be expressed as the following formula: .

[0078] Among them, represents the first loss, represents the true class on the key point the corresponding first-class probability, represents the second-class probability corresponding to the class other than the true class on the key point .

[0079] In some embodiments, the negative value of the first loss is taken to optimize the object detection model with the goal of minimizing the first loss.

[0080] In some embodiments, the object detection model is optimized with the goal of maximizing the first loss and the second loss.

[0081] Among them, the second loss is used to characterize the difference between the first-class probability of the key point and the first-class probability of the neighboring points of the key point.

[0082] Among them, the neighboring points refer to other pixel points or feature points within a local area (such as a circular or rectangular window) centered on the key point. For example: the neighboring points include the four points on the diagonal of the rectangular area centered on the key point, and the four points in the up, down, left, and right directions of the key point.

[0083] Optionally, the distance between the domain point and the key point is the default value, or it is a manually set value, which can be flexibly adjusted according to specific requirements.

[0084] Exemplarily, the first-class probability on the key point is determined, and the first-class probability of at least one neighboring point of the key point is determined. At least one second difference between the first-class probability and the first-class probability of at least one neighboring point is determined, and the sum of the absolute values of at least one second difference is used as the second loss.

[0085] For example, within the prediction box (pred_box), the key point has the first-class probability of pred_box , and the key point has 8 neighboring points, then the second loss can be expressed as the following formula: .

[0086] Among them, represents the second loss, Represents the first-class probability at the key point, Represents the first-class probability at the j-th neighboring point of the key point.

[0087] In some embodiments, the second loss is taken as a negative value, and the object detection model is optimized with the goal of minimizing the second loss.

[0088] In some embodiments, the prediction box parameters of the object detection model are adjusted to maximize the first loss and / or the second loss. Specifically, through the chain rule, the gradient of the prediction box parameters of the object detection model is passed to the convolutional network of the object detection model to update the convolutional kernel weights and achieve the optimization of the object detection model.

[0089] In some embodiments, during the optimization training process of the object detection model, losses such as the first loss, the second loss, the classification loss, and the regression loss are optimized simultaneously.

[0090] Exemplarily, through the backpropagation algorithm, the gradients of the first loss, the second loss, the classification loss, and the regression loss are accumulated and passed to each layer of the network of the object detection model to guide the update of the parameters of the object detection model, so that the object detection model can accurately distinguish the target categories to be recognized, reduce type misdetection, suppress noise, and improve the detection accuracy of the object detection model.

[0091] As Figure 3 shown, the object detection method provided by the embodiments of the present application includes: Step S301: The computing device acquires the image to be detected.

[0092] Step S302: The computing device inputs the image to be detected into the object detection model, and the object detection model predicts and outputs the object detection result corresponding to the image to be detected.

[0093] Among them, the object detection model is trained based on the training method of the above object detection model.

[0094] Exemplarily, the object detection result is a heat map predicted and generated based on the image to be detected, and different category objects in the image to be detected are labeled by the heat map in terms of region and category.

[0095] In summary, the technical solution provided by the embodiments of the present application, during the training process of the object detection model, by introducing and maximizing the first loss and / or the second loss, increases the difference between the true type and other types at the key points, and increases the difference between the key points and the neighboring points, effectively solving the problems of type misdetection and noise interference. At the same time, when the object detection model processes the image, it performs alignment and fusion of multi-scale features to solve the multi-scale problem, significantly improving the accuracy, robustness, and generalization ability of the object detection model.

[0096] As Figure 4As shown in the figure, the training device for the object detection model provided by the present application may include a prediction module 401 and an optimization module 402. Among them, the prediction module 401 is used to execute Figure 2 the operation of step S201 in the method illustrated, and the optimization module 402 is used to execute Figure 2 the operations of step S202 and step S203 in the method illustrated.

[0097] As Figure 5 shown in the figure, the object detection device provided by the present application may include an acquisition module 501 and a detection module 502. Among them, the acquisition module 501 is used to execute Figure 3 the operation of step S301 in the method illustrated, and the detection module 502 is used to execute Figure 3 the operation of step S302 in the method illustrated.

[0098] The above mainly introduces the solutions provided by the embodiments of the present application from the perspective of methods. To implement the above functions, the video quality assessment device or computing device includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0099] The embodiments of the present application can, according to the above video quality assessment method, exemplarily divide the functional modules of the video quality assessment device or computing device. For example, the video quality assessment system or computing device may include each functional module corresponding to each functional division, or two or more functions may be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0100] As Figure 6 shown in the figure, the computing device provided by the embodiments of the present application may include a processor 601, a bus 602, a communication interface 603, and a memory 604. The processor 601, the memory 604, and the communication interface 603 communicate with each other through the bus 602. It should be understood that the present application does not limit the number of processors and memories in the network device.

[0101] The bus 602 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or a Universal Serial Bus (USB), etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 6 only one line is used to represent it in Figure 6 , but it does not mean that there is only one bus or one type of bus. The bus 602 can include a path for transmitting information between various components of the network device (for example, the memory 604, the processor 601, and the communication interface 603).

[0102] The processor 601 can include any one or more of processors such as a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Micro Processor (MP), or a Digital Signal Processor (DSP).

[0103] The memory 604 can include volatile memory, such as Random Access Memory (RAM). The processor 601 can also include non-volatile memory, such as Read-Only Memory (ROM), flash memory, a Hard Disk Drive (HDD), or a Solid State Drive (SSD).

[0104] The communication interface 603 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the network device and other devices or a communication network.

[0105] The memory 604 stores executable program code, and the processor 601 executes the executable program code to respectively implement the functions of the foregoing method embodiments. That is, the memory 604 stores instructions for executing the above video quality assessment method.

[0106] The training method, object detection method, and computing device provided by the embodiments of the present application can be applied in vehicles. Vehicles can also be referred to as transportation means (vehicle), mobile carrier, electric vehicle (EV), hybrid electric vehicle (HEV), plug-in hybrid electric vehicle (PHEV), fuel cell vehicle (FCV), autonomous vehicle, intelligent and connected vehicle (ICV), driverless vehicle, etc.

[0107] In the embodiments of the present application, the vehicle can be a sedan, a sport utility vehicle (SUV), a truck, an electric vehicle, a motorcycle, a tricycle, a special vehicle (such as an ambulance, a fire truck, a police car, etc.), a driverless taxi, an intelligent and connected bus, an autonomous logistics vehicle, an electric truck, etc. In addition, the method is also applicable to various special vehicles, such as agricultural vehicles, mining vehicles, forestry vehicles, airport vehicles, port vehicles, etc. The present application does not make specific limitations in this regard.

[0108] The training method, object detection method, and computing device provided by the embodiments of the present application can also be applied to fields such as autonomous driving, robot navigation, and medical image processing, and have broad application prospects and high practical value.

[0109] The embodiments of the present application also provide a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the training method or object detection method of the object detection model provided by the above method embodiments.

[0110] Optionally, the computer-readable storage medium can be a non-temporary computer-readable storage medium. For example, the non-temporary computer-readable storage medium can be ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0111] The embodiments of the present application also provide a computer program product, which includes a computer program or instruction. When the computer program or instruction is executed by a processor, the training method or object detection method of the object detection model provided by the above method embodiments is implemented.

[0112] It should be noted that when one or more instructions in the above computer-readable storage medium or in the computer program product are executed by the processor of the computing device, the various processes of the above method embodiments are implemented, and the same technical effects as the above method can be achieved. To avoid repetition, they will not be elaborated here.

[0113] From the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0114] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0115] The units described as separate components may or may not be physically separated. The components displayed as units may be a physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0116] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0117] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs.

[0118] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A training method for a target detection model, characterized in that, The method includes: Using an object detection model to determine a heat map corresponding to a sample image, where each position point in the heat map shows the class probabilities of multiple object classes; Based on the true class of the target to be recognized in the heat map, determining key points from multiple position points included in the true bounding box of the target to be recognized; Optimizing the object detection model with the goal of maximizing a first loss; Wherein, the first loss is used to characterize the difference between the first class probability corresponding to the true class among the key points and the second class probability not corresponding to the true class.

2. The training method of the object detection model according to claim 1, wherein, The determining key points from multiple position points included in the true bounding box of the target to be recognized based on the true class of the target to be recognized in the heat map includes: Determining the position point with the maximum first class probability among the multiple position points included in the true bounding box as the key point.

3. The training method of the object detection model according to claim 1, characterized in that The optimizing the object detection model with the goal of maximizing the first loss includes: Optimizing the object detection model with the goal of maximizing the first loss and a second loss; Wherein, the second loss is used to characterize the difference between the first class probability of the key point and the first class probabilities of the neighborhood points of the key point.

4. The training method of the object detection model according to claim 1, wherein The calculation process of the first loss includes: Determining the first class probability on the key point; Determining the second class probabilities corresponding to at least one class other than the true class on the key point; Respectively determining at least one first difference between the first class and the at least one second class probability, and taking the sum of the absolute values of the at least one first difference as the first loss.

5. The training method of the object detection model according to claim 3, wherein The calculation process of the second loss includes: Determining the first class probability on the key point; Determining the first class probabilities of at least one neighborhood point of the key point; Respectively determining at least one second difference between the first class probability and the first class probabilities of at least one neighborhood point, and taking the sum of the absolute values of the at least one second difference as the second loss.

6. The training method of the object detection model according to claim 1, characterized in that The object detection model includes a feature extraction layer, a feature fusion layer, and an object detection layer; The using an object detection model to determine a heat map corresponding to a sample image includes: Using the feature extraction layer to determine multi-scale feature maps of the sample image; Using the feature fusion layer to perform feature alignment and fusion on the multi-scale feature maps to obtain a target feature map; Using the object detection layer to determine the heat map corresponding to the target feature map.

7. The training method of the object detection model according to claim 6, characterized in that, The using the feature fusion layer to perform feature alignment and fusion on the multi-scale feature maps to obtain a target feature map includes: Using the feature fusion layer to perform downsampling on the multi-scale feature maps through a pooling window to align the sizes of the multi-scale feature maps; Performing splicing and fusion on the aligned multi-scale feature maps to obtain the target feature map.

8. A target detection method, characterized in that, The object detection method includes: Obtaining an image to be detected; Inputting the image to be detected into an object detection model, and predicting and outputting an object detection result corresponding to the image to be detected through the object detection model; wherein, the object detection model is trained according to the object detection model training method according to any one of claims 1 to 7.

9. A training device for an object detection model, characterized in that, The device includes: A prediction module, configured to use an object detection model to determine a heat map corresponding to a sample image, wherein each position point in the heat map shows the class probabilities of multiple object classes; An optimization module, configured to determine key points from multiple position points included in the ground truth box of the target to be recognized based on the true class of the target to be recognized in the heat map; The optimization module is further configured to optimize the object detection model with the goal of maximizing a first loss; Wherein the first loss is used to characterize the difference between a first class probability corresponding to the true class and a second class probability not corresponding to the true class among the key points.

10. A target detection device, characterized in that, The device includes: An acquisition module, configured to acquire an image to be detected; A detection module, configured to input the image to be detected into an object detection model, and predict and output an object detection result corresponding to the image to be detected through the object detection model; wherein, the object detection model is trained by the training method of the object detection model according to any one of claims 1 to 7.

11. A computing device, characterized in that, The computing device trains an object detection model by using the training method of the object detection model according to any one of claims 1 to 7, or performs object detection by using the object detection method according to claim 8.

12. A vehicle, characterized in that, The vehicle includes a vehicle body and the computing device according to claim 11; or trains an object detection model by using the training method of the object detection model according to any one of claims 1 to 7; or performs object detection by using the object detection method according to claim 8.

Citation Information

Patent Citations

  • Systems and methods for mitigation bias in machine learning model output

    CA3111911A1

  • Key point detection method and device and computer readable storage medium

    CN111223143A

  • Age prediction model training method and device, equipment and storage medium

    CN115293260A

  • Image detection method and device

    CN117593738A

  • Model training method and device based on pre-annotated data, electronic equipment and medium

    CN118821970A