Dual-head decoupling alignment full-scene target detection method, system, device and medium
Through the full-scene object detection method with double-headed decoupling and alignment, the double-headed monocular object detection model and improved Soft-NMS function are used to solve the missed detection problem of monocular 3D object detection in harsh scenarios, and improve the target recognition accuracy and reliability of the autonomous driving system.
Patent Information
- Application Number
- CN202211170474.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-09-22
AI Technical Summary
The existing monocular 3D object detection method has reduced detection accuracy in harsh scenarios such as long-distance targets, occlusion targets, and truncated targets, resulting in missed detection problems and affecting the practicality and reliability of the autonomous driving system.
The full-scene object detection method with double-headed decoupling and alignment is adopted. By obtaining the RGB image of the vehicle-mounted monocular camera, pre-processing and post-processing, the preset double-headed monocular object detection model is used for detection. The model structure includes a feature extraction network and a double-headed object detection network, and the improved Soft-NMS function is used to filter and align the detection results.
It improves the accuracy and reliability of the autonomous driving system in the recognition of all scenario targets, reduces the missed detection rate in harsh scenarios, and improves the real-time and detection performance of the model.
Smart Images

Figure CN115410181B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of autonomous driving and computer vision, and in particular to a full-scene target detection method, system, device and medium with dual-head decoupling alignment. Background Art
[0002] With the development of technologies such as artificial intelligence and big data analysis, the level of automobile autonomous driving has been continuously improved, which has greatly facilitated people's travel and is expected to reduce safety hazards such as fatigue driving and drunk driving. The "Automobile Driving Automation Classification" standard points out that starting from the third level of conditional automated driving, the automobile system should have the function of performing all dynamic driving tasks, that is, it should have environmental perception and decision-making control functions. Among them, the environmental perception function requires the system to use on-board sensors and vehicle networking to accurately and quickly obtain the position information of vehicles around the road. Therefore, road vehicle target detection technology for autonomous driving has received extensive attention and research.
[0003] According to the type of data used, existing autonomous driving target detection methods can be divided into three categories: radar point cloud-based, binocular stereo image-based, and monocular RGB image-based. Compared with pure visual methods, the target positioning accuracy of target detection methods using laser radar is high, and depth information can be directly obtained. For example, Chinese patent CN109597087B discloses a 3D target detection method based on point cloud data, which uses a deep convolutional neural network to extract the fusion perception and recognition of the target of interest in the point cloud data and image data; the binocular method uses the left and right camera images as model input to infer target information, such as Chinese patent CN114332790A discloses a binocular vision 3D target detection method, which uses a stereo matching algorithm to extract the parallax information of the left and right images for depth estimation; the monocular method directly uses a single RGB image as input, which is easier to achieve real-time detection of road targets, such as Chinese patent CN111369617A discloses a 3D target detection method based on a monocular view of a convolutional neural network.
[0004] Among the above methods, factors such as the reduced resolution of long-distance targets, high prices, and large computing power requirements of LiDAR limit the application of point cloud methods; pure vision methods only use binocular or monocular cameras, which have the advantages of high cost-effectiveness and high frame rate of sensors, and the target detection algorithm has lower requirements for the process, installation, and calibration of monocular cameras than binocular cameras. However, it is difficult to improve the detection accuracy of pure vision methods, and it is even more challenging for monocular methods. Depth estimation of monocular images is an ill-posed problem, especially in harsh scenarios such as objects are far away from the camera, objects occlude each other, and objects appear at the edge of the field of view. The accuracy of existing monocular 3D target detection methods for object pose estimation is significantly reduced, which affects the practicality and reliability of autonomous driving systems. Summary of the invention
[0005] The purpose of the present invention is to provide a dual-head decoupled and aligned full-scene target detection method, system, device and medium to solve the problem of missed detection of existing monocular 3D target detection methods in harsh scenarios such as long-distance targets, occluded targets, and truncated targets. The present invention can improve the full-scene target recognition accuracy and reliability of the autonomous driving system.
[0006] In order to achieve the above object, the present invention adopts the following technical scheme:
[0007] The full-scene object detection method with dual-head decoupling alignment includes the following steps:
[0008] SI: Get the original RGB image to be detected taken by the on-board monocular camera in real time;
[0009] S2: preprocessing the original RGB image to be detected to obtain an adaptively scaled image;
[0010] S3: inputting the adaptively scaled image into a preset dual-head monocular target detection model to obtain redundant target parameter prediction values output by the preset dual-head monocular target detection model;
[0011] S4: Post-process the redundant target parameter prediction values to obtain a high-confidence detection result of the original RGB image to be detected.
[0012] Furthermore, the S2 specifically includes:
[0013] S2.1: performing edge filling on the original RGB image to be detected to obtain a high-resolution image; wherein the resolution of the original RGB image to be detected is not higher than the resolution of the high-resolution image;
[0014] S2.2: Perform RGB value normalization processing on the high-resolution image to obtain an adaptively scaled image.
[0015] Furthermore, the dual-head monocular target detection model in S3 is pre-set through two steps of model structure construction and weight offline training;
[0016] The model structure includes: a feature extraction network and a dual-head target detection network, wherein the backbone of the feature extraction network is a DLA-34 deep aggregation network, and the DLA-34 deep aggregation network uses deformable convolution to extract features of targets of interest; the dual-head target detection network includes a feature transition network and two detection heads, and the feature transition network combines Ghost convolution with deformable convolution. The structures of the two detection heads both include a convolution layer, a batch normalization layer, an activation function layer and a convolution layer, and the two detection heads use different prediction methods for target attribute parameters.
[0017] Furthermore, the target attribute parameters include target depth, target center and target posture, and the two detection heads use different prediction methods for the target attribute parameters, specifically including:
[0018] In terms of target depth, the two detection heads use mean variance prediction and exponential prediction respectively;
[0019] In terms of target center, the two detection heads use two-dimensional center prediction method and three-dimensional projection center prediction method respectively;
[0020] In terms of target posture, the two detection heads use direct prediction and MultiBin discrete prediction respectively.
[0021] Furthermore, the weight offline training is specifically as follows:
[0022] The parameters of the feature extraction network and the dual-head target detection network are jointly trained using the historical data set and the public data set to obtain a preset dual-head monocular target detection model, wherein the loss function used for joint training is as follows:
[0023]
[0024] Where L is the joint training loss, I is the adaptive scaling image, i represents the detection head number of the loss, i = 1, 2, L i,kpt , L i,3D , L i,2D are the key point loss, 3D box loss, and 2D box loss of the prediction result of the i-th detection head, respectively. dis is the parameter decoupled alignment loss of the two detection head outputs, φ f ,φ i are the learnable parameters of the feature extraction network and the i-th detection head, respectively.
[0025] Furthermore, obtaining the redundant target parameter prediction value output by the preset dual-head monocular target detection model in S3 includes:
[0026] The two detection heads of the preset dual-head monocular target detection model both output redundant prediction values of the category, size, posture and position of the full-scene target in the adaptively scaled image.
[0027] Furthermore, the S4 specifically includes:
[0028] S4.1: Using the internal parameters of the on-board monocular camera, the redundant target parameter prediction values output by the preset dual-head monocular target detection model are projected and transformed to obtain the redundant prediction values of the category, size, posture and position of the full scene target in the original RGB image coordinate system to be detected;
[0029] S4.2: The improved Soft-NMS function is used to filter the redundant prediction values of the size, posture and position of the full scene target in the coordinate system of the original RGB image to be detected, and the redundant target parameter prediction values with confidence lower than the preset value are filtered out to obtain the high confidence detection result of the original RGB image to be detected. The expression of the improved Soft-NMS function is as follows:
[0030]
[0031] Among them, s i is the confidence score of the detection result of the i-th full-scene target in the original RGB image to be detected, B M , B i are the maximum confidence target 3D projection frame and the i-th target 3D projection frame, respectively, z M 、z i are the maximum confidence target depth and the i-th target depth, τ z is the target depth threshold, σ and γ are both constants, and IoU(·, ·) is the intersection over union function of the 3D projection box.
[0032] The full-scene object detection system with dual-head decoupling and alignment includes:
[0033] The data acquisition module is used to obtain the original RGB image to be detected taken by the vehicle-mounted monocular camera in real time;
[0034] A preprocessing module, used for preprocessing the original RGB image to be detected to obtain an adaptively scaled image;
[0035] A prediction module, used for inputting the adaptively scaled image into a preset dual-head monocular target detection model to obtain a redundant target parameter prediction value output by the preset dual-head monocular target detection model;
[0036] The post-processing module is used to post-process the redundant target parameter prediction values to obtain a high-quality detection result of the original RGB image to be detected.
[0037] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the full-scene target detection method of dual-head decoupling alignment when executing the computer program.
[0038] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the full-scene target detection method with dual-head decoupling and alignment.
[0039] Compared with the prior art, the present invention has the following beneficial technical effects:
[0040] 1) Using monocular RGB images as input and a single-stage 3D object detection model improves the model reasoning speed and the real-time performance of the autonomous driving system;
[0041] 2) In order to promote the convergence of the parameters of the dual-head monocular target detection model when training on a large data set, a joint training loss is used to optimize both the two-dimensional and three-dimensional parameters of the target parameters, and the output results of the dual detection heads are aligned in a parameter decoupling manner, ensuring the performance of the optimized detection model;
[0042] 3) In order to overcome the problem of degraded target detection performance in harsh scenarios, a dual detection head setting is adopted to predict the target to be detected from different aspects, and the improved Soft-NMS is used to align the detection results, which greatly reduces the missed detection rate of full-scene target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings in the specification are used to provide further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0044] Figure 1 A schematic diagram of a flow chart of a full-scene target detection method with dual-head decoupling alignment provided by an embodiment of the present invention;
[0045] Figure 2 A structural diagram of a dual-head monocular target detection model for a full-scene target detection method with dual-head decoupling alignment provided by an embodiment of the present invention;
[0046] Figure 3 An improved Soft-NMS principle diagram of the full-scene target detection method with dual-head decoupling alignment provided in an embodiment of the present invention;
[0047] Figure 4 A schematic diagram of the structure of a full-scenario target detection system provided by an embodiment of the present invention;
[0048] Figure 5 A diagram of detection results of the full-scene target detection method with dual-head decoupling alignment provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0049] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0050] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0051] The technical solution of the embodiments of the present application is described in detail below with reference to the accompanying drawings.
[0052] S1: Obtain the original RGB image to be detected taken by the on-board monocular camera in real time.
[0053] In actual applications, the car's autonomous driving system uses the on-board monocular camera to obtain real-time visual information of all-scene targets such as pedestrians, vehicles, obstacles on the road around the vehicle, etc., inputs the target detection model for detection, and controls the car's driving state according to parameter information such as target type, position, and posture to achieve the purpose of autonomous driving.
[0054] S2: Preprocess the original RGB image to be detected to obtain an adaptively scaled image.
[0055] The specific steps are:
[0056] S2.1: performing edge filling on the original RGB image to be detected to obtain a high-resolution image; wherein the resolution of the original RGB image to be detected is not higher than the resolution of the high-resolution image;
[0057] S2.2: Perform RGB value normalization processing on the high-resolution image to obtain an adaptively scaled image.
[0058] This step uniformly adjusts the resolution of the original image to be detected and normalizes the RGB values to reduce the model's sensitivity to scene changes and facilitate the model to batch process real-time road images.
[0059] S3: Input the adaptively scaled image into a preset dual-head monocular target detection model to obtain a target parameter prediction value output by the preset dual-head monocular target detection model.
[0060] The preset dual-head monocular target detection model is pre-set through two steps: model structure construction and weight offline training.
[0061] like Figure 2 As shown, the preset dual-head monocular target detection model in step S3 is a single-stage monocular 3D target detection model, and its structure includes: a feature extraction network and a dual-head target detection network, wherein the backbone of the feature extraction network is a DLA-34 deep aggregation network, and the DLA-34 deep aggregation network uses deformable convolution to extract the features of the target of interest; the dual-head target detection network includes a feature transition network and two detection heads, and the feature transition network combines Ghost convolution with deformable convolution. The two detection heads structurally include a convolution layer, a batch normalization layer, an activation function layer and a convolution layer, and the two detection heads use different prediction methods for target depth, target center and target posture parameters, specifically including: in terms of target depth, the two detection heads respectively use a mean variance prediction method and an exponential prediction method; in terms of target center, the two detection heads respectively use a two-dimensional center prediction method and a three-dimensional projection center prediction method; in terms of target posture, the two detection heads respectively use a direct prediction method and a MultiBin discrete prediction method.
[0062] The preset dual-head monocular target detection model determines its weight parameters by offline training, and the training method is as follows:
[0063] The feature extraction network and the dual-head target detection network parameters are jointly trained using historical data sets and public data sets to obtain a preset dual-head monocular target detection model, wherein the loss function used for joint training is as follows:
[0064]
[0065] Where L is the joint training loss, I is the adaptive scaling image, i represents the detection head number of the loss, i = 1, 2, L i,kpt , L i,3D , L i,2D are the key point loss, 3D box loss, and 2D box loss of the prediction result of the i-th detection head, respectively. dis is the parameter decoupled alignment loss of the two detection head outputs, φ f ,φ i are the learnable parameters of the feature extraction network and the i-th detection head, respectively.
[0066] The step S3 of obtaining the redundant target parameter prediction value output by the dual-head monocular target detection model includes:
[0067] The two detection heads of the preset dual-head monocular target detection model both output redundant prediction values of the category, size, posture and position of the full-scene target in the adaptively scaled image.
[0068] S4: Post-process the redundant target parameter prediction values to obtain a high-confidence detection result of the original RGB image to be detected.
[0069] The specific steps of step S4 include:
[0070] S4.1: Using the internal parameters of the on-board monocular camera, the redundant target parameter prediction values output by the preset dual-head monocular target detection model are projected and transformed to obtain the redundant prediction values of the category, size, posture and position of the full scene target in the original RGB image coordinate system to be detected;
[0071] S4.2: The improved Soft-NMS function is used to filter the redundant prediction values of the size, posture, and position of the full scene target in the original RGB image coordinate system to be detected, and the redundant target parameter prediction values with confidence lower than the preset value (0.3) are filtered out to obtain high-quality detection results of the original RGB image to be detected. The expression of the improved Soft-NMS function is as follows:
[0072]
[0073] Among them, s i is the confidence score of the detection result of the i-th full-scene target in the original RGB image to be detected, B M , B i are the maximum confidence target 3D projection frame and the i-th target 3D projection frame, respectively, z M 、z i are the maximum confidence target depth and the i-th target depth, τ z is the target depth threshold, σ and γ are both constants, and IoU(·, ·) is the intersection over union function of the 3D projection box.
[0074] like Figure 3 As shown in Figure 1, the purpose of improving the Soff-NMS function is to retain the effective detection results of the dual detection heads, delete the redundant results, and complement the missed detection results, that is, to retain the detection results of targets with a long distance, filter out the redundant detection results of the same target, and retain the detection results of targets with a close distance.
[0075] Corresponding to the aforementioned application function implementation method embodiment, the present application also provides a full-scenario target detection system for autonomous driving and corresponding embodiments.
[0076] Figure 4 The schematic diagram of the structure of the full-scene target detection system provided by the present invention includes:
[0077] The data acquisition module is used to obtain the original RGB image to be detected taken by the vehicle-mounted monocular camera in real time;
[0078] A preprocessing module, used for preprocessing the original RGB image to be detected to obtain an adaptively scaled image;
[0079] A prediction module, used for inputting the adaptively scaled image into a preset dual-head monocular target detection model to obtain a redundant target parameter prediction value output by the preset dual-head monocular target detection model;
[0080] The post-processing module is used to post-process the redundant target parameter prediction values to obtain a high-quality detection result of the original RGB image to be detected.
[0081] The specific execution operation mode of each module in the system has been described in detail in the embodiment of the method provided by the present invention, and will not be described in detail here.
[0082] In order to further demonstrate the significant substantial effects of the present invention, the present invention is further described in detail below in conjunction with specific embodiments:
[0083] In this embodiment, the target detection method provided by the present invention is compared with typical monocular 3D target detection methods such as MonoFlex, RTM3D, and MonoDIS, and the KITTI dataset widely used in the field of 3D target detection is used as a pair to verify the method. The evaluation indicators are the average accuracy of 3D detection AP|3D and the average accuracy of position prediction AP|BEV. The intersection-over-union ratio threshold of the target detection result is 0.7. The detection object is a panoramic road vehicle target with three detection difficulties (simple, medium, and difficult) divided by the KITTI dataset. The results are shown in Table 1.
[0084] Table 1 Road vehicle target detection results
[0085]
[0086]
[0087] It can be seen that compared with the existing methods, the method provided by the present invention has improved the average accuracy of 3D detection and the average accuracy of position prediction in three difficult scenarios. Figure 5 The detection result diagram of the method provided by the present invention is shown. Each row shows a detection scene. The first two columns are the results of a single detection head. After Soft-NMS alignment, the problems of missed detection, redundant detection, low-quality detection, etc. in harsh scenarios are solved, indicating the effectiveness of the method provided by the present invention in full-scene road vehicle target detection.
[0088] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0089] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0090] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit its protection scope. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that after reading the present invention, those skilled in the art can still make various changes, modifications or equivalent substitutions to the specific implementation methods of the invention, but these changes, modifications or equivalent substitutions are all within the protection scope of the pending claims of the invention.
Claims
1. The full-scene target detection method with dual-head decoupling alignment is characterized by: The following steps are involved: S1: Obtain the original RGB image to be detected taken by the on-board monocular camera in real time; S2: preprocessing the original RGB image to be detected to obtain an adaptively scaled image; S3: inputting the adaptively scaled image into a preset dual-head monocular target detection model to obtain redundant target parameter prediction values output by the preset dual-head monocular target detection model; The dual-head monocular target detection model in S3 is pre-set through two steps: model structure construction and weight offline training; The model structure includes: a feature extraction network and a dual-head target detection network, wherein the backbone of the feature extraction network is a DLA-34 deep aggregation network, and the DLA-34 deep aggregation network uses deformable convolution to extract features of the target of interest; the dual-head target detection network includes a feature transition network and two detection heads, the feature transition network combines Ghost convolution with deformable convolution, the structures of the two detection heads both include a convolution layer, a batch normalization layer, an activation function layer and a convolution layer, and the two detection heads use different prediction methods for target attribute parameters; S4: Post-processing the redundant target parameter prediction values to obtain a high-confidence detection result of the original RGB image to be detected; specifically including: S4.1: Using the internal parameters of the on-board monocular camera, the redundant target parameter prediction values output by the preset dual-head monocular target detection model are projected and transformed to obtain the redundant prediction values of the category, size, posture and position of the full scene target in the original RGB image coordinate system to be detected; S4.2: The improved Soft-NMS function is used to filter the redundant prediction values of the size, posture and position of the full scene target in the coordinate system of the original RGB image to be detected, and the redundant target parameter prediction values with confidence lower than the preset value are filtered out to obtain the high confidence detection result of the original RGB image to be detected. The expression of the improved Soft-NMS function is as follows: in, is the first The confidence score of the detection results of all scene targets, , are the maximum confidence target 3D projection box and the A target 3D projection frame, , are the maximum confidence target depth and Target depth, is the target depth threshold, , are constants, is the intersection-and-union function of the three-dimensional projection box.
2. The full-scene target detection method with dual-head decoupling alignment according to claim 1 is characterized in that: The S2 specifically includes: S2.1: performing edge filling on the original RGB image to be detected to obtain a high-resolution image; wherein the resolution of the original RGB image to be detected is not higher than the resolution of the high-resolution image; S2.2: Perform RGB value normalization processing on the high-resolution image to obtain an adaptively scaled image.
3. The full-scene target detection method with dual-head decoupling alignment according to claim 1 is characterized in that: The target attribute parameters include target depth, target center and target posture. The two detection heads use different prediction methods for the target attribute parameters, specifically including: In terms of target depth, the two detection heads use mean variance prediction and exponential prediction respectively; In terms of target center, the two detection heads use two-dimensional center prediction method and three-dimensional projection center prediction method respectively; In terms of target posture, the two detection heads use direct prediction and MultiBin discrete prediction respectively.
4. The full-scene target detection method with dual-head decoupling alignment according to claim 1 is characterized in that: The weight offline training is specifically as follows: The parameters of the feature extraction network and the dual-head target detection network are jointly trained using the historical data set and the public data set to obtain a preset dual-head monocular target detection model, wherein the loss function used for joint training is as follows: in, is the joint training loss, To adaptively scale images, Indicates the number of the detection head for the loss being sought, , , , Respectively The key point loss, 3D box loss, and 2D box loss of the detection head prediction results are Decouple the alignment loss for the parameters of the two detection head outputs, , They are the feature extraction network and the The learnable parameters of the detection head.
5. The full-scene target detection method with dual-head decoupling alignment according to claim 1 is characterized in that: The redundant target parameter prediction values output by the preset dual-head monocular target detection model are obtained in S3, including: The two detection heads of the preset dual-head monocular target detection model both output redundant prediction values of the category, size, posture and position of the full-scene target in the adaptively scaled image.
6. Dual-head decoupling and alignment full-scene object detection system, characterized by: include: The data acquisition module is used to obtain the original RGB image to be detected taken by the vehicle-mounted monocular camera in real time; A preprocessing module, used for preprocessing the original RGB image to be detected to obtain an adaptively scaled image; A prediction module, used for inputting the adaptively scaled image into a preset dual-head monocular target detection model to obtain a redundant target parameter prediction value output by the preset dual-head monocular target detection model; The dual-head monocular target detection model is pre-set through two steps: model structure construction and weight offline training; The model structure includes: a feature extraction network and a dual-head target detection network, wherein the backbone of the feature extraction network is a DLA-34 deep aggregation network, and the DLA-34 deep aggregation network uses deformable convolution to extract features of the target of interest; the dual-head target detection network includes a feature transition network and two detection heads, the feature transition network combines Ghost convolution with deformable convolution, the structures of the two detection heads both include a convolution layer, a batch normalization layer, an activation function layer and a convolution layer, and the two detection heads use different prediction methods for target attribute parameters; A post-processing module is used to post-process the redundant target parameter prediction values to obtain a high-quality detection result of the original RGB image to be detected; specifically, it includes: S4.1: Using the internal parameters of the on-board monocular camera, the redundant target parameter prediction values output by the preset dual-head monocular target detection model are projected and transformed to obtain the redundant prediction values of the category, size, posture and position of the full scene target in the original RGB image coordinate system to be detected; S4.2: The improved Soft-NMS function is used to filter the redundant prediction values of the size, posture and position of the full scene target in the coordinate system of the original RGB image to be detected, and the redundant target parameter prediction values with confidence lower than the preset value are filtered out to obtain the high confidence detection result of the original RGB image to be detected. The expression of the improved Soft-NMS function is as follows: in, is the first The confidence score of the detection results of all scene targets, , are the maximum confidence target 3D projection box and the A target 3D projection frame, , are the maximum confidence target depth and Target depth, is the target depth threshold, , are constants, is the intersection-and-union function of the three-dimensional projection box.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the full-scene target detection method with dual-head decoupling alignment are implemented as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the full-scene target detection method with dual-head decoupling alignment are implemented as described in any one of claims 1 to 5.
Citation Information
Patent Citations
A 3D target detection method based on point cloud data
CN109597087B
3D target detection method of monocular view based on convolutional neural network
CN111369617A
Binocular vision 3D target detection method, storage medium and terminal equipment
CN114332790A
Wide area covering narrow area multi-point key point reconnaissance detection technology
CN108181666A
Neural network based on fixed-point operation
CN108345939A