Target detection method and device, electronic equipment, storage medium and vehicle

By using an end-to-end multi-task target detection network, shared multi-scale target features are extracted and fused, solving the problem of two-dimensional and three-dimensional target recognition in fisheye image recognition and achieving efficient recognition of rich targets in parking scenarios.

CN121214404APending Publication Date: 2025-12-26CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511585942.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing technologies struggle to simultaneously achieve effective two-dimensional and three-dimensional recognition of a wide range of targets in parking scenarios using fisheye image recognition, and they also consume significant computational resources.

Method used

An end-to-end multi-task target detection network is adopted. By acquiring fisheye images without distortion correction, shared multi-scale target features are extracted and feature fusion is performed. Two-dimensional and three-dimensional target features are analyzed separately to achieve two-dimensional and three-dimensional target detection.

Benefits of technology

It improves the ability to identify a variety of targets in parking scenarios, reduces computational resource consumption, and enables effective identification of two-dimensional and three-dimensional targets around the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121214404A_ABST
    Figure CN121214404A_ABST
Patent Text Reader

Abstract

The invention relates to a target detection method and apparatus, an electronic device, a storage medium and a vehicle. The method comprises the steps of obtaining a to-be-detected fisheye image; extracting a plurality of shared target features of different scales from the to-be-detected fisheye image, wherein the shared target features are used for carrying out a two-dimensional target detection task and a three-dimensional target detection task; carrying out feature fusion on the shared target features of different scales to obtain two-dimensional task fusion features; extracting a plurality of three-dimensional target features of different scales from the shared multi-scale target features; performing feature fusion on the shared target features and the three-dimensional target features to obtain three-dimensional task fusion features; extracting two-dimensional target features from the two-dimensional task fusion features; extracting three-dimensional target features from the three-dimensional task fusion features; and respectively analyzing the two-dimensional target features and the three-dimensional target features to obtain corresponding target detection results. According to the method, the detection performance of two-dimensional and three-dimensional targets can be guaranteed, so that relatively complete parking scene environment sensing information can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of parking assistance technology, specifically to a target detection method, device, electronic equipment, storage medium, and vehicle. Background Technology

[0002] Fisheye image recognition boasts a wide field of view, enabling the identification of nearby targets around the vehicle. Typically, only four fisheye cameras are needed to completely cover the area surrounding the vehicle. Therefore, fisheye image recognition has become the mainstream technology for parking assistance and driver perception.

[0003] In related technologies, a single-task recognition architecture is typically employed. This involves using a 3D (three-dimensional) object detection network to directly predict 3D object information from undistorted fisheye images and to convert distorted graphic features into bird's-eye view features for 3D object recognition. Directly predicting 3D object information from undistorted fisheye images ensures no loss of visibility, but learning object features is challenging. Converting distorted graphic features into bird's-eye view features for 3D object recognition makes object features easier to learn, but it reduces the visibility range. Furthermore, both methods lack the ability to detect 2D (two-dimensional) objects in parking scenarios, requiring a separate 2D object detection model to compensate for other object recognition information in parking scenes.

[0004] Therefore, how to effectively identify a variety of targets in parking scenarios has become an urgent problem to be solved. Summary of the Invention

[0005] This application provides a target detection method to address the problem in the prior art of effectively identifying a wide range of targets in parking scenarios.

[0006] Accordingly, embodiments of this application also provide a target detection device, an electronic device, a storage medium, and a vehicle to ensure the implementation and application of the above methods.

[0007] To address the aforementioned problems, this application discloses a target detection method applied to an electronic device. The electronic device includes an end-to-end multi-task target detection network, where the tasks include two-dimensional target detection tasks and three-dimensional target detection tasks. The method includes: Acquire the fisheye image to be detected; Shared multi-scale target features are extracted from the fisheye image to be detected; the shared multi-scale target features include multiple shared target features of different scales, which are used for two-dimensional target detection tasks and three-dimensional target detection tasks. Feature fusion is performed on the shared target features at multiple different scales to obtain two-dimensional task fusion features; Three-dimensional multi-scale target features are extracted from the shared multi-scale target features; the three-dimensional multi-scale target features include multiple three-dimensional target features at different scales; The shared target features and the three-dimensional target features are fused to obtain three-dimensional task fusion features; Extract two-dimensional target features from the two-dimensional task fusion features; Extract three-dimensional target features from the three-dimensional task fusion features; The corresponding target detection results are obtained by analyzing the two-dimensional target features and the three-dimensional target features respectively.

[0008] Optionally, acquiring the fisheye image to be detected includes: Acquire raw fisheye images; the raw fisheye images include the target to be detected; The original fisheye image is scaled. The scaled original fisheye image is converted into a floating-point fisheye image; The floating-point fisheye image is normalized to generate the fisheye image to be detected.

[0009] Optionally, the feature fusion of the shared target features and the three-dimensional target features to obtain three-dimensional task fusion features includes: Determine low-level shared target features from the shared multi-scale target features; Determine the mid-level and high-level three-dimensional target features from the aforementioned three-dimensional multi-scale target features; The low-level shared target features, the mid-level 3D target features, and the high-level 3D target features are fused to obtain the 3D task fusion features.

[0010] Optionally, the feature fusion of the low-level shared target features, the mid-level 3D target features, and the high-level 3D target features to obtain the 3D task fusion features includes: The low-level shared target features are downsampled and fused with the mid-level three-dimensional target features to obtain the first three-dimensional fused features; The first three-dimensional fusion feature is downsampled and fused with the high-level three-dimensional target feature to obtain the second three-dimensional fusion feature. The second 3D fusion feature is upsampled and fused with the first 3D fusion feature to obtain the third 3D fusion feature. The first three-dimensional fusion feature is upsampled and fused with the low-level shared target feature to obtain the fourth three-dimensional fusion feature. The third 3D fusion feature is upsampled and then fused with the fourth 3D fusion feature to obtain the 3D task fusion feature.

[0011] Optionally, the two-dimensional task fusion features include two-dimensional anonymized task fusion features and two-dimensional common task fusion features, and the extraction of two-dimensional target features from the two-dimensional task fusion features includes: The two-dimensional desensitization task fusion features are used to perform two-dimensional desensitization target detection and extract two-dimensional desensitization target features; Two-dimensional common target detection is performed on the aforementioned two-dimensional common task fusion features to extract two-dimensional common target features; The step of extracting 3D target features from the 3D task fusion features includes: The three-dimensional target features are then used to perform three-dimensional target detection and extract the three-dimensional target features.

[0012] Optionally, the multi-task target features include two-dimensional desensitized target features, two-dimensional common target features, and three-dimensional target features; the step of parsing the two-dimensional target features and the three-dimensional target features respectively to obtain the corresponding target detection results includes: The detection box information of the two-dimensional desensitized target is obtained by parsing the features of the two-dimensional desensitized target; the two-dimensional desensitized target includes at least one of license plates and pedestrians; The detection bounding box information of the two-dimensional common target is obtained by parsing the features of the two-dimensional common target; The detection bounding box information of the three-dimensional target is obtained by parsing the features of the three-dimensional target; the three-dimensional target includes at least one of vehicles, pedestrians and cyclists.

[0013] Optionally, the step of parsing the two-dimensional desensitized target features to obtain the detection box information of the two-dimensional desensitized target includes: The target type of the two-dimensional desensitization target is determined by analyzing its features; Heatmap prediction is performed on the features of the two-dimensional desensitized target to obtain the uncalibrated center point of the two-dimensional desensitized target; The predicted center position offset of the two-dimensional desensitized target is used as the first offset. The coordinates of the uncalibrated center point of the two-dimensional desensitized target are calibrated based on the first offset to obtain the coordinates of the center point of the two-dimensional desensitized target; The detection frame information of the two-dimensional desensitized target is determined based on the center point coordinates of the target.

[0014] Optionally, the step of parsing the features of the two-dimensional common targets to obtain the detection box information of the two-dimensional common targets includes: Analyze the features of the two-dimensional common targets to determine their target types; Heatmap prediction is performed on the features of the two-dimensional common targets to obtain the uncalibrated center point of the two-dimensional common targets; The predicted center position offset of the common two-dimensional target is used as the second offset. The coordinates of the uncalibrated center point of the two-dimensional common target are calibrated according to the second offset to obtain the coordinates of the center point of the two-dimensional common target; The detection bounding box information of the two-dimensional common target is determined based on the center point coordinates of the two-dimensional common target.

[0015] Optionally, the step of parsing the three-dimensional target features to obtain the detection box information of the three-dimensional target includes: Heatmap prediction is performed on the features of the three-dimensional target to obtain the uncalibrated center point of the three-dimensional target; The predicted center position offset of the three-dimensional target is used as the third offset. The coordinates of the uncalibrated center point of the three-dimensional target are calibrated based on the third offset to obtain the coordinates of the center point of the three-dimensional target. The size parameters, center point camera coordinate system depth, and first heading angle of the three-dimensional target are determined by analyzing the three-dimensional target features. The detection frame parameters of the three-dimensional target are determined based on the center point coordinates of the three-dimensional target, the size parameters, the depth of the center point camera coordinate system, and the first heading angle.

[0016] Optionally, determining the detection box parameters of the three-dimensional target based on the center point coordinates of the three-dimensional target, the size parameters, the depth of the center point camera coordinate system, and the first heading angle includes: Obtain the calibration intrinsic and extrinsic parameters of the fisheye camera; the fisheye camera is used to acquire raw fisheye images; The three-dimensional center point coordinates of the three-dimensional target in the camera coordinate system are calculated based on the calibration intrinsic parameters of the fisheye camera, the center point coordinates of the three-dimensional target, and the depth of the center point camera coordinate system. The coordinates of the three-dimensional center point of the three-dimensional target in the camera coordinate system and the calibration extrinsic parameters of the fisheye camera are used to calculate the coordinates of the three-dimensional center point of the three-dimensional target in the vehicle coordinate system. The first heading angle is converted into a second heading angle in the vehicle coordinate system; the first heading angle is the heading angle of the three-dimensional target in the camera coordinate system; The detection frame parameters of the three-dimensional target are calculated based on the size parameters of the three-dimensional target, the coordinates of the three-dimensional center point in the vehicle coordinate system, and the second heading angle.

[0017] Optionally, the target detection method is applied to an electronic device, which includes a multi-task target detection network. The tasks include the two-dimensional target detection task and the three-dimensional target detection task. The target detection network includes a detection head module for the three-dimensional target detection task. The method includes: Training the target detection network includes: adding a corresponding two-dimensional detection auxiliary head to the detection head module of the three-dimensional target detection task; the two-dimensional detection auxiliary head is used to detect the two-dimensional features of the target in the three-dimensional target detection task.

[0018] This application also discloses a target detection device, the device comprising: The acquisition module is used to acquire the fisheye image to be detected; The first extraction module is used to extract shared multi-scale target features from the fisheye image to be detected; the shared multi-scale target features include multiple shared target features of different scales, which are used for two-dimensional target detection tasks and three-dimensional target detection tasks. The two-dimensional fusion module is used to fuse the shared target features at multiple different scales to obtain two-dimensional task fusion features; The second extraction module is used to extract three-dimensional multi-scale target features from the shared multi-scale target features; the three-dimensional multi-scale target features include multiple three-dimensional target features at different scales; A 3D fusion module is used to fuse the shared target features and the 3D target features to obtain 3D task fusion features; A two-dimensional target feature extraction module is used to extract two-dimensional target features from the two-dimensional task fusion features; A three-dimensional target feature extraction module is used to extract three-dimensional target features from the three-dimensional task fusion features; The parsing module is used to parse the two-dimensional target features and the three-dimensional target features respectively to obtain the corresponding target detection results.

[0019] This application also discloses an electronic device, including: a processor; and a memory for storing processor-executable instructions; The processor is configured to execute the instructions to implement the target detection method described above.

[0020] This application also discloses a computer-readable storage medium storing executable code thereon, which, when executed, causes a processor to perform the target detection method described above.

[0021] This application also discloses a vehicle that includes the electronic equipment described above.

[0022] Compared with the prior art, the embodiments of this application have the following advantages: In this embodiment, a fisheye image to be detected is acquired. Shared multi-scale target features are extracted from the fisheye image. These shared multi-scale target features include multiple shared target features at different scales and can be used for both 2D and 3D target detection tasks. Feature fusion of these shared target features at different scales yields 2D task fusion features. Similarly, 3D multi-scale target features can be extracted from the shared multi-scale target features. These 3D multi-scale target features include multiple 3D target features at different scales. Feature fusion of the shared target features and the 3D target features yields 3D task fusion features. Then, 2D and 3D target features are extracted from the obtained 2D and 3D task fusion features, respectively. Finally, the 2D and 3D target features are analyzed to obtain the corresponding target detection results. Through the above implementation process, the introduction of three-dimensional multi-scale target features can improve the difficulty of three-dimensional target detection caused by the lack of distortion correction in fisheye images, and enhance the ability to recognize the three-dimensional information of targets in subsequent target detection tasks. By extracting multi-task target features from two-dimensional task fusion features and three-dimensional task fusion features, the target detection results of two-dimensional and three-dimensional tasks can be obtained separately, ensuring the effective recognition of rich two-dimensional and three-dimensional targets around the vehicle. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart illustrating the steps of a target detection method provided in an embodiment of this application. Figure 2 This is a flowchart of another target detection method provided in the embodiments of this application; Figure 3 This is a flowchart illustrating the steps of another target detection method provided in the embodiments of this application; Figure 4 This is a diagram of a multi-task target detection network architecture provided in the embodiments of this application; Figure 5 This is a visualization of target detection results provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a target detection method apparatus provided in the embodiments of this application; Figure 7This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0025] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0026] Parking assistance systems are a type of automotive driving assistance technology that uses sensors and path planning to assist drivers in completing parking operations. Parking assistance perception aims to identify nearby targets around the vehicle to ensure safety during the parking process. Fisheye image recognition is currently the mainstream technology for parking assistance perception. It has a large field of view; typically, only four fisheye cameras are needed to completely cover the area around the vehicle, allowing the driver to see a 360-degree panoramic view of the surroundings on the display screen, eliminating any blind spots. Therefore, most parking assistance perception systems use the acquired fisheye images as input, feeding them into a 2D object detection network to detect various targets in the parking scene, providing target attributes and rough location information. Simultaneously, to accurately identify the positions of key dynamic targets in the parking scene, a 3D object detection network can be used to detect the 3D positions of dynamic targets on the fisheye images.

[0027] In related technologies, a single-task recognition architecture is typically employed, using independent 2D and 3D object detection networks to achieve object detection in parking scenes. This includes using a 3D object detection network to directly predict 3D object information from undistorted fisheye images and converting distorted graphic features into bird's-eye view features for 3D object recognition. Directly predicting 3D object information from undistorted fisheye images ensures no loss of visible range and reduces preprocessing steps for fisheye images; however, image distortion makes learning object features difficult. Converting distorted graphic features into bird's-eye view features for 3D object recognition makes object features easier to learn, but the spatial information loss due to the reduced visible range caused by distortion correction is irreparable. Furthermore, the perception range and performance of this bird's-eye view-based object detection are affected by the feature size and granularity of the bird's-eye view. Achieving accurate object recognition requires increasing the size of the bird's-eye view features and decreasing the granularity, which consumes significant deployment computing resources. Furthermore, both of the above object detection methods lack the ability to detect 2D objects in parking scenarios, requiring a separate 2D object detection model to compensate for the lack of other object recognition information in parking scenarios.

[0028] To address the aforementioned problems, this application proposes a target detection method applied to an electronic device. The electronic device includes an end-to-end multi-task target detection network, which performs multiple target detection tasks, including two-dimensional target detection tasks and three-dimensional target detection tasks. This method enables effective identification of a rich variety of two-dimensional and three-dimensional targets in parking scenarios. The electronic device may include mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks or personal digital assistants (PDAs), servers, network-attached storage (NAS), personal computers (PCs), etc.

[0029] Reference Figure 1 The diagram illustrates a flowchart of a target detection method provided in an embodiment of this application, which specifically includes the following steps: Step 101: Obtain the fisheye image to be detected; Among them, the fisheye image to be detected is an undistorted fisheye image, which includes the target to be detected.

[0030] Step 102: Extract shared multi-scale target features from the fisheye image to be detected; the shared multi-scale target features include multiple shared target features at different scales, which are used for two-dimensional target detection tasks and three-dimensional target detection tasks; Among them, shared multi-scale target features are shared features extracted from fisheye images to be detected that can be applied to multiple target detection tasks.

[0031] Step 103: Perform feature fusion on the shared target features at multiple different scales to obtain two-dimensional task fusion features; Among them, the two-dimensional task fusion features are the two-dimensional feature information of the target to be detected and other targets included in the fisheye image to be detected, which can be used for two-dimensional target detection tasks. Step 104: Extract three-dimensional multi-scale target features from the shared multi-scale target features; the three-dimensional multi-scale target features include three-dimensional target features of multiple different scales; Among them, the three-dimensional multi-scale target feature is a three-dimensional geometric feature of the target to be detected that is extracted from the shared multi-scale target feature and applied separately to the three-dimensional target detection task.

[0032] Step 105: Perform feature fusion on the shared target features and the three-dimensional target features to obtain three-dimensional task fusion features; Among them, the 3D task fusion feature is the 3D feature information of the target to be detected and other targets included in the fisheye image to be detected, which can be used for 3D target detection tasks.

[0033] Step 106: Extract two-dimensional target features from the two-dimensional task fusion features; Step 107: Extract three-dimensional target features from the three-dimensional task fusion features; Step 108: Analyze the two-dimensional target features and the three-dimensional target features respectively to obtain the corresponding target detection results.

[0034] Object detection can be performed using an end-to-end multi-task object detection network. The tasks can include 2D and 3D object detection. During object detection, an undistorted fisheye image of the target object can be acquired. Shared multi-scale object features are extracted from the fisheye image. These shared multi-scale object features include multiple shared object features at different scales and can be used for both 2D and 3D object detection tasks. Feature fusion of these shared object features at different scales yields 2D task fusion features. Similarly, 3D multi-scale object features can be extracted from the shared multi-scale object features. These 3D multi-scale object features include multiple 3D object features at different scales. Feature fusion of the shared object features and 3D object features yields 3D task fusion features. Then, 2D and 3D object features can be extracted from the obtained 2D and 3D task fusion features, respectively. Finally, the corresponding object detection results are obtained by parsing the 2D and 3D object features. Through the above implementation process, using the undistorted fisheye image to be detected, compared with using the distorted fisheye image for target detection, can significantly reduce the consumption of computing resources in the target detection process. In addition, by introducing three-dimensional multi-scale target features, the difficulty of three-dimensional target detection caused by the undistorted fisheye image can be improved, and the ability to recognize the three-dimensional information of the target in subsequent target detection tasks can be improved. By extracting two-dimensional target features and three-dimensional target features from the two-dimensional task fusion features and the three-dimensional task fusion features respectively, the target detection results of two-dimensional and three-dimensional tasks can be obtained separately. Therefore, it is not necessary to configure separate two-dimensional and three-dimensional target recognition networks to ensure effective recognition of rich two-dimensional and three-dimensional targets around the vehicle.

[0035] In one embodiment provided in this application, acquiring the fisheye image to be detected includes: Acquire raw fisheye images; the raw fisheye images include the target to be detected; The original fisheye image is scaled. The scaled original fisheye image is converted into a floating-point fisheye image; The floating-point fisheye image is normalized to generate the fisheye image to be detected.

[0036] The original fisheye image can be an undistorted fisheye image acquired by a fisheye camera mounted on the vehicle. The fisheye camera can be one or more fisheye cameras deployed around the vehicle body. If four fisheye cameras with a field of view (FOV) of 195° or more are deployed around the vehicle body, a Surround-View Camera System (SVCS) can be constructed, achieving 360° coverage around the vehicle without blind spots. The number of fisheye cameras deployed and the selected field of view are merely examples. This application embodiment can also deploy other numbers and other field of view fisheye cameras to acquire the original fisheye image as needed; this application embodiment does not impose any limitations on this.

[0037] After acquiring the raw fisheye image, the undistorted raw fisheye image can be input into a multi-task object detection network. This network can scale the raw fisheye image to a preset fixed size, reducing computational load and ensuring consistency of the fisheye images to be detected, thereby improving the efficiency of subsequent object detection processing. For example, the image resolution of the raw fisheye image can be uniformly reduced to an image size of 480 pixels in height and 640 pixels in width. Since the raw fisheye image output by the fisheye camera is usually an integer type image, the scaled raw fisheye image can be converted to a floating-point fisheye image, which helps improve the detection accuracy of subsequent object detection. For example, the data precision type of the floating-point fisheye image can be set according to actual needs; typically, it can be converted to a 32-bit floating-point fisheye image. Using a 32-bit floating-point fisheye image can better balance image precision and memory usage. Data processing at the vehicle-mounted system needs to balance real-time performance and data accuracy. By scaling and converting the original fisheye image, a rapid response to real-time image processing requirements can be achieved within the limited processing performance and computing resources of the vehicle-mounted system, while maintaining good image quality. After type conversion, the floating-point fisheye image can be normalized to map its values ​​to a preset range, such as [0,1], generating an undistorted fisheye image to be detected. This helps improve the generalization ability of the multi-task object detection network and reduces the risk of overfitting.

[0038] In one embodiment provided in this application, shared multi-scale target features can be extracted from the fisheye image to be detected, and three-dimensional multi-scale target features can be further extracted from the shared multi-scale target features. When extracting the multi-scale shared target features, comprehensive feature extraction is performed on both the two-dimensional and three-dimensional information contained in the fisheye image to be detected, ensuring that relevant features for both two-dimensional and three-dimensional target detection tasks are considered simultaneously. In a parking scenario, the shared multi-scale target features can include detailed information, geometric information, and global semantic information of various dynamic and static targets of interest in the parking scenario. Dynamic targets of interest in the parking scenario can include vehicles and pedestrians, while static targets can include various markers or obstacles such as water-filled barriers, traffic cones, and pillars. The shared multi-scale target features are not limited to targeted detection of two-dimensional or three-dimensional targets; they can extract more diverse and comprehensive feature information. Further extraction of three-dimensional multi-scale target features based on the shared multi-scale target features allows the extracted three-dimensional multi-scale target features to possess richer and more accurate global information. For example, the DLA (DeepLayer Aggregation) 34 network can be used as the shared backbone network for the shared feature extraction module to extract shared multi-scale target features. The DLA34 backbone network is a convolutional neural network structure with efficient feature aggregation capabilities, particularly suitable for tasks such as image recognition and segmentation. Using DLA-34 as the backbone network and fusing multi-layer features through deep aggregation technology allows for comprehensive aggregation learning of features within and outside the module at multiple scales. This is more conducive to extracting complex target features from fisheye images as shared multi-scale target features for subsequent 2D and 3D target detection tasks, solving the problems of uneven target scale and appearance deformation caused by distortion. Alternatively, a 3D feature extraction backbone network unique to 3D target detection tasks and without shared parameters can be used—a lightweight feature extraction backbone network exclusive to 3D detection tasks—to further extract the 3D geometric information of each target from the shared multi-scale target features as 3D multi-scale target features. Among them, the 3D feature extraction backbone network (a lightweight feature extraction backbone network exclusively for 3D detection tasks) extracts 3D multi-scale target features that are exclusive to 3D tasks and are only used for 3D tasks. The corresponding feature extraction backbone network can be a more lightweight network architecture than the shared backbone network. For example, the last two stages of the DLA34 network can be used to make it more suitable for running in resource-constrained environments such as vehicle terminals, which can reduce computational latency and improve the real-time performance of 3D multi-scale target features.For example, shared target features at a certain scale (e.g., one-eighth scale) from shared multi-scale target features can be input into a 3D feature extraction backbone network to extract higher-level (e.g., one-sixteenth and one-thirty-second scale) 3D target features. Increasing the network depth further expands the network's receptive field. The receptive field is the area that a neuron in a convolutional neural network can "see" in an input image. Because the targets of interest in shared multi-scale target features are richer and contain more diverse and precise low-dimensional features, extracting 3D multi-scale target features based on shared multi-scale target features through a 3D feature extraction backbone network can also improve the global contextual semantic learning ability of 3D targets, obtaining the global geometric features of the 3D targets of interest in the 3D task, thereby helping to improve the performance of 3D target detection.

[0039] In one embodiment provided in this application, two-dimensional task fusion features can be obtained by fusing shared multi-scale target features through the DLA UP module. The two-dimensional task can include a two-dimensional desensitized target detection task and a two-dimensional common target detection task. The two-dimensional desensitized target detection task is to detect targets in the fisheye image to be detected that require desensitization, including license plates, pedestrians, etc. The two-dimensional common target detection task is to detect various common targets in the fisheye image to be detected that are suitable for the application scenario. For example, in the case of a parking scenario, two-dimensional common targets can include traffic lights, signs, obstacles, ground markings, various parking-related facilities, etc. The shared multi-scale target features include multiple target features at different scales. During feature fusion, some low, medium, and high-level features can be selected for fusion, thereby reducing the consumption of computing resources. The feature levels from low to high correspond to feature image scales from large to small, and feature image resolutions from high to low; different levels focus on different features. Low-level features can be extracted from shallow backbone networks (such as the first two layers of the DLA34 network), including basic information such as edges, textures, and colors; mid-level features are located in the middle of the backbone network (such as the middle two layers of the DLA34 network), containing relatively complex local structures, such as texture combinations and target contours, which can represent the local geometric features of the detected target, such as lines and shapes; high-level features can be extracted from deep backbone networks (such as the last two layers of the DLA34 network), and can contain global semantic information such as the overall attributes, category information, position information, orientation information, and overall geometric information of the detected target.In the process of fusing shared target features at multiple different scales, we can start with the lowest-level features. The low-level shared target features are downsampled to the same scale (resolution) as the mid-level shared target features. The downsampled low-level and mid-level shared target features are then fused through element-wise addition or feature concatenation to obtain the first shared fused feature. This first shared fused feature is then downsampled to the same scale as the high-level shared target features. The downsampled first shared fused feature is then fused with the high-level shared target features through element-wise addition or feature concatenation to obtain the second shared fused feature. Finally, the second shared fused feature is upsampled to the same scale as the mid-level shared target features. The shared target features in the middle layer are at the same scale. The upsampled second shared fusion feature is fused with the first shared fusion feature through element-wise addition or feature concatenation to obtain the third shared fusion feature. The first shared fusion feature can be upsampled to the same scale as the shared target features in the lower layer, and then fused with the lower-level shared target features through element-wise addition or feature concatenation to obtain the fourth shared fusion feature. The third shared fusion feature is then upsampled to the same scale as the lower-level shared target features, and then fused with the fourth shared fusion feature through element-wise addition or feature concatenation to obtain the two-dimensional task fusion feature. For example, shared target features at three scales—one-eighth, one-sixteenth, and one-thirty-second of the original fisheye image size—can be selected from the shared multi-scale target features as low, middle, and high-level shared target features, respectively. These three scales of shared features are then fully fused through upsampling and downsampling, and finally, the fusion feature at the one-eighth scale is output as the two-dimensional task fusion feature. The larger the feature scale, the higher the resolution of the corresponding feature image, and the more detailed information it contains. Conversely, the smaller the feature scale, the lower the resolution of the corresponding feature image, but the richer the semantic information it contains. By fusing low-, medium-, and high-level shared target features from the shared multi-scale target features, we can cover feature information at different levels within the shared multi-scale target features while avoiding the excessive computational resource consumption that can result from full-scale fusion. This effectively reduces the computational load of the target detection network itself while aggregating multi-level features, resulting in two-dimensional task fusion features that combine detail and semantics. This significantly improves the target recognition ability and detection accuracy, and enhances the model's generalization ability in complex scenes. Furthermore, we can choose an eighth-scale as the scale of the output two-dimensional task fusion features. Since the eighth-scale resolution retains relatively rich feature details and consumes relatively few computational resources, the above implementation process can ensure target detection performance while reducing computational resource consumption.

[0040] Reference Figure 2This document illustrates a flowchart of another target detection method provided in an embodiment of this application. In one embodiment, the step of fusing the shared target features and the three-dimensional target features to obtain three-dimensional task fusion features includes: Step 201: Determine the low-level shared target features from the shared multi-scale target features; Here, low-level shared target features are those shared target features with a lower downsampling factor. For example, low-level shared target features can be shared target features with a scale of one-eighth of the original fisheye image size.

[0041] Step 202: Determine the mid-level three-dimensional target features and the high-level three-dimensional target features from the three-dimensional target features; Among them, the mid-level 3D target features and the high-level 3D target features are 3D target features with medium and high downsampling factors, respectively, in the 3D multi-scale target features. For example, the mid-level 3D target features can be 3D target features with a scale of one-sixteenth of the original fisheye image size, and the high-level 3D target features can be 3D target features with a scale of one-thirty-second of the original fisheye image size.

[0042] Step 203: Perform feature fusion on the low-level shared target features, the mid-level three-dimensional target features, and the high-level three-dimensional target features to obtain the three-dimensional task fusion features.

[0043] Low-level shared target features can be determined from shared multi-scale target features, while mid-level and high-level 3D target features can be determined from 3D multi-scale target features. Then, these low-level, mid-level, and high-level 3D target features are extracted and fused to obtain 3D task fusion features. For example, shared features at a scale of one-eighth the size of the original fisheye image can be extracted from shared multi-scale target features as low-level shared target features. 3D features at scales of one-sixteenth and one-thirty-second the size of the original fisheye image can be extracted from 3D multi-scale target features as mid-level and high-level 3D target features, respectively. These extracted features are then fused to obtain 3D task fusion features. Alternatively, low-resolution image features can be upsampled and then concatenated with high-resolution image features to achieve multi-scale feature fusion. The resulting 3D task fusion feature can have a scale of one-eighth the size of the original fisheye image. Using one-eighth as the scale of the 3D task fusion feature preserves relatively rich feature details while consuming relatively few computational resources. In practical applications, the scales of low-level shared target features, mid-level 3D target features, high-level 3D target features, and 3D task fusion features can all be set according to actual needs, and this application does not impose specific restrictions on this. Shared multi-scale target features are extracted by the backbone network, and the targets of interest are richer. Therefore, low-level shared target features have richer and more accurate high-resolution detail features, while low-resolution features have higher-dimensional context feature abstraction capabilities. Through the above implementation process, high-resolution low-level shared target features are fused with low-resolution mid-level and high-level 3D target features. This avoids the large computational resource consumption caused by point-by-point full-scale fusion of 2D and 3D multi-scale features, allowing the fused 3D task fusion features to contain more and more accurate target information. It also further enhances the high-dimensional information extraction capability, improves the difficulty in extracting target geometric information due to fisheye distortion in the original fisheye image, and helps improve subsequent 3D target detection performance. While reducing computational resource consumption by using the original fisheye image, it ensures the detection accuracy of 3D targets, achieving effective recognition of 3D targets.

[0044] In one embodiment provided in this application, the feature fusion of the low-level shared target features, the mid-level 3D target features, and the high-level 3D target features to obtain the 3D task fusion features includes: The low-level shared target features are downsampled and fused with the mid-level three-dimensional target features to obtain the first three-dimensional fused features; The first three-dimensional fusion feature is downsampled and fused with the high-level three-dimensional target feature to obtain the second three-dimensional fusion feature. The second 3D fusion feature is upsampled and fused with the first 3D fusion feature to obtain the third 3D fusion feature. The first three-dimensional fusion feature is upsampled and fused with the low-level shared target feature to obtain the fourth three-dimensional fusion feature. The third 3D fusion feature is upsampled and then fused with the fourth 3D fusion feature to obtain the 3D task fusion feature.

[0045] In the process of fusing low-level shared target features, mid-level 3D target features, and high-level 3D target features at different scales, we can start with the low-level shared target features. We downsample the low-level shared target features to the same scale (resolution) as the mid-level 3D target features, and then fuse the downsampled low-level shared target features with the mid-level 3D target features through element-wise addition or feature concatenation to obtain the first 3D fusion feature. Then, we further downsample the first 3D fusion feature to the same scale as the high-level 3D target features, and then fuse the downsampled first 3D fusion feature with the high-level 3D target features through element-wise addition or feature concatenation to obtain the second 3D fusion feature. Finally, we fuse the second 3D fusion feature... The 3D fusion features are upsampled to the same scale as the mid-level 3D target features. The upsampled second 3D fusion feature is then fused with the first 3D fusion feature through element-wise addition or feature concatenation to obtain the third 3D fusion feature. This process can be repeated by upsampling the first 3D fusion feature to the same scale as the low-level shared target features and fusing it with the low-level shared target features through element-wise addition or feature concatenation to obtain the fourth 3D fusion feature. Finally, the third 3D fusion feature is upsampled to the same scale as the low-level shared target features and fused with the fourth 3D fusion feature through element-wise addition or feature concatenation to obtain the 3D task fusion feature. Through this process, low-level shared target features, mid-level 3D target features, and high-level 3D target features can be fully fused, integrating the rich detail information in the low-level shared target features, the local target information in the mid-level 3D target features, and the precise global semantics in the high-level 3D target features. This helps improve the generalization ability of the target detection network in complex scenes, enhances the network's ability to locate and classify targets, and improves target detection accuracy.

[0046] In one embodiment provided in this application, the two-dimensional task fusion feature includes two-dimensional de-identified task fusion features and two-dimensional common task fusion features, and the extraction of two-dimensional target features from the two-dimensional task fusion features includes: The two-dimensional desensitization task fusion features are used to perform two-dimensional desensitization target detection and extract two-dimensional desensitization target features; Among them, the two-dimensional desensitization target features refer to the features related to the two-dimensional desensitization target in the fusion features of the two-dimensional desensitization task.

[0047] Two-dimensional common target detection is performed on the aforementioned two-dimensional common task fusion features to extract two-dimensional common target features; Among them, the two-dimensional common target features refer to the features related to the two-dimensional common targets in the two-dimensional common task fusion features.

[0048] The step of extracting 3D target features from the 3D task fusion features includes: The three-dimensional target features are then used to perform three-dimensional target detection and extract the three-dimensional target features.

[0049] Among them, the three-dimensional target features refer to the features related to the three-dimensional target in the three-dimensional task fusion features.

[0050] The 2D task fusion features include 2D desensitization task fusion features and 2D common task fusion features. 2D desensitization target detection can be performed on the 2D desensitization task fusion features to identify the target types of each target within the fusion features, thereby determining the desensitized targets required for the 2D desensitization task, and then extracting the relevant features of the corresponding desensitized targets as 2D desensitization target features. Similarly, 2D common task fusion features can be used to perform 2D common target detection, identifying the target types of each target within the fusion features, thereby determining the common targets required for the 2D common tasks, and then extracting the relevant features of the corresponding common targets as 2D common target features. 3D task fusion features can be used to perform 3D target detection, identifying the target types of each target within the 3D task fusion features, thereby determining the 3D targets required for the 3D tasks, and then extracting the relevant features of the corresponding 3D targets as 3D target features. Through the above implementation process, a multi-task target detection network is used to process the original fisheye image end-to-end, simultaneously executing 2D and 3D target detection tasks through a single network structure, including 2D desensitization target detection, 2D common target detection, and 3D target detection tasks. Two-dimensional desensitized target detection tasks can be used to desensitize information in acquired raw fisheye images. Two-dimensional common target detection tasks can be used to identify various common static and dynamic targets in intelligent driving scenarios, such as parking scenarios, providing attribute and rough position information of targets around the vehicle. Then, by combining signals from other downstream sensors, the target attributes and precise positions of common targets can be identified. Three-dimensional target detection tasks can be used to identify the 3D position, size, and orientation of dynamic 3D targets. For example, in parking scenarios, dynamic targets such as vehicles, pedestrians, and cyclists around the vehicle can be identified. The identified 3D information can provide important input signals for vehicle planning and control and in-vehicle display. By executing two-dimensional and three-dimensional target detection tasks and extracting multi-task target features from the fusion features of the two-dimensional and three-dimensional tasks, effective identification of rich targets in intelligent driving scenarios, such as parking scenarios, can be achieved, constructing relatively complete environmental perception information for parking scenarios.

[0051] refer to Figure 3 This document illustrates a flowchart of another target detection method provided in an embodiment of this application. In one embodiment, the multi-task target features include two-dimensional desensitized target features, two-dimensional common target features, and three-dimensional target features. The step of parsing the two-dimensional target features and the three-dimensional target features to obtain the corresponding target detection results includes: Step 301: Parse the features of the two-dimensional desensitized target to obtain the detection box information of the two-dimensional desensitized target; the two-dimensional desensitized target includes at least one of license plate and pedestrian; Among them, the two-dimensional desensitization target refers to the target object in the original fisheye image that needs to be desensitized in the corresponding intelligent driving scenario, including at least one of license plates and pedestrians.

[0052] Step 302: Parse the features of the two-dimensional common targets to obtain the detection box information of the two-dimensional common targets; Among them, common two-dimensional targets refer to common target objects that need to be identified in the original fisheye image under the corresponding intelligent driving scenario, including traffic lights, signs, obstacles (water barriers, pillars, tires, etc.), ground markings (lane lines, parking space lines, etc.), and various parking-related facilities.

[0053] Step 303: Analyze the three-dimensional target features to obtain the detection box information of the three-dimensional target; the three-dimensional target includes at least one of vehicles, pedestrians and cyclists.

[0054] Among them, the three-dimensional target refers to the target object in the original fisheye image that needs to be detected and identified in three dimensions under the corresponding intelligent driving scenario.

[0055] Parsing the features of two-dimensional desensitized targets yields their detection bounding boxes. Two-dimensional desensitized targets can include at least one of license plates and pedestrians. Parsing the features of common two-dimensional targets yields their detection bounding boxes. Parsing the features of three-dimensional targets yields their detection bounding boxes. Three-dimensional targets can include at least one of vehicles, pedestrians, and cyclists. Through the above implementation process, parsing the features of two-dimensional desensitized targets, common two-dimensional targets, and three-dimensional targets respectively yields their corresponding detection bounding boxes. This avoids the problems of lacking depth information in single-dimensional target detection and low accuracy in single-dimensional target detection, achieving effective detection of various targets in the original fisheye image. It also avoids the computational redundancy and high resource consumption caused by performing two-dimensional and three-dimensional target detection tasks separately using a single-task recognition architecture in existing technologies, making it more suitable for hardware-restricted scenarios such as in-vehicle terminals. The parsed detection bounding box information can be used as input signals for downstream functions to realize planning and control of intelligent driving scenarios and in-vehicle displays, enabling vehicles to achieve a comprehensive solution for varying degrees of safe and comfortable driving or parking functions in various traffic scenarios. In addition, the obtained detection bounding box information can be used to output visualized target detection results.

[0056] In one embodiment provided in this application, the step of parsing the two-dimensional desensitized target features to obtain the detection box information of the two-dimensional desensitized target includes: The target type of the two-dimensional desensitization target is determined by analyzing its features; Heatmap prediction is performed on the features of the two-dimensional desensitized target to obtain the uncalibrated center point of the two-dimensional desensitized target; Among them, heat Figure 1 This technique achieves target localization by generating a probability distribution map of the target's center point. It can be based on point target detection algorithms, such as CenterNet, to locate the target's center point.

[0057] The predicted center position offset of the two-dimensional desensitized target is used as the first offset. The coordinates of the uncalibrated center point of the two-dimensional desensitized target are calibrated based on the first offset to obtain the coordinates of the center point of the two-dimensional desensitized target; Among them, the coordinates of the center point of the two-dimensional desensitized target are the positions of the center point of the two-dimensional desensitized target in the feature map corresponding to the features of the two-dimensional desensitized target.

[0058] The detection frame information of the two-dimensional desensitized target is determined based on the center point coordinates of the target.

[0059] The target type of each two-dimensional desensitized target can be determined by parsing the two-dimensional desensitized target features. Then, a corresponding heatmap is generated based on the two-dimensional desensitized target features. Heatmap prediction is performed on the two-dimensional desensitized target features to obtain the probability value of each pixel as the center point of the target. Then, the probability values ​​corresponding to each pixel are filtered according to a preset confidence threshold, and pixels with a confidence value greater than the preset confidence threshold are considered as uncalibrated center points of the two-dimensional desensitized target. For example, the confidence threshold can be set to 0.5. The probability value of each pixel as the center point of the target, obtained after processing the heatmap feature map corresponding to the two-dimensional desensitized target features using softmax (normalized exponential function), ranges from 0 to 1. If the probability value of a pixel is greater than 0.5, it indicates that the pixel is the center point of a two-dimensional desensitized target; if the probability value is not greater than 0.5, it indicates that the pixel is not the center point of a two-dimensional desensitized target. The preset confidence threshold can be set according to actual needs, and this application does not impose specific restrictions on it. Since the resolution of the feature map is lower than that of the input original fisheye image, the predicted center position offset of the 2D desensitized target can be used as the first offset. Based on this first offset, the coordinates of the uncalibrated center point of the 2D desensitized target are calibrated to obtain the center point coordinates, thus achieving precise positioning of the 2D desensitized target's center point. For example, a Gaussian function can be used to map the floating-point coordinates of the 2D desensitized target's center point in the feature map to integer coordinates, and then combined with the feature map scale corresponding to the 2D desensitized target's features to obtain the accurate position of the 2D desensitized target's center point in the original fisheye image. After determining the center point coordinates of the 2D desensitized target, the corresponding detection box information of the 2D desensitized target can be determined based on the center point coordinates of each 2D desensitized target, that is, the width and height of the detection box of the 2D desensitized target. The detection box information of the 2D desensitized target can include the target type, center point coordinates, and the width and height of the 2D desensitized target's detection box. Through the above implementation process, the features of the two-dimensional desensitized target can be analyzed to obtain the target type, center point coordinates and detection box size corresponding to each two-dimensional desensitized target, thereby achieving accurate detection of each two-dimensional desensitized target included in the original fisheye image.

[0060] In one embodiment provided in this application, the step of parsing the two-dimensional common target features to obtain the detection box information of the two-dimensional common target includes: Analyze the features of the two-dimensional common targets to determine their target types; Heatmap prediction is performed on the features of the two-dimensional common targets to obtain the uncalibrated center point of the two-dimensional common targets; The predicted center position offset of the common two-dimensional target is used as the second offset. The coordinates of the uncalibrated center point of the two-dimensional common target are calibrated according to the second offset to obtain the coordinates of the center point of the two-dimensional common target; Among them, the coordinates of the center point of a common two-dimensional target are the positions of the center point of the common two-dimensional target in the feature map corresponding to the features of the common two-dimensional target.

[0061] The detection bounding box information of the two-dimensional common target is determined based on the center point coordinates of the two-dimensional common target.

[0062] The target type of each common two-dimensional target can be determined by analyzing its features. Then, a heatmap is generated based on these features, and heatmap prediction is performed to obtain the probability value of each pixel as the center point of a target. The probability values ​​of each pixel are then filtered according to a preset confidence threshold, and pixels with a confidence value greater than the threshold are considered uncalibrated center points of the common two-dimensional target. For example, the confidence threshold can be set to 0.5. The probability value of each pixel as the center point of a target, obtained after softmax processing of the heatmap feature map corresponding to the common two-dimensional target features, ranges from 0 to 1. If the probability value of a pixel is greater than 0.5, it indicates that the pixel is the center point of a common two-dimensional target; if the probability value is not greater than 0.5, it indicates that the pixel is not the center point of a common two-dimensional target. The preset confidence threshold can be set according to actual needs, and this application does not impose specific restrictions on it. Since the resolution of the feature map is lower than that of the input original fisheye image, the center position offset of the predicted 2D common target can be used as a second offset. The coordinates of the uncalibrated center point of the 2D common target are then calibrated based on this second offset to obtain the center point coordinates of the 2D common target, thereby achieving precise localization of the center point of the 2D common target. For example, a Gaussian function can be used to map the floating-point coordinates of the center point of the 2D common target in the feature map to integer coordinates, and then combined with the feature map scale corresponding to the 2D common target features to obtain the accurate position of the center point of the 2D common target in the original fisheye image. After determining the center point coordinates of the 2D common targets, the detection box information of the corresponding 2D common targets can be determined based on the center point coordinates of each 2D common target, that is, the width and height of the detection box of the 2D common target. The detection box information of the 2D common targets can include the target type, center point coordinates, and width and height of the detection box of the 2D common target. Through the above implementation process, the features of the 2D common targets can be parsed to obtain the target type, center point coordinates, and detection box size corresponding to each 2D common target, thereby achieving accurate detection of each 2D common target included in the original fisheye image.

[0063] In one embodiment provided in this application, the step of parsing the three-dimensional target features to obtain the detection box information of the three-dimensional target includes: Heatmap prediction is performed on the features of the three-dimensional target to obtain the uncalibrated center point of the three-dimensional target; The predicted center position offset of the three-dimensional target is used as the third offset. The coordinates of the uncalibrated center point of the three-dimensional target are calibrated based on the third offset to obtain the coordinates of the center point of the three-dimensional target. The coordinates of the center point of the 3D target are the positions of the center point of the 3D target in the feature map corresponding to the 3D target features.

[0064] The size parameters, center point camera coordinate system depth, and first heading angle of the three-dimensional target are determined by analyzing the three-dimensional target features. The size parameters of the 3D target can include its height, width, and length in the feature map corresponding to the 3D target features. The camera coordinate system is a 3D Cartesian coordinate system established with the optical center of the fisheye camera as the origin. Its Z-axis coincides with the optical axis and points in the imaging direction, while the X and Y axes are parallel to the horizontal and vertical directions of the image plane, respectively. The center point depth in the camera coordinate system (Z-axis) represents the straight-line distance of the center point of the 3D target along the camera's optical axis. The first heading angle represents the angle between the target's forward direction and the fisheye camera's optical axis in the camera coordinate system, and can be used to assess collision risk.

[0065] The detection frame parameters of the three-dimensional target are determined based on the center point coordinates of the three-dimensional target, the size parameters, the depth of the center point camera coordinate system, and the first heading angle.

[0066] The target type of each 3D target can be determined by parsing its features. Then, a heatmap is generated based on the 3D target features. Heatmap prediction is performed on the 3D target features to obtain the probability value of each pixel as the center point of the target. Then, the probability values ​​of each pixel are filtered according to a preset confidence threshold, and pixels with a confidence value greater than the preset threshold are considered uncalibrated center points of the 3D target. For example, the confidence threshold can be set to 0.5. The probability value of each pixel as the center point of the target, obtained after softmax processing of the heatmap feature map corresponding to the 3D target features, ranges from 0 to 1. If the probability value of a pixel is greater than 0.5, it indicates that the pixel is the center point of a 3D target; if the probability value is not greater than 0.5, it indicates that the pixel is not the center point of the 3D target. The preset confidence threshold can be set according to actual needs, and this application does not impose specific restrictions on it. Since the resolution of the feature map is lower than that of the input original fisheye image, the predicted center position offset of the 3D target can be used as a third offset. Based on this third offset, the coordinates of the uncalibrated center point of the 3D target are calibrated to obtain the center point coordinates of the 3D target. This allows for precise localization of the center point of the 3D target in the feature image corresponding to its features. For example, a Gaussian function can be used to map the floating-point coordinates of the 3D target's center point in the feature map to integer coordinates, thus obtaining the accurate position of the 3D target's center point in the original fisheye image. Then, based on the center point coordinates of the 3D target, the 3D target features can be analyzed to determine the corresponding size parameters, the center point camera coordinate system depth, and the first heading angle. Finally, the detection box parameters of the 3D target are determined based on the center point coordinates, size parameters, center point camera coordinate system depth, and the first heading angle in the camera coordinate system. Through the above implementation process, the features of three-dimensional targets can be analyzed to obtain the target type corresponding to each three-dimensional target, the center point coordinates on the feature image corresponding to the three-dimensional target features, the detection box size parameters, the center point camera coordinate system depth, and the first heading angle in the camera coordinate system. This enables accurate detection of each three-dimensional target included in the original fisheye image and further determines the detection box parameters of the three-dimensional target in the vehicle coordinate system.

[0067] In one embodiment provided in this application, determining the first heading angle corresponding to the three-dimensional target by parsing the three-dimensional target features includes: The orientation angle range of the three-dimensional target is divided into multiple prediction regions according to a preset angle interval, and the prediction region where the three-dimensional target is located is determined. The orientation angle of the three-dimensional target can be in the range of 0-360 degrees. The preset angle interval can be set according to actual needs. For example, it can be 30 degrees. This application does not impose any specific restrictions on this.

[0068] Predict the offset of the 3D target's location relative to the center angle of each prediction region; The heading angle information of the 3D target in the camera coordinate system is determined based on the offset and the predicted area.

[0069] The orientation angle range of a 3D target can be divided into multiple prediction regions according to a preset angular interval. The prediction region where the 3D target is located is determined, and the offset of the 3D target's position relative to the center angle of each prediction region is predicted. Then, based on the offset and the prediction region, the heading angle information of the 3D target in the camera coordinate system is determined. For example, the possible 360-degree orientation range of the 3D target can be divided into 12 regions at preset angular intervals of 30 degrees each. Classification detection is then performed on these 12 regions. Based on the classification detection results, the prediction region where the 3D target's orientation angle is located is determined. Finally, based on the predicted offset of the 3D target's position relative to the center angle of each prediction region and the prediction region where the 3D target's orientation angle is located, the heading angle of the 3D target in the camera coordinate system is determined.

[0070] In one embodiment provided in this application, determining the detection box parameters of the three-dimensional target based on the center point coordinates of the three-dimensional target, the size parameters, the depth of the center point camera coordinate system, and the first heading angle includes: Obtain the calibration intrinsic and extrinsic parameters of the fisheye camera; the fisheye camera is used to acquire raw fisheye images; The intrinsic calibration parameters of a fisheye camera describe its inherent characteristics and may include focal length, principal point, and distortion coefficients. The extrinsic calibration parameters of a fisheye camera may include rotation matrices and translation vectors. The rotation matrix, also called the camera extrinsic parameter matrix, is used to transform points in the world coordinate system to the camera coordinate system. The inverse of the camera extrinsic parameter matrix is ​​called the camera pose, which transforms points in the camera coordinate system to the world coordinate system.

[0071] The three-dimensional center point coordinates of the three-dimensional target in the camera coordinate system are calculated based on the calibration intrinsic parameters of the fisheye camera, the center point coordinates of the three-dimensional target, and the depth of the center point camera coordinate system. The coordinates of the three-dimensional center point of the three-dimensional target in the camera coordinate system and the calibration extrinsic parameters of the fisheye camera are used to calculate the coordinates of the three-dimensional center point of the three-dimensional target in the vehicle coordinate system. The vehicle coordinate system refers to a coordinate system formed when the vehicle is stationary, with the projection of the rear axle center onto the ground as the origin, and consisting of the forward-pointing X-axis, the left-pointing Y-axis, and the vertically upward Z-axis.

[0072] The first heading angle is converted into a second heading angle in the vehicle coordinate system; the first heading angle is the heading angle of the three-dimensional target in the camera coordinate system; Among them, the second heading angle is viewed clockwise when viewed in the positive direction of the Z / Y / X axes in the vehicle coordinate system.

[0073] The detection frame parameters of the three-dimensional target are calculated based on the size parameters of the three-dimensional target, the coordinates of the three-dimensional center point in the vehicle coordinate system, and the second heading angle.

[0074] The calibration intrinsic and extrinsic parameters of the fisheye camera mounted on the vehicle for acquiring raw fisheye images can be obtained. Then, based on the fisheye camera's calibration intrinsic parameters, the center point coordinates of the 3D target, and the depth of the center point in the camera coordinate system, the 3D center point coordinates of the 3D target in the camera coordinate system are calculated using matrix multiplication. Based on the 3D center point coordinates of the 3D target in the camera coordinate system and the calibration extrinsic parameters of the fisheye camera, a coordinate system transformation is performed on the 3D center point coordinates of the 3D target in the camera coordinate system using matrix multiplication to calculate the 3D center point coordinates of the 3D target in the vehicle coordinate system. According to the installation position of the fisheye cameras on the vehicle, a fixed transformation matrix between the camera coordinate system and the vehicle coordinate system for each fisheye is calculated. Using this fixed transformation matrix, the first heading angle of the 3D target in the camera coordinate system is transformed using matrix multiplication to obtain the second heading angle of the 3D target in the vehicle coordinate system. Finally, based on the 3D center point coordinates and second heading angle of the 3D target in the vehicle coordinate system, combined with the corresponding size parameters of the 3D target, the vertex coordinates of the detection box of the 3D target in the vehicle coordinate system can be calculated. For example, the coordinates of eight corner points of a 3D target detection box aligned with the vehicle equipped with a fisheye camera can be calculated based on the 3D center point coordinates of the 3D target in the vehicle coordinate system and the corresponding size parameters of the 3D target. Then, using a second heading angle, the coordinates of the eight corner points of the actual 3D target detection box after rotation according to the second heading angle can be calculated as the corner point coordinates of the 3D target detection box. The detection box parameters of the 3D target may include the target type corresponding to the 3D target, the 3D center point coordinates in the vehicle coordinate system, the corner point coordinates of the detection box, the size parameters, and the second heading angle. Through the above implementation process, the position, size, and heading angle information of the 3D target in the vehicle coordinate system can be obtained by using the calibration intrinsic and extrinsic parameters of the fisheye camera, based on the center point coordinates and size parameters of the 3D target on the original fisheye image, the depth of the camera coordinate system corresponding to the center point of the 3D target, and the first heading angle, through coordinate system transformation. This provides important input signals for vehicle planning and control and vehicle display.

[0075] For example, after obtaining the detection box information of two-dimensional desensitized targets, two-dimensional common targets, and three-dimensional targets, non-maximum suppression (NMS) can be applied to the obtained detection boxes according to different targets to filter out redundant target detection boxes.

[0076] For example, after obtaining the detection bounding box information for 2D desensitized targets, 2D common targets, and 3D targets, this information can be used as input to downstream functional modules for intelligent driving, particularly for implementing intelligent driving functions in parking scenarios. Additionally, the obtained detection bounding box information can be visualized in the original fisheye image, marking the detection bounding boxes for 2D desensitized targets, 2D common targets, and 3D targets in the original fisheye image, and labeling the target type corresponding to each detection bounding box as a label on the corresponding detection bounding box.

[0077] In one embodiment provided in this application, the task includes the two-dimensional target detection task and the three-dimensional target detection task, the target detection network includes a detection head module for the three-dimensional target detection task, and the method includes: Training the target detection network includes: adding a corresponding two-dimensional detection auxiliary head to the detection head module of the three-dimensional target detection task; the two-dimensional detection auxiliary head is used to detect the two-dimensional features of the target in the three-dimensional target detection task.

[0078] During the training of an object detection network, a corresponding 2D detection auxiliary head can be added to the detection head module of the 3D object detection task to detect the 2D features of the target in the 3D object detection task. The detection head is an important component of the object detection network. Its main function is to predict the location and category information of the target based on the features extracted by the model, thereby achieving the identification and localization of the detected target. It can utilize convolutional layers, fully connected layers, etc., to process the input features and ultimately output feature information such as the target's category probability and the coordinates of the detection box. Multi-task object detection networks can simultaneously perform 2D and 3D object detection tasks. However, the physical information of the targets predicted by 2D and 3D object detection tasks differs, and the loss functions of each task also have significant scale differences, which can affect the learning ability between tasks in a multi-task object detection network. Through the above implementation process, during the training of a multi-task object detection network, a two-dimensional detection auxiliary head corresponding to the category of the target in the three-dimensional object detection task can be added to the detection head module. This includes two-dimensional detection auxiliary heads for vehicle, pedestrian, and pillar (obstacle) categories. This allows the network to detect the two-dimensional features of the target in the three-dimensional object detection task. It also adds prediction information from similar two-dimensional tasks to the detection task corresponding to the three-dimensional target, enabling the three-dimensional task branch in the multi-task object detection network to learn some features of interest to the two-dimensional task. This helps enhance the shared backbone feature extraction module's ability to extract common features of two-dimensional targets, improving the learning ability of the multi-task object detection network and preventing the network from biasing towards learning three-dimensional feature information during training, which would degrade the detection performance of the two-dimensional task branch. By setting up the two-dimensional detection auxiliary head, backpropagation can optimize the extraction of shared target features, three-dimensional target features, and the generation of three-dimensional task fusion features. This makes the multi-task object detection network more likely to converge during training, avoiding significant performance deviations due to inconsistent information learned between tasks, and ultimately improving the network performance of the multi-task object detection network.

[0079] In one embodiment provided in this application, training the target detection network further includes: Initialize the learnable weight parameters for each of the tasks; Determine the loss value for each of the tasks; The total loss of the task is determined based on the loss value and the learnable weight parameters; The learnable weight parameters are adjusted based on the total loss.

[0080] Learnable weight parameters are a type of uncertainty-based learnable weight parameter that associates task weights with task uncertainty. In multi-task learning, the weights of each task are dynamically adjusted, optimizing model performance by modeling task uncertainty. Here, uncertainty refers to task-dependent or homoscedastic uncertainty, describing different values ​​for different tasks.

[0081] We can initialize the learnable weight parameters for each task, then determine the loss value for each task, and finally determine the total loss for the task based on the loss value and the learnable weight parameters. The formula for the total loss of a task is: , where L total For the total loss, These are the learnable weight parameters for task i. This is the loss value for task i. Then, the learnable weight parameters can be adjusted based on the total loss. Specifically, the gradient of the total loss with respect to the model parameters and uncertain weight parameters of the multi-task object detection network can be calculated through backpropagation, and the parameters can be updated using an optimizer. The loss weights of each task are inversely proportional to their corresponding uncertainty. For tasks with larger loss values, the corresponding uncertainty is high, and the trained multi-task object detection network will automatically reduce the weight of that task, thereby reducing its impact on the total loss. Through the above implementation process, an uncertainty-based dynamic weight adjustment strategy is introduced into the training process of the multi-task object detection network. The object detection network can be dynamically optimized based on the learnable weight parameters corresponding to each task, enabling the multi-task object detection network to automatically learn the uncertainty between 2D and 3D tasks and dynamically adjust the weight parameters corresponding to each task. This allows the trained object detection network to effectively cope with the uncertainty differences between tasks, adapt to dynamic changes in tasks, and reduce the impact of high-loss tasks on the total loss of the object detection network. Compared to complex manual weight parameter adjustment, the uncertainty-based dynamic weight adjustment strategy can automatically learn weights according to the distribution of training data, thereby accelerating the training speed of the multi-task object detection network and improving its overall performance.

[0082] Reference Figure 4 This diagram illustrates a multi-task target detection network architecture provided in an embodiment of this application, which specifically includes the following modules: The image input processing module performs scaling, type conversion, and image normalization operations on the original fisheye image to obtain the fisheye image to be detected. During the training phase of the multi-task object detection network, enhancement operations can also be performed on the normalized image, such as cropping and scaling, random flipping, and brightness adjustment, to enrich the training samples and improve the robustness of the multi-task object detection network.

[0083] The shared backbone feature extraction module extracts shared multi-scale target features from the input fisheye image to be detected, which are used in both 2D and 3D tasks. It can use the DLA34 (Deep Aggregation Network) as the backbone network to fully aggregate and learn features within the module and multi-scale features outside the module. DLA34 can achieve shared multi-scale target feature extraction from the fisheye image to be detected through multi-level downsampling using a recursive tree structure, and is commonly used in object detection tasks. This module is supervised by the loss functions of all 2D and 3D tasks.

[0084] The multi-scale feature fusion module for 2D desensitization tasks is used to fully fuse features at scales of 1 / 8, 1 / 16, and 1 / 32 of the original input image size from the shared multi-scale target features output by the shared backbone feature extraction module through upsampling, and finally outputs the fused features of the 2D desensitization task at the 1 / 8 scale.

[0085] The detection head module for the 2D desensitization task includes a detection head corresponding to the 2D desensitization task. This head is used to extract the 2D desensitized target features from the fused features of the 2D desensitization task using combined convolutional layers. For example, in a parking scenario, the size of the fused features for the 2D desensitization task can be 64 (number of channels) × 60 (feature map height) × 80 (feature map width). The 2D desensitized target can be a license plate or a pedestrian, which can be identified by a 1×1... The `conv` layer reduces the number of channels from 64 to 2, and adjusts the output heatmap feature size to 2 (number of channels) × 60 (feature map height) × 80 (feature map width). This means the number of heatmap feature channels is adjusted to 2, representing the 2D desensitized target (license plate or pedestrian). A 1×1 `conv` layer can be used to adjust the size of the length and width regression features of the output detection box to 4 (number of channels) × 60 (feature map height) × 80 (feature map width), correspondingly adjusting the number of feature map channels for the predicted 2D desensitized detection box width and height to 4, representing the X and Y coordinates of the two categories of 2D desensitized targets. Similarly, a 1×1 `conv` layer can be used to adjust the size of the output center point offset regression feature to 2 (number of channels) × 60 (feature map height) × 80 (feature map width), also correspondingly adjusting the number of feature map channels for the predicted center point offset in both directions to 2, representing the X and Y coordinates of the center point offset. By adjusting the number of feature channels, high-dimensional features can be transformed into the dimension set for each task's feature map, thus obtaining the corresponding 2D desensitized target features.

[0086] The multi-scale feature fusion module for common 2D target detection tasks is used to fully fuse features at scales of 1 / 8, 1 / 16, and 1 / 32 of the original input image size from the shared multi-scale target features output by the shared backbone feature extraction module, and finally outputs 2D common task fusion features at the 1 / 8 scale.

[0087] The detection head module for 2D common object detection tasks includes a detection head corresponding to the 2D common object detection task. It uses combined convolutional layers to extract 2D common object features from the fused features of the 2D common task. The network structure of this module is similar to that of the detection head module for 2D desensitization tasks; only the output dimensions of the convolutional layers for heatmap features and bounding box width and height features need to be set according to the number of 2D common object categories. Furthermore, since 2D common object detection tasks are typically demanding, the heatmap feature prediction in the detection head module can use Focal Loss for supervised learning to better handle class imbalance. The prediction of bounding box width and height features can use L1 loss (absolute loss function) for supervised learning. The method for extracting 2D common object features in the detection head module for 2D common object detection tasks is similar to that in the detection head module for 2D desensitization tasks; adjustments can be made accordingly, and this application will not repeat the description here.

[0088] A lightweight backbone feature extraction module, exclusive to the 3D detection task, is used to further extract the 3D geometric features of the target from the original fisheye image, based on the shared multi-scale target features extracted by the shared backbone feature extraction module. This results in 3D multi-scale target features, making the multi-task target detection network more robust in recognizing 3D information of targets during multi-task joint execution. This module can adopt a more lightweight network structure than the shared backbone feature extraction module, and the extracted 3D multi-scale target features are exclusive to the 3D detection task and do not participate in the execution of other 2D tasks.

[0089] The multi-scale feature fusion module for 3D object detection tasks is used to fully fuse the features at one-eighth scale of the original input image size in the shared multi-scale object features output by the shared backbone feature extraction module with the features at one-sixteenth and one-thirty-second scales in the three-dimensional multi-scale object features output by the lightweight backbone feature extraction module unique to 3D detection tasks, and finally output the three-dimensional task fusion features at one-eighth scale.

[0090] The detection head module for the 3D object detection task takes the 1 / 8 scale 3D task fusion features output by the multi-scale feature fusion module of the 3D object detection task as input and extracts 3D object features from the 3D task fusion features. This module can include a 3D center point location heatmap prediction head, a center point offset prediction head, a center point camera coordinate system depth prediction head, a camera coordinate system heading angle prediction head, and an object 3D scale prediction head. The network structures of the different task prediction heads are similar, all adjusting the high-dimensional features to the corresponding target information dimension by stacking two convolutional layers. The 3D center point location heatmap prediction head can use Focal Loss for supervised learning to better handle the class imbalance problem. The object 3D scale prediction head, which predicts the length, width, and height features of the detection box, can use the L1 loss function for supervised learning. The center point camera coordinate system depth prediction head can also use the L1 loss function for supervised learning. The camera coordinate system heading angle prediction head can use cross-entropy loss and L1 loss function for joint supervised learning. In the parking scenario, the 3D center point location heatmap prediction head can output 3 feature channels, representing the predicted 3D targets of three types: vehicles, pedestrians, and cyclists. The center point offset prediction head can have 2 feature channels, representing the precise center point offset in two directions. The center point camera coordinate system depth prediction head can have 1 feature channel. The target 3D scale prediction head can have 3 feature channels, representing the size information of the 3D target in the image, including the length, width, and height of the 3D target. The heading angle prediction head in the camera coordinate system can also include two task prediction heads: angle block classification prediction and angle offset prediction. The angle block classification prediction head is used to determine the angle region where the 3D target's orientation angle is located, and the angle offset prediction head is used to predict the offset of the 3D target relative to the center angle of each region. The number of feature channels of the heading angle prediction head can be set according to the number of angle regions. The specific adjustment method can be referred to the detection head module of the 2D desensitization task, which is implemented through 1×1 conv.

[0091] The output decoding module is used to parse the two-dimensional desensitized target features, two-dimensional common target features, and three-dimensional target features to obtain the corresponding target detection results, which are the detection boxes and target categories corresponding to the two-dimensional desensitized targets, two-dimensional common targets, and three-dimensional targets in the original fisheye image.

[0092] Reference Figure 4In parking scenarios, raw fisheye images acquired from various fisheye cameras mounted on the vehicle can be input into a multi-task object detection network. The raw fisheye images first pass through an image input processing module, which scales, converts, and normalizes the undistorted raw fisheye images. Then, the processed fisheye images to be detected are sent to a shared backbone feature extraction module. The shared backbone feature extraction module is responsible for extracting the target features of interest in the parking scenario from the images and outputting multi-scale features (shared multi-scale target features), which serve as inputs for 2D desensitized detection tasks, 3D object detection tasks, and 2D scene object detection tasks. In 2D desensitization and 2D common target detection tasks, shared multi-scale target features first pass through the multi-scale feature fusion modules of these two tasks. These modules are responsible for fusing the input multi-scale features to enhance the recognition ability of targets at different scales. Then, the fusion result (two-dimensional task fusion features) is sent to the 2D target detection head (the detection head module of the 2D desensitization task and the detection head module of the 2D common target detection task) to identify the defined target category and target location attributes. In 3D target detection tasks, shared multi-scale target features first pass through the 3D detection task-specific lightweight backbone feature extraction module. This module further extracts the geometric features of the 3D target (three-dimensional multi-scale target features). Then, the shared multi-scale target features and the three-dimensional multi-scale target features are sent to the multi-scale feature fusion module to obtain the three-dimensional task fusion features. Finally, the detection head module of the 3D target detection task is used to predict the relevant 3D information of the target. After obtaining the output results of the three tasks, they are sent to the output decoding module. This module can parse the corresponding 2D target detection box information and 3D target box information from the feature image of the corresponding task output.

[0093] Reference Figure 5 This image illustrates a visualization of target detection results provided in an embodiment of this application. Figure 5 The left side shows the detection results of the two-dimensional desensitized target detection task. The person_face appearing in the original fisheye image was effectively identified as a two-dimensional desensitized target, and the corresponding detection box was visualized in the fisheye image. Figure 5 The middle image shows the detection results of the 3D object detection task. Pedestrians and vehicles appearing in the original fisheye image were effectively identified as 3D objects, and the corresponding detection boxes of the 3D objects were visualized in the fisheye image. Figure 5 The right side shows the detection results of a common 2D target detection task. Several common 2D targets, including pedestrians, pillars, and front tires, were effectively identified, and the corresponding detection boxes were visualized in the fisheye image.

[0094] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily necessary for the embodiments of this application.

[0095] Reference Figure 6 The diagram shows a structural block diagram of an embodiment of a target detection device according to this application, which may specifically include the following modules: The acquisition module 601 is used to acquire the fisheye image to be detected; The first extraction module 602 is used to extract shared multi-scale target features from the fisheye image to be detected; the shared multi-scale target features include multiple shared target features of different scales, which are used for two-dimensional target detection tasks and three-dimensional target detection tasks. The two-dimensional fusion module 603 is used to perform feature fusion on the shared target features at multiple different scales to obtain two-dimensional task fusion features; The second extraction module 604 is used to extract three-dimensional multi-scale target features from the shared multi-scale target features; the three-dimensional multi-scale target features include multiple three-dimensional target features at different scales; The 3D fusion module 605 is used to perform feature fusion on the shared target features and the 3D target features to obtain 3D task fusion features; A two-dimensional target feature extraction module 606 is used to extract two-dimensional target features from the two-dimensional task fusion features; The three-dimensional target feature extraction module 607 is used to extract three-dimensional target features from the three-dimensional task fusion features; The parsing module 608 is used to parse the two-dimensional target features and the three-dimensional target features respectively to obtain the corresponding target detection results.

[0096] The acquisition module is further used for: Acquire raw fisheye images; the raw fisheye images include the target to be detected; The original fisheye image is scaled. The scaled original fisheye image is converted into a floating-point fisheye image; The floating-point fisheye image is normalized to generate the fisheye image to be detected.

[0097] The 3D fusion module includes: The first determining submodule is used to determine low-level shared target features from the shared target features; The second determining submodule is used to determine the middle-level three-dimensional target features and the high-level three-dimensional target features from the three-dimensional target features; The fusion submodule is used to perform feature fusion on the low-level shared target features, the mid-level 3D target features, and the high-level 3D target features to obtain the 3D task fusion features.

[0098] The fusion submodule is further used for: The low-level shared target features are downsampled and fused with the mid-level three-dimensional target features to obtain the first three-dimensional fused features; The first three-dimensional fusion feature is downsampled and fused with the high-level three-dimensional target feature to obtain the second three-dimensional fusion feature. The second 3D fusion feature is upsampled and fused with the first 3D fusion feature to obtain the third 3D fusion feature. The first three-dimensional fusion feature is upsampled and fused with the low-level shared target feature to obtain the fourth three-dimensional fusion feature. The third 3D fusion feature is upsampled and then fused with the fourth 3D fusion feature to obtain the 3D task fusion feature.

[0099] The two-dimensional task fusion features include two-dimensional desensitized task fusion features and two-dimensional common task fusion features. The two-dimensional target feature extraction module includes: The two-dimensional desensitization target feature extraction submodule is used to perform two-dimensional desensitization target detection on the fused features of the two-dimensional desensitization task and extract the two-dimensional desensitization target features. The two-dimensional common target feature extraction submodule is used to perform two-dimensional common target detection on the two-dimensional common task fusion features and extract two-dimensional common target features. The three-dimensional target feature extraction module includes: The three-dimensional target feature extraction submodule is used to perform three-dimensional target detection on the three-dimensional task fusion features and extract three-dimensional target features.

[0100] The two-dimensional target features include two-dimensional desensitized target features and two-dimensional common target features. The parsing module includes: The first parsing submodule is used to parse the features of the two-dimensional desensitized target to obtain the detection box information of the two-dimensional desensitized target; the two-dimensional desensitized target includes at least one of license plates and pedestrians; The second parsing submodule is used to parse the features of the two-dimensional common targets to obtain the detection box information of the two-dimensional common targets; The third parsing submodule is used to parse the three-dimensional target features to obtain the detection box information of the three-dimensional target; the three-dimensional target includes at least one of vehicles, pedestrians and cyclists.

[0101] The first parsing submodule is further configured to: The target type of the two-dimensional desensitization target is determined by analyzing its features; Heatmap prediction is performed on the features of the two-dimensional desensitized target to obtain the uncalibrated center point of the two-dimensional desensitized target; The predicted center position offset of the two-dimensional desensitized target is used as the first offset. The coordinates of the uncalibrated center point of the two-dimensional desensitized target are calibrated based on the first offset to obtain the coordinates of the center point of the two-dimensional desensitized target; The detection frame information of the two-dimensional desensitized target is determined based on the center point coordinates of the target.

[0102] The second parsing submodule is further used for: Analyze the features of the two-dimensional common targets to determine their target types; Heatmap prediction is performed on the features of the two-dimensional common targets to obtain the uncalibrated center point of the two-dimensional common targets; The predicted center position offset of the common two-dimensional target is used as the second offset. The coordinates of the uncalibrated center point of the two-dimensional common target are calibrated according to the second offset to obtain the coordinates of the center point of the two-dimensional common target; The detection bounding box information of the two-dimensional common target is determined based on the center point coordinates of the two-dimensional common target.

[0103] The third parsing submodule includes: A three-dimensional target feature heatmap prediction unit is used to perform heatmap prediction on the three-dimensional target features to obtain the uncalibrated center point of the three-dimensional target; The third offset prediction unit is used to predict the center position offset of the three-dimensional target as the third offset. A three-dimensional target center point calibration unit is used to calibrate the coordinates of the uncalibrated center point of the three-dimensional target according to the third offset, so as to obtain the center point coordinates of the three-dimensional target; A three-dimensional target feature parsing unit is used to parse the three-dimensional target features to determine the size parameters, center point camera coordinate system depth, and first heading angle of the three-dimensional target; A three-dimensional target detection box parameter determination unit is used to determine the detection box parameters of the three-dimensional target based on the center point coordinates of the three-dimensional target, the size parameters, the depth of the center point camera coordinate system, and the first heading angle.

[0104] The three-dimensional target detection box parameter determination unit is further used for: Obtain the calibration intrinsic and extrinsic parameters of the fisheye camera; the fisheye camera is used to acquire raw fisheye images; The three-dimensional center point coordinates of the three-dimensional target in the camera coordinate system are calculated based on the calibration intrinsic parameters of the fisheye camera, the center point coordinates of the three-dimensional target, and the depth of the center point camera coordinate system. The coordinates of the three-dimensional center point of the three-dimensional target in the camera coordinate system and the calibration extrinsic parameters of the fisheye camera are used to calculate the coordinates of the three-dimensional center point of the three-dimensional target in the vehicle coordinate system. The first heading angle is converted into a second heading angle in the vehicle coordinate system; the first heading angle is the heading angle of the three-dimensional target in the camera coordinate system; The detection frame parameters of the three-dimensional target are calculated based on the size parameters of the three-dimensional target, the coordinates of the three-dimensional center point in the vehicle coordinate system, and the second heading angle.

[0105] The target detection device is applied to an electronic device, which includes a multi-task target detection network. The tasks include the two-dimensional target detection task and the three-dimensional target detection task. The device also includes: The training module is used to train the target detection network, including: adding a corresponding two-dimensional detection auxiliary head to the detection head module of the three-dimensional target detection task; the two-dimensional detection auxiliary head is used to detect the two-dimensional features of the target to be detected in the three-dimensional target detection task.

[0106] The training module also includes: An initialization submodule is used to initialize the learnable weight parameters corresponding to each of the tasks. The loss value determination submodule is used to determine the loss value corresponding to each of the tasks. The total loss determination submodule is used to determine the total loss of the task based on the loss value and the learnable weight parameters. The adjustment submodule is used to adjust the learnable weight parameters based on the total loss.

[0107] This application provides a target detection device that can perform target detection through an end-to-end multi-task target detection network. The tasks can include two-dimensional target detection tasks and three-dimensional target detection tasks. During target detection, an undistorted fisheye image of the target can be acquired. Shared multi-scale target features are extracted from the fisheye image. These shared multi-scale target features include multiple shared target features at different scales and can be used for both two-dimensional and three-dimensional target detection tasks. Feature fusion of these shared target features at different scales yields two-dimensional task fusion features. Three-dimensional multi-scale target features can be extracted from the shared multi-scale target features. These three-dimensional multi-scale target features include multiple three-dimensional target features at different scales. Feature fusion of the shared target features and the three-dimensional target features yields three-dimensional task fusion features. Two-dimensional and three-dimensional target features can then be extracted from the obtained two-dimensional and three-dimensional task fusion features, respectively. The two-dimensional and three-dimensional target features are then analyzed to obtain the corresponding target detection results. Through the above implementation process, using the undistorted fisheye image to be detected, compared with using the distorted fisheye image for target detection, can significantly reduce the consumption of computing resources in the target detection process. In addition, by introducing three-dimensional multi-scale target features, the difficulty of three-dimensional target detection caused by the undistorted fisheye image can be improved, and the ability to recognize the three-dimensional information of the target in subsequent target detection tasks can be improved. By extracting two-dimensional target features and three-dimensional target features from the two-dimensional task fusion features and the three-dimensional task fusion features respectively, the target detection results of two-dimensional and three-dimensional tasks can be obtained separately. Therefore, it is not necessary to configure separate two-dimensional and three-dimensional target recognition networks to ensure effective recognition of rich two-dimensional and three-dimensional targets around the vehicle.

[0108] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0109] This application also provides an electronic device, such as... Figure 7 As shown, it includes a processor 701, a device interface 702, a memory 703, and a bus 704; Memory 703 is used to store computer programs; The processor 701 performs the above steps when executing the program stored in the memory 703.

[0110] The bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0111] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0112] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0113] This application also provides a computer-readable storage medium having executable code stored thereon, which, when executed by a processor, enables the processor to perform the target detection method of the foregoing embodiments.

[0114] This application also provides a vehicle that includes the aforementioned electronic equipment.

[0115] The algorithms and displays provided herein are not inherently related to any particular computer, virtual device, or other equipment. The structure required to construct such a device is obvious from the above description. Furthermore, this application is not directed to any particular programming language. It should be understood that the content of this application described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of this application.

[0116] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0117] Similarly, it should be understood that, in order to simplify this application and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of this application, various features of this application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this method of disclosure should not be construed as reflecting an intention that the claimed application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of this application.

[0118] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0119] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the sequencing device according to this application. This application can also be implemented as a device or apparatus program for performing part or all of the methods described herein. Such an implementation of this application can be stored on a computer-readable medium, or can take the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0120] It should be noted that the above embodiments are illustrative of this application and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

[0121] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0122] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.

[0123] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0124] It should be noted that the various data-related processes in the embodiments of this application are carried out in compliance with the relevant data protection laws and policies of the country where the location is located, and with the authorization granted by the owner of the corresponding device.

Claims

1. A target detection method, characterized in that, The method includes: Acquire the fisheye image to be detected; Shared multi-scale target features are extracted from the fisheye image to be detected; the shared multi-scale target features include multiple shared target features of different scales, which are used for two-dimensional target detection tasks and three-dimensional target detection tasks. Feature fusion is performed on the shared target features at multiple different scales to obtain two-dimensional task fusion features; Three-dimensional multi-scale target features are extracted from the shared multi-scale target features; the three-dimensional multi-scale target features include multiple three-dimensional target features at different scales; The shared target features and the three-dimensional target features are fused to obtain three-dimensional task fusion features; Extract two-dimensional target features from the two-dimensional task fusion features; Extract three-dimensional target features from the three-dimensional task fusion features; The corresponding target detection results are obtained by analyzing the two-dimensional target features and the three-dimensional target features respectively.

2. The method according to claim 1, characterized in that, The acquisition of the fisheye image to be detected includes: Acquire raw fisheye images; the raw fisheye images include the target to be detected; The original fisheye image is scaled. The scaled original fisheye image is converted into a floating-point fisheye image; The floating-point fisheye image is normalized to generate the fisheye image to be detected.

3. The method according to claim 1, characterized in that, The feature fusion of the shared target features and the three-dimensional target features to obtain three-dimensional task fusion features includes: Determine low-level shared target features from the shared target features; The mid-level and high-level three-dimensional target features are determined from the aforementioned three-dimensional target features; The low-level shared target features, the mid-level 3D target features, and the high-level 3D target features are fused to obtain the 3D task fusion features.

4. The method according to claim 3, characterized in that, The feature fusion of the low-level shared target features, the mid-level 3D target features, and the high-level 3D target features to obtain the 3D task fusion features includes: The low-level shared target features are downsampled and fused with the mid-level three-dimensional target features to obtain the first three-dimensional fused features; The first three-dimensional fusion feature is downsampled and fused with the high-level three-dimensional target feature to obtain the second three-dimensional fusion feature. The second 3D fusion feature is upsampled and fused with the first 3D fusion feature to obtain the third 3D fusion feature. The first three-dimensional fusion feature is upsampled and fused with the low-level shared target feature to obtain the fourth three-dimensional fusion feature. The third 3D fusion feature is upsampled and then fused with the fourth 3D fusion feature to obtain the 3D task fusion feature.

5. The method according to claim 1, characterized in that, The two-dimensional task fusion features include two-dimensional de-identified task fusion features and two-dimensional common task fusion features. Extracting two-dimensional target features from the two-dimensional task fusion features includes: The two-dimensional desensitization task fusion features are used to perform two-dimensional desensitization target detection and extract two-dimensional desensitization target features; Two-dimensional common target detection is performed on the aforementioned two-dimensional common task fusion features to extract two-dimensional common target features; The step of extracting 3D target features from the 3D task fusion features includes: The three-dimensional target features are then used to perform three-dimensional target detection and extract the three-dimensional target features.

6. The method according to claim 1, characterized in that, The two-dimensional target features include two-dimensional desensitized target features and two-dimensional common target features. The step of parsing the two-dimensional target features and the three-dimensional target features to obtain the corresponding target detection results includes: The detection box information of the two-dimensional desensitized target is obtained by parsing the features of the two-dimensional desensitized target; the two-dimensional desensitized target includes at least one of license plates and pedestrians; The detection bounding box information of the two-dimensional common target is obtained by parsing the features of the two-dimensional common target; The detection bounding box information of the three-dimensional target is obtained by parsing the features of the three-dimensional target; the three-dimensional target includes at least one of vehicles, pedestrians and cyclists.

7. The method according to claim 6, characterized in that, The step of parsing the two-dimensional desensitized target features to obtain the detection box information of the two-dimensional desensitized target includes: The target type of the two-dimensional desensitization target is determined by analyzing its features; Heatmap prediction is performed on the features of the two-dimensional desensitized target to obtain the uncalibrated center point of the two-dimensional desensitized target; The predicted center position offset of the two-dimensional desensitized target is used as the first offset. The coordinates of the uncalibrated center point of the two-dimensional desensitized target are calibrated based on the first offset to obtain the coordinates of the center point of the two-dimensional desensitized target; The detection frame information of the two-dimensional desensitized target is determined based on the center point coordinates of the target.

8. The method according to claim 6, characterized in that, The step of parsing the features of common two-dimensional targets to obtain detection bounding box information for common two-dimensional targets includes: Analyze the features of the two-dimensional common targets to determine their target types; Heatmap prediction is performed on the features of the two-dimensional common targets to obtain the uncalibrated center point of the two-dimensional common targets; The predicted center position offset of the common two-dimensional target is used as the second offset. The coordinates of the uncalibrated center point of the two-dimensional common target are calibrated according to the second offset to obtain the coordinates of the center point of the two-dimensional common target; The detection bounding box information of the two-dimensional common target is determined based on the center point coordinates of the two-dimensional common target.

9. The method according to claim 6, characterized in that, The process of parsing the 3D target features to obtain the detection bounding box information of the 3D target includes: Heatmap prediction is performed on the features of the three-dimensional target to obtain the uncalibrated center point of the three-dimensional target; The predicted center position offset of the three-dimensional target is used as the third offset. The coordinates of the uncalibrated center point of the three-dimensional target are calibrated based on the third offset to obtain the coordinates of the center point of the three-dimensional target. The size parameters, center point camera coordinate system depth, and first heading angle of the three-dimensional target are determined by analyzing the three-dimensional target features. The detection frame parameters of the three-dimensional target are determined based on the center point coordinates of the three-dimensional target, the size parameters, the depth of the center point camera coordinate system, and the first heading angle.

10. The method according to claim 9, characterized in that, The step of determining the detection box parameters of the 3D target based on the center point coordinates, the size parameters, the depth of the center point camera coordinate system, and the first heading angle includes: Obtain the calibration intrinsic and extrinsic parameters of the fisheye camera; the fisheye camera is used to acquire raw fisheye images; The three-dimensional center point coordinates of the three-dimensional target in the camera coordinate system are calculated based on the calibration intrinsic parameters of the fisheye camera, the center point coordinates of the three-dimensional target, and the depth of the center point camera coordinate system. The coordinates of the three-dimensional center point of the three-dimensional target in the camera coordinate system and the calibration extrinsic parameters of the fisheye camera are used to calculate the coordinates of the three-dimensional center point of the three-dimensional target in the vehicle coordinate system. The first heading angle is converted into a second heading angle in the vehicle coordinate system; the first heading angle is the heading angle of the three-dimensional target in the camera coordinate system; The detection frame parameters of the three-dimensional target are calculated based on the size parameters of the three-dimensional target, the coordinates of the three-dimensional center point in the vehicle coordinate system, and the second heading angle.

11. The method according to claim 1, characterized in that, The method is applied to an electronic device, the electronic device including a multi-task target detection network, the tasks including a two-dimensional target detection task and a three-dimensional target detection task, the target detection network including a detection head module for the three-dimensional target detection task, the method including: Training the target detection network includes: adding a corresponding two-dimensional detection auxiliary head to the detection head module of the three-dimensional target detection task; the two-dimensional detection auxiliary head is used to detect the two-dimensional features of the target in the three-dimensional target detection task.

12. A target detection device, characterized in that, The device includes: The acquisition module is used to acquire the fisheye image to be detected; The first extraction module is used to extract shared multi-scale target features from the fisheye image to be detected; the shared multi-scale target features include multiple shared target features of different scales, which are used for two-dimensional target detection tasks and three-dimensional target detection tasks. The two-dimensional fusion module is used to fuse the shared target features at multiple different scales to obtain two-dimensional task fusion features; The second extraction module is used to extract three-dimensional multi-scale target features from the shared multi-scale target features; the three-dimensional multi-scale target features include multiple three-dimensional target features at different scales; A 3D fusion module is used to fuse the shared target features and the 3D target features to obtain 3D task fusion features; A two-dimensional target feature extraction module is used to extract two-dimensional target features from the two-dimensional task fusion features; A three-dimensional target feature extraction module is used to extract three-dimensional target features from the three-dimensional task fusion features; The parsing module is used to parse the two-dimensional target features and the three-dimensional target features respectively to obtain the corresponding target detection results.

13. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the instructions to implement the target detection method as described in any one of claims 1 to 11.

14. A computer-readable storage medium having executable code stored thereon, which, when executed, causes a processor to perform the target detection method as described in any one of claims 1 to 11.

15. A vehicle, characterized in that, The vehicle includes the electronic equipment as described in claim 13.