Three-dimensional object detection and fusion method based on multi-sensor fusion
By constructing a three-dimensional object detection method for multi-sensor fusion, using cameras and lidar data to build bounding boxes, and combining deep learning models for multi-stage refined processing, the information redundancy and loss problems in multi-source data fusion are solved, and the accuracy and reliability of detection results are improved.
Patent Information
- Application Number
- CN202510780176.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing three-dimensional object detection and fusion methods based on multi-sensor fusion are difficult to achieve accurate fusion of multi-source data due to the complex heterogeneous sensor data, resulting in redundancy and loss of information, affecting the reliability of detection results.
Available multi-directional two-dimensional image data through the camera and acquiring point cloud data through the lidar, constructing two-dimensional and three-dimensional bounding boxes for matching and optimization, combining deep learning models for multi-stage refined processing, adopting a two-bounding box cross-verification mechanism and dynamic sorting and screening strategy to achieve optimized fusion of multi-source data.
It significantly improves the accuracy of three-dimensional object detection, solves the problem of low matching accuracy of multi-source data, and enhances the reliability and robustness of detection results.
Smart Images

Figure CN120299021B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a three-dimensional target detection and fusion method based on multi-sensor fusion. Background Art
[0002] In areas such as autonomous driving and intelligent monitoring, a single sensor's detection of three-dimensional targets faces problems such as insufficient information and limited accuracy. Multi-sensor fusion has become a key direction for improving detection performance.
[0003] Existing three-dimensional target detection and fusion methods based on multi-sensor fusion are difficult to make multi-source data accurate due to the complexity of heterogeneous sensor data. In addition, there are problems of information redundancy and loss in the fusion process, which leads to defects in fusion accuracy and target geometric information modeling, resulting in low reliability of detection results. Therefore, it is necessary to provide a three-dimensional target detection and fusion method based on multi-sensor fusion to solve the above problems. Summary of the Invention
[0004] In order to solve the above technical problems, a three-dimensional target detection and fusion method based on multi-sensor fusion is provided. This technical solution solves the existing three-dimensional target detection and fusion method based on multi-sensor fusion proposed in the above background technology. Due to the complexity of heterogeneous sensor data, it is difficult to make multi-source data accurate, and there are problems of information redundancy and loss in the fusion process, resulting in defects in fusion accuracy and target geometric information modeling, leading to low reliability of detection results.
[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0006] Three-dimensional target detection and fusion methods based on multi-sensor fusion include:
[0007] Acquire multi-directional two-dimensional image data through the camera, and simultaneously acquire point cloud data through the lidar;
[0008] Based on the multi-directional two-dimensional image data and point cloud data, a two-dimensional bounding box and a three-dimensional bounding box are constructed respectively;
[0009] Matching the two-dimensional bounding box with the three-dimensional bounding box to determine the preferred two-dimensional image and the preferred point cloud data;
[0010] Fusing the optimal two-dimensional image and the optimal point cloud data to obtain a two-dimensional fused image to be measured;
[0011] Obtaining the recognition accuracy of the two-dimensional fusion image to be tested;
[0012] The 2D fused images to be tested are sorted according to the recognition accuracy, and then the sorted 2D fused images to be tested are extracted in sequence. The preliminary 3D target detection results are determined by combining them with the corresponding preferred point cloud data.
[0013] Based on the preliminary 3D target detection results, the 3D target geometric information is set, and the standard 3D target detection results are determined synchronously based on the 3D target geometric information, the 2D fused image to be tested, and the corresponding preferred point cloud data.
[0014] In an optional embodiment, constructing a two-dimensional bounding box specifically includes:
[0015] Based on the multi-directional two-dimensional image data, traverse the corresponding cameras to determine the camera position information;
[0016] According to the camera position information, a reference plane is constructed, and the coordinates corresponding to the camera position information are connected in the reference plane to obtain the camera network area information;
[0017] Based on the camera network area information, the geometric center of gravity of the camera network area is determined as the origin to establish a three-dimensional space coordinate system;
[0018] Extracting multi-directional two-dimensional images from the multi-directional two-dimensional image data in sequence, and labeling the multi-directional two-dimensional images with orientation labels based on the three-dimensional space coordinate system and the corresponding camera position information;
[0019] Performing image processing on the multi-directional two-dimensional image to obtain a binary image corresponding to the multi-directional two-dimensional image;
[0020] Obtaining an area value of a pixel with a grayscale value of 255 in a binary image corresponding to the multi-directional two-dimensional image as a first reference value for selecting a two-dimensional bounding box;
[0021] Establishing an image coordinate system in the binary image corresponding to the multi-directional two-dimensional image with the largest first reference value selected in the two-dimensional bounding box, wherein the image coordinate system is a plane rectangular coordinate system, and the origin coordinates of the image coordinate system are the center point of the binary image;
[0022] Traverse the pixel points with a grayscale value of 255 in the binary image and obtain their coordinates (xi, yi) in the image coordinate system;
[0023] Determine the two pixel points with the largest absolute value of xi when xi is greater than 0 and less than 0, and use the mapped lengths of the two pixel points as the length of the interception window. Determine the two pixel points with the largest absolute value of yi when yi is greater than 0 and less than 0, and use the mapped heights of the two pixel points as the width of the interception window.
[0024] Using a capture window to capture all multi-directional two-dimensional images in the multi-directional two-dimensional image data to obtain a multi-directional two-dimensional sub-image set;
[0025] Extract multi-directional 2D sub-images from the multi-directional 2D sub-image set in sequence, and obtain the binary images corresponding to the multi-directional 2D sub-images. Traverse the pixels with a grayscale value of 255 in the binary image, then calculate the area of the white connected regions in the binary image, mark the area with the largest white connected region, and obtain the 2D bounding box reference area.
[0026] Get the edge points of the reference area of the 2D bounding box, synchronously traverse all edge points of the white connected area, connect the two edge points with the greatest distance, and get the height of the 2D bounding box;
[0027] Using the height of the 2D bounding box as the dividing line, set a straight line perpendicular to the dividing line to slide through the edge points on both ends of the dividing line in the white connected area. The sliding unit is 1 pixel. Then, the two edge points located farthest from each other on the straight line are connected to obtain the width of the 2D bounding box, completing the construction of the 2D bounding box.
[0028] In an optional embodiment, the construction of the three-dimensional bounding box specifically includes:
[0029] Identify and remove isolated points in the point cloud data using statistical methods, then merge the dense point cloud into a point set that retains the structure to obtain standard point cloud data;
[0030] The minimum cluster size is set. Based on the standard point cloud data, the HDBSCAN clustering algorithm is applied to identify dense areas as clusters and mark low-density points as noise. Each point is assigned a label, and the noise point label is -1, to obtain a 3D point cloud cluster.
[0031] Determine the number of clusters, calculate the minimum and maximum x-coordinates of all points in each cluster, process the y- and z-coordinates simultaneously, combine them into a bounding box, and form a three-dimensional bounding reference cuboid;
[0032] Use the 3D bounding reference cuboid as the 3D bounding box.
[0033] In an optional embodiment, matching the two-dimensional bounding box with the three-dimensional bounding box to determine the preferred two-dimensional image and the preferred point cloud data specifically includes:
[0034] Obtain the coordinates of the 3D bounding box, and simultaneously map the 3D bounding box to the corresponding multi-directional 2D image to obtain a selected 2D image;
[0035] Use the two-dimensional bounding box to slide and intercept the two-dimensional image to be selected to obtain the two-dimensional sub-image to be selected;
[0036] Traversing the candidate two-dimensional sub-images, calculating the number of three-dimensional bounding boxes in the candidate two-dimensional sub-images, and obtaining a first reference value of the preferred two-dimensional image;
[0037] summing the first reference values of the preferred two-dimensional images corresponding to the two-dimensional images to be selected to obtain a second reference value of the preferred two-dimensional image;
[0038] removing the candidate two-dimensional images with the largest and smallest second reference values of the preferred two-dimensional image, and taking the candidate two-dimensional image with the largest second reference value of the preferred two-dimensional image among the retained candidate two-dimensional images as the preferred two-dimensional image;
[0039] The three-dimensional point cloud cluster corresponding to the three-dimensional bounding box in the preferred two-dimensional image is used as the preferred point cloud data.
[0040] In an optional embodiment, fusing the preferred two-dimensional image and the preferred point cloud data to obtain the two-dimensional fused image to be measured specifically includes:
[0041] Determine the optimal two-dimensional image and the multi-directional two-dimensional image corresponding to the optimal point cloud data, and obtain the multi-directional two-dimensional image to be measured;
[0042] Obtain a binary image of the multi-directional two-dimensional image to be measured, and obtain the label value corresponding to the point in the corresponding preferred point cloud data;
[0043] Synchronously mapping points with the same label value to the binary image of the multi-directional two-dimensional image to be measured, and marking pixels with a grayscale value of 255 corresponding to the points with the same label value in the binary image of the multi-directional two-dimensional image to be measured with the same color;
[0044] The two-dimensional bounding box corresponding to the preferred two-dimensional image and the three-dimensional bounding box corresponding to the preferred point cloud data are integrated into the marked multi-directional two-dimensional image to be tested by using computer vision to obtain the two-dimensional fused image to be tested.
[0045] In an optional embodiment, obtaining the recognition accuracy of the two-dimensional fused image to be tested specifically includes:
[0046] Use the two-dimensional bounding box to slide and intercept the two-dimensional fused image to be tested to obtain the two-dimensional fused sub-image to be tested;
[0047] Obtaining the volume of the three-dimensional bounding box in each two-dimensional fused sub-image to be tested as a first recognition accuracy reference value;
[0048] Obtaining the volumes of all three-dimensional bounding boxes in the two-dimensional fused image to be tested as a second recognition accuracy reference value;
[0049] Obtaining a ratio of the first recognition accuracy reference value to its corresponding second recognition accuracy reference value as a third recognition accuracy reference value;
[0050] Setting a recognition accuracy reference threshold, and changing the grayscale value of the two-dimensional fused sub-image to be tested whose third recognition accuracy reference value is less than or equal to the recognition accuracy reference threshold to 0 in the two-dimensional fused image to be tested, to obtain a two-dimensional fused standard image to be tested;
[0051] The area occupied by the pixels whose grayscale values are not 0 in the two-dimensional fusion standard image to be tested is calculated, and the area is quantified to obtain the recognition accuracy of the two-dimensional fusion image to be tested.
[0052] In an optional embodiment, the sorting of the two-dimensional fused images to be tested according to recognition accuracy specifically includes:
[0053] Based on the recognition accuracy of the two-dimensional fused images to be tested, the two-dimensional fused images to be tested are sorted in descending order.
[0054] In an optional embodiment, sequentially extracting the sorted two-dimensional fused images to be tested and combining them with the corresponding preferred point cloud data to determine the preliminary three-dimensional target detection results specifically includes:
[0055] Set identification tags and determine the number of identification tags;
[0056] Sequentially extract n sorted 2D fusion images to be tested, and use the deep learning model to perform n inspections on the multi-directional 2D images corresponding to the 2D fusion images to be tested, to obtain preliminary inspection 2D bounding boxes and preliminary inspection 3D bounding boxes;
[0057] Traverse the initial inspection identification label corresponding to the initial inspection two-dimensional bounding box at the nth inspection, and determine the initial inspection identification label Label (n) at the nth inspection and the initial inspection identification label Label (n-1) at the n-1th inspection;
[0058] If the number of similarities and differences between Label (n) and Label (n-1) is greater than or equal to the number of identification labels, the corresponding two-dimensional fusion image to be tested during the nth detection is removed;
[0059] Continue to detect the remaining m two-dimensional fusion images to be tested until the detection of m two-dimensional fusion images to be tested is completed, and the number of similarities and differences between the initial detection identification label Label (m) at the m-th detection and the initial detection identification label Label (m-1) at the m-1-th detection is equal to 0, and the current result is used as the preliminary three-dimensional target detection result.
[0060] In an optional embodiment, the setting of three-dimensional target geometric information based on the preliminary three-dimensional target detection result specifically includes:
[0061] Based on the preliminary 3D target detection results, the target type is determined, and the target geometry information is set based on the target type, including the length, width, and height of the detected target.
[0062] In an optional embodiment, determining a standard 3D target detection result based on the 3D target geometric information, the 2D fused image to be detected, and the corresponding preferred point cloud data specifically includes:
[0063] Mapping the 3D target geometric information to the 2D fused image to be measured, and obtaining the 3D target geometric volume value in the 2D fused image to be measured;
[0064] Based on the two-dimensional fused image to be measured and the corresponding preferred point cloud data, obtaining the corresponding two-dimensional bounding box and three-dimensional bounding box;
[0065] The two-dimensional bounding box is used to intercept the two-dimensional fusion image to be tested, and the two-dimensional fusion standard sub-image to be tested is obtained again;
[0066] Based on the two-dimensional fusion standard sub-image to be tested, obtaining the recognition accuracy of the two-dimensional fusion image to be tested;
[0067] Obtaining the volume value of the 3D bounding box in the 2D fused standard sub-image to be tested;
[0068] If the difference between the volume value of the three-dimensional bounding box in the two-dimensional fusion standard sub-image to be tested and the geometric volume value of the three-dimensional target in the two-dimensional fusion image to be tested is the smallest, and the recognition accuracy of the two-dimensional fusion image to be tested is the largest, then the initial inspection recognition label of the initial inspection two-dimensional bounding box corresponding to the two-dimensional fusion image to be tested is obtained as the standard three-dimensional target detection result.
[0069] Compared with the prior art, the present invention has the following beneficial effects:
[0070] The proposed 3D object detection and fusion method based on multi-sensor fusion achieves fusion optimization of multi-source data through a dual bounding box cross-validation mechanism and a dynamic sorting and screening strategy, avoiding the contradiction between information redundancy and loss during the fusion process, and effectively retaining key features through multi-stage refined processing.
[0071] The proposed 3D target detection and fusion method based on multi-sensor fusion achieves accurate matching of multi-sensor data by constructing a 3D spatial coordinate system and using the HDBSCAN clustering algorithm. This solves the problem of low matching accuracy for multi-source data and significantly improves the reliability of heterogeneous sensor data fusion.
[0072] The three-dimensional target detection and fusion method based on multi-sensor fusion proposed in this scheme realizes self-correction of detection results through iterative detection and closed-loop verification system based on deep learning models, effectively suppresses noise interference through label consistency verification, and ultimately improves the accuracy of three-dimensional target detection in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1This is a flow chart of the three-dimensional target detection and fusion method based on multi-sensor fusion proposed in the present invention;
[0074] Figure 2 A flowchart for constructing a two-dimensional bounding box in the present invention;
[0075] Figure 3 A flowchart for constructing a three-dimensional bounding box in the present invention;
[0076] Figure 4 This is a flow chart for determining the preferred two-dimensional image and the preferred point cloud data in the present invention. DETAILED DESCRIPTION
[0077] The following description is intended to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are merely examples, and those skilled in the art may conceive of other obvious variations.
[0078] Reference Figure 1 - Figure 4 As shown in FIG, a three-dimensional target detection and fusion method based on multi-sensor fusion includes:
[0079] Acquire multi-directional two-dimensional image data through the camera, and simultaneously acquire point cloud data through the lidar;
[0080] Based on the multi-directional two-dimensional image data and point cloud data, a two-dimensional bounding box and a three-dimensional bounding box are constructed respectively;
[0081] Matching the two-dimensional bounding box with the three-dimensional bounding box to determine the preferred two-dimensional image and the preferred point cloud data;
[0082] Fusing the optimal two-dimensional image and the optimal point cloud data to obtain a two-dimensional fused image to be measured;
[0083] Obtaining the recognition accuracy of the two-dimensional fusion image to be tested;
[0084] The 2D fused images to be tested are sorted according to the recognition accuracy, and then the sorted 2D fused images to be tested are extracted in sequence. The preliminary 3D target detection results are determined by combining them with the corresponding preferred point cloud data.
[0085] Based on the preliminary 3D target detection results, the 3D target geometric information is set, and the standard 3D target detection results are determined synchronously based on the 3D target geometric information, the 2D fused image to be tested, and the corresponding preferred point cloud data.
[0086] Furthermore, a two-dimensional bounding box is constructed, specifically including:
[0087] Based on the multi-directional two-dimensional image data, traverse the corresponding cameras to determine the camera position information;
[0088] According to the camera position information, a reference plane is constructed, and the coordinates corresponding to the camera position information are connected in the reference plane to obtain the camera network area information;
[0089] Based on the camera network area information, the geometric center of gravity of the camera network area is determined as the origin to establish a three-dimensional space coordinate system;
[0090] Extracting multi-directional two-dimensional images from the multi-directional two-dimensional image data in sequence, and labeling the multi-directional two-dimensional images with orientation labels based on the three-dimensional space coordinate system and the corresponding camera position information;
[0091] Performing image processing on the multi-directional two-dimensional image to obtain a binary image corresponding to the multi-directional two-dimensional image;
[0092] Obtaining an area value of a pixel with a grayscale value of 255 in a binary image corresponding to the multi-directional two-dimensional image as a first reference value for selecting a two-dimensional bounding box;
[0093] Establishing an image coordinate system in the binary image corresponding to the multi-directional two-dimensional image with the largest first reference value selected in the two-dimensional bounding box, wherein the image coordinate system is a plane rectangular coordinate system, and the origin coordinates of the image coordinate system are the center point of the binary image;
[0094] Traverse the pixel points with a grayscale value of 255 in the binary image and obtain their coordinates (xi, yi) in the image coordinate system;
[0095] Determine the two pixel points with the largest absolute value of xi when xi is greater than 0 and less than 0, and use the mapped lengths of the two pixel points as the length of the interception window. Determine the two pixel points with the largest absolute value of yi when yi is greater than 0 and less than 0, and use the mapped heights of the two pixel points as the width of the interception window.
[0096] Using a capture window to capture all multi-directional two-dimensional images in the multi-directional two-dimensional image data to obtain a multi-directional two-dimensional sub-image set;
[0097] Extract multi-directional 2D sub-images from the multi-directional 2D sub-image set in sequence, and obtain the binary images corresponding to the multi-directional 2D sub-images. Traverse the pixels with a grayscale value of 255 in the binary image, then calculate the area of the white connected regions in the binary image, mark the area with the largest white connected region, and obtain the 2D bounding box reference area.
[0098] Get the edge points of the reference area of the 2D bounding box, synchronously traverse all edge points of the white connected area, connect the two edge points with the greatest distance, and get the height of the 2D bounding box;
[0099] Using the height of the 2D bounding box as the dividing line, set a straight line perpendicular to the dividing line to slide through the edge points on both ends of the dividing line in the white connected area. The sliding unit is 1 pixel. Then, the two edge points located farthest from each other on the straight line are connected to obtain the width of the 2D bounding box, completing the construction of the 2D bounding box.
[0100] Specifically, when constructing a 2D bounding box, a unified 3D coordinate system is established based on camera position information, achieving geometric alignment of multi-directional images. Furthermore, the system integrates data from multiple cameras, breaking through the field of view limitations of a single perspective and improving detection capabilities in occluded scenarios. Binarization processing, connected domain analysis, and dynamic window capture are used to locate the core area of the target. A 2D bounding box that closely follows the target's contour is generated by combining edge point mapping and a sliding traversal strategy. Area threshold screening and connected domain analysis effectively suppress interference from background noise and false targets. Pixel-level edge mapping and a sliding optimization strategy are employed to reduce bounding box positioning errors.
[0101] It is understood that sequentially extracting multi-directional 2D images from multi-directional 2D image data and assigning orientation labels to these multi-directional 2D images based on a 3D coordinate system and the corresponding camera position information specifically involves the following steps: First, the specific position and orientation of each camera in 3D space must be known. This is typically achieved through a calibration process, such as using a calibration plate or a known 3D structure to calculate the camera's intrinsic and extrinsic parameters. Constructing a 3D coordinate system: Selecting a reference point (such as the camera's origin or geometric center of gravity) as the origin of the coordinate system and defining the directions of the coordinate axes (such as the x, y, and z axes) to establish a unified 3D coordinate system. Image extraction and association: Sequentially extracting each image from the multi-directional image data and determining which camera captured each image and when. Image timestamps or sequence information are required. Coordinate transformation: For each image, transform it into a unified 3D coordinate system based on the position and orientation of its corresponding camera (converting the image's pixel coordinates to coordinates in the world coordinate system). Orientation label generation: Based on the transformed coordinate information, an orientation label (including the image's position (such as coordinates) and direction (such as orientation angle) in three-dimensional space) is generated for each image.
[0102] Furthermore, the construction of the 3D bounding box specifically includes:
[0103] Identify and remove isolated points in the point cloud data using statistical methods, then merge the dense point cloud into a point set that retains the structure to obtain standard point cloud data;
[0104] The minimum cluster size is set. Based on the standard point cloud data, the HDBSCAN clustering algorithm is applied to identify dense areas as clusters and mark low-density points as noise. Each point is assigned a label, and the noise point label is -1, to obtain a 3D point cloud cluster.
[0105] Determine the number of clusters, calculate the minimum and maximum x-coordinates of all points in each cluster, process the y- and z-coordinates simultaneously, combine them into a bounding box, and form a three-dimensional bounding reference cuboid;
[0106] Use the 3D bounding reference cuboid as the 3D bounding box.
[0107] Specifically, HDBSCAN constructs a minimum spanning tree (describing the distance relationships between points) of the point cloud, analyzes the tree's stability, and automatically merges or splits subtrees to form clusters. The algorithm controls the clustering granularity through parameters (such as the minimum cluster size), ultimately outputting a cluster label for each point (such as "belongs to cluster 1," "belongs to cluster 2," or "noise"). The operation steps include parameter setting: Setting a minimum cluster size (e.g., 100 points) is used by the algorithm to determine which density areas are considered independent clusters. Performing clustering: Running HDBSCAN automatically identifies dense areas as clusters and labels low-density points as noise. Extracting cluster labels: Each point is assigned a label; points with the same label belong to the same cluster; noise points are labeled -1. The number of clusters is determined automatically, aiming to directly obtain the number of valid clusters from the clustering results. First, points labeled as noise must be excluded. The number of unique values for the remaining labels is the total number of clusters. For example, if the labels are 0, 1, and 2, there are three clusters.
[0108] Furthermore, the two-dimensional bounding box and the three-dimensional bounding box are matched to determine the preferred two-dimensional image and the preferred point cloud data, specifically including:
[0109] Obtain the coordinates of the 3D bounding box, and simultaneously map the 3D bounding box to the corresponding multi-directional 2D image to obtain a selected 2D image;
[0110] Use the two-dimensional bounding box to slide and intercept the two-dimensional image to be selected to obtain the two-dimensional sub-image to be selected;
[0111] Traversing the candidate two-dimensional sub-images, calculating the number of three-dimensional bounding boxes in the candidate two-dimensional sub-images, and obtaining a first reference value of the preferred two-dimensional image;
[0112] summing the first reference values of the preferred two-dimensional images corresponding to the two-dimensional images to be selected to obtain a second reference value of the preferred two-dimensional image;
[0113] removing the candidate two-dimensional images with the largest and smallest second reference values of the preferred two-dimensional image, and taking the candidate two-dimensional image with the largest second reference value of the preferred two-dimensional image among the retained candidate two-dimensional images as the preferred two-dimensional image;
[0114] The three-dimensional point cloud cluster corresponding to the three-dimensional bounding box in the preferred two-dimensional image is used as the preferred point cloud data.
[0115] Specifically, the coordinates of the 3D bounding box are obtained and simultaneously mapped to the corresponding multi-directional 2D image to obtain the candidate 2D image. This is done by using the camera calibration parameters (intrinsic and extrinsic matrices) to project each vertex of the 3D bounding box from the world coordinate system onto the 2D image plane of the corresponding camera, generating a 2D projection frame. Based on the position and size of the projection frame, a region containing the complete 3D bounding box is captured from the original 2D image as the candidate 2D image. Sliding capture of the candidate 2D image using the 2D bounding box involves sliding a window with a fixed step size (e.g., 50 pixels) across the candidate 2D image to capture multiple sub-images of a fixed size (e.g., 256×256 pixels). This ensures that the sub-images cover different locations and scales of the target area, capturing local features. Calculating the number of 3D bounding boxes (the first reference value) involves object detection: applying a lightweight detection model (e.g., YOLOv5s) to each sub-image and counting the number of detected 3D bounding boxes. Counting logic: If a sub-image completely contains the core region of the 3D bounding box (e.g., the center point falls within the sub-image), the count is incremented by 1. The step of obtaining the second reference value includes cumulative statistics: adding the first reference values (number of three-dimensional bounding boxes) of all sub-images of the same candidate two-dimensional image to obtain the second reference value of the image. Example: If an image contains 3 sub-images, and 2, 1, and 3 bounding boxes are detected respectively, the second reference value is 6. Screening the preferred two-dimensional image includes outlier removal: removing the candidate images with the largest and smallest second reference values (extreme values may be caused by false detection or missed detection). Optimal selection: selecting the image with the highest second reference value from the remaining images as the preferred two-dimensional image. Extracting the preferred point cloud data includes label matching: based on the orientation label of the preferred two-dimensional image, reversely querying the point cloud cluster of the corresponding perspective in the original point cloud data. Data association: taking the matched point cloud cluster (including XYZ coordinates, reflection intensity and other information) as the preferred point cloud data.
[0116] As can be understood, projection transformation and sliding capture are used to select the 2D image with the highest degree of match to the 3D target from multi-camera data. Sliding window detection and statistics are used to filter out low-quality or irrelevant image areas (such as pure background or repeated areas). Sorting based on a second reference value prioritizes images containing complete target information, reducing redundant computation. Binding the 2D image detection results to the original point cloud data provides reliable feature input for subsequent fusion. The advantage lies in leveraging the high-resolution texture information from the camera and the depth information from the lidar to achieve cross-modal data alignment through projection mapping. By covering multiple viewpoints (such as front, side, and top views), missed detections due to target occlusion from a single viewpoint are reduced. The sliding window and outlier rejection mechanisms significantly reduce inefficient computation. Dynamic adjustment of projection parameters (such as updating extrinsic parameters based on vehicle motion) allows for adaptation to complex driving scenarios (such as bumps and sharp turns).
[0117] Furthermore, the preferred two-dimensional image and the preferred point cloud data are fused to obtain a two-dimensional fused image to be measured, which specifically includes:
[0118] Determine the optimal two-dimensional image and the multi-directional two-dimensional image corresponding to the optimal point cloud data, and obtain the multi-directional two-dimensional image to be measured;
[0119] Obtain a binary image of the multi-directional two-dimensional image to be measured, and obtain the label value corresponding to the point in the corresponding preferred point cloud data;
[0120] Synchronously mapping points with the same label value to the binary image of the multi-directional two-dimensional image to be measured, and marking pixels with a grayscale value of 255 corresponding to the points with the same label value in the binary image of the multi-directional two-dimensional image to be measured with the same color;
[0121] The two-dimensional bounding box corresponding to the preferred two-dimensional image and the three-dimensional bounding box corresponding to the preferred point cloud data are integrated into the marked multi-directional two-dimensional image to be tested by using computer vision to obtain the two-dimensional fused image to be tested.
[0122] Specifically, the integration of multi-directional 2D images involves image selection: extracting 2D images that match the selected point cloud data from different camera perspectives (e.g., front, side, and rear camera images). Coordinate alignment: Based on camera calibration parameters (intrinsic and extrinsic), the multi-directional images are uniformly projected into a global coordinate system to generate the multi-directional 2D image to be measured. Fusion strategies: Image stitching (e.g., weighted average fusion) or feature-level fusion (e.g., keypoint matching) is used to generate a comprehensive view. Binary image generation and label mapping involve binarization: applying adaptive threshold segmentation (e.g., Otsu's algorithm) to the integrated multi-directional images to generate a binary image (foreground objects are 255, background objects are 0). Label association: assigning a category label (e.g., "vehicle" or "pedestrian") to each point in the selected point cloud data, and mapping the point cloud labels to pixel coordinates in the 2D image through a projective transformation. Semantic labeling and visualization specifically involve color encoding: assigning unique colors to points with different label values (e.g., red for vehicles, blue for pedestrians), and highlighting the corresponding pixels in the binary image. Bounding box overlay: 2D bounding box: Draws a detected target rectangular box directly on the image (such as the green border). 3D bounding box projection: Projects the eight vertices of the 3D bounding box onto the 2D image using the camera projection model, connecting the vertices to generate a perspective image (such as the blue dashed box).
[0123] It is understandable that the purpose of integrating the texture features of the two-dimensional image with the geometric features of the point cloud is to form complementary information. The target category, position and spatial posture are intuitively displayed through color marking and bounding box overlay. The reliability of the detection is verified by using the consistency between the two-dimensional detection results (bounding box, label) and the point cloud data. The advantage is that a single image simultaneously presents two-dimensional semantic segmentation, three-dimensional spatial structure and multi-camera perspective information. Color coding and bounding box annotation make it easier for humans to quickly understand the detection results (such as distinguishing between vehicles and pedestrians). Multi-view data complementarity reduces the impact of occlusion (such as the side view image supplements the occluded area of the front view). The projection transformation and parallel labeling algorithm are accelerated by the GPU, with a processing delay of less than 30ms.
[0124] Furthermore, obtaining the recognition accuracy of the two-dimensional fused image to be tested specifically includes:
[0125] Use the two-dimensional bounding box to slide and intercept the two-dimensional fused image to be tested to obtain the two-dimensional fused sub-image to be tested;
[0126] Obtaining the volume of the three-dimensional bounding box in each two-dimensional fused sub-image to be tested as a first recognition accuracy reference value;
[0127] Obtaining the volumes of all three-dimensional bounding boxes in the two-dimensional fused image to be tested as a second recognition accuracy reference value;
[0128] Obtaining a ratio of the first recognition accuracy reference value to its corresponding second recognition accuracy reference value as a third recognition accuracy reference value;
[0129] Setting a recognition accuracy reference threshold, and changing the grayscale value of the two-dimensional fused sub-image to be tested whose third recognition accuracy reference value is less than or equal to the recognition accuracy reference threshold to 0 in the two-dimensional fused image to be tested, to obtain a two-dimensional fused standard image to be tested;
[0130] The area occupied by the pixels whose grayscale values are not 0 in the two-dimensional fusion standard image to be tested is calculated, and the area is quantified to obtain the recognition accuracy of the two-dimensional fusion image to be tested.
[0131] Specifically, in this step, the first reference value is calculated by calculating the volume (e.g., length × width × height) of the 3D bounding box within each sub-image, reflecting the actual spatial size of the target. The second reference value is calculated by calculating the total volume of all 3D bounding boxes in the entire fused image, reflecting the global target distribution density. The volume of each sub-image is divided by the first reference value to obtain a third reference value (i.e., local volume fraction). This aims to quantify the importance of local regions and distinguish key targets from background noise. When setting the threshold, a preset volume fraction threshold (e.g., 0.1) is used to filter out sub-images with a third reference value above this threshold. The grayscale values of sub-image regions below the threshold are reset to zero, generating a 2D fused standard image to be tested, retaining only the high-threshold regions. The area fraction of non-zero pixels in the standard image is calculated (e.g., 85%, which can be set empirically) as a quantitative indicator of recognition accuracy. The advantage of dynamic thresholding (e.g., based on volume fraction) is adaptive adjustment, avoiding the limitations of fixed thresholds for complex scenes. Through multi-level screening (volume threshold + area fraction), the output image is ensured to contain only high-confidence target regions. Supports custom thresholds and window parameters to adapt to different scenario requirements (for example, in dense scenarios, the threshold needs to be lowered to retain more information).
[0132] Furthermore, the two-dimensional fused images to be tested are sorted according to the recognition accuracy, specifically including:
[0133] Based on the recognition accuracy of the two-dimensional fused images to be tested, the two-dimensional fused images to be tested are sorted in descending order.
[0134] Furthermore, the sorted 2D fused images to be tested are extracted in sequence and combined with the corresponding preferred point cloud data to determine the preliminary 3D target detection results, specifically including:
[0135] Set identification tags and determine the number of identification tags;
[0136] Sequentially extract n sorted 2D fusion images to be tested, and use the deep learning model to perform n inspections on the multi-directional 2D images corresponding to the 2D fusion images to be tested, to obtain preliminary inspection 2D bounding boxes and preliminary inspection 3D bounding boxes;
[0137] Traverse the initial inspection identification label corresponding to the initial inspection two-dimensional bounding box at the nth inspection, and determine the initial inspection identification label Label (n) at the nth inspection and the initial inspection identification label Label (n-1) at the n-1th inspection;
[0138] If the number of similarities and differences between Label (n) and Label (n-1) is greater than or equal to the number of identification labels, the corresponding two-dimensional fusion image to be tested during the nth detection is removed;
[0139] Continue to detect the remaining m two-dimensional fusion images to be tested until the detection of m two-dimensional fusion images to be tested is completed, and the number of similarities and differences between the initial detection identification label Label (m) at the m-th detection and the initial detection identification label Label (m-1) at the m-1-th detection is equal to 0, and the current result is used as the preliminary three-dimensional target detection result.
[0140] Specifically, predefine category labels (such as "vehicle", "pedestrian", and "bicycle") based on the target type. Determine the maximum allowable label change threshold (such as allowing the difference in label type between two detections to be ≤1). Extract the sorted n two-dimensional fusion images to be tested in sequence, and use the deep learning model to perform n detections on the multi-directional two-dimensional images corresponding to the two-dimensional fusion images to be tested, and obtain the initial two-dimensional bounding box and the initial three-dimensional bounding box. Traverse the initial inspection identification label corresponding to the initial two-dimensional bounding box at the nth detection, and determine the initial inspection identification label Label (n) at the nth detection and the initial inspection identification label Label (n-1) at the n-1th detection. For example, use a deep learning model (such as Faster R-CNN) to detect the first two-dimensional fusion image to be tested after sorting, output the initial two-dimensional bounding box (position, size) and the initial three-dimensional bounding box (spatial coordinates), and record the label set Label (1). Detect the subsequent n images in sequence, and record the label set Label (n) each time.
[0141] It can be understood that the difference between two consecutive detection labels (such as the number of labels that have been added, disappeared, or changed in category) is calculated. If the difference is ≥ a set threshold (for example, the intersection of Label(n) and Label(n-1) < the total number of labels minus the threshold), the current image detection result is considered unstable and is discarded. The iteration terminates when the label sets of two consecutive detections are completely consistent (Label(m) = Label(m-1)). The final stable detection result (2D bounding box, 3D bounding box, label) is used as the preliminary 3D object detection result. The advantage is that the label consistency constraint ensures that the output results conform to physical laws (for example, vehicles will not be mistakenly detected as pedestrians). Low-reliability images are dynamically eliminated, reducing the computational load of subsequent processing.
[0142] Furthermore, based on the preliminary 3D target detection results, 3D target geometric information is set, specifically including:
[0143] Based on the preliminary 3D target detection results, the target type is determined, and the target geometry information is set based on the target type, including the length, width, and height of the detected target.
[0144] Furthermore, based on the 3D target geometric information, the 2D fused image to be detected, and the corresponding preferred point cloud data, a standard 3D target detection result is determined, specifically including:
[0145] Mapping the 3D target geometric information to the 2D fused image to be measured, and obtaining the 3D target geometric volume value in the 2D fused image to be measured;
[0146] Based on the two-dimensional fused image to be measured and the corresponding preferred point cloud data, obtaining the corresponding two-dimensional bounding box and three-dimensional bounding box;
[0147] The two-dimensional bounding box is used to intercept the two-dimensional fusion image to be tested, and the two-dimensional fusion standard sub-image to be tested is obtained again;
[0148] Based on the two-dimensional fusion standard sub-image to be tested, obtaining the recognition accuracy of the two-dimensional fusion image to be tested;
[0149] Obtaining the volume value of the 3D bounding box in the 2D fused standard sub-image to be tested;
[0150] If the difference between the volume value of the three-dimensional bounding box in the two-dimensional fusion standard sub-image to be tested and the geometric volume value of the three-dimensional target in the two-dimensional fusion image to be tested is the smallest, and the recognition accuracy of the two-dimensional fusion image to be tested is the largest, then the initial inspection recognition label of the initial inspection two-dimensional bounding box corresponding to the two-dimensional fusion image to be tested is obtained as the standard three-dimensional target detection result.
[0151] Specifically, the length, width, and height parameters of the 3D target (obtained through point cloud clustering, prior models, or the real world) are mapped to the 2D image plane to generate the projected coordinates of the 3D bounding box. Based on the projected 2D bounding box dimensions (e.g., length × width) and a preset height value (e.g., the average height measured by a LiDAR), the virtual volume of the 3D target in the 2D image is calculated. 2D bounding box extraction involves obtaining the 2D bounding box (position and size) of the detected target from the 2D fused image to be measured. 3D bounding box extraction involves obtaining the corresponding 3D bounding box (spatial coordinate range) from the preferred point cloud data. Standard sub-image cropping involves cropping the 2D fused image to be measured using the 2D bounding box to generate a sub-image containing only the core area of the target. The difference (e.g., absolute error or relative error) between the volume of the 3D bounding box in the standard sub-image and the original 3D target geometric volume is compared. The initial inspection label corresponding to the sub-image with the smallest volume difference and the highest recognition accuracy is selected as the final result.
[0152] It can be understood that the spatial rationality of the detection results is verified by the consistency of the projections of the 2D and 3D bounding boxes. Dual thresholds, volume difference and accuracy, are used to filter out geometrically unreasonable or ambiguous detection results. Even when the target is partially occluded or sensor noise interferes, detection stability can be maintained through sub-image cropping and iterative verification. The advantage lies in simultaneously constraining the geometric size (volume) and semantic label (category) of the target, avoiding single-modality misjudgments (such as misdetecting a truck as a car). For example, in autonomous driving, when the vehicle encounters heavy snow and the camera is blurred: first, a 3D bounding box projection of the target (such as the outline of a truck) is generated based on the lidar point cloud. Then, the core area containing the truck is cropped to eliminate snow noise interference. After performing an accuracy calculation, the detection model detected the truck in this sub-image with an accuracy of 89%.
[0153] Next, volume verification is performed. The difference between the truck volume in the sub-image (2.5 m³) and the global geometric volume (2.6 m³) is 3.8%, meeting the threshold requirement. The final result is output: the final label is "truck," and the system triggers a deceleration and avoidance decision.
[0154] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A three-dimensional target detection and fusion method based on multi-sensor fusion, characterized in that: include: Acquire multi-directional two-dimensional image data through the camera, and simultaneously acquire point cloud data through the lidar; Based on the multi-directional two-dimensional image data and point cloud data, a two-dimensional bounding box and a three-dimensional bounding box are constructed respectively; Matching the two-dimensional bounding box with the three-dimensional bounding box to determine the preferred two-dimensional image and the preferred point cloud data; Fusing the optimal two-dimensional image and the optimal point cloud data to obtain a two-dimensional fused image to be measured; Obtaining the recognition accuracy of the two-dimensional fusion image to be tested; The 2D fused images to be tested are sorted according to the recognition accuracy, and then the sorted 2D fused images to be tested are extracted in sequence. The preliminary 3D target detection results are determined by combining them with the corresponding preferred point cloud data. Based on the preliminary 3D target detection results, the 3D target geometric information is set, and the standard 3D target detection results are determined synchronously based on the 3D target geometric information, the 2D fused image to be tested, and the corresponding preferred point cloud data; The matching of the two-dimensional bounding box and the three-dimensional bounding box to determine the preferred two-dimensional image and the preferred point cloud data specifically includes: Obtain the coordinates of the 3D bounding box, and simultaneously map the 3D bounding box to the corresponding multi-directional 2D image to obtain a selected 2D image; Use the two-dimensional bounding box to slide and intercept the two-dimensional image to be selected to obtain the two-dimensional sub-image to be selected; Traversing the candidate two-dimensional sub-images, calculating the number of three-dimensional bounding boxes in the candidate two-dimensional sub-images, and obtaining a first reference value of the preferred two-dimensional image; summing the first reference values of the preferred two-dimensional images corresponding to the two-dimensional images to be selected to obtain a second reference value of the preferred two-dimensional image; removing the candidate two-dimensional images with the largest and smallest second reference values of the preferred two-dimensional image, and taking the candidate two-dimensional image with the largest second reference value of the preferred two-dimensional image among the retained candidate two-dimensional images as the preferred two-dimensional image; The three-dimensional point cloud cluster corresponding to the three-dimensional bounding box in the preferred two-dimensional image is used as the preferred point cloud data.
2. The three-dimensional target detection and fusion method based on multi-sensor fusion according to claim 1, characterized in that: The constructing of the two-dimensional bounding box specifically includes: Based on the multi-directional two-dimensional image data, traverse the corresponding cameras to determine the camera position information; According to the camera position information, a reference plane is constructed, and the coordinates corresponding to the camera position information are connected in the reference plane to obtain the camera network area information; Based on the camera network area information, the geometric center of gravity of the camera network area is determined as the origin to establish a three-dimensional space coordinate system; Extracting multi-directional two-dimensional images from the multi-directional two-dimensional image data in sequence, and labeling the multi-directional two-dimensional images with orientation labels based on the three-dimensional space coordinate system and the corresponding camera position information; Performing image processing on the multi-directional two-dimensional image to obtain a binary image corresponding to the multi-directional two-dimensional image; Obtaining an area value of a pixel with a grayscale value of 255 in a binary image corresponding to the multi-directional two-dimensional image as a first reference value for selecting a two-dimensional bounding box; Establishing an image coordinate system in the binary image corresponding to the multi-directional two-dimensional image with the largest first reference value selected in the two-dimensional bounding box, wherein the image coordinate system is a plane rectangular coordinate system, and the origin coordinates of the image coordinate system are the center point of the binary image; Traverse the pixel points with a grayscale value of 255 in the binary image and obtain their coordinates (xi, yi) in the image coordinate system; Determine the two pixel points with the largest absolute value of xi when xi is greater than 0 and less than 0, and use the mapped lengths of the two pixel points as the length of the interception window. Determine the two pixel points with the largest absolute value of yi when yi is greater than 0 and less than 0, and use the mapped heights of the two pixel points as the width of the interception window. Using a capture window to capture all multi-directional two-dimensional images in the multi-directional two-dimensional image data to obtain a multi-directional two-dimensional sub-image set; Extract multi-directional 2D sub-images from the multi-directional 2D sub-image set in sequence, and obtain the binary images corresponding to the multi-directional 2D sub-images. Traverse the pixels with a grayscale value of 255 in the binary image, then calculate the area of the white connected regions in the binary image, mark the area with the largest white connected region, and obtain the 2D bounding box reference area. Get the edge points of the reference area of the 2D bounding box, synchronously traverse all edge points of the white connected area, connect the two edge points with the greatest distance, and get the height of the 2D bounding box; Using the height of the 2D bounding box as the dividing line, set a straight line perpendicular to the dividing line to slide through the edge points on both ends of the dividing line in the white connected area. The sliding unit is 1 pixel. Then, the two edge points located farthest from each other on the straight line are connected to obtain the width of the 2D bounding box, completing the construction of the 2D bounding box.
3. The three-dimensional target detection and fusion method based on multi-sensor fusion according to claim 1, characterized in that: The construction of the three-dimensional bounding box specifically includes: Identify and remove isolated points in the point cloud data using statistical methods, then merge the dense point cloud into a point set that retains the structure to obtain standard point cloud data; The minimum cluster size is set. Based on the standard point cloud data, the HDBSCAN clustering algorithm is applied to identify dense areas as clusters and mark low-density points as noise. Each point is assigned a label, and the noise point label is -1, to obtain a 3D point cloud cluster. Determine the number of clusters, calculate the minimum and maximum x-coordinates of all points in each cluster, process the y- and z-coordinates simultaneously, combine them into a bounding box, and form a three-dimensional bounding reference cuboid; Use the 3D bounding reference cuboid as the 3D bounding box.
4. The three-dimensional target detection and fusion method based on multi-sensor fusion according to claim 1, characterized in that: The step of fusing the preferred two-dimensional image and the preferred point cloud data to obtain the two-dimensional fused image to be measured specifically includes: Determine the optimal two-dimensional image and the multi-directional two-dimensional image corresponding to the optimal point cloud data, and obtain the multi-directional two-dimensional image to be measured; Obtain a binary image of the multi-directional two-dimensional image to be measured, and obtain the label value corresponding to the point in the corresponding preferred point cloud data; Synchronously mapping points with the same label value to the binary image of the multi-directional two-dimensional image to be measured, and marking pixels with a grayscale value of 255 corresponding to the points with the same label value in the binary image of the multi-directional two-dimensional image to be measured with the same color; The two-dimensional bounding box corresponding to the preferred two-dimensional image and the three-dimensional bounding box corresponding to the preferred point cloud data are integrated into the marked multi-directional two-dimensional image to be tested by using computer vision to obtain the two-dimensional fused image to be tested.
5. The three-dimensional target detection and fusion method based on multi-sensor fusion according to claim 1, characterized in that: The obtaining of the recognition accuracy of the two-dimensional fused image to be tested specifically includes: Use the two-dimensional bounding box to slide and intercept the two-dimensional fused image to be tested to obtain the two-dimensional fused sub-image to be tested; Obtaining the volume of the three-dimensional bounding box in each two-dimensional fused sub-image to be tested as a first recognition accuracy reference value; Obtaining the volumes of all three-dimensional bounding boxes in the two-dimensional fused image to be tested as a second recognition accuracy reference value; Obtaining a ratio of the first recognition accuracy reference value to its corresponding second recognition accuracy reference value as a third recognition accuracy reference value; Setting a recognition accuracy reference threshold, and changing the grayscale value of the two-dimensional fused sub-image to be tested whose third recognition accuracy reference value is less than or equal to the recognition accuracy reference threshold to 0 in the two-dimensional fused image to be tested, to obtain a two-dimensional fused standard image to be tested; The area occupied by the pixels whose grayscale values are not 0 in the two-dimensional fusion standard image to be tested is calculated, and the area is quantified to obtain the recognition accuracy of the two-dimensional fusion image to be tested.
6. The three-dimensional target detection and fusion method based on multi-sensor fusion according to claim 1, characterized in that: The sorting of the two-dimensional fused images to be tested according to the recognition accuracy specifically includes: Based on the recognition accuracy of the two-dimensional fused images to be tested, the two-dimensional fused images to be tested are sorted in descending order.
7. The three-dimensional target detection and fusion method based on multi-sensor fusion according to claim 1, characterized in that: The step of sequentially extracting the sorted two-dimensional fused images to be tested and combining them with the corresponding preferred point cloud data to determine the preliminary three-dimensional target detection results specifically includes: Set identification tags and determine the number of identification tags; Sequentially extract n sorted 2D fusion images to be tested, and use the deep learning model to perform n inspections on the multi-directional 2D images corresponding to the 2D fusion images to be tested, to obtain preliminary inspection 2D bounding boxes and preliminary inspection 3D bounding boxes; Traverse the initial inspection identification label corresponding to the initial inspection two-dimensional bounding box at the nth inspection, and determine the initial inspection identification label Label (n) at the nth inspection and the initial inspection identification label Label (n-1) at the n-1th inspection; If the number of similarities and differences between Label (n) and Label (n-1) is greater than or equal to the number of identification labels, the corresponding two-dimensional fusion image to be tested during the nth detection is removed; Continue to detect the remaining m two-dimensional fusion images to be tested until the detection of m two-dimensional fusion images to be tested is completed, and the number of similarities and differences between the initial detection identification label Label (m) at the m-th detection and the initial detection identification label Label (m-1) at the m-1-th detection is equal to 0, and the current result is used as the preliminary three-dimensional target detection result.
8. The three-dimensional target detection and fusion method based on multi-sensor fusion according to claim 1, characterized in that: The step of setting the three-dimensional target geometric information based on the preliminary three-dimensional target detection result specifically includes: Based on the preliminary 3D target detection results, the target type is determined, and the target geometry information is set based on the target type, including the length, width, and height of the detected target.
9. The three-dimensional target detection and fusion method based on multi-sensor fusion according to claim 1, characterized in that: Determining a standard 3D target detection result based on the 3D target geometric information, the 2D fused image to be detected, and the corresponding preferred point cloud data specifically includes: Mapping the 3D target geometric information to the 2D fused image to be measured, and obtaining the 3D target geometric volume value in the 2D fused image to be measured; Based on the two-dimensional fused image to be measured and the corresponding preferred point cloud data, obtaining the corresponding two-dimensional bounding box and three-dimensional bounding box; The two-dimensional bounding box is used to intercept the two-dimensional fusion image to be tested, and the two-dimensional fusion standard sub-image to be tested is obtained again; Based on the two-dimensional fusion standard sub-image to be tested, obtaining the recognition accuracy of the two-dimensional fusion image to be tested; Obtaining the volume value of the 3D bounding box in the 2D fused standard sub-image to be tested; If the difference between the volume value of the three-dimensional bounding box in the two-dimensional fusion standard sub-image to be tested and the geometric volume value of the three-dimensional target in the two-dimensional fusion image to be tested is the smallest, and the recognition accuracy of the two-dimensional fusion image to be tested is the largest, then the initial inspection recognition label of the initial inspection two-dimensional bounding box corresponding to the two-dimensional fusion image to be tested is obtained as the standard three-dimensional target detection result.
Citation Information
Patent Citations
Target data fusion vehicle detection method and detection device
CN117115784A
Method and system for 2D and 3D target information fusion
CN117911817A