A time sequence fused three-dimensional target detection method, system, device and medium

By using a temporal fusion method, the encoded features of multiple viewpoint cameras are transformed by viewpoint conversion, feature superposition, and spatial transformation, which solves the problem of inaccurate depth in camera-based 3D object detection and improves the accuracy of 3D target detection.

CN117315645BActive Publication Date: 2026-05-08CHONGQING CHANGAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING CHANGAN TECH CO LTD
Filing Date
2023-09-27
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing camera-based 3D object detection methods suffer from inaccurate depth perception, resulting in a significant performance gap compared to LiDAR-based 3D object detection methods.

Method used

A temporal fusion method is adopted to perform viewpoint transformation, feature superposition, and feature space transformation on the coded features of multiple viewpoint cameras, and then process them through a 3D detection head to improve the accuracy of 3D target detection.

Benefits of technology

By fusing long-time encoded features, the accuracy of depth perception of 3D targets is improved, thus enhancing the accuracy of 3D target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117315645B_ABST
    Figure CN117315645B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a kind of three-dimensional target detection method, system, equipment and medium of timing fusion, method includes: the view angle of multiple encoding features based on perspective camera is converted to the view angle of first encoding feature in the multiple encoding features, obtains corresponding multiple target encoding features;The multiple target encoding features and the first encoding feature are superimposed, and the camera feature of the first encoding feature corresponding time is obtained;By the camera feature of multiple view angle cameras is converted in feature space, corresponding multiple space features are obtained;The multiple space features are input into three-dimensional detection head and are processed, and three-dimensional target detection result is obtained.It aims at improving the depth perception accuracy of three-dimensional target, and then improve the accuracy of three-dimensional target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of three-dimensional target detection technology, specifically to a time-series fusion method, system, device, and medium for three-dimensional target detection. Background Technology

[0002] Autonomous vehicles need to perceive the surrounding three-dimensional scene to drive on the road. Long-time fusion three-dimensional object detection is used to obtain the position and category information of objects in three-dimensional space. It is the foundation of autonomous driving perception system and plays an important guiding role in subsequent path planning, motion prediction and collision avoidance.

[0003] Inspired by the success of long-term temporal fusion-based 3D object detection algorithms using laser point clouds, a novel pseudo-LiDAR-based method for long-term temporal fusion-based 3D object detection using camera images has been proposed. This monocular or stereo vision long-term temporal fusion-based 3D object detection calculates disparity, reprojects the image into 3D space to obtain a pseudo-LiDAR, and then employs a high-precision long-term temporal fusion-based 3D object detection algorithm using laser point clouds. Due to the stability and low cost of visual sensors, camera-based 3D object detection has received widespread attention. However, compared to LiDAR-based 3D object detection methods, a significant performance gap remains because it exposes the well-known problem of inaccurate depth perception. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a temporal fusion method, system, device and medium for three-dimensional target detection, aiming to improve the accuracy of depth perception of three-dimensional targets, thereby improving the accuracy of three-dimensional target detection.

[0005] The first aspect of this invention provides a long-time-series fusion method for three-dimensional target detection, the method comprising:

[0006] The viewpoint of multiple coded features obtained from the viewpoint camera is converted into the viewpoint of the first coded feature among the multiple coded features to obtain the corresponding multiple target coded features;

[0007] The multiple target encoded features and the first encoded feature are superimposed to obtain the camera features at the time corresponding to the first encoded feature;

[0008] By performing feature space transformation on the camera features of multiple viewpoint cameras, the corresponding multi-view spatial features are obtained.

[0009] The multi-view spatial features are input into the 3D detection head for processing to obtain 3D target detection results.

[0010] Optionally, before inputting the multi-view spatial features into the 3D detection head for processing to obtain the 3D target detection result, the method further includes:

[0011] The perspectives of multiple multi-view spatial features are converted into the perspective of the first multi-view spatial feature among the multiple multi-view spatial features to obtain the corresponding multiple target multi-view spatial features.

[0012] The multiple target multi-view spatial features and the first multi-view spatial features are concatenated to obtain the temporal fusion features corresponding to the first multi-view spatial features at the time point.

[0013] The step of inputting the multi-view spatial features into the 3D detection head for processing to obtain 3D target detection results includes:

[0014] The temporal fusion features are input into the 3D detection head for processing to obtain 3D target detection results.

[0015] Optionally, the step of converting the viewpoint of multiple coded features obtained based on the viewpoint camera into the viewpoint of the first coded feature among the multiple coded features to obtain the corresponding multiple target coded features includes:

[0016] Determine the intrinsic parameters of the viewpoint camera and the extrinsic parameters of the viewpoint camera relative to the vehicle body coordinate system;

[0017] Determine a first coded feature obtained based on the viewpoint camera and multiple coded features at multiple times before and after the time corresponding to the first coded feature;

[0018] Based on the intrinsic parameters, the extrinsic parameters, and the first coding feature, the perspective of the multiple coding features is converted to the perspective of the first coding feature through homography transformation to obtain the corresponding multiple target coding features.

[0019] Optionally, the step of superimposing the plurality of target encoded features and the first encoded feature to obtain the camera features at the time corresponding to the first encoded feature includes:

[0020] The multiple target encoded features and the first encoded feature are superimposed to obtain the initial camera features at the time corresponding to the first encoded feature;

[0021] The initial camera features are subjected to dimensionality reduction convolution processing using a target convolution kernel to obtain the camera features at the corresponding time step.

[0022] Optionally, the step of obtaining corresponding multi-view spatial features by performing feature space transformation on the camera features of multiple viewpoint cameras includes:

[0023] Monocular depth estimation results are obtained by performing monocular depth estimation on camera features using a monocular depth estimation model.

[0024] The camera features are used to perform multi-view depth estimation using a multi-view depth estimation model to obtain the multi-view depth estimation results;

[0025] A depth map is obtained by summing the monocular depth estimation results and the multi-view depth estimation results;

[0026] By using columnar pooling, the contextual information in the depth maps of multiple view cameras and the corresponding multi-view depth estimation results is fused and the viewpoint is transformed to obtain the multi-view spatial features at the corresponding time.

[0027] Optionally, before obtaining the depth map by summing the monocular depth estimation results and the multi-view depth estimation results, the method further includes:

[0028] The depth centers from the monocular depth estimation results and the multi-view depth estimation results are input into a weighted network for processing to obtain a weighted map.

[0029] Based on the weight map and the depth center and depth range in the multi-view depth estimation results, the target multi-view depth estimation results are obtained by processing through a stereo depth model.

[0030] The process of obtaining a depth map by summing the monocular depth estimation results and the multi-view depth estimation results includes:

[0031] A depth map is obtained by summing the monocular depth estimation results and the target multi-view depth estimation results.

[0032] Optionally, the step of converting the perspective of multiple multi-view spatial features into the perspective of a first multi-view spatial feature among the multiple multi-view spatial features to obtain corresponding multiple target multi-view spatial features includes:

[0033] Determine the first multi-view spatial feature and multiple multi-view spatial features at multiple times before and after the time corresponding to the first multi-view spatial feature;

[0034] Based on the first multi-view spatial feature, the perspective of the multiple multi-view spatial features is converted into the perspective of the first multi-view spatial feature through homography transformation, thereby obtaining the corresponding multiple target multi-view spatial features.

[0035] Optionally, the step of inputting the temporal fusion features into the 3D detection head for processing to obtain the 3D target detection result includes:

[0036] The temporal fusion features from multiple time points are input into their respective 3D detection heads for processing to obtain multiple 3D target detection results.

[0037] Optionally, the step of inputting the multi-view spatial features into the 3D detection head for processing to obtain the 3D target detection result includes:

[0038] The multi-view spatial features at multiple time points are input into their respective 3D detection heads for processing to obtain multiple 3D target detection results.

[0039] Optionally, the training of the monocular depth estimation model includes:

[0040] By performing deep annotation on a preset number of sample camera features, a model training sample set and a model validation sample set are constructed.

[0041] The initial monocular depth estimation model is trained by inputting the camera features of each target sample in the model training sample set to obtain the corresponding depth prediction results.

[0042] Based on the depth prediction results, the corresponding depth annotations, and the preset loss function, determine whether the initial monocular depth estimation model has been successfully trained.

[0043] The initial monocular depth estimation model that has passed training is validated using the validation sample set.

[0044] If the verification is successful, the initial monocular depth estimation model will be determined as the monocular depth estimation model.

[0045] A second aspect of the present invention provides a long-time-series fusion-based three-dimensional target detection system, the system comprising:

[0046] The viewpoint conversion module is used to convert the viewpoint of multiple coded features obtained based on the viewpoint camera into the viewpoint of the first coded feature among the multiple coded features, thereby obtaining the corresponding multiple target coded features.

[0047] The feature overlay module is used to overlay the plurality of target encoded features and the first encoded feature to obtain the camera features at the time corresponding to the first encoded feature;

[0048] The feature space transformation module is used to obtain the corresponding multi-view spatial features by performing feature space transformation on the camera features of multiple view cameras;

[0049] The target detection module is used to input the multi-view spatial features into the three-dimensional detection head for processing to obtain three-dimensional target detection results.

[0050] A third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, it implements a long-time fusion three-dimensional target detection method as described in the first aspect of the present invention.

[0051] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a long-time-series fusion three-dimensional target detection method as described in the first aspect of the present invention.

[0052] This invention provides a long-time-series fusion method for 3D target detection. First, the viewpoints of multiple coded features obtained from multiple viewpoint cameras are converted into the viewpoint of a first coded feature among the multiple coded features, resulting in corresponding multiple target coded features. Then, the multiple target coded features and the first coded feature are superimposed to obtain the camera features corresponding to the first coded feature at that time. Next, feature space transformation is performed on the camera features from the multiple viewpoint cameras to obtain corresponding multi-view spatial features. Finally, the multi-view spatial features are input into a 3D detection head for processing to obtain the 3D target detection result. Therefore, the process of obtaining corresponding multi-view spatial features by performing feature space transformation on camera features from multiple viewpoint cameras involves depth estimation of 3D targets. This process uses feature space transformation of camera features obtained by superimposing encoded features from multiple time points, which considers long-term encoded features. The long-term encoded features are then fused to estimate the depth of 3D targets. Since the overlapping area of ​​the view frustum of multi-view cameras in autonomous driving is small, most of the stereo spatial information comes from the temporal information brought about by motion. Therefore, this invention can effectively improve the accuracy of depth perception of 3D targets by fusing long-term encoded features to estimate the depth of 3D targets, thereby improving the accuracy of the final 3D target detection. At the same time, the multi-view spatial features used for 3D target detection are multi-view camera features that fuse camera features from multiple viewpoint cameras. Therefore, by considering the relevant features of multiple viewpoint cameras, the accuracy of the final 3D target detection can be further improved. Attached Figure Description

[0053] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 The flowchart illustrates a long-time-series fusion method for three-dimensional target detection according to an embodiment of the present invention.

[0055] Figure 2 This is a network architecture diagram of the depth estimation module in a long-time-series fusion three-dimensional target detection method according to an embodiment of the present invention;

[0056] Figure 3 This is a schematic diagram of a long-time-series fusion three-dimensional target detection system according to an embodiment of the present invention. Detailed Implementation

[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] refer to Figure 1 , Figure 1 This is a flowchart illustrating a long-time-series fusion-based three-dimensional target detection method according to an embodiment of the present invention. The long-time-series fusion-based three-dimensional target detection method provided by the present invention, such as... Figure 1 As shown, the method includes:

[0059] Step S1: Convert the viewpoint of multiple coded features obtained from the viewpoint camera into the viewpoint of the first coded feature among the multiple coded features to obtain the corresponding multiple target coded features.

[0060] In this embodiment, multiple cameras mounted on the vehicle body acquire image data at multiple times, resulting in multiple time frames of image data for each camera. For all acquired image data, a 2D convolutional backbone network is used to encode image features for each frame, resulting in a corresponding encoded feature for each frame. The 2D convolutional backbone network performs shared-weight image feature encoding on the input image data for each frame. This 2D convolutional backbone network includes, but is not limited to, commonly used 2D convolutional network structures such as ResNet, EfficientNet, SwinTransformer, VoVNetV2, and ConvNeXt with defined depth levels, as well as a convolutional network architecture composed of a CBNet combined backbone network architecture and a feature pyramid (FPN) multi-scale fusion.

[0061] In this embodiment, after obtaining multiple coded features based on image data from multiple time frames acquired by a camera with the same viewpoint, the viewpoint of the coded features changes over time because the coded features of different time frames are obtained from image data of different time frames. In order to perform long-term fusion of the coded features and improve the accuracy of depth perception of subsequent 3D targets, this invention first performs viewpoint conversion on each coded feature to be fused to unify the coded features to the same viewpoint before performing long-term fusion. This viewpoint unification before long-term coded feature fusion allows the fused camera features to better cover the features in the image data of all time periods.

[0062] Specifically: After obtaining coded features from image data acquired by a camera from the same viewpoint, this coded feature is designated as the first coded feature. Then, a preset number of coded features before the corresponding time point and a preset number of coded features after the corresponding time point are acquired. Using the viewpoint of the first coded feature as a reference, the viewpoint of each acquired coded feature is converted to be the same as that of the first coded feature, thus obtaining multiple corresponding target coded features. The preset number can be set according to the actual application scenario and is not specifically limited here; for example, the preset number could be 5, 10, or 20, etc.

[0063] For example, the viewpoint camera 1 acquires 100 image data points at 100 time points. By encoding these 100 image data points, 100 corresponding encoded features are obtained, denoted as f1, f2, f3, ..., f100. Here, we take the encoded feature f20 at the 20th time point as an example to illustrate this. With a preset number of 5, we obtain the 5 encoded features f19, f18, f17, f16, and f15 corresponding to the time point before this first encoded feature, and obtain the first encoded feature... The five encoded features f21, f22, f23, f24, and f25 corresponding to the time step are then used as the reference for the first encoded feature. The perspectives of the ten encoded features f19, f18, f17, f16, f15, f21, f22, f23, f24, and f25 are then converted to the perspective of the first encoded feature to obtain the corresponding ten target encoded features ft19, ft18, ft17, ft16, ft15, ft21, ft22, ft23, ft24, and ft25.

[0064] Step S2: Overlay the multiple target coding features and the first coding feature to obtain the camera features at the time corresponding to the first coding feature.

[0065] In this embodiment, after obtaining multiple target coded features after viewpoint transformation, the viewpoints of these multiple target coded features are the same as those of the first coded feature. At this point, the multiple target coded features and the first coded feature are superimposed to obtain the camera feature at the time corresponding to the first coded feature. For example, continuing with the above example, the obtained 10 target coded features ft19, ft18, ft17, ft16, ft15, ft21, ft22, ft23, ft24, ft25 and the first coded feature f20 at the 20th time are superimposed to obtain the corresponding camera feature F20 at the 20th time.

[0066] Step S3: Obtain the corresponding multi-view spatial features by performing feature space transformation on the camera features of multiple view cameras.

[0067] In this embodiment, based on the same implementation method as described above, camera features at a specific moment can be calculated for each viewpoint camera. After obtaining the camera features of multiple viewpoint cameras mounted on the vehicle body at the same moment, depth estimation is performed on each camera feature at the same moment to obtain a depth map corresponding to each camera feature. Based on each camera feature and the corresponding depth map, feature space transformation is performed to convert it into a multi-view spatial feature in three-dimensional space. For example, the vehicle body is equipped with 6 viewpoint cameras. For each viewpoint camera, 6 camera features aF20, bF20, cF20, dF20, eF20, and fF20 corresponding to each viewpoint camera at the same moment can be calculated using the above implementation method. By performing feature space transformation on these 6 camera features, a multi-view spatial feature AF20 in three-dimensional space is obtained.

[0068] Step S4: Input the multi-view spatial features into the three-dimensional detection head for processing to obtain the three-dimensional target detection result.

[0069] In this embodiment, the multi-view spatial features obtained in step S3 are input into a 3D detection head for 3D target detection to obtain the corresponding 3D target detection results. The 3D target detection head DtHead includes, but is not limited to, anchor-point based PointPillars, centerpoint based heatmap prediction, and Transformer-based DETR3D and ransfusion detection heads.

[0070] This invention provides a long-time-series fusion method for 3D target detection. First, the viewpoints of multiple coded features obtained from multiple viewpoint cameras are converted into the viewpoint of a first coded feature among the multiple coded features, resulting in corresponding multiple target coded features. Then, the multiple target coded features and the first coded feature are superimposed to obtain the camera features corresponding to the first coded feature at that time. Next, feature space transformation is performed on the camera features from the multiple viewpoint cameras to obtain corresponding multi-view spatial features. Finally, the multi-view spatial features are input into a 3D detection head for processing to obtain the 3D target detection result. Therefore, the process of obtaining corresponding multi-view spatial features by performing feature space transformation on camera features from multiple viewpoint cameras involves depth estimation of 3D targets. This process uses feature space transformation of camera features obtained by superimposing encoded features from multiple time points, which considers long-term encoded features. The long-term encoded features are then fused to estimate the depth of 3D targets. Since the overlapping area of ​​the view frustum of multi-view cameras in autonomous driving is small, most of the stereo spatial information comes from the temporal information brought about by motion. Therefore, this invention can effectively improve the accuracy of depth perception of 3D targets by fusing long-term encoded features to estimate the depth of 3D targets, thereby improving the accuracy of the final 3D target detection. At the same time, the multi-view spatial features used for 3D target detection are multi-view camera features that fuse camera features from multiple viewpoint cameras. Therefore, by considering the relevant features of multiple viewpoint cameras, the accuracy of the final 3D target detection can be further improved.

[0071] In conjunction with the above embodiments, in one implementation, the present invention also provides a long-time-series fusion method for three-dimensional target detection. In this long-time-series fusion method for three-dimensional target detection, before step S4, the method further includes steps S401 to S402.

[0072] Step S401: Convert the perspective of multiple multi-view spatial features into the perspective of the first multi-view spatial feature among the multiple multi-view spatial features to obtain the corresponding multiple target multi-view spatial features.

[0073] In this embodiment, after obtaining the multi-view spatial features, these features are designated as the first multi-view spatial features. Then, a first preset number of multi-view spatial features before the corresponding time point and a first preset number of multi-view spatial features after the corresponding time point are acquired. Using the viewpoint of the first multi-view spatial feature as a reference, the viewpoints of each acquired multi-view spatial feature are converted to be the same as the viewpoint of the first multi-view spatial feature, thus obtaining multiple corresponding target multi-view spatial features. The first preset number can be set according to the actual application scenario and is not specifically limited here; for example, the preset number could be 5, 10, or 20, etc.

[0074] Step S402: Perform feature concatenation between the multiple target multi-view spatial features and the first multi-view spatial feature to obtain the temporal fusion feature corresponding to the first multi-view spatial feature at the given time.

[0075] In this embodiment, after obtaining multiple target multi-view spatial features after viewpoint transformation, the viewpoints of these multiple target multi-view spatial features are the same as those of the first multi-view spatial feature. At this point, the multiple target multi-view spatial features and the first multi-view spatial feature are concatenated to obtain the temporal fusion feature at the corresponding moment of the first multi-view spatial feature. This invention fuses multi-view spatial features from multiple moments to better utilize the properties of future frames, allowing the 3D detection head model to learn data information from future frames, enriching the model's representation, and thus making the final 3D target detection result more accurate. The specific meaning of the aforementioned future frame is: multi-view spatial features before a specific moment are past frame feature data for that specific moment, while multi-view spatial features after a specific moment are future frame feature data for that specific moment. Both past frame feature data before a specific moment and future frame feature data after a specific moment are feature data of events that have already occurred. For example, if multi-view spatial features have been obtained for 20 time points, then for the multi-view spatial features of the 5th time point among these 20 time points, the multi-view spatial features of the previous 4 time points are the feature data of the past frames of the multi-view spatial features of the 5th time point, and the multi-view spatial features of the last 10 time points are the feature data of the future frames of the multi-view spatial features of the 5th time point.

[0076] In this embodiment, step S4 includes: inputting the temporal fusion features into the three-dimensional detection head for processing to obtain the three-dimensional target detection result.

[0077] In this embodiment, the temporal fusion features obtained in step S402 are input into the three-dimensional detection head to perform three-dimensional target detection, and the corresponding three-dimensional target detection results are obtained.

[0078] In conjunction with the above embodiments, in one implementation, the present invention also provides a long-time-series fusion method for three-dimensional target detection. In this long-time-series fusion method for three-dimensional target detection, step S1 includes steps S11 to S13:

[0079] Step S11: Determine the intrinsic parameters of the view camera and the extrinsic parameters of the view camera relative to the vehicle body coordinate system.

[0080] In this embodiment, the encoded features are obtained by encoding image data acquired by a viewpoint camera. The acquired image data needs to be transformed based on intrinsic and extrinsic parameters to obtain the spatial position of the target in the image data relative to the vehicle. Therefore, when performing viewpoint transformation on the encoded features obtained from the image data, the camera's intrinsic parameters and the extrinsic parameters of the viewpoint camera relative to the vehicle's body coordinate system are required. Therefore, before performing viewpoint transformation on the encoded features, the camera's intrinsic parameters corresponding to the encoded features are determined, and the extrinsic parameters of the viewpoint camera relative to the vehicle's body coordinate system are also determined.

[0081] Step S12: Determine the first encoded feature obtained based on the viewpoint camera and multiple encoded features at multiple times before and after the time corresponding to the first encoded feature.

[0082] In this embodiment, a coded feature obtained by the viewpoint camera is determined as a first coded feature. Then, a preset number of coded features before the corresponding time point of the first coded feature are obtained, as well as a preset number of coded features after the corresponding time point of the first coded feature, thereby obtaining multiple coded features. All of these multiple coded features and the first coded feature are obtained by encoding image data obtained from the same viewpoint camera.

[0083] Step S13: Based on the intrinsic parameters, the extrinsic parameters, and the first coding feature, the perspective of the multiple coding features is converted to the perspective of the first coding feature through homography transformation to obtain the corresponding multiple target coding features.

[0084] In this embodiment, based on the memory of the view camera and the extrinsic parameters of the view camera relative to the vehicle body coordinate system determined in step S11, and based on the multiple encoded features obtained in step S12, the view of the multiple encoded features is converted to the same view as the first encoded feature through homography transformation, thereby obtaining multiple target encoded features corresponding to the multiple encoded features.

[0085] In conjunction with the above embodiments, in one implementation, the present invention also provides a long-time-series fusion method for three-dimensional target detection. In this long-time-series fusion method for three-dimensional target detection, step S2 includes steps S21 to S22:

[0086] Step S21: Overlay the multiple target coding features and the first coding feature to obtain the initial camera features at the time corresponding to the first coding feature.

[0087] In this embodiment, although step S2 improves depth perception accuracy by performing long-term temporal fusion of features through feature overlay, the overlay increases the amount of data processing. To reduce the amount of data processing while ensuring long-term temporal fusion of features, this invention proposes another implementation method for step S2, which includes steps S21 and S22. Specifically: First, the multiple target encoded features obtained in step S1 and the corresponding first encoded features are overlaid to obtain the initial camera features at the time corresponding to the first encoded features, and then step S22 is executed.

[0088] Step S22: Perform dimensionality reduction convolution processing on the initial camera features using the target convolution kernel to obtain the camera features at the corresponding time.

[0089] In this embodiment, the initial camera features obtained in step S21 are subjected to dimensionality reduction convolution processing using a 1×1 target convolution kernel to obtain the camera features at the time corresponding to the first encoded feature.

[0090] In conjunction with the above embodiments, in one implementation, the present invention also provides a long-time-series fusion method for three-dimensional target detection. In this long-time-series fusion method for three-dimensional target detection, step S3 includes steps S31 to S34:

[0091] Step S31: Perform monocular depth estimation on camera features using a monocular depth estimation model to obtain monocular depth estimation results.

[0092] In this embodiment, the camera features obtained in step S2 are input into the monocular depth estimation model for monocular depth estimation to obtain the monocular depth estimation results for each feature point in the camera features. The monocular depth estimation result for any feature point is a specific monocular depth estimate value, such as a monocular depth of 2 for a feature point.

[0093] Step S32: Perform multi-view depth estimation on the camera features using a multi-view depth estimation model to obtain the multi-view depth estimation result.

[0094] In this embodiment, the camera features obtained in step S2 are simultaneously input into a multi-view depth estimation model for multi-view depth estimation to obtain the multi-view depth estimation results for each feature point in the camera features. The multi-view depth estimation result for any feature point includes a depth center μ, a depth range σ, and contextual information C. The depth center refers to the depth estimate with the highest probability for that feature point, and the depth range is the set probability range of the depth estimate. By calculating the multi-view depth estimation results, a specific numerical multi-view depth estimate is obtained for each feature point. For example, the multi-view depth estimation result for a feature point includes a normal distribution with a depth center of 2 and a depth range of [0.7, 1], and the contextual information for that feature point. This indicates that the probability of predicting the depth of that feature point to be 2 is the highest. Simultaneously, all depth results within the normal distribution with probabilities between 0.7 and 1 are used to characterize the depth of that feature point. For example, the average of the depth results within the range of [0.7, 1] of the normal distribution is calculated, and the calculated average is determined as the multi-view depth estimate for that feature point. For instance, the multi-view depth estimate for that feature point is 1.6.

[0095] Step S33: Obtain a depth map by summing the monocular depth estimation results and the multi-view depth estimation results.

[0096] In this embodiment, after obtaining the monocular depth estimation results and multi-view depth estimation results of the camera features, the depth map corresponding to the camera features is obtained by summing the monocular depth estimation value of each feature point in the monocular depth estimation result with the multi-view depth estimation value of the corresponding feature point in the multi-view depth estimation result. Both the monocular and multi-view depth estimation results are obtained based on the same camera feature, and there is a one-to-one correspondence between the feature points in the monocular and multi-view depth estimation results. During the summation process, only the two corresponding feature points are summed.

[0097] Step S34: By using cylinder pooling, the depth maps of multiple view cameras and the contextual information in the corresponding multi-view depth estimation results are fused and the view is transformed to obtain the multi-view spatial features at the corresponding time.

[0098] In this embodiment, based on the same implementation method described above, the camera features of each view camera at the same time can be used to calculate a corresponding depth map. The depth maps of multiple view cameras at the same time and the contextual information in the multi-view depth estimation results at that same time are obtained. Then, cylinder pooling is used to fuse the obtained depth maps and contextual information and transform them into features under a unified BEV view centered on the origin of the vehicle coordinate system, thereby obtaining the multi-view spatial features at that same time.

[0099] In conjunction with the above embodiments, in one implementation, the present invention also provides a long-time-series fusion method for three-dimensional target detection. In this long-time-series fusion method for three-dimensional target detection, before step S33, the method further includes steps S3301 to S3302:

[0100] Step S3301: Input the depth centers from the monocular depth estimation results and the multi-view depth estimation results into a weighted network for processing to obtain a weighted map.

[0101] In this embodiment, the present invention determines the final depth estimation result of feature points based on the monocular depth estimation result and the multi-view depth estimation result of feature points. However, in actual calculation, it was found that the reliability of the multi-view depth estimation results of different feature points is different. Therefore, the present invention determines the weight of the multi-view depth estimation result of each feature point in determining its own final depth estimation result through a weight network, so as to obtain a more accurate depth estimation result based on the monocular depth estimation result and the multi-view depth estimation result of feature points. Therefore, for step S3 above, another implementation method for step S3 is proposed, which includes steps S321 to S322.

[0102] Specifically: such as Figure 3 As shown, the depth center μ from the monocular depth estimation results and multi-view depth estimation results obtained in steps S31 and S32, respectively, is input into the weight network for calculation to obtain a weight map. This weight map has a corresponding weight value for each feature point.

[0103] Step S3302: Based on the weight map and the depth center and depth range in the multi-view depth estimation result, process the data using a stereo depth model to obtain the target multi-view depth estimation result.

[0104] In this embodiment, the target multi-view depth estimation result is obtained by inputting the weight map obtained in step S3301 and the depth center μ and depth range σ in the multi-view depth estimation result obtained in step S332 into the stereo depth model for calculation. The target multi-view depth estimation result records the multi-view depth estimation values ​​of each feature point after weighting.

[0105] In this embodiment, step S33 includes: obtaining a depth map by summing the monocular depth estimation result and the target multi-view depth estimation result.

[0106] In this embodiment, after obtaining the monocular depth estimation results and multi-view depth estimation results of the camera features, the depth map corresponding to the camera features is obtained by summing the monocular depth estimation value of each feature point in the monocular depth estimation result with the multi-view depth estimation value of the corresponding feature point in the target multi-view depth estimation result. Both the monocular depth estimation result and the target multi-view depth estimation result are obtained based on the same camera features, and there is a one-to-one correspondence between the feature points in the monocular depth estimation result and the feature points in the target multi-view depth estimation result. During the summation process, only the two corresponding feature points are summed.

[0107] In conjunction with the above embodiments, in one implementation, the present invention also provides a long-time-series fusion method for three-dimensional target detection. In this long-time-series fusion method for three-dimensional target detection, step S401 includes steps S4011 to S4012:

[0108] Step S4011: Determine the first multi-view spatial feature and multiple multi-view spatial features at multiple times before and after the time corresponding to the first multi-view spatial feature.

[0109] In this embodiment, a multi-view spatial feature is determined as a first multi-view spatial feature. Then, a first preset number of multi-view spatial features before the time corresponding to the first multi-view spatial feature are obtained, and a first preset number of multi-view spatial features after the time corresponding to the first multi-view spatial feature are obtained, thereby obtaining multiple multi-view spatial features.

[0110] Step S4012: Based on the first multi-view spatial feature, convert the perspective of the multiple multi-view spatial features into the perspective of the first multi-view spatial feature through homography transformation to obtain the corresponding multiple target multi-view spatial features.

[0111] In this embodiment, based on the multiple multi-view spatial features obtained in step S4011, the viewpoints of the multiple multi-view spatial features are converted to be the same as the viewpoints of the first multi-view spatial feature through homography transformation, thereby obtaining multiple target multi-view spatial features corresponding to the multiple multi-view spatial features.

[0112] In conjunction with the above embodiments, in one implementation, this invention also provides a long-time-series fusion method for 3D target detection. In this long-time-series fusion method, the step of inputting the temporal fusion features into a 3D detection head for processing to obtain 3D target detection results includes: inputting the temporal fusion features from multiple time points into their respective corresponding 3D detection heads for processing to obtain multiple corresponding 3D target detection results.

[0113] In this embodiment, another implementation method for inputting the temporal fusion features into a 3D detection head for processing to obtain 3D target detection results is as follows: Multiple 3D detection heads are set up, and the temporal fusion features at multiple time points are respectively input into the corresponding multiple 3D detection heads for 3D target detection, obtaining multiple corresponding 3D target detection results. Based on the obtained multiple 3D target detection results, the final 3D target detection result is obtained to further improve the accuracy of 3D target detection. Specifically, the temporal fusion features at one time point are input into one corresponding 3D target detection head.

[0114] In conjunction with the above embodiments, in one implementation, this invention also provides a long-time-series fusion method for 3D target detection. In this long-time-series fusion method, the step of inputting the multi-view spatial features into a 3D detection head for processing to obtain 3D target detection results includes: inputting multi-view spatial features from multiple time points into their respective corresponding 3D detection heads for processing to obtain multiple corresponding 3D target detection results.

[0115] In this embodiment, another method for processing the multi-view spatial features into a 3D detection head to obtain 3D target detection results is as follows: Multiple 3D detection heads are set up, and the multi-view spatial features at multiple time points are respectively input into the corresponding multiple 3D detection heads for 3D target detection, obtaining multiple corresponding 3D target detection results. Based on the obtained multiple 3D target detection results, the final 3D target detection result is obtained to further improve the accuracy of 3D target detection. Specifically, the multi-view spatial features at one time point are input into one corresponding 3D target detection head.

[0116] In conjunction with the above embodiments, in one implementation, this invention also provides a long-time-series fusion method for 3D target detection. In this long-time-series fusion method for 3D target detection, the training of the monocular depth estimation model includes steps S01 to S05:

[0117] Step S01: Construct a model training sample set and a model validation sample set by performing depth annotation on a preset number of sample camera features.

[0118] In this embodiment, a preset number of sample camera features are first obtained, and depth annotation is performed on each feature point in the sample camera features. The preset number of target sample camera features after depth annotation are then divided into a model training sample set and a model validation sample set according to a set ratio. The preset number and the set ratio can be set according to the actual application scenario and are not specifically limited here.

[0119] Step S02: Train the initial monocular depth estimation model constructed by inputting the camera features of each target sample in the model training sample set to obtain the corresponding depth prediction results.

[0120] In this embodiment, the initial monocular depth estimation model is constructed by inputting the camera features of each target sample in the model training sample set to obtain the corresponding depth prediction results.

[0121] Step S03: Based on the depth prediction results, the corresponding depth annotations, and the preset loss function, determine whether the initial monocular depth estimation model has been successfully trained.

[0122] In this embodiment, based on the depth prediction results and the depth annotation of the target sample camera features in the input initial monocular depth estimation model and the preset loss function, it is determined whether the initial monocular depth estimation model has been successfully trained. After determining that the initial monocular depth estimation model has been successfully trained using the loss function, step S04 is executed.

[0123] Step S04: Validate the trained initial monocular depth estimation model using the validation sample set.

[0124] In this embodiment, the depth estimation accuracy of the trained initial monocular depth estimation model is verified by verifying the target sample camera features in the verification sample set.

[0125] Step S05: If the verification is successful, the initial monocular depth estimation model is determined as the monocular depth estimation model.

[0126] In this embodiment, if the depth estimation accuracy of the initial monocular depth estimation model that has been trained is verified, the initial monocular depth estimation model is determined as the final monocular depth estimation model to be applied.

[0127] A second aspect of the present invention provides a long-time-series fusion three-dimensional target detection system 300, the system 300 comprising:

[0128] The view conversion module 301 is used to convert the view of multiple coded features obtained based on the view camera into the view of the first coded feature among the multiple coded features, so as to obtain the corresponding multiple target coded features.

[0129] The feature overlay module 302 is used to overlay the plurality of target encoded features and the first encoded feature to obtain the camera features at the time corresponding to the first encoded feature;

[0130] The feature space transformation module 303 is used to obtain the corresponding multi-view spatial features by performing feature space transformation on the camera features of multiple view cameras;

[0131] The target detection module 304 is used to input the multi-view spatial features into the three-dimensional detection head for processing to obtain three-dimensional target detection results.

[0132] Optionally, the system 300 further includes:

[0133] The target multi-view spatial feature determination module is used to convert the perspective of multiple multi-view spatial features into the perspective of the first multi-view spatial feature among the multiple multi-view spatial features, so as to obtain the corresponding multiple target multi-view spatial features.

[0134] The feature stitching module is used to stitch together the multiple target multi-view spatial features and the first multi-view spatial feature to obtain the temporal fusion feature corresponding to the first multi-view spatial feature at the time point.

[0135] The target detection module 304 includes:

[0136] The first target detection module is used to input the temporal fusion features into the three-dimensional detection head for processing to obtain the three-dimensional target detection result.

[0137] Optionally, the viewpoint conversion module 301 includes:

[0138] The intrinsic and extrinsic parameter determination module is used to determine the intrinsic parameters of the viewpoint camera and the extrinsic parameters of the viewpoint camera relative to the vehicle body coordinate system.

[0139] A viewpoint conversion coding feature determination module is used to determine a first coding feature obtained based on the viewpoint camera and multiple coding features at multiple times before and after the time corresponding to the first coding feature;

[0140] The perspective conversion submodule is used to convert the perspective of the multiple coding features into the perspective of the first coding feature through homography transformation based on the intrinsic parameters, the extrinsic parameters and the first coding feature, so as to obtain the corresponding multiple target coding features.

[0141] Optionally, the feature overlay module 302 includes:

[0142] The first feature overlay module is used to overlay the plurality of target encoded features and the first encoded feature to obtain the initial camera features at the time corresponding to the first encoded feature;

[0143] The convolution processing module is used to perform dimensionality reduction convolution processing on the initial camera features using a target convolution kernel to obtain the camera features at the corresponding time.

[0144] Optionally, the feature space transformation module 303 includes:

[0145] The monocular depth estimation module is used to perform monocular depth estimation on camera features through a monocular depth estimation model to obtain monocular depth estimation results.

[0146] The multi-view depth estimation module is used to perform multi-view depth estimation on the camera features through a multi-view depth estimation model to obtain multi-view depth estimation results.

[0147] The summation module is used to obtain a depth map by summing the monocular depth estimation results and the multi-view depth estimation results;

[0148] The feature space transformation submodule is used to fuse and transform the contextual information in the depth maps of multiple view cameras and the corresponding multi-view depth estimation results through cylindrical pooling, so as to obtain the multi-view spatial features at the corresponding time.

[0149] Optionally, the system 300 further includes:

[0150] The weight map determination module is used to input the depth centers in the monocular depth estimation results and the multi-view depth estimation results into the weight network for processing to obtain a weight map.

[0151] The target multi-view depth estimation module is used to process the depth center and depth range in the weight map and the multi-view depth estimation result through a stereo depth model to obtain the target multi-view depth estimation result.

[0152] The summation module includes:

[0153] The summation submodule is used to obtain a depth map by summing the monocular depth estimation result and the target multi-view depth estimation result.

[0154] Optionally, the target multi-view spatial feature determination module includes:

[0155] The viewpoint switching multi-view spatial feature determination module is used to determine the first multi-view spatial feature and multiple multi-view spatial features at multiple times before and after the time corresponding to the first multi-view spatial feature.

[0156] The first perspective conversion module is used to convert the perspective of the multiple multi-view spatial features into the perspective of the first multi-view spatial features through homography transformation based on the first multi-view spatial features, thereby obtaining the corresponding multiple target multi-view spatial features.

[0157] Optionally, the first target detection module includes:

[0158] The first target detection submodule is used to input the temporal fusion features from multiple time points into their respective 3D detection heads for processing, thereby obtaining multiple 3D target detection results.

[0159] Optionally, the target detection module 304 includes:

[0160] The second target detection module is used to input the multi-view spatial features at multiple time points into their respective 3D detection heads for processing, and obtain the corresponding multiple 3D target detection results.

[0161] Optionally, the training of the monocular depth estimation model in the monocular depth estimation module includes the following steps: constructing a model training sample set and a model validation sample set by depth annotation of a preset number of sample camera features; inputting the target sample camera features from the model training sample set into the constructed initial monocular depth estimation model for training to obtain the corresponding depth prediction results; determining whether the initial monocular depth estimation model is successfully trained based on the depth prediction results, the corresponding depth annotations, and a preset loss function; validating the successfully trained initial monocular depth estimation model using the validation sample set; and determining the initial monocular depth estimation model as the monocular depth estimation model if the verification is successful.

[0162] A third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, it implements a long-time fusion three-dimensional target detection method as described in the first aspect of the present invention.

[0163] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a long-time-series fusion three-dimensional target detection method as described in the first aspect of the present invention.

[0164] As the system implementation is basically similar to the method implementation, it is described in a relatively simple way. For relevant details, please refer to the description of the method implementation.

[0165] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0166] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0167] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0168] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0169] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0170] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0171] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0172] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0173] The present invention provides a detailed description of a time-series fusion-based three-dimensional target detection method, system, device, and medium. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A three-dimensional target detection method based on long-term temporal fusion, characterized in that, The method includes: The viewpoint of multiple coded features obtained from the viewpoint camera is converted into the viewpoint of the first coded feature among the multiple coded features to obtain the corresponding multiple target coded features; The multiple target encoded features and the first encoded feature are superimposed to obtain the camera features at the time corresponding to the first encoded feature; By performing feature space transformation on the camera features of multiple viewpoint cameras, the corresponding multi-view spatial features are obtained. The multi-view spatial features are input into the 3D detection head for processing to obtain 3D target detection results; The step of obtaining corresponding multi-view spatial features by performing feature space transformation on camera features from multiple viewpoint cameras includes: Monocular depth estimation results are obtained by performing monocular depth estimation on camera features using a monocular depth estimation model. The camera features are used to perform multi-view depth estimation using a multi-view depth estimation model to obtain the multi-view depth estimation results; A depth map is obtained by summing the monocular depth estimation results and the multi-view depth estimation results; By using column pooling, the contextual information in the depth maps of multiple view cameras and the corresponding multi-view depth estimation results is fused and the viewpoint is transformed to obtain the multi-view spatial features at the corresponding time. Before obtaining the depth map by summing the monocular depth estimation results and the multi-view depth estimation results, the method further includes: The depth centers from the monocular depth estimation results and the multi-view depth estimation results are input into a weighted network for processing to obtain a weighted map. Based on the weight map and the depth center and depth range in the multi-view depth estimation results, the target multi-view depth estimation results are obtained by processing through a stereo depth model. The process of obtaining a depth map by summing the monocular depth estimation results and the multi-view depth estimation results includes: A depth map is obtained by summing the monocular depth estimation results and the target multi-view depth estimation results.

2. The three-dimensional target detection method with long-time fusion according to claim 1, characterized in that, Before inputting the multi-view spatial features into the 3D detection head for processing to obtain the 3D target detection result, the method further includes: The perspectives of multiple multi-view spatial features are converted into the perspective of the first multi-view spatial feature among the multiple multi-view spatial features to obtain the corresponding multiple target multi-view spatial features. The multiple target multi-view spatial features and the first multi-view spatial features are concatenated to obtain the temporal fusion features corresponding to the first multi-view spatial features at the time point. The step of inputting the multi-view spatial features into the 3D detection head for processing to obtain 3D target detection results includes: The temporal fusion features are input into the 3D detection head for processing to obtain 3D target detection results.

3. The three-dimensional target detection method with long-term temporal fusion according to claim 1, characterized in that, The step of converting the viewpoint of multiple coded features obtained based on the viewpoint camera into the viewpoint of the first coded feature among the multiple coded features, and obtaining the corresponding multiple target coded features, includes: Determine the intrinsic parameters of the viewpoint camera and the extrinsic parameters of the viewpoint camera relative to the vehicle body coordinate system; Determine a first coded feature obtained based on the viewpoint camera and multiple coded features at multiple times before and after the time corresponding to the first coded feature; Based on the intrinsic parameters, the extrinsic parameters, and the first coding feature, the perspective of the multiple coding features is converted to the perspective of the first coding feature through homography transformation to obtain the corresponding multiple target coding features.

4. The three-dimensional target detection method with long-time fusion according to claim 1, characterized in that, The step of superimposing the plurality of target encoded features and the first encoded feature to obtain the camera features at the time corresponding to the first encoded feature includes: The multiple target encoded features and the first encoded feature are superimposed to obtain the initial camera features at the time corresponding to the first encoded feature; The initial camera features are subjected to dimensionality reduction convolution processing using a target convolution kernel to obtain the camera features at the corresponding time step.

5. The three-dimensional target detection method with long-time fusion according to claim 2, characterized in that, The step of converting the perspective of multiple multi-view spatial features into the perspective of a first multi-view spatial feature among the multiple multi-view spatial features to obtain corresponding multiple target multi-view spatial features includes: Determine the first multi-view spatial feature and multiple multi-view spatial features at multiple times before and after the time corresponding to the first multi-view spatial feature; Based on the first multi-view spatial feature, the perspective of the multiple multi-view spatial features is converted into the perspective of the first multi-view spatial feature through homography transformation, thereby obtaining the corresponding multiple target multi-view spatial features.

6. The three-dimensional target detection method with long-time fusion according to claim 2, characterized in that, The step of inputting the temporal fusion features into the 3D detection head for processing to obtain 3D target detection results includes: The temporal fusion features from multiple time points are input into their respective 3D detection heads for processing to obtain multiple 3D target detection results.

7. The three-dimensional target detection method with long-time fusion according to claim 2, characterized in that, The step of inputting the multi-view spatial features into the 3D detection head for processing to obtain 3D target detection results includes: The multi-view spatial features at multiple time points are input into their respective 3D detection heads for processing to obtain multiple 3D target detection results.

8. The three-dimensional target detection method with long-term temporal fusion according to claim 1, characterized in that, The training of the monocular depth estimation model includes: By performing deep annotation on a preset number of sample camera features, a model training sample set and a model validation sample set are constructed. The initial monocular depth estimation model is trained by inputting the camera features of each target sample in the model training sample set to obtain the corresponding depth prediction results. Based on the depth prediction results, the corresponding depth annotations, and the preset loss function, determine whether the initial monocular depth estimation model has been successfully trained. The initial monocular depth estimation model that has passed training is validated using the validation sample set. If the verification is successful, the initial monocular depth estimation model will be determined as the monocular depth estimation model.

9. A long-time-series fusion three-dimensional target detection system, characterized in that, The system includes: The viewpoint conversion module is used to convert the viewpoint of multiple coded features obtained based on the viewpoint camera into the viewpoint of the first coded feature among the multiple coded features, thereby obtaining the corresponding multiple target coded features. The feature overlay module is used to overlay the plurality of target encoded features and the first encoded feature to obtain the camera features at the time corresponding to the first encoded feature; The feature space transformation module is used to obtain the corresponding multi-view spatial features by performing feature space transformation on the camera features of multiple view cameras; The target detection module is used to input the multi-view spatial features into the three-dimensional detection head for processing to obtain three-dimensional target detection results; The feature space transformation module includes: The monocular depth estimation module is used to perform monocular depth estimation on camera features through a monocular depth estimation model to obtain monocular depth estimation results. The multi-view depth estimation module is used to perform multi-view depth estimation on the camera features through a multi-view depth estimation model to obtain multi-view depth estimation results. The summation module is used to obtain a depth map by summing the monocular depth estimation results and the multi-view depth estimation results; The feature space transformation submodule is used to fuse and transform the contextual information in the depth maps of multiple view cameras and the corresponding multi-view depth estimation results through cylindrical pooling to obtain the multi-view spatial features at the corresponding time. The system also includes: The weight map determination module is used to input the depth centers in the monocular depth estimation results and the multi-view depth estimation results into the weight network for processing to obtain a weight map. The target multi-view depth estimation module is used to process the depth center and depth range in the weight map and the multi-view depth estimation result through a stereo depth model to obtain the target multi-view depth estimation result. The summation module includes: The summation submodule is used to obtain a depth map by summing the monocular depth estimation result and the target multi-view depth estimation result.

10. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements a long-time fusion three-dimensional target detection method as described in any one of claims 1 to 8.

11. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements a long-time-series fusion three-dimensional target detection method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Vehicle aerial view generation method and device, vehicle and storage medium

    CN115273022A

  • Automatic driving 3D target detection method and related device

    CN116259043A