A Multi-Perspective 3D Object Detection Method, System, and Device

By designing an uncertainty-aware projection network, the depth and camera parameter uncertainty are handled, the accuracy and efficiency of multi-view 3D object detection are improved, and the problem of failure to effectively deal with projection uncertainty in the prior art is solved.

CN117197255BActive Publication Date: 2025-07-18DEEP SPACE EXPLORATION LABORATORY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311155703.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-07
Publication Date
2025-07-18
Estimated Expiration
2043-09-07

AI Technical Summary

Technical Problem

The existing 3D object detection method based on multi-view cameras fails to effectively process uncertainty in images into 3D spatial projection, resulting in a decrease in detection accuracy.

Method used

An uncertainty-aware projection network is designed, and the depth uncertainty and camera parameter uncertainty are handled by two modules respectively. The first and second loss functions are used to constrain the 3D position encoding, and the final 3D position embedding is obtained through filtering and weighting summing.

Benefits of technology

It improves the accuracy and efficiency of 3D object detection and effectively deals with uncertainty problems in multi-view camera systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117197255B_ABST
    Figure CN117197255B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-view 3D object detection method, system and device, belonging to the field of computer vision, including: extracting features of an input image; processing an image 3D grid based on the extracted features to obtain 2D pixel coordinates of the image and 3D coordinates of grid points of the image; obtaining a 3D position encoding based on the 3D coordinates and the extracted features, and using a first loss function to constrain the 3D position encoding; predicting a depth map and a variance map from the extracted features, and using a second loss function to constrain the prediction results: Compared with the prior art, the present application separately considers depth uncertainty and camera parameter uncertainty through two modules. The present application considers the work on uncertainty in 3D object detection based on multi-view cameras, and the uncertainty-aware projection network can improve the performance of 3D object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a multi-view 3D object detection method, system and device. Background Art

[0002] Currently, 3D object detection methods based on multi-view cameras can be divided into two categories. The bird's-eye view based method estimates an explicit depth distribution to lift each image to generate a frustum-shaped point cloud, and then pools the frustum in height to a shared bird's-eye view plane. However, due to the image-to-bird's-eye view transformation, the bird's-eye view based method is complex and time-consuming. Another class of methods is called the sparse representation based method, which aims to detect 3D objects without a dense bird's-eye view representation. These methods first generate a uniform depth vector with predefined depth intervals for each point, and the two-dimensional coordinates on the image plane and the depth vector jointly form a 3D grid. The grid is transformed into 3D world space and further encoded into a 3D position embedding as a geometric prior for object queries to predict 3D bounding boxes. These methods are simpler than the bird's-eye view based methods, but their performance is significantly reduced due to the lack of explicit consideration of depth estimation. Although the above methods have made significant progress, they all ignore the uncertainty of the 2D image to 3D space projection. Therefore, this patent designs an uncertainty projection network to address this problem. Summary of the Invention

[0003] Aiming at the deficiencies of the prior art, the present invention proposes a multi-view 3D object detection method, system and device, which respectively consider depth uncertainty and camera parameter uncertainty through two modules. This application considers the work on uncertainty in 3D object detection based on multi-view cameras, and the uncertainty-aware projection network can improve the performance of 3D object detection.

[0004] The object of the present invention can be achieved by the following technical solutions:

[0005] In a first aspect, this application proposes a multi-view 3D object detection method, including:

[0006] Extract the features of the input image; process the image 3D grid based on the extracted features to obtain the 2D pixel coordinates of the image and the 3D coordinates of the grid points of the image;

[0007] Obtain a 3D position encoding based on the 3D coordinates and the extracted features, and constrain the 3D position encoding using a first loss function; predict a depth map and a variance map from the extracted features, and constrain the prediction results using a second loss function:

[0008] The mean of the 2D pixel coordinates and the predicted depth is combined to form 3D points and projected into 3D space to obtain a set of 3D points one; the 2D pixel coordinates and randomly sampled depth values are used to form 3D points and projected into 3D space to obtain a set of 3D points two; the set of 3D points one and the set of 3D points two are input into a position encoder to obtain a 3D position embedding one;

[0009] Based on the camera parameters of the image, the 3D position embedding one is filtered to obtain a 3D position embedding two;

[0010] The 3D position embedding one and the 3D position embedding two are weighted and summed to obtain the final 3D position embedding; the first loss function and the second loss function are weighted and summed to obtain the loss function of the final 3D position embedding;

[0011] Based on the final 3D position embedding, 3D position-aware features are calculated and optimized and trained according to the loss function of the final 3D position embedding to obtain a target query; based on the target query, the object category and the bounding box position are predicted.

[0012] In some embodiments, the acquisition of the 2D pixel coordinates of the image and the 3D coordinates of the grid points of the image includes the following steps:

[0013] Feature extraction is performed on the input image through an image encoder;

[0014] The camera view space is discretized into a 3D grid through a 3D position encoder, and each point in the grid is represented as: p j =(u×d j ,v×d j ,d j ), j = 1,..D, where (u, v) represents the pixel coordinates of the image plane, is a series of predefined depth values;

[0015] The 3D coordinates corresponding to each grid point can be calculated according to the projection matrix of each camera:

[0016] d j ·[u, v, 1] T =K i ·[x i,j ,y i,j ,z i,j ,1] T

[0017] where K i ∈R 3×4 is the projection matrix of the i-th camera.

[0018] In some embodiments, the 3D position encoding satisfies:

[0019]

[0020] Among them represents the set of coordinates of the points where the 3D mesh is projected into the real world under the i-th camera, is a 3D position encoder;

[0021] The first loss function is: L det = λ cls L cls + λ reg L reg

[0022] Among them, λ cls and λ reg are predefined constants, L cls is the classification loss, and L reg is the regression loss.

[0023] In some embodiments, the depth map and the variance map are obtained by continuously processing two convolutional blocks and then performing convolution; wherein the convolutional block sequentially performs convolution, regularization, and ReLU activation processing on the extracted features;

[0024] The second loss function is: L depth = |D μ - D p |;

[0025] Among them, D p is the sparse depth label obtained by projecting the sparse point cloud P cloud onto the image plane; D μ is the corresponding depth value in the depth map.

[0026] In some embodiments, the obtaining of the 3D position embedding one includes the following steps:

[0027] Regularize the predicted depth map and rewrite L depth as:

[0028]

[0029] Use the variance D σ to measure the importance of the two position embeddings of each pixel, and normalize D σ to the interval [0, 1]:

[0030] Among them are respectively the maximum and minimum values in D σ ; The weighted sum of the two 3D position embeddings gives the 3D position embedding one: PE depth = PE μ + λ σ ⊙ PE σ

[0031] where ⊙ represents element-wise multiplication.

[0032] In some embodiments, filtering the 3D position embedding one with the image-based camera parameters to obtain the 3D position embedding two includes the following steps:

[0033] Flatten the camera parameter K ∈ R 3×4 into a 12-dimensional vector, then feed it into a 3-layer MLP, and add a softmax layer to obtain a 9-dimensional vector; reshape this 9-dimensional vector into a 3×3 filter; jointly define these operations as φ, and the filter can be expressed as:

[0034]

[0035] The value at the center of the filter represents the probability of accurate camera parameters, and the other 8 values represent the probabilities of offsets in the surrounding 8 directions caused by camera parameter errors;

[0036] Calculate the weighted sum of the 3D position embedding according to the filter to obtain the 3D position embedding two:

[0037]

[0038] In some embodiments, the final 3D position embedding is: PE = PE depth + λ c PE cam ;

[0039] where λ c is a predefined constant;

[0040] The loss function of the final 3D position embedding is:

[0041] L = L det + λ depth L depth

[0042] where L det is the object detection loss function, i.e., the first loss function; L depth is the depth supervision loss function, i.e., the first loss function; λ depth is a predefined constant.

[0043] In a second aspect, the present application proposes a multi-view 3D object detection system, including;

[0044] Receiving module: receiving the input image;

[0045] Detection module: Extract the features of the input image; Process the 3D mesh of the image based on the extracted features to obtain the 2D pixel coordinates of the image and the 3D coordinates of the grid points of the image;

[0046] Depth estimation module: Obtain the 3D position encoding based on the 3D coordinates and the extracted features, and use the first loss function to constrain the 3D position encoding; Predict the depth map and the variance map for the extracted features, and use the second loss function to constrain the prediction results:

[0047] Depth uncertainty module: Combine the 2D pixel coordinates and the mean of the predicted depths to form 3D points and project them into the 3D space to obtain the first set of 3D points; Combine the 2D pixel coordinates and randomly sampled depth values to form 3D points and project them into the 3D space to obtain the second set of 3D points; Pass the first set of 3D points and the second set of 3D points through the position encoder to obtain the first 3D position embedding;

[0048] Filtering module: Filter the first 3D position embedding based on the camera parameters of the image to obtain the second 3D position embedding;

[0049] Weighting module: Perform weighted summation on the first 3D position embedding and the second 3D position embedding to obtain the final 3D position embedding; Perform weighted summation on the first loss function and the second loss function to obtain the loss function of the final 3D position embedding;

[0050] Output module: Calculate the 3D position-aware features based on the final 3D position embedding, and perform optimization training according to the loss function of the final 3D position embedding to obtain the target query; Predict the object category and the bounding box position based on the target query.

[0051] In a third aspect, the present application proposes a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The terminal device is characterized in that the memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, a multi-view 3D object detection method as described above is adopted.

[0052] In a third aspect, the present application proposes a computer-readable storage medium, in which a computer program is stored. When the computer program is loaded and executed by a processor, a multi-view 3D object detection method as described above is adopted.

[0053] Advantages of the present invention:

[0054] Compared with existing methods, this patent proposes a new uncertainty-aware projection network, which can improve accuracy while ensuring high efficiency. Secondly, this patent designs two modules to separately consider depth uncertainty and camera parameter uncertainty. This application considers the work on uncertainty in 3D object detection based on multi-view cameras, and the uncertainty-aware projection network can improve the performance of 3D object detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The present invention will be further described below with reference to the accompanying drawings.

[0056] Figure 1 Schematic diagram of the uncertainty-guided projection network of this application;

[0057] Figure 2 Schematic diagram of the depth detection module of this application;

[0058] Figure 3 Experimental data graph of the method of this application and PETRv2. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0060] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0061] A multi-view 3D object detection method includes:

[0062] Extract the features of the input image; process the image 3D grid based on the extracted features to obtain the 2D pixel coordinates of the image and the 3D coordinates of the grid points of the image;

[0063] Obtain a 3D position encoding based on the 3D coordinates and the extracted features, and use a first loss function to constrain the 3D position encoding; predict the depth map and variance map from the extracted features, and use a second loss function to constrain the prediction results:

[0064] Combine the mean of 2D pixel coordinates and predicted depth to form 3D points and project them into 3D space to obtain a set of 3D points one; form 3D points by combining 2D pixel coordinates and randomly sampled depth values and project them into 3D space to obtain a set of 3D points two; input the set of 3D points one and the set of 3D points two into a position encoder to obtain a 3D position embedding one;

[0065] Filter the 3D position embedding one based on the camera parameters of the image to obtain a 3D position embedding two;

[0066] Perform weighted summation on the 3D position embedding one and the 3D position embedding two to obtain the final 3D position embedding; perform weighted summation on the first loss function and the second loss function to obtain the loss function of the final 3D position embedding;

[0067] Calculate 3D position-aware features based on the final 3D position embedding and perform optimization training according to the loss function of the final 3D position embedding to obtain a target query; predict the object category and bounding box position based on the target query.

[0068] In some embodiments, the acquisition of the 2D pixel coordinates of the image and the 3D coordinates of the grid points of the image includes the following steps:

[0069] Extract features from the input image through an image encoder;

[0070] Discretize the camera view frustum space into a 3D grid through a 3D position encoder, and each point in the grid is represented as: p j =(u×d j ,v×d j ,d j ), j = 1,..D, where (u, v) represents the pixel coordinates of the image plane, is a series of predefined depth values;

[0071] The 3D coordinates corresponding to each grid point can be calculated according to the projection matrix of each camera:

[0072] d j ·[u, v, 1] T =K i ·[x i,j ,y i,j ,z i,j ,1] T

[0073] where K i ∈R 3×4 is the projection matrix of the i-th camera.

[0074] In some embodiments, the 3D position encoding satisfies:

[0075]

[0076] wherein represents the set of coordinates of the points projected from the 3D mesh onto the real world under the i-th camera, is a 3D position encoder;

[0077] The first loss function is: L det = λ clsLcl s + λ reg L reg

[0078] where λ cls and λ reg are predefined constants, L cls is the classification loss, and L reg is the regression loss.

[0079] In some embodiments, the depth map and the variance map are obtained by consecutive processing through two convolutional blocks followed by convolution; wherein the convolutional blocks sequentially perform convolution, regularization, and ReLU activation on the extracted features;

[0080] The second loss function is: L depth = |D μ - D p |;

[0081] where D p is the sparse depth label obtained by projecting the sparse point cloud P cloud onto the image plane; D μ is the corresponding depth value in the depth map.

[0082] In some embodiments, the obtaining of the 3D position embedding one includes the following steps:

[0083] Regularize the predicted depth map and rewrite L depth as:

[0084]

[0085] Use the variance D σ to measure the importance of the two position embeddings of each pixel, and normalize D σ to the interval [0, 1]:

[0086] where are respectively the maximum and minimum values in D σ ; The weighted sum of the two 3D position embeddings gives the 3D position embedding one: PE depth = PE μ + λ σ ⊙ PE σ

[0087] where ⊙ represents element-wise multiplication.

[0088] In some embodiments, filtering the 3D position embedding one with the image-based camera parameters to obtain the 3D position embedding two includes the following steps:

[0089] Flatten the camera parameter K ∈ R 3×4 into a 12-dimensional vector, then feed it into a 3-layer MLP, followed by a softmax layer, to obtain a 9-dimensional vector; reshape this 9-dimensional vector into a 3×3 filter; collectively define these operations as φ, and the filter can be expressed as:

[0090]

[0091] The value at the center of the filter represents the probability of the camera parameters being accurate, and the other 8 values represent the probabilities of offsets in the surrounding 8 directions caused by errors in the camera parameters;

[0092] Calculate the weighted sum of the 3D position embedding according to the filter to obtain the 3D position embedding two:

[0093]

[0094] In some embodiments, the final 3D position embedding is: PE = PE depth + λ c PE cam ;

[0095] where λ c is a predefined constant;

[0096] The loss function of the final 3D position embedding is:

[0097] L = L det + λ depth L depth

[0098] where L det is the object detection loss function, i.e., the first loss function; L depth is the depth supervision loss function, i.e., the first loss function; λ depth is a predefined constant.

[0099] Example 1: A multi-view 3D object detection method based on an uncertainty projection network with the structure as Figure 1 shown. This method consists of three parts: (1) a 3D object detection backbone network; (2) a depth estimation network; (3) an uncertainty-guided position encoder. The overall technology is as Figure 1 shown, and the training process is as follows:

[0100] (1) 3D object detection backbone network.

[0101] For the input image, PETR is used as the basic object detector. PETR includes an image encoder G, a 3D position encoder P, and a decoder D. The image encoder extracts features from the input N images. In the 3D position encoder, the camera view space is discretized into a 3D grid. Each point in the grid is represented as: p j =(u×d j , v×d j , d j ), j = 1,..D, where (u, v) represents the pixel coordinates in the image plane, is a series of predefined depth values. Then the 3D coordinates corresponding to each grid point can be calculated according to the projection matrix of each camera:

[0102] d j ·[u, v, 1] T = K i ·[x i,j , y i,j , z i,j , 1] T

[0103] where K i ∈R 3×4 is the projection matrix of the i-th camera. The 3D position encoding can be calculated as follows:

[0104]

[0105] where represents the set of coordinates of the points where the 3D grid is projected into the real world under the i-th camera, is the 3D position encoder. Then the 3D position-aware feature is calculated as follows:

[0106]

[0107] where is the feature of the image captured by the i-th camera. Finally, the object query is updated by the multi-head attention and feed-forward network of the decoder D interacting with the 3D position-aware feature. The updated query is used to predict the object category and bounding box position. The overall loss of 3D object detection can be expressed as:

[0108] L det = λ cls L cls + λ reg L reg

[0109] where, λcls and λ reg are predefined constants, L cls is the classification loss, L reg is the regression loss.

[0110] (2) Depth estimation network. For each input image, we designed a lightweight depth estimation module as shown in Figure 2 to predict a dense depth map and variance map where H F and W F are the width and height of the input feature map respectively. The input of the depth estimation module is the feature extracted by the image encoder G. After passing through two convolutional blocks, different convolutional layers are used to predict the depth map and variance map. The convolutional block includes a convolutional layer, a regularization layer, and a ReLU activation layer; that is, the feature extracted by the image encoder G first passes through the first convolutional block for convolution, regularization, and ReLU activation processing in sequence; then passes through the second convolutional block for convolution, regularization, and ReLU activation processing in sequence; finally, the predicted depth map and variance map are obtained through convolutional processing by two convolutional layers respectively;

[0111] The specific steps are that the feature extracted by the image encoder G, after passing through two convolutional blocks, obtains a feature F of H F ×W F ×C. For this feature, two convolutional layers with a convolutional kernel of 3*3 are used to predict the depth map depth and variance map respectively.

[0112] To better train the depth estimation module, this patent projects the sparse point cloud P cloud onto the image plane to obtain the sparse depth label D p , where the sparse point cloud is obtained by scanning with a radar sensor: the L1 loss function is used to constrain the depth estimation module:

[0113] L depth = |D μ - D p |

[0114] (3) Uncertainty-guided position encoder. This patent designs a new uncertainty-guided 3D position encoder P, which is divided into two components:

[0115] ① Depth uncertainty module. Modeling the possible depth distribution of pixels with a Gaussian distribution, we input two 3D point coordinates into two position encoders to generate different position embeddings.

[0116] (a) 2D pixel coordinates (u, v) and the predicted depth dμ The means form a 3D point (u, v, d μ ), where d μ ∈D μ . After projecting it into 3D space, the corresponding 3D point set one is obtained Input into the 3D position encoder to obtain the depth-guided 3D position embedding PE μ .

[0117] (b) For each pixel (u, v), randomly sample D depth values from a Gaussian distribution and sort them in ascending order; since 2D points with different depth values will be projected onto a ray in 3D space, sorting the depths in ascending order is to ensure that the projected 3D points are from near to far from the camera, which is convenient for the model to learn. Among them,

[0118] represents the Gaussian distribution formed by the element in the u-th row and v-th column with mean D and variance D μ in the u-th row and v-th column. The projected 3D point set two σ is used to generate the depth-uncertainty-guided position embedding PE . To regularize the predicted depth map, we further rewrite L σ h as follows: dept

[0119]

[0120] In this formula, pixels with difficult-to-predict depths will result in large variances. Since the depth uncertainty module we designed has the ability to predict depth uncertainty, we can use the variance D σ to measure the importance of the two position embeddings of each pixel. We normalize D σ to the interval [0, 1]:

[0121]

[0122] Among them, are the maximum and minimum values in D σ respectively. The weighted sum formula of the two 3D position embeddings is as follows:

[0123] PE depth = PE μ + λ σ ⊙ PE σ

[0124] ​where ⊙ represents element-wise multiplication. Additionally, when training a deep predictor with random samples, we cannot take the derivative with respect to the random variables, so backpropagation cannot be performed. Therefore, we use the reparameterization trick to reformulate the problem. Specifically, we first sample ε from the standard Gaussian distribution where ε is a number sampled from the standard Gaussian distribution ; then we calculate the sample using D μ + εD σ . The reparameterization trick can split the process of random sampling, which makes our deep predictor trainable.

[0125] ② Camera parameter uncertainty module. To model the camera parameter uncertainty, we predict a 3×3 filter based on the camera parameters of each image. Specifically, we tile the camera parameter K ∈ R 3×4 into a 12-dimensional vector, then feed it into a 3-layer MLP, followed by a softmax layer, to obtain a 9-dimensional vector. Reshape this 9-dimensional vector into a 3×3 filter. We jointly define these operations as φ, and the filter can be expressed as:

[0126]

[0127] The value at the center of the filter represents the probability that the camera parameters are accurate, and the other 8 values represent the probabilities of offsets in the 8 surrounding directions caused by camera parameter errors. Therefore, we calculate the weighted sum of the 3D position embeddings according to the filter, and this process can be regarded as a convolution operation:

[0128]

[0129] where PE cam is the position encoding guided by camera parameters;

[0130] Then the final 3D position embedding PE is:

[0131] PE = PE depth + λ c PE cam

[0132] where λ c is a predefined constant.

[0133] The loss of the final 3D position embedding PE model satisfies:

[0134] L = L det + λ depth L depth

[0135] where L det is the object detection loss function, Ldepth is the deep supervision loss function, and λ depth is a predefined constant.

[0136] The 3D position-aware feature of the final i-th image is:

[0137]

[0138] Input into a decoder with the same structure as PETR, and predict the detection box and category according to the object output. Use the overall loss L to train the object detector, depth predictor, and uncertainty-guided position encoder, and at the same time adjust the parameters of these three parts to minimize the loss, so as to optimize the model.

[0139] In this patent, the 3D object detection backbone network is an existing technology in the background art. Based on this technology, this patent adds a depth estimation module and designs a new uncertainty-guided 3D position encoder.

[0140] This application embodiment discloses a multi-view 3D object detection system, including:

[0141] Receiving module: Receive the input image;

[0142] Detection module: Extract the features of the input image; Process the 3D grid of the image based on the extracted features to obtain the 2D pixel coordinates of the image and the 3D coordinates of the grid points of the image;

[0143] Depth estimation module: Obtain the 3D position encoding based on the 3D coordinates and the extracted features, and use the first loss function to constrain the 3D position encoding; Predict the depth map and variance map for the extracted features, and use the second loss function to constrain the prediction results:

[0144] Depth uncertainty module: Combine the 2D pixel coordinates and the mean of the predicted depth to form a 3D point and project it into 3D space to obtain a 3D point set one; Combine the 2D pixel coordinates and a randomly selected depth value to form a 3D point and project it into 3D space to obtain a 3D point set two; Pass the 3D point set one and the 3D point set two through the position encoder to obtain a 3D position embedding one;

[0145] Filtering module: Filter the 3D position embedding one based on the camera parameters of the image to obtain a 3D position embedding two;

[0146] Weighting module: Weight and sum the 3D position embedding one and the 3D position embedding two to obtain the final 3D position embedding; Weight and sum the first loss function and the second loss function to obtain the loss function of the final 3D position embedding;

[0147] Output module: Calculate 3D position perception features based on the final 3D position embedding, and perform optimization training according to the loss function of the final 3D position embedding to obtain the query of the target; predict the object category and bounding box position based on the query of the target.

[0148] An embodiment of the present application also discloses a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, any one of the multi-view 3D object detection methods based on the uncertainty projection network in the above embodiments is adopted.

[0149] Among them, the terminal device can be a computer device such as a desktop computer, a laptop computer, or a cloud server, and the terminal device includes, but is not limited to, a processor and a memory. For example, the terminal device may also include an input / output device, a network access device, and a bus, etc.

[0150] Among them, the processor can adopt a central processing unit (CPU). Of course, according to the actual usage situation, other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be adopted. The general-purpose processor can adopt a microprocessor or any conventional processor, etc. The present application does not make any restrictions on this.

[0151] Among them, the memory can be an internal storage unit of the terminal device, for example, the hard disk or memory of the terminal device, or an external storage device of the terminal device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital card (SD), or a flash card (FC) equipped on the terminal device, etc. And the memory can also be a combination of the internal storage unit and the external storage device of the terminal device. The memory is used to store the computer program and other programs and data required by the terminal device. The memory can also be used to temporarily store the data that has been output or will be output. The present application does not make any restrictions on this.

[0152] Among them, through this terminal device, any one of the multi-view 3D object detection methods based on the uncertainty projection network in the above embodiments is stored in the memory of the terminal device, and is loaded and executed on the processor of the terminal device for convenient use.

[0153] An embodiment of the present application also discloses a computer-readable storage medium, and the computer-readable storage medium stores a computer program. When the computer program is executed by the processor, any one of the multi-view 3D object detection methods based on the uncertainty projection network in the above embodiments is adopted.

[0154] Among them, the computer program can be stored in a computer-readable medium. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some middleware form, etc. The computer-readable medium includes any entity or device, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the computer-readable medium includes but is not limited to the above components.

[0155] Among them, through this computer-readable storage medium, any one of the multi-view 3D object detection methods based on the uncertainty projection network in the above embodiments is stored in the computer-readable storage medium, and is loaded and executed on the processor to facilitate the storage and application of the above method.

[0156] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.

Claims

1. A multi-view 3D object detection method, characterized in that, Including: Extract the features of the input image; Process the 3D mesh of the image based on the extracted features to obtain the 2D pixel coordinates of the image and the 3D coordinates of the grid points of the image; Obtain the 3D position encoding based on the 3D coordinates and the extracted features, and use the first loss function to constrain the 3D position encoding; predict the depth map and variance map from the extracted features, and use the second loss function to constrain the prediction results: Form 3D points by combining the 2D pixel coordinates and the mean of the predicted depths and project them into 3D space to obtain the first set of 3D points; form 3D points by combining the 2D pixel coordinates and randomly sampled depth values and project them into 3D space to obtain the second set of 3D points; input the first set of 3D points and the second set of 3D points into the position encoder to obtain the first 3D position embedding; Filter the first 3D position embedding based on the camera parameters of the image to obtain the second 3D position embedding; Perform weighted summation on the first 3D position embedding and the second 3D position embedding to obtain the final 3D position embedding; perform weighted summation on the first loss function and the second loss function to obtain the loss function of the final 3D position embedding; Calculate the 3D position-aware features based on the final 3D position embedding and perform optimization training according to the loss function of the final 3D position embedding to obtain the target query; Predict the object category and bounding box position based on the target query; The obtaining of the first 3D position embedding includes the following steps: Regularized predicted depth map, the L depth Rewrite as: Use variance D σ to measure the importance of the two position embeddings of each pixel, and normalize D σ to the interval [0, 1]: where L depth is the deep supervision loss function, i.e., the second loss function; D μ is the corresponding depth value in the depth map, D p is the sparse point cloud P cloud projected onto the image plane to obtain a sparse depth label; are the minimum and maximum values in D σ respectively; The weighted sum of two 3D positional embeddings results in 3D positional embedding one: PE depth = PE μ + λ σ ⊙ PE σ Where ⊙ represents element-wise multiplication; The filtering of the first 3D position embedding based on the camera parameters of the image to obtain the second 3D position embedding includes the following steps: Let the camera parameter \(K\in\mathbb{R}\) 3×4 be flattened into a 12 - dimensional vector, which is then fed into a 3 - layer MLP followed by a softmax layer to obtain a 9 - dimensional vector; reshape this 9 - dimensional vector into a \(3\times3\) filter; collectively define these operations as \(\varphi\), and the filter is expressed as: The value at the center of the filter represents the probability that the camera parameters are accurate, and the other 8 values represent the probabilities of offsets in the 8 surrounding directions caused by camera parameter errors; Calculate the weighted sum of the 3D position embedding according to the filter to obtain the second 3D position embedding: The final 3D positional embedding is: PE = PE depth + λ c PE cam ; where λ c is a predefined constant, and PE cam is the position encoding guided by camera parameters; The loss function of the final 3D position embedding is: L = L det + λ depth L depth Among which L det is the target detection loss function, i.e., the first loss function; λ depth is a predefined constant.

2. The multi-view 3D object detection method according to claim 1, wherein The obtaining of the 2D pixel coordinates of the image and the 3D coordinates of the grid points of the image includes the following steps: Extract features from the input image through an image encoder; The camera view - table space is discretized into a 3D grid by a 3D position encoder, and each point in the grid is represented as: p j =(u × d j , v × d j , d j ), j = 1,..D, where (u, v) represents the pixel coordinates of the image plane, is a series of predefined depth values; The 3D coordinates corresponding to each grid point Calculated according to the projection matrix of each camera: d j ·[u, v, 1] T =K i ·[x i,j , y i,j , z i,j , 1] T where K i ∈R 3×4 is the projection matrix of the i-th camera.

3. The multi-view 3D object detection method according to claim 2, wherein The 3D position encoding satisfies: where represents the set of coordinates of the points projected from the 3D mesh onto the real world under the i-th camera, is the 3D position encoder, H F and W F are the width and height of the input feature map, respectively; The first loss function is: L det = λ cls L cls + λ reg L reg where λ cls and λ reg are predefined constants, L cls is the classification loss, and L reg is the regression loss.

4. The multi-view 3D object detection method according to claim 1, characterized in that The depth map and variance map are obtained by continuous processing of two convolutional blocks followed by convolution; where the convolutional block sequentially performs convolution, regularization, and ReLU activation processing on the extracted features; The second loss function is: L depth = |D μ - D p |.

5. A multi-view 3D object detection system, characterized in that, The system is used to implement a multi-view 3D object detection method according to any one of claims 1 to 4, and the system includes; Receiving module: Receive the input image; Detection module: Extract the features of the input image; Process the 3D mesh of the image based on the extracted features to obtain the 2D pixel coordinates of the image and the 3D coordinates of the grid points of the image; Depth estimation module: Obtain the 3D position encoding based on the 3D coordinates and the extracted features, and use the first loss function to constrain the 3D position encoding; predict the depth map and variance map from the extracted features, and use the second loss function to constrain the prediction results: Depth Uncertainty Module: Combine the 2D pixel coordinates and the mean of the predicted depth to form 3D points and project them into 3D space to obtain the first set of 3D points; form 3D points by combining the 2D pixel coordinates and randomly sampled depth values and project them into 3D space to obtain the second set of 3D points; pass the first set of 3D points and the second set of 3D points through a position encoder to obtain the first 3D position embedding; Filtering Module: Filter the first 3D position embedding based on the camera parameters of the image to obtain the second 3D position embedding; Weighting Module: Perform weighted summation on the first 3D position embedding and the second 3D position embedding to obtain the final 3D position embedding; perform weighted summation on the first loss function and the second loss function to obtain the loss function of the final 3D position embedding; Output Module: Calculate 3D position-aware features based on the final 3D position embedding, and perform optimization training according to the loss function of the final 3D position embedding to obtain the target query; predict the object category and bounding box position based on the target query.

6. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The memory stores a computer program that can run on a processor. When the processor loads and executes the computer program, it adopts a multi-view 3D object detection method according to any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is loaded and executed by the processor, it adopts a multi-view 3D object detection method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Construction method of multi-plane coding point cloud feature deep learning model based on pointpillars

    CN111612059A

  • Transformer-based multi-view target detection method and system

    CN113673425A