Three-dimensional target detection method for real-time perception of mobile environment
By adopting technical means of multi-level deep supervision and semantic supervision in the three-dimensional object detection method, combined with the use of mixed resolution grids, the problem of insufficient accuracy of three-dimensional object detection in complex mobile environments in the prior art is solved, and more efficient environmental perception and detection adaptability are achieved.
Patent Information
- Application Number
- CN202510071644.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The existing three-dimensional object detection methods have problems such as insufficient spatial characterization intuitiveness, insufficient utilization of timing information, and difficulty in fusion of multi-sensor data in complex and dynamic mobile environments, especially in terms of depth estimation and detection accuracy.
A three-dimensional object detection method for real-time perception of mobile environments is adopted. The image two-dimensional feature input depth prediction network and semantic prediction network are combined with multi-level deep supervision and semantic supervision to improve the utilization efficiency of depth information, and adapt to different detection needs through a hybrid resolution grid.
It significantly improves the accuracy and adaptability of three-dimensional object detection, can more accurately perceive the environment around mobile devices, and adapt to complex and dynamic mobile environment needs.
Smart Images

Figure CN119992051A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and in particular relates to a three-dimensional target detection method for real-time perception in a mobile environment. Background Art
[0002] With the rapid development of intelligent technology, mobile devices such as vehicles, automatic guided vehicles (AGVs), drones, and mobile robots have been widely used in industry, transportation, logistics, and service fields. These devices have extremely high requirements for real-time perception of the surrounding environment and target detection when performing tasks such as autonomous driving, intelligent transportation, environmental monitoring, and human-computer interaction. Especially in complex and dynamic mobile environments, accurate three-dimensional target detection is crucial for collision avoidance, path planning, and decision execution. For application scenarios such as vehicles, AGVs, drones, and mobile robots, a three-dimensional target detection method for real-time perception of mobile environments is studied, which can not only significantly improve the environmental perception ability of the equipment, but also provide important technical support for autonomous driving, intelligent logistics, and multi-field applications.
[0003] Traditional 3D target detection methods have shortcomings in terms of intuitive spatial representation, utilization of temporal information, and multi-sensor data fusion. For example, 2D image detection is easily affected by lighting conditions, occlusion, and changes in viewing angle, while the data sparsity and limited field of view of a single sensor often make it difficult to fully perceive the 3D information of the target.
[0004] In recent years, 3D target detection methods based on multi-sensor data fusion and efficient computing have gradually attracted attention. These methods integrate data from multiple sensors such as cameras and lidars, and combine bird's eye view (BEV) technology and deep learning models to provide more comprehensive and robust target perception capabilities in complex dynamic environments.
[0005] Compared with traditional methods, the BEV method can provide a more intuitive and unified environmental representation by uniformly converting multi-sensor data into a bird's-eye view, thus showing significant advantages in three-dimensional target detection, but there are still some shortcomings. On the one hand, the existing BEV three-dimensional target detection method usually only focuses on the absolute depth information of each point of the target, and the geometric structure characteristics of the target itself and the spatial distribution information between targets are not utilized, resulting in a lack of sufficient supervision of depth estimation, thereby affecting the detection accuracy. On the other hand, the existing BEV three-dimensional target detection method usually sets a grid of uniform size in the BEV space. Considering actual needs, more accurate detection results are usually required at locations closer to the self, and the uniform size grid setting makes the entire detection range at the same detection requirement level, which cannot be well adapted to actual detection needs. Taking the above factors into consideration, the present invention proposes a three-dimensional target detection method for real-time perception of mobile environments. Summary of the invention
[0006] The purpose of the present invention is to provide a three-dimensional target detection method for real-time perception in mobile environments, which makes full use of depth information to improve the detection effect and adapts to actual detection needs through a mixed resolution grid.
[0007] To achieve the above object, the technical solution adopted by the present invention is:
[0008] A three-dimensional target detection method for real-time perception in a mobile environment, comprising:
[0009] Input the multi-view images of the mobile device in the training set into the image backbone network to extract the two-dimensional features of the image;
[0010] The two-dimensional features of the image are input into the depth prediction network and the semantic prediction network respectively to obtain the depth prediction value and the semantic prediction value, and multi-level depth supervision is performed on the depth prediction value, and semantic supervision is performed on the semantic prediction value;
[0011] Project the two-dimensional features of the image into three-dimensional space to obtain BEV features;
[0012] The mixed-resolution grid is used to perform bilinear interpolation sampling on the BEV features to form new BEV features;
[0013] The new BEV features are input into the BEV feature encoding network, and then the 3D object detection results of the mobile device environment are obtained through the 3D detection head;
[0014] Perform detection supervision on the 3D target detection results, and combine multi-level depth supervision and semantic supervision to update the image backbone network, depth prediction network, semantic prediction network, BEV feature encoding network and 3D detection head;
[0015] New multi-view images of mobile devices are collected in real time and input into the trained image backbone network, depth prediction network, semantic prediction network, BEV feature encoding network and 3D detection head to obtain the corresponding three-dimensional target detection results, and the bird's-eye view detection results are obtained based on the three-dimensional target detection results.
[0016] Several optional methods are also provided below, but they are not intended to be additional limitations on the above-mentioned overall solution, but are merely further supplements or preferences. Under the premise that there are no technical or logical contradictions, each optional method can be combined with the above-mentioned overall solution separately, and multiple optional methods can also be combined.
[0017] Preferably, the multi-level depth supervision of the depth prediction value comprises:
[0018] (1) Semantically weighted point-level absolute depth supervision: The depth truth value and semantic truth value of the multi-view image of the mobile device are one-hot encoded and downsampled to the size of the image’s two-dimensional features to obtain the depth label value D lable and semantic label value S lable , then based on the depth prediction value D pred and the depth tag value D lable , calculate the Binary Cross Entropy loss as follows:
[0019]
[0020] Where L1 represents the semantically weighted point-level absolute depth loss of multi-view images on mobile devices, Indicates that the input is the depth prediction value D pred and the depth tag value D lable Binary CrossEntropy loss, ω s is the foreground background weight factor, ω1 is the foreground weight factor, ω2 is the background weight factor, S lable ≠0 means the pixel belongs to the foreground, S lable =0 means the pixel belongs to the background;
[0021] (2) Relative depth supervision at the target internal structure level: The pixel with the smallest absolute value of the internal depth prediction error of each target in the multi-view image of the mobile device is taken as the internal structure reference point θ of the target. The depth prediction value and depth label value of the internal structure reference point θ are recorded as and The absolute value of the depth prediction error is the absolute value of the difference between the depth prediction value and the depth label value. The depth prediction difference and the depth label difference of other pixels inside each target relative to the internal structure reference point θ are calculated, and the MSE loss of the two is calculated. The formula is as follows:
[0022]
[0023] Where L2 represents the relative depth loss of the target internal structure level of the multi-view image of the mobile device, Indicates that the input is a depth prediction difference Difference between the depth label and MSE loss;
[0024] (3) Relative depth supervision at the target spatial distribution level: In each multi-view image of a mobile device, the point with the smallest depth label value inside each target is taken as the representative point α of the target. The depth prediction value and depth label value of the representative point α are recorded as and Calculate the depth prediction value and the depth tag value The difference between the two is taken, and the representative point with the smallest absolute difference in each mobile device multi-view image is taken as the target spatial distribution reference point β of the mobile device multi-view image where the representative point is located. The depth prediction value and depth label value of the spatial distribution reference point β are recorded as and Then calculate the depth prediction difference and depth label difference of other representative points α in each mobile device multi-view image relative to the spatial distribution reference point β of the mobile device multi-view image, and calculate the MSE loss of the two. The formula is as follows:
[0025]
[0026] Where L3 represents the relative depth loss of the target spatial distribution level of multi-view images of mobile devices, Indicates that the input is a depth prediction difference Difference between the depth label and MSE loss;
[0027] (4) The total loss of deep supervision is calculated as follows:
[0028] L Depth =ε1*L1+ε2*L2+ε3*L3
[0029] Where, L Depth represents the total loss of deep supervision of multi-view images on mobile devices, ε1 represents the first-level loss coefficient, ε2 represents the second-level loss coefficient, and ε3 represents the third-level loss coefficient.
[0030] Preferably, the performing semantic supervision on the semantic prediction value comprises:
[0031] The semantic truth value of the multi-view image of the mobile device is one-hot encoded and downsampled to the size of the two-dimensional feature of the image to obtain the semantic label value S lable , based on the semantic prediction value S predand semantic label value S lable , calculate the Binary CrossEntropy loss as follows:
[0032]
[0033] Where, L Seg represents semantic loss, ε4 represents semantic loss coefficient, Indicates that the input is a semantic prediction value S pred and semantic label value S lable Binary Cross Entropy loss.
[0034] Preferably, projecting the two-dimensional features of the image into three-dimensional space to obtain the BEV features comprises:
[0035] Perform outer product of the image two-dimensional features and the depth prediction value to obtain the stretched image two-dimensional features;
[0036] The features whose depth prediction value is less than the depth threshold or whose semantic prediction value is less than the semantic threshold in the two-dimensional features of the stretched image are filtered;
[0037] The remaining features after filtering in the stretched two-dimensional features of the image are projected into the three-dimensional space to obtain the BEV features.
[0038] Preferably, the method of using a mixed resolution grid to perform bilinear interpolation sampling on the BEV feature to form a new BEV feature includes:
[0039] A plane rectangular coordinate system is established with the center of the grid plane of the BEV feature space as the origin;
[0040] The division for the x-axis is as follows: in the range of |x|≤x1, the division is based on the side length of each grid being d1; in the range of x1<|x|≤x2, the division is based on the side length of each grid being d2; in the range of x2<|x|≤x3, the division is based on the side length of each grid being d3, where x1 <x2<x3,d1<d2<d3;
[0041] The division for the y-axis is as follows: in the range of |y|≤y1, the division is based on the side length of each grid being d1; in the range of y1<|y|≤y2, the division is based on the side length of each grid being d2; in the range of y2<|y|≤y3, the division is based on the side length of each grid being d3, where x1=y1, x2=y2, x3=y3;
[0042] Based on the new grid obtained after division, bilinear interpolation sampling is performed on the BEV features to obtain new BEV features.
[0043] Preferably, the step of performing detection supervision on the three-dimensional target detection result includes:
[0044] Obtaining the category label and the 3D detection frame label marked by the multi-view image of the mobile device corresponding to the 3D target detection result;
[0045] Calculate the Gaussian focus loss as the category prediction loss based on the category label and the category prediction value in the 3D object detection result;
[0046] The L1 loss is calculated based on the 3D detection box label and the 3D detection box prediction value in the 3D object detection result as the 3D detection box prediction loss.
[0047] Preferably, the real-time acquisition of new multi-view images of mobile devices, input of trained image backbone network, depth prediction network, semantic prediction network, BEV feature encoding network and 3D detection head, and obtaining corresponding three-dimensional target detection results include:
[0048] Collect new multi-view images from mobile devices and input them into the trained image backbone network to extract the two-dimensional features of the image to be perceived;
[0049] Input the two-dimensional features of the image to be perceived into the trained depth prediction network and semantic prediction network respectively to obtain a depth prediction value and a semantic prediction value;
[0050] Projecting the two-dimensional features of the image to be sensed into the three-dimensional space to obtain the BEV features to be sensed;
[0051] A mixed-resolution grid is used to perform bilinear interpolation sampling on the BEV features to be sensed to form new BEV features to be sensed;
[0052] The new BEV features to be perceived are input into the trained BEV feature encoding network, and then the three-dimensional target detection results of the mobile device environment are obtained through the trained 3D detection head.
[0053] The present invention provides a three-dimensional target detection method for real-time perception in a mobile environment. The method makes full use of depth information through multi-level depth supervision including semantically weighted point-level absolute depth supervision, target internal structure-level relative depth supervision and target spatial distribution-level relative depth supervision. At the same time, by setting a mixed-resolution BEV grid, grids of different sizes are set according to the distance from their own center to adapt to different detection requirements without changing the detection range and the size of the BEV feature map. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A flowchart of a three-dimensional target detection method for real-time perception in a mobile environment according to the present invention applied in a training phase;
[0055] Figure 2 A schematic diagram of an embodiment of a multi-view image of a mobile device applied by the present invention;
[0056] Figure 3 A schematic diagram of an embodiment of mixed-resolution grid division of the present invention;
[0057] Figure 4 A flowchart of a three-dimensional target detection method for real-time perception in a mobile environment according to the present invention applied in the reasoning stage;
[0058] Figure 5 A schematic diagram of an embodiment of a three-dimensional target detection result and a bird's-eye view detection result of the present invention. DETAILED DESCRIPTION
[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0061] Embodiment 1: A three-dimensional target detection method for real-time perception in a mobile environment is applied in the training phase, such as Figure 1 As shown, the following steps are included:
[0062] Step 1: Input the multi-view images of the mobile device in the training set into the image backbone network to extract the image two-dimensional features (abbreviated as image features), denoted as F img ∈R N×C×H×W , where N represents the number of multi-view images of mobile devices in the training set, C represents the number of feature channels of the two-dimensional features of the image, and H and W represent the size of the feature map of the two-dimensional features of the image. In this embodiment, the backbone network adopts the ResNet50 neural network.
[0063] The multi-view images of the mobile device should be understood as obtaining one or more view images around the mobile device. If there are multiple view images, the multiple images corresponding to the multiple view images at the same time are used as the original input of the image backbone network. Figure 2As shown, the multi-view images of the mobile device obtained include images taken by image acquisition devices in front of the vehicle, in front of the left, in the rear left, in the rear, in the rear right and in the front right. In other embodiments, the viewing angles can be flexibly combined according to the application scenario and hardware conditions, such as only one viewing angle of the front, or three viewing angles of the front + left front + right front, or two viewing angles of the front + rear. If the mobile device is large or the camera viewing angle range is small, it can be expanded to eight viewing angles, etc. This embodiment does not limit this.
[0064] Step 2: Input the two-dimensional features of the image into the depth prediction network and the semantic prediction network respectively to obtain the depth prediction value and the semantic prediction value, perform multi-level depth supervision on the depth prediction value, and perform semantic supervision on the semantic prediction value.
[0065] Step 21: transform the image two-dimensional feature F img Send them to the depth prediction network and semantic prediction network respectively to get the depth prediction value D pred ∈R N×118×H×W and semantic prediction value S pred ∈R N×2×H×W , where 118 represents 118 different discrete depths, 2 represents the length of the corresponding dimension, and includes foreground probability and background probability, which are used to distinguish foreground from background. For example, (0.4, 0.6) means that the foreground probability of the pixel is 0.4 and the background probability is 0.6.
[0066] The discrete depth range in this embodiment is 1 to 60m, with an interval of 0.5m, and it should be noted that in other embodiments, the number, range and interval of discrete depths can be adjusted as needed. In addition, the depth prediction network and semantic prediction network in this embodiment are convolutional neural networks composed of basic convolution modules, which are not specifically limited in this embodiment.
[0067] Step 22: Depth prediction value D pred Multi-level deep supervision is performed, which includes semantically weighted point-level absolute depth supervision, target internal structure-level relative depth supervision, and target spatial distribution-level relative depth supervision, as follows:
[0068] (1) Semantically weighted point-level absolute depth supervision: The depth truth value and semantic truth value of the multi-view image of the mobile device are one-hot encoded and downsampled to the image feature map size to obtain the depth label value Dl able ∈R N×118×H×W and semantic label value S lable ∈R N×2×H×W , then based on the depth prediction value D pred and the depth tag value D lable , calculate the Binary CrossEntropy loss as follows:
[0069]
[0070] Where L1 represents the semantically weighted point-level absolute depth loss of multi-view images on mobile devices, Indicates that the input is the depth prediction value D pred and the depth tag value D lable Binary CrossEntropy loss, ω s is the foreground background weight factor, ω1 is the foreground weight factor, ω2 is the background weight factor, S lable ≠0 means the pixel belongs to the foreground, S lable =0 indicates that the pixel belongs to the background. In this embodiment, different weight factors ω are assigned to the foreground and background when calculating the loss. s , to balance a large number of background pixels and a small number of foreground pixels. Semantic weighted point-level absolute depth supervision only supervises the absolute depth error of each point. Points are independent of each other and emphasize the supervision of foreground pixels.
[0071] (2) Relative depth supervision at the target internal structure level: The pixel with the smallest absolute value of the internal depth prediction error of each target in the multi-view image of the mobile device is taken as the internal structure reference point θ of the target. The depth prediction value and depth label value of the internal structure reference point θ are recorded as and The absolute value of the depth prediction error is the absolute value of the difference between the depth prediction value and the depth label value. The depth prediction difference and the depth label difference of other pixels inside each target relative to the internal structure reference point θ are calculated, and the MSE loss of the two is calculated. The formula is as follows:
[0072]
[0073] Where L2 represents the relative depth loss of the target internal structure level of the multi-view image of the mobile device, Indicates that the input is a depth prediction difference Difference between the depth label and The relative depth supervision at the internal structure level of the target supervises the relative depth between the internal points of the target. All the internal points of each target in the same image are interrelated, making full use of the geometric structure information of the target itself.
[0074] (3) Relative depth supervision at the target spatial distribution level: In each multi-view image of a mobile device, the point with the smallest depth label value inside each target is taken as the representative point α of the target. The depth prediction value and depth label value of the representative point α are recorded as and Calculate the depth prediction value and the depth tag value The difference between , takes the representative point with the smallest absolute difference value in each mobile device multi-view image as the target spatial distribution reference point β of the mobile device multi-view image where the representative point is located (assuming that there are 6 multi-view images of the mobile device, there are 6 spatial distribution reference points β, and each β is only used for the image where it is located). The depth prediction value and depth label value of the spatial distribution reference point β are recorded as and Then calculate the depth prediction difference and depth label difference of other representative points α in each mobile device multi-view image relative to the spatial distribution reference point β of the mobile device multi-view image, and calculate the MSE loss of the two. The formula is as follows:
[0075]
[0076] Where L3 represents the relative depth loss of the target spatial distribution level of multi-view images of mobile devices, Indicates that the input is a depth prediction difference Difference between the depth label and The relative depth supervision at the target spatial distribution level supervises the relative depths between target representative points. The targets are interrelated and the spatial distribution information between targets is fully utilized.
[0077] (4) Calculate the total loss of deep supervision and assign different loss coefficients to each level:
[0078] L Depth =ε1*L1+ε2*L2+ε3*L3
[0079] Where, L Depth represents the total loss of deep supervision of multi-view images on mobile devices, ε1 represents the first-level loss coefficient, ε2 represents the second-level loss coefficient, and ε3 represents the third-level loss coefficient.
[0080] Step 23: Based on the semantic prediction value S pred and semantic label value S lable , calculate the Binary Cross Entropy loss and perform semantic supervision on the semantic prediction value:
[0081]
[0082] Where, L Seg represents semantic loss, ε4 represents semantic loss coefficient, Indicates that the input is a semantic prediction value S pred and semantic label value S lable Binary Cross Entropy loss.
[0083] Step 24: Calculate the total loss L of the two-dimensional features of the image Zas follows:
[0084] L Z =L Depth +L Seg
[0085] Step 3: Project the two-dimensional features of the image into three-dimensional space to obtain BEV features.
[0086] Based on the Lift-Splat-Shoot (LSS) algorithm, the specific process of obtaining the BEV feature in this embodiment is as follows: img The outer product is performed with the depth prediction value to obtain the two-dimensional features of the stretched image (an operation within the Lift-Splat-Shoot algorithm); in order to improve the projection features, the features whose depth prediction values are less than the depth threshold (set to 0.0085 in this embodiment, which can be adjusted) or whose semantic prediction values are less than the semantic threshold (set to 0.25 in this embodiment, which can be adjusted) in the two-dimensional features of the stretched image are filtered out; then the remaining features after filtering the two-dimensional features of the stretched image are projected into the three-dimensional space to obtain the BEV features, which are recorded as Among them C BEV Represents the number of feature channels, 128 represents the BEV feature map size, and the BEV feature map size can be selected according to actual conditions.
[0087] Step 4: Use a mixed-resolution grid to perform bilinear interpolation sampling on the BEV features to form new BEV features.
[0088] The grid of the original BEV feature space is composed of 128*128 squares with a side length of 0.8m. To adapt to the actual detection requirements, a mixed resolution grid is used to perform bilinear interpolation sampling on the BEV features without changing the number of grids and the detection range to form a new BEV feature, denoted as F′. BEV The specific operation is as follows: first, the grid of the original BEV feature space is re-divided to obtain a new grid, namely a mixed-resolution grid, and then bilinear interpolation sampling is performed on the original BEV feature based on the mixed-resolution grid to form a new BEV feature.
[0089] The specific process of dividing the mixed-resolution grid is as follows: A plane rectangular coordinate system is established with the center of the grid plane in the original BEV feature space as the origin. The division for the x-axis is as follows: within |x| ≤ x1, it is divided with each grid side length being d1; within x1 < |x| ≤ x2, it is divided with each grid side length being d2; within x2 < |x| ≤ x3, it is divided with each grid side length being d3, where x1 < x2 < x3 and d1 < d2 < d3. The division for the y-axis is as follows: within |y| ≤ y1, it is divided with each grid side length being d1; within y1 < |y| ≤ y2, it is divided with each grid side length being d2; within y2 < |y| ≤ y3, it is divided with each grid side length being d3, where x1 = y1, x2 = y2, and x3 = y3. x1 is the first division boundary value of the x-axis, x2 is the second division boundary value of the x-axis, x3 is the third division boundary value of the x-axis, y1 is the first division boundary value of the y-axis, y2 is the second division boundary value of the y-axis, y3 is the third division boundary value of the y-axis, d1 is the first length value, d2 is the second length value, and d3 is the third length value.
[0090] As Figure 3 shown, for the original BEV feature grid composed of 128 * 128 squares with a side length of 0.8m, this embodiment provides an excellent division example as follows: A plane rectangular coordinate system is established with the center of the grid plane as the origin. Within |x| ≤ 9.6, it is divided with each grid side length being 0.6m; within 9.6 < |x| ≤ 35.2, it is divided with each grid side length being 0.8m; within 35.2 < |x| ≤ 51.2, it is divided with each grid side length being 1m. The y-direction is divided similarly, and a total of 3 * 3 = 9 different-sized grids are formed. The mixed-resolution grid does not increase the total number of grids, so it does not affect the detection speed, and the local grid design with smaller grids closer and larger grids farther adapts to the actual detection requirements.
[0091] Step 5: Input the new BEV feature into the BEV feature encoding network, and then obtain the three-dimensional object detection result of the mobile device environment through the 3D detection head.
[0092] Input the new BEV feature F B ′ EV into the BEV feature encoding network, and then through the 3D detection head, the 3D detection head obtains the three-dimensional object detection result of the target in the three-dimensional space. The three-dimensional object detection result includes: the category prediction value and the three-dimensional detection box prediction value in the three-dimensional space centered on itself. Among them, the three-dimensional detection box prediction value includes the three-dimensional coordinates of the center of the detection box, the length, width, and height of the detection box, and the yaw angle of the detection box.
[0093] Step 6: Perform detection supervision on the three-dimensional object detection result, and update the image backbone network, depth prediction network, semantic prediction network, BEV feature encoding network, and 3D detection head by combining multi-level depth supervision and semantic supervision.
[0094] Detection supervision includes category prediction supervision and 3D detection box prediction supervision, specifically: obtaining the category label and 3D detection box label marked by the multi-view image of the mobile device corresponding to the 3D target detection result; calculating the Gaussian focal loss (GaussianFocalLoss) according to the category label and the category prediction value in the 3D target detection result as the category prediction loss; calculating the L1 loss (L1Loss) according to the 3D detection box label and the 3D detection box prediction value in the 3D target detection result as the 3D detection box prediction loss.
[0095] The total loss L of the image's two-dimensional features Z The final detection loss is added to the Gaussian focus loss and L1 loss, and according to the final detection loss, the gradient descent method is used to update the parameters of the image backbone network, depth prediction network, semantic prediction network, BEV feature encoding network and 3D detection head until the training is completed.
[0096] Embodiment 2: A three-dimensional target detection method for real-time perception in a mobile environment is applied in the inference stage, such as Figure 4 As shown, the following steps are included:
[0097] Step 1: Real-time acquisition of new mobile device multi-view image input, and the trained image backbone network extracts the image two-dimensional features to be perceived, denoted as F img ∈R N×C×H×W , where N represents the number of multi-view images of mobile devices in the training set, C represents the number of feature channels of the two-dimensional features of the image, and H and W represent the size of the feature map of the two-dimensional features of the image. In this embodiment, the backbone network adopts the ResNet50 neural network.
[0098] Step 2: Input the two-dimensional features of the image to be perceived into the trained depth prediction network and semantic prediction network respectively to obtain the depth prediction value and the semantic prediction value.
[0099] Step 3: Project the two-dimensional features of the image to be sensed into three-dimensional space to obtain the BEV features to be sensed.
[0100] Step 4: Use a mixed-resolution grid to perform bilinear interpolation sampling on the BEV features to be sensed to form new BEV features to be sensed.
[0101] Step 5: Input the new BEV features to be perceived into the trained BEV feature encoding network, and then obtain the three-dimensional target detection results of the mobile device environment through the trained 3D detection head, and further obtain the bird's-eye view detection results based on the three-dimensional target detection results.
[0102] The new BEV feature F B ′EV Input the BEV feature encoding network, and then pass it through the 3D detection head. The 3D detection head obtains the 3D target detection results of the target in the 3D space. The 3D target detection results include: category prediction value and 3D detection box prediction value in the 3D space centered on itself, where the 3D detection box prediction value includes the 3D coordinates of the center of the detection box, the length, width, height, and yaw angle of the detection box. The 8 corner points of the 3D detection box can be obtained, and the spatial coordinates are transformed by combining the camera external parameters and internal parameters. The 8 corner points in the 3D space are projected onto the 2D plane of the image to obtain the 3D target detection results on the image. In addition, the 3D space is flattened by removing the height dimension to obtain a bird's-eye view of the plane.
[0103] This embodiment provides a method for detecting a vehicle as a mobile device. Figure 5 As shown, Figure 5 The left side of the figure shows the 3D target detection results, including the 3D detection frames of each target. Different colors represent different categories of targets, such as Figure 5 The middle right side is a bird's-eye view of the detection effect.
[0104] For the specific limitations of the three-dimensional target detection method for real-time perception in a mobile environment in this embodiment, please refer to the limitations of the three-dimensional target detection method for real-time perception in a mobile environment in Example 1, which will not be repeated in this embodiment.
[0105] It is easy to understand that for ease of understanding, the present invention is divided into a training phase and an inference phase to describe a three-dimensional target detection method for real-time perception in a mobile environment. However, in actual applications, the training phase and the inference phase can exist separately or simultaneously, that is, Example 1 and Example 2 can be implemented separately, or Example 1 and Example 2 can be combined and implemented as a new overall embodiment.
[0106] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0107] The above-mentioned embodiments only express several implementation modes of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.
Claims
1. A three-dimensional target detection method for real-time perception in a mobile environment, characterized in that: The three-dimensional target detection method for real-time perception in a mobile environment includes: Input the multi-view images of the mobile device in the training set into the image backbone network to extract the two-dimensional features of the image; The two-dimensional features of the image are input into the depth prediction network and the semantic prediction network respectively to obtain the depth prediction value and the semantic prediction value, and multi-level depth supervision is performed on the depth prediction value, and semantic supervision is performed on the semantic prediction value; Project the two-dimensional features of the image into three-dimensional space to obtain BEV features; The mixed-resolution grid is used to perform bilinear interpolation sampling on the BEV features to form new BEV features; The new BEV features are input into the BEV feature encoding network, and then the 3D object detection results of the mobile device environment are obtained through the 3D detection head; Perform detection supervision on the 3D target detection results, and combine multi-level depth supervision and semantic supervision to update the image backbone network, depth prediction network, semantic prediction network, BEV feature encoding network and 3D detection head; New multi-view images of mobile devices are collected in real time and input into the trained image backbone network, depth prediction network, semantic prediction network, BEV feature encoding network and 3D detection head to obtain the corresponding three-dimensional target detection results, and the bird's-eye view detection results are obtained based on the three-dimensional target detection results.
2. The three-dimensional target detection method for real-time perception in mobile environment according to claim 1 is characterized in that: The multi-level depth supervision of the depth prediction value includes: (1) Semantically weighted point-level absolute depth supervision: The depth truth value and semantic truth value of the multi-view image of the mobile device are one-hot encoded and downsampled to the size of the image’s two-dimensional features to obtain the depth label value D lable and semantic label value S lable , then based on the depth prediction value D pred and the depth tag value D lable , calculate the Binary Cross Entropy loss as follows: Where L1 represents the semantically weighted point-level absolute depth loss of multi-view images on mobile devices, Indicates that the input is the depth prediction value D pred and the depth tag value D lable Binary CrossEntropy loss, ω s is the foreground background weight factor, ω1 is the foreground weight factor, ω2 is the background weight factor, S lable ≠0 means the pixel belongs to the foreground, S lable =0 means the pixel belongs to the background; (2) Relative depth supervision at the target internal structure level: The pixel with the smallest absolute value of the internal depth prediction error of each target in the multi-view image of the mobile device is taken as the internal structure reference point θ of the target. The depth prediction value and depth label value of the internal structure reference point θ are recorded as and The absolute value of the depth prediction error is the absolute value of the difference between the depth prediction value and the depth label value. The depth prediction difference and the depth label difference of other pixels inside each target relative to the internal structure reference point θ are calculated, and the MSE loss of the two is calculated. The formula is as follows: Where L2 represents the relative depth loss of the target internal structure level of the multi-view image of the mobile device, Indicates that the input is a depth prediction difference Difference between the depth label and MSE loss; (3) Relative depth supervision at the target spatial distribution level: In each multi-view image of a mobile device, the point with the smallest depth label value inside each target is taken as the representative point α of the target. The depth prediction value and depth label value of the representative point α are recorded as and Calculate the depth prediction value and the depth tag value The difference between the two is taken, and the representative point with the smallest absolute difference in each mobile device multi-view image is taken as the target spatial distribution reference point β of the mobile device multi-view image where the representative point is located. The depth prediction value and depth label value of the spatial distribution reference point β are recorded as and Then calculate the depth prediction difference and depth label difference of other representative points α in each mobile device multi-view image relative to the spatial distribution reference point β of the mobile device multi-view image, and calculate the MSE loss of the two. The formula is as follows: Where L3 represents the relative depth loss of the target spatial distribution level of multi-view images of mobile devices, Indicates that the input is a depth prediction difference Difference between the depth label and MSE loss; (4) The total loss of deep supervision is calculated as follows: L Depth =ε1*L1+ε2*L2+ε3*L3 Where, L Depth represents the total loss of deep supervision of multi-view images on mobile devices, ε1 represents the first-level loss coefficient, ε2 represents the second-level loss coefficient, and ε3 represents the third-level loss coefficient.
3. The three-dimensional target detection method for real-time perception in mobile environment according to claim 1 is characterized in that: The semantic supervision of the semantic prediction value includes: The semantic truth value of the multi-view image of the mobile device is one-hot encoded and downsampled to the size of the two-dimensional feature of the image to obtain the semantic label value S lable , based on the semantic prediction value S pred and semantic label value S lable , calculate the Binary CrossEntropy loss as follows: Where, L Seg represents semantic loss, ε4 represents semantic loss coefficient, Indicates that the input is a semantic prediction value S pred and semantic label value S lable Binary Cross Entropy loss.
4. The three-dimensional target detection method for real-time perception in mobile environment according to claim 1, characterized in that: The projecting of the two-dimensional features of the image into the three-dimensional space to obtain the BEV features includes: Perform outer product of the image two-dimensional features and the depth prediction value to obtain the stretched image two-dimensional features; The features whose depth prediction value is less than the depth threshold or whose semantic prediction value is less than the semantic threshold in the two-dimensional features of the stretched image are filtered; The remaining features after filtering in the stretched two-dimensional features of the image are projected into the three-dimensional space to obtain the BEV features.
5. The three-dimensional target detection method for real-time perception in mobile environment according to claim 1, characterized in that: The method of using a mixed resolution grid to perform bilinear interpolation sampling on the BEV feature to form a new BEV feature includes: A plane rectangular coordinate system is established with the center of the grid plane of the BEV feature space as the origin; The division for the x-axis is as follows: in the range of |x|≤x1, the division is based on the side length of each grid being d1; in the range of x1<|x|≤x2, the division is based on the side length of each grid being d2; in the range of x2<|x|≤x3, the division is based on the side length of each grid being d3, where x1 <x2<x3,d1<d2<d3; The division for the y-axis is as follows: in the range of |y|≤y1, the division is based on the side length of each grid being d1; in the range of y1<|y|≤y2, the division is based on the side length of each grid being d2; in the range of y2<|y|≤y3, the division is based on the side length of each grid being d3, where x1=y1, x2=y2, x3=y3; Based on the new grid obtained after division, bilinear interpolation sampling is performed on the BEV features to obtain new BEV features.
6. The three-dimensional target detection method for real-time perception in mobile environment according to claim 1, characterized in that: The detection and supervision of the three-dimensional target detection result includes: Obtaining the category label and the 3D detection frame label marked by the multi-view image of the mobile device corresponding to the 3D target detection result; Calculate the Gaussian focus loss as the category prediction loss based on the category label and the category prediction value in the 3D object detection result; The L1 loss is calculated based on the 3D detection box label and the 3D detection box prediction value in the 3D object detection result as the 3D detection box prediction loss.
7. The three-dimensional target detection method for real-time perception in mobile environment according to claim 1, characterized in that: The real-time acquisition of new multi-view images of mobile devices is input into the trained image backbone network, depth prediction network, semantic prediction network, BEV feature encoding network and 3D detection head to obtain corresponding three-dimensional target detection results, including: Collect new multi-view images from mobile devices and input them into the trained image backbone network to extract the two-dimensional features of the image to be perceived; Input the two-dimensional features of the image to be perceived into the trained depth prediction network and semantic prediction network respectively to obtain a depth prediction value and a semantic prediction value; Projecting the two-dimensional features of the image to be sensed into the three-dimensional space to obtain the BEV features to be sensed; A mixed-resolution grid is used to perform bilinear interpolation sampling on the BEV features to be sensed to form new BEV features to be sensed; The new BEV features to be perceived are input into the trained BEV feature encoding network, and then the three-dimensional target detection results of the mobile device environment are obtained through the trained 3D detection head.
Citation Information
Patent Citations
Weak supervision semantic segmentation method and application thereof
CN111462163A
Monocular BEV perception method through self-supervised depth estimation
CN117911484A
Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment
WO2024230038A1