Obstacle three-dimensional space occupation analysis method based on look-around image perception

Through the surround-view image perception technology, the surround-view camera collects multiple image data for multi-level preprocessing and feature extraction, and combined with the three-dimensional bilinear sampling technology, the three-dimensional occupancy grid points are generated, which solves the problems of high cost, high complexity and poor environmental adaptability of the existing environmental perception system, and achieves efficient and accurate obstacle identification and positioning.

CN120451941APending Publication Date: 2025-08-08ANHUI JIANGHUAI AUTOMOBILE GRP CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510621170.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing environmental perception systems are costly, have high complexity, poor environmental adaptability, and have limited field of view, large depth estimation errors, and weak dynamic object handling capabilities, making it difficult to achieve efficient and low-cost three-dimensional obstacle identification and positioning.

Method used

The three-dimensional spatial occupation analysis method of obstacles based on surround-view image perception is adopted, and multiple image data are collected through surround-view cameras, and multi-level preprocessing and feature extraction are performed. The three-dimensional bilinear sampling technology is used to fuse semantics and position information to generate 3-dimensional occupancy grid points to realize the distinction and prediction of dynamic and static obstacles.

Benefits of technology

It realizes low-cost and efficient three-dimensional environment perception, which can accurately distinguish and predict dynamic and static obstacles in real time, and improves the robustness and real-time detection performance of the perception system in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451941A_ABST
    Figure CN120451941A_ABST
Patent Text Reader

Abstract

The invention discloses an obstacle three-dimensional space occupation analysis method based on all-round view image perception, which comprises the following steps of: acquiring multi-path image data from an all-round view camera, and converting all-round view data into an image feature map through a multi-stage image preprocessing and encoding link; then, a three-dimensional bilinear sampling technology is utilized to add spatial dimension information to the feature maps, so that deep fusion and interaction of semantic and position information among different scales are achieved, multiple feature maps of different sizes are generated, on one hand, the feature maps are used for multi-scale supervised learning in the training stage, and on the other hand, the feature maps are used for multi-scale supervised learning; on the other hand, the required regression quantity is output through a specific detection mechanism in the reasoning stage, and therefore an automatic driving environment sensing scheme which is economical, efficient and excellent in performance is achieved. According to the invention, high-precision three-dimensional environment perception is realized only by means of the look-around camera system, dynamic and static obstacles in the environment can be effectively distinguished and predicted, the real-time detection performance is ensured, and meanwhile, the robustness of the perception system in a complex environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a method for analyzing the three-dimensional space occupancy of obstacles based on surround view image perception. Background Art

[0002] With the rapid development of autonomous driving technology, environmental perception, as one of its core technologies, has become crucial for driving safety. The capabilities of the environmental perception system directly determine the vehicle's understanding of its surroundings, which in turn influences driving strategies and overall safety. Currently, mainstream environmental perception technologies typically use high-precision sensors such as lidar and millimeter-wave radar to acquire three-dimensional information about the surrounding environment. However, the high cost of these sensors has severely hindered the widespread adoption and commercialization of autonomous driving technology.

[0003] In contrast, visual perception technology has attracted considerable attention in recent years due to its advantages, including low cost, rich information, and easy installation. Surround-view camera systems, in particular, offer a low-cost, yet information-rich, sensing solution that provides comprehensive environmental information, opening up new possibilities for autonomous driving systems. Surround-view cameras capture high-resolution images of the vehicle's surroundings, and using deep learning algorithms to mine these images for environmental information and achieve efficient and accurate semantic representation of the scene, making them a promising method for sensing the occupancy of both static and dynamic obstacles.

[0004] Currently, mainstream environmental perception systems rely primarily on LiDAR, supplemented by multi-sensor fusion solutions using cameras. Specifically, LiDAR acquires high-precision three-dimensional point cloud data, which is combined with image information captured by the camera to detect and locate obstacles. LiDAR provides precise depth information, while the camera compensates for the point cloud data's shortcomings in texture, color, and semantic information, thereby improving the overall performance of the perception system. However, these systems suffer from the following drawbacks:

[0005] (1) High cost: LiDAR is expensive, which significantly increases the overall cost of the vehicle and limits its application in mass-produced autonomous vehicles.

[0006] (2) High complexity: Multi-sensor fusion technology has extremely high requirements for hardware time synchronization and spatial calibration, which increases the complexity of the system and engineering difficulty.

[0007] (3) Poor environmental adaptability: Especially in bad weather such as rain and snow, the performance of lidar decreases significantly, which may lead to unreliable detection results.

[0008] In addition, the industry currently also has visual perception technologies based on monocular or binocular cameras. These methods utilize deep learning algorithms to process images, including semantic segmentation, depth estimation, and object detection, to identify and locate obstacles. Binocular vision obtains depth information through parallax calculations, while monocular vision relies on learning algorithms to estimate the three-dimensional structure of the scene. However, these methods also have many shortcomings:

[0009] (1) Limited field of view: The perception range of a single camera or dual cameras usually has blind spots and cannot cover all-round information around the vehicle.

[0010] (2) Large depth estimation error: The depth estimation of a monocular camera has scale uncertainty, and the error increases significantly, especially in long-distance scenes.

[0011] (3) Weak dynamic object processing capabilities: Traditional visual methods have limitations in detecting and predicting dynamic objects, making it difficult to accurately determine their occupancy status. Summary of the Invention

[0012] In view of the above, the present invention aims to provide a method for analyzing the three-dimensional space occupancy of obstacles based on surround image perception to solve the technical problems mentioned above.

[0013] The technical solution adopted in the present invention is as follows:

[0014] The present invention provides a method for analyzing the three-dimensional space occupancy of obstacles based on surround image perception, which includes:

[0015] Only the surround view image data of the obstacle is collected, and the surround view image data includes obstacle pictures from different perspectives collected by multiple cameras;

[0016] performing multi-stage preprocessing on the surround view image data;

[0017] Extracting image features of the pre-processed surround view image data, including: adding spatial dimension information to the pre-extracted image feature map through three-dimensional bilinear sampling;

[0018] Integrating the image features based on different levels to obtain fused interactive information of semantics and positions at different scales;

[0019] Based on the feature information fused between different layers, three-dimensional occupied grid points are generated to determine the location and category of the occupied points representing obstacles.

[0020] In at least one possible implementation manner, timestamps of images captured by multiple cameras at different viewing angles are aligned.

[0021] In at least one possible implementation, the multi-stage preprocessing includes at least:

[0022] The original images collected by multiple cameras are cropped and then enhanced. The multi-camera images are spliced along the channel dimension to form a single data structure containing all perspective information; and real value encoding is performed in combination with synchronized timestamps.

[0023] In at least one possible implementation manner, extracting image features from the preprocessed surround view image data specifically includes:

[0024] Perform feature pre-extraction on the surround view image data to obtain an image feature map;

[0025] Construct several voxel center points in 3D space and project the voxel center points to the size range of the image feature map;

[0026] Normalize the coordinates of the projected voxel center point to obtain a normalized three-dimensional projection point;

[0027] Combined with the 3D bilinear sampling method, the image features collected by cameras with different viewpoints are sampled onto a spatial feature map that matches the predetermined spatial grid size, and 3D voxel feature maps of different scales are obtained.

[0028] In at least one possible implementation manner, the integrating the image features based on different levels includes:

[0029] Perform convolution processing on different 3D voxel feature maps to obtain several category feature maps and several regression feature maps;

[0030] After the category feature map and the regression feature map are reshaped, they are matched with the corresponding real value encoding processing results, and the original feature maps for classification tasks and regression tasks are extracted and output.

[0031] In at least one possible implementation, generating a three-dimensional occupied grid of points to determine positions and categories of occupied points representing obstacles includes:

[0032] According to a preset category threshold, only the first original feature map output is filtered to obtain a confidence feature map and a corresponding mask, wherein the confidence feature map includes the filtered target object data points;

[0033] Convert the target object data points into discrete data points with real-world coordinate information and category labels to obtain the obstacle type prediction results;

[0034] Combined with the mask, the speed of the target object is calculated to obtain the speed prediction result of the obstacle;

[0035] After combining the type prediction results with the speed prediction results, they are converted into occupancy states in real three-dimensional space, and the probability of the occupied points representing the target obstacle in the three-dimensional occupied grid points is output.

[0036] Compared with the existing technology, the main design concept of the present invention is to first collect multi-channel image data from the surround-view camera and convert the surround-view data into an image feature map through multi-level image preprocessing and encoding. Then, 3D bilinear sampling technology is used to add spatial dimension information to the feature map to achieve deep fusion and interaction of semantic and position information between different scales, generating multiple feature maps of different sizes. On the one hand, it can be used for multi-scale supervised learning in the training phase, and on the other hand, a specific detection mechanism can be used to output the required regression quantity in the inference phase, thereby realizing a cost-effective and high-performance autonomous driving environment perception solution. The present invention relies solely on the surround-view camera system to achieve high-precision three-dimensional environmental perception, can effectively distinguish and predict dynamic and static obstacles in the environment, ensure the performance of real-time detection, and at the same time improve the robustness of the perception system in complex environments.

[0037] This paper demonstrates that by combining pure visual methods with deep learning technology, it achieves high-precision environmental perception, providing a low-cost, high-performance solution for autonomous driving systems. Compared to existing technologies (such as lidar solutions), this paper offers significant advantages in cost-effectiveness, real-time performance, and system reliability, and is of great significance in promoting the commercial application of autonomous driving technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described below with reference to the accompanying drawings, in which:

[0039] Figure 1 A flowchart of a method for analyzing the three-dimensional space occupancy of obstacles based on surround image perception provided by an embodiment of the present invention;

[0040] Figure 2 A schematic diagram of a network architecture for three-dimensional space occupancy of obstacles based on surround view image perception provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0042] The present invention proposes an embodiment of a method for analyzing the three-dimensional space occupation of obstacles based on surround image perception. Specifically, Figure 1shown, including:

[0043] Step S1: Only surround view image data is collected, and the surround view image data includes pictures from different perspectives collected by multiple cameras (it can be understood that during implementation, obstacle images in different states such as dynamic and static can be collected when the vehicle is running or parked).

[0044] Surround view images are collected by multiple wide-angle cameras arranged around the vehicle. Before inputting these image data into the model, it is necessary to ensure that the timestamps of all image data are synchronized. For camera image data with the same timestamp t, it can be expressed as ,in is the number of pictures from different perspectives.

[0045] Step S2: performing multi-stage preprocessing on the surround view image data.

[0046] At this stage, a series of pre-processing steps may be performed on the collected image data. For example, in some preferred embodiments of the present invention, these pre-processing steps include at least image data enhancement and true value encoding.

[0047] 1. Preprocessing of input image data

[0048] (1) Multi-channel camera images To crop the image, a fixed area of the original image is selected as the input of the model image branch. The overall cropping process can be described as follows:

[0049]

[0050] After cropping the image Perform data augmentation operations, such as randomly adjusting brightness and contrast and introducing noise, which can enhance the model's robustness to different lighting conditions and noise interference. This process can be formalized as an augmentation function, denoted as:

[0051]

[0052] (2) For multi-channel image data, a strategy of splicing along the channel dimension is adopted to construct a multi-channel input data structure. Specifically, the image data from different cameras are merged along the channel dimension to form a single tensor data representation containing all perspective information. This operation can be expressed as:

[0053]

[0054]

[0055] 2. Preprocessing of ground truth data

[0056] Combined with the previous article, we need to use the true value of the occ scattered points with the same timestamp t The ground truth box associated with this frame and available for annotation .in, , , cls represents the category integer value; , The following steps will be used to calculate the loss for training and the test metrics.

[0057] (1) Similar to the Crop operation of image cropping, the scatter points of OCC are also The true value range is restricted to adapt to the division of a specific grid size. The process can be expressed as follows:

[0058]

[0059] (2) The real value points after clipping Perform voxelization and convert it into a dense tensor representation:

[0060]

[0061] (3) In the true value encoding stage, the occ Gaussian heat map of the target is constructed by generating a Gaussian heat map for subsequent loss function calculation and indicator testing. This process involves the range-limited Point Cloud Range (PCR), grid size, and annotation box. The center coordinate (x, y) information and the two-dimensional Gaussian function GS. The occ Gaussian heat map of the target distribution is generated for the subsequent loss calculation. The overall data transformation process can be simplified as follows:

[0062]

[0063] Step S3: extracting image features of the pre-processed surround view image data, including: adding spatial dimension information to the pre-extracted image feature map through three-dimensional bilinear sampling.

[0064] After data preprocessing, key features can be extracted from the image, which can capture important properties such as color, texture, and shape of the object in the image tensor.

[0065] 1. Image feature pre-extraction. After completing the multi-channel image preprocessing described above, the input data is passed through the first four modules of the ResNet50 network. This process results in a 16-fold downsampling of feature map C4. Subsequently, C4 is upsampled by a factor of 2 and subjected to a double-layer convolution with a 1×1 convolution kernel to integrate information from the channel dimension. Finally, the processed C4 is fused with C3, which has been processed with a single convolution kernel, to obtain the 8-fold downsampled image feature map I, which is used for subsequent analysis and processing.

[0066] 2. Fusion of depth information. First, construct (Generate3DGrid) the voxel center points (3DGrids) in the 3D space, and then use the internal and external parameters to project (3DTo2D) these center points to the size range of the cropped feature map. Then, through the coordinate normalization process (NormalizeGridsTo2D), the normalized three-dimensional projection points 3DGridsInImg are obtained. After that, the 3D bilinear sampling (3DGridSample) method is used to sample each image feature map onto a spatial feature map that matches the predetermined spatial grid size to obtain a 3D voxel feature map, thereby integrating the image feature map I obtained in the pre-extraction link into the three-dimensional spatial information. This process can be expressed as:

[0067] 3DGrids = Generate3DGrid(PCR,Size)

[0068] 3DGridInImg = NormalizeGridsTo2D(3DTo2D(3DGrids))

[0069] = 3DGridSample(I, 3DGridInImg)

[0070] Subsequently, the 3D voxel feature map F I,i After dimension reshaping, it is compressed into the bird's-eye view (BEV) space, and the feature maps from cameras with different perspectives are merged to finally obtain a BEV feature map containing comprehensive perspective information, which will not be described in detail here.

[0071] Step S4: integrating the image features based on different levels to obtain fused interactive information of semantics and positions at different scales;

[0072] Specifically, in this step, the extracted image features are integrated through image feature hierarchy technology to obtain richer three-dimensional spatial occupancy information. In practice, the fusion process can be implemented through various strategies, including weighted averaging, feature concatenation, and feature splicing, which help improve the accuracy of perception tasks. For example, a method similar to FPN can be used to enable feature maps of different scales to interact with each other, which can be expressed as:

[0073] = ~FPN ( )

[0074] The main purpose of this link in the training phase is to extract and utilize a total of 7 feature maps of four different scales in parallel (for example and not limitation). After these feature maps are interactively processed, they can be used as the input of the OCC detection head (the OCC detection head consists of two consecutive 1x1 convolutional layers with a ReLU activation function in between, acting on the convolutional feature maps to extract the original output features for classification and regression tasks). The output set of these feature maps can be expressed as ;

[0075] The main purpose of this link in the inference stage is to derive the OCC feature map for dynamic and static obstacle target detection. ,It can be understood that this process is similar to the information interaction between different feature layers in the Feature Pyramid Network (FPN), which effectively enhances the model's ability to detect targets of different scales.

[0076] Step S5: Based on the feature information fused between different layers, the environment is perceived and three-dimensional occupied grid points are generated to determine the position and category of the occupied points representing obstacles.

[0077] Combine Figure 2 From the schematic network architecture diagram, it can be seen that the aforementioned steps S1 to S2 can be understood as the input in the model architecture, and steps S3 to S5 can be understood as the processing and output in the model architecture.

[0078] Therefore, those skilled in the art can understand that the model can be divided into the training phase and the inference phase. In particular, the aforementioned fusion feature information plays a key role in training and inference. The following is a detailed explanation with examples:

[0079] 1. Training phase

[0080] 7 3D voxel feature maps exported from Each detection head is connected to 7 specialized loss detection heads, which are responsible for multi-level supervised learning to calculate the loss. Each detection head is processed by convolution and outputs the feature maps of 5 categories arranged in channel order. (where the fifth category represents the background), and the regression feature maps related to the two velocities (the target point along the xy axis) This design enables the model to perform fine-grained analysis of features at different scales, thereby improving detection accuracy and robustness at multiple levels;

[0081] Next, in the embodiment of the present invention, the feature map and Split on the channel dimension C and height dimension D and reshape into a five-dimensional tensor, which can be uniformly represented as a tensor set This tensor will then be compared to the OCC true value set of different sizes after preprocessing with the same size as the index. Matching is performed and loss is calculated based on these matching pairs. This approach allows the model to accurately evaluate and optimize the performance of OCC detection at different scales.

[0082] In addition, for the two different tasks of classification and regression, the embodiment of the present invention adopts a targeted loss calculation strategy. In the classification task, the embodiment of the present invention adopts Focal Loss, which is a specially designed loss function that optimizes the model performance by reducing the weight of easy-to-classify samples and increasing the weight of difficult-to-classify samples, so that the classification loss can be expressed as In the regression task, the L1 loss function is used for regression supervision. This loss function is simple and has good robustness to outliers. The regression loss can be expressed as .

[0083] Therefore, in the multi-layer supervision loss of OCC, the default total loss is The loss distribution ratio is , and finally update the weights of the entire model through gradient backpropagation.

[0084] 2. OCC reasoning stage

[0085] In the inference phase, only the first feature map extracted is used for OCC prediction. , the specific steps are as follows:

[0086] In the above OCC training task, we can see that After convolution and dimension reshaping, a feature map containing 5 categories (including background) is obtained. And the regression feature map of 2 speeds ;

[0087] Based on the feature information after fusion between the above different layers, in the process of reasoning, the probability output can be obtained by first passing through the Sigmoid activation function, and the confidence can be filtered by setting the category threshold to 0.1 to obtain the filtered confidence feature map and the corresponding mask , combined with the previously defined parameters PCR, Size, and channel order, the filtered points are converted into discrete points with real-world coordinates and category labels ; For obstacle speed prediction, it is necessary to combine To obtain the target speed It should be pointed out here that the consideration of category and speed can effectively cover different obstacle targets in both static and dynamic real states.

[0088] Therefore, the prediction results of obstacle category and speed are combined, and the three-dimensional space decoder is used to convert these features into the occupancy state in the real three-dimensional space, and the occupancy point probability of dynamic and static obstacles in the three-dimensional occupied grid points is output, which can be specifically recorded as .

[0089] In summary, the main design concept of the present invention is to first collect multi-channel image data from the surround-view camera and convert the surround-view data into image feature maps through multi-level image preprocessing and encoding. Then, 3D bilinear sampling technology is used to add spatial dimension information to the feature map to achieve deep fusion and interaction of semantic and position information between different scales, generating multiple feature maps of different sizes. On the one hand, it can be used for multi-scale supervised learning in the training phase, and on the other hand, a specific detection mechanism can be used to output the required regression quantity in the inference phase, thereby realizing a cost-effective and high-performance autonomous driving environment perception solution. The present invention relies solely on the surround-view camera system to achieve high-precision three-dimensional environmental perception, can effectively distinguish and predict dynamic and static obstacles in the environment, ensure the performance of real-time detection, and at the same time improve the robustness of the perception system in complex environments.

[0090] Compared to the existing technologies mentioned above, (1) the present invention has a better latency advantage in real-time model reasoning, can achieve millisecond-level processing speed, and can handle complex environments in real time; (2) it accurately perceives the distribution and dynamic and static properties of obstacles in three-dimensional space, and provides accurate three-dimensional space occupancy information. (3) Adaptability advantage: it has strong adaptability to changes in lighting and maintains stable performance in complex scenes. The present invention can simply and effectively implement scene occupancy analysis, which is helpful for understanding the entire traffic environment, providing more accurate real-world obstacle conditions, and can better support subsequent downstream tasks, such as planning safe paths.

[0091] If the expressions expressing directions are mentioned in the embodiments of the present invention, they are relative concepts based on the embodiments. In addition, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c or a, b and c, where a, b, c can be single or multiple.

[0092] The above describes in detail the structure, features and effects of the present invention based on the embodiments shown in the drawings, but the above is only a preferred embodiment of the present invention. It should be noted that the technical features involved in the above embodiments and their preferred modes can be reasonably combined and matched into a variety of equivalent schemes by those skilled in the art without departing from or changing the design ideas and technical effects of the present invention; therefore, the scope of implementation of the present invention is not limited to what is shown in the drawings. Any changes made in accordance with the concept of the present invention, or modifications to equivalent embodiments with equivalent changes, which still do not exceed the spirit covered by the description and drawings, should be within the scope of protection of the present invention.

Claims

1. A method for analyzing the three-dimensional space occupancy of obstacles based on surround image perception, characterized in that: include: Only the surround view image data of the obstacle is collected, and the surround view image data includes obstacle pictures from different perspectives collected by multiple cameras; performing multi-stage preprocessing on the surround view image data; Extracting image features of the pre-processed surround view image data, including: adding spatial dimension information to the pre-extracted image feature map through three-dimensional bilinear sampling; Integrating the image features based on different levels to obtain fused interactive information of semantics and positions at different scales; Based on the feature information fused between different layers, three-dimensional occupied grid points are generated to determine the location and category of the occupied points representing obstacles.

2. The method for analyzing obstacle 3D space occupancy based on surround image perception according to claim 1, characterized in that: Align the timestamps of images captured by multiple cameras at different viewing angles.

3. The method for analyzing obstacle 3D space occupancy based on surround image perception according to claim 2, characterized in that: The multi-stage pretreatment includes at least: The original images collected by multiple cameras are cropped and then enhanced. The multi-camera images are spliced along the channel dimension to form a single data structure containing all perspective information; and real value encoding is performed in combination with synchronized timestamps.

4. The method for analyzing obstacle 3D space occupancy based on surround image perception according to claim 3, characterized in that: The image features of the pre-processed surround view image data are specifically extracted as follows: Perform feature pre-extraction on the surround view image data to obtain an image feature map; Construct several voxel center points in 3D space and project the voxel center points to the size range of the image feature map; Normalize the coordinates of the projected voxel center point to obtain a normalized three-dimensional projection point; Combined with the 3D bilinear sampling method, the image features collected by cameras with different viewpoints are sampled onto a spatial feature map that matches the predetermined spatial grid size, and 3D voxel feature maps of different scales are obtained.

5. The method for analyzing obstacle 3D space occupancy based on surround image perception according to claim 4, characterized in that: The integrating the image features based on different levels includes: Perform convolution processing on different 3D voxel feature maps to obtain several category feature maps and several regression feature maps; After the category feature map and the regression feature map are reshaped, they are matched with the corresponding real value encoding processing results, and the original feature maps for classification tasks and regression tasks are extracted and output.

6. The method for analyzing obstacle 3D space occupancy based on surround image perception according to any one of claims 1 to 5, characterized in that: Generating a three-dimensional occupied grid of points to determine the positions and categories of occupied points representing obstacles includes: According to a preset category threshold, only the first original feature map output is filtered to obtain a confidence feature map and a corresponding mask, wherein the confidence feature map includes the filtered target object data points; Convert the target object data points into discrete data points with real-world coordinate information and category labels to obtain the obstacle type prediction results; Combined with the mask, the speed of the target object is calculated to obtain the speed prediction result of the obstacle; After combining the type prediction results with the speed prediction results, they are converted into occupancy states in real three-dimensional space, and the probability of the occupied points representing the target obstacle in the three-dimensional occupied grid points is output.

Citation Information

Cited By

  • Unmanned aerial vehicle multi-dimensional environment perception obstacle avoidance method and system

    CN121433284A