Grid occupation prediction method, device, equipment, medium and product
By adjusting occupany grid sizes based on distance, the method balances efficiency and accuracy in environmental modeling, ensuring precise detection of nearby obstacles with reduced computational load.
Patent Information
- Application Number
- CN202510487671.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art is difficult to take into account both the prediction efficiency and accuracy requirements in the prediction of occupancy grid prediction, and fixed grid size leads to wasted computing resources or the prediction accuracy is reduced.
The multi-resolution occupancy grid prediction method is used to dynamically adjust the grid size of different spatial regions, use small grids to ensure high accuracy in the close-range areas, use large grids to reduce computing resource consumption in the long-range areas, and use feature pyramid networks and multi-occupancy prediction units to generate multi-scale occupancy grids.
With limited computing resources, the efficiency and accuracy balance of grid prediction is improved, the perception needs of different distance areas are adapted to ensure the reliable operation of the autonomous driving system.
Smart Images

Figure CN120318459A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and particularly relates to an occupancy grid prediction method, device, equipment, medium and product. Background Art
[0002] With the development of autonomous driving technology, the environmental perception system needs to understand the 3D scene around the vehicle in real time and accurately, including static obstacles such as curbs and buildings, and dynamic objects such as vehicles and pedestrians. Among them, the Occupancy Grid (OCC) has gradually become the core means of environmental three-dimensional modeling because it can represent obstacles of any shape.
[0003] However, when predicting the occupancy grid of the surrounding environment, it is difficult to balance the requirements of prediction efficiency and accuracy. Summary of the Invention
[0004] To overcome the problems existing in the related art, the present application provides an occupancy grid prediction method, device, equipment, medium and product.
[0005] According to the first aspect of any embodiment of the present application, an occupancy grid prediction method is provided. The method includes:
[0006] Obtain the environmental data of the vehicle's environment, and generate the BEV feature in the vehicle's environment according to the environmental data;
[0007] Based on the BEV feature, use different occupancy prediction units respectively to generate the prediction results of the occupancy grids in different spatial regions in the vehicle's environment. The prediction results are used to represent the occupancy information of the occupancy grids; the grid sizes of the occupancy grids in each spatial region are the same, and the grid sizes of the occupancy grids in different spatial regions increase with the increase of the distance from the vehicle.
[0008] According to the second aspect of any embodiment of the present application, an occupancy grid prediction device is provided. The device includes:
[0009] A generation module, configured to obtain the environmental data of the vehicle's environment, and generate the BEV feature in the vehicle's environment according to the environmental data;
[0010] A prediction module, configured to use different occupancy prediction units respectively based on the BEV feature to generate the prediction results of the occupancy grids in different spatial regions in the vehicle's environment. The prediction results are used to represent the occupancy information of the occupancy grids; the grid sizes of the occupancy grids in each spatial region are the same, and the grid sizes of the occupancy grids in different spatial regions increase with the increase of the distance from the vehicle.
[0011] According to a third aspect of any embodiment of the present application, there is provided an electronic device, including:
[0012] a processor;
[0013] a memory for storing instructions executable by the processor;
[0014] wherein, the processor realizes the method described in any embodiment of the present application by running the executable instructions.
[0015] According to a fourth aspect of any embodiment of the present application, there is provided a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the method described in any embodiment of the present application above is realized.
[0016] According to a fifth aspect of any embodiment of the present application, there is provided a computer program product, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the method described in any embodiment of the present application above is realized.
[0017] The technical solution provided by the present application may include the following beneficial effects:
[0018] According to the above embodiments, by acquiring environmental data of the vehicle's location and generating BEV features in the vehicle's environment based on the environmental data; based on the BEV features, different occupancy prediction units are respectively used to generate prediction results of occupancy grids in different spatial regions in the vehicle's environment. By dynamically adjusting the resolution ratio of occupancy grids in different spatial regions, occupancy grids with smaller grid sizes are used in the close-range spatial region to ensure high-precision perception requirements, and occupancy grids with larger grid sizes are used in the far-range spatial region to reduce computational resource consumption, improving the prediction efficiency of occupancy grids, thereby achieving a balance between prediction efficiency and accuracy requirements.
[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings herein are incorporated into the specification and constitute a part of the present application, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0021] Figure 1 is a flowchart of an occupancy grid prediction method shown according to an exemplary embodiment of the present application;
[0022] Figure 2 is a schematic diagram of a feature pyramid network shown according to an exemplary embodiment of the present application;
[0023] Figure 3It is a flowchart of another occupancy grid prediction method shown according to an exemplary embodiment of the present application;
[0024] Figure 4 It is a schematic structural diagram of an occupancy prediction network shown according to an exemplary embodiment of the present application;
[0025] Figure 5 It is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of the present application;
[0026] Figure 6 It is a block diagram of an occupancy grid prediction device shown according to an exemplary embodiment of the present application. Detailed implementation manners
[0027] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0028] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0029] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0030] When predicting the occupancy grid, currently, an occupancy grid with a fixed grid size is usually used to perform 3D modeling on the surrounding environment, and it is impossible to balance the occupancy grid prediction efficiency and accuracy requirements. If the grid size of the occupancy grid is too small, it will lead to too high a resolution of the occupancy grid in the surrounding environment, an increase in the number of occupancy grids to be processed, a significant increase in the amount of calculation and video memory occupancy, and it is difficult to meet the real-time requirement. Moreover, in the vehicle driving scenario, the far-distance area does not have extremely high accuracy requirements, resulting in a waste of computing resources.
[0031] If the grid size of the occupied grid is too large, it will lead to too low resolution of the occupied grid in the surrounding environment, resulting in a decrease in the prediction accuracy of nearby obstacles. Especially in narrow scenarios or when fine obstacle avoidance is required, it is difficult to meet the safety requirements. Moreover, for key targets such as suspended obstacles and small objects at close range (such as small obstacles on the road surface, pedestrians in the vehicle's blind spot, etc.), there are high-precision perception requirements, and the occupied grid with too large a grid size is prone to missing or misidentifying key targets.
[0032] To solve the above problems, the present application proposes an occupied grid prediction method. To further illustrate the present application, the following embodiments are provided:
[0033] Please refer to Figure 1 , Figure 1 FIG. is a flowchart of an occupied grid prediction method shown according to an exemplary embodiment of the present application. This occupied grid prediction method can be executed by a prediction system, and the prediction system can be applied to a vehicle or can also be applied to a server side such as a single server, a cluster server, a cloud server, etc. This environment perception method can also be executed by other systems or devices in different application scenarios, and the embodiments of the present application do not limit this.
[0034] As Figure 1 shown, the occupied grid prediction method may include the following steps:
[0035] Step 101: Obtain the environmental data of the vehicle's environment and generate BEV features in the vehicle's environment according to the environmental data.
[0036] In this step, the prediction system can obtain the environmental data of the vehicle's environment based on on-vehicle cameras, radars or other sensor devices. The environmental data is the original information of the vehicle's surrounding environment obtained through various sensors, providing basic information for subsequent occupied grid prediction.
[0037] The environmental data may include visual data collected by visual sensors and / or point cloud data collected by radars. The visual sensors can be panoramic cameras located in various directions of the vehicle, covering a 360° field of view around the vehicle.
[0038] The visual data is the original image sequence collected by the camera, which can contain RGB or grayscale pixel information and can be collected by six panoramic cameras respectively located in the front / rear / left / right / left front / right front directions of the vehicle, or can also be collected by the six panoramic cameras and four panoramic cameras for surround view. The embodiments of the present application do not limit this.
[0039] Environmental data can be feature-extracted, and based on the extracted features, BEV features in the environment where the vehicle is located can be generated. For example, through inverse perspective transformation, Lift-Splat-Shoot (LSS) based on deep learning, etc., two-dimensional environmental data can be mapped into BEV features; for another example, through voxelization, PointPillars, etc., three-dimensional environmental data can be mapped into BEV features, and so on.
[0040] Step 102: Based on the BEV features, use different occupancy prediction units respectively to generate prediction results of occupancy grids in different spatial regions in the environment where the vehicle is located. The prediction results are used to represent the occupancy information of the occupancy grids; the grid sizes of the occupancy grids in each spatial region are the same, and the grid sizes of the occupancy grids in different spatial regions increase as the distance from the vehicle increases.
[0041] In this step, based on the BEV features, multi-scale spatial analysis of the vehicle's surrounding environment can be performed through multiple independent different occupancy prediction units. Each occupancy prediction unit is responsible for processing the environmental information within a specific spatial region and outputting the prediction result of the occupancy grid in that spatial region. The occupancy prediction unit can be a Detection Head, a 3D segmentation network, etc.
[0042] The occupancy grid prediction results of the multi-level spatial regions can provide a comprehensive environmental understanding basis for subsequent autonomous driving decisions such as path planning and obstacle avoidance.
[0043] Among them, the environment where the vehicle is located can include at least two spatial regions. The grid sizes of the occupancy grids in each spatial region are the same, and the grid sizes of the occupancy grids in different spatial regions increase as the distance from the vehicle increases.
[0044] The division rules for each spatial region, for example, can be based on linear / non-linear increasing division of distance to determine each spatial region in the environment where the vehicle is located. The grid sizes of the occupancy grids in each spatial region from the center to the outside of the environment where the vehicle is located increase in turn; for another example, it can be based on dynamic division of semantic importance to determine each spatial region in the environment where the vehicle is located, and use the monitored semantic information to adjust the sizes of each spatial region in real time, and so on.
[0045] The grid size of the occupancy grid is the resolution of the entire occupancy grid, which is used to describe the fineness of the grid division. The grid size can be the division unit size of the occupancy grid in three-dimensional space, such as the voxel size (length × width × height) of the three-dimensional division unit in the voxel space.
[0046] The prediction result can be used to represent the occupancy information of the occupancy grid. The occupancy information can reflect whether the occupancy grid is occupied, and is used to distinguish the free area and the obstacle area in the vehicle's environment. The occupancy information can include the probability of occupancy or non-occupancy of the occupancy grid. The occupancy information can also include the occupancy classification information of the occupied occupancy grid, and the occupancy classification information includes but is not limited to drivable road surface, vehicle, building, natural vegetation, etc.
[0047] In one embodiment, before predicting the occupancy grid using the occupancy prediction unit, the BEV features collected at different times in time series can be fused in time series by means such as feature fusion based on a recurrent network, fusion based on an attention mechanism, and fusion based on motion compensation, and the multi-frame BEV features at consecutive times are fused to generate the fused BEV features.
[0048] As described above, by fusing the BEV features collected at different times in time series to generate the fused BEV features, and using the motion continuity to enhance the temporal consistency of the occupancy grid prediction, the occlusion and noise interference of single-frame data can be effectively alleviated, and the accuracy and stability of the occupancy grid prediction can be improved.
[0049] The occupancy grid prediction method of this embodiment obtains the environmental data of the vehicle's environment, generates the BEV features in the vehicle's environment according to the environmental data, and based on the BEV features, uses different occupancy prediction units respectively to generate the prediction results of the occupancy grids in different spatial regions in the vehicle's environment. By dynamically adjusting the resolution ratio of the occupancy grids in different spatial regions, the occupancy grids with smaller grid sizes are used in the close-range spatial region to ensure the high-precision perception requirements, and the occupancy grids with larger grid sizes are used in the far-range spatial region to reduce the consumption of computing resources, improve the prediction efficiency of the occupancy grid, and thus achieve the balance of prediction efficiency and accuracy requirements.
[0050] By adopting the multi-resolution grid division strategy, occupancy grids with higher resolution are used in the spatial region close to the vehicle to capture detailed information such as pedestrians and small obstacles; occupancy grids with lower resolution are used in the spatial region far from the vehicle, which can significantly reduce the computational complexity while ensuring the ability to detect key obstacles, so as to realize the real-time modeling of a large-scale environment under limited computing resources and provide an important guarantee for the reliable operation of the autonomous driving system in different scenarios.
[0051] In the foregoing embodiments, the multi-resolution occupancy grid prediction is realized through multi-scale spatial division and multiple occupancy prediction units. In the following embodiments, the prediction process of the occupancy grid will be described in more detail and can be applied to any of the above embodiments.
[0052] In one embodiment, the spatial region may include a first spatial region and a second spatial region. The grid size of the occupied grids in the first spatial region is smaller than the grid size of the occupied grids in the second spatial region.
[0053] According to the physical characteristics of the autonomous driving scenario, fine structures such as pedestrians and low obstacles need to be recognized in the near-field region of the vehicle, while only the outlines of large-scale objects such as vehicles and buildings need to be perceived in the far-field region. For the division of the spatial region, a dichotomy strategy of the near-field region (the first spatial region) and the far-field region (the second spatial region) can be adopted.
[0054] Exemplarily, the near-field region and the far-field region can be divided with the vehicle's front 6 meters, rear 3 meters, and 3 meters on each side as the boundaries. The near-field region uses occupied grids with a grid size of 5cm×5cm×40cm; the far-field region uses occupied grids with a grid size of 20cm×20cm×40cm.
[0055] Please refer to Figure 2 , Figure 2 which shows a schematic diagram of a Feature Pyramid Network. The Feature Pyramid Networks (FPN) performs multi-scale feature fusion on the input BEV features. First, feature sampling is performed through the bottom-up path, then path sampling is performed through the top-down path, and then feature fusion is performed through lateral connections, finally generating multi-layer feature maps.
[0056] The BEV features can be input into the Feature Pyramid Network to obtain the first feature and the second feature output by the Feature Pyramid Network. The feature dimension of the first feature is higher than that of the second feature. The first feature is the high-dimensional feature output by the high-dimensional feature layer, and the second feature is the low-dimensional feature output by the low-dimensional feature layer, corresponding to the spatial information density requirements of the near-field region and the far-field region respectively. The first feature retains more edge details and texture information and is suitable for processing high-resolution predictions in the near-field region. The second feature layer focuses on the macroscopic semantic features of the far-field region and is suitable for processing low-resolution predictions in the far-field region.
[0057] The occupancy prediction unit may include a first detection head and a second detection head. The first feature can be input into the first detection head to obtain the prediction result of the occupied grids in the first spatial region output by the first detection head. The second feature is input into the second detection head to obtain the prediction result of the occupied grids in the second spatial region output by the second detection head.
[0058] Among them, the network depth structures of the first detection head and the second detection head can be the same or different. Exemplarily, the first detection head can adopt at least two levels of residual module structures. Each level of structure can include n×n convolution, batch normalization, and ReLU activation. The first detection head finally outputs the prediction results of 120×184×8 occupancy grids with a grid size of 5cm×5cm×40cm.
[0059] The residual module structure of the second detection head can be less than that of the first detection head. The second detection head can output the prediction results of 96×104×8 occupancy grids with a grid size of 20cm×20cm×40cm. Through the differentiated network depth structures of the detection heads, the allocation of computing resources can be further balanced.
[0060] As described above, by adopting the feature pyramid network, the first feature and the second feature of different dimensions are generated and respectively input into the corresponding detection heads for prediction, so that the high-dimensional first feature is used for the fine prediction of the first spatial region at close range, and the low-dimensional second feature is used for the efficient inference of the second spatial region at long range, optimizing the computing efficiency while ensuring the accuracy requirements. Moreover, through the hierarchical feature processing of the feature pyramid network, the adaptability to targets of different scales can be enhanced.
[0061] In one embodiment, an occupancy prediction network can be used to generate the BEV feature in the environment where the vehicle is located and generate the prediction results of the occupancy grids in different spatial regions. The occupancy prediction network can include a feature acquisition unit and different occupancy prediction units. The feature acquisition unit can be a backbone network, which is used to generate the BEV feature in the environment where the vehicle is located based on the environmental data, and generate a global BEV feature representation by performing feature extraction and multi-view fusion on the environmental data.
[0062] The occupancy prediction network can be trained. The feature acquisition unit generates the BEV feature in the environment where the vehicle is located according to the pre-prepared sample environmental data. The BEV feature is input into different occupancy prediction units to obtain the sample prediction results of the occupancy grids in different spatial regions output by the occupancy prediction units. Based on the sample prediction results, the total loss is obtained. According to the total loss, the network parameters of the occupancy prediction network are adjusted by the backpropagation algorithm, so that the model gradually improves the prediction accuracy and consistency during the training process.
[0063] Among them, the total loss is the error between the sample prediction result and the true annotation of the sample environmental data. The total loss consists of multiple parts. The total loss can include: the probability distribution loss and the occupancy prediction losses respectively corresponding to each occupancy prediction unit.
[0064] The probability distribution loss is used to characterize the differences in probability distributions between individual spatial regions, reflecting the consistency of the probability distribution of the overall spatial region and ensuring that the prediction results in different spatial regions are consistent in terms of probability distribution. For example, the KL divergence (Kullback-Leibler divergence) can be used to measure the difference in probability distributions between near-field and far-field predictions, prompting the model to be coordinated between different regions.
[0065] The total loss can be determined based on the probability distribution loss and the weights corresponding to the occupancy prediction losses of each occupancy prediction unit. Continuing with the example where the occupancy prediction unit includes a first detection head and a second detection head, the loss function can be as shown in Equation 1 below, and the total loss can be determined according to Equation 1:
[0066] Total Loss =λ near * Near Loss +λ far * Far Loss +λ consist * C Loss Equation 1
[0067] Where, Total Loss represents the total loss of the sample prediction result, λ near represents the weight of the occupancy prediction loss of the first detection head, Near Loss represents the occupancy prediction loss of the first detection head, λ far represents the weight of the occupancy prediction loss of the second detection head, Far Loss represents the occupancy prediction loss of the second detection head, λ consist represents the weight of the probability distribution loss, C Loss represents the probability distribution loss.
[0068] The weights corresponding to the occupancy prediction losses of each occupancy prediction unit can be determined according to the spatial region corresponding to the occupancy prediction unit. For example, higher weights can be assigned to the predictions in the near-field region because it is more critical for autonomous driving decisions. Exemplarily, λ near can be set to 0.6, reflecting the higher importance of the near-field region for autonomous driving decisions, λ far can be set to 0.3, as the information in the far-field region is relatively coarse-grained and has a lower weight, and λ consist can be set to 0.1.
[0069] Exemplarily, the probability distribution loss can be determined according to Equation 2 below:
[0070]
[0071] Where, C Loss represents the probability distribution loss, KLDivergence Indicates the difference in probability distribution between the near-field region and the far-field region, which can ensure the consistency of the predicted occupancy probability distribution in the near-field region and the far-field region. Indicates the probability distribution of the first detection head in the near-field region. Indicates the probability distribution of the second detection head in the far-field region.
[0072] As described above, by introducing the probability distribution loss to constrain the output consistency of different detection heads, forcing the occupancy grid predictions in different spatial regions to maintain probability distribution alignment, the problem of probability distribution differences in multi-scale predictions is solved, and the synergy and robustness of cross-range predictions are improved.
[0073] In one embodiment, at least one of the following losses can be obtained according to the sample prediction results of each occupancy prediction unit: classification loss, regression loss, and orientation loss. Each sub-loss is for a different prediction task. According to the at least one loss and the corresponding loss weights, the occupancy prediction loss of the occupancy prediction unit is obtained.
[0074] Among them, the classification loss can be the Cross-Entropy Loss, which is used to measure the semantic classification accuracy of the occupancy grid (such as vehicles, pedestrians, roads, etc.). The regression loss can reflect the offset error between the predicted occupancy grid position and the true annotation position, and is used to optimize the position offset of the occupancy grid (such as the exact position of the vehicle bounding box). The orientation loss can be the Cosine Similarity Loss, which is used to optimize the prediction accuracy of the object orientation (such as the vehicle orientation).
[0075] Exemplarily, the occupancy prediction loss of the first detection head can be determined according to the following formula 3:
[0076]
[0077] Where Near Loss Indicates the occupancy prediction loss of the first detection head. Indicates the weight of the classification loss of the first detection head. Indicates the classification loss of the first detection head. Indicates the weight of the regression loss of the first detection head. Indicates the regression loss of the first detection head. Indicates the weight of the orientation loss of the first detection head. Indicates the orientation loss of the first detection head.
[0078] Exemplarily, the occupancy prediction loss of the second detection head can be determined according to the following formula 4:
[0079]
[0080] Among them, Far Loss represents the occupancy prediction loss of the first detection head, represents the weight of the classification loss of the second detection head, represents the classification loss of the second detection head, represents the weight of the regression loss of the second detection head, represents the regression loss of the second detection head, represents the weight of the orientation loss of the second detection head, represents the orientation loss of the second detection head.
[0081] As described above, by flexibly combining the classification loss, regression loss, and orientation loss, and differentiating the design of the loss and the corresponding loss weights according to the target characteristics of different spatial regions (such as emphasizing position accuracy in the near spatial region and focusing on target detection in the far spatial region), the network can adaptively optimize the key indicators, thereby taking into account both the detection prediction accuracy and efficiency.
[0082] In one embodiment, the environmental data may include visual data and point cloud data. Feature extraction can be performed on the visual data to obtain the visual features of the visual data. The visual features are two-dimensional features extracted from the visual data and are used to capture abstract information such as target contours and textures. Feature extraction is performed on the point cloud data to obtain the point cloud features of the point cloud data. The visual features and point cloud features are mapped to the three-dimensional BEV perspective to obtain the visual BEV features and point cloud BEV features. The visual BEV features and point cloud BEV features are fused to obtain the BEV features.
[0083] As described above, by separately performing feature extraction on the visual data and point cloud data and mapping the extracted visual features and point cloud features to the three-dimensional BEV perspective, the perspective deviation in the view or point cloud perspective can be avoided, combining the rich semantic information of vision and the precise geometric information of the point cloud, enhancing the generalization ability of the occupancy grid prediction in scenarios such as illumination changes and occlusion, and improving the accuracy and robustness of the occupancy grid prediction; by fusing the visual BEV features and point cloud BEV features after mapping to the three-dimensional BEV perspective, the high computational complexity of directly aligning multimodal data can be avoided, while retaining the complementary advantages of different sensors, improving the real-time performance of the occupancy grid prediction, and being more suitable for in-vehicle platforms with limited computing resources.
[0084] To further introduce the environmental perception process, Figure 3 The flowchart of another occupancy grid prediction method is shown. The occupancy grid prediction method may include the following steps:
[0085] Step 301: Obtain the visual data and point cloud data around the vehicle.
[0086] In this step, visual data and point cloud data around the vehicle can be obtained based on ten panoramic cameras and radars of the vehicle.
[0087] Step 302: Extract features from the visual data and map the extracted visual features to the three-dimensional BEV perspective.
[0088] In this step, the obtained visual data and point cloud data can be input into the trained occupancy prediction network. Please refer to Figure 4 , Figure 4 which shows a schematic structural diagram of an occupancy prediction network. Exemplarily, the occupancy prediction network may include: an image encoder 400, a perspective transformation module 401, a point cloud encoder 402, a splicing module 403, a BEV encoder 404, a temporal fusion module 405, a BEV decoder 406, a clipping module 407, an upsampling convolution module 408, a first detection head 409, and a second detection head 410.
[0089] The image encoder 400 can extract features from the visual data to obtain visual features. The perspective transformation module 401 maps the visual features to the three-dimensional BEV perspective to obtain the mapped visual BEV features (Cam-BEVFeat).
[0090] Step 303: Extract features from the point cloud data and map the extracted point cloud features to the three-dimensional BEV perspective.
[0091] In this step, the point cloud encoder 402 extracts features from the point cloud data to obtain the point cloud features of the point cloud data, maps the point cloud features to the three-dimensional BEV perspective to obtain the mapped point cloud BEV features (Lidar-BEVFeat), and the dimension of the mapped point cloud BEV features is aligned with the visual BEV features.
[0092] Step 304: Fuse the visual BEV features and the point cloud BEV features to obtain BEV features.
[0093] In this step, the splicing module 403 splices the point cloud BEV features and the visual BEV features along the channel dimension (Channel) to generate multimodal fusion features. The BEV encoder 404 further fuses the features through 3D convolution or Transformer to eliminate the sensor modality differences and generate the BEV features at the current time t.
[0094] Step 305: Perform temporal fusion on the BEV features collected at different times in time series to generate the fused BEV features.
[0095] In this step, the temporal fusion module 405 performs temporal fusion on the BEV features at consecutive times t-n,......, t-1, and t in time series to generate the fused BEV features.
[0096] Step 306: Use the Feature Pyramid Network to generate the first feature and the second feature corresponding to the BEV feature.
[0097] In this step, the BEV decoder 406 can input the fused BEV feature into the Feature Pyramid Network, and use the Feature Pyramid Network to generate a high-dimensional first feature and a low-dimensional second feature.
[0098] Step 307: Use the first detection head to generate the prediction result of the occupancy grid in the first spatial region.
[0099] In this step, the cropping module 407 can crop the first feature generated by the BEV decoder 406 to obtain the first feature corresponding to the first spatial region. The upsampling convolution module 408 performs interpolation upsampling on the cropped first feature to improve the feature resolution, and refines the upsampled first feature through a convolutional layer to eliminate artifacts. The first detection head 409 generates the prediction result of the occupancy grid in the first spatial region based on the processed first feature.
[0100] Step 308: Use the second detection head to generate the prediction result of the occupancy grid in the second spatial region.
[0101] In this step, the second detection head 410 can generate the prediction result of the occupancy grid in the second spatial region based on the second feature generated by the BEV decoder 406.
[0102] It can be understood that Figure 4 the network structure of the occupancy prediction network shown is only an example, and other network structures can also be adopted. The embodiments of the present application do not limit this.
[0103] Figure 5 FIG. is a schematic structural diagram of an electronic device according to an exemplary embodiment of the present application. The electronic device may be, for example, a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a personal digital assistant, a server, a smart home appliance, a vehicle-mounted computer, etc. Refer to Figure 5 , at the hardware level, the electronic device includes a processor 501, an internal bus 502, a network interface 503, a memory 504, and a non-volatile memory 505. Of course, other hardware required for other services may also be included. The processor 501 reads the corresponding computer program from the non-volatile memory 505 into the memory 504 and then runs it, forming an occupancy grid prediction device at the logical level. Of course, in addition to the software implementation method, the present application does not exclude other implementation methods, such as logic devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and may also be hardware or a logic device.
[0104] Figure 6 This is a block diagram of an occupancy grid prediction device shown according to an exemplary embodiment of the present application. Referring to Figure 6 , the device may include: a generation module 601 and a prediction module 602, where:
[0105] The generation module 601 is configured to obtain environmental data of the environment where the vehicle is located, and generate a BEV feature in the environment where the vehicle is located according to the environmental data;
[0106] The prediction module 602 is configured to, based on the BEV feature, respectively use different occupancy prediction units to generate prediction results of occupancy grids in different spatial regions in the environment where the vehicle is located, and the prediction results are used to represent the occupancy information of the occupancy grids; the grid size of the occupancy grids in each spatial region is the same, and the grid size of the occupancy grids in different spatial regions increases as the distance from the vehicle increases.
[0107] In an example, the spatial region includes: a first spatial region and a second spatial region; the grid size of the occupancy grids in the first spatial region is smaller than the grid size of the occupancy grids in the second spatial region; the occupancy prediction unit includes: a first detection head and a second detection head; when the prediction module 602 is configured to, based on the BEV feature, respectively use different occupancy prediction units to generate prediction results of occupancy grids in different spatial regions in the environment where the vehicle is located, it includes: inputting the BEV feature into a feature pyramid network to obtain a first feature and a second feature output by the feature pyramid network, where the feature dimension of the first feature is higher than the feature dimension of the second feature; inputting the first feature into the first detection head to obtain the prediction result of the occupancy grids in the first spatial region output by the first detection head; inputting the second feature into the second detection head to obtain the prediction result of the occupancy grids in the second spatial region output by the second detection head.
[0108] In an example, the device further includes a training module ( Figure 6(not shown), the training module is used to train the occupancy prediction network; the occupancy prediction network includes: a feature acquisition unit and the different occupancy prediction units; the feature acquisition unit is used to generate BEV features in the environment where the vehicle is located based on the environmental data; when used to train the occupancy prediction network, the training module includes: generating, by the feature acquisition unit, BEV features in the environment where the vehicle is located according to the sample environmental data; inputting the BEV features into the different occupancy prediction units to obtain sample prediction results of the occupancy grids of the different spatial regions output by the occupancy prediction units; based on the sample prediction results, obtaining a total loss, the total loss including: a probability distribution loss and occupancy prediction losses respectively corresponding to the respective occupancy prediction units; the probability distribution loss is used to characterize the difference in probability distribution between the respective spatial regions; adjusting the network parameters of the occupancy prediction network according to the total loss.
[0109] In one example, when used to obtain a total loss based on the sample prediction results, the training module includes: obtaining at least one of the following losses according to the sample prediction results of the respective occupancy prediction units: a classification loss, a regression loss, and an orientation loss; obtaining the occupancy prediction loss of the occupancy prediction unit according to the at least one loss and the corresponding loss weight.
[0110] In one example, before the prediction module 602 is used to generate prediction results of the occupancy grids in different spatial regions in the environment where the vehicle is located by using different occupancy prediction units based on the BEV features, it further includes: performing temporal fusion on the BEV features collected at different times in time series to generate a fused BEV feature.
[0111] In one example, the environmental data includes visual data collected by a visual sensor and point cloud data collected by a radar; when used to generate BEV features in the environment where the vehicle is located according to the environmental data, the generation module 601 includes: extracting features from the visual data to obtain visual features of the visual data; extracting features from the point cloud data to obtain point cloud features of the point cloud data; mapping the visual features and the point cloud features to a three-dimensional BEV perspective to obtain visual BEV features and point cloud BEV features; fusing the visual BEV features and the point cloud BEV features to obtain the BEV features.
[0112] The implementation processes of the functions and roles of the respective units in the above device are specifically detailed in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.
[0113] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0114] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory including instructions. The above instructions can be executed by a processor of the occupancy grid prediction device to implement the method described in any one of the above embodiments.
[0115] Among them, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc. This application does not limit this.
[0116] In an exemplary embodiment, a computer program product including a computer program / instructions is also provided. The above computer program / instructions can be executed by a processor of the occupancy grid prediction device to implement the method described in any one of the above embodiments.
[0117] The specific embodiments of this application are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0118] Those skilled in the art will readily conceive of other embodiments of this application after considering the specification and practicing the invention herein. This application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is only limited by the appended claims.
[0119] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. An occupied grid prediction method, characterized in that, The method includes: Obtaining environmental data of the vehicle's environment and generating a BEV feature in the vehicle's environment based on the environmental data; Based on the BEV feature, using different occupancy prediction units respectively to generate prediction results of occupancy grids in different spatial regions in the vehicle's environment, where the prediction results are used to represent the occupancy information of the occupancy grids; the grid sizes of the occupancy grids in each spatial region are the same, and the grid sizes of the occupancy grids in different spatial regions increase as the distance from the vehicle increases.
2. The method according to claim 1, characterized in that The spatial regions include: a first spatial region and a second spatial region; the grid size of the occupancy grids in the first spatial region is smaller than the grid size of the occupancy grids in the second spatial region; The occupancy prediction unit includes: a first detection head and a second detection head; The generating prediction results of occupancy grids in different spatial regions in the vehicle's environment by using different occupancy prediction units respectively based on the BEV feature includes: Inputting the BEV feature into a feature pyramid network to obtain a first feature and a second feature output by the feature pyramid network, where the feature dimension of the first feature is higher than the feature dimension of the second feature; Inputting the first feature into the first detection head to obtain the prediction result of the occupancy grid in the first spatial region output by the first detection head; Inputting the second feature into the second detection head to obtain the prediction result of the occupancy grid in the second spatial region output by the second detection head.
3. The method according to claim 1, wherein The method further includes: training the occupancy prediction network; The occupancy prediction network includes: a feature acquisition unit and the different occupancy prediction units; the feature acquisition unit is used to generate a BEV feature in the vehicle's environment based on the environmental data; The training the occupancy prediction network includes: Generating a BEV feature in the vehicle's environment by the feature acquisition unit according to sample environmental data; Inputting the BEV feature into the different occupancy prediction units to obtain sample prediction results of the occupancy grids in the different spatial regions output by the occupancy prediction units; Based on the sample prediction results, obtaining a total loss, where the total loss includes: a probability distribution loss and an occupancy prediction loss corresponding to each occupancy prediction unit respectively; the probability distribution loss is used to characterize the difference in probability distribution between different spatial regions; Adjusting the network parameters of the occupancy prediction network according to the total loss.
4. The method according to claim 3, wherein The obtaining the total loss based on the sample prediction results includes: Obtaining at least one of the following losses according to the sample prediction results of each occupancy prediction unit: classification loss, regression loss, direction loss; Obtaining the occupancy prediction loss of the occupancy prediction unit according to the at least one loss and the corresponding loss weight.
5. The method according to claim 1, wherein Before the generating prediction results of occupancy grids in different spatial regions in the vehicle's environment by using different occupancy prediction units respectively based on the BEV feature, the method further includes: Performing temporal fusion on BEV features collected at different times in time series to generate a fused BEV feature.
6. The method according to claim 1, wherein The environmental data includes visual data collected by a visual sensor and point cloud data collected by a radar; Generating the BEV features in the environment where the vehicle is located according to the environmental data includes: Performing feature extraction on the visual data to obtain visual features of the visual data; Performing feature extraction on the point cloud data to obtain point cloud features of the point cloud data; Mapping the visual features and the point cloud features to a three-dimensional BEV perspective to obtain visual BEV features and point cloud BEV features; Fusing the visual BEV features and the point cloud BEV features to obtain the BEV features.
7. An occupancy grid prediction device, characterized in that, The device includes: A generating module, configured to obtain environmental data of the environment where the vehicle is located, and generate BEV features in the environment where the vehicle is located according to the environmental data; A prediction module, configured to, based on the BEV features, respectively use different occupancy prediction units to generate prediction results of occupancy grids in different spatial regions in the environment where the vehicle is located, where the prediction results are used to represent the occupancy information of the occupancy grids; the grid sizes of the occupancy grids in each spatial region are the same, and the grid sizes of the occupancy grids in different spatial regions increase as the distance from the vehicle increases.
8. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor realizes the method according to any one of claims 1-6 by running the executable instructions.
9. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the instructions are executed by the processor, the method according to any one of claims 1-6 is realized.
10. A computer program product, on which a computer program / instructions are stored, characterized in that, When the computer program / instructions are executed by the processor, the method according to any one of claims 1-6 is realized.