Occupancy grid prediction method and apparatus, smart device, and storage medium

By adopting occupancy grid prediction methods and pillar-formed occupation information in the form of pillar, the problem of insufficient perception of long-tail obstacles in open-world traffic scenarios is solved, more efficient training and deployment is achieved, and resource occupancy and pressure of post-processing is reduced.

WO2025108121A1PCT designated stage expired Publication Date: 2025-05-30ANHUI NIO AUTONOMOUS DRIVING TECH CO LTD

Patent Information

Application Number
PCT/CN2024/131253
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-11-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The lack of perception of deformable, shaped, and unknown categories of obstacles in open-world traffic scenarios has led to the possibility of perception technology failing in long-tail problems, and a more robust representation method is urgently needed.

Method used

The occupancy raster prediction method is used to input the occupancy raster prediction model through pure visual images, and the occupancy results in the form of cylinder are generated using the model, and the occupancy information in the form of pillar is used in the training data processing stage to reduce the CPU resource occupancy and data transmission pressure of post-processing.

Benefits of technology

It reduces the training time and time delay of the automatic driving function during deployment, facilitates mass production vehicle deployment, and reduces the time-consuming and direct data transmission pressure of post-processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024131253_30052025_PF_FP_ABST
    Figure CN2024131253_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of autonomous driving, and in particular relates to an occupancy grid prediction method, an occupancy grid prediction apparatus for implementing the method, a computer storage medium for implementing the method, and a smart device having the occupancy grid prediction apparatus. The method comprises: acquiring a pure visual image of the surroundings of a vehicle body; inputting the pure visual image into an occupancy grid prediction model; and using the occupancy grid prediction model to determine an occupancy result of an object in the pure visual image as a model output. The model output is an occupancy result in a cylindrical form, and the occupancy grid prediction model is constructed on the basis of a training dataset comprising sample images and truth value information of the sample images. The truth value information is directly generated from point cloud data or converted from occupancy information in a voxel form generated on the basis of the point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

Occupancy grid prediction method and device, intelligent device and storage medium

[0001] This application claims priority to Chinese patent application No. 202311558724.3, filed on November 22, 2023, entitled “Occupancy grid prediction method and apparatus, intelligent device and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of autonomous driving, and more specifically to an occupancy grid prediction method, an occupancy grid prediction device for implementing the method, a computer storage medium for implementing the method, and an intelligent device equipped with the occupancy grid prediction device. Background Art

[0003] The safe operation of autonomous vehicles requires an accurate and comprehensive representation of the vehicle's surroundings. The three-dimensional (3D) vehicle perception module primarily consists of 3D object detection, multi-object tracking, and trajectory prediction. However, object-centric perception techniques can fail in open-world traffic scenarios where the object's shape or appearance is ambiguous. These obstacles, also known as long-tail obstacles, include deformable obstacles such as two-section trailers; irregular obstacles such as overturned vehicles; obstacles of unknown class such as road debris and garbage; and partially obscured objects. Therefore, a more robust representation of these long-tail obstacles is urgently needed. Occupancy grid prediction is considered a promising solution because it can provide the occupancy and motion of any location in the 3D space around the vehicle without knowing the object. In this way, occupancy grid prediction will become an important prerequisite for supporting downstream driving tasks such as vehicle collision avoidance and vehicle trajectory planning.

[0004] Some occupancy grid prediction schemes use voxel-based model output. This type of output takes up a lot of CPU resources during post-processing, which is not conducive to the deployment and mass production of edge devices. It also brings excessive data transmission pressure and high application costs.

[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field.

[0006] Summary of the Invention

[0007] To solve or at least alleviate one or more of the above problems, the following technical solutions are provided. The embodiments of the present application provide an occupancy grid prediction method, an occupancy grid prediction device for implementing the method, a computer storage medium for implementing the method, and an intelligent device equipped with the occupancy grid prediction device, which can reduce the training time and deployment delay of autonomous driving functions, facilitate deployment in mass-produced vehicles, and reduce the time consumption of post-processing and the pressure of direct data transmission.

[0008] According to a first aspect of the present application, an occupancy grid prediction method is provided, the method comprising the following steps: acquiring a purely visual image around a vehicle body; inputting the purely visual image into an occupancy grid prediction model; and using the occupancy grid prediction model to determine an occupancy result of a target in the purely visual image as a model output, wherein the model output is an occupancy result in the form of a cylinder, and the occupancy grid prediction model is constructed based on a training data set containing sample images and true value information of the sample images, and the true value information is directly generated by point cloud data or converted from voxel-form occupancy information generated based on point cloud data.

[0009] As an alternative or supplement to the above scheme, in an occupancy grid prediction method according to an embodiment of the present application, the occupancy information in voxel form is generated based on at least the following steps: dividing the three-dimensional space into regular voxel grids and quantizing the point cloud data within each voxel grid into occupancy information, wherein the occupancy information includes the occupancy status, semantic category, speed level, speed value, or a combination thereof of the voxel.

[0010] As an alternative or supplement to the above scheme, in an occupancy grid prediction method according to an embodiment of the present application, the true value information is converted from occupancy information in voxel form generated based on point cloud data, including: offline data processing of the occupancy information in voxel form to generate compressed cylindrical occupancy information, and the occupancy information in voxel form is generated based on one or more of the following items: single-frame laser point cloud data, time-series spliced ​​laser point cloud data, and point cloud data reconstructed by visual three-dimensional reconstruction.

[0011] As an alternative or supplement to the above scheme, in an occupancy grid prediction method according to an embodiment of the present application, offline data processing of the occupancy information in the form of voxels includes: selecting voxels with an occupancy status of occupied as occupied voxels; projecting the three-dimensional space into a two-dimensional grid; aggregating one or more occupied voxels falling into a two-dimensional grid into one or two cylinders; and generating occupancy information for each cylinder.

[0012] As an alternative or supplement to the above scheme, in an occupancy grid prediction method according to an embodiment of the present application, projecting the three-dimensional space into the two-dimensional grid includes: in the vehicle coordinate system, dividing the x and y axes into a regular two-dimensional grid according to a first resolution, wherein the positive direction of the x-axis is the forward direction of the vehicle, and the positive direction of the y-axis is the left side of the vehicle; and projecting the three-dimensional space in which the xy plane falls within the two-dimensional grid into the two-dimensional grid, wherein the xy plane is the plane formed by the x-axis and the y-axis.

[0013] As an alternative or supplement to the above scheme, in an occupancy grid prediction method according to an embodiment of the present application, offline data processing of the occupancy information in the form of voxels also includes: filtering out occupied voxels of unnecessary types based on semantic categories; and / or filtering out occupied voxels whose z-axis coordinates in the vehicle coordinate system are less than a first height threshold or greater than a second height threshold.

[0014] As an alternative or supplement to the above scheme, in an occupancy grid prediction method according to an embodiment of the present application, generating occupancy information for each column includes: if the number of occupied voxels falling into the column is greater than or equal to a third quantity threshold, marking the column as occupied, and marking the upper and lower boundaries of the height of the column as the maximum height and minimum height of one or more occupied voxels falling therein, respectively; and if the number of occupied voxels falling into the column is less than the third quantity threshold, marking the column as unoccupied, and marking the upper and lower boundaries of the height of the column as 0.

[0015] As an alternative or supplement to the above scheme, in an occupancy grid prediction method according to an embodiment of the present application, the occupancy information of the column includes the coordinate information of the column, and the coordinate information includes one of the following: the coordinates of the center point of the column and the height of the column in the vehicle coordinate system; the upper and lower boundaries of the height of the column in the vehicle coordinate system; the coordinates of the center point of the column, the height of the column and the height above the ground in the vehicle coordinate system; or the upper and lower boundaries of the height of the column and the height above the ground in the vehicle coordinate system.

[0016] As an alternative or supplement to the above scheme, in an occupancy grid prediction method according to an embodiment of the present application, generating occupancy information for each column includes: marking the speed of the column as the average speed of one or more occupied voxels falling within the column; or if the average speed of one or more occupied voxels falling within the column is greater than or equal to a fourth speed threshold, marking the speed of the column as dynamic, otherwise marking it as static.

[0017] As an alternative or supplement to the above scheme, in an occupancy grid prediction method according to an embodiment of the present application, generating occupancy information of each column includes: determining the semantic category of the column using a voting mechanism based on the semantic category of one or more occupied voxels falling within the column.

[0018] As an alternative or supplement to the above solution, in an occupancy grid prediction method according to an embodiment of the present application, two-dimensional convolution is used in the head part of the occupancy grid prediction model.

[0019] According to a second aspect of the present application, an occupancy grid prediction device is provided, comprising: a memory; a processor; and a computer program stored on the memory and executable on the processor, wherein the execution of the computer program enables any one of the occupancy grid prediction methods described in the first aspect of the present application to be executed.

[0020] According to a third aspect of the present application, a smart device is provided, which includes the occupancy grid prediction device according to the second aspect of the present application.

[0021] According to a fourth aspect of the present application, a computer storage medium is provided, wherein the computer storage medium includes instructions, and the instructions, when run, execute any one of the occupancy grid prediction methods described in the first aspect of the present application.

[0022] According to one or more embodiments of the present application, the occupancy grid prediction scheme uses pure visual images as input, and uses pillar-form occupancy information in the training data processing stage (for example, compressing voxel-form occupancy information into pillar form, or directly generating pillar-form true values ​​from point cloud data), thereby directly generating pillar-form occupancy results in the model output stage. Compared with the voxel-form occupancy results, this scheme greatly reduces the amount of central processing unit (CPU) resources occupied in the post-processing process and the pressure of direct data transmission, and reduces the time delay during deployment without affecting network inference during deployment, thereby facilitating the deployment and mass production of edge devices (for example, computing power-constrained devices such as mobile phones and vehicle-mounted chips). BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The above and / or other aspects and advantages of the present application will become clearer and easier to understand through the following description of various aspects in conjunction with the accompanying drawings, in which the same or similar elements are represented by the same reference numerals. In the drawings:

[0024] FIG1 is a schematic flow chart of an occupancy grid prediction method 10 according to one or more embodiments of the present application;

[0025] FIG2 is a schematic flowchart of voxel to pillar format conversion according to one or more embodiments of the present application;

[0026] FIG3 is a schematic diagram illustrating how voxel aggregation is expressed as a pillar according to one or more embodiments of the present application;

[0027] FIG4 is a schematic diagram of a vehicle's uphill field of view according to one or more embodiments of the present application;

[0028] FIG5 is a schematic diagram of a network design according to one or more embodiments of the present application;

[0029] FIG6 is a schematic diagram of model input and output according to one or more embodiments of the present application;

[0030] FIG7 is a schematic block diagram of an occupancy grid prediction device 70 according to one or more embodiments of the present application; and

[0031] FIG8 is a schematic block diagram of a smart device 80 according to one or more embodiments of the present application. DETAILED DESCRIPTION

[0032] The description of the following specific embodiments is merely exemplary in nature and is not intended to limit the disclosed technology or the application and use of the disclosed technology. In addition, there is no intention to be bound by any express or implied theory presented in the foregoing technical field, background technology or the following specific embodiments.

[0033] In the following detailed description of the embodiments, numerous specific details are set forth to provide a more thorough understanding of the disclosed technology. However, it will be apparent to one of ordinary skill in the art that the disclosed technology can be practiced without these specific details. In other instances, well-known features are not described in detail to avoid unnecessarily complicating the description.

[0034] Terms such as "comprising" and "including" indicate that in addition to the units and steps directly and explicitly stated in the specification, the technical solution of the present application does not exclude the presence of other units and steps that are not directly or explicitly stated. Terms such as "first" and "second" do not indicate the order of units in terms of time, space, size, etc., but are only used to distinguish between the units. The technology of the present application is generally used in electric vehicles, including but not limited to pure electric vehicles (BEVs), hybrid electric vehicles (HEVs), fuel cell vehicles (FCEVs), etc.

[0035] An automated driving system (ADS) is a system that continuously performs all dynamic driving tasks (DDT) within its operational domain design (ODD). Specifically, the system is only allowed to fully assume the task of autonomous vehicle control under specified appropriate driving scenarios. When the vehicle meets the ODD conditions, the system is activated, replacing the human driver as the vehicle's primary driver. The DDT refers to the continuous lateral (left and right steering) and longitudinal motion control (acceleration, deceleration, and constant speed) of the vehicle, as well as the detection and response to objects and events in the vehicle's driving environment. The ODD refers to the conditions under which the automated driving system can operate safely. These conditions can include geographic location, road type, speed range, weather, time of day, and national and local traffic laws and regulations.

[0036] The Vision Occupancy Network is a visual perception technology based on deep learning. It can be used as a supplement or alternative to LiDAR obstacle avoidance in autonomous driving. Occupancy network technology is based on visual signals. On top of traditional 3D target recognition capabilities, it understands and processes spatial information in a voxel (i.e., 3D spatial unit) manner. Due to the added perception of voxel occupancy, the perception system can restore the traversable area of ​​3D physical space with high fidelity. It does not need to consider what the object is, but only whether the voxel is occupied. This fundamentally avoids the problem of traditional vision missing objects in the training set, greatly improving the generalization ability of the model and better adapting to different scenarios and environments. In addition, compared to the sparse and discontinuous point cloud data generated by LiDAR, the information content collected by high-definition cameras is richer, enabling the occupancy network to better fuse 3D geometric information with semantic information, helping the vehicle to more accurately restore the 3D scene. However, some occupancy grid prediction schemes based on pure visual occupancy networks use voxel-format model outputs. The huge amount of data that accompanies this fine data will bring huge transmission pressure whether it is directly transmitted to the downstream or used for visualization of mass-produced vehicles. In addition, the model output in voxel format will occupy more CPU resources during post-processing, which is not conducive to the deployment and mass production of edge devices. In this regard, the present application proposes to use pillar-format occupancy information in the training data processing stage (for example, compressing voxel-format occupancy information into pillar format, or directly generating pillar-format true values ​​from point cloud data), so as to directly generate pillar-format occupancy results in the model output stage, so as to reduce the CPU resources and transmission pressure occupied in post-processing, and reduce the time delay during deployment without affecting network inference during deployment.

[0037] Hereinafter, various exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings.

[0038] Referring to the drawings below, FIG1 is a schematic flow chart of an occupancy grid prediction method 10 according to one or more embodiments of the present application.

[0039] As shown in FIG. 1 , in step S110 , a purely visual image of the surroundings of the vehicle body is acquired.

[0040] Exemplarily, according to one or more embodiments of the present application, multiple video acquisition units are used to perform multi-perspective acquisition of environmental information around the vehicle body and output corresponding multi-perspective video data. Multiple video acquisition units can be respectively set at different preset positions on the vehicle, and each video acquisition unit corresponds to a perspective. Exemplarily, the video data collected by the multi-perspective video acquisition units (for example, the left front camera, the forward camera, and the right front camera) can also be subjected to frame extraction processing, and a series of pure visual images with timestamps are generated. In this article, cameras, lenses, video cameras, cameras, etc. all refer to devices that can obtain images or images within the coverage range. They have similar meanings and are interchangeable. This application does not impose any restrictions on this.

[0041] In step S120, the purely visual image acquired in step S110 is input into an occupancy grid prediction model, and in step S130, the occupancy grid prediction model is used to determine the occupancy of objects in the purely visual image as a model output. The occupancy grid prediction model is constructed based on a training dataset containing sample images and ground truth information for the sample images, where the ground truth information is generated directly from point cloud data or converted from voxel-based occupancy information generated based on the point cloud data.

[0042] It should be noted that the occupancy result output in step S130 is in pillar format, not voxel format. As mentioned above, the model output in voxel format will generate a huge amount of data, which will put pressure on subsequent processing and transmission, and is not conducive to the deployment and mass production of edge devices. In this regard, in terms of the amount of post-processing data input, since the occupancy result in pillar format is only the height of the pillar on the z-axis, it saves several times the data processing of the z-axis voxels compared to the data in voxel format.

[0043] Furthermore, when ground truth information is converted from voxel-based occupancy information generated from point cloud data, this solution places the voxel-to-pillar conversion process within the training data processing, rather than in the network inference or post-processing steps. This approach reduces post-processing time and direct transmission pressure without affecting network inference during deployment.

[0044] Specifically, optionally, the pillar-format occupancy result can be achieved by one or more of the following steps S131-S133. It should be noted that steps S131-S133 apply only to the case where the ground truth information is converted from voxel-format occupancy information generated based on point cloud data. This application does not impose any specific restrictions on the method of generating pillar-format or voxel-format occupancy information from point cloud data.

[0045] In step S131 , in the pre-processing stage of training data, the three-dimensional space is divided into regular voxel grids and the point cloud data in each voxel grid is quantized into occupancy information.

[0046] Voxel is the abbreviation of volume pixel, which is the smallest unit of data located on a regular grid in three-dimensional space. Step S131 aims to realize voxelization of point cloud data. Voxelization is the process of using voxels to approximate the spatial structure and geometric shape of a scene or object. The basic principle of voxelization is: create a 3D stereo grid on the input point cloud data, then set the voxel resolution, and evenly divide the 3D stereo grid with the voxel resolution as the basic unit to form a set of regular 3D cubes (that is, voxel grids). Then, in each 3D cube, the center point or centroid point of all the point cloud data in the cube is used to approximate all the points in the cube, and the corresponding voxel form occupancy information is obtained after grid processing. As an alternative to step S131, voxel form occupancy information can also be directly obtained from public data sets. Occupancy information in voxel form may include one or more of the following: occupancy state of the voxel (e.g., occupied, unoccupied, unknown), semantic category (e.g., vehicle, pedestrian, tree), speed level (e.g., static, dynamic), and speed value.

[0047] Optionally, the occupancy information in voxel form may be derived from any one or a combination of multiple items of single-frame laser point cloud data, time-series spliced ​​laser point cloud data, and visual three-dimensional reconstructed point cloud data.

[0048] Next, in step S132 , offline data processing is performed on the occupancy information in the voxel form to generate compressed occupancy information in the pillar form.

[0049] Step S132 is intended to achieve voxel-to-pillar format conversion. As described above, the format conversion in step S132 occurs during the training data processing, rather than during the network inference step or post-processing step. The format conversion process according to one or more embodiments of the present application will be described in detail below with reference to Figure 2.

[0050] As shown in FIG2 , in step S210 , a plane rasterization process is first performed. Exemplarily, the plane of z=0 is first defined as a reference plane in the vehicle coordinate system, and the x and y axes are divided into a regular two-dimensional grid (for example, each grid size is 0.2m×0.2m) according to a first resolution (for example, 0.2m) on the reference plane, and then the three-dimensional space in which the xy plane falls within the two-dimensional grid is projected onto the two-dimensional grid. The vehicle coordinate system in this article is a coordinate system established with the vehicle itself as the reference system. In the vehicle coordinate system, the center of mass of the vehicle is often selected as the origin (or the midpoint of the rear axle of the vehicle), the forward direction of the vehicle is defined as the positive direction of the x-axis, the left side of the vehicle is defined as the positive direction of the y-axis, and the top of the vehicle is defined as the positive direction of the z-axis.

[0051] Optionally, step S210 further includes: selecting a voxel with an occupied status of occupied as an occupied voxel (occ_voxel); and aggregating one or more occ_voxels falling into a two-dimensional grid into one or two pillars and generating occupancy information for each pillar. Generally, for a two-dimensional grid, it is sufficient to aggregate the occ_voxels falling therein into one pillar; however, when dealing with special suspended obstacles (for example, double-layer objects), since one pillar cannot fully and accurately describe the convexity and concavity in the z-axis direction occupied by the two-dimensional grid, the occ_voxel falling into the two-dimensional grid can be expressed as two pillars. Figure 3 shows a schematic diagram of aggregating and expressing voxels into one pillar according to one or more embodiments of the present application. It should be noted that the voxel shown in the figure is occ_voxel, and the expression of voxel and pillar is not limited to the specific shapes and annotation information shown in the figure.

[0052] In step S220 , the voxels are filtered. In one example, unnecessary types of occ_voxels may be filtered out based on semantic categories, for example, occ_voxels with semantic meanings of ground are filtered out.

[0053] In another example, the reserved range of occ_voxel can be determined based on factors such as vehicle height, field of view, slope, etc. For example, from the perspective of vehicle autonomous driving safety and obstacle avoidance, if the vehicle is traveling at a speed of 80km / h and -3.5m / s 2When driving at a deceleration rate, the forward obstacle avoidance distance is approximately 75m. Considering the uphill field of view, for example, when the road slope is 10%, as shown in Figure 4, the horizontal distance in this scenario is approximately 75m, and the vertical distance is approximately 7.5m. Given an additional height of approximately 2m (i.e., vehicle height), the vertical angle is approximately 10.204° and the maximum vertical obstacle avoidance distance is approximately 9.5m. Further considering the left and right rearward fields of view when driving on a slope, the minimum vertical obstacle avoidance distance is approximately -2.5m. In summary, a z-axis perception range of approximately -2.5m to 9.5m is required. However, designing a pure vision network to cover this large range (e.g., a z-axis span of approximately 12m) faces the following challenges: excessive interference in the z-axis direction, such as height limit poles on the road and the ceiling of an underground parking lot; and due to the excessive amount of redundant data, the inference phase consumes a large amount of computing power if a 3D convolutional network is used. To this end, one or more embodiments of the present application filter out occ_voxels whose z-axis coordinates in the vehicle coordinate system are less than the first height threshold z1 or greater than the second height threshold z2, leaving only occ_voxels with z∈[z1, z2]. For example, since the height of the vehicle is approximately 2m, only occ_voxels with z∈[-1m, 2m] can be left to filter out irrelevant information interference while ensuring the obstacle avoidance function. Specifically, the advantages of this type of design are: it can remove most of the suspended obstacles, reducing the necessity of expressing the occ_voxels that fall into a two-dimensional grid as two pillars, and because the vehicle's own height is limited, it does not affect safe obstacle avoidance; after reducing the interfering voxels, the network is easier to learn with the same number of parameters.

[0054] In step S230, it is determined whether the number of occ_voxels falling within the pillar is greater than or equal to a third threshold. If the number of occ_voxels is greater than or equal to the third threshold, the pillar is marked as occupied, and the upper and lower boundaries of the pillar's height are marked as the maximum height and minimum height of the one or more occupied voxels falling therein, respectively (step S240). If the number of occupied voxels falling within the pillar is less than the third threshold, the pillar is marked as unoccupied, and the upper and lower boundaries of the pillar's height are marked as 0 (step S250). For example, assume that each occ_voxel that falls within a two-dimensional grid is aggregated and expressed as one pillar. Assume that the voxel resolution is 0.1m, the grid resolution is 0.2m, and only the occ_voxels z∈[-1m, 2m] are left in the previous step. There are 120 voxels in each grid. If the number of remaining occ_voxels in each grid is less than 3, the pillar is marked as unoccupied; otherwise, it is marked as occupied, and the upper and lower boundaries of the pillar height correspond to the maximum and minimum heights of the occ_voxels within it, respectively. The height here can be the absolute height in the vehicle coordinate system or relative to the ground.

[0055] In step S260 , the occupancy information of the pillar is calculated. Exemplarily, the occupancy information includes one or more of the following: coordinate information, semantic category, speed level, and speed value.

[0056] Exemplarily, the coordinate information of the pillar includes one of the following: the coordinates of the pillar center point and the pillar height (pillar center, pillar height) in the vehicle coordinate system; the upper and lower boundaries of the pillar height (pillar height max, pillar height min) in the vehicle coordinate system; the coordinates of the pillar center point, the pillar height and the ground height (pillar center, pillar height, ground height) in the vehicle coordinate system; and the upper and lower boundaries of the pillar height and the ground height (pillar height max, pillar height min, ground height) in the vehicle coordinate system.

[0057] For example, the speed of the pillar can be marked as the average speed of the occ_voxels falling within it. Alternatively, if the average speed of the occ_voxels falling within the pillar is greater than or equal to a fourth speed threshold, the speed of the pillar can be marked as dynamic, otherwise it can be marked as static.

[0058] For example, a pillar's semantic category can be determined using a voting mechanism based on the semantic categories of the occ_voxels within the pillar. The voting mechanism in this article refers to a hard voting mechanism, where the final decision is determined according to the majority rule. For example, if a pillar contains three occ_voxels, two of which are classified as electric vehicles and one as pedestrian, the majority rule determines the semantic category of the pillar as electric vehicles.

[0059] Returning to FIG. 1 , after the voxel-to-pillar conversion is achieved, in step S133 , a grid prediction model is trained using a training dataset containing sample images and true value information of the sample images.

[0060] FIG5 shows a schematic diagram of a network design according to one or more embodiments of the present application. Exemplarily, in the network design, the 2D backbone part and the 2D-3D conversion module can adopt the BEV Depth scheme, the BEV Former scheme, the BEV Det scheme, etc. Since the data has been converted into pillar form, the head part in the model can adopt 2D convolution instead of 3D convolution, wherein the more convolutions in the head part, the more times the amount of computation saved by this scheme relative to the output of the voxel model. In addition, since the pillar form saves several times the amount of data processing of the z-axis voxel compared to the voxel form, the amount of computation of the head part is significantly reduced.

[0061] Figure 6 illustrates a schematic diagram of model input and output according to one or more embodiments of the present application. The three views shown in the top portion of Figure 6 correspond to 2D purely visual images captured by the vehicle's left front-facing camera, front-facing camera, and right front-facing camera, respectively. The bottom portion of Figure 6 illustrates the pillar-like occupancy results output by the grid prediction model, where the ego vehicle is identified by a cube, and objects in the purely visual image (e.g., a large truck, a tree) are identified as pillars.

[0062] The occupancy grid prediction method 10 according to one or more embodiments of the present application reduces the significant transmission pressure that would otherwise be incurred when the model output data is directly transmitted downstream or visualized in mass-produced vehicles. Furthermore, because the model output of the occupancy grid prediction method 10 is directly in pillar form, the process of post-processing the voxel output to generate a pillar output can be omitted. Furthermore, when the post-processing process does not require high data precision and there is a lot of redundant information, for example, when only obstacle edge information needs to be retained or when time series fusion processing is required, the computational complexity of the pillar form is much smaller than that of the voxel form.

[0063] Figure 7 is a schematic block diagram of an occupancy grid prediction device 70 according to one or more embodiments of the present application. The occupancy grid prediction device 70 includes a memory 710, a processor 720, and a computer program 730 stored on the memory 710 and executable on the processor 720. The execution of the computer program 730 enables the occupancy grid prediction method 10 shown in Figure 1 to be executed. Exemplarily, the occupancy grid prediction device 70 can be part of the electronic control unit (ECU) of the vehicle system, or the control unit of the autonomous driving system. Exemplarily, the occupancy grid prediction device 70 can also be a control unit in other terminals (e.g., smartphones, or other edge devices with limited computing power) that can communicate with the vehicle.

[0064] Figure 8 is a schematic block diagram of an intelligent device 80 according to one or more embodiments of the present application. The intelligent device 80 has an occupancy grid prediction device 70 as shown in Figure 7. In some embodiments of the present application, the intelligent device 80 further includes at least one sensor, such as a video acquisition unit, which is used to perceive information. The sensor is communicatively connected to any type of processor mentioned in the present application. Optionally, the intelligent device 80 may also include an automatic driving system, which is used to guide the intelligent device to drive independently or assist in driving. The processor communicates with the sensor and / or the automatic driving system to complete the method described in any of the above embodiments. Exemplarily, the intelligent device 80 may include driving equipment, smart cars, electric cars, robots and other devices.

[0065] In addition, as described above, the present application can also be implemented as a computer storage medium having a program stored therein for causing a computer to execute the occupancy grid prediction method 10 shown in FIG1 . Here, as the computer storage medium, various computer storage media such as disks (e.g., magnetic disks, optical disks, etc.), cards (e.g., memory cards, optical cards, etc.), semiconductor memories (e.g., ROMs, non-volatile memories, etc.), and tapes (e.g., magnetic tapes, cassettes, etc.) can be used.

[0066] In the applicable situation, the combination of hardware, software or hardware and software can be used to realize the various embodiments provided by the application. Moreover, in the applicable situation, without departing from the scope of the application, the various hardware components and / or software components set forth herein can be combined into a composite component comprising software, hardware and / or both. In the applicable situation, without departing from the scope of the application, the various hardware components and / or software components set forth herein can be divided into a subcomponent comprising software, hardware or both. In addition, in the applicable situation, it is contemplated that the software component can be implemented as a hardware component, and vice versa.

[0067] Software according to the present application (such as program code and / or data) can be stored on one or more computer storage media. It is also contemplated that the software identified herein can be implemented using one or more general or special computers and / or computer systems, networked and / or otherwise. Where applicable, the order of the various steps described herein can be changed, combined into composite steps and / or divided into sub-steps to provide the features described herein.

[0068] The relevant user personal information that may be involved in the various embodiments of this application is strictly in accordance with the requirements of laws and regulations, following the principles of legality, legitimacy and necessity, and based on the reasonable purposes of business scenarios, to process the personal information that users actively provide during the use of products / services or generated due to the use of products / services, as well as the personal information obtained with the user's authorization.

[0069] The user personal information processed in this application will vary depending on the specific product / service scenario and will be based on the specific scenario in which the user uses the product / service. This may involve the user's account information, device information, driving information, vehicle information, or other related information. The applicant will treat the user's personal information and its processing with a high degree of diligence.

[0070] This application attaches great importance to the security of user personal information and has taken reasonable and feasible security protection measures that comply with industry standards to protect user information and prevent personal information from being accessed, disclosed, used, modified, damaged or lost without authorization.

[0071] The embodiments and examples set forth herein are provided to best illustrate embodiments according to the present application and its specific applications, and thereby enable those skilled in the art to make and use the present application. However, those skilled in the art will appreciate that the above description and examples are provided for ease of illustration and example only. The descriptions set forth are not intended to be exhaustive of all aspects of the present application or to limit the present application to the precise forms disclosed.

Claims

1. A method for predicting an occupied grid, characterized in that: The method comprises the following steps: Get a purely visual image of the vehicle's surroundings; inputting the visual-only image into an occupancy grid prediction model; and The occupancy grid prediction model is used to determine the occupancy result of the target in the pure visual image as the model output, wherein the model output is the occupancy result in the form of a cylinder, and the occupancy grid prediction model is constructed based on a training data set including sample images and true value information of the sample images, and the true value information is directly generated by point cloud data or converted from occupancy information in the form of voxels generated based on point cloud data.

2. The occupancy grid prediction method according to claim 1, wherein: The voxel-based occupancy information is generated based on at least the following steps: The three-dimensional space is divided into regular voxel grids and the point cloud data in each voxel grid is quantified into occupancy information, wherein the occupancy information includes an occupancy state, a semantic category, a speed level, a speed value, or a combination thereof of the voxel.

3. The occupancy grid prediction method according to claim 1 or 2, wherein: The true value information is converted from the occupancy information in the form of voxels generated based on the point cloud data and includes: Offline data processing is performed on the occupancy information in the form of voxels to generate compressed occupancy information in the form of cylinders, and the occupancy information in the form of voxels is generated based on one or more of the following items: single-frame laser point cloud data, time-series spliced ​​laser point cloud data, and point cloud data reconstructed by visual three-dimensional reconstruction.

4. The occupancy grid prediction method according to claim 3, wherein: Offline data processing of the occupancy information in voxel form includes: Selecting voxels whose occupancy status is occupied as occupied voxels; Project the three-dimensional space into a two-dimensional grid; Aggregating one or more occupied voxels that fall into a two-dimensional grid into one or two cylinders; and Generate occupancy information for each column.

5. The occupancy grid prediction method according to claim 4, wherein: Projecting the three-dimensional space into the two-dimensional grid comprises: In the vehicle coordinate system, the x-axis and the y-axis are divided into a regular two-dimensional grid according to the first resolution, wherein the positive direction of the x-axis is the forward direction of the vehicle, and the positive direction of the y-axis is the left side of the vehicle; and The three-dimensional space in which the xy plane falls within the two-dimensional grid is projected onto the two-dimensional grid, wherein the xy plane is a plane formed by the x-axis and the y-axis.

6. The occupancy grid prediction method according to claim 4, wherein: Offline data processing of the occupancy information in voxel form further includes: Filtering out occupied voxels of unwanted types based on semantic categories; and / or Occupied voxels whose z-axis coordinates in the ego-vehicle coordinate system are less than the first height threshold or greater than the second height threshold are filtered out.

7. The occupancy grid prediction method according to claim 4, wherein: Generating the occupancy information of each column includes: If the number of occupied voxels falling within the cylinder is greater than or equal to a third number threshold, marking the cylinder as occupied, and marking the upper and lower boundaries of the height of the cylinder as the maximum height and minimum height of one or more occupied voxels falling therein, respectively; and If the number of occupied voxels falling into the cylinder is less than a third number threshold, the cylinder is marked as unoccupied, and the upper and lower boundaries of the height of the cylinder are marked as zero.

8. The occupancy grid prediction method according to claim 4, wherein: The occupancy information of the column includes coordinate information of the column, and the coordinate information includes one of the following: The coordinates of the center point of the cylinder and the height of the cylinder in the vehicle coordinate system; The upper and lower boundaries of the column height in the vehicle coordinate system; The center coordinates of the cylinder, the height of the cylinder and the height above the ground in the vehicle coordinate system; or The upper and lower boundaries of the cylinder height and the ground height in the vehicle coordinate system.

9. The occupancy grid prediction method according to claim 4, wherein: Generating the occupancy information of each column includes: The velocity of the cylinder is labeled as the average velocity of one or more occupied voxels falling within the cylinder. degree; or If the average speed of one or more occupied voxels falling within the column is greater than or equal to a fourth speed threshold, the speed of the column is marked as dynamic, otherwise it is marked as static.

10. The occupancy grid prediction method according to claim 4, wherein: The occupancy information generated for each column includes: The semantic category of the cylinder is determined using a voting mechanism based on the semantic categories of one or more occupied voxels falling within the cylinder.

11. The occupancy grid prediction method according to claim 1, wherein: Two-dimensional convolution is used in the head part of the occupancy grid prediction model.

12. An occupancy grid prediction device, characterized in that: The invention comprises: a memory; a processor; and a computer program stored in the memory and executable on the processor, wherein the execution of the computer program enables the occupancy grid prediction method according to any one of claims 1 to 11 to be executed.

13. A smart device, characterized in that: A device for predicting an occupied grid according to claim 12 is provided.

14. A computer storage medium, characterized in that: The computer storage medium comprises instructions, which when executed, perform the occupancy grid prediction method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Occupancy grid prediction method and device, intelligent equipment and storage medium

    CN117274941A

  • Quick three-dimensional grid construction system and method, medium, equipment and unmanned vehicle

    CN113160398A

  • Driving scene simulation method, system and equipment based on three-dimensional occupation grid and medium

    CN116452766A

Cited By

  • Target detection method, device and equipment

    CN120807899A

  • Semantic map occupation prediction method and device, storage medium and electronic device

    CN120997431A

  • Method and device for segmenting and extracting rod-shaped object target

    CN121505393A

  • Three-dimensional perception model training and application method, device, equipment and robot

    CN122114046A