Vehicle and three-dimensional occupancy prediction method thereof, and training method of bidirectional decoder

The bidirectional decoder uses the contextual information of future frames to correct the prediction of the current frame in vehicle 3D occupancy prediction, solving the problem of reduced accuracy caused by missing or occluded historical frames and achieving higher prediction accuracy.

CN120808309APending Publication Date: 2025-10-17CHERY AUTOMOBILE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510934405.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In the existing technology, the missing or occluded historical frames lead to reduced accuracy of vehicle 3D occupancy prediction, resulting in problems of false detection or missed detection.

Method used

A bidirectional decoder is used to perform feature interaction on the 3D fusion features, the initial query vectors of the current frame and the future frame to generate a target query vector. The 3D occupancy prediction result is generated through 3D reconstruction, and the prediction of the current frame is corrected using the context information of the future frame.

Benefits of technology

The accuracy of 3D occupancy prediction is improved, which avoids false detection or missed detection due to occlusion of historical frames and improves the accuracy of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808309A_ABST
    Figure CN120808309A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a vehicle, a three-dimensional occupancy prediction method thereof and a training method of a bidirectional decoder, and the three-dimensional occupancy prediction method of the vehicle comprises the steps: sensing a three-dimensional space where the vehicle is located, and obtaining a multi-frame image, the multi-frame image comprising a current frame image and a preset number of historical frame images; generating a three-dimensional fusion feature based on the two-dimensional feature of the multi-frame image; performing feature interaction on the three-dimensional fusion feature, the initial query vector of the current frame and the initial query vector of the future frame by using a bidirectional decoder to obtain a target query vector of the current frame and a target query vector of the future frame; and performing three-dimensional reconstruction based on the target query vector of the current frame to generate a three-dimensional occupation prediction result of the current frame, and performing three-dimensional reconstruction based on the target query vector of the future frame to generate a three-dimensional occupation prediction result of the future frame. According to the invention, the technical problem of low accuracy of 3D occupancy prediction of the current frame by using the historical frame is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the fields of automatic driving, artificial intelligence and image processing, in particular, relate to a vehicle and a three-dimensional occupancy prediction method thereof and a training method of a bidirectional decoder. BACKGROUND

[0002] In the field of automatic driving, especially in the scene of environment perception, 3D (three-dimensional) occupancy prediction is an important link to realize safe driving of a vehicle. The 3D occupancy prediction can refer to a process of predicting the occupancy state of different voxels in the environment around the vehicle, that is, predicting the 3D occupancy state. Through the 3D occupancy prediction, the automatic driving system can help to identify the obstacles around the vehicle, so as to reasonably plan the driving path of the vehicle, and the vehicle can safely and efficiently drive while avoiding obstacles.

[0003] In order to realize the 3D occupancy prediction, currently, a plurality of deep learning models can be used for 3D occupancy prediction, and the 3D occupancy state of a current frame is usually predicted by using historical frames. However, due to many disturbances in the actual driving process of the vehicle, the historical frames can be missing or occluded, and in this case, the accuracy of the 3D occupancy prediction result is reduced, and there are problems of false detection or missed detection.

[0004] At present, there is no good solution to the above problems. SUMMARY

[0005] Embodiments of the present application provide a vehicle and a three-dimensional occupancy prediction method thereof and a training method of a bidirectional decoder, to at least solve the technical problem of low accuracy of 3D occupancy prediction of a current frame by using historical frames.

[0006] According to an aspect of an embodiment of the present application, a three-dimensional occupancy prediction method of a vehicle is provided, comprising: perceiving a three-dimensional space where the vehicle is located to obtain a plurality of frames of images, wherein the plurality of frames of images include a current frame of image and a preset number of historical frames of images; generating three-dimensional fusion features based on two-dimensional features of the plurality of frames of images; using a bidirectional decoder to perform feature interaction on the three-dimensional fusion features, an initial query vector of the current frame and an initial query vector of a future frame to obtain a target query vector of the current frame and a target query vector of the future frame, wherein the future frame is a next frame of the current frame; performing three-dimensional reconstruction based on the target query vector of the current frame to generate a three-dimensional occupancy prediction result of the current frame, and performing three-dimensional reconstruction based on the target query vector of the future frame to generate a three-dimensional occupancy prediction result of the future frame, wherein the three-dimensional occupancy prediction result is used to represent whether a plurality of voxels in the three-dimensional space correspond to an occupancy probability greater than a preset probability, and the occupancy probability is a probability that the voxel is occupied by an object.

[0007] Further, the three-dimensional fusion feature, the initial query vector of the current frame and the initial query vector of the future frame are interacted by using the bidirectional decoder to obtain the target query vector of the current frame and the target query vector of the future frame, including: performing attention processing on the three-dimensional fusion feature and the initial query vector of the current frame to obtain the attention feature of the current frame; performing attention processing on the three-dimensional fusion feature and the initial query vector of the future frame to obtain the attention feature of the future frame; and performing attention processing on the attention feature of the current frame and the attention feature of the future frame to obtain the target query vector of the current frame and the target query vector of the future frame.

[0008] Further, the attention processing is performed on the attention feature of the current frame and the attention feature of the future frame to obtain the target query vector of the current frame and the target query vector of the future frame, including: splicing the attention feature of the current frame and the attention feature of the future frame to obtain spliced features; performing attention processing on the spliced features to obtain target attention features; and splitting the target attention features to obtain the target query vector of the current frame and the target query vector of the future frame.

[0009] Further, based on the two-dimensional features of the multiple frames of images, the three-dimensional fusion feature is generated, including: respectively extracting features of the multiple frames of images to obtain two-dimensional features of the multiple frames of images; respectively performing three-dimensional reconstruction on the two-dimensional features of the multiple frames of images to obtain three-dimensional features of the multiple frames of images; and using a space-time encoder to fuse the three-dimensional features of the multiple frames of images to obtain the three-dimensional fusion feature.

[0010] Further, the three-dimensional space where the vehicle is located is perceived to obtain the multiple frames of images, including: obtaining historical frame images collected by a sensor on the vehicle; and in response to the number of historical frame images collected by the sensor being less than a preset number, generating a preset number of historical frame images based on the historical frame images collected by the sensor.

[0011] According to another aspect of the embodiments of the present application, a training method of a bidirectional decoder is also provided, including: obtaining training data corresponding to a three-dimensional space in which a vehicle is located, wherein the training data includes: a plurality of image samples, a three-dimensional occupancy labeling result of a current frame, and a three-dimensional occupancy labeling result of a future frame, and the three-dimensional occupancy labeling result is used to represent whether a voxel in the three-dimensional space is occupied by an object; generating a three-dimensional sample feature based on a two-dimensional sample feature of the plurality of image samples; performing feature interaction on the three-dimensional sample feature, a first query vector of the current frame, and a first query vector of the future frame by using an initial decoder to obtain a second query vector of the current frame and a second query vector of the future frame; performing three-dimensional reconstruction based on the second query vector of the current frame to generate a three-dimensional occupancy prediction result of the current frame, and performing three-dimensional reconstruction based on the second query vector of the future frame to generate a three-dimensional occupancy prediction result of the future frame, wherein the three-dimensional occupancy prediction result is used to represent whether a plurality of occupancy probabilities corresponding to voxels in the three-dimensional space are greater than a preset probability, and the occupancy probability is a probability that a voxel is occupied by an object; and adjusting parameters of the initial decoder based on the three-dimensional occupancy prediction result of the current frame, the three-dimensional occupancy labeling result of the current frame, the three-dimensional occupancy prediction result of the future frame, and the three-dimensional occupancy labeling result of the future frame to obtain the bidirectional decoder.

[0012] Further, the feature interaction on the three-dimensional sample feature, the first query vector of the current frame, and the first query vector of the future frame by using the initial decoder to obtain the second query vector of the current frame and the second query vector of the future frame includes: performing attention processing on the three-dimensional sample feature and the first query vector of the current frame to obtain an attention feature of the current frame; performing attention processing on the three-dimensional sample feature and the first query vector of the future frame to obtain an attention feature of the future frame; and performing attention processing on the attention feature of the current frame and the attention feature of the future frame to obtain the second query vector of the current frame and the second query vector of the future frame.

[0013] Further, the attention processing on the attention feature of the current frame and the attention feature of the future frame to obtain the second query vector of the current frame and the second query vector of the future frame includes: splicing the attention feature of the current frame and the attention feature of the future frame to obtain spliced features; performing attention processing on the spliced features to obtain target attention features; and splitting the target attention features to obtain the second query vector of the current frame and the second query vector of the future frame.

[0014] Further, the generation of the three-dimensional sample feature based on the two-dimensional sample feature of the plurality of image samples includes: performing feature extraction on the plurality of image samples respectively to obtain the two-dimensional sample feature of the plurality of image samples; performing three-dimensional reconstruction on the two-dimensional sample feature of the plurality of image samples respectively to obtain three-dimensional features of the plurality of images; and fusing the three-dimensional features of the plurality of images by using a space-time encoder to obtain the three-dimensional sample feature.

[0015] Further, based on the two-dimensional sample features of the multi-frame image samples, the three-dimensional sample features are generated, including: determining a current training round of the initial decoder; in a case that the current training round is less than or equal to a first preset round, generating the three-dimensional sample features based on the two-dimensional sample features of the multi-frame image samples; in a case that the current training round is greater than the first preset round, performing random mask processing on the multi-frame image samples to obtain multi-frame processed image samples, and generating the three-dimensional sample features based on the two-dimensional sample features of the multi-frame processed image samples.

[0016] Further, the random mask processing on the multi-frame image samples to obtain multi-frame processed image samples includes: determining a random mask probability based on the current training round; performing random mask processing on any one historical frame image sample in the multi-frame image samples based on the random mask probability to obtain the multi-frame processed image samples.

[0017] Further, the training data corresponding to the three-dimensional space where the vehicle is located is obtained, including: in response to the number of historical frame image samples in the multi-frame image samples being less than a preset number, generating a preset number of historical frame image samples based on the historical frame image samples.

[0018] Further, based on the three-dimensional occupancy prediction result of the current frame, the three-dimensional occupancy labeling result of the current frame, the three-dimensional occupancy prediction result of the future frame, and the three-dimensional occupancy labeling result of the future frame, the parameters of the initial decoder are adjusted to obtain a bidirectional decoder, including: based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy labeling result of the current frame, constructing a classification prediction loss function value and an edge perception loss function value of the current frame; based on the three-dimensional occupancy prediction result of the future frame and the three-dimensional occupancy labeling result of the future frame, constructing a classification prediction loss function value and an edge perception loss function value of the future frame; based on the three-dimensional features of the multi-frame image samples, constructing a sparse regularization loss function value; based on the predicted depth distribution of the multi-frame image and the preset depth distribution of the multi-frame image, constructing a depth perception loss function value; weighting the classification prediction loss function value and the edge perception loss function value of the current frame, the classification prediction loss function value and the edge perception loss function value of the future frame, the sparse regularization loss function value, and the depth perception loss function value to obtain a total loss function value of the initial decoder; adjusting the parameters of the initial decoder based on the total loss function value to obtain the bidirectional decoder.

[0019] Further, the classification prediction loss function value and the edge perception loss function value of the current frame, the classification prediction loss function value and the edge perception loss function value of the future frame, the sparse regularization loss function value, and the depth perception loss function value are weighted to obtain a total loss function value of the initial decoder, including: constructing a consistency loss function value based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame; and weighting the classification prediction loss function value and the edge perception loss function value of the current frame, the classification prediction loss function value and the edge perception loss function value of the future frame, the sparse regularization loss function value, the depth perception loss function value, and the consistency loss function value to obtain the total loss function value.

[0020] Further, the consistency loss function value is constructed based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame, including: determining a current training round of the initial decoder; and in a case where the current training round is greater than a second preset round, constructing the consistency loss function value based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame.

[0021] According to another aspect of the embodiments of the present application, a vehicle path planning method is also provided, including: perceiving a three-dimensional space where a vehicle is located to obtain multiple frames of images, wherein the multiple frames of images include a current frame of image and a preset number of historical frame of images; generating three-dimensional fusion features based on two-dimensional features of the multiple frames of images; performing feature interaction on the three-dimensional fusion features, an initial query vector of the current frame, and an initial query vector of a future frame by using a bidirectional decoder to obtain a target query vector of the current frame and a target query vector of the future frame, wherein the future frame is a next frame of the current frame; performing three-dimensional reconstruction based on the target query vector of the current frame to generate a three-dimensional occupancy prediction result of the current frame, and performing three-dimensional reconstruction based on the target query vector of the future frame to generate a three-dimensional occupancy prediction result of the future frame, wherein the three-dimensional occupancy prediction result is used to represent whether a plurality of voxels in the three-dimensional space correspond to an occupancy probability greater than a preset probability, and the occupancy probability is a probability that a voxel is occupied by an object; and generating a driving path of the vehicle based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame.

[0022] According to another aspect of the embodiments of the present application, a vehicle is also provided, including: a memory storing an executable program; and a processor configured to execute the program, wherein the program performs the method in the embodiments of the present application when executed.

[0023] According to another aspect of the embodiments of the present application, a computer readable storage medium is also provided, including a stored executable program, wherein the computer readable storage medium controls a device where the computer readable storage medium is located to perform the method in the embodiments of the present application when the executable program is executed.

[0024] According to another aspect of the embodiments of the present application, a computer program product is also provided, which comprises a computer program, and the computer program implements the method in each of the embodiments of the present application when executed by a processor.

[0025] According to another aspect of the embodiments of the present application, a computer program product is also provided, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program implements the method in each of the embodiments of the present application when executed by a processor.

[0026] According to another aspect of the embodiments of the present application, a computer program is also provided, and the computer program implements the method in each of the embodiments of the present application when executed by a processor.

[0027] In the embodiments of the present application, a three-dimensional space where a vehicle is located is perceived to obtain multiple frames of images; three-dimensional fusion features are generated based on two-dimensional features of the multiple frames of images; the three-dimensional fusion features, an initial query vector of a current frame and an initial query vector of a future frame are interacted by using a bidirectional decoder to obtain a target query vector of the current frame and a target query vector of the future frame; three-dimensional reconstruction is performed based on the target query vector of the current frame to generate a three-dimensional occupancy prediction result of the current frame, and three-dimensional reconstruction is performed based on the target query vector of the future frame to generate a three-dimensional occupancy prediction result of the future frame, thereby achieving the purpose of 3D occupancy prediction. It is easy to note that in the process of 3D occupancy prediction, the three-dimensional fusion features, the initial query vector of the current frame and the initial query vector of the future frame can be interacted by using the bidirectional decoder, that is, the initial query vector of the current frame can be corrected by using the context information of the future frame, so that the 3D occupancy prediction of the current frame considers the 3D occupancy state of the historical frame and the 3D occupancy state of the future frame at the same time, and no longer only depends on the 3D occupancy state of the historical frame, thereby avoiding false detection or missed detection caused by the historical frame being blocked, achieving the technical effect of improving the accuracy of 3D occupancy prediction, and solving the technical problem of low accuracy of 3D occupancy prediction of the current frame by using the historical frame. BRIEF DESCRIPTION OF DRAWINGS

[0028] The accompanying drawings, which are included to provide a further understanding of the present application and are incorporated in and constitute a part of this application, illustrate embodiments of the present application and serve to explain the present application. In the drawings:

[0029] Figure 1 is a flowchart of a three-dimensional occupancy prediction method of a vehicle according to an embodiment of the present application;

[0030] Figure 2 is a flowchart of a training method of a bidirectional decoder according to an embodiment of the present application;

[0031] Figure 3 is a schematic diagram of an optional random mask processing according to an embodiment of the application;

[0032] Figure 4 is a flowchart of an optional three-dimensional occupancy prediction method of a vehicle according to an embodiment of the application;

[0033] Figure 5 is a flowchart of an optional Gaussian-voxel hybrid representation flow according to an embodiment of the application;

[0034] Figure 6 is a flowchart of an optional decoding process of a bidirectional decoder according to an embodiment of the application;

[0035] Figure 7 is a flowchart of an optional processing process of a bidirectional attention mechanism according to an embodiment of the application;

[0036] Figure 8 is a flowchart of a vehicle path planning method according to an embodiment of the application;

[0037] Figure 9 is a schematic diagram of a three-dimensional occupancy prediction device of a vehicle according to an embodiment of the application;

[0038] Figure 10 is a schematic diagram of a training device of a bidirectional decoder according to an embodiment of the application;

[0039] Figure 11 is a schematic diagram of a vehicle path planning device according to an embodiment of the application. DETAILED DESCRIPTION

[0040] In order to enable persons skilled in the art to better understand the application scheme, the technical solutions in the embodiments of the application will be described clearly and completely below in conjunction with the drawings in the embodiments of the application. Obviously, the described embodiments are only a part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by persons skilled in the art without creative labor should belong to the scope of protection of the application.

[0041] It should be noted that the terms "first", "second", and the like in the description and in the claims of the present application and the above-described accompanying drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or apparatus including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to such a process, method, product, or apparatus.

[0042] According to an embodiment of the present application, a three-dimensional occupancy prediction method of a vehicle is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0043] In the present embodiment, a three-dimensional occupancy prediction method of a vehicle is provided, Figure 1 is a flowchart of a three-dimensional occupancy prediction method of a vehicle according to an embodiment of the present application, which can be executed by a vehicle or a server, and in the present embodiment, the execution by a vehicle is taken as an example for illustration. As shown in Figure 1 the flowchart includes the following steps:

[0044] In step S102, the three-dimensional space in which the vehicle is located is perceived to obtain a plurality of frames of images, wherein the plurality of frames of images include a current frame of image and a preset number of historical frames of images.

[0045] The vehicle in the above steps can be a vehicle that needs to perform 3D occupancy prediction. From the perspective of driving mode, it can be an autonomous vehicle or a driver-driven vehicle; from the perspective of energy, it can be a fuel vehicle, a pure electric vehicle, or a hybrid vehicle; and from the perspective of specific type, it can be a car, a truck, a bus, etc., but is not limited thereto.

[0046] The three-dimensional space in the above steps can be the real environment in which the vehicle is located, i.e., the environment around the vehicle, especially the environment in front of the vehicle, which often contains static obstacles such as road signs, buildings, trees, etc., and dynamic obstacles such as other vehicles, pedestrians, animals, etc., but is not limited thereto.

[0047] The multiple frames of images in the above steps can be obtained by continuously shooting the vehicle's surroundings multiple times through the on-board camera while the vehicle is driving. Therefore, based on the shooting time, it can be determined that the last image shot is the image shot at the current moment, that is, the last image is the current frame image, and the images before this image are all historical frame images. In this case, it can be determined that the number of historical frame images is the same as the preset number. In order to ensure the accuracy of the 3D occupancy prediction, the larger the preset number here, the better. However, the more historical frame images there are, the more time it takes to process, which will lead to a decrease in prediction efficiency. Therefore, the preset number can be determined based on the actual processing performance of the vehicle through experiments. For example, the preset number can be 1-3, but is not limited to this. It should be noted that when the vehicle just starts to drive, the number of images of the vehicle's surroundings shot by the on-board camera is limited. In order to meet the needs of 3D occupancy prediction, the historical frame images can be supplemented through certain image processing strategies. For example, at the moment a vehicle starts, the onboard camera can only capture one image, which is the current frame image, but lacks historical frame images. Therefore, it is possible to consider capturing an image of the vehicle's surroundings with the onboard camera when the vehicle is powered on, and using this image as the historical frame image. Multiple copies of this image can then be made through image replication to obtain a preset number of historical frames. It should also be noted that, given the interference of the vehicle's surrounding environment, images captured by the onboard camera may be missing or blocked, making them unsuitable for use as historical frame images. This can result in the number of historical frame images being less than the preset number. In this case, a strategy can be employed to complete the historical frame images, such as replicating or interpolating the perceived historical frame images, to ensure that the preset number of historical frame images is present when 3D occupancy prediction is performed. Furthermore, vehicles are often equipped with multiple onboard cameras with different perspectives. To ensure that multiple frames accurately reflect the vehicle's surroundings, multiple onboard cameras can be used to capture images simultaneously, and the multi-perspective images captured at the same time can be stitched together to form a single frame image.

[0048] In an optional embodiment, after the vehicle begins driving, an onboard camera can capture images of the vehicle's surroundings, capturing at least one frame of imagery. This imagery is then processed using a specific image processing strategy to produce a multi-frame image consisting of a current frame image and a preset number of historical frame images. If only the current frame image is captured, and no historical frame images are available, an image captured of the vehicle's surroundings when the vehicle was powered on can be read and replicated multiple times to obtain a preset number of historical frame images. If a small number of historical frame images are captured, these historical frame images can be processed by replication or linear interpolation to obtain a preset number of historical frame images.

[0049] In this application, considering a special case, the vehicle-mounted camera only collects one historical frame image, or reads the image collected by the vehicle-mounted camera when the vehicle is powered on, at this time it can be considered that the multi-frame image frame only contains the current frame image and one historical frame image, in this case, the historical frame image can not be expanded, but the one historical frame image is used for 3D occupancy prediction, this case is called single frame backup mode. In order to improve the accuracy of 3D occupancy prediction in this mode, during the training of the bidirectional decoder, training data can be constructed for this mode, and the constructed training data can be used for training to ensure that the bidirectional decoder can decode features for the above-mentioned case.

[0050] Step S104, generating three-dimensional fusion features based on two-dimensional features of multiple frame images.

[0051] The two-dimensional features in the above steps can be extracted from the multi-frame image, which can represent the features of objects at different positions in the vehicle surrounding environment. In order to more accurately represent the real situation of the vehicle surrounding environment, the two-dimensional features can be multi-scale image features, but not limited to this. The three-dimensional fusion features in the above steps can be obtained based on the two-dimensional features, which can represent the features of objects at different positions in the vehicle surrounding environment, and can accurately reflect the accurate position of the object in the three-dimensional space.

[0052] In an optional embodiment, since the image captured by the vehicle-mounted camera is a two-dimensional image, and the vehicle surrounding environment is often a three-dimensional space, the two-dimensional features can be converted into three-dimensional fusion features through various algorithms. For example, the two-dimensional features can be three-dimensionally reconstructed based on depth information to obtain three-dimensional fusion features, for example, the depth information in the two-dimensional features can be predicted through a deep learning model to generate three-dimensional fusion features, for example, the two-dimensional features can be converted into three-dimensional fusion features through coordinate transformation, Gaussian-voxel hybrid representation and the like. The specific generation method of the three-dimensional fusion features is not limited in this embodiment, and can be determined according to actual needs.

[0053] Step S106, using the bidirectional decoder to perform feature interaction on the three-dimensional fusion features, the initial query vector of the current frame and the initial query vector of the future frame to obtain the target query vector of the current frame and the target query vector of the future frame, wherein the future frame is the next frame of the current frame.

[0054] The bidirectional decoder in the above step can be a decoder pre-trained based on an attention mechanism. The bidirectional herein means that the 3D occupancy state of the historical frame can be used to predict the 3D occupancy state of the current frame and further predict the 3D occupancy state of the future frame, that is, the next frame of the current frame, and the 3D occupancy state of the future frame can also be used to predict the 3D occupancy state of the current frame. Therefore, by training the bidirectional decoder, a "past-present-future" bidirectional space-time dependence model can be constructed, and the historical frame and the future context are considered simultaneously in the 3D occupancy prediction process of the current frame. The specific training process of the bidirectional decoder can adopt the training method described in the following embodiments, which will not be described here.

[0055] The initial query vectors (including the initial query vector of the current frame and the initial query vector of the future frame) in the above step can refer to the query vectors in the attention mechanism, which are used to represent the 3D occupancy state and can be continuously learned and updated in the process iteratively performed by the bidirectional decoder, so as to obtain the target query vectors that can more accurately reflect the 3D occupancy state. It should be noted that the initial query vector can be a query vector obtained by initialization. The initialization can be to set the initial query vector as a random vector, or to generate the initial query vector according to the 3D occupancy prediction result obtained by the last prediction, which can be determined according to actual needs.

[0056] In an optional embodiment, the three-dimensional fusion feature, the initial query vector of the current frame and the initial query vector of the future frame can be interacted by the bidirectional decoder through the mutual attention mechanism, that is, the three-dimensional fusion feature, the initial query vector of the current frame and the initial query vector of the future frame are directly input into the bidirectional decoder, the initial query vector of the future frame is adjusted based on the context information of the historical frame and the current frame by the bidirectional decoder, and the initial query vector of the current frame is adjusted based on the context information of the historical frame and the future frame by the bidirectional decoder, so as to obtain the target query vector of the current frame and the target query vector of the future frame output by the bidirectional decoder. Alternatively, the three-dimensional fusion feature and the initial query vector of the current frame can be interacted by the bidirectional decoder to obtain the updated query vector of the current frame, so as to achieve the purpose of predicting the 3D occupancy state of the current frame based on the 3D occupancy state of the historical frame, and the three-dimensional fusion feature and the initial query vector of the future frame can be interacted by the bidirectional decoder to obtain the updated query vector of the future frame, so as to achieve the purpose of predicting the 3D occupancy state of the future frame based on the 3D occupancy state of the historical frame, and finally the updated query vector of the current frame and the updated query vector of the future frame can be interacted by the bidirectional decoder to obtain the target query vector of the current frame and the target query vector of the future frame, so as to achieve the purpose of predicting the 3D occupancy state of the future frame based on the 3D occupancy state of the current frame and predicting the 3D occupancy state of the current frame based on the 3D occupancy state of the future frame.

[0057] In another optional embodiment, the bidirectional decoder can be used to facilitate the feature interaction between the initial query vector of the current frame and the initial query vector of the future frame, and the target query vector of the current frame and the target query vector of the future frame output by the bidirectional decoder.

[0058] In step S108, the target query vector of the current frame is used for three-dimensional reconstruction to generate a three-dimensional occupancy prediction result of the current frame, and the target query vector of the future frame is used for three-dimensional reconstruction to generate a three-dimensional occupancy prediction result of the future frame. The three-dimensional occupancy prediction result is used to represent whether the occupancy probability of a plurality of voxels in the three-dimensional space is greater than a preset probability. The occupancy probability is the probability that a voxel is occupied by an object.

[0059] The three-dimensional occupancy prediction result in the above step can be a 3D occupancy state obtained by three-dimensional reconstruction based on the target query vector. The 3D occupancy state is predicted by the target query vector and is not the real occupancy state of a plurality of voxels in the three-dimensional space. Therefore, the occupancy probability can be used to indicate whether a voxel is occupied by an object. That is, the probability that a voxel is occupied by an object can be determined based on the target query vector, and then the voxel can be indicated as being occupied by an object or not by comparing with the preset probability.

[0060] The preset probability in the above step can be a preset occupancy probability threshold. If the occupancy probability of a voxel is greater than the preset probability, it indicates that the voxel is occupied by an object. If the occupancy probability of a voxel is less than or equal to the preset probability, it indicates that the voxel is not occupied by an object.

[0061] In an optional embodiment, the target query vector is an information representation manner of the feature space. In order to reduce the computational complexity and improve the processing efficiency, the resolution of the target query vector is usually low, and the format of the target query vector does not match the 3D occupancy prediction result. Therefore, the target query vectors of the current frame and the future frame need to be processed by three-dimensional reconstruction to ensure that the scale, details, structure and format of the 3D occupancy prediction result meet the real needs of the environment around the vehicle. Optionally, the specific process of three-dimensional reconstruction can be implemented by using feature mapping, upsampling and other operations provided in related technologies. The specific operation process is not limited in the present application and can be adjusted according to actual needs. After the target query vector of the current frame or the future frame is three-dimensionally reconstructed, the occupancy probability that each voxel in the three-dimensional space is occupied by an object can be obtained, and then the occupancy probability of each voxel is compared with the preset probability to determine whether each voxel is occupied by an object, thereby obtaining the three-dimensional occupancy prediction result of the current frame or the future frame.

[0062] Furthermore, the 3D occupancy prediction results can be applied to the path planning process of autonomous driving, so that obstacles in the vehicle's surrounding environment can be identified and avoidance plans can be made in advance, improving vehicle driving safety.

[0063] Based on the method provided in the above embodiments of the present application, the three-dimensional space in which the vehicle is located is perceived to obtain multiple frames of images; based on the two-dimensional features of the multiple frames of images, three-dimensional fusion features are generated; a bidirectional decoder is used to perform feature interaction on the three-dimensional fusion features, the initial query vector of the current frame, and the initial query vector of the future frame to obtain a target query vector of the current frame and a target query vector of the future frame; three-dimensional reconstruction is performed based on the target query vector of the current frame to generate a three-dimensional occupancy prediction result of the current frame, and three-dimensional reconstruction is performed based on the target query vector of the future frame to generate a three-dimensional occupancy prediction result of the future frame, thereby achieving the purpose of 3D occupancy prediction. It is easy to notice that in the 3D occupancy prediction process, the bidirectional decoder can be used to perform feature interaction on the three-dimensional fusion features, the initial query vector of the current frame and the initial query vector of the future frame. That is, the context information of the future frame can be used to correct the initial query vector of the current frame, so that the 3D occupancy prediction of the current frame considers the 3D occupancy status of the historical frame and the 3D occupancy status of the future frame at the same time, and no longer relies solely on the 3D occupancy status of the historical frame, thereby avoiding false detection or missed detection due to occlusion of the historical frame, achieving the technical effect of improving the accuracy of 3D occupancy prediction, and solving the technical problem of low accuracy of 3D occupancy prediction of the current frame using historical frames.

[0064] In the above embodiment of the present application, a bidirectional decoder is used to perform feature interaction on the three-dimensional fusion feature, the initial query vector of the current frame and the initial query vector of the future frame to obtain the target query vector of the current frame and the target query vector of the future frame, including: performing attention processing on the three-dimensional fusion feature and the initial query vector of the current frame to obtain the attention feature of the current frame; performing attention processing on the three-dimensional fusion feature and the initial query vector of the future frame to obtain the attention feature of the future frame; performing attention processing on the attention feature of the current frame and the attention feature of the future frame to obtain the target query vector of the current frame and the target query vector of the future frame.

[0065] In an optional embodiment, while decoding the target query vector of the current frame and the target query vector of the future frame from the three-dimensional fusion feature, a mutual attention mechanism can be introduced to constrain the 3D occupancy prediction processes of the two frames with each other. In the above-mentioned embodiments of the present application, the bidirectional decoder can include an independent decoding module and a bidirectional interaction module. The processing procedure of decoding by using the two modules is as follows: first, the independent decoding module is used to independently decode the current frame and the future frame, that is, the mutual attention mechanism is used to process the three-dimensional fusion feature and the initial query vector of the current frame to obtain the attention feature of the current frame, and the mutual attention mechanism is used to process the three-dimensional fusion feature and the initial query vector of the future frame to obtain the attention feature of the future frame; then, the bidirectional interaction module is used to perform bidirectional interaction on the current frame and the future frame, that is, the mutual attention mechanism is used to process the attention feature of the current frame and the attention feature of the future frame, so that the target query vector of the current frame and the target query vector of the future frame can be obtained.

[0066] Based on the above technical solution, the mutual information exchange between the current frame and the future frame can be realized by introducing the mutual attention mechanism, the context of the historical frame and the future frame is considered in the 3D occupancy prediction process of the current frame, and the 3D occupancy prediction accuracy is improved.

[0067] In the above-mentioned embodiments of the present application, the attention mechanism is used to process the attention feature of the current frame and the attention feature of the future frame to obtain the target query vector of the current frame and the target query vector of the future frame, which includes: the attention feature of the current frame and the attention feature of the future frame are spliced to obtain spliced features; the attention mechanism is used to process the spliced features to obtain target attention features; and the target attention features are split to obtain the target query vector of the current frame and the target query vector of the future frame.

[0068] In an optional embodiment, in the process of processing the attention features of the two frames by using the attention mechanism, a bidirectional attention mechanism can be introduced. In the above-mentioned embodiments of the present application, the bidirectional interaction module can include a feature splicing module, a multi-head attention module and a feature splitting module. The processing procedure of bidirectional attention processing by using the three modules is as follows: first, the attention features of the two frames are spliced to obtain spliced features; then, the mutual attention mechanism is used to process the spliced features to obtain target attention features; finally, the target attention features are split to obtain the target query vector of the current frame and the target query vector of the future frame. It should be noted that the feature splicing and the feature splitting can be two mutual processes, and the specific implementation mode is not limited in the present application.

[0069] Based on the technical scheme, the bidirectional attention mechanism is introduced to realize the enhancement of the query vectors of two frames and the bidirectional propagation of information, and the continuity and robustness of the 3D occupancy prediction result in the time and space dimensions are improved.

[0070] In the above embodiments of the present application, based on the two-dimensional features of multiple frames of images, three-dimensional fusion features are generated, including: performing feature extraction on the multiple frames of images respectively to obtain two-dimensional features of the multiple frames of images; performing three-dimensional reconstruction on the two-dimensional features of the multiple frames of images respectively to obtain three-dimensional features of the multiple frames of images; and fusing the three-dimensional features of the multiple frames of images using a space-time encoder to obtain three-dimensional fusion features.

[0071] In an optional embodiment, a deep learning model (such as a lightweight convolutional neural network) can be pre-trained as a feature extractor, and the feature extractor can be used to extract features of each frame of image, extract multi-scale image features as two-dimensional features of each frame. Then the two-dimensional features of each frame of image can be converted into a mixed representation of voxels in three-dimensional space by a 3D scene reconstruction module, that is, the three-dimensional features of each frame of image are obtained. Finally, in order to consider that directly splicing the three-dimensional features of multiple frames of images will cause dimension explosion, therefore, in order to achieve the purpose of dimension reduction, the space-time encoder can be used to fuse the three-dimensional features of multiple frames of images (such as summation, mutual attention processing, etc.), thereby obtaining three-dimensional fusion features.

[0072] Through the above technical scheme, in order to avoid the lack of depth information of two-dimensional features, which leads to the reduction of the accuracy of 3D occupancy prediction, the two-dimensional features can be reconstructed into three-dimensional features by three-dimensional reconstruction, and the dimension can be reduced by feature fusion, thereby achieving the effect of ensuring the accuracy of 3D occupancy prediction and improving the prediction efficiency of 3D occupancy prediction.

[0073] In the above embodiments of the present application, the three-dimensional space where the vehicle is located is perceived to obtain multiple frames of images, including: obtaining historical frames of images collected by a sensor on the vehicle; in response to the number of historical frames of images collected by the sensor being less than a preset number, generating a preset number of historical frames of images based on the historical frames of images collected by the sensor.

[0074] The sensor in the above step can be a vehicle-mounted camera installed on the vehicle, but is not limited thereto, and can also be other sensors capable of collecting visual images.

[0075] In an optional embodiment, the vehicle-mounted camera can be used to collect images of the three-dimensional environment in which the vehicle is located to obtain at least one frame of image. However, due to the complexity of the environment around the vehicle, the image collected by the vehicle-mounted camera can be blocked. In actual situations, the number of historical frame images collected by the vehicle-mounted camera can be less than the preset number. For this situation, the preset number of historical frame images can be generated based on the historical frame images collected by the sensor. Here, the historical frame images can be directly copied to achieve this, or the historical frame images can be linearly interpolated. The present application does not make specific limitations here, and the selection can be made according to actual needs.

[0076] The following will be described by taking the preset number as 3. Under normal circumstances, the multiple images collected by the vehicle-mounted camera can be represented as I t-3 , I t-2 , I t-1 , I t , wherein I t represents the current frame image, I t-3 , I t-2 and I t-1 represent historical frame images. Assuming that the historical frame images collected by the sensor only include I t-2 and I t-1 , but lack I t-3 , in this case, I t-2 can be copied, and the copied image can be used as I t-3 . Assuming that the historical frame images collected by the sensor only include I t-3 and I t-1 , but lack I t-2 , in this case, I t-3 and I t-1 can be linearly interpolated, and the generated image can be used as I t-2 . Assuming that the historical frame images collected by the sensor only include any one frame of image, in this case, a single frame backup mode can be enabled, for example, the frame image can be directly copied, and the copied image can be used as the missing image. It should be noted that for the single frame backup mode, a single historical frame of training data can also be specially constructed during the training of the bidirectional decoder, and the training data can be used for training, so that even in the case of only one historical frame image, the feature decoding and 3D occupancy prediction can also be accurately performed.

[0077] It should be noted here that since the vehicle-mounted camera can only obtain the current frame image when it first captures the environment around the vehicle, the historical frame image is missing. In order to avoid this situation, a frame of image can be collected by the vehicle-mounted camera when the vehicle is on, which can be used as a historical frame image. Next, a single frame backup mode can be used for processing, which is not limited here.

[0078] Through the technical solution, the historical frame images of the preset number are constructed based on the historical frame images collected by the sensor, so as to expand the application scenario of the application and improve the accuracy of 3D occupancy prediction.

[0079] The embodiment also provides a training method of a bidirectional decoder, Figure 2 is a flowchart of a training method of a bidirectional decoder according to an embodiment of the application. The method can be executed by a server, and a vehicle can perceive a three-dimensional space where the vehicle is located, obtain multiple frame images, and upload the server as training data. After the server trains the bidirectional decoder, the server can be distributed to the vehicle and deployed on the vehicle. As shown in Figure 2 The flowchart includes the following steps:

[0080] In step S202, training data corresponding to a three-dimensional space where a vehicle is located is obtained, wherein the training data includes: multiple frame image samples, a three-dimensional occupancy calibration result of a current frame, and a three-dimensional occupancy calibration result of a future frame. The three-dimensional occupancy calibration result is used to represent whether a voxel in the three-dimensional space is occupied by an object.

[0081] The vehicle in the above steps can be a vehicle that needs to perform 3D occupancy prediction. From the perspective of driving mode, it can be an autonomous vehicle or a driver-driven vehicle. From the perspective of energy, it can be a fuel vehicle, an electric vehicle, or a hybrid vehicle. From the perspective of specific type, it can be a car, a truck, a bus, etc., but is not limited thereto.

[0082] The three-dimensional space in the above steps can be the real environment where the vehicle is located, i.e., the environment around the vehicle, especially the environment in front of the vehicle. There are often static obstacles in the environment, such as road signs, buildings, trees, etc. There are also dynamic obstacles, such as other vehicles, pedestrians, animals, etc., but not limited thereto.

[0083] The multiple frame image samples in the above steps can be images obtained by intercepting a video collected during the driving process of the vehicle. The video stream can be obtained by continuously shooting the environment around the vehicle through the vehicle-mounted camera during the driving process of the vehicle. Therefore, based on the shooting time, it can be determined that the image shot last time is the image shot at the current time, i.e., the last image is the current frame image, and the images before the image are historical frame images.

[0084] The three-dimensional occupancy calibration result can be a 3D occupancy state obtained by manual calibration or machine calibration. The three-dimensional occupancy calibration result can be used as the true value result in the subsequent bidirectional decoder training process. The specific calibration process is not described here.

[0085] In an optional embodiment, the vehicle can capture a video of the three-dimensional space in which the vehicle is located in real time during driving, and store the video stream to the server, so that the server can frame the video stream to obtain a preset number of historical frame images, a current frame image and future frame images. Then, the 3D occupancy states of the current frame and the future frame can be labeled by manual labeling, so as to obtain the 3D occupancy labeling results of the two frames. Finally, the current frame image, the historical frame image and the 3D occupancy labeling results of the two frames can be used as the final training data.

[0086] In step S204, three-dimensional sample features are generated based on two-dimensional sample features of the multiple frame images.

[0087] In step S206, the initial decoder is used to perform feature interaction on the three-dimensional sample features, the first query vector of the current frame and the first query vector of the future frame, to obtain a second query vector of the current frame and a second query vector of the future frame.

[0088] The initial decoder in the above steps can be a bidirectional decoder before the training starts. The initial decoder and the bidirectional decoder have the same model structure, but the model parameters of the initial decoder are usually initial values which need to be updated during the training process.

[0089] In step S208, three-dimensional reconstruction is performed based on the second query vector of the current frame to generate a three-dimensional occupancy prediction result of the current frame, and three-dimensional reconstruction is performed based on the second query vector of the future frame to generate a three-dimensional occupancy prediction result of the future frame. The three-dimensional occupancy prediction result is used to represent whether the occupancy probability of a plurality of voxels in the three-dimensional space is greater than a preset probability. The occupancy probability is the probability that a voxel is occupied by an object.

[0090] It should be noted that for a deep learning model, the model training process is similar to the inference process, that is, the implementation process of steps S203 to S208 is similar to the implementation process of steps S104 to S108, and therefore, the present application will not be repeated here.

[0091] Optionally, the initial decoder is used to perform feature interaction on the three-dimensional sample features, the first query vector of the current frame and the first query vector of the future frame, to obtain a second query vector of the current frame and a second query vector of the future frame, including: performing attention processing on the three-dimensional sample features and the first query vector of the current frame to obtain an attention feature of the current frame; performing attention processing on the three-dimensional sample features and the first query vector of the future frame to obtain an attention feature of the future frame; and performing attention processing on the attention feature of the current frame and the attention feature of the future frame to obtain the second query vector of the current frame and the second query vector of the future frame.

[0092] Optionally, the attention features of the current frame and the attention features of the future frame are subjected to attention processing to obtain a second query vector of the current frame and a second query vector of the future frame, including: splicing the attention features of the current frame and the attention features of the future frame to obtain spliced features; subjecting the spliced features to attention processing to obtain target attention features; and splitting the target attention features to obtain the second query vector of the current frame and the second query vector of the future frame.

[0093] Optionally, based on the two-dimensional sample features of the multiple image samples, three-dimensional sample features are generated, including: respectively performing feature extraction on the multiple image samples to obtain two-dimensional sample features of the multiple image samples; respectively performing three-dimensional reconstruction on the two-dimensional sample features of the multiple image samples to obtain three-dimensional features of the multiple images; and fusing the three-dimensional features of the multiple images by using a space-time encoder to obtain the three-dimensional sample features.

[0094] In step S210, based on the three-dimensional occupancy prediction result of the current frame, the three-dimensional occupancy calibration result of the current frame, the three-dimensional occupancy prediction result of the future frame, and the three-dimensional occupancy calibration result of the future frame, the parameters of the initial decoder are adjusted to obtain the bidirectional decoder.

[0095] In an optional embodiment, the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy calibration result of the current frame can be compared to determine whether the three-dimensional occupancy prediction result of the current frame is accurate, and similarly, the three-dimensional occupancy prediction result of the future frame and the three-dimensional occupancy calibration result of the future frame can be compared to determine whether the three-dimensional occupancy prediction result of the future frame is accurate. Further, to avoid the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame being inconsistent or changing abruptly, the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame can also be compared to determine whether the three-dimensional occupancy prediction results of the two frames are consistent. If the three-dimensional occupancy prediction result of the current frame or the future frame is inaccurate, or the three-dimensional occupancy prediction results of the two frames are inconsistent, the parameters of the initial decoder can be continuously adjusted. If the three-dimensional occupancy prediction result of the current frame or the future frame is accurate, and the three-dimensional occupancy prediction results of the two frames are consistent, it can be determined that the initial decoder training is completed, and the trained decoder is the bidirectional decoder.

[0096] In another optional embodiment, the total loss function of the initial decoder can be constructed based on the three-dimensional occupancy prediction result of the current frame, the three-dimensional occupancy calibration result of the current frame, the three-dimensional occupancy prediction result of the future frame, and the three-dimensional occupancy calibration result of the future frame. The construction process of the total loss function can consider multiple aspects, for example, whether the three-dimensional occupancy prediction result is accurate, whether the three-dimensional occupancy prediction results of the two frames are consistent, and the like. Therefore, whether the parameters of the initial decoder need to be adjusted can be determined based on the total loss function value, and after the training ends, the bidirectional decoder can be obtained.

[0097] Based on the method provided in the above embodiments of the present application, the three-dimensional fusion features, the initial query vector of the current frame and the initial query vector of the future frame can be interacted by using the bidirectional decoder, that is, the initial query vector of the current frame can be corrected by using the context information of the future frame, so that the 3D occupancy prediction of the current frame considers the 3D occupancy state of the historical frame and the 3D occupancy state of the future frame at the same time, and no longer depends only on the 3D occupancy state of the historical frame, thereby avoiding false detection or missed detection due to the occlusion of the historical frame, achieving the technical effect of improving the accuracy of 3D occupancy prediction, and solving the technical problem of low accuracy of 3D occupancy prediction of the current frame using the historical frame.

[0098] In the above embodiments of the present application, the three-dimensional sample features are generated based on the two-dimensional sample features of the multi-frame image samples, including: determining the current training round of the initial decoder; in the case that the current training round is less than or equal to the first preset round, generating the three-dimensional sample features based on the two-dimensional sample features of the multi-frame image samples; in the case that the current training round is greater than the first preset round, performing random mask processing on the multi-frame image samples to obtain multi-frame processed image samples, and generating the three-dimensional sample features based on the two-dimensional sample features of the multi-frame processed image samples.

[0099] The first preset round described above can be a preset training round, for example, the first preset round can be 40 rounds, 50 rounds, 60 rounds, 70 rounds, etc.

[0100] In an optional embodiment, in order to ensure that the bidirectional decoder encounters frame loss / frame extraction and the like in the actual inference process, and still can be better predicted by the 3D occupancy, the multi-frame image samples can be processed through random mask processing to simulate the above phenomenon. However, random mask processing will affect the learning of the bidirectional decoder on the space-time mapping, and will affect the accuracy of the 3D occupancy prediction, so it is not recommended to use the random mask processing process in the initial stage of the initial decoder training. In view of the above two cases, a first preset round can be preset, and in the case that the current training round of the initial decoder is less than or equal to the first preset round, that is, in the initial stage of the initial decoder training, the multi-frame image samples can be directly used as training data for training; in the case that the current training round of the initial decoder is greater than the first preset round, that is, in the middle stage of the initial decoder training, the multi-frame image samples can be randomly masked, and the multi-frame processed image samples can be used as training data for training.

[0101] Based on the above technical solution, by judging the relationship between the current training round of the initial decoder and the first preset round, it is determined whether to enable random mask processing, thereby improving the robustness of the bidirectional decoder.

[0102] In the above embodiments of the present application, the multi-frame image samples are randomly masked to obtain multi-frame processed image samples, including: determining a random mask probability based on the current training round; randomly masking any one historical frame image sample in the multi-frame image samples based on the random mask probability to obtain the multi-frame processed image samples.

[0103] The random mask probability in the above step can be a dynamic probability value determined based on the current training round, which is used to determine the probability of random mask processing of the historical frame image. The random mask probability can be in the range of 0-30%, but is not limited thereto, and can be determined according to actual needs.

[0104] In an optional embodiment, a random mask probability can be randomly determined from the above value range based on the current training round, and a historical frame image sample can be randomly masked based on the random mask probability to obtain a processed historical frame image. The current frame image, the processed historical frame image, and the historical frame image without random mask processing can be used as multi-frame processed image samples.

[0105] In another optional embodiment, based on the current training round, starting from the minimum value within the above-mentioned value range, the value can be increased according to a certain increment to determine the random mask probability, and a historical frame image sample can be randomly masked based on the random mask probability to obtain a processed historical frame image, and the current frame image, the processed historical frame image, and the historical frame image that has not been randomly masked are used as multi-frame processed image samples.

[0106] The following is an example of 4 frames of image samples. Figure 3 As shown, in training round 1, image sample t-3 can be randomly masked to obtain mask t-3, and mask t-3, image sample t-2, image sample t-1, and image sample t can be used as training data. In training round 2, image sample t-2 can be randomly masked to obtain mask t-2, and mask t-2, image sample t-3, image sample t-1, and image sample t can be used as training data. In training round 3, image sample t-1 can be randomly masked to obtain mask t-1, and mask t-1, image sample t-3, image sample t-2, and image sample t can be used as training data.

[0107] Based on the above technical solution, by determining the random mask probability based on the current training round and randomly masking a historical frame image sample based on the random mask probability, the uncertainty of the real data is simulated and the training efficiency is improved.

[0108] In the above embodiment of the present application, obtaining training data corresponding to the three-dimensional space in which the vehicle is located includes: in response to the number of historical frame image samples in multiple frame image samples being less than a preset number, generating a preset number of historical frame image samples based on the historical frame image samples.

[0109] The implementation process of the above technical solution is similar to the process of processing the historical frame images collected by the sensor in the previous embodiment, and will not be described in detail here.

[0110] In the above embodiments of the present application, based on the three-dimensional occupancy prediction result of the current frame, the three-dimensional occupancy labeling result of the current frame, the three-dimensional occupancy prediction result of the future frame, and the three-dimensional occupancy labeling result of the future frame, the parameters of the initial decoder are adjusted to obtain the bidirectional decoder, including: based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy labeling result of the current frame, constructing the classification prediction loss function value and the edge perception loss function value of the current frame; based on the three-dimensional occupancy prediction result of the future frame and the three-dimensional occupancy labeling result of the future frame, constructing the classification prediction loss function value and the edge perception loss function value of the future frame; based on the three-dimensional features of the multiple image samples, constructing the sparse regularization loss function value; based on the predicted depth distribution of the multiple images and the preset depth distribution of the multiple images, constructing the depth perception loss function value; weighting the classification prediction loss function value and the edge perception loss function value of the current frame, the classification prediction loss function value and the edge perception loss function value of the future frame, the sparse regularization loss function value, and the depth perception loss function value to obtain the total loss function value of the initial decoder; based on the total loss function value, adjusting the parameters of the initial decoder to obtain the bidirectional decoder.

[0111] The classification prediction loss function value in the above step is used to quantify the difference degree of the occupancy probability of each voxel, is used to supervise the occupancy probability of each voxel, and is the most direct supervision signal in the bidirectional decoder training process. The edge perception loss function value in the above step is used to quantify the difference degree of the edge (object surface), and can be obtained by increasing the weight on the edge voxel. The sparse regularization loss function value in the above step is used to quantify the sparsity of the three-dimensional features obtained through three-dimensional reconstruction, and can be determined by the number of non-zero voxels. The depth perception loss function value in the above step is used to quantify the difference degree of the depth distribution in the process of obtaining the three-dimensional features through three-dimensional reconstruction. The preset depth distribution in the above step can be obtained by depth estimation on the multiple images, and the specific implementation process can adopt the depth estimation algorithm in the related art, which is not described herein.

[0112] In an optional embodiment, in order to better fuse the multiple loss function values and balance the influence of different loss function values on the bidirectional decoder training process, the weight coefficients corresponding to different loss function values can be set in advance according to experience, the multiple loss function values are weighted based on the weight coefficients to obtain the total loss function value, and whether the initial decoder is successfully trained is determined based on the total loss function value. If there is no training result, the parameters of the initial decoder are adjusted and the training is continued until the initial decoder is successfully trained to obtain the bidirectional decoder. Optionally, the weight coefficient of the classification prediction loss function value can be 1, the weight coefficient of the edge perception loss function value can be 0.3, the weight coefficient of the sparse regularization loss function value can be 0.01, and the weight coefficient of the depth perception loss function value can be 0.1.

[0113] Based on the above technical solution, by calculating the classification prediction loss function value, the edge perception loss function value, the sparse regularization loss function value and the depth perception loss function value, the effect of constructing the total loss function value of the bidirectional decoder is realized, and the accuracy of 3D occupancy prediction is improved.

[0114] In the above embodiments of the present application, the classification prediction loss function value and the edge perception loss function value of the current frame, the classification prediction loss function value and the edge perception loss function value of the future frame, the sparse regularization loss function value and the depth perception loss function value are weighted to obtain the total loss function value of the initial decoder, including: based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame, constructing a consistency loss function value; the classification prediction loss function value and the edge perception loss function value of the current frame, the classification prediction loss function value and the edge perception loss function value of the future frame, the sparse regularization loss function value, the depth perception loss function value and the consistency loss function value are weighted to obtain the total loss function value.

[0115] The consistency loss function value in the above step is used to quantify the consistency of the three-dimensional occupancy prediction results of the two frames in time sequence, for example, a static object is the same in the three-dimensional occupancy prediction results of the two frames, while a dynamic object should conform to the motion trajectory in the three-dimensional occupancy prediction results of the two frames.

[0116] In an optional embodiment, different methods can be used to calculate the consistency between the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame for different types of objects, thereby obtaining the consistency loss function value, which can be implemented by a general algorithm in the present application, which is not limited in the present application. It should be noted that most objects in three-dimensional space are static objects, therefore, a constraint weight can be added in the consistency loss function value calculation process to increase the weight of static objects.

[0117] Similarly, the weight coefficient corresponding to the consistency loss function value can be set according to experience in advance, and the weight coefficient is used to weight the plurality of loss function values to obtain the total loss function value, and the total loss function value is used to determine whether the initial decoder is successfully trained, if there is no training result, the parameters of the initial decoder are adjusted and the training is continued until the initial decoder is successfully trained to obtain the bidirectional decoder. Optionally, the weight coefficient of the consistency loss function value can be 0.5, but is not limited thereto.

[0118] Based on the above technical solution, by calculating the classification prediction loss function value, the edge perception loss function value, the sparse regularization loss function value, the depth perception loss function value and the consistency loss function value, the effect of constructing the total loss function value of the bidirectional decoder is realized, and the accuracy of 3D occupancy prediction is improved.

[0119] In the above embodiments of the application, the consistency loss function value is constructed based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame, including: determining the current training round of the initial decoder; in the case that the current training round is greater than the second preset round, constructing the consistency loss function value based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame.

[0120] The second preset round described above can be a preset training round, for example, the second preset round can be 80 rounds, 90 rounds, 100 rounds, 110 rounds, 120 rounds, etc.

[0121] In an optional embodiment, the bidirectional decoder can be trained in multiple stages, and a second preset round can be preset. In the case that the current training round of the initial decoder is less than or equal to the second preset round, that is, in the middle stage of the initial decoder training, the total loss function value of the bidirectional decoder can be constructed by calculating the classification prediction loss function value, the edge perception loss function value, the sparse regularization loss function value and the depth perception loss function value, without introducing the consistency loss function value; in the case that the current training round of the initial decoder is greater than the second preset round, that is, in the last stage of the initial decoder training, the total loss function value of the bidirectional decoder can be constructed by calculating the classification prediction loss function value, the edge perception loss function value, the sparse regularization loss function value, the depth perception loss function value and the consistency loss function value.

[0122] Based on the above technical solution, by judging the relationship between the current training round of the initial decoder and the second preset round, it is determined whether to introduce the consistency loss function value, so as to improve the continuity of the time series prediction and avoid overfitting of the bidirectional decoder.

[0123] The following will be described in detail in combination with Figures 4 to 7 A preferred embodiment of the application will be described in detail. The application can use bidirectional time series modeling and Gaussian-voxel hybrid representation technology to realize high-precision and efficient 3D occupancy prediction. As shown in the following figure, the processing flow is as follows: Figure 4

[0124] Obtain input frames t-3-t: four consecutive frames of images captured by a vehicle-mounted camera can be received, and the four image sequences can be represented as {I t-3 , I t-2 , I t-1 , I t}, and the corresponding 3D occupancy labels can be manually labeled, which can be represented as {O t-3 , O t-2 , O t-1 , O t ​} and during training, O t+1 as the supervision signal of future frames. In the case of missing historical frames, it can be automatically detected and compensated, for example, I t-3 missing, I t-2 copy padding; I t-2 missing, it can be generated by linear difference; only 1 frame exists, single-frame backup mode can be enabled.

[0125] Feature extraction using feature extractor: a lightweight convolutional neural network can be used as a feature extractor, and multi-scale image features can be extracted using the feature extractor.

[0126] Feature enhancement using random mask module: a random mask module can be used to dynamically mask historical frames, simulating real frame extraction scenarios. The random mask probability can increase with the increase of training rounds, and the increase of the random mask probability between two training rounds can be 0.01, and the maximum value can be 0.3.

[0127] 3D scene reconstruction: including 2D to 3D projection and Gaussian-voxel hybrid representation two processing flows, as shown in Figure 5 , the 2D to 3D projection process can be a depth distribution prediction for 2D features, and the 2D features are converted to 3D features based on the predicted depth distribution, that is, the initial voxel construction process is realized. Then execute the Gaussian-voxel hybrid representation process, Gaussian diffusion is performed on the initial voxel, and key point detection is performed, and further local detail enhancement is performed, so as to obtain the hybrid representation output.

[0128] Feature fusion using spatio-temporal encoder: after obtaining the hybrid representation output of each frame of image, the hybrid representation data of 4 frames of image can be mixed to obtain the final encoded feature.

[0129] Bidirectional spatio-temporal modeling using bidirectional decoder: first, two learnable query vectors can be obtained, which are current frame query Q t and future frame query Q t+1 , then through the bidirectional decoder, the current frame and future frame features are decoded from the encoded features, and at the same time, the mutual attention mechanism is introduced to constrain each other, and finally the current frame prediction Occ t and future frame prediction Occ t+1 are output. As shown in Figure 6 , the bidirectional decoder can be divided into two processing stages, the first stage is independent decoding, that is, using historical frames and current frames to predict the current frame, and using historical frames and current frames to predict the future frame (forward guiding process), the second stage is bidirectional interaction, that is, the future frame prediction is used as a reverse constraint in the current frame prediction process. As shown in Figure 7The second stage can be implemented by using a bidirectional attention mechanism, first splicing the current frame features and the future frame features, then inputting the spliced features into a multi-head attention for mutual attention processing, and finally obtaining the enhanced current features and the enhanced future features through feature splitting.

[0130] In the above embodiment, the total loss of the bidirectional decoder can be constructed based on the basic prediction loss, the consistency loss and the regularization loss, and specifically can include: constructing the basic prediction loss using the current frame prediction and the current frame true value (artificially labeled label), the weight value of the loss being λ1, constructing the basic prediction loss using the future frame prediction and the future frame true value (artificially labeled label), the weight value of the loss being λ2, constructing the consistency loss using the current frame prediction and the future frame prediction, the weight value of the loss being λ3, constructing the depth perception loss using the predicted depth distribution and the true depth distribution, the weight value of the loss being λ4, constructing the edge perception loss using the current frame prediction and the current frame true value, the weight value of the loss being λ5, constructing the edge perception loss using the future frame prediction and the future frame true value, the weight value of the loss being λ6, and constructing the sparse regularization loss using the number of zero voxels, the weight value of the loss being λ7. Then the final total loss can be obtained through weighted processing.

[0131] In actual training, the above losses can be weighted and adjusted as needed. For example, the basic prediction loss should be dominant, the consistency loss and the depth perception loss should be secondary, and the edge perception loss and the sparse regularization loss should be auxiliary. It should be noted that the depth perception loss is only used at the beginning of training (the first 50 epochs), because the depth prediction mainly affects the 2D to 3D projection, and once the projection module is stable, the loss can be removed in subsequent training.

[0132] In addition, the present application can use a multi-stage training mode to train the bidirectional decoder, which can be divided into three stages. Stage 1 is the first 50 rounds, during which the random mask module can be closed to facilitate the bidirectional decoder to learn the basic space-time mapping; stage 2 is the 51-100 rounds, during which the random mask module can be turned on; and stage 3 is after the 101st round, during which the consistency loss can be added to the total loss. It should be noted that the consistency loss can also be directly added in stage 1 and stage 2.

[0133] Through the above technical solutions, the following technical effects can be achieved: a "past-present-future" bidirectional space-time dependence model is established, the accuracy of the occupancy prediction for the current frame and the future frame is significantly improved; under the same frame loss rate / frame extraction scheme, the prediction stability and time delay jitter tolerance are significantly improved, the anti-interference ability and the tolerance of the system to incomplete input are improved; flexible historical frame input can be supported, and the prediction results of two frames can be output, the deployment is improved, and the computing efficiency of the vehicle-mounted level is maintained.

[0134] In the embodiment, a vehicle path planning method is also provided, Figure 8 is a flowchart of a vehicle path planning method according to an embodiment of the present application, as shown in the figure, the flow includes the following steps: Figure 8

[0135] In step S802, a three-dimensional space where the vehicle is located is perceived to obtain multiple frames of images, wherein the multiple frames of images include a current frame of image and a preset number of historical frame of images.

[0136] In step S804, three-dimensional fusion features are generated based on two-dimensional features of the multiple frames of images.

[0137] In step S806, the three-dimensional fusion features, an initial query vector of the current frame and an initial query vector of a future frame are interacted by using a bidirectional decoder to obtain a target query vector of the current frame and a target query vector of the future frame, wherein the future frame is a next frame of the current frame.

[0138] In step S808, three-dimensional reconstruction is performed based on the target query vector of the current frame to generate a three-dimensional occupancy prediction result of the current frame, and three-dimensional reconstruction is performed based on the target query vector of the future frame to generate a three-dimensional occupancy prediction result of the future frame, wherein the three-dimensional occupancy prediction result is used to represent whether a plurality of voxels in the three-dimensional space correspond to an occupancy probability greater than a preset probability, and the occupancy probability is a probability that the voxel is occupied by an object.

[0139] It should be noted that the implementation process of steps S802 to S208 is similar to that of steps S102 to S108, and therefore the present application will not be repeated here.

[0140] In step S810, a driving path of the vehicle is generated based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame.

[0141] In an optional embodiment, the 3D occupancy prediction result can be applied to the path planning process of autonomous driving, that is, the driving path of the vehicle can be generated based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame, so that the vehicle can drive according to the driving path, obstacles in the environment around the vehicle can be identified and avoided in advance, and the safety of the vehicle driving is improved.

[0142] ​It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0143] According to the embodiments of the present application, a three-dimensional occupancy prediction device of a vehicle is provided. It should be noted that the device can be used to execute the three-dimensional occupancy prediction method of the vehicle. As shown in the Figure 9 The device can include the following modules:

[0144] The first perception module 92 is configured to perceive a three-dimensional space in which the vehicle is located to obtain a plurality of images, wherein the plurality of images include a current frame image and a preset number of historical frame images.

[0145] The first generation module 94 is configured to generate three-dimensional fusion features based on two-dimensional features of the plurality of images.

[0146] The first interaction module 96 is configured to perform feature interaction on the three-dimensional fusion features, an initial query vector of the current frame and an initial query vector of a future frame by using a bidirectional decoder to obtain a target query vector of the current frame and a target query vector of the future frame, wherein the future frame is a next frame of the current frame.

[0147] The first prediction module 98 is configured to perform three-dimensional reconstruction based on the target query vector of the current frame to generate a three-dimensional occupancy prediction result of the current frame, and perform three-dimensional reconstruction based on the target query vector of the future frame to generate a three-dimensional occupancy prediction result of the future frame, wherein the three-dimensional occupancy prediction result is used to represent whether a plurality of voxels in the three-dimensional space correspond to an occupancy probability greater than a preset probability, and the occupancy probability is a probability that the voxel is occupied by an object.

[0148] In the above embodiments of the present application, the first interaction module 96 is further configured to perform attention processing on the three-dimensional fusion features and the initial query vector of the current frame to obtain an attention feature of the current frame, perform attention processing on the three-dimensional fusion features and the initial query vector of the future frame to obtain an attention feature of the future frame, and perform attention processing on the attention feature of the current frame and the attention feature of the future frame to obtain the target query vector of the current frame and the target query vector of the future frame.

[0149] In the above embodiments of the present application, the first interaction module 96 is further configured to concatenate the attention feature of the current frame and the attention feature of the future frame to obtain a concatenated feature; perform attention processing on the concatenated feature to obtain a target attention feature; and split the target attention feature to obtain the target query vector of the current frame and the target query vector of the future frame.

[0150] In the above embodiments of the present application, the first generation module 94 is further configured to perform feature extraction on the multiple frames of images respectively to obtain two-dimensional features of the multiple frames of images; perform three-dimensional reconstruction on the two-dimensional features of the multiple frames of images respectively to obtain three-dimensional features of the multiple frames of images; and fuse the three-dimensional features of the multiple frames of images by using a space-time encoder to obtain three-dimensional fused features.

[0151] In the above embodiments of the present application, the first perception module 92 is further configured to acquire historical frame images collected by a sensor on the vehicle; and in response to a quantity of the historical frame images collected by the sensor being less than a preset quantity, generate a preset quantity of historical frame images based on the historical frame images collected by the sensor.

[0152] According to embodiments of the present application, a training device of a bidirectional decoder is also provided. It should be noted that the device can be used to execute the training method of the bidirectional decoder. As shown in FIG. 10, the device can include the following modules: Figure 10

[0153] The acquisition module 102 is configured to acquire training data corresponding to a three-dimensional space in which the vehicle is located, wherein the training data includes: multiple frames of image samples, a three-dimensional occupancy labeling result of a current frame, and a three-dimensional occupancy labeling result of a future frame, and the three-dimensional occupancy labeling result is used to represent whether a voxel in the three-dimensional space is occupied by an object.

[0154] The second generation module 104 is configured to generate three-dimensional sample features based on two-dimensional sample features of the multiple frames of image samples.

[0155] The second interaction module 106 is configured to perform feature interaction on the three-dimensional sample features, a first query vector of the current frame, and a first query vector of the future frame by using an initial decoder to obtain a second query vector of the current frame and a second query vector of the future frame.

[0156] The second prediction module 108 is configured to perform three-dimensional reconstruction based on the second query vector of the current frame to generate a three-dimensional occupancy prediction result of the current frame, and perform three-dimensional reconstruction based on the second query vector of the future frame to generate a three-dimensional occupancy prediction result of the future frame, wherein the three-dimensional occupancy prediction result is used to represent whether a plurality of voxels in the three-dimensional space correspond to an occupancy probability greater than a preset probability, and the occupancy probability is a probability that a voxel is occupied by an object.

[0157] ​The training module 110 is used to adjust the parameters of the initial decoder based on the 3D occupancy prediction result of the current frame, the 3D occupancy calibration result of the current frame, the 3D occupancy prediction result of the future frame, and the 3D occupancy calibration result of the future frame to obtain a bidirectional decoder.

[0158] In the above embodiment of the present application, the second interaction module 106 is also used to perform attention processing on the three-dimensional sample features and the first query vector of the current frame to obtain the attention features of the current frame; perform attention processing on the three-dimensional sample features and the first query vector of the future frame to obtain the attention features of the future frame; perform attention processing on the attention features of the current frame and the attention features of the future frame to obtain the second query vector of the current frame and the second query vector of the future frame.

[0159] In the above embodiment of the present application, the second interaction module 106 is also used to splice the attention features of the current frame and the attention features of the future frame to obtain a spliced ​​feature; perform attention processing on the spliced ​​feature to obtain a target attention feature; and split the target attention feature to obtain a second query vector of the current frame and a second query vector of the future frame.

[0160] In the above embodiment of the present application, the second generation module 104 is also used to extract features from multiple frame image samples respectively to obtain two-dimensional sample features of the multiple frame image samples; perform three-dimensional reconstruction on the two-dimensional sample features of the multiple frame image samples respectively to obtain three-dimensional features of the multiple frame images; and use the space-time encoder to fuse the three-dimensional features of the multiple frame images to obtain three-dimensional sample features.

[0161] In the above embodiment of the present application, the second generation module 104 is also used to determine the current training round of the initial decoder; when the current training round is less than or equal to the first preset round, three-dimensional sample features are generated based on the two-dimensional sample features of the multi-frame image samples; when the current training round is greater than the first preset round, random masking is performed on the multi-frame image samples to obtain multi-frame processed image samples, and three-dimensional sample features are generated based on the two-dimensional sample features of the multi-frame processed image samples.

[0162] In the above embodiment of the present application, the second generation module 104 is also used to determine the random mask probability based on the current training round; based on the random mask probability, any historical frame image sample in the multiple frame image samples is randomly masked to obtain multiple frame processed image samples.

[0163] In the above embodiment of the present application, the acquisition module 102 is further configured to generate a preset number of historical frame image samples based on the historical frame image samples in response to the number of historical frame image samples in the multiple frame image samples being less than a preset number.

[0164] In the foregoing embodiments of the present application, the training module 110 is further configured to construct a classification prediction loss function value and an edge perception loss function value of the current frame based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy labeling result of the current frame; construct a classification prediction loss function value and an edge perception loss function value of the future frame based on the three-dimensional occupancy prediction result of the future frame and the three-dimensional occupancy labeling result of the future frame; construct a sparse regularization loss function value based on the three-dimensional features of the multiple image samples; construct a depth perception loss function value based on the predicted depth distribution of the multiple images and the preset depth distribution of the multiple images; and perform weighted processing on the classification prediction loss function value and the edge perception loss function value of the current frame, the classification prediction loss function value and the edge perception loss function value of the future frame, the sparse regularization loss function value, and the depth perception loss function value to obtain a total loss function value of the initial decoder; and adjust the parameters of the initial decoder based on the total loss function value to obtain the bidirectional decoder.

[0165] In the foregoing embodiments of the present application, the training module 110 is further configured to construct a consistency loss function value based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame; and perform weighted processing on the classification prediction loss function value and the edge perception loss function value of the current frame, the classification prediction loss function value and the edge perception loss function value of the future frame, the sparse regularization loss function value, the depth perception loss function value, and the consistency loss function value to obtain a total loss function value.

[0166] In the foregoing embodiments of the present application, the training module 110 is further configured to determine a current training round of the initial decoder; and in a case where the current training round is greater than a second preset round, construct a consistency loss function value based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame.

[0167] According to embodiments of the present application, a vehicle path planning device is also provided. It should be noted that the device can be used to execute the vehicle path planning method described above. As shown in FIG. 8, the device can include the following modules: Figure 11

[0168] The second perception module 112 is configured to perceive a three-dimensional space in which the vehicle is located to obtain multiple image frames, wherein the multiple image frames include a current image frame and a preset number of historical image frames.

[0169] The third generation module 114 is configured to generate three-dimensional fusion features based on two-dimensional features of the multiple image frames.

[0170] The third interaction module 116 is configured to perform feature interaction on the three-dimensional fusion features, the initial query vector of the current frame, and the initial query vector of the future frame by using the bidirectional decoder to obtain a target query vector of the current frame and a target query vector of the future frame, wherein the future frame is a next frame of the current frame.​

[0171] The third prediction module 118 is configured to perform three-dimensional reconstruction based on the target query vector of the current frame to generate a three-dimensional occupancy prediction result of the current frame, and perform three-dimensional reconstruction based on the target query vector of the future frame to generate a three-dimensional occupancy prediction result of the future frame, where the three-dimensional occupancy prediction result is used to represent whether a plurality of voxels in a three-dimensional space correspond to an occupancy probability greater than a preset probability, and the occupancy probability is a probability that the voxel is occupied by an object.

[0172] The planning module 1110 is configured to generate a driving path of the vehicle based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame.

[0173] Embodiments of the present application also provide a vehicle, including a memory storing an executable program, and a processor configured to execute the program, where the program is executed to perform the method in any of the embodiments of the present application.

[0174] Embodiments of the present application also provide a computer readable storage medium including a stored executable program, where the executable program is executed to control a device where the computer readable storage medium is located to perform the method in any of the embodiments of the present application.

[0175] Embodiments of the present application also provide a computer program product including a computer program, where the computer program is executed by a processor to implement the method in any of the embodiments of the present application.

[0176] Embodiments of the present application also provide a computer program product including a non-volatile computer readable storage medium for storing a computer program, where the computer program is executed by a processor to implement the method in any of the embodiments of the present application.

[0177] Embodiments of the present application also provide a computer program, where the computer program is executed by a processor to implement the method in any of the embodiments of the present application.

[0178] In the above-described embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0179] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units can be a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.

[0180] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0181] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0182] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0183] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.

Claims

1. A three-dimensional vehicle occupancy prediction method, characterized in that: include: Perceiving the three-dimensional space in which the vehicle is located to obtain multiple frames of images, wherein the multiple frames of images include: a current frame image and a preset number of historical frame images; Generating three-dimensional fusion features based on the two-dimensional features of the multiple frames of images; Performing feature interaction on the three-dimensional fusion feature, the initial query vector of the current frame, and the initial query vector of a future frame using a bidirectional decoder to obtain a target query vector of the current frame and a target query vector of the future frame, wherein the future frame is a frame next to the current frame; Three-dimensional reconstruction is performed based on the target query vector of the current frame to generate a three-dimensional occupancy prediction result of the current frame, and three-dimensional reconstruction is performed based on the target query vector of the future frame to generate a three-dimensional occupancy prediction result of the future frame, wherein the three-dimensional occupancy prediction result is used to represent whether the occupancy probability corresponding to multiple voxels in the three-dimensional space is greater than a preset probability, and the occupancy probability is the probability that the voxel is occupied by an object.

2. The method according to claim 1, characterized in that The step of performing feature interaction on the three-dimensional fusion feature, the initial query vector of the current frame, and the initial query vector of the future frame using a bidirectional decoder to obtain the target query vector of the current frame and the target query vector of the future frame includes: Performing attention processing on the three-dimensional fusion feature and the initial query vector of the current frame to obtain an attention feature of the current frame; performing attention processing on the three-dimensional fusion feature and the initial query vector of the future frame to obtain an attention feature of the future frame; Performing attention processing on the attention features of the current frame and the attention features of the future frame to obtain a target query vector of the current frame and a target query vector of the future frame; Preferably, performing attention processing on the attention features of the current frame and the attention features of the future frame to obtain the target query vector of the current frame and the target query vector of the future frame includes: Splicing the attention feature of the current frame and the attention feature of the future frame to obtain a spliced ​​feature; Performing attention processing on the spliced ​​features to obtain target attention features; The target attention feature is split to obtain a target query vector of the current frame and a target query vector of the future frame.

3. The method according to claim 1 or 2, characterized in that Generating three-dimensional fusion features based on the two-dimensional features of the multiple frames of images includes: Performing feature extraction on the multiple frames of images respectively to obtain two-dimensional features of the multiple frames of images; Reconstructing the two-dimensional features of the multiple frames of images into three dimensions to obtain the three-dimensional features of the multiple frames of images; fusing the three-dimensional features of the multiple frames of image using a spatiotemporal encoder to obtain the three-dimensional fused features; Preferably, the sensing of the three-dimensional space in which the vehicle is located to obtain multiple frames of images includes: Obtain historical frame images collected by sensors on the vehicle; In response to the number of historical frame images collected by the sensor being less than the preset number, the preset number of historical frame images is generated based on the historical frame images collected by the sensor.

4. A method for training a bidirectional decoder, characterized in that: include: Obtaining training data corresponding to the three-dimensional space in which the vehicle is located, wherein the training data includes: multiple frames of image samples, a three-dimensional occupancy calibration result of a current frame, and a three-dimensional occupancy calibration result of a future frame, wherein the three-dimensional occupancy calibration result is used to indicate whether a voxel in the three-dimensional space is occupied by an object; Generating three-dimensional sample features based on the two-dimensional sample features of the multiple frames of image samples; Performing feature interaction on the three-dimensional sample feature, the first query vector of the current frame, and the first query vector of the future frame using an initial decoder to obtain a second query vector of the current frame and a second query vector of the future frame; performing three-dimensional reconstruction based on the second query vector of the current frame to generate a three-dimensional occupancy prediction result for the current frame, and performing three-dimensional reconstruction based on the second query vector of the future frame to generate a three-dimensional occupancy prediction result for the future frame, wherein the three-dimensional occupancy prediction result is used to indicate whether an occupancy probability corresponding to a plurality of voxels in the three-dimensional space is greater than a preset probability, the occupancy probability being a probability that the voxel is occupied by an object; Based on the 3D occupancy prediction result of the current frame, the 3D occupancy calibration result of the current frame, the 3D occupancy prediction result of the future frame, and the 3D occupancy calibration result of the future frame, the parameters of the initial decoder are adjusted to obtain a bidirectional decoder.

5. The method according to claim 4, characterized in that Generating three-dimensional sample features based on the two-dimensional sample features of the multiple frames of image samples includes: determining a current training round of the initial decoder; generating the three-dimensional sample features based on the two-dimensional sample features of the multiple frames of image samples when the current training round is less than or equal to the first preset round; When the current training round is greater than the first preset round, performing random mask processing on the multiple frames of image samples to obtain multiple frames of processed image samples, and generating the three-dimensional sample features based on the two-dimensional sample features of the multiple frames of processed image samples; Preferably, performing random mask processing on the multiple frames of image samples to obtain multiple frames of processed images includes: determining a random mask probability based on the current training round; Random mask processing is performed on any one historical frame image sample in the multiple frame image samples based on the random mask probability to obtain the multiple frame processed image samples.

6. The method according to claim 4, characterized in that The adjusting the parameters of the initial decoder based on the 3D occupancy prediction result of the current frame, the 3D occupancy calibration result of the current frame, the 3D occupancy prediction result of the future frame, and the 3D occupancy calibration result of the future frame to obtain a bidirectional decoder includes: Constructing a classification prediction loss function value and an edge perception loss function value of the current frame based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy calibration result of the current frame; Constructing a classification prediction loss function value and an edge perception loss function value of the future frame based on the three-dimensional occupancy prediction result of the future frame and the three-dimensional occupancy calibration result of the future frame; Constructing a sparse regularization loss function value based on the three-dimensional features of the multi-frame image samples; Constructing a depth perception loss function value based on the predicted depth distribution of the multiple frames of image and the preset depth distribution of the multiple frames of image; Performing weighted processing on the classification prediction loss function value and the edge perception loss function value of the current frame, the classification prediction loss function value and the edge perception loss function value of the future frame, the sparse regularization loss function value, and the depth perception loss function value to obtain a total loss function value of the initial decoder; Parameters of the initial decoder are adjusted based on the total loss function value to obtain the bidirectional decoder.

7. The method according to claim 6, characterized in that The weighted processing of the classification prediction loss function value and the edge perception loss function value of the current frame, the classification prediction loss function value and the edge perception loss function value of the future frame, the sparse regularization loss function value, and the depth perception loss function value to obtain the total loss function value of the initial decoder includes: Constructing a consistency loss function value based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame; Performing weighted processing on the classification prediction loss function value and the edge perception loss function value of the current frame, the classification prediction loss function value and the edge perception loss function value of the future frame, the sparse regularization loss function value, the depth perception loss function value, and the consistency loss function value to obtain the total loss function value; Preferably, constructing a consistency loss function value based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame includes: determining a current training round of the initial decoder; In a case where the current training round is greater than a second preset round, the consistency loss function value is constructed based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame.

8. A vehicle path planning method, characterized in that: include: Perceiving the three-dimensional space in which the vehicle is located to obtain multiple frames of images, wherein the multiple frames of images include: a current frame image and a preset number of historical frame images; Generating three-dimensional fusion features based on the two-dimensional features of the multiple frames of images; performing feature interaction on the three-dimensional fusion feature, the initial query vector of the current frame, and the initial query vector of a future frame using a bidirectional decoder to obtain a target query vector of the current frame and a target query vector of the future frame, wherein the future frame is a frame next to the current frame; performing three-dimensional reconstruction based on the target query vector of the current frame to generate a three-dimensional occupancy prediction result for the current frame, and performing three-dimensional reconstruction based on the target query vector of the future frame to generate a three-dimensional occupancy prediction result for the future frame, wherein the three-dimensional occupancy prediction result is used to indicate whether an occupancy probability corresponding to a plurality of voxels in the three-dimensional space is greater than a preset probability, the occupancy probability being a probability that the voxel is occupied by an object; A driving path of the vehicle is generated based on the three-dimensional occupancy prediction result of the current frame and the three-dimensional occupancy prediction result of the future frame.

9. A vehicle, characterized in that: include: a memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 8 when running.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the storage medium is located is controlled to execute the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Occupancy prediction model training method, related method, device, equipment and medium

    CN122049887A