Semantic Occupancy Grid Prediction Method, Electronic Device, and Storage Medium
By training the rain removal model and combining the 3D occupation grid prediction method with Hungarian algorithm and Kalman filtering, the problem of instability in the occupation grid prediction of the autonomous driving system in Yutianxia is solved, and the perception ability and system security are improved.
Patent Information
- Application Number
- CN202510512112.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing autonomous driving technology is insufficient in complex weather conditions, especially in rainy days, and the accuracy and stability of grid predictions are caused by degradation of perception capabilities and affecting the safety and decision-making performance of autonomous driving systems.
The 3D occupancy raster prediction model is trained using the rain removal model, combining Hungarian algorithm and Kalman filtering, and dynamically adjusting the model parameters to ensure the stability and consistency of the prediction results.
It improves the timeliness and stability of the autonomous driving perception module, enhances the environmental perception capability under complex weather conditions, and ensures the safe operation of the autonomous driving system.
Smart Images

Figure CN120032335B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving perception, and more specifically, to a semantic occupancy grid prediction method, an electronic device, and a storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, the requirements for autonomous driving technology in terms of perception accuracy, environmental adaptability, and safety are becoming increasingly stringent. Especially in complex weather conditions such as rainy days, the perception and understanding of driving scenarios have become a key challenge for autonomous driving systems to achieve safe driving. Occupancy grid, as a data structure that divides the environmental space into multiple small cells and marks the occupancy probability of each cell by an object, is an important tool for environmental modeling in autonomous driving. Therefore, efficient occupancy grid modeling and stability optimization in the driving environment to improve the perception ability in complex weather have become an important direction of technological development. This not only requires the model to accurately and quickly understand the semantic information of the surrounding environment, but also needs to maintain high stability and reliability under various conditions to ensure the safe operation of the autonomous driving system.
[0003] In related technologies, one type of method uses single-view or multi-view two-dimensional images for scene reconstruction and grid prediction, and these methods can achieve certain results in simple scenarios. For example, a monocular vision-based method captures and processes images through a single camera, and infers occupancy grid information through image segmentation and feature extraction. This type of method performs well in a clear field of view, but when encountering complex scenarios such as rainy days, there will be raindrops in the image or the raindrops cause the image to be blurred, and the image quality drops severely, resulting in a significant decrease in the accuracy and stability of model prediction.
[0004] Another type of method fuses lidar and multi-view image data. Although it has high geometric accuracy, due to the complex processing of sensor data and large computational overhead, it cannot meet the requirements of real-time prediction. For example, lidar can provide accurate spatial position information through point cloud data, and combining multi-view image data can enhance the perception ability of the model. However, this method requires a large amount of computing resources, especially when processing a large amount of point cloud data, which usually requires the support of high-performance computing devices. At the same time, there are still obvious deficiencies in the consistency and stability of multi-frame prediction results of existing models, which are prone to prediction fluctuations, thereby affecting the decision-making and planning performance of autonomous driving systems. Summary of the Invention
[0005] The purpose of the present invention is to provide a semantic occupancy grid prediction method, an electronic device, and a storage medium, which can improve the stability of semantic occupancy grid prediction.
[0006] To achieve the above purpose, the technical solutions adopted in the embodiments of the present application are as follows:
[0007] In a first aspect, an embodiment of the present application provides a semantic occupancy grid prediction method, and the method includes:
[0008] Train a de-raining model to obtain a trained de-raining model;
[0009] Obtain a 3D occupancy grid prediction model constructed based on the trained de-raining model;
[0010] Input a continuous image frame to be processed into the 3D occupancy grid prediction model to obtain a semantic occupancy grid prediction result of the continuous image frame to be processed;
[0011] Match the semantic occupancy grid prediction result of the continuous image frame to be processed based on the Hungarian algorithm to obtain a first semantic occupancy grid prediction result of multiple time frames with an association relationship;
[0012] Correct the first semantic occupancy grid prediction result based on the Kalman filter, and determine a stability parameter of the corrected first semantic occupancy grid prediction result based on a verification model;
[0013] When the stability parameter meets the output condition, use the corrected first semantic occupancy grid prediction result as the output of the continuous image frame to be processed;
[0014] When the stability parameter does not meet the output condition, adjust the parameters of the 3D occupancy grid prediction model;
[0015] Based on the adjusted 3D occupancy grid prediction model, return to execute inputting the continuous image frame to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction result of the continuous image frame to be processed until the stability parameter meets the output condition.
[0016] In an optional implementation manner, the step of training the de-raining model to obtain a trained de-raining model includes:
[0017] Obtain rainy-day images that meet preset conditions from a driving dataset;
[0018] Generate a semantic segmentation image for the rainy-day image based on a generative algorithm;
[0019] Optimize the semantic segmentation image based on an optimal transport strategy;
[0020] Train the de-raining model based on the optimized semantic segmentation image to obtain a trained de-raining model.
[0021] In an optional implementation manner, the step of obtaining a 3D occupancy grid prediction model constructed based on the trained de-raining model includes:
[0022] Input the rainy-day image of the real scene into the trained rain-removing model to obtain multiple rain-removed images from multiple perspectives;
[0023] Calibrate and synchronize the time of the multiple rain-removed images respectively to obtain multiple first images from multiple perspectives with consistent time;
[0024] For each first image, based on the semantic features of the first image, the high-dimensional feature representation tensor of the first image, and the vehicle geometric information, obtain the BEV feature of the first image;
[0025] For each BEV feature, determine the current BEV feature and the historical BEV feature corresponding to the BEV feature;
[0026] Generate a three-dimensional occupancy descriptor based on the current BEV feature and the historical BEV feature to complete the construction of the 3D occupancy grid prediction model.
[0027] In an alternative embodiment, the step of matching the semantic occupancy grid prediction results of the consecutive image frames to be processed based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with an association relationship includes:
[0028] For each semantic occupancy grid prediction result, determine the historical semantic occupancy grid prediction result of the semantic occupancy grid prediction result;
[0029] Determine the first parameter information of the semantic occupancy grid prediction result and the second parameter information of the historical semantic occupancy grid prediction result;
[0030] Based on the first parameter information and the second parameter information, determine the difference value between the semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result;
[0031] The Hungarian algorithm processes each of the difference values to obtain the first semantic occupancy grid prediction results of multiple time frames with an association relationship.
[0032] In an alternative embodiment, the step of determining the stable parameters of the corrected first semantic occupancy grid prediction result includes:
[0033] For each of the corrected first semantic occupancy grid prediction results, determine the historical first semantic occupancy grid prediction result corresponding to the corrected first semantic occupancy grid prediction result;
[0034] Determine the confidence stability value, position stability value, size stability value, and orientation stability value between the corrected first semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result;
[0035] Determine the stability parameters of the corrected first semantic occupancy grid prediction result based on the confidence stability value, position stability value, size stability value, and orientation stability value.
[0036] In an alternative embodiment, the confidence stability value satisfies the following formula:
[0037] ;
[0038] where is the confidence stability value, is the first confidence, is the second confidence, and are the 99% and 1% percentiles of the confidence;
[0039] The position stability value satisfies the following formula:
[0040] ;
[0041] where is the position stability value, is the first position information, is the second position information;
[0042] The size stability value satisfies the following formula:
[0043] ;
[0044] where is the size stability value, is the first size information, is the second size information;
[0045] The orientation stability value satisfies the following formula:
[0046] ;
[0047] where is the orientation stability value, is the first orientation information, is the second orientation information, is the orientation angle corresponding to the first orientation information, is the orientation angle corresponding to the second orientation information.
[0048] In an alternative embodiment, when the stability parameters of the corrected first semantic occupancy grid prediction result do not meet the output conditions, the steps of adjusting the parameters of the 3D occupancy grid prediction model include:
[0049] When the stability parameter of the corrected first semantic occupancy grid prediction result does not meet the output condition, calculate a center position loss based on the first position information and the second position information, where the first position information and the second position information indicate the center position;
[0050] Determine the total number of time frames of the consecutive image frames to be processed;
[0051] Calculate a position offset loss based on the first position information, the second position information, and the total number of time frames;
[0052] Calculate a size loss based on the first size information and the second size information;
[0053] Calculate an orientation loss based on the first orientation information and the second orientation information;
[0054] Adjust the parameters of the 3D occupancy grid prediction model based on the center position loss, the position offset loss, the size loss, and the orientation loss.
[0055] In a second aspect, an embodiment of the present application provides a semantic occupancy grid prediction device, where the device includes:
[0056] A training module for training a rain removal model to obtain a trained rain removal model;
[0057] An acquisition module for acquiring a 3D occupancy grid prediction model constructed based on the trained rain removal model;
[0058] A processing module for inputting consecutive image frames to be processed into the 3D occupancy grid prediction model to obtain a semantic occupancy grid prediction result of the consecutive image frames to be processed; matching the semantic occupancy grid prediction result of the consecutive image frames to be processed based on the Hungarian algorithm to obtain a first semantic occupancy grid prediction result of multiple time frames with an association relationship; correcting the first semantic occupancy grid prediction result based on Kalman filtering, and determining a stability parameter of the corrected first semantic occupancy grid prediction result based on a verification model; when the stability parameter meets the output condition, using the corrected first semantic occupancy grid prediction result as the output of the consecutive image frames to be processed; when the stability parameter does not meet the output condition, adjusting the parameters of the 3D occupancy grid prediction model; based on the adjusted 3D occupancy grid prediction model, returning to execute inputting the consecutive image frames to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction result of the consecutive image frames to be processed until the stability parameter meets the output condition.
[0059] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the semantic occupancy grid prediction method are implemented.
[0060] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the semantic occupancy grid prediction method are implemented.
[0061] The present application has the following beneficial effects:
[0062] In the present application, a trained rain removal model is obtained by training a rain removal model, a 3D occupancy grid prediction model is constructed based on the trained rain removal model, a continuous image frame to be processed is input into the 3D occupancy grid prediction model, and a semantic occupancy grid prediction result of the continuous image frame to be processed is obtained. Based on the Hungarian algorithm, the semantic occupancy grid prediction results of the continuous image frame to be processed are matched to obtain a first semantic occupancy grid prediction result of multiple time frames with an association relationship. The first semantic occupancy grid prediction result is corrected based on the Kalman filter, and the stability parameter of the corrected first semantic occupancy grid prediction result is determined based on the verification model. When the stability parameter meets the output condition, the corrected first semantic occupancy grid prediction result is used as the output of the continuous image frame to be processed. When the stability parameter does not meet the output condition, the parameters of the 3D occupancy grid prediction model are adjusted. Based on the adjusted 3D occupancy grid prediction model, the process returns to execute inputting the continuous image frame to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction result of the continuous image frame to be processed until the stability parameter meets the output condition. It can dynamically adjust the parameters of the 3D occupancy grid prediction model, ensure the continuous optimization of the prediction result, thereby improving the output occupancy grid prediction result, and further effectively improving the timeliness and stability of the autonomous driving perception module. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0064] Figure 1 It is a block diagram of the electronic device provided by the embodiment of the present invention;
[0065] Figure 2 It is one of the flow diagrams of a semantic occupancy grid prediction method provided by the embodiment of the present invention;
[0066] Figure 3 It is the second flowchart diagram of a semantic occupancy grid prediction method provided by an embodiment of the present invention;
[0067] Figure 4 It is the third flowchart diagram of a semantic occupancy grid prediction method provided by an embodiment of the present invention;
[0068] Figure 5 It is the fourth flowchart diagram of a semantic occupancy grid prediction method provided by an embodiment of the present invention;
[0069] Figure 6 It is the fifth flowchart diagram of a semantic occupancy grid prediction method provided by an embodiment of the present invention;
[0070] Figure 7 It is the sixth flowchart diagram of a semantic occupancy grid prediction method provided by an embodiment of the present invention;
[0071] Figure 8 It is the structural block diagram of a semantic occupancy grid prediction device provided by an embodiment of the present invention. Detailed implementation manners
[0072] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the figures herein can be arranged and designed in a variety of different configurations.
[0073] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but is merely representative of selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of the present invention.
[0074] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, it need not be further defined and explained in subsequent figures.
[0075] In the description of the present invention, it should be noted that if terms such as "upper", "lower", "inner", "outer", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings or the orientation or positional relationship in which the product of the invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0076] In addition, terms such as "first" and "second" are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0077] In the description of the present application, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "installed", "connected", and "linked" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0078] Through a large number of studies, it is found that the existing high-efficiency occupancy grid modeling and stability optimization in the driving environment are not satisfactory.
[0079] In view of the discovery of the above problems, this embodiment provides a semantic occupancy grid prediction method, an electronic device, and a storage medium. It can train a rain removal model to obtain a trained rain removal model, construct a 3D occupancy grid prediction model based on the trained rain removal model, input a continuous image frame to be processed into the 3D occupancy grid prediction model to obtain a semantic occupancy grid prediction result of the continuous image frame to be processed, match the semantic occupancy grid prediction result of the continuous image frame to be processed based on the Hungarian algorithm to obtain a first semantic occupancy grid prediction result of multiple time frames with an association relationship, correct the first semantic occupancy grid prediction result based on the Kalman filter, and determine the stability parameter of the corrected first semantic occupancy grid prediction result based on the verification model. When the stability parameter meets the output condition, the corrected first semantic occupancy grid prediction result is used as the output of the continuous image frame to be processed. When the stability parameter does not meet the output condition, the parameters of the 3D occupancy grid prediction model are adjusted, and based on the adjusted 3D occupancy grid prediction model, return to execute inputting the continuous image frame to be processed into the 3D occupancy grid prediction model to obtain a semantic occupancy grid prediction result of the continuous image frame to be processed until the stability parameter meets the output condition. It can dynamically adjust the parameters of the 3D occupancy grid prediction model to ensure the continuous optimization of the prediction result, thereby improving the output occupancy grid prediction result, and further effectively improving the timeliness and stability of the autonomous driving perception module. The solution provided in this embodiment will be elaborated in detail below.
[0080] This embodiment provides an electronic device that can predict semantic occupancy grids. In a possible implementation, the electronic device may be a user terminal. For example, the electronic device may be, but is not limited to, a server, a smart phone, a personal computer (PC), a tablet computer, a personal digital assistant (PDA), a mobile internet device (MID), etc.
[0081] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of the electronic device 100 provided by an embodiment of the present application. The electronic device 100 may further include more or fewer components than those shown in Figure 1 or have a different configuration from that shown in Figure 1 . Figure 1 Each component shown in
[0082] can be implemented by hardware, software, or a combination thereof.
[0083] The electronic device 100 includes a semantic occupancy grid prediction device 110, a memory 120, and a processor 130.
[0084] Among them, the memory 120 can be, but is not limited to, a Random Access Memory (RAM), a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electric Erasable Programmable Read-Only Memory (EEPROM), etc. Among them, the memory 120 is used to store a program, and after receiving an execution instruction, the processor 130 executes the program.
[0085] Please refer to Figure 2 , Figure 2 for a flowchart of a semantic occupancy grid prediction method for an electronic device 100 applied to Figure 1 . The following will elaborate on each step included in the method in detail.
[0086] S201: Train a rain removal model to obtain a trained rain removal model.
[0087] S202: Obtain a 3D occupancy grid prediction model constructed based on the trained rain removal model.
[0088] S203: Input the continuous image frames to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction result of the continuous image frames to be processed.
[0089] S204: Match the semantic occupancy grid prediction result of the continuous image frames to be processed based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction result of multiple time frames with an associated relationship.
[0090] S205: Correct the first semantic occupancy grid prediction result based on the Kalman filter, and determine the stability parameter of the corrected first semantic occupancy grid prediction result based on the verification model.
[0091] S206: When the stability parameter meets the output condition, use the corrected first semantic occupancy grid prediction result as the output of the continuous image frames to be processed.
[0092] S207: When the stability parameter does not meet the output condition, adjust the parameters of the 3D occupancy grid prediction model.
[0093] S208: Based on the adjusted 3D occupancy grid prediction model, return the semantic occupancy grid prediction result of the continuous image frames to be processed by inputting them into the 3D occupancy grid prediction model until the stability parameter meets the output condition.
[0094] The main function of the rain removal model is to remove interferences such as raindrops and rain streaks in images or videos and restore a clear background image. The rain removal model can help the vehicle's vision system more accurately identify roads, pedestrians, and obstacles in rainy days.
[0095] There are various ways to train the rain removal model. In one implementation: A large number of paired rainy images and clean images are required as training data. Operations such as cropping, scaling, and normalizing the images are performed. Data diversity can also be increased through data augmentation techniques (such as rotation, flipping, adding noise, etc.). Commonly used convolutional neural networks (CNNs) or generative adversarial networks (GANs) are used. For example, PReNet is an efficient recursive network, and the rain removal effect can be further improved by introducing an attention mechanism. Design a multi-layer convolutional network to extract image features, and optimize the model through specific loss functions (such as L1 loss, SSIM loss). Usually, L1 loss, SSIM loss, or a combination of both is used to measure the rain removal effect. The Adam optimizer is commonly used for parameter updates. Set the number of training epochs, learning rate adjustment strategy, and monitor the model performance using a validation set to complete the training of the rain removal model.
[0096] Based on the trained rain-removed images, construct a 3D occupancy grid prediction model. Using the pre-trained rain removal model, preprocess multi-view images, extract high-dimensional image features, and generate a three-dimensional occupancy grid descriptor through cascaded voxel decoding to accurately represent the objects and scene details in the three-dimensional space. Use the Hungarian algorithm to match the semantic occupancy grid prediction results of the continuous image frames to be processed, further apply smoothing algorithms such as Kalman filtering to correct the first semantic occupancy grid prediction result, and quantify the stability of the semantic occupancy grid by calculating stability indicators considering confidence, position, size, and orientation. Determine the stability parameter of the corrected first semantic occupancy grid prediction result, select an optimization strategy according to the feedback result of the stability parameter, iteratively train the 3D occupancy grid prediction model, and re-evaluate the stability index to ensure the high stability and accuracy of the occupancy grid prediction result output by the 3D occupancy grid prediction model in the spatio-temporal domain.
[0097] In another implementation of training the rain removal model to obtain a trained rain removal model, as Figure 3 shown, it includes the following steps:
[0098] S301: Obtain rainy images that meet the preset conditions from the driving dataset.
[0099] S302: Generate a semantic segmentation image for the rainy-day image based on a generative algorithm.
[0100] S303: Optimize the semantic segmentation image based on an optimal transport strategy.
[0101] S304: Train the rain removal model based on the optimized semantic segmentation image to obtain a trained rain removal model.
[0102] In the driving dataset, sunny-day and rainy-day images can be selected, and invalid or unlabeled images can be removed to ensure the reliability of the dataset and provide guarantee for subsequent model training. To further improve the generalization ability of the rain removal model, the rainy-day images are screened again, and more challenging samples are selected to help the model better handle complex rainy-day scenarios in the real world. The images after the second screening are used as the rainy-day images that meet the preset conditions. Among them, the images that meet the preset conditions cover various characteristics such as dense raindrops, thick fog, and complex light reflections, ensuring their representativeness and high difficulty, which helps the algorithm learn more effective features during the training process, thereby improving its performance in practical applications.
[0103] The rainy-day images are divided into a training set and a test set to facilitate the evaluation of the model's performance. Generating a semantic segmentation image for the rainy-day image based on a generative algorithm can generate semantic segmentation labels and save the semantic segmentation results for direct use in the subsequent training process, which can greatly accelerate the training process of the rain removal model and also provide strong support for subsequent image processing tasks.
[0104] Optimizing the semantic segmentation image based on an optimal transport strategy can optimize the weight distribution of image features, thereby significantly improving the effect of contrast learning.
[0105] There are various implementation methods for constructing a 3D occupancy grid prediction model based on the trained rain removal model. In one implementation method, as Figure 4 shown, it includes the following steps:
[0106] S401: Input the rainy-day image of the real scene into the trained rain removal model to obtain multiple de-rained images from multiple perspectives.
[0107] S402: Calibrate and synchronize the time of each de-rained image respectively to obtain multiple first images from multiple perspectives with consistent time.
[0108] S403: For each first image, based on the semantic features of the first image, the high-dimensional feature representation tensor of the first image, and the vehicle geometry information, obtain the BEV feature of the first image.
[0109] S404: For each BEV feature, determine the current BEV feature and the historical BEV feature corresponding to the BEV feature.
[0110] S405: Generate a three-dimensional occupancy descriptor based on the current BEV feature and the historical BEV feature to complete the construction of the 3D occupancy grid prediction model.
[0111] The specific method for calibrating and time-synchronizing multiple de-rained images to obtain multiple first images at multiple perspectives with consistent time is as follows: Based on the internal and external parameters of the camera, calibrate the multiple de-rained images at multiple perspectives to obtain the calibrated multiple de-rained images at multiple perspectives. Perform time synchronization on the calibrated multiple de-rained images at multiple perspectives to obtain multiple first images at multiple perspectives with consistent time.
[0112] The specific method for obtaining the BEV feature of each first image based on the semantic feature of the first image, the high-dimensional feature representation tensor of the first image, and the vehicle geometry information is as follows: For each first image, obtain the low-order feature and the semantic feature of the first image, and based on the low-order feature and the semantic feature, obtain the high-dimensional feature representation tensor of the first image. For multiple first images, obtain the BEV feature based on the high-dimensional feature representation tensors corresponding to the respective first images and the vehicle geometry information.
[0113] The specific method for determining the current BEV feature and the historical BEV feature corresponding to the BEV feature for each BEV feature, and generating a three-dimensional occupancy descriptor based on the current BEV feature and the historical BEV feature to complete the construction of the 3D occupancy grid prediction model is as follows: For each BEV feature, use the BEV feature as the current BEV feature, and obtain the historical BEV feature corresponding to the BEV feature. Perform spatial transformation, unify to the same coordinate system, and feature alignment processing on the current BEV feature and the historical BEV feature corresponding to the current BEV feature in sequence to obtain the aligned current BEV feature and the aligned historical BEV feature. Based on the aligned current BEV feature and the corresponding historical BEV feature, obtain the temporal change feature. Based on the temporal change feature, supplement or delete the aligned current BEV feature to obtain a new current BEV feature. Based on the new current BEV feature and the aligned historical feature, determine the semantic occupancy grid movement trend and scene change. Based on the multi-head attention mechanism, the 3D cross-attention mechanism, the new current BEV feature, the aligned historical feature, the semantic occupancy grid movement trend, and the scene change, obtain the final BEV feature after fusing the temporal sequence of the new current BEV feature. Based on the final BEV feature, generate a three-dimensional occupancy descriptor to complete the construction of the 3D occupancy grid prediction model.
[0114] Among them, the three-dimensional occupancy descriptor is used to describe the prediction result of the semantic occupancy grid.
[0115] Use the rainy-day images obtained from the real scene as the input to the trained rain-removal model. The rainy-day images obtained from the real scene are multi-view rainy-day images acquired by a vehicle. The multi-view rainy-day images include multiple rainy-day images from different perspectives. The trained rain-removal model outputs multi-view rain-removed images of the rainy-day images in the real scene, that is, multiple rain-removed images from different perspectives.
[0116] Perform correction and time synchronization on the images of each perspective through image preprocessing, and complete the geometric calibration of the multi-view images based on the internal and external parameters of the camera to obtain complete surrounding environment information.
[0117] Use the internal parameters of the camera, such as the focal length ( fx , fy ), the principal point ( cx , cy ), and the distortion coefficients ( k 1, k 2, p 1, p 2) to perform radial and tangential distortion correction on the images of each perspective, and obtain the calibrated multi-view rain-removed images.
[0118] The method for performing time synchronization on the calibrated multi-view rain-removed images can be: based on the image acquisition timestamps, select the frame with the closest time or perform interpolation, estimate the inter-frame motion and adjust the time offset in post-processing, and finally obtain the first multi-view images with consistent time. That is, multiple first images corresponding to different perspectives.
[0119] For each first image, obtain the low-order features and semantic features of the first image, and based on the low-order features and semantic features, obtain the high-dimensional feature representation tensor of the first image. Among them, the semantic features indicate the labels of the objects in the first image.
[0120] Exemplarily, use a convolutional neural network (CNN) to extract low-order features (such as edges, textures, colors, etc.). Gradually extract higher-level features through multiple layers of convolution (such as architectures like ResNet, VGG, etc.). Perform global relationship modeling on the low-order features and semantic features through a multi-head attention mechanism, and finally obtain the high-dimensional feature representation tensors of the first images corresponding to each perspective.
[0121] Exemplarily, when the first image includes the first image corresponding to the first perspective, the first image corresponding to the second perspective, and the first image corresponding to the third perspective, respectively determine the first high-dimensional feature representation of the first image corresponding to the first perspective, determine the second high-dimensional feature representation of the first image corresponding to the second perspective, and the third high-dimensional feature representation of the first image corresponding to the third perspective. Determine the vehicle geometric information, that is, taking the vehicle as the origin, construct a three-dimensional coordinate system, and input the high-dimensional feature representations at the same time into this three-dimensional coordinate system, that is, input the first high-dimensional feature representation, the second high-dimensional feature representation, and the third high-dimensional feature representation into this three-dimensional coordinate system to obtain the BEV feature, that is, realize the fusion of images from multiple perspectives to obtain the fused BEV feature. The fused BEV feature carries the three-dimensional spatial features of vehicle geometric information and semantic features.
[0122] When processing each BEV feature, regard the currently processed BEV feature as the current BEV feature, and obtain the historical BEV feature of the current BEV feature.
[0123] When the rainy-day image in the real scene contains 3s of image frames and contains images from multiple perspectives at 3 different times, when processing the images from multiple perspectives at 2s, obtain the current BEV feature. Then the BEV feature corresponding to the 1st second is the historical BEV feature corresponding to the current BEV feature. When processing the images from multiple perspectives at the 3rd second, obtain the current BEV feature. Then the BEV feature corresponding to the 1st second and the BEV feature corresponding to the 2nd second are the historical BEV features corresponding to the current BEV feature.
[0124] Perform spatial transformation, unify to the same coordinate system, and feature alignment processing on the current BEV feature and the historical BEV feature in sequence to obtain the aligned current BEV feature and the aligned historical BEV feature.
[0125] The time alignment technology can be used to perform spatial transformation on the historical features, unify the historical BEV feature and the current BEV feature to the same coordinate system, and align the current BEV feature and the historical BEV feature through the three-dimensional deformable attention mechanism to obtain the aligned current BEV feature and the aligned historical BEV feature.
[0126] Based on the aligned current BEV feature and the corresponding historical BEV feature, obtain the temporal change feature to calibrate the current BEV feature.
[0127] The method for determining the temporal change features can be as follows: calculate the difference between the current BEV feature and each historical BEV feature, perform temporal modeling on the historical BEV features using 3D convolution, model the historical BEV features using an RNN neural network, perform global temporal modeling on the historical BEV feature sequence using a Transformer encoder, concatenate the current BEV feature and the temporal change features along the channel dimension, dynamically fuse the current BEV feature and the temporal change features using an attention mechanism, use a residual connection to retain the information of the current BEV feature, finally obtain the temporal change features, determine whether to supplement or delete the current BEV feature based on the temporal change features, and supplement or delete the current BEV feature based on the temporal change features, so as to achieve the calibration of the current BEV feature based on the historical BEV features and obtain a new current BEV feature.
[0128] Input the new current BEV feature and the aligned BEV feature into a recurrent neural network, and the recurrent neural network captures the semantic occupancy grid motion trend and scene changes of the object in consecutive frames through iterative calculations of time steps.
[0129] The recurrent neural network can combine the features of historical frames to learn the motion trajectory of the object and the dynamic changes of the scene. Since the recurrent neural network can retain long-term dependency information when processing time series, combined with more complex structures such as LSTM or GRU, it can predict the motion trend of the object and ensure the dynamic consistency and stability of the semantic occupancy grid between different time frames.
[0130] The multi-head attention mechanism performs global relationship modeling on the new current BEV feature, and the 3D cross-attention mechanism performs spatio-temporal fusion of the aligned historical BEV feature and the new current BEV feature. Considering the semantic occupancy grid motion trend, the motion trend can be fused with the spatio-temporal features, and considering the scene changes, the features after fusing the scene change information with the motion trend are further fused, so as to generate high-quality temporally fused BEV features and provide strong support for tasks such as object detection and behavior prediction.
[0131] After feature processing, the self-attention mechanism is used to help capture the global dependencies between time frames, especially between different time steps, and can focus on key information regions. The 3D cross-attention mechanism is used to align the BEV features at different time frames and perspectives, further improving the model's ability to model the object motion trend and scene changes. Through this precise alignment and fusion, the model can reduce the prediction fluctuations caused by environmental changes or motion, enabling the semantic occupancy grid motion trend and scene changes to be captured more stably and accurately.
[0132] Generate a 3D occupancy descriptor from the above - fused final BEV features. The 3D occupancy descriptor is a feature representation method used to represent the occupancy state of objects or spaces in a 3D scene. It provides a fine - grained description of the scene by quantifying the physical 3D scene into a structured grid map, with each cell having a semantic label or occupancy state.
[0133] There are multiple implementation methods for matching prediction boxes within multiple time frames based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with an associated relationship. In one implementation method, as Figure 5 shown, it includes the following steps:
[0134] S501: For each semantic occupancy grid prediction result, determine the historical semantic occupancy grid prediction result of the semantic occupancy grid prediction result.
[0135] S502: Determine the first parameter information of the semantic occupancy grid prediction result and the second parameter information of the historical semantic occupancy grid prediction result.
[0136] S503: Based on the first parameter information and the second parameter information, determine the difference value between the semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result.
[0137] S504: The Hungarian algorithm processes each difference value to obtain the first semantic occupancy grid prediction results of multiple time frames with an associated relationship.
[0138] Input the image data to be processed into the constructed 3D occupancy grid prediction model, and obtain the semantic occupancy grid prediction results of the consecutive image frames to be processed. Based on the Hungarian algorithm, the semantic occupancy grids at different time stamps are associated to obtain the first semantic occupancy grid prediction results of multiple time frames with an associated relationship.
[0139] Based on the constructed 3D occupancy grid prediction model, information such as the position, size, and category of the object is extracted from each frame to construct the current semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result. By calculating the first parameter information of the current semantic occupancy grid prediction result and the second parameter information of the historical semantic occupancy grid prediction result, the difference value between the semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result is obtained, and this difference value is used to evaluate the matching cost. The Hungarian algorithm is used to minimize these costs to find the optimal match between each pair of semantic occupancy grid prediction results, thereby establishing the association between objects in multiple time frames. Finally, the matching results of each time frame are output to form the time series trajectory of the object, that is, the first semantic occupancy grid prediction results of multiple time frames with an association relationship. This not only ensures the association consistency of the same object between different time frames but also effectively reduces the prediction fluctuations caused by motion or environmental changes, improving the stability and accuracy of object tracking.
[0140] Based on the first parameter information and the second parameter information, the implementation method for determining the difference value between the semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result can be:
[0141] The first parameter information includes the position (x1, y1, z1), size (l1, w1, h1), and orientation (θ1) of the semantic occupancy grid prediction result, and the second parameter information includes the position (x2, y2, z2), size (l2, w2, h2), and orientation (θ2) of the historical semantic occupancy grid prediction result.
[0142] The calculation formula for the difference value between the semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result is as follows:
[0143] Cost(B1,B2)=||x1 - x2|| + ||y1 - y2|| + ||z1 - z2|| + ||l1 - l2|| + ||ω1 - ω2|| + ||h1 - h2|| + ||θ1 - θ2||.
[0144] Where, B1 represents the semantic occupancy grid prediction result, B2 represents the historical semantic occupancy grid prediction result, the first parameter information includes the position (x1, y1, z1), size (l1, w1, h1), and orientation (θ1) of the semantic occupancy grid prediction result, and the second parameter information includes the position (x2, y2, z2), size (l2, w2, h2), and orientation (θ2) of the historical semantic occupancy grid prediction result.
[0145] That is, the matching cost is evaluated by calculating the Euclidean distance between the central positions of the semantic occupancy grid prediction results and the historical semantic occupancy grid prediction results, the change in size, and the consistency of the class labels. The Hungarian algorithm is used to minimize these costs to find the optimal match between each pair of prediction boxes, thereby establishing the association between objects in multiple time frames.
[0146] The Hungarian algorithm performs global optimization matching on the difference values of the semantic occupancy grid prediction results, ensuring the association consistency of the same semantic occupancy grid prediction results among multiple time frames and minimizing the prediction fluctuations caused by motion or environmental changes. After the matching is completed, the time association results of each semantic occupancy grid are output, that is, the first semantic occupancy grid prediction result, for subsequent correction and stability optimization.
[0147] The first semantic occupancy grid prediction result is corrected by the Kalman filter smoothing algorithm to improve the stability and consistency of the semantic occupancy grid detection in the time series.
[0148] By extracting the geometric features of the first semantic occupancy grid prediction result, including information such as the central position, size, and orientation, the association between multiple time frames is established.
[0149] Among them, the geometric features of the first semantic occupancy grid prediction result include low-order geometric features and high-order dynamic information. The low-order geometric features, such as position coordinates and size changes, can be obtained through simple geometric algebra calculations. The high-order dynamic information, such as the change in the orientation angle, is further quantified by combining the motion trajectory of the semantic occupancy grid. Subsequently, based on the associated semantic occupancy grid trajectory, the Kalman filter is applied to smoothly predict the position of the semantic occupancy grid.
[0150] The application of the Kalman filter in the position smoothing prediction of the first semantic occupancy grid prediction result is achieved by jointly modeling the historical trajectory of the first semantic occupancy grid prediction result and the current observed first semantic occupancy grid prediction result for position smoothing.
[0151] For the size change, the average size of the first semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result of the first semantic occupancy grid prediction result is used for scale constraint to ensure the consistency of the size of the first semantic occupancy grid prediction result in the time series.
[0152] The size of each frame is compared with the sizes of its adjacent frames, and the size of the current frame is corrected by calculating the average size of historical frames. This can effectively reduce the size fluctuations caused by detection errors or short-term environmental changes. To further optimize the smooth transition of the size, size consistency regularization is introduced, which restricts the drastic fluctuations of the size in the time series and ensures that the size change of the first semantic occupancy grid prediction result is within a reasonable range. In addition, the IoU method is used to optimize the size consistency. By calculating the overlap degree between the first semantic occupancy grid prediction result of the current frame and the first semantic occupancy grid prediction results of adjacent frames, it is ensured that the sizes of the first semantic occupancy grid prediction results of adjacent frames are consistent, reducing the deviation of size prediction. A high IoU value indicates a small size change, meeting the stability requirements of the object in the time dimension. In this way, the smooth correction and consistency optimization of the size are guaranteed, thus enhancing the stability and continuity of the semantic occupancy grid in the time dimension.
[0153] Regarding the orientation information of the first semantic occupancy grid prediction result, in combination with the historical first semantic occupancy grid prediction results of the first semantic occupancy grid prediction result, the orientation angle of the current first semantic occupancy grid prediction result is smoothly corrected. If the orientation change between adjacent frames exceeds the threshold, it is restricted within a reasonable range to avoid the influence of drastic angle jumps.
[0154] Through the above geometric correction and time series smoothing methods, the corrected first semantic occupancy grid prediction results are aligned to a unified stable state.
[0155] There are multiple ways to determine the stable parameters of the corrected first semantic occupancy grid prediction result. In one implementation, as Figure 6 shown, it includes the following steps:
[0156] S601: For each corrected first semantic occupancy grid prediction result, determine the corresponding historical first semantic occupancy grid prediction result of the corrected first semantic occupancy grid prediction result.
[0157] S602: Determine the confidence stability value, position stability value, size stability value, and orientation stability value between the corrected first semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result.
[0158] S603: Based on the confidence stability value, position stability value, size stability value, and orientation stability value, determine the stable parameters of the corrected first semantic occupancy grid prediction result.
[0159] Specifically: Determine the first confidence level of the corrected first semantic occupancy grid prediction result, determine the second confidence level of the historical semantic occupancy grid prediction result, based on the first confidence level and the second confidence level, determine the confidence level stability value of the corrected first semantic occupancy grid prediction result, determine the first position information of the corrected first semantic occupancy grid prediction result and the second position information of the historical first semantic occupancy grid prediction result, and based on the first position information and the second position information, determine the position stability value of the corrected first semantic occupancy grid prediction result.
[0160] The position stability value is measured by calculating the intersection over union (IoU) of the position information (center point) of two corrected first semantic occupancy grid prediction results. The position-corrected semantic occupancy grid reduces the offset of the center point.
[0161] Determine the first size information of the corrected first semantic occupancy grid prediction result and the second size information of the historical first semantic occupancy grid prediction result, based on the first size information and the second size information, determine the size stability value of the corrected first semantic occupancy grid prediction result, determine the first orientation information of the corrected first semantic occupancy grid prediction result and the second orientation information of the historical first semantic occupancy grid prediction result, and based on the first orientation information and the second orientation information, determine the orientation stability value of the corrected first semantic occupancy grid prediction result. Based on the confidence level stability value, the position stability value, the size stability value, and the orientation stability value, determine the stability parameter of the corrected first semantic occupancy grid prediction result.
[0162] The confidence level stability value satisfies the following formula:
[0163] ;
[0164] Where is the confidence level stability value, is the first confidence level, is the second confidence level, and are the 99% and 1% percentiles of the confidence level;
[0165] The position stability value satisfies the following formula:
[0166] ;
[0167] Where is the position stability value, is the first position information, is the second position information;
[0168] The size stability value satisfies the following formula:
[0169] ;
[0170] Among them, is the dimensional stability value, is the first dimensional information, is the second dimensional information;
[0171] The orientation stability value satisfies the following formula:
[0172] ;
[0173] Among them, is the orientation stability value, is the first orientation information, is the second orientation information, is the orientation angle corresponding to the first orientation information, is the orientation angle corresponding to the second orientation information.
[0174] The orientation information indicates the coordinate information of the front and rear of the vehicle, and the orientation angle is calculated based on the orientation information.
[0175] The final stability parameter is:
[0176] .
[0177] When the stability parameter of the corrected first semantic occupancy grid prediction result does not meet the output conditions, there are various implementation methods for adjusting the parameters of the 3D occupancy grid prediction model. In one implementation method, as Figure 7 shown, it includes the following steps:
[0178] S701: Determine the total number of time frames of the continuous image frames to be processed.
[0179] S702: Calculate the position offset loss based on the first position information, the second position information, and the total number of time frames.
[0180] S703: Calculate the dimensional loss based on the first dimensional information and the second dimensional information.
[0181] S704: Calculate the orientation loss based on the first orientation information and the second orientation information.
[0182] S705: Adjust the parameters of the 3D occupancy grid prediction model based on the center position loss, the position offset loss, the dimensional loss, and the orientation loss.
[0183] The center position loss satisfies the following formula:
[0184] ;
[0185] is the center position loss, is the central position of the first position information of the corrected first semantic occupancy grid prediction result; is the observation matrix, usually the identity matrix in one-dimensional or two-dimensional space, is the Kalman gain.
[0186] The position offset loss satisfies the following formula:
[0187] ;
[0188] T is the total number of time frames, is the position offset loss, is the first position information, is the second position information.
[0189] The size loss satisfies the following formula:
[0190] ; ;
[0191] and are the width and height of each frame respectively.
[0192] The orientation loss satisfies the following formula:
[0193] ;
[0194] When is greater than the preset threshold, the angle is corrected to keep it changing smoothly.
[0195] ;
[0196] is the first orientation information, is the second orientation information.
[0197] Based on the center position loss, position offset loss, size loss, and orientation loss, the parameters of the 3D occupancy grid prediction model are adjusted to obtain the 3D occupancy grid prediction model with adjusted parameters.
[0198] Input the continuous image frames to be processed into the 3D occupancy grid prediction model after parameter adjustment to obtain the semantic occupancy grid prediction results of the continuous image frames to be processed; match the semantic occupancy grid prediction results of the continuous image frames to be processed based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with an associated relationship; correct the first semantic occupancy grid prediction results based on the Kalman filter, and determine the stability parameters of the corrected first semantic occupancy grid prediction results based on the verification model; when the stability parameters meet the output conditions, then use the corrected first semantic occupancy grid prediction results as the output of the continuous image frames to be processed; when the stability parameters meet the output conditions, then output the corrected first semantic occupancy grid prediction results.
[0199] Based on the stability parameters as the termination condition for model optimization to comprehensively quantify the spatio-temporal stability performance of the 3D occupancy grid prediction model in a dynamic scenario. Specifically, for the stability sub-indicators of confidence, position, size, and orientation, calculate the change ranges before and after optimization respectively, and comprehensively evaluate whether the model performance meets the preset requirements, that is, whether the stability parameters meet the requirement of SI>0.8, where SI is the stability parameter.
[0200] When the stability parameter is greater than 0.8, it is determined that the stability parameter meets the output conditions.
[0201] Please refer to Figure 8 In addition, the embodiment of the present application also provides a semantic occupancy grid prediction device 110 applied to Figure 1 the electronic device 100. The semantic occupancy grid prediction device 110 includes:
[0202] A training module 111, configured to train the de-raining model to obtain a trained de-raining model;
[0203] An acquisition module 112, configured to construct a 3D occupancy grid prediction model based on the trained de-raining model;
[0204] The processing module 113 is configured to input consecutive image frames to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction result of the consecutive image frames to be processed; match the semantic occupancy grid prediction results of the consecutive image frames to be processed based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with an associated relationship; correct the first semantic occupancy grid prediction results based on the Kalman filter, and determine the stability parameters of the corrected first semantic occupancy grid prediction results based on the verification model; when the stability parameters meet the output conditions, use the corrected first semantic occupancy grid prediction results as the output of the consecutive image frames to be processed; when the stability parameters do not meet the output conditions, adjust the parameters of the 3D occupancy grid prediction model; based on the adjusted 3D occupancy grid prediction model, return to execute inputting the consecutive image frames to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction result of the consecutive image frames to be processed until the stability parameters meet the output conditions.
[0205] This application also provides an electronic device 100, which includes a processor 130 and a memory 120. When the computer-executable instructions stored in the memory 120 are executed by the processor 130, the semantic occupancy grid prediction method is implemented.
[0206] This application embodiment also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by the processor 130, the semantic occupancy grid prediction method is implemented.
[0207] In the embodiments provided in this application, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the apparatus, method, and computer program product according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0208] In addition, each functional module in various embodiments of the present application may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0209] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the said element.
[0210] As described above, these are only various implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A semantic occupancy grid prediction method, characterized in that, The method includes: Training a deraining model to obtain a trained deraining model; Constructing a 3D occupancy grid prediction model based on the trained deraining model; Inputting the continuous image frames to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction results of the continuous image frames to be processed; Matching the semantic occupancy grid prediction results of the continuous image frames to be processed based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with an associated relationship; Correcting the first semantic occupancy grid prediction results based on the Kalman filter and determining the stability parameters of the corrected first semantic occupancy grid prediction results based on the verification model; When the stability parameters meet the output conditions, using the corrected first semantic occupancy grid prediction results as the output of the continuous image frames to be processed; When the stability parameters do not meet the output conditions, adjusting the parameters of the 3D occupancy grid prediction model; Based on the adjusted 3D occupancy grid prediction model, return to execute inputting the continuous image frames to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction results of the continuous image frames to be processed until the stability parameters meet the output conditions.
2. The method according to claim 1, wherein The step of training the deraining model to obtain a trained deraining model includes: Obtaining rainy-day images that meet the preset conditions from the driving dataset; Generating semantic segmentation images for the rainy-day images based on the generative algorithm; Optimizing the semantic segmentation images based on the optimal transport strategy; Training the deraining model based on the optimized semantic segmentation images to obtain a trained deraining model.
3. The method according to claim 1, wherein The step of constructing a 3D occupancy grid prediction model based on the trained deraining model includes: Inputting the rainy-day images of the real scene into the trained deraining model to obtain multiple derained images from multiple perspectives; Calibrating and time-synchronizing the multiple derained images respectively to obtain multiple first images from multiple perspectives with consistent time; For each first image, obtaining the BEV feature of the first image based on the semantic feature of the first image, the high-dimensional feature representation tensor of the first image, and the vehicle geometry information; For each BEV feature, determining the current BEV feature and the historical BEV feature corresponding to the BEV feature; Generating a three-dimensional occupancy descriptor based on the current BEV feature and the historical BEV feature to complete the construction of the 3D occupancy grid prediction model.
4. The method according to claim 1, wherein The step of matching the semantic occupancy grid prediction results of the continuous image frames to be processed based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with an associated relationship includes: For each semantic occupancy grid prediction result, determining the historical semantic occupancy grid prediction result of the semantic occupancy grid prediction result; Determining the first parameter information of the semantic occupancy grid prediction result and the second parameter information of the historical semantic occupancy grid prediction result; Based on the first parameter information and the second parameter information, determining the difference value between the semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result; The Hungarian algorithm processes each of the said difference values to obtain a first semantic occupancy grid prediction result of multiple time frames with an associated relationship.
5. The method according to claim 1, wherein The steps of determining the stability parameters of the corrected first semantic occupancy grid prediction result include: For each of the corrected first semantic occupancy grid prediction results, determining the corresponding historical first semantic occupancy grid prediction result of the corrected first semantic occupancy grid prediction result; Determining the confidence stability value, position stability value, size stability value, and orientation stability value of the corrected first semantic occupancy grid prediction result and the historical first semantic occupancy grid prediction result; Based on the confidence stability value, position stability value, size stability value, and orientation stability value, determining the stability parameters of the corrected first semantic occupancy grid prediction result.
6. The method according to claim 5, wherein The confidence stability value satisfies the following formula: ; Among them, is the confidence stability value, is the first confidence level, is the second confidence level, and are the 99% and 1% percentiles of the confidence level; The position stability value satisfies the following formula: ; Among them, is the position stability value, is the first position information, is the second position information; The size stability value satisfies the following formula: ; Among them, is the dimensional stability value, is the first dimensional information, is the second dimensional information; The orientation stability value satisfies the following formula: ; wherein, is the stable value of the orientation, is the first orientation information, is the second orientation information, is the orientation angle corresponding to the first orientation information, is the orientation angle corresponding to the second orientation information.
7. The method according to claim 6, wherein The steps of adjusting the parameters of the 3D occupancy grid prediction model when the stability parameters of the corrected first semantic occupancy grid prediction result do not meet the output conditions include: When the stability parameters of the corrected first semantic occupancy grid prediction result do not meet the output conditions, calculating a center position loss based on the first position information and the second position information, where the first position information and the second position information indicate the center position; Determining the total number of time frames of the consecutive image frames to be processed; Calculating a position offset loss based on the first position information, the second position information, and the total number of time frames; Calculating a size loss based on the first size information and the second size information; Calculating an orientation loss based on the first orientation information and the second orientation information; Adjusting the parameters of the 3D occupancy grid prediction model based on the center position loss, the position offset loss, the size loss, and the orientation loss.
8. An electronic device, characterized in that, Comprising a memory and a processor, the memory stores a computer program, characterized in that when the processor executes the computer program, the steps of the method according to any one of claims 1-7 are implemented.
9. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Target motion state detection method and device, mobile device and storage medium
CN116309693A
Lightweight occupancy grid prediction method and system based on large model self-labeling
CN118823139A