Semantic occupancy grid prediction method, electronic equipment and storage medium
By training the rain removal model and building a 3D occupancy raster prediction model, combined with the scoop-yali algorithm and Kalman filtering, the problem of insufficient prediction stability of the autonomous driving system under complex weather conditions is solved, and the timeliness and stability of the perception module is improved.
Patent Information
- Application Number
- CN202510512112.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
AI Technical Summary
In complex weather conditions, it is difficult to achieve stable semantic occupancy grid prediction in rainy days, resulting in the impact of perceptual accuracy and safety.
By training the rain-destroying model, a 3D occupancy raster prediction model is constructed, and combined with Hungarian algorithms and Kalman filtering, the model parameters are dynamically adjusted to ensure the stability of the prediction results.
The timeliness and stability of the autonomous driving perception module is improved, ensuring that high accuracy and reliability of semantic occupancy grid prediction can be achieved under complex weather conditions.
Smart Images

Figure CN120032335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving perception technology, and in particular to a semantic occupancy grid prediction method, electronic device and storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, the requirements for autonomous driving technology in terms of perception accuracy, environmental adaptability and safety are becoming increasingly stringent. Especially in complex weather conditions such as rainy days, the perception and understanding of driving scenes have become a key challenge for autonomous driving systems to achieve safe driving. Occupancy grids, as a data structure that divides the environmental space into multiple small units and marks the probability of each unit being occupied by an object, are an important tool for environmental modeling in autonomous driving. Therefore, efficient occupancy grid modeling and stability optimization in driving environments to improve perception capabilities in complex weather have become an important direction for technological development. This requires not only that the model can accurately and quickly understand the semantic information of the environment, but also that it needs to maintain high stability and reliability under various conditions to ensure the safe operation of the autonomous driving system.
[0003] In the related art, one type of method uses single-view or multi-view two-dimensional images for scene reconstruction and grid prediction. These methods can achieve certain results in simple scenes. For example, the monocular vision-based method captures and processes images through a single camera, and infers the occupied grid information through image segmentation and feature extraction. This type of method performs well in a clear field of view, but when encountering complex scenes such as rainy days, there will be raindrops in the image or raindrops will cause the image to be blurred, and the image quality will be seriously reduced, resulting in a significant decrease in the accuracy and stability of the model prediction.
[0004] Another type of method combines lidar and multi-view image data. Although it has high geometric accuracy, it cannot meet the needs of real-time prediction due to the complex processing of sensor data and high computational overhead. For example, lidar can provide accurate spatial position information through point cloud data, and combined with multi-view image data, it can enhance the perception ability of the model. However, this method requires a lot of computing resources, especially when processing large amounts of point cloud data, which usually requires the support of high-performance computing devices. At the same time, existing models still have obvious deficiencies in the consistency and stability of multi-frame prediction results, which can easily lead to prediction fluctuations, thereby affecting the decision-making and planning performance of the autonomous driving system. Summary of the invention
[0005] The object of the present invention is to provide a semantic occupancy grid prediction method, electronic device and storage medium, which can improve the stability of semantic occupancy grid prediction.
[0006] In order to achieve the above purpose, the technical solution adopted in the embodiment of the present application is as follows: In a first aspect, an embodiment of the present application provides a semantic occupancy grid prediction method, the method comprising: Train the rain removal model to obtain a trained rain removal model; Obtain a 3D occupancy grid prediction model based on the trained rain removal model; Inputting the continuous image frames to be processed into the 3D occupancy grid prediction model to obtain semantic occupancy grid prediction results of the continuous image frames to be processed; Based on the Hungarian algorithm, the semantic occupancy grid prediction results of the continuous image frames to be processed are matched to obtain the first semantic occupancy grid prediction results of multiple time frames with associated relationships; Correcting the first semantic occupancy grid prediction result based on Kalman filtering, and determining a stability parameter of the corrected first semantic occupancy grid prediction result based on a verification model; When the stability parameter satisfies the output condition, the modified first semantic occupancy grid prediction result is used as the output of the continuous image frame to be processed; When the stability parameters do not meet the output conditions, adjusting the 3D occupancy grid prediction model parameters; Based on the adjusted 3D occupancy grid prediction model, return to execute inputting the continuous image frames to be processed into the 3D occupancy grid prediction model to obtain semantic occupancy grid prediction results of the continuous image frames to be processed until the stability parameters meet the output conditions.
[0007] In an optional implementation manner, the step of training the rain removal model to obtain a trained rain removal model includes: Obtain rainy day images that meet preset conditions from the driving dataset; Generating a semantic segmentation image for the rainy day image based on a generative algorithm; Optimizing the semantic segmentation image based on an optimal transmission strategy; The rain removal model is trained based on the optimized semantic segmentation image to obtain a trained rain removal model.
[0008] In an optional implementation, the step of acquiring a 3D occupancy grid prediction model based on a trained rain removal model comprises: Inputting a rainy day image of a real scene into the trained deraining model to obtain multiple derained images under multiple viewing angles; Calibrate and time-synchronize the plurality of rain-removed images respectively to obtain a plurality of first images at a plurality of viewing angles with consistent time; For each first image, based on the semantic features of the first image, the high-dimensional feature representation tensor of the first image, and the vehicle geometry information, obtain a BEV feature of the first image; For each BEV feature, determining a current BEV feature and a historical BEV feature corresponding to the BEV feature; A three-dimensional occupancy descriptor is generated based on the current BEV features and the historical BEV features to complete the construction of a 3D occupancy grid prediction model.
[0009] In an optional implementation, the step of matching the semantic occupancy grid prediction results of the consecutive image frames to be processed based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with associated relationships includes: For each semantic occupancy grid prediction result, determining a historical semantic occupancy grid prediction result of the semantic occupancy grid prediction result; Determine first parameter information of the semantic occupancy grid prediction result and second parameter information of the historical semantic occupancy grid prediction result; Determining a difference value between the semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result based on the first parameter information and the second parameter information; The Hungarian algorithm processes each of the difference values to obtain first semantic occupancy grid prediction results of multiple time frames with associated relationships.
[0010] In an optional implementation, the step of determining the stability parameters of the modified first semantic occupancy grid prediction result includes: For each of the revised first semantic occupancy grid prediction results, determining a historical first semantic occupancy grid prediction result corresponding to the revised first semantic occupancy grid prediction result; Determining a confidence stability value, a position stability value, a size stability value, and an orientation stability value of the first revised semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result; Based on the confidence stability value, the position stability value, the size stability value, and the orientation stability value, stability parameters of the modified first semantic occupancy grid prediction result are determined.
[0011] In an optional implementation, the confidence stability value satisfies the following formula: ; in, is the confidence stability value, is the first confidence level, is the second confidence level, and are the 99% and 1% percentiles of confidence; The position stability value satisfies the following formula: ; in, is the position stability value, is the first location information, is the second location information; The dimensional stability value satisfies the following formula: ; in, is the dimensional stability value, is the first size information, is the second size information; The orientation stability value satisfies the following formula: ; in, Towards a stable value, is the first direction information, is the second orientation information, is the orientation angle corresponding to the first orientation information, is the orientation angle corresponding to the second orientation information.
[0012] In an optional implementation, when the stability parameters of the modified first semantic occupancy grid prediction result do not meet the output condition, the step of adjusting the parameters of the 3D occupancy grid prediction model parameters comprises: When the stability parameter of the modified first semantic occupancy grid prediction result does not meet the output condition, calculating the center position loss based on the first position information and the second position information, wherein the first position information and the second position information indicate the center position; Determine the total number of time frames of the continuous image frames to be processed; Calculating a position offset loss based on the first position information, the second position information and the total number of time frames; Calculating a size loss based on the first size information and the second size information; Calculating the orientation loss based on the first orientation information and the second orientation information; Based on the center position loss, position offset loss, size loss and orientation loss, parameters of the 3D occupancy grid prediction model parameters are adjusted.
[0013] In a second aspect, an embodiment of the present application provides a semantic occupancy grid prediction device, the device comprising: A training module is used to train the deraining model to obtain a trained deraining model; An acquisition module is used to acquire a 3D occupancy grid prediction model based on a trained rain removal model; A processing module is used to input the continuous image frames to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction results of the continuous image frames to be processed; match the semantic occupancy grid prediction results of the continuous image frames to be processed based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with associated relationships; correct the first semantic occupancy grid prediction results based on the Kalman filter, and determine the stability parameters of the corrected first semantic occupancy grid prediction results based on the verification model; when the stability parameters meet the output conditions, the corrected first semantic occupancy grid prediction results are used as the output of the continuous image frames to be processed; when the stability parameters do not meet the output conditions, the parameters of the 3D occupancy grid prediction model are adjusted; based on the adjusted 3D occupancy grid prediction model, return to execute the continuous image frames to be processed input into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction results of the continuous image frames to be processed until the stability parameters meet the output conditions.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the semantic occupancy grid prediction method when executing the computer program.
[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the semantic occupancy grid prediction method when executed by a processor.
[0016] This application has the following beneficial effects: The present application obtains a trained rain removal model by training a rain removal model, obtains a 3D occupancy grid prediction model based on the trained rain removal model, inputs the continuous image frames to be processed into the 3D occupancy grid prediction model, obtains semantic occupancy grid prediction results of the continuous image frames to be processed, matches the semantic occupancy grid prediction results of the continuous image frames to be processed based on the Hungarian algorithm to obtain first semantic occupancy grid prediction results of multiple time frames with associated relationships, corrects the first semantic occupancy grid prediction results based on the Kalman filter, and determines the stability parameters of the corrected first semantic occupancy grid prediction results based on the test model. When the stability parameters meet the output conditions, the corrected first semantic occupancy grid prediction results are used as the output of the continuous image frames to be processed. When the stability parameters do not meet the output conditions, the parameters of the 3D occupancy grid prediction model are adjusted. Based on the adjusted 3D occupancy grid prediction model, the execution is returned to input the continuous image frames to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction results of the continuous image frames to be processed until the stability parameters meet the output conditions. It can dynamically adjust the parameters of the 3D occupancy grid prediction model to ensure continuous optimization of the prediction results, thereby improving the output occupancy grid prediction results, and thus effectively improving the timeliness and stability of the autonomous driving perception module. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 A block diagram of an electronic device provided by an embodiment of the present invention; Figure 2 One of the flow charts of a semantic occupancy grid prediction method provided by an embodiment of the present invention; Figure 3 A second flow chart of a semantic occupancy grid prediction method provided by an embodiment of the present invention; Figure 4 A third flow chart of a semantic occupancy grid prediction method provided by an embodiment of the present invention; Figure 5 A fourth flow chart of a semantic occupancy grid prediction method provided by an embodiment of the present invention; Figure 6 A fifth flow chart of a semantic occupancy grid prediction method provided by an embodiment of the present invention; Figure 7A sixth flow chart of a semantic occupancy grid prediction method provided by an embodiment of the present invention; Figure 8 A structural block diagram of a semantic occupancy grid prediction device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0020] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0022] In the description of the present invention, it should be noted that if the terms "upper", "lower", "inside", "outside", etc. appear to indicate an orientation or position relationship, they are based on the orientation or position relationship shown in the accompanying drawings, or are the orientation or position relationship in which the product of the invention is usually placed when used. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0023] In addition, the terms “first”, “second”, etc., if used, are merely used to distinguish between the descriptions and should not be understood as indicating or implying relative importance.
[0024] In the description of this application, it should also be noted that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0025] After extensive research, it was found that the existing efficient occupancy grid modeling and stability optimization in driving environments are not effective.
[0026] In view of the discovery of the above problems, the present embodiment provides a semantic occupancy grid prediction method, an electronic device and a storage medium, which can obtain a trained rain removal model by training a rain removal model, obtain a 3D occupancy grid prediction model based on the trained rain removal model, input the continuous image frames to be processed into the 3D occupancy grid prediction model, obtain the semantic occupancy grid prediction results of the continuous image frames to be processed, match the semantic occupancy grid prediction results of the continuous image frames to be processed based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with an associated relationship, and match the semantic occupancy grid prediction results of the continuous image frames to be processed based on the Kalman filter. The first semantic occupancy grid prediction result is corrected, and the stability parameters of the corrected first semantic occupancy grid prediction result are determined based on the test model. When the stability parameters meet the output conditions, the corrected first semantic occupancy grid prediction result is used as the output of the continuous image frame to be processed. When the stability parameters do not meet the output conditions, the 3D occupancy grid prediction model parameters are adjusted. Based on the adjusted 3D occupancy grid prediction model, the execution is returned to input the continuous image frame to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction results of the continuous image frame to be processed until the stability parameters meet the output conditions. The 3D occupancy grid prediction model parameters can be dynamically adjusted to ensure the continuous optimization of the prediction results, thereby improving the output occupancy grid prediction results, thereby effectively improving the timeliness and stability of the autonomous driving perception module. The scheme provided in this embodiment is described in detail below.
[0027] This embodiment provides an electronic device that can predict a semantic occupancy grid. In a possible implementation, the electronic device can be a user terminal, for example, the electronic device can be, but is not limited to, a server, a smart phone, a personal computer (PC), a tablet computer, a personal digital assistant (PDA), a mobile Internet device (MID), etc.
[0028] Please refer to Figure 1 , Figure 1 1 is a schematic diagram of the structure of the electronic device 100 provided in the embodiment of the present application. The electronic device 100 may also include Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0029] The electronic device 100 includes a semantic occupancy grid prediction device 110 , a memory 120 , and a processor 130 .
[0030] The components of the memory 120 and the processor 130 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The semantic occupancy grid prediction device 110 includes at least one software function module that can be stored in the memory 120 in the form of software or firmware or fixed in the operating system (OS) of the electronic device 100. The processor 130 is used to execute the executable modules stored in the memory 120, such as the software function modules and computer programs included in the semantic occupancy grid prediction device 110.
[0031] The memory 120 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc. The memory 120 is used to store a program, and the processor 130 executes the program after receiving an execution instruction.
[0032] Please refer to Figure 2 , Figure 2 For application Figure 1 A flowchart of a semantic occupancy grid prediction method for an electronic device 100 is provided, and the method including each step is described in detail below.
[0033] S201: Train the rain removal model to obtain a trained rain removal model.
[0034] S202: Acquire a 3D occupancy grid prediction model based on the trained rain removal model.
[0035] S203: Inputting the continuous image frames to be processed into a 3D occupancy grid prediction model to obtain semantic occupancy grid prediction results of the continuous image frames to be processed.
[0036] S204: Matching the semantic occupancy grid prediction results of the continuous image frames to be processed based on the Hungarian algorithm to obtain first semantic occupancy grid prediction results of multiple time frames with associated relationships.
[0037] S205: Correcting the first semantic occupancy grid prediction result based on the Kalman filter, and determining a stable parameter of the corrected first semantic occupancy grid prediction result based on the verification model.
[0038] S206: When the stability parameter satisfies the output condition, the modified first semantic occupancy grid prediction result is used as the output of the continuous image frame to be processed.
[0039] S207: When the stability parameters do not meet the output conditions, the 3D occupancy grid prediction model parameters are adjusted.
[0040] S208: Based on the adjusted 3D occupancy grid prediction model, return to execute inputting the continuous image frames to be processed into the 3D occupancy grid prediction model to obtain semantic occupancy grid prediction results of the continuous image frames to be processed until the stable parameters meet the output conditions.
[0041] The main function of the rain removal model is to remove interference such as raindrops and rain streaks in images or videos and restore a clear background image. The rain removal model can help the vehicle's visual system more accurately identify roads, pedestrians and obstacles on rainy days.
[0042] There are many ways to train a rain removal model. In one implementation, a large number of paired rainy images and clean images are required as training data. The images are cropped, scaled, normalized, and other operations are performed. Data diversity can also be increased through data augmentation techniques (such as rotation, flipping, adding noise, etc.). Convolutional neural networks (CNNs) or generative adversarial networks (GANs) are commonly used. For example, PReNet is an efficient recursive network that can further improve the rain removal effect by introducing an attention mechanism. A multi-layer convolutional network is designed to extract image features, and the model is optimized through a specific loss function (such as L1 loss, SSIM loss). L1 loss, SSIM loss, or a combination of the two are usually used to measure the rain removal effect. The Adam optimizer is often used for parameter updates. The training rounds, learning rate adjustment strategy, and validation set are set to monitor model performance to complete the training of the rain removal model.
[0043] A 3D occupancy grid prediction model is constructed based on the trained derained image. The pre-trained derained model is used to preprocess the multi-view images, extract the high-dimensional features of the images, and generate a three-dimensional occupancy grid descriptor through cascade voxel decoding to accurately characterize the objects and scene details in the three-dimensional space. The semantic occupancy grid prediction results of the continuous image frames to be processed are matched using the Hungarian algorithm, and the first semantic occupancy grid prediction results are further corrected by applying smoothing algorithms such as Kalman filtering. The stability index considering confidence, position, size and orientation is calculated to quantify the stability of the semantic occupancy grid. The stability parameters of the corrected first semantic occupancy grid prediction results are determined. According to the feedback results of the stability parameters, an optimization strategy is selected, the 3D occupancy grid prediction model is iteratively trained, and the stability index is re-evaluated to ensure the high stability and accuracy of the occupancy grid prediction results output by the 3D occupancy grid prediction model in the spatiotemporal domain.
[0044] Another implementation method of training the deraining model to obtain the trained deraining model is as follows: Figure 3 As shown, the following steps are included: S301: Acquire rainy day images that meet preset conditions from a driving data set.
[0045] S302: Generate a semantic segmentation image for the rainy day image based on a generative algorithm.
[0046] S303: Optimizing the semantic segmentation image based on the optimal transmission strategy.
[0047] S304: training a rain removal model based on the optimized semantic segmentation image to obtain a trained rain removal model.
[0048] In the driving data set, we can select sunny and rainy images and remove invalid or unlabeled images to ensure the reliability of the data set and provide guarantee for subsequent model training. In order to further improve the generalization ability of the deraining model, we conduct a secondary screening of rainy images and select more challenging samples to help the model better cope with complex rainy scenes in the real world. The images after the secondary screening are used as rainy images that meet the preset conditions. Among them, the images that meet the preset conditions cover a variety of characteristics such as dense raindrops, diffuse fog, and complex light reflections, ensuring their representativeness and high difficulty, which helps the algorithm learn more effective features during the training process, thereby improving its performance in practical applications.
[0049] The rainy images are divided into training sets and test sets to facilitate the evaluation of the model's performance. Based on the generative algorithm, semantic segmentation images are generated for rainy images, and semantic segmentation labels are generated. The semantic segmentation results are saved for direct use in subsequent training processes, which can greatly accelerate the training process of the deraining model and provide strong support for subsequent image processing tasks.
[0050] The semantic segmentation images are optimized based on the optimal transmission strategy, thereby optimizing the weight distribution of image features and significantly improving the effect of contrastive learning.
[0051] There are many ways to implement a 3D occupancy grid prediction model based on a trained rain removal model. In one implementation, Figure 4 As shown, the following steps are included: S401: Input a rainy day image of a real scene into a trained rain removal model to obtain multiple rain removal images under multiple viewing angles.
[0052] S402: calibrating and time-synchronizing the multiple derained images respectively to obtain multiple first images at multiple viewing angles that are consistent in time.
[0053] S403: For each first image, based on the semantic features of the first image, the high-dimensional feature representation tensor of the first image, and the vehicle geometry information, a BEV feature of the first image is obtained.
[0054] S404: For each BEV feature, determine a current BEV feature and a historical BEV feature corresponding to the BEV feature.
[0055] S405: Generate a three-dimensional occupancy descriptor based on the current BEV features and the historical BEV features to complete the construction of the 3D occupancy grid prediction model.
[0056] The specific method of calibrating and time-synchronizing multiple de-rained images to obtain multiple first images at multiple perspectives with consistent time is as follows: based on the internal and external parameters of the camera, calibrate multiple de-rained images at multiple perspectives to obtain multiple calibrated de-rained images at multiple perspectives. Time-synchronize the multiple de-rained images at the calibrated multiple perspectives to obtain multiple first images at multiple perspectives with consistent time.
[0057] For each first image, based on the semantic features of the first image and the high-dimensional feature representation tensor of the first image and the vehicle geometry information, the specific method of obtaining the BEV features of the first image is as follows: for each first image, the low-order features and semantic features of the first image are obtained, and based on the low-order features and the semantic features, the high-dimensional feature representation tensor of the first image is obtained; for multiple first images, the BEV features are obtained based on the high-dimensional feature representation tensor corresponding to each first image and the vehicle geometry information.
[0058] For each BEV feature, determine the current BEV feature and the historical BEV feature corresponding to the BEV feature, and generate a three-dimensional occupancy descriptor based on the current BEV feature and the historical BEV feature to complete the construction of the 3D occupancy grid prediction model. The specific method is as follows: for each BEV feature, use the BEV feature as the current BEV feature, and obtain the historical BEV feature corresponding to the BEV feature. The current BEV feature and the historical BEV feature corresponding to the current BEV feature are spatially transformed, unified to the same coordinate system, and feature aligned in sequence to obtain the aligned current BEV feature and the aligned historical BEV feature. Based on the aligned current BEV feature and the corresponding historical BEV feature, a temporal change feature is obtained. Based on the temporal change feature, the aligned current BEV feature is supplemented or deleted to obtain a new current BEV feature. Based on the new current BEV feature and the aligned historical features, the semantic grid motion trend and scene change are determined. Based on the multi-head attention mechanism, 3D cross-attention mechanism, new current BEV features, aligned historical features, semantic occupancy grid motion trends and scene changes, the final BEV features after the fusion time sequence of the new current BEV features are obtained. Based on the final BEV features, a three-dimensional occupancy descriptor is generated to complete the construction of the 3D occupancy grid prediction model.
[0059] Among them, the three-dimensional occupancy descriptor is used to describe the prediction results of the semantic occupancy grid.
[0060] The rainy day images obtained from real scenes are used as the input of the trained rain removal model. The rainy day images obtained from real scenes are multi-perspective rainy day images obtained by vehicles. The multi-perspective rainy day images include multiple rainy day images from different perspectives. The trained rain removal model outputs multi-perspective rain removal images of the rainy day images from real scenes, that is, multiple rain removal images from different perspectives.
[0061] Through image preprocessing, the image of each perspective is corrected and time-synchronized, and the geometric calibration of multi-perspective images is completed according to the internal and external parameters of the camera to obtain complete surrounding environment information.
[0062] Using camera intrinsics, such as focal length ( fx , fy )、Main point( xx , cy ), distortion coefficient ( k 1, k 2, p 1, p 2) Perform radial and tangential distortion correction on the image of each perspective to obtain a calibrated multi-perspective derained image.
[0063] The method of performing time synchronization on the calibrated multi-view de-rained images can be: based on the image acquisition timestamp, select the frame closest in time or interpolate, estimate the inter-frame motion and adjust the time offset in post-processing, and finally obtain the multi-view first images with consistent time, that is, multiple first images corresponding to different viewpoints.
[0064] For each first image, low-order features and semantic features of the first image are obtained, and based on the low-order features and semantic features, a high-dimensional feature representation tensor of the first image is obtained, wherein the semantic features indicate labels of each object in the first image.
[0065] For example, a convolutional neural network (CNN) is used to extract low-level features (such as edges, textures, colors, etc.). Multi-layer convolution (such as ResNet, VGG, etc.) is used to gradually extract higher-level features. The low-level features and semantic features are globally modeled through a multi-head attention mechanism, and finally a high-dimensional feature representation tensor corresponding to the first image of each perspective is obtained.
[0066] Exemplarily, the first image includes a first image corresponding to a first perspective, a first image corresponding to a second perspective, and a first image corresponding to a third perspective, and a first high-dimensional feature representation of the first image corresponding to the first perspective, a second high-dimensional feature representation of the first image corresponding to the second perspective, and a third high-dimensional feature representation of the first image corresponding to the third perspective are respectively determined to determine the vehicle geometry information, that is, a three-dimensional coordinate system is constructed with the vehicle as the origin, and each high-order feature representation within the same time is put into the three-dimensional coordinate system, that is, the first high-dimensional feature representation, the second high-dimensional feature representation and the third high-dimensional feature representation are put into the three-dimensional coordinate system to obtain BEV features, that is, to achieve fusion of images from multiple perspectives to obtain fused BEV features, and the fused BEV features carry three-dimensional spatial features of vehicle geometry information and semantic features.
[0067] When processing each BEV feature, the currently processed BEV feature is used as the current BEV feature, and the historical BEV feature of the current BEV feature is obtained.
[0068] When the rainy day image of the real scene contains 3s of image frames and 3 groups of images at multiple perspectives at different times, when the images at multiple perspectives of 2s are processed, the current BEV feature is obtained, and the BEV feature corresponding to the 1s is the historical BEV feature corresponding to the current BEV feature. When the images at multiple perspectives of the 3rd second are processed, the current BEV feature is obtained, and the BEV feature corresponding to the 1s and the BEV feature corresponding to the 2s are the historical BEV features corresponding to the current BEV feature.
[0069] The current BEV features and the historical BEV features are sequentially spatially transformed, unified to the same coordinate system, and feature aligned to obtain the aligned current BEV features and the aligned historical BEV features.
[0070] Time alignment technology can be used to perform spatial transformation on historical features, and the historical BEV features and current BEV features can be unified into the same coordinate system. The current BEV features and historical BEV features can be aligned through a three-dimensional deformable attention mechanism to obtain aligned current BEV features and aligned historical BEV features.
[0071] Based on the aligned current BEV features and the corresponding historical BEV features, the time series variation features are obtained to calibrate the current BEV features.
[0072] The method for determining the time series change feature is to calculate the difference between the current BEV feature and each historical BEV feature, use 3D convolution to perform time series modeling on the historical BEV features, use RNN neural network to model the historical BEV features, use Transformer encoder to perform global time series modeling on the historical BEV feature sequence, splice the current BEV feature and the time series change feature along the channel dimension, use the attention mechanism to dynamically fuse the current BEV feature and the time series change feature, use residual connection to retain the information of the current BEV feature, and finally obtain the time series change feature, determine to supplement or delete the current BEV feature based on the time series change feature, and supplement or delete the current BEV feature based on the time series change feature, so as to realize the calibration of the current BEV feature based on the historical BEV feature and obtain the new current BEV feature.
[0073] The new current BEV features and the aligned BEV features are input into the recurrent neural network, which captures the semantic grid motion trends and scene changes of objects in consecutive frames through iterative calculation of time steps.
[0074] Recurrent neural networks can combine the features of historical frames to learn the movement trajectory of objects and the dynamic changes of scenes. Since recurrent neural networks can retain long-term dependency information when processing time series, combined with more complex structures such as LSTM or GRU, they can predict the movement trend of objects and ensure the dynamic consistency and stability of the semantic occupancy grid between different time frames.
[0075] The multi-head attention mechanism is to model the global relationship of the new current BEV features. The 3D cross attention mechanism is to spatially and temporally fuse the aligned historical BEV features with the new current BEV features. Considering the motion trend of the semantically occupied grid, the motion trend can be fused with the spatial and temporal features. Considering the scene changes, the scene change information and the features fused with the motion trend are further fused to generate high-quality temporal fusion BEV features, providing strong support for tasks such as target detection and behavior prediction.
[0076] After feature processing, the self-attention mechanism is used to help capture the global dependencies between time frames, especially between different time steps, so as to focus on key information areas. The 3D cross-attention mechanism is used to align the BEV features in different time frames and perspectives, further improving the model's ability to model object motion trends and scene changes. Through this precise alignment and fusion, the model can reduce prediction fluctuations caused by environmental changes or movement, so that the motion trends and scene changes of the semantic occupancy grid can be captured more stably and accurately.
[0077] The final BEV features after the above fusion are used to generate a three-dimensional occupancy descriptor through the decoder. The three-dimensional occupancy descriptor is a feature representation method used to represent the occupancy state of objects or spaces in a three-dimensional scene. It provides a fine-grained description of the scene by quantizing the physical 3D scene into a structured grid map with each cell carrying a semantic label or occupancy state.
[0078] There are many ways to achieve the matching of prediction boxes in multiple time frames based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with associated relationships. In one implementation, Figure 5 As shown, the following steps are included: S501: For each semantic occupancy grid prediction result, determine a historical semantic occupancy grid prediction result of the semantic occupancy grid prediction result.
[0079] S502: Determine first parameter information of a semantic occupancy grid prediction result and second parameter information of a historical semantic occupancy grid prediction result.
[0080] S503: Determine a difference value between a semantic occupancy grid prediction result and a historical semantic occupancy grid prediction result based on the first parameter information and the second parameter information.
[0081] S504: The Hungarian algorithm processes each difference value to obtain a first semantic occupancy grid prediction result of multiple time frames with associated relationships.
[0082] The image data to be processed is input into the constructed 3D occupancy grid prediction model, and the semantic occupancy grid prediction results of the continuous image frames to be processed are obtained. Based on the Hungarian algorithm, the semantic occupancy grids at different timestamps are associated to obtain the first semantic occupancy grid prediction results of multiple time frames with associated relationships.
[0083] Based on the constructed 3D occupancy grid prediction model, the position, size, category and other information of the object are extracted from each frame to construct the current semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result. By calculating the first parameter information of the current semantic occupancy grid prediction result and the second parameter information of the historical semantic occupancy grid prediction result, the difference between the semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result is obtained, and the difference is used to evaluate the matching cost. The Hungarian algorithm is used to minimize these costs and find the optimal match between each pair of semantic occupancy grid prediction results, so as to establish the association between objects in multiple time frames. Finally, the matching results of each time frame are output to form the time series trajectory of the object, that is, the first semantic occupancy grid prediction results of multiple time frames with association relationships. This not only ensures the consistency of association between the same object in different time frames, but also effectively reduces the prediction fluctuations caused by motion or environmental changes, and improves the stability and accuracy of object tracking.
[0084] Based on the first parameter information and the second parameter information, the difference value between the semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result may be determined by: The first parameter information includes the position of the semantic grid prediction result (x 1 , y 1 , z 1 )、Size(l 1 , w 1 ,h 1 ), direction (θ 1 ), the second parameter information includes the position of the historical semantics occupying the grid prediction result (x 2 , y 2 , z 2 )、Size(l 2 , w 2 , h 2 ), direction (θ 2 ).
[0085] The calculation formula for the difference between the semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result is as follows: Cost(B 1 ,B 2 )=||x1-x2||+||y 1 -y 2 ||+||z1 -z 2 ||+||l 1 -l 2 ||+||ω 1 -ω 2 ||+||h1-h2||+||θ 1 -θ 2 ||.
[0086] Among them, B 1 represents the semantic occupancy grid prediction result, B 2 Indicates the historical semantics occupying the grid prediction result. The first parameter information includes the position of the semantics occupying the grid prediction result (x 1 , y 1 , z 1 )、Size(l 1 , w 1 , h 1 ), direction (θ 1 ), the second parameter information includes the position of the historical semantics occupying the grid prediction result (x 2 , y 2 , z 2 )、Size(l 2 , w 2 , h 2 ), direction (θ 2 ).
[0087] That is, the matching cost is evaluated by calculating the Euclidean distance between the center position of the semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result, the change in size, and the consistency of the category label. The Hungarian algorithm is used to minimize these costs and find the optimal match between each pair of prediction boxes, thereby establishing associations between objects in multiple time frames.
[0088] The Hungarian algorithm performs global optimization matching on the difference values of the semantic occupancy grid prediction results to ensure the correlation consistency of the same semantic occupancy grid prediction results across multiple time frames and minimize the prediction fluctuations caused by movement or environmental changes. After the matching is completed, the time correlation result of each semantic occupancy grid is output, that is, the first semantic occupancy grid prediction result, for subsequent correction and stability optimization.
[0089] The prediction result of the first semantic occupancy grid is corrected by the Kalman filter smoothing algorithm to improve the stability and consistency of the semantic occupancy grid detection in the time series.
[0090] The correlation between multiple time frames is established by extracting the geometric characteristics of the first semantic occupancy grid prediction results, including information such as center position, size and orientation.
[0091] The geometric characteristics of the first semantic occupancy grid prediction result include low-order geometric characteristics and high-order dynamic information. The low-order geometric characteristics, such as position coordinates and size changes, can be obtained through simple geometric algebra calculations, and the high-order dynamic information, such as orientation angle changes, is further quantified in combination with the motion trajectory of the semantic occupancy grid. Subsequently, based on the associated semantic occupancy grid trajectory, the Kalman filter is applied to smoothly predict the position of the semantic occupancy grid.
[0092] The application of Kalman filtering in the position smoothing prediction of the first semantic occupancy grid prediction result realizes the position smoothing processing by jointly modeling the historical trajectory of the first semantic occupancy grid prediction result and the first semantic occupancy grid prediction result of the current observation.
[0093] In view of the size change, the average size of the first semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result of the first semantic occupancy grid prediction result is used to perform scale constraint to ensure that the size of the first semantic occupancy grid prediction result remains consistent in time series.
[0094] The size of each frame is compared with the size of the previous and next frames, and the size of the current frame is corrected by calculating the mean size of the historical frames. This can effectively reduce the size fluctuation caused by detection errors or short-term environmental changes. In order to further optimize the smooth transition of the size, size consistency regularization is introduced to limit the drastic fluctuation of the size in the time series and ensure that the size change of the first semantic occupancy grid prediction result is within a reasonable range. In addition, the IoU method is used to optimize the size consistency. By calculating the overlap between the first semantic occupancy grid prediction result of the current frame and the first semantic occupancy grid prediction result of the previous and next frames, the size of the first semantic occupancy grid prediction result of each frame is ensured to be consistent between adjacent frames, reducing the deviation of size prediction. The area with a high IoU value indicates that the size change is small, which meets the stability requirements of the object in the time dimension. In this way, the smooth correction and consistency optimization of the size are guaranteed, thereby improving the stability and continuity of the semantic occupancy grid in the time dimension.
[0095] For the orientation information of the first semantic occupancy grid prediction result, the orientation angle of the current first semantic occupancy grid prediction result is smoothly corrected in combination with the historical first semantic occupancy grid prediction results of the first semantic occupancy grid prediction result. If the orientation change between adjacent frames exceeds the threshold, it is limited to a reasonable range to avoid the impact of drastic angle jumps.
[0096] Through the above-mentioned geometric correction and time series smoothing methods, the corrected first semantic occupancy grid prediction results are aligned to a unified stable state.
[0097] There are many implementations for determining the stability parameters of the modified first semantic occupancy grid prediction result. In one implementation, for example, Figure 6 As shown, the following steps are included: S601: For each corrected first semantic occupancy grid prediction result, determine a historical first semantic occupancy grid prediction result corresponding to the corrected first semantic occupancy grid prediction result.
[0098] S602: Determine the confidence stability value, position stability value, size stability value, and orientation stability value of the revised first semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result.
[0099] S603: Determine stability parameters of the modified first semantic occupancy grid prediction result based on the confidence stability value, the position stability value, the size stability value, and the orientation stability value.
[0100] Specifically: determine a first confidence of a revised first semantic occupancy grid prediction result, determine a second confidence of a historical semantic occupancy grid prediction result, determine a confidence stability value of the revised first semantic occupancy grid prediction result based on the first confidence and the second confidence, determine first position information of the revised first semantic occupancy grid prediction result and second position information of the historical first semantic occupancy grid prediction result, and determine a position stability value of the revised first semantic occupancy grid prediction result based on the first position information and the second position information.
[0101] The position stability is measured by calculating the intersection over union (IoU) of the position information (center points) of the two corrected first semantic occupancy grid prediction results. The position corrected semantic occupancy grid is to reduce the offset of the center point.
[0102] Determine first size information of a revised first semantic occupancy grid prediction result and second size information of a historical first semantic occupancy grid prediction result, determine a size stability value of the revised first semantic occupancy grid prediction result based on the first size information and the second size information, determine first orientation information of the revised first semantic occupancy grid prediction result and second orientation information of the historical first semantic occupancy grid prediction result, determine an orientation stability value of the revised first semantic occupancy grid prediction result based on the first orientation information and the second orientation information, and determine stability parameters of the revised first semantic occupancy grid prediction result based on the confidence stability value, the position stability value, the size stability value and the orientation stability value.
[0103] The confidence stability value satisfies the following formula: ; in, is the confidence stability value, is the first confidence level, is the second confidence level, and are the 99% and 1% percentiles of confidence; The position stability value satisfies the following formula: ; in, is the position stability value, is the first location information, is the second location information; The dimensional stability value satisfies the following formula: ; in, is the dimensional stability value, is the first size information, is the second size information; The orientation stability value satisfies the following formula: ; in, Towards a stable value, is the first direction information, is the second orientation information, is the orientation angle corresponding to the first orientation information, is the orientation angle corresponding to the second orientation information.
[0104] The orientation information indicates coordinate information of the front and rear ends of the vehicle, and the orientation angle is calculated based on the orientation information.
[0105] The final stable parameters are: .
[0106] When the stable parameters of the modified first semantic occupancy grid prediction result do not meet the output condition, there are multiple implementations of adjusting the parameters of the 3D occupancy grid prediction model parameters. In one implementation, Figure 7 As shown, the following steps are included: S701: Determine the total number of time frames of continuous image frames to be processed.
[0107] S702: Calculate the position offset loss based on the first position information, the second position information and the total number of time frames.
[0108] S703: Calculate the size loss based on the first size information and the second size information.
[0109] S704: Calculate the orientation loss based on the first orientation information and the second orientation information.
[0110] S705: Adjust the parameters of the 3D occupancy grid prediction model parameters based on the center position loss, the position offset loss, the size loss, and the orientation loss.
[0111] The center position loss satisfies the following formula: ; is the center position loss, The first semantics after correction occupies the center position of the first position information of the grid prediction result; is the observation matrix, usually the identity matrix in one or two dimensions, is the Kalman gain.
[0112] The position offset loss satisfies the following formula: ; T is the total number of time frames, is the position offset loss, is the first location information, It is the second location information.
[0113] The size loss satisfies the following formula: ; ; and are the width and height of each frame respectively.
[0114] The orientation loss satisfies the following formula: ; when When it is greater than the preset threshold, the angle is corrected to keep it changing smoothly.
[0115] ; is the first direction information, It is the second direction information.
[0116] Based on the center position loss, position offset loss, size loss and orientation loss, the parameters of the 3D occupancy grid prediction model parameters are adjusted to obtain the adjusted 3D occupancy grid prediction model.
[0117] The continuous image frames to be processed are input into the 3D occupancy grid prediction model after parameter adjustment to obtain the semantic occupancy grid prediction results of the continuous image frames to be processed; the semantic occupancy grid prediction results of the continuous image frames to be processed are matched based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with associated relationships; the first semantic occupancy grid prediction results are corrected based on the Kalman filter, and the stability parameters of the corrected first semantic occupancy grid prediction results are determined based on the verification model; when the stability parameters meet the output conditions, the corrected first semantic occupancy grid prediction results are used as the output of the continuous image frames to be processed; when the stability parameters meet the output conditions, the corrected first semantic occupancy grid prediction results are output.
[0118] Based on the stability parameter as the termination condition of model optimization, the spatiotemporal stability performance of the 3D occupancy grid prediction model in dynamic scenes is fully quantified. Specifically, for the stability sub-indicators of confidence, position, size and orientation, the change range before and after optimization is calculated respectively, and the model performance is comprehensively evaluated to see whether it meets the preset requirements, that is, whether the stability parameter meets the requirement of SI>0.8, where SI is the stability parameter.
[0119] When the stability parameter is greater than 0.8, it is determined that the stability parameter meets the output condition.
[0120] Please refer to Figure 8 The present application embodiment also provides a method for applying Figure 1 The semantic occupancy grid prediction device 110 of the electronic device 100 includes: A training module 111 is used to train the rain removal model to obtain a trained rain removal model; An acquisition module 112 is used to construct a 3D occupancy grid prediction model based on the trained rain removal model; The processing module 113 is used to input the continuous image frames to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction results of the continuous image frames to be processed; match the semantic occupancy grid prediction results of the continuous image frames to be processed based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with associated relationships; correct the first semantic occupancy grid prediction results based on the Kalman filter, and determine the stability parameters of the corrected first semantic occupancy grid prediction results based on the verification model; when the stability parameters meet the output conditions, the corrected first semantic occupancy grid prediction results are used as the output of the continuous image frames to be processed; when the stability parameters do not meet the output conditions, the parameters of the 3D occupancy grid prediction model are adjusted; based on the adjusted 3D occupancy grid prediction model, return to execute the input of the continuous image frames to be processed into the 3D occupancy grid prediction model to obtain the semantic occupancy grid prediction results of the continuous image frames to be processed until the stability parameters meet the output conditions.
[0121] The present application also provides an electronic device 100, which includes a processor 130 and a memory 120. The memory 120 stores computer executable instructions, and when the computer executable instructions are executed by the processor 130, the semantic occupancy grid prediction method is implemented.
[0122] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by the processor 130, the semantic occupancy grid prediction method is implemented.
[0123] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or the flowchart, and the combination of boxes in the block diagram and / or the flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0124] In addition, each functional module in each embodiment of the present application can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk.
[0125] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0126] The above are only various implementations of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A semantic occupancy grid prediction method, characterized in that: The method comprises: Train the rain removal model to obtain a trained rain removal model; Build a 3D occupancy grid prediction model based on the trained rain removal model; Inputting the continuous image frames to be processed into the 3D occupancy grid prediction model to obtain semantic occupancy grid prediction results of the continuous image frames to be processed; Based on the Hungarian algorithm, the semantic occupancy grid prediction results of the continuous image frames to be processed are matched to obtain the first semantic occupancy grid prediction results of multiple time frames with associated relationships; Correcting the first semantic occupancy grid prediction result based on Kalman filtering, and determining a stability parameter of the corrected first semantic occupancy grid prediction result based on a verification model; When the stability parameter satisfies the output condition, the modified first semantic occupancy grid prediction result is used as the output of the continuous image frame to be processed; When the stability parameters do not meet the output conditions, adjusting the 3D occupancy grid prediction model parameters; Based on the adjusted 3D occupancy grid prediction model, return to execute inputting the continuous image frames to be processed into the 3D occupancy grid prediction model to obtain semantic occupancy grid prediction results of the continuous image frames to be processed until the stability parameters meet the output conditions.
2. The method according to claim 1, characterized in that The step of training the rain removal model to obtain a trained rain removal model includes: Obtain rainy day images that meet preset conditions from the driving dataset; Generating a semantic segmentation image for the rainy day image based on a generative algorithm; Optimizing the semantic segmentation image based on an optimal transmission strategy; The rain removal model is trained based on the optimized semantic segmentation image to obtain a trained rain removal model.
3. The method according to claim 1, characterized in that The step of constructing a 3D occupancy grid prediction model based on the trained rain removal model includes: Inputting a rainy day image of a real scene into the trained deraining model to obtain multiple derained images under multiple viewing angles; Calibrate and time-synchronize the plurality of rain-removed images respectively to obtain a plurality of first images at a plurality of viewing angles with consistent time; For each first image, based on the semantic features of the first image, the high-dimensional feature representation tensor of the first image, and the vehicle geometry information, obtain a BEV feature of the first image; For each BEV feature, determining a current BEV feature and a historical BEV feature corresponding to the BEV feature; A three-dimensional occupancy descriptor is generated based on the current BEV features and the historical BEV features to complete the construction of a 3D occupancy grid prediction model.
4. The method according to claim 1, characterized in that: The step of matching the semantic occupancy grid prediction results of the continuous image frames to be processed based on the Hungarian algorithm to obtain the first semantic occupancy grid prediction results of multiple time frames with associated relationships includes: For each semantic occupancy grid prediction result, determining a historical semantic occupancy grid prediction result of the semantic occupancy grid prediction result; Determine first parameter information of the semantic occupancy grid prediction result and second parameter information of the historical semantic occupancy grid prediction result; Determining a difference value between the semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result based on the first parameter information and the second parameter information; The Hungarian algorithm processes each of the difference values to obtain first semantic occupancy grid prediction results of multiple time frames with associated relationships.
5. The method according to claim 1, characterized in that The step of determining the stability parameters of the modified first semantic occupancy grid prediction result comprises: For each of the revised first semantic occupancy grid prediction results, determining a historical first semantic occupancy grid prediction result corresponding to the revised first semantic occupancy grid prediction result; Determining a confidence stability value, a position stability value, a size stability value, and an orientation stability value of the first revised semantic occupancy grid prediction result and the historical semantic occupancy grid prediction result; Based on the confidence stability value, the position stability value, the size stability value, and the orientation stability value, stability parameters of the modified first semantic occupancy grid prediction result are determined.
6. The method according to claim 5, characterized in that The confidence stability value satisfies the following formula: ; in, is the confidence stability value, is the first confidence level, is the second confidence level, and are the 99% and 1% percentiles of confidence; The position stability value satisfies the following formula: ; in, is the position stability value, is the first location information, is the second location information; The dimensional stability value satisfies the following formula: ; in, is the dimensional stability value, is the first size information, is the second size information; The orientation stability value satisfies the following formula: ; in, Towards a stable value, is the first direction information, is the second orientation information, is the orientation angle corresponding to the first orientation information, is the orientation angle corresponding to the second orientation information.
7. The method according to claim 5, characterized in that When the stability parameters of the modified first semantic occupancy grid prediction result do not meet the output condition, the step of adjusting the parameters of the 3D occupancy grid prediction model parameters comprises: When the stability parameter of the modified first semantic occupancy grid prediction result does not meet the output condition, calculating the center position loss based on the first position information and the second position information, wherein the first position information and the second position information indicate the center position; Determine the total number of time frames of the continuous image frames to be processed; Calculating a position offset loss based on the first position information, the second position information and the total number of time frames; Calculating a size loss based on the first size information and the second size information; Calculating an orientation loss based on the first orientation information and the second orientation information; Based on the center position loss, position offset loss, size loss and orientation loss, parameters of the 3D occupancy grid prediction model parameters are adjusted.
8. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of the method according to any one of claims 1 to 7 when executing the computer program.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Pose estimation method of robot, robot and computer storage medium
CN112363158A
Predicted trajectory correction method and device based on occupancy grid and automatic driving vehicle
CN115933654A
Obstacle trajectory prediction method and device, equipment and medium
CN116309689A
Target motion state detection method and device, mobile device and storage medium
CN116309693A
Lightweight occupancy grid prediction method and system based on large model self-labeling
CN118823139A