A method and device for generating point cloud based on millimeter wave and visual image fusion
By generating target supervised point clouds and combining millimeter wave and visual feature information, the decoder is trained to predict and generate occluded target point clouds, solving the problem of insufficient perception accuracy in traditional sensors in complex scenarios, and achieving high-precision and efficient point cloud generation.
Patent Information
- Application Number
- CN202411657436.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-11-19
AI Technical Summary
Traditional single sensors are difficult to meet the needs of high-precision perception in complex scenarios, especially in occlusion and low-light conditions that cannot accurately build point clouds.
By acquiring lidar data from historical occlusion scenes, the target supervised point cloud is generated, and combining the characteristic information of millimeter wave radar and vision cameras, the training decoder predicts the generation of occlusion target point clouds and unoccluded target point clouds.
It improves the integrity and accuracy of point clouds, enhances the system's perception ability in complex environments and harsh conditions, and ensures the efficiency and reliability of point cloud generation.
Smart Images

Figure CN119168890B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly to a method and device for generating point clouds based on the fusion of millimeter waves and visual images. Background Art
[0002] With the rapid development of technologies such as autonomous driving, intelligent robots, and augmented reality, the requirement for the accuracy of environmental perception is getting higher and higher. In these scenarios, it is crucial to accurately identify obstacles, vehicles, pedestrians, and other targets in the surrounding environment.
[0003] In practical applications, lidar and vision cameras are usually used as sensors for object detection. Lidar obtains three-dimensional point cloud data of the surrounding environment by emitting laser pulses and receiving their reflected signals, which has the advantages of high precision and high density, but is costly and easily affected by weather conditions, and cannot accurately complete the construction of point clouds when there are occlusions. Vision cameras provide rich visual information by capturing images, with low cost and certain adaptability to environmental changes, but their performance will be greatly reduced when there is insufficient light or occlusion. Therefore, traditional single sensors are difficult to meet the requirements of high-precision perception in complex scenarios and cannot meet the usage standards. Summary of the Invention
[0004] In view of this, this application provides a method and device for generating point clouds based on the fusion of millimeter waves and visual images, so as to generate high-precision point clouds in complex scenarios.
[0005] Specifically, this application is implemented through the following technical solutions:
[0006] The first aspect of this application provides a method for generating point clouds based on the fusion of millimeter waves and visual images, and the method includes:
[0007] Obtain lidar data of a historical occlusion scene, and generate target supervised point clouds based on the lidar data;
[0008] Wherein, the point cloud density of the occluded target in the lidar data is less than a first preset threshold, the point cloud density of the occluded target in the target supervised point clouds is greater than the first preset threshold, and the point cloud contour of the occluded target in the target supervised point clouds is complete;
[0009] Identify the historical occlusion scene based on a millimeter wave radar and a vision camera respectively, and obtain historical millimeter wave features and historical visual features;
[0010] Using the target supervised point clouds as decoder supervision, input the historical millimeter wave features and the historical visual features into a decoder, and train the decoder to predict and generate occluded target point clouds and unoccluded target point clouds;
[0011] Identify the occlusion scene to be recognized based on millimeter-wave radar and vision camera, obtain millimeter-wave features and vision features, and predict and generate the scene point cloud of the occlusion scene to be recognized based on the millimeter-wave features, the vision features and the trained decoder.
[0012] The second aspect of this application provides a point cloud generation device based on the fusion of millimeter-wave and vision images. The device includes an acquisition module, an identification module, a prediction module and a generation module. Among them,
[0013] The acquisition module is used to acquire lidar data of historical occlusion scenes and generate target supervised point clouds based on the lidar data.
[0014] Among them, the point cloud density of the occluding target in the lidar data is less than a first preset threshold, the point cloud density of the occluding target in the target supervised point cloud is greater than the first preset threshold, and the point cloud contour of the occluding target in the target supervised point cloud is complete.
[0015] The identification module is used to respectively identify the historical occlusion scene based on the millimeter-wave radar and the vision camera to obtain historical millimeter-wave features and historical vision features.
[0016] The prediction module is used to use the target supervised point cloud as decoder supervision, input the historical millimeter-wave features and the historical vision features into the decoder, and train the decoder to predict and generate occluding target point clouds and non-occluding target point clouds.
[0017] The generation module is used to identify the occlusion scene to be recognized based on the millimeter-wave radar and the vision camera, obtain millimeter-wave features and vision features, and predict and generate the scene point cloud of the occlusion scene to be recognized based on the millimeter-wave features, the vision features and the trained decoder.
[0018] The method and device for generating point cloud based on millimeter-wave and visual image fusion provided by this application first obtain the lidar data of the historical occlusion scene, and generate target supervised point cloud based on the lidar data; among them, the point cloud density of the occluded target in the lidar data is less than the first preset threshold, the point cloud density of the occluded target in the target supervised point cloud is greater than the first preset threshold, and the point cloud contour of the occluded target in the target supervised point cloud is complete; furthermore, the historical occlusion scene is respectively recognized based on the millimeter-wave radar and the visual camera to obtain historical millimeter-wave features and historical visual features; then, using the target supervised point cloud as the decoder supervision, the historical millimeter-wave features and the historical visual features are input into the decoder, and the decoder is trained to predict and generate occluded target point cloud and unoccluded target point cloud; finally, the millimeter-wave radar and the visual camera are used to recognize the to-be-recognized occlusion scene to obtain millimeter-wave features and visual features, and the scene point cloud of the to-be-recognized occlusion scene is predicted and generated based on the millimeter-wave features, visual features and the trained decoder. The method provided by the present invention uses the point cloud information of the lidar as the supervision data during training, and comprehensively uses the image information of the millimeter-wave radar and the visual camera to generate the complete point cloud information of the occlusion scene, improving the integrity and accuracy of the finally generated point cloud. Specifically, the lidar is good at detecting the boundaries of objects, so the contour information obtained by the lidar is used as the boundary and position hint for point cloud generation, improving the accuracy of the point cloud predicted and generated by the encoder and decoder. Further, the millimeter-wave radar can obtain clear image information under harsh conditions such as low light, improving the adaptability of the point cloud generation method and being able to complete the acquisition of unoccluded information in any case. Finally, the visual camera can quickly obtain clear image information at a low cost. The method provided by the present invention utilizes the advantages of three types of information acquisition means, providing sufficient input image information for the input of the encoder and decoder prediction, and being able to help the encoder and decoder quickly locate the relevant image regions and contour information, improving the efficiency and accuracy of point cloud generation. In this way, through the penetration ability of the millimeter-wave radar, the millimeter-wave features of the occluded part can be effectively obtained, and at the same time, through the high-level semantic information of the visual camera, the visual features containing rich context information can be effectively obtained. Further, the occluded target point cloud and unoccluded target point cloud generated by the millimeter-wave features and visual features not only contain the information of the occlusion scene, but also incorporate the semantic information of the visual image, making the generated point cloud closer to the real to-be-recognized occlusion scene. Further, the supervision of the target supervised point cloud can ensure the accuracy and robustness of the scene point cloud. In this way, by combining the millimeter-wave radar and the visual image, the information of the occluded part and other parts of the to-be-recognized occlusion scene can be fully fused, ensuring the accuracy and reliability of the scene point cloud. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flowchart of the first embodiment of the method for generating point cloud based on millimeter-wave and visual image fusion provided by this application;
[0020] Figure 2 This is a hardware structure diagram of a point cloud generation device based on millimeter - wave and visual image fusion for the point cloud generation device of the present application in a point cloud generation device based on millimeter - wave and visual image fusion;
[0021] Figure 3 This is a schematic structural diagram of Embodiment 1 of the point cloud generation device based on millimeter - wave and visual image fusion provided by the present application. Specific embodiments
[0022] Here, exemplary embodiments will be described in detail, and their examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application.
[0023] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0024] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0025] The present application provides a method and a device for generating a point cloud based on millimeter - wave and visual image fusion, so as to generate a high - precision point cloud in a complex scene.
[0026] The following specific embodiments are given to introduce the technical solutions of the present application in detail.
[0027] Figure 1 This is a flowchart of Embodiment 1 of the method for generating a point cloud based on millimeter - wave and visual image fusion provided by the present application. Please refer to Figure 1 , the method provided in this embodiment may include:
[0028] S101. Obtain lidar data of a historical occlusion scene, and generate a target supervised point cloud based on the lidar data;
[0029] Among them, the point cloud density of the occluded target in the lidar data is less than the first preset threshold, the point cloud density of the occluded target in the target supervised point cloud is greater than the first preset threshold, and the point cloud contour of the occluded target in the target supervised point cloud is complete.
[0030] Specifically, a historical occlusion scene refers to a scene in an environmental perception task where, due to the occlusion of other objects, the detector cannot directly obtain complete data. In other words, the targets in these scenes may be partially or completely invisible due to being occluded by other objects. It should be noted that in the scene of object recognition, accurately identifying and understanding the targets in these occlusion scenes is crucial for ensuring the safety and reliability of the system.
[0031] When specifically implemented, all previously detected scenes can be screened, and the part that meets the requirements of the occlusion scene is determined as the historical occlusion scene.
[0032] Furthermore, the lidar data is obtained through a lidar sensor. It should be noted that a lidar is a sensor that calculates distance and speed by emitting laser beams and measuring the reflection time.
[0033] When specifically implemented, the lidar can emit laser pulses towards the historical occlusion scene and receive the signals reflected from the object surfaces in the historical occlusion scene. By measuring the round-trip time of the laser pulses, the lidar can calculate the distance to each point in the surrounding environment and record these points in the form of three-dimensional coordinates in combination with the angle and height information of the laser emission.
[0034] Furthermore, the target supervised point cloud is generated based on the lidar data. It should be noted that the target supervised point cloud contains all of the historical occlusion scene. By supervising through the target supervised point cloud, it can be ensured that the supervised target is as close as possible to the target supervised point cloud, thus being closer to the real historical occlusion scene.
[0035] When specifically implemented, the specific value of the first preset threshold is determined according to actual needs and is not limited in this embodiment. For example, in one embodiment, the first preset threshold can be determined according to the lowest point cloud density value required for object detection.
[0036] Furthermore, due to occlusion, the point cloud density of the corresponding occluded target in the lidar data is low. The point cloud density at all positions in the target supervised point cloud should be greater than the first preset threshold. Therefore, after converting the lidar data into a point cloud, point cloud filling is required to prevent problems such as overly sparse point clouds and incomplete contours.
[0037] It should be noted that lidar data is used to supervise the data of millimeter-wave radar and vision camera during the training process. It can be understood that when generating point clouds, accurate point cloud generation can be completed only through the data of radar millimeter waves and vision cameras.
[0038] A specific embodiment is given below to introduce in detail the process of obtaining target-supervised point clouds:
[0039] (1) Split the lidar data into a specified number of target regions, and determine the target density of the point clouds within each target region.
[0040] Specifically, the specific value of the specified number is set according to actual needs. In this embodiment, it is not limited herein. For example, in one embodiment, the specified number can be set to 6 according to the number of surfaces of the lidar data.
[0041] It should be noted that when splitting the lidar data into target regions, it is necessary to ensure that the size and shape of each target region are exactly the same.
[0042] A specific embodiment is given below to introduce in detail the process of splitting lidar data:
[0043] Step 1: Obtain the geometric center and corner points of the target object in the lidar data.
[0044] Specifically, obtain the targets included in the lidar data. It should be noted that the specific types included in the target object are determined according to actual needs. In this embodiment, it is not limited herein. In specific implementation, the targets to be measured can include the background, pedestrians, vehicles, other objects, etc. in the historical occlusion scene.
[0045] In specific implementation, the targets included in the lidar data can be obtained based on traditional object detection methods or based on neural network object detection methods. In this embodiment, it is not limited herein.
[0046] Furthermore, perform object detection on the lidar data to obtain all the target objects therein. It should be noted that the detected target objects can be labeled with 3D boxes, and a corresponding target object can be completely divided into the 3D box by a 3D box.
[0047] Furthermore, parse the data corresponding to the target object into a series of three-dimensional coordinate points, and the position and shape of the target object can be represented by these points.
[0048] Step 2: Using the geometric center as the vertex of the target region, generate multiple bases based on the corner points, and generate the target region based on the vertex and the bases; wherein, the number of the bases is equal to the specified number; the shapes of all the target regions are the same.
[0049] Specifically, the sizes and shapes of the bases formed by the corner points are the same, and the distances from the geometric center are also the same, which can ensure that all the target regions are consistent.
[0050] When specifically implemented, generate the corresponding number of bases according to the specified number of corner points, and connect the corner points of each base to the geometric center to form the target region. For example, in an embodiment, if the specified number is 6, then 6 are generated based on 8 corner points, and then these 6 bases are connected to the geometric center to form 6 complete three-dimensional target regions. At this time, the shape of the target region is a triangular pyramid.
[0051] It should be noted that the base can be determined based on the smallest 3D box of the framed column target object. When specifically implemented, a 3D box can be generated based on the traditional 3D box generation method, and the 8 corner points of the generated 3D box are connected to the geometric center to form 6 triangular pyramids to obtain the target region.
[0052] The point cloud generation method based on the fusion of millimeter wave and visual image provided in this embodiment can obtain target regions with exactly the same size and shape by obtaining the combination center and corner points. In this way, through the target region to complete the three-dimensional shape modeling, it is possible to better understand the geometric structure and form in the historical occlusion scene, and further provide strong support for the completion of the target supervised point cloud. Further, by dividing the target region, the enhancement of the point cloud can be completed inside the lidar data, which ensures the accuracy and reliability of the data processing and reduces the calculation cost.
[0053] (2) Determine the target regions corresponding to the target density less than the first preset threshold as the regions to be enhanced, and increase the number of point clouds in the regions to be enhanced to obtain the target supervised point cloud.
[0054] Specifically, calculate the number of point clouds in the target region, and divide the number of point clouds by the volume of the target region to obtain the corresponding target density. It should be noted that each target region has a corresponding target density.
[0055] Further, determine the target regions with a target density lower than the target density as the regions to be enhanced. When specifically implemented, the traditional point cloud thickening method or the neural network point cloud thickening method can be used to increase the number of point clouds in the regions to be enhanced.
[0056] The following gives a specific embodiment to introduce in detail the process of obtaining the target supervised point cloud:
[0057] Step 1: Determine the target area corresponding to the target density greater than the first preset threshold as the sample area; wherein, the sample area is of the same type as the area to be enhanced.
[0058] In specific implementation, determine the target area with the target density greater than the first preset threshold as the sample area. It can be understood that the number of point clouds in the sample area is large and the point cloud density is dense.
[0059] It should be noted that the sample area is of the same type as the area to be enhanced, with the same size and shape, and the same type of specific data.
[0060] Step 2: Splice the point clouds in the sample area into the area to be enhanced in proportion to obtain the target supervised point cloud.
[0061] Specifically, the part with a low point cloud density in the area to be enhanced needs to be supplemented, and the point clouds at the corresponding positions of the area to be enhanced and the sample area are spliced and copied into the sample area. It should be noted that the proportion of the number of point clouds spliced each time is determined according to actual needs and is not limited in this embodiment. For example, in one embodiment, 30% of the point clouds at position A in the sample area can be extracted and spliced into position A in the area to be enhanced.
[0062] Furthermore, after filling the point clouds at all positions in the area to be enhanced, the target supervised point cloud is obtained.
[0063] The point cloud generation method based on the fusion of millimeter wave and visual image provided in this embodiment splices the point clouds in the area to be enhanced with the sample area where the point cloud density meets the requirements, so that the point cloud density in the area to be enhanced meets the requirements. In this way, the target supervised point cloud can contain rich data, effectively reflect all the information in the area to be recognized and occluded area, and provide high-quality supervision information. Furthermore, by splicing the point clouds in the sample area into the area to be enhanced, the boundary contour in each target area is ensured to be complete, preventing edge blurring, and improving the accuracy and reliability of point cloud generation.
[0064] The point cloud generation method based on the fusion of millimeter - wave and visual images provided in this embodiment first splits the lidar data into a specified number of target regions, determines the target density of the point cloud in each target region, then determines the target regions corresponding to the target density less than the first preset threshold as the regions to be enhanced, increases the number of point clouds in the regions to be enhanced, and obtains the target supervised point cloud. In this way, by enhancing the lidar data and complementing the point cloud of occluded targets, the overall quality of the point cloud can be significantly improved. The enhanced point cloud L is more complete in shape contour, has a higher density, and can better reflect the target situation in the real environment. Further, through the high - quality target supervised point cloud, a more accurate and reliable supervision signal can be provided for point cloud generation, which helps to improve the accuracy and robustness of point cloud generation.
[0065] S102. Identify the historical occlusion scene based on the millimeter - wave radar and the visual camera respectively to obtain the historical millimeter - wave feature and the historical visual feature.
[0066] Specifically, a millimeter - wave radar is a radar sensor operating in the millimeter - wave frequency band. The millimeter - wave radar uses electromagnetic waves for detection and ranging, and has the characteristics of stable detection performance, long operating distance, and good environmental adaptability.
[0067] In specific implementation, when the electromagnetic wave emitted by the millimeter - wave radar encounters a target, it will be reflected and received by the radar. By measuring the time difference and phase difference between the transmitted wave and the reflected wave, information such as the distance and azimuth angle of the target can be calculated.
[0068] Further, the visual camera collects the light emitted from the surrounding and irradiated onto the photosensitive surface through the camera lens, thereby generating a clear image. It should be noted that the visual camera has the advantages of low price and high environmental image resolution.
[0069] In specific implementation, the specific models of the millimeter - wave radar and the visual camera are determined according to actual needs. In this embodiment, no limitation is imposed on this. For example, in one embodiment, a monocular camera can be selected as the visual camera, or a binocular camera can be selected as the visual camera.
[0070] Further, feature extraction is respectively performed on the data obtained by shooting the millimeter - wave radar and the visual camera to obtain the millimeter - wave feature and the visual feature.
[0071] The following gives a specific embodiment to introduce in detail the process of obtaining the millimeter - wave feature:
[0072] (1) Detect the historical occlusion scene based on the millimeter - wave radar to obtain the historical millimeter - wave information.
[0073] Specifically, after detecting a historical occlusion scenario with a millimeter-wave radar, the corresponding historical millimeter-wave information is obtained. The received historical millimeter-wave information is processed through a radar signal processing algorithm to extract parameters such as the distance and angle of different targets in the historical occlusion scenario, and these parameters are parsed to form the historical millimeter-wave information. For example, in one embodiment, a millimeter-wave radar is used to detect an object after vehicle occlusion in a historical scenario, and the reflected signal is processed through an FFT algorithm to obtain the historical millimeter-wave information corresponding to the object.
[0074] (2) Dimensionally reduce and map the historical millimeter-wave information into a bird's-eye view space to obtain two-dimensional data.
[0075] Specifically, the bird's-eye view space is a representation space that projects three-dimensional space information onto a two-dimensional plane. It should be noted that in the bird's-eye view space, the positions and shapes of targets such as obstacles can be represented more intuitively.
[0076] Furthermore, after dimensional reduction mapping, the historical millimeter-wave radar information is converted into two-dimensional data in the bird's-eye view space, and these data are usually represented in the form of a two-dimensional matrix or image. In specific implementation, each pixel or grid cell in the bird's-eye view space corresponds to a position in the bird's-eye view space, and there is corresponding information at each position.
[0077] In specific implementation, the points in the historical millimeter-wave information can be converted into the coordinate system of the bird's-eye view space, and positioned on the plane of the bird's-eye view space according to the distance and angle information of the target.
[0078] (3) Extract features from the two-dimensional data to obtain the historical millimeter-wave features.
[0079] In specific implementation, the millimeter-wave features corresponding to the two-dimensional data can be extracted based on traditional feature extraction methods or based on neural network feature extraction methods. In this embodiment, no limitation is imposed on this. For example, in one embodiment, the two-dimensional data is input into a convolutional network CNN for feature extraction to obtain the historical millimeter-wave features. Another example, in another embodiment, the two-dimensional data is input into a pre-trained feature extraction Radar Encoder module, and the historical millimeter-wave features are extracted through the Radar Encoder module.
[0080] The point cloud generation method based on the fusion of millimeter - wave and visual images provided in this embodiment first detects historical occlusion scenarios based on a millimeter - wave radar to obtain historical millimeter - wave information. Then, the historical millimeter - wave information is dimension - reduced and mapped into the bird's - eye view space to obtain two - dimensional data. Finally, feature extraction is performed on the two - dimensional data to obtain historical millimeter - wave features. In this way, by using key information such as distance and angle information provided by the millimeter - wave radar and obtaining historical millimeter - wave features through the optimization process of a deep - learning model, the accuracy and precision of target detection can be improved. Further, the acquisition of historical millimeter - wave features enables the information of the millimeter - wave radar to be fused with the information of other sensors, thereby comprehensively utilizing the advantages of different sensors and improving the overall performance and robustness of the system. Further, due to the multipath reflection and certain penetration ability of the millimeter - wave radar, targets that are occluded or out of sight can be detected. Therefore, by extracting historical millimeter - wave features, the perception ability of the system for complex environments can be significantly enhanced. In this way, by obtaining accurate historical millimeter - wave features, it is possible to better adapt to different environmental and scenario requirements and improve the accuracy and reliability of point cloud generation.
[0081] A specific embodiment is given below to introduce in detail the process of obtaining visual features:
[0082] (1) Detect the historical occlusion scenario based on the visual camera to obtain historical visual information.
[0083] Specifically, the visual camera captures the light of the surrounding environment through the lens to form an image. It can be understood that the image captured by the visual camera contains rich color, texture, and shape information.
[0084] It should be noted that the historical visual information can exist in the form of an RGB image, and the static and dynamic information in the historical occlusion scenario is saved through the historical visual information.
[0085] (2) Perform feature extraction on the historical visual information to obtain primary features.
[0086] In specific implementation, the primary features corresponding to the historical visual information can be extracted based on traditional feature extraction methods or based on neural - network - based feature extraction methods. In this embodiment, no limitation is imposed on this. For example, in one embodiment, the historical visual information is input into a convolutional network CNN for feature extraction to obtain primary features. For another example, in another embodiment, the historical visual information is input into a pre - trained feature extraction CameraEncoder module, and the primary features are extracted through the Camera Encoder module.
[0087] (3) Generate a predicted depth map based on the primary features.
[0088] Specifically, the distance information of an object in a three-dimensional scene is inferred from the primary features by means of depth estimation. In specific implementation, a predicted depth map can be generated based on the visual cues in the primary features.
[0089] It should be noted that a predicted depth map can be generated based on traditional depth estimation methods or neural network-based depth estimation methods. In this embodiment, no limitation is imposed on this.
[0090] (4) Based on a preset calibration matrix and the predicted depth map, map the primary features into the bird's-eye view space to obtain the historical visual features.
[0091] Specifically, the calibration matrix can convert the coordinate system of the visual camera and the coordinate system of the bird's-eye view space. After processing the predicted depth map based on the calibration matrix, the historical visual features in the bird's-eye view space are obtained. It should be noted that the calibration matrix contains the geometric transformation parameters required to convert the points in the camera image into the bird's-eye view space.
[0092] In specific implementation, use the calibration matrix to perform coordinate transformation on each primary feature, and convert it from the visual camera coordinate system to the bird's-eye view space coordinate system. It should be noted that the position of the pixel in the bird's-eye view space can be calculated using the depth information.
[0093] The point cloud generation method based on the fusion of millimeter wave and visual image provided in this embodiment first detects the historical occlusion scene based on the visual camera to obtain historical visual information, then extracts features from the historical visual information to obtain primary features, then depth-estimates the primary features to generate a predicted depth map, and then maps the primary features into the bird's-eye view space based on a preset calibration matrix and the predicted depth map to obtain historical visual features. In this way, by capturing the light of the occlusion environment to be recognized by the visual camera to form an image, the key information in the image, such as edges, corners, texture features, etc., can be further highlighted, thereby providing rich environmental information for subsequent processing. Further, the historical visual features not only contain the appearance information of the target, but also implicitly contain key information such as the shape, size, and position of the target. Further, as an important information output by the visual camera, the historical visual features can improve the reliability and accuracy of point cloud generation. Further, during the acquisition process of the historical visual features, through multiple steps of processing such as feature extraction, depth estimation, and coordinate transformation, the influence of noise and interference can be reduced, and the stability and reliability of the historical visual features can be improved. In this way, by obtaining accurate historical visual features, different environmental and scene requirements can be better adapted, and the accuracy and reliability of point cloud generation can be improved.
[0094] S103. Using the target supervised point cloud as the decoder supervision, input the historical millimeter-wave feature and the historical visual feature into the decoder, and train the decoder to predict and generate occluded target point clouds and unoccluded target point clouds.
[0095] Specifically, the decoder is a complex neural network structure responsible for decoding high-level feature information into specific point cloud data. In specific implementation, inside the decoder, through operations such as multi-layer convolution, transposed convolution, and upsampling, the input historical millimeter-wave feature and historical visual feature are gradually expanded into point cloud data with spatial dimensions.
[0096] In the specific implementation of this step, an initial decoder can be trained based on the historical millimeter-wave feature and the historical visual feature to obtain a trained decoder.
[0097] The following gives a specific embodiment to introduce in detail the process of obtaining occluded target point clouds and unoccluded target point clouds:
[0098] (1) Amplify the discriminative features of the historical millimeter-wave feature and the historical visual feature to obtain the occluding features of the occluded target; amplify the common features of the historical millimeter-wave feature and the historical visual feature to obtain the common shared features.
[0099] Specifically, by amplifying the discriminative features of the historical millimeter-wave feature and the historical visual feature, the unique parts between the two can be determined. By amplifying the common features of the historical millimeter-wave feature and the historical visual feature, the commonalities between the two can be determined.
[0100] The following gives a specific embodiment to introduce in detail the process of obtaining discriminative features and common features:
[0101] Step 1: Differentially amplify the historical millimeter-wave feature and the historical visual feature to obtain the discriminative features.
[0102] Specifically, perform a differential operation on the millimeter-wave radar feature and the historical visual feature to highlight the different parts between the two. It should be noted that these different parts mainly reflect the target information that the millimeter-wave radar can detect but the visual camera may not be able to see due to occlusion, that is, the occluding features of the occluded object.
[0103] Furthermore, the feature obtained by differential amplification is determined as the discriminative feature.
[0104] Step 2: Sum and amplify the historical millimeter-wave feature and the historical visual feature to obtain the common features.
[0105] Specifically, a summation operation is performed on the millimeter-wave radar features and the historical visual features to retain the same information between the two. It should be noted that this same information mainly reflects the unoccluded parts of the scene, including unoccluded objects and the background, that is, the same features.
[0106] Furthermore, the feature obtained by summation amplification is determined as the same feature.
[0107] The point cloud generation method based on millimeter-wave and visual image fusion provided in this embodiment obtains the difference features by differentially amplifying the historical millimeter-wave features and the historical visual features, and obtains the same features by summation amplification of the historical millimeter-wave features and the historical visual features. In this way, through differential amplification, the occluded targets can be acquired, and accurate occluded target point clouds can be generated. Furthermore, through summation amplification, a comprehensive point cloud representation including all visible targets and background information in the scene can be generated. In this way, through the same features and difference features, the comprehensiveness and accuracy of environmental perception can be improved.
[0108] (2) Process the difference features based on the decoder to obtain the occluded target point cloud; process the same features based on the decoder to obtain the unoccluded target point cloud; wherein, the unoccluded target point cloud includes the point cloud of the unoccluded object and the point cloud of the background.
[0109] Specifically, the difference features obtained after differential amplification are input into the decoder, and the decoder decodes the difference features, maps the high-level difference features to the specific point cloud coordinates and attributes in the three-dimensional space, and generates the occluded target point cloud.
[0110] Furthermore, the same features obtained after summation amplification are input into the decoder, and the decoder decodes the difference features, maps the high-level same features to the specific point cloud coordinates and attributes in the three-dimensional space, and generates the unoccluded target point cloud.
[0111] It should be noted that since the decoder is trained with the target supervised point cloud as the supervision, when the decoder generates the occluded target point cloud and the unoccluded target point cloud, the decoder can ensure that the generated occluded target point cloud and unoccluded target point cloud are close to the actual scene.
[0112] Specifically, the same features obtained after summation amplification can be directly input into the decoder, and with the target supervised point cloud as the supervision information, the decoder directly predicts and generates the point cloud of the unoccluded target.
[0113] A specific embodiment is given below to introduce in detail the process of obtaining the occluded target point cloud:
[0114] (1) Determine the position information of the occluded target corresponding to the occluded feature and generate the region of interest.
[0115] Specifically, first, by analyzing the occlusion features obtained after differential amplification, the position information of the occlusion target in the scene is determined.
[0116] In specific implementation, a classifier, a regressor, a clustering algorithm, etc. can be used to identify the position corresponding to the occlusion feature and determine it as the region of interest.
[0117] It should be noted that the region of interest can include the position corresponding to the occlusion feature and the region of data within a certain range around it.
[0118] The occlusion feature corresponding to the occlusion target can be input into the feature detector to locate the region of interest.
[0119] (2) Input the features within the region of interest into the decoder, and based on the decoder, obtain high-dimensional target coordinates and target attributes.
[0120] In specific implementation, the features within the region of interest can be extracted based on traditional feature extraction methods or based on neural network feature extraction methods. In this embodiment, no limitation is imposed on this.
[0121] Furthermore, the features within the extracted region of interest are input into the decoder, and the decoder decodes and transforms these features to generate high-dimensional target coordinates and target attributes.
[0122] (3) Denoise the target coordinates and the target attributes based on the target supervised point cloud to generate the occlusion target point cloud.
[0123] Specifically, the target coordinates and the target attributes are supervised by the target supervised point cloud to generate a high-quality occlusion target point cloud.
[0124] In specific implementation, the features of the region of interest can be directly input into the decoder, and the target supervised point cloud is used as the supervision information, and the decoder directly predicts and generates the occlusion target point cloud.
[0125] The method for generating point cloud based on the fusion of millimeter wave and visual image provided by this embodiment first determines the position information of the occluded target corresponding to the occlusion feature, generates the region of interest, and then inputs the features within the region of interest into the decoder. Based on the decoder, high-dimensional target coordinates and target attributes are obtained. Finally, denoising processing is performed on the target coordinates and target attributes based on the target-supervised point cloud to generate the occluded target point cloud. In this way, the computational amount of subsequent processing is reduced through the region of interest, and the attention is concentrated on the occluded target, ensuring the accuracy and reliability of the occluded target point cloud. Further, by analyzing the occlusion feature and determining the position information of the occluded target, the occluded target can be more accurately identified. In this way, through the accurate and reliable occluded point cloud, the reliability and accuracy of the point cloud generation corresponding to the occlusion scene to be recognized can be ensured.
[0126] In this step, the decoder is trained through the target-supervised point cloud, historical millimeter wave features, and historical visual features to obtain a trained decoder.
[0127] S104. Identify the occlusion scene to be recognized based on the millimeter wave radar and the visual camera to obtain millimeter wave features and visual features, and predict and generate the scene point cloud of the occlusion scene to be recognized based on the millimeter wave features, the visual features, and the trained decoder.
[0128] Specifically, the occlusion scene to be recognized is the same as the historical occlusion scene, both of which are scenes containing occlusion. In specific implementation, when point cloud generation is required, the corresponding scene is determined as the occlusion scene to be recognized.
[0129] Further, by detecting the occlusion scene to be recognized based on the millimeter wave radar and the visual camera, corresponding millimeter wave features and visual features can be generated.
[0130] Further, after extracting the millimeter wave features and visual features corresponding to the occlusion scene to be recognized, the millimeter wave features and visual features are input into the trained decoder, and the scene point cloud of the occlusion scene to be recognized is generated through the decoder.
[0131] It can be understood that the scene point cloud includes the point cloud of the unoccluded object in the occlusion scene to be recognized and also includes the point cloud of the occluded object. In specific implementation, the decoder determines the point cloud of the unoccluded object and the point cloud of the occluded object respectively by combining the millimeter wave features and the visual features, and combines the two to obtain the scene point cloud.
[0132] Further, since the point cloud of the unoccluded object and the point cloud of the occluded object can be generated in different coordinate systems or perspectives, they need to be aligned first. In specific implementation, operations such as coordinate transformation, rotation, and translation can be performed to ensure the spatial consistency of the point cloud of the unoccluded object and the point cloud of the occluded object.
[0133] Further, the point cloud of the unoccluded object and the point cloud of the occluded object can be fused based on traditional point cloud fusion methods, or can be fused based on neural network-based point cloud fusion methods. In this embodiment, no limitation is imposed thereon. During specific implementation, a scene point cloud containing all information of the occluded scene to be recognized can be generated through superposition, weighted average, or more complex fusion algorithms.
[0134] The point cloud generation method based on millimeter-wave and visual image fusion provided by this application first obtains lidar data of a historical occluded scene and generates a target supervised point cloud based on the lidar data. Among them, the point cloud density of the occluded target in the lidar data is less than a first preset threshold, the point cloud density of the occluded target in the target supervised point cloud is greater than the first preset threshold, and the point cloud contour of the occluded target in the target supervised point cloud is complete. Furthermore, the historical occluded scene is respectively recognized based on a millimeter-wave radar and a visual camera to obtain historical millimeter-wave features and historical visual features. Then, using the target supervised point cloud as decoder supervision, the historical millimeter-wave features and the historical visual features are input into the decoder to train the decoder to predict and generate an occluded target point cloud and an unoccluded target point cloud. Finally, the millimeter-wave radar and the visual camera are used to recognize the occluded scene to be recognized to obtain millimeter-wave features and visual features, and the scene point cloud of the occluded scene to be recognized is predicted and generated based on the millimeter-wave features, the visual features, and the trained decoder. In this way, through the penetration ability of the millimeter-wave radar, the millimeter-wave features of the occluded part can be effectively obtained. At the same time, through the high-level semantic information of the visual camera, the visual features containing rich context information can be effectively obtained. Further, the unoccluded point cloud is determined based on the same features of the millimeter-wave features and the visual features, and the occluded point cloud is determined based on the different features of the millimeter-wave features and the visual features. In this way, by combining the occluded point cloud and the unoccluded point cloud, the occluded target can be effectively detected and recognized. The penetration ability of the millimeter-wave radar makes up for the deficiency of the visual image in the case of occlusion, while the high-level semantic information of the visual image helps to improve the detection accuracy. Further, when generating the occluded point cloud and the unoccluded point cloud, a data augmentation method is used to obtain a target supervised point cloud with a complete and dense shape contour to generate a target training sample to train the decoder, ensuring the accuracy and reliability of point cloud generation. In this way, the finally generated scene point cloud contains complete information of all targets in the scene, helps to reduce misidentifications caused by a complex environment, and ensures the reliability and accuracy of point cloud generation.
[0135] Corresponding to the foregoing embodiment of a point cloud generation method based on millimeter-wave and visual image fusion, this application also provides an embodiment of a point cloud generation device based on millimeter-wave and visual image fusion.
[0136] An embodiment of the point cloud generation device based on millimeter-wave and visual image fusion in this application can be applied to the point cloud generation device based on millimeter-wave and visual image fusion. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of the point cloud generation device based on millimeter-wave and visual image fusion where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. At the hardware level, as Figure 2 shown, it is a hardware structure diagram of the point cloud generation device based on millimeter-wave and visual image fusion where the point cloud generation device based on millimeter-wave and visual image fusion in this application is located. In addition to Figure 2 the processor, memory, network interface, and non-volatile memory shown, the point cloud generation device where the device is located in the embodiment usually includes other hardware according to the actual functions of the point cloud generation device based on millimeter-wave and visual image fusion, which will not be elaborated here.
[0137] Figure 3 It is a schematic structural diagram of Embodiment 1 of the point cloud generation device based on millimeter-wave and visual image fusion provided in this application. Please refer to Figure 3 , the device provided in this embodiment includes an acquisition module 310, an identification module 320, a prediction module 330, and a generation module 340; where
[0138] The acquisition module 310 is used to acquire lidar data of a historical occlusion scene and generate a target supervised point cloud based on the lidar data;
[0139] Among them, the point cloud density of the occluding target in the lidar data is less than a first preset threshold, the point cloud density of the occluding target in the target supervised point cloud is greater than the first preset threshold, and the point cloud contour of the occluding target in the target supervised point cloud is complete;
[0140] The identification module 320 is used to respectively identify the historical occlusion scene based on the millimeter-wave radar and the visual camera to obtain historical millimeter-wave features and historical visual features;
[0141] The prediction module 330 is used to use the target supervised point cloud as decoder supervision, input the historical millimeter-wave features and the historical visual features into the decoder, and train the decoder to predict and generate occluding target point clouds and non-occluding target point clouds;
[0142] The generation module 340 is used to identify a to-be-identified occlusion scene based on the millimeter-wave radar and the visual camera to obtain millimeter-wave features and visual features, and predict and generate the scene point cloud of the to-be-identified occlusion scene based on the millimeter-wave features, the visual features, and the trained decoder.
[0143] The device of this embodiment can be used to execute Figure 1 the steps of the method embodiment shown. The specific implementation principle and process are similar, and will not be elaborated here.
[0144] Optionally, the obtaining module 310 is specifically configured to split the lidar data into a specified number of target regions, and determine the target density of the point cloud in each target region;
[0145] The obtaining module 310 is further specifically configured to determine the target region corresponding to the target density less than the first preset threshold as the region to be enhanced, increase the number of point clouds in the region to be enhanced, and obtain the target supervised point cloud.
[0146] Optionally, the obtaining module 310 is further specifically configured to determine the target region corresponding to the target density greater than the first preset threshold as the sample region; wherein, the type of the sample region is the same as that of the region to be enhanced;
[0147] The obtaining module 310 is further specifically configured to splice the point clouds in the sample region into the region to be enhanced according to a proportion, and obtain the target supervised point cloud.
[0148] Optionally, the obtaining module 310 is further specifically configured to obtain the geometric center and corner points of the target object in the lidar data;
[0149] The obtaining module 310 is further specifically configured to use the geometric center as the vertex of the target region, generate multiple bases based on the corner points, and connect the vertex and the corner points of the base to generate the target region; wherein, the number of the bases is equal to the specified number; the shapes of the respective target regions are the same.
[0150] Optionally, the recognition module 320 is specifically configured to detect the historical occlusion scene based on the millimeter-wave radar, and obtain historical millimeter-wave information;
[0151] The recognition module 320 is further specifically configured to dimensionally reduce and map the historical millimeter-wave information into the bird's-eye view space to obtain two-dimensional data;
[0152] The recognition module 320 is further specifically configured to extract features from the two-dimensional data to obtain the historical millimeter-wave features.
[0153] Optionally, the recognition module 320 is further specifically configured to detect the historical occlusion scene based on the vision camera, and obtain historical vision information;
[0154] The recognition module 320 is further specifically configured to extract features from the historical vision information to obtain primary features;
[0155] The recognition module 320 is further specifically configured to generate a predicted depth map based on the primary features;
[0156] The recognition module 320 is further specifically configured to map the primary features into the bird's-eye view space based on a preset calibration matrix and the predicted depth map to obtain the historical visual features.
[0157] Optionally, the prediction module 330 is specifically configured to amplify the distinct features between the historical millimeter-wave features and the historical visual features to obtain the occlusion features of the occluded target; amplify the identical features between the historical millimeter-wave features and the historical visual features to obtain the identical common features;
[0158] The prediction module 330 is further specifically configured to process the distinct features based on the decoder to obtain the point cloud of the occluded target; process the identical features based on the decoder to obtain the point cloud of the unoccluded target; wherein the point cloud of the unoccluded target includes the point cloud of the unoccluded object and the point cloud of the background.
[0159] Optionally, the prediction module 330 is further specifically configured to determine the position information of the occluded target corresponding to the occlusion features and generate an interested region;
[0160] The prediction module 330 is further specifically configured to input the features within the interested region into the decoder and obtain the high-dimensional target coordinates and target attributes based on the decoder;
[0161] The prediction module 330 is further specifically configured to perform denoising processing on the target coordinates and the target attributes based on the target supervised point cloud to generate the point cloud of the occluded target.
[0162] Optionally, the prediction module 330 is further specifically configured to differentially amplify the historical millimeter-wave features and the historical visual features to obtain the distinct features;
[0163] The prediction module 330 is further specifically configured to sum and amplify the historical millimeter-wave features and the historical visual features to obtain the identical features.
[0164] Please continue to refer to Figure 2 , this application further provides a point cloud generation device based on the fusion of millimeter-wave and visual images, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of any of the methods provided in the first aspect of this application are implemented.
[0165] This application further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of any of the methods provided in this application are implemented.
[0166] For the implementation processes of the functions and roles of each unit in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0167] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0168] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the scope of protection of this application.
Claims
1. A point cloud generation method based on millimeter wave and visual image fusion, characterized in that: The method comprises: Acquire laser radar data of historical occlusion scenes, and generate a target supervision point cloud based on the laser radar data; The step of acquiring laser radar data of a historical occlusion scene and generating a target supervision point cloud based on the laser radar data includes: Splitting the laser radar data into a specified number of target areas, and determining the target density of the point cloud within each of the target areas; The target area corresponding to the target density less than the first preset threshold is determined as the area to be enhanced, and the number of point clouds in the area to be enhanced is increased to obtain the target supervision point cloud; The point cloud density of the occluded target in the laser radar data is less than a first preset threshold, the point cloud density of the occluded target in the target supervision point cloud is greater than the first preset threshold, and the point cloud contour of the occluded target in the target supervision point cloud is complete; Based on the millimeter wave radar and the visual camera, the historical occlusion scene is respectively identified to obtain the historical millimeter wave feature and the historical visual feature; Using the target supervision point cloud as a decoder supervision, inputting the historical millimeter wave features and the historical visual features into a decoder, and training the decoder to predict and generate an occluded target point cloud and an unoccluded target point cloud; The method uses the target supervision point cloud as a decoder supervision, inputs the historical millimeter wave features and the historical visual features into a decoder, and trains the decoder to predict and generate an occluded target point cloud and an unoccluded target point cloud, including: amplifying the distinguishing features of the historical millimeter wave features and the historical visual features to obtain the occlusion features of the occluded target; Processing the distinguishing features based on the decoder to obtain the occluded target point cloud; The amplifying the distinguishing features of the historical millimeter wave features and the historical visual features includes: differentially amplifying the historical millimeter wave feature and the historical visual feature to obtain the distinguishing feature; Based on the millimeter wave radar and the visual camera, the occluded scene to be identified is identified to obtain millimeter wave features and visual features, and based on the millimeter wave features, the visual features and the trained decoder prediction, a scene point cloud of the occluded scene to be identified is generated.
2. The method according to claim 1, characterized in that The step of determining the target area corresponding to the target density less than the first preset threshold as the area to be enhanced, increasing the number of point clouds in the area to be enhanced, and obtaining the target supervision point cloud includes: Determine the target area corresponding to the target density greater than the first preset threshold as a sample area; wherein the sample area is of the same type as the area to be enhanced; The point cloud in the sample area is spliced into the area to be enhanced in proportion to obtain the target supervised point cloud.
3. The method according to claim 1, characterized in that The step of dividing the laser radar data into a specified number of target areas and determining the target density of the point cloud in each target area includes: Obtaining the geometric center and corner points of the target object in the laser radar data; The geometric center is used as the vertex of the target area, multiple bottom surfaces are generated based on the corner points, and the target area is generated by connecting the vertex and the corner points of the bottom surfaces; wherein the number of the bottom surfaces is equal to the specified number; and the shapes of the target areas are the same.
4. The method according to claim 1, characterized in that: The millimeter wave radar and the visual camera are used to respectively identify the historical occlusion scenes to obtain the historical millimeter wave features and the historical visual features, including: Detecting the historical occlusion scene based on the millimeter wave radar, obtaining historical millimeter wave information; Mapping the historical millimeter wave information into a bird's-eye view space by reducing the dimension of the historical millimeter wave information to obtain two-dimensional data; Feature extraction is performed on the two-dimensional data to obtain the historical millimeter wave features.
5. The method according to claim 1, characterized in that The millimeter wave radar and the visual camera are used to respectively identify the historical occlusion scenes to obtain the historical millimeter wave features and the historical visual features, including: Detecting the historical occlusion scene based on the visual camera to obtain historical visual information; Extracting features from the historical visual information to obtain primary features; generating a predicted depth map based on the primary features; Based on a preset calibration matrix and the predicted depth map, the primary features are mapped into a bird's-eye view space to obtain the historical visual features.
6. The method according to claim 1, characterized in that The method uses the target supervision point cloud as a decoder supervision, inputs the historical millimeter wave features and the historical visual features into a decoder, and trains the decoder to predict and generate an occluded target point cloud and an unoccluded target point cloud, including: amplifying the same features of the historical millimeter wave features and the historical visual features to obtain the same common features; Based on the decoder processing the same features, the unobstructed target point cloud is obtained; wherein the unobstructed target point cloud includes the point cloud of the unobstructed object and the point cloud of the background.
7. The method according to claim 1, characterized in that The processing of the distinguishing features based on the decoder to obtain the occluded target point cloud includes: Determine the position information of the occlusion target corresponding to the occlusion feature, and generate a region of interest; Inputting the features in the region of interest into the decoder, and obtaining high-dimensional target coordinates and target attributes based on the decoder; The target coordinates and the target attributes are denoised based on the target supervision point cloud to generate the occluded target point cloud.
8. The method according to claim 6, characterized in that Amplifying the same features of the historical millimeter wave features and the historical visual features to obtain the same features includes: The historical millimeter wave feature and the historical visual feature are summed and amplified to obtain the same feature.
9. A point cloud generation device based on millimeter wave and visual image fusion, characterized in that: The device comprises an acquisition module, an identification module, a prediction module and a generation module; wherein, The acquisition module is used to acquire the laser radar data of the historical occlusion scene and generate the target supervision point cloud based on the laser radar data; The step of acquiring laser radar data of a historical occlusion scene and generating a target supervision point cloud based on the laser radar data includes: Splitting the laser radar data into a specified number of target areas, and determining the target density of the point cloud within each of the target areas; The target area corresponding to the target density less than the first preset threshold is determined as the area to be enhanced, and the number of point clouds in the area to be enhanced is increased to obtain the target supervision point cloud; The point cloud density of the occluded target in the laser radar data is less than a first preset threshold, the point cloud density of the occluded target in the target supervision point cloud is greater than the first preset threshold, and the point cloud contour of the occluded target in the target supervision point cloud is complete; The recognition module is used to respectively recognize the historical occlusion scene based on the millimeter wave radar and the visual camera to obtain the historical millimeter wave features and the historical visual features; The prediction module is used to use the target supervision point cloud as a decoder supervision, input the historical millimeter wave features and the historical visual features into a decoder, and train the decoder to predict and generate an occluded target point cloud and an unoccluded target point cloud; The method uses the target supervision point cloud as a decoder supervision, inputs the historical millimeter wave features and the historical visual features into a decoder, and trains the decoder to predict and generate an occluded target point cloud and an unoccluded target point cloud, including: amplifying the distinguishing features of the historical millimeter wave features and the historical visual features to obtain the occlusion features of the occluded target; Processing the distinguishing features based on the decoder to obtain the occluded target point cloud; The amplifying the distinguishing features of the historical millimeter wave features and the historical visual features includes: differentially amplifying the historical millimeter wave feature and the historical visual feature to obtain the distinguishing feature; The generation module is used to identify the occluded scene to be identified based on the millimeter wave radar and the visual camera, obtain millimeter wave features and visual features, and generate a scene point cloud of the occluded scene to be identified based on the millimeter wave features, the visual features and the trained decoder prediction.