A dynamic scene perception method, a perception system and an intelligent driving device
By combining frame cameras and event cameras and using depth models and radar point cloud training methods, the difficulties of frame cameras and LiDAR in recognition under adverse weather and high-speed scenarios have been solved, achieving accurate target recognition in complex environments and improving recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510028385.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-01-08
AI Technical Summary
Existing frame cameras and lidar suffer from insufficient accuracy and fusion difficulties when identifying targets in adverse weather conditions and high-speed, high-dynamic scenarios, resulting in limited identification capabilities and increased uncertainty.
By combining frame cameras and event cameras, depth information of events is identified through a depth model, and a depth model is trained using radar point clouds to reconstruct clear images. By combining the advantages of event cameras and frame cameras, accurate target identification can be achieved.
In high-speed, high-dynamic scenarios, it can accurately identify targets in the scene, improve recognition accuracy and robustness, avoid the shortcomings of traditional fusion solutions, and enhance perception capabilities in complex environments.
Smart Images

Figure CN120070847B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image analysis technology, and more specifically, relates to a dynamic scene perception method and perception system, and intelligent driving equipment. Background Technology
[0002] Target perception technology is an important part of intelligent driving and a safety guarantee for vehicles. Currently, commonly used perception methods include frame camera imaging, lidar detection, and combinations of the two.
[0003] Due to the limitations of physical mechanisms, the performance of traditional frame cameras deteriorates significantly in adverse weather conditions such as heavy rain, heavy snow, and dense fog, affecting the accuracy of object recognition and vehicle safety; and in high dynamic scenes, blind spots such as underexposure and overexposure, as well as high-speed blurring problems may occur.
[0004] LiDAR suffers from drawbacks such as large size, high energy consumption, and high cost. Furthermore, due to its low frame rate, it is prone to motion distortion, sparse point cloud data, and even missed detections when processing fast-moving objects, especially small targets, thus limiting its recognition capabilities. In addition, in adverse weather conditions such as rain, snow, fog, haze, and sandstorms, the laser beam is scattered, leading to a significant deterioration in ranging accuracy.
[0005] The fusion scheme of frame camera and LiDAR faces objective difficulties due to the differences in data attributes between RGB data and 3D point clouds, and there is currently a lack of a large-scale 3D perception model that can unify point clouds and images. If the two are processed separately, the AI system lacks good interpretability, and the effective fusion distance of LiDAR is relatively short, resulting in completely different confidence levels between the two. This often leads to the problem of "not knowing which to trust," failing to achieve tight coupling and deep integration. In general target perception tasks, while the sensors bring in more information, they also introduce more uncertainty.
[0006] Therefore, how to accurately identify targets in high-speed, high-dynamic scenarios is a technical problem that urgently needs to be solved. Summary of the Invention
[0007] In view of the above-mentioned defects or improvement needs of the existing technology, the present invention provides a dynamic scene perception method and perception system, and intelligent driving equipment, the purpose of which is to accurately identify targets in high-speed and highly dynamic scenes.
[0008] To achieve the above objectives, the present invention provides a dynamic scene perception method, comprising:
[0009] Acquire the RGB image obtained by the frame camera at the current exposure time, and a series of events output by the event camera during the current exposure time of the frame camera;
[0010] The depth information of each event is identified using a depth model; the depth model is a model trained with events labeled with depth information, and the depth information labels are the depth information of the corresponding training events extracted based on radar point clouds.
[0011] The optimization model is solved to reconstruct the sharp image at the current moment. The goal of the optimization model is to make the event image and the downsampled result of the sharp image convolve and add image noise to approximate the RBG image. The event image is the image obtained by double integration of a series of events output during the current exposure. The downsampling is to reduce the resolution of the sharp image to the resolution of the event camera.
[0012] Based on the reconstructed clear images and depth information of events, the system can perceive targets in the current scene.
[0013] Optionally, the method further includes training the deep model before using it to identify the depth information of each event, the training step including:
[0014] Acquire LiDAR point clouds and a series of training events output by the event camera during the acquisition of LiDAR point clouds;
[0015] Extract target point cloud data from each frame of LiDAR point cloud;
[0016] The laser point cloud data of the same target in multiple consecutive frames are aligned to the current frame and then superimposed by the point cloud registration algorithm to obtain the high-density point cloud data of the corresponding target in the current frame.
[0017] By combining the high-density point cloud data of all targets in the current frame, the high-density point cloud of the current frame is obtained;
[0018] The high-density point cloud of the current frame is cropped to fit within the field of view of the event camera;
[0019] The clipped high-density point cloud is voxelized into a 3D placeholder mesh;
[0020] The 3D placeholder mesh is projected onto the shooting surface of the event camera, and voxels whose intensity does not reach the predetermined threshold are removed. The remaining voxels are used as the 3D placeholder ground truth of the current frame.
[0021] The depth information of each training event in the current frame is obtained by combining the 3D placeholder ground truth of the current frame; the depth model is trained by taking the training events as input and the corresponding depth information as labels.
[0022] Optionally, the loss functions used to train the deep model include Dice loss and Focal loss.
[0023] Optionally, the target is an obstacle that obstructs travel, including dynamic obstacles and non-ground static obstacles.
[0024] Optionally, the reconstructed clear image and the depth information of the event are mapped using a hash table.
[0025] The present invention also provides a dynamic scene perception system, comprising:
[0026] The acquisition unit is used to obtain the RBG image obtained by the frame camera at the current exposure time, as well as a series of events output by the event camera during the current exposure of the frame camera;
[0027] A depth information recognition unit is used to identify the depth information of each event using a depth model; the depth model is a model trained with events labeled with depth information, and the depth information label is the depth information of the corresponding training event extracted based on radar point cloud.
[0028] A sharp image reconstruction unit is used to solve an optimization model to reconstruct a sharp image at the current moment. The objective of the optimization model is to make the event image and the downsampled result of the sharp image convolve and add image noise to approximate the RBG image. The event image is an image obtained by double integration of a series of events output during the current exposure. The downsampling is to reduce the resolution of the sharp image to the resolution of the event camera.
[0029] The analysis unit is used to perceive targets in the current scene based on the reconstructed clear image and depth information of events.
[0030] Optionally, it also includes:
[0031] A training unit is used to train the deep model before using the deep model to identify the depth information of each event.
[0032] The present invention also provides an intelligent driving device, which includes a driving body and a frame camera, an event camera, and a dynamic scene perception system as described above, mounted on the driving body.
[0033] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any of the preceding claims.
[0034] The present invention also provides a computer program product, including a computer program or instructions, wherein the computer program or instructions, when executed by a processor, implement the steps of the method described in any of the preceding claims.
[0035] In summary, compared with the prior art, the technical solutions conceived in this invention have the following main advantages:
[0036] 1. The dynamic scene perception method proposed in this invention, on the one hand, combines radar and event cameras to extract depth information of event streams from radar point clouds, obtaining event streams with depth information labels, which are then used as a training set to train a depth model. Since both event streams and point clouds have a "point cloud" format, they can be fused effectively. The depth model trained using this method can accurately identify the depth information of events. In application, there is no need to analyze radar point cloud data again; simply inputting the event stream into the depth model can quickly and accurately obtain the depth information of each event, adapting to rapidly changing, highly dynamic scenes. On the other hand, this invention combines event cameras and frame cameras. Because event cameras possess high-speed, high-dynamic characteristics lacking in both LiDAR and frame cameras, they have strong dynamic processing potential. Given the sparse and fast nature of event streams, the event images themselves possess complete contour information. Fusing the event images with RGB images can obtain RGB detailed texture information. The combination of these two is flexible and sensitive, enabling clear imaging in high-speed, highly dynamic scenes. After obtaining clear images and depth information, targets in the scene can be accurately identified. Overall, the dynamic scene perception method proposed in this invention can accurately identify targets in high-speed, highly dynamic scenes, such as obstacles.
[0037] 2. Traditional pre-fusion of LiDAR and frame camera is limited by the lack of a unified 3D perception model of point cloud and image; post-fusion is caught in a dilemma because the AI system does not have good interpretability and the effective fusion distance of LiDAR is relatively short. This invention uses only LiDAR for depth supervision, avoiding the bottleneck effect of LiDAR in multi-source fusion system.
[0038] 3. The modal differences between 3D point cloud data and RGB image frame data are significant, and directly fusing the two requires a huge leap. However, both 3D point cloud and event stream have the form of "point cloud", and both RGB image and event stream are two-dimensional. This invention uses event stream as an intermediate modality between 3D point cloud and RGB image, filling the gap between the two.
[0039] 4. Based on the current mainstream sensing solutions, this invention enables the event camera to have more accurate depth information; relying on the high speed and high dynamic range of the event camera, this invention has faster sensing speed and better adaptability to complex external conditions compared with existing sensing devices (such as RGB frame cameras and LiDAR) (such as large dynamic range, more sensitive to low light, and robust to rainy environments). Attached Figure Description
[0040] Figure 1 This is a flowchart of the steps of a dynamic scene perception method in one embodiment of the present invention;
[0041] Figure 2This is a flowchart of the steps for training a deep model in one embodiment of the present invention. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0043] Example 1
[0044] This invention provides a dynamic scene perception method, such as Figure 1 The diagram shown is a flowchart of the dynamic scene perception method according to an embodiment of the present invention. The order of the steps is merely an example and is not intended to be limiting. The steps of the method are described in detail below.
[0045] Step S1: Obtain the RBG image exposed by the frame camera at the current moment, and a series of events output by the event camera during the current exposure of the frame camera.
[0046] Specifically, intelligent driving devices will be equipped with frame cameras and event cameras. Frame cameras can image the entire scene simultaneously, but their frame rate is limited, resulting in poor imaging performance in high-speed, high-dynamic scenes. Event cameras, on the other hand, only sense dynamically changing objects, offering advantages such as low latency and high sensitivity, but they cannot image the entire scene simultaneously.
[0047] In this step, RGB images captured by the frame camera and event streams output by the event camera are continuously acquired. Subsequent steps fuse the data from the frame camera and the event camera, and simultaneously identify the depth information of the event, enabling accurate target identification in high-speed, high-dynamic scenes.
[0048] Step S2: Use a depth model to identify the depth information of each event; the depth model is a model trained with events labeled with depth information, and the depth information labels are the depth information of the corresponding training events extracted based on radar point clouds.
[0049] Traditional methods typically fuse 3D point cloud data and RGB images to obtain scene images with depth information for identifying obstacles and other objects within the scene. However, due to the significant modal differences between 3D point cloud data and RGB image frame data, tight coupling and deep combination are not feasible. The information obtained from their fusion has considerable uncertainty, thus affecting recognition accuracy.
[0050] In this invention, event streams from event cameras and radar point clouds are fused. Since both event streams and point clouds are in the form of "point clouds," they can be fused effectively. Depth information for each event in the event stream during the same period can be extracted from the point cloud. This extracted depth information is used as depth information labels for the events. By training a depth model with these labeled events, a depth model capable of accurately identifying event depth information can be obtained. In application, there is no need to analyze radar point cloud data; simply inputting the event stream into the depth model allows for rapid acquisition of depth information for each event, adapting to rapidly changing, highly dynamic scenarios.
[0051] In one embodiment, the method further includes training a deep model, such as [example missing], before using the deep model to identify the depth information of each event. Figure 2 The diagram shown is a flowchart of the steps for training a deep model according to an embodiment of the present invention. The training steps are described in detail below.
[0052] Step S21: Acquire the LiDAR point cloud and a series of training events output by the event camera during the acquisition of the LiDAR point cloud.
[0053] Step S22: Extract the target point cloud data from each frame of the lidar point cloud.
[0054] Specifically, with the goal of identifying obstacles in the scene, the k-th frame of the original LiDAR point cloud is denoted as P. k ={p i} k Let i = 1, 2, 3, ..., n, where i is the i-th point cloud data point. Identify non-ground static point clouds within the point cloud. and dynamic point clouds Where r is the index of non-ground static obstacles (such as plants, buildings, etc.) and m is the index of dynamic obstacles (such as pedestrians, vehicles, etc.).
[0055] Step S23: Align the laser point cloud data of multiple consecutive frames of the same target to the current frame using the point cloud registration algorithm and then superimpose them to obtain the high-density point cloud data of the corresponding target in the current frame.
[0056] Specifically, the ICP method can be used to estimate the relative attitude drift between consecutive frames. Then, the static point clouds of frames t to t+m are registered with the static point cloud of frame t+m+1, and then superimposed with the static point cloud of frame t+m+1 to obtain a high-density static point cloud. Similarly, the dynamic point cloud of frames t to t+m is registered with the dynamic point cloud of frame t+m+1, and then superimposed with the dynamic point cloud of frame t+m+1 to obtain a high-density dynamic point cloud.
[0057] Step S24: Combine the high-density point cloud data of all targets in the current frame to obtain the high-density point cloud of the current frame.
[0058] Specifically, the high-density static point cloud and high-density dynamic point cloud of each object are stitched together to obtain the high-density off-ground lidar point cloud of each frame, denoted as . gt P k .
[0059] Step S25: Crop the high-density point cloud of the current frame so that it is within the field of view of the event camera.
[0060] Specifically, each is based on the external calibration parameters. gt P k Convert to the event camera coordinate system, and then truncate according to the event camera's detection range and field of view (FOV).
[0061] Step S26: Convert the clipped high-density point cloud into a three-dimensional placeholder mesh.
[0062] Step S27: Project the 3D placeholder mesh onto the shooting surface of the event camera, remove voxels whose intensity does not reach the predetermined threshold, and use the remaining voxels as the 3D placeholder ground truth of the current frame.
[0063] Specifically, the point cloud is then voxelized into a 3D placeholder mesh and projected onto the FV plane (front view of the event camera) to match the data format of the event camera tensor. Finally, voxels are filtered based on their intensity in the event camera tensor, retaining only those exceeding a predetermined threshold, which are ultimately used as the 3D placeholder ground truth. gt V k The empirical thresholds should be adjusted according to the actual situation. For the one-sided threshold, it can be set to 1 if it is greater than a certain value, and 0 otherwise. This is because there are far more dense point clouds than event points. As long as there is at least a one-to-one mapping, it is sufficient. This threshold can be set a little larger to establish a one-to-one mapping based on the closest relationship with the event point.
[0064] Step S28: Combine the 3D placeholder ground truth of the current frame to obtain the depth information of each training event in the current frame; use the training events as input and the corresponding depth information as labels to train the depth model.
[0065] Specifically, the depth information of each training event in the current frame is obtained by combining the ground truth of the 3D placeholder in the current frame, resulting in an event stream with depth information labels; a depth-supervised network is constructed, and the network is trained using the event stream with depth information labels, so that it learns the relationship between events and depth information.
[0066] The loss functions include Dice loss and Focal loss.
[0067] The loss function is designed as follows:
[0068]
[0069] Where, λ F It is a balance L Dice and L Focal The empirical parameters for the two types of loss are i, which is the index of the upsampled feature layer. In this embodiment, the deep model uses the Unet deep model, which has a total of 3 upsampled feature layers (excluding the last output layer), and their indices correspond to 0, 1, and 2 respectively.
[0070] Step S3: Solve the optimization model to reconstruct the sharp image at the current moment. The goal of the optimization model is to make the convolution of the downsampled result of the event image and the sharp image and the addition of image noise approximate the RGB image. The event image is the image obtained by double integration of a series of events output during the current exposure period. Downsampling is to reduce the resolution of the sharp image to the resolution of the event camera.
[0071] Specifically, the clear image at time t is defined as I. t Clear image I t The object to be reconstructed is defined as follows: First, a double integral is performed on the event stream output by the event camera during the exposure period of the frame camera at time t to obtain the event image E. t The clear image to be reconstructed t Downsampling is performed to reduce its resolution to that of the event camera, resulting in a downsampled image PI. t P represents the downsampling operation, and then E t and PI t Perform convolution operation to make E t ⊙PI t +ε approaches the RGB image Y t , where ε represents Gaussian white noise.
[0072] In this step, the RBG image is a known blurred image. The optimization objective determines the relationship between the blurred image, the event stream, and the instantaneous sharp image. Since the blurred image and the event stream are known, the instantaneous sharp image can be solved, that is, the feature fusion of the event stream and the RBG image is realized, and a continuous sharp image with less influence from motion blur is reconstructed.
[0073] Because event cameras possess high speed and high dynamic range characteristics that are lacking in both LiDAR and frame cameras, they have strong dynamic processing potential. Under the sparse and fast characteristics of event streams, the event images themselves have complete contour information, allowing the event images to be fused with RGB images to obtain RGB detailed texture information. The combination of the two is flexible and sensitive, enabling clear imaging in high-speed and high-dynamic scenes.
[0074] Step S4: Based on the reconstructed clear image and the depth information of the event, perceive the target in the current scene.
[0075] Specifically, step S2 obtains the depth information of the event stream, and step S3 fuses the event stream and the RGB image to obtain a clear image. After obtaining the clear image and depth information, post-processing, such as BEV and NERF processing, is performed to accurately identify obstacles in the scene. In practice, the clear image and the depth information of the events are mapped using a hash table.
[0076] The dynamic scene perception method proposed in this invention, on the one hand, combines radar and event cameras to extract depth information of event streams from radar point clouds, obtaining event streams with depth information labels, which are then used as a training set to train a depth model. Since both event streams and point clouds have a "point cloud" format, they can be well fused. The depth model trained using this method can accurately identify the depth information of events. In application, there is no need to analyze radar point cloud data again; simply inputting the event stream into the depth model allows for rapid and accurate acquisition of the depth information of each event, adapting to rapidly changing, highly dynamic scenes. On the other hand, this invention combines event cameras and frame cameras. Event cameras possess high-speed, high-dynamic characteristics lacking in both LiDAR and frame cameras, making them highly capable of dynamic processing. Given the sparse and fast nature of event streams, the event images themselves possess complete contour information. Fusing the event images with RGB images allows for the acquisition of RGB detail texture information. The combination of these two methods is flexible and sensitive, enabling clear imaging even in high-speed, highly dynamic scenes. After obtaining clear images and depth information, targets in the scene can be accurately identified. Overall, the dynamic scene perception method proposed in this invention can accurately identify targets in high-speed, highly dynamic scenes, such as obstacles.
[0077] Example 2
[0078] The present invention also relates to a dynamic scene perception system, comprising:
[0079] The acquisition unit is used to obtain the RBG image obtained by the frame camera at the current exposure time, as well as a series of events output by the event camera during the current exposure of the frame camera;
[0080] The depth information recognition unit is used to identify the depth information of each event using a depth model. The depth model is a model trained with events labeled with depth information, and the depth information labels are the depth information of the corresponding training events extracted based on radar point clouds.
[0081] The sharp image reconstruction unit is used to solve the optimization model to reconstruct the sharp image at the current moment. The goal of the optimization model is to make the downsampled result of the event image and the sharp image approach the RGB image after convolving and adding image noise. The event image is the image obtained by double integration of a series of events output during the current exposure. Downsampling is to reduce the resolution of the sharp image to the resolution of the event camera.
[0082] The analysis unit is used to perceive targets in the current scene based on the reconstructed clear image and depth information of events.
[0083] Furthermore, the perception system also includes a training unit, through which a deep model is trained in advance.
[0084] Specifically, the above sensing system can be used to implement the sensing method in Embodiment 1, and each unit therein can be used to implement the corresponding steps of the sensing method in Embodiment 1. For details, please refer to the above description, which will not be repeated here.
[0085] Example 3
[0086] This invention also relates to an intelligent driving device, such as an unmanned vehicle, which includes a driving body and a frame camera, an event camera, and the dynamic scene perception system described in Embodiment 2, all mounted on the driving body. This intelligent driving device can accurately perceive obstacles and targets in high-speed, high-dynamic scenes, improving the safety of intelligent driving.
[0087] Example 4
[0088] The present invention also relates to a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0089] Specifically, the memory may include high-speed random access memory, as well as non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.
[0090] Example 5
[0091] This invention provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of the method described in the above embodiments of this invention.
[0092] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification. It should be noted that the terms "in one embodiment," "for example," and "again" are intended to illustrate the present invention and are not intended to limit the present invention.
[0093] The above embodiments merely illustrate several implementation methods of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A dynamic scene perception method, characterized in that, include: Acquire the RGB image obtained by the frame camera at the current exposure time, and a series of events output by the event camera during the current exposure time of the frame camera; Use a depth model to identify the depth information of each event; The depth model is a model trained with events labeled with depth information, and the depth information labels are the depth information of the corresponding training events extracted based on radar point clouds. The optimization model is solved to reconstruct the sharp image at the current moment. The goal of the optimization model is to make the event image and the downsampled result of the sharp image convolve and add image noise to approximate the RBG image. The event image is the image obtained by double integration of a series of events output during the current exposure. The downsampling is to reduce the resolution of the sharp image to the resolution of the event camera. Based on the reconstructed clear images and depth information of events, the system can perceive targets in the current scene.
2. The dynamic scene perception method as described in claim 1, characterized in that, The method further includes training the deep model before using it to identify the depth information of each event. The training steps include: Acquire LiDAR point clouds and a series of training events output by the event camera during the acquisition of LiDAR point clouds; Extract target point cloud data from each frame of LiDAR point cloud; The laser point cloud data of the same target in multiple consecutive frames are aligned to the current frame and then superimposed by the point cloud registration algorithm to obtain the high-density point cloud data of the corresponding target in the current frame. By combining the high-density point cloud data of all targets in the current frame, the high-density point cloud of the current frame is obtained; The high-density point cloud of the current frame is cropped to fit within the field of view of the event camera; The clipped high-density point cloud is voxelized into a 3D placeholder mesh; The 3D placeholder mesh is projected onto the shooting surface of the event camera, and voxels whose intensity does not reach the predetermined threshold are removed. The remaining voxels are used as the 3D placeholder ground truth of the current frame. The depth information of each training event in the current frame is obtained by combining the 3D placeholder ground truth of the current frame; the depth model is trained by taking the training events as input and the corresponding depth information as labels.
3. The dynamic scene perception method as described in claim 2, characterized in that, The loss functions used to train the deep model include Dice loss and Focal loss.
4. The target perception method in dynamic scenes as described in any one of claims 1 to 3, characterized in that, The target is an obstacle that impedes movement, including dynamic obstacles and non-ground static obstacles.
5. The dynamic scene perception method as described in any one of claims 1 to 3, characterized in that, The reconstructed clear image and the depth information of the event are mapped using a hash table.
6. A dynamic scene perception system, characterized in that, include: The acquisition unit is used to obtain the RBG image obtained by the frame camera at the current exposure time, as well as a series of events output by the event camera during the current exposure of the frame camera; The depth information recognition unit is used to identify the depth information of each event using a depth model; The depth model is a model trained with events labeled with depth information, and the depth information labels are the depth information of the corresponding training events extracted based on radar point clouds. A sharp image reconstruction unit is used to solve an optimization model to reconstruct a sharp image at the current moment. The objective of the optimization model is to make the event image and the downsampled result of the sharp image convolve and add image noise to approximate the RBG image. The event image is an image obtained by double integration of a series of events output during the current exposure. The downsampling is to reduce the resolution of the sharp image to the resolution of the event camera. The analysis unit is used to perceive targets in the current scene based on the reconstructed clear image and depth information of events.
7. The dynamic scene perception system as described in claim 6, characterized in that, Also includes: A training unit is used to train the deep model before using the deep model to identify the depth information of each event.
8. An intelligent driving device, characterized in that, It includes a driving body and a frame camera, an event camera, and a dynamic scene perception system as described in claim 6 or 7, mounted on the driving body.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.
10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Depth estimation method based on laser radar and event camera fusion
CN114359744A
Motion detection method, device and equipment based on event camera and laser radar
CN115588042A