Dynamic scene sensing method, sensing system and intelligent driving equipment
By combining frame cameras and event cameras, using depth models and optimization models, target recognition in high-speed and high-dynamic scenarios is achieved, solving the recognition difficulties of traditional technologies under harsh conditions, and improving recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510028385.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-08
AI Technical Summary
In high-speed and high-dynamic scenarios, traditional frame cameras and lidars find it difficult to accurately identify targets, especially in severe weather conditions, where there are problems such as underexposure, overexposure, motion distortion and perception omission.
The dynamic scene perception method is adopted, combining the frame camera and event camera, and the depth information of the event is recognized through the depth model, and the clear image is reconstructed through the optimization model to achieve accurate recognition of the goal.
In high-speed and high-dynamic scenarios, the targets in the scene can be accurately identified, which improves the recognition accuracy and system robustness, and adapts to complex external conditions.
Smart Images

Figure CN120070847A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to image analysis, and more specifically, relates to a dynamic scene perception method and perception system, and an intelligent driving device. Background Art
[0002] Target perception technology is an important link and safety guarantee in vehicle intelligent driving. Currently, commonly used perception methods include frame camera imaging, lidar detection, and a combination of the two.
[0003] Due to the limitations of physical mechanisms, the performance of traditional frame camera imaging is significantly reduced in severe weather conditions such as heavy rain, heavy snow, and dense fog, affecting the accuracy of object recognition and the safety of the vehicle; and under-exposure, over-exposure blind spots and high-speed blur problems will occur in high-dynamic scenes.
[0004] LiDAR has problems such as large size, high energy consumption, and high cost. In addition, due to the low frame rate of LiDAR, when dealing with fast-moving objects, especially motion detection and recognition of small targets, it is prone to motion distortion, sparse point clouds, and even perception omissions, so the recognition ability is limited. In addition, in severe weather conditions such as rain, snow, haze, and dust, the laser beam of LiDAR will be scattered, resulting in a significant deterioration in ranging accuracy.
[0005] The fusion solution of frame camera and LiDAR has objective fusion difficulties due to the difference in data attributes between RGB data and 3D point cloud, and there is currently a lack of a large 3D perception model that can unify point cloud and image. If the two are processed separately, the confidence of the two is completely different because the AI system does not have good interpretability and the effective fusion distance of LiDAR is short, which often leads to the problem of "not knowing who to trust". Tight coupling and deep combination cannot be achieved. In general target perception tasks, sensors bring in more information while also bringing more uncertainty.
[0006] Therefore, how to accurately identify targets in high-speed, high-dynamic scenarios is a technical problem that needs to be solved urgently. Summary of the invention
[0007] In view of the above defects or improvement needs of the prior art, the present invention provides a dynamic scene perception method and perception system, and an intelligent driving device, the purpose of which is to accurately identify targets in high-speed and high-dynamic scenes.
[0008] To achieve the above object, the present invention provides a dynamic scene perception method, which includes:
[0009] Get the RBG image obtained by the frame camera at the current moment of exposure, and a series of events output by the event camera during the current exposure period of the frame camera;
[0010] Identify the depth information of each event using a depth model; the depth model is a model trained with events labeled with depth information, and the depth information label is the depth information corresponding to the training event extracted from the radar point cloud.
[0011] Solve the optimization model to reconstruct the clear image at the current moment. The solution objective of the optimization model is to make the event image convolved with the downsampled result of the clear image and added with image noise approach the RGB image. The event image is an image obtained by double integrating a series of events output during the current exposure. The downsampling is to reduce the resolution of the clear image to the resolution of the event camera.
[0012] Perceive the targets in the current scene based on the reconstructed clear image and the depth information of the events.
[0013] Optionally, the method further includes training the depth model before identifying the depth information of each event using the depth model. The training steps include:
[0014] Obtain the lidar point cloud and a series of training events output by the event camera during the acquisition of the lidar point cloud.
[0015] Extract the target point cloud data from each frame of the lidar point cloud.
[0016] Align the continuous multi-frame lidar point cloud data of the same target to the current frame through a point cloud registration algorithm and then stack them to obtain the high-density point cloud data of the corresponding target in the current frame.
[0017] Integrate the high-density point cloud data of all targets in the current frame to obtain the high-density point cloud of the current frame.
[0018] Crop the high-density point cloud of the current frame so that it is within the field of view of the event camera.
[0019] Voxelize the cropped high-density point cloud into a three-dimensional occupancy grid.
[0020] Project the three-dimensional occupancy grid onto the shooting plane of the event camera, remove the voxels whose intensity does not reach the predetermined threshold, and the remaining voxels are used as the three-dimensional occupancy ground truth of the current frame.
[0021] Combine the three-dimensional occupancy ground truth of the current frame to obtain the depth information of each training event in the current frame; use the training event as the input and the corresponding depth information as the label to train the depth model.
[0022] Optionally, the loss function used to train the depth model includes Dice loss and Focal loss.
[0023] Optionally, the target is an obstacle hindering driving, and the obstacle includes a dynamic obstacle and a non-ground static obstacle.
[0024] Optionally, a mapping relationship is constructed between the reconstructed clear image and the depth information of the event in the form of a hash table.
[0025] The present invention also provides a dynamic scene perception system, which includes:
[0026] An acquisition unit for obtaining the RBG image obtained by the frame camera at the current moment of exposure, and a series of events output by the event camera during the current exposure of the frame camera;
[0027] A depth information recognition unit for recognizing the depth information of each event by using a depth model; the depth model is a model trained with events with depth information labels, and the depth information labels are the depth information of the corresponding training events extracted based on radar point clouds;
[0028] A clear image reconstruction unit for solving an optimization model to reconstruct the clear image at the current moment. The solution target of the optimization model is to make the event image convolved with the downsampled result of the clear image and added with image noise approach the RBG image. The event image is an image obtained by double integrating a series of events output during the current exposure, and the downsampling is to reduce the resolution of the clear image to the resolution of the event camera;
[0029] An analysis unit for perceiving the target in the current scene based on the reconstructed clear image and the depth information of the event.
[0030] Optionally, it further includes:
[0031] A training unit for training the depth model before recognizing the depth information of each event by using the depth model.
[0032] The present invention also provides an intelligent driving device, which includes a driving body and a frame camera, an event camera and the above-mentioned dynamic scene perception system loaded on the driving body.
[0033] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0034] The present invention also provides a computer program product, including a computer program or instruction. When the computer program or instruction is executed by a processor, the steps of the method described in any one of the above are implemented.
[0035] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the present invention mainly has the following beneficial effects:
[0036] 1. The dynamic scene perception method proposed by the present invention, on the one hand, combines a radar and an event camera to extract the depth information of the event stream from the radar point cloud, obtaining an event stream with depth information tags as a training set to train a depth model. Since both the event stream and the point cloud have the form of "point cloud", they can be better fused. The depth model trained by this method can accurately identify the depth information of events. When applied, there is no need to analyze the radar point cloud data anymore. Just input the event stream into the depth model, and the depth information of each event can be quickly and accurately obtained to adapt to fast and highly dynamic changing scenes. On the other hand, the present invention combines an event camera and a frame camera. Since the event camera has high-speed and high-dynamic characteristics that are lacking in both lidar and frame cameras, this gives it strong dynamic processing potential. Under the characteristics of sparse and fast event streams, the event image itself has complete contour information. By fusing the event image with the RGB image, RGB detail texture information can be obtained. The combination of the two is flexible and sensitive, and clear imaging can be obtained in high-speed and high-dynamic scenes. After obtaining a clear image and depth information, the targets in the scene can be accurately identified. Generally speaking, using the dynamic scene perception method proposed by the present invention, the targets in the scene can be accurately identified in fast and highly dynamic changing scenes, such as identifying obstacles, etc.
[0037] 2. The traditional front fusion of lidar and frame cameras is limited by the lack of a unified three-dimensional perception large model for point clouds and images; the back fusion is in a dilemma because the AI system does not have good interpretability and the effective fusion distance of lidar is short. The present invention only uses lidar as depth supervision, avoiding the short-board effect of lidar in a multi-source fusion system.
[0038] 3. The three-dimensional point cloud data and the RGB image frame data have large modal differences, and directly fusing the two requires a huge leap. However, both the three-dimensional point cloud and the event stream have the form of "point cloud", and both the RGB image and the event stream are in two-dimensional form. The present invention uses the event stream as an intermediate modality between the three-dimensional point cloud and the RGB image, filling the gap between the two.
[0039] 4. Based on the current mainstream perception solutions, the present invention enables the event camera to have more accurate depth information; relying on the high-speed and high-dynamic characteristics of the event camera, the present invention has faster sensing speed and better adaptability to complex external conditions (such as large dynamic range, more sensitive to low light, robust in rainy environments, etc.) compared with existing sensing devices (such as RGB frame cameras, lidar). BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flowchart of the steps of the dynamic scene perception method in an embodiment of the present invention;
[0041] Figure 2It is a flowchart of the steps for training a deep model in an embodiment of the present invention. Detailed implementation manners
[0042] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0043] Embodiment 1
[0044] The present invention provides a dynamic scene perception method, as Figure 1 shown is a flowchart of the steps of the dynamic scene perception method in an embodiment of the present invention, and the order of the steps is only an example and is not limited thereto. The steps of the method will be specifically introduced below.
[0045] Step S1, obtain the RBG image obtained by the frame camera at the current moment of exposure, and a series of events output by the event camera during the current exposure of the frame camera.
[0046] Specifically, a frame camera and an event camera are installed on the intelligent driving device. Among them, the frame camera can image the entire scene at the same time, but its frame rate is limited, and the imaging effect in high-speed and high-dynamic scenes is poor. The event camera, on the other hand, only senses dynamically changing objects and has the advantages of low latency and high sensitivity, but it cannot image the entire scene at the same time.
[0047] In this step, continuously obtain the RBG image captured by the frame camera and the event stream output by the event camera. In the subsequent steps, by fusing the data of the frame camera and the event camera and simultaneously identifying the depth information of the events, the target can be accurately identified in high-speed and high-dynamic scenes.
[0048] Step S2, use a deep model to identify the depth information of each event; the deep model is a model trained with events with depth information labels, and the depth information labels are the depth information of the corresponding training events extracted based on the radar point cloud.
[0049] In the traditional solution, usually the three-dimensional point cloud data and the RGB image are fused to obtain a scene image with depth information to identify obstacles in the scene, etc. However, due to the too large difference in the data modalities of the three-dimensional point cloud data and the RGB image frame data, it is impossible to achieve tight coupling and deep combination, and the information obtained by fusing the two has relatively large uncertainty, thus affecting the recognition accuracy.
[0050] In the present invention, the event stream of the event camera and the radar point cloud are fused. Since both the event stream and the point cloud have the form of "point cloud", they can be fused well. Depth information of each event in the contemporaneous event stream can be extracted from the point cloud. The extracted depth information is used as the depth information label of the event. Training a depth model with the events with depth information labels can obtain a depth model that can accurately identify the depth information of the events. In application, there is no need to analyze the radar point cloud data anymore. Only by inputting the event stream into the depth model, the depth information of each event can be quickly obtained to adapt to fast high-dynamic change scenarios.
[0051] In one embodiment, before using the depth model to identify the depth information of each event, the method further includes training the depth model, as Figure 2 shown in the flowchart of the steps of training the depth model in an embodiment of the present invention. The following will introduce the training steps in detail.
[0052] Step S21, obtain the lidar point cloud and a series of training events output by the event camera during the acquisition of the lidar point cloud.
[0053] Step S22, extract the target point cloud data in each frame of the lidar point cloud.
[0054] Specifically, aiming at identifying scene obstacles, the k-th frame of the original lidar point cloud is denoted as P k ={p i} k , i = 1, 2, 3,... n, where i is the i-th point cloud data in the point cloud. Identify the non-ground static point cloud and the dynamic point cloud in the point cloud. Among them, r is the index of non-ground static obstacles (such as plants, buildings, etc.), and m is the index of dynamic obstacles (such as pedestrians, vehicles, etc.).
[0055] Step S23, align and stack the continuous multi-frame lidar point cloud data of the same target to the current frame through a point cloud registration algorithm to obtain the high-density point cloud data of the corresponding target in the current frame.
[0056] Specifically, the ICP method can be used to estimate the relative pose drift between consecutive frames, and then register the static point cloud of the t-th to t+m-th frames with the static point cloud of the t+m+1-th frame and then stack it with the static point cloud of the t+m+1-th frame to obtain a high-density static point cloud Similarly, register the dynamic point cloud of the t-th to t+m-th frames with the dynamic point cloud of the t+m+1-th frame and then stack it with the dynamic point cloud of the t+m+1-th frame to obtain a high-density dynamic point cloud
[0057] Step S24: Integrate the high-density point cloud data of all targets in the current frame to obtain the high-density point cloud of the current frame.
[0058] Specifically, splice the high-density static point cloud and high-density dynamic point cloud of each object to obtain the high-density ground laser radar point cloud of each frame, denoted as gt P k 。
[0059] Step S25: Crop the high-density point cloud of the current frame so that it is within the field of view of the event camera.
[0060] Specifically, according to the external calibration parameters, each gt P k is converted to the event camera coordinate system, and then truncated according to the detection range and field of view (FOV) of the event camera.
[0061] Step S26: Voxelize the cropped high-density point cloud into a three-dimensional occupancy grid.
[0062] Step S27: Project the three-dimensional occupancy grid onto the imaging plane of the event camera, and remove the voxels whose intensity does not reach the predetermined threshold. The remaining voxels are used as the three-dimensional occupancy ground truth of the current frame.
[0063] Specifically, the point cloud is then voxelized into a three-dimensional occupancy grid and projected onto the FV plane (front view of the event camera) to match the data format of the event camera tensor. Finally, the voxels are filtered according to the intensity of the corresponding voxels in the event camera tensor, and only those voxels that exceed the predetermined threshold are retained, which are ultimately used as the three-dimensional occupancy ground truth gt V k 。Among them, the empirical threshold should be adjusted according to the actual situation. For the unilateral threshold, if it is greater than a certain value, it is set to 1, otherwise it is set to 0. Because the densified point cloud is much more than the event points, as long as at least one-to-one mapping is ensured, this threshold can be set appropriately larger to establish a one-to-one mapping according to the nearest relationship with the event points.
[0064] Step S28: Combine the three-dimensional occupancy ground truth of the current frame to obtain the depth information of each training event in the current frame; use the training event as the input and the corresponding depth information as the label to train the depth model.
[0065] Specifically, combine the three-dimensional occupancy ground truth of the current frame to obtain the depth information of each training event in the current frame, and obtain an event stream with depth information labels; construct a depth supervision network, and use the event stream with depth information labels to train the network so that it learns the association between events and depth information.
[0066] Among them, the loss function includes Dice loss and Focal loss.
[0067] The loss function is designed as:
[0068]
[0069] Among them, λ F is an empirical parameter for balancing the Dice losses of L Focal and L
[0070] Step S3, solve the optimization model to reconstruct the clear image at the current moment. The solution objective of the optimization model is to make the event image convolved with the downsampled result of the clear image and add image noise to approach the RGB image. The event image is the image obtained by double integrating a series of events output during the current exposure period, and the downsampling is to reduce the resolution of the clear image to the resolution of the event camera.
[0071] Specifically, define the clear image at time t as I t , and the clear image I t is the object to be reconstructed. First, double integrate the event stream output by the event camera during the exposure period of the frame camera at time t to obtain the event image E t . Downsample the clear image I t to be reconstructed to reduce its resolution to the resolution of the event camera, obtaining the downsampled image PI t , where P represents the downsampling operation. Then, perform a convolution operation on E t and PI t so that E t ⊙PI t +ε approaches the RGB image Y t , where ε represents Gaussian white noise.
[0072] In this step, the RGB image is the determined blurred image. The optimization objective determines the association between the blurred image, the event stream, and the instantaneous clear image. Since the blurred image and the event stream are known, the instantaneous clear image can be solved, that is, the feature fusion of the event stream and the RGB image is realized, and a continuous clear image less affected by motion blur is reconstructed.
[0073] Since the event camera has high-speed and high-dynamic characteristics that lidar and frame cameras both lack, this makes its dynamic processing potential very strong. Under the characteristics of sparse and fast event streams, the event image itself has complete contour information. By fusing the event image with the RGB image, RGB detail texture information can be obtained. The combination of the two is flexible and sensitive, and clear imaging can be obtained in high-speed and high-dynamic scenarios.
[0074] Step S4: Based on the reconstructed clear image and the depth information of the event, perceive the targets in the current scene.
[0075] Specifically, the depth information of the event stream can be obtained through Step S2. The event stream and the RGB image are fused through Step S3 to obtain a clear image. After obtaining the clear image and the depth information, post-processing is performed, such as BEV and NERF processing, so as to accurately identify the obstacles in the scene. In specific operations, a mapping relationship is constructed between the clear image and the depth information of the event in the form of a hash table.
[0076] For the dynamic scene perception method proposed by the present invention, on the one hand, by combining a radar and an event camera, the depth information of the event stream is extracted from the radar point cloud, and the event stream with depth information tags is obtained and used as a training set to train a depth model. Since both the event stream and the point cloud have the form of "point cloud", they can be better fused. The depth model trained by using this method can accurately identify the depth information of the event. In application, there is no need to analyze the radar point cloud data anymore. Only by inputting the event stream into the depth model, the depth information of each event can be quickly and accurately obtained to adapt to the fast and highly dynamic changing scene. On the other hand, the present invention combines an event camera and a frame camera. Since the event camera has the high-speed and high-dynamic characteristics that both the lidar and the frame camera lack, its dynamic processing potential is very strong. Under the characteristics of sparse and fast event streams, the event image itself has complete contour information. By fusing the event image with the RGB image, the RGB detail texture information can be obtained. The combination of the two is flexible and sensitive, and clear imaging can be obtained in a high-speed and high-dynamic scene. After obtaining the clear image and the depth information, the targets in the scene can be accurately identified. Generally speaking, by using the dynamic scene perception method proposed by the present invention, the targets in the scene can be accurately identified in a fast and highly dynamic changing scene, such as identifying obstacles, etc.
[0077] Embodiment 2
[0078] The present invention also relates to a dynamic scene perception system, including:
[0079] An acquisition unit, configured to obtain the RGB image obtained by the frame camera at the current moment of exposure, and a series of events output by the event camera during the current exposure of the frame camera;
[0080] A depth information recognition unit, configured to use a depth model to recognize the depth information of each event; the depth model is a model trained with events with depth information tags, and the depth information tags are the depth information of the corresponding training events extracted based on the radar point cloud;
[0081] A clear image reconstruction unit is configured to solve an optimization model to reconstruct a clear image at the current moment. The solution objective of the optimization model is to make the convolution of the event image and the downsampled result of the clear image and the addition of image noise approach the RGB image. The event image is an image obtained by double integrating a series of events output during the current exposure period. Downsampling is to reduce the resolution of the clear image to the resolution of the event camera.
[0082] An analysis unit is configured to perceive the targets in the current scene based on the reconstructed clear image and the depth information of the events.
[0083] Furthermore, the perception system further includes a training unit for pre-training a depth model through the training unit.
[0084] Specifically, the above perception system can be used to implement the perception method in Embodiment 1, and each unit therein can be used to implement the corresponding steps of the perception method in Embodiment 1. For details, reference can be made to the above introduction and will not be elaborated here.
[0085] Embodiment 3
[0086] The present invention also relates to an intelligent driving device, such as an autonomous vehicle, which includes a driving body and a frame camera, an event camera, and the dynamic scene perception system introduced in Embodiment 2 mounted on the driving body. The intelligent driving device can accurately perceive the obstacle targets in the scene in a high-speed and high-dynamic scene, improving the safety guarantee of intelligent driving.
[0087] Embodiment 4
[0088] The present invention also relates to a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0089] Specifically, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0090] Embodiment 5
[0091] The embodiment of the present invention provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the method in the above embodiments of the present invention.
[0092] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification. It should be noted that the phrases "in one embodiment of the present invention", "for example", "again, for example", etc. are intended to illustrate the present invention rather than to limit the present invention.
[0093] The above embodiments only express several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent application. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present invention, several variations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A dynamic scene perception method, characterized in that: include: Get the RBG image obtained by the frame camera at the current moment of exposure, and a series of events output by the event camera during the current exposure period of the frame camera; Use deep models to identify the depth information of each event; The depth model is a model trained with events with depth information labels, and the depth information labels are depth information of corresponding training events extracted based on radar point clouds; Solve the optimization model to reconstruct the clear image at the current moment, wherein the solution goal of the optimization model is to make the event image and the down-sampling result of the clear image convolve and add image noise to approach the RBG image, wherein the event image is an image obtained by double integration of a series of events output during the current exposure period, and the down-sampling is to reduce the resolution of the clear image to the resolution of the event camera; Based on the reconstructed clear image and the depth information of the event, the target in the current scene is perceived.
2. The dynamic scene perception method according to claim 1, characterized in that: The method further includes, before using the deep model to identify the depth information of each event, training the deep model, the training step comprising: Acquire a lidar point cloud and a series of training events output by an event camera during acquisition of the lidar point cloud; Extract target point cloud data from each frame of lidar point cloud; The continuous multi-frame laser point cloud data of the same target are aligned to the current frame through the point cloud registration algorithm and then superimposed to obtain the high-density point cloud data of the corresponding target in the current frame; The high-density point cloud data of all targets in the current frame are integrated to obtain the high-density point cloud of the current frame; Crop the high-density point cloud of the current frame to make it within the field of view of the event camera; Voxelize the cropped high-density point cloud into a three-dimensional placeholder grid; Project the 3D occupancy grid onto the shooting surface of the event camera, remove the voxels whose intensity does not reach the predetermined threshold, and use the remaining voxels as the true value of the 3D occupancy of the current frame; The depth information of each training event in the current frame is obtained by combining the three-dimensional occupancy true value of the current frame; the depth model is trained by taking the training event as input and the corresponding depth information as a label.
3. The dynamic scene perception method according to claim 2, characterized in that: The loss functions used to train the deep model include Dice loss and Focal loss.
4. The target perception method in a dynamic scene according to any one of claims 1 to 3, characterized in that: The target is an obstacle that hinders driving, and the obstacle includes a dynamic obstacle and a non-ground static obstacle.
5. The dynamic scene perception method according to any one of claims 1 to 3, characterized in that: The reconstructed clear image and the depth information of the event are mapped in the form of a hash table.
6. A dynamic scene perception system, characterized in that: include: An acquisition unit, used for the RBG image obtained by the frame camera at the current moment of exposure, and a series of events output by the event camera during the current exposure period of the frame camera; A depth information recognition unit, used to recognize depth information of each event using a depth model; The depth model is a model trained with events with depth information labels, and the depth information labels are depth information of corresponding training events extracted based on radar point clouds; A clear image reconstruction unit, used for solving an optimization model to reconstruct a clear image at a current moment, wherein the optimization model is solved to make the event image approach the RBG image after convolving the down-sampling result of the clear image and adding image noise, wherein the event image is an image obtained by double integration of a series of events output during the current exposure period, and the down-sampling is to reduce the resolution of the clear image to the resolution of the event camera; The analysis unit is used to perceive the target in the current scene based on the reconstructed clear image and the depth information of the event.
7. The dynamic scene perception system according to claim 6, characterized in that: Also includes: A training unit is used to train the depth model before using the depth model to identify the depth information of each event.
8. An intelligent driving device, characterized in that: It comprises a moving body, a frame camera and an event camera mounted on the moving body, and a dynamic scene perception system as claimed in claim 6 or 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Depth estimation method based on laser radar and event camera fusion
CN114359744A
Motion detection method, device and equipment based on event camera and laser radar
CN115588042A
Event camera guided high-dynamic high-speed scene imaging method and system
CN117830126A
Heterogeneous multi-scale event camera monocular depth estimation optimization method
CN118537384A
Image processing method, and related apparatus and system
WO2023185693A1