Depth measurement method and device based on active binocular event camera

CN119152008BActive Publication Date: 2026-08-11TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-17
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0008]本发明提供一种基于主动式双目事件相机的深度测量方法及装置,以解决基于深度学习方法虽能够从双目事件流获取稠密深度图,但是均是被动式双目视觉,在无纹理区域和暗场景中会导致深度图不准确,且计算复杂度高等问题

Benefits of technology

[0025] The depth measurement method and device based on an active binocular event camera proposed in this invention can effectively utilize the high temporal resolution provided by the event camera to perform real-time depth measurement of high-speed scenes. At the same time, active structured light has higher measurement accuracy than passive vision in low-light environments and weak texture scenes. It has broad application potential in fields involving high-speed depth measurement such as UAV system positioning and obstacle avoidance, 3D reconstruction, and virtual reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119152008B_ABST
    Figure CN119152008B_ABST
Patent Text Reader

Abstract

This invention relates to the field of 3D visual depth measurement technology, and particularly to a depth measurement method and apparatus based on an active binocular event camera. The method includes: using the binocular event camera to acquire active infrared structured light emitted towards a target 3D scene to generate two event stream data; adaptively dividing the two event stream data to obtain two discrete event blocks; using a preset binocular disparity matcher to perform feature matching on the two discrete event blocks to calculate two pixel-by-pixel disparity maps; and calculating a depth measurement value based on the two pixel-by-pixel disparity maps and a preset baseline of the binocular event camera. This solves the problems of inaccurate depth maps and high computational complexity in textureless areas and dark scenes, although deep learning-based methods can obtain dense depth maps from binocular event streams.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of three-dimensional visual depth measurement technology, and in particular to a depth measurement method and apparatus based on an active binocular event camera. Background Technology

[0002] 3D scene depth measurement utilizes different light wavelengths to measure the distance from the scene to the sensor. Currently, the main measurement methods include LiDAR, millimeter-wave radar, and visual sensors. Each sensor has different performance parameters and is suitable for scene depth measurement tasks with varying detection accuracy and range in different environments. LiDAR and millimeter-wave radar typically operate at depth measurement frequencies of 30-90Hz, but their lower frequencies are insufficient for high-speed obstacle avoidance tasks. Furthermore, the size and weight of these sensors are unsuitable for lightweight, agile drones. Using visual sensors for depth measurement is the most common method, offering advantages such as high-resolution real-scene recording, non-contact measurement, and minimal interference from magnetic fields or sensor interactions. Currently, 3D depth measurement is primarily based on the image frame paradigm, while passive binocular depth measurement methods suffer from low accuracy in textureless areas and dark environments. In contrast, active stereo imaging obtains more accurate depth maps by projecting infrared texture patterns. However, the traditional frame-based imaging paradigm limits the depth frame rate of traditional active stereo cameras (e.g., 30FPS for Kinect V2 and 33FPS for RealSense D435). Traditional active stereo cameras face significant challenges in depth perception in high-speed scenarios. For example, in the high-speed obstacle avoidance maneuvers of drones, serious collisions may occur within the short time between two adjacent depth frames.

[0003] Event cameras, or bio-inspired vision sensors, differ from traditional frame-based cameras in their operation. Instead of capturing frames at a finite, fixed rate, event cameras operate by detecting intensity changes in each pixel and providing asynchronous events with high temporal resolution (microseconds). The event stream is described as a spatiotemporally sparse stream of events, offering advantages such as high temporal resolution, high dynamic range, and low power consumption compared to traditional fixed-frame-rate cameras. Therefore, event cameras provide a viable solution for tasks such as autonomous driving vision sensors, UAV vision sensors, and robot visual navigation and localization, especially for high-speed depth measurement tasks in agile robots that require fast, low-latency processing of visual information.

[0004] Currently, most binocular depth measurements based on event cameras rely on passive stereo camera systems, which can lead to inaccurate depth maps in textureless areas and dark scenes. In fact, passive event stereo methods depend on feature matching, which can be difficult or impossible to perform in areas lacking texture or with low contrast. While some structured light systems with a single event camera attempt to achieve high-speed depth measurement, active monocular depth methods may not always match the accuracy of stereo depth systems in certain scenarios.

[0005] Therefore, how to draw on the neural networks and visual sampling and processing mechanisms of biological visual systems to propose a depth measurement method based on active binocular events and establish a novel binocular event structured light device for high-speed visual depth measurement has become a core problem that urgently needs to be solved in the field of 3D depth measurement.

[0006] Depth measurement methods based on binocular event cameras can be broadly categorized into traditional model optimization methods and deep learning methods. Traditional model optimization methods optimize the solution for matching disparity by finding local or global relationships between two event streams. While these methods can acquire sparse or semi-dense depth maps in real-time, they struggle to obtain globally dense depth maps. Deep learning methods can acquire dense depth maps from binocular event streams, but these methods rely on passive binocular vision, leading to inaccurate depth maps in textureless regions and dark scenes. Furthermore, these deep learning methods improve depth measurement accuracy by designing complex network structures, rarely focusing on reducing computational complexity. In other words, these deep learning algorithms require significant amounts of GPU memory, making them impractical for mobile robots or mobile devices. Therefore, designing lightweight binocular matching neural networks is a pressing issue for depth measurement based on binocular event cameras.

[0007] Increasingly, work is employing event cameras combined with active infrared structured light for high-speed depth measurement. Active infrared structured light sources are generally classified into three types: point, line, and 2D speckle patterns. For example, Brandli et al. first integrated a laser line projector and an event camera for 3D reconstruction. Martel et al. combined a laser source with an event-based stereo setup. Muglikar et al. designed a structured light system that included a laser point projector and an event camera for depth estimation. Huang et al. proposed a structured light system using an event camera and a laser pattern projector for high-speed 3D scanning. While these systems, using traditional model-based optimization methods, achieve high-speed depth measurement, they generate sparse depth maps with limited accuracy. Currently, there are also methods integrating deep learning models into event-based structured light systems to generate dense depth maps, but the computational complexity is too high. Therefore, there is a pressing need for an active event-based stereo structured light system capable of generating dense depth maps using a lightweight depth matching model for depth measurement. Summary of the Invention

[0008] This invention provides a depth measurement method and apparatus based on an active binocular event camera to solve the problems that although deep learning-based methods can obtain dense depth maps from binocular event streams, they are all passive binocular vision, which leads to inaccurate depth maps in textureless areas and dark scenes, and also has high computational complexity.

[0009] A first aspect of the present invention provides a depth measurement method based on an active binocular event camera, comprising the following steps: using the binocular event camera to acquire active infrared structured light emitted toward a target 3D scene to generate two event stream data; adaptively dividing the two event stream data to obtain two discrete event blocks; using a preset binocular disparity matcher to perform feature matching on the two discrete event blocks to calculate two pixel-by-pixel disparity maps; and calculating a depth measurement value based on the two pixel-by-pixel disparity maps and a preset baseline of the binocular event camera.

[0010] Optionally, before acquiring the active infrared structured light emitted towards the target 3D scene using a binocular event camera, the method further includes:

[0011] The infrared spot pattern, infrared band, infrared emission frequency, and amplitude of the active infrared structured light are determined based on the target depth measurement task.

[0012] Based on the infrared light spot pattern, the infrared light band, the infrared light emission frequency and amplitude, the active infrared structured light is emitted into the target three-dimensional scene using an active infrared structured light emitter.

[0013] Optionally, the infrared spot pattern includes dotted active infrared structured light, linear active infrared structured light, and lattice active infrared structured light.

[0014] Optionally, the step of using a binocular event camera to acquire active infrared structured light emitted towards the target 3D scene to generate two event stream data includes:

[0015] The transformation matrix of the binocular event camera is calibrated using a checkerboard pattern to obtain the baseline-corrected binocular event camera.

[0016] The baseline-corrected binocular event camera is triggered synchronously by a hardware clock to acquire active infrared structured light emitted toward the target 3D scene, thereby generating time-synchronized two-channel event stream data.

[0017] Optionally, the two event streams are either a single-modal event stream or a hybrid multimodal visual stream, wherein,

[0018] The single-modal event stream consists of the infrared and visible light intensity changes of the target 3D scene perceived by the binocular event camera, and the light intensity change pixels are recorded asynchronously as events.

[0019] The hybrid multimodal visual stream consists of infrared structured light images and light intensity changes of the target 3D scene simultaneously perceived and recorded by the binocular event camera.

[0020] Optionally, the preset binocular parallax matcher is any one of an optimized matching solver, an artificial neural network solver, and a spiking neural network solver.

[0021] A second aspect of the present invention provides a depth measurement device based on an active binocular event camera, comprising: an active infrared structured light emitter for emitting active infrared structured light into a target three-dimensional scene; a binocular event camera, including a left event camera and a right event camera, for acquiring the active infrared structured light to generate two event stream data; a partitioning module for adaptively partitioning the two event stream data to obtain two discrete event blocks; a binocular matching calculation module for performing feature matching on the two discrete event blocks to calculate two pixel-by-pixel disparity maps; and an output module for calculating depth measurement values ​​based on the two pixel-by-pixel disparity maps and a preset baseline of the binocular event camera.

[0022] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the depth measurement method based on an active binocular event camera as described in the above embodiments.

[0023] A fourth aspect of the present invention provides a computer program product, which, when executed by a processor, implements the depth measurement method based on an active binocular event camera as described above.

[0024] A fifth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the depth measurement method based on an active binocular event camera as described above.

[0025] The depth measurement method and device based on an active binocular event camera proposed in this invention can effectively utilize the high temporal resolution provided by the event camera to perform real-time depth measurement of high-speed scenes. At the same time, active structured light has higher measurement accuracy than passive vision in low-light environments and weak texture scenes. It has broad application potential in fields involving high-speed depth measurement such as UAV system positioning and obstacle avoidance, 3D reconstruction, and virtual reality.

[0026] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0027] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0028] Figure 1 A flowchart illustrating a depth measurement method based on an active binocular event camera, provided as an embodiment of the present invention;

[0029] Figure 2 This is a schematic diagram of the structure of the binocular event camera acquiring active infrared structured light emitted toward the target three-dimensional scene, as provided in an embodiment of the present invention.

[0030] Figure 3 This is a schematic diagram of a binocular event structured light prototype device for high-speed visual depth measurement provided in an embodiment of the present invention.

[0031] Figure 4 This is a schematic diagram illustrating the principle of the depth measurement method for an active binocular event camera provided in an embodiment of the present invention.

[0032] Figure 5 A schematic diagram of a lightweight event stereo matching neural network provided in an embodiment of the present invention;

[0033] Figure 6 This is a performance diagram of each test sequence of Active-Event-Stereo provided in the embodiments of the present invention;

[0034] Figure 7This is a schematic diagram comparing the performance of various binocular stereo matching algorithms provided in the embodiments of the present invention;

[0035] Figure 8 This is a schematic diagram illustrating the visualization results of each test sequence of Active-Event-Stereo provided in an embodiment of the present invention;

[0036] Figure 9 This is a visual comparison diagram of the results of the embodiment of the present invention with other methods;

[0037] Figure 10 This is a visualization result of an embodiment of the present invention and a Realsense D455 depth map;

[0038] Figure 11 A block diagram illustrating a depth measurement device based on an active binocular event camera, provided in an embodiment of the present invention.

[0039] Figure 12 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0040] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0041] The following description, with reference to the accompanying drawings, illustrates a depth measurement method and apparatus based on an active binocular event camera according to embodiments of the present invention.

[0042] Figure 1 This is a schematic flowchart of a depth measurement method based on an active binocular event camera, provided as an embodiment of the present invention.

[0043] like Figure 1 As shown, the depth measurement method based on an active binocular event camera includes the following steps:

[0044] In step S101, a binocular event camera is used to acquire active infrared structured light emitted toward the target 3D scene to generate two event stream data.

[0045] In some embodiments, before acquiring the active infrared structured light emitted toward the target 3D scene using a binocular event camera, the method further includes:

[0046] The infrared spot pattern, infrared band, infrared emission frequency, and amplitude of the active infrared structured light are determined based on the target depth measurement task.

[0047] Based on the infrared spot pattern, infrared light band, infrared light emission frequency and amplitude, an active infrared structured light emitter is used to emit active infrared structured light into the target three-dimensional scene.

[0048] In some embodiments, a binocular event camera is used to acquire active infrared structured light emitted toward a target 3D scene to generate two event stream data, including:

[0049] The transformation matrix of the binocular event camera is calibrated using a checkerboard pattern to obtain the baseline-corrected binocular event camera.

[0050] The baseline-corrected binocular event camera is triggered synchronously by a hardware clock to acquire active infrared structured light emitted toward the target 3D scene, thereby generating time-synchronized two-channel event stream data.

[0051] In actual implementation, such as Figure 2 As shown, in this embodiment of the invention, a binocular event camera consisting of two event cameras capable of sensing infrared structured light effects and an active infrared structured light emitter are pre-selected as the active binocular event acquisition device. Each event camera typically has a spectral sensing range of 300nm-1000nm in its quantum effect curve, meaning it can sense both visible and infrared light. The active infrared structured light spectrum is generally around 850nm (similar to the infrared projector in the Intel RealSense series cameras). Active infrared structured light sources are typically divided into three types (i.e., point, line, and 2D speckle patterns), with the infrared light band, emission frequency, and amplitude selected based on the actual depth measurement task. Therefore, in principle, the active infrared structured light emitted by the active infrared structured light emitter to the target 3D scene includes, but is not limited to: infrared spot mode, which can be selected according to the range of the scene depth measurement area to emit point, line, or dot matrix active infrared structured light; infrared light band, which is selected according to the working environment of scene depth measurement; and infrared light emission frequency and amplitude, which are set according to the output frequency of scene depth measurement to meet the needs of dynamic light intensity changes of the event camera.

[0052] Furthermore, after determining the infrared spot pattern, infrared band, infrared emission frequency, and amplitude of the active infrared structured light according to the target depth measurement task, the transformation matrix of the binocular event camera is calibrated using a checkerboard pattern to obtain the baseline-corrected binocular event camera.

[0053] Furthermore, active infrared structured light with a predetermined infrared spot pattern, infrared band, infrared emission frequency, and amplitude is emitted into the target 3D scene. The baseline-corrected binocular event camera is synchronously triggered by a hardware clock to collect the active infrared structured light and generate two event stream data.

[0054] For example, such as Figure 3 As shown, the active binocular event acquisition device can be mainly achieved by integrating two DAVIS346 event cameras and an infrared speckle projector. This acquisition device is only a prototype system for active binocular event data acquisition and measurement. The system can be further integrated and miniaturized. This active binocular event acquisition device can sense static scenes and generate dynamic events by adjusting the frequency or intensity of the laser signal generator. Furthermore, to enhance the depth measurement capability of the active binocular event acquisition device in high-speed motion scenes, the current flagship binocular depth camera, RealSense D455, can also be used. To allow for comparative analysis of different cameras from a unified perspective, the different sensors need to be spatiotemporally synchronized. For time calibration, the active binocular event acquisition device of this embodiment can use a ROS system to synchronize the time-stamped event camera and the RealSense D455 sensor. For spatial calibration, the active binocular event acquisition device of this embodiment can use a standard checkerboard pattern for baseline correction of the DAVIS346 camera and perform spatial registration between the left-side event camera and the RealSense D455 sensor.

[0055] In step S102, the two event stream data are adaptively divided to obtain two discrete event blocks.

[0056] In actual execution, the continuous event stream in the two event stream data is adaptively divided according to the speed of scene movement to obtain two discrete event blocks. The two event stream data are either single-modal event streams or hybrid multimodal visual streams. The single-modal event stream is the light intensity change pixels of the target 3D scene perceived by the binocular event camera in the infrared and visible light bands, recorded as asynchronous events. The hybrid multimodal visual stream is the infrared structured light image and light intensity change of the target 3D scene perceived and recorded simultaneously by the binocular event camera.

[0057] In step S103, a preset binocular disparity matcher is used to perform feature matching on the two discrete event blocks to calculate the two pixel-by-pixel disparity maps.

[0058] The preset binocular disparity matcher is any one of the following: an optimized matching solver, an artificial neural network solver, and a spiking neural network solver.

[0059] In practical implementation, any one of the following solvers—an optimized matching solver, an artificial neural network solver, or a spiking neural network solver—can be selected to perform feature matching on the two discrete event blocks, thereby calculating the disparity maps for each pixel. Specifically, when the traditional model optimized matching solver is selected, the disparity is optimized by representing the features of the two event camera blocks and then using classic algorithms such as local dense matching and global dense matching. When the artificial neural network solver is selected, an end-to-end deep artificial neural network is designed, and the disparity is learned through inference driven by the construction of binocular structured light disparity estimation data. When the spiking neural network solver is selected, a bio-inspired low-power spiking neural network is designed, and the disparity is learned through inference by mining the spatiotemporal relationship between the two event data.

[0060] For example, such as Figure 4-5 As shown, a lightweight event stereo matching neural network (i.e., a pre-defined binocular disparity matcher) is designed, mainly comprising four modules: an event representation module, a lightweight feature extraction module, a dynamic interactive construction module for calculating the matching cost volume, and an encoder-decoder output module. The lightweight feature extraction module and the dynamic interactive construction module for calculating the matching cost volume are the main modules. Specifically, two event streams are acquired and adaptively divided into two event streams according to the measurement task requirements, resulting in discrete event bins. Event features are represented for each event bin, and an event tensor is output. A lightweight MobileNet is designed as the backbone network of the event stereo matching neural network, providing compact and powerful features. A channel interaction module is designed to calculate the matching cost of the two feature streams, i.e., the cost of matching corresponding pixels between two event streams at slightly different viewpoints. Finally, the constructed matching cost is output to the encoder-decoder output module to predict a dense depth map.

[0061] Among them, the feature representation module: the asynchronous spatiotemporal event number presents a sparse discrete point array in three-dimensional space in the temporal and spatial domains, which is different from the structured image frame, so it cannot be directly compatible with the existing mainstream deep learning binocular matchers. Existing neuromorphic vision algorithms generally need to take two operations: (1) divide the continuous event stream into discrete pulse blocks in the temporal domain according to a fixed temporal window length or a fixed number of events, such as a fixed temporal length of 30 milliseconds, which can be equivalent to an image sequence of about 33 Hz; (2) design an asynchronous event signal representation method to encode the asynchronous event signal in the discrete pulse block into a feature that can be compatible with the existing mainstream deep learning algorithms. The asynchronous event signal representation can be formally defined as:

[0062]

[0063] In the formula, S L (x n ,yn ,t n The ) indicates the asynchronous event stream output by the event camera on the left. It is the event vector. k(x,y,t) is the kernel function, which can be a hand-designed function or a representation learned by an artificial neural network or a spiking neural network.

[0064] In summary, event representation is designed to make asynchronous events compatible with existing image frame paradigms in binocular matching algorithms. Current mainstream event representation methods include event images, voxel grids, and event embedding.

[0065] Lightweight Feature Extraction Module: To ensure a balance between accuracy and speed in depth measurement, this invention employs the lightweight MobileNet as the backbone of the binocular event matching network. Currently, MobileNet v1, MobileNet v2, and MobileNet v3 are the mainstream lightweight modules, and none of them use traditional convolutional operations. For example, MobileNet v1 uses depthwise separable convolution followed by pointwise convolution to achieve the spatial dimensions of standard convolution while reducing computation. Compared to standard convolution operations, depthwise separable convolution in MobileNet v1 achieves an 8 to 9-fold reduction in computation by effectively utilizing 3x3 kernels. To further improve the accuracy of MobileNet v1 while maintaining considerable computational complexity, MobileNet v2 introduces a linear bottleneck and inverse residuals. The entire structure of MobileNet v3 uses NAS neural network search, rather than manual prior design, resulting in limited performance improvement in downstream tasks. Therefore, this invention uses lightweight MobileNet v1 or MobileNet v2 operations to replace standard 2D convolutions in the feature extraction module and encoder-decoder module of the lightweight binocular event matching network.

[0066] The matching cost volume module: The matching cost volume represents the matching cost between pixels under different disparities and is a key step in the deep learning-based disparity calculation process. Typically, the matching volume is a 3D or 4D tensor constructed based on the similarity between the left and right side features. For example, the matching cost volume of a 3D tensor can be formally defined as:

[0067]

[0068] In the formula, d is the disparity of binocular stereo matching. It is a similarity measurement function. In the era of deep learning, This is typically achieved through operations such as aggregation or calculating correlations.L and F R These are the characteristics of the outputs of the left and right backbone networks, respectively.

[0069] Finally, a dynamic interaction strategy is designed to construct the matching cost volume, which is achieved by analyzing the two feature streams F. L and F R Dynamic interaction is performed during training, and then the two-way features are generated after the interaction. and The aggregation process can be formally described as follows:

[0070]

[0071]

[0072] In the formula, and These are the dynamic interaction features of the left and right visual streams, respectively, W. R and W L These are the weight parameters corresponding to the left and right visual streams, respectively. The dynamic interaction described above determines whether to perform feature interaction based on the channel threshold parameters after normalizing the two visual features.

[0073] The matching cost volume C generated above d The data is directly input into the encoder-decoder module, which predicts the disparity map of each pixel through regression. The disparity map of each pixel can have multiple depth map modes, which can be selected according to task requirements, including: dense depth map, which obtains the depth information of each pixel in the scene, and can meet the high-precision dense measurement and reconstruction tasks of the scene; semi-dense depth map, which obtains the depth information of local pixels in the scene, and meets the trade-off between accuracy and inference speed for specific scenes; and sparse depth map, which obtains the depth information of a few variable pixels in the scene, and meets the tasks of asynchronous high-speed depth measurement.

[0074] In step S104, the depth measurement value is calculated based on the two-way pixel disparity maps and the preset baseline of the binocular event camera.

[0075] The following quantitative and qualitative evaluation of the depth measurement method based on an active binocular event camera proposed in the embodiments of the present invention demonstrates its beneficial effects.

[0076] Experimental Data Preparation: An Active-Event-Stereo stereo matching dataset was established. First, 119 infrared video sequences were acquired, taking into account factors such as velocity distribution, lighting conditions, and scene diversity. This dataset provides event streams for stereo pairs and 23.8k synchronized ground truth labels. Finally, the dataset was divided into 16,000 sequences for training, 3,800 for validation, and 400 for testing.

[0077] Experimental parameter settings: The maximum disparity was set to 192. All networks were trained for 50 epochs on an NVIDIA GeForce RTX 3090 GPU using the Adam optimizer with a learning rate of 0.001. For training loss, L1 loss was used to measure the absolute difference between the predicted map and the label, and gradient loss was used to improve the smoothness of the predicted map. Furthermore, the accuracy in the stereo matching task was evaluated using mean pixel error (EPE), root mean square error (RMSE), the percentage of pixel disparity errors less than 3 pixels and 0.05 of the ground truth (D1-all), and the percentage of pixel disparity errors greater than 1 pixel, 2 pixels, and 3 pixels. Model parameters (#Params) and runtime (ms) were used to evaluate computational complexity.

[0078] Quantitative experimental results:

[0079] 1) Performance of each sequence in the dataset. For example... Figure 6 As shown, the performance of each test sequence of Active-Event-Stereo is demonstrated. It can be seen that the average pixel error (EPE) of the embodiment of the present invention is 1.223, and the inference speed on the NVIDIA GeForce RTX 3090 GPU reaches 21.9ms.

[0080] 2) Performance comparison of different methods. For example... Figure 7 As shown, the comparison results between the embodiments of the present invention and various existing binocular stereo matching algorithms are presented. It can be seen that the embodiments of the present invention have strong advantages in disparity estimation accuracy and inference speed.

[0081] Qualitative visualization results:

[0082] 1) Visualization results of each sequence in the dataset. For example... Figure 8 As shown, the structured light solution of this invention can acquire high-quality dense parallax maps in both indoor and outdoor scenes, and performs excellently under varying lighting conditions. Particularly in dark environments, the active event stereo solution with structured light of this invention outperforms passive stereo solutions.

[0083] 2) Visual comparison results of each method. For example... Figure 9 As shown, SGM based on the traditional matching optimization model can only generate sparse disparity maps. The embodiments of the present invention can obtain high-quality dense disparity maps in indoor and outdoor scenes, and even under changing lighting conditions.

[0084] 3) Depth measurement results in high-speed scenes. To verify the solution for high-speed depth perception in this embodiment of the invention, the camera prototype of this embodiment was compared with a traditional active stereo camera (i.e., RealSense D455). Figure 10 As shown, the camera prototype of this invention can obtain high-quality disparity maps in high-speed scenes, while the RealSense D455 clearly performs poorly at 90 FPS. This is because the RGB images of the RealSense D455 may exhibit motion blur. Conversely, this invention achieves 150 FPS in disparity estimation at a resolution of 346×260.

[0085] The depth measurement method based on an active binocular event camera proposed in this embodiment of the invention can effectively utilize the high temporal resolution provided by the event camera to perform real-time depth measurement of high-speed scenes; at the same time, active structured light has higher measurement accuracy than passive vision in low-light environments and weak texture scenes; it has broad application potential in fields involving high-speed depth measurement such as UAV system positioning and obstacle avoidance, 3D reconstruction, and virtual reality.

[0086] Next, with reference to the accompanying drawings, a depth measurement device based on an active binocular event camera according to an embodiment of the present invention is described.

[0087] Figure 11 This is a block diagram of a depth measurement device based on an active binocular event camera according to an embodiment of the present invention.

[0088] like Figure 11 As shown, the depth measurement device 110 based on an active binocular event camera includes: an active infrared structured light emitter 1101, a binocular event camera 1102, a segmentation module 1103, a binocular matching calculation module 1104, and an output module 1105.

[0089] The active infrared structured light emitter 1101 emits active infrared structured light into the target 3D scene. The binocular event camera 1102, comprising a left and a right event camera, acquires the active infrared structured light to generate two event streams. The segmentation module 1103 adaptively segments the two event streams to obtain two discrete event blocks. The binocular matching calculation module 1104 performs feature matching on the two discrete event blocks to calculate two pixel-by-pixel disparity maps. The output module 1105 calculates depth measurements based on the two pixel-by-pixel disparity maps and a preset baseline from the binocular event camera.

[0090] It should be noted that the foregoing explanation of the depth measurement method embodiment based on an active binocular event camera also applies to the depth measurement device based on an active binocular event camera in this embodiment, and will not be repeated here.

[0091] The depth measurement device based on an active binocular event camera proposed in this embodiment of the invention can effectively utilize the high temporal resolution provided by the event camera to perform real-time depth measurement of high-speed scenes; at the same time, active structured light has higher measurement accuracy than passive vision in low-light environments and weak texture scenes; it has broad application potential in fields involving high-speed depth measurement such as UAV system positioning and obstacle avoidance, 3D reconstruction, and virtual reality.

[0092] Figure 12 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. The electronic device may include:

[0093] The memory 1201, the processor 1202, and the computer program stored on the memory 1201 and executable on the processor 1202.

[0094] When the processor 1202 executes the program, it implements the depth measurement method based on an active binocular event camera provided in the above embodiments.

[0095] Furthermore, electronic devices also include:

[0096] Communication interface 1203 is used for communication between memory 1201 and processor 1202.

[0097] The memory 1201 is used to store computer programs that can run on the processor 1202.

[0098] The memory 1201 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0099] If the memory 1201, processor 1202, and communication interface 1203 are implemented independently, then the communication interface 1203, memory 1201, and processor 1202 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 12 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0100] Optionally, in a specific implementation, if the memory 1201, processor 1202, and communication interface 1203 are integrated on a single chip, then the memory 1201, processor 1202, and communication interface 1203 can communicate with each other through an internal interface.

[0101] The processor 1202 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0102] This invention also provides a computer program product, which, when executed by a processor, implements the above-described depth measurement method based on an active binocular event camera.

[0103] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described depth measurement method based on an active binocular event camera.

[0104] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0105] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0106] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0107] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0108] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0109] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0110] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0111] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A depth measurement method based on an active binocular event camera, characterized in that, Includes the following steps: A binocular event camera is used to acquire active infrared structured light emitted toward the target 3D scene to generate two event stream data. The two event stream data are adaptively divided to obtain two discrete event blocks; The two discrete event blocks are feature-matched using a preset binocular disparity matcher to calculate two pixel-by-pixel disparity maps, specifically including: The preset binocular disparity matcher employs a lightweight event-based stereo matching neural network, comprising: an event representation module, a lightweight feature extraction module, a dynamic interactive module for constructing and calculating the matching cost volume, and an encoder-decoder output module. It acquires two event streams of data and adaptively divides them according to the measurement task requirements to obtain discrete event blocks. Each event block is then characterized by event features, and an event tensor is output. A lightweight MobileNet is used as the backbone network of the event-based stereo matching neural network to extract features. A channel interaction module is designed to calculate the matching cost of the two feature streams, matching the cost of corresponding pixels between two event streams from different viewpoints. Finally, the constructed matching cost is output to the encoder-decoder output module to predict a dense depth map. Depth measurements are calculated based on the two pixel-by-pixel disparity maps and the preset baseline of the binocular event camera.

2. The depth measurement method based on active binocular event camera according to claim 1, wherein, Before using a binocular event camera to acquire active infrared structured light emitted towards the target 3D scene, the process also includes: The infrared spot pattern, infrared band, infrared emission frequency, and amplitude of the active infrared structured light are determined based on the target depth measurement task. Based on the infrared light spot pattern, the infrared light band, the infrared light emission frequency and amplitude, the active infrared structured light is emitted into the target three-dimensional scene using an active infrared structured light emitter.

3. The depth measurement method based on active binocular event camera according to claim 2, wherein, The infrared light spot patterns include dotted active infrared structured light, linear active infrared structured light, and lattice active infrared structured light.

4. The depth measurement method based on active binocular event camera according to claim 1, wherein, The method of using a binocular event camera to acquire active infrared structured light emitted towards the target 3D scene to generate two event stream data includes: The transformation matrix of the binocular event camera is calibrated using a checkerboard pattern to obtain the baseline-corrected binocular event camera. The baseline-corrected binocular event camera is triggered synchronously by a hardware clock to acquire active infrared structured light emitted toward the target 3D scene, thereby generating time-synchronized two-channel event stream data.

5. The active binocular event camera based depth measurement method of claim 1, wherein, The two event streams are either single-modal event streams or hybrid multimodal visual streams, wherein... The single-modal event stream consists of the infrared and visible light intensity changes of the target 3D scene perceived by the binocular event camera, and the light intensity change pixels are recorded asynchronously as events. The hybrid multimodal visual stream consists of infrared structured light images and light intensity changes of the target 3D scene simultaneously perceived and recorded by the binocular event camera.

6. The depth measurement method based on an active binocular event camera according to claim 1, characterized in that, The preset binocular parallax matcher is any one of the following: an optimized matching solver, an artificial neural network solver, and a spiking neural network solver.

7. A depth measurement device based on an active binocular event camera, characterized in that, include: An active infrared structured light emitter is used to emit active infrared structured light into a target 3D scene. A binocular event camera, comprising a left event camera and a right event camera, is used to acquire the active infrared structured light to generate two event stream data; The partitioning module is used to adaptively partition the two event stream data to obtain two discrete event blocks; The binocular matching calculation module is used to perform feature matching on the two discrete event blocks to calculate the two pixel-by-pixel disparity maps, specifically including: A preset binocular disparity matcher is used to perform feature matching on the two discrete event blocks. The preset binocular disparity matcher employs a lightweight event stereo matching neural network, which includes: an event representation module, a lightweight feature extraction module, a dynamic interactive construction and calculation matching cost volume module, and an encoder-decoder output module. It collects two event stream data and adaptively divides the two event stream data according to the measurement task requirements to obtain discrete event blocks. Each event block is then characterized by event features and an event tensor is output. A lightweight MobileNet is used as the backbone network of the event stereo matching neural network to extract features. A channel interaction module is designed to calculate the matching cost of the two feature streams, matching the cost of corresponding pixels between two event streams from different viewpoints. Finally, the constructed matching cost is output to the encoder-decoder output module to predict a dense depth map. The output module is used to calculate depth measurement values ​​based on the two-way pixel disparity maps and the preset baseline of the binocular event camera.

8. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, the processor executing the program to implement the depth measurement method based on an active binocular event camera as described in any one of claims 1-6.

9. A computer program product, characterized in that, When executed by a processor, the computer program / instruction implements the depth measurement method based on an active binocular event camera as described in any one of claims 1-6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the depth measurement method based on an active binocular event camera as described in any one of claims 1-6.