Binocular pulse depth data acquisition method and device for super-high-speed motion scene

By acquiring and processing binocular pulse signals using a neuromorphic vision sensor, and combining feature extraction and attention mechanism models, the difficulty of depth estimation in high-speed motion scenes using traditional cameras is solved, achieving clear depth estimation of both dynamic and static objects.

CN115422964BActive Publication Date: 2025-12-12PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210860369.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-21
Publication Date
2025-12-12
Estimated Expiration
2042-07-21

AI Technical Summary

Technical Problem

Traditional cameras struggle to perform effective depth estimation in high-speed motion scenes, especially under conditions of high-speed motion blur, overexposure, or low light, where their performance deteriorates sharply.

Method used

A neuromorphic vision sensor is used to acquire binocular pulse signals, which are accumulated into integral frames by a pulse encoder. The hourglass model is used for feature extraction, and the feature point matching relationship is predicted by combining self-attention and cross-attention mechanism models. Finally, the predicted depth map is obtained by disparity regression calculation.

Benefits of technology

It achieves clear depth estimation for both dynamic and static objects in high-speed motion scenes, solving the problem of depth estimation difficulties for traditional cameras under conditions of high-speed motion blur, overexposure, or low light.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115422964B_ABST
    Figure CN115422964B_ABST
Patent Text Reader

Abstract

The application relates to a binocular pulse depth data acquisition method and device for a super-high-speed motion scene. The method comprises the following steps: acquiring a binocular pulse signal in a motion scene; extracting features of the binocular pulse signal to obtain pulse signal dimension reduction features in the motion scene; estimating a binocular pulse feature point matching relationship in the motion scene according to the pulse signal dimension reduction features and an attention mechanism model; and performing parallax regression calculation on the binocular pulse feature point matching relationship to determine an estimated depth map in the motion scene. The application can effectively solve the problem that a traditional camera cannot perform depth estimation in a high-speed motion blur, overexposure or low-light motion scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, and more particularly, to a binocular pulse depth data acquisition method and device for a super-high-speed motion scene. BACKGROUND

[0002] Depth estimation is a long-standing and challenging problem that is used in a variety of applications such as autonomous driving, robotics, and medical diagnosis. In fact, traditional frame-based cameras have limitations in depth estimation in fast motion scenes, resulting in a sharp decline in performance when using blurred images. SUMMARY

[0003] The present application provides a binocular pulse depth data acquisition method and device for a super-high-speed motion scene. To have a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This part is not a general review, nor is it intended to determine the key / important elements or delineate the scope of protection of these embodiments. Its only purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.

[0004] In a first aspect, the present application provides a binocular pulse depth data acquisition method for a super-high-speed motion scene, which comprises:

[0005] Acquiring a binocular pulse signal in a motion scene;

[0006] Extracting features from the binocular pulse signal to obtain pulse signal dimensionality reduction features in the motion scene;

[0007] Estimating binocular pulse feature point matching relationships in the motion scene according to the pulse signal dimensionality reduction features and an attention mechanism model;

[0008] Calculating the disparity regression of the binocular pulse feature point matching relationships to determine the estimated depth map in the motion scene.

[0009] Optionally, the acquiring of the binocular pulse signal in the motion scene comprises:

[0010] Acquiring the binocular pulse signal in the motion scene through a neuromorphic vision sensor.

[0011] Optionally, the extracting of the features from the binocular pulse signal to obtain the pulse signal dimensionality reduction features in the motion scene comprises:

[0012] Accumulating the binocular pulse signal into an integral frame through a pulse encoder;

[0013] input the integral frame into the hourglass model, and output the pulse signal dimension reduction feature in the motion scene.

[0014] Optionally, the attention mechanism model comprises a self-attention mechanism and a cross-attention mechanism.

[0015] The self-attention mechanism calculates a self-feature attention matrix on the same pulse signal, and captures the internal correlation between the pulse signal and the pulse signal dimension reduction feature through the self-feature attention matrix.

[0016] The cross-attention mechanism calculates a cross-feature attention matrix on the binocular pulse signals, and captures the feature correlation between the pixel coordinate points of the binocular pulse signals through the cross-feature attention matrix.

[0017] The self-attention mechanism and the cross-attention mechanism are alternately connected to form the attention mechanism model.

[0018] Optionally, the method further comprises:

[0019] inputting the pulse signal dimension reduction feature into the attention mechanism model;

[0020] taking the cross-feature attention matrix output by the last round of the attention mechanism model as the binocular pulse feature point matching relationship.

[0021] Optionally, the method further comprises:

[0022] performing disparity regression calculation on the binocular pulse feature point matching relationship to obtain the disparity in the motion scene.

[0023] According to the disparity, the estimated depth map in the motion scene is derived.

[0024] Optionally, the method further comprises:

[0025] collecting the real depth map in the motion scene through clock synchronization technology while collecting the binocular pulse signals;

[0026] comparing the real depth map with the estimated depth map to determine the accuracy of the estimated depth map.

[0027] In a second aspect, the embodiments of the present application provide a binocular pulse depth data acquisition device for a super-high-speed motion scene, which comprises:

[0028] The collection module is configured to collect a binocular pulse signal in a motion scene.

[0029] The feature extraction module is configured to extract features of the binocular pulse signal to obtain pulse signal dimension reduction features in the motion scene.

[0030] The matching relationship determination module is configured to estimate binocular pulse feature point matching relationships in the motion scene according to the pulse signal dimension reduction features and an attention mechanism model.

[0031] The depth determination module is configured to perform disparity regression calculation on the binocular pulse feature point matching relationships to determine an estimated depth map in the motion scene.

[0032] In a third aspect, an embodiment of the present application provides a computer storage medium, which stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and performing the method steps described above.

[0033] In a fourth aspect, an embodiment of the present application provides a terminal, which can include a processor and a memory, wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and performing the method steps described above.

[0034] The technical scheme provided by the embodiment of the present application can include the following beneficial effects:

[0035] In the embodiment of the present application, the binocular pulse depth data collection method and device for the ultra-high speed motion scene. The binocular pulse signal in the motion scene is collected by a neuromorphic vision sensor, the features of the binocular pulse signal are extracted, the pulse signal dimension reduction features in the motion scene are obtained, the binocular pulse feature point matching relationships in the motion scene are estimated according to the pulse signal dimension reduction features and an attention mechanism model, and the disparity regression calculation is performed on the binocular pulse feature point matching relationships to determine the estimated depth map in the motion scene. The present application can effectively solve the problem that the traditional camera cannot perform depth estimation in the motion scene of high-speed motion blur, overexposure or low light; and the depth estimation problem of dynamic objects and static objects can be solved by using the neuromorphic vision sensor.

[0036] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0037] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the specification, serve to explain the principles of the present application.

[0038] Figure 1is a flowchart of a binocular pulse depth data acquisition method for an ultra-high-speed motion scene provided by an embodiment of the present application.

[0039] Figure 2 is a framework diagram of a binocular pulse depth data acquisition method for an ultra-high-speed motion scene provided by an embodiment of the present application.

[0040] Figure 3 is a framework diagram of an attention mechanism model of a binocular pulse depth data acquisition method for an ultra-high-speed motion scene provided by an embodiment of the present application.

[0041] Figure 4 is a device schematic diagram of a binocular pulse depth data acquisition device for an ultra-high-speed motion scene provided by an embodiment of the present application.

[0042] Figure 5 is a terminal schematic diagram provided by an embodiment of the present application. DETAILED DESCRIPTION

[0043] The following description and drawings are illustrative of the specific embodiments of the present application and are not intended to limit the generality of the present application.

[0044] It should be clear that the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0045] The following description refers to the accompanying drawings. Unless otherwise indicated, same numbers in different drawings indicate same or similar elements. The implementations described in the following example embodiments are not meant to represent all implementations consistent with the present application. Rather, they are merely examples of systems and methods consistent with some aspects of the present application as detailed in the appended claims.

[0046] In the description of the present application, it should be understood that the terms "first", "second", etc. are only for the purpose of description, and cannot be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances. In addition, in the description of the present application, "multiple" means two or more, unless otherwise specified. "And / or", which describes the relationship between the associated objects, means that there can be three relationships, for example, A and / or B can mean that there are three cases of A alone, A and B together, and B alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.

[0047] Please refer to Figures 1-3A flowchart of a binocular pulse depth data acquisition method for a super-high-speed motion scene is provided for the embodiments of the present application. As shown in Figures 1-3 the method of the embodiments of the present application can include the following steps:

[0048] S100, acquiring a binocular pulse signal in a motion scene. The motion scene is a super-high-speed motion scene, and the binocular pulse signal is a binocular pulse array.

[0049] The S100 includes acquiring the binocular pulse signal in the motion scene by a neuromorphic visual sensor.

[0050] In the embodiments of the present application, the binocular pulse signal acquired by the neuromorphic visual sensor is a binocular pulse signal stimulated and recorded by the light intensity of the motion scene. The pulse signal is a sparse discrete point set in the space-time domain, and the light intensity information in the motion scene is recorded and represented in the form of the pulse signal. The pulse signal is a three-dimensional pulse array D(x, y, t), and D(x, y, t) represents whether a pulse is triggered at the time t coordinate (x, y). If a pulse signal is just triggered, a digital signal P = 1 (P represents the pulse signal) is output, otherwise P = 0. The pulse camera included in the neuromorphic visual sensor refers to a pulse array with a size of 20000x400x250 (t, x, y) output per second. The specific pulse signal is shown in the pulse data part of Figure 2 .

[0051] In the embodiments of the present application, the neuromorphic visual sensor is a visual sensor that simulates the mechanism of the firing of retinal ganglion cells and the sensitivity of retinal photoreceptor cells to brightness changes, and the spatio-temporal pulse signal fired is a spatio-temporal sparse pulse signal. The light intensity is encoded as sparse and discrete spatio-temporal peaks by simulating the fovea; the super-high-speed full-time visual sensor included in the neuromorphic visual sensor is different from the event camera (DVS), which generates a pulse when the photon accumulation of each pixel reaches a threshold; since the neuromorphic visual sensor perceives the absolute brightness, not the brightness change under the super-high sampling frequency (i.e. 20,000 Hz), this frameless imaging paradigm brings the ability to reconstruct high-speed visual textures; compared with traditional fixed frame rate cameras, it has advantages such as high temporal resolution, high dynamic range and low power consumption. The neuromorphic visual sensor is suitable for processing high-speed visual tasks and has great application potential in the fields of autonomous driving, robot visual navigation positioning, etc.

[0052] In the embodiments of the present application, the neuromorphic visual sensor includes a super-high-speed full-time visual sensor and a dynamic visual sensor. In actual application, the former is mainly used, and the latter is mainly used for comparison.

[0053] The dynamic vision sensor adopts a differential sampling mechanism to perceive the change of light intensity in a motion scene, generates an asynchronous pulse signal, and has the advantages of high time resolution, high dynamic range, and low power consumption. The dynamic vision sensor can overcome the motion blur of a traditional camera in a high-speed motion scene, the overexposure scene and the weak exposure scene of a light-sensitive scene, and the unclear imaging of the dynamic vision sensor. The dynamic vision sensor includes a DVS (Dynamic Vision Sensor), a DAVIS, an ATIS, a Celex, and the like. The dynamic vision sensor can select a DAVIS346 with a resolution of 346*260 and a sampling frequency of 12Meps.

[0054] The ultra-high-speed all-time vision sensor adopts an integral sampling mechanism to perceive the absolute light intensity in a motion scene, and generates a synchronous pulse signal in a fixed time. The ultra-high-speed all-time vision sensor has the characteristics of high time resolution, clear texture, and high dynamic range. The ultra-high-speed all-time vision sensor can overcome the high-speed motion blur of a traditional camera and has the function of high dynamic imaging. More specifically, the neuromorphic vision sensor can select a retinal sensor with a resolution of 400*250 and a sampling frequency of 20000HZ pulse plane.

[0055] In the embodiments of the present application, the neuromorphic vision sensor has more advantages than the dynamic vision sensor. In a high-speed motion blur, overexposure, or low light motion scene, the use of the dynamic vision sensor can only solve the depth estimation problem of dynamic objects; the use of the ultra-high-speed all-time vision sensor can not only solve the depth estimation problem of dynamic objects, but also solve the depth estimation problem of static objects. For example, on a highway, the dynamic vision sensor can capture dynamic cars but cannot capture static roads; the neuromorphic vision sensor can not only capture dynamic cars, but also capture static roads. The objects captured by the dynamic vision sensor and the neuromorphic vision sensor are clear enough and do not have the problem of blur.

[0056] In S200, feature extraction is performed on the binocular pulse signal to obtain a pulse signal dimension reduction feature in the motion scene. The original pulse signal is characterized by a feature extraction algorithm designed based on the characteristics of the pulse signal. Specifically, the feature extraction algorithm includes:

[0057] The binocular pulse signal is accumulated into an integral frame by a pulse encoder. The representation form of the binocular pulse signal includes a frame form: the pulse signal is accumulated into an integral frame in a fixed time.

[0058] Figure 2In the neuromorphic coding portion, L represents the left pulse signal, R represents the right pulse signal, N represents the number of binocular pulse signals, C represents the number of layers or channels, H represents the height of the pulse signal, and W represents the width of the pulse signal; accumulating the number of pulses of different durations can represent different temporal information. In order to cover the changes of the pulse signal over a period of time, this application accumulates the time for each channel of a three-dimensional pulse signal at different intervals, and the accumulation time step within each channel range can be arbitrary.

[0059] The integral frame is input into the hourglass model, and the dimensionality reduction features of the pulse signal in the motion scene are output. In this embodiment, the hourglass model is formed by an hourglass structure feature extractor, and the backbone network uses the hourglass model to extract and reduce the dimensionality of the integral frame to obtain the dimensionality reduction features of the pulse signal, thus realizing multi-level spatiotemporal feature extraction of the pulse signal.

[0060] like Figure 2 As shown in the neuromorphic coding section, when performing convolution and dimensionality reduction operations on the pulse signals, the number of binocular pulse signals becomes 2N, the number of layers or channels becomes C0, C1 and C2 respectively, the height of the pulse signals becomes H / 4, H / 8 and H / 16 respectively, and the width of the pulse signals becomes W / 4, W / 8 and W / 16 respectively.

[0061] S300, based on the pulse signal dimensionality reduction features and attention mechanism model, predict the matching relationship of binocular pulse feature points in the motion scene.

[0062] The attention mechanism model includes a self-attention mechanism and a cross-attention mechanism. The self-attention mechanism calculates a self-feature attention matrix on the same pulse signal, capturing the internal correlation between the pulse signal and its dimensionality-reduced features. The cross-attention mechanism calculates a cross-feature attention matrix on the binocular pulse signals, capturing the feature correlation between pixel coordinates between the binocular pulse signals.

[0063] Both the self-attention mechanism and the cross-attention mechanism are solved using three linear transformation matrices: query value, key value, and value value. In the self-attention mechanism model, the query value, key value, and value value all come from the same pulse signal, while in the cross-attention mechanism, the query value comes from the target pulse signal, and the key value and value value come from the source pulse signal.

[0064] The self-attention mechanism and the cross-attention mechanism are alternately connected to form the attention mechanism model. In this embodiment, the attention mechanism model uses multi-head attention, which enables the attention mechanism model to notice different aspects of the pulse signal, that is, the dimensionality reduction features of the pulse signal are divided into 8 groups and input into 8 attention heads respectively.

[0065] In this embodiment of the application, S300 includes: inputting the dimensionality reduction features of the pulse signal into the attention mechanism model; and using the last round of the cross-feature attention matrix output by the attention mechanism model as the binocular pulse feature point matching relationship.

[0066] The binocular pulse signal is the left and right pulse signal, and the dimensionality reduction feature of the pulse signal is also the dimensionality reduction feature of the left and right pulse signal. The feature correlation between the pixel coordinates of the left and right pulse signals captured by the cross-feature attention matrix obtained by calculating the attention matrix of the dimensionality reduction features of the left and right pulse signals through the attention mechanism model is the matching relationship between the left and right pixels.

[0067] The matching relationship between the left and right pixels captured by the cross-feature attention matrix in the last round, also known as the correlation between the left and right pulse signals, is the matching relationship of the binocular pulse feature points.

[0068] exist Figure 2 In this process, the positional encoding is input into the attention mechanism model so that the dimensionality reduction features of the left pulse signal and the right pulse signal can be distinguished when operating within the attention mechanism model.

[0069] like Figure 3 As shown, the attention mechanism model extracts attention weights by alternately calculating self-attention and cross-attention between the dimensionality-reduced features of the left and right pulse signals using a multi-head attention mechanism: The dimensionality-reduced features of the left and right pulse signals are input into the first-layer self-attention mechanism, which outputs a self-feature attention matrix. This self-feature attention matrix is ​​then input into the first-layer cross-attention mechanism, which outputs a cross-feature attention matrix. This cross-feature attention matrix is ​​then input into the second-layer self-attention mechanism, which outputs a self-feature attention matrix, and so on. The self-feature attention mechanism can be input into the Q-th layer cross-attention mechanism, outputting the final layer's cross-feature attention matrix; the Q-th layer can be 6 or 8 layers. The cross-feature attention matrix is ​​the attention weight output by the attention mechanism model.

[0070] The self-attention mechanism includes a normalized layer and an 8-head self-attention mechanism, while the cross-attention mechanism includes a normalized layer and an 8-head cross-attention mechanism.

[0071] S400, performing disparity regression calculation on the binocular pulse feature point matching relationship to determine the estimated depth map under the motion scene. S400 includes:

[0072] The disparity regression calculation is performed on the binocular pulse feature point matching relationship to obtain the disparity in the motion scene; based on the disparity, the estimated depth map in the motion scene is derived.

[0073] In the embodiment of the present application, the matching relationship between the left and right pixels captured by the last round of the cross feature attention matrix output by the attention mechanism model is subjected to winner-takes-all disparity regression calculation, that is, the pixel with the maximum matching prediction value of the matching relationship between the left and right pixels is selected as the best matching pixel block, and the disparity is obtained through the best matching pixel block.

[0074] According to the inverse relationship that the smaller the depth is, the greater the disparity in the left and right cameras is, the dense estimated depth map of the present application can be obtained by using disparity regression, as shown in Figure 2

[0075] In the embodiment of the present application, the maximum matching prediction value of the matching relationship between the left and right pixels is the maximum attention weight, and the maximum attention weight is taken as the regression attention weight.

[0076] In the embodiment of the present application, the method further comprises:

[0077] At the same time of collecting the binocular pulse signals, the real depth map under the motion scene is collected through clock synchronization technology; the real depth map is the depth map directly collected by the camera.

[0078] The synchronous collection device of the embodiment of the present application comprises two neuromorphic visual sensors and one ZED depth camera, the left and right pulse signals are collected through the two neuromorphic visual sensors; the real depth map is collected through the ZED depth camera.

[0079] The NTP network clock synchronization technology is adopted, the binocular pulse signals and the real depth map under the motion scene are synchronously collected through the synchronous collection device, and the space-time synchronization label is formed. The space-time synchronization label: according to the feature points of the manually labeled binocular pulse signals and the real depth map, the mapping matrix of the binocular pulse signals and the real depth map is determined, and the mapping synchronization operation of the binocular pulse signals and the real depth map is realized.

[0080] The real depth map is compared with the estimated depth map to determine the accuracy of the estimated depth map.

[0081] In the embodiment of the present application, the method in the prior art depends on the "cost volume", which is usually compressed to a manually predefined difference range, and there is a problem of fixed disparity range. This problem limits the generation of the binocular camera configuration in different scenes by the existing method. The attention mechanism model of the present application reconsiders this problem from the order to order angle, which can show good performance while eliminating the cost batch construction.

[0082] ​In the embodiment of the present application, the problem that the frame-based depth estimation technology cannot directly use the new data format as input due to the composition of the pulse signal through three-dimensional sparse and discrete space-time pulse is solved. Although satisfactory performance can be obtained by using the reconstructed frame, the first image reconstruction stage and the depth estimation stage are very complex and time-consuming, and it is difficult to fully utilize the space-time properties of the pulse signal.

[0083] In the embodiment of the present application, the binocular pulse depth data acquisition method for a super-high-speed motion scene is provided. The binocular pulse signal in the motion scene is acquired, the feature extraction is performed on the binocular pulse signal, and the pulse signal dimension reduction feature in the motion scene is obtained. According to the pulse signal dimension reduction feature and the attention mechanism model, the binocular pulse feature point matching relationship in the motion scene is estimated. The disparity regression calculation is performed on the binocular pulse feature point matching relationship, and the estimated depth map in the motion scene is determined. The problem that the traditional camera cannot perform depth estimation in the motion scene such as high-speed motion blur, overexposure or low light is effectively solved.

[0084] The following is an embodiment of the device of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.

[0085] Please refer to Figure 4 which shows a structure schematic diagram of a binocular pulse depth data acquisition device for a super-high-speed motion scene provided by an exemplary embodiment of the present application. The device 1 comprises an acquisition module 10, a feature extraction module 20, a matching relationship determination module 30 and a depth determination module 40.

[0086] The acquisition module 10 is used to acquire the binocular pulse signal in the motion scene.

[0087] The feature extraction module 20 is used to perform feature extraction on the binocular pulse signal, and obtain the pulse signal dimension reduction feature in the motion scene.

[0088] The matching relationship determination module 30 is used to estimate the binocular pulse feature point matching relationship in the motion scene according to the pulse signal dimension reduction feature and the attention mechanism model.

[0089] The depth determination module 40 is used to perform disparity regression calculation on the binocular pulse feature point matching relationship, and determine the estimated depth map in the motion scene.

[0090] It should be noted that the binocular pulse depth data acquisition device for the ultra-high speed motion scene provided in the above embodiment is in the process of executing the binocular pulse depth data acquisition method for the ultra-high speed motion scene, and only the division of the above functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the binocular pulse depth data acquisition device for the ultra-high speed motion scene provided in the above embodiment and the binocular pulse depth data acquisition method for the ultra-high speed motion scene belong to the same concept, and the implementation process is described in detail in the method embodiment, which will not be repeated here.

[0091] The serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0092] In the embodiments of the present application, the binocular pulse depth data acquisition device for the ultra-high speed motion scene. The binocular pulse signal under the motion scene is collected, the feature extraction of the binocular pulse signal is performed, the pulse signal dimension reduction feature under the motion scene is obtained, the binocular pulse feature point matching relationship under the motion scene is estimated according to the pulse signal dimension reduction feature and the attention mechanism model, the disparity regression calculation of the binocular pulse feature point matching relationship is performed, and the estimated depth map under the motion scene is determined. The present application can effectively solve the problem that the traditional camera cannot perform depth estimation in the motion scene such as high-speed motion blur, overexposure or low light.

[0093] The present application also provides a computer readable medium having program instructions stored thereon, which, when executed by a processor, implement the binocular pulse depth data acquisition method for the ultra-high speed motion scene provided by each of the above method embodiments.

[0094] The present application also provides a computer program product containing instructions, which, when running on a computer, causes the computer to execute the binocular pulse depth data acquisition method for the ultra-high speed motion scene of each of the above method embodiments.

[0095] Please refer to Figure 5 The present application provides a terminal structure schematic diagram. As shown in Figure 5 The terminal 1000 can include at least one processor 1001, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.

[0096] The communication bus 1002 is used to realize the connection communication between the components.

[0097] The user interface 1003 can include a display screen, a camera, and optionally a standard wired interface and a wireless interface.

[0098] The network interface 1004 can optionally include a standard wired interface and a wireless interface (e.g., a WI-FI interface).

[0099] The processor 1001 can include one or more processing cores. The processor 1001 connects various parts of the electronic device 1000 through various interfaces and lines, executes various functions of the electronic device 1000 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 1005, and calling data stored in the memory 1005. Optionally, the processor 1001 can be implemented in at least one of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 1001 can be integrated with a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes operating systems, user interfaces, and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 1001, but can be realized by a separate chip.

[0100] The memory 1005 can include a random access memory (RAM) and can also include a read-only memory (ROM). Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 1005 can include a program storage area and a data storage area, where the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store data involved in the above-mentioned various method embodiments, etc. The memory 1005 can also be at least one storage device located away from the aforementioned processor 1001. As shown in Figure 5 The memory 1005 as a computer storage medium can include an operating system, a network communication module, a user interface module, and an availability analysis application program of vehicle running trajectory data.

[0101] In Figure 5 In the terminal 1000 shown, the user interface 1003 is mainly used to provide an interface for user input and obtain user input data; and the processor 1001 can be used to call the binocular pulse depth data acquisition application program for a super-high-speed motion scene stored in the memory 1005 and specifically perform the following operations:

[0102] Acquire binocular pulse signals in a motion scene;

[0103] While acquiring the binocular pulse signals, acquire a real depth map in the motion scene through a clock synchronization technology;

[0104] Extract features from the binocular pulse signals to obtain pulse signal dimension reduction features in the motion scene;

[0105] According to the pulse signal dimension reduction features and an attention mechanism model, estimate a binocular pulse feature point matching relationship in the motion scene;

[0106] Perform a disparity regression calculation on the binocular pulse feature point matching relationship to determine an estimated depth map in the motion scene;

[0107] Compare the real depth map with the estimated depth map to determine the accuracy of the estimated depth map.

[0108] In one embodiment, the processor 1001, when performing the acquisition of the binocular pulse signals in the motion scene, specifically performs the following operations:

[0109] acquire the binocular pulse signals under the motion scene through a neuromorphic visual sensor.

[0110] In one embodiment, the processor 1001, in performing the feature extraction on the binocular pulse signals and acquiring the pulse signal dimension reduction features under the motion scene, specifically performs the following operations:

[0111] accumulate the binocular pulse signals into integral frames through a pulse encoder;

[0112] input the integral frames into an hourglass model, and output the pulse signal dimension reduction features under the motion scene.

[0113] In one embodiment, the processor 1001, in performing the estimation of the binocular pulse feature point matching relationship under the motion scene according to the pulse signal dimension reduction features and an attention mechanism model, specifically performs the following operations:

[0114] The attention mechanism model includes a self-attention mechanism and a cross-attention mechanism; wherein,

[0115] The self-attention mechanism calculates a self-feature attention matrix on the same pulse signal, and captures the internal correlation of the pulse signal and the pulse signal dimension reduction features through the self-feature attention matrix;

[0116] The cross-attention mechanism calculates a cross-feature attention matrix on the binocular pulse signals, and captures the feature correlation of the pixel coordinate points between the binocular pulse signals through the cross-feature attention matrix.

[0117] The self-attention mechanism and the cross-attention mechanism are alternately connected to form the attention mechanism model.

[0118] input the pulse signal dimension reduction features into the attention mechanism model;

[0119] The last round of the cross-feature attention matrix output by the attention mechanism model is taken as the binocular pulse feature point matching relationship.

[0120] In one embodiment, the processor 1001, in performing the disparity regression calculation on the binocular pulse feature point matching relationship to determine the estimated depth map under the motion scene, specifically performs the following operations:

[0121] perform the disparity regression calculation on the binocular pulse feature point matching relationship to obtain the disparity under the motion scene;

[0122] According to the disparity, the estimated depth map under the motion scene is derived.

[0123] In the embodiment of the present application, the binocular pulse depth data acquisition method for a super-high-speed motion scene is provided. The binocular pulse signals in a motion scene are acquired, the feature extraction is performed on the binocular pulse signals, and the pulse signal dimension reduction features in the motion scene are obtained. According to the pulse signal dimension reduction features and an attention mechanism model, the binocular pulse feature point matching relationship in the motion scene is estimated. The disparity regression calculation is performed on the binocular pulse feature point matching relationship, and the estimated depth map in the motion scene is determined. The present application can effectively solve the problem that the traditional camera cannot perform depth estimation in a high-speed motion blur, overexposure, low light, and other motion scenes.

[0124] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The program can be stored in a computer readable storage medium, and when the program is executed, the processes of the above-mentioned embodiments of each method can be included. The storage medium can be a magnetic disc, an optical disc, a read-only memory, or a random access memory, etc.

[0125] The above only discloses the preferred embodiments of the present application, and of course cannot limit the scope of the rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope of the present application.

Claims

1. A binocular pulse depth data acquisition method for ultra-high-speed motion scenarios, characterized in that, Includes the following steps: Acquire binocular pulse signals in motion scenarios; Feature extraction is performed on the binocular pulse signal to obtain the dimensionality reduction features of the pulse signal in the motion scene; The pulse signal dimensionality reduction features include left pulse signal dimensionality reduction features and right pulse signal dimensionality reduction features; Based on the pulse signal dimensionality reduction features and attention mechanism model, the binocular pulse feature point matching relationship in the motion scene is predicted. The multi-layer structure of the attention mechanism model includes, in the i-th layer: a first normalization layer and a first cross-attention mechanism corresponding to the dimensionality reduction features of the left pulse signal, a second normalization layer, a second cross-attention mechanism, and a third normalization layer corresponding to the dimensionality reduction features of the right pulse signal; i is greater than or equal to 1 and less than or equal to the total number of layers corresponding to the multi-layer structure. The outputs of the first and second normalized layers are residually concatenated and then output to the second cross-attention mechanism. The output of the second cross-attention mechanism is connected to the input of the third normalized layer. The outputs of the first and third normalized layers are residually concatenated and then output to the first cross-attention mechanism. The output of the first cross-attention mechanism is residually concatenated with the input of the first normalized layer to obtain the output of the i-th layer corresponding to the dimensionality reduction feature of the left pulse signal. The output of the third normalized layer is the output of the i-th layer corresponding to the dimensionality reduction feature of the right pulse signal. The method involves performing disparity regression calculation on the binocular pulse feature point matching relationship to determine the estimated depth map under the motion scene; this includes: selecting the pixel with the largest matching prediction value between the left and right pixels as the best matching pixel block, and obtaining the disparity through the best matching pixel block; based on the inverse relationship that the smaller the depth, the larger the disparity in the left and right cameras, the method uses disparity regression to obtain the estimated depth map under the motion scene.

2. The binocular pulse depth data acquisition method according to claim 1, characterized in that, The acquisition of binocular pulse signals in motion scenarios includes: The binocular pulse signals in the motion scene are acquired using a neuromorphic visual sensor.

3. The binocular pulse depth data acquisition method according to claim 1, characterized in that, The step of extracting features from the binocular pulse signals to obtain the dimensionality-reduced features of the pulse signals in the motion scene includes: The binocular pulse signals are accumulated into integral frames using a pulse encoder; The integral frame is input into the hourglass model, and the dimensionality reduction features of the pulse signal in the motion scene are output.

4. The binocular pulse depth data acquisition method according to claim 3, characterized in that, The attention mechanism model includes self-attention and cross-attention mechanisms; wherein... The self-attention mechanism calculates a self-feature attention matrix on the same pulse signal, and captures the internal correlation between the pulse signal and the dimensionality-reduced features of the pulse signal through the self-feature attention matrix. The cross-attention mechanism calculates a cross-feature attention matrix on the binocular pulse signal and captures the feature correlation between pixel coordinates of the binocular pulse signal through the cross-feature attention matrix. The self-attention mechanism and the cross-attention mechanism are alternately connected to form the attention mechanism model.

5. The binocular pulse depth data acquisition method according to claim 4, characterized in that, The step of estimating the binocular pulse feature point matching relationship in the motion scene based on the pulse signal dimensionality reduction features and attention mechanism model includes: The dimensionality reduction features of the pulse signal are input into the attention mechanism model; The cross-feature attention matrix output by the attention mechanism model in the last round is used as the binocular pulse feature point matching relationship.

6. The binocular pulse depth data acquisition method according to claim 1, characterized in that, The step of performing disparity regression calculation on the binocular pulse feature point matching relationship to determine the estimated depth map in the motion scene includes: The disparity regression calculation is performed on the binocular pulse feature point matching relationship to obtain the disparity in the motion scene; Based on the parallax, a predicted depth map for the motion scene is derived.

7. The binocular pulse depth data acquisition method according to claim 1, characterized in that, The method further includes: While acquiring the binocular pulse signals, a real depth map of the motion scene is acquired using clock synchronization technology; The accuracy of the estimated depth map is determined by comparing the actual depth map with the estimated depth map.

8. A binocular pulse depth data acquisition device for ultra-high-speed motion scenarios, characterized in that, Includes the following steps: The acquisition module is used to acquire binocular pulse signals in motion scenarios; The feature extraction module is used to extract features from the binocular pulse signal to obtain the dimensionality reduction features of the pulse signal in the motion scene; The pulse signal dimensionality reduction features include left pulse signal dimensionality reduction features and right pulse signal dimensionality reduction features; The matching relationship determination module is used to predict the matching relationship of binocular pulse feature points in the motion scene based on the dimensionality reduction features of the pulse signal and the attention mechanism model. The multi-layer structure of the attention mechanism model includes, in the i-th layer: a first normalization layer and a first cross-attention mechanism corresponding to the dimensionality reduction features of the left pulse signal, a second normalization layer, a second cross-attention mechanism, and a third normalization layer corresponding to the dimensionality reduction features of the right pulse signal; i is greater than or equal to 1 and less than or equal to the total number of layers corresponding to the multi-layer structure. The outputs of the first and second normalized layers are residually concatenated and then output to the second cross-attention mechanism. The output of the second cross-attention mechanism is connected to the input of the third normalized layer. The outputs of the first and third normalized layers are residually concatenated and then output to the first cross-attention mechanism. The output of the first cross-attention mechanism is residually concatenated with the input of the first normalized layer to obtain the output of the i-th layer corresponding to the dimensionality reduction feature of the left pulse signal. The output of the third normalized layer is the output of the i-th layer corresponding to the dimensionality reduction feature of the right pulse signal. The depth determination module is used to perform disparity regression calculation on the matching relationship of the binocular pulse feature points to determine the estimated depth map under the motion scene; including: selecting the pixel with the largest matching prediction value between the left and right pixels as the best matching pixel block, and obtaining the disparity through the best matching pixel block; according to the inverse relationship that the smaller the depth, the larger the disparity in the left and right cameras, the estimated depth map under the motion scene is obtained by disparity regression.

9. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the method steps as claimed in any one of claims 1-7.

10. A terminal, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed the method steps as claimed in any one of claims 1-7.

Citation Information

Patent Citations

  • Joint target detection method and device based on video frames and pulse array signals

    CN110427823A

  • Depth information measurement method based on binocular event camera

    CN113222945A