Optical flow estimation method, device, medium and computer equipment based on brain-like pulses

By combining event camera and frame camera data, using pulsed neural networks and convolutional neural networks to extract features, and generating optical flow fields through residual networks and optical flow refinement networks, the problem of low accuracy of optical flow estimation in the prior art is solved, and more efficient optical flow estimation is achieved.

CN119169045BActive Publication Date: 2025-05-16INST OF AUTOMATION CHINESE ACAD OF SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411283206.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-05-16
Estimated Expiration
2044-09-12

AI Technical Summary

Technical Problem

There are two problems with the existing optical flow estimation method based on pulsed neural networks: one is that it is impossible to record the internal motion information of large blocks of brightness close to objects in the event camera data, and can only record edge motion information; the other is that it does not have performance advantages compared to convolutional neural networks when processing frame camera data.

Method used

By obtaining event camera data and frame camera data, the features are extracted using pulse neural networks and convolutional neural networks, splicing and transforming features through residual networks, and finally using optical flow refinement networks to generate optical flow field.

Benefits of technology

The accuracy of optical flow estimation is improved, the advantages of event cameras and frame cameras are fully utilized, and the performance of optical flow estimation is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119169045B_ABST
    Figure CN119169045B_ABST
Patent Text Reader

Abstract

The present disclosure provides an optical flow estimation method, device, medium and computer equipment based on brain-like pulses. The optical flow estimation method includes: acquiring event camera data and frame camera data; extracting a first feature from the event camera data through a pulse neural network; extracting a second feature from the frame camera data through a convolutional neural network; concatenating the first feature and the second feature to obtain a third feature and transforming the third feature using a residual network to obtain a transformed feature; and performing optical flow refinement on the transformed feature, the feature extracted from at least one layer of the pulse neural network except the output layer, and the feature extracted from at least one layer of the convolutional neural network except the output layer using an optical flow refinement network to generate an optical flow field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision, and in particular, to a method, apparatus, medium and computer equipment for optical flow estimation based on brain-like pulses. Background Art

[0002] Optical flow estimation processes the image sequence on the imaging plane and solves the motion information of the corresponding pixels in two consecutive frames to obtain the optical flow field between adjacent frames. The optical flow field describes the most basic information about the movement of objects, surfaces or edges caused by the movement of the observation camera. Due to the consistency of the object's movement, the outline of the object and the structural information of the scene can be obtained from the optical flow field, and the motion information of the object in the field of view and the posture information of the observation camera can be further inferred.

[0003] Optical flow field is of great significance in the field of computer vision and image processing, and is widely used in tasks such as target detection and tracking, motion compensation coding, image segmentation, scene depth estimation, 3D reconstruction and autonomous driving. There are three main methods for optical flow estimation: convolutional neural network (CNN), recurrent neural network (RNN), and spiking neural network (SNN).

[0004] Since spiking neural networks have a natural advantage in processing event camera data, optical flow estimation algorithms based on spiking neural networks have received widespread attention. However, there are two problems with current spiking neural network methods: First, when using spiking neural networks to process event camera data, the working principle of the event camera makes it impossible to record the internal motion information of large blocks of objects with close brightness in the picture, and can only record the motion information of the edges of such objects, which is not conducive to optical flow estimation; second, when using spiking neural networks to process frame camera data, it does not have a performance advantage compared to convolutional neural networks. Summary of the invention

[0005] One of the objectives of the present invention is to provide an optical flow estimation method that can effectively improve the accuracy of optical flow estimation.

[0006] According to one aspect of the present disclosure, an optical flow estimation method includes: acquiring event camera data and frame camera data; extracting a first feature from the event camera data through a pulse neural network; extracting a second feature from the frame camera data through a convolutional neural network; concatenating the first feature and the second feature to obtain a third feature, and transforming the third feature using a residual network to obtain a transformed feature; performing optical flow refinement on the transformed feature, features extracted from at least one layer of the pulse neural network except the output layer, and features extracted from at least one layer of the convolutional neural network except the output layer using an optical flow refinement network to generate an optical flow field.

[0007] As an example, the residual network can introduce cross-layer connections, and the third feature and the feature extracted from the i-th layer of the residual network are concatenated and input into the i+1-th layer of the residual network. The feature extracted from the i-th layer is concatenated with the feature extracted from the output layer of the residual network as a conversion feature, where i+1 is less than the total number of layers of the residual network.

[0008] As an example, the optical flow refinement network can introduce cross-layer connections, and the features extracted by at least one layer of the pulse neural network and the features extracted by at least one layer of the convolutional neural network are spliced ​​with the output of the non-output layer of the optical flow refinement network in hierarchical correspondence.

[0009] As an example, the optical flow refinement network may include multiple deconvolution layers, multiple transposed convolution layers and multiple optical flow estimation layers, the multiple transposed convolution layers are correspondingly connected to the multiple optical flow estimation layers, the previous transposed convolution layer among the multiple transposed convolution layers is connected to the previous optical flow estimation layer among the multiple optical flow estimation layers, the output of the previous deconvolution layer among the multiple deconvolution layers, the output of the previous optical flow estimation layer of the multiple optical flow estimation layers, the features extracted by the corresponding layer in at least one layer of the pulse neural network, and the features extracted by the corresponding layer in at least one layer of the convolutional neural network are spliced ​​and input into the next transposed convolution layer among the multiple transposed convolution layers and the next deconvolution layer among the multiple deconvolution layers, respectively, and the multiple optical flow estimation layers respectively output optical flow fields of different resolutions.

[0010] As an example, the conversion features can be input to a first deconvolution layer among multiple deconvolution layers and a first transposed convolution layer among multiple transposed convolution layers, the first transposed convolution layer is connected to a first optical flow estimation layer among multiple optical flow estimation layers, the output of the first deconvolution layer, the output of the first optical flow estimation layer, the output of the corresponding layer of the pulse neural network, and the output of the corresponding layer of the convolutional neural network are spliced ​​and input to a second deconvolution layer among multiple deconvolution layers and a second transposed convolution layer among multiple transposed convolution layers, the second transposed convolution layer is connected to a second optical flow estimation layer among multiple optical flow estimation layers, the output of the second deconvolution layer, the output of the second optical flow estimation layer, the output of the corresponding layer of the pulse neural network, and the output of the corresponding layer of the convolutional neural network are spliced ​​and input to a third deconvolution layer among multiple deconvolution layers and a third transposed convolution layer among multiple transposed convolution layers, and the optical flow estimation layer outputs optical flow fields of different resolutions.

[0011] As an example, the first feature and the second feature may be concatenated in a first dimension to obtain a third feature.

[0012] As an example, the step of acquiring frame camera data may include: extracting frame camera data including two adjacent frame images and corresponding timestamps from the frame camera data; the step of acquiring event camera data may include: extracting event data occurring between the two timestamps based on the two acquired timestamps; sorting the extracted event data by time and grouping them according to polarity; and equally dividing the data in each group and then projecting them to obtain event camera data.

[0013] According to a second aspect of the present disclosure, an optical flow estimation device is provided, the optical flow estimation device comprising: an acquisition module, acquiring event camera data and frame camera data; a feature extraction module, extracting a first feature from the event camera data through a pulse neural network and extracting a second feature from the frame camera data through a convolutional neural network; a residual unit, transforming a third feature formed by concatenating the first feature and the second feature to obtain a transformed feature; and an optical flow refinement module, performing optical flow refinement on the transformed feature, features extracted by at least one layer of the pulse neural network except the output layer, and features extracted by at least one layer of the convolutional neural network except the output layer using an optical flow refinement network to generate an optical flow field.

[0014] According to a third aspect of the present disclosure, a computer device is provided, the computer device comprising a memory and a processor, the memory storing instructions or programs, when the instructions or programs are loaded and executed by the processor, the processor is prompted to execute the above optical flow estimation method.

[0015] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a program or instruction, and when the instruction or program is loaded and executed by a processor, the processor is prompted to execute the above-mentioned optical flow estimation method.

[0016] According to the optical flow estimation method of the embodiment of the present disclosure, the frame camera data and the event camera data are used for optical flow estimation at the same time, that is, the event camera data is processed using a pulse neural network and the frame camera data is processed using a convolutional neural network, so as to make full use of the respective advantages of the two cameras and improve the optical flow estimation performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings.

[0018] Figure 1 is a flowchart of an optical flow estimation method according to an embodiment of the present disclosure;

[0019] Figure 2 is a flowchart showing the steps of acquiring event camera data according to an embodiment of the present disclosure;

[0020] Figure 3is a block diagram illustrating an optical flow estimation apparatus according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] The following detailed description is provided to help gain a comprehensive understanding of the methods, devices and / or systems described herein. However, the order of operations described herein is only an example and is not limited to those orders set forth herein, but may be equivalently replaced or changed except for operations that must occur or be performed in a specific order. In addition, for greater clarity and simplicity, the description of content known in the art will be omitted or simplified.

[0022] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which the present disclosure belongs after understanding the present disclosure. Unless explicitly defined as such herein, terms (such as those defined in general dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and should not be interpreted in an idealized or overly formal manner.

[0023] Unless otherwise specified, the same reference numerals generally refer to the same elements (e.g., components, steps, and methods). Reference numerals described in previous embodiments may be omitted if they appear again in later embodiments. In addition, the technical features described in different or the same embodiments may be combined in any manner, as long as the combined embodiments or technical solutions are complete and can solve the technical problems of the present application or achieve the technical effects described or not described in the present disclosure but can be determined based on the above complete technical solutions.

[0024] It should be noted that, in the absence of conflicts between the various embodiments, these embodiments and their features can be combined with each other. The following is a brief description of the terms used in this disclosure.

[0025] Convolutional Neural Network: Convolutional Neural Network (CNN) is a deep learning model widely used in image recognition, video analysis, natural language processing, etc. Convolutional Neural Network (CNN) is designed to automatically learn spatial hierarchical features from input data (usually two-dimensional images), and mainly includes convolutional layers, pooling layers, and fully connected layers.

[0026] Spiking neural network: Spiking neural network (SNN) is composed of spiking neurons, which are interconnected through synapses. The information in the spiking neural network is transmitted in the form of spikes. Each spike is a discrete event, representing that the neuron emitted an electrical signal at a certain moment. Neurons accumulate input signals. When the accumulated potential exceeds a certain threshold, the neuron generates a pulse and transmits this pulse to other neurons connected to it. The neuron then enters a short reset period. Spiking neural networks calculate and transmit information through discrete pulses. The characteristics of spiking neural networks make them particularly suitable for processing event camera data. Compared with traditional frame cameras, event cameras have the advantages of high dynamic range and low latency, and work better in high-speed motion and large light ratio scenes. These advantages can be exploited by using spiking neural networks to obtain better optical flow estimation results.

[0027] Residual Network (ResNet): Residual network is a deep learning model that aims to solve the problem of gradient vanishing or gradient exploding in deep neural networks, as well as the degradation problem when training deep networks. Residual network can achieve this by introducing residual blocks, which enables the network to effectively train deeper layers.

[0028] Optical flow refinement network: The optical flow refinement network is a network model for improving the quality of the optical flow field after the optical flow field is preliminarily estimated, and may include network models that implement various functions.

[0029] According to the optical flow estimation method of the embodiment of the present disclosure, the accuracy of optical flow estimation is improved by processing event camera data and frame camera data respectively using a pulse neural network and a convolutional neural network. The embodiment of the present disclosure is described below in conjunction with the accompanying drawings.

[0030] Figure 1 is a flowchart of an optical flow estimation method according to an embodiment of the present disclosure; Figure 2 is a flowchart showing the steps of acquiring event camera data according to an embodiment of the present disclosure; Figure 3 is a block diagram illustrating an optical flow estimation apparatus according to an embodiment of the present disclosure.

[0031] Reference Figure 1 The optical flow estimation method disclosed in the present invention includes step S10, step S20, step S30, step S40 and step S50.

[0032] In step S10 , event camera data and frame camera data are acquired.

[0033] The event camera data here can be the data acquired by the event camera, and the frame camera data can be the data acquired by the frame camera. The frame camera captures a frame image of the entire scene at a fixed time interval, that is, each frame is a complete two-dimensional image, recording the brightness information of all pixels at that moment. The data output of the frame camera can be a two-dimensional grayscale image, usually with a fixed frame rate (such as 30 frames / second or 60 frames / second).

[0034] Event cameras are visual sensors that mimic the properties of the retina in a biological eye. Instead of capturing images at fixed time intervals, event cameras record events when the brightness of a pixel changes by more than a certain threshold. Therefore, event cameras record pixel-level change information rather than complete frames. Each event contains temporal and spatial information (i.e., when and where the brightness change occurred), as well as the direction of the brightness change (i.e., polarity (e.g., brightening or dimming)).

[0035] As an example, the processed event camera data and frame camera data may be directly read, that is, the event camera data and frame camera data are the event camera data and frame camera data that have been processed.

[0036] As an example, the raw data acquired by the event camera and the frame camera may be pre-processed to obtain the event camera data and the frame camera data.

[0037] The step of acquiring frame camera data may include: extracting frame camera data including two adjacent frame images and corresponding time stamps from the frame camera data.

[0038] As an example, the step of acquiring frame camera data may also include preprocessing the data, such as scaling, cropping, normalization, etc., to facilitate input into the convolutional neural network.

[0039] Frame cameras capture image frames continuously, each of which can be a complete image. These images are usually stored in digital form, such as JPEG, PNG or RAW format. Each frame can be accompanied by a timestamp, which records the exact time when the image was captured. The timestamp is usually high-precision, in microseconds or nanoseconds. The way to obtain the timestamp depends on the hardware and software environment used. As an example, for some frame cameras, the timestamp is directly embedded in the metadata of the image file. As an example, the timestamp can be read through a software development kit (SDK) or an application program interface (API). Event camera data can contain 2 groups of 10 images each.

[0040] Reference Figure 2 , the step of acquiring event camera data may include step S11, step S12 and step S13.

[0041] In step S11 , event data occurring between the two acquired time stamps is extracted based on the two acquired time stamps.

[0042] That is, first we need to determine two timestamps, which correspond to the time points of two adjacent frames captured by the frame camera. We can read the event data that occurred between these two time points from the event stream of the event camera. The event data usually contains the time of occurrence, spatial location (i.e., x, y coordinates), and polarity (increase or decrease in brightness) of each event.

[0043] In step S12, the extracted event data are sorted by time and grouped according to polarity.

[0044] All events can be sorted from smallest to largest in terms of occurrence time. This ensures that events are processed in chronological order. Then, the sorted event data is divided into two groups based on polarity: one group is events with increasing brightness (positive polarity) and the other group is events with decreasing brightness (negative polarity).

[0045] For each set of (i.e. positive and negative) event data, you can divide it into several subsets evenly according to time. For example, if there are 100 events between two timestamps, you can choose to divide them into 10 parts, each with 10 events.

[0046] In step S13, the data in each group is equally divided and then projected to obtain the event camera data.

[0047] The event data in each subset can be projected onto a two-dimensional image to form an "event image". As an example, the event data in each subset is projected into an image according to its coordinate information. If the same coordinate contains multiple events, only one is recorded during projection. The event camera data may include coordinates, timestamps, and polarity.

[0048] In step S20, a first feature is extracted from the event camera data by a spiking neural network. The first feature here may be a feature output by an output layer of the spiking neural network.

[0049] As an example, the event camera data can be input to the spiking neural network. The input vector specification of the spiking neural network is the same as the vector specification obtained after preprocessing. The representation of the spiking neuron is shown in the following equation (1):

[0050] Formula (1)

[0051] Formula (1) is the pulse neuron potential update formula, v l [ n ] represents the neuron potential at that moment, o l-1[n] represents the output of the previous layer of neurons connected to the current neuron at that moment, w l represents the connection weight, V th,pos represents the forward pulse release voltage, V th,neg Represents the reverse pulse release voltage.

[0052] The first feature can be extracted by the spiking neural network, where the first feature can be a vector (for example, a multidimensional tensor). 256 5 Taking the 4-channel event camera information tensor of 4 as an example, the first feature can be 16 16 256. The format (eg, tensor format) of the first feature output by the spiking neural network may vary depending on the design and architecture of the spiking neural network.

[0053] The above-mentioned spiking neural network is a pre-trained spiking neural network. The spiking neural network can be pre-trained by existing means. The training method may include supervised learning, unsupervised learning, transfer learning, etc. The number of layers of the spiking neural network and the convolutional neural network is not specifically limited. Both the spiking neural network and the convolutional neural network can have at least four layers, and the at least four layers here are parallel profiles. As an example, refer to Figure 3 , both the pulse neural network and the convolutional neural network can have four layers.

[0054] In step S30, a second feature is extracted from the frame camera data by a convolutional neural network. The second feature here may be a feature output by an output layer of the convolutional neural network.

[0055] The convolutional neural network may include multiple convolutional layers, pooling layers, and fully connected layers, and is a pre-trained convolutional neural network. For example, a large amount of labeled data can be used to train the convolutional neural network, and the network weights can be optimized by the back propagation algorithm. The cross entropy loss function can be used as the objective function of the classification task, and the convolutional neural network can be trained using optimization algorithms such as gradient descent. The pulse neural network and convolutional neural network used in the present disclosure are only examples and are not specifically limited.

[0056] In step S40, the first feature and the second feature are concatenated to obtain a third feature, and the third feature is transformed using a residual network to obtain a transformed feature.

[0057] The first feature and the second feature can be concatenated in the first dimension to obtain a third feature. As an example, the vectors (i.e., the extracted features) output by the spiking neural network and the convolutional neural network are both multi-dimensional tensors, and the two can be concatenated in the first dimension to obtain a third tensor.

[0058] As an example, the residual network here may adopt a ResNet architecture, such as ResNet-18, ResNet-34, ResNet-50, etc. The conversion feature may be a feature output by the residual network, or a feature output by an output layer of the residual network.

[0059] The residual network can introduce cross-layer connections. For example, the third feature (i.e., input feature or input vector) and the feature extracted from the i-th layer of the residual network are concatenated and input to the i+1-th layer of the residual network, and the feature extracted from the i-th layer is concatenated with the feature extracted from the output layer of the residual network as the conversion feature, where i+1 is less than the total number of layers of the residual network.

[0060] When performing cross-layer connections of the residual network, the input vector can be spliced ​​to the later layers. For example, taking a six-layer network as an example, the third feature (i.e., the input vector) can be input to the 3rd, 4th, or 5th layer of the residual network, and the features extracted from this layer are spliced ​​with the features extracted from the 6th layer as the conversion features (i.e., the output of the residual network).

[0061] Reference Figure 3 Taking a 4-layer network as an example, the input vector is concatenated with the output vector of the second layer and then input into the third layer. The output vector of the second layer is concatenated with the output vector of the fourth layer to form the output of the residual network.

[0062] In step S50, an optical flow refinement network is used to perform optical flow refinement on the conversion features, the features extracted from at least one layer of the pulse neural network except the output layer, and the features extracted from at least one layer of the convolutional neural network except the output layer to generate an optical flow field.

[0063] As an example, in each layer of the optical flow refinement network, the features extracted by the corresponding layers of the previous spiking neural network and convolutional neural network can be spliced ​​with the output of the corresponding layer of the optical flow refinement network and input into the next layer. Therefore, cross-layer connections can also be introduced in the optical flow refinement module to splice the features extracted by each layer of the spiking neural network and convolutional neural network with the output of the corresponding layer of the optical flow refinement module.

[0064] As an example, the optical flow refinement network can introduce cross-layer connections, and the features extracted by at least one layer of the pulse neural network and the features extracted by at least one layer of the convolutional neural network are spliced ​​with the output of the non-output layer of the optical flow refinement network in hierarchical correspondence.

[0065] The optical flow refinement network may include multiple deconvolution layers, multiple transposed convolution layers and multiple optical flow estimation layers. Multiple transposed convolution layers may be connected to multiple optical flow estimation layers accordingly, the previous transposed convolution layer in the multiple transposed convolution layers is connected to the previous optical flow estimation layer in the multiple optical flow estimation layers, the output of the previous deconvolution layer in the multiple deconvolution layers, the output of the previous optical flow estimation layer in the multiple optical flow estimation layers, the features extracted by the corresponding layer in at least one layer of the pulse neural network and the features extracted by the corresponding layer in at least one layer of the convolution neural network are spliced ​​and respectively input to the next transposed convolution layer in the multiple transposed convolution layers and the next deconvolution layer in the multiple deconvolution layers, and the multiple optical flow estimation layers may respectively output optical flow fields of different resolutions.

[0066] As an example, the conversion feature can be input to the first deconvolution layer in the multiple deconvolutions and the first transposed convolution layer in the multiple transposed convolution layers, the first transposed convolution layer is connected to the first optical flow estimation layer in the multiple optical flow estimation layers, the output of the first deconvolution layer, the output of the first optical flow estimation layer, the output of the corresponding layer of the pulse neural network, and the output of the corresponding layer of the convolution neural network are spliced ​​and input to the second deconvolution layer in the multiple deconvolution layers and the second transposed convolution layer in the multiple transposed convolution layers, the second transposed convolution layer is connected to the second optical flow estimation layer in the multiple optical flow estimation layers, the output of the second deconvolution layer, the output of the second optical flow estimation layer, the output of the corresponding layer of the pulse neural network, and the output of the corresponding layer of the convolution neural network are spliced ​​and input to the third deconvolution layer in the multiple deconvolution layers and the third transposed convolution layer in the multiple transposed convolution layers. If the optical flow refinement network has more layers, it can be deduced in the same way.

[0067] For details, please refer to Figure 3 , the spiking neural network input is 256 256 5 4-channel event camera information tensor of 4. The spiking neural network has four spiking neural network layers (indicated by dark yellow, the first layer saccu1, the second layer saccu2, the third layer saccu3 and the fourth layer saccu4), which can be composed of a layer of SIF neurons with the same number of channels as the sconv layer with the same number. The spiking neural network can also include four convolutional layers (indicated by dark blue, the first convolutional layer sconv1, the second convolutional layer sconv2, the third convolutional layer sconv3 and the fourth convolutional layer sconv4), whose convolution kernel size can be 3 3, the step size can be 2, the number of input channels is the number of channels of the tensor passed in the previous layer, the number of output channels of sconv1 can be 32, except for sconv1, the number of output channels can be twice the number of input channels, each convolution layer is connected to a batch normalization layer, the output of the normalization layer is passed to the next level of convolution layer and the same level of pulse neural network layer, the output of the fourth level of pulse neural network layer is the output of the pulse neural network, that is, 16 16 A tensor of 256.

[0068] Convolutional neural network input can be 256 256 2 dual-channel frame camera information tensor. The convolutional neural network has four convolutional layers (indicated by green, the first convolutional layer aconv1, the second convolutional layer aconv2, the third convolutional layer aconv3 and the fourth convolutional layer aconv4), and their convolution kernel sizes can all be 3 3, the stride can be 2, the number of input channels is the number of channels of the tensor passed in the previous layer, the number of output channels of aconv1 can be 32, except for aconv1, the number of output channels is twice the number of input channels, each convolution layer is connected to a batch normalization layer, the normalization layer is connected to a leaky rectified linear unit layer, and then connected to the next level of convolution layer (if any), the output of the fourth level of convolution layer is the output of the convolutional neural network, that is, 16 16 A tensor of 256.

[0069] Reference Figure 3 , the optical flow refinement module may include three deconvolution layers (represented by light yellow, the first deconvolution layer deconv4, the second deconvolution layer deconv3, and the third deconvolution layer deconv2), four optical flow estimation layers (represented by orange, the first optical flow estimation layer predict_flow4, the second optical flow estimation layer predict_flow3, the third optical flow estimation layer predict_flow2, and the fourth optical flow estimation layer predict_flow1), and four transposed convolution layers (represented by light blue, the first transposed convolution layer upsampled_flow4, the second transposed convolution layer upsampled_flow3, the third transposed convolution layer upsampled_flow2, and the fourth transposed convolution layer upsampled_flow1).

[0070] As an example, each deconvolution layer can be composed of a transposed convolution layer, a batch normalization layer, and a leaky rectified linear unit layer connected in sequence, and each optical flow estimation layer can be composed of a batch normalization layer and a convolution kernel with an output channel of 2 and a kernel size of 1. 1 convolutional layers are connected sequentially.

[0071] The output of the residual network is input to the first deconvolution layer deconv4 and the first transposed convolution layer upsampled_flow4. The output of the first deconvolution layer deconv4 is concatenated with the output of the third layer saccu3 of the pulse neural network and the output of the third convolution layer aconv3 and the output of the first optical flow estimation layer predict_flow4, and then input to the second deconvolution layer deconv3 and the second transposed convolution layer upsampled_flow3; the output of the first transposed convolution layer upsampled_flow4 is input to the first optical flow estimation layer predict_flow4, and the first optical flow estimation layer predict_flow4 outputs flow4, that is, the tensor size is 32 32 2. Resolution is 32 32 optical flow estimation results (i.e., optical flow field). The subsequent connection method is similar to that of flow3 (64 64 2), flow2 (128 128 2) and flow1 (256 256 2) Optical flow estimation results at three resolutions. flow1 is the output of the optical flow refinement network. Its two channels are the optical flow estimation values ​​of the corresponding pixels in the x-axis and y-axis directions. The optical flow estimation result can be the same as the resolution of the input image.

[0072] Reference Figure 3 The optical flow estimation device of the present disclosure may include an acquisition module, a feature extraction module, a residual unit and an optical flow refinement module.

[0073] The acquisition module can acquire event camera data and frame camera data; the feature extraction module extracts a first feature from the event camera data through a pulse neural network and extracts a second feature from the frame camera data through a convolutional neural network; the residual unit transforms a third feature formed by concatenating the first feature and the second feature to obtain a transformed feature; the optical flow refinement module uses an optical flow refinement network to perform optical flow refinement on the transformed feature, features extracted from at least one layer of the pulse neural network except the output layer, and features extracted from at least one layer of the convolutional neural network except the output layer to generate an optical flow field.

[0074] Although not shown, the acquisition module, feature extraction module, residual unit and optical flow refinement module can perform the above steps accordingly, which will not be described in detail here.

[0075] It should be noted that the above-mentioned optical flow estimation method according to the exemplary embodiment of the present disclosure can completely rely on the operation of computer programs or instructions to realize the corresponding functions, that is, each device corresponds to each step in the functional architecture of the computer program, so that the entire system is called through a special software package (for example, lib library) to realize the corresponding functions.

[0076] On the other hand, when the systems, units or modules shown in the drawings are implemented in software, firmware, middleware or microcode, the program code or code segments for performing the corresponding operations may be stored in a computer-readable medium such as a storage medium, so that at least one processor or at least one computing device may perform the corresponding operations by reading and running the corresponding program code or code segments. In addition, the computer-readable medium or storage medium may cause the processor to execute the above-mentioned optical flow estimation method when the computer program is executed by the processor.

[0077] The present disclosure may provide a computer-readable storage medium storing a program or instruction. When the instruction or program is loaded and executed by a processor, the processor is prompted to execute the above optical flow estimation method.

[0078] Examples of computer-readable storage media include read-only memory (ROM), random-access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random-access memory (RAM), dynamic random-access memory (DRAM), static random-access memory (SRAM), flash memory, nonvolatile memory, and the like.

[0079] The present disclosure may provide a computer device, which includes a memory and a processor. The memory stores instructions or programs. When the instructions or programs are loaded and run by the processor, the processor is prompted to execute the above optical flow estimation method.

[0080] The modules and units of the present disclosure can be implemented by a neural network model. Functions associated with the neural network model can be executed by a non-volatile memory, a volatile memory, and a processor.

[0081] The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors, such as a central processing unit (CPU), an application processor (AP), etc., a processor used only for graphics (such as a graphics processing unit (GPU), a visual processing unit (VPU) and / or an AI-specific processor (such as a neural processing unit (NPU)).

[0082] One or more processors control the processing of input data according to predefined operating rules or models stored in non-volatile memory and volatile memory. The predefined operating rules or artificial intelligence models can be provided by training or learning. Here, providing by learning means that by applying a learning algorithm to multiple learning data, a predefined operating rule or neural network model with desired characteristics is formed. Learning can be performed in the device itself that executes the neural network model according to the embodiment, and / or can be implemented by a separate server / device / system.

[0083] The optical flow estimation method according to the embodiment of the present disclosure integrates the respective advantages of the event camera and the frame camera, and through feature stitching, enables the model to capture dynamic changes while retaining static image details, thereby improving the accuracy of optical flow estimation.

[0084] According to the optical flow estimation method of the embodiment of the present disclosure, through the fusion of multimodal features, the model can comprehensively consider static images and dynamic event data, thereby achieving deeper information exchange and complementarity at the feature level, and improving the quality of optical flow estimation.

[0085] The optical flow estimation method according to the embodiment of the present disclosure helps to retain and utilize feature information at different scales, ensuring that multi-level information can be fully considered in the process of generating the final optical flow field, thereby improving the robustness and accuracy of optical flow estimation.

[0086] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An optical flow estimation method, characterized in that: The optical flow estimation method comprises: Acquire event camera data and frame camera data, wherein the event camera data is data acquired by the event camera, and the frame camera data is data acquired by the frame camera; extracting a first feature from the event camera data by a spiking neural network; Extracting a second feature from the frame camera data by a convolutional neural network; Concatenating the first feature and the second feature to obtain a third feature, and transforming the third feature using a residual network to obtain a transformed feature; performing optical flow refinement on the conversion features, the features extracted from at least one layer of the spiking neural network except the output layer, and the features extracted from at least one layer of the convolutional neural network except the output layer using an optical flow refinement network to generate an optical flow field, The residual network introduces a cross-layer connection, the third feature and the feature extracted from the i-th layer of the residual network are concatenated and input into the i+1-th layer of the residual network, and the feature extracted from the i-th layer is concatenated with the feature extracted from the output layer of the residual network as the conversion feature, wherein i+1 is less than the total number of layers of the residual network. The optical flow refinement network introduces cross-layer connections, and the features extracted by the at least one layer of the pulse neural network and the features extracted by the at least one layer of the convolutional neural network are spliced ​​with the output of the non-output layer of the optical flow refinement network in hierarchical correspondence.

2. The optical flow estimation method according to claim 1, characterized in that: The optical flow refinement network includes multiple deconvolution layers, multiple transposed convolution layers and multiple optical flow estimation layers, the multiple transposed convolution layers are correspondingly connected to the multiple optical flow estimation layers, the previous transposed convolution layer among the multiple transposed convolution layers is connected to the previous optical flow estimation layer among the multiple optical flow estimation layers, the output of the previous deconvolution layer among the multiple deconvolution layers, the output of the previous optical flow estimation layer of the multiple optical flow estimation layers, the features extracted by the corresponding layer of the at least one layer of the pulse neural network and the features extracted by the corresponding layer of the at least one layer of the convolutional neural network are spliced ​​and respectively input into the next transposed convolution layer among the multiple transposed convolution layers and the next deconvolution layer among the multiple deconvolution layers, and the multiple optical flow estimation layers respectively output optical flow fields of different resolutions.

3. The optical flow estimation method according to claim 2, characterized in that: The converted features are input to the first deconvolution layer among the multiple deconvolution layers and the first transposed convolution layer among the multiple transposed convolution layers, the first transposed convolution layer is connected to the first optical flow estimation layer among the multiple optical flow estimation layers, the output of the first deconvolution layer, the output of the first optical flow estimation layer, the output of the corresponding layer of the pulse neural network, and the output of the corresponding layer of the convolutional neural network are spliced ​​and input to the second deconvolution layer among the multiple deconvolution layers and the second transposed convolution layer among the multiple transposed convolution layers, the second transposed convolution layer is connected to the second optical flow estimation layer among the multiple optical flow estimation layers, the output of the second deconvolution layer, the output of the second optical flow estimation layer, the output of the corresponding layer of the pulse neural network, and the output of the corresponding layer of the convolutional neural network are spliced ​​and input to the third deconvolution layer among the multiple deconvolution layers and the third transposed convolution layer among the multiple transposed convolution layers, and the optical flow estimation layer outputs optical flow fields of different resolutions.

4. The optical flow estimation method according to claim 1, characterized in that: The first feature and the second feature are concatenated in a first dimension to obtain a third feature.

5. The optical flow estimation method according to claim 1, characterized in that: The step of acquiring frame camera data includes: extracting frame camera data including two adjacent frame images and corresponding time stamps from the frame camera data; The step of acquiring event camera data includes: extracting event data occurring between the two acquired time stamps based on the two acquired time stamps; sorting the extracted event data by time and grouping them according to polarity; and equally dividing the data in each group and then projecting them to obtain the event camera data.

6. An optical flow estimation device, characterized in that: The optical flow estimation device comprises: An acquisition module, acquiring event camera data and frame camera data, wherein the event camera data is data acquired by the event camera, and the frame camera data is data acquired by the frame camera; A feature extraction module that extracts a first feature from the event camera data by a spiking neural network and a second feature from the frame camera data by a convolutional neural network, and concatenates the first feature and the second feature to obtain a third feature; The residual unit transforms the third feature using the residual network to obtain a transformed feature; an optical flow refinement module, performing optical flow refinement on the conversion features, the features extracted from at least one layer of the spiking neural network except the output layer, and the features extracted from at least one layer of the convolutional neural network except the output layer using an optical flow refinement network to generate an optical flow field, The residual network introduces a cross-layer connection, the third feature and the feature extracted from the i-th layer of the residual network are concatenated and input into the i+1-th layer of the residual network, and the feature extracted from the i-th layer is concatenated with the feature extracted from the output layer of the residual network as the conversion feature, wherein i+1 is less than the total number of layers of the residual network. The optical flow refinement network introduces cross-layer connections, and the features extracted by the at least one layer of the pulse neural network and the features extracted by the at least one layer of the convolutional neural network are spliced ​​with the output of the non-output layer of the optical flow refinement network in hierarchical correspondence.

7. A computer device, characterized in that: The computer device comprises a memory and a processor, wherein the memory stores instructions or programs, and when the instructions or programs are loaded and executed by the processor, the processor is prompted to perform the optical flow estimation method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program or an instruction, which, when loaded and executed by a processor, causes the processor to execute the optical flow estimation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Synthetic aperture imaging method fusing event camera and traditional optical camera

    CN114862732A

  • Bionic vision sensor optical flow prediction method based on hybrid neural network

    CN115170687A