Optical flow field calculation system based on complementary neuromorphic vision
Through an optical flow field calculation system based on complementary neuromorphic vision, the problems of limited applicability, poor stability, susceptibility to interference and complex calculations in the prior art are solved, and high-speed and stable dense optical flow calculation is realized, which is suitable for applications such as robot control and autonomous driving in complex environments.
Patent Information
- Application Number
- PCT/CN2024/116947
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-09-04
- Publication Date
- 2025-06-19
AI Technical Summary
The existing optical flow field computing technology has problems such as limited applicability, poor stability, susceptibility to interference and complex calculations, making it difficult to achieve high-speed and stable dense optical flow calculations in complex environments.
An optical flow field calculation system based on complementary neuromorphic vision is adopted, and the complementary neuromorphic vision sensor outputs time difference data and spatial differential data. By optimizing the target calculation unit and optical flow field solution calculation unit, combining multi-scale image pyramids and average high velocity constraints in energy forms, iterative dense optical flow estimation is achieved.
It improves the applicability, stability and anti-interference ability of optical flow field calculation, reduces the calculation complexity, and achieves a calculation speed of more than 1,000fps. It is suitable for robot control, autonomous driving and industrial monitoring and other fields.
Smart Images

Figure CN2024116947_19062025_PF_FP_ABST
Abstract
Description
Optical flow field calculation system based on complementary neuromorphic vision
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 2023117316299, filed on December 15, 2023, entitled “Optical flow field calculation system based on complementary neuromorphic vision”, which is incorporated herein by reference in its entirety. Technical Field
[0003] The present application relates to the field of computer vision technology, and in particular to an optical flow field calculation system based on complementary neuromorphic vision. Background Art
[0004] The optical flow field is the corresponding representation of the motion field of an object in space on the image plane. It can be used to approximate the true motion field in space and is applicable to situations where the camera is moving. The optical flow field not only carries information about the object's motion but also about the three-dimensional structure of the scene. Optical flow calculations can detect moving objects without knowing any scene information. Therefore, optical flow calculations play a fundamental role in pattern recognition, computer vision, and other image processing fields.
[0005] Major optical flow algorithms include gradient-based algorithms, typically the LK optical flow method, the HS optical flow method, matching-based algorithms, and algorithms based on energy, phase, and neural dynamics. However, existing algorithms are subject to estimation errors caused by dynamic objects, parallax, and inconsistency. At the sensor level, traditional optical flow cameras also require strong lighting conditions and are very sensitive to geometry, light intensity variations, and surface reflectivity. Consequently, existing technologies suffer from limited applicability, poor stability, susceptibility to interference, and computational complexity.
[0006] Summary of the Invention
[0007] The present application provides an optical flow field calculation system based on complementary neuromorphic vision to address the defects of the existing technology such as limited applicability, poor stability, susceptibility to interference and complex calculation. The present application has good applicability, stability and anti-interference performance in the fields of robot control, autonomous driving, industrial monitoring, etc. The calculation data is directly provided by the complementary neuromorphic vision sensor, without the need for additional data preprocessing, thereby reducing the complexity of the calculation.
[0008] The present application provides an optical flow field calculation system based on complementary neuromorphic vision, comprising: a complementary neuromorphic vision sensor for outputting temporal difference data and spatial difference data; an optimization target calculation unit for determining an optimization target for optical flow field calculation based on the temporal difference data and the spatial difference data; the temporal difference data and the spatial difference data are processed by a preset data transformation algorithm; and an optical flow field solution calculation unit for calculating a multi-scale optical flow field solution based on the optimization target.
[0009] According to an optical flow field calculation system based on complementary neuromorphic vision provided by the present application, the multi-scale optical flow field solution is a dense or sparse optical flow field solution of a multi-resolution image implemented by an image pyramid.
[0010] According to an optical flow field calculation system based on complementary neuromorphic vision provided by the present application, the preset data transformation algorithm is a combination of one or more of a congruent transformation algorithm, a geometric transformation algorithm, a linear combination algorithm, a neural network encoding algorithm and a data compression algorithm.
[0011] According to the optical flow field calculation system based on complementary neuromorphic vision provided by the present application, it also includes: a noise reduction unit for performing correlation noise reduction on the temporal difference data and the spatial difference data.
[0012] According to the optical flow field calculation system based on complementary neuromorphic vision provided by the present application, it also includes: a downsampling unit for downsampling the temporal difference data and the spatial difference data.
[0013] According to an optical flow field calculation system based on complementary neuromorphic vision provided by the present application, the optimization target calculation unit is specifically used to perform mapping coordinate calculation based on the temporal difference data and the spatial difference data, and set a local constrained regularization term to obtain the optimization target of the optical flow field calculation.
[0014] According to an optical flow field calculation system based on complementary neuromorphic vision provided by the present application, the optical flow field solution calculation unit is specifically used to perform differential calculation according to the optimization target, and iteratively calculate the multi-scale optical flow field solution according to the number of iterations of the preset multi-scale optical flow field solution.
[0015] According to an optical flow field calculation system based on complementary neuromorphic vision provided by the present application, it also includes: a verification unit, used to verify the iterative function of the optical flow field according to the multi-scale optical flow field solution; a first execution unit, used to use the multi-scale optical flow field solution as the final optical flow field solution if the iterative function passes the verification; a second execution unit, used to update the multi-scale optical flow field solution if the iterative function fails the verification, and send the updated multi-scale optical flow field solution to the optimization target calculation unit.
[0016] According to an optical flow field calculation system based on complementary neuromorphic vision provided by the present application, the verification unit is specifically used to determine the iterative function based on the multi-scale optical flow field solution; when the iterative function is less than a preset threshold and the loop parameter is less than 0, the iterative function terminates; the iterative function is the error or distance of the optical flow vector at the same position between two iterations.
[0017] According to an optical flow field calculation system based on complementary neuromorphic vision provided by the present application, the second execution unit is specifically used to upsample the multi-scale optical flow field solution as the initial value of the optical flow field solution of the next scale when the iterative function fails to pass the verification, and send the initial value of the optical flow field solution of the next scale to the optimization target calculation unit.
[0018] The present application provides an optical flow field calculation system based on complementary neuromorphic vision, comprising a complementary neuromorphic vision sensor, an optimization target calculation unit, and an optical flow field solution calculation unit. The complementary neuromorphic vision sensor outputs time difference data and spatial difference data; the optimization target calculation unit determines the optimization target for optical flow field calculation based on the time difference data and spatial difference data; the time difference data and spatial difference data are processed by a preset data transformation algorithm; and the optical flow field solution calculation unit calculates multi-scale optical flow field solutions based on the optimization target. This system has good applicability, stability, and anti-interference properties in applications such as robot control, autonomous driving, and industrial monitoring. The computational data is directly provided by the complementary neuromorphic vision sensor, eliminating the need for additional data preprocessing and reducing computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] FIG1 is a schematic diagram of the structure of an optical flow field calculation system based on complementary neuromorphic vision provided by an embodiment of the present application;
[0021] FIG2 is a schematic diagram of the principle of an optical flow field calculation system based on complementary neuromorphic vision provided by an embodiment of the present application;
[0022] FIG3 is a schematic diagram of the principle of geometric transformation provided by an embodiment of the present application;
[0023] FIG4 is a schematic diagram of the principle of correlation filtering provided by an embodiment of the present application;
[0024] FIG5 is a schematic diagram showing the application principle of the CVS provided in an embodiment of the present application in tracking abnormal targets in a robot based on the dense optical flow algorithm. DETAILED DESCRIPTION
[0025] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions in this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0026] Optical flow describes the motion of image brightness. Classic optical flow calculations assume that the brightness of a moving object remains constant over a short period of time, and that the velocity vector field within a given area changes slowly. Therefore, optical flow can estimate the deformation between two images and, in turn, calculate the velocity field in the field of view, yielding motion information.
[0027] Motion perception is a key problem in computer vision, aiming to detect and estimate various motions from data collected by visual sensors (such as cameras or lidar). A crucial component in fields such as robotic control, industrial inspection, and surveillance, motion perception is typically achieved through techniques such as optical flow calculation using monocular or multi-camera systems, inertial navigation using gyroscopes and accelerometers, and lidar-based SLAM (Simulated Local Area Mapping). These methods are widely used in applications such as video analysis, robotic navigation, autonomous driving systems, and virtual and augmented reality. Inertial navigation solutions offer high autonomy, continuity, and speed, but suffer from large errors and difficulty in long-term operation. Vision-based solutions, while free of long-term error accumulation, require significantly more computation and are highly dependent on lighting conditions and environmental texture. In recent years, vision-led or even purely vision-based motion perception has become mainstream in fields such as autonomous driving, and extensive research is underway to integrate inertial navigation with visual motion perception.
[0028] The basic principle of current mainstream image sensors is frame-based capture and recording, achieved through active pixel arrays. Most optical flow cameras belong to this category of visual sensors. Active pixel sensors can only process color images arranged in a pixel matrix image frame format. They offer advantages such as high color reproduction, high resolution, and high image quality. However, the dynamic range of the image signals they acquire is limited, and the capture speed is slow. Event cameras, also known as dynamic visual sensors, are a new type of imaging system. Unlike traditional cameras, which use a shutter to control the frame rate and record light intensity on a per-frame basis, event cameras are sensitive to the rate of change of light intensity. Each pixel independently records the change in the logarithm of the light intensity at that pixel, generating a positive or negative pulse when the change exceeds a threshold. The asynchronous nature of event cameras makes them unconstrained by shutter constraints and possesses extremely high temporal resolution (frame rates of approximately 10,000 to 100,000 fps, compared to the approximately 100 fps of traditional cameras). Combined with their sensitivity to change, they are naturally suitable for tasks such as motion detection. Another camera called DAVIS combines traditional active pixel sensors with event cameras, which can record both single-frame images and event information. It has the advantages of high spatial resolution of traditional cameras and high temporal resolution of event cameras.
[0029] Major optical flow algorithms include gradient-based algorithms, typically LK optical flow, HS optical flow, matching-based algorithms, energy-based algorithms, phase-based algorithms, and neural dynamics algorithms. The most popular optical flow algorithms are LK (Lucas-Kanade) and HS (Horn-Schunck) optical flow. HS optical flow is an image registration algorithm that performs point-by-point matching on a specific region within an image, calculating the offsets of all points to form a dense optical flow field for image registration. However, due to the need to calculate all points, the HS optical flow algorithm is computationally intensive and time-inefficient. The LK optical flow algorithm, on the other hand, uses sparse optical flow registration, registering a set of distinct feature points within the image. Compared to the HS algorithm, it uses less computation and is more time-efficient. The LK algorithm also offers relatively stable and reliable tracking. To improve the feature matching performance of the LK algorithm, a pyramid-based LK optical flow method has been studied. This method processes the image layer by layer, resizing the image, and solves the problem layer by layer to improve computational accuracy. At present, research is also ongoing on positioning algorithms that combine optical flow detection block matching methods with inertial measurement units.
[0030] Although optical flow calculation based on visual sensors has gained importance and key applications in motion perception, existing optical flow methods still have many drawbacks. First, algorithmically, optical flow estimation methods are subject to estimation errors caused by dynamic objects, parallax, and image inconsistencies. At the sensor level, traditional optical flow cameras also require strong lighting conditions and are very sensitive to geometry, lighting changes, and surface reflectivity. Therefore, implementing a system that can quickly calculate dense optical flow in open environments remains a challenging task and is crucial for robotic motion control.
[0031] In the prior art, the slight motion of the event camera itself can encode the edge information (i.e., spatial gradient) of rigid objects into the time axis, making it possible to estimate optical flow. However, the optical flow estimation method based on event cameras is not robust and still has the following drawbacks:
[0032] (1) Limited applicability: Most event camera optical flow estimation methods are only applicable to specific scenes or motions and are difficult to apply to complex or diverse environments. In particular, neural network-based optical flow estimation methods still lack generalization.
[0033] (2) Poor stability: Event camera optical flow estimation methods are challenging in terms of sparsity and noise. The sparsity of event cameras may lead to inaccuracies in optical flow calculations, while noise also affects the stability of optical flow calculations.
[0034] (3) Susceptibility to interference: The event camera optical flow estimation method is easily disturbed by environmental factors such as local lighting properties and mirror reflection, resulting in a decrease in accuracy.
[0035] (4) Large amount of data: Since event cameras can generate high-speed data streams, the amount of data they output is very large, and relevant algorithms are required to effectively process and analyze this data.
[0036] In order to solve the technical problems existing in the prior art, the present application uses data from a new type of complementary neuromorphic vision sensor to provide an iterative, high-speed and real-time dense optical flow estimation method. The present application also supports the use of traditional neuromorphic sensors to form a similar multi-channel complementary neuromorphic data structure through a spectroscopic system. The present application develops an iterative real-time dense optical flow algorithm based on complementary neuromorphic data, combines a multi-scale image pyramid, and solves the dense optical flow field with the same resolution as the image through the average high-speed constraint in the form of energy and the basic equation of optical flow, and ensures a computing speed of more than 1000fps. At the same time, it supports the use of the spatiotemporal constraints of complementary cameras and binocular complementary cameras to solve the camera's own motion, background motion and target motion, and achieve more refined motion perception.
[0037] Human vision is currently the most complete visual system, capable of achieving dynamic range, speed, resolution, image quality, and more that far surpass those of artificial visual systems at the cost of extremely low bandwidth and power consumption. The most important reason for this is that humans decompose visual information into complementary information streams and process them in parallel and asynchronously in the brain, enabling efficient responses to both static, detailed information and dynamic, rapid changes. Similar to the human visual system, this application proposes a processing framework and data structure for complementary visual information streams. The complementary theory, drawing on the characteristics of human vision, requires that the acquired data belong to multiple different pathways, and that the data structure of each pathway should be complementary in different properties. These properties with complementary characteristics are called primitives. Primitives include temporal resolution (fast and slow complementarity), spatial resolution (high and low complementarity), color (complementarity in spectral sensitivity ranges such as color, grayscale, infrared, and ultraviolet), sensitivity (i.e., high and low complementarity in the response coefficient to light intensity), response modality (integrated intensity or differential change), and data accuracy (high and low complementarity). CVS (complementary vision sensor) is a new type of neuromorphic vision sensor. Complementary cameras are based on the theory of complementary perception. This theory, inspired by the characteristics of human vision, requires that acquired data be distributed across multiple pathways, with the data structures of each pathway complementing each other in various ways. These complementary properties are called primitives. Primitives include visual data's temporal resolution (fast and slow complementarity), spatial resolution (high and low complementarity), color (complementarity across spectral sensitivity ranges, such as color, grayscale, infrared, and ultraviolet), sensitivity (high and low complementarity in the response coefficient to light intensity), response modality (integrated intensity or differential change), and data accuracy (high and low complementarity). CVS is characterized by its ability to output multiple data modalities from a single vision chip, modeled after the human retina. By constructing a hybrid pixel arrangement within the image sensor and designing a hybrid data readout circuit, CVS can output RGB, spatially differential, and temporally differential data from the same CMOS chip, encoding this information in different modalities. This information is transmitted using separate data pathways. These different modal data have strong complementary properties, including in terms of sampling accuracy, sampling speed, dynamic range, sensitivity, color range, spatial resolution, etc. These pathways are different from each other and can complement each other to ensure the integrity of the information. Therefore, they are also called complementary cameras. Unlike traditional image sensors, CVS is closer to the information processing mechanism of the human retina. It simulates the computing mode of retinal ganglion cells at the chip level, supports isotropic and anisotropic center-periphery structures, and can be used for tasks such as motion detection, scene segmentation, and target tracking. This bio-inspired sensor design method is expected to significantly reduce computational complexity and improve computational efficiency. The multimodal output of CVS allows it to be used in combination with different neural networks to build an end-to-end visual system.It can also be applied to multi-sensor fusion, working in conjunction with other non-visual sensors to provide more robust and intelligent environmental perception capabilities. CVS technology is expected to promote the practical application of visual algorithms in a wide range of fields such as autonomous driving, service robots, and intelligent monitoring.
[0038] Please refer to Figure 1, which is a schematic diagram of the structure of an optical flow field calculation system based on complementary neuromorphic vision provided in an embodiment of the present application.
[0039] Please refer to FIG2 , which is a schematic diagram showing the principle of an optical flow field calculation system based on complementary neuromorphic vision provided in an embodiment of the present application.
[0040] The present application provides an optical flow field calculation system based on complementary neuromorphic vision, comprising: a complementary neuromorphic vision sensor for outputting temporal difference data and spatial difference data; an optimization target calculation unit for determining an optimization target for optical flow field calculation based on the temporal difference data and the spatial difference data; the temporal difference data and the spatial difference data are processed by a preset data transformation algorithm; and an optical flow field solution calculation unit for calculating a multi-scale optical flow field solution based on the optimization target.
[0041] The data structure of this application requires multiple data paths, where at least one primitive between each path is complementary. This data can be directly provided by the CVS camera. Specifically, full-resolution SD (spatial difference data) and TD (temporal difference data) are input. Both SD and TD are high-precision data (greater than or equal to 4 bits of precision), with length and width both being W*H. SD and TD are directly output by the CVS camera.
[0042] The system has good applicability, stability and anti-interference performance in applications such as robot control, autonomous driving, and industrial monitoring. The computing data is directly provided by complementary neuromorphic vision sensors, without the need for additional data preprocessing, thus reducing the complexity of the calculation.
[0043] Based on the above embodiment:
[0044] As a preferred embodiment, the multi-scale optical flow field solution is a dense or sparse optical flow field solution of a multi-resolution image implemented by an image pyramid.
[0045] An image pyramid is a multi-scale representation of an image and an effective structure for interpreting images at multiple resolutions. An image pyramid is a collection of images arranged in a pyramidal shape, each originating from the same original image, with decreasing resolution. It is obtained by sequential downsampling until a certain termination condition is reached. The base of the pyramid is a high-resolution representation of the image being processed, while the top is a low-resolution approximation. The layers of an image can be likened to a pyramid: the higher the level, the smaller the image and the lower the resolution.
[0046] Please refer to FIG3 , which is a schematic diagram of the principle of geometric transformation provided in an embodiment of the present application.
[0047] As a preferred embodiment, the preset data transformation algorithm is a combination of one or more of a congruent transformation algorithm, a geometric transformation algorithm, a linear combination algorithm, a neural network encoding algorithm and a data compression algorithm.
[0048] Specifically, the gradients in the x, y, and t directions are derived using the geometric constraints of the CVS pixel array. Since the differential pixel arrangement of a CVS camera may not be in the positive x and y directions, for ease of calculation, the SD is first projected onto the x and y directions of the coordinate plane.
[0049] Since CVS cameras often use complementary pixel arrays, the SD of most pixels is calculated in a 45-degree direction rather than parallel to the pixel coordinate system. The SD data in a CVS camera is obtained by subtracting the voltage of the pixels in the even-numbered rows from the voltage of the pixels in the odd-numbered rows (counting from 0), and is expressed as: SD in the four directions of upper left, lower left, upper right, and lower right, respectively. ul ,SD ll ,SD ur ,SD lr , and TD is generated by each pixel and can be directly used as I traw For I xraw and I yraw , can be derived as follows:
[0050] First, for even rows, each target I xraw , I yraw There are four available SD values:
[0051] For odd rows, 1 to 4 SD values in its 8-neighborhood are used to derive the value. Here, SD(dx,dy) is the SD value with a (dx,dy) deviation from the target position. The case of 1 value is shown as R0 at (0,0) in the figure above:
[0052] The case of 2 values is such as R1 at (2,0)
[0053] The case of 4 values is R1 at (2,2)
[0054] At this point, all I in TD position xraw and I yrawFurthermore, if the optical flow information needs to be supplemented at the RGB position, bicubic interpolation can be performed.
[0055] Please refer to FIG4 , which is a schematic diagram of the principle of correlation filtering provided in an embodiment of the present application.
[0056] As a preferred embodiment, it further includes: a noise reduction unit, which is used to perform correlation noise reduction on the time difference data and the spatial difference data.
[0057] Specifically, since TD and SD have shot noise, the shot noise is eliminated by median filtering, and the continuity of the optical flow can also be enhanced.
[0058] Since TD and SD have fixed pattern noise, they are eliminated by mean filtering.
[0059] Considering the continuity of actual motion, it is generally assumed that there are no spatiotemporal isolated points in TD. Therefore, correlation filtering is performed on TD frames with T>0 to exclude isolated TD values that do not have other valid values in their spatiotemporal neighborhood. The specific operation is to determine whether the TD values in the surrounding spatiotemporal cube (x±dx, y±dy, t±dt) at each time point (x, y, t) contain non-zero values. If all are zero, the point is set to zero. A similar operation is performed on the spatial difference data SD.
[0060] As a preferred embodiment, it further includes: a downsampling unit, which is used to perform downsampling processing on the time difference data and the spatial difference data.
[0061] V x ,V y , SD and TD are downsampled to get Gradient map of resolution, denoted as TD i ,SD i .
[0062] As a preferred embodiment, the optimization target calculation unit is specifically used to perform mapping coordinate calculation based on the time difference data and the spatial difference data, and set a regularization term of local constraints to obtain the optimization target of the optical flow field calculation.
[0063] Determining an optimization target for optical flow field calculation according to the time difference data and the space difference data includes: determining the optimization target for optical flow field calculation according to the time difference data and the space difference data based on a first preset formula; the first preset formula is:
[0064] Due to the sparsity of event data, the continuity requirement of optical flow, and the interference of noise, the optical flow data of existing technologies is difficult to use for accurate motion perception. In this embodiment, the energy function ε is introduced as the optimization target to ensure the consistency of optical flow while avoiding the computational pathology caused by sparsity. Here, SD and TD are mapped to the difference value I in the x direction. x , the difference value in the y direction I y , and the differential value I in the t direction t , so the optimization objective is as shown in the first preset formula: ε(v x ,v y )=(I x v x +I y v y +I t ) 2 +λ 2 C 2
[0065] Among them, ε(v x ,v y ) is the optimization target; I x Mapping time difference data and space difference data into differential values in the x direction; I y Mapping time difference data and space difference data into difference values in the y direction; I t The time difference data and space difference data are mapped to the difference value in the t direction; t is time; C is the local constraint constant; λ is the Lagrange multiplier; v x is the solution of the optical flow field in the x direction; v y is the solution of the optical flow field in the y direction.
[0066] C is used to minimize the difference between the mean of the estimated x and y velocities and the actual calculated values. This serves as a continuity constraint for rigid body motion at each spatiotemporal coordinate, while also ensuring smoothness in the neighborhood of the optical flow calculation. λ, an adjustable parameter in the energy function optimization, can be set to 1. As λ increases, the optimization results tend toward global optical flow consistency.
[0067] As a preferred embodiment, the optical flow field solution calculation unit is specifically used to perform differential calculation according to the optimization target, and iteratively calculate the multi-scale optical flow field solution according to the preset number of iterations of the multi-scale optical flow field solution.
[0068] Specifically, for ε(v x ,v y ) calculate the total differential and make it 0, that is, the multi-scale optical flow field solution is shown in the second preset formula:
[0069] Among them, vx [i+1] is the solution of the multi-scale optical flow field in the x direction, v y [i+1] is the solution of the multi-scale optical flow field in the y direction, and i is the number of iterations of the solution.
[0070] As a preferred embodiment, it also includes: a verification unit, used to verify the iterative function of the optical flow field based on the multi-scale optical flow field solution; a first execution unit, used to use the multi-scale optical flow field solution as the final optical flow field solution if the iterative function passes the verification; a second execution unit, used to update the multi-scale optical flow field solution if the iterative function fails the verification, and send the updated multi-scale optical flow field solution to the optimization target calculation unit.
[0071] As a preferred embodiment, the verification unit is specifically used to determine the iterative function based on the multi-scale optical flow field solution; when the iterative function is less than a preset threshold and the loop parameter is less than 0, the iterative function terminates; the iterative function is the error or distance of the optical flow vector at the same position between two iterations.
[0072] As a preferred embodiment, the second execution unit is specifically used to upsample the multi-scale optical flow field solution as the initial value of the optical flow field solution of the next scale when the iterative function fails to pass the verification, and send the initial value of the optical flow field solution of the next scale to the optimization target calculation unit.
[0073] Specifically, the iterative calculation process of the present application takes the upper limit of the relative distance or the upper limit of the number of iteration rounds as the iteration end point.
[0074] It is worth noting that due to I x ,I y ,I t All are given directly by the hardware, so there is no need for preprocessing.
[0075] This application further incorporates multi-scale optimization to avoid the aperture effect in optical flow calculations. By observing the definition of C, it can be seen that the velocity vector involved in this calculation can be observed at a smaller resolution.
[0076] TD i ,SD i Substitute the second preset formula and verify the iterative function.
[0077] The iterative function is: |(V x ,V y )[i+1]-(V x ,V y )[i]|
[0078] If the iteration function is less than the preset threshold and the loop parameter is less than 0, the iteration function passes the verification;
[0079] To prevent multi-scale optical flow from ultimately converging to the result of single-scale optical flow, as the number of iterations increases (i.e., the resolution of the optical flow image increases), the threshold is increased while the number of loops is reduced. The simplest strategy is to increase the threshold linearly with the logarithm and decrease the number of loops linearly. The specific strategy depends on the degree of object fineness that the actual optical flow needs to perceive.
[0080] If the iterative function is not less than the preset threshold, or if the iterative function is less than the preset threshold and the loop parameter is not less than 0, the iterative function fails verification; the loop parameter is reduced by 1 after each iterative calculation.
[0081] When the iteration function is less than a preset threshold and the loop parameter n is less than 0, the multi-scale optical flow field solution is taken as the final optical flow field solution.
[0082] When the iterative function is not less than the preset threshold, i=i+1, and i is substituted into the second preset formula, and the iteration is repeated until the iterative function passes the verification.
[0083] When the iteration function is less than the preset threshold and the loop parameter is not less than 0, it indicates that V x ,V y The current size is less than or equal to the original image (I x ,I y ,I t ) size, for V x ,V y Perform bicubic upsampling to obtain V with 2 times the resolution x ,V y , until the iterative function passes the check.
[0084] Please refer to Figure 5, which is a schematic diagram of the application principle of CVS based on the dense optical flow algorithm in robot abnormal target tracking provided by an embodiment of the present application.
[0085] Based on this high-speed dense optical flow algorithm and a CVS camera, dense optical flow data can help robots detect dynamic obstacles in their surroundings and calculate their speed and direction. This information can be used to plan the robot's path to avoid collisions. By analyzing dense optical flow data, the position and trajectory of moving objects can be tracked. This is extremely useful for robots tracking targets, performing pursuit tasks, and performing surveillance and security applications.
[0086] Since CVS cameras often have three different data modalities: RGB, SD, and TD. RGB data has high accuracy but a low frame rate. Therefore, a detector can be used to screen out suspicious objects in the RGB data. When the object first appears, the optical flow at the corresponding position of the dense optical flow field generated under SD and TD is obtained based on the RoI. The velocity of each point of the target can be calculated for point-by-point tracking. This method can significantly shorten the tracking delay of abnormal targets that suddenly appear in the field of view. Based on the calculation that the RGB frame rate is 30fps and the TD and SD frame rates are above 1500fps, this method can reduce the target tracking delay from 33ms to below 3ms on the edge and to below 1ms on the host.
[0087] This method enables real-time optical flow computation for complementary neuromorphic vision sensors. Because this algorithm uses vectorized computation (all operators can be accelerated in parallel), data is directly supported by a single CVS camera, and there is little need for additional data preprocessing, it can compute full-resolution dense optical flow at extremely high frame rates, exceeding 300 fps on edge computers. This will greatly expand the application scenarios of CVS. It can be directly applied in robotic vision environments with high speed and motion accuracy requirements, such as high-speed drone control, autonomous driving, and industrial monitoring.
[0088] Visual odometry: Dense optical flow data can be used to calculate the robot’s motion relative to its previous position. By tracking optical flow patterns in the environment, the distance and direction the robot has moved in space can be estimated, enabling autonomous navigation and map construction.
[0089] Motion Feedback and Control: Dense optical flow data can provide real-time feedback about the motion of an object or environment. This data can be used to control the attitude, velocity, and position of robots, vehicles, or aircraft. By analyzing the direction and intensity of optical flow patterns, control commands can be adjusted in real time to achieve accurate motion control.
[0090] Visual servoing: Dense optical flow data can be used as a feedback loop in visual servo control systems. Optical flow data provides information about the motion of objects in the environment and can be used to adjust the orientation, distance, and speed of a robot or machine vision system to keep the object or camera stable or on a specific tracking trajectory.
[0091] In the aforementioned tasks, the frame rate of dense optical flow is often low, typically 10 to 30 fps. However, the iteration speed requirement of the control loop in robotic control is closely related to the robot's operating speed. In systems such as high-speed drones, the closed-loop control speed often exceeds 100 or even 1000 fps. The system constructed with this method can directly apply dense optical flow in such high-speed environments.
[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. An optical flow field calculation system based on complementary neuromorphic vision, comprising: A complementary neuromorphic vision sensor for outputting temporal difference data and spatial difference data; An optimization target calculation unit, used to determine an optimization target for optical flow field calculation according to the temporal difference data and the spatial difference data; The time difference data and the space difference data are processed by a preset data transformation algorithm; The optical flow field solution calculation unit is used to calculate the multi-scale optical flow field solution according to the optimization target.
2. The optical flow field calculation system based on complementary neuromorphic vision according to claim 1, wherein: The multi-scale optical flow field solution is a dense or sparse optical flow field solution of a multi-resolution image implemented by an image pyramid.
3. The optical flow field calculation system based on complementary neuromorphic vision according to claim 1, wherein: The preset data transformation algorithm is a combination of one or more of a congruent transformation algorithm, a geometric transformation algorithm, a linear combination algorithm, a neural network encoding algorithm and a data compression algorithm.
4. The optical flow field calculation system based on complementary neuromorphic vision according to claim 1, further comprising: The denoising unit is used to perform correlation denoising on the temporal differential data and the spatial differential data.
5. The optical flow field calculation system based on complementary neuromorphic vision according to claim 1, further comprising: The down-sampling unit is used to perform down-sampling processing on the time difference data and the space difference data.
6. The optical flow field calculation system based on complementary neuromorphic vision according to claim 1, wherein: The optimization target calculation unit is specifically used to perform mapping coordinate calculation according to the time difference data and the space difference data, and set a regularization term of a local constraint to obtain the optimization target of the optical flow field calculation.
7. The optical flow field calculation system based on complementary neuromorphic vision according to claim 1, wherein: The optical flow field solution calculation unit is specifically used to perform differential calculation according to the optimization target, and iteratively calculate the multi-scale optical flow field solution according to the preset number of iterations of the multi-scale optical flow field solution.
8. The optical flow field calculation system based on complementary neuromorphic vision according to any one of claims 1 to 7, further comprising: A verification unit, used for verifying the iterative function of the optical flow field according to the multi-scale optical flow field solution; A first execution unit, configured to use the multi-scale optical flow field solution as a final optical flow field solution if the iterative function passes verification; The second execution unit is used to update the multi-scale optical flow field solution when the iterative function fails to pass the verification, and send the updated multi-scale optical flow field solution to the optimization target calculation unit.
9. The optical flow field calculation system based on complementary neuromorphic vision according to claim 8, wherein: The verification unit is specifically used to determine the iterative function according to the multi-scale optical flow field solution; When the iterative function is less than a preset threshold and the loop parameter is less than 0, the iterative function terminates; The iteration function is the error or distance of the optical flow vector at the same position between two iterations.
10. The optical flow field calculation system based on complementary neuromorphic vision according to claim 8, wherein: The second execution unit is specifically used to upsample the multi-scale optical flow field solution as the initial value of the optical flow field solution of the next scale when the iterative function fails to pass the verification, and send the initial value of the optical flow field solution of the next scale to the optimization target calculation unit.
Citation Information
Patent Citations
Bionic vision sensor optical flow prediction method based on hybrid neural network
CN115170687A
SLAM system based on vehicle-mounted multi-view camera and deep neural network
CN116664621A
Optical flow field calculation system based on complementary neuromorphic vision
CN117893578A
Method of Estimating Relative Motion Using a Visual-Inertial Sensor
US20180075609A1
Edge-aware spatio-temporal filtering and optical flow estimation in real time
US20180324465A1
Cited By
Motion information extraction and fast optical flow calculation system based on memristor
CN120707601A
Marine vortex intelligent sensing and tracking space-time fusion method
CN121615490A
Air target automatic detection and following lock positioning method
CN121721653A