An Image Static Region Extraction Method and Device Based on Optical Flow and View Geometry Constraints

Through the image static area extraction method based on optical flow and view geometric constraints, the problem of low positioning and mapping accuracy in a dynamic environment is solved, and higher positioning and mapping accuracy and real-time performance are achieved.

CN114782499BActive Publication Date: 2025-06-20HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210463260.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-28
Publication Date
2025-06-20
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

Traditional visual SLAM technology has low positioning and mapping accuracy in dynamic environments, making it difficult to effectively remove interference from dynamic targets.

Method used

The image static area extraction method based on optical flow and view geometric constraints is adopted. Through the steps of feature point matching, optical flow calculation, counter-pole geometric checksum clustering, dynamic target areas are removed and the image static area is extracted.

Benefits of technology

It effectively improves the positioning and mapping accuracy of visual SLAM in a dynamic environment, overcomes the problems of instability and low accuracy of traditional SLAM under dynamic target interference, and achieves real-time and general effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782499B_ABST
    Figure CN114782499B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for extracting static regions of an image based on optical flow and view geometry constraints. The acquired image is preprocessed by filtering and feature corner points with direction information are extracted; after preprocessing the feature point set, the feature point set is matched, the sparse optical flow method is used to calculate the optical flow vector field, and static feature points are initially separated; then the static feature points are further verified using epipolar geometry constraints to optimize the separation result of the static feature points; finally, the non-static feature points are clustered, their pixel regions are divided, and after image morphological processing, the static regions of the image are output and provided to visual SLAM for positioning and map construction. The present invention combines feature optical flow and geometric constraints, solves the interference of dynamic targets to the visual SLAM system, and improves the positioning and mapping accuracy of the visual SLAM system in a dynamic environment; at the same time, it can ensure the real-time performance of the system and realize a reliable, fast, high-precision, and low-latency visual positioning and mapping system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and mobile robot positioning, and particularly relates to a method and device for extracting static regions of images based on optical flow and view geometry constraints. Background Art

[0002] In recent years, computer vision and robotics have become a hot research direction. Among them, the most fundamental research is the positioning problem of the robot itself. The Simultaneous Localization and Mapping (SLAM) technology is widely applied to the positioning, navigation, and obstacle avoidance of robots. There is a rich variety of sensors for robots. Visual sensors have the advantages of low cost and large amount of information, and are widely used in the SLAM technology. Traditional simultaneous localization and mapping technologies focus on the research of ideal static scenarios without moving targets. Dynamic targets in the scenario can cause deviations in the calculation of the camera pose, resulting in inaccurate positioning of the entire visual SLAM system. The pure static assumption is not applicable in the real environment, and SLAM applicable to dynamic environments is a research hotspot of this technology.

[0003] With the development of deep learning technology, deep learning methods have good effects in object recognition and semantic segmentation. However, most deep learning networks require GPU acceleration support to achieve a real-time detection effect, which does not conform to the actual hardware situation of current mobile robots. Currently, the recognition methods using computer vision are mostly designed for specific scenarios, with poor generality and low accuracy. Summary of the Invention

[0004] The purpose of the present invention is to propose a method and device for extracting static regions of images based on optical flow and view geometry constraints in view of the deficiencies of the prior art, which are applied to a visual SLAM system, can remove the interference of dynamic targets to the SLAM system, and improve the positioning and mapping accuracy of the visual SLAM system in dynamic environments.

[0005] The purpose of the present invention is achieved through the following technical solutions: In the first aspect, the present invention provides a method for extracting static regions of images based on optical flow and view geometry constraints, and the method includes the following steps:

[0006] Step 1: Convert a plurality of color images to be processed into grayscale images;

[0007] Step 2: After performing Gaussian filtering on each frame of grayscale image, perform feature extraction, and at the same time use image pyramids for processing, extract FAST corner points and perform uniform processing, calculate the direction information of FAST corner points using the gray centroid method, and a FAST corner point and its direction information form a feature point. The feature points extracted from each frame of grayscale image form a feature point set.

[0008] Step 3: Filter the feature point sets of two adjacent grayscale images to be matched by using the feature point direction information and Euclidean distance difference between two adjacent grayscale images, so as to reduce the search range of matching.

[0009] Step 4: Perform optical flow calculation on the filtered feature point sets of two adjacent frames to obtain an optical flow vector field; for the calculated optical flow vector field, remove the mismatches based on the intensity and direction consistency, cluster to obtain a background optical flow vector field, and simultaneously obtain the feature point set of the background optical flow vector field.

[0010] Step 5: Further verify the matching situation for the feature point set of the background optical flow vector field by using view geometry constraints, use the Progressive Sampling Consensus (PROSAC) method to solve the essential matrix, and perform epipolar geometry verification to further optimize the matching result.

[0011] Step 6: After the view geometry constraints, cluster the remaining unmatched or mismatched feature points in two adjacent grayscale images, perform denoising processing after clustering, extract the clustered point sets with the number of feature points in each category greater than the number threshold, and initially identify them as the feature point sets on dynamic objects. Use the minimum convex deformation to enclose each dynamic object feature point set. When the convex hull area is less than the area threshold, it is considered a misjudgment, and the point sets with areas greater than the area threshold are retained.

[0012] Step 7: Superimpose the pixel positions where each convex hull is located on the binary image to remove the dynamic object area and obtain the remaining static area of the image.

[0013] Further, in the above Step 2, the specific process of performing Gaussian filtering on the image is as follows: perform weighted averaging on the entire image, specifically as follows:

[0014] The value of each pixel point in the image is obtained by weighted averaging of itself and other pixel values in its neighborhood. The Gaussian filtering formula is as follows:

[0015]

[0016] where σ is the standard deviation of the values of all pixel points in the neighborhood, and x and y are pixel coordinates.

[0017] Further, in the step 2, the specific process of FAST corner point homogenization is as follows: The entire image is divided into blocks. The corner points detected in each block are sorted according to the magnitude of their V values, and n corner points with V values greater than the detection threshold t are retained. When the number of corner points detected in a block is less than n, the detection threshold t of this block is reduced and the corner points are detected again; to prevent corner points from clustering in local areas, the adjacent point elimination strategy is used. The entire image is processed with a sliding window, and the corner point with the largest V value within the sliding window is retained, thereby eliminating adjacent points and avoiding corner point clustering; the calculation formula of the corner point score function is as follows:

[0018]

[0019] where p is the central pixel value, t is the detection threshold, pixel values are pixel values, corresponding to N adjacent pixels of the pixel center.

[0020] Further, in the step 3, for feature point matching, the displacement estimation formula of two feature points is as follows:

[0021]

[0022] where is the estimated displacement, f(x p ,y p ,t k ) represents the p-th feature point of the k-th frame image, x p and y p respectively represent the x coordinate and y coordinate of the p-th feature point in the image pixel coordinate system, t k represents the k-th frame image in the time series of the image stream, d1 and d2 are the offsets of the feature point in the image pixel coordinate system, and B represents the set of feature points of this frame.

[0023] The determination formula for the matching of two feature points is as follows:

[0024]

[0025] where t is the set determination threshold. When T = 1, it indicates that the feature points match.

[0026] Further, in the step 4, the specific process of clustering the background optical flow vector field is as follows: According to the fact that the optical flows of points on the same target are approximately the same, the optical flows are clustered; the measure function for measuring the similarity of two optical flows (u1, v1) and (u2, v2) is:

[0027]

[0028] where η is the order of the error. Set a relatively small threshold D thOnce the measure function D between two optical flow vectors is less than D th these two optical flows will be merged into one class, and the average optical flow of the class will be calculated; when determining whether an optical flow vector can be merged into a certain class, it is necessary to measure the measure function between the optical flow and the average optical flow of the class.

[0029] Further, in step 4, the specific steps of the method for removing false matches based on intensity and direction consistency are as follows:

[0030] (1) Let the set of the lengths of the connecting lines of the matching point pairs be map d , and the set of its slopes be map k ;

[0031] (2) Statistically analyze the length d of the connecting line of each pair of matching point pairs ij ∈map d and the slope k ij ∈map k . Select the lengths in map d for segmented statistics, and select the average value of the length segment with the largest number of length samples as the length reference, denoted as D. Select the slopes in map k for segmented statistics, and select the average value of the slope segment with the largest number of slope samples as the slope reference, denoted as K;

[0032] (3) Set the error range of the matching connecting line length as m; set the error range of the matching connecting line slope as n.

[0033] (4) When D - m ≤ d ij ≤ D + m, retain the connecting line of this matching point pair, otherwise delete it; similarly, when K - n ≤ k ij ≤ K + n, retain the connecting line of this matching point pair, otherwise delete it.

[0034] Further, in step 5, the epipolar geometric constraint of the view geometric constraint is described as:

[0035] For any pixel point on image I1 its corresponding epipolar line l on image I2 i : A i x + B i y + C i =0 and the matching point satisfy the following formula:

[0036]

[0037] where d is the set threshold and F is the fundamental matrix, in and are respectively the jth matching point of the latter frame In the image pixel coordinate system, the x-coordinate and y-coordinate are expressed in a normalized manner due to the scale uncertainty of the depth information.

[0038] Furthermore, in step 6, the specific operation of denoising the clustered feature point set is as follows:

[0039] Each cluster point includes k feature points (x1, y1), (x2, y2), ..., (x i ,y i ),…,(x k ,y k ), then the estimated value of (μ,σ) is calculated according to the following formula:

[0040]

[0041]

[0042] d i =(x i -c x ) 2 +(y i -c y ) 2 ,

[0043]

[0044] Where μ represents the mean of the normal distribution, σ represents the standard deviation of the normal distribution, if d i ≥(μ+3σ), then delete the point (x i ,y i ).

[0045] In a second aspect, the present invention provides an image static area extraction device based on optical flow and view geometry constraints, comprising a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, the image static area extraction method based on optical flow and view geometry constraints is implemented.

[0046] In a third aspect, the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the image static area extraction method based on optical flow and view geometry constraints.

[0047] The beneficial effects of the present invention are as follows: In view of the situation that traditional SLAM cannot overcome the interference of dynamic targets, the SLAM method for removing dynamic targets based on RGBD sensors proposed by the present invention effectively overcomes the disadvantages of traditional SLAM methods, which are unstable and have low accuracy under the interference of dynamic targets, and effectively improves the accuracy under the interference of dynamic targets. At the same time, the method has universality and can achieve real-time effects. In terms of pixel-level segmentation methods, the method proposed by the present invention has a small amount of calculation and good segmentation effects. Different from the mainstream method of using deep learning to process dynamic targets, the present invention has the advantages of small calculation amount, good real-time performance, and good effect of removing dynamic targets, which is conducive to deployment and application on mobile robots. It can effectively improve the accuracy and robustness of tracking using RGBD sensors, and improve the positioning and mapping accuracy of visual SLAM in dynamic scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic structural diagram of the method for extracting static regions of an image based on optical flow and view geometry constraints of the present invention.

[0049] Figure 2 It is a schematic diagram of the optical flow filtering method based on intensity and direction consistency of the present invention.

[0050] Figure 3 It is a schematic diagram of the relationship between 3D points and images in consecutive frames in the epipolar constraint.

[0051] Figure 4 It is a structural diagram of a device for extracting static regions of an image based on optical flow and view geometry constraints of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0053] Due to the static assumption of traditional simultaneous localization and mapping technology, it is difficult for visual SLAM technology to be widely applied in real scenarios. After realizing the extraction of the static part of the environment and excluding the interference of dynamic objects, it can be more effectively deployed in the real environment.

[0054] The present invention provides a method for extracting static regions of an image based on optical flow and view geometry constraints. The acquired image is pre-filtered and feature corner points with direction information are extracted. After preprocessing the feature point set, the feature points are matched, the sparse optical flow method is used to calculate the optical flow vector field, and the static feature points are initially separated. Then, the static feature points are further verified using epipolar geometry constraints to optimize the separation result of the static feature points. Finally, the non-static feature points are clustered, their pixel regions are divided, and after image morphological processing, the static region of the image is output and provided for visual SLAM for positioning and map construction. This method has a small computational amount and good generality for visual SLAM, removes the interference of dynamic targets, can be applied to mobile robots in real time, and is beneficial to the stable operation of the SLAM system. As Figure 1 shown, the method includes the following steps:

[0055] Step 1: A color image is acquired through a sensor, and the three-channel color image is converted into a grayscale image.

[0056] The specific process is as follows: The sensor video stream is read as a color image sequence, the image size is set to 640x480, and each frame of the three-channel color image is converted into a single-channel grayscale image through the open-source image processing opencv library.

[0057] Step 2: After Gaussian filtering each frame of the grayscale image, feature extraction is performed. At the same time, image pyramid processing is used to extract FAST corner points and perform homogenization processing. The FAST corner point direction information is calculated using the gray centroid method. The FAST corner point and its direction information form a feature point, and the feature points extracted from each frame of the grayscale image form a feature point set.

[0058] Further, in step 2, the following sub-steps are specifically included:

[0059] Step 2-1: The entire image is weighted-averaged. The value of each pixel point is obtained by weighted-averaging itself and other pixel values in the neighborhood. In the present invention, the adjacent neighborhood is set as a 7×7 grayscale image block centered on the pixel. The Gaussian filtering formula is as follows:

[0060]

[0061] where σ is the standard deviation of the values of all pixel points in the neighborhood, and x, y are pixel coordinates.

[0062] Step 2-2: Remove the meaningless image edges around. The image edge part contains the diffusion effect at the energy-free places in the surrounding space, and the information reflected by the image is inaccurate.

[0063] Step 2-3: Scale the image size and use the Gaussian image pyramid for image downsampling. Since the assumption of continuous image change restricts the pose transformation amount of sequence images for estimating motion to be small, which does not conform to most scenarios in the actual process, using the Gaussian image pyramid can effectively solve this problem. After decomposing the image into low-resolution layers, the motion amount between sequence frame images will become small enough to meet the constraint conditions of optical flow calculation, enabling the detection of sequence images with rapid changes.

[0064] Step 2-4: Extract FAST corner points and perform feature homogenization processing. To ensure the uniform distribution of feature points, divide the entire image into blocks. The corner points detected in each block are sorted according to the magnitude of their V values, and the n corner points with relatively larger V values are retained. When the number of corner points detected in a block is less than n, reduce the detection threshold t for that block and re-detect the corner points. Thus, the corner points with significant features in each block are detected and retained, initially achieving the uniformity of corner point distribution. However, the phenomenon of corner point clustering in local areas may still occur, so a neighboring point elimination strategy is added, that is, use a 7×7 sliding window to process the entire image. If there is more than one corner point in the sliding window, only retain the corner point with the largest V value, thereby eliminating neighboring points and avoiding corner point clustering. The calculation formula for the corner point detection score function is as follows:

[0065]

[0066] where p is the central pixel value, t is the detection threshold, pixel values are pixel values corresponding to N adjacent pixels of the pixel center.

[0067] Step 2-5: Use the gray centroid method to calculate the direction information of each FAST corner point to complement the missing rotational invariance of the FAST corner points. The specific method is: taking the feature point as the center, calculate the centroid position within its neighborhood N, and then construct a vector with the feature point as the starting point and the centroid as the ending point. The direction of this vector is the direction of this feature point.

[0068] Step 3: Filter the feature point sets of two adjacent grayscale images to be matched through the direction information and Euclidean distance difference of the feature points between two adjacent grayscale images, reducing the search range of matching.

[0069] Step 3-1: Use the direction information and Euclidean distance difference of feature points to pre-filter the feature point sets to be matched to reduce the amount of matching calculation. First, set the Euclidean distance radius for feature point matching. When the pixel distance between two feature points in their respective frame images is greater than the set threshold, the feature points to be matched are removed from the point set. When the main direction angle difference between two feature points is greater than the set threshold, the feature points to be matched are removed from the point set.

[0070] Step 3-2, Feature point matching. First, for the feature points F i ,y j ,t k ) on the image f(x q , find all possible matching points F i ,y j ,t k+1 ) of the image f(x p ). Then, according to the matching criterion, find the most suitable point, calculate the disparity between the two points, and then calculate the optical flow of the feature points according to the sampling frequency.

[0071] Assume that the variation range of the disparity (d1, d2) is d1 ∈ [-r, r], d2 ∈ [-r, r]. Then each feature point in the image f(x i ,y j ,t k ) should search within the following range of the image f(x i ,y j ,t k+1 ), that is, a square area centered on (x p ,y p ) with a perimeter of 2r. The feature points within the square area are used as the initial matching candidate feature points to find all possible matching relationships.

[0072] Step 3-3, Use the following matching relationship for feature point matching. The displacement estimation formula for two feature points is as follows:

[0073]

[0074] where is the estimated displacement, f(x p ,y p ,t k ) represents the p-th feature point of the k-th frame image, x p and y p respectively represent the x coordinate and y coordinate of the p-th feature point in the image pixel coordinate system, t k represents the k-th frame image in the time series of the image stream, d1 and d2 are the offsets of the feature points in the image pixel coordinate system, and B represents the set of feature points of this frame.

[0075] The decision formula for the matching points is as follows:

[0076]

[0077] where t is the set decision threshold. When T = 1, it means the feature points are matched.

[0078] Step 4: Perform optical flow calculation on the filtered feature point sets of two adjacent frames to obtain an optical flow vector field; for the calculated optical flow vector field, remove the mismatches based on intensity and direction consistency to obtain a background optical flow vector field, and at the same time obtain the feature point set of the background optical flow vector field.

[0079] Step 4-1: Calculate the optical flow field. According to the feature point sets matched in the above steps, use the Lucas-Kanade (LK) optical flow to calculate the sparse optical flow field. For a pixel located at (x, y) at time t, its grayscale can be written as I(x, y, t). For the pixel located at (x, y) at time t, assume that at time t + dt, it moves to (x + dx, y + dy). The following formula can be obtained:

[0080] I(x + dx, y + dy, t + dt) = I(x, y, t)

[0081] Perform Taylor expansion on the left side and retain the first-order terms. The formula can be transformed into the following formula:

[0082]

[0083] Assume that the grayscale value remains unchanged, then

[0084]

[0085] where dx / dt is the moving speed of the pixel on the x-axis, and dy / dt is the speed on the y-axis. Denote them as u and v. At the same time is the gradient of the image in the x direction at this point, and the other term is the gradient in the y direction, denoted as I x , I y . Denote the change in the grayscale image with respect to time as I t , and write it in matrix form as:

[0086]

[0087] In the LK optical flow method, we assume that the pixels within a certain window have the same motion. Consider a window of size w×w, which contains w 2 number of pixels. Since the pixels within this window have the same motion, we thus have w 2 equations:

[0088]

[0089] The equation can be rewritten as:

[0090]

[0091] Through the above equations, the optical flow motion vector field of the image feature points can be solved.

[0092] Step 4-2: Remove mismatches based on length and direction consistency. In the results of initial feature matching, there are often some incorrect matching connection lines caused by mismatches. For the matching point p 1i (x,y) in image I1(x,y) and its corresponding matching point p 2j (u,v) in image I2(u,v) have the following two characteristics: First, the distance d ij between correct matching point pairs is equal; Second, the slope k ij of the straight line determined by correct matching point pairs is equal.

[0093] As Figure 2 shown, two solid dots are a pair of matching points, and the line segment is the matching connection line between the matching points. From the geometric characteristics of the matching connection line, it can be seen that the distances between each pair of matching points are equal and the slopes are equal, which are represented by black solid line segments in Figure 2 . Figure 2 In

[0094] , the two virtual line segments are removed because their distances are significantly longer than those of the black matching point pairs and their slopes are significantly different from those of the black matching point pairs. Considering the errors in actual calculations, the distances and slopes only need to be approximately equal. 1i (x,y) and p 2j (u,v), the distance d ij is calculated as follows:

[0095]

[0096] The slope k ij of the two points is calculated as follows

[0097]

[0098] where i ∈ [1, n], j ∈ [1, n], and n is a positive integer, which is the number of matching point pairs.

[0099] The specific steps of this method are as follows:

[0100] (1) Let the set of the lengths of the connection lines of the matching point pairs be map d , and its slope set be map k ;

[0101] (2) Count the length d ij of each pair of connection lines of the matching point pairs ∈ map d , slope k ij ∈ map k . Select the lengths in map d for segmented statistics, and select the average value of the length segment with the largest number of length samples as the length benchmark, denoted as D. Select map kPerform piecewise statistics on the slopes in it, and select the average value of the slope segment with the largest number of slope samples as the benchmark of the slope, denoted as K;

[0102] (3) Set the error range of the matching connection length to m; set the error range of the matching connection slope to n.

[0103] (4) When D - m ≤ d ij ≤ D + m, retain the connection of this matching point pair, otherwise delete it; similarly, when K - n ≤ k ij ≤ K + n, retain the connection of this matching point pair, otherwise delete it.

[0104] Step 4-3: According to the fact that the optical flows of points on the same target are approximately the same, cluster the optical flows. Among them, the background optical flow has the characteristic of a large optical flow area. At the same time, combined with the actual scene, the relative distribution closer to the image edge is more likely to be the background optical flow. The measure function for measuring the similarity of two optical flows is:

[0105]

[0106] Among them, η is the order of the error. Set a relatively small threshold D th , once the measure function D < D th between two optical flow vectors, merge these two optical flows into one class and calculate the average optical flow of the class. When judging whether an optical flow vector can be merged into a certain class, it is necessary to measure the measure function of the optical flow and the average optical flow of the class.

[0107] Step 5: Further verify the matching situation for the feature point set of the background optical flow vector field using view geometry constraints, and use the Progressive Sampling Consensus (PROSAC) method to solve the essential matrix for further optimizing the matching result through epipolar geometry verification.

[0108] Step 5-1: Use the PROSAC method to solve the essential matrix. The PROSAC algorithm is an improved algorithm of RANSAC, mainly used for the estimation problem of image matching models. It is a sorting method based on the sample selection strategy. The algorithm can effectively reduce the randomness in the sampling process, improve the sampling probability of correct sample data, thereby reducing the number of iterations of the algorithm and improving the timeliness.

[0109] Step 5-2, Epipolar geometry verification. Due to the movement of the video acquisition device, both moving objects and the actually stationary background in the video sequence will move, but their motion patterns are different. That is, between successive frames, the pixel points of the moving object and the background have different displacement vectors. The background is stationary. Relative to the camera, the displacement amounts of most background corner points at the same moment between successive frames can be regarded as consistent. Most background corner points in successive frames generally satisfy the epipolar geometry constraint conditions. The corner points corresponding to the moving object usually do not satisfy the epipolar geometry constraint conditions due to the large displacement amounts.

[0110] The epipolar geometry constraint can be expressed as Figure 3 shown in the figure. For two consecutive image frames I1 and I2, O1 and O2 are the optical centers of the two image frames respectively, and the line connecting O1 and O2 is called the baseline. The projections of the 3D point P on frames I1 and I2 are p1 and p2 respectively. The plane determined by P, O1, and O2 is the epipolar plane π, and the intersection lines of π with the image planes I1 and I2 are the epipolar lines, which are l1 and l2 in the figure. The intersection points of the baseline with the image planes I1 and I2 are the epipoles e1 and e2.

[0111] As Figure 3 shown in the figure, assuming that the feature point p1 in I1 is known and we try to find the corresponding feature point p2 in I2. At this time, if the depth information is unknown, we can only infer that point P is on the extension line of O1, p1, and its specific position is unknown. From this, we can infer the corresponding movement of p2 on the epipolar line l2 in I2, which is called the epipolar constraint. The physical meaning of the epipolar constraint is that O1, P, and O2 are coplanar. In addition, it also describes the spatial position relationship between two matching points, which can be expressed by the fundamental matrix F as:

[0112] p2 T Fp1 = 0

[0113] where p1 and p2 are the pixel positions corresponding to two pixel points in the image respectively. First, use the basic 8-point method as the initial value to estimate the fundamental matrix F, and use the PROSAC method to evaluate the F obtained by each sampling until a stable fundamental matrix is obtained. Based on the calculated fundamental matrix, perform multiple iterations, continuously filtering out the assumed inliers that are too far from the epipolar line, that is, trying to make the proportion of background corner points among all assumed inliers higher, so that more background corner points in the estimated fundamental matrix satisfy the constraint, that is, improve the accuracy of the fundamental matrix. Until the sum of the distances of all inliers from the epipolar line is less than a certain threshold, use the obtained inlier set to calculate the finally corrected fundamental matrix F. After solving the fundamental matrix F, it can be determined whether any pair of feature points p1, p2 in the image frame belong to dynamic or static according to whether they satisfy the epipolar line constraint.

[0114] For any pixel point on image I1 its corresponding epipolar line l on image I2 i :Ai x + B i y + C i = 0 and the matching points satisfy the following formula:

[0115]

[0116] where d is a set threshold, and F is the fundamental matrix, in, and are respectively the x - coordinate and y - coordinate of the j - th matching point in the subsequent frame in the image pixel coordinate system. 1 is the expression for normalization due to the scale uncertainty of depth information.

[0117] After view geometric constraints, cluster the remaining unmatched or mismatched feature points in two adjacent grayscale images. After clustering, perform denoising processing, extract the point sets after clustering where the number of feature points in each class is greater than the quantity threshold, and preliminarily identify them as the feature point sets on dynamic objects. Use the minimum convex deformation to enclose each dynamic object feature point set. When the convex hull area is less than the area threshold, it is identified as a misjudgment, and retain the point sets larger than the area threshold;

[0118] Step 6 - 1: Cluster the remaining unmatched or mismatched feature points according to the method described in Step 4 - 2.

[0119] Step 6 - 2: The feature points in each class obtained from the previous clustering should be similar, but due to the influence of noise in the actual environment, there may be noise points. These noise points generally show as "outliers". Remove these outlier points. We assume that the distances from the points in the same class to the class center follow a normal distribution with parameters (μ, σ). For the points whose distances to the class center are greater than or equal to (μ, 3σ), delete them from this class. The specific operation is as follows:

[0120] Each class of clustered points includes k feature points (x1, y1), (x2, y2), …, (x i , y i ), …, (x k , y k ). Then the estimated values of (μ, σ) are calculated according to the following formula:

[0121]

[0122]

[0123] d i = (x i - c x ) 2 + (yi -c y ) 2 ,

[0124]

[0125] where μ represents the mean of the normal distribution, σ represents the standard deviation of the normal distribution. If d i ≥(μ + 3σ), then the point (x i , y i ) is deleted from this class.

[0126] Step 6 - 3: Use the minimum convex hull algorithm to generate a minimum convex polygon around each class of feature points. Before generating the minimum convex polygon, judge the number of feature points in each class after clustering. If the number of feature points is greater than the set threshold, it is regarded as a dynamic object; if it is less than the set threshold, it is a misjudgment. Use the minimum convex polygon to judge the size of the convex polygon area. If the area is less than the set area, it is regarded as a misjudgment and ignored, and the point set with an area greater than the area threshold is retained.

[0127] Step 7: Superimpose the pixel positions where each convex hull is located onto the binary image, and perform a closing operation of image morphology on the binary image to cover the edges of the dynamic objects and remove the dynamic object areas, obtaining the remaining static area of the image.

[0128] Corresponding to the embodiment of the method for extracting the static area of an image based on optical flow and view geometry constraints described above, the present invention also provides an embodiment of an apparatus for extracting the static area of an image based on optical flow and view geometry constraints.

[0129] See Figure 4 , an apparatus for extracting the static area of an image based on optical flow and view geometry constraints provided by an embodiment of the present invention includes a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it is used to implement the method for extracting the static area of an image based on optical flow and view geometry constraints in the above embodiment.

[0130] The embodiment of the apparatus for extracting the static area of an image based on optical flow and view geometry constraints of the present invention can be applied to any device with data processing capabilities. The any device with data processing capabilities can be a device or apparatus such as a computer. The apparatus embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful apparatus, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non - volatile memory into the memory for operation. From the hardware level, as Figure 4 shown, it is a hardware structure diagram of any device with data processing capabilities where the apparatus for extracting the static area of an image based on optical flow and view geometry constraints of the present invention is located. Except for Figure 4In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities where the device in the embodiment is located may generally include other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated herein.

[0131] The implementation processes of the functions and roles of each unit in the above device are specifically described in detail in the implementation processes of the corresponding steps in the above method, which will not be elaborated herein.

[0132] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0133] The embodiment of the present invention further provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the method for extracting static regions of an image based on optical flow and view geometry constraints in the above embodiments.

[0134] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by the device with data processing capabilities, and can also be used to temporarily store the data that has been output or will be output.

[0135] The above are only the preferred embodiments of the present invention. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention by using the methods and technical content disclosed above, or modify it into equivalent embodiments with equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the technical solution of the present invention still fall within the scope of protection of the technical solution of the present invention.

Claims

1. An image static region extraction method based on optical flow and view geometry constraints, characterized in that, The method includes the following steps: Step 1: Convert a plurality of color images to be processed into grayscale images; Step 2: After performing Gaussian filtering on each frame of grayscale image, perform feature extraction, and at the same time use image pyramid processing to extract FAST corner points and perform homogenization processing. Use the gray centroid method to calculate the direction information of FAST corner points. The FAST corner points and their direction information form a feature point, and the feature points extracted from each frame of grayscale image form a feature point set; Step 3: Filter the feature point sets of two adjacent frames of grayscale images to be matched through the direction information and Euclidean distance difference of the feature points of two adjacent frames, and reduce the search range of matching; Step 4: Perform optical flow calculation on the filtered feature point sets of two adjacent frames to obtain an optical flow vector field; for the calculated optical flow vector field, remove false matches based on intensity and direction consistency, cluster to obtain a background optical flow vector field, and at the same time obtain a feature point set of the background optical flow vector field; Step 5: Further verify the matching situation for the feature point set of the background optical flow vector field through view geometry constraints, use the Progressive Sampling Consensus (PROSAC) method to solve the essential matrix, and perform epipolar geometry verification to further optimize the matching result; the epipolar geometry constraint of view geometry constraint is described as: For any pixel point on the image I1 its corresponding epipolar line l on the image I2 i :A i x + B i y + C i = 0 and the matching point satisfy the following formula: where d is the set threshold value, and F is the fundamental matrix, in and are respectively the x - coordinate and y - coordinate of the j - th matching point in the subsequent frame in the image pixel coordinate system. 1 is the expression for normalization due to the scale uncertainty of the depth information; Step 6: After view geometry constraints, cluster the remaining unmatched or mismatched feature points in two adjacent frames of grayscale images, perform denoising processing after clustering, extract the clustered point sets with the number of feature points in each category greater than the number threshold, and initially identify them as feature point sets on dynamic objects. Use the minimum convex deformation to enclose each dynamic object feature point set. When the convex hull area is less than the area threshold, it is considered a misjudgment, and the point set with an area greater than the area threshold is retained; Step 7: Superimpose the pixel positions where each convex hull is located on the binary image to remove the dynamic object area and obtain the remaining static area of the image.

2. The image static region extraction method based on optical flow and view geometry constraints according to claim 1, characterized in that, In the said Step 2, the specific process of performing Gaussian filtering on the image is as follows: perform weighted averaging on the entire image, specifically as follows: The value of each pixel point in the image is obtained by weighted averaging of itself and other pixel values in the neighborhood. The Gaussian filtering formula is as follows: where σ is the standard deviation of the values of all pixel points in the neighborhood, and x, y are pixel coordinates.

3. The image static region extraction method based on optical flow and view geometry constraints according to claim 1, characterized in that, In the said Step 2, the specific process of FAST corner point homogenization is as follows: Divide the entire image into blocks, sort the detected corner points in each block according to the magnitude of their V values, and retain n corner points with V values greater than the detection threshold t. When the number of detected corner points in a block is less than n, reduce the detection threshold t for this block and re-detect the corner points; To prevent corner points from clustering in local areas, use the adjacent point elimination strategy, process the entire image with a sliding window, and retain the corner point with the largest V value within the sliding window, thereby eliminating adjacent points and avoiding corner point clustering; the calculation formula of the corner point score function is as follows: where p is the central pixel value, t is the detection threshold, pixel values are pixel values, corresponding to N adjacent pixels of the pixel center.

4. The image static region extraction method based on optical flow and view geometry constraints according to claim 1, characterized in that, In step 3, feature point matching is performed, and the displacement estimation formula for two feature points is as follows: Among them To estimate the displacement, f(x p , y p , t k ) represents the p-th feature point of the k-th frame image, where xp and yp represent the x-coordinate and y-coordinate of the p-th feature point in the image pixel coordinate system, and t k represents the k-th frame image in the image sequence over time, d1 and d2 are the offsets of the feature point in the image pixel coordinate system, and B represents the set of feature points of this frame; The determination formula for the matching of two feature points is as follows: where t is the set determination threshold. When T = 1, it indicates that the feature points are matched.

5. The image static region extraction method based on optical flow and view geometry constraints according to claim 1, characterized in that, In step 4, the specific process of clustering the background optical flow vector field is as follows: According to the fact that the optical flows of points on the same target are approximately the same, the optical flows are clustered; the measure function for measuring the similarity of two optical flows (u1, v1) and (u2, v2) is: where η is the order of the error; a relatively small threshold D is set th , once the measure function D between two optical flow vectors is less than D th these two optical flows are combined into one class, and the average optical flow of the class is calculated; when determining whether an optical flow vector can be merged into a certain class, it is necessary to measure the measure function of this optical flow and the average optical flow of the class.

6. The image static region extraction method based on optical flow and view geometry constraints according to claim 1, characterized in that, In step 4, the specific steps of the method for removing false matches based on intensity and direction consistency are as follows: (1) Let the set of the lengths of the connecting lines of the matching point pairs be map d , and the set of its slopes be map k ; (2) Statistically calculate the length d of the line connecting each pair of matching points ij ∈map d and the slope k ij ∈map k , select map d to perform segmented statistics on the lengths, select the average value of the length segment with the largest number of length samples as the benchmark for the length, denoted as D, and select map k to perform segmented statistics on the slopes, select the average value of the slope segment with the largest number of slope samples as the benchmark for the slope, denoted as K; (3) Set the error range of the length of the matching connection line to m; set the error range of the slope of the matching connection line to n; When D - m ≤ d ij ≤ D + m, retain the connection line of this pair of matching points, otherwise delete it; similarly, when K - n ≤ k ij ≤ K + n, retain the connection line of this pair of matching points, otherwise delete it.

7. A method for extracting static regions of an image based on optical flow and view geometry constraints, characterized in that, In step 6, the specific operation for denoising the set of feature points after clustering is as follows: Each class of clustering points includes k feature points (x1, y1), (x2, y2), …, (x i , y i ), …, (x k , y k ). Then, the estimated value of (μ, σ) is calculated according to the following formula: d i = (x i - c x ) 2 + (y i - c y ) 2 , where μ represents the mean of the normal distribution and σ represents the standard deviation of the normal distribution. If d i ≥(μ + 3σ), then the point (x i , y i ) is removed from the class.

8. An apparatus for extracting static regions of an image based on optical flow and view geometry constraints, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the processor executes the executable code, it implements the method for extracting the static region of an image based on optical flow and view geometry constraints as described in any one of claims 1-7.

9. A computer-readable storage medium, on which a program is stored, characterized in that, When the program is executed by the processor, it implements the method for extracting the static region of an image based on optical flow and view geometry constraints as described in any one of claims 1-7.

Citation Information

Patent Citations

  • SLAM method based on improved semantic optical flow method in dynamic environment

    CN112446885A

  • Visual positioning method and device based on visual map

    WO2022002039A1