A dynamic feature elimination method and its corner detection algorithm

The dynamic feature removal method using Mask-RCNN and geometric analysis addresses the challenge of dynamic objects in feature extraction, improving the accuracy of visual SLAM and 3D reconstruction by isolating static features from dynamic ones.

CN113256663BActive Publication Date: 2025-05-23SHENZHEN YIJIAHE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110520048.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-13
Publication Date
2025-05-23
Estimated Expiration
2041-05-13

AI Technical Summary

Technical Problem

Existing feature extraction algorithms like Harris, SIFT, and SURF are designed for static environments and fail to effectively handle dynamic objects, which interfere with feature detection in dynamic scenes, posing challenges for applications such as visual SLAM and 3D reconstruction.

Method used

A dynamic feature removal method using Mask-RCNN for object segmentation, combined with geometric analysis of multiple frames, to identify and exclude dynamic features by creating a mask for static regions, employing algorithms like Harris, SIFT, and SURF for point detection.

Benefits of technology

Effectively removes dynamic features in dynamic environments, enhancing the stability of visual SLAM and 3D reconstruction by accurately distinguishing and excluding dynamic objects and their associated features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113256663B_ABST
    Figure CN113256663B_ABST
Patent Text Reader

Abstract

The present invention discloses a dynamic feature elimination method and a corner detection algorithm thereof, a dynamic feature elimination method, comprising the steps of: (1) segmenting the moving object in the image by Mask-RCNN algorithm; (2) obtaining the dynamic object associated with the moving object obtained in step (1) by image RGB-D data and multi-view geometry algorithm; (3) combining steps (1) and (2) to obtain the dynamic features of the image, and setting the pixel points of all dynamic feature areas to 0, and performing pixel expansion at the edge of the area. The present invention integrates a variety of methods to provide important support for visual slam or visual three-dimensional reconstruction. The present invention combines AI methods with traditional geometric methods to solve the influence of dynamic objects on feature extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to a dynamic feature elimination method and a corner point detection algorithm thereof. Background Art

[0002] The current mainstream feature extraction algorithms include Harris corner point, SIFT, SURF, Fast and other corner point detection algorithms. These algorithms are based on the traditional method framework, that is, they assume that the environment is static, and all points in a certain area (window) have the same running trend, and the elements in the window have the same gradient in the x and y directions, so the application scope of these points is mostly static scenes. The starting point is to find representative pixels in a single image frame by frame.

[0003] The processes of traditional feature extraction algorithms are slightly different. Let's take the Harris corner point extraction algorithm as an example to illustrate: The core of the Harris algorithm is to find the point where the gradient change of the image pixel in the x and y directions is the most obvious in a local area. The process of the Harris algorithm is:

[0004] Step 1. Use horizontal and vertical difference operators to filter each pixel of the image to obtain the gradients in the x and y directions and construct the matrix M. Step 2. Perform Gaussian smoothing filtering on the four elements of M to eliminate some unnecessary isolated points and protrusions to obtain a new matrix M. Step 3. Next, use M to calculate the corner point response function R corresponding to each pixel. Step 4. Suppress local maximum values ​​and select their maximum values ​​at the same time. Perform non-maximum suppression and expansion in a fixed area of ​​the point to obtain the matrix R. Step 5. In the matrix R, if R(i,j) is greater than a certain threshold and R(i,j) is a local maximum in a certain field, it is considered to be a corner point.

[0005] Through the above steps, we can extract the points in the image where the feature changes are more obvious. The premise is that it is in a static environment. When calculating the response function R, we assume that the points in the window have the same movement trend, and do not consider the interference of dynamic objects. If there are dynamic objects introduced into the system, the calculation of the above response function R will be affected. At the same time, dynamic objects themselves are a kind of interference. The features extracted from them are also motion features, which need to be removed in some application scenarios.

[0006] In visual SLAM or visual 3D reconstruction, dynamic feature points brought by dynamic objects pose a great challenge to the stability of the system. In order to remove dynamic feature points, exploring static feature point extraction algorithms has practical application significance. Summary of the invention

[0007] Purpose of the invention: In view of the above-mentioned shortcomings, the present invention combines the characteristics of traditional feature extraction algorithms that only rely on a single frame to propose a dynamic feature elimination method and its corner detection algorithm, which is based on multi-frame image data and assisted by an AI segmentation algorithm to obtain a mask of non-dynamic areas for dynamic environment feature extraction.

[0008] Technical solution:

[0009] A dynamic feature elimination method comprises the steps of:

[0010] (1) Segment the moving objects in the image using the Mask-RCNN algorithm;

[0011] (2) obtaining a dynamic object associated with the moving object obtained in step (1) through image RGB-D data and multi-view geometry algorithm;

[0012] (3) Combine steps (1) and (2) to obtain the dynamic features of the image, set all pixels in the dynamic feature area to 0, and perform pixel expansion at the edge of the area.

[0013] The step (2) is specifically:

[0014] (21) selecting a number of frames of image data as key frames for determining dynamic features;

[0015] (22) Find the correspondence between the corresponding feature points on each frame, that is, the position x′ of the feature point x in each key frame in the current frame;

[0016] (23) According to the projection relationship and the change of the camera's posture between each frame, the depth Zproj of the feature point x′ in space is obtained, and the angle α corresponding to each feature point in the moving process is calculated. If α is greater than 30°, the feature point is discarded;

[0017] (24) Based on the pixel correspondence, the depth Z′ of the feature point x′ is obtained from the depth data collected by the camera;

[0018] (25) The projection error of the feature point is calculated according to ΔZ=Z′-Zproj, and whether the feature point is blocked is determined according to whether ΔZ is within the error threshold τz, that is, whether the feature point is on a dynamic object.

[0019] In the step (21), 2 to 5 frames of image data are selected as key frames.

[0020] In the step (25), the error threshold τz is set to 0.5 m.

[0021] A corner point detection algorithm using the above-mentioned dynamic feature elimination method is characterized by comprising the steps of:

[0022] Step 1, using the above-mentioned dynamic feature elimination method to eliminate the dynamic features in the image;

[0023] Step 2: Perform corner detection on the image obtained in Step 1 using a corner detection algorithm.

[0024] The corner detection algorithm adopts Harris corner, SIFT, SURF or Fast.

[0025] Beneficial effects: In the substation scene, people often pass by. People and the operating devices in their hands may be moving objects. When processing such scenes, visual slam needs to remove the feature points corresponding to the moving people and the books in their hands. The present invention combines multiple methods to provide important support for visual slam or visual three-dimensional reconstruction. The present invention combines AI methods with traditional geometric methods to solve the impact of dynamic objects on feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is the algorithm flow chart of the present invention.

[0027] Figure 2 This is the Mask-RCNN network structure diagram.

[0028] Figure 3 is an example image for corner detection using existing technology.

[0029] Figure 4 This is an example image for performing corner detection using the corner detection algorithm of the present invention. DETAILED DESCRIPTION

[0030] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments.

[0031] Figure 1 It is the algorithm flow chart of the present invention. Figure 1 As shown, the dynamic feature elimination method of the present invention comprises the steps of:

[0032] (1) Perform semantic segmentation on the image through Mask-RCNN to segment the moving objects in the image;

[0033] Among them, the Mask-RCNN network structure diagram is as follows Figure 2 As shown in the figure, Mask-RCNN is an instance segmentation algorithm. Mask-RCNN is a very flexible framework that can add different branches to complete different tasks. It can complete multiple tasks such as target classification, target detection, semantic segmentation, instance segmentation, and human posture recognition.

[0034] The Mask-RCNN network is decomposed into the following three modules: Faster-rcnn, ROIAlign and FCN;

[0035] Faster-rcnn is a baseline algorithm in the field of target detection. It consists of four core modules: feature extraction network, ROI generation, ROI classification, and ROI regression. The definitions of each module are as follows: Feature extraction network: Its function is to obtain important features of different targets in the image. It is generally composed of convolutional layers, activation functions, and pooling layers. Some pre-trained networks (VGG, Inception, Resnet, etc.) are often used. The result obtained is called a feature map; ROI generation: Make multiple candidate ROIs (here 9) at each point of the obtained feature map, and then use the classifier to distinguish these ROIs into background and foreground, and use the regressor to make preliminary adjustments to the positions of these ROIs; ROI classification: In the RPN stage, it is used to distinguish between foreground (overlapping with the real target and its overlapping area is greater than 0.5) and background (not overlapping with any target or its overlapping area is less than 0.1); in the Fast-rcnn stage, it is used to distinguish different types of targets (cats, dogs, people, etc.); ROI regression: In the RPN stage, preliminary adjustments are made, and precise adjustments are made in the Fast-rcnn stage;

[0036] ROIAlign operation In order to obtain a feature map of a fixed size (for example, 7X7), ROIAlign technology does not use quantization operations. So how do we deal with these floating-point numbers? Our solution is to use the "bilinear interpolation" algorithm. Bilinear interpolation is a relatively good image scaling algorithm. It makes full use of the four real pixel values ​​around the virtual point in the original image (such as the floating-point number 20.56, the pixel positions are all integer values, there are no floating-point values) to jointly determine a pixel value in the target image, that is, the pixel value corresponding to the virtual position point 20.56 can be estimated.

[0037] The FCN algorithm is a representative of classic semantic segmentation algorithms. It can quickly and accurately segment targets in images. It contains multiple convolutional pooling layers and fully connected layers. When processing data, it first convolves and pools the image to continuously reduce the size of its feature map; then it performs a deconvolution operation, which is similar to restoring the size of the map. The continuously increasing feature map is used for the final pixel-level classification, thereby achieving accurate segmentation of the input image.

[0038] The Mask-RCNN algorithm is used to classify moving objects in the image and segment moving objects such as people, cats, and dogs in the image area.

[0039] (2) Determine the dynamic object part by combining the geometric model;

[0040] Combining geometric models to assist in determining dynamic objects is an important step in eliminating dynamic features, because the dynamic features in the image not only include the moving object itself, but also the physical objects associated with it may also move with the dynamic object. To find an effective method for eliminating dynamic features, it is necessary to consider whether the dynamic object holds other physical objects. The present invention uses RGB-D data and multi-view geometry algorithms to further eliminate dynamic objects. In step (1), the Mask-RCNN network has been used to segment possible moving objects in the image. Due to the actual visual slam usage scenario, dynamic objects only consider pedestrians (if the network is trained, other animals are also similarly added to the moving object set). Deeply understand the segmented dynamic object mask and the corresponding depth information at the pixel level, and determine whether there are static objects following the moving object, as follows:

[0041] Select several frames of image data as key frames to judge dynamic features. According to experience, when the number of image frames is greater than 5, a non-dynamic object may also be considered as a dynamic object due to the parallax generated by movement. Therefore, the present invention selects 2 to 5 frames of image data as key frames;

[0042] Assuming that the camera is moving, the feature point x on the image moves to the position x′. The two are triangulated to obtain the coordinate point X of the feature point in space, and the angle α of the corresponding movement of the line connecting the feature point on the image and the coordinate point X in space during the movement is calculated. It is believed that when the angle α is greater than 30°, the feature point will not participate in the dynamic feature judgment queue, because the parallax caused by the movement of the camera may cause non-dynamic objects to be mistakenly judged as dynamic objects.

[0043] Determine whether the feature point x is on a dynamic object, and determine whether the feature point is on a dynamic object based on the depth error of the corresponding feature point:

[0044] The depth of feature point x' calculated by the camera's motion posture is Zproj, and the depth of feature point x' found pixel by pixel by the depth camera is Z'. It can be predicted that if there is no occlusion of dynamic objects, the error ΔZ = Z'-Zproj of these two values ​​should be within the error threshold τz, and τz takes an empirical value of 0.5m.

[0045] According to the above description, the specific implementation method is:

[0046] Step 1, find the corresponding relationship between feature points between frames, that is, the position x′ of the feature point x in each key frame in the current frame;

[0047] Step 2: According to the projection relationship and the change of the camera's posture between each frame, the depth Zproj of the feature point x' in space is obtained, and the angle α corresponding to the feature point during the movement is calculated. If α is greater than 30°, it is discarded.

[0048] Step 3, based on the pixel correspondence, obtain the depth Z′ of the feature point x′ from the depth data collected by the camera;

[0049] Step 4, calculate the projection error of the feature point according to ΔZ=Z′-Zproj;

[0050] Step 5: Determine whether the feature is occluded based on the error threshold τz.

[0051] Determine dynamic objects through geometric models and segment out more parts that are likely to be dynamic objects;

[0052] (3) Feature extraction combined with static mask:

[0053] Through steps (1) and (2), the dynamic objects in the image have been segmented out relatively accurately, and a highly accurate mask matrix of static objects has been obtained. The mask after segmenting the dynamic objects will be combined with the traditional feature extraction algorithm to solve two problems: one is the assumption during feature extraction (the points in the window have the same motion trend, and the gradient changes in the x and y directions are consistent). If some points in the window are dynamic, the gradient change relationship does not hold; if all the points in the window are points on dynamic objects, the extracted features meet the assumptions, but the feature points are on dynamic objects, and such points also need to be eliminated.

[0054] The algorithm flow after synthesis is as follows Figure 1 As shown, more specifically:

[0055] Step 1. For the current frame image input, according to the mask matrix that removes moving objects and dynamic objects, set the pixels of all dynamic feature areas of the image to 0 to ensure that the feature points will not be in this area, and expand the pixels in the edge area to prevent the feature points from falling around the mask matrix;

[0056] The example of pixel expansion is as follows: for example, the pixels are 0, 0, 5, 6. 0, 0 is the result after dynamic object culling, and 5, 6 is the original pixel value. Then the expanded pixels are 0, 6, 5, 6, and the elements are folded at the edge.

[0057] After the elements of the dynamic object area of ​​the image are set to 0, feature points will appear on this edge. Because the gradient has changed dramatically, it is necessary to expand the elements in the edge area to prevent gradient jumps at the edge and avoid extracting feature points at the edge.

[0058] Step 2, use the horizontal and vertical difference operators to filter each pixel of the image to obtain the x and y direction gradients and construct the matrix M;

[0059] Step 3: Perform Gaussian smoothing filtering on the four elements of M to eliminate some unnecessary isolated points and protrusions and obtain a new matrix M.

[0060] Step 4. Next, use M to calculate the corner point response function R corresponding to each pixel. The Harris corner point response function R and the Shi-Tomasi corner point response function R are different in expression.

[0061] Step 5, local maximum suppression, select its maximum value at the same time, perform non-maximum suppression and expansion in a fixed area of ​​the point to obtain the matrix R;

[0062] Step 6. In the matrix R, if R(i,j) is greater than a certain threshold and R(i,j) is a local maximum in a certain area, it is considered a corner point.

[0063] The present invention can not only minimize the influence of dynamic objects on feature extraction at the algorithm theory level, but also effectively remove feature points on the dynamic objects or on the "appendages" following the dynamic objects. Figure 3 and Figure 4 The feature extraction effect of the present invention is demonstrated.

[0064] The preferred embodiments of the present invention are described in detail above, but the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations (such as quantity, shape, position, etc.) can be made to the technical scheme of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.

Claims

1. A dynamic feature removal method, Features: Includes steps: (1) Segment the moving objects in the image using the Mask-RCNN algorithm; (2) Obtaining a dynamic object that moves in association with the moving object obtained in step (1) through image RGB-D data and multi-view geometry algorithm; Specifically: (21) Selecting several frames of image data as key frames for dynamic feature determination; (22) Find the corresponding relationship between the feature points on each frame, that is, the feature points in each key frame x At the position of the current frame, get the corresponding feature point in the current frame x ′; (23) Through RGB-D data, feature points are obtained based on the projection relationship and the change of camera posture between frames. x The depth of the space Zproj , and calculate the angle of motion between each feature point and the line connecting its corresponding point in space during the movement α , α When it is greater than 30°, the feature point is discarded; (24) Based on the pixel correspondence, feature points are obtained from the depth data collected by the camera. x Depth of ' Z ′; (25) According to ∆Z = Z ' -Zproj Calculate the projection error of the feature points and use ∆Z Is it within the error threshold? τz Within the range, determine whether the feature point is blocked, that is, whether the feature point is on the dynamic object; (3) Combine steps (1) and (2) to obtain the dynamic features of the image, set all pixels in the dynamic feature area to 0, and expand the pixels at the edge of the area.

2. The dynamic feature elimination method according to claim 1, Features: In the step (21), 2 to 5 frames of image data are selected as key frames.

3. The dynamic feature elimination method according to claim 1, Features: In step (25), the error threshold τ z The value is 0.5m.

4. A corner detection algorithm using the dynamic feature elimination method according to any one of claims 1 to 3, Features: Includes steps: Step 1, removing dynamic features from the image using the dynamic feature removal method described in any one of claims 1 to 3; Step 2: Perform corner detection on the image obtained in Step 1 using a corner detection algorithm.

5. The corner detection algorithm according to claim 4, Features: The corner detection algorithm adopts Harris corner, SIFT, SURF or Fast.