Image processing method and system based on feature point detection

By combining SuperPoint and MaskR-CNN networks, using IMU data to optimize feature point detection, the problem of feature point distinction and reconstruction error in dynamic environments is solved, and high-precision feature point matching and map construction are achieved.

CN120279174APending Publication Date: 2025-07-08ANHUI UNIV OF TECH SCI & TECH PARK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337968.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively distinguish between static and dynamic feature points in a dynamic environment, resulting in an increase in reconstruction errors. Feature points that rely solely on image information are easily affected by light changes and occlusion, making it difficult to meet the needs of high-precision scene modeling.

Method used

Feature point detection is performed in combination with SuperPoint network and MaskR-CNN network, inter-pose constraints are used using IMU data, and dynamic feature points are eliminated through semantic segmentation and reprojection error optimization, static feature points are retained, and three-dimensional reconstruction is performed.

Benefits of technology

It improves the robustness and matching accuracy of feature points, adapts to lighting and viewing angle changes, and improves the reconstruction accuracy and map construction accuracy in dynamic scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279174A_ABST
    Figure CN120279174A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and system based on feature point detection, and belongs to the technical field of image processing. The input image frame is preprocessed; inputting the processed image frames into a SuperPoint network, and extracting features of the image frames; extracting point features by using a semantic segmentation method; performing line feature detection by using an IMU (Inertial Measurement Unit), extracting line features, and constructing re-projection errors of the point features and the line features; and constructing and drawing a local map according to the camera pose, updating, and stopping updating when all image frames are read. According to the method, through static feature point screening and multi-modal fusion, the reliability of image feature extraction and the precision of three-dimensional reconstruction are remarkably improved, the method can be operated in complex scenes such as dense dynamic objects, weak textures and illumination variation, and stable positioning and mapping capabilities are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image processing method and system based on feature point detection. Background Technique

[0002] With the rapid development of computer vision and deep learning technologies, image feature point-based processing methods have been widely applied in fields such as autonomous driving, augmented reality, and robot navigation. Traditional feature point extraction methods such as SIFT and ORB can maintain good robustness under image scale and rotation changes, but when dealing with dynamic scenes, high-noise environments, or complex texture images, there are often problems such as low feature point repetition rate and decreased matching accuracy. In recent years, deep learning models have gradually been introduced into the field of image feature extraction. For example, the SuperPoint network can simultaneously generate key point positions and descriptors through end-to-end training, significantly improving the stability and expression ability of feature extraction. In addition, the fusion of semantic segmentation and IMU data further broadens the boundaries of image processing. By combining multi-modal information, more refined image understanding and 3D reconstruction can be achieved. However, how to effectively distinguish static and dynamic feature points in a dynamic environment and improve the reconstruction accuracy in multi-frame observations remains a difficult point in current research.

[0003] However, most traditional image feature extraction methods ignore the interference of dynamic objects on feature points, resulting in an increase in reconstruction errors and difficulty in meeting the dual requirements of real-time performance and accuracy. Secondly, simply relying on image information to extract feature points is easily affected by external factors such as illumination changes and occlusion, resulting in unstable feature points and descriptor matching failures. Although some deep learning methods perform excellently in static scenes, they do not fully consider the screening and optimization of feature points in a dynamic environment. The extracted feature points are mixed with dynamic object information and are difficult to meet the requirements of high-precision scene modeling. Summary of the Invention

[0004] The purpose of the present invention is to provide an image processing method and system based on feature point detection to solve the problems raised in the above background technique.

[0005] To solve the above technical problems, the present invention provides the following technical solutions:

[0006] An image processing method based on feature point detection, the method includes the following steps: Step S1: Preprocess the input image frame; Step S2: Input the processed image frame into the SuperPoint network to extract the features of the image frame; Step S3: Use the semantic segmentation method to extract point features; Step S4: Use the IMU to detect line features, extract line features, and construct the reprojection error of point features and line features; Step S5: Construct and update a local map according to the camera pose, and stop updating when all image frames are read.

[0007] As a preferred solution of the image processing method based on feature point detection according to the present invention, the file paths of two images are read in sequence, and the image pixels are normalized.

[0008] According to the input requirements of the SuperPoint network, the image size is adjusted; generally, it is required that the input image has a fixed size. The specific size depends on the training settings of the model and the application scenario. Common input image sizes may be 256×256 pixels, 512×512 pixels, etc. If the size of the input image is inconsistent with the size required by the model, preprocessing operations such as scaling or cropping need to be performed.

[0009] The adjusted image is converted into a tensor using a deep learning framework.

[0010] As a preferred solution of the image processing method based on feature point detection according to the present invention, a SuperPoint network architecture is built, and the pre-trained weight parameters are imported.

[0011] Using the SuperPoint network, the positions, descriptors, and confidence levels of the key points in the image are extracted.

[0012] A confidence threshold is preset. If the confidence level of a key point is greater than or equal to the confidence threshold, the key point is marked as a high-quality feature point, and all high-quality feature points are screened out.

[0013] It should be noted that the confidence score output by the SuperPoint network is a key indicator for measuring the quality of feature points. The confidence score represents the network's evaluation of the accuracy of the extracted feature points. The higher the score, the more reliable the network believes the feature point is, and the more likely it is to be a real, stable, and representative feature. In subsequent image matching, pose estimation, and other links, it can provide more accurate and effective information. Such high-confidence feature points belong to high-quality feature points.

[0014] As a preferred solution of the image processing method based on feature point detection according to the present invention, the MaskR-CNN network is used to perform semantic segmentation on the image and generate a mask of the dynamic object.

[0015] Based on the mask of the dynamic object, it is checked whether the feature point falls within the mask. Specifically: if the pixel position corresponding to the coordinate of the feature point is within the area covered by the mask of the dynamic object, it is determined that the feature point falls within the mask of the dynamic object; if the pixel position corresponding to the coordinate of the feature point is not within the area covered by the mask of the dynamic object, it is determined that the feature point does not fall within the mask of the dynamic object.

[0016] Extract dynamic feature points and retain static feature points; the dynamic feature points represent the feature points that fall within the mask of the dynamic object, and the static feature points represent the feature points that do not fall within the mask of the dynamic object.

[0017] It should be noted that the dynamic feature points refer to the feature points located on the dynamic object. Since the position and pose of the dynamic object change continuously over time, the positions and attributes of these dynamic feature points in different image frames will also change accordingly. In the visual SLAM algorithm, dynamic feature points are not conducive to constructing a stable map and accurately estimating the camera pose because they have a high degree of uncertainty and may introduce errors. Therefore, they usually need to be identified and excluded during the algorithm processing; the static feature points are a concept opposite to the dynamic feature points and refer to the feature points located on the static object. The position and pose of the static object are relatively fixed in the image sequence, so the positions and attributes of the static feature points on it are also relatively stable in different image frames. In the visual SLAM algorithm, the static feature points are an important basis for constructing the map, estimating the camera pose, and performing three-dimensional reconstruction. They can provide stable and reliable information for the algorithm, helping to improve the accuracy and stability of the algorithm.

[0018] As a preferred solution of the image processing method based on feature point detection according to the present invention, the reprojection error of the static feature points is constructed by using the observations within the sliding window:

[0019] e point (T i ,P j ) = z ij -π(T i ,P j );

[0020] Among them, e point (T i ,P j ) represents the reprojection error of the point feature, T i represents the pose of the i-th camera, P j represents the j-th feature point, z ij represents the observed position of the feature point P j in the image of the i-th camera, π(T i ,P j ) is the pixel coordinate of the feature point P j projected onto the camera T i ;

[0021] For line features, the error is constructed using the endpoint observations:

[0022] e line (T i ,L k ) = z ik -π line(T i ,L k );

[0023] Among them, e line (T i ,L k ) represents the reprojection error of the line feature, L k represents the k-th line feature, z ik represents the observation position of the line feature L k in the i-th camera image, π line (T i ,L k ) represents the pixel coordinates of the line feature L k projected onto the camera T i .

[0024] As a preferred solution of the image processing method based on feature point detection described in the present invention, add inter-frame pose constraints:

[0025] e imu (T i ,T i+1 ) = ΔT imu - f(T i ,T i+1 );

[0026] Among them, e imu (T i ,T i+1 ) represents the inter-frame pose constraint error based on the IMU, T i+1 represents the pose of the (i + 1)-th camera, ΔT imu represents the pose change measured by the IMU from the i-th frame to the (i + 1)-th frame, f(T i ,T i+1 ) represents the pose change function from the i-th frame to the (i + 1)-th frame obtained according to visual observations;

[0027] Use a non-linear optimization tool to minimize the objective function.

[0028] As a preferred solution of the image processing method based on feature point detection described in the present invention, through multi-frame observations, triangulate the static feature points to obtain three-dimensional coordinates;

[0029] Combine the dynamic feature detection results and only triangulate the static feature points;

[0030] For line features, use multi-frame observations of their endpoints for three-dimensional reconstruction;

[0031] Merge the local map points within the sliding window into the global map.

[0032] An image processing system based on feature point detection, the system includes: a preprocessing module for preprocessing the input image frame; a feature extraction module for inputting the processed image frame into the SuperPoint network to extract the features of the image frame; a semantic segmentation extraction module for using semantic segmentation methods to extract point features; an error calculation module for using an IMU to detect line features, extract line features, and construct the reprojection error of point features and line features; an update module for constructing and updating a local map according to the camera pose and stopping the update when all image frames are read.

[0033] Compared with the prior art, the beneficial effects achieved by the present invention are: in the image processing method and system based on feature point detection provided by the present invention, by introducing superpoint network feature extraction, it has high robustness and real-time performance, can adapt to changes in illumination, perspective, and scale, and improve the quality of feature points and descriptors. Eliminate the feature points on dynamic objects to avoid the interference of false matches on positioning and mapping. Combine point and line features to generate a high-quality local map, which is suitable for dynamic scenarios and structured environments. Use the local optimization strategy of the sliding window to ensure that the system meets the requirements of real-time applications and ensure the accuracy of camera pose estimation and map construction. Description of the Drawings

[0034] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification, and are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention.

[0035] Figure 1 It is a schematic diagram of the steps of an image processing method based on feature point detection of the present invention;

[0036] Figure 2 It is the absolute trajectory error of mapping;

[0037] Figure 3 It is the relative pose error of mapping. Detailed Embodiments

[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0039] Please refer to Figure 1 , in the first embodiment: An image processing method based on feature point detection is provided, and the method includes the following steps:

[0040] Step S1: Preprocess the input image frame, aiming to provide a standardized data format for subsequent feature extraction, reduce noise interference, and improve the algorithm efficiency. Image frame preprocessing is the basic link of visual processing, directly affecting the accuracy and stability of subsequent feature extraction.

[0041] Step S1.1: Read the file paths of two images in sequence and perform normalization processing on the image pixels.

[0042] Step S1.2: Adjust the image size according to the input requirements of the SuperPoint network. Generally, it is required that the input image has a fixed size. The specific size depends on the training settings and application scenarios of the model. Common input image sizes may be 256×256 pixels, 512×512 pixels, etc. If the size of the input image is inconsistent with the size required by the model, preprocessing operations such as scaling or cropping need to be performed.

[0043] Step S1.3: Use the deep learning framework to convert the adjusted image into a tensor.

[0044] In the present invention, data is usually stored and processed in the form of tensors. Converting the image into a tensor facilitates fast calculation and parallel processing in the neural network. The dimensions of the image tensor are usually (batch_size, channels, height, width), where batch_size represents the number of images processed at one time, channels represents the number of channels of the image (such as 3 channels for RGB images), and height and width represent the height and width of the image respectively. In this algorithm, the processed image is converted into a tensor form of (1, 3, 256, 256) to adapt to the input requirements of the SuperPoint network.

[0045] Step S2: Input the processed image frame into the SuperPoint network to extract the features of the image frame.

[0046] Step S2.1: Build the SuperPoint network architecture and import the pre-trained weight parameters.

[0047] Step S2.2: Use the SuperPoint network to extract the positions, descriptors, and confidence levels of the key points in the image.

[0048] The SuperPoint network extracts features from the input image through structures such as convolutional layers and pooling layers. The key point positions represent the coordinates of the pixels with significant features in the image. The descriptors are used to describe the feature information of the key points for subsequent feature matching, and the confidence reflects the reliability degree of the key points. The key point positions output by the network are represented in the form of (x, y) coordinates, the descriptor is a 128-dimensional vector, and the confidence is a value between 0 and 1. The higher the value, the more reliable the key point.

[0049] Step S2.3: Preset a confidence threshold. If the confidence of a key point is greater than or equal to the confidence threshold, then mark the key point as a high-quality feature point and filter out all high-quality feature points.

[0050] Step S3: Extract point features using the semantic segmentation method.

[0051] Step S3.1: Use the Mask R-CNN network to perform semantic segmentation on the image and generate a mask for the dynamic object;

[0052] Step S3.2: Based on the mask of the dynamic object, check whether the feature point falls within the mask. Specifically: If the pixel position corresponding to the coordinate of the feature point is within the area covered by the dynamic object mask, then it is determined that the feature point falls within the mask of the dynamic object; if the pixel position corresponding to the coordinate of the feature point is not within the area covered by the dynamic object mask, then it is determined that the feature point does not fall within the mask of the dynamic object.

[0053] Step S3.3: Extract dynamic feature points and retain static feature points; the dynamic feature points represent the feature points that fall within the mask of the dynamic object, and the static feature points represent the feature points that do not fall within the mask of the dynamic object.

[0054] Step S4: Use the IMU to detect line features, extract line features, and construct the reprojection error of point features and line features.

[0055] Step S4.1: Use the observations within the sliding window to construct the reprojection error of the static feature points:

[0056] e point (T i ,P j )=z ij -π(T i ,P j );

[0057] Where, e point (T i ,P j ) represents the reprojection error of the point feature, T i represents the pose of the i-th camera, P jDenote the j-th feature point, z ij Denote the feature point P j The observed position in the i-th camera image, π(T i , P j ) is the feature point P j Projected onto the camera T i The pixel coordinates;

[0058] Step S4.2: For line features, construct the error using the endpoint observations:

[0059] e line (T i , L k ) = z ik - π line (T i , L k );

[0060] Among them, e line (T i , L k ) represents the reprojection error of the line feature, L k Represents the k-th line feature, z ik Represents the observed position of the line feature L k In the i-th camera image, π line (T i , L k ) represents the line feature L k Projected onto the camera T i The pixel coordinates.

[0061] Step S4.3: Add the inter-frame pose constraint:

[0062] e imu (T i , T i+1 ) = ΔT imu - f(T i , T i+1 );

[0063] Among them, e imu (T i , T i+1 ) represents the inter-frame pose constraint error based on the IMU, T i+1 Represents the pose of the (i + 1)-th camera, ΔT imu Represents the pose change from the i-th frame to the (i + 1)-th frame measured by the IMU, f(T i , T i+1 ) represents the pose change function from the i-th frame to the (i + 1)-th frame obtained from visual observations;

[0064] Step S4.4: Use a non-linear optimization tool to minimize the objective function.

[0065] Step S5: Construct and update a local map based on the camera pose, and stop the update when all image frames have been read.

[0066] Step S5.1: Through multi-frame observations, perform triangulation calculations on the static feature points to obtain three-dimensional coordinates;

[0067] Step S5.2: Combine the dynamic feature detection results and only perform triangulation on the static feature points;

[0068] Step S5.3: For line features, use multi-frame observations of their endpoints for three-dimensional reconstruction;

[0069] Step S5.4: Merge the local map points within the sliding window into the global map.

[0070] Furthermore, this embodiment also provides an image processing system based on feature point detection. The system includes: a preprocessing module for preprocessing the input image frames; a feature extraction module for inputting the processed image frames into the SuperPoint network to extract the features of the image frames; a semantic segmentation extraction module for extracting point features using semantic segmentation methods; an error calculation module for detecting line features using an IMU, extracting line features, and constructing the reprojection errors of point features and line features; an update module for constructing and updating a local map based on the camera pose, and stopping the update when all image frames have been read.

[0071] Furthermore, Figure 2 is the absolute trajectory error for mapping, Figure 3 is the relative pose error for mapping. Experiments have shown that in a dynamic scene, after removing dynamic features, the pose estimation error is significantly reduced; the map combining points and lines shows higher geometric consistency in weakly textured areas (such as walls), improving the quality of map construction.

[0072] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0073] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image processing method based on feature point detection, characterized in that The method includes the following steps: Step S1: Preprocess the input image frame; Step S2: Input the processed image frame into the SuperPoint network to extract the features of the image frame; Step S3: Extract point features using the semantic segmentation method; Step S4: Use the IMU to detect line features, extract line features, and construct the reprojection errors of point features and line features; Step S5: Construct and update the local map according to the camera pose, and stop updating when all image frames are read.

2. The image processing method based on feature point detection according to claim 1, wherein, The specific implementation process of step S1 includes: Step S1.1: Read the file paths of two images in sequence and normalize the image pixels; Step S1.2: Adjust the image size according to the input requirements of the SuperPoint network; Step S1.3: Use the deep learning framework to convert the adjusted image into a tensor.

3. The image processing method based on feature point detection according to claim 2, wherein The specific implementation process of step S2 includes: Step S2.1: Build the SuperPoint network architecture and import the pre-trained weight parameters; Step S2.2: Use the SuperPoint network to extract the positions, descriptors, and confidence levels of the key points in the image; Step S2.3: Preset the confidence threshold. If the confidence level of a key point is greater than or equal to the confidence threshold, mark the key point as a high-quality feature point and filter out all high-quality feature points.

4. An image processing method based on feature point detection according to claim 3, characterized in that, The specific implementation process of step S3 includes: Step S3.1: Use the Mask R-CNN network to perform semantic segmentation on the image and generate a mask for dynamic objects; Step S3.2: Based on the mask of the dynamic object, check whether the feature point falls within the mask. Specifically, if the pixel position corresponding to the coordinate of the feature point is within the area covered by the dynamic object mask, it is determined that the feature point falls within the mask of the dynamic object; if the pixel position corresponding to the coordinate of the feature point is not within the area covered by the dynamic object mask, it is determined that the feature point does not fall within the mask of the dynamic object; Step S3.3: Extract dynamic feature points and retain static feature points; the dynamic feature points represent the feature points that fall within the mask of the dynamic object, and the static feature points represent the feature points that do not fall within the mask of the dynamic object.

5. An image processing method based on feature point detection according to claim 4, wherein The specific implementation process of step S4 includes: Step S4.1: Use the observations within the sliding window to construct the reprojection error of the static feature points; e point (T i ,P j ) = z ij -π(T i ,P j ); Among them, e point (T i , P j ) represents the reprojection error of the point feature. T i represents the pose of the i-th camera, and P j represents the j-th feature point. z ij represents the observed position of the feature point P j in the image of the i-th camera. π(T i , P j ) is the pixel coordinate where the feature point P j is projected onto the camera T i ; Step S4.2: For line features, use the endpoint observations to construct the error; e line (T i ,L k ) = z ik -π line (T i ,L k ); Among them, e line (T i , L k ) represents the reprojection error of the line feature, L k represents the k-th line feature, z ik represents the observation position of the line feature L k in the i-th camera image, π line (T i , L k ) represents the projection of the line feature L k onto the pixels of camera T i .

6. The image processing method based on feature point detection according to claim 5, wherein The specific implementation process of step S4 also includes: Step S4.3: Add inter-frame pose constraints; e imu (T i ,T i+1 ) = ΔT imu -f(T i ,T i+1 ); Among them, e imu (T i , T i+1 ) represents the inter-frame pose constraint error based on the IMU. T i+1 represents the pose of the (i + 1)-th camera, and ΔT imu represents the pose change from the i-th frame to the (i + 1)-th frame measured by the IMU. f(T i , T i+1 ) represents the pose change function from the i-th frame to the (i + 1)-th frame obtained according to visual observations; Step S4.4: Use the nonlinear optimization tool to minimize the objective function.

7. An image processing method based on feature point detection according to claim 5, characterized in that The specific implementation process of step S5 includes: Step S5.1: Through multi-frame observations, perform triangulation calculations on the static feature points to obtain three-dimensional coordinates; Step S5.2: Combine the dynamic feature detection results and only perform triangulation on the static feature points; Step S5.3: For line features, use the multi-frame observations of their endpoints for three-dimensional reconstruction; Step S5.4: Merge the local map points within the sliding window into the global map.

8. An image processing system based on feature point detection, which executes an image processing method based on feature point detection as described in any one of claims 1-7, characterized in that, The system includes: A preprocessing module for preprocessing the input image frames; A feature extraction module for inputting the processed image frames into the SuperPoint network to extract the features of the image frames; A semantic segmentation extraction module for extracting point features using semantic segmentation methods; An error calculation module for detecting line features using an IMU, extracting line features, and constructing reprojection errors of point features and line features; An update module for constructing and updating a local map based on the camera pose and stopping the update when all image frames have been read.