Scene matching navigation method based on region segmentation and feature optimization

Through the scene matching navigation method based on area segmentation and feature optimization, the segmentation and feature optimization problems of flat image areas in helicopter navigation are solved, high-precision navigation and positioning are achieved, and the stability and accuracy of the navigation system are improved.

CN120274748APending Publication Date: 2025-07-08NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510363900.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing scene matching navigation methods lack rapid segmentation and feature optimization of image flat areas in helicopter navigation, resulting in mismatch problems and affecting positioning accuracy.

Method used

Using a method based on region segmentation and feature optimization, edge strength is calculated through the Scharr operator, image segmentation is performed by combining open and closed operations and connected area calculation, feature points are optimized using ORB algorithm and Harris response value, and matching point pairs are screened in combination with RANSAC algorithm to achieve high-precision navigation.

Benefits of technology

Effectively segment the flat image and rich texture areas, optimize feature points, improve positioning accuracy and robustness, reduce mismatch, and improve the stability and accuracy of the navigation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120274748A_ABST
    Figure CN120274748A_ABST
Patent Text Reader

Abstract

The invention discloses a scene matching navigation method based on region segmentation and feature optimization, and the method comprises the steps: 1, generating a reference image, and reading a real-time image transmitted by an airborne image sensor; 2, forming an edge intensity image model of the two images; 3, preliminarily realizing image region segmentation, and segmenting a flat region and a texture-rich region in the image; step 4, filling the region with rich textures in the step 3, and completing image region segmentation; 5, extracting feature points of the reference image and the real-time image by using an ORB algorithm based on the mask matrix; step 6, retaining a predetermined number of feature points, and forming feature descriptors; 7, re-projection errors of the rough matching point pairs are calculated, a preset number of matching point pairs are reserved to serve as positioning point pairs, and a homography matrix is solved based on an RANSAC algorithm according to the positioning point pairs; and 8, calculating navigation coordinates according to a pixel positioning result and the additional information of the reference image, and outputting a navigation result to complete scene matching navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of navigation systems and relates to a scene matching navigation method based on region segmentation and feature optimization. Background Art

[0002] Common navigation systems for helicopters include the Global Navigation Satellite System (GNSS), Inertial Navigation System (INS), Scene Matching Navigation System (SMNS), etc. Among them, the most commonly used navigation system is the Global Navigation Satellite System. However, with the increasing development of anti-satellite technology and electromagnetic interference technology, satellite navigation faces unprecedented challenges, and autonomous navigation technology needs to be used to ensure flight safety. However, inertial navigation has the problem of drift error accumulation, and relying solely on the inertial navigation system to achieve autonomous navigation cannot meet the stability and accuracy requirements of helicopter navigation. Therefore, other autonomous navigation technologies are needed for assistance. New autonomous navigation technologies have gradually become a research hotspot. Among them, scene matching navigation technology has become an important direction for the future development of helicopter autonomous navigation due to its low dependence on external signals, strong adaptability to complex confrontation environments, and the ability to be combined with other navigation technologies.

[0003] Scene matching navigation technology is a technology that realizes autonomous positioning and navigation by matching sensor data with an existing reference map. However, due to the flexible flight characteristics of helicopters, scene matching navigation often faces complex and changing task scenarios. In such application scenarios, traditional scene matching navigation methods based on feature matching often have the problem of false matching, resulting in a decrease in positioning accuracy. There have been relevant technical studies at home and abroad to solve the false matching problem. For the outlier problem, the research team of the Institute of Automation, Chinese Academy of Sciences proposed a method for robust feature correspondence of outliers using a pruning attention graph neural network and a matching layer; for the false matching problem caused by factors such as locally similar structures, the research team of Beihang University proposed a SIFT feature point matching method based on nearest neighbor support feature points; for the problem of feature point description accuracy, the research team of Moscow State University proposed a phase consistency metric method calculated near image key points.

[0004] Aerial scene images can be divided into flat areas and areas with rich textures. There is little research on feature optimization methods for feature false matching problems in image flat areas. Currently, the research on scene matching navigation methods lacks a scene matching navigation method based on region segmentation and feature optimization, lacks a method for quickly judging and segmenting image flat areas, and cannot effectively optimize the feature points of false matching in image flat areas, thus affecting the matching positioning accuracy. Summary of the Invention

[0005] The purpose of the present invention is to provide a scene matching navigation method based on region segmentation and feature optimization.

[0006] The technical solution adopted by the present invention is as follows: A scene matching navigation method based on region segmentation and feature optimization, comprising the following steps:

[0007] Step 1: Generate a reference image according to the indicated position of the helicopter airborne inertial navigation system, read the additional information of the reference image, and read the real-time image transmitted by the airborne image sensor, where the additional information is the positioning parameters of the image;

[0008] Step 2: Calculate the edge intensity of the reference image and the real-time image based on the Scharr operator to form an edge intensity image model of the two images;

[0009] Step 3: Initially realize image region segmentation according to the edge intensity and image opening and closing operations, and segment the flat regions and texture-rich regions in the image;

[0010] Step 4: Fill the flat regions in Step 3 according to the image connected region calculation method, and fill the texture-rich regions in Step 3 according to the image closing operation method to complete the image region segmentation;

[0011] Step 5: Generate a mask matrix according to the image region segmentation result, and extract the feature points of the reference image and the real-time image based on the mask matrix using the ORB algorithm;

[0012] Step 6: Optimize the feature points of the reference image and the real-time image based on the Harris response value calculation method, retain half of the number of feature points, and the response values of the retained feature points are higher than those of the non-retained feature points, and form a feature descriptor;

[0013] Step 7: Perform feature matching on the feature points retained in Step 6 to form a rough matching point pair, calculate the reprojection error of the rough matching point pair, retain a predetermined number of matching point pairs as positioning point pairs, and the reprojection error of the retained matching point pairs is less than that of the non-retained matching point pairs. Solve the homography matrix based on the positioning point pairs using the RANSAC algorithm;

[0014] Step 8: Complete image pixel positioning through homography transformation, solve the navigation coordinates according to the pixel positioning result and the additional information of the reference image, and output the navigation result to complete the scene matching navigation.

[0015] Further, Step 1 is specifically:

[0016] The reference image database is prepared in advance. Based on the navigation information of the helicopter-borne inertial navigation system at the current moment, the approximate range of the current position is obtained, and the reference image for matching is selected. The additional information of the reference image is the positioning parameters of the image, including the Mercator projection coordinates corresponding to the upper-left pixel point and the projection coordinate difference corresponding to one pixel grid. The real-time image is an RGB optical image taken by the optical camera carried by the helicopter from the aerial perspective. When reading, the RGB optical image is converted into a grayscale image.

[0017] Further, the specific steps of step 2 are as follows:

[0018] The convolution kernels K of the Scharr operator x and K y are shown in formulas (1) and (2) as follows:

[0019]

[0020]

[0021] For the reference image and the real-time image, in combination with the convolution kernels K x and K y calculate the gray-scale gradient changes at each pixel position of the two images respectively:

[0022] G x (x, y) = I(x, y) * K x (3)

[0023] G y (x, y) = I(x, y) * K y (4)

[0024] In formulas (3) and (4), I(x, y) is the gray-scale value at the position (x, y) of the image, * is the convolution operation, which means scanning the image with the convolution kernel, G x (x, y) is the horizontal gradient component at the position (x, y) of the image, indicating the degree of gray-scale change of the image in the horizontal direction; G y (x, y) is the vertical gradient component at the position (x, y) of the image, indicating the degree of gray-scale change of the image in the vertical direction;

[0025] According to formulas (3) and (4), in combination with the convolution kernels K x and K y calculate the gradients G x and G y of the image in the horizontal and vertical directions, revealing the intensity of pixel gray-scale value change in the x and y directions of the image. Use G x (x, y) and G y (x, y) to synthesize the gradient amplitude:

[0026]

[0027] Combine the horizontal gradient component and the vertical gradient component to synthesize the gradient magnitude G(x, y). Assume that the edge intensity distribution of the image is an m×n matrix M e , then the edge intensity distribution matrix M e is as shown in Equation (6):

[0028]

[0029] In order to apply the image edge intensity distribution in image processing and facilitate subsequent region segmentation based on the Scharr edge intensity, a proportional scaling method is used to linearly normalize M e to construct an edge intensity model Normalize the Scharr edge intensity amplitude within the range of image grayscale display;

[0030] Assume that the maximum and minimum edge intensities of M e are G max and G min . Let the normalized maximum edge intensity be assigned G max = 255, and the normalized minimum edge intensity be assigned G min = 0. Normalize the edge intensity to the range of [0, 255]. Then, the normalization formula for the element G e in the edge intensity distribution matrix M i is as shown in Equation (7):

[0031]

[0032] Based on Equation (7), the normalized edge intensity model is as shown in Equation (8):

[0033]

[0034] Set the part with edge intensity higher than the edge intensity threshold to 255, and the part with edge intensity lower than the threshold to 0. Binarize to obtain the edge intensity image model I e .

[0035] Furthermore, the specific steps of Step 3 are as follows:

[0036] The opening and closing operations rely on the structuring element (SE) to perform image processing operations and control the specific behaviors of the erosion operation and the dilation operation. Among them, the erosion operation moves the structuring element SE on the image and replaces the gray value at the current pixel position with the minimum value in the local area of the image, as shown in Equation (9):

[0037]

[0038] In formula (9), z represents the current pixel position, and SE z represents the set when the anchor point of the structural element SE is at the z position. Only when all pixels of SE are within the foreground range of the image I, the pixel at the z position will be retained in the erosion result;

[0039] The dilation operation moves the structural element SE on the image and replaces the gray value at the current pixel position with the maximum value in the local area of the image, as shown in formula (10):

[0040]

[0041] In formula (10), SE z represents the set when the anchor point of the structural element SE is at the z position. When any pixel in SE intersects with the foreground, the pixel at the z position will be retained in the dilation result;

[0042] According to the above erosion operation and dilation operation, perform opening and closing operations on the edge intensity image model I e Integrate the discrete edge pixel points based on the image edge information of I e to extract the flat regions in the image model, thereby achieving the effect of region segmentation. First, perform an image opening operation on I e by first performing an erosion operation on the image and then a dilation operation, as shown in formula (11):

[0043]

[0044] First, remove some pixels in the foreground region of I e to smooth the boundary. The foreground region is the region with rich texture. Then, restore the edge of the foreground region lost due to the erosion operation through the dilation operation;

[0045] Then, perform an image closing operation on I e by first performing a dilation operation on the image and then an erosion operation, as shown in formula (12):

[0046]

[0047] First, remove some pixels in the background region of I e to fill small regions in the background and connect adjacent foreground regions. The background region is the flat region. Then, restore the edge of the background region lost due to the dilation operation through the erosion operation to smooth the boundary and complete the preliminary image region segmentation based on the image opening and closing operations.

[0048] Further, the specific content of step 4 is as follows:

[0049] In connectivity calculation, the pixel connectivity judgment criterion is as follows: based on whether the gray values of the eight pixels in the up, down, left, right, and four diagonal directions of the current pixel in the image are the same as that of the current pixel, it is judged whether all the surrounding pixels are connected to it;

[0050] Use the Two-Pass method to complete the connected component labeling and region merging of the edge intensity image model I e after the opening and closing operations. In the first pass, by traversing each pixel of I e , when the pixel is in the foreground area, the gray relationship of the neighborhood pixels is calculated in three cases:

[0051] (1) If all neighborhood pixels are in the background area, assign a new label L to the current pixel z i ;

[0052] (2) If there is a pixel with a labeled L i in the neighborhood, the current pixel z inherits the label L of the existing pixel in the neighborhood i ;

[0053] (3) If there are multiple pixels with different labeled tags in the neighborhood, assign any one of the tags to the current pixel z, and record the equivalence relationship between these tags;

[0054] In the second pass, according to the labeling results of the first pass, the equivalent labeling of each connected region is merged. Based on the labeling categories generated by the first pass results, during the second pass, each pixel label is replaced with the representative label in the equivalent class. Usually, the final connected region label is used. Among them, the equivalent class refers to the set of region labels with the same connection relationship. Based on the Two-Pass method, all pixels belonging to the same connected region in the edge intensity image model I e are labeled with the same label, and the connected region calculation is completed;

[0055] After quickly obtaining the connected region calculation results using the Two-Pass method, retain the blocks in I e where the connected region is larger than 100 pixel squares, and then construct a new structural element SE' (for example, set the size to 30 * 30 pixels), perform an image closing operation once to filter out small background blocks, and complete the image region segmentation.

[0056] Further, the specific step 5 is as follows:

[0057] Assume that after the image region segmentation in step 4, the extracted flat region is represented by R f , and the texture-rich region is represented by R d ;

[0058] According to the flat region Rf Create the corresponding mask matrix R f_mask , let R f_mask For the elements at the pixel positions corresponding to the R f region in the matrix, the element values are 0, and for the elements at the pixel positions corresponding to the non-R f region, the element values are 255, thus completing the construction of the mask matrix R f_mask , forming the mask matrix R f_mask As shown in formula (13):

[0059]

[0060] Based on the mask matrix R f_mask Omit the feature detection of the flat region on the original image, create an ORB feature detection function, and introduce the mask matrix R after setting the detection parameters f_mask , and extract the set of feature points of the reference image and the real-time image.

[0061] Furthermore, the specific content of step 6 is as follows:

[0062] Record the pixel positions (x, y) and index information of all FAST feature points of the real-time image and the reference image, forming a FAST feature point pixel position vector;

[0063] According to the pixel position vector, construct a structure tensor matrix M at each feature point pixel position (x, y) to describe the gradient distribution within the window. Among them, the calculation of the structure tensor matrix M is as shown in formula (14):

[0064]

[0065] where w(x, y) is a weighting function, which is a Gaussian weight, I x and I y respectively represent the gradients of the current pixel position in the x direction and y direction, W is a local window in the image used to calculate the information of surrounding pixels, and M is represented by the following three elements A, B, and C:

[0066]

[0067]

[0068]

[0069] According to the structure tensor matrix M, calculate the magnitude of the Harris response value R at the corresponding position of each feature point, and record the index information of this position. Among them, the calculation of the response value R is as shown in formula (18):

[0070] R = detM - k(traceM) 2 (18)

[0071] detM = λ1λ2 = AB - C 2 (19)

[0072] trace M = λ1 + λ2 = A + B (20)

[0073] Where λ1 and λ2 are the eigenvalues of matrix M, representing the direction and magnitude of the image gradient change, detM represents the determinant of matrix M, traceM represents the trace of matrix M, and k is an empirical constant with a value ranging from 0.04 to 0.06;

[0074] Then store the feature points and the Harris response values at the corresponding pixel positions as key - value pairs, sort them in descending order of the response values, set the retention quantity N, and according to the sorting result, filter and retain the top N feature points with the highest response values (rounded up to the nearest integer);

[0075] Finally, update the original feature point set, use the BRIEF algorithm for feature description, form the feature descriptor sets of the real - time image and the reference image, simply called the feature set, and represent them as P = {p1, p2.., p i ,...p m} and P′ = {p1', p2'...p i '..p n '}, respectively, where p i and p i ' represent the feature descriptors in the two sets respectively.

[0076] Furthermore, the specific content of step 7 is as follows:

[0077] Create a brute - force matcher. Since the ORB feature descriptor is used for matching, set the brute - force matching parameter to the Hamming distance mode. Take the descriptor of a feature in the first feature set P = {p1, p2.., p i ,...p m} and match it with the descriptors of all features in the second feature set P′ = {p1', p2'...p i '..p n '}. According to the matching result, return the feature point with the minimum Hamming distance, that is, the feature point with the highest matching degree, to complete the matching of a single feature point. Then loop through all the feature points to complete the matching of all feature points and form the rough - matching point pairs;

[0078] Based on the rough - matching point pairs, use RANSAC to filter and solve the rough - matching homography matrix H1. Use H1 to represent the transformation relationship between the real - time image and the reference image based on the rough - matching point pairs, and predict the position of the predicted points of the real - time image feature points on the reference image;

[0079] Solve the reprojection error \(e\) of each pair of rough matching point pairs according to the pixel positions of the prediction points and the matching points i , let \(q\) i and be the observed point and the prediction point on an image. For a two-dimensional image, the reprojection error \(e\) on the reference image is the Euclidean distance between \(q\) i and , as shown in formula (21);

[0080]

[0081] where is the pixel coordinate of the observed point, is the pixel coordinate of the prediction point. In the homogeneous coordinate system, the coordinate of the matching point on the reference image is \((x′, y′, 1)\), and the coordinate of the prediction point on the reference image is The calculation of the reprojection error \(e\) i is shown in formula (22):

[0082]

[0083] where is the homogeneous coordinate component for normalization, taking

[0084] According to the above reprojection error calculation process, obtain the matching performance evaluation of all rough matching point pairs;

[0085] According to the magnitude of the reprojection error \(e\) i of each pair of rough matching point pairs, sort all matching point pairs from high to low. When the number of rough matching point pairs is greater than \(2n\) pairs, retain the first \(n\) matching point pairs (rounded up). When the number of rough matching point pairs is less than \(2n\) pairs, retain the first 50% matching point pairs (rounded up). At this time, the retained matching point pairs are the fine matching point pairs;

[0086] Based on the fine matching point pairs, use RANSAC to filter and solve the fine matching homography matrix \(H2\), predict the position of the predicted points of the real-time image feature points on the reference image, and then solve the reprojection error of each pair of fine matching point pairs;

[0087] According to the magnitude of the reprojection error of each pair of fine matching point pairs, retain 80% of the fine matching point pairs with the smallest reprojection error (rounded up). At this time, the retained matching point pairs are the positioning point pairs; then use RANSAC to filter and solve the positioning homography matrix \(H3\) based on the positioning point pairs.

[0088] Furthermore, step 8 is specifically as follows:

[0089] The homography matrix can achieve point-to-point mapping from one image to another. Assuming that the matching points in the real-time image and the reference image are in homogeneous coordinates (u, v, 1) and (u', v', 1), the homography matrix can describe the mapping relationship of this matching point pair, as shown in formula (23):

[0090]

[0091] Among them, H3 represents the positioning homography matrix. After obtaining the predicted pixel coordinates of the central pixel point of the real-time image on the reference image according to the homography transformation relationship, the pixel coordinate positioning of the two images is completed;

[0092] In the tiff-format reference image, the corresponding relationship of Mercator projection coordinates is realized through the origin and row and column parameters. The origin pixel coordinates (0, 0) of the reference image are calibrated as a Mercator projection coordinate (x0, y0), and then according to the row parameter P w and the column parameter P h Solve the projection coordinates corresponding to the pixel coordinates (u', v'):

[0093]

[0094] In formula (24), (x0, y0) is the projection coordinate of the upper left corner pixel of the reference image, (u', v') are the pixel coordinates, P w and P h are the row parameter and column parameter of the reference image, representing the pixel resolution, (x, y) are the projection coordinates after the transformation of (u', v'). Based on the above calculations, the transformation of the reference image pixel coordinates to the Mercator projection coordinates is completed;

[0095] The origin of the Mercator projection is set as the intersection point (λ0, φ0) of the equator and the prime meridian, where λ0 = 0°, φ0 = 0°. The conversion of longitude and latitude coordinates (λ, φ) to projection coordinates (x, y) is as shown in formula (25) and formula (26):

[0096] x = j·(λ - λ0) (25)

[0097]

[0098] Among them, j is the scale factor of the projection, which is the radius of the earth. In the WGS-84 coordinate system, j ≈ 6378137m;

[0099] Based on the above coordinate transformation calculations, the conversion of image pixel coordinates to longitude and latitude coordinates is realized, the navigation positioning coordinate solution is completed, and the navigation result is output.

[0100] The beneficial effects of the present invention are as follows: (1) The scene matching navigation method based on region segmentation and feature optimization proposed by the present invention designs an image region segmentation method based on opening and closing operations and connected region calculation, which can quickly and completely segment the flat region and texture-rich region of the image, form a mask matrix to optimize the image feature extraction process, retain the feature points of the texture-rich region, and improve the robustness of image feature points in complex scenes; (2) Based on the commonly used ORB feature point detection, the present invention introduces the Harris response value calculation to optimize the feature points, effectively optimizes the corner degree of the feature points, improves the distinctiveness of the feature points, and reduces the generation of false matches; (3) For the feature matching results, a matching point pair optimization method based on reprojection error calculation is designed. By screening and retaining the matching point pairs with strong overall consistency based on the reprojection error size of the matching point pairs for positioning calculation, the positioning accuracy can be effectively improved; (4) Simulation experiments show that compared with the traditional scene matching navigation method, the scene matching navigation method based on region segmentation and feature optimization proposed by the present invention can greatly improve the positioning accuracy, has little impact on real-time performance, improves the robustness of the scene matching navigation system, and in different matching scenarios, the method of the present invention shows better matching and positioning accuracy.

[0101] In addition to the objectives, features, and advantages described above, the present invention has other objectives, features, and advantages. The present invention will be further described in detail below with reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0102] Figure 1 It is a schematic diagram for pixel connectivity calculation.

[0103] Figure 2 It is a schematic diagram of the image region distribution after region segmentation.

[0104] Figure 3 It is a simulation experiment image data diagram.

[0105] Figure 4 It is the effect diagram of region segmentation realized by the method of the present invention.

[0106] Figure 5 It is the image matching effect diagram based on different algorithms.

[0107] Figure 6 It is the eastward positioning error diagram.

[0108] Figure 7 It is the northward positioning error diagram.

[0109] Figure 8 It is the total positioning error diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0110] The present invention will be described in detail below with reference to the drawings and specific embodiments.

[0111] A scene matching navigation method based on region segmentation and feature optimization, comprising the following steps:

[0112] Step 1: Generate a reference image according to the position indicated by the helicopter airborne inertial navigation system, read the additional information of the reference image, and read the real-time image transmitted by the airborne image sensor, where the additional information is the positioning parameters of the image;

[0113] Step 2: Calculate the edge intensity of the reference image and the real-time image based on the Scharr operator to form an edge intensity image model of the two images;

[0114] Step 3: Initially implement image region segmentation according to the edge intensity and image opening and closing operations, and segment the flat regions and texture-rich regions in the image;

[0115] Step 4: Fill the flat regions in Step 3 according to the image connected region calculation method, and fill the texture-rich regions in Step 3 according to the image closing operation method to complete the image region segmentation;

[0116] Step 5: Generate a mask matrix according to the image region segmentation result, and use the ORB algorithm to extract the feature points of the reference image and the real-time image based on the mask matrix;

[0117] Step 6: Optimize the feature points of the reference image and the real-time image based on the Harris response value calculation method, retain a predetermined number of feature points, where the response values of the retained feature points are higher than those of the un-retained feature points, and form a feature descriptor;

[0118] Step 7: Perform feature matching on the feature points retained in Step 6 to form rough matching point pairs, calculate the reprojection error of the rough matching point pairs, retain a predetermined number of matching point pairs as positioning point pairs, where the reprojection error of the retained matching point pairs is less than that of the un-retained matching point pairs, and solve the homography matrix based on the positioning point pairs using the RANSAC algorithm;

[0119] Step 8: Complete image pixel positioning through homography transformation, solve the navigation coordinates according to the pixel positioning result and the additional information of the reference image, and output the navigation result to complete the scene matching navigation.

[0120] Further, the specific content of Step 1 is:

[0121] The reference image database is prepared in advance. Based on the navigation information of the helicopter's airborne inertial navigation system at the current moment, the approximate range of the current position is obtained, and the reference image for matching is selected. The additional information of the reference image is the positioning parameters of the image, including the Mercator projection coordinates corresponding to the upper left pixel point and the projection coordinate difference corresponding to one pixel grid. The real-time image is an RGB optical image taken by an optical camera carried by the helicopter from an aerial perspective. When reading, the RGB optical image is converted into a grayscale image.

[0122] Further, the specific steps of step 2 are as follows:

[0123] The convolution kernel K of the Scharr operator x and K y are shown in formulas (1) and (2) as follows:

[0124]

[0125]

[0126] For the reference image and the real-time image, combined with the convolution kernel K x and K y calculate the gray-scale gradient changes at each pixel position of the two images respectively:

[0127] G x (x, y) = I(x, y) * K x (3)

[0128] G y (x, y) = I(x, y) * K y (4)

[0129] In formulas (3) and (4), I(x, y) is the gray-scale value at the position (x, y) of the image, * is the convolution operation, which means scanning the image with the convolution kernel, G x (x, y) is the horizontal gradient component at the position (x, y) of the image, indicating the degree of gray-scale change of the image in the horizontal direction; G y (x, y) is the vertical gradient component at the position (x, y) of the image, indicating the degree of gray-scale change of the image in the vertical direction;

[0130] According to formulas (3) and (4), combined with the convolution kernel K x and K y calculate the gradients G x and G y of the image in the horizontal and vertical directions, revealing the intensity of pixel gray-scale value change in the x and y directions of the image. Use G x (x, y) and G y (x, y) to synthesize the gradient amplitude:

[0131]

[0132] Combine the horizontal gradient component and the vertical gradient component to synthesize the gradient magnitude G(x, y). Assume that the edge intensity distribution of the image is an m×n matrix M e , then the edge intensity distribution matrix M e is as shown in Equation (6):

[0133]

[0134] In order to apply the image edge intensity distribution in image processing and facilitate subsequent region segmentation based on the Scharr edge intensity, a proportional scaling method is used to linearly normalize M e to construct an edge intensity model to standardize the Scharr edge intensity amplitude within the range of image grayscale display;

[0135] Assume that the maximum and minimum edge intensities of M e are G max and G min . Let the normalized maximum edge intensity be assigned G max = 255, and the normalized minimum edge intensity be assigned G min = 0. Standardize the edge intensity within the numerical range of [0, 255]. Then, the normalization formula for the element G e in the edge intensity distribution matrix M i is as shown in Equation (7):

[0136]

[0137] Based on Equation (7), the normalized edge intensity model is as shown in Equation (8):

[0138]

[0139] Let the part with edge intensity higher than the edge intensity threshold (for example, take the edge intensity threshold = 10) be set to 255, and the part with edge intensity lower than the threshold be set to 0. Binarize to obtain the edge intensity image model I e .

[0140] Furthermore, the specific steps of step 3 are as follows:

[0141] Opening and closing operations rely on a structuring element (SE) to perform image processing operations and control the specific behaviors of erosion and dilation operations. Among them, the erosion operation moves the structuring element SE on the image and replaces the grayscale value at the current pixel position with the minimum value in the local area of the image, as shown in Equation (9):

[0142]

[0143] In formula (9), z represents the current pixel position, and SE z represents the set when the anchor point of the structural element SE (the reference point of the structural element, generally the center point) is at the z position. Only when all pixels of SE are within the foreground range of the image I, the pixel at the z position will be retained in the erosion result;

[0144] The dilation operation moves the structural element SE on the image and replaces the gray value at the current pixel position with the maximum value in the local area of the image, as shown in formula (10):

[0145]

[0146] In formula (10), SE z represents the set when the anchor point of the structural element SE is at the z position. When any one pixel in SE intersects with the foreground, the pixel at the z position will be retained in the dilation result;

[0147] According to the above erosion operation and dilation operation, perform opening and closing operations on the edge intensity image model I e Integrate the discrete edge pixel points based on the image edge information of I e and extract the flat areas in the image model, so as to achieve the effect of region segmentation. First, perform an opening operation on I e Perform an erosion operation on the image first and then a dilation operation, as shown in formula (11):

[0148]

[0149] First, remove some pixels in the foreground area of I e to smooth the boundary. The foreground area is the area with rich texture. Then, restore the edge of the foreground area lost due to the erosion operation through the dilation operation;

[0150] Then perform a closing operation on I e Perform a dilation operation on the image first and then an erosion operation, as shown in formula (12):

[0151]

[0152] First, remove some pixels in the background area of I e to fill the small areas in the background and connect adjacent foreground areas. The background area is the flat area. Then, restore the edge of the background area lost due to the dilation operation through the erosion operation to smooth the boundary and complete the preliminary image region segmentation based on the image opening and closing operations.

[0153] Furthermore, step 4 is specifically as follows:

[0154] The present invention calculates the connectivity between pixels using an 8-connected connectivity matrix C, as Figure 1 shown. In the connectivity calculation, the pixel connectivity judgment criterion is: based on whether the gray values of the eight pixels in the up, down, left, right, and four diagonal directions of the current pixel in the image are the same as that of the current pixel, it is determined whether the surrounding pixels are all connected to it;

[0155] The two-pass method is used to complete the connected component labeling and region merging of the edge intensity image model I e after the opening and closing operations. In the first scan, by traversing each pixel of I e , when the pixel is in the foreground area, the gray relationship of the neighborhood pixels is calculated in three cases:

[0156] (1) If all neighborhood pixels are in the background area, a new label L is assigned to the current pixel z i ;

[0157] (2) If there is a pixel with a marked L i in the neighborhood, the current pixel z inherits the label L of the existing pixel in the neighborhood i ;

[0158] (3) If there are multiple pixels with different marked labels in the neighborhood, any one of the labels is assigned to the current pixel z, and the equivalence relationship between these labels is recorded;

[0159] During the second scan, based on the marking results of the first scan, the equivalent markings of each connected region are merged. Based on the marking categories generated by the first scan results, during the second scan, each pixel marking is replaced with the representative marking in the equivalent class. Usually, the final connected region marking is used. Among them, the equivalent class refers to a set of region markings with the same connection relationship. Based on the two-pass method, all pixels belonging to the same connected region in the edge intensity image model I e are marked with the same label, and the connected region calculation is completed;

[0160] After quickly obtaining the connected region calculation result using the two-pass method, the blocks in I e with a connected region larger than 100 pixel squares are retained, and then a new structural element SE' (for example, set the size to 30*30 pixels) is constructed, and an image closing operation is performed once to filter out small background blocks and complete the image region segmentation.

[0161] Furthermore, step 5 is specifically as follows:

[0162] Assume that after the image region segmentation in step 4, the extracted flat region is denoted as Rf It is indicated that the texture-rich area is represented by R d It is indicated that the distribution of the image areas after domain segmentation is as Figure 2 shown;

[0163] According to the flat area R f Create the corresponding mask matrix R f_mask , let the elements at the pixel positions corresponding to the R f_mask area in the R f matrix be 0, and the elements at the pixel positions corresponding to the non-R f area be 255, complete the construction of the mask matrix R f_mask , and form the mask matrix R f_mask as shown in formula (13):

[0164]

[0165] Based on the mask matrix R f_mask Omit the feature detection of the flat area on the original image, create an ORB feature detection function, introduce the mask matrix R after setting the detection parameters f_mask , and extract the feature point sets of the reference image and the real-time image.

[0166] Furthermore, the specific content of step 6 is as follows:

[0167] Record the pixel positions (x, y) and index information of all FAST feature points of the real-time image and the reference image to form a FAST feature point pixel position vector;

[0168] According to the pixel position vector, construct a structure tensor matrix M at each feature point pixel position (x, y) to describe the gradient distribution within the window. Among them, the calculation of the structure tensor matrix M is as shown in formula (14):

[0169]

[0170] Among them, w(x, y) is a weighting function, which is a Gaussian weight, I x and I y respectively represent the gradients of the current pixel position in the x direction and the y direction. W is a local window in the image, which is used to calculate the information of surrounding pixels. Represent M with the following three elements A, B, and C:

[0171]

[0172]

[0173]

[0174] According to the structure tensor matrix M, calculate the magnitude of the Harris response value R at the corresponding positions of each feature point, and record the position index information. Among them, the calculation of the response value R is shown in formula (18):

[0175] R = detM - k(traceM) 2 (18)

[0176] detM = λ1λ2 = AB - C 2 (19)

[0177] trace M = λ1 + λ2 = A + B (20)

[0178] Among them, λ1 and λ2 are the eigenvalues of the matrix M, representing the direction and magnitude of the image gradient change. detM represents the determinant of the matrix M, traceM represents the trace of the matrix M, and k is an empirical constant, with a value ranging from 0.04 to 0.06;

[0179] Then store the Harris response values of the feature points and the corresponding pixel positions as key-value pairs, sort them in descending order of the response values, set the retention quantity N, and according to the sorting result, filter and retain the top N feature points with the highest response values (rounded up);

[0180] Finally, update the original feature point set, use the BRIEF algorithm for feature description, form the feature descriptor sets of the real-time image and the reference image, referred to as the feature set, and represent them as P = {p1, p2.., p i ,...p m} and P′ = {p1', p2'...p i '..p n '}, respectively. Among them, p i and p i ' represent the feature descriptors in the two sets respectively.

[0181] Furthermore, the specific content of step 7 is as follows:

[0182] Create a brute-force matcher. Since the ORB feature descriptor is used for matching, set the brute-force matching parameter to the Hamming distance mode. Take the descriptor of a feature in the first feature set P = {p1, p2.., p i ,...p m}, and match it with the descriptors of all features in the second feature set P′ = {p1', p2'...p i '..p n '}. According to the matching result, return the feature point with the smallest Hamming distance, that is, the feature point with the highest matching degree, to complete the matching of a single feature point. Then loop through all the feature points to complete the matching of all feature points and form the rough matching point pairs;

[0183] Using RANSAC screening to solve the rough-matching homography matrix H1 based on the rough-matching point pairs. H1 represents the transformation relationship between the real-time image and the reference image based on the rough-matching point pairs, and predicts the position of the feature points in the real-time image on the reference image;

[0184] Solve the reprojection error e of each pair of rough-matching point pairs according to the pixel positions of the predicted points and the matching points i Let q i and be the observed point and the predicted point on an image. For a two-dimensional image, the reprojection error e on the reference image is the Euclidean distance between q i and , as shown in formula (21);

[0185]

[0186] wherein, is the pixel coordinate of the observed point, is the pixel coordinate of the predicted point. In the homogeneous coordinate system, the coordinate of the matching point on the reference image is (x′, y′, 1), and the coordinate of the predicted point on the reference image is The reprojection error e i is calculated as shown in formula (22):

[0187]

[0188] wherein, is the homogeneous coordinate component for normalization, taking

[0189] According to the above reprojection error calculation process, obtain the matching performance evaluation of all rough-matching point pairs;

[0190] According to the reprojection error e i of each pair of rough-matching point pairs, sort all the matching point pairs from high to low. When the number of rough-matching point pairs is greater than 2n pairs, keep the first n matching point pairs (rounded up). When the number of rough-matching point pairs is less than 2n pairs, keep the first 50% matching point pairs (rounded up). At this time, the retained matching point pairs are the fine-matching point pairs;

[0191] Using RANSAC screening to solve the fine-matching homography matrix H2 based on the fine-matching point pairs, predict the position of the feature points in the real-time image on the reference image, and then solve the reprojection error of each pair of fine-matching point pairs;

[0192] According to the magnitude of the reprojection error of each pair of accurately matched points, retain 80% of the pairs of accurately matched points with the smallest reprojection error (rounded up), and the retained matched point pairs at this time are the positioning point pairs; then use RANSAC to filter and solve the positioning homography matrix H3 based on the positioning point pairs.

[0193] Further, the specific steps of step 8 are as follows:

[0194] The homography matrix can achieve point-to-point mapping from one image to another. Assuming that the homogeneous coordinates of the matched points in the real-time image and the reference image are (u, v, 1) and (u', v', 1), respectively, the homography matrix can describe the mapping relationship of this pair of matched points, as shown in formula (23):

[0195]

[0196] Among them, H3 represents the positioning homography matrix. After obtaining the predicted pixel coordinates of the center pixel point of the real-time image on the reference image according to the homography transformation relationship, the pixel coordinate positioning of the two images is completed;

[0197] The corresponding relationship of the Mercator projection coordinates stored in the tiff-format reference image is realized through the origin and row and column parameters. The reference image origin pixel coordinates (0, 0) are calibrated as a Mercator projection coordinate (x0, y0), and then according to the row parameter P w and the column parameter P h Solve the projection coordinates corresponding to the pixel coordinates (u', v'):

[0198]

[0199] In formula (24), (x0, y0) is the projection coordinate of the upper left corner pixel of the reference image, (u', v') is the pixel coordinate, P w and P h are the row parameter and column parameter of the reference image, representing the pixel resolution, and (x, y) is the projection coordinate after the transformation of (u', v'). Based on the above calculations, the transformation from the reference image pixel coordinates to the Mercator projection coordinates is completed;

[0200] The origin of the Mercator projection is set as the intersection point (λ0, φ0) of the equator and the prime meridian, where λ0 = 0°, φ0 = 0°. The conversion of the longitude and latitude coordinates (λ, φ) to the projection coordinates (x, y) is as shown in formula (25) and formula (26):

[0201] x = j·(λ - λ0) (25)

[0202]

[0203] Among them, j is the projection scale factor, and is the radius of the earth. In the WGS-84 coordinate system, j ≈ 6378137 m;

[0204] Based on the above coordinate transformation calculation, the conversion from image pixel coordinates to longitude and latitude coordinates is realized, the navigation positioning coordinate calculation is completed, and the navigation result is output.

[0205] To verify the effectiveness of the method proposed in the present invention, a simulation verification is carried out on the scene matching navigation method based on region segmentation and feature optimization.

[0206] The data processing platform of the present invention uses a CPU of AMD Ryzen 7 5800H, with 8 cores and 16 threads, a main frequency of 3.2 GHz, and a memory of 16 GB DDR4, with a dual-channel memory design, which can meet the computing requirements of the simulation experiment of the scene matching navigation method based on region segmentation and feature optimization.

[0207] The software platform runs on the Windows 11 operating system. The development environment is selected as Visual Studio 2022. The core algorithms and functional modules are written in the C++ language, using the modern C++ standard (C++17 and above). To implement the key functions of scene matching navigation, the present invention introduces the open-source library OpenCV 4.8.0 for image processing and computer vision-related functions. The experimental data used is a UAV real-shot dataset.

[0208] The UAV real-shot image dataset is formed by shooting and making with an airborne optical camera of the UAV. The shooting height is 110 m, and the size of the shot images is set to 3000×2000 pixels. There are large image differences between the real-shot images at a height of 110 m and the directly downward reference images, which can simulate the real-time images of helicopter scene matching navigation in complex scenarios. The reference map is constructed from a large number of real-shot images based on the Dajiang Zhitu software. Cropping and processing the reference map can simulate the reference images read by helicopter scene matching navigation.

[0209] A total of 50 groups of image data are taken from the UAV real-shot image dataset in the simulation experiment. The corresponding relationships of each group of images can be calibrated in advance. The UAV GNSS RTK data at the time of real-time image shooting is used as the true position of the central pixel of the real-time image. Each pixel of the reference image in tiff format can read the longitude and latitude position, and the pixel correspondence between the two images can be realized, so as to obtain the corresponding pixel coordinates as the matching positioning true value.

[0210] According to the above experimental platform and experimental data design, the simulation experiment image database of the present invention is as Figure 3 shown.

[0211] The simulation image data of helicopter scene matching navigation in complex scenarios used in the present invention is asFigure 3 As shown, the reference image uses a directly downward viewing angle image, denoted by the suffix 0; the real-time image uses a complex and variable image, denoted by the suffixes 1 to 5. The above data can simulate image differences such as scale changes, viewing angle changes, and lighting changes faced during helicopter flight, and can meet the simulation verification requirements of image matching algorithms in complex environments.

[0212] Taking the 040 image data as an example, after extracting the image edge intensity information based on the Scharr operator, the method proposed in the present invention is used to achieve region segmentation, and the effect is as Figure 4 shown.

[0213] Figure 4 In, the black area represents the flat area segmented by the method of the present invention, and a feature detection mask matrix is constructed in the black area to complete the region segmentation of the image.

[0214] Let the calculation method for realizing scene matching navigation and positioning based on the traditional ORB algorithm be simply called the ORB algorithm; the calculation method for realizing scene matching navigation and positioning by only using the feature optimization method without using the region segmentation method be simply called the HR-ORB algorithm; the calculation method for realizing scene matching navigation and positioning by the method proposed in the present invention be simply called the SHR-ORB algorithm.

[0215] The ORB algorithm, HR-ORB algorithm, and SHR-ORB algorithm are respectively used to perform matching and positioning calculations on 50 groups of images, where there are a total of 10 matching scenarios, and each matching scenario has 5 groups of images with different shooting positions. Selecting typical image matching experiments among them, the distribution of positioning point pairs of different algorithms is as Figure 5 shown.

[0216] Figure 5 In, the red box marks the point pairs in the flat area, and such point pairs are usually mis-matching point pairs or point pairs with too dense distribution. The yellow box marks the matching situation in the corresponding area after optimization by the SHR-ORB algorithm. From Figure 5 it can be seen that the ORB algorithm and HR-ORB algorithm are severely interfered by the feature points in the flat area, resulting in mis-matching point pairs and dense matching point pairs, and the overall matching point pairs are not stable enough. In contrast, the SHR-ORB algorithm has stronger overall consistency of matching point pairs. This algorithm can realize region segmentation processing, increase the global constraint of feature points, reduce the detection of feature points in the flat area, and thus optimize the robustness of image matching point pairs.

[0217] In order to verify the improvement effect of the SHR-ORB algorithm on the positioning accuracy, the positioning errors of the three algorithms in 50 groups of matching and positioning experiments are statistically analyzed, and the average absolute error of image matching for each scenario is solved to form an error curve as Figures 6 - 8 shown.

[0218] From Figures 6 - 8It can be seen from the error curve that the fluctuation ranges of the eastward positioning error and the northward positioning error of the SHR-ORB algorithm are small, which reflects the stability of the matching positioning of the SHR-ORB algorithm. Compared with other algorithms, the SHR-ORB algorithm significantly improves the image matching positioning accuracy and has a small total positioning error.

[0219] To quantify the experimental data, record the eastward positioning error, the northward positioning error and the total positioning error of the three algorithms in 10 scenarios, as shown in Table 1, Table 2 and Table 3:

[0220] Table 1 Comparison list of the eastward positioning errors of each algorithm in different scenarios

[0221]

[0222] Table 2 Comparison list of the northward positioning errors of each algorithm in different scenarios

[0223]

[0224] Table 3 Comparison list of the total positioning errors of each algorithm in different scenarios

[0225]

[0226] It can be seen from the data in Table 3 that in the image matching scenario of the on-board camera of the unmanned aerial vehicle, compared with other algorithms, the SHR-ORB algorithm designed in the present invention can significantly reduce the image matching positioning error. The average positioning error of the ORB algorithm in 50 groups of image matching experiments is 3.806 meters, the average positioning error of the HR-ORB algorithm in 50 groups of image matching experiments is 3.392 meters, and the average positioning error of the SHR-ORB algorithm in 50 groups of image matching experiments is 2.392 meters. The positioning accuracy of the SHR-ORB algorithm is improved by 29.48% on the basis of the HR-ORB algorithm and 37.15% on the basis of the ORB algorithm, and the overall positioning error of the SHR-ORB algorithm is small, verifying the effectiveness and accuracy of the method of the present invention.

[0227] Use the ORB algorithm, the HR-ORB algorithm and the SHR-ORB algorithm to perform matching positioning calculations on 50 groups of images respectively, and count the calculation time consumed in the feature detection stage and the matching positioning stage of the 50 groups of matching positioning experiments of each algorithm, and form the average calculation time consumption statistics of each stage as shown in Table 4:

[0228] Table 4 Average time consumption list of image matching of each algorithm

[0229]

[0230] As can be seen from the experimental data in Table 4, the time consumption of each stage of the SHR-ORB algorithm is close to that of the HR-ORB algorithm, indicating that the introduction of the region segmentation method does not affect the computational time consumption. Moreover, by optimizing the feature points and matching point pairs, the computational time consumption of subsequent matching and positioning is reduced, and the overall impact on the real-time performance of the scene matching navigation system is small.

[0231] As can be seen from the above experimental data, the scene matching navigation method based on region segmentation and feature optimization proposed in the present invention can effectively improve the problem of false matching, improve the positioning accuracy of scene matching navigation without affecting the real-time performance, ensure the positioning performance of the scene matching navigation system, and provide an effective solution to the problem of false matching caused by redundant features in the flat area of the image.

[0232] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A scene matching navigation method based on region segmentation and feature optimization, characterized in that Including the following steps: Step 1: Generate a reference image according to the position indicated by the helicopter's airborne inertial navigation system, read the additional information of the reference image, and read the real-time image transmitted by the airborne image sensor. The additional information is the positioning parameters of the image; Step 2: Calculate the edge intensity of the reference image and the real-time image based on the Scharr operator to form an edge intensity image model of the two images; Step 3: Initially realize image region segmentation according to the edge intensity and image opening and closing operations, and segment the flat regions and texture-rich regions in the image; Step 4: Fill the flat regions in Step 3 according to the image connected region calculation method, and fill the texture-rich regions in Step 3 according to the image closing operation method to complete the image region segmentation; Step 5: Generate a mask matrix according to the image region segmentation result, and use the ORB algorithm to extract the feature points of the reference image and the real-time image based on the mask matrix; Step 6: Optimize the feature points of the reference image and the real-time image based on the Harris response value calculation method, retain a predetermined number of feature points, and the response values of the retained feature points are higher than those of the un-retained feature points, and form a feature descriptor; Step 7: Perform feature matching on the feature points retained in Step 6 to form a set of rough matching point pairs, calculate the reprojection error of the rough matching point pairs, retain a predetermined number of matching point pairs as positioning point pairs, and the reprojection error of the retained matching point pairs is less than that of the un-retained matching point pairs. Solve the homography matrix based on the RANSAC algorithm according to the positioning point pairs; Step 8: Complete image pixel positioning through homography transformation, calculate the navigation coordinates according to the pixel positioning result and the additional information of the reference image, and output the navigation result to complete scene matching navigation.

2. The scene matching navigation method based on region segmentation and feature optimization according to claim 1, wherein The specific content of Step 1 is as follows: The reference image database is prepared in advance. Obtain the approximate range of the current position according to the navigation information of the helicopter's airborne inertial navigation system at the current moment, select the reference image for matching. The additional information of the reference image is the positioning parameters of the image, including the Mercator projection coordinates corresponding to the upper left corner pixel point and the projection coordinate difference corresponding to one pixel grid. The real-time image is an RGB optical image taken by an optical camera carried by the helicopter from an aerial perspective. When reading, convert the RGB optical image into a grayscale image.

3. The scene matching navigation method based on region segmentation and feature optimization according to claim 2, characterized in that The specific content of Step 2 is as follows: Convolution kernel K of the Scharr operator x and K y As shown in formulas (1) and (2): For the reference image and the real-time image, combined with the convolution kernels K x and K y Calculate the gray-scale gradient changes at each pixel position of the two images respectively: G x (x,y) = I(x,y) * K x (3) G y (x,y) = I(x,y) * K y (4) In formulas (3) and (4), I(x, y) is the gray value at the position (x, y) of the image, and * is the convolution operation, which means scanning the image with the convolution kernel, G x (x, y) is the horizontal gradient component at the position (x, y) of the image, indicating the degree of gray change of the image in the horizontal direction; G y (x, y) is the vertical gradient component at the position (x, y) of the image, indicating the degree of gray change of the image in the vertical direction; According to formula (3) and formula (4), combined with the convolution kernels K x and K y calculate the gradients G x and G y of the image in the horizontal and vertical directions, revealing the intensity of the pixel gray value change in the x and y directions of the image. Use G x (x,y) and G y (x,y) to synthesize the gradient magnitude: Combine the horizontal gradient component and the vertical gradient component to synthesize the gradient magnitude G(x, y). Let the edge intensity distribution of the image be an m×n matrix M e , then the edge intensity distribution matrix M e is as shown in Equation (6): In order to apply the image edge intensity distribution in image processing and facilitate subsequent region segmentation based on the Scharr edge intensity, a scaling method is used for M e Linearly normalize to construct an edge intensity model Normalize the Scharr edge intensity amplitude within the image grayscale display range; Let M e The maximum and minimum edge intensities of max are G min and G max . Let the normalized maximum edge intensity be assigned as G min = 255, and the normalized minimum edge intensity be assigned as G e = 0. After normalizing the edge intensity to the numerical range of [0, 255], the normalization formula for the element G i in the edge intensity distribution matrix M e is as shown in formula (7): The normalized edge intensity model is obtained based on Equation (7). As shown in Equation (8): Set the part with edge intensity higher than the edge intensity threshold to 255, and the part with edge intensity lower than the threshold to 0. Then perform binarization to obtain the edge intensity image model I e .

4. The scene matching navigation method based on region segmentation and feature optimization according to claim 3, characterized in that The specific content of Step 3 is as follows: The opening and closing operations rely on a structuring element (SE) to perform image processing operations, controlling the specific behaviors of the erosion operation and the dilation operation. Among them, the erosion operation moves the structuring element SE on the image and replaces the gray value at the current pixel position with the minimum value in the local area of the image, as shown in formula (9): In formula (9), z represents the current pixel position, and SE z represents the set when the anchor point of the structural element SE is at the z position. The pixel at the z position will be retained in the erosion result if and only if all pixels of SE are within the foreground range of the image I; The dilation operation moves the structuring element SE on the image and replaces the gray value at the current pixel position with the maximum value in the local area of the image, as shown in formula (10): In formula (10), SE z represents the set when the anchor point of the structural element SE is at the z position. When any pixel in SE intersects with the foreground, the pixel at the z position will be retained in the dilation result; Based on the above corrosion operation and dilation operation, for the edge strength image model I e perform opening and closing operations, and based on the image edge information of I e integrate the discrete edge pixel points, extract the flat regions in the image model, so as to achieve the effect of region segmentation. First, perform an image opening operation on I e perform an erosion operation on the image first and then a dilation operation, as shown in formula (11): First, remove some pixels in the foreground region of I e through an erosion operation, which smooths the boundary. The foreground region is the region with rich texture. Then, restore the edges of the foreground region that were lost due to the erosion operation through a dilation operation; Then perform an image closing operation on I e by first performing a dilation operation on the image and then an erosion operation, as shown in Equation (12): First, remove some pixels in the background area of I e to fill small areas of the background and connect adjacent foreground areas. The background area is the flat area. Then, restore the edges of the background area lost due to the dilation operation through an erosion operation to smooth the boundaries, completing the preliminary image region segmentation based on image opening and closing operations.

5. The scene matching navigation method based on region segmentation and feature optimization according to claim 4, characterized in that The specific content of Step 4 is as follows: In the connectivity calculation, the pixel connectivity judgment criterion is: According to whether the gray values of the eight pixels in the up, down, left, right, and four diagonal directions of the current pixel in the image are the same as those of the current pixel, judge whether the surrounding pixels are all connected to it; The connected component labeling and region merging of the edge intensity image model I after morphological opening and closing operations are completed using the Two-Pass method. e In the first pass, by traversing each pixel of I e When the pixel is in the foreground region, the gray-level relationships of its neighboring pixels are calculated in three cases: (1) If all neighboring pixels are located in the background region, assign a new label L to the current pixel z i ; (2) If there is a pixel in the neighborhood that has been labeled L i then the current pixel z inherits the label L of the existing pixel in the neighborhood i ; (3) If there are multiple pixels with different marked labels in the neighborhood, assign any one of these labels to the current pixel z, and record the equivalence relationship between these labels; During the second scan, based on the labeling results of the first scan, equivalent label merging is performed on each connected region. Based on the label categories generated from the first scan results, during the second scan, each pixel label is replaced with the representative label in the equivalent class. Usually, the final connected region label is used. Here, the equivalent class refers to a set of region labels with the same connection relationship. Based on the two-scan method, for the edge intensity image model I e all pixels belonging to the same connected region in it are labeled with the same label, and the calculation of the connected region is completed; After quickly obtaining the calculation results of connected regions using the secondary scanning method, retain the blocks in I e where the area of the connected regions is greater than 100 square pixels. Then, construct a new structural element SE' and perform an image closing operation once to filter out small background blocks and complete image region segmentation.

6. The scene matching navigation method based on region segmentation and feature optimization according to claim 5, characterized in that The specific steps of step 5 are as follows: After the image region segmentation in step 4, let the extracted flat region be represented by R f and the region with rich texture be represented by R d . According to the flat area R f Create the corresponding mask matrix R f_mask , let R f_mask In the matrix R f The element value at the pixel position corresponding to the area is 0, and the element value at the pixel position corresponding to the non-R f area is 255, and the construction of the mask matrix R f_mask is completed to form the mask matrix R f_mask As shown in formula (13): Based on the mask matrix R f_mask Omit the feature detection of flat areas on the original image, create an ORB feature detection function, and introduce the mask matrix R after setting the detection parameters f_mask , and extract the set of feature points of the reference image and the real-time image.

7. The scene matching navigation method based on region segmentation and feature optimization according to claim 6, characterized in that The specific steps of step 6 are as follows: Record the pixel positions (x, y) and index information of all FAST feature points of the real-time image and the reference image to form a FAST feature point pixel position vector; According to the pixel position vector, construct a structure tensor matrix M at each feature point pixel position (x, y) to describe the gradient distribution within the window. Among them, the calculation of the structure tensor matrix M is shown in formula (14): where w(x, y) is a weighting function, a Gaussian weight, I x and I y represent the gradients of the current pixel position in the x and y directions respectively, W is a local window in the image for calculating information of surrounding pixels, and M is represented by the following three elements A, B, and C: According to the structure tensor matrix M, calculate the magnitude of the Harris response value R at each corresponding position of the feature points, and record the position index information. Among them, the calculation of the response value R is shown in formula (18): R = detM - k(traceM) 2 (18) detM = λ1λ2 = AB - C 2 (19) traceM = λ1 + λ2 = A + B (20) Among them, λ1 and λ2 are the eigenvalues of matrix M, representing the direction and magnitude of the image gradient change, detM represents the determinant of matrix M, traceM represents the trace of matrix M, and k is an empirical constant with a value ranging from 0.04 to 0.06; Then store the feature points and the Harris response values of the corresponding pixel positions as key-value pairs, sort them in descending order of the response values, set the retention quantity N, and according to the sorting result, screen and retain half of the N feature points with the highest response values; Finally, update the original set of feature points, and use the BRIEF algorithm for feature description to form the feature descriptor sets of the real-time image and the reference image, which are represented by P = {p1, p2.., p i ,...p m} and P′ = {p1', p2'...p i '..p n '}, respectively. Among them, p i and p i ' represent the feature descriptors in the two sets respectively.

8. The scene matching navigation method based on region segmentation and feature optimization according to claim 7, characterized in that The specific steps of step 7 are as follows: Create a brute-force matcher. Since ORB feature descriptors are used for matching, set the brute-force matching parameter to Hamming distance mode. Take the descriptor of one feature from the first feature set P = {p1, p2.., p i ,...p m}, and match it with the descriptors of all features in the second feature set P' = {p1', p2'...p i '..p n '}. Return the feature point with the minimum Hamming distance based on the matching result, that is, the feature point with the highest matching degree, to complete the matching of a single feature point. Then loop through all feature points to complete the matching of all feature points and form rough matching point pairs; Based on the rough matching point pairs, use RANSAC to screen and solve the rough matching homography matrix H1. Use H1 to represent the transformation relationship between the real-time image and the reference image based on the rough matching point pairs, and predict the predicted point positions of the real-time image feature points on the reference image; Solve the reprojection error e of each pair of coarsely matched point pairs according to the pixel positions of the prediction points and the matched points i Let q i and be the observed point and the predicted point on an image. For a two-dimensional image, the reprojection error e on the reference image is the Euclidean distance between q i and as shown in formula (21); Among them, is the pixel coordinate of the observation point, is the pixel coordinate of the prediction point. In the homogeneous coordinate system, the coordinate of the matching point on the reference image is (x′, y′, 1), and the coordinate of the prediction point on the reference image is The reprojection error e i is calculated as shown in formula (22): Among them, is the homogeneous coordinate component for normalization, taking According to the above reprojection error calculation process, obtain the matching performance evaluation of all rough matching point pairs; According to the reprojection error e of each pair of coarsely matched points i sort all the matched point pairs from high to low according to the size. When the number of coarsely matched point pairs is greater than 2n pairs, retain the first n matched point pairs. When the number of coarsely matched point pairs is less than 2n pairs, retain the first 50% of the matched point pairs. The retained matched point pairs at this time are the finely matched point pairs; Based on the fine matching point pairs, use RANSAC to screen and solve the fine matching homography matrix H2, predict the predicted point positions of the real-time image feature points on the reference image, and then solve the reprojection error of each pair of fine matching point pairs; According to the magnitude of the reprojection error of each pair of fine matching point pairs, retain 80% of the fine matching point pairs with the smallest reprojection error. At this time, the retained matching point pairs are the positioning point pairs; then use RANSAC to screen and solve the positioning homography matrix H3 based on the positioning point pairs.

9. The scene matching navigation method based on region segmentation and feature optimization according to claim 8, characterized in that, The specific steps of step 8 are as follows: The homography matrix can achieve a point-to-point mapping from one image to another. Suppose the matching points in the real-time image and the reference image are in homogeneous coordinates (u, v, 1) and (u', v', 1), then the homography matrix can describe the mapping relationship of this matching point pair, as shown in formula (23): Among them, H3 represents the positioning homography matrix. After obtaining the predicted pixel coordinates of the center pixel point of the real-time image on the reference image according to the homography transformation relationship, the pixel coordinate positioning of the two images is completed; The Mercator projection coordinate correspondence stored in the TIFF format reference image is realized through the origin and row and column parameters. The reference image origin pixel coordinates (0, 0) are calibrated as a Mercator projection coordinate (x0, y0), and then according to the row parameter P w and the column parameter P h calculate the projection coordinates corresponding to the pixel coordinates (u', v'): In formula (24), (x0, y0) are the projected coordinates of the pixel at the upper left corner of the reference image, (u', v') are the pixel coordinates, P w and P h are the row parameter and column parameter of the reference image, representing the pixel resolution, (x, y) are the projected coordinates after the transformation of (u', v'), and based on the above calculations, the transformation from the pixel coordinates of the reference image to the Mercator projection coordinates is completed; The origin of the Mercator projection is set as the intersection point (λ0, φ0) of the equator and the prime meridian, where λ0 = 0°, φ0 = 0°. The conversion of the longitude and latitude coordinates (λ, φ) to the projection coordinates (x, y) is shown in formula (25) and formula (26): x = j·(λ - λ0) (25) Among them, j is the projection scale factor, which is the radius of the earth. In the WGS-84 coordinate system, j ≈ 6378137 m; Based on the above coordinate transformation calculation, the conversion from image pixel coordinates to longitude and latitude coordinates is realized, the navigation positioning coordinate solution is completed, and the navigation result is output.