Image feature point matching method and system for endoscopic surgery examination scene

By using deep learning feature descriptor extraction network and self-attention mechanism in endoscopic surgical scenarios, the matching accuracy reduction caused by low texture, parallax changes and motion blur is solved, and a higher matching accuracy and positioning reliability of the surgical navigation system are achieved.

CN120164005APending Publication Date: 2025-06-17SUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510281411.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In endoscopic surgery scenarios, the prior art is difficult to effectively solve the lack of feature distinction in low-texture environments, limitations of local features extracted in fixed windows, and motion fuzzy interference, resulting in a decrease in endoscopic image matching accuracy and affecting the positioning reliability of the surgical navigation system.

Method used

A feature descriptor extraction network based on deep learning is adopted, and the source image and target image of the endoscope are input to the feature extraction network with shared weights, and a multi-channel feature descriptor diagram is obtained. Through steps such as coarse matching, fine matching and false matching filtering, the sampling window size is adaptively adjusted, and fine matching optimization is used to use self-attention and cross-attention mechanisms.

Benefits of technology

It improves the accuracy of image feature point matching in endoscopic surgery scenarios, enhances the adaptability and accuracy of feature matching, improves the positioning reliability of the surgical navigation system, and provides more accurate visual guidance support for surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164005A_ABST
    Figure CN120164005A_ABST
Patent Text Reader

Abstract

The invention provides an image feature point matching method and system for an endoscopic surgery examination scene, and relates to the technical field of medical treatment and computer science, and the method comprises the steps: inputting a source image and a target image into a feature extraction network sharing the weight, and obtaining a multi-channel feature descriptor graph; then obtaining source image feature point coordinates and sampling to obtain a descriptor vector, and obtaining a rough matching result through correlation calculation; filtering mismatching by using a random sampling consistent algorithm, and estimating a camera pose transformation matrix; adjusting the size of a sampling window according to the depth proportion relationship; fine matching optimization is carried out by using an attention mechanism; and finally, finely adjusting the coordinates according to the matching probability graph, and outputting an optimized matching result. According to the method, the image matching problem in an endoscopic surgery scene is effectively solved, the interval multi-frame image matching accuracy is remarkably improved, a denser three-dimensional point cloud can be reconstructed when a motion structure is recovered, in-vivo cavity texture can be better observed in an assisted mode, and surgery navigation positioning reliability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of medical and computer science technologies, and particularly relates to an image feature point matching method and system for an endoscopic surgical examination scenario. Background Art

[0002] Minimally invasive surgery has become an important alternative to traditional open surgery due to its advantages of small trauma, high efficiency, and remarkable curative effect. During the surgery, the endoscopic visual system provides key real-time visual guidance for surgeons, directly affecting the accuracy and safety of surgical instrument operation. In recent years, with the rapid development of artificial intelligence technology, computer-aided diagnosis systems have gradually been applied to the field of endoscopic examination, assisting doctors to more efficiently identify lesion areas through image processing and machine learning technologies, and improving the diagnostic accuracy. However, in minimally invasive surgery, the real-time motion trajectory tracking and spatial positioning of the endoscope still face major challenges, which corely rely on extracting stable feature points from continuous endoscopic image sequences and performing fast matching to achieve accurate estimation of the camera pose.

[0003] In the prior art, endoscopic surgical navigation mainly follows traditional computer vision methods, including the following steps:

[0004] 1. Feature point detection: Using operators such as SIFT, SURF, FAST, etc. to detect significant positions such as corner points and edge intersection points in the image, but there is a trade-off between computational complexity and stability for different operators;

[0005] 2. Feature descriptor extraction: Analyzing the gradient distribution or texture features in the area around the feature points to generate a numerical description vector, but the existing methods are difficult to cope with the low texture characteristics of endoscopic images;

[0006] 3. Feature matching: Calculating the similarity of descriptors through brute-force matching or approximate algorithms, and it is difficult to balance the matching speed and accuracy;

[0007] 4. Matching point screening: Using geometric constraint algorithms such as RANSAC to eliminate false matches, but its effectiveness depends on the initial matching quality.

[0008] However, the above methods have significant defects in the endoscopic surgical scenario:

[0009] 1. Insufficient feature discrimination in low-texture environments: The surface of internal organs lacks significant texture, resulting in too high similarity of feature descriptor vectors and a significant increase in the false match rate;

[0010] 2. Limitations of fixed-window extraction: Traditional methods use fixed windows to extract local features without considering the parallax changes caused by endoscopic movement (such as the scaling of the feature area when the camera distance changes), affecting the robustness of the descriptor;

[0011] 3. Motion blur interference: During surgical operations, the shaking of the endoscope easily causes image blurring, further reducing the accuracy of feature point detection and matching.

[0012] The deficiencies of the existing technologies directly lead to a decrease in the matching accuracy of endoscopic images, thereby affecting the positioning reliability of the surgical navigation system. Therefore, there is an urgent need for a feature matching technology optimized for endoscopic surgical scenarios to solve the core problems brought about by low texture, parallax changes, and motion blur, and to provide more accurate visual guidance support for clinical practice. Summary of the Invention

[0013] To this end, embodiments of the present invention provide an image feature point matching method and system for endoscopic surgical examination scenarios, which are used to solve the problems in the existing technologies that in endoscopic surgical scenarios, due to the low-texture environment, the lack of feature distinctiveness, the limitations of extracting local features with a fixed window, motion blur interference, etc., the matching accuracy of endoscopic images is reduced, affecting the positioning reliability of the surgical navigation system.

[0014] To solve the above problems, embodiments of the present invention provide an image feature point matching method for endoscopic surgical examination scenarios, and the method includes:

[0015] S1: Input the source image and the target image of the endoscope into a feature extraction network with shared weights to obtain multi-channel feature descriptor maps of the source image and the target image;

[0016] S2: Obtain the feature point coordinates of the source image based on a feature point detection operator, and sample in the source feature descriptor map to obtain source feature point descriptor vectors;

[0017] S3: Calculate the correlation between the source feature point descriptor vectors and the target feature descriptor map to generate a response value heat map on the target image, and extract the coordinates corresponding to the maximum response value in the heat map as the rough matching result;

[0018] S4: Filter out the mismatches in the rough matching result through the Random Sample Consensus (RANSAC) algorithm, and estimate the camera pose transformation matrix based on the remaining matching point pairs;

[0019] S5: Adaptively adjust the local sampling window sizes of the source image and the target image according to the depth ratio relationship of the matching point pairs;

[0020] S6: Based on the adjusted window sizes, perform feature descriptor block sampling on the local regions of the source image and the target image respectively, and use the self-attention mechanism and the cross-attention mechanism for fine matching optimization;

[0021] S7: Fine-tune the target image coordinates according to the matching probability map in the fine matching stage, and output the finally optimized feature point matching result.

[0022] Preferably, in step S1, the feature extraction network is a feature extraction network with dual - head output, where one output head is used to generate a de - blurred feature map, and the other output head is used to generate a multi - channel feature descriptor map with the same size as the input image.

[0023] Preferably, in step S5, the formula for adaptively adjusting the window size is:

[0024]

[0025] where, is the initial window size on the source image, s i is the scale factor calculated according to the depth ratio relationship of the matching point pairs, is the adjusted window size on the target image.

[0026] Preferably, in step S6, the fine - matching optimization includes:

[0027] Enhancing the correlation of feature descriptors within the local region through the self - attention mechanism;

[0028] Capturing the correspondence between the feature descriptor blocks of the source image and the target image through the cross - attention mechanism.

[0029] Preferably, before performing fine - matching optimization using the self - attention mechanism and the cross - attention mechanism, sample the target feature descriptor block to the size of the source feature descriptor block.

[0030] Preferably, in step S7, calculate the expectation according to the matching probability map in the fine - matching stage, and fine - tune the target image coordinates through the expectation result, where the formula for the expectation is:

[0031]

[0032] where, represents the target image coordinates of the i - th feature point after fine - tuning; P i represents the initial coordinates of the i - th feature point; is the adjusted window size on the target image, u, v are pixel coordinate variables; M i (u, v) represents the matching probability map; represents the sum of the matching probabilities of all pixel positions within the window.

[0033] Preferably, the method further includes designing a multi - stage loss function composed of the loss in the coarse - matching stage, the loss in the fine - matching stage, and the loss of the de - blurring module, for training the feature extraction network, where the multi - stage loss function is:

[0034]

[0035] Among them

[0036]

[0037]

[0038] In the formula represents the multi-stage loss function; λ1, λ2, λ3 represent adjustable hyperparameters; represents the loss in the coarse matching stage; represents the heat response map generated in the coarse matching stage; represents the heat response value at position P on the target image; α is a scaling factor; P gt represents the true matching coordinates on the target image; represents the heat response value at the true matching coordinates; P i heat response map the i-th position in; represents the loss in the fine matching stage; ||||2 represents the L2 norm; P i,gt represents the true coordinates of the i-th feature point in the target image; P′ i represents the predicted coordinates of the i-th feature point obtained after the fine matching stage; represents the loss of the deblurring module; is the mean square error between the output image of the feature extraction network and the label image; is the adversarial loss; G is the generator in the feature extraction head; D is the discriminator; represents the mathematical expectation; A~p actual (A) indicates that the variable A follows the probability distribution p of the real samples actual (A), A represents an image sample randomly drawn from the set of real clear images; B~p blurry (B) indicates that the variable B follows the probability distribution p of the blurred images blurry (B), B represents a sample randomly drawn from the set of blurred images.

[0039] The embodiment of the present invention also provides an image feature point matching system for the endoscopic surgical examination scenario. This system is used to implement the above-mentioned image feature point matching method for the endoscopic surgical examination scenario, and specifically includes:

[0040] A feature extraction module, which is used to input the source image and the target image of the endoscope into a feature extraction network with shared weights to obtain multi-channel feature descriptor maps of the source image and the target image;

[0041] A source feature point processing module, which is used to obtain the feature point coordinates of the source image based on a feature point detection operator and sample source feature point descriptor vectors in the source feature descriptor map;

[0042] A coarse matching module, which is used to calculate the correlation between the source feature point descriptor vector and the target feature descriptor map, generate a response value heat map on the target image, and extract the coordinates corresponding to the maximum response value in the heat map as the coarse matching result;

[0043] A false match filtering and matrix estimation module, which is used to filter false matches in the coarse matching result by the random sample consensus algorithm and estimate the camera pose transformation matrix based on the remaining matching point pairs;

[0044] A window adjustment module, which is used to adaptively adjust the local sampling window sizes of the source image and the target image according to the depth ratio relationship of the matching point pairs;

[0045] A fine matching optimization module, which is used to perform feature descriptor block sampling on the local regions of the source image and the target image respectively based on the adjusted window size, and use the self-attention mechanism and the cross-attention mechanism for fine matching optimization;

[0046] A coordinate fine-tuning and output module, which is used to fine-tune the target image coordinates according to the matching probability map in the fine matching stage and output the finally optimized feature point matching result.

[0047] An embodiment of the present invention further provides an electronic device, which includes a processor, a memory and a bus system. The processor and the memory are connected through the bus system. The memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the above-mentioned image feature point matching method for the endoscopic surgery inspection scenario.

[0048] An embodiment of the present invention further provides a computer storage medium, which stores a computer software product. The computer software product includes several instructions for causing a computer device to execute the above-mentioned image feature point matching method for the endoscopic surgery inspection scenario.

[0049] It can be seen from the above technical solutions that the present invention application has the following beneficial effects:

[0050] (1) Higher matching accuracy: Aiming at the problem of insufficient feature discrimination in the low-texture environment of the endoscopic surgery scenario, the present invention designs a feature descriptor extraction network based on deep learning. By inputting the source image and the target image of the endoscope into the feature extraction network with shared weights, a multi-channel feature descriptor map is obtained. Combining a series of subsequent steps such as coarse matching and fine matching, when performing image pair matching at intervals of multiple frames on the clinical endoscopic surgery inspection dataset, a better matching accuracy than the traditional method is achieved, improving the accuracy of endoscopic image matching.

[0051] (2) Strong adaptive processing ability: Given the limitations of traditional fixed-window extraction of local features, the present invention divides the complete feature matching framework into two parts: rough matching and fine matching. After obtaining the initial matching coordinate results through rough matching, the depth ratio relationship is calculated based on visual geometry, and the local sampling window sizes of the source image and the target image are adaptively adjusted, solving the problem of feature description of local regions in traditional methods, better adapting to different endoscopic examination scenarios, and improving the adaptability and accuracy of feature matching.

[0052] (3) Assisting surgical observation and navigation: Since image motion blur problems will occur during the surgical operation process, the feature extraction network of the present invention is designed with dual outputs. One output head is used to obtain a deblurred feature map, so that the other output head in the feature extraction network obtains a feature descriptor map with higher discrimination for feature points in the motion-blurred area. When restoring the motion structure of the same endoscopic video sequence, a denser three-dimensional point cloud can be reconstructed, which helps doctors better observe the surface texture feature information in the narrow internal cavity, improves the reliability of the positioning of the surgical navigation system, and provides strong support for the precise operation of the surgery. Description of the Drawings

[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly describe the drawings required in the embodiments. By referring to the drawings, the features and advantages of the present invention will be more clearly understood. The drawings are schematic and should not be construed as limiting the present invention in any way. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:

[0054] Figure 1 It is a flowchart of an image feature point matching method for an endoscopic surgical examination scenario provided in the embodiment;

[0055] Figure 2 It is a schematic diagram of the experimental results of the comparison method according to three basic indicators under different data sets in the embodiment;

[0056] Figure 3 It is a schematic diagram of the effect comparison in feature point matching, feature depth estimation, and point cloud reconstruction using SIFT, Liu, and the method proposed in the present invention in the embodiment;

[0057] Figure 4 It is a schematic diagram of the feature point matching effect when comparing five methods at multiple-frame intervals in two data set scenarios in the embodiment;

[0058] Figure 5 It is a schematic diagram of the feature matching results of different methods in four data set scenarios in the embodiment;

[0059] Figure 6 It is a block diagram of an image feature point matching system for an endoscopic surgical examination scenario provided in the embodiment. Specific implementation manners

[0060] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0061] Embodiment 1

[0062] To solve the problems in the prior art that in the endoscopic surgical scenario, due to the low-texture environment, the feature distinctiveness is insufficient, there are limitations in extracting local features with a fixed window, motion blur interference, etc., which further cause the decline of the endoscopic image matching accuracy and affect the positioning reliability of the surgical navigation system. As Figure 1 shown, an image feature point matching method for an endoscopic surgical examination scenario is proposed in the embodiment of the present invention. The method includes:

[0063] S1: Input the source image and the target image of the endoscope into a feature extraction network with shared weights to obtain multi-channel feature descriptor maps of the source image and the target image;

[0064] S2: Obtain the feature point coordinates of the source image based on a feature point detection operator, and sample in the source feature descriptor map to obtain source feature point descriptor vectors;

[0065] S3: Calculate the correlation between the source feature point descriptor vectors and the target feature descriptor map to generate a response value heat map on the target image, and extract the coordinates corresponding to the maximum response value in the heat map as the rough matching result;

[0066] S4: Filter out the mismatches in the rough matching result through the random sample consensus algorithm, and estimate the camera pose transformation matrix based on the remaining matching point pairs;

[0067] S5: Adaptively adjust the local sampling window sizes of the source image and the target image according to the depth ratio relationship of the matching point pairs;

[0068] S6: Based on the adjusted window sizes, perform feature descriptor block sampling on the local regions of the source image and the target image respectively, and use the self-attention mechanism and the cross-attention mechanism for fine matching optimization;

[0069] S7: Fine-tune the target image coordinates according to the matching probability map in the fine matching stage, and output the finally optimized feature point matching result.

[0070] As can be seen from the above technical solution, in an image feature point matching method for endoscopic surgery examination scenarios according to the present invention, first, the source image and the target image are input into a feature extraction network with shared weights to solve the problem of insufficient feature discrimination in low-texture environments, and a multi-channel feature descriptor map is obtained, providing a basis for accurate matching. Then, a feature point detection operator is used to obtain the coordinates of the feature points in the source image and sample them to obtain descriptor vectors. Next, a heat map is generated through correlation calculation to determine the rough matching result, and the random sample consensus algorithm is used to filter out the mismatches and estimate the camera pose transformation matrix to improve the matching reliability. After that, the sampling window size is adaptively adjusted according to the depth ratio relationship to overcome the limitation of feature extraction with a fixed window. Then, based on the sampling with the adjusted window, the self-attention and cross-attention mechanisms are used for fine matching optimization. Finally, the coordinates of the target image are fine-tuned according to the matching probability map, and the optimized result is output. Experimental verification on the clinical dataset shows that this method has a higher accuracy when matching image pairs with multiple frames in between, and can reconstruct a denser three-dimensional point cloud in the motion structure recovery, facilitating the observation of the texture of the narrow internal cavity and improving the reliability of the positioning of the surgical navigation system.

[0071] In step S1, the source image and the target image of the endoscope are input into a feature extraction network with shared weights to obtain the multi-channel feature descriptor maps of the source image and the target image.

[0072] Given a video clip of an endoscopic surgery examination scenario, the goal is to complete the feature point matching between adjacent images or multiple frames of images at intervals, so as to finally achieve an accurate estimation of the camera pose.

[0073] Specifically, the source image I of the endoscope is obtained s and the target image I t . These images are usually in RGB format and have a size of H×W×3 (h represents the image height and W represents the image width). The source image I s and the target image I t are input into a dual-head output feature extraction network with shared weights. One output head generates a deblurred feature map, which is implicitly trained to help the network learn and obtain a more discriminative feature descriptor map for the feature points in the motion-blurred area at the other output head; the other output head generates a multi-channel feature descriptor map with the same size as the input image, denoted as the source feature descriptor map F s and the target feature descriptor map F t respectively. These descriptor maps will be used for subsequent feature point matching operations.

[0074] In step S2, based on the feature point detection operator, the coordinates of the feature points in the source image are obtained, and the source feature point descriptor vectors are sampled from the source feature descriptor map.

[0075] Using traditional feature point detection operators, such as SIFT, SURF, FAST, etc., or by selecting according to human experience, feature point detection is performed on the source image to obtain the coordinate of the feature points of the source image. According to these coordinates, pixel-level feature descriptor vector sampling is performed in the source feature descriptor map F s to obtain the source feature point descriptor vector f i (i represents the index of the feature points in the detected source image), and the size of each source feature point descriptor vector f i is 1×1×C (C represents the feature channel length of the feature descriptor).

[0076] In step S3, the source feature point descriptor vector is calculated with the target feature descriptor map to generate a heat map of response values on the target image, and the coordinates corresponding to the maximum response value in the heat map are extracted as the rough matching result.

[0077] Specifically, the source feature point descriptor vector f i is calculated with the target feature descriptor map F t , which is specifically implemented through a convolution operation. After calculation, a heat map of response values on the target image is generated.

[0078] The response value of each pixel in the heat map reflects the correlation degree between the source feature point descriptor and the corresponding target feature descriptor. The coordinates with the maximum response value in the heat map are extracted and used as the rough matching result of the source image feature points on the target image. The following formula is used to calculate the obtained rough matching coordinates:

[0079]

[0080] where (u j , v j ) represents the coordinates on the target image, and argmax represents the image coordinates of obtaining the maximum value from the single-channel heat response value map, thereby obtaining the image correspondence in the rough matching stage.

[0081] After calculating the rough matching correspondence, position encoding is performed on the source feature descriptor map and the target feature descriptor map {F s , F t}, and the source feature descriptor map and the target feature descriptor map {F ps , F pt} after adding position dependence are prepared for the subsequent fine matching stage.

[0082] In step S4, the random sample consensus algorithm is used to filter the mismatches in the rough matching results, and the camera pose transformation matrix is estimated based on the remaining matching point pairs.

[0083] Specifically, during the transition from the rough matching stage to the fine matching stage, the present invention uses the Random Sample Consensus algorithm (RANSAC) to filter out the mismatches in the rough matching results, ensuring a reliable transformation matrix is estimated from the remaining correct matching results. For any pair of matching points in the filtered rough matching results, the reliable transformation matrix obtained by estimation can be used to transform from the source image coordinate position to the target image coordinate position, that is, the following formula is satisfied:

[0084] d j c j = d i R·c i + ηT;

[0085] where R and T represent the transformation from the camera coordinate system of the source image to the target image coordinate system, d i and d j respectively represent the depth values of the source point and the target point relative to the same three-dimensional space coordinate, and η is the scaling ratio. c i and c j are respectively defined as and where (u i , v i ), (u j , v j ) respectively represent the coordinate pairs of the matching points in the source image and the target image, K i and K j are defined as the camera internal parameters of the corresponding camera images. In the same endoscopic surgical examination scenario, the same camera is used for image acquisition, so K i = K j . After derivation, it can be obtained that:

[0086]

[0087] where, represents calculating the average value of the entire vector, represents the element-wise division of two vectors, and ∧ represents the outer product.

[0088] In step S5, according to the depth ratio relationship of the matching point pairs, the local sampling window sizes of the source image and the target image are adaptively adjusted.

[0089] By estimating the characteristic depth ratio of the two camera viewpoints, the scale factor s i can be obtained. Using the calculated scale factor, the window size of the local area sampling can be adjusted adaptively. By initializing the local window size defined for sampling on the source image, according to the scale factor s i the local window size for sampling on the target image can be obtained. The calculation formula for adaptively adjusting the window size is as follows:

[0090]

[0091] Among them, is the initial window size on the source image, s i is the scale factor calculated according to the depth ratio relationship of the matching point pairs, is the adjusted window size on the target image. In this way, the sampling window can be adaptively adjusted under different depths to better capture image features.

[0092] In step S6, based on the adjusted window size, feature descriptor blocks are sampled for local regions of the source image and the target image respectively, and fine matching optimization is performed using the self-attention mechanism and the cross-attention mechanism.

[0093] Specifically, entering the fine matching stage, the source feature descriptor map and the target feature descriptor map {F ps , F pt} after position encoding, the filtered initial feature matching results, and the corresponding local sampling window sizes are obtained. Centered on the initial paired feature point coordinates on the source image and the target image, local regions are sampled according to the calculated corresponding window sizes, and thus the source feature descriptor block and the adaptively adjusted target feature descriptor block are obtained accordingly In this more refined matching stage, we give priority to the correlation of local region features and use the attention mechanism to accurately capture intricate visual features. The attention mechanism only operates on the cropped source feature descriptor block and the adaptively adjusted target feature descriptor block instead of running on the global feature map, thus ensuring computational efficiency. The attention module consists of a self-attention mechanism and a cross-attention mechanism. Among them, the self-attention mechanism pays more attention to the feature information between each feature descriptor vector in the image, and the cross-attention mechanism pays more attention to the feature information between the corresponding feature descriptor vectors of the paired images. It should be noted that before the source feature descriptor block and the target feature descriptor block enter the attention mechanism, the target feature descriptor needs to be sampled to the size of the source feature descriptor block, aiming to avoid the problem of video memory space explosion caused by storing target feature descriptor vectors of different sizes. After passing through the attention mechanism, the processed source and target feature descriptor vectors are obtained By extracting the central feature descriptor vector at and performing a correlation calculation with , the formula for calculating the local correlation is as follows:

[0094]

[0095] Among them, τ represents the temperature coefficient of the softMax() function, which is used to adjust the smoothness of the output probability distribution map. It means to perform cumulative calculation along the feature channel direction. Thus, the matching probability map of the source feature descriptor to the local target feature descriptor map is obtained.

[0096] In step S7, the target image coordinates are finely adjusted according to the matching probability map in the fine matching stage, and the final optimized feature point matching result is output.

[0097] Specifically, by calculating the expected result of the matching probability map, the initial matching result on the target image is finely adjusted, and finally the paired relationship between the optimized source image and the target image is obtained. The process of calculating the expectation is as follows:

[0098]

[0099] Among them, represents the target image coordinates of the i-th feature point after fine adjustment; P i represents the initial coordinates of the i-th feature point; is the adjusted window size on the target image, and u, v are pixel coordinate variables; M i (u, v) represents the matching probability map; represents the sum of the matching probabilities of all pixel positions within the window.

[0100] In the final coordinate fine adjustment stage, by adjusting the adaptive window ratio of the target feature descriptor block, the scaling factor lost during the sampling process of the target feature descriptor block is restored.

[0101] In order to maintain the concept of matching from coarse to fine and combine the deblurring feature extraction head, the present invention designs an overall loss function composed of three parts: the loss in the coarse matching stage, the loss in the fine matching stage, and the loss of the deblurring module, which is a multi-stage loss function for training the feature extraction network.

[0102] First, a thermal response guidance mechanism is adopted to supervise the coarse matching stage to ensure that the correct matching coordinates generate the highest response value, while the non-matching regions show lower response values. To effectively guide the learning process, we define a loss term L c , which can supervise the network to generate a strong response at the correct matching position. The loss term can be expressed as:

[0103]

[0104] Among them, represents the loss in the coarse matching stage; represents the thermal response map generated in the coarse matching stage; represents the thermal response value at position P on the target image; α is the scale factor; P gt represents the true matching coordinates on the target image; represents the thermal response value at the true matching coordinates; P i thermal response map the i-th position in;

[0105] Secondly, to achieve hierarchical fine-grained matching optimization, a reprojection loss is introduced here. The Euclidean distance between the GT position and the refined coordinate P' after the coarse-grained and fine-grained matching stages is calculated using the L2 criterion, thereby obtaining the reprojection loss L i P' i between, and thus the reprojection loss L is obtained f as follows:

[0106]

[0107] where, represents the loss in the fine matching stage; ||||2 represents the L2 norm; P i,gt represents the true coordinates of the i-th feature point in the target image; P' i represents the predicted coordinates of the i-th feature point obtained after the fine matching stage.

[0108] Finally, to generate a robust and discriminative feature descriptor for blurred images in dynamic scenes, the deblurring module adopts the design concept of an adversarial generative network. This architecture encourages the hidden layer to capture the pixel-level differences between the blurred image and the clear image, while adversarial training promotes the learning of the clear image distribution by the feature extractor. The calculation formula of this loss function is as follows:

[0109]

[0110] where, is the mean square error between the output image of the feature extraction network and the label image; is the adversarial loss, and its definition is as follows:

[0111]

[0112] where, G is the generator designed in the first feature extraction head; D is the discriminator, which determines whether the input data comes from real samples or is generated by the generator.; represents the mathematical expectation; A~p actual (A) indicates that the variable A follows the probability distribution p of real samples actual (A), A represents an image sample randomly drawn from the set of real clear images; B~p blurry (B) indicates that the variable B follows the probability distribution p of blurred images blurry(B), where B represents a sample randomly drawn from the set of blurred images.

[0113] Adding up all these loss terms with tunable hyperparameters λ1, λ2, and λ3, the resulting total loss function is expressed as follows:

[0114]

[0115] Before starting the training, the tunable hyperparameters λ1, λ2, and λ3 are artificially defined and adjusted. By changing the order of magnitude between them, the possible impact on the convergence of the loss function is observed. After the hyperparameters are determined, λ1, λ2, and λ3 are kept fixed during the training process, and at the same time, other parameters in the feature extraction network are continuously optimized to gradually reduce the multi-stage loss function, thereby optimizing the performance of the feature extraction network and improving the accuracy of image feature point matching.

[0116] The experimental results part of the present invention includes comparative experimental data graphs and visualization schematic diagrams of comparative experimental results, specifically as follows: Figure 2 It is a schematic diagram of the experimental results of the comparative method according to three basic indicators under different data sets; Figure 3 It is a schematic diagram of the effect comparison in feature point matching, feature depth estimation, and point cloud reconstruction using SIFT, Liu, and the method proposed in the present invention respectively; Figure 4 It is a schematic diagram of the feature point matching effect comparison of five methods at multiple frames interval in two data set scenarios; Figure 5 It is a schematic diagram of the feature matching results of different methods in four data set scenarios. These charts clearly show the performance of the method of the present invention in different scenarios and data sets, as well as the comparison with other methods, further verifying the superiority and reliability of the present invention in image feature point matching in the endoscopic surgery inspection scenario.

[0117] Embodiment 2

[0118] As Figure 6 shown, the present invention provides an image feature point matching system for endoscopic surgery inspection scenarios. This system is used to implement the image feature point matching method for endoscopic surgery inspection scenarios in the above Embodiment 1, and specifically includes:

[0119] A feature extraction module 100, configured to input the source image and the target image of the endoscope into a feature extraction network with shared weights to obtain multi-channel feature descriptor graphs of the source image and the target image;

[0120] A source feature point processing module 200, configured to obtain the coordinate of the feature point of the source image based on a feature point detection operator, and sample a source feature point descriptor vector in the source feature descriptor graph;

[0121] The rough matching module 300 is used to calculate the correlation between the source feature point descriptor vector and the target feature descriptor map, generate a heat map of response values on the target image, and extract the coordinates corresponding to the maximum response value in the heat map as the rough matching result;

[0122] The false matching filtering and matrix estimation module 400 is used to filter false matches in the rough matching result through the random sample consensus algorithm and estimate the camera pose transformation matrix based on the remaining matching point pairs;

[0123] The window adjustment module 500 is used to adaptively adjust the local sampling window sizes of the source image and the target image according to the depth ratio relationship of the matching point pairs;

[0124] The fine matching optimization module 600 is used to perform feature descriptor block sampling on the local regions of the source image and the target image respectively based on the adjusted window sizes, and use the self-attention mechanism and the cross-attention mechanism for fine matching optimization;

[0125] The coordinate fine-tuning and output module 700 is used to fine-tune the target image coordinates according to the matching probability map in the fine matching stage and output the finally optimized feature point matching result.

[0126] An image feature point matching system for an endoscopic surgery inspection scenario in this embodiment is used to implement the aforementioned image feature point matching method for an endoscopic surgery inspection scenario. Therefore, the specific implementation manners in the image feature point matching system for an endoscopic surgery inspection scenario can be seen in the embodiment part of the aforementioned image feature point matching method for an endoscopic surgery inspection scenario. For example, the feature extraction module 100, the source feature point processing module 200, the rough matching module 300, the false matching filtering and matrix estimation module 400, the window adjustment module 500, the fine matching optimization module 600, and the coordinate fine-tuning and output module 700 are respectively used to implement steps S1, S2, S3, S4, S5, S6, and S7 in the aforementioned image feature point matching method for an endoscopic surgery inspection scenario. Therefore, the specific implementation manners can refer to the descriptions of the corresponding various part embodiments. To avoid redundancy, they will not be elaborated here.

[0127] Embodiment III

[0128] An embodiment of the present invention provides an electronic device. The electronic device includes a processor, a memory, and a bus system. The processor and the memory are connected through the bus system. The memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the aforementioned image feature point matching method for an endoscopic surgery inspection scenario.

[0129] Embodiment IV

[0130] An embodiment of the present invention provides a computer storage medium, which stores a computer software product. The computer software product includes a number of instructions for causing a computer device to execute the above-mentioned image feature point matching method for an endoscopic surgical examination scenario.

[0131] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0132] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0133] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0134] Obviously, the above embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.

Claims

1. An image feature point matching method for endoscopic surgery inspection scenes, characterized in that: include: S1: Input the source image and target image of the endoscope into the feature extraction network with shared weights to obtain the multi-channel feature descriptor graphs of the source image and the target image; S2: Obtain the feature point coordinates of the source image based on the feature point detection operator, and sample the source feature point descriptor vector in the source feature descriptor map; S3: Calculate the correlation between the source feature point descriptor vector and the target feature descriptor map, generate a response value heat map on the target image, and extract the coordinates corresponding to the maximum response value in the heat map as the rough matching result; S4: Filter out the mismatches in the rough matching results through the random sampling consensus algorithm, and estimate the camera pose transformation matrix based on the remaining matching point pairs; S5: Adaptively adjust the local sampling window size of the source image and the target image according to the depth ratio relationship of the matching point pairs; S6: Based on the adjusted window size, feature descriptor blocks are sampled in local areas of the source image and the target image respectively, and self-attention mechanism and cross-attention mechanism are used for fine matching optimization; S7: Fine-tune the target image coordinates according to the matching probability map in the fine matching stage, and output the final optimized feature point matching results.

2. The image feature point matching method for endoscopic surgery inspection scenes according to claim 1 is characterized in that: In step S1, the feature extraction network is a dual-head output feature extraction network, wherein one output head is used to generate a deblurred feature map, and the other output head is used to generate a multi-channel feature descriptor map with the same size as the input image.

3. The image feature point matching method for endoscopic surgery inspection scenes according to claim 1 is characterized in that: In step S5, the formula for adaptively adjusting the window size is: in, is the initial window size on the source image, s i is the scaling factor calculated based on the matching point to depth ratio, is the adjusted window size on the destination image.

4. The image feature point matching method for endoscopic surgery inspection scenes according to claim 1 is characterized in that: In step S6, the fine matching optimization includes: Enhance the correlation of feature descriptors in local regions through self-attention mechanism; The correspondence between the source and target image feature descriptor blocks is captured through a cross-attention mechanism.

5. The image feature point matching method for endoscopic surgery inspection scenes according to claim 1 is characterized in that: Before using the self-attention mechanism and the cross-attention mechanism for fine-tuning the matching, the target feature descriptor block is downsampled to the size of the source feature descriptor block.

6. The image feature point matching method for endoscopic surgery inspection scenes according to claim 1 is characterized in that: In step S7, the expectation is calculated according to the matching probability map in the fine matching stage, and the target image coordinates are fine-tuned according to the expected result, wherein the calculation formula of the expectation is: in, represents the target image coordinates of the i-th feature point after fine-tuning; P i Represents the initial coordinates of the i-th feature point; is the adjusted window size on the target image, u, v are pixel coordinate variables; M i (u,v) represents the matching probability map; Represents the sum of the matching probabilities of all pixel positions within the window.

7. The image feature point matching method for endoscopic surgery inspection scenes according to claim 1 is characterized in that: The method further includes designing a multi-stage loss function consisting of a coarse matching stage loss, a fine matching stage loss, and a deblurring module loss for training the feature extraction network, wherein the multi-stage loss function is: in In the formula, represents a multi-stage loss function; λ1, λ2, λ3 represent adjustable hyperparameters; represents the loss in the rough matching stage; represents the thermal response map produced in the coarse matching stage; represents the thermal response value at position P on the target image; α is the scale factor; P gt Represents the true matching coordinates on the target image; represents the thermal response value at the true matching coordinate; P i Thermal response diagram The i-th position in ; represents the loss of the fine matching stage; || ||2 represents the L2 norm; P i,gt Represents the true coordinates of the i-th feature point in the target image; P′ i Represents the predicted coordinates of the i-th feature point obtained after the fine matching stage; represents the deblurring module loss; is the mean square error between the output image of the feature extraction network and the label image; is the adversarial loss; G is the generator in the feature extraction head; D is the discriminator; represents mathematical expectation; A~p actual (A) indicates that variable A follows the probability distribution p of the real sample actual (A), A represents an image sample randomly selected from a set of real clear images; B~p blurry (B) indicates that variable B follows the probability distribution p of the blurred image. blurry (B), B represents a sample randomly selected from the blurred image set.

8. An image feature point matching system for endoscopic surgery inspection scenes, characterized in that: The system is used to implement the image feature point matching method for endoscopic surgery inspection scenes according to any one of claims 1 to 7, specifically comprising: A feature extraction module is used to input the source image and the target image of the endoscope into a feature extraction network with shared weights to obtain a multi-channel feature descriptor map of the source image and the target image; A source feature point processing module is used to obtain the feature point coordinates of the source image based on the feature point detection operator, and to obtain the source feature point descriptor vector by sampling in the source feature descriptor map; A coarse matching module is used to calculate the correlation between the source feature point descriptor vector and the target feature descriptor map, generate a response value heat map on the target image, and extract the coordinates corresponding to the maximum response value in the heat map as the coarse matching result; The mismatch filtering and matrix estimation module is used to filter mismatches in the rough matching results through a random sampling consensus algorithm, and estimate the camera pose transformation matrix based on the remaining matching point pairs; A window adjustment module is used to adaptively adjust the local sampling window size of the source image and the target image according to the depth ratio relationship of the matching point pairs; A fine matching optimization module is used to sample feature descriptor blocks in local areas of the source image and the target image based on the adjusted window size, and perform fine matching optimization using the self-attention mechanism and the cross-attention mechanism; The coordinate fine-tuning and output module is used to fine-tune the target image coordinates according to the matching probability map in the fine matching stage and output the final optimized feature point matching results.

9. An electronic device, characterized in that: The electronic device includes a processor, a memory and a bus system, the processor and the memory are connected through the bus system, the memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the image feature point matching method for endoscopic surgical examination scenes as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that: The computer storage medium stores a computer software product, which includes a number of instructions for enabling a computer device to execute the image feature point matching method for endoscopic surgery inspection scenes as described in any one of claims 1 to 7.