Endoluminal image blood vessel branch point extraction method based on combination of mixed attention mechanism and U-Net

Through the hybrid attention mechanism combined with the U-Net network, the problem of insufficient robustness and efficiency of vascular branch point feature extraction in the lumen environment is solved, stable and accurate vascular branch point detection is achieved, and the accuracy and efficiency of the SLAM system are improved.

CN120340075APending Publication Date: 2025-07-18HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510386150.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional methods have problems with insufficient robustness and efficiency in the feature extraction of vascular branch points in the lumen environment, especially the lack of feature points caused by the loss of texture features in the endoscopic field of view, which affects the accuracy and stability of the SLAM system, and the difficulty in labeling gold standard data leads to difficulty in model training.

Method used

The hybrid attention mechanism is used to combine the U-Net network, and the Det network and the Des network are constructed through a shared encoder, combined with multi-scale feature extraction and upsampling operations, Dice-Det loss function and Triplet-Des loss function based on branch feature weights are designed, transfer learning and semi-supervised training are introduced, and pseudo-gold standard data sets are used for training, enhancing the accuracy and robustness of vascular branch point detection.

Benefits of technology

It realizes the stable and accurate extraction of blood vessel branch points in the lumen environment, improves the robustness and generalization ability of feature extraction, and a feature extraction method suitable for the lumen environment, which improves the accuracy and efficiency of the SLAM system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_4
    Figure SMS_4
  • Figure SMS_5
    Figure SMS_5
  • Figure SMS_6
    Figure SMS_6
Patent Text Reader

Abstract

The invention discloses a lumen image blood vessel branch point extraction method combining a mixed attention mechanism with U-Net, and belongs to the technical field of medical digital image processing. The method sequentially comprises the following steps: sharing an encoder in a U-Net network architecture to construct a Det network and a Des network; the method comprises the following steps: performing pre-training on a U-Net network by adopting a SuperRetina data set to obtain a pre-trained network model; constructing an MIS image pseudo-gold standard branch feature data set for training by adopting a transfer learning method and utilizing the marked public data set Vivo and the unmarked Hamlyn center endoscope data set; designing a dynamic up-sampling module based on a mixed attention mechanism, introducing a network architecture to carry out bilinear up-sampling, and dynamically adjusting a key area of up-sampling; a Dice-Det loss function based on a branch feature weight and a Triplet-Des loss function based on a branch point and dynamic weight weighting are designed, so that the model can pay more attention to a blood vessel branch region, and key branch features can be learned more quickly. According to the method, end-to-end combined detection of the blood vessel branch points and the descriptors can be stably and accurately achieved, and good robustness and generalization are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The method for extracting vascular branch points in lumen images by combining a hybrid attention mechanism with U-Net of the present invention belongs to the technical field of medical digital image processing. Background Technique

[0002] Through years of research, it has been found that the improvement of traditional feature extraction methods cannot parallelly ensure the improvement of accuracy and speed. In the Simultaneous Localization and Mapping (SLAM) system with a running speed of 30 milliseconds per frame, the ORB algorithm occupies two-thirds of the computing time. The feature-based monocular Visual Odometry (VO) needs to extract significant and repeatable features, and it requires robust feature tracking of image frames. The lumen environment has special textures and lacks stable features. Under the endoscopic view, there are hundreds of thousands of pixels, and the traditional method extracts few feature points, which is not enough for robust calculation of camera movement. With the development of deep learning, more and more researchers have begun to try to use deep learning technology to solve the problem of feature extraction.

[0003] Currently, mainstream methods such as SuperPoint lack attention to lumen texture features. Although MagicPoint shows surprisingly excellent performance, there are still problems with difficult domain adaptation. As a result, after such deep learning methods are introduced into the SLAM system, there are problems such as a significant decrease in tracking accuracy or even inability to work in the application of the lumen environment. Gold standard data is crucial for the training of prediction networks and algorithm evaluation. Since such supervised algorithms require a large amount of gold standard features for training, the gold standard branch features of MIS images are complex, and large-scale manual annotation is difficult and time-consuming. Considering that transfer learning can transfer the trained model parameters to a new model to help the new model train, thereby accelerating and optimizing the learning efficiency of the model and improving the generalization ability of the model.

[0004] Research is carried out on the texture features - vascular branch points in the human lumen environment, and the problem of pseudo-gold standard vascular branch features is solved based on the theoretical basis and research results of the method for extracting and matching vascular feature points with enhanced vascular structure by double-branch weighted fusion. Deep learning can simultaneously meet the improvement of accuracy and efficiency. Summary of the Invention

[0005] In view of the above technical background, the applicant has designed a method for extracting vascular branch points in lumen images by combining a hybrid attention mechanism with U-Net. The object of the present invention is achieved as follows:

[0006] 1. A method for extracting vascular branch points in lumen images by combining a hybrid attention mechanism with U-Net, characterized by comprising the following steps:

[0007] Step a: Construct a lumen image vascular bifurcation point extraction network framework that combines a hybrid attention mechanism with U-Net, and extract and describe the vascular bifurcation points in the lumen image in a single forward pass by combining multi-scale feature extraction and upsampling operations. The steps are as follows:

[0008] Step a1: Share the encoder in the U-Net network architecture to construct the Det network and the Des network, which are divided into three modules: the shared encoder, the Det decoder, and the Des decoder.

[0009] Step a2: Input the lumen image I, extract and describe the vascular bifurcation points by combining multi-scale feature extraction, and generate a probability map P and a descriptor tensor D.

[0010] Step a3: To address the training challenges, use a semi-supervised training algorithm and train with a small number of pseudo gold standard labeled images. At the same time, introduce the semi-supervised learning strategy PKE of progressive expansion to automatically detect highly reliable vascular bifurcation points in retinal images.

[0011] Step a4: Introduce a double matching strategy to improve the accuracy of bifurcation point detection.

[0012] Step b: Construct a vascular bifurcation feature pseudo gold standard and dataset based on the transfer learning strategy. The steps are as follows:

[0013] Step b1: Use the SuperRetina dataset to pre-train the U-Net network to obtain the SuperRetina pre-trained model. Initialize the network model weights with parameters and perform global fine-tuning.

[0014] Step b2: Pay attention to the source domain retinal dataset D R and the learning task T R , the target domain lumen dataset D V and the learning task T V , and obtain the knowledge in D R and T R through training to improve the prediction function f V (·).

[0015] Step b3: Use five public datasets Vivo with 20 manually marked ones and the unlabeled Hamlyn Center endoscopic dataset, denoted as dataset A and dataset B respectively. A is labeled as the test set, and images with prominent blood vessels and no obvious occlusion are selected from B to construct a MIS image pseudo gold standard bifurcation feature dataset for training.

[0016] Step b4: Perform data augmentation by applying random affine transformations (such as rotation, translation, scaling, etc.) to the images.

[0017] Step c. Design a dynamic upsampling module based on a hybrid attention mechanism and introduce it into the network framework in step a. DySample combines CBAM to replace the traditional upsampling module, so as to enhance the feature expression ability and efficiency of the lumen vascular branch point detection network.

[0018] Step c1. The input feature map is processed by a sampling point generator to generate a fixed offset, initialize the offset position, and introduce a dynamic range factor for further optimization to generate a sampling set δ. The feature map X is upsampled through the grid_sample function and the positions in the sampling set δ to obtain the upsampled feature map x'.

[0019] Step c2. For the feature map X' in step c1, first perform weighting through a channel attention module to generate a channel attention feature map X".

[0020] Step c3. Weight the feature map X" in step c2 through a spatial attention module to generate a spatial attention feature map Y.

[0021] Step d. Design a Dice-Det loss function based on branch feature weights and introduce it into the network framework in step a to make the network pay more attention to the vascular branch area.

[0022] Step d1. Introduce a weighted term weight tensor weight_map, which is designed according to the predicted branch point positions. According to the probability map P output in step a2, corresponding weights are assigned to the probabilities of each pixel, and higher weights are assigned to the pixels with higher probabilities at the predicted branch points.

[0023] Step d2. The weight map in step d1 is passed to the loss function as a weighting coefficient. By weighting the branch point area, the pixel points around the branch points have higher weights, so that the model pays more attention to the branch point area during training.

[0024] Step d3. For N samples in a batch, calculate the loss value of each sample respectively, and then add and average these loss values to obtain the overall average loss value, which is used to measure the overall performance of the model on this batch and avoid the excessive influence of the extreme loss value of a single sample on the training process.

[0025] Step d4. To ensure the overall consistency between P I and its geometric transformation P I’ calculate the Dice-Det loss between P and P H which is called

[0026] Step e. Design a Triplet-Des loss function based on branch point and dynamic weight weighting and introduce it into the network framework in step a to enable the model to learn key branch features faster.

[0027] Step e1: The Det decoder in step a1 automatically provides branch points, enabling the Dec decoder to be trained in a fully self-supervised manner. To prevent the generation of trivial solutions, the Triplet Loss is adopted to ensure that the descriptor decoder can effectively learn from unlabeled data.

[0028] Step e2: Define the Anchor as the current blood vessel branch point, the Positive as other points in the similar or the same blood vessel, and the Negative as different branch points or non-branch points. Use the cosine distance as the distance metric function.

[0029] Step e3: Introduce the hard negative mining strategy. Define the distances of two types of negative samples, one is the randomly selected negative sample, and the other is the most difficult-to-pair negative sample, and take their mean as the negative term.

[0030] Step e4: Introduce weight dynamic adjustment to the contribution of each triplet, enabling the model to learn key branch features faster.

[0031] Step e5: To further optimize the feature space structure, encourage the orthogonality between negative samples, prevent descriptors from being too similar, and introduce the GOR regularization term. Combine the Triplet-Des loss function and the GOR regularization term to form the final Des loss function, and combine it with the branch point decoder loss function for joint training.

[0032] Beneficial effects:

[0033] The method of the present invention constructs a Det network and a Des network by sharing an encoder in the U-Net network architecture; pre-trains the U-Net network using the SuperRetina dataset to obtain a pre-trained network model, which helps the model obtain better initial weights and improve the detection and description performance of the current task; uses the transfer learning method to construct a pseudo gold standard branch feature dataset of MIS images for training using the labeled public dataset Vivo and the unlabeled Hamlyn Center endoscopic dataset, so that the transfer learning model can extract more representative features in a richer data background; designs a dynamic upsampling module based on a hybrid attention mechanism to introduce into the network architecture for bilinear upsampling, dynamically adjusts the key areas of upsampling, more effectively utilizes spatial and channel features, reduces redundant information and noise, and enhances the feature expression ability and efficiency of the lumen vascular branch point detection network; designs a Dice-Det loss function based on branch feature weights and a Triplet-Des loss function based on branch points and dynamic weight weighting, so that the model can pay more attention to the vascular branch area and learn key branch features faster. On open-source datasets with high authority in the field of lumen image processing such as the Hamlyn dataset and the Vivo dataset, through quantitative analysis of multiple indicators, it is proved that the method of the present invention can accurately extract a large number of vascular branch feature points, has robustness and generalization, and is a feature extraction method suitable for the lumen environment. Description of the Drawings

[0034] Figure 1 It is a flowchart of the lumen image vascular branch point extraction scheme combining the hybrid attention mechanism in the method of the present invention with U-Net.

[0035] Figure 2 It is a network framework diagram of the lumen image vascular branch point extraction combining the hybrid attention mechanism in the method of the present invention with U-Net.

[0036] Figure 3 It is an example diagram of the dataset construction in the method of the present invention.

[0037] Figure 4 It is a comparison diagram of the training loss curves of different module combinations with the transfer learning strategy in the method of the present invention.

[0038] Figure 5 It is a comparison diagram of the training loss curves of each experimental scheme in the method of the present invention.

[0039] Figure 6 It is a visualization comparison diagram of the matching results of the clip5 dataset in the method of the present invention. Detailed Implementation Modes

[0040] The following describes the implementation manners of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0041] The following further describes in detail the specific implementation manners of the present invention with reference to the accompanying drawings.

[0042] The method flow of this specific implementation manner is as Figure 1 shown and includes the following steps:

[0043] Step a, construct a lumen image vascular branch point extraction network framework combining a hybrid attention mechanism and U-Net, as Figure 2 shown, and extract and describe the vascular branch points in the lumen image in a single forward pass by combining multi-scale feature extraction and upsampling operations; including the following steps:

[0044] Step a1, share an encoder in the U-Net network architecture to construct a Det network and a Des network, which are divided into three modules: a shared encoder, a Det decoder, and a Des decoder.

[0045] Step a2, the input is a lumen image I of H×W. Combine multi-scale feature extraction and upsampling modules to extract and describe vascular branch points, generate a probability map P and a descriptor tensor D. The network model parameters are shown in Table 1. In the self-supervised learning framework, the Des network shares the encoder of the Det network and generates a d-dimensional descriptor for each pixel. This dense output result is achieved through interpolation techniques. Apply NMS to P to obtain a binary mask as the final detection result, so as to remove duplicate detection results while retaining key points.

[0046] Table 1 Network model parameters

[0047]

[0048]

[0049] Step a3, to address the training challenges, use a semi-supervised training algorithm and train with a small number of pseudo gold standard labeled images. At the same time, introduce the Progressive Key Point Expansion (PKE) method to solve the problem of incomplete pseudo gold standards, as Figure 3As shown below. First, construct the geometric mapping I' = H(I) of image I, where H is a specific homography transformation. Input I' into the network to obtain its probability prediction map P', and obtain the branch point map P of I through back-projection. H = H -1 (P'). After performing NMS on P obtained for I, for each point (x, y) in P N , if (P H ) x,y > 0.5, then add it to Det t (such as the green branch points). Expand the labeled set B by gradually adding new branch points. During training, use the union B t of the initial branch point set B0 (blue branch points) and the branch point set Det t detected in the t-th round for this round of training. Given that in the initial stage of training, Det t may contain incorrect branch points, and an excellent detector should be able to detect the same branch points at different perspectives and scales. To improve the accuracy of branch point detection, perform geometric matching to identify branch points that can be repeatedly detected from a given image and its projective transformation.

[0050] Step a4: Construct a dual matching strategy by introducing a content-based matching mechanism. For each point (x, y) in Det t , obtain its corresponding descriptor by directly sampling the output of the descriptor decoder, thus obtaining the descriptor set Des t . Correspondingly, extract Des t from I' according to H(Det t ). For each descriptor in Des t , perform a nearest neighbor search on Des t ' using it as a query item. Only when the spatial corresponding point (x', y') of the point (x, y) passes the ratio test, the point (x, y) will be retained in Det t . Suppose a branch point A in the non-vascular area is detected in I, and a branch point C is geometrically matched in I' after a specific homography transformation. The non-vascular area lacks specificity in visual appearance, meaning that the descriptors extracted from these areas are relatively close. Therefore, even if C is the best match for A in the descriptor space, its second-best match is not sufficient to pass Lowe's ratio test.

[0051] Step b: Construct a vascular branch feature pseudo-ground truth and dataset based on the transfer learning strategy; including the following steps:

[0052] Step b1: Pre-train the U-Net network using the SuperRetina dataset to obtain the SuperRetina pre-trained model. Initialize the weights of the network model with the pre-trained model parameters of the retinal image dataset and perform global fine-tuning.

[0053] Step b2: Focus on the source domain retinal dataset D R and the learning task T R 、the target domain lumen dataset D V and the learning task T V ,and improve the prediction function f R (·) by learning the knowledge in D R and T V . During the training process, the parameters in the model are denoted as Det R and Des R before migration, and the parameters in the model after migrating to the target domain are denoted as Det V and Des V . The overall process of model migration is expressed as:

[0054]

[0055] The key advantage of transfer learning is that by leveraging the feature extraction capabilities already obtained in the SuperRetina pre-trained model, it helps the model obtain better initial weights, improve the detection and description performance of the current task, achieve efficient and high-quality key point detection, and avoid the time and data overhead of training the model from scratch.

[0056] Step b3: Use five public datasets with 20 manually marked Vivo and unannotated Hamlyn Center endoscopic / endoscopic video datasets, denoted as dataset A and dataset B respectively, to construct a MIS image pseudo-gold standard branch feature dataset for training. Comprehensively analyze the image characteristics of the two public datasets. As shown in Table 2, A is used as the test set with annotations. For dataset B, select images with prominent blood vessels and no obvious occlusion. The available annotated datasets are Rectified01, Rectified06, Rectified08, and Rectified09. Among them, for Rectified08 and Rectified09, the images containing surgical instruments and jittery scenes are removed as duplicate scenes of the three scenarios. Therefore, a small number of images are selected for each scenario and included in the dataset. 6 and 14 images with prominent blood vessels and no obvious occlusion are selected respectively. Since the dataset is made from continuously shot videos, to prevent overfitting during training due to too many similar scene images, for Rectified01 and Rectified06, one image is selected every 50 frames and included in the dataset. The training set has a total of 140 images, the validation set has a total of 66 images, the unannotated dataset has a total of 933 images, and the test set has a total of 242 images. The ratio of the training set to the validation set is set to 7:3.

[0057] Table 2 Details of the lumen dataset

[0058]

[0059] Step b4: Perform data augmentation by applying random affine transformations (such as rotation, translation, scaling, etc.) to the images. As Figure 3 shown, expand the training dataset so that the model can see images with different deformations. Let the model learn these changes during training to improve its adaptability to unknown data.

[0060] The mathematical description of the affine transformation is as follows. The affine transformation maps a point (x, y) in a two-dimensional space to a new point (x', y') through a 2×3 matrix. The expression is:

[0061]

[0062] where A, B, C, D: define rotation, scaling, and shear operations, t x , t y : define translation operations. t x >0: move to the right. t y >0: move down. Rotating the angle θ around the center point (c x , c y ) of the image, the expression is:

[0063]

[0064] In proportion s x , s y The expression for enlarging or reducing an image is:

[0065]

[0066] s x > 1: Horizontal enlargement. s y > 1: Vertical enlargement. 0 < s x , s y < 1: Reducing the image. Stretching or compressing the image along the x-axis or y-axis, the expression is:

[0067]

[0068] Step c, Design a dynamic upsampling module based on a hybrid attention mechanism and introduce the network framework in step a. DySample combines CBAM to replace the traditional upsampling module to enhance the feature expression ability and efficiency of the lumen blood vessel branch point detection network.

[0069] Step c1, The input feature map is processed by a sampling point generator to generate a fixed offset, initialize the offset position, and introduce a dynamic range factor for further optimization to generate a sampling set δ. Different from the traditional convolution kernel-based dynamic upsampling method, DySample is designed from the perspective of point sampling, splitting a point into multiple points to achieve a clearer edge. Input feature map After being processed by the sampling point generator, a sampling set δ of size 2 × H2 × W2 is generated. Specifically, first, a fixed offset Offset is generated through the combination of a linear layer and pixel shuffling technology Linear, and the expression is:

[0070]

[0071] Among them, the output channel is 2gs 2 The linear layer, and the offset contains offsets in the horizontal and vertical directions. Then, the position of the offset is initialized InitPos, and the expression is:

[0072]

[0073] where s is the upsampling ratio factor. On the static basis, the dynamic range factor is further optimized by introducing variable parameters in the linear layer and pixel shuffling technology. First, a dynamic range factor 0.25σ is generated using the Sigmoid function to adjust the offset Scope, and the expression is:

[0074] Scope = W Scope · X + b Scope

[0075] Then, the pixel shuffle method is used to reshape it into an offset of 2g×sh×sw. The OffsetFinal expression is:

[0076] OffsetFinal = σ(Scope)·Offset + InitPos

[0077] Normalize the sum of the offset OffsetFinal and the original sampling grid G to obtain a sampling set δ of 2g×sh×sw. The expression is:

[0078]

[0079] where N = [W, H] is the normalization factor, and 2g represents the x and y coordinates. Upsample the feature map X through the grid_sample function and the positions in the sampling set δ to obtain the upsampled feature map X'. The expression is:

[0080] X' = grid_sample(X, δ)

[0081] where,

[0082] Step c2: For the feature map X' in step c1, first perform weighting through the channel attention module to generate the channel attention feature map X''.

[0083] Step c3: Perform weighting on the feature map x'' in step c2 through the spatial attention module to generate the spatial attention feature map Y.

[0084] Step c2: For the feature map X' in step c1, first perform weighting through the channel attention module to generate the channel attention feature map X''. To effectively calculate the channel attention, first compress the spatial dimension of the input feature map, and at the same time use average pooling and max pooling operations to capture different spatial context descriptors, which are fed into a shared multi-layer perceptron (MLP) to generate the channel attention map. The expression is:

[0085]

[0086] The expressions for global average pooling and max pooling are respectively:

[0087]

[0088] Then the pooling result expression is:

[0089]

[0090] The pooling result is implemented through two layers of convolution:

[0091] MLP(X′) = W1·ReLU(W0·X′ + b1) + b2, X′ ∈ {AvgOut, MaxOut}

[0092] Sum the pooling results and apply Sigmoid activation:

[0093]

[0094] Finally, apply the channel attention weight to the feature map:

[0095]

[0096] Step c3: Weight the feature map X″ from step c2 through a spatial attention module to generate a spatial attention feature map X″, obtaining the final feature map Y. To calculate the spatial attention, first apply average pooling and max pooling operations along the channel axis, and then concatenate them to generate an efficient feature descriptor. Apply a convolutional layer on this descriptor to generate the spatial attention map. Its expression is:

[0097]

[0098] Perform average and max pooling on X″ along the channel dimension:

[0099]

[0100] Then the pooling results are:

[0101]

[0102] Concatenate the pooling results and generate spatial attention weights through convolution:

[0103]

[0104] Finally, apply the spatial attention weight to the feature map:

[0105]

[0106] Step d: Design a Dice-Det loss function based on branch feature weights and introduce it into the network framework in step a to make the network pay more attention to the vascular branch region.

[0107] Step d1: Introduce a weighted term weight tensor weight_map, which is designed according to the predicted branch point positions. Based on the probability map output in step a2, assign corresponding weights to the probabilities of each pixel, with higher weights assigned to the predicted branch points and lower weights assigned to the lower probabilities. The Dice loss function estimates the loss value by normalizing the overlap degree between the predicted value and the ground truth. Its expression is:

[0108] L Dice = 1 - Dice

[0109] The Dice coefficient ranges from [0, 1]. It is 1 when two sets completely overlap and 0 when they do not overlap at all. The smaller the loss value, the higher the degree of overlap. The expression is:

[0110]

[0111] In the detection of vascular branch points, the goal is usually to identify small vascular branches, endpoints, and other tiny structures, while the standard Dice Loss is mainly used to process larger structures (such as an entire organ or region). The expression is:

[0112]

[0113] where g i and p i are the ground truth value and the predicted value of the i-th pixel respectively.

[0114] In step d2, the weight maps of step d1 are passed to the loss function as weighting coefficients. By weighting the branch point area, the pixels around the branch point have higher weights, so that the model pays more attention to the branch point area during training. The probability map (Predicted Probability Map) generated by the feature point detection network represents the probability that each pixel is a feature point, and the probability value is between [0, 1]. Among them, it contains the learned ground truth points, that is, the known positions of vascular branch points. According to the probability map output by the model, the probability of each pixel is given a corresponding weight, and the pixel color (white or red) at the predicted branch point is significantly different from the pixels in other areas (black).

[0115] If the probability of a certain position being a branch point is high (close to 1), a higher weight can be given. If the probability is low, a smaller weight is given. In this way, during the training process, the model can pay more attention to the positions of real feature points and give corresponding weights according to the predicted probability, so that the model can pay appropriate attention to difficult areas (such as edges and low-probability areas). The weight map is passed to the loss function as a weighting coefficient. By weighting the branch point area, the pixels around the branch point have higher weights, so that the model pays more attention to the branch point area during training, making the influence of these small areas on the loss function greater. The loss function expression is:

[0116]

[0117] where w iis the i-th value of the weight map. To avoid the denominator being zero, a smoothing factor ∈ is introduced, which usually takes a small value, such as 10 -3 .

[0118] Since network training usually processes data in batches, the loss function expression for N samples i is as follows:

[0119]

[0120] where g ij is the true value of the j-th pixel in the i-th sample, p ij is the predicted value of the j-th pixel in the i-th sample, and w ij is the weight value corresponding to the weight map. If the weight map is None, it is defaulted to 1.

[0121] Step d3: For the N samples in the batch, calculate the loss value of each sample separately, then add these loss values and take the average to obtain the overall average loss value, which is used to measure the overall performance of the model on this batch and avoid the excessive influence of extreme loss values of individual samples on the training process. To maintain stable gradient updates during model training, the losses of all samples are averaged for loss reduction. Averaging the loss values helps avoid gradient instability caused by differences in batch size or extremely large or small losses of individual samples. The expression is:

[0122]

[0123] That is,

[0124]

[0125] Step d4: To ensure the overall consistency between P I and its geometric transformation P I’ , calculate the Dice-Det loss between P and P H , which is called Dice-Det loss based on B t and H The expression is:

[0126]

[0127] Step e: Design a Triplet-Des loss function based on branch points and dynamic weight weighting and introduce it into the network framework in step a, enabling the model to learn key branch features faster.

[0128] Step e1: The Det decoder in step a1 automatically provides branch points, so the Dec decoder is trained in a fully self-supervised manner. At the same time, to avoid the trivial solution of generating constant descriptors, the Triplet Loss, which is often used in self-supervised learning, is selected, enabling the descriptor decoder to easily learn from unlabeled data.

[0129] Step e2: Define Anchor as the current blood vessel branch point, Positive as other points in the similar or the same blood vessel, and Negative as different branch points or non-branch points. Use the cosine distance as the distance metric function. The Triplet Loss optimizes the feature representation through the triplet (Anchor, Positive, Negative), especially in diverse blood vessel branch point data. This ensures that the features between the same category (positive samples) are more similar, and the features between different categories (negative samples) are more different. For the blood vessel branch points, the triplets are designed as follows: Anchor: the current blood vessel branch point. Positive: other points similar to or in the same blood vessel as this branch point. Negative: branch points or non-branch points different from this branch point. The expression is:

[0130] L Triplet = max(d(Anchor, Positive) - d(Anchor, Negative) + margin, 0)

[0131] where d(x, y) is the distance between x and y, and margin is a hyperparameter. It has a good optimization effect on the feature embedding space, can clearly push similar blood vessel branch points closer and dissimilar points farther away, better handle the spatial and topological characteristics of blood vessel network data, and can obtain a powerful feature representation in the case of self-supervised learning.

[0132] Ideally, for each key point (x, y) detected in I, the descriptor decoder outputs the same at the corresponding position relative to I'. The Triplet Loss makes the distance between paired key points less than the distance between unpaired key points. Input I and I' into the network respectively to obtain the full-size descriptor tensors D and D'. For each element (x, y) in P N let D x,y be its descriptor. Since (x, y) and (x', y') should be paired, the distance d x,y between their descriptors must be reduced. The distance metric function d(x1, x2) uses the cosine distance, which is transformed into a distance based on similarity. The expression is:

[0133]

[0134] Step e3: Introduce the hard negative mining strategy. Define the distances of two types of negative samples. One is the randomly selected negative sample, and the other is the most difficult-to-match negative sample. Take their mean as the negative term. After excluding (x', y'), since directly using all negative samples may lead to low optimization efficiency, the hard negative mining strategy is introduced. For the two types of negative samples, d rand represents the descriptor distance between (x, y) and a randomly selected point in H(P N ), and d hard represents the descriptor distance between (x, y) and the most difficult-to-match point in H(P N ). The expression is:

[0135]

[0136] The requirement is relatively easy to meet. Using alone cannot obtain descriptors with good discrimination. At the same time, since the network lacks the ability to generate good descriptors in the early training stage, using alone will make the network difficult to train. Therefore, using only d rand or d hard as the negative term in the triplet loss is problematic. To solve this problem, the mean of and is used as the negative term. The descriptor loss L Des expression is:

[0137]

[0138] where m > 0 is the predefined positive and negative sample spacing.

[0139] Step e4: Introduce weights to dynamically adjust the contribution of each triplet, enabling the model to learn key branch features faster. The loss function expression is:

[0140]

[0141] where w i is the weight of each triplet, dynamically adjusted based on the current loss size.

[0142] Step e5: To further optimize the feature space structure, encourage the orthogonality between negative samples, and prevent descriptors from being too similar, introduce the GOR regularization term, the expression is:

[0143] gor rand =E[((x,y)·H(P N )) rand )) 2

[0144] gor​hard = E[((x, y) · H(P N )) hard )) 2

[0145] Minimize the square of the dot product between the negative samples and the paired points, and the GOR regularization loss expression is obtained as follows:

[0146] L gor = β · (gor rand + gor hard )

[0147] where β is the weight coefficient, which is used to control the contribution degree of the GOR regularization term to the total loss function.

[0148] Combining the two parts of the loss, the final Des loss function expression is:

[0149]

[0150] Finally, the model is jointly trained by minimizing the combined loss function of the branch point decoder and the descriptor decoder (H changes for each mini-batch), and the combined loss function is defined as:

[0151]

[0152] The method of the present invention uses the PyTorch deep learning framework for training, and the basic configuration is Anaconda3 and PyCharm. The Adam (Adaptive Moment Estimation) optimizer is used for initialization. Before conducting the performance analysis experiment, wandb is used to optimize the hyperparameters of the model, and the optimal hyperparameter combination is determined: the initial learning rate is 1e-4, the batch size is 4, and the model training effect is the best and fully converges when the total number of iterations is 300.

[0153] To verify the influence of different module combinations with the transfer learning strategy on the model performance, ablation experiments are designed for training and testing. The tests include the model based on the transfer learning strategy, the model based on the transfer learning strategy combined with the Dysample module, and the model based on the transfer learning strategy combined with the dynamic upsampling module based on the hybrid attention mechanism. The experimental comparison of each module shows the total training loss, feature extraction loss, descriptor loss, and average number of learned features as Figure 4 shown. The Dysample module has a significant effect on improving the model convergence speed and generalization ability. On this basis, the combination of the dynamic upsampling module based on the hybrid attention mechanism significantly reduces the descriptor loss and the total loss. Comparing the average number of learned features, this design enhances the vascular feature selection ability, helps retain more effective features, effectively improves the descriptor quality, and is more suitable for the vascular feature matching task with high description ability and high precision.​

[0154] Among them, when designing the dynamic upsampling module based on the hybrid attention mechanism, ablation experiments were also conducted for training and testing to verify the effectiveness of different design schemes and their impact on the model, as Figure 5 shown. Each scheme is as follows:

[0155] 1: Both the det network and the des network introduce the dynamic upsampling module based on the hybrid attention mechanism, and the Dysample module adopts the lp mode (linear projection first, then pixel shuffling);

[0156] 2: Both the det network and the des network introduce the dynamic upsampling module based on the hybrid attention mechanism. At the same time, the Dysample module of the det network adopts the lp mode, and the Dysample module of the des network adopts the pl mode (pixel shuffling first, then linear projection);

[0157] 3: Both the det network and the des network introduce the dynamic upsampling module based on the hybrid attention mechanism, and the Dysample module adopts the pl mode;

[0158] 4: Both the det network and the des network introduce the dynamic upsampling module based on the hybrid attention mechanism. At the same time, the Dysample module of the det network adopts the pl mode, and the Dysample module of the des network adopts the lp mode;

[0159] 5: Only the det network introduces the dynamic upsampling module based on the hybrid attention mechanism;

[0160] 6: Only the des network introduces the dynamic upsampling module based on the hybrid attention mechanism.

[0161] The results show that although the performance of Scheme 1 in feature extraction loss is not obvious, it shows the best performance in descriptor loss and overall loss, with the fastest convergence speed and the lowest final loss. The trade-off according to the task requirements shows that this scheme effectively optimizes the global features and is suitable for scenarios pursuing global convergence efficiency and tasks with highly discriminative features.

[0162] To test the feature extraction and matching capabilities of the model, the method of the present invention was compared with the SuperRetina model using the test set. The feature point matching strategy in ORB-SLAM3 was adopted for the matching algorithm. First, KNN matching and ratio testing were performed, and then double geometric verification was carried out. At the same time, the homography matrix and fundamental matrix were calculated, and the optimal geometric model was selected for matching. Lowe's ratio threshold was set to 0.65, the homography RANSAC threshold was 3.0, the fundamental matrix RANSAC threshold was 1.0, and the RANSAC confidence was 0.9.

[0163] Subjectively, as Figure 6As shown, the matching effects of the base images and a random image in two models in different datasets are presented. The number of successful matches of the SuperRetina model is relatively small, and there are obvious mis-matches and incorrect feature extractions in non-content areas (such as the areas marked by the white boxes). In contrast, the method of the present invention can extract more and more stable vascular branch points. As can be seen from the comparison of the areas marked by the yellow box icons,

[0164] Objectively, as shown in Table 2, the total number of matches, the number of successful matches, and the matching accuracy of the vascular feature extractions of the two models are compared. It can be seen that the method of the present invention is superior in both quantity and accuracy. At the same time, referring to FIRE (Fundus Image Registration Evaluation), an AUC (area under the curve) evaluation dataset is constructed. According to the difficulty of matching, the five test datasets are divided into three groups, namely 99 groups of Easy, 70 groups of Mod, and 21 groups of Hard, and the average values of the three groups of results are used as the overall indicators. AUC estimates the cumulative correct rate of all samples N at different thresholds δ, quantifies the overall performance through the integral area, and reflects that the overall performance of the method of the present invention is better. The expression is: k It estimates the cumulative correct rate of all samples N at different thresholds δ, quantifies the overall performance through the integral area, and reflects that the overall performance of the method of the present invention is better. The expression is:

[0165]

[0166] Table 3 Comparison of Feature Extraction Evaluations in Different Datasets

[0167]

[0168]

Claims

1. A method for extracting vascular branch points in lumen images by combining a hybrid attention mechanism with U-Net, characterized in that, Including the following steps: Step a: Construct a lumen image vascular bifurcation point extraction network framework that combines a hybrid attention mechanism with U-Net, and extract and describe the vascular bifurcation points in the lumen image in a single forward pass by combining multi-scale feature extraction and upsampling operations. The steps are as follows: Step a1: Share the encoder in the U-Net network architecture to construct the Det network and the Des network, which are divided into three modules: the shared encoder, the Det decoder, and the Des decoder. Step a2: Input the lumen image I, extract and describe the vascular bifurcation points by combining multi-scale feature extraction, and generate the probability map P and the descriptor tensor D. Step a3: To address the training challenges, use a semi-supervised training algorithm and train with a small number of pseudo gold standard labeled images. At the same time, introduce the semi-supervised learning strategy PKE of progressive augmentation to automatically detect highly reliable vascular bifurcation points in retinal images. Step a4: Introduce a dual matching strategy to improve the accuracy of bifurcation point detection. Step b: Construct a vascular bifurcation feature pseudo gold standard and dataset based on the transfer learning strategy. The steps are as follows: Step b1: Use the SuperRetina dataset to pre-train the U-Net network to obtain the SuperRetina pre-trained model. Initialize the network model weights with parameters and perform global fine-tuning. Step b2: Pay attention to the source domain retinal dataset D R and the learning task T R , the target domain lumen dataset D V and the learning task T V , and improve the prediction function f R and T R by training with the knowledge in D V (·). Step b3: Use five public datasets Vivo with 20 manually marked and the unlabeled Hamlyn Center endoscopic dataset, denoted as dataset A and B respectively. A is labeled as the test set, and images with prominent blood vessels and no obvious occlusion are selected from B to construct a MIS image pseudo gold standard bifurcation feature dataset for training. Step b4: Perform data augmentation by applying random affine transformations (such as rotation, translation, scaling, etc.) to the images. Step c: Design a dynamic upsampling module based on the hybrid attention mechanism and introduce it into the network framework in step a. DySample combines CBAM to replace the traditional upsampling module to enhance the feature expression ability and efficiency of the lumen vascular bifurcation point detection network. Step c1: The input feature map is processed by the sampling point generator to generate a fixed offset, initialize the offset position, and further optimize it by introducing a dynamic range factor to generate the sampling set δ. The feature map X is upsampled through the grid_sample function and the positions in the sampling set δ to obtain the feature map X'. Step c2: For the feature map X' in step c1, first perform weighting through the channel attention module to generate the channel attention feature map X''. Step c3: Weight the feature map X'' in step c2 through the spatial attention module to generate the spatial attention feature map Y. Step d: Design a Dice-Det loss function based on the bifurcation feature weight and introduce it into the network framework in step a to make the network pay more attention to the vascular bifurcation area. Step d1: Introduce the weighted term weight tensor weight_map, which is designed according to the predicted bifurcation point positions. According to the probability map P output in step a2, corresponding weights are assigned to the probabilities of each pixel, with higher weights assigned to the predicted bifurcation points and smaller weights assigned to the lower probabilities. Step d2: The weight maps of step d2 and step d1 are passed to the loss function as weighting coefficients. By weighting the branch point region, the pixels around the branch point are given higher weights, so that the model pays more attention to the branch point region during training. Step d3: For the N samples in the batch, the loss value of each sample is calculated separately, and then these loss values are added and averaged to obtain the overall average loss value, which is used to measure the overall performance of the model on this batch, and at the same time avoid the excessive influence of the extreme loss value of a single sample on the training process. Step d4: To ensure the overall consistency between P I and its geometric transformation P I’ calculate the Dice-Det loss between P and P H which is called Step e: Design a Triplet-Des loss function based on branch points and dynamic weight weighting and introduce it into the network framework in step a, so that the model can learn key branch features faster. Step e1: The Det decoder in step a1 automatically provides branch points, enabling the Dec decoder to be trained in a fully self-supervised manner. To prevent the generation of trivial solutions, the Triplet Loss is adopted to ensure that the descriptor decoder can effectively learn from unlabeled data. Step e2: Define the Anchor as the current blood vessel branch point, the Positive as other points in the similar or the same blood vessel, and the Negative as different branch points or non-branch points. Use the cosine distance as the distance metric function. Step e3: Introduce the hard negative mining strategy, define the distances of two types of negative samples, one is the randomly selected negative sample, and the other is the most difficult-to-pair negative sample, and take their mean as the negative term. Step e4: Introduce weight dynamic adjustment to the contribution of each triplet, so that the model can learn key branch features faster. Step e5: To further optimize the feature space structure, encourage the orthogonality between negative samples, prevent the descriptors from being too similar, and introduce the GOR regularization term. Combine the Triplet-Des loss function and the GOR regularization term to form the final Des loss function, and combine it with the branch point decoder loss function for joint training.