Filamentous Structure Topology-Aware Fine Feature Extraction Method and System Based on Large Models
By optimizing the big model architecture and introducing discrete Morse theory, combined with the uncertainty quantization model of topological skeleton diagram, the morphological cognitive bias and topological integrity problems of image big model in filament structure recognition are solved, and the accuracy and reliability of fine feature extraction and segmentation are achieved.
Patent Information
- Application Number
- CN202510648590.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-20
AI Technical Summary
When existing image models deal with slender and complex topological filament structures, it is difficult to accurately extract features, and it is easy to misidentify them as block structures, resulting in inaccurate segmentation results. In addition, traditional methods are prone to problems such as line breakage and pseudo-branching during segmentation, which cannot effectively maintain topological integrity.
By optimizing the architecture of the basic large model, inserting the hint layer of trainable parameters and fine-tuning it in vertical domains, combining discrete Morse theory and uncertainty quantization model of topological skeleton diagrams, identifying and connecting saddle points and maximum points, building a stable manifold, and optimizing segmented images.
The fine feature extraction and segmentation of filamentous structures is realized, and the morphological cognitive bias and topological integrity problems of image large models in filamentous structure recognition is solved, and the accuracy and reliability of segmentation are improved.
Smart Images

Figure CN120182624B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method and system for fine feature extraction of filamentous structure topological perception based on a large model. Background Art
[0002] In fields such as medical imaging, remote sensing mapping, and integrated circuits, slender and topologically complex filamentous structures such as blood vessels, nerve fibers, road networks, and chip circuits are widespread. The accurate extraction of their features is crucial for subsequent analysis such as accurate segmentation. However, when traditional feature extraction methods process such targets, it is often difficult to accurately capture their branch continuity, filament distribution structure, and topological relationships. Current image feature extraction technologies mainly focus on the feature recognition of block objects or continuous regions in natural scenes, and there is relatively insufficient research on the feature extraction and segmentation of filamentous structures such as blood vessels that are slender, messy, have dense branch intersections, and complex topological structures.
[0003] The emergence of large models has brought many breakthrough developments to the field of deep learning. Due to their powerful feature extraction capabilities, excellent generalization performance, and efficient transfer learning characteristics, their applications in image feature extraction and segmentation have gradually become widespread. However, the vast majority of existing image large models are mainly oriented towards the recognition of general targets, and their feature extraction mechanisms tend to divide structures in the form of blocks or regions. When directly applied to the feature recognition of filamentous targets, image large models have obvious morphological cognitive biases. They cannot extract reasonable features, resulting in the inability to output filamentous segmentation results subsequently. Instead, they will output block regions generated when filaments intersect. That is, in the process of key feature extraction, image large models are extremely likely to regard filamentous structures as lines for segmenting block structures, rather than regarding them as the segmentation main body. This limitation severely restricts the application effect of large models in the feature extraction and segmentation of fine structures such as blood vessels and nerve fibers. In addition, most current deep learning feature extraction and segmentation methods are based on pixel-level recognition modes. When segmenting filamentous targets, problems such as line breaks and adhesions are likely to occur, and background noise interference is also likely to cause incorrect segmentation problems such as pseudo branches, which will directly lead to the damage of topological integrity and the phenomenon that the topological structure does not conform to the actual situation. Summary of the Invention
[0004] The object of the present invention is to propose a filamentous structure topology-aware fine feature extraction method and system based on a large model. First, it optimizes the original architecture of the basic large model for image feature extraction and segmentation, and retrains it on a filamentous structure image dataset using a parameter-efficient fine-tuning method, enabling it to adapt to and handle the problems of filamentous structure fine feature extraction and segmentation. Secondly, discrete Morse theory is used to analyze the output results of the large model for filamentous structure feature extraction, and a stable manifold at the topological level is constructed by finding and connecting saddle points and maximum points. Subsequently, an uncertainty quantification model for the topological skeleton graph is constructed, enabling it to analyze false positive and false negative structures in the topological skeleton. The original output content of the large model for filamentous structure feature extraction is optimized by combining the topological skeleton graph and its uncertainty graph.
[0005] The present invention is implemented through the following technical solutions.
[0006] A filamentous structure topology-aware fine feature extraction method based on a large model, comprising:
[0007] Regarding the morphological characteristics of the filamentous structure, for the original architecture of the basic large model for image feature extraction and segmentation, freeze the pre-trained weights of the basic large model and insert a prompt layer with trainable parameters, and use the training set to perform vertical domain fine-tuning on the basic large model and save the best results of the validation set to obtain a large model for filamentous structure feature extraction;
[0008] Use the large model for filamentous structure feature extraction to predict the original image of the filamentous structure to obtain an initial likelihood map with the foreground probability as the likelihood value, calibrate the initial likelihood map to obtain a calibrated likelihood map, and analyze the calibrated likelihood map through discrete Morse theory to identify critical points and construct a manifold;
[0009] Judge the survival period of each manifold through persistent homology, remove the manifolds with a survival period lower than the set threshold to obtain a stable manifold, and use the union of all stable manifolds as the topological skeleton graph of the filamentous structure;
[0010] Construct an uncertainty quantification model for the topological skeleton graph, train the uncertainty quantification model by perturbing the stable manifold data, and use the trained uncertainty quantification model to predict the confidence of each stable manifold in the topological skeleton graph to form an uncertainty graph;
[0011] Optimize the segmented image predicted by the large model for filamentous structure feature extraction by combining the topological skeleton graph and the corresponding uncertainty graph.
[0012] Further preferably, a combined method of conditional prediction value offset, adaptive temperature adjustment, and interference filtering is used to calibrate the initial likelihood map to obtain a calibrated likelihood map.
[0013] The present invention also provides a filamentous structure topology-aware fine feature extraction system based on a large model, comprising:
[0014] A large model for filamentous structure feature extraction is used to predict the original image of the filamentous structure to obtain an initial likelihood map with foreground probability as the likelihood value, and predict the original image of the filamentous structure to obtain a segmentation image;
[0015] A calibration module is used to calibrate the initial likelihood map to obtain a corrected likelihood map;
[0016] A topological awareness module is used to analyze the corrected likelihood map, identify critical points and construct a manifold;
[0017] A persistent homology module is used to construct a topological skeleton graph, judge the survival period of each manifold through persistent homology, remove the manifolds with a survival period lower than a set threshold to obtain stable manifolds, and use the union of all stable manifolds as the topological skeleton graph of the filamentous structure;
[0018] An uncertainty quantification module, with an uncertainty quantification model of the topological skeleton graph built-in, is used to predict the confidence of each stable manifold in the topological skeleton graph to form an uncertainty map.
[0019] Further preferably, the large model for filamentous structure feature extraction is based on a basic large model for image feature extraction and segmentation, removes the prompt encoder in the basic large model, and cancels the four forms of external prompts, namely prompt masks, prompt points, prompt boxes, and prompt texts. A parameter-trainable prompt layer is added after each ViT model of the image encoder in the graph basic model, and the mask decoder in the basic large model is replaced with a parameter-trainable prediction head.
[0020] Further preferably, the calibration module consists of a conditional prediction value offset link, an adaptive temperature adjustment link, and an interference filtering link connected in series; perform a conditional prediction value offset on the initial likelihood value in the initial likelihood map output by the large model for filamentous structure feature extraction, process the likelihood value obtained from the conditional prediction value offset through adaptive temperature adjustment, and the interference filtering link regularizes the likelihood value obtained from the adaptive temperature adjustment link and filters out the likelihood values whose normalized values are less than a threshold. The filtered likelihood values will be set to 0, and after processing, the corrected likelihood map is obtained.
[0021] Further preferably, the topological awareness module uses discrete Morse theory to analyze the corrected likelihood map. It regards the foreground probability in the corrected likelihood map as a terrain function and constructs a discrete gradient vector field on it; in the discrete gradient vector field, non-critical points will all flow to critical points through V-paths. Discrete Morse theory will start from the isovalue neighborhood points of saddle points and follow the discrete gradient vector field to generate V-paths. The V-paths terminate at maximum points, and then all V-paths flowing to the same maximum point are merged to form the manifold corresponding to that maximum point.
[0022] Further preferably, the uncertainty quantification module first adds a random perturbation to the distribution of the terrain function to obtain a noisy terrain function . The uncertainty quantification module uses a random walk algorithm to reconstruct a new manifold on . The construction starts from a saddle point and moves one pixel at a time until a maximum point is reached. When generating a new manifold, the uncertainty quantification module performs m random perturbations to generate a set of random perturbations , where represents the m-th random perturbation, to obtain a set of terrain functions , where represents the terrain function of the m-th random perturbation. Finally, a set of manifold variants is obtained, where
[0023] represents the manifold variant obtained from the m-th random perturbation. Subsequently, persistent homology is used to remove manifolds with a lifespan below a set threshold, leaving only stable manifolds. Then, each set of stable manifolds is constructed into a graph structure, where each node in the graph structure represents a stable manifold. When there is non-zero overlap between stable manifolds, i.e., two stable manifolds have a shared area and an intersection, an edge is constructed between the nodes. The uncertainty quantification module is trained on the constructed graph structure to learn a method for quantifying the degree of uncertainty. After training, it is used to predict an uncertainty map.
[0024] 1. By using an efficient large model parameter fine-tuning method, freezing the weights of the base large model and inserting a prompt layer with trainable parameters, while retaining the general feature extraction ability, vertical domain adaptation for filamentous structure feature recognition is achieved, solving the problems that the base large model is difficult to handle filamentous fine feature extraction tasks and the need for frequent manual formulation of prompt words, and at the same time avoiding the high cost of retraining required by traditional methods.
[0025] 2. Combining conditional prediction value offset, adaptive temperature adjustment, and interference filtering techniques to calibrate the initial likelihood map output by the filamentous structure feature extraction large model, solving the problem that it is impossible to determine any saddle point and extreme point due to the overly polarized foreground and background probabilities predicted by the filamentous structure feature extraction large model, enabling better integration of the filamentous structure feature extraction large model and the subsequent topological perception module.
[0026] 3. Introduce the topology-aware technology guided by discrete Morse theory, extract saddle points and maximum points in the corrected likelihood map and draw V paths, and then construct a stable manifold to explicitly preserve topological features such as connectivity and branching in the filamentous structure. It overcomes the defect of the traditional pixel-level feature extraction method's insufficient ability to maintain geometric morphology.
[0027] 4. Construct an uncertainty quantification model for the topological structure, train this uncertainty quantification model by perturbing the stable manifold data, effectively identify false positive and false negative structures in the topological skeleton, provide interpretable error feedback for subsequent optimization, and thus enhance the reliability of topology awareness in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is the flowchart of the method of the present invention.
[0029] Figure 2 is the schematic diagram of the system structure of the present invention.
[0030] Figure 3 is the initial likelihood value heat map output by the large model for filamentous structure feature extraction.
[0031] Figure 4 is the grayscale image of the eye retina blood vessels used for testing and marked with foreground and background coordinates as prompt words.
[0032] Figure 5 is the segmentation effect diagram of Scheme 10 in the comparative experiment.
[0033] Figure 6 is the segmentation effect diagram of Scheme 11 in the comparative experiment.
[0034] Figure 7 is the topological skeleton diagram without using the persistent homology method to remove manifolds with short durations.
[0035] Figure 8 is the topological skeleton diagram finally obtained by the topology awareness module.
[0036] Figure 9 is the display diagram of the topological manifold with an uncertainty degree greater than 40%.
[0037] Figure 10 is the three-dimensional visualization diagram of the confidence values of each structure output by the uncertainty quantification module. DETAILED DESCRIPTION OF THE INVENTION
[0038] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0039] Refer to Figure 1 , a method for fine feature extraction of filamentous structure topology awareness based on a large model, the steps are as follows:
[0040] Step 1: Construct a filamentous structure image dataset and divide it into a training set, a validation set, and a test set: Collect high-resolution original images of filamentous structures and label images of the same resolution to construct an initial filamentous structure image dataset. Augment the filamentous structure image dataset through data augmentation methods of geometric transformation and photometric distortion, and then divide the filamentous structure image dataset into a training set, a validation set, and a test set;
[0041] Step 2: Construct a large model for filamentous structure feature extraction: For the morphological features of filamentous structures, use the original architecture of the basic large model for image feature extraction and segmentation. Freeze the pre-trained weights of the basic large model and insert a hint layer with trainable parameters. Use the training set to perform vertical domain fine-tuning on the basic large model and save the best results of the validation set to obtain a large model for filamentous structure feature extraction;
[0042] Step 3: Extract the initial likelihood map: Use the large model for filamentous structure feature extraction to predict the original images of the test set to obtain an initial likelihood map with the foreground probability as the likelihood value;
[0043] Step 4: Calibration: Use a combined method of conditional prediction value offset, adaptive temperature adjustment, and interference filtering to calibrate the initial likelihood map to obtain a corrected likelihood map;
[0044] Step 5: Construct a manifold: Analyze the corrected likelihood map through discrete Morse theory, identify critical points such as saddle points and maximum points therein, and construct V paths connecting the critical points to form a manifold;
[0045] Step 6: Persistent homology processing and construction of a topological skeleton graph: Determine the survival period of each manifold through persistent homology, remove the manifolds with a survival period lower than the set threshold to obtain stable manifolds, and use the union of all stable manifolds as the topological skeleton graph of the filamentous structure;
[0046] Step 7: Construct an uncertainty quantification model for the topological skeleton graph, train the uncertainty quantification model by perturbing the stable manifold data, and use the trained uncertainty quantification model to predict the confidence of each stable manifold in the topological skeleton graph to form an uncertainty graph;
[0047] Step 8: Combine the topological skeleton graph and the corresponding uncertainty graph to optimize the segmented image predicted by the large model for filamentous structure feature extraction.
[0048] Figure 2 Shows the basic framework of a topological perception fine feature extraction system for filamentous structures based on a large model, which is mainly composed of five parts: a large model for filamentous structure feature extraction, a calibration module, a topological perception module, a persistent homology module, and an uncertainty quantification module.
[0049] To overcome the two major problems that traditional large foundation models are difficult to effectively identify filamentous structures and require additional external prompts, the present invention optimizes the traditional architecture of the large foundation model. First, the prompt encoder in the traditional image feature extraction and segmentation large model is removed, and external prompts in the four forms of prompt masks, prompt points, prompt boxes, and prompt texts are cancelled. The image encoder in the large foundation model consists of a series of serially connected Vision Transformer (ViT) models. However, in order to enable the large model to automatically extract prompt information from image features and replace manual annotation, the present invention adds a parameter-trainable prompt layer after each Vision Transformer (ViT) model in the image encoder to replace the function of the prompt encoder in the large foundation model. Since the prompt encoder is mainly used to process data in the form of text sequences such as image pixel coordinates, prompt box coordinates, and prompt texts, while the prompt layer is mainly used to process the output vectors of the ViT model, and their objectives are inconsistent, the present invention redesigned the internal structure of the prompt layer, and its data processing process can be represented by formula (1):
[0050] (1);
[0051] Wherein, is the number of the ViT model in the image encoder, represents the output of the th prompt layer, represents the output value of the th ViT model, represents layer normalization operation, represents a dynamic gating function, which is responsible for determining the weight coefficient of local features, represents local feature extraction operation, which is responsible for capturing topological structures at the detail level such as filamentous structure branch points, represents global feature extraction operation, which is responsible for modeling the long-range continuity of the main trunk of the filamentous structure, represents Hadamard product.
[0052] Secondly, the present invention replaces the mask decoder in the original architecture of the large foundation model with a parameter-trainable prediction head. The prediction head does not need to receive prompt features from the mask encoder and can directly convert the output of the image encoder to obtain the initial likelihood map of the original image. The process of obtaining the initial likelihood map of the original image through the filamentous structure feature extraction large model can be represented by formula (2):
[0053] (2);
[0054] Wherein, I represents the original image, , represents the image patch position encoding module, represents the feedforward function of the i-th ViT model, represents the n-times upsampling process, represents a channel compression operation, and ∘ represents a function composition operation.
[0055] Subsequently, an efficient parameter fine-tuning method is used to train the large-scale model for filamentous structure feature extraction on the filamentous structure image data. During fine-tuning, all modules except the hint layer and the prediction head are loaded with the pre-trained weights of the basic large-scale model and then the weight parameters are frozen. As the fine-tuning proceeds, the hint layer and the prediction head will gradually learn the unique characteristics of filamentous structures that are different from block-like objects, enabling the large-scale model for filamentous structure feature extraction to accurately capture the topological connectivity and morphological details of slender targets such as blood vessels and nerves without destroying the general representation capabilities of the basic large-scale model.
[0056] The calibration module is used to solve the problem of severe polarization of the probability distribution in the initial likelihood graph directly output by the large model for filamentous structure feature extraction. That is, the likelihood value after regularization (value range [0,1]) is extremely close to the two boundary values of 0 or 1, resulting in the problem that there are no extreme points and saddle points in the likelihood graph. The extremely extreme likelihood values of the polar distribution will cause the gradient field to degrade severely. The numerical difference between the likelihood values near the boundary value 0 is extremely small, and the same is true for the likelihood values near the boundary value 1. This makes the spatial gradient The modulus approaches zero. Such a flattened gradient field cannot meet the requirements of the discrete Morse theory for critical point detection: "local extreme points require a certain degree of fluctuation, and saddle points require a clear saddle-shaped transition."
[0057] The calibration module consists of a series of conditional prediction value offset links, adaptive temperature adjustment links, and interference filtering links. It first performs conditional prediction value offset on the values in the initial likelihood map (initial likelihood values) output by the large model for filamentous structure feature extraction. The purpose is to avoid the initial likelihood values from showing a systematic and large negative distribution, which causes the likelihood values after regularization using the sigmoid function to be compressed to an extreme range close to zero, thereby causing the problem of gradient disappearance and topological feature degradation. Conditionality means that the offset processing is only performed when more than 80% of the initial likelihood values are negative numbers less than -2. This process can be expressed by formula (3):
[0058] (3);
[0059] in, represents the initial likelihood value of the large model for extracting filamentous structure features, represents the likelihood value obtained from the conditional prediction value offset link, Indicates the offset value, represents the distribution of offset values, is the key parameter for controlling the offset distribution range. is a constant representing the variation amplitude of the offset, represents the probability that the value is less than -2.
[0060] The calibration module then processes the likelihood value obtained from the conditional prediction value offset through adaptive temperature adjustment. This method first sets a high confidence threshold and a low confidence threshold . Values greater than are in the high confidence region, values less than are in the low confidence region, and values between the two are in the intermediate region. These three different regions will be processed by different adjustment methods. The high confidence region will use temperature to compress extreme probabilities and avoid gradient disappearance; the low confidence region will use temperature to enhance probabilities close to 0 and restore potential signals; the intermediate region uses linear interpolation temperature to enhance the discrimination of intermediate values. This process can be expressed by formula (4):
[0061] (4);
[0062] where, is the likelihood value after further adaptive temperature adjustment, is a regularization function that maps values to a fixed interval and then compares them with the threshold to enhance the generalization of the processing process for values with different distributions;
[0063] The interference filtering link in the calibration module regularizes the likelihood value obtained from the adaptive temperature adjustment link and filters out the likelihood values whose regularized values are less than a threshold. These likelihood values will be set to 0 to avoid the interference of low likelihood value fluctuations on the identification of extreme points and saddle points. After processing, a corrected likelihood map with a more reasonable numerical distribution can be obtained.
[0064] The topology-aware module uses discrete Morse theory to analyze the corrected likelihood map. It regards the foreground probability in the corrected likelihood map as a terrain function and constructs a discrete gradient vector field on it, which can be expressed by formula (5):
[0065] (5);
[0066] where p and q both represent the coordinates of pixel points, f(p) and f(q) respectively represent the corrected likelihood values corresponding to pixel points p and q, and N(p) represents the neighborhood range of pixel point p, Point to the pixel point q with the largest corrected likelihood value among all pixel points in the neighborhood of p that are greater than f(p). If there is no point in the neighborhood of p with a corrected likelihood value greater than f(p), it means that the gradient of the pixel point p is 0, that is, ∇f(p)=0, and it belongs to a critical point (maxima and saddle points are collectively called critical points). For a critical point p, if there are at least two pixel points in the neighborhood with values equal to f(p), it means that the critical point p is a saddle point, otherwise it belongs to a maximum point.
[0067] In the discrete gradient vector field, non-critical points will flow to critical points through V paths. Discrete Morse theory starts from the isovalue neighborhood points of saddle points and follows the discrete gradient vector field to generate V paths. These paths terminate at maximum points, and then all V paths flowing to the same maximum point are merged to form the manifold corresponding to that maximum point, thereby converting local gradient information into a global topological structure.
[0068] Next, the persistence homology module will determine the lifespan of each manifold. The calculation method is to subtract the terrain function corresponding value when the manifold appears from the terrain function corresponding value when the manifold disappears (there are two cases: the structure disappears or it merges into a larger manifold). Subsequently, manifolds with a lifespan lower than the set threshold of 0.05 are filtered out to obtain stable manifolds, and then the union of all stable manifolds is used as the topological skeleton graph.
[0069] The stable manifolds generated by discrete Morse theory are unique. This process may lead to the generation of incorrect topological structures due to the distribution of corrected likelihood values in the corrected likelihood map. Therefore, the present invention designs an uncertainty quantification module to predict the credibility of stable manifolds, thereby removing unreliable branches. The uncertainty quantification module first adds a random perturbation to the distribution of the terrain function to obtain a noisy terrain function , where . After adding the random perturbation, it will cause changes in the formation process of V paths and manifolds. The uncertainty quantification module uses a random walk algorithm to reconstruct new manifolds on . When constructing, it starts from a saddle point and moves one pixel at a time until it reaches a maximum point. To avoid the change in the distribution of the terrain function caused by the perturbation, which may cause the random walk to deviate from the maximum point as the end point, this module introduces a distance constraint as a penalty term in the random walk algorithm. When the random walk deviates from the end point, the quality value will decrease. And the random walk algorithm will select the neighboring pixel point with the largest quality value as the next destination for walking. The process of reconstructing manifolds on the noisy terrain function can be expressed by formula (6):
[0070] (6);
[0071] Where c represents the current pixel position, represents a pixel position in the neighborhood O. Q represents the quality function, which is used to evaluate the quality of all neighboring pixels of the current pixel c, that is, to quantify the priority of each candidate point. The neighboring pixel with the largest quality value is the next destination of the random walk algorithm. c m represents the maximum point, γ represents the weight of the distance constraint, f n ( ) represents the noisy terrain function f n in the corresponding value.
[0072] When generating a new manifold, the uncertainty quantification module will perform m random perturbations to generate a set of random perturbations , represents the m-th random perturbation, so as to obtain a set of terrain functions , represents the terrain function of the m-th random perturbation. Finally, a set of manifold variants is obtained, represents the manifold variant obtained by the m-th random perturbation. Subsequently, the manifolds with short lifetimes are removed by the persistent homology method, and only the stable manifolds are retained. Then, each set of stable manifolds is constructed into a graph structure. Each node in the graph structure represents a stable manifold. When there is a non-zero overlap between stable manifolds, that is, when two stable manifolds have a shared area and an intersection, an edge is constructed between the nodes. The uncertainty quantification module is trained on the constructed graph structure to learn the method of quantifying the degree of uncertainty. After training, it will be used to predict the uncertainty map. The loss function L UQ during the training process of the uncertainty quantification module is defined as shown in formula (7):
[0073] (7);
[0074] Where represents the parameters of the uncertainty quantification module, K represents the set of all stable manifolds to be evaluated, that is, the entire graph structure, and the elements in the set are the nodes in the graph structure. |K| represents the number of elements in the set, k represents the stable manifold number, represents the probability that the uncertainty quantification model predicts that the k-th stable manifold belongs to the true structure. is the soft label of the k-th stable manifold, which represents the overlap ratio between the k-th stable manifold and the true structure in the label image, and can be represented by , where y is the binarized true structure, mask is the binary mask of the current manifold, and ⊙ is the Hadamard product. is the logarithmic variance of the prediction value of the uncertainty quantification module, which is used to quantify the uncertainty.
[0075] In the embodiments of the present invention, the eye retina vascular type images are used as the dataset, which includes two types: three-channel original images and one-channel label images. Geometric transformations such as rotation, scaling, and shearing are used; photometric distortions such as brightness adjustment, color jitter, and contrast enhancement are performed to augment the dataset size through data augmentation.
[0076] In the embodiments, the pre-trained Segment Anything is used as the basic large model. The segmentation effect is measured using 7 metrics, namely accuracy (Acc), recall, specificity, intersection over union (IoU), Dice coefficient (Dice), information measure (BM), and Matthews correlation coefficient (MCC).
[0077] The initial likelihood map obtained by processing the eye retina vascular images with the filamentous structure feature extraction large model in the present invention is as Figure 3 shown, Figure 3 The heat map form is used to visually display the likelihood map. Among them, the pixel points with colors closer to white have larger likelihood values, indicating a greater possibility of being the foreground, and vice versa, the pixel points closer to black have a greater possibility of being the background. Figure 3 The values on the color bar on the right represent the numerical values corresponding to the colors here. Since the initial likelihood values are not regularized, its maximum likelihood value is 10.446 and the minimum likelihood value is -13.968. It can be seen from Figure 3 that the filamentous structure contour is already relatively obvious, indicating that the large model can better extract the detailed features of the fine filamentous structure of the retina blood vessels.
[0078] To further verify the effectiveness of the large model for filamentous structure feature extraction proposed in the present invention, a comparative experiment was designed. The large model proposed in the present invention was compared with 11 segmentation schemes of the other 4 large models on 7 evaluation indicators. Among them, the large model Segment Anything-H is based on the ViT-Huge architecture and is trained on the SA-1B dataset containing 11M images + 1.1B masks; Segment Anything-B is based on the ViT-Base architecture and is a lightweight segmentation image large model dedicated to reducing the demand for computing resources; SAM2 introduces multi-scale feature fusion and long-range attention mechanisms and has strong global perception ability. These three large models all rely on manually designed prompt words. In the comparative experiment, 5 foreground point coordinates and 2 background point coordinates were uniformly designed for them as prompt words. The foreground coordinates are (78,277), (122,409), (316,426), (275,85), (121,167) respectively; the background coordinates are (392,409), (333,325) respectively. In addition, these three large models can produce multiple inconsistent segmentation results in one segmentation. In this experiment, the segmentation results were sorted according to the segmentation result weights, and the top 3 segmentation results with the highest weights were selected. Schemes 1, 2, and 3 come from the segmentation structures with the 1st, 2nd, and 3rd highest weights in one segmentation of the Segment Anything-H model; schemes 4, 5, and 6 come from Segment Anything-B; schemes 9, 10, and 11 come from SAM2. DINOv2 is a ViT-type model trained by self-supervised learning, and the extraction of general features is trained through image-level contrast learning. Scheme 7 is the segmentation result of this model. Scheme 12 is the segmentation index value at the background level of the large model of the present invention; scheme 13 is the index value at the foreground level. The results of the comparative experiment are shown in Table 1, and only the index values at the foreground level are listed for the first 11 schemes.
[0079] Table 1 Comparative Experiment on the Segmentation Effect of Large Models
[0080]
[0081] It can be seen from the data in Table 1 that the large model obtained in the present invention has excellent effects on both the background and the foreground, and the values of the four evaluation indicators of intersection over union, Dice coefficient, information measure, and Matthews correlation coefficient are much higher than those of the other 11 segmentation schemes. The recall rates of foreground pixel points in schemes 3, 4, and 9 are as high as over 0.97, which indicates that when the corresponding large models segmented the eye retina blood vessel images, they did not identify the filamentous structures therein, but regarded the entire eye area as the segmentation result, which belongs to obvious invalid segmentation.
[0082] Figure 4Shows an image of the retinal blood vessels of the eye used for testing. It was originally a three-channel color image and was grayscale-converted into a black-and-white image. The black five-pointed stars represent the foreground point coordinates marked in the prompt, with a total of 5; the black inverted triangles represent the background point coordinates marked in the prompt, with a total of 2. This figure clearly shows the approximate original form of the fine features of the filamentous structure planned to be processed by the present invention. It can be seen from the figure that the filamentous structure contained in the image is relatively weak, especially at the filamentous ends. Figure 5 Shows the segmentation visualization results of Scheme 10, which has the highest values among the first 11 comparison schemes in terms of four evaluation indicators: intersection over union, Dice coefficient, information measure, and Matthews correlation coefficient. Figure 6 Then shows the segmentation visualization results of Scheme 11, where the black part is the background area and the white part is the foreground area. From Figure 5 、 Figure 6 it can be seen that for general image segmentation large models, it is difficult to avoid the segmentation inertia in the form of blocks, and the output segmentation results are not in the filamentous form, further reflecting the necessity of the basic large model optimization and fine-tuning method proposed by the present invention.
[0083] If directly using discrete Morse theory to analyze the initial likelihood map output by the filamentous structure feature extraction large model, no extreme points and saddle points can be determined. At this time, the output topological skeleton map contains 0 V paths and stable manifolds. After calibrating the values in the initial likelihood map using a combined method of conditional prediction value offset, adaptive temperature adjustment, and interference filtering, the obtained topological skeleton map is as Figure 7 shown, which contains many redundant horizontal lines. Further using the persistent homology method to filter out the manifolds with extremely short durations and only retaining the stable manifolds to obtain the final topological skeleton map, as Figure 8 shown, where the values on the horizontal and vertical axes represent the coordinates of the pixel points.
[0084] Subsequently, in this embodiment, the uncertainty quantification module is used to analyze the degree of uncertainty of the obtained topological skeleton map, and only the manifolds with an uncertainty degree greater than 40% (i.e., a confidence level lower than 60%) are retained and visually displayed. The result is as Figure 9 shown. Figure 10 Then uses three-dimensional visualization to show the confidence level of the topological skeleton map. The larger the value, the higher the confidence level of the structure. For the convenience of display, some extreme values have been removed during visualization. The scales on the x and y axes represent the pixel point coordinates, and the z axis represents the confidence level value corresponding to the pixel point. From Figure 10 it can be seen that the confidence level of the obvious main trunk structure in the eye retinal blood vessel image is relatively high, while the confidence level of the relatively weak filamentous structure ends is relatively low. Finally, the topological skeleton map and the uncertainty map are combined to enhance the final prediction performance of the filamentous structure feature extraction large model.
[0085] This embodiment provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the method for fine feature extraction of filamentous structure topology perception based on a large model is implemented.
[0086] The above-described invention only expresses the implementation manners of the embodiments of the present invention, and thus cannot be construed as a limitation on the scope of the invention patent, nor is it a limitation on the structure of the embodiments of the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present invention, several changes and improvements can still be made, and these all belong to the protection scope of the embodiments of the present invention.
Claims
1. A method for extracting fine features of filamentous structure topology perception based on a large model, characterized in that Including: For the morphological characteristics of the filamentous structure, the original architecture of the basic large model for image feature extraction and segmentation. Freeze the pre-trained weights of the basic large model and add a trainable prompt layer after each ViT model in the image encoder of the basic large model. The prompt layer includes local feature extraction operations and global feature extraction operations. The weights of the extracted local features and global features are controlled by a dynamic gating function for summation, and finally, layer normalization is performed on the summed features; Use the training set to perform vertical domain fine-tuning on the basic large model and save the best results of the validation set to obtain a filamentous structure feature extraction large model; Use the filamentous structure feature extraction large model to predict the original image of the filamentous structure to obtain an initial likelihood map with the foreground probability as the likelihood value, calibrate the initial likelihood map to obtain a corrected likelihood map, analyze the corrected likelihood map through discrete Morse theory, identify critical points and construct a manifold; Judge the survival period of each manifold through persistent homology, remove the manifolds with a survival period lower than the set threshold to obtain stable manifolds, and use the union of all stable manifolds as the topological skeleton map of the filamentous structure; Construct an uncertainty quantification model for the topological skeleton map, train the uncertainty quantification model by perturbing the stable manifold data, and use the trained uncertainty quantification model to predict the confidence of each stable manifold in the topological skeleton map to form an uncertainty map; Combine the topological skeleton map and the corresponding uncertainty map to optimize the segmented image predicted by the filamentous structure feature extraction large model.
2. The method for extracting fine features of filamentous structure topology perception according to claim 1, wherein Use the parameter fine-tuning method to train the filamentous structure feature extraction large model on the filamentous structure image data. When fine-tuning, all other modules except the prompt layer and the prediction head are loaded with the pre-trained weights of the basic large model and then the weight parameters are frozen.
3. The method for extracting fine features of filamentous structure topology perception according to claim 1, wherein, The processing process of the prompt layer for data is expressed by formula (1): (1); Among them, is the number of the ViT model in the image encoder, represents the output of the th prompt layer, represents the output value of the th ViT model, represents the layer normalization operation, represents the dynamic gating function, which is responsible for determining the weight coefficients of local features, represents the local feature extraction operation, represents the global feature extraction operation, represents the Hadamard product.
4. The method for extracting fine features of filamentous structure topology perception according to claim 1, characterized in that The process of using the filamentous structure feature extraction large model to predict the original image of the filamentous structure to obtain an initial likelihood map with the foreground probability as the likelihood value is expressed by formula (2): (2); Among them, I represents the original image, , represents the image block position encoding module, represents the feed-forward function of the i-th ViT model, represents the n-fold upsampling process, represents the channel compression operation, represents the function composition operation.
5. The topological perception fine feature extraction method for filamentous structures according to claim 2, characterized in that Perform conditional prediction value offset on the initial likelihood value in the initial likelihood map output by the filamentous structure feature extraction large model, process the likelihood value obtained by conditional prediction value offset through adaptive temperature adjustment, the interference filtering link regularizes the likelihood value obtained by the adaptive temperature adjustment link, and filters out the likelihood values whose regularized values are less than a threshold. The filtered likelihood values will be set to 0. After processing, the corrected likelihood map is obtained; Conditional prediction value offset is expressed by formula (3): (3); Among them, represents the initial likelihood value of the filamentous structure feature extraction large model, represents the likelihood value obtained from the conditional prediction value offset link, represents the offset value, represents the distribution of the offset value, is a key parameter for controlling the distribution range of the offset amount, is a constant representing the change amplitude of the offset amount, represents the probability that the value is less than -2; Adaptive temperature adjustment is expressed by formula (4): (4); wherein, is the likelihood value after further adaptive temperature adjustment, is the regularization function, is the high confidence threshold, is the low confidence threshold, is the temperature used for the high confidence region, is the temperature used for the low confidence region.
6. A fine feature extraction system for topological perception of filamentous structures based on large models, characterized in that Including: The filamentous structure feature extraction large model is used to predict the original image of the filamentous structure to obtain an initial likelihood map with the foreground probability as the likelihood value, and predict the original image of the filamentous structure to obtain a segmentation image; the filamentous structure feature extraction large model is based on the basic large model for image feature extraction and segmentation, removes the prompt encoder in the basic large model, and cancels the external prompts in the four forms of prompt masks, prompt points, prompt boxes, and prompt texts. A trainable prompt layer is added after each ViT model in the image encoder of the basic large model, and the mask decoder in the basic large model is replaced with a trainable prediction head; the prompt layer includes local feature extraction operations and global feature extraction operations, and the weights of the extracted local features and global features are controlled by a dynamic gating function for summation, and finally layer normalization operations are performed on the summed features; The calibration module is used to calibrate the initial likelihood map to obtain a corrected likelihood map; The topological awareness module is used to analyze the corrected likelihood map, identify critical points and construct manifolds; The persistent homology module is used to construct a topological skeleton graph, judge the survival period of each manifold through persistent homology, remove the manifolds with a survival period lower than the set threshold to obtain stable manifolds, and use the union of all stable manifolds as the topological skeleton graph of the filamentous structure; The uncertainty quantification module has an uncertainty quantification model of the topological skeleton graph built in, and is used to predict the confidence of each stable manifold in the topological skeleton graph to form an uncertainty graph.
7. The filamentous structure topology-aware fine feature extraction system according to claim 6, characterized in that The calibration module consists of a conditional prediction value offset link, an adaptive temperature adjustment link, and an interference filtering link connected in series; the conditional prediction value offset is performed on the initial likelihood value in the initial likelihood map output by the filamentous structure feature extraction large model, and the likelihood value obtained from the conditional prediction value offset is processed through adaptive temperature adjustment. The interference filtering link regularizes the likelihood value obtained from the adaptive temperature adjustment link and filters out the likelihood values with a value less than a threshold after regularization. The filtered likelihood values will be set to 0, and the corrected likelihood map is obtained after processing; Among them, the conditional prediction value offset is expressed by formula (3): (3); Among them, represents the initial likelihood value of the filamentous structure feature extraction large model, represents the likelihood value obtained from the conditional prediction value offset link, represents the offset value, represents the distribution of the offset value, is the key parameter for controlling the distribution range of the offset amount, is a constant representing the change amplitude of the offset amount, represents the probability that the value is less than -2; The adaptive temperature adjustment is expressed by formula (4): (4); Among them, is the likelihood value after further adaptive temperature adjustment, is the regularization function, is the high confidence threshold, is the low confidence threshold, is the temperature used in the high confidence region, is the temperature used in the low confidence region.
8. The filamentous structure topology-aware fine feature extraction system according to claim 6, wherein The topological awareness module uses discrete Morse theory to analyze and correct the likelihood map, which regards the foreground probability in the corrected likelihood map as a terrain function , and constructs a discrete gradient vector field on it , which is represented by formula (5): (5); where p and q both represent the coordinates of pixel points, f(p) and f(q) respectively represent the corrected likelihood values corresponding to pixel points p and q, and N(p) represents the neighborhood range of pixel point p. It points to the pixel point q with the largest corrected likelihood value among all pixel points in the neighborhood of p that are greater than f(p). If there is no pixel point in the neighborhood of p with a corrected likelihood value greater than f(p), it means that the gradient of pixel point p is 0, that is , which belongs to the critical point. For the critical point p, if there are at least two pixel points in the neighborhood whose values are equal to f(p), it means that the critical point p is a saddle point; otherwise, it belongs to the maximum point. In a discrete gradient vector field, non-critical points will flow to critical points through V-paths. Discrete Morse theory starts from the isovalue neighborhood points of saddle points and follows the discrete gradient vector field to generate V-paths. The V-paths terminate at maximum points, and then all V-paths flowing to the same maximum point are merged to form the manifold corresponding to that maximum point.
9. The filamentous structure topology-aware fine feature extraction system according to claim 8, wherein The uncertainty quantification module first adds a random perturbation to the distribution of the terrain function to obtain a noisy terrain function . The uncertainty quantification module then uses a random walk algorithm to reconstruct a new manifold on . The reconstruction starts from a saddle point and moves one pixel at a time until it reaches a maximum point. The process of reconstructing the manifold on the noisy terrain function is expressed by Equation (6): (6); Among them, c represents the current pixel position, represents the position of a pixel point in the neighborhood O; Q represents the quality function, which is used to evaluate the quality of all neighboring pixel points of the current pixel c, that is, to quantify the priority of each candidate point. The neighboring pixel point with the largest quality value is the next destination of the random walk algorithm; c m represents the maximum point, γ represents the weight of the distance constraint, f n ( ) represents the value corresponding to the noisy terrain function f n in ; When generating a new manifold, the uncertainty quantification module will perform m random perturbations to generate a set of random perturbations , denotes the m-th random perturbation, thereby obtaining a set of terrain functions , denotes the terrain function of the m-th random perturbation, and finally obtains a set of manifold variants , denotes the manifold variant obtained from the m-th random perturbation; subsequently, the manifolds with a survival period lower than the set threshold are removed by the persistent homology method, and only the stable manifolds are retained; then each stable manifold set is constructed into a graph structure, where each node in the graph structure represents a stable manifold, and when there is a non-zero overlap between the stable manifolds, that is, when two stable manifolds have a shared area and an intersection, an edge is constructed between the nodes; the uncertainty quantification module is trained on the constructed graph structure to learn the quantification method of the uncertainty degree, and after the training is completed, it will be used to predict the uncertainty map.
10. The filamentous structure topology-aware fine feature extraction system according to claim 9, wherein The loss function L during the training process of the uncertainty quantification module UQ is defined as shown in Equation (7): (7); Among them, represents the parameters of the uncertainty quantification module, K represents the set of all stable manifolds to be evaluated, that is, the entire graph structure, and the elements in the set are the nodes in the graph structure; |K| represents the number of elements in the set, and k represents the stable manifold number. represents the probability that the uncertainty quantification model predicts that the k-th stable manifold belongs to the true structure. is the soft label of the k-th stable manifold, indicating the overlapping ratio of the k-th stable manifold and the true structure in the label image. , y is the binarized true structure, mask is the binary mask of the current manifold, and ⊙ is the Hadamard product. is the logarithmic variance of the prediction value of the uncertainty quantification module.
Citation Information
Patent Citations
Urban rail transit deformation monitoring method and system based on three-dimensional laser scanning
CN118500281A
Method for Detecting Submerged Ferromagnetic Objects and System for Detecting Submerged Ferromagnetic Objects
RU2015121872A