Multi-granularity progressive optimization molecular image segmentation method based on large model

Through the progressive segmentation method of particle size from coarse to fine, combined with the SAM model and the SegRefiner diffusion model, the problem of missegment and labeled data dependence in pathological image segmentation is solved, and efficient and accurate pathological image segmentation is achieved, especially high-quality segmentation of multi-center hemorrhoid tissue.

CN120339318APending Publication Date: 2025-07-18NANJING HOSPITAL OF TCM +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510492026.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing pathological image segmentation techniques are prone to missegment when dealing with complex backgrounds or uneven light, rely on a large amount of labeled data and have limited generalization capabilities, especially when dealing with rare or new types of lesions.

Method used

The multi-grained progressive optimization molecular image segmentation method based on large models is adopted. Through the progressive segmentation method of particle size from coarse to fine, the SAM model is first used to generate a preliminary coarse-grained mask image, and then the second particle size segmentation is performed using the SegRefiner diffusion model. Combined with morphological operations and contour refinement post-processing technology, boundary blur and noise interference are gradually eliminated.

Benefits of technology

It improves the accuracy and efficiency of pathological image segmentation, reduces the dependence on fine labeled data, and can effectively restore fine structural details such as tiny blood vessels and cell gaps in complex pathological images, and generates high-quality segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339318A_ABST
    Figure CN120339318A_ABST
Patent Text Reader

Abstract

The invention provides a multi-granularity progressive optimization molecular image segmentation method based on a large model, and belongs to the technical field of medical image segmentation. The method comprises the steps of firstly obtaining a pathological tissue image of a patient; inputting the pathological tissue image into an image segmentation framework SAM model for first granularity segmentation, and generating a preliminary coarse granularity mask image; and inputting the obtained coarse-grained mask image into a fine-grained refining network based on a diffusion model to carry out second-time granularity segmentation, and finally outputting a molecular image segmentation result with clear contour boundaries of different organizations and complete internal details. According to the method, the segmentation quality of the pathological image of the multi-center hemorrhoid tissue is improved through a progressive segmentation method with the granularity from coarse to fine, the accuracy and efficiency of image segmentation are improved, and meanwhile, the dependence on fine annotation data is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and particularly relates to a multi-granularity progressive optimization molecular imaging segmentation method based on a large model. Background Art

[0002] In disease diagnosis and research, pathological images serve as important evidence, and the accuracy and efficiency of their analysis directly affect medical decisions. The pathological image segmentation technology has emerged, aiming to accurately divide target regions such as different tissues and cells in pathological images, providing key support for subsequent quantitative analysis, disease diagnosis, and treatment plan formulation. Pathological image segmentation is of crucial significance for disease diagnosis, condition assessment, etc. Currently, the following technical means are mainly adopted for pathological image segmentation:

[0003] 1) Traditional image processing methods: Such methods include threshold segmentation, edge detection, and region growing, etc. Threshold segmentation distinguishes the target and background in the image by setting one or more gray thresholds. This method is simple to operate and has high computational efficiency, but in the face of complex backgrounds or uneven illumination, it is prone to missegmentation. Edge detection focuses on regions with significant gray changes in the image to outline the target contour; however, this method is very sensitive to noise and may result in discontinuous or false edges. Region growing is a method that starts from selected seed points and gradually expands the region according to the similarity between pixels; although it can overcome some problems to a certain extent, for cases where there are multiple tissue types intertwined, it is often difficult to find suitable seed points, resulting in unsatisfactory segmentation results with relatively low accuracy and efficiency.

[0004] 2) Machine learning-based methods: In recent years, with the development of machine learning, especially deep learning, methods such as support vector machines (SVM), random forests, and convolutional neural networks (CNN) have been widely used in pathological image segmentation. These methods build models by learning a large number of labeled samples to achieve automated image segmentation. Nevertheless, they still face some challenges: First, the acquisition of high-quality labeled data is both time-consuming and laborious, and may be affected by subjective factors; second, even with a large amount of training data, for those highly heterogeneous pathological images, the generalization ability of the model is still limited, especially when dealing with rare or new types of lesions, and it performs poorly and depends on fine-labeled data. Summary of the Invention

[0005] Aiming at the deficiencies of the existing technology, the present invention provides a multi-granularity progressive optimization molecular imaging segmentation method based on a large model. By means of a progressive segmentation method with gradually finer granularity, the segmentation quality of pathological images of multi-centered hemorrhoid tissues is improved, the accuracy and efficiency of image segmentation are increased, and at the same time, the dependence on fine-labeled data is reduced.

[0006] To achieve the above technical objectives, the present invention adopts the following technical solutions:

[0007] A multi-granularity progressive optimization molecular image segmentation method based on a large model, comprising the following steps:

[0008] Step S1: Obtain the pathological tissue image of the patient;

[0009] Step S2: Input the pathological tissue image into the image segmentation framework SAM model for the first granularity segmentation to generate a preliminary coarse-grained mask image;

[0010] Step S3: Input the obtained coarse-grained mask image into the fine-grained refinement network based on the diffusion model for the second granularity segmentation, and finally output the molecular image segmentation result with different tissues.

[0011] Preferably, the tissue types for the first granularity segmentation and the second granularity segmentation include: inflammation, thrombus, edema, and capillaries.

[0012] Preferably, in step S2, the process of performing the first granularity segmentation on the pathological tissue image by the SAM model includes the following steps:

[0013] Step S21: Input the pathological tissue image to be segmented into the SAM model, and the SAM model includes an image encoder, a prompt encoder, and a mask decoder;

[0014] Step S22: Process the input pathological tissue image through the image encoder to convert it into an image feature vector;

[0015] Step S23: According to the segmentation requirements, input prompt information on the pathological tissue image, and the prompt forms include point prompts, box prompts, and text prompts describing the target;

[0016] Step S24: Convert the input prompt information into a prompt vector through the prompt encoder;

[0017] Step S25: Fuse the image feature vector and the prompt vector through the mask decoder, and decode to generate a preliminary coarse-grained mask image.

[0018] Preferably, in step S22, the image encoder uses the ViT model to process the pathological tissue image to convert it into an image feature vector, and the processing process includes the following steps:

[0019] Step S221: Input the pathological tissue image into the ViT model, perform a convolutional image patch embedding operation to obtain a feature map;

[0020] Step S222: Add positional encoding to the obtained feature map;

[0021] Step S223: The feature map embedded with positional encoding undergoes feature extraction through multiple Transformer blocks;

[0022] Step S224: The processed feature map is reduced in dimension through multi-layer convolution operations, and finally an image feature vector is obtained.

[0023] Preferably, in the process of performing Step S25 to generate a preliminary coarse-grained mask image, first, an output token and a mask token are generated by a mask decoder to represent the current prediction state of the model, and a concatenation operation is performed on the output token and the mask token to obtain a fused feature vector. Then, the fused feature vector, the image feature vector, and the prompt vector are input into the mask decoder for deep fusion and decoding, and finally a preliminary coarse-grained mask image is generated.

[0024] Preferably, after performing Step S25, generating a preliminary coarse-grained mask image can evaluate each mask in the predicted coarse-grained mask image through a hypernetwork, and output the quality score of each mask.

[0025] Preferably, in Step S3, the rough mask image generated by the SAM model is subjected to a second granularity segmentation by using the SegRefiner diffusion model.

[0026] Preferably, using the SegRefiner diffusion model to complete the second granularity segmentation of the coarse-grained mask image in Step S3 includes the following steps:

[0027] Step 31: The coarse-grained mask image obtained in Step S2 is input into the SegRefiner diffusion model, and a random noise vector is introduced through the forward diffusion process of the SegRefiner diffusion model to simulate the initial noise state, and a noisy mask image is obtained as the training data of the SegRefiner diffusion model;

[0028] Step S32: The noisy mask image is input into the SegRefiner diffusion model, and through the reverse diffusion process of the SegRefiner diffusion model, learn to restore the noisy mask image to the coarse-grained mask image, and obtain a trained SegRefiner diffusion model;

[0029] Step S33: The coarse-grained mask image obtained in Step S2 is input into the trained SegRefiner diffusion model, and through the reverse diffusion process, the rough mask image is subjected to a second granularity segmentation, and finally the molecular image segmentation result with different tissues is output.

[0030] Preferably, the SegRefiner diffusion model realizes the reconstruction of a one-dimensional noise binary sequence by introducing binary diffusion. The expression formulas for the forward diffusion process and the reverse diffusion process of the SegRefiner diffusion model are as follows:

[0031] q(x t |x t-1 ) = B(x t |x t-1 (1 - β t ) + 0.5β t )

[0032] p θ (x t |x t-1 ) = B(x t |f b (x t , t))

[0033] Among them, q(x t |x t-1 ) represents the transition probability of the forward diffusion process. x t and x t-1 represent the states or values at time step t or t - 1 respectively, and take two values in the binary diffusion model: 0 or 1; B(x t |x t-1 (1 - β t ) + 0.5β t ) is the specific function of B(x t |μ), and B(x t |μ) is the probability mass function of the Bernoulli distribution, where μ is the success probability, that is, the probability that x t = 1. For a given μ, the probability mass function of the Bernoulli distribution is: B(x t |μ) = μ^x t (1 - μ)^(-x t ); β t is the hyperparameter used at time step t in the forward diffusion process and belongs to the interval (0, 1); p θ (x t |x t-1 ) represents the transition probability of the reverse diffusion process, and f b (x t , t) is the model for predicting the Bernoulli probability, which is used to estimate the probability that x t = 1 at time step t; θ is the model adjustment parameter.

[0034] Preferably, for the reverse diffusion process of the SegRefiner diffusion model in step 32, it includes the following steps:

[0035] Step S321: Noise prediction: Use the SegRefiner diffusion model to predict the noise part to be removed in the next stage based on the current masked image with noise.

[0036] Step S322: Update the mask image: Based on the prediction result, adjust the current masked image, remove the predicted noise components, and make the masked image gradually approach the real details.

[0037] Step S323: Repeat the above Step S321 and Step S322 until the preset number of iterations is reached and then stop, and output the trained SegRefiner diffusion model.

[0038] Compared with the prior art, the beneficial effects produced by the present invention are as follows:

[0039] 1) The present invention provides a progressive segmentation method with gradually finer granularity. First, use the SAM model to generate a preliminary coarse-grained masked image, so as to quickly capture the general outlines and structures of different types of tissues and cells in the histopathological image. Input the obtained coarse-grained masked image into the fine-grained refinement network based on the diffusion model, and finally output a pathological structure image with clear contour boundaries and complete internal details of different tissues. Compared with the traditional single-granularity segmentation method, the multi-granularity segmentation effect is better, improving the segmentation quality of the pathological images of multi-centered hemorrhoid tissues and indicating the segmentation accuracy.

[0040] 2) When performing the first granularity segmentation, the method of the present invention uses the general image segmentation framework SAM model to generate a preliminary coarse-grained masked image, so as to quickly capture the general outlines and structures of different types of tissues and cells in the histopathological image. Different from the traditional algorithms that rely on a large amount of labeled data for training and are only used for segmenting specific types of objects, the SAM model has strong semantic understanding ability. It can extract high-level abstract features from few-sample histopathological images and flexibly determine the segmentation targets according to the input prompts (such as points, boxes) provided by users, greatly reducing the dependence on fine-labeled data and saving a large amount of labeling costs.

[0041] 3) In the second grain segmentation stage of the present invention, a reverse diffusion processing mechanism based on the SegRefiner diffusion model is introduced. Combining morphological operations and contour refinement post-processing techniques, through an iterative process of multi-round noise prediction and mask update, the boundary blurring and noise interference generated by the rough segmentation are gradually eliminated. Compared with traditional segmentation methods, such as the PaNet classification network and other methods, the present invention can effectively restore the fine structural detail information features such as tiny blood vessels and cell gaps in complex pathological images, solve the problem that the existing methods are prone to produce burr-like artifacts and detail loss at tissue junctions, improve the accuracy of microscopic structures such as the morphology of capillary branches and the boundaries of inflammatory infiltration regions, and the generated segmentation images have high diversity and accuracy, obtaining high-quality segmentation results. Description of the Drawings

[0042] Figure 1 It is a flowchart of a multi-grain progressive optimization molecular imaging segmentation method based on a large model according to an embodiment of the present invention;

[0043] Figure 2 It is an image segmentation result diagram of the input pathological tissue image processed by the multi-grain progressive optimization molecular imaging segmentation method based on a large model according to an embodiment of the present invention. Detailed Embodiments

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] Embodiment 1

[0046] Combined with the attached Figure 1 - Figure 2 As shown, the embodiment of the present invention provides a multi-grain progressive optimization molecular imaging segmentation method based on a large model, including the following steps:

[0047] Step S1: Obtain the pathological tissue image of the patient;

[0048] Step S2: Input the pathological tissue image into the image segmentation framework SAM model for the first grain segmentation to generate a preliminary rough-grain mask image;

[0049] Step S3: Input the obtained rough-grain mask image into the fine-grain refinement network based on the diffusion model for the second grain segmentation, and finally output the molecular imaging segmentation result with different tissues.

[0050] Among them, the first granularity segmentation is the coarse-grained segmentation of the pathological tissue image, generating a coarse-grained mask image; the second granularity segmentation is the fine-grained segmentation, that is, further segmenting the input coarse-grained mask image, and outputting a pathological structure image with clear contour boundaries of different tissues and complete internal details.

[0051] The segmentation method of the present invention first uses the general image segmentation framework SAM (Segment Anything Model) model to generate a preliminary coarse-grained mask image, so as to quickly capture the general contours and structures of different types of tissues and cells in the histopathological image. Then, in order to further improve the accuracy of tissue pathology region segmentation and refine tissue contours, the present invention inputs the obtained coarse-grained mask image into a fine-grained refinement network based on a diffusion model, and finally outputs a pathological structure image with clear contour boundaries of different tissues and complete internal details.

[0052] The present invention uses an artificial intelligence large language model combined with deep learning technology to perform high-precision segmentation on the pathological images of multi-centered hemorrhoid tissues at the molecular level. By automatically analyzing information such as cell morphology and tissue structure in the pathological images, combined with patient characteristics, environmental factors, treatment plans, surgical anesthesia time, wound size, etc., it reveals the molecular mechanism of the onset of multi-centered and multi-regional hemorrhoid diseases and the mechanism of related factors affecting the curative effect after surgery.

[0053] Example 2

[0054] Combined with the attached Figure 1 - Figure 2 As shown, the embodiment of the present invention provides a multi-granularity progressive optimization molecular imaging segmentation method based on a large model, including the following steps:

[0055] Step S1: Obtain the pathological tissue image of the patient;

[0056] Step S2: Input the pathological tissue image into the image segmentation framework SAM model for the first granularity segmentation to generate a preliminary coarse-grained mask image;

[0057] Step S3: Input the obtained coarse-grained mask image into a fine-grained refinement network based on a diffusion model for the second granularity segmentation, and finally output the molecular imaging segmentation result with different tissues.

[0058] In this embodiment, the tissue types of the first granularity segmentation and the second granularity segmentation include: inflammation, thrombosis, edema, capillaries, etc.

[0059] In the embodiment, in step S2, the process of performing the first granularity segmentation on the pathological tissue image by the SAM model includes the following steps:

[0060] Step S21: Input the pathological tissue image to be segmented into the SAM model, which includes an image encoder, a prompt encoder, and a mask decoder.

[0061] Step S22: Process the input pathological tissue image through the image encoder to convert it into an image feature vector.

[0062] Step S23: According to the segmentation requirements, input prompt information on the pathological tissue image. The prompt forms include point prompts, box prompts, and text prompts describing the target. Among them, a point prompt is to select some key points on the image as prompt information; a box prompt is to use a rectangular box to select the target area of interest; a text prompt is to describe the target by inputting text.

[0063] Step S24: Convert the input prompt information into a prompt vector through the prompt encoder.

[0064] Step S25: Fuse the image feature vector and the prompt vector through the mask decoder, and decode to generate a preliminary coarse-grained mask image.

[0065] In the embodiment, in step S22, the image encoder uses a ViT (Vision Transformer) model to process the pathological tissue image and convert it into an image feature vector. The processing process includes the following steps:

[0066] Step S221: Input the pathological tissue image into the ViT model and perform a convolutional image patch embedding operation to obtain a feature map. In the specific implementation process, image patches of size 16×16 can be used to process the input image, and the stride of the convolutional operation is set to 16. After this operation, the size of the feature map will be reduced by 16 times, and the number of channels of the feature map will be mapped from the original 3 channels to 768 channels.

[0067] Step S222: Add positional encoding to the obtained feature map. In the specific implementation process, the positional encoding is a learnable parameter matrix. When the model is initialized, all elements of this matrix are set to 0. The purpose of adding positional encoding is to enable the model to perceive the position information of each image patch in the original image.

[0068] Step S223: The feature map with positional encoding embedded is subjected to feature extraction through multiple Transformer blocks. In the specific implementation process, 16 Transformer blocks can be used for feature extraction. Among them, 12 Transformer blocks adopt the self-attention module based on window partitioning. The self-attention module based on window partitioning will divide the feature map into 14×14 windows, and then perform self-attention calculations within each window. This method helps the model capture local features and information. The other 4 Transformer blocks adopt global attention modules, and these 4 global attention modules are interspersed between the self-attention modules based on window partitioning. The global attention module enables the model to perform self-attention calculations within the entire range of the feature map, thereby capturing the global features and long-range dependencies of the image.

[0069] Step S224: The processed feature map is reduced in dimension through multiple convolutional operations, and finally an image feature vector is obtained. In the implementation process, the feature map processed by 16 Transformer blocks will go through two more convolutional operations to reduce the number of channels of the feature map from 768 to 256. The result obtained after this step is the final image embedding result, that is, it is transformed into a lower-dimensional image feature vector. The key feature information extracted from the pathological tissue image after being processed by the ViT model is the final image feature vector.

[0070] In the embodiment, in the process of executing step S25 to generate a preliminary coarse-grained mask image, first, a mask decoder is used to generate output tokens and mask tokens to represent the current prediction state of the model, and a concatenation operation is performed on the output tokens and mask tokens to obtain a fused feature vector. Then, the fused feature vector, the image feature vector, and the prompt vector are input into the mask decoder for deep fusion and decoding, and finally a preliminary coarse-grained mask image is generated. Through this step, effective integration of different types of information can be achieved, enabling the model to comprehensively consider information from different sources for more accurate segmentation. This combination method allows the model to not only rely on the features of the image itself (expressed by the image feature vector) but also combine the user's intention (expressed by the prompt vector) and the current prediction state of the model (reflected by the output tokens and mask tokens) when generating the final segmentation result. Through this process, the accuracy and granularity of segmentation can be significantly improved, especially when dealing with complex or subtle pathological tissue images. In short, effective fusion of multiple information sources is achieved, enhancing the model's ability to understand the image content and the accuracy and detail level of its segmentation results.

[0071] In this embodiment, after performing step S25, generating a preliminary coarse-grained mask image can be evaluated by a hyper-network specifically for each mask in the predicted coarse-grained mask image, and the quality score of each mask is output. The hyper-network completes this task through a small multi-layer perceptron (MLP). It takes the representation of the mask as input and then outputs the quality score of each mask. The specific implementation process is to take the representation form of each mask as input, and the output is a vector whose length is equal to the number of masks plus one. Among them, the first element is actually a useless prediction (i.e., the case where the mask is zero). Then, this vector is used as the weighted mask feature to calculate the quality score of the mask.

[0072] In step S2 of the present invention, by using the SAM model to generate a preliminary coarse-grained mask image, it can quickly capture the general outlines and structures of different types of tissues and cells in histopathological images, improve the accuracy and efficiency of image segmentation, and at the same time reduce the dependence on finely annotated data. It has the following advantages compared with traditional methods: 1) Reducing the dependence on finely annotated data: Different from traditional algorithms that rely on a large amount of annotated data for training and only segment specific types of objects, the SAM model has strong semantic understanding capabilities. It can extract high-level abstract features from few-sample histopathological images and flexibly determine the segmentation target according to the input prompts provided by the user (such as points, boxes), greatly reducing the dependence on finely annotated data and saving a large amount of annotation costs. 2) Flexible prompting mechanism: The SAM model allows users to specify the regions of interest in various ways (such as point prompts, box prompts, text prompts describing the target). This flexibility enables SAM to adapt to different application scenarios and requirements, improving the generality and applicability of the segmentation task. 3) Efficient image encoder: The image encoder in the SAM model adopts the ViT model, which divides the image into multiple small blocks, performs embedding operations on them, and then uses Transformer blocks for feature extraction. This method can not only capture global information but also effectively process local details, providing rich feature representations for subsequent mask decoding. 4) In the mask decoding stage, the SAM model not only relies on the image feature vector but also combines the prompt vector for fusion processing. This method can better integrate the intention information from the user, thereby improving the accuracy and pertinence of segmentation.

[0073] Embodiment 3

[0074] Combined with the attached Figure 1 - Figure 2 As shown, the embodiment of the present invention provides a multi-granularity progressive optimization molecular imaging segmentation method based on a large model, including the following steps:

[0075] Step S1: Obtain the pathological tissue image of the patient;

[0076] Step S2: Input the pathological tissue image into the image segmentation framework SAM model for the first granularity segmentation to generate a preliminary coarse-grained mask image;

[0077] Step S3: Input the obtained coarse-grained mask image into the fine-grained refinement network based on the diffusion model for the second granularity segmentation, and finally output the molecular image segmentation result with different tissues.

[0078] Because the SAM model cannot comprehensively and finely segment the boundaries and structures of each tissue in medical image segmentation, especially for irregular pathological tissue images, the present invention takes the coarse-grained pathological tissue image obtained by the SAM model as the input of the diffusion model for further second granularity segmentation.

[0079] In the embodiment, in step S3, the coarse-grained mask image generated by the SAM model is subjected to the second granularity segmentation by using the SegRefiner diffusion model. The pixels in the rough mask are gradually converted into a fine state by the SegRefiner diffusion model, so as to accurately correct the mispredicted regions in the rough mask. Completing the second granularity segmentation of the coarse-grained mask image in step S3 by using the SegRefiner diffusion model includes the following steps:

[0080] Step 31: Input the coarse-grained mask image obtained in step S2 into the SegRefiner diffusion model, and introduce a random noise vector through the forward diffusion process of the SegRefiner diffusion model to simulate the initial noise state, and obtain a noisy mask image as the training data of the SegRefiner diffusion model;

[0081] Step S32: Input the noisy mask image into the SegRefiner diffusion model, and learn to restore the noisy mask image to the original image through the reverse diffusion process of the SegRefiner diffusion model to obtain a trained SegRefiner diffusion model;

[0082] Step S33: Input the coarse-grained mask image obtained in step S2 into the trained SegRefiner diffusion model, and through the reverse diffusion process, perform the second granularity segmentation on the rough mask image, and finally output the molecular image segmentation result with different tissues.

[0083] The SegRefiner diffusion model realizes the reconstruction of the one-dimensional noise binary sequence by introducing binary diffusion. The expression formulas of the forward diffusion process and the reverse diffusion process of the SegRefiner diffusion model are:

[0084] q(x t |x t-1) = B(x t |x t-1 (1 - β t ) + 0.5β t ) (1)

[0085] p θ (x t |x t-1 ) = B(x t |f b (x t , t)) (2)

[0086] Among them, q(x t |x t-1 ) represents the transition probability of forward diffusion. x t and x t-1 represent the state or value at time step t or t - 1 respectively, and take two values in the binary diffusion model: 0 or 1; B(x t |x t-1 (1 - β t ) + 0.5β t ) is the specific function of B(x t |μ), and B(x t |μ) is the probability mass function of the Bernoulli distribution, where μ is the success probability (i.e., the probability that x t = 1). For a given μ, the probability mass function of the Bernoulli distribution is: B(x t |μ) = μ^x t (1 - μ)^(-x t ); β t is the hyperparameter used at time step t in the forward diffusion process and belongs to the interval (0, 1). The parameter β t controls the transition probability from x t-1 to x t . A smaller β t means less noise is added to the system. This formula describes the forward diffusion process of generating x t-1 from x t . Specifically, x t converts to a new state with a certain probability according to the state of the previous step x t-1 and the hyperparameter β t . x t-1 (1 - β t ) + 0.5β t calculates the probability that x t = 1, combining the state information of the previous step and a random noise component (regulated by β t ), so that as t increases, the state gradually becomes more uncertain or "more noisy".

[0087] where p θ (x t |x t-1 ) represents the transition probability of reverse diffusion, and f b (x t , t) is a model for predicting Bernoulli probability, which is used to estimate the probability that x t = 1 at time step t; θ is a model adjustment parameter, and by adjusting θ, p θ (x t |x t-1 ) is made as close as possible to the distribution of the real data. This model makes predictions based on the current state x t and time step t. The reverse diffusion process describes starting from a noisy state and gradually removing the noise to reconstruct the original clear image segmentation result.

[0088] In a specific implementation, the forward diffusion process can be defined as a discrete random variable that transitions between multiple states, and the state transition matrix Q t is used to characterize this process:

[0089] [Q t m,n = q(x t = n|x t-1 = m) (3)

[0090] where the element in the m-th row and n-th column of the transition matrix Q t represents the probability of transitioning from state m to state n. Specifically, it is the probability that the state at time step t - 1 transitions from m to state n at time step t. For a binary diffusion model, m and n can only take values of 0 or 1. Therefore, the state transition matrix Q t is actually a 2x2 matrix.

[0091] In this embodiment, for the forward diffusion process of the SegRefiner diffusion model described in step 31, specifically:

[0092] In the forward diffusion process, by gradually reducing the accuracy of the fine mask M fine (which can also be called the ground truth mask), it is transformed into a coarse mask M coarse . Among them, the fine mask is a mask for accurately segmenting and labeling objects or regions in the image, containing very detailed and accurate information; while the coarse mask is a relatively rough segmentation of objects or regions in the image, losing some detailed information.

[0093] In the specific implementation process, the fine mask m0 = M fine and the coarse mask m T = M coarse ​, during the conversion process, there is a series of intermediate stages, which can be divided by time steps. At any intermediate time step, an intermediate mask m is generated. t , which is in the transition stage between M fine and M coarse . Define that each pixel in m t has two states: fine state and rough state. Therefore, the forward diffusion process is formulated as the state transition of pixels between these two states. Pixels in the fine state will retain their values from the fine mask M fine , and in the rough state, they will take the values in the rough mask M coarse .

[0094] The present invention proposes a new conversion sampling module to formulate this process. During the forward diffusion process, the conversion sampling module takes the mask m t-1 at the previous moment, the rough mask m T and the state transition probability as inputs, and outputs the converted mask m T , which is the rough mask finally obtained by the model at the time step t = T. Among them, for the mask m t-1 at the previous moment, at the beginning stage of the forward diffusion process, it is the initial fine mask m0, the preliminary coarse-grained mask image generated by the SAM model; in the middle stage of the forward diffusion process, for each subsequent time step t, the mask m t-1 at the previous moment is generated by the diffusion process of the previous time step t-1. Specifically, it is obtained by applying the state transition matrix Q t to the state m t-2 of the previous step. The rough mask m T is gradually made rough from the fine mask m0 through multiple diffusion steps. The state transition probability describes the probability that each pixel in m t-1 transitions to the rough state, and is calculated through the expression formula of the forward diffusion process. The module first performs Gumbel-max sampling according to the given state transition probability to determine which pixels will undergo state transitions and obtain the pixels that need to be converted. Then, the pixels that need to be converted will take values from the rough mask m T , while the pixels that do not need to be converted will keep their values unchanged and still maintain the values in the mask at the previous moment.

[0095] Note that the conversion sample module represents a one-way process, and only "conversion to the rough state" occurs. The one-way nature ensures that the forward process will converge to M coarse , although each step is completely random. This is a significant difference between the SegRefiner model and previous diffusion models. Finally, the forward diffusion process converges to random noise.

[0096] The present invention formulates the above process by introducing a binary random variable x. (where t represents the time step variable, t ∈ [0, T]) is a binary one-hot vector used to represent the intermediate mask m t and the state of the pixel (i, j) in it, and respectively set and to represent the fine state and the coarse state. Therefore, the forward diffusion process can be formulated as:

[0097]

[0098] where Q t is the state transition matrix at time step t, and β t is the hyperparameter at time step t, belonging to the interval (0, 1), which controls the transition probability from the fine state to the coarse state, and 1 - β t is the probability that the pixel remains in the fine state. A gradually increasing sequence can be set, such as β1 < β2 <... < β T , to ensure that as the diffusion steps progress, more pixels will transition to the coarse state. The form of Q t clearly shows the unidirectionality, that is, all pixels in the coarse state will never transition back to the fine state. According to the above formula, the marginal distribution can be formulated as:

[0099]

[0100] where, In view of this, we can obtain the intermediate mask m t at any intermediate time step t without having to sample q(x t |x t-1 ) step by step, which can speed up the training.

[0101] In this embodiment, for the reverse diffusion process of the SegRefiner diffusion model in step 32, it includes the following steps:

[0102] Step S321: Noise prediction: Use the SegRefiner diffusion model to predict the noise part to be removed in the next stage according to the current noisy mask image;

[0103] Step S322: Update the mask image: Based on the prediction result, adjust the current mask image, remove the predicted noise components, and make the mask image gradually approach the real details;

[0104] Step S323: Repeat the above steps S321 and S322 until the preset number of iterations is reached and then stop, and output the trained SegRefiner diffusion model.

[0105] In the specific implementation process, the reverse diffusion process is to convert the rough mask m T The process of gradually transforming into a fine mask m0. Since the fine mask m0 and the reverse state transition probability are unknown, the denoising diffusion probability model (DDPM) is followed to train a neural network f parameterized by θ. θ To predict the fine mask Its formula is:

[0106]

[0107] Where I is the input image, t is the time step, and m t is the intermediate mask; θ is the model adjustment parameter; Represents the predicted refined mask; express Corresponding confidence score. In order to obtain the reverse state transition probability, according to formula (5), formula (6) and Bayes' theorem, firstly, the posterior formula of time step t-1 is formulated as:

[0108]

[0109] Among them, x0 is the fine state, x t and x t-1 Represents the state or value at time step t or t-1 respectively; Q t is the state transfer matrix at time step t; Q t The transposed matrix of That is, the multiplication of the state transfer matrix from time step 1 to time step t. During model training, the fine state x0 is set to [1,0], representing the true value of the pixel. During inference, x0 is unknown and the predicted fine mask is may not be completely accurate. Due to the confidence score represents the level of certainty that the model predicts the correctness of each pixel, can also be interpreted as the probability of being in the fine state. Therefore, by simply thresholding, we can obtain The state of each pixel in is:

[0110]

[0111] in, A fine mask used to represent the prediction The state of pixel (i,j) in Represents the predicted fine mask The corresponding confidence score; when it has a higher confidence score, the state value indicating that they are in the fine state, and conversely, the state value However, in this one-hot form, the value of the state transition probability is only determined by the predefined hyperparameter θ, which will cause significant information loss. The present invention retains the soft transition and sets Under this setting, the reverse diffusion process can be re-expressed as:

[0112]

[0113] where, is the reverse state transition matrix at time step t; is the hyperparameter in the reverse diffusion process at time step t, represents the predicted fine mask at time step t; represents the corresponding confidence score. Using the above reversed state transition probability, with m t and as inputs, the conversion sampling module can convert a part of the pixels to the fine state at each time step t, thereby correcting the wrong prediction.

[0114] According to the given m t and its corresponding image, we first initialize all pixels to be in the rough state, that is Then perform the operations: (1) Forward pass to obtain the maximum value and (2) Calculate the reverse state transition matrix and obtain x t-1 . (3) Based on x t-1 , m t and calculate the refinement mask m t-1 . Iterate the process (1)-(3) until we obtain the fine mask m0.

[0115] In addition, after multiple rounds of iterative reverse diffusion processing, the refinement of the mask image is greatly achieved, but it may contain subtle artifacts and irregular regions. The segmentation result can be further optimized by introducing post-processing techniques to ensure that the boundaries between different tissue regions are smoother and more natural. In the subsequent post-processing stage, two main processing methods are mainly used. On the one hand, basic morphological operations such as dilation, erosion, opening, and closing are introduced. With these operations, small noise points existing in the mask image can be effectively removed, holes in the image can be filled, and the boundaries can be smoothed to improve the segmentation effect. On the other hand, contour detection technology is used to accurately locate the edges of the target object, and the detected edges are refined and smoothed through algorithms such as the Sobel operator and the Canny edge detector to ensure that the edges of the output mask image are clear and coherent.

[0116] In order to further improve the accuracy of histopathological region segmentation and refine the tissue contour, the obtained coarse-grained mask image is input into a fine-grained refinement network based on a diffusion model, and finally a pathological structure image with clear contour boundaries and complete internal details is output. In the second granularity segmentation stage, a reverse diffusion processing mechanism based on the SegRefiner diffusion model is introduced, combined with morphological operations and contour refinement post-processing techniques. By modeling the noise reduction process, through an iterative process of multi-round noise prediction and mask update, the boundary blur and noise interference generated by the coarse segmentation are gradually eliminated, so as to achieve the effect of extracting and strengthening the local details of the pathological tissue structure from the coarse-grained mask information, significantly improving the overall quality and practicality of the segmentation in pathological images. Compared with traditional single-stage segmentation methods, this technology can effectively restore fine structural features such as tiny blood vessels and cell gaps in complex pathological images, solve the problem of burr-like artifacts and detail loss easily generated at tissue junctions in existing methods, and improve the accuracy of microscopic structures such as the morphology of capillary branches and the boundaries of inflammatory infiltration areas.

[0117] The above are only examples of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the scope of the application of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-granularity progressive optimization method for molecular image segmentation based on a large model, characterized in that, It includes the following steps: Step S1: Obtain the pathological tissue image of the patient; Step S2: Input the pathological tissue image into the SAM model for the first granularity segmentation to generate a preliminary coarse-grained mask image; Step S3: Input the obtained coarse-grained mask image into the fine-grained refinement network based on the diffusion model for the second granularity segmentation, and finally output the molecular imaging segmentation result with different tissues.

2. The multi-granularity progressive optimization molecular imaging segmentation method based on a large model according to claim 1, wherein The tissue types for the first granularity segmentation and the second granularity segmentation include: inflammation, thrombus, edema, and capillaries.

3. The multi-granularity progressive optimization molecular imaging segmentation method based on a large model according to claim 1, characterized in that, In step S2, the process of performing the first granularity segmentation on the pathological tissue image by the SAM model includes the following steps: Step S21: Input the pathological tissue image to be segmented into the SAM model, and the SAM model includes an image encoder, a prompt encoder, and a mask decoder; Step S22: Process the input pathological tissue image through the image encoder to convert it into an image feature vector; Step S23: According to the segmentation requirements, input prompt information on the pathological tissue image, and the prompt forms include point prompts, box prompts, and text prompts describing the target; Step S24: Convert the input prompt information into a prompt vector through the prompt encoder; Step S25: Fuse the image feature vector and the prompt vector through the mask decoder, and decode to generate a preliminary coarse-grained mask image.

4. A multi-granularity progressive optimization molecular imaging segmentation method based on a large model according to claim 3, characterized in that, In step S22, the image encoder uses the ViT model to process the pathological tissue image and convert it into an image feature vector. The processing process includes the following steps: Step S221: Input the pathological tissue image into the ViT model in the image encoder for convolution-based image patch embedding operation to obtain a feature map; Step S222: Add position encoding to the obtained feature map; Step S223: The feature map with embedded position encoding is subjected to feature extraction through multiple Transformer blocks; Step S224: Reduce the dimension of the processed feature map through multi-layer convolution operations, and finally obtain an image feature vector.

5. A multi-granularity progressive optimization molecular imaging segmentation method based on a large model according to claim 4, wherein, In the process of performing step S25 to generate a preliminary coarse-grained mask image, first generate an output token and a mask token through the mask decoder to express the current prediction state of the model, and perform a concatenation operation on the output token and the mask token to obtain a fused feature vector. Then, input the fused feature vector, the image feature vector, and the prompt vector into the mask decoder for deep fusion and decoding, and finally generate a preliminary coarse-grained mask image.

6. A multi-granularity progressive optimization molecular imaging segmentation method based on a large model according to claim 4, characterized in that, After performing step S25, the generated preliminary coarse-grained mask image is evaluated by a hypernetwork for each mask in the predicted coarse-grained mask image, and the quality score of each mask is output.

7. A multi-granularity progressive optimization molecular imaging segmentation method based on a large model according to claim 1, characterized in that In step S3, the rough mask image generated by the SAM model in step S2 is subjected to the second granularity segmentation by using the SegRefiner diffusion model.

8. A multi-granularity progressive optimization molecular imaging segmentation method based on a large model according to claim 7, characterized in that, Using the SegRefiner diffusion model to complete the second granularity segmentation of the coarse-grained mask image in step S3 includes the following steps: Step 31: Input the coarsened mask image obtained in step S2 into the SegRefiner diffusion model. Introduce a random noise vector through the forward diffusion process of the SegRefiner diffusion model to simulate the initial noise state, and obtain a noisy mask image as the training data of the SegRefiner diffusion model; Step S32: Input the noisy mask image into the SegRefiner diffusion model. Through the reverse diffusion process of the SegRefiner diffusion model, learn to restore the noisy mask image to the coarsened mask image, and obtain a trained SegRefiner diffusion model; Step S33: Input the coarsened mask image obtained in step S2 into the trained SegRefiner diffusion model. Through the reverse diffusion process, perform a second granularity segmentation on the rough mask image, and finally output the molecular image segmentation result with different tissues.

9. A multi-granularity progressive optimization molecular imaging segmentation method based on a large model according to claim 8, characterized in that, The SegRefiner diffusion model realizes the reconstruction of a one-dimensional noise binary sequence by introducing binary diffusion. The expression formulas for the forward diffusion process and the reverse diffusion process of the SegRefiner diffusion model are: q(x t |x t-1 ) = B(x t |x t-1 (1 - β t ) + 0.5β t ) p θ (x t |x t-1 ) = B(x t |f b (x t , t)) where q(x t |x t-1 ) represents the transition probability of the forward diffusion process, x t and x t-1 represent the state or value at time step t or t - 1 respectively, taking two values in the binary diffusion model: 0 or 1; B(x t |x t-1 (1 - β t ) + 0.5β t ) is the specific function of B(x t |μ), B(x t |μ) is the probability mass function of the Bernoulli distribution, where μ is the success probability, that is, the probability that x t = 1. For a given μ, the probability mass function of the Bernoulli distribution is: B(x t |μ) = μx t (1 - μ)-x t ; β t is the hyperparameter used at time step t in the forward diffusion process, belonging to the interval (0, 1); p θ (x t |x t-1 ) represents the transition probability of the reverse diffusion process, f b (x t , t) is the model for predicting the Bernoulli probability, used to estimate the probability that x t = 1 at time step t; θ is the model adjustment parameter.

10. A multi-granularity progressive optimization molecular imaging segmentation method based on a large model according to claim 8, characterized in that For the reverse diffusion process of the SegRefiner diffusion model in step 32, it includes the following steps: Step S321: Noise prediction: Use the SegRefiner diffusion model to predict the noise part to be removed in the next stage according to the current noisy mask image; Step S322: Update the mask image: Based on the prediction result, adjust the current mask image and remove the predicted noise components; Step S323: Repeat the above steps S321 and S322 until the preset number of iterations is reached and then stop, and output the trained SegRefiner diffusion model.