Pathological image instance segmentation network framework based on diffusion model and region converter

By introducing a diffusion model and region transformer into the pathological image instance segmentation network, the problems of high-resolution processing difficulties and poor generalization of the model in the prior art are solved, and the instance segmentation effect with high precision and low computational complexity is achieved.

CN120182602APending Publication Date: 2025-06-20熊瞻
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510302888.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art faces problems such as difficulty in high-resolution image processing, insufficient generation of candidate areas without anchor points, high computational complexity, and loss of fine-grained information in the segmentation of pathological image instances, resulting in poor generalization and underfitting of the model.

Method used

The pathological image instance segmentation network framework based on diffusion model and region transformer is adopted, and the anchor-free noise candidate areas are generated in the target classification and target box regression network by introducing the diffusion model, and combined with the equalization sampling algorithm and the attention mechanism of the region transformer, adaptive feature extraction and noise removal are achieved.

Benefits of technology

It improves the generalization and segmentation accuracy of the network framework, avoids underfitting, can efficiently process pathological images with different stainings, and extracts robust example features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182602A_ABST
    Figure CN120182602A_ABST
Patent Text Reader

Abstract

The invention discloses a pathological image instance segmentation network framework based on a diffusion model and a region converter, which is high in segmentation precision and fast in calculation speed, and can perform instance segmentation on various types of organization structures in different stained pathological images. The method comprises a residual error backbone network, an FPN network, a target classification network, a target frame regression network and an instance segmentation network. Wherein a diffusion model is introduced into the target classification network and the target frame regression network, an anchor-point-free noise candidate area can be generated, generalization of a network framework is improved, meanwhile, a balanced sampling algorithm is combined, feature expressions of rare instances can be mined from common instances in a dominant position, and the underfitting phenomenon is avoided; a region converter is introduced into an instance segmentation network, global context information and local detail information can be obtained based on an attention mechanism of the region converter, noise is gradually removed by adopting adaptive feature fusion, and instance features with robustness and rich semantics are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a pathological image instance segmentation network framework based on a diffusion model and a region transformer. Background Art

[0002] Machine vision is one of the important technical branches of the artificial intelligence industry, and its core technologies include semantic segmentation, object detection, and instance segmentation. Among them, instance segmentation gives both semantic class labels and instance labels to each pixel point in an image, so that a pixel-level instance mask can be synthesized to locate different instances in different category regions of the image. The purpose of instance segmentation is to obtain information about each individual in the image and extract multi-level interaction information, such as correlation information at the semantic level and instance level, high-level environmental interaction, and global contextual features, so as to better analyze the scene information in the image. It is a basic component of practical applications such as image retrieval, pathological diagnosis, autonomous driving, video surveillance, and robotics, and is one of the important directions to improve productivity. Currently, in the actual application scenario of pathological diagnosis, the input pathological images are usually high-resolution images, with uneven numbers of instances of different categories, and large variations in the scales (shapes or sizes) or texture features of instances of the same category, making the instance feature expression based on deep learning still face major challenges.

[0003] At the present stage, there are mainly two ideas in the research of instance segmentation. The first is the top-down method based on object detection, which simplifies the instance segmentation problem into three serial sub-problems: first, use object detection to decompose the image into regions of different instances, then predict the binary instance label of each pixel in the region, and finally predict the semantic category of the object represented by the instance mask. The second is the bottom-up method based on semantic segmentation, which describes instance segmentation as two serial sub-problems: first, use semantic segmentation to decompose the image into regions of different categories, and then use methods such as clustering to find the boundary information of different instances in the region.

[0004] The bottom-up semantic segmentation-based network directly extracts the features of pixel points, which can retain the detailed information of scene objects, but has the following disadvantages: (a) The number of points processed each time is very limited, such as hundreds of thousands of pixel points (512x512). However, real scenes usually contain millions of points (1024x1024), which means that the semantic segmentation-based network can only process a part of the input image, thus seriously affecting the expression of global features and context information of large-scale instances (such as large blood vessels); (b) When extracting pixel-level features, it is difficult to ensure the balance of the number between instances of different categories, which easily leads to the loss of information of instances that do not often appear in the scene; (c) In practical applications, traditional methods are usually used to extract instance boundary information, resulting in the sub-optimality of the bottom-up method network. Therefore, the current mainstream instance segmentation methods all adopt a top-down strategy. They first use the object detection module to crop the regions that potentially contain instances from the original image, which can ignore a large amount of meaningless background information, so that high-resolution images can be directly processed and the global context information of instances at different scales can be extracted; then, an efficient sampling method is used to generate category-balanced samples from the extracted potential regions to ensure that the network does not ignore rare instances; finally, the region category recognition and binary instance mask prediction sub-modules can be integrated into the object detection network end-to-end, so that the optimal network structure can be obtained by using the optimization algorithm of deep learning. However, the object detection-based network uses predefined anchor boxes, making the potential region extraction network independent of the input data, resulting in a serious reduction in the generalization of the network; secondly, due to limited computing resources, the network must use downsampling to reduce the computational complexity, which will cause serious loss of fine-grained information of small target instances in the scene. In addition, by cascading multiple convolutional kernels to obtain a large enough receptive field, the extraction of global context information becomes inefficient and rough, reducing the accuracy of extracting non-local features of large target instances. Finally, the inefficient sampling algorithm may lead to unbalanced samples between instances of different categories, making the model tend to learn the dominant instances (i.e., the most frequently occurring instances), resulting in under-fitting of the model. Summary of the Invention

[0005] The present invention provides a pathological image instance segmentation network framework based on a diffusion model and a region transformer to overcome the above problems existing in the prior art. The pathological image instance segmentation network framework based on the diffusion model and the region transformer of the present invention can perform instance segmentation on various types of tissue structures in pathological images with different stainings, and has outstanding characteristics such as high accuracy, fast calculation, and small application scenario limitations; it includes a residual backbone network, an FPN network, a target classification network, a target box regression network, and an instance segmentation network. A diffusion model is introduced into the target classification network and the target box regression network, which can generate anchor-free noise candidate regions, improve the generalization of the network framework, and at the same time, combined with an equalization sampling algorithm, it can mine the feature expressions of rare instances from common instances that are in the dominant position, avoiding the occurrence of underfitting; a region transformer is introduced into the instance segmentation network, which can obtain global context information and local detail information based on the attention mechanism of the region transformer, and use adaptive feature fusion to gradually remove noise to obtain robust and semantically rich instance features.

[0006] The technical solution of this application is as follows:

[0007] A pathological image instance segmentation network framework based on a diffusion model and a region transformer, including a residual backbone network, an FPN network, a target classification network, a target box regression network, and an instance segmentation network; the target classification network and the target box regression network are implemented based on a diffusion model, and the instance segmentation network is implemented based on a region transformer; the instance segmentation network framework is obtained by training with historical pathological images; the training process is as follows: S1. Calibrate the pixel-level contour, category information, and target box position of the target to be segmented in the historical pathological image to obtain the original pathological label image data; then, perform preprocessing of image enhancement on the historical pathological image; S2. Input the preprocessed pathological image data D into the instance segmentation network framework, and the backbone network and the FPN network perform feature extraction to obtain a multi-scale feature list P; S3. Randomly generate noise candidate regions through the diffusion model, crop the multi-scale feature list P to obtain a region query Q, and then input it into the region transformer, and remove noise through the attention mechanism; S4. Randomly select according to categories in the denoised region query R to generate balanced samples, and input them into the target classification network, the target box regression network, and the instance segmentation network to obtain the prediction results of classification, bounding box regression, and instance segmentation masks; S5. Calculate the loss function between the prediction results in step S4 and the original pathological label image data in step S1, and then use the stochastic gradient descent method for backpropagation to update the parameters of the instance segmentation network framework until the error between the prediction result and the true label value meets the requirements, and the training is completed.

[0008] Compared with the prior art, the instance segmentation network framework for pathological images based on the diffusion model and the region transformer introduces three new mechanisms according to the top-down object detection-based method, and has high segmentation accuracy and generalization: (1) A diffusion model is introduced into the object classification network and the object box regression network, which can adaptively generate anchor-free noise candidate regions according to the input data; (2) Balanced sampling is performed between instances of different classes, so that the network framework can mine the feature expressions of rare instances from the dominant common instances, avoiding the occurrence of under-fitting; (3) A region transformer is introduced into the instance segmentation network, which can realize the adaptive fusion and complementation of fine-grained and coarse-grained information based on the cross-attention mechanism of the region transformer, so as to gradually remove noise, obtain robust and rich fine-grained information and non-local context information, and provide support for efficiently extracting multi-scale features for target instances of different sizes. The instance segmentation network framework of the present application can perform instance segmentation on various category organizational structures in pathological images with different stainings, and has outstanding characteristics such as high accuracy, fast calculation, and small application scenario limitations.

[0009] As an optimization, in the aforementioned instance segmentation network framework for pathological images based on the diffusion model and the region transformer, in step S1, the preprocessing process of image enhancement is specifically as follows: First, multiple regions of interest are selected in the pathological image and classified according to the pathological labels of the tissues in the image to generate corresponding RoI pathological image data; then, the RoI image data is regionally cropped according to random scale parameters; then, the cropped image data is randomly flipped horizontally or vertically; finally, contrast transformation, HSV color space transformation (specifically: keeping the hue H unchanged, performing exponential operations on the saturation S and brightness V components), and random Gaussian noise are performed on the flipped pathological image data. By strengthening specific regions in the image through the above steps, high-quality input data can be provided for the instance segmentation network framework, which is beneficial to the smooth progress of subsequent image segmentation tasks.

[0010] As an optimization, in the aforementioned instance segmentation network framework for pathological images based on the diffusion model and the region transformer, in step S2, first, the preprocessed pathological image data D is subjected to feature extraction by the residual backbone network to obtain a multi-scale feature list F; specifically: The preprocessed pathological image data D is input into the residual backbone network and output after passing through stage 0, stage 1, stage 2, stage 3, and stage 4 in sequence, and feature maps F1, F2, F3, F4, and F5 are obtained; the length and width of F1 are 1 / 4 of D, the length and width of F2 are 1 / 2 of F1, the length and width of F3 are 1 / 2 of F2, the length and width of F4 are 1 / 2 of F3, and the length and width of F5 are 1 / 2 of F4; the multi-scale feature list Among them, stage 0 includes a convolutional layer and a pooling layer, and stages 1 to 4 include a residual module, a normalization module, a convolutional layer, and a pooling layer; the residual module is composed of a convolutional layer with a 1×1 convolutional kernel and a convolutional layer with a 3×3 convolutional kernel.

[0011] Then, the feature list F is input into the FPN network, and an enhanced multi-scale feature list P is generated through the top-down path, the bottom-up path, and the lateral connection. The FPN network can achieve a balance between accuracy and efficiency in multi-scale object detection through a hierarchical feature fusion mechanism, which is beneficial to improving the accuracy and speed of feature extraction.

[0012] As an optimization, in the foregoing pathological image instance segmentation network framework based on a diffusion model and a region transformer, the diffusion model is divided into a forward diffusion process and a reverse diffusion process; in step S3, first, in the forward diffusion process, Gaussian noise is gradually added t times (1≤t≤1000) to the true bounding box to obtain a random noise candidate region ; then, in the reverse diffusion process, the multi-scale feature list P is cropped using the random noise candidate region step to obtain the region query Q; then, the attention weight is calculated between the multi-scale feature list P and the region query Q through the attention mechanism to denoise Q. The diffusion model generates new image data by gradually adding noise in the forward diffusion process and removing noise in the reverse diffusion process, and the overall process is stable and the image generation quality is high.

[0013] As an optimization, in the foregoing pathological image instance segmentation network framework based on a diffusion model and a region transformer, the calculation formula of the loss function is: , is the classification loss, is the target detection bounding box regression loss, is the instance segmentation loss.

[0014] ; is the normalization weight, ; is the probability of predicting the k-th candidate target region as a target, ;

[0015] , ; is the normalization weight, is the loss function weight, is the coordinate vector of the true target annotation box, , representing the 4 parameterized coordinates of the detected target annotation box;

[0016] ; is the predicted binary mask, and is the ground truth binary mask.

[0017] As an optimization, in the above-mentioned pathological image instance segmentation network framework based on the diffusion model and the region transformer, in step S5, the backpropagation process is as follows: starting from the output layer, the gradient of the loss function with respect to each weight is calculated layer by layer forward through the chain rule and passed to the previous layer for parameter update. In this application, using the chain rule for backpropagation can simplify the gradient calculation and improve the calculation efficiency.

[0018] Furthermore, the formula for parameter update is: ; where is the learning rate, and G is the gradient value calculated in the previous layer.

[0019] As an optimization, in the above-mentioned pathological image instance segmentation network framework based on the diffusion model and the region transformer, in step S5, if the error between the prediction result and the ground truth label value does not meet the requirements, the same number of preprocessed pathological image data is input into the instance segmentation network framework for processing again, and steps S2 - S5 are repeated to iteratively update the parameters of the instance segmentation network framework until the maximum number of iterations. Through multiple iterations, the instance segmentation network framework gradually optimizes its parameters, thereby improving the generalization ability of the network framework on the data and reducing subsequent segmentation errors.

[0020] As an optimization, in the above-mentioned pathological image instance segmentation network framework based on the diffusion model and the region transformer, in step S1, the historical pathological images are divided into a training set, a validation set, and a test set; when training the instance segmentation network framework, the training set is used to train the instance segmentation network framework, and at the same time, the validation set is used to detect the performance change curve during the training process. When overfitting is detected, the training is stopped, the hyperparameters of the instance segmentation network framework are adjusted, and then the training is restarted; after the training is completed, the test set is used to verify the performance and generality of the trained instance segmentation network framework. At this time, the acquisition of the instance segmentation network framework not only depends on its training effect on the historical pathological image dataset, but also needs to supervise the training process through the validation set and verify the training results through the test set, thereby ensuring the generalization ability and segmentation accuracy of the network framework. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is the training flow chart of the pathological image instance segmentation network framework in this application;

[0022] Figure 2 is the schematic diagram of the pathological image instance segmentation network framework in this application;

[0023] Figure 3 It is a schematic diagram of the target classification network and the target box regression network in this application;

[0024] Figure 4 It is a schematic diagram of the instance segmentation network in this application;

[0025] Figure 5 It is a curve graph showing the performance change during the training of the pathological image instance segmentation network in this application;

[0026] Figure 6 It is the intermediate and final result diagrams when using the pathological image instance segmentation network framework of the present invention to perform instance segmentation on kidney pathological images. Specific implementation manners

[0027] The following further illustrates this application in conjunction with the drawings and embodiments, but it does not serve as a basis for limiting this application. The content not detailed in the following embodiments is all common technical knowledge in the art.

[0028] In order to solve the problem that the current instance segmentation method cannot adaptively generate anchor boxes according to the input data, resulting in poor model generalization ability, and in the case of limited computing resources, resulting in the loss of fine-grained information of small target instances in the scene and rough non-local features of the extracted large target instances, the present invention designs a pathological image instance segmentation network framework based on diffusion models (Diffusion Models) and region transformers (Transformer) by introducing three new mechanisms (diffusion model, attention mechanism, balanced sampling) in the top-down method based on object detection. It has good generalization performance and can perform instance segmentation on various category organizational structures in pathological images with different stainings, and has the characteristics of high segmentation accuracy, fast calculation, and small application scenario limitations.

[0029] Embodiment:

[0030] Refer to Figures 1 to 4 , the pathological image instance segmentation network framework of this application based on diffusion models and region transformers includes a residual backbone network, an FPN network, a target classification network, a target box regression network, and an instance segmentation network; the target classification network and the target box regression network are implemented based on diffusion models, and the instance segmentation network is implemented based on region transformers; the instance segmentation network framework is trained through historical pathological images; the training process is as follows.

[0031] S1. Calibrate the pixel-level contours, category information, and target box positions of the required segmentation targets in the historical pathological images to obtain the original pathological label image data (i.e., the true label values).

[0032] Then, the historical pathological images are randomly divided into a training set (80%), a validation set (10%), and a test set (10%) according to the ratio of 8:1:1, ensuring no information leakage between the three data sets.

[0033] Then, perform preprocessing of image enhancement on the training set: (1) Select multiple regions of interest (RoIs) on the pathological images (WSIs), and classify them according to the pathological labels of the tissues in the images to generate corresponding RoI pathological image data; (2) Crop the RoI pathological image data according to random scale parameters; (3) Randomly flip the cropped pathological image data horizontally or vertically; (4) Perform three pixel-level transformations on the flipped pathological image data, namely contrast transformation, HSV color space transformation (keeping the hue H unchanged and performing exponential operations on the saturation S and brightness V components), and random Gaussian noise, to obtain the pathological image data D.

[0034] S2. Input the pathological image data D into the instance segmentation network framework. First, use the residual backbone network to extract features from the preprocessed pathological image data D to obtain a multi-scale feature list. The data processing flow of the residual backbone network is: input -> stage 0 -> stage 1 -> stage 2 -> stage 3 -> stage 4 -> output; among them, only convolutional layers and pooling layers are used in stage 0, while stages 1 to 4 adopt a similar structure composed of residual modules, normalization modules, convolutional layers, and pooling layers; among them, stages 1 and 4 include 3 residual modules, stage 2 includes 4 residual modules, and stage 3 includes 6 residual modules; each residual module is composed of a convolutional layer with a convolution kernel of 1×1 and a convolutional layer with a convolution kernel of 3×3 connected in series.

[0035] The specific process of feature extraction is as follows:

[0036] In the first step, input the pathological image data D into stage 0. First, pass through a convolutional layer with 64 convolution kernels of 7x7 and a stride of 2. The number of channels of the output feature map is 64; then pass through the BN layer (Batch Normalization) and the activation function RELU; finally, perform a max-pooling operation with a size of 3x3 and a stride of 2 to obtain the intermediate feature map F1 (the length and width are 1 / 4 of D).

[0037] In the second step, input the feature map F1 into stage 1. First, use 64 1x1 convolution kernels to reduce (dimensionality reduction) the number of channels of the input feature map, thereby reducing the computational amount. Then, extract local features through 64 3x3 convolution kernels. Finally, use 64 1x1 convolution kernels to restore the number of channels to the original number of channels (dimensionality increase) to obtain the residual R1; add F1 and R1, and then pass through the RELU activation function and a 2x2 pooling operation to obtain the feature map F2 (the length and width are 1 / 2 of F1).

[0038] In the third step, similar to the second step, input the feature map F2 into stage 2 to obtain the feature map F3; input the feature map F3 into stage 3 to obtain the feature map F4; input the feature map F4 into stage 4 to obtain the feature map F5; the length and width of F3 are 1 / 2 of F2, the length and width of F4 are 1 / 2 of F3, and the length and width of F5 are 1 / 2 of F4.

[0039] Then, input the feature list into the FPN network, and generate an enhanced multi-scale feature list through the top-down pathway, bottom-up pathway, and lateral connections .

[0040] S3. Use the forward process of the diffusion model to gradually add t times of Gaussian noise (1 ≤ t ≤ 1000) to the real bounding boxes (Ground-Truth boxes) to obtain random noise candidate regions ; then, in the reverse process, use the noise candidate regions to crop the multi-scale feature list P to obtain the region query , and input it into the region transformer; calculate the attention weights between the multi-scale feature list P and the region query Q through the attention mechanism, and denoise Q to obtain the region query adaptively generated according to the data.

[0041] S4. Randomly select by category in the denoised region query R to generate balanced samples , (randomly sample the samples of each category to an equal number through the balanced sampling algorithm; thus, the imbalance problem between different category targets can be balanced, ensuring that each category has enough candidate regions to participate in the subsequent result prediction process), and input them into the object classification network, object box regression network, and instance segmentation network to obtain the prediction results of classification, bounding box regression, and instance segmentation mask.

[0042] S5. Calculate the loss function between the prediction results and the original pathological label image data, and the calculation formula is: , where is the classification loss, is the object detection bounding box regression loss, and is the instance segmentation loss.

[0043] ; is the normalized weight, (cross entropy loss function); is the probability of predicting the k-th candidate target region as the target, ;

[0044] , (smoothL1 function); is the normalized weight, is the weight of the loss function, is the coordinate vector of the true target annotation box, , representing the 4 parameterized coordinates of the detected target annotation box;

[0045] ; is the predicted binary mask, is the true binary mask.

[0046] Then, the backpropagation is performed using the stochastic gradient descent method to continuously update the parameters of the instance segmentation network framework until the error between the prediction result and the true label value meets the requirements.

[0047] If the error between the prediction result and the true label value does not meet the requirements, the same number of preprocessed pathological image data is input into the instance segmentation network framework again for processing, and steps S2 - S5 are repeated to iteratively update the parameters of the instance segmentation network framework until the maximum number of iterations, and the training is completed.

[0048] The update formula is: , is the learning rate, and G is the gradient value calculated in the previous layer.

[0049] The backpropagation calculates the gradient layer by layer from the output layer to the input layer through the chain rule, including

[0050] Output layer gradient: the gradient of the loss with respect to the output, and the formula is ( );

[0051] Feed-forward network (FFN) gradient: calculate the gradient of the loss with respect to the FFN weights and biases;

[0052] Multi-Head Attention gradient: calculate the gradient of the loss with respect to the attention weights (such as the matrices of query Q, key K, and value V);

[0053] LayerNorm gradient: calculate the gradient of the loss with respect to the normalization parameters (gamma and beta);

[0054] Embedding layer gradient: Calculate the gradient of the loss with respect to the weights of the word vectors.

[0055] During training, use the validation set to detect the performance change curve during the training process. Specifically, after training one epoch, let the current network framework run an evaluation operation on the validation set to obtain the intermediate performance results, and draw the performance change curve based on these results. The performance change curve is shown in Figure 5 , where the X-axis is the number of training epochs, and the Y-axis is the difference between the training accuracy and the validation accuracy, that is, the mAP (mean average precision). When overfitting is detected, stop training, adjust the hyperparameters of the network framework, and then retrain.

[0056] S6. Input the test set into the instance segmentation network framework after training is completed to obtain the average precision, required time for instance segmentation, and the final instance segmentation result, so as to verify the performance and generalization of the trained instance segmentation network framework. In this embodiment, compare the average precision, required time, and segmentation result with the State-of-the-Art (SOTA) algorithm. If it exceeds this algorithm, it proves that the trained network framework has better performance and generalization.

[0057] When applying the above-mentioned pathological image instance segmentation network framework based on the diffusion model and region transformer of the present application to pathological image instance segmentation, the instance segmentation method specifically includes the following steps: I. Perform preprocessing of image enhancement on the input pathological image (as described in step S1); II. Input the preprocessed pathological image data into the instance segmentation network framework of the present application to obtain the instance segmentation result. See Figure 2 , in this embodiment, the applicant uses the above instance segmentation method to perform instance segmentation on kidney pathological images to obtain the instance segmentation result. The intermediate and final result diagrams of the instance segmentation are as shown in Figure 6 . In addition, after comparison with other algorithms, the present application requires less time and has higher average precision.

[0058] The above general description of the invention involved in the present application and the description of its specific implementation manners should not be understood as a limitation on the technical solution of the invention. Those skilled in the art can, based on the disclosure of the present application, without departing from the constituent elements of the invention involved, add, subtract, or combine the disclosed technical features in the above general description or / and specific implementation manners (including embodiments) to form other technical solutions within the protection scope of the present application.

Claims

1. A network framework for pathological image instance segmentation based on diffusion model and region transformer, characterized by: It includes a residual backbone network, an FPN network, a target classification network, a target frame regression network and an instance segmentation network; the target classification network and the target frame regression network are implemented based on a diffusion model, and the instance segmentation network is implemented based on a region transformer; the instance segmentation network framework is obtained by training historical pathological images; The training process is as follows: S1. Calibrate the pixel-level contour, category information and target frame position of the target to be segmented in the historical pathological image to obtain the original pathological label image data; then, perform image enhancement preprocessing on the historical pathological image; S2, input the preprocessed pathological image data D into the instance segmentation network framework, and extract features by the residual backbone network and the FPN network to obtain a multi-scale feature list P; S3, randomly generate noise candidate regions through the diffusion model, crop the multi-scale feature list P, obtain the region query Q, and input it into the region transformer to remove noise through the attention mechanism; S4, randomly select in the denoised region query R by category to generate balanced samples, and input them into the target classification network, target box regression network and instance segmentation network to obtain the prediction results of classification, box regression and instance segmentation mask; S5. Calculate the loss function of the predicted result and the original pathological label image data, and then use the stochastic gradient descent method for back propagation to update the parameters of the instance segmentation network framework until the error between the predicted result and the true label value meets the requirements and the training is completed.

2. The pathological image instance segmentation network framework based on diffusion model and region transformer according to claim 1, characterized in that: In step S1, the preprocessing process of image enhancement is as follows: first, multiple regions of interest are selected in the pathological image, and they are classified according to the pathological labels of the tissues in the image to generate corresponding RoI pathological image data; then, the RoI pathological image data is regionally cropped according to a random scale parameter; then, the cropped pathological image data is randomly flipped in the horizontal or vertical direction; finally, the flipped pathological image data is subjected to contrast transformation, HSV color space transformation and random Gaussian noise.

3. The pathological image instance segmentation network framework based on diffusion model and region transformer according to claim 2, characterized in that: In the step S2, the preprocessed pathological image data D is firstly subjected to feature extraction through the residual backbone network to obtain a multi-scale feature list F; the feature list F is then input into the FPN network to generate an enhanced multi-scale feature list P through a top-down path, a bottom-up path and a lateral connection.

4. The pathological image instance segmentation network framework based on diffusion model and region transformer according to claim 3 is characterized in that: The preprocessed pathological image data D is input into the residual backbone network, and is output after going through stage 0, stage 1, stage 2, stage 3, and stage 4 in sequence, and feature maps F1, F2, F3, F4, and F5 are obtained; the length and width of F1 is 1 / 4 of D, the length and width of F2 is 1 / 2 of F1, the length and width of F3 is 1 / 2 of F2, the length and width of F4 is 1 / 2 of F3, and the length and width of F5 is 1 / 2 of F4; Multi-scale feature list ; Among them, stage 0 includes a convolution layer and a pooling layer, and stages 1 to 4 include a residual module, a normalization module, a convolution layer and a pooling layer; the residual module is composed of a convolution layer with a convolution kernel of 1×1 and a convolution layer with a convolution kernel of 3×3.

5. The pathological image instance segmentation network framework based on diffusion model and region transformer according to claim 3, characterized in that: The diffusion model is divided into a forward diffusion process and a reverse diffusion process; in step S3, first, in the forward diffusion process, the real boundary box Gradually add t times of Gaussian noise (1≤t≤1000) to obtain a random noise candidate area ; Then in the reverse diffusion process, the random noise candidate region step is used Crop the multi-scale feature list P to obtain the region query Q; Then, the attention weights are calculated between the multi-scale feature list P and the region query Q through the attention mechanism to denoise Q.

6. The pathological image instance segmentation network framework based on diffusion model and region transformer according to claim 1, characterized in that: In step S5, the loss function is calculated as: ; in, is the classification loss, ; is the normalized weight, ; To predict the probability of the kth candidate target area as the target, ; is the bounding box regression loss for object detection, , ; is the normalized weight, is the loss function weight, , represents the four parameterized coordinates of the detected target annotation box, The coordinate vector of the real target annotation box; is the instance segmentation loss, ; is the predicted binary mask, is the true binary mask.

7. The pathological image instance segmentation network framework based on diffusion model and region transformer according to claim 6, characterized in that: In step S5, the back propagation process is: starting from the output layer, the gradient of the loss function for each weight is calculated layer by layer through the chain rule, and passed to the previous layer.

8. The pathological image instance segmentation network framework based on diffusion model and region transformer according to claim 7, characterized in that: In step S5, the formula for parameter update is: ;in, is the learning rate, and G is the gradient value calculated from the previous layer.

9. The pathological image instance segmentation network framework based on diffusion model and region transformer according to claim 1, characterized in that: In step S5, if the error between the predicted result and the true label value does not meet the requirement, the same amount of preprocessed pathological image data is input into the instance segmentation network framework for processing again, and steps S2 to S5 are repeated to iteratively update the parameters of the instance segmentation network framework until the maximum number of iterations is reached.

10. The pathological image instance segmentation network framework based on diffusion model and region transformer according to claim 1, characterized in that: In the step S1, the historical pathological images are divided into a training set, a validation set and a test set; when the instance segmentation network framework is trained, the training set is used for training, and at the same time, the validation set is used to detect the performance change curve during the training process. When an overfitting phenomenon is detected, the training is stopped, and the hyperparameters of the instance segmentation network framework are adjusted and then retrained; After training is completed, the test set is used to verify the performance and versatility of the trained instance segmentation network framework.