Overlapped cervical cytoplasm region segmentation method based on deep learning and conditional diffusion model

By transforming overlapping cervical cell segmentation into a conditional diffusion model generation task, and using the non-overlapping portions to generate a complete cytoplasmic mask, the problem of overlapping cervical cytoplasm segmentation in existing technologies is solved, achieving more accurate and stable segmentation results and improving the accuracy and biological rationality of cervical cytoplasm segmentation.

CN120931686APending Publication Date: 2025-11-11WUHAN UNIV
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202511058983.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing methods for segmenting overlapping cervical cytoplasm have limited effectiveness in handling complex overlapping cell regions. They lack effective integration of prior knowledge of cervical cell morphology, fail to fully utilize the complementarity of frequency and spatial domain information, lack sampling strategies targeting the morphological characteristics of cervical cells, and suffer from insufficient overlapping cell samples, thus limiting the algorithm's generalization ability.

Method used

The task of segmenting overlapping cervical cells is transformed into a conditional diffusion model generation completion task. The visible non-overlapping parts of the cervical cytoplasm are used as conditions, and the corresponding overlapping parts are generated through the conditional diffusion model. A U-shaped network structure, a cervical cell morphology prior module, and a two-stream attention mechanism are constructed. An adaptive multi-scale combined loss function and a hierarchical classifier free-guided sampling strategy are designed to achieve the generation of complete cytoplasmic masks.

Benefits of technology

By effectively avoiding the difficulties of directly segmenting overlapping cytoplasmic regions and making full use of prior morphological knowledge of cervical cytoplasm, more accurate segmentation of overlapping cervical cytoplasm is achieved through image generation and completion problems. This improves segmentation accuracy and smoothness, conforms to biological morphological laws, and enhances the reliability and stability of the results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931686A_ABST
    Figure CN120931686A_ABST
Patent Text Reader

Abstract

The invention discloses an overlapped cervical cytoplasm region segmentation method based on deep learning and a conditional diffusion model, and relates to the technical field of artificial intelligence analysis of medical images. According to the method, accurate segmentation of the overlapped cytoplasm region in the cervical cell image is realized through a morphological prior guided conditional diffusion process. The method comprises the following steps: constructing a multi-scale cervical cytoplasm mask pair image; designing a cytoplasm specific data enhancement and preprocessing process; building a multi-branch cervical cell morphology perception condition diffusion network; using a self-adaptive multi-scale combination loss function to optimize training; and a hierarchical classifier is adopted to freely guide sampling for reasoning. According to the method, the frequency domain and space domain features are fused, a cellular morphology and statistics priori knowledge base is established, and a strategy of generating complete cytoplasm by adopting non-overlapped parts is adopted, so that the problem that the traditional method is difficult to segment in complex backgrounds and overlapped regions is successfully solved, and reliable technical support is provided for early screening of cervical cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence analysis technology for medical images, and in particular to a method for segmenting overlapping cervical cytoplasmic regions based on deep learning and conditional diffusion models. Background Technology

[0002] Cervical cancer is the fourth most common malignant tumor among women worldwide, and cervical cytology is an important means of screening for cervical cancer and precancerous lesions. In cytological analysis, the morphological characteristics of the cytoplasm are one of the key indicators for judging the degree of cellular abnormalities. Due to the natural distribution characteristics of cells during the staining and slide preparation process, there is often a large overlap between cells, making it difficult to clearly identify cytoplasmic boundaries. This has become a significant challenge for the development of automated cell morphology analysis.

[0003] Segmentation of overlapping cervical cytoplasm is a highly complex problem in medical image segmentation and one of the major challenges faced by cervical cancer-assisted screening and automated diagnostic systems. In traditional Pap smear or liquid-based cytology preparation processes, cervical cells often exhibit a multi-layered stacking distribution, forming a complex overlapping phenomenon. This overlap primarily occurs in the cervical cytoplasm, which often presents a semi-transparent, multi-layered stacked characteristic. This feature makes it difficult for traditional segmentation methods and existing deep learning-based methods to achieve ideal segmentation results for overlapping cervical cytoplasm, thus hindering the development of cervical cancer-assisted diagnosis and automated screening.

[0004] From a mathematical perspective, the problem of segmenting overlapping cervical cells can be expressed as follows: Given an image I∈RH×W×C containing multiple overlapping cervical cells, the goal is to segment each cervical cell i to obtain its complete cytoplasmic mask Mi∈{0,1}H×W, accurately segmenting the complete cytoplasmic boundary even in overlapping regions. Traditional machine learning and deep learning methods employ a direct segmentation strategy, using complex feature engineering, network structures, and morphological operations to distinguish overlapping cervical cells. This method can be represented as finding the corresponding mapping function: f:I→{M1,M2,...,M... n However, this method faces inherent difficulties because in overlapping cytoplasmic regions, it is difficult to determine the specific affiliation of overlapping pixels based solely on local image features, especially when there are many overlapping cells, complex overlapping structures, and semi-transparent cytoplasm. Even professional pathologists find it difficult to accurately segment overlapping cytoplasm manually.

[0005] Existing methods for cervical cytoplasm segmentation are mainly divided into traditional machine learning methods and deep learning methods. Traditional machine learning methods rely on prior knowledge of the image to extract features manually, and then combine various methods and stages for processing, including region-based methods (such as superpixels, mean shift, region growing, watershed algorithms, etc.), pixel intensity-based methods (such as thresholding, morphological operations, K-means clustering, etc.), and contour-based methods (such as active contouring, edge detection, level sets, etc.). Although these methods provide good interpretability, their segmentation effect in complex overlapping cell regions is limited.

[0006] In recent years, deep learning methods have made significant progress in the field of medical image segmentation. However, while semantic segmentation methods (such as U-Net and FCN) can classify image pixels into cytoplasm, nucleus, and background, they cannot effectively distinguish individual cell instances within overlapping regions. Instance segmentation methods (such as Mask R-CNN) address the issue of a single pixel potentially belonging to multiple cells by generating instance masks, making them more suitable for handling overlapping cytoplasmic regions. Furthermore, there are methods based on Generative Adversarial Networks (GANs), such as Cell-GAN, which achieve cell segmentation through adversarial training between the generator and discriminator. Recently, general-purpose segmentation models (such as SAM and Med-SAM) have also demonstrated potential for zero-shot segmentation of medical images.

[0007] However, existing segmentation techniques for overlapping cytoplasmic regions still have the following problems: (1) lack of a model architecture that effectively combines prior knowledge of cervical cell morphology; (2) failure to fully utilize the complementarity of frequency domain and spatial domain information; (3) insufficient modeling of the semi-transparent properties of overlapping regions; (4) lack of sampling strategies for cervical cell morphology characteristics; and (5) insufficient overlapping cell samples in existing datasets, which limits the generalization ability of the algorithm. Summary of the Invention

[0008] As a novel paradigm of generative models, diffusion models demonstrate superior image generation and completion capabilities by simulating a thermodynamic reverse diffusion process. Compared to traditional generative adversarial networks (GANs), diffusion models offer advantages such as stable training, high generation diversity, and consistent quality. Combining diffusion models with conditional generation allows the use of prior information to guide the image generation process, providing a new approach to solving the problem of cytoplasmic overlap segmentation.

[0009] This invention proposes an innovative problem transformation approach, converting the task of segmenting overlapping cervical cells into a conditional diffusion model generation and completion task. Specifically, it utilizes the visible non-overlapping portions of cervical cytoplasm as conditions, and generates the corresponding overlapping portions using a conditional diffusion model based on cervical cytoplasm morphology, thereby achieving complete segmentation of overlapping cervical cells. This transformed method can be represented as: g:Mpartial →M full , of which M partial It is the non-overlapping portion of the cervical cytoplasm, M full It is the generated complete cervical cytoplasmic mask portion.

[0010] The core advantage of this problem-solving approach lies in:

[0011] (1) It can effectively avoid the difficulty of directly segmenting overlapping cytoplasmic regions;

[0012] (2) Make full use of prior knowledge of cervical cell morphology;

[0013] (3) Transform the segmentation problem into an image generation and completion problem that the conditional diffusion model excels at.

[0014] Specifically, this is achieved through the following techniques. This invention provides a method for segmenting overlapping cervical cytoplasmic regions based on deep learning and a conditional diffusion model, comprising the following steps:

[0015] Construct images of cervical cytoplasmic mask pairs and divide them into training, validation, and test sets;

[0016] The data in the training set, validation set, and test set are augmented and standardized.

[0017] A conditional diffusion model is constructed, comprising a U-shaped network structure, a cervical cell morphology prior module, and a two-stream attention mechanism. The U-shaped network structure includes a feature extraction encoder, a morphology-aware intermediate block, a feature extraction decoder, and a final output layer. Adaptive instance normalization is performed to achieve conditional control and feature fusion. Features are adjusted based on the mask of non-overlapping parts to statistically analyze characteristics. Frequency domain and spatial domain features are fused.

[0018] Design an adaptive multi-scale combined loss function, including diffusion reconstruction loss, shape prior loss, multi-scale edge consistency loss, frequency domain structural similarity loss, morphological consistency loss, and overlapping region specificity loss; construct the total loss function;

[0019] Training the conditional diffusion model;

[0020] A hierarchical classifier-guided sampling strategy is used for inference to obtain the complete cytoplasmic mask.

[0021] Furthermore, methods for constructing cervical cytoplasmic mask pairs for images include:

[0022] From the pathological images and their annotation files of the ISBI 2015 cervical cell dataset, the complete cytoplasm mask of cervical cells is extracted, and the non-overlapping part mask is calculated; the non-overlapping part mask is used as the input condition of the conditional diffusion model, and the complete mask is used as the target mask of the conditional diffusion model.

[0023] The non-overlapping portion mask and the complete mask extracted from a single cervical cell are combined to form a multi-scale cervical cytoplasmic mask image; the image is then layered according to morphological complexity and degree of overlap; the layering results are divided into training set, validation set, and test set; cell samples of various degrees of overlap and morphological complexity are evenly distributed in the training set, validation set, and test set.

[0024] The mask pairs are uniformly processed to a size of 256×256 pixels.

[0025] Furthermore, the methods for augmenting and standardizing the data in the training set, validation set, and test set include:

[0026] The training set, validation set, and test set data are sequentially subjected to morphological deformation, random threshold binarization, rotation, flipping, shearing, scaling, adversarial morphological noise processing, and transparency simulation enhancement processing.

[0027] The enhanced mask image is then normalized. The normalization process includes grayscale conversion, normalization, and edge smoothing; a mask comparison sample library is then constructed.

[0028] Furthermore, the morphological deformation method includes: applying elastic transformation to cervical cells, with the deformation intensity parameter dynamically adjusted according to the morphological complexity and overlap of the cervical cell cytoplasm; the morphological complexity includes regular shapes, complex shapes, and irregular shapes; the overlap is divided into mild overlap, moderate overlap, and severe overlap.

[0029] Random thresholding binarization methods include: using a Gaussian distribution of N(0.5,0.05) to perform random thresholding binarization on an image with a morphologically deformed mask;

[0030] The methods of rotation, flipping, shearing and scaling include: randomly rotating the image of the mask that has been randomly thresholded and binarized from 0 to 360°, flipping it horizontally, flipping it vertically, performing random shearing transformation and micro-scaling;

[0031] Methods for handling adversarial morphological noise include: adding random morphological noise to the cytoplasmic edge region of an image that has been rotated, flipped, clipped, and scaled; using a morphological gradient algorithm to extract the cytoplasmic edge region; adding random noise with an amplitude of no more than 3 pixels to the edge region; and applying morphological opening and closing operations to make the noise have a natural transition effect.

[0032] Transparency simulation methods include calculating transparency blending using spatially varying alpha values ​​α(x,y) in the overlapping regions of cervical cell cytoplasm.

[0033] Furthermore, in the conditional diffusion model, the U-shaped network structure adopts a four-stage U-shaped structure; the feature extraction encoder and feature extraction decoder each contain four levels, each level including several residual blocks; each residual block contains two 3×3 convolutional layers, group normalization, and SiLU activation function;

[0034] The feature extraction encoder includes two residual blocks, temporal embedding, conditional fusion, multi-scale downsampling, and skip connections. Each multi-scale downsampling layer consists of two residual blocks and one downsampling layer. The number of channels in the feature extraction encoder is 64, 128, 256, and 512, respectively. The downsampling uses a convolution operation with a stride of 2.

[0035] The morphology-aware intermediate block includes four components: a multi-receptive-field dilated convolution module, a multi-head self-attention mechanism, a cross-attention module, and a channel attention mechanism. The multi-receptive-field dilated convolution module uses dilated convolutions with dilation rates of 1, 2, 4, and 8. The cross-attention module uses the features from the feature extraction encoder as a query and interacts with the features extracted by the U-shaped network structure. The channel attention mechanism uses a compressed excitation network to adjust the channel weights.

[0036] The feature extraction decoder includes four upsampling modules, instance normalization, residual connections, spatial-frequency domain fusion, feature fusion, and morphological guidance. Each upsampling module includes an upsampling layer and two residual blocks, which are fused with the corresponding encoder features through skip connections. Upsampling is performed using deconvolution.

[0037] The cervical cell morphology prior module includes a global morphology description generator, an edge consistency enhancer, a topology maintainer, and a morphological statistics prior encoder, which are used to extract and encode the morphological features, edge properties, topology, and shape statistics of the cytoplasm of cervical cells.

[0038] The dual-stream attention mechanism takes the feature map output by the U-shaped network structure as input, and after their respective feature extraction and enhancement, outputs fused features; it is used to capture morphological feature streams and texture feature streams, and integrates these two types of features through gating fusion.

[0039] Furthermore, the specific method of the dual-stream attention mechanism includes: taking the feature map output by the U-shaped network structure in the conditional diffusion model as input to obtain the morphological feature stream and the texture feature stream, respectively;

[0040] For the morphological feature flow pathway, multi-scale dilated convolution is used to extract morphological features; an edge-aware attention mechanism is used to enhance the learning of cytoplasmic boundary features; global branch-local branch feature integration is used to fuse morphological information at different scales; an elliptical shape prior enhancement module is used to introduce the enhanced and normalized training set mask into the image to an elliptical shape template set that conforms to the shape of cervical cells. Through deformable convolution and input features, the template set is adjusted to incorporate the prior knowledge of the elliptical shape of cervical cells into the feature extraction process, and finally, the morphological feature data of cervical cells is obtained.

[0041] For the texture feature flow pathway, multi-band texture features are extracted by frequency domain filtering; directional sensitive convolutional layers are used to capture directional texture features of the cytoplasm; nonlocal averaging mechanism is used to capture the dependencies of long-distance textures; and transparency-aware convolutional layers are used to process the semi-transparent properties of overlapping areas, ultimately obtaining texture feature data of cervical cells.

[0042] The gated feature fusion mechanism adjusts the weight ratio of morphological feature flow and texture feature flow through a gating function, adjusts the fusion parameters through the complexity and overlap of cell morphology, preserves the original cytoplasmic feature information through residual connections, and enhances the sensitivity to cytoplasmic overlapping regions through feature recalibration.

[0043] Furthermore, in the cervical cytoplasmic morphology prior module, the global morphological description generator is used to extract and encode the morphological features of the cervical cell cytoplasmic mask image after step S2 through global pooling and multilayer perceptron network.

[0044] The edge consistency enhancer is used to extract edge features at different scales, and the feature map is modulated through the edge response map to enhance the expression of the edge region;

[0045] The topology preserver is used to extract the skeleton of the mask through a morphological thinning algorithm, and the topology of the cytoplasm is preserved through skeleton-guided feature enhancement; distance transformation weighting is used to ensure that the weight distribution of the skeleton region conforms to the morphological characteristics of the cytoplasm.

[0046] Using the aforementioned morphological statistical prior encoder, principal component analysis of cytoplasmic shape is constructed based on the enhanced and standardized cervical cell dataset. Low-dimensional representation of the shape is extracted, and the shape prior encoding of cervical cells is fused with the feature map through deformable convolution.

[0047] Furthermore, in the adaptive multi-scale combined loss function, the diffusion loss uses weighted L2 loss to calculate the difference between the predicted noise and the actual noise at each time step, and its weight coefficients are adjusted according to the noise level and the location of the overlapping area.

[0048] The shape prior loss, combined with boundary-weighted binary cross-entropy loss and Tversky loss, strengthens the constraints on cytoplasmic shape.

[0049] The multi-scale edge consistency loss compares the edge features of the predicted mask and the real mask at multiple scales, especially the consistency of the boundary features of the overlapping region.

[0050] The frequency domain structural similarity loss evaluates the structural similarity between the predicted mask and the true mask in the frequency domain space;

[0051] The morphological consistency loss, combined with the modified Hausdorff distance metric loss and topological consistency loss, constrains the morphological rationality of the segmentation results by comparing morphological features.

[0052] The overlapping region-specific loss is designed for overlapping regions and includes a loss function, including mean square error loss and transparency consistency loss, to confirm whether the transparency transition of the predicted mask in the overlapping region is natural and conforms to optical characteristics.

[0053] Furthermore, the hierarchical classifier's free-guided sampling strategy includes: dividing the sampling process into a coarse stage, a refinement stage, and a fine stage; the guidance strength used in the coarse stage is ω=3.5; the guidance strength used in the refinement stage is ω=2.0; and the guidance strength used in the fine stage is ω=1.5.

[0054] Conditional and unconditional noise predictions are generated. Based on the guidance intensity of the current stage, the corresponding ω value is selected according to the current sampling stage, and the weighted combination of noise predictions is calculated. DDIM (Denoising Diffusion Implicit Model) is used for sampling to calculate the noise state of the next step. The known area is kept unchanged through masking operations, and the generation process is regularized by applying cervical cell cytoplasmic morphology constraints.

[0055] Based on the current noise level and the results of previous sampling steps, the conditional noise prediction in the conditional diffusion model is adjusted, and a special noise adjustment method is applied to the overlapping area to enhance the generation of the semi-transparent overlapping part.

[0056] Post-processing and rationality verification of the final results;

[0057] The post-processing includes binarization, conditional consistency correction, morphological post-processing, and morphological smoothing.

[0058] The rationality verification includes verification of shape parameters, topology, and the naturalness of transitions in overlapping regions, and constraint correction is performed when the verification falls within the unreasonable range.

[0059] The present invention also provides a system for segmenting overlapping cervical cytoplasmic regions based on deep learning and a conditional diffusion model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the steps of any of the methods described above.

[0060] The core innovation of this invention lies in transforming the problem of overlapping cervical cytoplasm segmentation into a conditional diffusion generation task. It predicts and generates a complete cytoplasmic mask using known non-overlapping cytoplasmic regions. This method avoids the difficulty of directly identifying and segmenting overlapping regions, offering a completely new perspective on solving the overlapping cell segmentation problem in the field of medical cell image segmentation.

[0061] This invention integrates the latest deep learning technologies and biomedical expertise, and specifically designs a multi-branch morphology-aware conditional diffusion network architecture for the task of overlapping cervical cytoplasm segmentation. This architecture captures morphological and textural features of the cytoplasm separately through a two-stream attention mechanism, encodes the shape statistics of cervical cytoplasm through a cytoplasmic morphology prior module, and fully utilizes complementary information from different domains through a frequency and spatial domain feature fusion module.

[0062] The present invention also provides a system for segmenting overlapping cervical cytoplasmic regions based on deep learning and conditional diffusion models, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the system is able to implement the steps of any of the methods described above.

[0063] Compared with the prior art, the advantages of the present invention are:

[0064] 1. This invention proposes a novel approach to solve the problem of segmenting overlapping cervical cytoplasm. Unlike traditional methods, this method starts from the visible non-overlapping cytoplasmic regions and uses a conditional diffusion model to generate a complete cytoplasmic mask, effectively overcoming the difficulty of directly segmenting overlapping regions.

[0065] 2. A multi-branch cervical cell morphology perception conditional diffusion network architecture was designed. Through a two-stream attention mechanism and a cytoplasmic morphology prior module, the morphological features and texture characteristics of cervical cytoplasm were effectively captured, making the generated complete cytoplasmic mask more consistent with biological morphological laws.

[0066] 3. It integrates spatial and frequency domain information and introduces a specific loss for overlapping regions. Through a dynamic weight adjustment mechanism, it adaptively balances the contributions of each loss component, thereby improving the model's accuracy and smoothness in segmenting the boundaries of overlapping regions.

[0067] 4. By using three-stage hierarchical sampling and adaptive noise adjustment, while ensuring global shape consistency, the generation quality of local details is enhanced, making the inference results more consistent with the natural morphology of cervical cytoplasm.

[0068] 5. A verification and correction process based on cell morphology characteristics was introduced to ensure that the segmentation results conform to the morphological characteristics and topological structure of cervical cytoplasm, thereby improving the reliability and stability of the results. Attached Figure Description

[0069] Figure 1 The flowchart illustrates the overall process of the overlapping cervical cytoplasmic region segmentation method based on deep learning and conditional diffusion models, as shown in this embodiment.

[0070] Figure 2 This is a diagram of a multi-branch morphological sensing conditional diffusion network architecture targeting cervical cells.

[0071] Figure 3 This is a structural diagram of the two-stream attention mechanism.

[0072] Figure 4 This is a structural diagram of the prior module for cytoplasmic morphology.

[0073] Figure 5 This is a schematic diagram of the adaptive multi-scale combined loss function.

[0074] Figure 6 Flowchart of the free-guided sampling strategy for hierarchical classifiers.

[0075] Figure 7 Here is an example image for input.

[0076] Figure 8 This is a schematic diagram of the output. Detailed Implementation

[0077] The technical solution of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0078] Example

[0079] The overlapping cervical cytoplasmic region segmentation method provided in this embodiment is based on deep learning and a conditional diffusion model. The overall process is as follows: Figure 1 As shown. In each step of the following method, "original features" refers to the features before the corresponding operation. Here, non-local features refer to the output feature map obtained after the complete calculation operation by the non-local averaging module of this section.

[0080] S1. Construct multi-scale cervical cytoplasmic mask pairs for images.

[0081] S11. Extract the complete cytoplasmic mask (M) of cervical cells from the pathological images and annotation files of the ISBI 2015 cervical cell dataset. full ), calculate to obtain the mask of the non-overlapping portion (M) partial ). Among them, the non-overlapping portion mask M partial Represents the non-overlapping region of the cytoplasm; complete mask M full It represents the complete morphology of the cytoplasm.

[0082] Specifically, the method for calculating the mask of the non-overlapping portion is as follows:

[0083] (1) According to the following formula, based on the existing annotation files of the ISBI 2015 cervical cell dataset, extract the overlapping region O between cervical cytoplasm. i .

[0084] .

[0085] Where i represents the i-th cell, j represents the j-th cell, and U j≠i This represents the non-overlapping union region of cells i and j.

[0086] (2) Generate the non-overlapping portion mask M according to the following formula. partial .

[0087] .

[0088] Among them, the non-overlapping portion mask As input condition, complete mask As a target mask.

[0089] S12. Combine the non-overlapping mask extracted from a single cervical cell with the complete mask to form a multi-scale cervical cytoplasm mask image; perform layering according to morphological complexity and degree of overlap; divide the layering results into training set, validation set and test set.

[0090] The non-overlapping portion mask and the complete mask together constitute the mask pair image. These training, validation, and test sets are used for the subsequent training, validation, and testing of the conditional diffusion model, ultimately predicting the mapping relationship of the complete cytoplasm from the known non-overlapping cytoplasmic portions.

[0091] (1) Layering according to morphological complexity

[0092] The main calculation is the degree of regularity of the cytoplasm morphology of each cervical cell; it is divided into regular shape (which can be approximately fitted as an ellipse), complex shape (which differs from the fitted ellipse), and irregular shape (which is highly irregular and cannot be fitted as an ellipse).

[0093] (2) Layering according to the degree of overlap

[0094] The main calculation involves determining the proportion of overlapping area to the total area of ​​intact cytoplasm in each cervical cell (overlap percentage). Based on the overlap percentage, it is categorized into mild overlap (overlap percentage <30%), moderate overlap (overlap percentage 30%-50%), and severe overlap (overlap percentage 50%-70%).

[0095] (3) By random sampling, the above-mentioned stratification results of morphological complexity and overlap are divided into training set, validation set and test set; and ensure that the distribution of cell samples with different overlap and morphological complexity in the training set, validation set and test set is balanced.

[0096] The final dataset contains 4783 masks for the training set, 540 masks for the validation set, and 270 masks for the test set.

[0097] S13. Center and standardize the images using masks from the training, validation, and test sets to ensure the integrity of boundary information.

[0098] Specifically, the centering and standardization process involves uniformly standardizing the mask pairs in the training, validation, and test sets to a size of 256×256 using distance-preserving interpolation, ensuring that cells are centered within the mask pair image. This method preserves the complete boundary information of the cytoplasm and avoids edge truncation.

[0099] S2, masks in the training set, validation set, and test set are used to enhance and standardize the image data.

[0100] S21. Enhance the masked images (i.e., masked images of cervical cell cytoplasm) in the training, validation, and test sets. The enhancement process includes, in sequence: morphological deformation, random thresholding binarization, rotation, flipping, shearing, scaling, adversarial morphological noise reduction, and transparency simulation.

[0101] (1) Morphological deformation: The elastic transformation of cervical cells is applied. The deformation intensity parameter is dynamically adjusted according to the complexity and overlap of the cytoplasmic morphology of cervical cells.

[0102] Specifically, based on the stratification results of cervical cell morphological complexity in step S12 above, a smaller deformation intensity parameter is used for cells with regular shapes (e.g., round or oval), and a larger deformation intensity parameter is used for cells with irregular shapes.

[0103] Specifically, based on the stratification results of the degree of overlap of cervical cells in step S12 above, for areas with severe overlap, the deformation strength parameter is reduced to maintain the authenticity of the overlap boundary.

[0104] Specifically, cytoplasmic morphological deformation α effective The calculation method is as follows:

[0105] ;

[0106] Where, α base This is the basic deformation strength parameter, typically set to a random value within the range of 5-10 pixels; w overlap It is the overlapping region weighting factor, linearly mapped to the range [0,1] based on the overlap ratio; f shape It is the shape complexity adjustment factor, with 0.7-0.9 for regular shapes, 0.9-1.1 for complex shapes, and 1.1-1.3 for irregular shapes.

[0107] (2) Random threshold binarization: Using Gaussian distribution sampling, random threshold binarization is performed on the mask image after morphological deformation to enhance the robustness of the subsequent conditional diffusion model to different binarization conditions and staining variations.

[0108] Specifically, this embodiment uses a Gaussian distribution sampling of N(0.5,0.05) and introduces an adaptive threshold selection based on the Otsu method.

[0109] The Gaussian distribution sampling formula is:

[0110] ;

[0111] Where t is the threshold parameter obtained by Gaussian distribution sampling, μ is the mean, and σ 2 Let Variance be the variance.

[0112] The Otsu adaptive threshold integration formula is:

[0113] ;

[0114] Among them, t otsu β is the optimal threshold calculated by the Otsu method; β is the mixed weighting coefficient, which is usually set to a random value between 0.3 and 0.7.

[0115] (3) Rotation, flipping, shearing and scaling: The mask that has been randomly thresholded and binarized is rotated randomly from 0 to 360°, and flipped horizontally and vertically.

[0116] The specific implementation method is as follows: ① Randomly determine the rotation angle and apply the rotation transformation matrix to perform the rotation; ② Randomly determine whether to perform horizontal flip and vertical flip.

[0117] To better simulate actual cell distribution, random shearing transformations (±10°) and small scaling factors (e.g., scaling by 0.9–1.1 times) were introduced. Specifically, the shearing transformations (±10°) and scaling factors (0.9–1.1) were applied randomly. All transformation parameters were randomly determined during the training of the subsequent conditional diffusion model to maximize data diversity.

[0118] (4) Adversarial morphological noise processing: Random morphological noise is added to the cytoplasmic edge regions in the image after rotation, flipping, shearing and scaling to simulate edge uncertainty in the staining and imaging process.

[0119] The specific implementation methods include: using morphological gradient algorithm operations to extract the edge regions of the cytoplasm of cervical cells in the cytoplasm mask image; adding random noise with an amplitude of no more than 3 pixels to the edge regions; and using morphological opening and closing operations to make the noise have a natural transition effect.

[0120] (5) Transparency simulation: In order to simulate the transparency change characteristics of the overlapping area of ​​cervical cell cytoplasm and enhance the ability of subsequent conditional diffusion model to perceive the semi-transparent overlapping area, the spatially varying Alpha value α(x,y) is applied to the overlapping area to make the overlapping boundary present natural optical properties.

[0121] The basic transparency blending model calculation formula is:

[0122] ;

[0123] Here, M1(x,y) and M2(x,y) represent the complete cytoplasm masks of two overlapping cervical cells, respectively.

[0124] For overlapping regions Spatially varying alpha mixing is employed. The formula for calculating the adaptive alpha value α(x,y) is:

[0125] ;

[0126] in, and These are the distances to the boundaries of the two cells, respectively. It is the disturbance factor (usually 0.05-0.1). It is a local noise function.

[0127] S22. Normalize the enhanced mask image. The normalization process includes grayscale conversion, normalization, and edge smoothing.

[0128] Normalization refers to normalizing to the range [0,1]. Edge smoothing is achieved by using a Gaussian filter with a radius of 0.5 pixels to reduce the impact of jagged edges and quantization errors.

[0129] S23. Construct a mask contrast sample library after data augmentation (simply reconstruct the mask pairs directly from the processed images) to ensure sample diversity and balance in the training process.

[0130] S3. Construct a conditional diffusion model.

[0131] Conditional diffusion model is a cell morphology model built on the basis of U-Net architecture, targeting the unique morphological characteristics of cervical cells and overlapping segmentation tasks.

[0132] The basic conditional diffusion model includes a morphological prior module, a two-stream attention mechanism, a U-shaped network structure, an adaptive multi-scale combined loss function, and a hierarchical classifier to achieve free-guided sampling. A schematic diagram of the specific structure is shown below. Figure 2 As shown.

[0133] S31. Establish a U-shaped network structure

[0134] The U-shaped network structure employs a four-stage U-shaped architecture, comprising a feature extraction encoder, a morphology-aware intermediate block, and a feature extraction decoder. The feature extraction encoder and feature extraction decoder each contain four layers (…). Figure 2 (Multi-scale downsampling in the model), each level includes multiple residual blocks; each residual block contains two 3×3 convolutional layers, group normalization, and SiLU activation function.

[0135] The mathematical expression for the residual block is: .

[0136] Where F is a nonlinear transformation function containing two 3×3 convolutional layers, group normalization, and SiLU activation, t is the temporal embedding, c is the conditional embedding, Proj is the projection layer when the number of input and output channels are not the same, and H... in These are input features.

[0137] (1) Feature extraction encoder: such as Figure 2As shown, the algorithm includes two residual blocks, temporal embedding, conditional fusion, multi-scale downsampling, and skip connections. Each multi-scale downsampling layer consists of two residual blocks and one downsampling layer. The number of channels in the feature extraction encoder is 64, 128, 256, and 512, respectively. Downsampling uses a convolution operation with a stride of 2.

[0138] Compared to max pooling, this structure of the feature extraction encoder preserves more feature information.

[0139] (2) Shape-sensing intermediate block ( Figure 2 (Not shown in the diagram): It includes four components: multi-receptive-field dilated convolution module, multi-head self-attention mechanism, cross-attention module, and channel attention mechanism.

[0140] ① Multi-receptive-field dilated convolution module: Uses dilated convolutions with dilation rates of 1, 2, 4 and 8 to capture multi-scale contextual information.

[0141] ② Multi-head self-attention mechanism: A multi-head self-attention mechanism is used to calculate the global dependencies between features.

[0142] The formula for calculating the multi-head self-attention mechanism is:

[0143] .

[0144] Where Q, K, and V represent the query, key, and value matrices, respectively, obtained from the input features through linear transformation; d k It is the dimension of the key vector. This is to alleviate the vanishing gradient problem; T represents matrix transpose; the softmax function normalizes the attention weights into a probability distribution.

[0145] The multi-head self-attention mechanism splits the input into multiple heads, calculates the attention of each head separately, and then concatenates the results. The calculation formula is as follows:

[0146] .

[0147] .

[0148] Where MultiHead represents multi-head self-attention mechanism, Concat represents concatenation operation, and head i W represents the output of the attention head for the i-th head, where Attention represents the computation of the attention mechanism. i Q W i K W i V and W O It is a learnable weight matrix.

[0149] ③ Cross-attention module: The features of the feature extraction encoder are used as queries and interact with the features extracted by the backbone network (U-shaped network structure).

[0150] ④ Channel attention mechanism: Use a compressed excitation network to dynamically adjust channel weights.

[0151] The compressed excitation network refers to the use of global average pooling to capture channel statistics, and then using a multilayer perceptron network to learn the interdependencies between channels and generate channel attention weights.

[0152] The method for dynamically adjusting channel weights is as follows: first, global average pooling is used to extract the global description of each channel; then, two fully connected networks (compression layer and activation layer) are used to learn the relationship between channels; finally, the sigmoid activation function is used to generate channel weights in the range of 0-1, and the original features (features before the channel attention mechanism) are weighted by channel.

[0153] (3) Feature extraction decoder: such as Figure 2 As shown, it includes four upsampling modules, instance normalization, residual connections, spatial-frequency domain fusion, feature fusion, and morphological guidance. Each upsampling module contains an upsampling layer and two residual blocks, which are fused with the corresponding encoder features via skip connections. Upsampling employs deconvolution operations. Compared to standard quadratic linear interpolation, this approach provides better feature recovery capabilities.

[0154] (4) Final output layer ( Figure 2 (Not shown in the diagram): It consists of one set of convolutional layers, which generate prediction noise and shape priors, and then apply the optimization processing module of the final output layer for optimization before the final output.

[0155] The optimization methods include edge smoothing, small connected region removal, and morphological constraints on the image obtained from the feature extraction decoder to ensure that the final output cytoplasmic mask has a reasonable shape, smooth boundaries, and a complete topological structure. Small connected regions are defined as those with an area less than 1% of the total area.

[0156] S32, Design of a priori module for cervical cell morphology

[0157] The cervical cell morphology prior module includes a global morphological description generator, an edge consistency enhancer, a topology maintainer, and a morphological statistical prior encoder; used to extract and encode the morphological features, edge properties, topology, and shape statistics of the cervical cell cytoplasm. For example... Figure 2 and 4 As shown, the specific method is as follows.

[0158] (1) Extraction and encoding of morphological features

[0159] Using global pooling and a multilayer perceptron network, morphological features of the cervical cell cytoplasm mask image after step S2 are extracted and encoded. These morphological features include global morphological attributes such as overall shape, size, and roundness, which help guide the generation of features representing the overall shape of the cytoplasm.

[0160] The specific steps include: compressing the obtained feature map (the processed masked image of S2) into a 1×1 global representation using adaptive average pooling; performing nonlinear transformation through a three-layer perceptron network; and using the SiLU activation function and Dropout regularization to improve generalization ability.

[0161] The three-layer perceptron network is:

[0162] ;

[0163] .

[0164] Where σ() is the SiLU activation function, W i Let F be the weight matrix of the corresponding layer, and b be the input features. i This is the bias term. Dropout regularization with probability p=0.2 is applied after the first and second layers.

[0165] (2) Enhanced edge consistency

[0166] Edge features at different scales are extracted to construct an edge response map; the feature map is modulated by the edge response map to enhance the expression of the edge region and ensure that the generated cytoplasmic boundary can be natural and coherent.

[0167] Specifically, the steps include: simultaneously using the Sobel edge detection operator and the Canny edge detection operator to extract edge features of cervical cell cytoplasm at different scales and sensitivities, calculating gradient magnitude and direction, and constructing a multi-channel edge response map.

[0168] ①Sobel edge detection operator: calculates the image gradient using convolution kernels in two directions as shown in the following formula.

[0169] ;

[0170] .

[0171] Where I is the input image, * denotes the convolution operation, and K x and K y These are the horizontal and vertical Sobel kernels, respectively.

[0172] Gradient calculation uses the Sobel kernel to calculate the gradient magnitude and direction.

[0173] ②Canny edge detection

[0174] First, Gaussian smoothing (using convolution kernels to filter the image) is performed, and gradient calculation is performed using the Sobel edge detection operator. Then, non-maximum suppression and double thresholding are performed. Finally, edge tracking and hysteresis are performed to connect strong edges with weak edges connected to the strong edges.

[0175] Specifically, Gaussian smoothing filters the image using a Gaussian kernel with a standard deviation of 0.5-1.5. Non-maximum suppression preserves local maxima along the gradient direction. Double thresholding uses two thresholds (the lower threshold is typically 1 / 2 or 1 / 3 of the higher threshold) to distinguish between strong edges, weak edges, and non-edges. Edge tracking preserves weak edges connected to strong edges and discards the rest.

[0176] The edge response map primarily modulates the feature map through an attention mechanism, and its calculation formula is as follows:

[0177] .

[0178] Among them, E norm This is a normalized edge response map, where λ is the enhancement coefficient (usually set to 0.5), and F is the original feature. edge These are the features after edge enhancement. This edge consistency enhancement mechanism ensures that the model can generate cytoplasmic masks with clear boundaries.

[0179] (3) Topology preservation

[0180] The cytoplasmic mask skeleton of cervical cells is extracted using a morphological thinning algorithm. Skeleton-guided feature enhancement is then employed to preserve the topological structure of the generated cytoplasm. Distance transform weighting is used to ensure that the weight distribution of the skeleton region conforms to the morphological characteristics of the cytoplasm.

[0181] Morphological thinning is an image processing technique based on mathematical morphology, and it is an iterative process. In this embodiment, the morphological learning algorithm maintains connectivity by gradually removing mask edge pixels, ultimately obtaining a skeleton 1 pixel wide. The specific steps of the morphological thinning algorithm include:

[0182] ① Initialize the label matrix and label all foreground pixels;

[0183] ② Iteratively execute the refinement steps, with each iteration including two sub-iterations; in each sub-iteration, check the 8-neighborhood pattern of each foreground pixel.

[0184] If a certain condition is met (i.e., maintaining connectivity and not being an endpoint), then the pixel is marked for deletion; after the iteration ends, all pixels marked for deletion are removed; finally, the above steps are repeated until no pixels are deleted.

[0185] The formula for calculating the distance transformation weighting is:

[0186] ;

[0187] Where D(x,y) is the distance from pixel (x,y) to the nearest skeleton point, and σ is a parameter that controls the weight decay rate (generally set to 1 / 3 of the average skeleton width). This topology preservation mechanism ensures that the generated cytoplasmic mask maintains a morphologically consistent topology with the real cytoplasm.

[0188] (4) Morphological statistical prior coding

[0189] Based on the cervical cell dataset processed in steps S1 and S2, principal component analysis (PCA) of cervical cell cytoplasmic shape is constructed to extract low-dimensional representations of the shape. The shape prior encoding of cervical cells is then fused with the feature map through deformable convolution.

[0190] Specifically, principal component analysis was performed on the cytoplasmic masks of all cervical cells in the entire cervical cell dataset obtained from S1 and S2 to extract the main shape patterns. The specific formula is expressed as follows:

[0191] .

[0192] Among them, E i It is the eigenvector of the shape, α i is the weighting coefficient, μ is the average shape, and k is the number of principal components used.

[0193] In this embodiment, deformable convolution is a convolution operation that adaptively adjusts the position of the convolution kernel sampling points based on input features; deformable convolution is used to fuse the prior knowledge of cervical cell shape with the feature map. The fusion formula is:

[0194] .

[0195] Where DCN represents deformable convolutional network operation; s prior It is a priori encoding of cervical cell shape, generated by principal component analysis (PCA).

[0196] S33. Design a dual-stream attention mechanism.

[0197] The dual-stream attention mechanism receives feature maps from the U-shaped network structure of the conditional diffusion model as input. After their respective feature extraction and enhancement, they output fused features for use by the subsequent decoder. For example... Figure 3 As shown, firstly, a two-stream attention mechanism is used to capture morphological feature streams and texture feature streams; then, these two types of features are integrated through gated fusion. The specific method is as follows.

[0198] (1) Feature input and feature separation

[0199] Using the feature map output by the U-shaped network structure in the conditional diffusion model as input, morphological feature flow and texture feature flow are obtained respectively.

[0200] (2) Morphological characteristics of the circulation pathway

[0201] The morphological feature stream focuses on extracting the shape, contour, and structural features of cervical cells. Its input is the 512-channel feature map output from the last layer of the encoder.

[0202] ① Multi-scale dilated convolution is used to extract morphological features.

[0203] Multiscale dilated convolution can expand the receptive field without increasing the number of parameters. Specifically, 3×3 convolution kernels with different dilation rates (1, 2, 4, 8) are used to extract multiscale morphological features of the cytoplasm.

[0204] ② Employ an edge-aware attention mechanism to enhance the learning of cytoplasmic boundary features.

[0205] Specifically, first, the Sobel edge detection operator is used to extract the edge feature map of the input features, and the gradients in the horizontal and vertical directions are calculated. The edge intensity map is obtained through the gradient magnitude. Then, the edge intensity map is normalized and used as the attention weight, which is then multiplied element-wise with the original features.

[0206] The edge-aware attention mechanism makes the conditional diffusion model pay more attention to the cytoplasmic boundary region.

[0207] ③ A global-local feature integration module is used to fuse morphological information at different scales.

[0208] The global-local feature integration module contains two parallel global branches and a local branch.

[0209] Global branch: The entire feature map is compressed into a channel description vector through global average pooling, capturing the overall morphological information of the cervical cell cytoplasm.

[0210] Local branches: Parallel processing using multi-scale convolutions (including 1×1, 3×3, 5×5 and 7×7 convolution kernels) captures point features, local features and regional features of different ranges respectively.

[0211] Adaptive Feature Fusion: By learning trainable fusion weights, the result of global feature expansion is combined with local features at various scales in a weighted manner. The fusion weights are dynamically generated based on the input features through a lightweight multilayer perceptron network, and softmax normalization is used to ensure that the weight sums to 1.

[0212] ④ An elliptical prior enhancement module is used to enhance the extraction of typical morphological features of cervical cell cytoplasm.

[0213] The elliptical shape prior enhancement module introduces an elliptical shape template set that conforms to the shape of cervical cells into the training set obtained from steps S1 and S2. In subsequent training processes, the template set is adjusted through deformable convolution and input features, integrating the prior knowledge of the elliptical shape of cervical cells into the feature extraction process.

[0214] This section (i.e., using a pre-defined elliptical template to extract prior knowledge) is similar to the subsequent description and processing of morphological statistical prior coding (i.e., analyzing and coding all images to extract prior knowledge).

[0215] Finally, through the above-mentioned morphological feature flow similar processing, morphological feature data of cervical cells were obtained.

[0216] (2) Texture feature flow path

[0217] ① Frequency domain filtering is used to extract multi-band texture features.

[0218] Specifically, firstly, the feature map is transformed into a frequency domain representation using a Fast Fourier Transform (FFT). Then, three sets of "frequency bands + filters" (low frequency, mid frequency, and high frequency) are used to extract features at different frequencies. Finally, an Inverse Fourier Transform is used to transform the feature map back into the spatial domain, ultimately capturing the multi-band texture features of the cytoplasm of cervical cells.

[0219] ② An orientation-sensitive convolutional layer is used to capture the directional texture features of the cytoplasm.

[0220] Specifically, multiple sets of directional convolution kernels (0°, 45°, 90°, 135°, etc.) are used to enhance texture features in different directions to adapt to the directional tissue structure in the cytoplasm of cervical cells.

[0221] ③ Use a nonlocal averaging module to capture long-distance texture dependencies.

[0222] The nonlocal averaging module implements a spatial attention mechanism that can model the relationship between any two spatial locations.

[0223] Specifically, firstly, query, key, and value features are generated through three 1×1 convolutions. Then, the spatial dimension is flattened, and a similarity matrix is ​​calculated between all position pairs; each element in the similarity matrix represents the feature similarity between two positions. Secondly, the similarity matrix is ​​normalized using softmax and used as attention weights to weighted aggregate the value features (another mapping representation of the input features). Thus, the output feature at each position integrates information from all other positions in the image, and the weights (the attention weights obtained by normalizing the similarity matrix using softmax) are determined by feature similarity. Finally, the non-local features (the output feature map obtained after the complete computation of the non-local averaging module) are added to the original features through residual connections, preserving local information while enhancing the ability to model long-distance dependencies.

[0224] ④ Apply transparency-sensing convolutional layers to specifically handle the translucent properties of overlapping cytoplasm regions of cervical cells.

[0225] Specifically, firstly, a transparency prediction branch is added, which learns the transparency value α at each location through an independent convolutional layer to obtain a transparency map. Secondly, the learned transparency map is used as an additional input channel, concatenated with the original features, and then fed into the main convolutional layer (the backbone of the texture feature flow extraction model), allowing the convolution operation to consider both feature information and transparency information simultaneously. Finally, according to the basic transparency hybrid model calculation formula defined in step S21 above, the predicted transparency value is constrained and adjusted to ensure that it conforms to physical optical properties. This design can explicitly model and utilize the semi-transparent properties of overlapping regions, improving the segmentation accuracy of overlapping cytoplasm.

[0226] Finally, through the same processing of the texture feature flow described above, texture feature data of cervical cells are obtained.

[0227] (3) The gated feature fusion mechanism is to fuse the data output after processing the morphological feature stream and the texture feature stream to obtain the fused feature; and add the fused feature to the original feature before processing the morphological feature stream and the texture feature stream to obtain the final output feature.

[0228] Specifically, the gating feature fusion method is as follows:

[0229] ① Through a learnable gating function G(F) m ,F t ), used to dynamically adjust the weight ratio of the two feature flows.

[0230] .

[0231] Among them, F m and F tW represents morphological features and textural features, respectively. g For learnable weight parameters, It is the sigmoid activation function.

[0232] ② The fusion parameters are adjusted based on the complexity and overlap of cell morphology.

[0233] The specific method for adjusting the fusion parameters is as follows: An additional lightweight network (an independent convolutional network containing two 3×3 convolutional layers, one global average pooling layer, and two parallel fully connected branches) is used to evaluate the morphological complexity and overlap of the current cervical cells, ultimately outputting a morphological complexity score and an overlap score. Based on the evaluation results of the morphological complexity score and overlap score, the parameters of the gating function are dynamically adjusted.

[0234] The specific method of dynamic adjustment is as follows: for cells with complex and irregular shapes, the weight of morphological feature flow is enhanced; for heavily overlapping regions, the weight of texture feature flow is enhanced. This dynamic adjustment method ensures that the gating weights are always within an effective range, while dynamically adjusting the relative importance of the two features according to cell characteristics.

[0235] ③ Residual connections are used to preserve the original cytoplasmic characteristics of cervical cells.

[0236] The residual connection method is as follows: First, residual connections are set in each processing module (i.e., each component of the morphological feature stream and the texture feature stream) within the feature stream (obtained in step S33). Second, in the feature fusion stage, the data from the morphological feature stream and the texture feature stream are weighted by a gating function to obtain the fused features. Finally, in the overall structure, the original features of the morphological feature stream and the texture feature stream before processing are added to the fused features to obtain the final output features. This multi-level residual connection facilitates gradient propagation and feature preservation.

[0237] ④ Perform feature recalibration to enhance sensitivity to overlapping cytoplasmic regions.

[0238] The specific method for feature recalibration is as follows: an adaptive instance normalization method is used to dynamically adjust the statistical distribution of features based on the degree of overlap of cervical cell cytoplasm; this is used to enhance the features of overlapping regions and improve the model's sensitivity to overlapping regions.

[0239] S34. Perform adaptive instance normalization to achieve conditional control and feature fusion; adjust features based on the mask of non-overlapping parts to statistically analyze characteristics. The core formula is:

[0240] .

[0241] Here, γ(c) and β(c) are scaling and offset parameters learned from the conditional feature c, and μ(x) and σ(x) are the mean and standard deviation of the feature map x, respectively.

[0242] The adaptive instance normalization method enables the distribution of feature maps to be dynamically adjusted according to conditional information, thereby enhancing the adaptability of the conditional diffusion model to the cytoplasmic morphology of different cervical cells.

[0243] S35. Perform frequency domain and spatial domain feature fusion: Extract global features in the frequency domain through fast Fourier transform, and then fuse them with spatial domain features.

[0244] The specific steps include: First, using 2D Fast Fourier Transform to extract the amplitude spectrum and phase spectrum from the feature map; then, applying a frequency domain attention mechanism to adjust the weights of different frequency components; second, using Inverse Fast Fourier Transform to convert the adjusted frequency domain features back to the spatial domain, dynamically integrating the frequency domain features and spatial domain features.

[0245] The dynamic integration method of frequency domain features and spatial domain features is achieved through a gating mechanism:

[0246] .

[0247] Where λ is the learned fusion weight, which is dynamically adjusted as the network trains, F spatial It is a spatial domain feature, F freq It is a frequency domain feature.

[0248] This method of fusing frequency domain and spatial domain features fully leverages the complementarity of the two representation spaces, enhancing the global understanding of cervical cytoplasmic morphology and capturing local details.

[0249] S4. Design an adaptive multi-scale combined loss function and train an optimized conditional diffusion model, such as... Figure 5 As shown.

[0250] S41. Construct an adaptive multi-scale combined loss function, including the construction of diffusion reconstruction loss, shape prior loss, multi-scale edge consistency loss, frequency domain structural similarity loss, morphological consistency loss, and overlapping region specificity loss.

[0251] (1) Constructing diffusion reconstruction loss: The difference between the predicted noise (generated by the final output layer of step S31) and the real noise at each time step is calculated using weighted L2 loss. Its dynamic weight coefficients are dynamically adjusted according to the noise level (the intensity of the noise added to the original image at a specific time step t of the diffusion process) and the location of the overlapping area.

[0252] The formula for calculating the diffusion loss function is as follows:

[0253] .

[0254] Where N is the batch size, It is noise in the model's predictions; It is real noise added to the original image; x t i Is the i-th sample at time step Noisy image; c i It is a conditional input (mask of the non-overlapping portion); w i It is a dynamic weighting coefficient calculated based on noise level and the location of overlapping areas.

[0255] The dynamic weighting coefficient is calculated using the following formula, which dynamically adjusts based on noise level and overlapping region location:

[0256] .

[0257] Where α is the base weight, β is the overlap region enhancement coefficient (usually 0.5), and o i It is the mask for the overlapping area.

[0258] (2) Constructing shape prior loss: combining boundary weighted binary cross-entropy loss (BCE) w ) and Tversky loss, to strengthen the constraint on cytoplasmic shape;

[0259] The formula for calculating the shape prior loss function is as follows:

[0260] .

[0261] Among them, y pred and y true These are the predicted mask and the true mask, respectively; α(s) is a weight dynamically adjusted based on the cell morphological complexity s; the Tversky loss assigns different weights to false positives and false negatives using two parameters β=0.3 and γ=0.7, respectively.

[0262] Boundary Weighted Binary Cross-Entropy Loss (BCE) w The formula for calculating ) is:

[0263] .

[0264] The formula for calculating Tversky loss is:

[0265] .

[0266] Where N is the batch size; w b (i) is the boundary weight, which is calculated from the distance of the pixel to the nearest boundary; y true iy is the true label of the i-th sample; pred i It is the predicted probability of the i-th sample.

[0267] (3) Construct a multi-scale edge consistency loss: compare the edge features of the predicted mask and the real mask at multiple scales, paying particular attention to the consistency of the boundary features of the overlapping region.

[0268] Specifically, first, the original scale, half-scale, and quarter-scale are selected with weights of 0.5, 0.3, and 0.2, respectively. Then, the Sobel edge detection operator, Canny edge detection, and morphological gradient algorithm are used to extract edge features from the predicted mask and the ground truth mask. The edge loss weight is increased for overlapping regions.

[0269] The formula for calculating multi-scale edge consistency loss is:

[0270] .

[0271] Where S is the scale index, w s It is the weight of each scale, L edge s This represents the single-scale edge consistency loss.

[0272] The single-scale edge consistency loss is calculated as follows:

[0273] ;

[0274] .

[0275] Where E is the edge detection operation, and N is the number of pixels; w e (i) is the edge weight, calculated based on the overlapping region mask and the overlapping region enhancement coefficient β of O(i).

[0276] (4) Construct frequency domain structural similarity loss: evaluate the structural similarity between the predicted mask and the real mask in the frequency domain space.

[0277] The formula for calculating frequency domain structural similarity loss is:

[0278] .

[0279] Among them, y pred and y true These are the predicted mask and the true mask, respectively. F represents the Fourier transform, and SSIM... freq It is a frequency domain structural similarity index.

[0280] Frequency-domain SSIM, compared to spatial-domain SSIM, is better at capturing global structural features and is more sensitive to the overall morphological assessment of the cytoplasm. Its calculation formula is as follows:

[0281] .

[0282] Where, μ X and μ Y It is the mean of the amplitude spectrum. and It is the variance of the amplitude spectrum, σ XY C1 and C2 are the covariance of the amplitude spectrum and the stability constants.

[0283] (5) Construct morphological consistency loss: Combine the modified Hausdorff distance metric loss and topological consistency loss, and constrain the morphological rationality of the segmentation results by comparing morphological features.

[0284] Morphological consistency loss L morph The calculation formula is:

[0285] .

[0286] Where α is the balance coefficient (usually set to 0.5), L Hausdorff It is the modified Hausdorff distance loss, L topo It is a loss of topological consistency.

[0287] The modified Hausdorff distance metric loss is defined as:

[0288] .

[0289] Where p represents a point on the predicted mask boundary, q represents a point on the true mask boundary, and ∂y pred Denotes the boundary of the prediction mask, ∂y ture Represents the boundary of the real mask.

[0290] Topological consistency loss is used to evaluate the topological similarity between the predicted mask and the true mask, including properties such as connectivity and the number of holes. The formula is defined as follows:

[0291] .

[0292] Among them, CC(y) pred ) and CC(y true H(y) is the number of connected components. pred H(y) represents the number of holes in the generated mask. true ) represents the number of holes in the real mask, and w is the weighting coefficient. This morphological consistency loss ensures that the generated mask maintains topological consistency with the real mask.

[0293] (6) Constructing overlapping region-specific loss: Design a special loss function for overlapping regions, including mean squared error loss and transparency consistency loss.

[0294] The formula for calculating the specificity loss of the overlapping region is:

[0295] ;

[0296] Where O is the overlapping region mask, L trans It is the transparency consistency loss, λ trans It is the weighting coefficient (usually set to 0.7).

[0297] Transparency consistency loss measurement is used to evaluate whether the transparency transition of the predicted mask in the overlapping area is natural and conforms to optical properties.

[0298] S42. Construct the total loss function: The combined loss function is the weighted sum of all losses in step S41 above, and the weight coefficients are dynamically adjusted with time steps.

[0299] The formula for calculating the total loss function is:

[0300] .

[0301] Among them, w diff (t) represents the weighting coefficient of the diffusion reconstruction loss, w shape (t) represents the weighting coefficient of the shape prior loss, w edge w represents the weighting coefficients of the multi-scale edge consistency loss. freq (t) represents the weighting coefficients of the frequency domain structural similarity loss, w morph (t) represents the weighting coefficient of the morphological consistency loss, w overlap (t) represents the weighting coefficients for the overlap region-specific loss; L diff L represents the diffusion-reconstruction loss. shape L represents the shape prior loss. edge L represents the multi-scale edge consistency loss. freq L represents the frequency domain structural similarity loss. morph L represents the loss of morphological consistency. overlap This indicates a loss specific to the overlapping region.

[0302] Specifically, the weighting coefficient w i(t) is dynamically adjusted with time step t, and appropriate parameters will be learned during the training process of the model with a given threshold. In the early stage of training (t<0.3), more emphasis is placed on diffusion reconstruction loss and shape constraint loss, that is, the weights of reconstruction loss and shape constraint loss are set to larger values ​​in the early stage of training; (2) In the middle stage of training (0.3≤t<0.7), more emphasis is placed on multi-scale edge consistency loss and frequency domain structural similarity loss, and more attention is paid to edge details and structural consistency; (3) In the later stage of training (t≥0.7), more emphasis is placed on morphological consistency loss and overlapping area specificity loss, and more attention is paid to the morphological rationality of the generated mask and the processing of overlapping areas.

[0303] S43, Training Conditional Diffusion Model

[0304] In the dataset usage strategy section, the training set is used for iterative updates and learning of model parameters. The data order is randomly shuffled at each epoch to ensure training randomness. Augmented data samples are generated in real-time through data augmentation and normalization processes in step S2. The validation set is used for evaluation every 5 epochs to monitor the training process, prevent overfitting, and adjust the learning rate based on validation set performance. The test set is used after training, stratified according to morphological complexity and overlap, to test and evaluate the performance of the final model.

[0305] The training method for the conditional diffusion model is as follows:

[0306] (1) Set the AdamW optimizer, with an initial learning rate of 1e-4, a weight decay parameter of 1e-5, and use cosine annealing to schedule the learning rate. The minimum learning rate is set to 5e-6, and the warm-up period is 10 epochs.

[0307] (2) Training is performed using an incremental batch size approach. The initial training batch size is set to 8, and it is increased to 16 in the middle of training (100th epoch); the total training cycle is 300 epochs. Training can be terminated early when the performance on the validation set stops improving for 20 consecutive epochs.

[0308] The conditional diffusion model uses 1000 noise steps and employs linear noise scheduling.

[0309] The training steps for a conditional diffusion model include:

[0310] (1) Data loading and preprocessing: Load the non-overlapping part mask (condition) and the complete mask (target) of the cytoplasm from the cervical cell dataset and apply random data augmentation.

[0311] (2) Forward propagation: Add noise to the target mask, randomly select time steps, and use the conditional diffusion model to predict the noise.

[0312] (3) Loss calculation: Calculate the loss of each component (6 loss functions) and combine them in a weighted manner to form the total loss.

[0313] (4) Backpropagation: Calculate the gradient and update the model parameters.

[0314] (5) Evaluation: Periodically evaluate the model performance on the validation set and save the best model.

[0315] S5. A hierarchical classifier-guided sampling method is used to obtain a complete cervical cell cytoplasmic mask. The process is as follows: Figure 6 As shown.

[0316] S51, Initialize sampling

[0317] (1) Use the instance segmentation model to extract the non-overlapping cervical cytoplasmic region (non-overlapping part mask) as conditional input.

[0318] The instance segmentation model is based on the Mask R-CNN architecture and is pre-trained on the training set of images using the cervical cytoplasm mask constructed in step S1.

[0319] During training, the original cervical cell image is used as input, and the mask of the non-overlapping region is used as the target output to optimize the model's ability to identify and extract non-overlapping cytoplasmic regions. The backbone network of the model adopts ResNet-50, and the RPN (Region Proposal Network) configuration is adjusted for the scale range of cervical cells. The anchor box size is set to [16, 32, 64, 128] pixels, and a multi-scale feature fusion strategy of Mask R-CNN is introduced to improve the detection accuracy of cells of different sizes.

[0320] (2) Based on the extracted non-overlapping part mask, generate initial condition information. The initial condition information is processed by the feature extraction encoder in step S3.

[0321] Specifically, firstly, the non-overlapping portion mask is input into a feature extraction encoder containing four levels (64, 128, 256, and 512 channels respectively); then, after downsampling, multi-scale morphological and textural features are extracted. The encoded feature vector contains morphological features, textural features, and positional information of the cytoplasm, serving as conditional information to guide the diffusion process in the conditional diffusion pore.

[0322] (3) Initialize the random noise mask as the initial state. The random noise is sampled using the standard normal distribution N(0,1).

[0323] (4) Set the sampling steps and guiding intensity parameter ω for DDIM (Denoising Diffusion Implicit Model). Diffusion sampling adopts the DDIM framework, and the total number of sampling steps is set to 50 to balance generation quality and speed. 50 points are uniformly sampled from T=1000: t∈{1000, 980, 960, ..., 20}.

[0324] The intensity parameter ω is dynamically adjusted at different sampling stages to control the degree of influence of conditional information on the generation process.

[0325] S52. Divide the sampling process into three stages: coarse, fine, and refined, and design corresponding hierarchical sampling methods for each stage.

[0326] (1) Roughing stage (steps 1-20, corresponding to time steps t=1000 to t=600): Use a larger guiding intensity (guiding intensity parameter ω=3.5). This stage emphasizes the consistency of the overall shape, that is, in this stage, we mainly focus on the overall morphological structure of the cytoplasm of cervical cells to ensure that the overall shape conforms to the general morphological characteristics of the cytoplasm.

[0327] (2) Refinement stage (steps 21-40, corresponding to time steps t=580 to t=200): Use medium guidance intensity (guidance intensity parameter ω=2.0). This stage is used to balance shape and edge details, that is, to start focusing on the boundary and detailed features of the cytoplasm at this stage, and gradually improve the local structure of the cytoplasm.

[0328] (3) Refinement stage (steps 41-50, corresponding to time steps t=180 to t=20): Use a smaller guidance intensity (guidance intensity parameter ω=1.5). This stage focuses on local details, that is, optimizing details, especially the smoothness and coherence of the cytoplasmic edges, and the fine processing of overlapping areas.

[0329] The formula for calculating conditional noise prediction under classifier free guidance is:

[0330] .

[0331] in, It is conditional noise prediction. It is an unconditional noise prediction, and ω is the guiding strength parameter.

[0332] S53. Perform the following steps in each sampling step.

[0333] (1) Generate conditional noise prediction and unconditional noise prediction.

[0334] Conditional noise prediction The conditional information c (containing coded features of the non-overlapping mask) is used as input and generated through the conditional diffusion model built in S3. Unconditional noise prediction. Generated using empty conditions (all-zero vectors) as input. Both share the same U-shaped network structure parameters of the conditional diffusion model.

[0335] (2) Based on the guidance intensity of the current stage, combined with conditional noise prediction and unconditional noise prediction, the calculation formula for classifier free guidance in S52 is used. The corresponding ω value is selected according to the current sampling stage, and the weighted combination noise prediction is calculated. .

[0336] (3) Use DDIM (Denoising Diffusion Implicit Model) sampling to calculate the noise state for the next step.

[0337] The DDIM sampling formula is:

[0338] .

[0339] Where, x t This is the current noise state, x t-1 This is the next noise state; α t and α t-1 It is the cumulative diffusion coefficient; It is the noise in the model prediction, σ t is the noise level at time step t, z is the random noise sampled from the standard normal distribution, and c is the conditional input.

[0340] (4) Use masking operations to keep the known area unchanged.

[0341] Specifically, for the regions (non-overlapping parts) that have already been determined in the non-overlapping mask, the original values ​​are used directly instead of the values ​​generated from the diffusion process to ensure that the generated results are consistent with the conditional inputs.

[0342] (5) Apply cervical cell cytoplasmic morphology constraints to regularize the generation process.

[0343] Specifically, based on the morphological characteristics of cervical cell cytoplasm (such as connectivity, smoothness, shape constraints, etc.), the generated intermediate states (intermediate states generated at each time step t during the noise reduction generation process of the conditional diffusion model) are fine-tuned to ensure that they conform to the morphological rules of cervical cell cytoplasm.

[0344] S54. Design an adaptive noise adjustment method.

[0345] (1) Adjust the conditional noise prediction in the conditional diffusion model dynamically based on the current noise level and the results of previous sampling steps.

[0346] Specifically, adaptive noise conditioning calculates a noise adjustment factor based on the generation results of the previous sampling steps and the current noise level; and dynamically adjusts the predicted noise to enhance the stability and quality of the generation process.

[0347] (2) Apply special noise adjustment methods to the overlapping areas to enhance the generation of semi-transparent overlapping parts.

[0348] Specifically, a special adaptive noise adjustment method is designed to address the unique characteristics of overlapping regions, thereby enhancing the modeling capability of semi-transparent regions and ensuring that the generated overlapping regions have a natural transition and reasonable boundaries.

[0349] The formula for calculating adaptive noise is:

[0350] .

[0351] in, It is a conditional noise prediction guided freely by a classifier; It is a conditional noise prediction based on morphological prior information; λ t It is a mixing coefficient value that changes over time, gradually decreasing from close to 1 in the early noise reduction steps to about 0.7 in the later noise reduction steps.

[0352] S55. Perform post-processing and rationality verification on the final result.

[0353] (1) Apply binarization and set the threshold to 0.5.

[0354] Specifically, the continuous value mask (representing the probability that each pixel belongs to the cytoplasm) generated by the conditional diffusion model is binarized with a threshold of 0.5 to convert the pixel values ​​to 0 or 1, thus obtaining a clear cytoplasm mask.

[0355] (2) Perform condition consistency correction to ensure that the known area remains unchanged.

[0356] Specifically, the generated complete cytoplasmic mask is compared with the input non-overlapping mask to ensure that the known regions (non-overlapping parts) remain unchanged in the generated result, correcting any possible inconsistencies.

[0357] (3) Perform morphological post-processing, including the removal of small connected regions and the filling of holes.

[0358] Specifically, small connected regions with an area less than 1% of the total area are removed, and holes with an area less than 0.5% of the total area are filled to ensure that the topology of the generated mask is reasonable.

[0359] (4) Use morphological smoothing to refine the cytoplasmic boundary.

[0360] Specifically, a morphological smoothing operation with a radius of 1 pixel (usually a combination of opening and closing operations) is used to make the cytoplasmic boundary smoother and more natural, removing possible jagged edges.

[0361] (5) Verify the rationality of the segmentation results based on all prior knowledge of cytoplasmic morphology, and make constraint corrections if necessary.

[0362] Specifically, based on a pre-established cervical cytoplasm database, the shape parameters, roundness, elongation, topology, and naturalness of transitions in overlapping regions of the cytoplasmic morphological results generated by the conditional diffusion model are verified to be within reasonable ranges, and abnormal regions are appropriately corrected.

[0363] Through the above steps, the present invention can predict the complete cytoplasmic morphology from known non-overlapping cytoplasmic regions in cervical cell images, thereby effectively solving the problem of segmenting overlapping cytoplasm in cervical cells.

[0364] Figure 7 and Figure 8 The images show a comparison of the input and output of the conditional diffusion model used in this invention for segmenting overlapping cervical cells. It can be seen that the method provided by this invention can generate a complete cytoplasmic mask (output) from a partially non-overlapping cervical cell mask (input). The generated complete cytoplasmic mask has smooth and natural mask boundaries, and its morphology conforms to the biological characteristics of cytoplasm, demonstrating the model's powerful generative capabilities.

[0365] This invention provides an automated cervical cytology analysis system for analyzing cervical cell images and screening for cervical cancer. The method can process cervical cell smear images captured under a microscope, achieving accurate segmentation of overlapping cytoplasm, providing a foundation for subsequent cell classification and diagnosis. Furthermore, this method has great potential for further application to other types of cell overlap segmentation tasks, such as blood cells and nerve cells.

[0366] The core innovation of this invention lies in the effective integration of conditional diffusion models and prior knowledge of cell morphology. Through conditional diffusion models, adaptive multi-scale combined loss functions, and hierarchical classifiers, it achieves free-guided sampling, improving the segmentation accuracy of overlapping cytoplasmic regions. This method can be applied to automated cervical cytology analysis systems, improving the accuracy of cell detection and providing reliable technical support for early cervical cancer screening.

[0367] The above detailed embodiments describe the implementation of the present invention; however, the present invention is not limited to the specific details described in the above embodiments. Within the scope of the claims and technical concept of the present invention, various simple modifications and changes can be made to the technical solution of the present invention, and these simple modifications all fall within the protection scope of the present invention.

Claims

1. A method for segmenting overlapping cervical cytoplasmic regions based on deep learning and conditional diffusion models, characterized in that, Includes the following steps: Construct images of cervical cytoplasmic mask pairs and divide them into training, validation, and test sets; The data in the training set, validation set, and test set are augmented and standardized. A conditional diffusion model is constructed, comprising a U-shaped network structure, a cervical cell morphology prior module, and a two-stream attention mechanism. The U-shaped network structure includes a feature extraction encoder, a morphology-aware intermediate block, a feature extraction decoder, and a final output layer. Adaptive instance normalization is performed to achieve conditional control and feature fusion. Features are adjusted based on the mask of non-overlapping parts to statistically analyze characteristics. Frequency domain and spatial domain features are fused. Design an adaptive multi-scale combined loss function, including diffusion reconstruction loss, shape prior loss, multi-scale edge consistency loss, frequency domain structural similarity loss, morphological consistency loss, and overlapping region specificity loss; construct the total loss function; Training the conditional diffusion model; A hierarchical classifier-guided sampling strategy is used for inference to obtain the complete cytoplasmic mask.

2. The method for segmenting overlapping cervical cytoplasmic regions based on deep learning and conditional diffusion model according to claim 1, characterized in that, Methods for constructing cervical cytoplasmic mask pairs for images include: From the pathological images and their annotation files of the ISBI 2015 cervical cell dataset, the complete cytoplasm mask of cervical cells is extracted, and the non-overlapping part mask is calculated; the non-overlapping part mask is used as the input condition of the conditional diffusion model, and the complete mask is used as the target mask of the conditional diffusion model. The non-overlapping portion mask and the complete mask extracted from a single cervical cell are combined to form a multi-scale cervical cytoplasmic mask image; the image is then layered according to morphological complexity and degree of overlap; the layering results are divided into training set, validation set, and test set; cell samples of various degrees of overlap and morphological complexity are evenly distributed in the training set, validation set, and test set. The mask pairs are uniformly processed to a size of 256×256 pixels.

3. The method for segmenting overlapping cervical cytoplasmic regions based on deep learning and conditional diffusion model according to claim 1, characterized in that, Methods for augmenting and standardizing the data in the training, validation, and test sets include: The training set, validation set, and test set data are sequentially subjected to morphological deformation, random threshold binarization, rotation, flipping, shearing, scaling, adversarial morphological noise processing, and transparency simulation enhancement processing. The enhanced mask image is then normalized. The normalization process includes grayscale conversion, normalization, and edge smoothing; a mask comparison sample library is then constructed.

4. The method for segmenting overlapping cervical cytoplasmic regions based on deep learning and conditional diffusion model according to claim 3, characterized in that, The morphological deformation method includes: applying elastic transformation to cervical cells, with the deformation intensity parameter dynamically adjusted according to the morphological complexity and overlap of the cervical cell cytoplasm; the morphological complexity includes regular shapes, complex shapes, and irregular shapes; the overlap is divided into mild overlap, moderate overlap, and severe overlap. Random thresholding binarization methods include: using a Gaussian distribution of N(0.5,0.05) to perform random thresholding binarization on an image with a morphologically deformed mask; The methods of rotation, flipping, shearing and scaling include: randomly rotating the image of the mask that has been randomly thresholded and binarized from 0 to 360°, flipping it horizontally, flipping it vertically, performing random shearing transformation and micro-scaling; Methods for handling adversarial morphological noise include: adding random morphological noise to the cytoplasmic edge region of an image that has been rotated, flipped, clipped, and scaled; using a morphological gradient algorithm to extract the cytoplasmic edge region; adding random noise with an amplitude of no more than 3 pixels to the edge region; and applying morphological opening and closing operations to make the noise have a natural transition effect. Transparency simulation methods include calculating transparency blending using spatially varying alpha values ​​α(x,y) in the overlapping regions of cervical cell cytoplasm.

5. The method for segmenting overlapping cervical cytoplasmic regions based on deep learning and conditional diffusion model according to claim 1, characterized in that, In the conditional diffusion model, the U-shaped network structure adopts a four-stage U-shaped structure; the feature extraction encoder and feature extraction decoder each contain four levels, each level including several residual blocks; each residual block contains two 3×3 convolutional layers, group normalization, and SiLU activation function; The feature extraction encoder includes two residual blocks, temporal embedding, conditional fusion, multi-scale downsampling, and skip connections. Each multi-scale downsampling layer consists of two residual blocks and one downsampling layer. The number of channels in the feature extraction encoder is 64, 128, 256, and 512, respectively. The downsampling uses a convolution operation with a stride of 2. The morphology-aware intermediate block includes four components: a multi-receptive-field dilated convolution module, a multi-head self-attention mechanism, a cross-attention module, and a channel attention mechanism. The multi-receptive-field dilated convolution module uses dilated convolutions with dilation rates of 1, 2, 4, and 8. The cross-attention module uses the features from the feature extraction encoder as a query and interacts with the features extracted by the U-shaped network structure. The channel attention mechanism uses a compressed excitation network to adjust the channel weights. The feature extraction decoder includes four upsampling modules, instance normalization, residual connections, spatial-frequency domain fusion, feature fusion, and morphological guidance. Each upsampling module includes an upsampling layer and two residual blocks, which are fused with the corresponding encoder features through skip connections. Upsampling is performed using deconvolution. The cervical cell morphology prior module includes a global morphology description generator, an edge consistency enhancer, a topology maintainer, and a morphological statistics prior encoder, which are used to extract and encode the morphological features, edge properties, topology, and shape statistics of the cytoplasm of cervical cells. The dual-stream attention mechanism takes the feature map output by the U-shaped network structure as input, and after its respective feature extraction and enhancement, outputs a fused feature. It is used to capture morphological feature streams and texture feature streams, and integrate these two types of features through gated fusion.

6. The method for segmenting overlapping cervical cytoplasmic regions based on deep learning and conditional diffusion model according to claim 5, characterized in that, The specific method of the dual-stream attention mechanism includes: taking the feature map output by the U-shaped network structure in the conditional diffusion model as input, and obtaining the morphological feature stream and the texture feature stream respectively; For the morphological feature flow pathway, multi-scale dilated convolution is used to extract morphological features; an edge-aware attention mechanism is used to enhance the learning of cytoplasmic boundary features; global branch-local branch feature integration is used to fuse morphological information at different scales; an elliptical shape prior enhancement module is used to introduce the enhanced and normalized training set mask into the image to an elliptical shape template set that conforms to the shape of cervical cells. Through deformable convolution and input features, the template set is adjusted to incorporate the prior knowledge of the elliptical shape of cervical cells into the feature extraction process, and finally, the morphological feature data of cervical cells is obtained. For the texture feature flow pathway, multi-band texture features are extracted by frequency domain filtering; directional sensitive convolutional layers are used to capture directional texture features of the cytoplasm; nonlocal averaging mechanism is used to capture the dependencies of long-distance textures; and transparency-aware convolutional layers are used to process the semi-transparent properties of overlapping areas, ultimately obtaining texture feature data of cervical cells. The gated feature fusion mechanism adjusts the weight ratio of morphological feature flow and texture feature flow through a gating function, adjusts the fusion parameters through the complexity and overlap of cell morphology, preserves the original cytoplasmic feature information through residual connections, and enhances the sensitivity to cytoplasmic overlapping regions through feature recalibration.

7. The method for segmenting overlapping cervical cytoplasmic regions based on deep learning and conditional diffusion model according to claim 5, characterized in that, In the cervical cytoplasmic morphology prior module, the global morphological description generator is used to extract and encode the morphological features of the cervical cell cytoplasmic mask image after step S2 through global pooling and multilayer perceptron network. The edge consistency enhancer is used to extract edge features at different scales, and the feature map is modulated through the edge response map to enhance the expression of the edge region; The topology preserver is used to extract the skeleton of the mask through a morphological thinning algorithm, and the topology of the cytoplasm is preserved through skeleton-guided feature enhancement; distance transformation weighting is used to ensure that the weight distribution of the skeleton region conforms to the morphological characteristics of the cytoplasm. Using the aforementioned morphological statistical prior encoder, principal component analysis of cytoplasmic shape is constructed based on the enhanced and standardized cervical cell dataset. Low-dimensional representation of the shape is extracted, and the shape prior encoding of cervical cells is fused with the feature map through deformable convolution.

8. The method for segmenting overlapping cervical cytoplasmic regions based on deep learning and conditional diffusion model according to claim 1, characterized in that, In the adaptive multi-scale combined loss function, the diffusion loss uses weighted L2 loss to calculate the difference between the predicted noise and the actual noise at each time step, and its weight coefficients are adjusted according to the noise level and the location of the overlapping area. The shape prior loss, combined with boundary-weighted binary cross-entropy loss and Tversky loss, strengthens the constraints on cytoplasmic shape. The multi-scale edge consistency loss compares the edge features of the predicted mask and the real mask at multiple scales, especially the consistency of the boundary features of the overlapping region. The frequency domain structural similarity loss evaluates the structural similarity between the predicted mask and the true mask in the frequency domain space; The morphological consistency loss, combined with the modified Hausdorff distance metric loss and topological consistency loss, constrains the morphological rationality of the segmentation results by comparing morphological features. The overlapping region-specific loss is designed for overlapping regions and includes a loss function, including mean square error loss and transparency consistency loss, to confirm whether the transparency transition of the predicted mask in the overlapping region is natural and conforms to optical characteristics.

9. The method for segmenting overlapping cervical cytoplasmic regions based on deep learning and conditional diffusion model according to claim 1, characterized in that, The hierarchical classifier's free-guided sampling strategy includes: dividing the sampling process into a coarse stage, a refinement stage, and a fine stage; the guidance strength used in the coarse stage is ω=3.5; the guidance strength used in the refinement stage is ω=2.0; and the guidance strength used in the fine stage is ω=1.

5. Conditional and unconditional noise predictions are generated. Based on the guidance intensity of the current stage, the corresponding ω value is selected according to the current sampling stage, and the weighted combination of noise predictions is calculated. DDIM (Denoising Diffusion Implicit Model) is used for sampling to calculate the noise state of the next step. The known area is kept unchanged through masking operations, and the generation process is regularized by applying cervical cell cytoplasmic morphology constraints. Based on the current noise level and the results of previous sampling steps, the conditional noise prediction in the conditional diffusion model is adjusted, and a special noise adjustment method is applied to the overlapping area to enhance the generation of the semi-transparent overlapping part. Post-processing and rationality verification of the final results; The post-processing includes binarization, conditional consistency correction, morphological post-processing, and morphological smoothing. The rationality verification includes verification of shape parameters, topology, and the naturalness of transitions in overlapping regions, and constraint correction is performed when the verification falls within the unreasonable range.

10. A system for segmenting overlapping cervical cytoplasmic regions based on deep learning and a conditional diffusion model, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the steps of the method according to any one of claims 1-9.

Citation Information

Cited By

  • Optical flow estimation method and system

    CN121120683A

  • Automatic bacterial colony identifying and counting method based on machine learning segmentation

    CN121482784A

  • Multi-stage cell culture-oriented adaptive mechanism fusion living cell density prediction method

    CN121583333A

  • Feces sample microscopic image detection method and system based on morphological prior constraint

    CN121884016A

  • Stool sample microscopic image detection method and system based on morphological prior constraint

    CN121884016B