Method and system for processing cell image based on DistSegNet model

Through image processing technology based on DistSegNet model, combined with Pix2Pix GAN and U-Net models, the problem of difficult balance between accuracy and morphological authenticity in WSI-level cell detection and segmentation is solved, and high-precision and reliable nuclear segmentation are achieved.

CN120219284APending Publication Date: 2025-06-27LICHUANG DIAGNOSTIC TECHNOLOGY (SUZHOU) CO LTD +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510158452.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to balance accuracy and morphological authenticity in WSI-level cell detection and segmentation, especially when processing high-resolution large-scale images.

Method used

Image processing technology based on the DistSegNet model is adopted, combined with the Pix2Pix GAN model for image enhancement, segmentation is used with the classic U-Net model, and cell contour detection is optimized through morphological operations and watershed algorithms.

Benefits of technology

It significantly improves the accuracy and reliability of cell nuclear segmentation, improves the overall performance of cell detection and segmentation, and ensures accurate extraction of cell profiles in case of cell density or adhesion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219284A_ABST
    Figure CN120219284A_ABST
Patent Text Reader

Abstract

The invention provides a method and a system for processing a cell image based on a DistSegNet model. The method for detecting the cells based on the DistSegNet model image processing technology comprises the following steps: preprocessing a collected cell image; performing cell segmentation on the cell image based on a DistSegNet model image processing technology to generate a cell nucleus region; based on the segmented cell nucleus region, generating a background region, a foreground region and an edge region by adopting morphological operation; calculating the shortest Euclidean distance from each background region pixel to the nearest foreground region pixel, and segmenting the cell image by using a watershed algorithm; performing result optimization on the segmentation map corresponding to each cell instance label to obtain a filtered segmentation label map; and extracting the outer contour of the intra-nucleus region and the outer contour of the surrounding region of each cell from the segmentation mark graph, and storing the outer contours to a corresponding result file. According to the invention, the accuracy and accuracy of a WSI-level cell detection result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular, to a method and system for detecting cells based on the DistSegNet model image processing technology, and a method and system for stitching the cell image segmentation results. Background Art

[0002] With the development of digital pathology, the whole slide image (WSI) technology has become an important tool in pathological image analysis. The WSI technology can record the entire pathological section with high resolution, providing a basis for the storage, analysis, and sharing of large-scale pathological images. However, due to the complex cell morphology, high density, blurred boundaries in pathological images, and the diversity of microscope imaging conditions, achieving high-precision cell detection and segmentation has always been a major technical challenge in this field.

[0003] Currently, the application of deep learning technology, especially convolutional neural networks (CNNs), has made remarkable progress in the field of automated processing of pathological images. For example, the U-Net-based segmentation method is widely used in medical image segmentation tasks due to its simple structure and excellent performance. In addition, generative adversarial networks (GANs) have also shown powerful capabilities in medical image enhancement and segmentation in recent years. For another example, the Pix2Pix GAN model can generate high-quality enhanced images through adversarial learning between the generator and the discriminator, making the cell nucleus center more prominent and the boundary clearer. At the same time, some methods based on traditional image processing technologies (such as morphological operations and watershed algorithms) are still widely used. These methods have certain advantages in capturing global features and enhancing boundary characteristics.

[0004] However, currently, it is still difficult to balance the accuracy and morphological authenticity of cell detection and segmentation methods at the WSI level. Especially when dealing with high-resolution large-scale images, how to combine the powerful modeling ability of deep learning with traditional image processing technologies to optimize the accuracy and morphological authenticity of cell contour detection is still an unsolved technical challenge. Summary of the Invention

[0005] The technical problem solved by the technical solution of the present invention is: how to improve the accuracy and precision of the cell detection results at the WSI level.

[0006] To solve the above technical problem, the technical solution of the present invention provides a method for detecting cells based on the DistSegNet model image processing technology, including:

[0007] Preprocess the collected cell images;

[0008] Based on the DistSegNet model image processing technology, perform cell segmentation on the cell images to generate the cell nucleus region;

[0009] Based on the segmented cell nucleus region, use morphological operations to generate the background region, foreground region, and edge region;

[0010] Calculate the shortest Euclidean distance from each pixel in the background region to the nearest pixel in the foreground region, and use the watershed algorithm to segment the cell images;

[0011] Optimize the results of the segmentation map corresponding to each cell instance label to obtain a filtered segmentation label map;

[0012] Extract the outer contour of the nuclear region and the outer contour of the surrounding region of each cell from the segmentation label map, and save them to the corresponding result file.

[0013] Optionally, the collected cell images are whole-slide images, and the preprocessing of the collected cell images includes:

[0014] Segment the whole-slide image into several image patches, each patch image contains local cell information, and retain the position information of the graphic patch in the global whole-slide image according to a preset naming format;

[0015] Upload the cell section image to be processed to the server, save the image to a specified folder, and convert the section image to grayscale format.

[0016] Optionally, the performing cell segmentation on the cell images based on the DistSegNet model image processing technology to generate the cell nucleus region includes:

[0017] Load the preprocessed cell images, and input the preprocessed cell images into the Pix2Pix GAN model;

[0018] Use the Pix2Pix GAN model to enhance the preprocessed cell images;

[0019] Use the classical U-Net model to segment the enhanced image to output the maximum value obtained by the class probability of the segmentation image pixels, so as to generate a class mask for cell segmentation.

[0020] Optionally, the Pix2Pix GAN model includes: a generator and a discriminator; the using the Pix2Pix GAN model to enhance the preprocessed cell images includes:

[0021] By combining traditional convolutional layers with the Res-CBAM module, the ability to capture key information is enhanced using residual connections and attention mechanisms, while background noise is filtered to design the generator of the Pix2Pix GAN model;

[0022] Based on the multi-scale method of Transformer, the Patch size is dynamically adjusted to capture local details and global morphology, and the enhanced image is used as the input of the discriminator; the input grayscale image is sliced into Patches of different sizes, and the features of each Patch are mapped to a preset dimension through linear projection, and a learnable position encoding PosEncoding is introduced to design the discriminator of the Pix2Pix GAN model;

[0023] Through the interactive training of the generator and the discriminator, the generator generates the enhanced image through feature extraction and enhancement modules, and the discriminator, based on the multi-scale Transformer structure, determines whether the enhanced image output by the generator is a real image. Through adversarial training, the generator is pushed to generate high-quality enhanced images that more conform to the real data distribution, and finally the enhanced image is output.

[0024] Optionally, the enhanced image is segmented using the classic U-Net model to output the maximum value of the category probability of the segmented image pixels to generate a category mask for cell segmentation, including:

[0025] Let the enhanced grayscale image output by the Pix2Pix GAN model to the U-Net model be Y, with dimensions H×W. Three types of segmentation masks are output through the classic U-Net model, respectively representing: Z1 is the main body of the cell nucleus, Z2 is the edge region, and Z3 is the background; then the structure of the U-Net is set as:

[0026] There is an encoder suitable for extracting multi-scale features F of the input image Y through multiple convolutional layers and downsampling layers l+1 , and there is:

[0027] F l+1 =σ(W l F l +b l ), l∈{1,2,...,L e}

[0028] where W l is the convolutional kernel weight of the l-th layer, b l is the bias term, σ is the activation function, l takes 1,2,...,L e , L e is the total number of layers of the encoder, and F l is the multi-scale feature of the current input image Y being extracted;

[0029] The encoder is also adapted to perform downsampling through a max pooling operation Pool:

[0030]

[0031] representing the downsampled image features of the output;

[0032] A decoder is provided, adapted to gradually restore high-resolution features through upsampling and convolution operations:

[0033]

[0034] where UpSample is the upsampling operation, Concat is the skip connection, combining the output high-resolution encoder features F down l ;

[0035] The decoder is also adapted to output results after the last convolutional layer, and generate three types of segmentation results c in combination with an activation function. c can take values of 1, 2, 3, which can respectively represent whether the pixel belongs to Z1, Z2, and Z3:

[0036]

[0037] where S i,j,c is the original score that the image pixel position (i, j) belongs to class c, and the output result P i,j,c is the probability distribution that the pixel position (i, j) belongs to class c after being normalized by the activation function.

[0038] Optionally, using the classical U-Net model to segment the enhanced image to output the maximum value of the class probabilities of the segmentation image pixels, and generate a class mask for cell segmentation, further includes:

[0039] Defining a combined loss function L seg :

[0040] L seg = L CE + λL Dice

[0041] where L CE is the cross-entropy loss function, L Dice is the Dice loss function, λ is the loss weight, and the parameter value of the loss weight λ is preset;

[0042] Using the cross-entropy loss function L CE to optimize the accuracy of the output result P i,j,c for class prediction, it can be designed:

[0043]

[0044] In this formula, P i,j,c represents the predicted probability value belonging to the c-th class at the pixel position (i, j) in the image. Specifically:

[0045] (i, j) represents the pixel coordinate position in the image, referring to the pixel in the i-th row and the j-th column. c is the class index, indicating the class to which the pixel belongs, which may be different types of cells, background, or other regions. P i,j,c is the predicted probability that the pixel belongs to class c.

[0046] refers to the "true label" or "target label," that is, the true class label at the position (i, j) in the image, representing the true value that the pixel actually belongs to class c. Usually, the true label is obtained through annotation and will be used as a reference for comparing the model's prediction results during the training process.

[0047] The cross-entropy loss function L CE optimizes the model by comparing the P i,j,c predicted by the model and the true label The goal of the cross-entropy loss function L CE is to minimize the gap between the predicted probability and the true label, thereby improving the classification accuracy.

[0048] The loss function L Dice is then used to improve the fineness of the segmentation edge and can be designed as:

[0049]

[0050] Take the maximum value obtained by the output image pixel point (i, j) according to the class probability P i,j,c to generate a class mask for cell segmentation; represents the true value that the pixel point (i, j) belongs to class c.

[0051] Optionally, based on the segmented cell nucleus region, morphological operations are used to generate a background region, a foreground region, and an edge region, including:

[0052] First, perform a dilation operation to increase the pixel of the cell nucleus boundary for marking the potential background region;

[0053] Then, perform an erosion operation to reduce the pixel of the cell nucleus boundary for generating a high-confidence foreground region;

[0054] Mark the region with the difference between the foreground and the background as the edge region for subsequent segmentation.

[0055] Optionally, let the binary image of the cell nucleus instance be II, where II(x1, y1) = 1 represents the target region and II(x1, y1) = 0 represents the background, and x1, y1 represent the horizontal and vertical serial numbers of the pixels in the binary image II;

[0056] The increasing of the cell nucleus boundary pixels through the dilation operation includes:

[0057] Define the background dilation B k1 (II):

[0058]

[0059] where K k1 is a structuring element of size k1×k1, represents the dilation operation, and k1 is the dilation kernel size;

[0060] Set each target region to be marked with an independent label l, and the background is 0;

[0061] The decreasing of the cell nucleus boundary pixels through the erosion operation includes:

[0062] Define the instance erosion F m (II):

[0063] F m (II) = II - K m

[0064] where "-" represents the erosion operation here, and K m is a structuring element of size m×m, and m is the erosion operation coefficient;

[0065] The marking of the region with the difference between the foreground and the background as the edge region includes:

[0066] Calculate the total foreground region F all (II), which is the union of all instance foregrounds:

[0067]

[0068] Mark the region with the difference between the foreground and the background as the edge region U k1 ;

[0069] U k1 = B k1 (II) - F all (II).

[0070] Optionally, the calculating of the shortest Euclidean distance from each background region pixel to the nearest foreground region pixel and the segmentation of the cell image using the watershed algorithm includes:

[0071] Calculate each background region pixel P b(x b , y b ) to the nearest foreground region pixel P f (x f , y f ) is the shortest distance D(P b , P f ):

[0072]

[0073] For each background pixel P b , find all foreground pixel sets F(P b ) within its 5×5 neighborhood, and take the minimum value as the final distance d min (P b ), there is:

[0074]

[0075] If F(P b ) is empty, then d min (P b ) is set to infinity;

[0076] Define the distance map D(x a , y a ), where the value of each background pixel (x b , y b ) is equal to its distance to the nearest foreground pixel:

[0077]

[0078] Among them, Background represents the set of background pixels, and Foregound represents the set of foreground pixels;

[0079] Generate a distance map based on D(x a , y a ) to segment the cell image;

[0080] Segment the cell image using the watershed algorithm based on the labels of the foreground, background, edge region, and distance map:

[0081] Define the initial label M(x c , y c ) of the watershed algorithm, the foreground region label F = F all (II), the background region label B = B k1 (II), the edge region label E = U k1 , there is:

[0082]

[0083] Among them, the foreground region of the pixel (x c , y c ) is marked as 1, the background region is marked as 2, and the edge region is marked as 0;

[0084] Denote that each pixel (x e , y e ) in the unknown region E(x e , y e ) is assigned the nearest marked value, then there is:

[0085]

[0086] Optionally, the result optimization of the segmentation map corresponding to each cell instance label to obtain a filtered segmentation label map includes:

[0087] Perform connectivity analysis on the instance segmentation map corresponding to each cell instance label to generate a label map L(x f1 , y f1 ), let lf ∈ {0, 1, 2, …, N f}, lf represents the label of the connected region, and lf is a natural number greater than zero;

[0088] Define the area A lf of each connected region. The area A f1 , y f1 ) of each connected region is calculated by counting the number of pixels belonging to the label lf in the label map L(x lf . The calculation formula is:

[0089]

[0090] The area filtering condition is:

[0091]

[0092] Lf'(x f1 , y f1 ) is the filtered segmentation label map.

[0093] Optionally, the extraction of the outer contour of the nuclear region and the outer contour of the surrounding region of each cell from the segmentation label map and saving them to the corresponding result file includes:

[0094] Extract the outer contour of the nuclear region and the outer contour of the surrounding region of each cell from the segmentation label map based on the following formula:

[0095] Let the output label map be M'(x e , y e ), then from the label map:

[0096] Extract the outer contour C of the nuclear region of Ig in each cell nb (Ig) is:

[0097] C nb (Ig) = {(x g , y g ) | (x g , y g ) ∈ M'(x e , y e ) = 1}

[0098] The outer contour C of the surrounding region of Ig in each cell cb (Ig) is:

[0099] C cb (Ig) = {(x h , y h ) | (x h , y h ) ∈ M'(x e , y e ) = 2}

[0100] Record and save the corresponding results to a JSON file.

[0101] To solve the above technical problems, the technical solution of the present invention also provides a method for stitching the cell image segmentation results, including:

[0102] Obtain input data, where the input data includes: the original slice image file and the segmentation image result after detecting cells by the method of detecting cells based on the above-mentioned DistSegNet model image processing technology, and the original slice image file includes: several slice image blocks;

[0103] Horizontally stitch the image blocks in each row in sequence, while merging the overlapping region masks between the image blocks, removing redundant parts and optimizing the overlapping regions, load the image blocks in each row from the specified path and record the segmentation image results of the image blocks, check the overlapping situation of the masks in the overlapping regions, update the merged masks and remove redundant masks;

[0104] Use the IoU metric to measure the overlapping situation between two masks, and evaluate the overlapping situation of the masks based on the IoU metric;

[0105] Draw an image based on the overlapping situation of all masks to generate the stitched image of the cell image segmentation result.

[0106] Optionally, the step of using the IoU metric to measure the overlapping situation between two masks and evaluating the overlapping situation of the masks based on the IoU metric includes:

[0107] Evaluate the similarity between two mask polygons P1 and P2, and the calculation formula is as follows:

[0108]

[0109] Where: Area of Intersection represents the area of the overlapping region of the two mask polygons P1 and P2, and Area of Union represents the area of the union region of the two mask polygons P1 and P2;

[0110] For two mask polygons P1 and P2, calculate their IoU index. If IoU≥threshold, delete the mask Intersection with the smaller area; if IoU<threshold, calculate the Intersection part and perform the difference set operation: threshold is the preset index threshold;

[0111] P1' = P1 - Intersection

[0112] P2' = P2 - Intersection

[0113] Retain the non-overlapping regions on both sides.

[0114] To solve the above technical problems, the technical solution of the present invention also provides a system for detecting cells based on the image processing technology of the DistSegNet model, including: a preprocessing module, a first segmentation module, a second segmentation module, a third segmentation module, an optimization processing module, and an extraction module;

[0115] The preprocessing module is suitable for preprocessing the collected cell images;

[0116] The first segmentation module is suitable for segmenting the cell images based on the image processing technology of the DistSegNet model to generate the cell nucleus region;

[0117] The second segmentation module is suitable for generating the background region, the foreground region, and the edge region by using morphological operations based on the segmented cell nucleus region;

[0118] The third segmentation module is suitable for calculating the shortest Euclidean distance from each pixel in the background region to the nearest pixel in the foreground region, and segmenting the cell image by using the watershed algorithm;

[0119] The optimization processing module is suitable for optimizing the results of the segmentation map corresponding to each cell instance label to obtain a filtered segmentation label map;

[0120] The extraction module is suitable for extracting the outer contour of the nuclear region and the outer contour of the surrounding region of each cell from the segmentation label map and saving them to the corresponding result file.

[0121] To solve the above technical problems, the technical solution of the present invention also provides a system for stitching the results of cell image segmentation, which is characterized by including: an image input module, an image stitching module, an overlapping area processing module, and a contour merging and output module;

[0122] The image input module is adapted to obtain input data, and the input data includes: an original slice image file and a JSON file of the segmentation image result after detecting cells by the method of cell detection based on the above-mentioned DistSegNet model image processing technology. The original slice image file includes: a plurality of slice image blocks;

[0123] The image stitching module is adapted to horizontally stitch the image blocks in each row in sequence, and at the same time merge the overlapping area masks between the image blocks, remove redundant parts and optimize the overlapping areas, load the image blocks in each row and record the segmentation image results of the image blocks from the specified path, check the overlapping situation of the masks in the overlapping areas, update the merged masks and remove redundant masks;

[0124] The overlapping area processing module is adapted to use the IoU index to measure the overlapping situation between two masks, and evaluate the overlapping situation of the masks based on the IoU index;

[0125] The contour merging and output module is adapted to draw an image based on the overlapping situation of all masks, and generate a stitched image of the cell image segmentation result.

[0126] The beneficial effects of the technical solution of the present invention at least include:

[0127] The technical solution of the present invention adopts the DistGenSeg-Net deep learning network, uses the Pix2Pix GAN model for image enhancement, and through the image generation ability of the generative adversarial network, effectively improves the contrast and texture detail performance of the cell nucleus region, especially significantly suppresses noise interference in complex backgrounds, and provides higher-quality input images for the segmentation task. Subsequently, a modified U-Net model is used to segment the enhanced image. Through its encoder-decoder structure, multi-scale features are captured, and the image is segmented into three categories: the cell nucleus main body, the edge region, and the background. The above design significantly improves the accuracy and reliability of cell nucleus segmentation. Even in the case of cell density or adhesion, the cell contours can still be accurately extracted, improving the overall performance of cell detection and segmentation.

[0128] The technical solution of the present invention enhances the structural information of the background region by introducing the Euclidean distance transformation, provides a more accurate basis for foreground and background separation, and effectively improves the morphological authenticity of the cell detection result.

[0129] The technical solution of the present invention judges the overlapping area during the splicing process by introducing the IoU index, and combines parallel processing to effectively remove redundant parts, ensuring the integrity and accuracy of the spliced image. At the same time, the splicing speed and efficiency are improved, making the large-scale WSI image splicing more efficient and accurate. Description of the Drawings

[0130] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects and advantages of the present invention will become more obvious:

[0131] Figure 1 It is a schematic flowchart of a method for detecting cells based on the image processing technology of the DistSegNet model provided by the technical solution of the present invention;

[0132] Figure 2 It is a schematic diagram of the process steps for detecting and segmenting the cell nucleus of a cell image using the image processing technology of the DistSegNet model in the method for detecting cells based on the image processing technology of the DistSegNet model provided by the technical solution of the present invention;

[0133] Figure 3 It is a schematic flowchart of a method for splicing the segmentation results of cell images provided by the technical solution of the present invention;

[0134] Figure 4 It is a schematic structural diagram of a system for detecting cells based on the image processing technology of the DistSegNet model provided by the technical solution of the present invention;

[0135] Figure 5 It is a schematic structural diagram of a system for splicing the segmentation results of cell images provided by the technical solution of the present invention;

[0136] Figure 6 It is a schematic example diagram of a method for detecting cells based on the image processing technology of the DistSegNet model provided by the technical solution of the present invention;

[0137] Figure 7 It is a comparison schematic diagram of the cell segmentation experimental results generated by the method for detecting cells based on the image processing technology of the DistSegNet model provided by the technical solution of the present invention with the original image and the cell detection and processing results of the prior art respectively;

[0138] Figure 8 It is a schematic example diagram of a method for splicing the segmentation results of cell images provided by the technical solution of the present invention;

[0139] Figure 9Schematic diagram for comparing the stitching effect of the bright-field image segmentation result with the original image after the cell detection and stitching based on the DistSegNet model image processing technology in the technical solution of the present invention;

[0140] Figure 10 Schematic diagram for comparing the stitching effect of the fluorescence image segmentation result with the original image after the cell detection and stitching based on the DistSegNet model image processing technology in the technical solution of the present invention. Detailed implementation manners

[0141] In order to better present the technical solution of the present invention clearly, the present invention will be further described below with reference to the accompanying drawings.

[0142] The existing DistSegNet model (Distributed Semantic Segmentation Network) is a deep network for image semantic segmentation and has been applied to multiple fields including medical image analysis, remote sensing image processing, and autonomous driving. For example, in medical image analysis, it can be used to identify and segment lesion areas to assist doctors in diagnosis and treatment; in remote sensing image processing, it can be used to identify and segment target areas to provide support for applications such as military reconnaissance and environmental monitoring; in autonomous driving, it can be used to identify and segment roads and obstacles to assist vehicle navigation and autonomous driving.

[0143] The existing DistSegNet model is based on VGG-16 and adopts an encoder-decoder structure, with the characteristic of symmetric left and right network layers, capable of recording pooling positions, using indexed max-pooling upsampling, and being able to directly put data back to the corresponding positions during deconvolution, better retaining boundary feature information. However, the DistSegNet model needs to further improve the segmentation accuracy.

[0144] The technical solution of the present invention further combines the existing DistSegNet model with the convolutional neural network of deep learning to achieve more complex image segmentation tasks. An efficient encoder-decoder structure, attention mechanism, etc. can be adopted to enable the model to better capture semantic information in the image and improve the segmentation effect.

[0145] In an embodiment of the technical solution of the present invention, as Figure 1 shown, a method for detecting cells based on the DistSegNet model image processing technology is provided, including the following steps:

[0146] S10, preprocess the collected cell images.

[0147] According to step S10 of this embodiment, the collected cell images are Whole Slide Image (WSI) images. The preprocessing of the collected cell images includes the following processes: I. Segment the Whole Slide Image (WSI) image into image patches of size 512×512 (unit: pixel). Each sub-block image contains local cell information and retains its position information in the WSI global according to the naming format of {r}-{c}-{num}, where r represents the row index, c represents the column index, and num represents the patch number. II. Upload the cell slice image to be processed to the server and save the image to a specified folder. III. Convert the slice image into grayscale format, where: the slice image is divided into fluorescence images and bright-field images; after the fluorescence image is directly grayscaled, images with weak background signals (such as a too high proportion of pixels with a gray value less than 45) are filtered out according to the pixel brightness distribution, and the signal significant area is retained; for the bright-field image, the RGB image is converted to the HSI color model, the saturation component (S component) is extracted and enhanced to highlight the contrast of the cell area, and then the enhanced image is converted to a grayscale image for processing. IV. Standardize the grayscale image. V. Check whether the file size of each image is less than 5KB. If it meets the requirement, delete it.

[0148] Continuing to refer to Figure 1 , the method for detecting cells based on the DistSegNet model image processing technology in this embodiment further includes the steps:

[0149] S11, perform cell segmentation on the cell image based on the DistSegNet model image processing technology to generate a cell nucleus region.

[0150] In step S11 of this embodiment, the specific process of performing cell segmentation on the cell image based on the DistSegNet model image processing technology can refer to Figure 2 , including the steps:

[0151] Step S110, load the preprocessed cell image and input the preprocessed cell image into the Pix2Pix GAN model.

[0152] In step S110, specifically, the image after grayscale conversion in step S10 is input into the Pix2Pix GAN model. Among them, the used Pix2Pix GAN model is a deep learning image conversion model based on a conditional generative adversarial network, mainly composed of a generator and a discriminator. In this embodiment, the Pix2Pix GAN model is used to implement the image conversion task. In this step, the grayscale image is input into the Pix2Pix GAN model.

[0153] Continue to refer to Figure 2 , the specific process of cell segmentation of cell images based on the image processing technology of the DistSegNet model further includes:

[0154] Step S111, use the Pix2Pix GAN model to enhance the preprocessed cell image.

[0155] Specifically, the Pix2Pix GAN model includes a generator and a discriminator. In the Pix2Pix GAN model of step S111:

[0156] The design of its generator can combine traditional convolutional layers with the Res-CBAM module, use residual connections and attention mechanisms to enhance the ability to capture key information, and filter background noise at the same time. CBAM (Convolutional Block Attention Module) is a lightweight convolutional attention module that combines channel and spatial attention mechanism modules.

[0157] Let the input image be X, assuming the input image X ∈ H×W×3, where H is the image height and W is the image width. It is mapped to a feature map F(X) with C channels through the initial convolutional layer. The channel attention mechanism is to split the H×W×C feature map into H×W 1x1xC small feature maps. Each small feature map undergoes a MaxPool (maximum pooling) and an AvgPool (average pooling), then a shared weight operation, and finally added to obtain a 1x1xC channel attention, which is then multiplied by the input feature map, that is, H×W×C, to obtain the attention-mapped feature map M c (X), there is:

[0158]

[0159] Among them: σ (Sigma) represents an activation function, usually expressed as the Sigmoid function. The formula of the Sigmoid function can be:

[0160] sigma(x) = 1 / 1+e^{-x}

[0161] The Sigmoid function compresses the input value to the range [0,1], and is often used to generate the probability representation of weights or activation values;

[0162] The MLP is a Multi-Layer Perceptron, which is a neural network module composed of multiple fully connected layers (FC layers). The MLP is used to process the feature map after the pooling operation and generate the weights of channel attention. Usually, the MLP will receive the pooled feature map (such as Avg Pool(X) or MaxPool(X)) and generate an attention weight of 1×1×C (is this a multiplication sign?) through multiple fully connected layers;

[0163] W1 is usually the weight of the second fully connected layer in the MLP, which is used to generate channel attention. Its role is to convert the intermediate feature map processed by the MLP into a vector equal to the number of channels C;

[0164] W0 is the weight of the first fully connected layer in the MLP, which is used to process the pooled features Figure X avg c or X max c , and map it to an intermediate feature space;

[0165] X avg c is the result obtained by performing average pooling (AvgPool) on the input features Figure X . At each spatial position, average pooling is performed on the channel dimension to generate a feature map of 1×1×C;

[0166] X max c is the result obtained by performing max pooling (MaxPool) on the input features Figure X . Similar to X avg c , max pooling is performed on each spatial position to generate a feature map of 1×1×C.

[0167] The purpose of the algorithm represented by the above formula is to enable the model to weight the input feature map according to the importance of different channels by learning channel attention, so as to enhance useful information in the finally generated feature map M c (X).

[0168] The spatial attention mechanism is to obtain two feature maps of H×W×1 through max pooling and average pooling of the output result of channel attention, then splice the two feature maps through the Concat operation, and change them into a 1-channel feature map through a 7×7 convolution (f 7×7 ), and then obtain the feature map of spatial attention after being processed by the sigmoid function. Finally, multiply the output result by the input feature map to obtain the mapped feature map M s (X), there is:

[0169]

[0170] The meanings of the following parameters or functions need to be supplemented:

[0171] Among them:

[0172] Xavg S is the result obtained by performing average pooling (AvgPool) on the feature map output by the channel attention mechanism. Specifically, it performs average pooling on the feature map at each position (i.e., H×W×C) in the channel dimension to obtain a feature map of H×W×1, representing the average channel information at each spatial position.

[0173] Xmax S is the result obtained by performing max pooling (Max Pool) on the feature map output by the channel attention mechanism. Similarly, it performs max pooling on the feature map at each spatial position in the channel dimension to obtain a feature map of H×W×1, representing the maximum channel information at each spatial position.

[0174] In the spatial attention mechanism, Xavg S and Xmax S these two feature maps will go through a concatenation (Concat) operation and be merged into a feature map of H×W×2. Then, through a 7×7 convolution (f 7×7 ), the number of channels is compressed to 1, and finally, after being processed by σ (sigmoid activation function), the spatial attention feature map is generated. Finally, this spatial attention map is multiplied by the input feature map to obtain the weighted spatial feature map M s (X).

[0175] The purpose of this process is to enable the model to weight the spatial information of the feature map according to the importance of the spatial position, thereby enhancing the attention to key regions.

[0176] The discriminator is designed based on the multi-scale method of Transformer, dynamically adjusting the Patch size to capture local details and global morphology. The enhanced grayscale image Y∈H×W×1 is used as the input of the discriminator. The image Y is sliced into Patches of different sizes, namely P, 2P, 8P, and the features of each Patch are mapped to the D dimension through linear projection (i.e., Linearproj(Y)), and learnable position encoding PosEncoding is introduced:

[0177] S = LinearProj(Y) + PosEncoding

[0178] Each projected sequence S is used as input, and multi-scale features are extracted through Transformer blocks. For the high-resolution stage (e.g., P ≤ 16), grid self-attention is adopted, which is calculated only within the local grid to avoid the high cost of calculating global self-attention.

[0179] Specifically, the feature sequence is divided into multiple grids G, and self-attention operation Attention grid (Q, K, V):

[0180]

[0181] is performed within each grid, where Q, K, and V are matrices composed of queries, keys, and values in sequence, and M grid is a position-related mask matrix within the grid.

[0182] Among them:

[0183] T represents the transpose operation, which usually exchanges rows and columns of a matrix. In the formula, QK^T represents the multiplication of the query matrix Q and the transpose of the key matrix K, aiming to calculate the similarity between the query and the key.

[0184] d k is the dimension of the key vector, representing the number of columns of the key matrix K. In attention calculation, the result of QK^T is usually divided by the square root of d k , which is to prevent the calculation result from being too large in a larger dimension and ensure the numerical stability of the calculation.

[0185] Mgrid is a position-related mask matrix used to represent the relative relationship between different positions within the grid. The mask matrix is usually used to adjust or mask certain attention weights. When dividing the grid, the relationship between different positions can be modeled according to a specific structure.

[0186] V is the value matrix, representing the value associated with each key. During calculation, after calculating the similarity between the query matrix Q and the key matrix K, the result is applied to the value matrix V, and finally the weighted value matrix is obtained. This step is a key step in calculating the output in the standard self-attention mechanism.

[0187] Softmax is a commonly used normalization function that transforms each element of the input vector into a probability distribution. The formula is:

[0188] sigma(x) = 1 / (1 + e^{-x})

[0189] In the attention mechanism, the Softmax function is used to normalize the dot product result of the query and the key, making the sum of the attention weights at each position equal to 1, which is convenient for weighting the values.

[0190] For the low-resolution stage (e.g., P > 16), global self-attention is adopted to capture long-range dependencies Attention standard '(Q, K, V):

[0191]

[0192] The output of each Transformer block is S output :

[0193] S output S = MLP(LayerNorm(Attention(S))) + S

[0194] Where:

[0195] Attention here refers to the standard self-attention mechanism. By calculating the similarity between the query (Q) and the key (K), then applying Softmax for normalization, and finally performing a weighted sum on the value (V), the attention output is obtained. In the standard self-attention formula, the dot product of the query and the key is normalized and the attention weights corresponding to each query are obtained through Softmax.

[0196] S usually represents the input features or states, which can be the output from the previous layer. It serves as the input for the self-attention operation and undergoes layer normalization (LayerNorm) and subsequent MLP processing to obtain the final output. S can be understood as the input of the current Transformer block.

[0197] LayerNorm is a normalization method used in neural networks. It normalizes all features for each sample and is usually used to accelerate training and improve stability.

[0198] MLP is a multi-layer perceptron, referring to a network composed of multiple fully connected layers. Usually, MLP is used to perform non-linear transformations on the input features to extract higher-level representations. In this case, MLP is used to process the features after layer normalization and self-attention processing.

[0199] The output features S of each scale S output After average pooling, they are concatenated to form multi-scale fusion features S fusion :

[0200] S fusion S = Concat((AvgPool(S output )))

[0201] Among them, Concat is a concatenation operation used to connect multiple feature maps or tensors in a certain dimension. In this formula, AvgPool(Soutput) is the feature map obtained through average pooling, and the concatenation (concat) may form a multi-scale fused feature map Sfusion together with other features (such as features of different scales). The concatenation operation enables features of different scales to be combined, providing richer information.

[0202] Introduce a classification token and generate the discriminant result CLS through a global Transformer block (TransformerBlock). output :

[0203] CLS output = TransformerBlock(CLS token + S fusion )

[0204] Where:

[0205] CLS token is a special token used for classification tasks in the Transformer model. It is added to the input feature sequence and passed through each layer of the Transformer. Finally, this classification token will carry the global information of the entire input sequence and be used to generate the final output of the classification. CLS is the abbreviation of "classification" and represents the "classification token" in the classification task.

[0206] S fusion is the multi-scale fused feature, which means that the output features of multiple scales are merged together through the concatenation operation. It usually contains feature information from different scales or different processing stages. Through this fusion, the model can better handle details at different levels.

[0207] Finally, the discriminator outputs a probability value D(Y) for the enhanced image Y, indicating whether it is a real image:

[0208] D(Y)=σ(MLP(CLS output ))

[0209] Through the interactive training of the generator and the discriminator, the generator generates the enhanced image Y through a series of feature extraction and enhancement modules. The discriminator, based on the multi-scale Transformer structure, determines whether the Y output by the generator is a real image. Through adversarial training, it drives the generator to generate high-quality enhanced images that are more in line with the real data distribution. The enhanced image is finally passed as input to the subsequent segmentation stage.

[0210] In other embodiments, the generator of the Pix2Pix GAN model used in this step may specifically adopt a U-Net structure to convert the input contour map into a real image through encoding and decoding; the discriminator uses a conditional discriminator PatchGAN to judge the authenticity of the generated image under the condition of the contour map. The Pix2Pix GAN model uses L1 regularization to help improve the quality of the generated image, making the generated image clearer. At the same time, the model also adopts the Dropout technology in the generator model. By adding Dropout in the model layer of the generator, the generated data has a certain degree of randomness.

[0211] Continue to refer to Figure 2 , the specific process of cell segmentation of cell images based on the DistSegNet model image processing technology further includes the steps:

[0212] S112, using a classic U-Net model to segment the enhanced image;

[0213] In step S112, assume that the input enhanced grayscale image is Y, with dimensions of H×W. Three types of segmentation masks are output through the classic U-Net model, which respectively represent: Z1: the main body of the cell nucleus; Z2: the edge region; Z3: the background.

[0214] Specifically, the encoder-decoder structure of U-Net is set as:

[0215] The encoder extracts multi-scale features F of the input image Y through multiple convolutional layers and downsampling layers l+1 , there is:

[0216] F l+1 =σ(W l F l +b l ), l∈{1,2,...,L e}

[0217] Among them, W l is the convolutional kernel weight of the l-th layer, b l is the bias term, σ is the activation function (Re LU), l takes 1,2,...,L e , L e is the total number of layers of the encoder, and F l is the multi-scale feature of the current input image Y being extracted.

[0218] Downsampling is achieved through the max pooling operation Pool:

[0219]

[0220] Represents the downsampled image features of the output.

[0221] The decoder gradually restores the high-resolution features through upsampling and convolutional operations:

[0222]

[0223] Among them, UpSample is the upsampling operation, and Concat is the skip connection operation, which combines the output high-resolution encoder features F up l .

[0224] Among them:

[0225] F down l Represents the downsampled feature map of the l-th layer. The downsampling operation is usually completed through max pooling or convolution with a stride greater than 1, which is used to reduce the spatial size of the feature map, thereby extracting higher-level features. In the U-Net model of this embodiment, the encoder gradually extracts multi-scale feature information through downsampling, and these downsampled feature maps are usually passed to the decoder and concatenated with the upsampled feature maps of the decoder through skip connections to restore finer-grained spatial information.

[0226] The final segmentation output is that the output of the decoder passes through the last convolutional layer and combines with the Softmax activation function to generate three types of segmentation results c. c can take 1, 2, 3, which can respectively represent whether the pixel belongs to Z1 (nucleus main body), Z2 (edge region), and Z3 (background):

[0227]

[0228] Among them, S i,j,c is the original score that the pixel position (i, j) belongs to the category c, and P i,j,c is the probability distribution that the pixel position (i, j) belongs to the category c after Softmax normalization. exp is the power function of the base e of the natural logarithm, and K represents the total number of categories. Specifically, K is used as an accumulator to traverse all categories (here are 3 categories) and is used for the normalization process in the Softmax calculation.

[0229] The role of Softmax is to convert the original score S{i,j,c} of each pixel position into the probability that the pixel belongs to each category. The denominator part in the formula:

[0230] represents the sum of the exponential scores of all categories, ensuring that the result is a valid probability distribution, that is, the sum of the probabilities that each pixel belongs to different categories is 1.

[0231] In this scenario, K = 1, 2, 3, corresponding to the three categories of the cell nucleus body, the edge region, and the background respectively.

[0232] More specifically, the settings of the U-Net model parameters are as follows:

[0233] Assume that the input image size is H = 512, W = 512, the convolutional kernel size k = 3, the convolutional stride s = 1, the activation function is ReLU, and the upsampling method is bilinear interpolation. To optimize the segmentation result, define the combined loss function L seg :

[0234] L seg = L CE + λL Dice

[0235] where L CE is the cross-entropy loss function, L Dice is the Dice loss function, and λ is the loss weight.

[0236] More specifically, the cross-entropy loss function L CE is used to optimize the accuracy of class prediction and can be designed as:

[0237]

[0238] where the i value of the pixel point (i, j) can take values from 1 to H, the j value can take values from 1 to W, the class c can take values from 1 to 3, and P i,j,c is the probability distribution that the pixel position (i, j) belongs to class c after Softmax normalization. refers to the true class label of the pixel position (i, j), indicating the true value that this pixel point actually belongs to class c. More specifically,

[0239] In this formula, P i,j,c represents the predicted probability value that the pixel position (i, j) in the image belongs to the c-th class. Specifically:

[0240] (i, j) represents the pixel coordinate position in the image, referring to the pixel in the i-th row and the j-th column. c is the class index, indicating the class to which this pixel point belongs, which may be different types of cells, the background, or other regions. P i,j,c is the predicted probability that this pixel point belongs to class c.

[0241] Refers to the "true label" or "target label", that is, the true class label of the position (i, j) in the image, representing the true value that the pixel actually belongs to class c. Usually, the true label is obtained through annotation and will be used as a reference for comparing the model prediction results during training.

[0242] Cross-entropy loss function L CE By comparing the predicted P of the model i,j,c and the true label to optimize the model. The cross-entropy loss function L CE aims to minimize the gap between the predicted probability and the true label, thereby improving the classification accuracy.

[0243] Dice loss function L Dice is used to improve the fineness of the segmentation edge and can be designed as:

[0244]

[0245] The parameter value of the loss weight λ can be set according to actual needs to control the weights of cross-entropy and Dice loss. In this embodiment, it can be set to 0.01.

[0246] Through the above steps, the step process of cell segmentation of the cell image in step S11 of this embodiment is implemented, and finally the maximum value of P obtained by P according to the category probability is output to generate a category mask for cell segmentation. i,j,c to generate a category mask for cell segmentation.

[0247] Continue to refer to Figure 1 , the method for detecting cells based on the image processing technology of the DistSegNet model in this embodiment further includes the steps:

[0248] S12, based on the segmented nucleus region, use morphological operations to generate a background region, a foreground region, and an edge region.

[0249] In step S12, the process of the morphological operation is specifically as follows:

[0250] First, increase the nucleus boundary pixels through dilation operation to mark the potential background region. Use the binary map of the nucleus instance as II, where II(x1, y1)=1 represents the target region and II(x1, y1)=0 represents the background. x1 and y1 represent the horizontal and vertical serial numbers of the pixels in the binary map II, and then define the background dilation B k1 (II):

[0251]

[0252] where K k1 is a structuring element of size k1×k1, Denotes the dilation operation. When k1 is set to 3, it is a small-scale dilation for adjacent cell boundary detection; when k1 is set to 5, it is a large-scale dilation for cell cluster segmentation. k1 is the range dilation coefficient for adjacent cell boundary detection.

[0253] Then, the erosion operation is used to reduce the nuclear boundary pixels to generate a foreground region with high confidence. Using the given instance graph II of nuclear segmentation, where each target region is labeled with an independent label l and the background is 0. Define the instance erosion F m (II):

[0254] F m (II) = II - K m

[0255] where "-" represents the erosion operation here, and K m is a structuring element of size m×m, and m is set to 2. The total foreground region F all (II) is the union of all instance foregrounds:

[0256]

[0257] The region with the difference between the foreground and the background is marked as the edge region U k1 , which is used for subsequent segmentation.

[0258] U k1 = B k1 (II) - F all (II)

[0259] where k1 is the dilation kernel size.

[0260] Continuing to refer to Figure 1 , the method for detecting cells based on the DistSegNet model image processing technology in this embodiment further includes the steps:

[0261] S13, calculating the shortest Euclidean distance from each background region pixel to the nearest foreground region pixel, and segmenting the cell image using the watershed algorithm.

[0262] In step S13, the shortest distance from each background region pixel to the nearest foreground region pixel is calculated using the Euclidean Distance. The Euclidean distance is a way to calculate the straight-line distance between two points and is commonly used for calculating the distance between points in space. For a background pixel P b (x b , y b ) and P f (x f , y f ) in a two-dimensional space, the formula for calculating its Euclidean distance is:

[0263]

[0264] Therefore, for each background pixel P b , find all foreground pixel sets F(P b ) within its 5×5 neighborhood, take the minimum value as the final distance, and generate a distance map.

[0265] The closest distance d min is calculated by the formula:

[0266]

[0267] If F(P b ) is empty, then d min (P b ) is set to infinity.

[0268] Define the distance map D(x a , y a ), where the value of each background pixel (x b , y b ) is equal to its distance to the closest foreground pixel:

[0269]

[0270] where Background is defined as the set of background pixels, and Foreground is defined as the set of foreground pixels.

[0271] Generate the labels of the foreground, background, edge region, and distance map through the above steps.

[0272] In step S13, the specific process of segmenting the cell image using the watershed algorithm includes:

[0273] Use the labels of the foreground, background, edge region, and distance gradient map for cell segmentation. Consider the image as a topographic map, simulate the flow of water starting from the labeled foreground and background regions, expand along the path of the "water flow", form a dividing line between the "foreground" and the "background", and finally generate the segmented regions to refine the cell boundaries and assign the pixels in the unknown edge region to the foreground or background.

[0274] Among them, in the definition of the input data, the foreground region label (F = F all (II)): a binary image, where the foreground region is labeled with non-zero values and the background is labeled with 0. The background region label (B = B k1 (II)): a binary image, where the background region is labeled with non-zero values and the foreground is labeled with 0. The edge region (E = U k1 ): subtract the foreground from the dilated background, defined as the unknown pixel region. The distance gradient map (G = D(xa , y a )): It is defined as the shortest Euclidean distance map from background pixels to foreground pixels.

[0275] Define the initial marker M(x c , y c ), (x c , y c ) as the horizontal and vertical coordinates of the pixel, and integrate the foreground and background regions:

[0276]

[0277] Let the foreground region be marked as 1, the background region be marked as 2, and the edge region be marked as 0 (unknown region).

[0278] In this embodiment, the specific process of watershed segmentation is as follows:

[0279] Regard the distance gradient map G(x d , y d ) = D(x a , y a ) as the terrain height map, where the gradient value represents the "height" of the pixel: H(x d , y d ) = G(x d , y d ). Start simulating the expansion of water flow along the path with the lowest gradient from the marked foreground and background regions. Define the expansion condition of the water flow as for each pixel (x d , y d ) and its neighborhood N(x d ’, y d ’), find the neighborhood that satisfies the following conditions:

[0280] H(x d ’, y d ’) < H(x d , y d )

[0281] where N(x d , y d ’) is the neighborhood function of the pixel (x d , y d ), that is, in the watershed algorithm, it is used to define the neighborhood range of the current pixel. N is used to represent the set of neighborhood pixels of a pixel in the process of watershed expansion, and is used to judge the condition of water flow expansion. When the water flow meets from different marked regions (such as foreground and background), a dividing line is formed, and the pixel is assigned to the nearest marked region.

[0282] For each pixel (x e , y e ) in the unknown region E(xe , y e ) is assigned the nearest marker value (foreground or background):

[0283]

[0284] The output of the watershed algorithm is the marker map M'(x e , y e ), indicating the belonging of each pixel (x e , y e ). The dividing line represents the refinement of the cell boundary through the junction pixels of the marked regions.

[0285] Continue to refer to Figure 1 , the method for detecting cells based on the DistSegNet model image processing technology in this embodiment further includes the steps:

[0286] S14. Optimize the results of the segmentation map corresponding to each cell instance label to obtain a filtered segmentation marker map.

[0287] In step S14, the result optimization of the segmentation map corresponding to the cell instance label specifically includes the following process:

[0288] Perform connectivity analysis on the instance segmentation map corresponding to each cell instance label to generate a marker map L(x f1 , y f1 ), let lf ∈ {0, 1, 2, …, N f}, lf represents the label of the connected region, and lf is a natural number greater than zero.

[0289] Define the area A of each connected region lf , and calculate the area A of each connected region by counting the number of pixels belonging to the label lf in the marker map L(x f1 , y f1 ). The calculation formula is: lf .

[0290]

[0291] Define the area upper limit threshold A max as 600 to filter out too large regions. The area filtering condition is:

[0292]

[0293] Lf'(x f1 , y f1 ) is the filtered segmentation marker map. Through the operation of step S14, multiple cell nuclei can be prevented from being misrecognized as one.

[0294] During steps S13 and S14, (xa , y a ), (x b , y b )(x c , y c ), (x d , y d ), (x e , y e ), (x f , y f ), (x f1 , y f1 ), (x d ’, y d ’) is the neighborhood pixel of (x d , y d ).

[0295] Continue to refer to Figure 1 , the method for detecting cells based on the DistSegNet model image processing technology in this embodiment further includes the steps of:

[0296] S15, extract the outer contour of the nuclear region and the outer contour of the surrounding region of each cell from the segmentation label map, and save them to the corresponding result file.

[0297] Specifically, in step S15,

[0298] The outer contour of the nuclear region and the outer contour of the surrounding region of each cell can be extracted from the segmentation label map based on the following formula:

[0299] Let the output label map be M'(x e , y e ), then from the label map:

[0300] Extract the outer contour C nb (Ig) of the nuclear region of each cell Ig as:

[0301] C nb (Ig) = {(x g , y g ) | (x g , y g ) ∈ M'(x e , y e ) = 1}

[0302] The outer contour C cb (Ig) of the surrounding region of each cell Ig is:

[0303] C cb (Ig) = {(xh , y h ) | (x h , y h ) ∈ M'(x e , y e ) = 2}

[0304] Record and save the corresponding results to a JSON file. In step S15: (x g , y g ) represents the pixel values of the nuclear region contour image of each cell Ig, which are respectively the x - coordinate value and y - coordinate value of the pixel points of the nuclear region contour image in the image coordinate system in sequence; (x h , y h ) represents the pixel values of the outer contour image of the surrounding region of each cell Ig, which are respectively the x - coordinate value and y - coordinate value of the pixel points of the outer contour image of the cell surrounding region in the image coordinate system in sequence.

[0305] The JSON file contains the unique identifier of each cell inst, incrementing from 1, the set x of the x - coordinates of the nuclear region contour points x e , the set y of the y - coordinates of the nuclear region contour points y e , the set x of the x - coordinates of the surrounding region contour points x h , the set y of the y - coordinates of the surrounding region contour points y h , for subsequent analysis.

[0306] Based on the image - processing technology of the DistSegNet model in this embodiment for cell detection, the method can finally obtain the cell image segmentation results, that is, including: the outer contour of the nuclear region and the outer contour of the surrounding region C nb (Ig), and the outer contour of the surrounding region C cb (Ig) of each cell Ig, and further record and save the corresponding results of the inner and outer contours C nb (Ig) of the cell nucleus and the contour C cb (Ig) of the cell surrounding region to a JSON file.

[0307] In this embodiment, the processing process of the steps of this embodiment can be further batch - processed and performance - optimized, specifically including parallel processing and logging.

[0308] More specifically, the multiprocessing.Pool can be used to concurrently execute the DistSegNet segmentation tasks, with each process processing different images to improve the computing efficiency. At the same time, the time consumption of each step can be statistically analyzed and logged to facilitate performance monitoring and optimization.

[0309] The above technical solution of this embodiment uses the DistGenSeg-Net deep learning network, enhances images using the Pix2Pix GAN model, and effectively improves the contrast and texture details of the cell nucleus region through the image generation ability of the generative adversarial network. Especially in complex backgrounds, it significantly suppresses noise interference and provides higher-quality input images for the segmentation task. Subsequently, the modified U-Net model is used to segment the enhanced image, and its encoder-decoder structure captures multi-scale features to segment the image into three categories: the main body of the cell nucleus, the edge region, and the background. The above design significantly improves the accuracy and reliability of cell nucleus segmentation. Even in cases where cells are dense or adhered, it can still accurately extract cell contours and improve the overall performance of cell detection and segmentation.

[0310] Moreover, this embodiment also enhances the structural information of the background region by introducing the Euclidean distance transform, provides a more accurate basis for foreground and background separation, and effectively improves the morphological authenticity of cell detection results.

[0311] An embodiment of the technical solution of the present invention also provides a method for stitching the segmentation results of cell images as shown in Figure 3 and includes the following steps:

[0312] Step S20, obtain input data, where the input data includes: the original slice image file and a JSON file recording the corresponding segmentation image results after cell detection. The original slice image file includes: several slice image blocks.

[0313] Among them, the input data includes the original slice image file cut from the Whole Slide Imaging (WSI) image and the JSON file of the corresponding segmentation image information after cell detection. The original slice is a small tile cropped from the whole slide image (Whole Slide Image, WSI) taken by a microscope. The file format is {r}-{c}-{num}.png, and each file name represents its position in the whole slice, where r represents the row index (Row Index), c represents the column index (Column Index), and num represents the tile number. The file storage path is specified by the variable imgs_dir_path. The Json file is the segmentation result of each image block, with the file format of *.json, and each file name is the same as that of its corresponding image file, only the extension is different. The file storage path is specified by the variable json_dir_path. Each JSON file contains the segmentation results of one or more cells, and only the x_cb and y_cb data are used in this task.

[0314] Continue to refer to Figure 3The method for splicing the cell image segmentation results in this embodiment also includes:

[0315] Step S21, horizontally splicing the image blocks of each row in order, merging the overlapping area masks between the image blocks, removing redundant parts and optimizing the overlapping area, loading the image blocks of each row from the specified path and recording the image segmentation results of the image blocks, checking whether the masks are overlapping in the overlapping area, updating the merged masks and removing redundant masks.

[0316] It should be noted that, in this embodiment:

[0317] The segmentation result of each image block is the cell instance segmentation map within the block area, that is, the instance label (l k ). The segmentation results are usually saved in the form of images, where the pixel values ​​correspond to the unique identity of the cell instance or the background label. The overlap mask refers to the pixel area shared between two adjacent image blocks. This area may contain segmentation results from both blocks and needs to be optimized.

[0318] The mask records the instance label of each pixel in the segmentation result. If the instance labels segmented from different image blocks in the overlapping area conflict, trade-offs and optimizations are required.

[0319] Load the original image and its corresponding segmentation result for each image block from the specified path.

[0320] The segmentation result provides an instance label for each pixel, which provides a basis for subsequent mask merging. For adjacent image blocks, determine the range of their horizontal overlapping areas. In the overlapping area, generate a mask for each pixel. The mask contains the following information: Instance label: records the segmentation result label of each pixel (l k ); Source tag: records whether the pixel comes from block (A) or block (B); Check for conflicts: In overlapping areas, it may happen that two image blocks assign different labels to the same pixel (l k ), a conflict occurs; a label is assigned in one image block, and a label is not assigned in the other image block. If the segmentation result of a pixel only exists in one block, the result is directly retained. If the segmentation result of the pixel exists in both blocks: a trade-off is made based on the iou threshold comparison process

[0321] The concatenated mask is a global segmentation result map, where the label of each pixel represents its instance segmentation affiliation. Overlapping areas are optimized and conflict-processed to ensure that there are no redundant labels.

[0322] Specifically, in step S21, the image patches in each row can be horizontally stitched in sequence. Meanwhile, the overlapping region masks between images are merged, redundant parts are removed, and the overlapping regions are optimized. Load the image patches in each row and their masks (JSON files) from the specified path. Determine the overlapping region (with a width overlap of 256) between every two adjacent image patches. Within the overlapping region, check whether there is an intersection in the masks according to the intersection over union ratio of the overlapping part area. If the ratio is greater than the set threshold of 0.05, keep the larger mask and remove the smaller mask.

[0323] The updated masks are merged into the final result, and the processed image patches are stitched into a whole row in sequence. The horizontally stitched row images are vertically stitched into a complete large image in sequence, and at the same time, the vertical overlapping region masks between rows are merged. Determine the vertical overlapping region (with a height overlap of 256) between every two rows. Similarly, check whether there is an intersection in the masks within the overlapping region, update the merged masks, and remove redundant masks.

[0324] Continue to refer to Figure 3 , the method for stitching the cell image segmentation results in this embodiment further includes:

[0325] Step S22, use the IoU metric to measure the overlap between two masks, and evaluate the overlap of the masks based on the IoU metric.

[0326] According to step S22, the Intersection over Union (IoU) metric can be used to measure the overlap between two masks. IoU (Intersection over Union) is used to evaluate the similarity between two mask polygons, and the calculation formula is:

[0327]

[0328] Among them: Area of Intersection represents the area of the overlapping region of the two mask polygons, and Area of Union represents the area of the union region of the two mask polygons. For two mask polygons P1 and P2, calculate their IoU metric. Here, the threshold is set to 0.05. If IoU≥threshold, delete the mask Intersection with the smaller area. If IoU<threshold, calculate the Intersection part and perform the difference operation:

[0329] P1' = P1 - Intersection

[0330] P2' = P2 - Intersection

[0331] By retaining the non-overlapping regions on both sides and introducing the IoU metric, the overlapping situation of the masks can be evaluated more precisely, enhancing the reliability and accuracy of mask processing during the stitching process.

[0332] Continue to refer to Figure 3 , the method for stitching the cell image segmentation results in this embodiment further includes:

[0333] Step S23, draw an image based on the overlapping situation of all masks to generate a stitched image of the cell image segmentation results.

[0334] In the final stitched image, all mask information is unified and integrated and drawn on the image to generate an output image containing the complete content and the segmentation contours of each target region. Integrate multiple mask point sets into one image, ensuring that the masks do not overlap with each other while maintaining the consistency of spatial positions. The final stitched image is final_img, and the merged mask point set is merged_masks, in the format {mask_id:[(x1,y1),(x2,y2),...]}. Traverse the masks in merged_masks and use the cv2.circle method to draw the points on the image. The drawing color is set to red (0,0,255). Save the drawn image final_img_with_masks as a file for subsequent display and analysis.

[0335] In this embodiment, by introducing the IoU metric to judge the overlapping regions during the stitching process and combining parallel processing, redundant parts are effectively removed, ensuring the integrity and accuracy of the stitched image. At the same time, the stitching speed and efficiency are improved, making the stitching of large-scale WSI images more efficient and accurate.

[0336] Based on the above embodiments, the technical solution of the present invention further provides a Figure 4The system for detecting cells based on the image processing technology of the DistSegNet model shown in the figure includes: a preprocessing module, a first segmentation module, a second segmentation module, a third segmentation module, an optimization processing module, and an extraction module. Among them: the preprocessing module is adapted to execute step S10, that is, preprocess the collected cell images; the first segmentation module is adapted to execute step S11, that is, perform cell segmentation on the cell images based on the image processing technology of the DistSegNet model to generate a cell nucleus region; the second segmentation module is adapted to execute step S12, that is, based on the segmented cell nucleus region, use morphological operations to generate a background region, a foreground region, and an edge region; the third segmentation module is adapted to execute step S13, that is, calculate the shortest Euclidean distance from each background region pixel to the nearest foreground region pixel, and use the watershed algorithm to segment the cell images; the optimization processing module is adapted to execute step S14, that is, optimize the segmentation map corresponding to each cell instance label to obtain a filtered segmentation label map; the extraction module is adapted to execute step S15, that is, extract the outer contour of the intranuclear region and the outer contour of the surrounding region of each cell from the segmentation label map and save them to the corresponding result file.

[0337] Based on the above embodiments, the technical solution of the present invention also provides a Figure 5 system for stitching the segmentation results of cell images shown in the figure, including: an image input module, an image stitching module, an overlapping region processing module, a contour merging and output module. Among them: the image input module is adapted to execute step S20, that is, obtain input data, and the input data includes: an original slice image file and a JSON file recording the corresponding segmentation image results after cell detection. The original slice image file includes: several slice image blocks; the image stitching module is adapted to execute step S21, that is, horizontally stitch the image blocks in each row in sequence, and at the same time merge the overlapping region masks between the image blocks, remove redundant parts and optimize the overlapping regions, load the image blocks in each row and the segmentation image results of the image blocks from the specified path, check the overlapping situation of the masks in the overlapping regions, update the merged masks and remove redundant masks; the overlapping region processing module is adapted to execute step S22, that is, use the IoU index to measure the overlapping situation between two masks, and evaluate the overlapping situation of the masks based on the IoU index; the contour merging and output module is adapted to execute step S23, that is, draw an image based on the overlapping situation of all masks to generate a stitched image of the segmentation results of the cell images.

[0338] The specific execution process of the above system can refer to the content recorded in the above embodiments of the technical solution of the present invention, and will not be elaborated here.

[0339] Based on the method for detecting cells using the image processing technology of the DistSeNet model provided in this embodiment, the technical solution of the present invention provides an application example as shown in Figure 6 as follows.

[0340] Figure 6 The method for detecting cells using the image processing technology of the DistSegNet model as shown in includes the following steps:

[0341] I. Image preprocessing.

[0342] The specific process of image preprocessing is as follows:

[0343] 1. Cut the whole slide image (WSI) image:

[0344] Cut the WSI image through a sliding window. The size of each image block is 512×512, and there is an overlapping area of 256 pixels between the tiles. At the same time, record the position information of each image block in the WSI global.

[0345] 2. Upload the slice image:

[0346] Upload the cell slice image to be processed to a specified folder for processing.

[0347] 3. Image grayscale conversion and enhancement:

[0348] Directly perform grayscale conversion on the fluorescence image, and filter out images with weak signals according to the pixel brightness distribution (for example, images with a too high proportion of pixels with a grayscale value lower than 45 will be filtered), and retain the significant part of the signal.

[0349] For bright-field images, first convert the RGB image to the HSI color space model, extract and enhance the saturation (S component) to improve the contrast of the cell region, and then convert the enhanced image to a grayscale image.

[0350] 4. Image normalization:

[0351] Perform normalization processing on the grayscale image to unify the dynamic range of the image data.

[0352] 5. File filtering:

[0353] Delete invalid image files with a file size less than 5KB.

[0354] II. Cell segmentation based on the DistGenSeg-Net deep learning model

[0355] Use the DistGenSeg-Net model to complete cell segmentation. The specific steps are as follows:

[0356] 1. Image loading and input:

[0357] The grayscale image is input into the Pix2Pix GAN model for preprocessing.

[0358] 2. Image enhancement:

[0359] The generator adopts a combination of traditional convolutional layers and a residual CBAM module (Res-CBAM). Through residual connections and attention mechanisms, it enhances the ability to capture key regions and effectively suppresses background noise.

[0360] The discriminator is based on a multi-scale Transformer architecture, dynamically adjusting the Patch size to capture local details and global morphology. In the high-resolution stage, the Grid Self-Attention mechanism is used to enhance the ability to analyze complex regions.

[0361] The enhanced image is used as the input for subsequent segmentation.

[0362] 3. Initial segmentation:

[0363] The classical U-Net model is used for initial segmentation. Through the encoder-decoder structure, it captures multi-scale features, divides the image into three categories: nucleus, edge region, and background, and outputs the masked image of the initial segmentation.

[0364] III. Morphological operations.

[0365] Morphological methods are used to further refine the segmentation of the foreground and background. The specific steps are as follows:

[0366] 1. Generating foreground and background regions:

[0367] The nucleus boundary is expanded through dilation operations to mark potential background regions;

[0368] The nucleus boundary is shrunk through erosion operations to generate a foreground region with high confidence;

[0369] The difference region between the foreground and background is marked as the edge region.

[0370] 2. Calculating the distance map:

[0371] For each background pixel, calculate the shortest distance to the nearest foreground pixel. In the 5×5 neighborhood around each background pixel, find all foreground pixels and calculate the Euclidean distance, taking the minimum value to generate the distance map.

[0372] The formula for calculating the Euclidean distance is:

[0373]

[0374] Among them, (x1, y1) and (x2, y2) are the coordinates of two pixel points in the two-dimensional plane of the image.

[0375] IV. Distance transformation and watershed algorithm segmentation.

[0376] Regarding the image as a topographic map, by simulating the spreading behavior of water flow, starting from the foreground and background marker points, "water flow paths" are generated, and a dividing line is formed between the foreground and the background, finally completing the refinement of the cell boundary.

[0377] V. Region processing and repair.

[0378] Perform connectivity analysis on the segmented image, eliminate segmented regions with too large an area to avoid misidentifying multiple cell nuclei as one; clean small artifact regions with an area less than 5 KB to reduce the interference of invalid segmentation.

[0379] VI. Result recording and saving.

[0380] Record the contour point set of each cell, including the coordinates inside and around the cell nucleus; save the segmentation result as a JSON file for subsequent analysis.

[0381] This application example can also use a multiprocessing pool (multiprocessing.Pool) to implement parallel processing of the segmentation task. Each process is responsible for a different image, significantly improving the processing efficiency.

[0382] The specific steps of this application example can refer to the above embodiments, and this application example will not be elaborated here.

[0383] Based on the above application example of the method for detecting cells using the DistSegNet model-based image processing technology, a schematic diagram showing the comparison of the segmentation results of cell images with the existing technology cell detection and segmentation results can be referred to Figure 7 , and it can be seen that the cell detection results under the technical solution of the present invention can have a clearer and more accurate effect diagram.

[0384] Based on the method for stitching the segmentation results of cell images provided in this embodiment, the technical solution of the present invention provides an application example as Figure 8 shown.

[0385] Figure 8 The method for stitching the segmentation results of cell images shown in

[0386] includes sequentially calling the following modules: an image input module, an image stitching module, an overlapping region processing module, and a contour merging and output module. Among them:

[0387] Image stitching module: Stitch the small image patches in each row according to the position order and process the overlapping regions between adjacent patches.

[0388] Overlapping region processing module: Measure the overlapping degree of the masks by the Intersection over Union (IoU) metric and decide whether to retain or perform a difference set operation on the masks according to a threshold. Specifically, in image processing and mask merging, IoU measures the ratio of the area of intersection of two regions (usually masks or bounding boxes) to the area of their union. The specific calculation formula is as follows:

[0389]

[0390] If the proportion of the overlapping area of the two masks in the area of the smaller mask exceeds the threshold, remove the smaller mask; otherwise, retain the non-overlapping parts on both sides. Finally, vertically stitch the horizontally stitched row images to form a complete large image.

[0391] Contour merging and output module: Integrate the mask contours in the stitched image and draw them on the final image to generate the final stitched image containing the complete content and the mask of each target region.

[0392] For the specific implementation manner of this application example module, reference can also be made to other embodiments of the technical solution of the present invention, which will not be elaborated here.

[0393] In the application example of the method for stitching the cell image segmentation results of the technical solution of the present invention, by introducing the IoU metric to judge the overlapping regions in the stitching process and combining parallel processing, redundant parts are effectively removed, ensuring the integrity and accuracy of the stitched image. At the same time, the stitching speed and efficiency are improved, making the stitching of large-scale WSI images more efficient and accurate. For the effects of stitching the segmentation results of bright-field images and fluorescence images in this application example, reference can be made to Figure 9 and Figure 10 . It can be seen that whether it is a bright-field image or a fluorescence image, using the technical solution of the present invention to stitch the cell image segmentation results can ensure the integrity and accuracy of the stitched image.

[0394] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific implementation manners, and those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A method for detecting cells based on DistSegNet model image processing technology, characterized in that: include: Preprocessing the collected cell images; Based on the DistSegNet model image processing technology, cell images are segmented to generate cell nucleus regions; Based on the segmented cell nucleus region, morphological operations are used to generate background region, foreground region and edge region; Calculating the shortest Euclidean distance from each background area pixel to the nearest foreground area pixel, and segmenting the cell image using a watershed algorithm; Optimize the segmentation map corresponding to each cell instance label to obtain a filtered segmentation label map; The outer contour of the nuclear region and the outer contour of the surrounding region of each cell are extracted from the segmentation mark map and saved in the corresponding result file.

2. The method for detecting cells based on the DistSegNet model image processing technology according to claim 1, characterized in that: The collected cell image is a full-view slice image, and the preprocessing of the collected cell image includes: Segmenting the full-view slice image into a plurality of image blocks, each of which contains local cell information, and retaining the position information of the graphic block in the global full-view slice image according to a preset naming format; Upload the cell slice images to be processed to the server, save the images to the specified folder, and convert the slice images into grayscale format.

3. The method for detecting cells based on the DistSegNet model image processing technology according to claim 1, characterized in that: The cell image is segmented based on the DistSegNet model image processing technology to generate a cell nucleus region, including: Loading the preprocessed cell image, and inputting the preprocessed cell image into the Pix2Pix GAN model; Performing image enhancement on the preprocessed cell image using a Pix2Pix GAN model; The enhanced image is segmented using the classic U-Net model to output the maximum value of the category probability of the segmented image pixels to generate a category mask for cell segmentation.

4. The method for detecting cells based on the DistSegNet model image processing technology as claimed in claim 3, characterized in that: The Pix2Pix GAN model includes: a generator and a discriminator; and using the Pix2Pix GAN model to perform image enhancement on the preprocessed cell image includes: By combining traditional convolutional layers with the Res-CBAM module, the residual connection and attention mechanism are used to enhance the ability to capture key information while filtering background noise to design the generator of the Pix2Pix GAN model; Based on the Transformer multi-scale method, the patch size is dynamically adjusted to capture local details and global morphology, so that the enhanced image can be used as the input of the discriminator. The input grayscale image is divided into patches of different sizes, and the features of each patch are mapped to the preset dimension through linear projection. The learnable position encoding PosEncoding is introduced to design the discriminator of the Pix2Pix GAN model. Through interactive training of the generator and the discriminator, the generator generates the enhanced image through feature extraction and enhancement modules, and the discriminator determines whether the enhanced image output by the generator is a real image based on the multi-scale Transformer structure. Through adversarial training, the generator is driven to generate high-quality enhanced images that are more in line with the real data distribution, and finally the enhanced image is output.

5. The method for detecting cells based on the DistSegNet model image processing technology as claimed in claim 3, characterized in that: The method uses the classic U-Net model to segment the enhanced image to output the maximum value of the pixel points of the segmented image according to the category probability to generate a category mask for cell segmentation, including: Assume that the enhanced grayscale image output by the Pix2Pix GAN model to the U-Net model is Y, with a dimension of H×W. The classic U-Net model outputs three types of segmentation masks, which respectively represent: Z1 is the main body of the cell nucleus, Z2 is the edge area, and Z3 is the background; then the structure of the U-Net is set as: An encoder is provided, which is suitable for extracting multi-scale features F of the input image Y through multiple convolutional layers and downsampling layers l+1 ,have: F l+1 =σ(W l F l +b l ),l∈{1,2,...,L e } Among them, W l is the convolution kernel weight of the lth layer, b l is the bias term, σ is the activation function, and l is 1, 2, ..., L e , L e is the total number of encoder layers, F l is the multi-scale feature of the currently extracted input image Y; The encoder is also suitable for downsampling through a maximum pooling operation Pool: Represents the output downsampled image features; There is a decoder that is suitable for gradually recovering high-resolution features through upsampling and convolution operations: Among them, UpSample is the upsampling operation, Concat is the skip connection, and the output high-resolution encoder feature F is combined down l ; The decoder is also suitable for outputting the result after passing through the last convolution layer, and generating three types of segmentation results c in combination with the activation function. c can be 1, 2, and 3, which can represent whether the pixel belongs to Z1, Z2, and Z3 respectively: Among them, S i,j,c is the original score of the image pixel position (i, j) belonging to category c, and the output result P i,j,c is the probability distribution of the pixel position (i, j) belonging to category c after being normalized by the activation function.

6. The method for detecting cells based on the DistSegNet model image processing technology as claimed in claim 5, characterized in that: The method uses the classic U-Net model to segment the enhanced image to output the maximum value of the pixel points of the segmented image according to the category probability to generate a category mask for cell segmentation, and also includes: Define the combined loss function L seg : THE seg =L CE +λL Dice Among them, L CE is the cross entropy loss function, L Dice is the Dice loss function, λ is the loss weight, and the parameter value of the loss weight λ is preset; Use the cross entropy loss function L CE Optimize output result P i,j,c The accuracy of category prediction can be designed as follows: Loss function L Dice It is used to improve the fineness of the segmentation edge and can be designed as follows: Output image pixel (i, j) according to category probability P i,j,c The maximum value is obtained to generate the class mask for cell segmentation; Indicates the true value that the pixel (i, j) belongs to category c.

7. The method for detecting cells based on the DistSegNet model image processing technology according to claim 1, characterized in that: The method of using morphological operations to generate a background region, a foreground region, and an edge region based on the segmented cell nucleus region includes: First, the cell nucleus boundary pixels are increased through dilation operation to mark the potential background area; Then, the cell nucleus boundary pixels are reduced through corrosion operation to generate high-confidence foreground areas; The area where the foreground and background differ is marked as the edge area for subsequent segmentation.

8. The method for detecting cells based on the DistSegNet model image processing technology as claimed in claim 7, characterized in that: Let the binary image of the cell nucleus instance be II, where II(x1,y1)=1 represents the target area, II(x1,y1)=0 represents the background, and x1,y1 represent the horizontal and vertical numbers of the pixels in the binary image II; The step of increasing the cell nucleus boundary pixels by dilation operation comprises: Define background expansion B k1 (II): Where K k1 is a structuring element of size k1×k1, represents the dilation operation, k1 is the dilation kernel size; Set each target area to be marked with an independent label l and the background to be 0; The method of reducing the cell nucleus boundary pixels by the corrosion operation includes: Define the instance corrosion F m (II): F m (II)=II-K m Here, "-" indicates the corrosion operation, K m is a structural element of size m×m, where m is the corrosion operation coefficient; The step of marking the area where the foreground and the background differ as the edge area comprises: Calculate the total foreground area F all (II) is the union of all instance foregrounds: The area where the foreground and background differ is marked as the edge area U k1 ; U k1 =B k1 (II)-F all (II).

9. The method for detecting cells based on the DistSegNet model image processing technology according to claim 1, characterized in that: The method of calculating the shortest Euclidean distance from each background area pixel to the nearest foreground area pixel and segmenting the cell image using a watershed algorithm comprises: Calculate each background area pixel P b (x b ,y b ) to the nearest foreground pixel P f (x f ,y f )'s shortest distance D(P b ,P f ): For each background pixel P b , find all foreground pixel sets F(P b ), take the minimum value as the final distance d min (P b ),have: If F(P b ) is empty, then d min (P b ) is set to infinity; Define the distance graph D(x a ,y a ), where each background pixel (x b ,y b ) is equal to its distance to the nearest foreground pixel: Among them, Background represents the background pixel set, and Foregound represents the foreground pixel set; Based on D(x a ,y a ) generating a distance map to segment the cell image; The cell image is segmented using a watershed algorithm based on the markings of foreground, background, edge region, and distance map: Define the initial mark M(x) of the watershed algorithm c ,y c ), the foreground area is marked with F = F all (II), background area label B = B k1 (II), edge region marking E=U k1 ,have: Among them, the pixel (x c ,y c ) is marked as 1 for the foreground area, 2 for the background area, and 0 for the edge area; Let the unknown area E(x e ,y e ) for each pixel (x e ,y e ) is assigned the most recent tag value, then:

10. The method for detecting cells based on DistSeg model image processing technology according to claim 1, characterized in that: The result optimization of the segmentation map corresponding to each cell instance label to obtain a filtered segmentation label map includes: Perform connectivity analysis on the instance segmentation map corresponding to each cell instance label to generate a label map L(x f1 ,y f1 ), let lf∈{0,1,2,…,N f }, lf represents the label of the connected region, and lf is a natural number greater than zero; Define the area A of each connected region lf , by statistically marking the graph L(x f1 ,y f1 The area A of each connected region is calculated by the number of pixels belonging to label lf in lf , the calculation formula is: The area filter conditions are: Lf'(x f1 ,y f1 ) is the filtered segmentation label map.

11. The method for detecting cells based on the DistSegNet model image processing technology according to claim 1, characterized in that: The process of extracting the outer contour of the nuclear region and the outer contour of the surrounding region of each cell from the segmentation mark map and saving them to the corresponding result file includes: The outer contour of the nuclear region and the outer contour of the surrounding region of each cell are extracted from the segmentation mark map based on the following formula: Let the output labeled graph be M'(x e ,y e ), then from the labeled graph: Extract the outer contour of the nuclear region C of each cell Ig nb (Ig) is: C nb (Ig)={(x g ,y g )|(x g ,y g )∈M'(x e ,y e )=1} The outer contour of the area surrounding each cell Ig cb (Ig) is: C cb (Ig)={(x h ,y h )|(x h ,y h )∈M'(x e ,y e )=2} Record and save the corresponding results to a JSON file.

12. A method for splicing cell image segmentation results, characterized in that: include: Obtain input data, where the input data includes: an original slice image file and a segmentation image result after detecting cells by the method for detecting cells based on the DistSegNet model image processing technology as described in any one of claims 1 to 11. The original slice image file includes: a plurality of slice image blocks; Horizontally splice the image blocks in each row in sequence, and at the same time merge the overlapping region masks between the image blocks, remove redundant parts and optimize the overlapping regions. Load the image blocks in each row and record the segmentation image result of the image block from the specified path, check the overlapping situation of the masks in the overlapping region, update the merged mask and remove the redundant masks; Use the IoU metric to measure the overlapping situation between two masks, and evaluate the overlapping situation of the masks based on the IoU metric; Draw an image based on the overlapping situation of all masks to generate a spliced image of the cell image segmentation result.

13. The method for splicing cell image segmentation results according to claim 12, characterized in that: The using the IoU metric to measure the overlapping situation between two masks and evaluating the overlapping situation of the masks based on the IoU metric includes: Evaluate the similarity between two mask polygons P1 and P2, and the calculation formula is: Where: Area of Intersection represents the area of the overlapping region of the two mask polygons P1 and P2, and Area of Union represents the area of the union region of the two mask polygons P1 and P2; For two mask polygons P1 and P2, calculate their IoU metric. If IoU≥threshold, delete the mask Intersection with the smaller area; if IoU<threshold, calculate the Intersection part and perform a difference set operation: threshold is a preset metric threshold; P1' = P1 - Intersection P’2 = P2 - Intersection Retain the non-overlapping regions on both sides.

14. A system for detecting cells based on DistSegNet model image processing technology, characterized in that: Include: A preprocessing module, a first segmentation module, a second segmentation module, a third segmentation module, an optimization processing module, and an extraction module; The preprocessing module is suitable for preprocessing the collected cell images; The first segmentation module is suitable for segmenting the cell image based on the DistSegNet model image processing technology to generate a cell nucleus region; The second segmentation module is suitable for generating a background region, a foreground region, and an edge region by using morphological operations based on the segmented cell nucleus region; The third segmentation module is suitable for calculating the shortest Euclidean distance from each background region pixel to the nearest foreground region pixel, and segmenting the cell image by using the watershed algorithm; The optimization processing module is suitable for optimizing the result of the segmentation map corresponding to each cell instance label to obtain a filtered segmentation label map; The extraction module is suitable for extracting the outer contour of the intranuclear region and the outer contour of the surrounding region of each cell from the segmentation label map and saving them to the corresponding result file.

15. A system for splicing cell image segmentation results, characterized in that: Include: An image input module, an image splicing module, an overlapping region processing module, a contour merging and output module; The image input module is adapted to obtain input data, the input data including: an original slice image file and a JSON file of a segmented image result after cells are detected by the method for detecting cells based on the DistSegNet model image processing technology as described in any one of claims 1 to 11, the original slice image file including: a plurality of slice image blocks; The image stitching module is adapted to stitch the image blocks of each row horizontally in sequence, merge the overlapping area masks between the image blocks, remove the redundant parts and optimize the overlapping area, load the image blocks of each row from the specified path and record the image segmentation results of the image blocks, check whether the masks are overlapping in the overlapping area, update the merged masks and remove the redundant masks; The overlapping area processing module is adapted to use an IoU index to measure the overlap between two masks, and to evaluate the overlap of the masks based on the IoU index; The outline merging and outputting module is suitable for drawing an image based on the overlapping conditions of all masks to generate a spliced ​​image of the cell image segmentation result.

Citation Information

Cited By

  • Ginkgo leaf extract state real-time monitoring method based on image processing

    CN120876482A

  • Cell segmentation method and device based on spatial omics sequencing, storage medium and program product

    CN121582926A

  • Image splicing and edge optimization method and system based on Ashlar registration

    CN122023114A