Cell growth state detection method based on image segmentation
By using a hierarchical division feature encoder and a cascade feature fusion network in cell growth state detection, combined with a decoder of the hollow convolution layer, the cell contamination problems caused by large detection errors, low efficiency and manual sampling in the prior art are solved, and high-precision cell image instance segmentation and accurate cell growth state detection are achieved.
Patent Information
- Application Number
- CN202510553849.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The prior art has problems such as large errors, low efficiency, and manual sampling lead to cell contamination and destruction in cell growth state detection, and lacks the ability to model hierarchically and deep semantic understanding of information of different particle sizes.
Using a cell growth state detection method based on image segmentation, the coarse-grained characteristics of cell blocks and fine-grained characteristics of cell boundaries are extracted through a hierarchical division feature encoder, and combined with the cascade feature fusion network and the hollow convolutional layer in the decoder, to achieve high-precision cell image instance segmentation.
It improves the accuracy and robustness of cell growth state detection, realizes efficient and accurate cell image instance segmentation, and improves the training efficiency and automation of the model.
Smart Images

Figure CN120071350A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cell image processing and segmentation, and particularly relates to a method for detecting the growth state of cells based on image segmentation. Background Art
[0002] During the process of cell culture, it is necessary to observe the growth state of cells in real time, and adjust the operations in a timely manner according to the growth state of cells to ensure the smooth progress of cell culture experiments. At present, there are two main methods for observing the growth state of cells: the first method uses a microscope to observe the growth state of cells with the human eye; the second method evaluates data such as the number of cells and the survival rate in the sample by manually opening the lid to take samples and then using a flow cytometer. The observation results obtained by the first method, such as cell survival rate, have large errors, consume a lot of manpower, and have low efficiency. The second method is likely to cause contamination and damage to the cells in the existing culture containers during the process of manual sampling and counting, which to a certain extent affects the normal growth process of cells. With the development of artificial intelligence technology, machine learning algorithms and deep learning algorithms have gradually been applied to the field of cell growth state detection.
[0003] The Chinese invention patent with the application number CN202410755124.4 discloses a method, device, equipment, medium and product for segmenting tumor tissue cell images, and the specific steps are as follows: Step 1, obtain a tumor tissue cell image; Step 2, input the tumor tissue cell image into a segmentation model, and the segmentation model outputs a boundary map and a foreground map of the tumor tissue cells, wherein the segmentation model includes a feature encoder, a context extractor and a feature decoder, and the skip connection between the feature encoder and the feature decoder adds an ASPP module and / or a CBAM attention mechanism, and the feature decoder splices and upsamples the feature maps extracted by the feature encoder and the context extractor, as well as the attention-weighted feature maps calculated by the CBAM attention mechanism to obtain the boundary map and the foreground map of the tumor tissue cells; Step 3, fuse the boundary map and the foreground map to obtain the segmentation result of the tumor tissue cell image. Although this method improves the segmentation accuracy by introducing the ASPP module and the CBAM attention mechanism, its feature extraction process still encodes at a single scale, making it difficult to fully consider the cell block structure and edge details, and lacking the ability to hierarchically model different granularity information. At the same time, this method does not consider the semantic task-assisted optimization strategy, which limits the training efficiency and generalization ability of the model.
[0004] The Chinese invention patent with the application number CN202410275835.1 discloses a cervical liquid-based thin-layer cell image segmentation method based on the SAM segmentation model, which specifically includes the following steps: Step 1, for the prompt information required for cell segmentation in the cervical liquid-based thin-layer cell image, after determining the cell region, construct a cell region color brightness histogram in a sampling manner to obtain the seed points of the cell nucleus. At the same time, obtain the auxiliary points of the cytoplasm around the seed points of the cell nucleus, and use the seed points of the cell nucleus and their corresponding auxiliary points as prompt points; Step 2, use the SAM segmentation model to process the cervical liquid-based thin-layer cell image to obtain the mask region corresponding to the cell region, and further optimize and update the mask region according to the contour of the mask region to remove holes and noise, so that the segmented cell region is complete. However, this method relies on manually constructing prompt points and performs segmentation guided by the seed points; therefore, there are problems of strong dependence on prompt information and high manual participation, which limits the automation degree and real-time performance of the system. At the same time, the mask optimization process repairs the contour based on heuristic rules, lacking deep semantic understanding and context modeling, which affects the accuracy and robustness of the segmentation results in complex images. Summary of the Invention
[0005] In view of the above problems, the present invention proposes a method for detecting the cell growth state based on image segmentation, including the following processes: S1, Collect cell growth state images in real time; S2, Input the image into the trained hierarchical division feature encoder to obtain the coarse-grained features of the cell block and the fine-grained features of the cell boundary; S3, Input the coarse-grained features of the cell block and the fine-grained features of the cell boundary into the trained cascaded feature fusion network to obtain the deep interaction fusion features; the cascaded feature fusion network includes the first feature fusion network and the second feature fusion network; the first feature fusion network fuses the coarse-grained features of the cell block and the fine-grained features of the cell boundary in the hierarchical division features output by the hierarchical division feature encoder to obtain the fusion features; the second feature fusion network fuses the fusion features, the coarse-grained features of the cell block, and the fine-grained features of the cell boundary again to obtain the deep interaction fusion features; S4, Input the deep interaction fusion features into the trained decoder to perform instance segmentation on the cell image to obtain the instance segmentation result of the cell image; the decoder introduces a dilated convolution layer on the basis of the classical decoder to increase the receptive field without reducing the resolution of the feature map; S5, Calculate the number and survival rate of target cells according to the instance segmentation result of the cell image.
[0006] Preferably, the hierarchical division feature encoder includes a cell block feature rough extraction network and a cell boundary feature fine-grained encoding network; the cell block feature rough extraction network divides the image into 16*16 sub-blocks, and based on the self-attention layer and MixFNN feature extraction is performed to obtain the coarse-grained features of the cell blocks; the cell boundary feature fine-grained encoding network then performs refined extraction on the cell boundaries, divides the image into 64*64 sub-blocks, and then further extracts the fine-grained features through the encoding layer and feed-forward neural network based on Transformer to obtain the fine-grained features of the cell boundaries; finally, the outputs of the two networks are concatenated to obtain the hierarchical division features output by the hierarchical division feature encoder.
[0007] Preferably, the process of constructing the dataset for training the hierarchical division feature encoder, the cascaded feature fusion network, and the decoder includes: first, collecting image data containing target cells and performing denoising and normalization processing to obtain standard cell image data as ; Using the LabelMe instance segmentation tool to label the target cells to obtain an instance segmentation dataset; that is, drawing an instance bounding box for each cell instance, for each cell instance, assigning a unique instance ID to all pixels within the labeled area, and marking the instance ID label of each cell; marking the non-target cell area as the "background" label and annotating it with the RGB value 255, 255, 255; for the initial image the instance label result is obtained and the instance data is obtained ; the initial image and the instance data together constitute the instance segmentation dataset; Based on the instance-labeled image, script processing is performed to generate a semantic segmentation image, which is combined with the initial image to construct a semantic segmentation dataset; that is, the RGB of the labeled instances in the image is processed by OpenCV to be the consistent RGB pixel value 0, 0, 0; thus, a semantic label map is obtained the semantic label data and the initial image together constitute the semantic segmentation dataset.
[0008] Preferably, the specific data processing process of the cell block feature rough extraction network includes: Based on the cell block feature rough extraction network, the input image is roughly divided into 16*16 sub-blocks, and then each 16*16 sub-block is subjected to operations of the Conv3 convolutional layer and the max pooling layer to extract the low-level features of the image, including texture and color; After that, a self-attention layer is used to learn the relationship between different regions in the image, that is: ; Among them, , and represent matrices obtained by performing fully connected layer linear processing on respectively, is the normalization coefficient, represents the softmax function; is the feature after self-attention extraction; After that, the input features are non-linearly mapped using the hybrid feed-forward neural network MixFNN to obtain enhanced features ; including processing using multi-scale convolutional layers of Conv3, Conv5, and Conv7, and using Resnet residual connections to avoid the problem of gradient disappearance, that is: ; ; Among them, and represent convolution operations with convolution kernels of 5*5 and 7*7 respectively, represents the channel splicing operation; represents the ReLU activation function; is the multi-scale residual feature; After that, to enhance the feature extraction and encoding capabilities, multiple self-attention layers and MixFNN layers are repeatedly stacked to achieve deeper feature learning and encoding; the output of each layer will be processed as the input of the next layer; after passing through multiple self-attention layers and MixFNN layers, a cell block coarse-grained feature containing rich information is finally formed .
[0009] Preferably, the data processing process of the cell boundary feature fine-grained encoding network includes: First, the input image will be divided into 64*64 sub-blocks, and the fine-grained features in the sub-blocks will be preliminarily extracted through the Conv5 convolution operation and the average pooling layer Avgpool: ; Among them, is the extracted feature, is the convolution layer function with a convolution kernel of 5*5, is the bias term, is the average pooling operation; After that, the encoder layer of Transformer is used to further refine the extraction of features: ; Among them, is the feature representation after extraction, Represents the encoding layer processing function of the Transformer architecture; Subsequently, the features extracted by the encoding layer are further non-linearly mapped through the feed-forward neural network FNN, and each layer of the feed-forward neural network consists of two fully connected layers and a ReLU activation function; Finally, the fine-grained features of the cell boundary and the coarse-grained features of the cell block are concatenated to provide a more complete feature representation: ; Among them, is the concatenated feature and also the output result of the hierarchical division feature encoder.
[0010] Preferably, the specific data processing process of the cascaded feature fusion network is as follows: The first feature fusion network performs feature fusion on the coarse-grained features of the cell block and the fine-grained features of the cell boundary The first feature fusion network includes a bilinear interpolation upsampling layer, a channel concatenation layer, a 1×1 convolutional layer, a ReLU activation function layer, and a batch normalization layer; First, the coarse-grained features of the cell block are input into the bilinear interpolation upsampling layer to perform bilinear interpolation on the feature channels of and upsample to the same size as to obtain the padded coarse-grained features of the cell block ; Secondly, and are input into the channel concatenation layer for feature concatenation to obtain the concatenated feature ; Then, is input into the 1×1 convolutional layer to reduce the number of channels to reduce the computational complexity and obtain the dimensionality-reduced feature , and then is successively input into the ReLU activation function layer and the batch normalization layer to perform batch normalization operations on the feature channels to obtain the fused feature The second feature fusion network performs deep feature fusion on the fused feature output by the first feature fusion network, the coarse-grained features of the cell block and the fine-grained features of the cell boundary The second feature fusion network includes a bilinear interpolation upsampling layer, a channel concatenation layer, a SENet layer, a 3×3 convolutional layer, a ReLU activation function layer, and a batch normalization layer; First, the bilinear interpolation upsampling layer is used again to upsample the coarse-grained features of the cell block to the fine-grained feature Same size; secondly, the features after the first fusion , the coarsely grained features of the upsampled cell blocks and the finely grained features of the cell boundaries are concatenated to obtain the concatenated features ; On the concatenated features , the SENet module is applied to adaptively re-weight the importance of each channel feature to obtain the features ; The is input into a 3×3 convolutional layer to further integrate the feature information and reduce the number of channels to obtain the features , and then the is output through the ReLU activation function and batch normalization to obtain the deeply fused feature map .
[0011] Preferably, the decoder network is designed based on a classical encoder structure, including a transposed convolutional layer and an activation function layer, and a dilated convolutional layer is introduced. The convolutional kernels of the transposed convolution and the dilated convolution are both 3×3, and the number of channels is gradually reduced from 256 to 64; the specific data processing process includes: The is input into the first transposed convolutional layer for upsampling to obtain the downsampled features with a feature map size of W / 8×H / 8 , The number of channels of is 128; then the is input into the ReLU activation function layer for non-linear transformation to obtain the non-linear features ; Again, the is input into the dilated convolutional layer with a dilation rate of 2, and the deep features with a size of H / 8×W / 8 are output ; Again, the is successively input into the ReLU activation function layer and the second transposed convolutional layer to obtain the high-dimensional features with a size of W×H , and the second transposed convolutional layer is responsible for restoring the size of the feature map; Again, the ; Among them, represents the probability that the pixel at the position ([[]] i, j ) in the cell image belongs to the k-th cell instance, is the result of the Mask branch convolutional layer processing at the position ([[]] i, j) The pixel corresponds to the unnormalized score of the k-th cell instance; Finally, Input the sigmoid activation function layer to obtain the predicted value of the cell instance category to which the pixel located at the cell image position ( i, j ) belongs. The predicted values of the cell instance categories to which all the pixel points at all positions on the cell image belong form a set .
[0012] Preferably, during the training process of the network model, design an auxiliary transfer optimization training strategy, match a cell image semantic segmentation task that is simpler than this task as an auxiliary task for the cell image instance segmentation task, so that the model simultaneously solves the cell image instance segmentation task and the cell image semantic segmentation task during training, and the two tasks share a dual-task loss function for parameter update; the dual-task loss function is specifically as follows: ; where is the loss function of the cell image instance segmentation task, is the classification loss function of the cell image semantic segmentation, and are respectively and weights; The specific expression of is: is the pixel scale in the cell image; represents the weight at the -th pixel position; represents the predicted value of the cell instance category to which the -th pixel point in belongs; represents the cell instance category to which the -th pixel point in the cell image belongs; The specific expression of is: where is the number of cell categories, is the pixel scale in the cell image, is one-hot encoding. If the predicted category of the -th pixel point in the cell image is the same as the cell category , then , otherwise it is , is the -th pixel point belonging to the cell category The probability.
[0013] Compared with the prior art, the present invention has the following beneficial effects: 1. Improve the accuracy of cell growth state detection: By designing a hierarchical feature encoder, the present invention can extract the coarse-grained features of cell blocks and the fine-grained features of cell boundaries respectively. This multi-level feature extraction method helps to more accurately identify cell morphology and boundaries, thereby improving the accuracy of cell growth state detection; 2. Enhance the robustness of the model: The use of a cascaded feature fusion network enables the model to integrate feature information at different scales, improves the adaptability to scale changes in cell images, and enhances the robustness of the model. This is particularly important for processing cell images of different sizes and morphologies; 3. Achieve efficient and accurate instance segmentation of cell images: By introducing a dilated convolutional layer in the decoder part, the receptive field of the network is effectively expanded, while maintaining the spatial resolution of the feature map, improving the model's ability to capture detailed information and context information, thereby achieving efficient and accurate instance segmentation of cell images; 4. Accelerate the training convergence speed of the model: By designing a dual-task collaborative training strategy, combining the cell image instance segmentation task with the cell image semantic segmentation task, and using the context information provided by the semantic segmentation task to assist the instance segmentation task, the training convergence speed of the model is accelerated, and the training efficiency is improved; 5. Automation and real-time performance: Once the method of the present invention is trained and deployed in the background, it can achieve real-time detection of the cell growth state. This avoids the time-consuming and inefficient manual observation in traditional methods, as well as the pollution and damage that may be caused by manual sampling, and improves the automation degree and efficiency of cell culture experiments; 6. Wide application prospects: This method is not only applicable to the detection of the growth state of specific types of cells such as tumor cells and cervical liquid-based thin-layer cells, but can also be extended to other types of cell culture experiments. Brief Description of the Drawings
[0014] In order to more clearly illustrate the technical solutions of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following description is only one embodiment of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0015] Figure 1 It is the overall flowchart of the present invention.
[0016] Figure 2 It is the network structure diagram of the hierarchical feature encoder of the present invention.
[0017] Figure 3 This is the structural diagram of the cascaded feature fusion network of the present invention.
[0018] Figure 4 This is the structural diagram of the decoder network of the present invention.
[0019] Figure 5 This is an example diagram of cell image preprocessing according to an embodiment of the present invention.
[0020] Figure 6 This is the experimental result diagram according to an embodiment of the present invention. Detailed implementation manners
[0021] The overall process of a method for detecting cell growth state based on image segmentation proposed by the present invention is as Figure 1 shown. The specific implementation includes the following processes: Dataset construction: First, collect image data containing target cells, and perform denoising and normalization processing; use the LabelMe instance segmentation tool to label the target cells to obtain an instance segmentation dataset; based on the instance-labeled images, perform script processing to generate semantic segmentation images, and construct a semantic segmentation dataset together with the original images; Design a hierarchical division feature encoder: including a rough extraction network for cell block features and a fine-grained encoding network for cell boundary features; the rough extraction network for cell block features divides the image into 16*16 sub-blocks, and based on the self-attention layer and MixFNN performs feature extraction to obtain coarse-grained features of cell blocks; the fine-grained encoding network for cell boundary features then performs refined extraction of the cell boundary, divides the image into 64*64 sub-blocks, and then further extracts fine-grained features through an encoding layer based on Transformer and a feed-forward neural network to obtain fine-grained features of cell boundaries; finally, the outputs of the two networks are concatenated to obtain the output of the hierarchical division feature encoder, which is the hierarchical division feature; Design a cascaded feature fusion network: including a first feature fusion network and a second feature fusion network; the first feature fusion network fuses the coarse-grained features of cell blocks and the fine-grained features of cell boundaries in the hierarchical division features output by the hierarchical division feature encoder to obtain fusion features; the second feature fusion network fuses the fusion features, the coarse-grained features of cell blocks, and the fine-grained features of cell boundaries again to obtain deeply interactive fusion features; Design a decoder: Introduce a dilated convolution layer on the basis of a classic decoder to significantly increase the receptive field without reducing the resolution of the feature map, and achieve accurate instance segmentation of cell images; Design a dual-task collaborative training strategy to train the hierarchical division feature encoder, the cascaded feature fusion network, and the decoder, and improve the convergence speed of the model during training; Model deployment: Deploy the trained hierarchical division feature encoder, cascaded feature fusion network, and decoder in the background to achieve real-time detection of cell growth status.
[0022] The following further illustrates the specific implementation process of the present invention in conjunction with specific embodiments.
[0023] I. Construction of cell image dataset First, collect initial image data using an optical microscope , and a total of groups of images are collected, thus obtaining the initial cell image dataset ; and ; Then, use median filtering to denoise the image. For each pixel point in the image , the denoised result is: ; Among them, represents the neighborhood pixel point values centered on the pixel point ; is the median operation; linearly maps the pixel values of the image to the range through Min-Max normalization, that is: ; Among them, is the value after normalization of the pixel point , and are the minimum and maximum value functions respectively; denoising and normalization processing are performed on all images in , and the standard cell image data obtained is ; Use the professional annotation tool LabelMe to annotate each cell instance in the image; specifically, the operator needs to draw an instance bounding box for each cell instance. For each cell instance, all pixels within the annotation area are assigned a unique instance ID, and the instance ID label of each cell is marked (for example, "cell 1", "cell 2", and different ID labels are annotated through different RGB values); in addition, the non-target cell area is marked with the "background" label and annotated with RGB values of 255, 255, 255; therefore, for the initial image the instance label result is obtained, and the instance data is obtained; the initial image and the instance data together constitute the instance dataset; Based on the obtained instance data Through OpenCV, the RGB values of the labeled instances in the image are processed to be consistent RGB pixel values of 0, 0, 0, thus obtaining a semantic label map Semantic label data and the initial image together constitute a semantic dataset
[0024] II. Design of the hierarchical division feature encoder network structure The hierarchical division feature encoder respectively conducts rough feature division and refined feature encoding on the image, enabling the model to perform feature encoding on the image to be segmented at different levels. Specifically, the present invention designs a rough extraction network for cell block features and a fine-grained encoding network for cell boundary features, enabling the two branch networks to extract features at different levels, thereby improving the efficiency and accuracy of feature encoding. Its network structure is as Figure 2 shown, and the specific steps include 1. Design a rough extraction network for cell block features The input image is roughly divided into 16*16 sub-blocks, and then each 16*16 sub-block is subjected to a Conv3 convolutional layer and a max pooling layer Maxpool operation to extract the low-level features of the image (including information such as texture and color). This process is expressed as ; where is the extracted feature, is the convolutional layer function with a convolutional kernel of 3*3, is the bias term, is the max pooling operation After that, in order to capture the overall division information of the blocks and the context information between cells, and enhance the long-range dependence relationship in the image, the present invention uses a self-attention layer to learn the relationship between different regions in the image, that is ; where , and represent the matrices obtained by linearly processing through fully connected layers respectively, is the normalization coefficient, represents the softmax function is the feature after self-attention extraction The input feature is non-linearly mapped using a mixed feed-forward neural network (MixFNN) to effectively transfer the position information of the image sub-blocks and obtain the enhanced feature Specifically, this process includes processing using multi-scale convolutional layers of Conv3, Conv5, and Conv7, and using Resnet residual connections to avoid the problem of gradient disappearance, that is: ; ; Among them, and respectively represent convolutional operations with convolutional kernels of 5*5 and 7*7, represents the channel concatenation operation; represents the ReLU activation function; is the multi-scale residual feature; After that, in order to further enhance the feature extraction and encoding capabilities, multiple self-attention layers and MixFNN layers are repeatedly stacked to achieve deeper feature learning and encoding. The output of each layer will be processed as the input of the next layer; after passing through multiple self-attention layers and MixFNN layers, a coarse-grained feature of the cell block containing rich information is finally formed , providing a basis for subsequent feature fusion and instance segmentation tasks.
[0025] 2. Design a refined feature encoding network: Its goal is to refine the extraction of features related to cell boundaries in the image; first, the input image will be divided into sub-blocks of 64*64, and the fine-grained features in the sub-blocks will be initially extracted through the Conv5 convolutional operation and the average pooling layer Avgpool: ; Among them, is the extracted feature, is the convolutional layer function with a convolutional kernel of 5*5, is the bias term, is the average pooling operation; After that, taking advantage of the effectiveness of the encoder layer of the Transformer in learning cross-view image details and context information, and finally enhancing the representation of features in global information, the encoding layer of the Transformer architecture is adopted in the present invention to further refine the extraction of features: ; Among them, is the extracted feature representation, represents the processing function of the encoding layer of the Transformer architecture; The features extracted by the encoding layer are further non-linearly mapped through a feed-forward neural network (FNN), and each layer of the feed-forward neural network consists of two fully connected layers and a ReLU activation function. This process includes: ; Among them, represents a fully connected operation, is the deep feature after non-linear mapping; In order to improve the accuracy and depth of cell boundary features, the present invention stacks multiple self-attention layers and feed-forward neural network layers to form a deep network structure. The output of each layer is processed as the input of the next layer. Repeatedly stacking these layers helps to extract more complex and fine-grained features and improve the performance of the model, and finally obtain the fine-grained cell boundary features output by the refined feature encoding network ; Finally, the fine-grained cell boundary features are concatenated with the coarse-grained cell block features to provide a more complete feature representation: ; Among them, is the concatenated feature and also the output result of the hierarchical division feature encoder.
[0026] III. Design of Cascade Feature Fusion Network The cascade feature fusion network performs cascade feature fusion on the coarse-grained cell block features output by the hierarchical division feature encoder and the fine-grained cell boundary features so that the model can more accurately capture and understand the multi-level boundary information in the cell image, thereby improving the cell boundary segmentation accuracy of the model and the robustness to the scale change of the cell image through multi-scale information integration; Specifically, the cascade feature fusion network includes a first feature fusion network and a second feature fusion network, and its network structure is as Figure 3 shown. The steps of inputting the coarse-grained cell block features and the fine-grained cell boundary features into the cascade feature fusion network are as follows: The first feature fusion network performs feature fusion on the coarse-grained cell block features and the fine-grained cell boundary features . The first feature fusion network includes a bilinear interpolation upsampling layer, a channel concatenation layer, a 1×1 convolutional layer, a ReLU activation function layer, and a batch normalization layer; first, the coarse-grained cell block features are input into the bilinear interpolation upsampling layer to perform bilinear interpolation on the feature channels of to upsample to the same size as to obtain the padded coarse-grained cell block features ; secondly, is concatenated with The input channel splicing layer performs feature splicing to obtain spliced features ; Secondly, Input into a 1×1 convolutional layer to reduce the number of channels and computational complexity to obtain dimensionality-reduced features , then Input into the ReLU activation function layer and batch normalization layer in sequence to perform batch normalization on the feature channels to obtain fused features ; The second feature fusion network performs deep feature fusion on the fused features output by the first feature fusion network as well as the coarse-grained features of the cell block and the fine-grained features of the cell boundary ; The second feature fusion network includes a bilinear interpolation upsampling layer, a channel splicing layer, a SENet layer, a 3×3 convolutional layer, a ReLU activation function layer, and a batch normalization layer; First, use the bilinear interpolation upsampling layer again to upsample the coarse-grained features of the cell block to fine-grained features of the same size; Secondly, the features after the first fusion , the upsampled coarse-grained features of the cell block and the fine-grained features of the cell boundary are spliced to obtain spliced features ; Apply the SENet module to the spliced features to adaptively re-weight the importance of each channel feature to obtain features ; Input into a 3×3 convolutional layer to further integrate feature information and reduce the number of channels to 256 to obtain features , then outputs the deep fusion feature map through the ReLU activation function and batch normalization .
[0027] IV. Decoder Network Design The classical encoder structure is mainly composed of a transposed convolutional layer and an activation function layer, and it is difficult to capture context information. To solve this problem, the present invention introduces a dilated convolutional layer in the classical decoder, which allows the network to effectively expand the receptive field of the network without reducing the spatial resolution of the feature map, effectively improving the model's ability to capture detailed information and context information. In the present invention, the convolutional kernels of the transposed convolution and the dilated convolution are both 3×3, and the number of channels gradually decreases from 256 to 64. The network structure is as Figure 4 shown; Input the obtained deep interaction fusion feature map into the decoder network. First, Input the first transposed convolutional layer for upsampling to obtain a dimensionality-reduced feature with a feature map size of W / 8 × H / 8 , with 128 channels; Subsequently, input it into the ReLU activation function layer for non-linear transformation to obtain non-linear features ; Again, input it into the dilated convolutional layer with a dilation rate of 2. Since the padding is also 2, the dilated convolution does not change the size of the feature map, and the output is a depth feature with a size of H / 8 × W / 8 ; Again, input it into the ReLU activation function layer and the second transposed convolutional layer in sequence to obtain a high-dimensional feature with a size of W × H , and the second transposed convolutional layer is responsible for restoring the size of the feature map. Again, input it into the Mask branch convolutional layer to generate an instance cell image segmentation mask, that is, the probability that each pixel in the cell image belongs to a different cell instance. What the Mask branch convolutional layer outputs is a three-dimensional tensor containing the probabilities corresponding to each pixel for each cell instance. The specific calculation formula is as follows: ; where, represents the probability that the pixel at the position ([[]] i, j ) in the cell image belongs to the k-th cell instance, is the unnormalized score corresponding to the k-th cell instance of the pixel at the position ([[]] ) in the cell image after the Mask branch convolutional layer processes i, j ; Finally, input it into the sigmoid activation function layer to obtain the predicted value of the cell instance category to which the pixel at the position ([[]] i, j ) in the cell image belongs. The predicted values of the cell instance categories to which all the pixels at all positions in the cell image belong form a set .
[0028] V. Training Strategy Design Since the cell image instance segmentation task to be solved by the present invention requires accurate segmentation of different types of cells and also segmentation of cells of different instances of the same type, and due to the large number of cell instances in the cell image and the situation of cell boundary intersections, the difficulty of the cell image instance segmentation task is increased, and the traditional training strategy based on the cross-entropy loss function makes it difficult for the model to converge rapidly during training. To accelerate the convergence speed of the model during training, the present invention proposes an auxiliary transfer optimization training strategy, which matches a cell image semantic segmentation task simpler than this task as an auxiliary task for the cell image instance segmentation task, so that the model simultaneously solves the cell image instance segmentation task and the cell image semantic segmentation task during training, and the two tasks share a dual-task loss function for parameter update. Semantic segmentation provides context information about the content of the cell image by distinguishing pixels of different cell classes in the cell image, and this context information helps instance segmentation to more accurately identify and segment different instances of cells of the same class; Dual-task loss function Specifically as follows: ; Wherein is the loss function of the cell image instance segmentation task, is the classification loss function of the cell image semantic segmentation, and are respectively and weights of; The specific expression of is: ; is the pixel scale in the cell image; represents the weight at the th pixel position; represents the predicted value of the cell instance category to which the th pixel point in belongs; represents the cell instance category to which the th pixel point in the cell image belongs; The specific expression of is: ; Wherein is the number of cell classes, is the pixel scale in the cell image, is a one-hot encoding. If the predicted category of the th pixel point in the cell image is the same as the cell category , then , otherwise , is the probability that the th pixel belongs to the cell category .
[0029] VI. Model Deployment and Experiments Deploy the trained hierarchical classification feature encoder, cascaded feature fusion network, and decoder in the background; Input the collected cell images into the hierarchical classification feature encoder to obtain the coarse-grained features of cell blocks and the fine-grained features of cell boundaries; Input the coarse-grained features of cell blocks and the fine-grained features of cell boundaries into the cascaded feature fusion network to obtain the deep fusion features; Input the deep fusion features into the decoder to obtain the instance segmentation results of cell images; Calculate the number and survival rate of target cells based on the instance segmentation results of cell images; To verify the performance of the method proposed in the present invention, the inventor made a dataset containing 6,000 cell pictures. The training set and the test set were divided according to a ratio of 7:3, where the training set contained 4,200 cell images and the test set contained 1,800 cell images. The method proposed in the present invention was compared with YOLOV8 and Mask-RNN, and the comparison of the MPA (class average pixel accuracy) of different models is shown in Table 1. In the training set, the accuracy of the method of the present invention for instance segmentation of cell images was significantly better than that of YOLOV8 and Mask-RNN (P < 0.002, P < 0.002); in the test set, the MPA of the method of the present invention was still the highest, consistent with its performance in the training set, and significantly better than YOLOV8 and Mask-RNN (P < 0.002, P < 0.002). It is worth noting that the difference between the MPA of the method of the present invention on the test set and the training set was small, while YOLOV8 and Mask-RNN performed better in the training set but worse in the test set, indicating that YOLOV8 and Mask-RNN had weak generalization ability on the test set, and the method of the present invention had strong generalization ability on the test set.
[0030] Table 1 Comparison of MPA between different models
[0031] The experimental results that can be presented are as Figure 5 and Figure 6 shown. According to Figure 5 , it can be observed that after image preprocessing, the line features at the cell edges become clearer, which makes the subsequent model segmentation work more efficient and accurate. After image preprocessing, standardized and denoised cell image data are obtained, thus greatly improving the performance and reliability of the cell segmentation task.
[0032] From Figure 6 It can be seen that the instance segmentation method proposed by the present invention can accurately perform precise instance segmentation on different types of target cells, and provides reliable data support for subsequent medical staff to detect the cell growth state. This method not only improves the accuracy of segmentation, but also ensures the accurate identification of cell states, providing strong technical support for clinical applications.
[0033] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0034] Although the specific implementation manners of the present invention have been described above, they do not limit the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A cell growth state detection method based on image segmentation, characterized in that: The process includes: S1, real-time acquisition of cell growth status images; S2, input the image into the trained hierarchical feature encoder to obtain the coarse-grained features of cell blocks and the fine-grained features of cell boundaries; S3, inputting the coarse-grained features of the cell blocks and the fine-grained features of the cell boundaries into the trained cascade feature fusion network to obtain deep interactive fusion features; the cascade feature fusion network includes a first feature fusion network and a second feature fusion network; The first feature fusion network fuses the cell block coarse-grained features and the fine-grained features of the fine-packet boundary in the hierarchical division features output by the hierarchical division feature encoder to obtain a fusion feature; The second feature fusion network fuses the fusion features, the coarse-grained features of the cell blocks, and the fine-grained features of the fine-packet boundaries again to obtain deep interactive fusion features; S4, inputting the deep interactive fusion features into the trained decoder to perform instance segmentation on the cell image to obtain the cell image instance segmentation result; the decoder introduces a hole convolution layer on the basis of the classic decoder to increase the receptive field without reducing the resolution of the feature map; S5, calculate the number and survival rate of target cells based on the cell image instance segmentation results.
2. The method for detecting cell growth status based on image segmentation according to claim 1, characterized in that: The hierarchical feature encoder includes a cell block feature coarse extraction network and a cell boundary feature fine-grained encoding network; the cell block feature coarse extraction network divides the image into 16*16 sub-blocks and uses the self-attention layer and MixFNN Feature extraction is performed to obtain coarse-grained features of cell blocks; the cell boundary feature fine-grained encoding network performs refined extraction of cell boundaries, divides the image into 64*64 sub-blocks, and then further extracts fine-grained features through the Transformer-based encoding layer and feedforward neural network to obtain fine-grained features of cell boundaries; finally, the outputs of the two networks are spliced to obtain the hierarchical partition feature encoder output hierarchical partition features.
3. The cell growth state detection method based on image segmentation according to claim 1, characterized in that: The dataset construction process for training the hierarchical feature encoder, cascade feature fusion network and decoder includes: firstly, collecting image data containing target cells, and performing denoising and standardization to obtain standard cell image data as ; The target cells were labeled using the LabelMe instance segmentation tool to obtain an instance segmentation dataset. That is, an instance bounding box was drawn for each cell instance. For each cell instance, all pixels in the labeled area were assigned a unique instance ID, and the instance ID label of each cell was marked. Non-target cell areas were marked as "background" labels and annotated with RGB values of 255, 255, 255. For the initial image, Label the instance , and get the instance data ; Initial image With instance data Together they form an instance segmentation dataset; Based on the instance-annotated image, script processing is performed to generate a semantic segmentation image, and a semantic segmentation dataset is constructed together with the initial image; that is, the RGB of the labeled instance in the image is processed by OpenCV to a consistent RGB pixel value of 0, 0, 0; thus obtaining a semantic label map , semantically labeled data and the initial image Together they form a semantic segmentation dataset.
4. A cell growth state detection method based on image segmentation as claimed in claim 2, characterized in that: The specific data processing process of the cell block feature rough extraction network includes: Based on the cell block feature rough extraction network, the input image is roughly divided into 16*16 sub-blocks, and then each 16*16 sub-block is operated through the Conv3 convolution layer and the maximum pooling layer to extract the low-level features of the image, including texture and color; Afterwards, a self-attention layer is used to learn the relationship between different regions in the image, namely: ; in, , and Indicates respectively The matrix obtained by linear processing of the fully connected layer, is the normalization coefficient, represents the softmax function; is the feature extracted by self-attention; Afterwards, the mixed feedforward neural network MixFNN is used to perform nonlinear mapping on the input features to obtain enhanced features ; Including the use of multi-scale convolutional layers of Conv3, Conv5 and Conv7, and the use of Resnet residual connections to avoid the gradient vanishing problem, namely: ; ; in, and They represent convolution operations with convolution kernels of 5*5 and 7*7 respectively. Indicates channel splicing operation; ReLU activation function. is the multi-scale residual feature; Afterwards, in order to enhance the feature extraction and encoding capabilities, multiple self-attention layers and MixFNN layers are repeatedly stacked to achieve deeper feature learning and encoding; the output of each layer is processed as the input of the next layer; after multiple self-attention layers and MixFNN layers, a coarse-grained feature of the cell block containing rich information is finally formed. .
5. The method for detecting cell growth status based on image segmentation according to claim 2, characterized in that: The data processing process of the cell boundary feature fine-grained encoding network includes: First, the input image will be divided into 64*64 sub-blocks, and the fine-grained features in the sub-blocks will be initially extracted through the Conv5 convolution operation and the average pooling layer Avgpool: ; in, is the extracted feature, is the convolution layer function with a convolution kernel of 5*5. is the bias term, is the average pooling operation; After that, the Transformer encoder layer is used to refine the extracted features: ; in, is the extracted feature representation, Represents the encoding layer processing function of the Transformer architecture; The features extracted by the encoding layer are then further nonlinearly mapped through the feedforward neural network FNN, and each layer of the feedforward neural network consists of two fully connected layers and a ReLU activation function; Finally, the cell boundary fine-grained features Coarse-grained features of cell blocks Concatenate to provide a more complete feature representation: ; in, It is the concatenated feature and also the output result of the hierarchical feature encoder.
6. A cell growth state detection method based on image segmentation as claimed in claim 1, characterized in that: The specific data processing process of the cascade feature fusion network is: The first feature fusion network performs coarse-grained feature fusion on the cell block. and cell boundary fine-grained features To perform feature fusion, the first feature fusion network includes a bilinear interpolation upsampling layer, a channel splicing layer, a 1×1 convolution layer, a ReLU activation function layer, and a batch normalization layer; first, the cell block coarse-grained features Input bilinear interpolation upsampling layer pair The bilinear interpolation of the feature channels will be Upsample to The same size, to obtain the coarse-grained features of the filled cell blocks Secondly, and Input channel splicing layer to perform feature splicing to obtain splicing features ; Then, Input 1×1 convolution layer to reduce the number of channels to reduce computational complexity and obtain dimensionality reduction features , then Input the ReLU activation function layer and batch normalization layer in sequence to perform batch normalization operations on the feature channels to obtain fusion features ; The second feature fusion network combines the fusion features output by the first feature fusion network and the coarse-grained characteristics of cell blocks and cell boundary fine-grained features Perform deep feature fusion; The second feature fusion network includes a bilinear interpolation upsampling layer, a channel splicing layer, a SENet layer, a 3×3 convolution layer, a ReLU activation function layer, and a batch normalization layer. First, the bilinear interpolation upsampling layer is used again to convert the cell block coarse-grained features into Upsampling to fine-grained features The same size; secondly, the features after the first fusion , coarse-grained features of cell blocks after upsampling and fine-grained features of cell boundaries Perform splicing to obtain splicing features ; Features after splicing The SENet module is applied to adaptively reweight the importance of each channel feature to obtain the feature ;Will Input to a 3×3 convolutional layer to further integrate feature information and reduce the number of channels to obtain features , and then Output deep fusion feature map through ReLU activation function and batch normalization .
7. The method for detecting cell growth status based on image segmentation according to claim 1, characterized in that: The decoder network is designed based on the classic encoder structure, including a transposed convolution layer and an activation function layer, and a dilated convolution layer is introduced. The convolution kernel size of the transposed convolution and dilated convolution is 3×3, and the number of channels is gradually reduced from 256 to 64. The specific data processing process includes: Will Input the first transposed convolutional layer for upsampling to obtain a dimension reduction feature with a feature map size of W / 8×H / 8 , The number of channels is 128; then Input ReLU activation function layer for nonlinear transformation to obtain nonlinear features ; Again, Input to the hole convolution layer with a dilation rate of 2, and output the deep features of size H / 8×W / 8 ; Again, Input the ReLU activation function layer and the second transposed convolution layer in sequence to obtain high-dimensional features of size W×H , the second transposed convolutional layer is responsible for restoring the size of the feature map; again, The Mask branch convolutional layer is input to generate the instance cell image segmentation mask, that is, the probability that each pixel in the cell image belongs to a different cell instance. The Mask branch convolutional layer outputs a three-dimensional tensor containing the probability that each pixel corresponds to each cell instance. The specific calculation formula is as follows: ; in, Indicates the position of the cell image ( i,j ) belongs to the kth cell instance, It is the Mask branch convolution layer pair After processing, the cell image position ( i,j ) corresponds to the unnormalized fraction of the kth cell instance; Finally Input sigmoid activation function layer to obtain the cell image position ( i,j ) belongs to the cell instance category prediction value, and the cell instance category prediction values of all pixel points at all positions on the cell image form a set .
8. The cell growth state detection method based on image segmentation as claimed in claim 1, characterized in that: During the network model training process, an auxiliary transfer optimization training strategy is designed to match a cell image semantic segmentation task that is simpler than the cell image instance segmentation task as an auxiliary task, so that the model can solve the cell image instance segmentation task and the cell image semantic segmentation task at the same time during training. The two tasks share a dual-task loss function for parameter update; the dual-task loss function The details are as follows: ; in is the loss function for the cell image instance segmentation task, is the classification loss function for semantic segmentation of cell images, and They are and The weight of The specific expression is: ; is the pixel scale in the cell image; Indicates The weight of the pixel position; express Middle The predicted value of the cell instance category to which the pixel belongs; Represents the cell image The cell instance category to which the pixel belongs; The specific expression is: ; in is the number of cell types, is the pixel scale in the cell image, is a one-hot encoding. If The predicted category and cell category of pixels Same, then , otherwise , For the pixels belong to the cell class probability.
Citation Information
Patent Citations
Cervical liquid-based thin layer cell image segmentation method based on SAM segmentation model
CN117876401B
Tumor tissue cell image segmentation method, device, equipment, medium and product
CN118447041A
Cell image segmentation method based on large language model and deep watershed algorithm
CN119169282A
Cell counting method and device and computer equipment
CN119887649A
Method and apparatus for image processing and visualization for analyzing cell kinematics in cell culture
US20190130161A1
Cited By
System and method applied to cell detection and recognition
CN121904758A
Stem cell image convergence degree calculation method and system based on U-Net network
CN122200652A
A stem cell image convergence degree calculation method and system based on a U-Net network
CN122200652B