High-content image binuclear cell automatic identification method based on two-stage deep learning
Through the dual-stage deep learning method, combined with the YOLOv8 and ResNet50 models, high-connotation imaging technology is used to achieve accurate detection and classification of binuclear cells, solving the problems of manual counting, strong subjectivity, and insufficient detection accuracy in the existing technology, significantly improving detection efficiency and accuracy.
Patent Information
- Application Number
- CN202510036847.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art has problems such as time-consuming manual counting, strong subjectivity, and insufficient detection accuracy of commercial software in the detection and classification of binuclear cells.
The dual-stage deep learning method is adopted, combined with YOLOv8 and ResNet50 models, and cell images are collected through high-connotation imaging technology, and rough and fine detection are carried out to achieve accurate detection and classification of binuclear cells.
It significantly improves the efficiency and accuracy of nucleus segmentation and classification, reduces manual intervention, enhances the feasibility of high-throughput detection, and provides a fast, accurate and highly automated binuclear cell detection solution.
Smart Images

Figure CN120107960A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical image processing and deep learning technology, and in particular to a method for automatically identifying binucleated cells in high-content images using dual-stage deep learning. Background Art
[0002] With the continuous development of digital high-content imaging technology, automated biomedical image analysis technology has played an important role in medical research, toxicology research and clinical diagnosis. Traditional cell nucleus segmentation and classification methods usually rely on manual setting of parameters for specific tasks, which is difficult to be widely applied to various high-content image analysis tasks. With the advancement of deep learning technology, many new methods have been applied to the field of biomedical image segmentation, but existing technologies still face many challenges in complex scenarios such as cell nucleus aggregation and overlap. This technology can effectively segment and classify cell nuclei in high-content images accurately, providing accurate cell nucleus contours and category information for biomedical and toxicology researchers. This is of great significance for improving the accuracy of image analysis, and can assist researchers in more accurate analysis and data interpretation, reduce the workload of manual analysis, ensure the accuracy and consistency of analysis results, and provide basic data support for personalized medicine, drug screening and toxicology research, helping doctors and researchers to formulate personalized treatment plans or evaluate toxic effects according to the specific biomedical characteristics of patients or experimental models, and improve the effectiveness of research and diagnosis.
[0003] High-content imaging technology can obtain high-resolution images and image information of multiple parameters at the same time, showing the microstructure and details of cells and organelles. This high-quality image is of great significance for studying the morphology, structure and function of cells, and helps scientists to deeply understand the physiological and pathological state of cells. Combining high-content imaging technology with deep learning for rapid and accurate binucleate cell detection is of great value for simultaneously obtaining multiple toxicity indicators for high-throughput compound toxicity screening. In order to realize large-scale cell experiments, binucleate cells can be quickly identified and classified, providing reliable data support for drug screening and toxicity assessment.
[0004] In the process of toxicological research, in vitro screening based on the cell level is an important analytical method for quickly screening the abnormal effects of new pollutants, chemicals, drugs and other substances on organisms or exploring toxicological mechanisms. The changes in different types of cell subtypes, especially the frequency of binucleated cells, are closely related to the early warning of potential toxic reactions. Traditional binucleated cell detection mainly relies on manual microscopy, which is not only time-consuming but also lacks objective consistency. In addition, in the whole-section images of microscopes, there are many cell subtypes. Due to various factors such as low staining quality, extensive nuclear overlap, and slight morphological differences within and between types, the detection of these cells has always been challenging. In recent years, image-based high-throughput and high-content cell imaging technologies have provided new development opportunities for various life science disciplines. High-content imaging technology can quickly, batch and automatically capture cell, subcellular or tissue images. These images have the characteristics of high resolution and high throughput, which can capture the fine structure and dynamic changes of cells.
[0005] Existing analysis methods can usually only meet the requirements of high-throughput or high-content imaging. On the one hand, super-resolution imaging ensures high content, but its time cost greatly limits the analysis throughput; on the other hand, if high-throughput analysis is to be achieved, time costs can only be saved by reducing resolution, which will result in many subcellular structures being unable to be effectively resolved, limiting the development of toxicological research. Although there are commercial software such as Harmony and MetaXpress, which can quantify cell phenotypes in high-content imaging and batch-convert image information into numerical information, it is still challenging to achieve accurate and efficient universal binucleate cell detection due to the complex and diverse image features of binucleate cells in microscopes. Therefore, more robust and general automated processing methods are needed to improve high-content imaging analysis technology.
[0006] Traditional image processing algorithms have been used for binucleated cell detection. These methods use improved watershed algorithms, seed region growth, iterative threshold segmentation based on cell size, shape (aspect ratio, relative concave-convex depth, etc.) and color feature rules to extract regions of interest. However, these traditional image processing methods are usually limited by factors such as the small nucleus size outside the micronucleus, poor contrast between the nucleus and the cytoplasm, micronucleus aggregation, and staining impurities. At present, there have been some binucleated cell detection works based on deep learning, but these methods are mainly single cell type detection, or can only perform rough binucleated cell classification (normal / abnormal), and the recognition accuracy is low. In actual application scenarios, the complex microscope environment makes it very difficult to design a robust algorithm, and the detection needs to be further accelerated to meet the needs of rapid detection of large sample omics data.
[0007] In recent years, deep learning has been applied to the detection of binucleated cells. Combining a spatial transformer network with a general convolutional neural network (CNN) to assist in recognition, it can correct the image to fill the entire field of view, creating a better environment for image recognition of abnormal binucleated cells. In addition, the relatively small number of abnormal binucleated cells and the high cost of manual data annotation make the classification of binucleated cells more suitable for transfer learning technology. Various deep CNNs initialized by transfer learning on ImageNet are compared to classify abnormal binucleated cell images and normal binucleated cell images. The above work only focuses on coarse-grained detection. In the detection of abnormal binucleated cells, few works involve fine detection of further typing of binucleated cells. Most of the work focuses on the detection of normal binucleated cells or micronuclei, which is not comprehensive enough. Therefore, the performance of existing deep learning related works in binucleated cell detection and classification is still limited.
[0008] Currently, there are many classic convolutional neural network architectures for deep learning, among which the deep residual network (ResNet) performs outstandingly in image classification tasks. For example, ResNet50 is optimized based on the VGG19 network, and implements residual learning by introducing a short-circuit connection mechanism. Its key improvements include directly using convolutions with a step size of 2 for downsampling, and replacing the traditional fully connected layer with a global pooling layer. This design significantly improves model performance while maintaining network complexity. A major feature of the ResNet50 model is that when the size of the feature map is reduced by half, the number of feature channels will double accordingly, thereby effectively improving the feature expression capability. In addition, with the introduction of the short-circuit mechanism, ResNet50 significantly reduces the gradient vanishing problem, and the classification error rate on the ImageNet test set is only 3.6%.
[0009] YOLOv8 is an important version in the YOLO series, with multiple innovations and significant advantages. Compared with earlier versions, YOLOv8 has significantly improved target detection accuracy, speed, and robustness. Its target detection performance is further optimized, continuing the advantages of end-to-end design and integrating a more flexible and efficient convolutional neural network architecture. Similar to other YOLO models, YOLOv8 divides the input image into grids, and the grid units are responsible for detecting the center point position of the target within their range. At the same time, YOLOv8 uses a lighter model architecture and dynamic task allocation strategy to improve detection speed and accuracy. YOLOv8's convolutional neural network can efficiently extract image features and output prediction results in combination with classification tasks, thereby achieving fast and accurate target detection. The main advantages of YOLOv8 are also reflected in the following three aspects: (1) With an end-to-end design, the model is more concise and efficient, which significantly improves the inference speed; (2) By performing convolution operations on the entire image, the receptive field is larger during detection, which can effectively reduce misjudgments caused by background interference; (3) YOLOv8 shows excellent robustness in transfer learning tasks and is suitable for a variety of target detection scenarios. Summary of the invention
[0010] Based on the above technical problems of time-consuming manual counting, strong subjectivity, and insufficient detection accuracy of commercial software in the prior art, a two-stage deep learning method for automatic identification of binucleated cells in high-content images is provided, which is used for automatic detection and identification of binucleated cells in high-content imaging technology. It not only overcomes the errors caused by low-resolution images in traditional methods, significantly improves the efficiency and accuracy of cell nucleus segmentation and classification, but also has broad application prospects in biomedical microscopic image analysis, high-throughput toxicity screening, environmental toxicology research and other fields. This technology provides an efficient and reliable solution for automated high-content image cell nucleus identification, and promotes the further development of computational toxicology.
[0011] The technical means adopted by the present invention are as follows:
[0012] A two-stage deep learning method for automatic identification of binucleated cells in high-content images, comprising:
[0013] S1. Using high-content imaging technology to collect cell images, obtain cell image data sets, and preprocess the cell image data sets;
[0014] S2. Based on the YOLOv8 deep learning network, a YOLOv8 binuclear cell coarse detection model is constructed, and the binuclear cell coarse detection model is trained using the training set and the validation set;
[0015] S3, sending the cell image to be detected into the trained YOLOv8 binuclear cell coarse detection model, optimizing the trained YOLOv8 binuclear cell coarse detection model, and obtaining the prediction result;
[0016] S4, processing the prediction result obtained in step S2 to obtain a binucleate cell image block in the rough detection stage;
[0017] S5, based on the binucleated cell image block, cropping and normalizing the binucleated cell region of interest in the label;
[0018] S6. Based on the diffusion model, the data processed in step S5 is enhanced to generate a binucleated cell ROI image dataset;
[0019] S7. Based on the ResNet50 network, a ResNet50 binuclear cell detection model is constructed, and the binuclear cell ROI image is sent to the binuclear cell detection model for training to obtain the prediction result;
[0020] S8. Integrate the prediction result obtained in step S7 with the cell nucleus classification result in the rough detection stage of step S4, and output the final classification result.
[0021] Furthermore, step S1 specifically includes:
[0022] S11. Obtain high-content image data through high-throughput nuclear and cytomic experiments;
[0023] S12, performing normalization correction on the high-content image data obtained by high-content imaging, including normalization of image size and pixel size, and normalization of image brightness;
[0024] S13. Divide the normalized and corrected dataset into training set, validation set and test set, and manually annotate the location and category of binucleated cells in each image.
[0025] Furthermore, step S2 specifically includes:
[0026] Cells of different phenotypes were grouped and classified based on morphological parameters, and classification predictions were performed using the nuclear morphological parameters and fluorescence expression intensity extracted from high-content images combined with the YOLOv8 deep learning network.
[0027] Further, in step S2, the YOLOv8 dual-core cell coarse detection model achieves efficient target detection through the collaborative work of the backbone network, the neck, and the detection head, as follows:
[0028] The backbone network performs preliminary feature extraction through downsampling, C2f module, SPPF module and convolution layer;
[0029] The neck is subjected to feature fusion and enhancement through upsampling, feature concatenation and C2f modules;
[0030] The detection head uses multi-layer convolution stacking and combines box and cls loss functions to accurately predict the location and category of the target.
[0031] Further, step S3 specifically includes:
[0032] The test set images are input into the YOLOv8 binuclear cell coarse detection model trained in step S2. The prediction results output by the YOLOv8 binuclear cell coarse detection model are the prediction bounding boxes, category probabilities, and confidence scores of different phenotype cell subpopulations, respectively. Among them:
[0033] The predicted bounding box includes the position and size information of each type of cell nucleus bounding box, expressed as [x_min, y_min, x_max, y_max] or center point format [x_center, y_center, width, height]; the bounding box information is stored in the form of a floating point array;
[0034] The category probability is the probability distribution of the cell nucleus category to which each bounding box belongs. The probability distribution is represented in the form of a floating-point array and the output is a vector containing the probability value of each category;
[0035] The confidence score is the confidence score of each predicted bounding box, and the confidence score represents the probability that the predicted bounding box contains the object in the form of a floating-point value.
[0036] Further, step S4 specifically includes:
[0037] Step S41: According to the evaluation results of the validation set, the parameters of the YOLOv8 binuclear cell coarse detection model are tuned to improve the detection performance;
[0038] Step S42: Perform experimental evaluation on the test set, calculate the recall rate, precision, F1 score and mAP@0.5 value of the YOLOv8 binuclear cell coarse detection model, and finally visualize the results.
[0039] Furthermore, step S5 specifically includes:
[0040] S51, screening out normal and abnormal cells from the binucleated cell image blocks extracted in the rough detection stage, including binucleated cells with micronuclei, nuclear buds, and nucleoplasmic bridges, and performing normalization processing on the screened images;
[0041] S52, cropping the image block of the binucleate cell label in the result of step S51 into a uniform size of 1280×1280 pixels, and placing the cropped image block in the middle of the image.
[0042] Further, step S6 specifically includes:
[0043] S61. Use the existing labeled data to train the diffusion model, and through the gradual learning process of noise addition and denoising, let the diffusion model generate new samples that conform to the cell morphology distribution;
[0044] S62, using a trained diffusion model to generate binucleated cell images with diverse features from random noise or partial noise images, the generated binucleated cell images showing wide variability in cell nucleus size, shape and abnormal features;
[0045] S63. Low-quality samples are eliminated through manual screening or automatic quality assessment tools, and only high-quality data are retained. The high-quality data are combined with the original data to generate a binucleated cell ROI image dataset.
[0046] Furthermore, in step S7, the main structure of the ResNet50 network is a residual unit, including a convolutional layer, a pooling layer, and a normalization layer.
[0047] Compared with the prior art, the present invention has the following advantages:
[0048] 1. The present invention provides a dual-stage deep learning method for automatic identification of binucleated cells in high-content images, which can achieve accurate detection and classification of high-content imaging cell nuclei, significantly improving detection efficiency and reliability. High-content imaging technology has the advantages of high resolution and high throughput, and can capture the fine structure of cells and their dynamic changes.
[0049] 2. The present invention provides a two-stage deep learning method for automatically identifying binucleated cells in high-content images. By combining the YOLOv8 deep learning model and the ResNet50 model, a two-stage binucleated cell automatic detection and classification method is designed; in the coarse detection stage, the YOLOv8 model performs preliminary positioning and coarse classification of binucleated cells with efficient detection capabilities, and its improved feature extraction modules, such as C2f and SPPF, enhance the ability to express cell nuclear features, greatly improving the detection speed and accuracy. In the fine detection stage, the ResNet50 model is used for deep feature extraction to further achieve accurate classification of normal binucleated cells and abnormal binucleated cells, including micronuclei, nuclear buds, and nucleoplasmic bridges. At the same time, the introduction of data enhancement technology (diffusion model) expands the training data set by generating binucleated cell images with diverse features, effectively improving the robustness and adaptability of the model.
[0050] 3. The dual-stage deep learning high-content image binucleate cell automatic identification method provided by the present invention is not only applicable to a variety of complex cell image scenes, but also can show excellent performance in high-throughput toxicity screening, providing a fast, accurate and highly automated solution for biomedical research. By integrating high-content imaging technology and dual-stage deep learning models, the present invention improves data processing efficiency while ensuring the reliability of test results, providing important support for large-scale cell research, drug screening and environmental toxicity assessment.
[0051] Based on the above reasons, the present invention can be widely promoted in fields such as biomedical image processing and deep learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0053] Figure 1 The figure is a flow chart of the method of the present invention.
[0054] Figure 2 A diagram of the YOLOv8 network structure provided in an embodiment of the present invention.
[0055] Figure 3 This is a visualization result of cell nucleus detection and identification provided by an embodiment of the present invention.
[0056] Figure 4 This is a diagram of the ResNet50 network structure provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0057] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0058] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0059] The present invention designs a two-stage deep learning method based on the combination of YOLOv8 and ResNet50 models for automatic detection and identification of binucleated cells using high-content imaging technology, aiming to solve the problems of time-consuming manual counting, strong subjectivity, and insufficient detection accuracy of commercial software in the prior art. The two-stage detection method includes a coarse detection stage and a fine detection stage:
[0060] In the coarse detection stage: transfer learning technology is used to train the YOLOv8 binuclear cell coarse detection model through manually annotated data, optimize the loss function, and continuously improve the model detection accuracy, so that the model can simultaneously predict the cell position and category in one forward propagation, greatly improve the processing speed, significantly reduce manual intervention, and enhance the feasibility of high-throughput detection. Based on the preliminary detection results of the YOLOv8 binuclear cell coarse detection model, binuclear cell images are cropped and scaled, and the diffusion model is used as a data enhancement technology to construct a new binuclear cell image dataset.
[0061] Second, in the fine detection stage, the ResNet50 binuclear cell fine detection model is used to perform deep feature extraction to further classify binuclear cells in the coarse detection results. This model can accurately distinguish normal binuclear cells from three abnormal binuclear cells, including binuclear cells with micronuclei, nuclear buds, and nucleoplasmic bridges.
[0062] Through the collaboration of the above two-stage model, the present invention can efficiently process large-scale high-content image data while ensuring high accuracy of cell subtype classification. Compared with traditional methods, the present invention has significant advantages in processing speed and accuracy. It not only simplifies the detection process and reduces the dependence on manual intervention, but also provides an efficient, accurate and reliable technical solution for automated toxicity screening.
[0063] To achieve the above object, the technical solution of the present invention is implemented as follows:
[0064] like Figure 1As shown, the present invention provides a method for automatically identifying binucleated cells in high-content images using a two-stage deep learning method, comprising:
[0065] S1. Using high-content imaging technology to collect cell images, obtain cell image data sets, and preprocess the cell image data sets;
[0066] S2. Based on the YOLOv8 deep learning model, a binucleated cell coarse detection model is constructed, and the binucleated cell coarse detection model is trained using the training set and the validation set;
[0067] S3, sending the cell image to be detected into the trained binucleate cell coarse detection model, optimizing the trained binucleate cell coarse detection model, and obtaining a prediction result;
[0068] S4, processing the prediction result obtained in step S2 to obtain a binucleate cell image block in the rough detection stage;
[0069] S5, based on the binucleated cell image block, the binucleated cell region of interest (ROI) in the label is cropped and normalized;
[0070] S6. Based on the diffusion model, the data processed in step S5 is enhanced to generate a binucleated cell ROI image dataset;
[0071] S7. Based on the ResNet50 network, a binuclear cell detection model is constructed, and the binuclear cell ROI image is sent to the binuclear cell detection model for training to obtain the prediction result;
[0072] S8. Integrate the prediction result obtained in step S7 with the cell nucleus classification result in the rough detection stage of step S4, and output the final classification result.
[0073] In specific implementation, as a preferred embodiment of the present invention, step S1 specifically includes:
[0074] S11. Obtain high-content image data through high-throughput nuclear and cytomic experiments;
[0075] S12, performing normalization correction on the high-content image data obtained by high-content imaging, including normalization of image size and pixel size, and normalization of image brightness;
[0076] S13, dividing the normalized and corrected data set into a training set, a validation set and a test set, and manually marking the position and category of binucleated cells in each image. In this embodiment, 445 images of human lung cancer cells (A549) are divided into a training set (314 images), a validation set (43 images) and a test set (88 images).
[0077] In specific implementation, as a preferred embodiment of the present invention, step S2 specifically includes:
[0078] Cells of different phenotypes were grouped and classified based on morphological parameters, and classification predictions were performed using the nuclear morphological parameters and fluorescence expression intensity extracted from high-content images combined with the YOLOv8 deep learning network.
[0079] In specific implementation, as a preferred embodiment of the present invention, in step S2, the binuclear cell coarse detection model works together through the backbone network (Backbone), the neck (Neck) and the detection head (Head) to achieve efficient target detection, such as Figure 2 As shown, the details are as follows:
[0080] The backbone network performs preliminary feature extraction through downsampling, C2f module, SPPF module and convolution layer (Conv);
[0081] The neck is subjected to feature fusion and enhancement through upsampling, feature concatenation and C2f modules;
[0082] The detection head uses multi-layer convolution stacking and combines box and cls loss functions to accurately predict the location and category of the target.
[0083] In this embodiment, this structure of the YOLOv8 deep learning network ensures the high performance and accuracy of the YOLOv8 binuclear cell coarse detection model. As shown in Table 1, this embodiment uses the YOLOv8x deep learning network, adopts the transfer learning method, and trains the YOLOv8 binuclear cell coarse detection model based on the labeled data, and optimizes the loss function to improve the detection accuracy. The YOLOv8 binuclear cell coarse detection model adopts a single-stage detection method, which can directly predict the location and category of the target in one forward propagation; at the same time, the model has an end-to-end training method, which simplifies the model training process.
[0084] Table 1 Key parameters of YOLOv8 binuclear cell coarse detection model
[0085]
[0086] In step S2, the training phase is implemented by the PyTorch deep learning package and trained on a workstation equipped with an RTX 4090 GPU, using the Adam optimizer, with the learning rate set to 1e-3, the weight decay to 0.0005, and the batch size set to 8.
[0087] In specific implementation, as a preferred embodiment of the present invention, step S3 specifically includes:
[0088] The test set images are input into the YOLOv8 binuclear cell coarse detection model trained in step S2. The prediction results output by the YOLOv8 binuclear cell coarse detection model are the prediction bounding boxes, category probabilities, and confidence scores of different phenotype cell subpopulations, respectively. Among them:
[0089] The predicted bounding box includes the position and size information of the bounding box of each type of cell nucleus (especially binucleated cells), expressed as [x_min, y_min, x_max, y_max] or center point format [x_center, y_center, width, height]; the bounding box information is stored in the form of a floating point array;
[0090] The category probability is the probability distribution of the cell nucleus category to which each bounding box belongs. The probability distribution is represented in the form of a floating-point array and the output is a vector containing the probability value of each category;
[0091] The confidence score is the confidence score of each predicted bounding box, and the confidence score represents the probability that the predicted bounding box contains the object in the form of a floating-point value.
[0092] In specific implementation, as a preferred embodiment of the present invention, step S4 specifically includes:
[0093] Step S41: According to the evaluation results of the validation set, the parameters of the YOLOv8 binuclear cell coarse detection model are tuned to improve the detection performance;
[0094] Step S42: Perform experimental evaluation on the test set, calculate the recall rate, precision, F1 score and mAP@0.5 value of the YOLOv8 binuclear cell coarse detection model, and finally visualize the results. Figure 3 shown.
[0095] In specific implementation, as a preferred embodiment of the present invention, step S5 specifically includes:
[0096] S51, screening out normal and abnormal cells from the binucleated cell image blocks extracted in the rough detection stage, including binucleated cells with micronuclei, nuclear buds, and nucleoplasmic bridges, and performing normalization processing on the screened images;
[0097] S52, cropping the image block of the binucleate cell label in the result of step S51 into a uniform size of 1280×1280 pixels, and placing the cropped image block in the middle of the image.
[0098] In specific implementation, as a preferred embodiment of the present invention, step S6 specifically includes:
[0099] S61. Use the existing labeled data to train the diffusion model, and through the gradual learning process of noise addition and denoising, let the diffusion model generate new samples that conform to the cell morphology distribution;
[0100] S62, using a trained diffusion model to generate binucleated cell images with diverse features from random noise or partial noise images, the generated binucleated cell images showing wide variability in cell nucleus size, shape and abnormal features;
[0101] S63. Low-quality samples are eliminated through manual screening or automatic quality assessment tools, and only high-quality data are retained. The high-quality data are combined with the original data to generate a binucleated cell ROI image dataset.
[0102] In specific implementation, as a preferred embodiment of the present invention, in step S7, the main structure of the ResNet50 network is a residual unit, and its structure is as follows: Figure 4 As shown, it includes convolutional layer, pooling layer and normalization layer. In this embodiment, the initial learning rate selected in the ResNet50 network training is 0.01, the optimizer adopts the stochastic gradient descent SGD algorithm, the number of iterations (epoch) is set to 50, and the batch size (batch size) is set to 16;
[0103] Example 1: Toxicity screening of plastic additives
[0104] In the plastics industry, plastic additives (such as plasticizers, antioxidants, etc.) are widely used, but their potential toxicity has not been fully evaluated. In order to evaluate the hazards of these additives to the environment and organisms, especially the effects on cells, high-content imaging technology combined with dual-stage deep learning detection technology can be used for efficient toxicity screening. The specific process is as follows:
[0105] First, high-content imaging technology is used to capture images of cells treated with additives to form a high-resolution image dataset. Then, the images are roughly detected based on the YOLOv8 binuclear cell coarse detection model to quickly screen out potential binuclear cells and mark abnormal areas. Next, the ResNet50 binuclear cell fine detection model is used to carefully classify the screened cells to distinguish between normal and abnormal binuclear cells, such as pathological features such as micronuclei, nuclear buds, or nucleoplasmic bridges. By analyzing the changes in cell subtypes after treatment with different additives, the toxic effects of these additives can be quickly evaluated, providing data support for further environmental risk assessment. This method can not only efficiently process large amounts of cell image data, but also accurately identify the potential carcinogenicity, mutagenicity and other toxicological effects of additives, greatly improving the speed and reliability of toxicity screening.
[0106] Example 2: Cytotoxicity Screening in New Drug Development
[0107] In the process of new drug development, drug safety assessment is crucial, especially the screening of drug-induced cytotoxicity and potential genetic damage. High-content imaging technology combined with a two-stage deep learning detection method can provide an efficient and accurate solution for cytotoxicity screening in drug development. The specific process is as follows:
[0108] High-content imaging technology is used to obtain high-resolution images of cells after drug treatment, capturing changes in cell morphology and intracellular structure. The YOLOv8 binuclear cell coarse detection model is used for preliminary coarse detection to quickly locate binuclear cells and mark areas where pathological changes may exist. Next, the ResNet50 binuclear cell fine detection model is used for detailed feature extraction to identify abnormalities such as micronuclei, nuclear buds, and nucleoplasmic bridges in cells, which are usually associated with DNA damage, chromosome breaks, and gene mutations. This automated detection method can quickly evaluate the toxic effects of drugs on cells in the early stages of drug screening, especially providing early warnings when possible mutagenic, carcinogenic, or teratogenic effects are discovered. Compared with traditional manual testing, this method not only improves screening efficiency, but also provides more accurate data support for drug safety, helping R&D teams accelerate the screening and optimization of drug candidates.
[0109] Example 3: Acute toxicity testing in chemical safety assessment
[0110] The safety assessment of chemicals is a key component of environmental and health risk management, especially before new chemicals enter the market, they need to undergo rigorous toxicity testing. In order to quickly and accurately assess the acute toxicity of chemicals, high-content imaging technology and dual-stage deep learning detection technology can be combined to automatically detect the damage of chemicals to cells. Assuming that a chemical (such as pesticides, industrial solvents, etc.) needs to undergo acute toxicity testing, the specific process is as follows:
[0111] First, high-content imaging technology was used to process cell images after chemical exposure to obtain a large amount of cell data. The YOLOv8 binuclear cell coarse detection model was used for preliminary detection to identify all binuclear cells and screen out suspected damaged cell areas. Then, the ResNet50 binuclear cell fine detection model was applied to subclassify the cells and analyze the morphological changes of the cells, especially the appearance of micronuclei, nuclear buds, and nucleoplasmic bridges. These cell abnormalities are related to DNA damage and chromosomal instability caused by chemicals. This technology can efficiently evaluate the acute toxicity of chemicals at different concentrations, quickly detect the toxic effects of chemicals on cells, and provide detailed safety data based on the type and degree of cell damage. This method not only speeds up the toxicity testing process, but also reduces the need for animal experiments through high-throughput screening, which is helpful for environmental safety assessments of chemicals and regulatory compliance inspections.
[0112] In summary, the present invention achieves efficient and accurate detection of binucleated cells by combining high-content imaging technology with a two-stage deep learning model (YOLOv8 and ResNet50), providing a new toxicity screening and analysis method. This technology can not only meet clinical diagnosis needs, but also help accelerate toxicity assessment, reduce experimental costs, and provide more accurate results.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A two-stage deep learning method for automatic identification of binucleated cells in high-content images, characterized in that: include: S1. Using high-content imaging technology to collect cell images, obtain cell image data sets, and preprocess the cell image data sets; S2. Based on the YOLOv8 deep learning model, a YOLOv8 binuclear cell coarse detection model is constructed, and the YOLOv8 binuclear cell coarse detection model is trained using the training set and the validation set; S3, sending the cell image to be detected into the trained YOLOv8 binuclear cell coarse detection model, optimizing the trained binuclear cell coarse detection model, and obtaining the prediction result; S4, processing the prediction result obtained in step S2 to obtain a binucleate cell image block in the rough detection stage; S5, based on the binucleated cell image block, cropping and normalizing the binucleated cell region of interest in the label; S6. Based on the diffusion model, the data processed in step S5 is enhanced to generate a binucleated cell ROI image dataset; S7. Based on the ResNet50 network, a ResNet50 binuclear cell detection model is constructed, and the binuclear cell ROI image is sent to the ResNet50 binuclear cell detection model for training to obtain the prediction result; S8. Integrate the prediction result obtained in step S7 with the cell nucleus classification result in the rough detection stage of step S4, and output the final classification result.
2. The method for automatic identification of binucleated cells in high-content images by two-stage deep learning according to claim 1, characterized in that: Step S1 specifically includes: S11. Obtain high-content image data through high-throughput nuclear and cytomic experiments; S12, performing normalization correction on the high-content image data obtained by high-content imaging, including normalization of image size and pixel size, and normalization of image brightness; S13. Divide the normalized and corrected dataset into training set, validation set and test set, and manually annotate the location and category of binucleated cells in each image.
3. The method for automatic identification of binucleated cells in high-content images by two-stage deep learning according to claim 1, characterized in that: Step S2 specifically includes: Cells of different phenotypes were grouped and classified based on morphological parameters, and classification predictions were performed using the nuclear morphological parameters and fluorescence expression intensity extracted from high-content images combined with the YOLOv8 deep learning network.
4. The method for automatic identification of binucleated cells in high-content images by two-stage deep learning according to claim 3, characterized in that: In step S2, the YOLOv8 dual-core cell coarse detection model achieves efficient target detection through the collaborative work of the backbone network, the neck, and the detection head, as follows: The backbone network performs preliminary feature extraction through downsampling, C2f module, SPPF module and convolution layer; The neck is subjected to feature fusion and enhancement through upsampling, feature concatenation and C2f modules; The detection head uses multi-layer convolution stacking and combines box and cls loss functions to accurately predict the location and category of the target.
5. The method for automatic identification of binucleated cells in high-content images using a two-stage deep learning method according to claim 1, characterized in that: Step S3 specifically includes: The test set images are input into the YOLOv8 binuclear cell coarse detection model trained in step S2. The prediction results output by the YOLOv8 binuclear cell coarse detection model are the prediction bounding boxes, category probabilities, and confidence scores of different phenotype cell subpopulations, respectively. Among them: The predicted bounding box includes the position and size information of each type of cell nucleus bounding box, expressed as [x_min, y_min, x_max, y_max] or center point format [x_center, y_center, width, height]; the bounding box information is stored in the form of a floating point array; The category probability is the probability distribution of the cell nucleus category to which each bounding box belongs. The probability distribution is represented in the form of a floating-point array and the output is a vector containing the probability value of each category; The confidence score is the confidence score of each predicted bounding box, and the confidence score represents the probability that the predicted bounding box contains the object in the form of a floating-point value.
6. The method for automatic identification of binucleated cells in high-content images using a two-stage deep learning method according to claim 1, characterized in that: Step S4 specifically includes: Step S41: According to the evaluation results of the validation set, the parameters of the YOLOv8 binuclear cell coarse detection model are tuned to improve the detection performance; Step S42: Perform experimental evaluation on the test set, calculate the recall rate, precision, F1 score and mAP@0.5 value of the YOLOv8 binuclear cell coarse detection model, and finally visualize the results.
7. The method for automatic identification of binucleated cells in high-content images using dual-stage deep learning according to claim 1, characterized in that: Step S5 specifically includes: S51, screening out normal and abnormal cells from the binucleated cell image blocks extracted in the rough detection stage, including binucleated cells with micronuclei, nuclear buds, and nucleoplasmic bridges, and performing normalization processing on the screened images; S52, cropping the image block of the binucleate cell label in the result of step S51 into a uniform size of 1280×1280 pixels, and placing the cropped image block in the middle of the image.
8. The method for automatic identification of binucleated cells in high-content images using dual-stage deep learning according to claim 1, characterized in that: Step S6 specifically includes: S61. Use the existing labeled data to train the diffusion model, and through the gradual learning process of noise addition and denoising, let the diffusion model generate new samples that conform to the cell morphology distribution; S62, using a trained diffusion model to generate binucleated cell images with diverse features from random noise or partial noise images, the generated binucleated cell images showing wide variability in cell nucleus size, shape and abnormal features; S63. Low-quality samples are eliminated through manual screening or automatic quality assessment tools, and only high-quality data are retained. The high-quality data are combined with the original data to generate a binucleated cell ROI image dataset.
9. The method for automatic identification of binucleated cells in high-content images using dual-stage deep learning according to claim 1, characterized in that: In step S7, the main structure of the ResNet50 network is a residual unit, including a convolutional layer, a pooling layer, and a normalization layer.
Citation Information
Cited By
Optical fiber abnormal state early warning method and system based on multi-data fusion
CN120910806A