Training method of micronucleus image deep learning detection model
By building a basic and fine-tuning database and using an improved YOLOv11s model and data augmentation technology, the problems of high labeled data cost and overfitting in the training of the micronucleus image detection model were solved, and the model's recognition ability was transferred and performance was improved.
Patent Information
- Application Number
- CN202510909022.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-10
AI Technical Summary
The cost of labeling data in the training of existing micronucleus image deep learning detection models is high, the model is severely overfitted, and the recognition ability has weak portability, making it difficult to adapt to the data needs of different acquisition equipment and hospitals.
A basic database was constructed by screening micronucleus images, and an improved YOLOv11s model with the MSAA module was added. The model was trained on a combination of basic and fine-tuning datasets, and data augmentation and multiple loss functions were used to optimize model performance.
It reduces the need for labeled data, improves the model's recognition, portability, and adaptability, and enhances the model's training performance and generalization.
Smart Images

Figure CN120766059A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of neural network model training, and particularly relates to a training method of a micro-nucleus image deep learning detection model. BACKGROUND
[0002] In the prior art, the cost of micro-nucleus image labeling data is high, and supervised learning requires a large number of labeled samples, and the model is trained in full quantity through the labeled samples. Meanwhile, since the content of the micro-nucleus image is relatively simple and the target types to be recognized are less, the modern neural network model has a large number of parameters, and overfitting is easily generated during training, and the generalization of the model is difficult to control. In addition, when the model is fine-tuned to adapt to different collection devices and data of different hospitals, the overfitting of the neural network will cause the model to forget, that is, after training on dataset 1, fine-tuning with data of dataset 2 will cause the performance of the final model to decrease in recognizing dataset 1.
[0003] Therefore, there is an urgent need for a training method of a micro-nucleus image deep learning detection model which can reduce the demand for labeled data, realize model recognition ability migration, and ensure the performance of the model. SUMMARY
[0004] Therefore, there is an urgent need for a training method of a micro-nucleus image deep learning detection model which can reduce the demand for labeled data, realize model recognition ability migration, and ensure the performance of the model.
[0005] A training method of a micro-nucleus image deep learning detection model, comprising the following steps: a plurality of micro-nucleus images are selected from a micro-nucleus image database, and a basic database is constructed based on the plurality of micro-nucleus images, the basic database comprising the same number of double-nucleus cell images and single-nucleus cell images, the micro-nucleus images being color images or black-and-white images; a plurality of fine-tuning images are collected based on a target hospital system, and a fine-tuning database is constructed based on the plurality of fine-tuning images; the basic database and the fine-tuning database are labeled and data augmentation is performed to obtain a basic data set and a fine-tuning data set; a deep learning model framework is built based on a PyTorch neural network library, an improved YOLOv11s model is used to construct an initial micro-nucleus detection model, and the improved YOLOv11s model adds an MSAA module in the neck network; and the initial micro-nucleus detection model is trained in sequence using the basic data set and the fine-tuning data set to obtain a target micro-nucleus detection model.
[0006] In one embodiment, the data augmentation method includes rotation, mirroring, up-down flipping, random tone adjustment, random image saturation adjustment, and random brightness adjustment.
[0007] In one embodiment, the improved YOLOv11s model includes: a backbone network, a neck network and a detection head; the backbone network is used for feature extraction, including a CBS module, a combination layer of four consecutive CBS modules and C3k2 modules, an SPPF module and a C2PSA module; wherein, the C2PSA module is used to divide the input feature map into two parts, one part directly passes to retain the original information, and the other part is processed by the PSA attention module, and the two parts are spliced and fused. The PSA attention module enhances the model's perception of target details by dynamically adjusting feature weights; the neck network is connected to the feature map through upsampling to perform feature fusion of the feature map, including: a first upsampling layer, a combination layer composed of a splicing layer, an MSAA layer and a C3k2 module, a second upsampling layer, a combination layer composed of a splicing layer, an MSAA layer and a C3k2 module, and two layers of a combination layer composed of a CBS module, a splicing layer, an MSAA layer and a C3k2 module; the detection head is provided with multiple layers for generating the final prediction result based on the processed feature map.
[0008] In one embodiment, the backbone network has 3 input channels and the image size is resized to 640*640, including: the first layer, CBS module, the number of filters is 64, the convolution kernel size is 3*3, the step size is 2, and the padding is 0; the second layer, CBS module, the number of filters is 128, the convolution kernel size is 3*3, the step size is 2, and the padding is 0; the third layer, C2f module, the number of channels is 256, the channel expansion coefficient is set to 0.25, and the number of Bottleneck blocks is 1; the fourth layer, CBS module, the number of filters is 256, the convolution kernel size is 3*3, the step size is 2, and the padding is 0; the fifth layer, C2f module, the number of channels is 512, the channel expansion coefficient is set to 0.25, and the number of Bottleneck blocks is 1; the sixth layer, CBS Module, number of filters 512, convolution kernel size 3*3, step size 2, padding 0; seventh layer, C3k2 module, number of channels 512, channel expansion coefficient defaults to 0.5, number of C3k modules is 1; eighth layer, CBS module, number of filters 1024, convolution kernel size 3*3, step size 2, padding 0; ninth layer, C3k2 module, number of channels 1024, channel expansion coefficient defaults to 0.5, number of C3k blocks is 1; tenth layer, SPPF module, number of filters 1024, convolution kernel size of convolution layer is 1*1, step size 1, padding 0, pooling kernel size is 5*5; eleventh layer, C2PSA module, number of filters 1024, number of PSABlock blocks is 1, channel expansion coefficient defaults to 0.5.
[0009] In one embodiment, the neck network includes: a first layer, an upsampling layer, a sampling coefficient of 2; a second layer, a splicing layer, a splicing layer of the first layer and the seventh layer of the backbone; a third layer, an MSAA module; a fourth layer, a C2f module, a filter number of 512, a channel expansion coefficient of 0.5, and a Bottleneck block number of 1; a fifth layer, an upsampling layer, a sampling coefficient of 2; a sixth layer, a splicing layer, a splicing layer of the fourth layer and the fifth layer of the backbone network; a seventh layer, an MSAA module; an eighth layer, a C2 module, a filter number of 256, a channel expansion coefficient of 0.5, and a Bottleneck block number of 1; a ninth layer, a convolutional layer, a filter number of The number of filters is 256, the convolution kernel size is 3*3, the step size is 2, and the padding is 0; the tenth layer is the splicing layer, which splices the seventh layer and the third layer; the eleventh layer is the MSAA module; the twelfth layer is the C2f module, the number of filters is 512, the channel expansion coefficient is set to 0.5, and the number of Bottleneck blocks is 1; the thirteenth layer is the convolution layer, the number of filters is 512, the convolution kernel size is 3*3, the step size is 2, and the padding is 0; the fourteenth layer is the splicing layer, which splices the tenth layer and the eleventh layer of the backbone network; the fifteenth layer is the MSAA module; the sixteenth layer is the C3k2 module, the number of channels is 1024, the channel expansion coefficient defaults to 0.5, and the number of C3k blocks is 1.
[0010] In one embodiment, the detection head includes a bounding box prediction network and an object classification prediction network; wherein the bounding box prediction network includes: a first layer, a Conv layer, the Conv layer includes a convolution layer, a BN layer and an activation function, the convolution kernel size is 3*3, and the activation function is SiLU; the second layer, the Conv layer, is the same as the first layer; the third layer, the convolution layer, the convolution kernel size is 1, and the padding is 0; the object classification prediction network includes: a first layer, composed of a DWConv layer and a Conv layer, the Conv layer includes a convolution layer, a BN layer and an activation function, the convolution kernel size is 3*3, the DWConv layer is a depth-separable convolution layer, the convolution kernel size is 3*3, and the step size is 1; the second layer, composed of a DWConv layer and a Conv layer, is the same as the first layer; the third layer, the convolution layer, the convolution kernel size is 1, and the padding is 0.
[0011] In one embodiment, the method further includes: using a classification loss function, a positioning loss function, and a distribution focus loss function to perform loss calculation of the improved YOLOv11s model.
[0012] In one embodiment, the initial micronucleus detection model is trained in sequence using the basic dataset and the fine-tuning dataset to obtain a target micronucleus detection model, including: training the initial micronucleus detection model using the basic dataset, and stopping the training when a set first early stopping condition is met, to obtain a preliminarily trained micronucleus detection model; and retraining the preliminarily trained micronucleus detection model using the fine-tuning dataset, and stopping the training when a set second early stopping condition is met, to obtain a target micronucleus detection model.
[0013] Compared with the existing technology, the advantages and beneficial effects of the present invention are as follows: a number of micronucleus images are obtained by screening the micronucleus database and used to construct a basic database, which contains an equal number of binuclear cell images and mononuclear cell images, and the micronucleus images can be color or black and white images. By collecting a variety of micronucleus images as training data for the model, a more comprehensive model can be obtained through training; a number of fine-tuning images are collected based on the target hospital system and used to construct a fine-tuning database. Model training through the fine-tuning database can achieve adaptation to the needs of the target hospital; the basic database and the fine-tuning database are annotated and data enhancement is performed to obtain a basic data set and fine-tuning data for subsequent model training, and through Data augmentation can ensure the validity of training data and improve the training performance of the model; a deep learning model framework is built based on the PyTorch neural network library, and the improved YOLOv11s model is used to construct the initial micronucleus detection model. The improved YOLOv11s model adds an MSAA module to the neck network, and the MSAA module ensures the adequacy of feature extraction and multi-scale information fusion; the basic data set and fine-tuning data set are used to train the initial micronucleus detection model in sequence to obtain the target micronucleus detection model. Through the basic database and fine-tuning database, the demand for labeled data is reduced, the migration of model recognition capabilities is realized, and the performance of the trained model can be ensured to adapt to actual needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 1 is a flow chart of a method for training a micronucleus image deep learning detection model in one embodiment;
[0015] Figure 2 is a schematic diagram of a micronucleus image in one embodiment;
[0016] Figure 3 A schematic diagram of a micronucleus image after annotation in one embodiment;
[0017] Figure 4 Schematic diagram of the structure of the improved Yolo11s network model in one embodiment;
[0018] Figure 5 Schematic diagram of the structure of the C3k2 module in one embodiment;
[0019] Figure 6 Schematic diagram of the structure of a C2PSA module in one embodiment;
[0020] Figure 7 Schematic diagram of the structure of the SPPF module in one embodiment;
[0021] Figure 8 is a schematic structural diagram of an MSAA module in one embodiment;
[0022] Figure 9 is a schematic structural diagram of a detection head in one embodiment;
[0023] Figure 10 FIG. 4 is a graph showing the relationship between the number of iterations and the loss in one embodiment. DETAILED DESCRIPTION
[0024] Before describing the specific embodiments of the present invention, the overall concept of the present invention is described as follows:
[0025] The present invention is mainly developed based on the training process of the micronucleus detection model. Currently, the micronucleus detection model requires a large amount of labeled data, and the model's recognition ability has weak transferability and low performance.
[0026] Therefore, the present invention proposes a training method for a deep learning detection model of micronucleus images, which obtains a number of micronucleus images by screening a micronucleus database and uses them to construct a basic database. The basic database contains an equal number of binuclear cell images and mononuclear cell images, and the micronucleus images can be color or black and white images. By collecting a variety of micronucleus images as training data for the model, a more comprehensive model can be obtained through training; based on the target hospital system, a number of fine-tuning images are collected and used to construct a fine-tuning database. Model training through the fine-tuning database can achieve adaptation to the needs of the target hospital; the basic database and the fine-tuning database are annotated and data enhancement is performed to obtain a basic data set and fine-tuning data for subsequent model training. Data enhancement can ensure the validity of training data and improve the training performance of the model; a deep learning model framework is built based on the PyTorch neural network library, and the improved YOLOv11s model is used to construct the initial micronucleus detection model. The improved YOLOv11s model adds an MSAA module to the neck network, and the MSAA module ensures the adequacy of feature extraction and multi-scale information fusion; the basic data set and the fine-tuning data set are used to train the initial micronucleus detection model in turn to obtain the target micronucleus detection model. Through the basic database and the fine-tuning database, the demand for labeled data is reduced, the migration of model recognition capabilities is realized, and the performance of the training model can be ensured to adapt to actual needs.
[0027] After introducing the overall concept of the present invention, in order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0028] In one embodiment, Figure 1 As shown, a training method for a micronucleus image deep learning detection model is provided, comprising the following steps:
[0029] Step S110 , a plurality of micronucleus images are screened from the micronucleus image database, and a basic database is constructed based on the plurality of micronucleus images. The basic database includes an equal number of binuclear cell images and mononuclear cell images, and the micronucleus images are color images or black and white images.
[0030] Specifically, based on an existing micronucleus image database, several micronucleus images with few impurities and clear images are screened out, and a basic database is constructed based on the selected micronucleus images. The micronucleus images can be images of binucleated cells and mononucleated cells, and can be either color images or black and white images. The screened micronucleus images have an equal number of mononucleated cell images and binucleated cell images. For example, 3000 micronucleus images are screened, of which 1500 are images of binucleated cells and 1500 are images of mononucleated cells.
[0031] Step S120 : collecting a number of fine-tuning images based on the target hospital system, and constructing a fine-tuning database based on the number of fine-tuning images.
[0032] Specifically, in order to adapt to the needs of the target hospital, a number of fine-tuning images can be collected from the target hospital system. The collected fine-tuning images can be mononuclear cell images and / or mononuclear cell images. The fine-tuning images are as follows: Figure 2 As shown in the figure, a fine-tuning database is constructed based on several fine-tuning images collected, so that the model can be trained based on the fine-tuning database, and the trained AI-assisted diagnosis system can be deployed in the target hospital to achieve adaptation of the model to the target hospital.
[0033] Step S130 : annotate the basic database and the fine-tuning database, and perform data enhancement to obtain a basic data set and a fine-tuning data set.
[0034] Specifically, all images in the base database and fine-tuning database are annotated with micro-kernels, e.g. Figure 3 As shown in the figure, the cell nuclei in the image are annotated with bounding boxes. Data augmentation is performed on the annotated images. This introduces variability into the training data to help the model generalize better to unseen data. When performing data augmentation, the Ultralytics software package can be used to perform image augmentation on all images.
[0035] Among them, data enhancement methods include rotation, mirroring, upside down flipping, random hue adjustment, random image saturation adjustment and random brightness adjustment.
[0036] Specifically, when performing data enhancement, image enhancement methods such as random rotation, mirroring, upside-down flipping, random hue adjustment, random image saturation adjustment, and random brightness adjustment can be used to ensure the validity of the data in the dataset and improve the performance of the trained model.
[0037] In step S140 , a deep learning model framework is built based on the PyTorch neural network library, and an improved YOLOv11s model is used to construct an initial micronucleus detection model. The improved YOLOv11s model adds an MSAA module to the neck network.
[0038] Specifically, when constructing the micronucleus detection model, a deep learning model framework was built based on the PyTorch neural network library. The model selected was YOLOv11s. Based on the YOLOv11s model, its neck network was improved and the MSAA (Multi-Scale Attention Aggregation) module was added to take into account both detection speed and accuracy. The initial micronucleus detection model was constructed based on the improved YOLOv11s model.
[0039] Among them, the improved YOLOv11s model includes: a backbone network, a neck network and a detection head. The backbone network is used for feature extraction, including a CBS module, a combination layer of four consecutive CBS modules and C3k2 modules, an SPPF module and a C2PSA module; among them, the C2PSA module is used to divide the input features into two parts, one part is directly passed to retain the original information, and the other part is processed by the PSA attention module, and the two parts are spliced and fused. The PSA attention module enhances the model's perception of target details by dynamically adjusting the feature weights; the neck network is connected to the feature map through upsampling to perform feature fusion of the feature map, including: a first upsampling layer, a combination layer composed of a splicing layer, an MSAA layer and a C3k2 module, a second upsampling layer, a combination layer composed of a splicing layer, an MSAA layer and a C3k2 module, and two layers of a combination layer composed of a CBS module, a splicing layer, an MSAA layer and a C3k2 module; the detection head is provided with multiple layers for generating the final prediction result based on the processed feature map.
[0040] Specifically, if Figure 4 As shown in the figure, the improved YOLOv11s network structure consists of three parts: backbone network (Backbone), neck network (Neck) and detection head (Head).
[0041] Backbone is responsible for feature extraction and uses a series of convolution and deconvolution layers. It also uses residual connections and bottleneck structures to reduce the size of the network and improve performance. YOLOv11s uses C3k2 modules to handle feature extraction at different stages of the backbone. The structure of the C3k2 module is as follows: Figure 5 As shown in the figure, the C3k2 module is an optimized version of the traditional cross-stage partial network CSPNet (Cross Stage Partial Network) structure of YOLOv11s. Its core goal is to improve the efficiency of feature extraction through convolution design and flexible parameter configuration. Its structural features include: a part of the shallow features can be directly passed, and the other part relies on parameter settings to select multiple C3k (parameter is True) or Bottleneck (parameter is False) to perform deep processing on the features, and finally the two parts of the features are fused to achieve an effective combination of multi-level features. When the C3k module is selected, the entire module is Figure 5 When the Bottleneck module is selected for the C3k2 shown, the entire module is converted to Figure 5 C2f shown.
[0042] The backbone network is used for feature extraction and includes a CBS module, a combination of four consecutive CBS modules and a C3k2 module, an SPPF module, and a C2PSA module. The C2PSA module is used to split the input feature map into two parts. One part is directly passed on to retain the original information, while the other part is processed by the PSA attention module and the two parts are spliced and fused. The PSA attention module enhances the model's perception of target details by dynamically adjusting feature weights.
[0043] The C2PSA (Cross Stage Partial with Pyramid Squeeze Attention) module is a core module introduced in YOLOv11s. By combining the CSP structure with the attention mechanism, it significantly improves the feature extraction capability of the model. The structure of the C2PSA module is as follows: Figure 6 As shown in the figure, its core design includes: CSP segmentation processing, which divides the input feature map into two parts. One part is directly passed to retain the original information, and the other part is processed by the Pyramid Squeeze Attention (PSA) attention module, and finally spliced and fused to achieve effective combination of multi-level features.
[0044] The PSA attention mechanism dynamically adjusts feature weights to enhance the model's ability to perceive object details, making it particularly suitable for complex scenes. This method uses parallel convolution with multi-scale kernels, such as 3×3, 5×5, and 7×7, to extract multi-scale features and expand the receptive field. The channel-wise attention mechanism, called Squeeze-and-Excitation (SE), dynamically weights channel features to enhance the response of important channels. Compared to traditional attention mechanisms, C2PSA significantly improves its ability to focus on complex occluded objects and key areas through multi-scale convolution and channel weighting. Furthermore, its lightweight design reduces computational overhead by 50%, while the PSA mechanism introduces only a small number of parameters, maintaining overall high efficiency. The C2PSA module demonstrates outstanding performance in tasks such as oriented object detection, balancing performance and computational efficiency, providing strong support for object detection in complex scenes. In micronucleus images, where multiple micronuclei are present and impurities overlap and occlude each other, the C2PSA module can effectively improve the detection of micronuclei.
[0045] The structure of the SPPF (Spatial Pyramid Pooling Fast, efficient spatial pyramid improvement technology) module is as follows Figure 7 As shown in the figure, it includes convolutional layers, three maximum pooling layers, splicing layers and convolutional layers, which are used to perform pooling operations of different scales, splicing feature maps of different scales together, and improving the detection ability of targets of different sizes.
[0046] The Neck part is connected to the feature map through upsampling, realizing the feature fusion of large, medium and small feature maps, which helps to identify targets of different sizes. In order to connect and enhance feature transfer, the Neck part of the Yolo11 network is improved. The MSAA module is added after the concat layer to compensate for the problems of insufficient feature extraction or insufficient multi-scale information fusion that may exist when splicing features across layers.
[0047] The head part is the part of the model responsible for generating the final prediction. In object detection, it generates bounding boxes and classifies the objects within the bounding boxes.
[0048] Among them, the input channel of the backbone network is 3, the image size is resized to 640*640, including: the first layer, CBS module, the number of filters is 64, the convolution kernel size is 3*3, the step size is 2, and the padding is 0; the second layer, CBS module, the number of filters is 128, the convolution kernel size is 3*3, the step size is 2, and the padding is 0; the third layer, C2f module, the number of channels is 256, the channel expansion factor is set to 0.25, and the number of Bottleneck blocks is 1; the fourth layer, CBS module, the number of filters is 256, the convolution kernel size is 3*3, the step size is 2, and the padding is 0; the fifth layer, C2f module, the number of channels is 512, the channel expansion factor is set to 0.25, and the number of Bottleneck blocks is 1; the sixth layer, CBS module, The number of filters is 512, the convolution kernel size is 3*3, the step size is 2, and the padding is 0; the seventh layer, the C3k2 module, the number of channels is 512, the channel expansion coefficient is 0.5 by default, and the number of C3k modules is 1; the eighth layer, the CBS module, the number of filters is 1024, the convolution kernel size is 3*3, the step size is 2, and the padding is 0; the ninth layer, the C3k2 module, the number of channels is 1024, the channel expansion coefficient is 0.5 by default, and the number of C3k blocks is 1; the tenth layer, the SPPF module, the number of filters is 1024, the convolution kernel size of the convolution layer is 1*1, the step size is 1, the padding is 0, and the convolution kernel size of the pooling layer is 5*5; the eleventh layer, the C2PSA module, the number of filters is 1024, the number of PSABlock blocks is 1, and the channel expansion coefficient is 0.5 by default.
[0049] Specifically, the Bottleneck structure reduces the number of parameters and computation by introducing a narrower layer in the middle layer of the network.
[0050] The neck network includes: the first layer, an upsampling layer with a sampling coefficient of 2; the second layer, a splicing layer, which is the splicing layer of the first layer and the seventh layer of the backbone; the third layer, an MSAA module; the fourth layer, a C2f module, with 512 filters, a channel expansion coefficient of 0.5, and a number of Bottleneck blocks of 1; the fifth layer, an upsampling layer with a sampling coefficient of 2; the sixth layer, a splicing layer, which is the splicing layer of the fourth layer and the fifth layer of the backbone network; the seventh layer, an MSAA module; the eighth layer, a C2 module, with 256 filters, a channel expansion coefficient of 0.5, and a number of Bottleneck blocks of 1; the ninth layer, a convolutional layer, with 256 filters. , convolution kernel size 3*3, step size 2, padding 0; the tenth layer, splicing layer, splicing the seventh layer and the third layer; the eleventh layer, MSAA module; the twelfth layer, C2f module, the number of filters is 512, the channel expansion coefficient is set to 0.5, and the number of Bottleneck blocks is 1; the thirteenth layer, convolution layer, the number of filters is 512, the convolution kernel size is 3*3, the step size is 2, and the padding is 0; the fourteenth layer, splicing layer, splicing the tenth layer and the eleventh layer of the backbone network; the fifteenth layer, MSAA module; the sixteenth layer, C3k2 module, the number of channels is 1024, the channel expansion coefficient defaults to 0.5, and the number of C3k blocks is 1.
[0051] like Figure 8 The figure shows the structure of the MSAA module, which is used to connect and enhance feature transfer to make up for the problems of insufficient feature extraction or insufficient multi-scale information fusion that may exist in cross-layer splicing features.
[0052] Among them, the detection head includes a bounding box prediction network and an object classification prediction network; among them, the bounding box prediction network includes: the first layer, Conv layer, the Conv layer includes a convolution layer, a BN layer and an activation function, the convolution kernel size is 3*3, and the activation function is SiLU; the second layer, Conv layer, is the same as the first layer; the third layer, convolution layer, the convolution kernel size is 1, and padding is 0; the object classification prediction network includes: the first layer, composed of a DWConv layer and a Conv layer, the Conv layer consists of a convolution layer, a BN layer and an activation function, the convolution kernel size is 3*3, the DWConv layer is a depth-separable convolution layer, the convolution kernel size is 3*3, and the stride is 1; the second layer, composed of a DWConv layer and a Conv layer, is the same as the first layer; the third layer, convolution layer, the convolution kernel size is 1, and padding is 0.
[0053] Specifically, the structure of the detection head is as follows Figure 9 As shown, the bounding box is composed of Figure 9 The bounding box prediction network in the upper part makes predictions, and the object classification within the box is determined by Figure 9 The object classification prediction network in the middle and lower parts performs classification prediction.
[0054] Step S150 : Using the basic dataset and the fine-tuning dataset to train the initial micronucleus detection model in sequence to obtain a target micronucleus detection model.
[0055] Specifically, the Ultralytics package integrates the YOLOv11 network. Using the Ultralytics package, you can quickly and easily import network structures and databases, and provide a rich set of enhancement and hyperparameter settings. During training, you only need to set the desired network model, batch size (the number of samples fed into the model during each forward and backward pass during deep learning model training), image size, enhancement parameters, and hyperparameters to train the model with one click.
[0056] Among them, step S150 includes: using the basic data set to train the initial micronucleus detection model, and stopping the training when the set first early stopping condition is met, to obtain the micronucleus detection model after the initial training; using the fine-tuning data set to re-train the micronucleus detection model after the initial training, and stopping the training when the set second early stopping condition is met, to obtain the target micronucleus detection model.
[0057] Specifically, the initial micronucleus detection model is preliminarily trained using the data in the basic dataset. The 3,000 basic dataset images can be divided into an 8:2 ratio, with 80% of the images used for model training and 20% for model validation. During training, the epoch (iterations) is set to 500 and the batch size is set to 4. At the same time, the tensorboard visualization tool suite is used to observe the training process. When the model loss no longer decreases or the number of iterations is reached, the training is terminated and the model is saved. Figure 10 As shown in the figure, the loss no longer decreases around the 160th epoch. Therefore, the model saved at the 158th or 159th epoch can be selected as the basic model required for the next step.
[0058] The data in the fine-tuning dataset is divided into an 8:2 ratio, where 80% of the images are used for model training and 20% of the images are used for model verification. The epoch is set to 500 and the batch size is set to 4. The basic model obtained after training the basic dataset is imported and fine-tuned. During training, the tensorboard visualization tool suite is used to observe the training process. When it is found that the model loss no longer decreases or the number of iterations is reached, the training is terminated and the model is saved to obtain the target micronucleus detection model.
[0059] Through two early stopping techniques, the training time of the model is shortened, and the model has good generalization. The trained target micronucleus detection model can have good performance on both the basic dataset and the fine-tuning dataset.
[0060] Among them, during training, the classification loss function, positioning loss function and distribution focus loss function are used to calculate the loss of the improved YOLOv11s model.
[0061] Specifically, the classification loss function, localization loss function, and distribution focus loss function are used to calculate the loss of the improved YOLOv11s model. By combining multiple loss functions, a target micronucleus detection model with better training results can be obtained. Among them, the classification loss function can handle multi-classification problems and has high computational efficiency. The localization loss function is used for target regression detection, has a certain degree of robustness to outliers, and can maintain a good convergence speed. The distribution focus loss function can better handle the problem of class imbalance. Micronucleus cells occupy a small position in the image and are small targets. The detection of small targets is more difficult. The use of the distribution focus loss function can enhance the model's attention to difficult samples and improve the detection rate of micronucleus cells.
[0062] In this embodiment, a number of micronucleus images are obtained by screening the micronucleus database and used to construct a basic database. The basic database contains an equal number of binuclear cell images and mononuclear cell images, and the micronucleus images can be color or black and white images. By collecting a variety of micronucleus images as training data for the model, a more comprehensive model can be obtained through training; a number of fine-tuning images are collected based on the target hospital system and used to construct a fine-tuning database. Model training through the fine-tuning database can achieve adaptation to the needs of the target hospital; the basic database and the fine-tuning database are annotated, and data enhancement is performed to obtain a basic data set and fine-tuning data for subsequent model training, and data enhancement can ensure The effectiveness of training data is improved to improve the training performance of the model; a deep learning model framework is built based on the PyTorch neural network library, and the improved YOLOv11s model is used to build the initial micronucleus detection model. The improved YOLOv11s model adds an MSAA module to the neck network, and the MSAA module ensures the adequacy of feature extraction and multi-scale information fusion; the basic data set and the fine-tuning data set are used to train the initial micronucleus detection model in sequence to obtain the target micronucleus detection model. Through the basic database and the fine-tuning database, the demand for labeled data is reduced, the migration of model recognition capabilities is realized, and the performance of the training model can be ensured to adapt to actual needs.
[0063] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0064] Obviously, those skilled in the art should understand that the modules or steps of the present invention described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices. Alternatively, they can be implemented using program codes executable by the computing device, so that they can be stored in a computer storage medium (ROM / RAM, magnetic disk, optical disk) and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than herein, or they can be made into individual integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module for implementation. Therefore, the present invention is not limited to any specific combination of hardware and software.
[0065] The above content is a further detailed description of the present invention in conjunction with specific embodiments, and the specific implementation of the present invention cannot be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A training method for a micronucleus image deep learning detection model, characterized in that: The following steps are involved: A plurality of micronucleus images are obtained by screening a micronucleus image database, and a basic database is constructed based on the plurality of micronucleus images, wherein the basic database includes an equal number of binuclear cell images and mononuclear cell images, and the micronucleus images are color images or black and white images; Collecting a number of fine-tuning images based on the target hospital system, and constructing a fine-tuning database based on the number of fine-tuning images; Annotating the basic database and the fine-tuning database, and performing data enhancement to obtain a basic dataset and a fine-tuning dataset; A deep learning model framework was built based on the PyTorch neural network library, and an initial micronucleus detection model was constructed using an improved YOLOv11s model that added an MSAA module to the neck network. The basic data set and the fine-tuning data set are used to train the initial micronucleus detection model in sequence to obtain a target micronucleus detection model.
2. The method for training a micronucleus image deep learning detection model according to claim 1, characterized in that: The data enhancement methods include rotation, mirroring, upside-down flipping, random hue adjustment, random image saturation adjustment and random brightness adjustment.
3. The training method of a micronucleus image deep learning detection model according to claim 1, characterized in that: The improved YOLOv11s model includes: a backbone network, a neck network and a detection head; The backbone network is used for feature extraction and includes a CBS module, a combination layer of four consecutive CBS modules and C3k2 modules, an SPPF module, and a C2PSA module. The C2PSA module is used to divide the input feature map into two parts, one of which is directly passed to retain the original information, and the other is processed by the PSA attention module, and the two parts are spliced and fused. The PSA attention module enhances the model's perception of target details by dynamically adjusting feature weights. The neck network is connected to the feature map through upsampling to perform feature fusion of the feature map, including: a first upsampling layer, a combination layer consisting of a splicing layer, an MSAA layer and a C3k2 module, a second upsampling layer, a combination layer consisting of a splicing layer, an MSAA layer and a C3k2 module, and two layers of combination layers consisting of a CBS module, a splicing layer, an MSAA layer and a C3k2 module; The detection heads are provided in plurality and are used to generate a final prediction result according to the processed feature map.
4. The method for training a micronucleus image deep learning detection model according to claim 3, characterized in that: The backbone network has 3 input channels and the image size is resized to 640*640. It includes: the first layer, CBS module, 64 filters, 3*3 convolution kernel size, 2 steps, and 0 padding; the second layer, CBS module, 128 filters, 3*3 convolution kernel size, 2 steps, and 0 padding; the third layer, C2f module, 256 channels, 0.25 channel expansion coefficient, and 1 Bottleneck block; the fourth layer, CBS module, 256 filters, 3*3 convolution kernel size, 2 steps, and 0 padding; the fifth layer, C2f module, 512 channels, 0.25 channel expansion coefficient, and 1 Bottleneck block; the sixth layer, CBS module, The number of filters is 512, the convolution kernel size is 3*3, the stride is 2, and the padding is 0; the seventh layer, the C3k2 module, the number of channels is 512, the channel expansion coefficient is 0.5 by default, and the number of C3k modules is 1; the eighth layer, the CBS module, the number of filters is 1024, the convolution kernel size is 3*3, the stride is 2, and the padding is 0; the ninth layer, the C3k2 module, the number of channels is 1024, the channel expansion coefficient is 0.5 by default, and the number of C3k blocks is 1; the tenth layer, the SPPF module, the number of filters is 1024, the convolution kernel size of the convolution layer is 1*1, the stride is 1, the padding is 0, and the convolution kernel size of the pooling layer is 5*5; the eleventh layer, the C2PSA module, the number of filters is 1024, the number of PSABlock blocks is 1, and the channel expansion coefficient is 0.5 by default.
5. The method for training a micronucleus image deep learning detection model according to claim 4, characterized in that: The neck network includes: the first layer, an upsampling layer, with a sampling coefficient of 2; the second layer, a splicing layer, which is a splicing layer of the first layer and the seventh layer of the backbone; the third layer, an MSAA module; the fourth layer, a C2f module, with 512 filters, a channel expansion coefficient of 0.5, and a number of Bottleneck blocks of 1; the fifth layer, an upsampling layer, with a sampling coefficient of 2; the sixth layer, a splicing layer, which is a splicing layer of the fourth layer and the fifth layer of the backbone network; the seventh layer, an MSAA module; the eighth layer, a C2 module, with 256 filters, a channel expansion coefficient of 0.5, and a number of Bottleneck blocks of 1; the ninth layer, a convolutional layer, with 256 filters , convolution kernel size 3*3, step size 2, padding 0; the tenth layer, splicing layer, splicing the seventh layer and the third layer; the eleventh layer, MSAA module; the twelfth layer, C2f module, the number of filters is 512, the channel expansion coefficient is set to 0.5, and the number of Bottleneck blocks is 1; the thirteenth layer, convolution layer, the number of filters is 512, the convolution kernel size is 3*3, the step size is 2, and the padding is 0; the fourteenth layer, splicing layer, splicing the tenth layer and the eleventh layer of the backbone network; the fifteenth layer, MSAA module; the sixteenth layer, C3k2 module, the number of channels is 1024, the channel expansion coefficient defaults to 0.5, and the number of C3k blocks is 1.
6. The method for training a micronucleus image deep learning detection model according to claim 3, characterized in that: The detection head includes a bounding box prediction network and an object classification prediction network; The bounding box prediction network includes: the first layer, the Conv layer, which includes a convolution layer, a BN layer, and an activation function, with a convolution kernel size of 3*3 and an activation function of SiLU; the second layer, the Conv layer, which is the same as the first layer; the third layer, the convolution layer, with a convolution kernel size of 1 and a padding of 0; The object classification prediction network includes: the first layer, consisting of a DWConv layer and a Conv layer, the Conv layer consists of a convolution layer, a BN layer and an activation function, the convolution kernel size is 3*3, the DWConv layer is a depth-separable convolution layer, the convolution kernel size is 3*3, and the step size is 1; the second layer, consisting of a DWConv layer and a Conv layer, is the same as the first layer; the third layer is a convolution layer, the convolution kernel size is 1, and the padding is 0.
7. The method for training a micronucleus image deep learning detection model according to claim 1, characterized in that: Also includes: The classification loss function, positioning loss function and distribution focus loss function are used to calculate the loss of the improved YOLOv11s model.
8. The method for training a micronucleus image deep learning detection model according to claim 1, characterized in that: The initial micronucleus detection model is trained in sequence using the basic dataset and the fine-tuning dataset to obtain a target micronucleus detection model, including: The initial micronucleus detection model is trained using the basic data set, and the training is stopped when a set first early stopping condition is met, thereby obtaining a micronucleus detection model after preliminary training; The fine-tuning dataset is used to retrain the micronucleus detection model after preliminary training, and the training is stopped when the set second early stopping condition is met to obtain the target micronucleus detection model.
Citation Information
Patent Citations
Incremental learning method and device based on deep learning detection task and storage medium
CN115496203A
Worker safety behavior detection method and system, storage medium and electronic equipment
CN117636266A
Training method of colorectal cancer microsatellite state detection model
CN118247548A
Unmanned aerial vehicle aerial photography mangrove forest target detection method based on improved YOLOv11
CN120182871A