Saccharomyces cerevisiae contamination detection method based on improved YOLOv10n image recognition
By combining microscopic imaging technology and improving the YOLOv10n model, the morphological characteristics of Saccharomyces cerevisiae and Candida cerevisiae are trained, which solves the problem of insufficient rapid and accurate detection of Saccharomyces cerevisiae in the prior art, and achieves rapid, low cost and high accuracy of Saccharomyces cerevisiae.
Patent Information
- Application Number
- CN202510277815.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has insufficient fast and low-cost methods in the detection of Saccharomyces cerevisiae, which leads to the detection results relying on experience, taking a long time, and misjudgment and false detection.
Through the combination of microscopic imaging technology and object detection, the morphological characteristics of Saccharomyces cerevisiae and Candida are trained using the improved YOLOv10n model to achieve real-time detection of the bacterial infection of Saccharomyces cerevisiae.
The rapid and accurate detection of the bacteria-contaminated conditions of Saccharomyces cerevisiae is achieved, with short time and low cost. The improved YOLOv10n model can operate efficiently even under resource constraints, and the detection accuracy and recall rate are significantly improved.
Smart Images

Figure CN120220145A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting the contamination of Saccharomyces cerevisiae by improved YOLOv10n image recognition, belonging to the technical field of biological detection. Background Art
[0002] Saccharomyces cerevisiae is a single-celled fungus and is the most valuable species in various industrial applications. This is mainly because of its Crabtree effect, which is characterized in that Saccharomyces cerevisiae does not utilize the respiratory mechanism to metabolize sugars and promote biomass growth even under aerobic conditions, but produces ethanol and other two-carbon compounds through pyruvate. In the fermentation of Saccharomyces cerevisiae, Candida is a common contaminant, which affects the fermentation process and yield. At the same time, in recent years, the wine industry has adopted the fermentation of mixed yeast inoculants, and the presence of Candida as a mixed fermenting agent can change the sensory characteristics of wine. Therefore, the dynamic change activities of Candida in the fermentation of Saccharomyces cerevisiae need to be detected in real time and accurately.
[0003] Common detection methods include: direct microscopy: taking a sample of the fermentation broth and directly smearing it, performing Gram staining and then microscopic examination. Candida appears as Gram-positive, oval or round spores and pseudohyphae under the microscope; culture method: inoculating the sample on a suitable medium, such as Sabouraud medium or Candida chromogenic medium, culturing it at 25-28 °C for 48 hours, and observing the colony growth. Candida albicans usually presents blue-green on the Candida chromogenic medium, while Candida tropicalis presents blue-gray or iron-blue; molecular biology method: using polymerase chain reaction (PCR) technology to amplify the Candida DNA in the sample and detecting it through specific probes; immunological method: detecting the Candida antigen in the sample through methods such as enzyme-linked immunosorbent assay (ELISA).
[0004] Microscopic examination is the most commonly used detection method, but this method also has some disadvantages. The results of microscopic examination depend on the experience and skills of the examiner, and different examiners may have different judgments on the same sample, which may lead to misjudgment of the results; microscopic examination cannot distinguish whether the Candida in the sample is a live bacterium or a dead bacterium, which may affect the accurate assessment of the infection status in some cases; microscopic examination requires multiple steps such as sample preparation, staining, and observation, and the process is relatively cumbersome and time-consuming. The culture method examination takes a long time, is complex in operation, and is easily affected by other microorganisms. The molecular biology method requires high costs, is technically complex, and requires strict personnel training. The immunological method has low sensitivity and cannot detect low-concentration microorganisms. Currently, there is a lack of a rapid and low-cost method for detecting the contamination of Saccharomyces cerevisiae that can be used in production. Summary of the Invention
[0005] To address the deficiencies of the existing technology, the present invention combines microscopic imaging technology with target detection. Based on the fact that Saccharomyces cerevisiae exhibits different morphologies under the influence of Candida contamination during fermentation, further label classification is carried out, and deep learning technology is used to train the different morphologies, so as to judge the contamination degree of Saccharomyces cerevisiae by Candida and the changes in its own physiological state.
[0006] The present invention is achieved through the following technical solutions:
[0007] The first object of the present invention is to provide a method for detecting Candida contamination in Saccharomyces cerevisiae based on improved YOLOv10n image recognition, including the following steps:
[0008] S1. According to the morphological characteristics of Saccharomyces cerevisiae and Candida at different growth stages, and in accordance with the budding mode, the yeast morphology is divided into the G1 phase without budding, the G2 phase in which the bud body has not formed a single cell and is less than one-fifth of the mother cell, the M phase in which a single-cell bud body is formed, unilateral multi-budding in which two or more bud bodies appear at one end of the mother cell, two-end single-budding in which a single bud body appears at both ends of the mother cell, two-end multi-budding in which two or more bud bodies appear at both ends of the mother cell, and a bud chain with budding at the distal pole of the bud body;
[0009] S2. Prepare a data set, label the Saccharomyces cerevisiae cell images and Candida cell images according to the classification in step S1, and divide the labeled image data set into a training set and a validation set;
[0010] S3. Construct an improved YOLOv10n model based on YOLOv10n, including a backbone feature extraction network, a feature fusion network, and a detection head;
[0011] S4. Use the training set prepared in step S2 to train the improved YOLOv10n model to obtain the best training model; input the images in the validation set into the best training model obtained by training for verification and then save the best model;
[0012] S5. Use the best model to detect the image to be detected and judge the Candida contamination situation of Saccharomyces cerevisiae.
[0013] In an embodiment of the present invention, a multi-scale feature sharing pyramid convolution module is introduced between the shallow feature layer and the deep feature layer of the backbone network in the backbone feature extraction network to extract the features of Saccharomyces cerevisiae and Candida cells from multiple scales.
[0014] In an embodiment of the present invention, in step S4, the multi-scale feature sharing pyramid convolution module consists of 6 CBSs, k represents the convolution kernel size, and d represents the dilation multiple; for the input The feature map first undergoes processing by CBS to obtain an intermediate feature map F1, and then F1 is input into the dilated convolutional layer; in each layer, the output feature map is fused with the features of the next dilated convolutional layer; the output feature maps of four dilated convolutions, F 12 、F 123 、F 1234 and F 12345 are obtained respectively. After splicing these four feature maps, a new feature map is obtained through CBS, and its expression is shown in the following formula (1):
[0015]
[0016] where represents the splicing operation, W represents the width of the input feature map, H represents the height of the input feature map, and C in represents the number of channels of the input feature map; F out represents the output feature map.
[0017] In an embodiment of the present invention, a channel prior convolutional attention module is introduced in the last layer of the backbone feature extraction network to support the dynamic allocation of attention weights in the channel and spatial dimensions, enabling the model to pay more attention to the differences between Saccharomyces cerevisiae cells and Candida albicans cells.
[0018] In an embodiment of the present invention, the channel prior convolutional attention module first parallelizes the input feature map F∈R C×H×W through global pooling and max pooling respectively, flattens them into one-dimensional vectors through MLP, and then performs sigmoid activation functions respectively to obtain channel weight parameters ω1 N×C×1×1 and ω2 N×C×1×1 . According to the obtained weights, they are added and then multiplied by the initial feature map F∈R C×H×W to obtain an enhanced channel information feature map F CA ; as shown in the following formula (2):
[0019] F CA =σ(MLP(AvgPool(F))+MLP(MaxPool(F))) (2)
[0020] where MLP is a multi-layer perceptron and σ is the Sigmoid activation function;
[0021] Then, F CA is input into depthwise separable convolutions of different scale sizes for feature extraction and fusion to obtain F SA ; then, a 1×1 convolutional structure is used to adjust the dimension of the fused feature map to obtain F′ SA , as shown in the following formula (3):
[0022]
[0023] Among them, DWConv represents depthwise separable convolution;
[0024] Finally, multiply F CA and F' SA element-wise to obtain the enhanced feature fusion map F o ∈R C×H×W as shown in the following formula (4):
[0025]
[0026] Where represents the dot product operation; F o : the output feature map; F CA : the feature after channel attention processing; F' SA : the feature after spatial attention processing.
[0027] In one embodiment of the present invention, focal loss is introduced into the classification loss of the detection head and dynamically combined with the cross-entropy loss through a weight factor.
[0028] Original loss function: In the YOLOv10n baseline model, the classification loss uses cross-entropy (CE), and the bounding box regression loss uses CIoU. The formula is: classification loss = CE, regression loss = CIoU. Improved loss function: To solve the problem of sample imbalance, focal loss (FL) is introduced into the classification loss and dynamically combined with the cross-entropy loss through a weight factor: Cross-entropy (CE) is the original loss function, and focal loss (FL) is the newly introduced supplementary term. The weight factor controls the contribution ratio of the two losses (e.g., = 0.5 means equal-weight mixing). During training, the model calculates both CE and FL simultaneously and optimizes the combined loss according to the weighted sum. Summary: Introduction of Focal Loss: Combined with the original cross-entropy loss, it improves the learning ability for hard samples and minority classes through weighted mixing.
[0029] In one embodiment of the present invention, the calculation method of the loss function used during the training of the model is as follows:
[0030] L = u × CE + (1 - u) × FL (7)
[0031] Where L represents the total loss (Total Loss), the final optimization objective, which is obtained by the weighted sum of the two losses. CE represents the cross-entropy loss of classification, FL represents the focal loss of classification, and u is the weight factor between the two losses, and its value range is between 0 and 1.
[0032] In one embodiment of the present invention, the calculation formula of the cross-entropy loss is as follows:
[0033]
[0034] Among them, CE represents the cross-entropy loss for classification, x represents the sample, y represents the label, a represents the predicted output, and n represents the total number of samples.
[0035] In one implementation of the present invention, the calculation formula of the focal loss is as follows:
[0036] FL = α(1 - a) γ ×C(6)
[0037] Among them, FL represents the focal loss for classification, α represents the sample weight factor, a represents the predicted output, γ represents the adjustment factor, C represents the cross-entropy loss term, usually written as -log(P t ), where P t is the predicted probability of the model for the target class.
[0038] The second object of the present invention is to provide a computer device, which includes a processor and a memory. The memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to implement the method for detecting contaminated yeast in Saccharomyces cerevisiae based on improved YOLOv10n image recognition.
[0039] Advantages of the present invention:
[0040] By combining microscopic imaging technology with target detection, the present invention can detect the dynamic changes between Saccharomyces cerevisiae and Candida albicans in real time, can detect the main contaminant Candida albicans during fermentation, is time-consuming short, convenient and fast. The improved YOLOv10n has only 2.45MB of parameters and 6.6GFLOPs, which is convenient to be deployed in small portable detection devices and can operate efficiently even under resource-constrained conditions.
[0041] In the present invention, the average detection accuracies of Saccharomyces cerevisiae and Candida albicans are 95.0% and 87.7% respectively. Compared with the existing YOLOv10n, the improved model has improved the precision, recall rate, and mAP@0.5 by 1.2%, 1.2%, and 2.0% respectively. It has improved by 1.6% in mAP@0.5:0.95. The improved YOLOv10n has only 2.45MB of parameters and 6.6GFLOPs, which is more convenient to be deployed in small portable detection devices and can operate efficiently even under resource-constrained conditions. Description of the Drawings
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0043] Figure 1 Changes in the growth and budding rate of Saccharomyces cerevisiae at different concentrations of Candida albicans; (a) cell concentration; (b) budding rate;
[0044] Figure 2 Morphological characteristics of co-culture of Saccharomyces cerevisiae and Candida albicans;
[0045] Figure 3 Improved YOLOv10 model;
[0046] Figure 4 MPSC module;
[0047] Figure 5 CPCA module. Detailed implementation manners
[0048] The following will further elaborate on the present invention patent in combination with specific examples. These implementation cases are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0049] The following will elaborate in detail on the technical solutions of the present invention in combination with specific embodiments. In the following embodiments, unless otherwise specified, the reagents, materials, and equipment used can be obtained from commercial channels, or prepared by conventional methods, or commonly used in this industry.
[0050] Embodiment 1:
[0051] Candida tropicalis ATCC20336 and Saccharomyces cerevisiae W3031A were taken out from the frozen glycerol tubes, streaked on YPD plates, and after culturing for 48 h, single colonies were picked and transferred into culture tubes containing YPD liquid medium. They were placed in a shaker and cultured at 30 °C and 200 rpm for 12 h. After the culture was completed, the culture was transferred to a centrifuge tube, centrifuged at 5000 rpm for 1 min, the supernatant was poured out after centrifugation, washed twice with PBS solution buffer, and sterilized YPD liquid medium was added. The activated yeast culture solution of Saccharomyces cerevisiae and Candida tropicalis were mixed at ratios of 1:1, 1:100, and 100:1 respectively to co-culture Saccharomyces cerevisiae and Candida tropicalis; the yeast culture solutions of Saccharomyces cerevisiae and Candida tropicalis at the same ratios were mixed with YPD broth respectively as the single culture groups of Saccharomyces cerevisiae and Candida tropicalis. The culture solution was placed at 30 °C and 200 rpm for culture, and samples were taken every 4 hours to measure the cell number with a hemocytometer. The above experiments were repeated in parallel 3 times. As Figure 1 shown, within 12 hours, Candida had no obvious inhibitory effect on the growth of Saccharomyces cerevisiae. After 12 hours, compared with the single culture of Saccharomyces cerevisiae, the growth of Saccharomyces cerevisiae in the co-culture was significantly inhibited, and the higher the relative proportion of the initial concentration of Candida, the stronger the inhibitory effect on the growth of Saccharomyces cerevisiae. The budding rate of Saccharomyces cerevisiae was inhibited by Candida. Compared with the single culture, the inhibition of the budding rate of Saccharomyces cerevisiae by Candida was positively correlated with the initial concentration. At the 12th hour (the initial concentration ratio of Saccharomyces cerevisiae and Candida was 1:100, 1:1, 100:1), the inhibition rates of budding of Saccharomyces cerevisiae were 36.4%, 3.2%, and 19.4% respectively, indicating that the higher the relative proportion of Candida, the stronger the inhibitory effect on Saccharomyces cerevisiae.
[0052] All morphological characteristic changes were observed under the microscope. The cell classifications of Saccharomyces cerevisiae and Candida were respectively as Figure 2As shown in the figure. First, during the G1 phase of the first cycle of yeast cells, the cell volume increases, essential proteins and organelles are synthesized, and the cells carry out growth and metabolic activities in the G1 phase to prepare for DNA replication. At this time, the cell morphology shows a single cell without budding. During the G2 phase, the cells continue to grow: it is a gap period after DNA synthesis and before cell division. The cells will continue to grow during this period, further synthesize proteins and other cell components to prepare for cell division. At this time, the yeast cells begin to bud, and the budding is obvious, but the bud is less than one-fifth of the mother cell and has not become an independent cell. During the M phase, cell division is obvious, and the yeast cells form obvious single cells. Yeast cells can bud from the same end or both ends, and the number of budding cells is also different. The budding at each end may be one, two, or even more. We name the morphology where two or more buds appear at a single end of the mother cell as unilateral multi-budding (SUM, CUM); when single buds appear at both ends of the mother cell, we call it two-end single budding (STB, CTB); when two or more buds appear at both ends of the mother cell, we call it two-end multi-budding (SMB, CMB); after the mother cell buds again, its bud becomes the mother cell and will bud again to form a chain structure. We call this method bud chain (SB, CB). The cell morphology of Saccharomyces cerevisiae is diverse, commonly round, oval, or elliptical, and some are cylindrical, lemon-shaped, etc.; Candida cells are often oval or round in liquid culture, and there is a certain overlap in their shapes. Both reproduce asexually by budding and form a connection structure between the mother cell and the daughter cell. We classify the similar structures of the two to find out their different structural characteristics. For example, the two-end multi-budding (SMB) of Saccharomyces cerevisiae and the two-end multi-budding (CMB) of Candida are both budding irregularly at both ends of the cell and are difficult to distinguish. However, through the classification of this application and training with a large number of data pictures, the differences can be found in the similar places, so as to better distinguish Saccharomyces cerevisiae and Candida, and further judge the degree of contamination of Saccharomyces cerevisiae by Candida and the changes in its own physiological state.
[0053] Example 2:
[0054] Due to problems such as the similar morphology of the main contaminant Candida and Saccharomyces cerevisiae, small differences between different budding methods, small budding cells, and serious imbalance in data samples, there will be many missed detections and misdetections, and it takes a lot of time for the testers to conduct a review. To address the above problems, in this study, YOLOv10n was used as the baseline model, and an improved YOLOv10n was designed to detect Saccharomyces cerevisiae and contaminants. The improved YOLOv10 model is as Figure 3 shown.
[0055] First, the Multi-Feature Pyramid SharedConv (MPSC) module is introduced. By introducing convolutional layers with different dilation rates, it can extract the features of Saccharomyces cerevisiae and contaminating bacteria cells at multiple scales. The MPSC module is limited between specific levels of the backbone network. The specific design is as follows: The MPSC module aims to extract features from multiple scales through convolutional layers with different dilation rates and improve efficiency by combining the shared convolution design. Deep features (Stage5) contain global semantic information but have low resolution; shallow features (Stage3) have high resolution but less semantic information. The MPSC module expands the receptive field through multi-dilation rate convolution while retaining local details, making it suitable for introducing cross-level feature transfer. Introducing the MPSC module between the last few layers (between Stage3 and Stage5) of the backbone network can enhance the detection ability of minute contaminating Candida yeast.
[0056] The MPSC module can not only capture the local and global features of contaminated bacteria and Saccharomyces cerevisiae cells at different scales but also use convolutions with multiple dilation rates to expand the model's receptive field, thereby improving the model's feature extraction ability. Moreover, the MPSC module adopts the shared convolutional layer design, which can reduce redundant calculations and improve the inference efficiency of the model. In the present invention, introducing MPSC into the YOLOv10n model mainly focuses on enhancing the feature extraction ability, but the specific implementation involves the following multi-faceted optimizations: (1) Enhancing the core mechanism of feature extraction: By processing features with different receptive fields through parallel paths, capturing local details and global context information, improving the detection effects of small and large targets, and achieving multi-scale feature fusion; The multi-path design may combine dilated convolution, deformable convolution, etc. to retain spatial information; The traditional single-path structure may lose details due to successive downsampling layer by layer, while the multi-path structure retains more original information through branches, thereby reducing information loss. (2) Balancing efficiency and performance: Similar to the divide-and-conquer strategy of CSPNet, splitting the feature map and processing it on different paths before fusion, reducing redundant calculations, controlling the number of parameters while improving performance, and achieving computational optimization; Integrating depthwise separable convolution or grouped convolution to reduce the computational cost, which is suitable for mobile deployment and realizes lightweight design. (3) Enhancing the robustness of features: After adding MPSC, the model can extract features from multiple angles and paths for the image, which makes the model more robust to various changes and noises in the image. MPSC can better capture the key features of the target through feature extraction and fusion of different paths, reducing the impact of these factors on feature extraction, thereby improving the stability and accuracy of the model.
[0057] Introduce the Channel prior convolutional attention (CPCA) module at the last layer of the backbone. By adopting a multi-scale depth convolutional module, it supports the dynamic allocation of attention weights in both the channel and spatial dimensions, pays more attention to the differences between yeast cells and contaminated bacteria, and can suppress the interference of background noise on model recognition. Due to the severe sample imbalance, the Focal Loss function is introduced to make the model pay more attention to the categories with fewer samples and improve the model accuracy. The Focal Loss function not only pays more attention to the categories with fewer samples but also reflects in aspects such as focusing on difficult-to-classify samples: (1) Solve the sample imbalance problem: When we perform image recognition to detect yeast, in addition to the positive samples such as Saccharomyces cerevisiae and Candida albicans, there are also negative samples such as dust, other bacteria, or microorganisms. The number of negative samples corresponding to the background area is much larger than the number of positive samples containing the target. The Focal Loss function can reduce the weights of a large number of simple and easy-to-classify negative samples by introducing a modulation factor, and relatively increase the weights of positive samples and the categories with fewer numbers (unilateral multi-budding SUM, CUM categories, and bipolar multi-budding SMB, CMB categories), so that the model pays more attention to these categories with fewer samples during the training process, thereby improving the detection or classification performance of the minority categories. (2) Focus on difficult samples: The bud chains (SB) formed by Saccharomyces cerevisiae and STB (bipolar multi-budding of Saccharomyces cerevisiae) have very similar morphologies and are difficult to distinguish. The difference between them lies in the budding sites, and it is difficult for the model to classify them correctly. At this time, the Focal Loss function automatically adjusts the weights according to the predicted probability of the samples. For samples that are easy to classify (predicted probability close to 1), the value of the loss function will be reduced; while for samples that are difficult to classify (predicted probability close to 0.5), the value of the loss function will relatively increase. In this way, the model will be more inclined to learn the features of these difficult samples and improve the generalization ability of the model.
[0058] The MPSC module is as Figure 4 shown, consisting of 6 CBS (Conv Batch Normalization and SiLU). k represents the kernel size, and d represents the dilation factor. For the input feature map, it is first processed by CBS to obtain the intermediate feature map F1, and then F1 is input into the dilated convolutional layer. In each layer, the output feature map is not only local but also fused with the features of the next dilated convolutional layer. Four output feature maps F 12 、F 123 、F 1234 and F 12345 can be obtained respectively. After splicing these four feature maps, a new feature map is obtained through CBS, and its expression is shown in the following formula (1):
[0059]
[0060] Among them represents the splicing operation, W represents the width of the input feature map, H represents the height of the input feature map, and C in represents the number of channels of the input feature map; F out represents the output feature map (Output Feature Map), which is used for subsequent tasks (such as classification and regression in object detection), integrates features at different levels, and has both high-resolution details and deep semantic information. Its structural diagram is as shown in Figure 3 the figure. The MPSC module can not only capture the local and global features of contaminated bacteria and Saccharomyces cerevisiae cells at different scales, but also use convolutions with multiple dilation rates to expand the receptive field of the model, thereby improving the feature extraction ability of the model.
[0061] The CPCA module is as shown in Figure 5 the figure. First, it parallelizes the input feature map F∈R C×H×W through global pooling and max pooling, flattens them into one-dimensional vectors through MLP respectively, and then obtains the channel weight parameters ω1 N×C×1×1 and ω2 N×C×1×1 through the sigmoid activation function respectively. According to the obtained weights, they are added and then multiplied by the initial feature map F∈R C×H×W to obtain the enhanced channel information feature map F CA ; as shown in the following formula (2):
[0062] F CA =σ(MLP(AvgPool(F)) + MLP(MaxPool(F))) (2)
[0063] where MLP is a multi-layer perceptron and σ is the Sigmoid activation function. Then, F CA is input into depthwise separable convolutions of different scales for feature extraction and fusion to obtain F SA ; then, a 1×1 convolution structure is used to adjust the dimension of the fused feature map to obtain F′ SA , as shown in the following formula (3):
[0064]
[0065] where DWConv represents the depthwise separable convolution. Finally, F CA and F′ SA are element-wise multiplied to obtain the enhanced feature fusion map F o ∈R C×H×W , as shown in the following formula (4):
[0066]
[0067] Among them represents the dot product operation, that is, the corresponding position elements of two feature maps are multiplied, which is used to fuse the attention weights of channels and spaces;
[0068] F o : Output Feature Map, which represents the final feature after attention or fusion operations.
[0069] F CA : The feature after channel attention processing. Usually, global pooling and fully connected layers are used to generate channel weights to enhance the information of important channels.
[0070] F′ SA : The feature after spatial attention processing (possibly adjusted, such as transposed or dimensional transformation), which is used to enhance the important spatial regions in the feature map.
[0071] While retaining the channel prior knowledge, the CPCA module also effectively extracts spatial features. This structure uses multi-scale depthwise separable convolution to achieve dynamic weight allocation in both channel and spatial dimensions, thereby reducing the noise interference of the background.
[0072] The focal loss is essentially an improvement of the cross-entropy loss. The original loss function: In the YOLOv10n baseline model, the classification loss uses cross-entropy (CE), and the bounding box regression loss uses CIoU. The formula is: classification loss = CE, regression loss = CIoU. The improved loss function: To solve the problem of sample imbalance, Focal Loss (FL) is introduced into the classification loss and dynamically combined with the cross-entropy loss through the weight factor u: L = u×CE+(1 - u)×FL, where cross-entropy (CE) is the original loss function, and focal loss (FL) is the newly introduced supplementary term. The weight factor u controls the contribution ratio of the two losses (e.g., u = 0.5 means equal-weight mixing). During training, the model calculates CE and FL simultaneously and sums them up weighted according to u, and finally optimizes the mixed loss. Introduction of Focal Loss: Combined with the original cross-entropy loss, it improves the learning ability for difficult samples and minority classes through weighted mixing. In a binary classification task, for a sample, the calculation method of its cross-entropy loss is as follows:
[0073]
[0074] Among them, x represents the sample, y represents the label, a represents the predicted output, and n represents the total number of samples. The magnitude of a can reflect the degree of difficulty in classifying the sample. The larger a is, the easier it is to distinguish the sample; conversely, the smaller a is, the more difficult it is to classify the sample. The value of a for each sample will change as the model is trained. When training the model, due to the small number of samples, the classification effect of the model on the minority-class samples is poor, and they belong to difficult-to-classify samples. Therefore, a sample weight factor α is introduced to increase the proportion of minority-class samples in the loss function. To enable the model to also pay attention to these samples, is introduced to adjust the weights of the samples, and the fast scaling characteristic of the power function is used to dynamically adjust the rate at which the weights of simple samples decrease. Therefore, the calculation formula for the focal loss can be obtained:
[0075] FL = α(1 - a) γ ×C (6)
[0076] γ is used to adjust the weights of difficult-to-classify samples, with high loss weights for difficult-to-classify samples and low loss weights for easy-to-classify samples. α is used to adjust the weights for calculating losses of different-class samples. Usually, the weights of the majority class are low, and the weights of the minority class are high. By setting appropriate values of γ and α, the losses are dynamically adjusted to control the attention to difficult-to-classify and easy-to-classify samples, enhance the model's learning of difficult-to-classify samples, and weaken the influence of easy-to-classify samples. The weight relationship between the cross-entropy loss and FocalLoss is controlled by the parameter u. The calculation method of the loss function used during the model training process is as follows:
[0077] L = u×CE + (1 - u)×FL (7)
[0078] Among them, L represents the total loss, CE represents the cross-entropy loss of classification, FL represents the focal loss of classification, and u is the weight factor between the two losses, and its value range is between 0 and 1. When the value of u is 1, it means that only the cross-entropy loss is used for model training; when the value of u is 0, it means that only the focal loss is used for model training.
[0079] Example 3:
[0080] To comprehensively evaluate the improved YOLOv10 algorithm for the current mainstream single-stage object detection algorithm models, a detailed comparison is made with the currently popular object detection algorithms YOLOv5, YOLOv6, and YOLOv8. The results are shown in Table 2.
[0081] YOLOv5: Developed by the Ultralytics team, implemented based on PyTorch, it is an improved version of YOLOv3 / v4, mainly released through code and community documentation. Literature: Glenn Jocher. (2020). YOLOv5: State-of-the-Art Object Detection at 1.7ms. Code repository: Ultralytics YOLOv5 GitHub, Technical documentation: Ultralytics YOLOv5 Docs. Background of YOLOv6: Proposed by the Meituan team, focusing on industrial-level object detection, as shown in the specific literature: YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications. Background of YOLOv8: An upgraded version of the Ultralytics team after YOLOv5, optimizing the model structure and training strategy. Ultralytics. (2023). YOLOv8: The Latest Version of YOLO for Real-Time Object Detection. Background of YOLOv10: Proposed by teams such as Tsinghua University, as shown in the specific literature: Wang, W., Liao, L., Zhao, F., et al. (2024). YOLOv10: Real-Time End-to-End Object Detection. arXiv:2405.14458.
[0082] Table 2 compares different YOLO series algorithms
[0083]
[0084] The improved YOLOv10 algorithm has higher precision, recall, and mAP@0.5 than other YOLO series models. Compared with YOLOv10n, the improved model has increased precision, recall, and mAP@0.5 by 1.2%, 1.2%, and 2.0% respectively. It has increased by 1.6% in mAP@0.5:0.95. The improved YOLOv10n has only 2.45MB of parameters and 6.6 GFLOPs, making it more convenient to be deployed on small portable detection devices and can operate efficiently even under resource constraints.
[0085] The embodiments provided above are not intended to limit the scope covered by the present invention, nor are the described steps intended to limit their execution order. Obvious improvements made by those skilled in the art to the present invention in combination with the existing common general knowledge also fall within the protection scope defined by the claims of the present invention.
Claims
1. A method for detecting yeast contamination based on improved YOLOv10n image recognition, characterized in that: The steps include: S1. Based on the morphological characteristics of Saccharomyces cerevisiae and Candida albicans at different growth stages, the yeast morphology is divided into the G1 phase without budding, the G2 phase in which the bud has not formed a single cell and is smaller than one-fifth of the mother cell, the M phase in which the bud forms a single cell, the unilateral multiple budding with two or more buds at one end of the mother cell, the two-end single budding with a single bud at both ends of the mother cell, the two-end multiple budding with two or more buds at both ends of the mother cell, and the bud chain with buds budding at the distal pole. S2, preparing a data set, annotating the Saccharomyces cerevisiae cell images and the Candida albicans cell images according to the classification in step S1, and dividing the annotated image data set into a training set and a validation set; S3. Build an improved YOLOv10n model based on YOLOv10n, including a backbone feature extraction network, a feature fusion network, and a detection head; S4, using the training set prepared in step S2 to train the improved YOLOv10n model to obtain the best training model; input the images in the verification set into the best training model obtained by training, verify it, and then save the best model; S5. Use the best model to detect the image to be detected and determine the contamination status of brewer's yeast.
2. The method for detecting bacterial contamination of brewer's yeast according to claim 1, characterized in that: A multi-scale feature sharing pyramid convolution module is introduced between the shallow feature layer and the deep feature layer of the backbone network in the backbone feature extraction network to extract the features of Saccharomyces cerevisiae and Candida albicans cells from multiple scales.
3. The method for detecting bacterial contamination of brewer's yeast according to claim 2, characterized in that: The multi-scale feature sharing pyramid convolution module consists of 6 CBS, k represents the convolution kernel size, d represents the expansion multiple; for input The feature map of is first processed by CBS to obtain the intermediate feature map F1, and then F1 is input into the dilated convolution layer; in each layer, the output feature map is fused with the feature of the next dilated convolution layer; four dilated convolution output feature maps F are obtained respectively. 12 、F 123 、F 1234 and F 12345 , these four feature maps are concatenated and then CBS is used to obtain a new feature map, which is expressed as follows (1): in represents the concatenation operation, W represents the width of the input feature map, H represents the height of the input feature map, and C in Indicates the number of channels of the input feature map; F out Represents the output feature map.
4. The method for detecting bacterial contamination of brewer's yeast according to claim 1, characterized in that: A channel prior convolutional attention module is introduced in the last layer of the backbone feature extraction network to support the dynamic allocation of attention weights in channel and spatial dimensions, so that the model pays more attention to the differences between Saccharomyces cerevisiae cells and Candida albicans cells.
5. The method for detecting bacterial contamination of brewer's yeast according to claim 4, characterized in that: The channel prior convolutional attention module first inputs the feature map F∈R C×H×W After global pooling and maximum pooling in parallel, they are flattened into one-dimensional vectors through MLP, and then the sigmoid activation function is performed to obtain the channel weight parameter ω1 N×C×1×1 and ω2 N×C×1×1 , add the obtained weights and then add them to the initial feature map F∈R C×H×W Multiply them together to get the enhanced channel information feature map F CA ; As shown in the following formula (2): F CA =σ(MLP(AvgPool(F))+MLP(MaxPool(F))) (2) Where MLP is a multi-layer perceptron and σ is a Sigmoid activation function; Then, F CA Input different scales of depth-separable convolution to extract features and fuse them to get F SA ; Then, use the 1×1 convolution structure to adjust the dimension of the fused feature map to obtain F′ SA , as shown in the following formula (3): Where DWConv stands for depth-wise separable convolution; Finally, F CA and F′ SA The enhanced feature fusion graph F is obtained by element-by-element multiplication o ∈R C×H×W , as shown in the following formula (4): in Indicates the dot product operation; F o : Output feature map; F CA : Features after channel attention processing; F′ SA : Features after spatial attention processing.
6. The method for detecting bacterial contamination of brewer's yeast according to claim 1, characterized in that: The focal loss is introduced into the classification loss of the detection head and is dynamically combined with the cross entropy loss through a weight factor.
7. The method for detecting bacterial contamination of brewer's yeast according to claim 6, characterized in that: The total loss function during model training is calculated as follows: L=u×CE+(1-u)×FL (7) Among them, L represents the total loss, CE represents the cross entropy loss of classification, FL represents the focal loss of classification, and u is the weight factor between the two losses, and its value range is between 0 and 1.
8. The method for detecting bacterial contamination of brewer's yeast according to claim 7, characterized in that: The calculation formula for cross entropy loss is as follows: Among them, CE represents the cross entropy loss of classification, x represents the sample, y represents the label, a represents the predicted output, and n represents the total number of samples.
9. The method for detecting bacterial contamination of brewer's yeast according to claim 7, characterized in that: The calculation formula of focal loss is as follows: FL=α(1-a) γ ×C (6) Among them, FL represents the focal loss of classification, α represents the sample weight factor, a represents the predicted output, γ represents the adjustment factor, and C represents the cross entropy loss term.
10. A computer device, characterized in that: The computer device includes a processor and a memory, the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to implement the brewer's yeast contamination detection method based on improved YOLOv10n image recognition as described in any one of claims 1 to 9.