A method for detecting parasitic eggs
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 邛崃市疾病预防控制中心(邛崃市卫生监督所)
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]本发明的目的在于提供一种寄生虫卵检测方法,其解决了现有技术中存在的罕见虫卵识别精度低、样本类别分布极度不均衡、模型在复杂显微背景下鲁棒性差以及计算资源消耗与检测性能难以兼顾的问题
[0028](1)本发明提供的寄生虫卵检测方法,通过引入深度度量学习中的监督对比损失,改变传统目标检测仅依赖全连接层进行线性分类的局限性,有效提升寄生虫卵检测的综合精度。
Smart Images

Figure CN122531007A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image or video processing technology, and in particular to a method for detecting parasite eggs. Background Technology
[0002] In existing technologies, some attempts are made to combine traditional machine learning with deep learning algorithms. While this can improve specificity under certain conditions, it is often difficult to effectively integrate the advantages of both when the outputs of the two algorithms are inconsistent. This can even lead to a decrease in model accuracy, and running multiple algorithms increases time costs. On the other hand, some technologies improve the detection ability of small targets by improving network structure, such as introducing attention mechanisms or lightweight convolutional layers. However, deep learning models are highly dependent on the quality of the dataset. Due to the extremely low infection rate of some parasites, image data of rare eggs is scarce, and the training dataset suffers from severe class imbalance, resulting in a generally low recognition rate of rare eggs by existing models. Summary of the Invention
[0003] The purpose of this invention is to provide a method for detecting parasite eggs, which solves the problems of low accuracy in identifying rare parasite eggs, extremely uneven distribution of sample categories, poor robustness of models under complex microscopic backgrounds, and difficulty in balancing computational resource consumption and detection performance in the existing technology.
[0004] The technical solution of the present invention:
[0005] This invention provides a method for detecting parasite eggs, comprising the following steps:
[0006] S1. Acquire microscopic image data of parasite eggs, perform target localization and category labeling, and construct a target detection dataset;
[0007] S2. Construct training, validation, and test sets based on the object detection dataset;
[0008] S3. Construct an improved YOLO26n model that supports deep metric learning;
[0009] S4. During the training phase of the improved YOLO26n model, extract the embedding feature vectors of positive samples, construct global class labels, and establish a one-to-one correspondence between the embedding vectors of positive samples and the true class.
[0010] S5. Construct a supervised contrastive loss function;
[0011] S6. Construct a total loss function by combining classification loss, regression loss, distribution focus loss, and supervised contrast loss;
[0012] S7. Train the improved YOLO26n model to obtain an optimized parasite egg detection model.
[0013] Furthermore, in S2, a class-balanced sampling strategy and composite data augmentation processing are performed during the training set construction process, wherein,
[0014] The aforementioned category-balanced sampling strategy specifically includes: counting the total number of samples of each category of insect eggs in the target detection dataset, calculating the sampling weight of each category relative to the largest sample category, and performing random sampling with replacement on the smaller sample categories according to the sampling weights when constructing training batches;
[0015] The composite data augmentation process includes Mosaic augmentation and Mixup augmentation.
[0016] Furthermore, in S3, the improved YOLO26n model employs depthwise separable convolutions in the backbone or neck network of the single-stage detection architecture, and adds a parallel embedding head on top of the classification head and regression head.
[0017] Furthermore, the structure of the embedding head sequentially includes: a first linear layer, a BatchNorm layer, a SiLU activation function layer, a second linear layer, and an L2 regularization layer.
[0018] Furthermore, in step S4, the specific process includes:
[0019] A target assignment strategy is used to match predicted bounding boxes with ground truth bounding boxes to generate a foreground mask.
[0020] Extract the positive sample embedding tensor based on the position index indicated by the foreground mask;
[0021] Calculate the global index of each positive sample within the entire batch using the target allocation index and the batch image index vector;
[0022] The corresponding real category label is extracted from the global category label list based on the global index.
[0023] Furthermore, in S5, the sample distribution is constrained in the metric space by calculating the cosine similarity between samples of the same class and samples of different classes within a batch.
[0024] Furthermore, in S5, the supervised contrast loss function introduces a temperature parameter to adjust the smoothness of the Softmax probability distribution.
[0025] Furthermore, in step S6, the total loss function performs a weighted summation of the classification loss, regression loss, distribution focus loss, and supervised comparison loss using preset loss weight hyperparameters.
[0026] Furthermore, in step S7, during the training of the improved YOLO26n model, the hyperparameters are adjusted. If the performance of the improved YOLO26n model does not improve for a specified number of consecutive rounds, then the training is stopped.
[0027] Based on the above technical features, the beneficial effects of the present invention are as follows:
[0028] (1) The parasite egg detection method provided by the present invention introduces supervised contrast loss in deep metric learning, which changes the limitation of traditional target detection that relies solely on fully connected layers for linear classification, and effectively improves the overall accuracy of parasite egg detection.
[0029] (2) The parasite egg detection method provided by the present invention combines class-balanced sampling at the data end with supervised contrast loss at the algorithm end, so that the model no longer blindly favors the common class with a large sample size. The supervised contrast loss constrains the samples in the embedding space by clustering, so that rare parasite eggs can form independent feature clusters even when the sample size is small, thereby ensuring the high sensitivity of the model to rare pathogens. It effectively solves the problem caused by the imbalance of parasite egg sample categories and also solves the problem of high false negative rate in clinical detection.
[0030] (3) The parasite egg detection method provided by the present invention is optimized by depth-separable convolution, so that the high-performance automatic parasite egg detection algorithm can run smoothly on resource-limited terminal devices such as smartphones and portable microscopes. Attached Figure Description
[0031] Figure 1 A flowchart illustrating the method of this invention;
[0032] Figure 2 A schematic diagram illustrating the principle of the method of this invention;
[0033] Figure 3 A schematic diagram of the depthwise separable convolutional structure of the model in the method of this invention is provided.
[0034] Figure 4 This is a schematic diagram of the embedded head structure of the model in the method of the present invention. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0036] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0037] It is important to note that parasitic infections are a significant challenge in global public health. Medical worms and protozoa, after infecting the human body, continuously consume nutrients, cause mechanical damage or space-occupying lesions, leading to serious consequences such as malnutrition, anemia, stunted growth and development, and organ damage in the host. Currently, clinical diagnosis of parasitic diseases mainly relies on microscopic examination of fecal samples, manually observing the morphology and structure of the eggs to determine the species and assess the degree of infection. However, manual microscopy is highly dependent on the professional skills of the laboratory personnel and the time available for interpreting slides. With the overall decline in parasitic infection rates, primary healthcare institutions are gradually lacking in microscopic examination experience, while imported cases due to population movement are increasing. This leads to difficulties in the practical application of manual microscopy, such as low efficiency and a high risk of missed diagnoses.
[0038] Example 1
[0039] Please refer to the above as well. Figures 1-4 The present invention provides a method for detecting parasite eggs, comprising the following steps:
[0040] S1. Acquire microscopic image data of parasite eggs, perform target localization and category labeling, and construct a target detection dataset;
[0041] S2. Construct training, validation, and test sets based on the object detection dataset;
[0042] S3. Construct an improved YOLO26n model that supports deep metric learning;
[0043] S4. During the training phase of the improved YOLO26n model, extract the embedding feature vectors of positive samples, construct global class labels, and establish a one-to-one correspondence between the embedding vectors of positive samples and the true class.
[0044] S5. Construct a supervised contrastive loss function;
[0045] S6. Construct a total loss function by combining classification loss, regression loss, distribution focus loss, and supervised contrast loss;
[0046] S7. Train the improved YOLO26n model to obtain an optimized parasite egg detection model.
[0047] It should be noted that the parasite egg detection method provided by this invention combines deep metric learning with a target detection model. The model can learn similarity metrics from the data, thereby achieving intra-class compactness and inter-class separation, which enhances the model's understanding of the similarity relationship between samples, making the model's feature generalization ability stronger and its robustness to changes in data distribution stronger. This effectively alleviates the identification accuracy problem caused by the imbalance of parasite egg categories, while also solving the problems of limited application scenarios and the limited number of identification types in existing detection technologies.
[0048] It should be noted that this invention combines the YOLO26n model and deep metric learning algorithm, and adds supervised contrastive loss SupConLoss to train the model, which enhances the model's ability to discriminate features. Compared with ordinary object detection models, it can significantly improve the detection accuracy of the model, and also has good detection results for rare insect eggs.
[0049] Furthermore, the implementation steps of this method are explained in detail:
[0050] In step S1, microscopic image data of parasite eggs are acquired, target localization and category labeling are performed, and a target detection dataset is constructed. In detail, a wide range of parasite egg images are collected, including photos of eggs taken under a microscope, images downloaded from the Internet, and images taken from books. These images include at least soil-transmitted nematode eggs, foodborne trematode eggs, schistosome eggs, tapeworm eggs, and other targets with different morphological characteristics. Then, a target detection dataset based on the microscopic images of parasite eggs is constructed. Finally, the collected parasite egg images are labeled using labeling software, such as LabelImg software.
[0051] In step S2, a training set, a validation set, and a test set are constructed based on the object detection dataset. The object detection dataset is divided into the training set, validation set, and test set in a 7:2:1 ratio. In detail, a class balance sampling strategy and composite data augmentation processing are performed during the construction of the training set. The class balance sampling strategy specifically includes: counting the total number of samples of each class of insect eggs in the object detection dataset, calculating the sampling weight of each class relative to the largest sample class, and performing random sampling with replacement on the small sample class according to the sampling weight when constructing the training batch. The composite data augmentation processing includes Mosaic augmentation and Mixup augmentation.
[0052] In step S3, an improved YOLO26n model supporting deep metric learning is constructed. Specifically, the improved YOLO26n model employs depthwise separable convolutions in the backbone or neck network of a single-stage detection architecture, and adds a parallel embedding head on top of the classification and regression heads. That is, the head network includes a classification head, a regression head, and an embedding head, as shown below. Figure 4As shown, the embedding head structure includes, in sequence: a first linear layer, a BatchNorm layer, a SiLU activation function layer, a second linear layer, and an L2 regularization layer. The improved YOLO26n model simultaneously outputs the embedded features. The results are for classification and regression, where D is the embedding dimension. It should be noted that an embedding header is added to the output header of the YOLO26n model, such as... Figure 3 As shown, depthwise separable convolution is used to replace ordinary convolution, thereby reducing the number of parameters.
[0053] In some possible embodiments, during the inference phase of the improved YOLO26n model, the embedding head can be removed or frozen, leaving only the optimized classification head and regression head to perform real-time detection. Since the backbone network has already learned more discriminative feature representations constrained by the metric space during training, the classification head faces feature inputs with more pronounced inter-class differences and stronger intra-class consistency during actual inference. This allows the improved YOLO26n model to effectively enhance its ability to capture rare eggs without increasing inference time, and effectively reduce the misdiagnosis rate caused by complex backgrounds (such as fecal residue, bubbles, fibers, etc.).
[0054] In some possible implementations, the first linear layer maps the feature vector output by the neck network to an intermediate dimension; the batch normalization layer is used to accelerate convergence and prevent gradient vanishing; the SiLU activation function introduces a nonlinear transformation to enhance feature representation; the second linear layer projects the features to the final embedding dimension D; and the L2 regularization layer maps the output embedding vector to a unit hypersphere, so that in subsequent metric calculations, the dot product of two embedding vectors is equivalent to cosine similarity, thereby providing a standardized input representation for deep metric learning.
[0055] In detail, the embedding head module receives feature map input from the neck network, sequentially passes it through a first linear layer for dimensionality transformation, and then enters a batch normalization layer to perform data normalization processing, thereby stabilizing gradient flow during training. Next, a SiLU activation function layer is used to introduce non-linear expressive power, and then a second linear layer projects the features onto a preset embedding space dimension. Finally, an L2 regularization layer maps the output feature vector onto a unit hypersphere. This process ensures that the output embedding vectors have consistent length, allowing for cosine similarity measurement through simple dot product operations in subsequent calculations, providing high-quality feature input for deep metric learning.
[0056] In step S4, during the training phase of the improved YOLO26n model, the embedding feature vectors of positive samples are extracted using the embedding head, and global category labels are constructed based on the matching relationship between positive samples and the true target, establishing a one-to-one correspondence between the positive sample embedding vectors and the true categories; the specific process includes:
[0057] A target assignment strategy is used to match predicted bounding boxes with ground truth bounding boxes, generating a foreground mask. Specifically, foreground embedding features and corresponding labels are constructed. The positions of the embedding features of all positive samples in the current batch are obtained through the foreground mask. ,in This indicates that the a-th anchor box in the b-th image is a positive sample, B is the batch size, and A is the number of anchor boxes.
[0058] The positive sample embedding tensor is extracted based on the position index indicated by the foreground mask. Specifically, then, from the embedding features... Extract the embedding features corresponding to all positive samples to obtain the positive sample embedding tensor;
[0059] The global index of each positive sample within the entire batch is calculated using the target allocation index and the batch image index vector. Specifically, the true target index matched by the anchor box of each positive sample is obtained. This is achieved by using the target allocation index and the batch image index vector to obtain the global index of the target within the entire batch. The target allocation index... , This represents the local index of the real target matched by the anchor box [b, a] within its respective image; where, the batch image index vector... , This represents the total number of targets in the batch, with each element representing the image index to which that target belongs;
[0060] Extract the corresponding real category label from the global category label list based on the global index, where the global category label... Arranged by image.
[0061] In some possible embodiments, step S4 further includes:
[0062] First, count the number of real targets in each image. ;
[0063] Then, the starting offset of each image in the global target list is calculated. and the set of positive sample locations ;
[0064] Next, for each positive sample Extraction: Embedded Vector Local target index Global target index Corresponding category tags ;
[0065] Finally, the positive sample embedding tensor is calculated. and the corresponding tags .
[0066] In step S5, a supervised contrastive loss function is constructed. By calculating the cosine similarity between samples of the same class and samples of different classes within a batch, constraints are applied to the sample distribution in the metric space. The supervised contrastive loss function introduces a temperature parameter to adjust the smoothness of the Softmax probability distribution. Specifically, a supervised contrastive loss function, SupConLoss, for deep metric learning is constructed for the YOLO26 model.
[0067]
[0068] Where N is the batch size, e is the normalized embedding vector after feature extraction, P(i) is the set of positive sample indices of anchor point i, A(i) is the set of all sample indices except anchor point i, and τ is the temperature hyperparameter.
[0069] It should be noted that in step S5, a YOLO26n model supporting deep metric learning is constructed, the detection head is modified, and features such as... Figure 4 The embedding layer shown generates an embedding feature tensor. This embedded feature tensor is normalized and returned along with the detection results. Further, the loss function of the YOLO26n model is modified by adding SupConLoss loss. SupConLoss uses cosine similarity to measure the similarity between two embedding vectors, and the calculation formula is as follows:
[0070]
[0071] Then, the loss function calculates cosine similarity, with the ultimate goal of adjusting the orientation of vectors in high-dimensional space. For samples of the same class, the loss function forces their vector directions to tend towards the same direction (0° angle). For samples of different classes, the loss function forces their vector directions to tend towards orthogonality or opposite directions (as large as possible angle).
[0072] It should be noted that the supervised contrastive loss function increases the dot product between embedding vectors of similar samples while decreasing the dot product between dissimilar samples. This forces the improved YOLO26n model to cluster features of eggs belonging to the same parasite species together in the feature space and pushes features of eggs of different species as far apart as possible. As a result, the improved YOLO26n model can accurately distinguish between eggs with extremely similar morphologies (such as hookworm eggs and certain types of straight nematodes) by learning subtle discriminative features.
[0073] It should be noted that the loss function measures cosine similarity by calculating the dot product of the anchor vector and the positive sample vector, and combines the temperature hyperparameter to adjust the contrast intensity of the probability distribution. By minimizing the supervised contrast loss, the improved YOLO26n model clusters the features of the same type of parasite eggs (such as different morphologies of Ascaris eggs) within a very small angle range in the embedding space, while pushing the features of different types of eggs in orthogonal or opposite directions. This allows the improved YOLO26n model to learn key discriminative information such as the texture of the egg surface, the thickness of the egg shell, and the structure of the egg contents, rather than relying solely on the outline of the target. Even when faced with interference such as fecal residue and bubbles, it still achieves accurate feature recognition.
[0074] In step S6, a total loss function is constructed by combining the classification loss, regression loss, distribution focus loss, and supervised comparison loss. The total loss function performs a weighted summation of the classification loss, regression loss, distribution focus loss, and supervised comparison loss using preset loss weight hyperparameters. Specifically, the total loss function is:
[0075]
[0076] in, It is a classification loss, used to determine whether the target exists and its category; It is the bounding box regression loss, used to optimize the coordinate accuracy of the predicted box; α is the distributed focus loss, used to optimize the localization accuracy of bounding boxes in complex backgrounds; α, β, γ, and δ are hyperparameters, which are used to control the importance of classification loss, regression loss, distributed focus loss, and supervised contrast loss. Specifically, the weights of each loss term are assigned through preset hyperparameters to ensure that the improved YOLO26n model optimizes detection performance while taking into account the representation quality of the metric space.
[0077] In step S7, the improved YOLO26n model is trained to obtain an optimized parasite egg detection model. During the training of the improved YOLO26n model, hyperparameters are adjusted. If the performance of the improved YOLO26n model does not improve for a specified number of consecutive rounds, training is stopped. The implementation steps are detailed using an experimental example:
[0078] First, a data loader is built to randomly extract images by category, ensuring that each batch contains images of each category.
[0079] Next, the improved YOLO26n model is trained using the training set constructed in step S2, and data augmentation and sample expansion are performed using color enhancement, geometric transformation Mosaic, and Mixup techniques. Mosaic enhancement is used to randomly scale, crop, and arrange four randomly selected training images and then stitch them together to form a synthetic image containing complex backgrounds and multiple targets, thereby significantly improving the improved YOLO26n model's ability to capture targets at the edge of the microscope's field of view. Mixup enhancement is used to linearly fuse two images and their corresponding labels proportionally, thereby increasing the continuity of the improved YOLO26n model in the feature space and preventing the improved YOLO26n model from experiencing a decline in generalization performance due to overfitting during training.
[0080] Then, the model's performance was evaluated during training using the validation set constructed in step S2. All experiments were conducted on the same computer: an Intel(R) Xeon(R) Platinum 8474C CPU and an RTX 4090D GPU with 24GB of VRAM; the training environment was PyTorch 2.8.0 + cu128, and Python version 3.9.5. During the experiments, the epochs were set to 300, the batch size to 32, and the MuSGD optimizer with a cosine annealing learning rate strategy was used, with an initial learning rate of 0.01. Pre-trained weights were used during training; if the model's performance did not improve for 30 consecutive epochs, training was stopped to save time.
[0081] Finally, an optimized model for detecting parasite eggs was obtained.
[0082] It should be noted that during the training of the improved YOLO26n model, supervised contrastive loss is used. The generated gradients not only act on the embedding head, but are also propagated back to the neck network and the backbone network through the backpropagation algorithm. Furthermore, when extracting features, the backbone network not only considers the target's localization and classification information, but also extracts highly discriminative metric features. In this process, by dynamically adjusting the weights, the proportion of metric learning can be appropriately increased in the early stage of training to quickly widen the gap between classes, while in the later stage of training, the focus is on optimizing the localization accuracy, thereby achieving the optimal balance of detection model performance.
[0083] In some possible embodiments, the method provided by the present invention further includes model evaluation, with the evaluation metrics being: mean Average Precision (mAP), number of parameters, and GigaFloating Point Operations (GFLOPs) used to evaluate experimental results. mAP reflects the overall performance of the algorithm, and it is calculated by integrating the resulting PR curve with recall (R) as the x-axis and precision (P) as the y-axis. Recall represents the percentage of correctly predicted positive samples in the total number of samples, and precision represents the percentage of correctly predicted positive samples in all detected positive samples. Wherein:
[0084] The formula for calculating recall rate is:
[0085]
[0086] The formula for calculating accuracy is:
[0087]
[0088] The formula for calculating average accuracy is:
[0089]
[0090] The formula for calculating the mean precision is:
[0091]
[0092] in, Indicates the order after sorting by confidence level. The number of correctly predicted boxes in each prediction box. Indicates the preceding The number of incorrectly predicted boxes in a prediction box (including IoU less than the threshold, class errors, and duplicate detections). Indicate category Total number of real frames, It is a category of .
[0093] It should be noted that this method uses the YOLO26n model as a foundation and combines supervised contrastive loss with deep metric learning for training. Deep metric learning combines deep learning with metric learning to directly learn similarity metrics from the data. The model improves its detection capability by understanding the similarity relationships between samples, especially for rare insect eggs with limited training data. Comparative experiments yielded some results, as shown in the table below:
[0094] Model Parameter quantity / M GFLOPs mAP@0.5 / % mAP@0.5:0.95 / % YOLOv8n 3.0 8.2 86.1 77.9 YOLO26n 2.5 5.8 87.6 79.9 YOLO26n-SupConLoss 2.6 6.2 89.6 81.5
[0095] As shown in the table above, the improved YOLO26n model has fewer parameters than the more common YOLOv8n model, but higher mAP@0.5 and mAP@0.5:0.95. On the insect egg dataset, the YOLO26n model provided by this method exhibits stronger detection capabilities. The improved YOLO26n model with added metric learning shows a 2% improvement in mAP@0.5 and a 1.6% improvement in mAP@0.5:0.95 compared to the ordinary YOLO26n model. Therefore, the improved YOLO26n model provided by this method achieves further improvements in detection performance, while the increase in parameters and GFLOPs is not significant.
[0096] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the scope of the invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A method for detecting parasite eggs, characterized in that, Includes the following steps: S1. Acquire microscopic image data of parasite eggs, perform target localization and category labeling, and construct a target detection dataset; S2. Construct training, validation, and test sets based on the object detection dataset; S3. Construct an improved YOLO26n model that supports deep metric learning; S4. During the training phase of the improved YOLO26n model, extract the embedding feature vectors of positive samples, construct global class labels, and establish a one-to-one correspondence between the embedding vectors of positive samples and the true class. S5. Construct a supervised contrastive loss function; S6. Construct a total loss function by combining classification loss, regression loss, distribution focus loss, and supervised contrast loss; S7. Train the improved YOLO26n model to obtain an optimized parasite egg detection model.
2. The detection method according to claim 1, characterized in that, In step S2, a class balance sampling strategy and composite data augmentation processing are executed during the training set construction process. The aforementioned category-balanced sampling strategy specifically includes: counting the total number of samples of each category of insect eggs in the target detection dataset, calculating the sampling weight of each category relative to the largest sample category, and performing random sampling with replacement on the smaller sample categories according to the sampling weights when constructing training batches; The composite data augmentation process includes Mosaic augmentation and Mixup augmentation.
3. The detection method according to claim 1, characterized in that, In S3, the improved YOLO26n model employs depthwise separable convolutions in the backbone or neck network of the single-stage detection architecture, and adds a parallel embedding head on top of the classification head and regression head.
4. The detection method according to claim 3, characterized in that, The structure of the embedding head includes, in sequence: a first linear layer, a BatchNorm layer, a SiLU activation function layer, a second linear layer, and an L2 regularization layer.
5. The detection method according to claim 1, characterized in that, In S4, the specific process includes: A target assignment strategy is used to match predicted bounding boxes with ground truth bounding boxes to generate a foreground mask. Extract the positive sample embedding tensor based on the position index indicated by the foreground mask; Calculate the global index of each positive sample within the entire batch using the target allocation index and the batch image index vector; The corresponding real category label is extracted from the global category label list based on the global index.
6. The detection method according to claim 1, characterized in that, In S5, the sample distribution is constrained in the metric space by calculating the cosine similarity between samples of the same class and samples of different classes within a batch.
7. The detection method according to claim 1, characterized in that, In step S5, the supervised contrast loss function introduces a temperature parameter to adjust the smoothness of the Softmax probability distribution.
8. The detection method according to claim 1, characterized in that, In step S6, the total loss function performs a weighted summation of the classification loss, regression loss, distribution focus loss, and supervised comparison loss using preset loss weight hyperparameters.
9. The detection method according to claim 1, characterized in that, In step S7, during the training of the improved YOLO26n model, the hyperparameters are adjusted. If the performance of the improved YOLO26n model does not improve for a specified number of consecutive rounds, the training is stopped.