Improved YOLOv5s model construction method, small target detection method and system

By improving the YOLOv5s model, the Focal-SIOU loss function, multi-head self-attention module, shuffle attention module and cross-modal image segmentation module were introduced, which solved the problem of lymphocyte detection difficulty in pathological diagnosis of Sjogren's syndrome, and achieved more efficient and accurate diagnostic results.

CN117475434BActive Publication Date: 2025-06-06YANTAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311651968.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2025-06-06
Estimated Expiration
2043-12-05

AI Technical Summary

Technical Problem

During the pathological diagnosis of Sjogren's syndrome, it is difficult to accurately detect lymphocytes, resulting in frequent misdiagnosis and misdiagnosis, and it is difficult for the existing technology to achieve efficient and accurate diagnosis.

Method used

Improve the construction method of the YOLOv5s model, and optimize the model to improve the accuracy of lymphocyte detection by introducing Focal-SIOU loss function, multi-head self-attention module, shuffle attention module and cross-modal image segmentation module.

Benefits of technology

It improves the accuracy of lymphocyte detection tasks and can more accurately assist the pathological diagnosis of lip biopsy of Sjogren's syndrome, reducing the occurrence of missed diagnosis and misdiagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117475434B_ABST
    Figure CN117475434B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image processing technology, and specifically discloses a construction method of an improved YOLOv5s model, a small target detection method and a system, wherein the small target detection method comprises the following steps: obtaining an image of a region to be detected and performing preprocessing; obtaining a friction force map of the region to be detected and performing preprocessing; improving the YOLOv5s model to obtain an optimized YOLOv5s model; inputting the preprocessed image into the optimized YOLOv5s model, detecting small targets in the image, and counting the number of small targets in a single region; comparing the number of small targets in a single region with a set threshold, and when the threshold is exceeded, screening out the region and marking it. The present technical solution is adopted to improve the accuracy of small target detection tasks by improving the YOLOv5s model, and is used to assist in the pathological diagnosis of labial gland biopsy of Sjögren's syndrome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing and relates to a construction method of an improved YOLOv5s model, a small target detection method and a system. Background Art

[0002] Sjögren's syndrome (SS) is a chronic inflammatory autoimmune disease characterized by lymphocyte proliferation and progressive damage to exocrine glands. In addition to primarily affecting the salivary and lacrimal glands, it can also affect multiple organ systems such as the lungs, kidneys, skin, and blood, and is often associated with other systemic immune diseases such as rheumatoid arthritis (RA) and systemic lupus erythematosus (SLE).

[0003] At present, the cause of Sjögren's syndrome is still unclear, and may involve genetics, viral infection, sex hormone levels and other factors. It is worth noting that Sjögren's syndrome is not rare. However, many patients have limited knowledge of Sjögren's syndrome, which leads to delayed medical treatment and delayed disease. If early diagnosis and systematic treatment are carried out to improve prognosis, most of the disease can be well controlled.

[0004] In the diagnosis process of Sjögren's syndrome, pathologists need to review each pathological section one by one under different magnification microscopes. This process is not only lengthy and time-consuming, but also due to the subjective heterogeneity of pathologists at different levels, missed diagnoses and misdiagnoses often occur. Therefore, accurate and efficient pathological diagnosis has become a huge challenge for pathologists. In the era of big data, artificial intelligence has been widely used in medical image-assisted diagnosis. With the rapid development of digital pathology technology, artificial intelligence-assisted pathological diagnosis technology has gradually emerged. At present, in various tumors such as lung cancer and breast cancer, AI-assisted pathological diagnosis is not only efficient, stable and highly repeatable, but its level is comparable to that of professional physicians. However, research on the pathological diagnosis of labial gland biopsy in Sjögren's syndrome has not been reported. Summary of the invention

[0005] The purpose of the present invention is to provide a method for constructing an improved YOLOv5s model, a small target detection method and system, so as to improve the accuracy of small target detection tasks and assist in the pathological diagnosis of labial gland biopsy in Sjögren's syndrome.

[0006] In order to achieve the above object, the basic scheme of the present invention is: a method for constructing an improved YOLOv5s model, comprising the following steps:

[0007] Use Focal-SIOU loss function to replace CIOU loss function of YOLOv5s model;

[0008] Introduce the multi-head self-attention module into the skeleton network part of the YOLOv5s model;

[0009] In the neck part of the YOLOv5s model, the shuffle attention module is introduced;

[0010] A cross-modal image segmentation module is introduced after the shuffle attention module, wherein the cross-modal image segmentation module includes an image feature extraction module, a friction feature extraction module, and a feature fusion module;

[0011] Remove the large object detection head of the YOLOv5s model to obtain the optimized YOLOv5s model.

[0012] The working principle and beneficial effects of this basic scheme are as follows: Since lymphocytes are small and difficult to distinguish, this technical scheme makes four improvements to YOLOv5 to improve the accuracy of lymphocyte detection tasks, so as to achieve the purpose of accurately detecting lymphocytes and assisting pathological diagnosis.

[0013] Furthermore, the method of replacing the CIOU loss function of the YOLOv5s model with the Focal-SIOU loss function is as follows:

[0014] Assume that α is the angle between the predicted box and the real box coordinate center, let C h , C w Respectively represent the horizontal distance and vertical distance between the coordinate centers of the predicted box and the real box, then the straight-line distance σ and the angle α between the coordinate centers of the two are:

[0015]

[0016]

[0017] According to the angle α, the angle loss Λ is defined as:

[0018]

[0019] The distance loss Δ of SIOU is:

[0020]

[0021]

[0022]

[0023] γ=2-Λ

[0024] in, Represent the center point coordinates of the real box and the predicted box respectively; ρ x , ρ y It represents the distance loss factor between the real box and the predicted box in both width and height directions; γ is the angle loss factor.

[0025] The angle loss is integrated into the distance loss function. When the angle α is closer to 45 degrees, the contribution of the angle loss is greater. When the angle α is closer to 0, the contribution of the angle loss is smaller and degenerates into the distance loss. At this time, C w With C h It represents the maximum distance between the predicted box and the true box, not the distance between their center points.

[0026] The shape loss Ω of SIOU is:

[0027]

[0028]

[0029]

[0030] Among them, θ represents the attention to shape loss, and its value needs to be adjusted according to the specific data set; w, h, w gt ,h gt Represents the width and height of the predicted box and the real box respectively; ω w ,ω h Represents the shape loss factor in the width direction and the height direction.

[0031] After integrating the three indicator losses, SIOU's regression box loss function L SIOU for:

[0032]

[0033] Among them, IOU is the intersection-union ratio between the real box and the predicted box;

[0034] SIOU takes the angle loss into account and optimizes the model performance.

[0035] When regressing the bounding box of a predicted object, the process suffers from the problem of imbalanced training samples. In an image, there are fewer high-quality anchor boxes with small regression errors than low-quality anchor boxes with large errors. Poor-quality anchor boxes produce excessively large gradients, which negatively affects the training process. To address this issue, Focal-loss is integrated with SIOU to distinguish between high-quality and low-quality anchor boxes. This helps improve the accuracy of regression. The Focal-SIOU loss function is:

[0036] L Focal-SIOU =IOUγL SIOU

[0037] Among them, γ represents the attention to IOU, and its value range is greater than 0. The larger the value of γ is, the more attention the loss function pays to IOU; the closer the value is to 0, the less attention the loss function pays to IOU and gradually degenerates to L SIOU IOU represents the intersection-over-union ratio between the real box and the predicted box; L SIOU Represents the loss function of SIOU.

[0038] Furthermore, the method of introducing the multi-head self-attention module into the backbone module of the YOLOv5s model is:

[0039] The C3 module of the original YOL0v5s network is restructured and integrated with a multi-head self-attention layer;

[0040] The multi-head self-attention layer adds position encoding to make it position sensitive;

[0041] Define the number of heads in the multi-head self-attention layer. The input is first generated by point convolution to generate query vector q, key vector k and value vector v, while R h , R w Each represents a positional code extracted from the height and width;

[0042] After the position encoding performs the corresponding element addition operation, the position vector r is generated. r and q are matrix multiplied to generate the corresponding content-position vector qr T , and q and k perform matrix multiplication to generate the corresponding content-content vector qk T ;

[0043] qr T With qk T The corresponding elements are added, and after passing through the softmax layer, matrix multiplication is performed with v to finally obtain the output feature z.

[0044] The multi-head self-attention module is easy to build, the number of parameters is slightly reduced, and the detection accuracy is significantly improved.

[0045] Furthermore, the cross-modal image segmentation module obtains a color image of the area to be detected and a force map formed by the friction force of the area to be detected, wherein the force map uses grayscale values ​​to represent the magnitude of the friction force at different positions of the area to be detected, and adjusts the two images to the same resolution, and has Sjögren's disease segmentation label information in each image, and divides the image data pairs into a training set, a validation set, and a test set;

[0046] The image feature extraction module extracts features from the color image of the area to be detected to obtain single-modal color image features;

[0047] The friction force feature extraction module extracts features from the force map of the detection area to obtain single-mode friction force features;

[0048] The feature fusion module includes a first gating module (Relu function), a second gating module and a fusion network. The first gating module obtains color image features and processes the output part greater than the image threshold. The second gating module obtains friction features and processes the output part greater than the friction threshold. The fusion network fuses the features output by the first gating module and the second gating module, and uses the friction features as a channel of the output image.

[0049] Furthermore, in the bottleneck part of the YOLOv5s model, the method of introducing the shuffle attention module is as follows:

[0050] The input feature map dimension is c*h*w. The input is split into g groups along the channel dimension c, and the dimension of each group is c / g*h*w.

[0051] Each group will be split into two branches along the channel dimension again, and the dimension of each branch becomes c / 2g*h*w;

[0052] The two branches generate their own feature maps through the spatial attention module and the channel attention module respectively, helping the model to focus on the detection target and its location information;

[0053] After extracting information, the two feature maps are concatenated and the dimension becomes c / g*h*w. After all g groups have extracted features, they are concatenated again to get the output. The dimension of the output is still c*h*w, which remains the same as the dimension of the input.

[0054] The output is reordered through the channel reorganization function to ensure the information flow between different groups;

[0055] The channel attention mechanism in the ShuffleAttention module first averages the input to obtain a set of channel-related statistics. After this set of statistics is linearly transformed and activated by the sigmoid function, it is multiplied with the corresponding elements of the original input to obtain the output result of the collected position information.

[0056] The spatial attention mechanism used in the Shuffle attention module first group-normalizes the input to obtain spatially related statistics. This set of statistics is linearly transformed and passed through the sigmoid activation function, and then multiplied with the corresponding elements of the original input to obtain the output result of the collected target information.

[0057] Shuffle attention reduces the number of parameters and computational cost of the attention mechanism while incorporating feature information from both channel and spatial dimensions, improving the detection accuracy of the detector.

[0058] The method to further remove the large object detection head of the YOLOv5s model is:

[0059] Since the sample targets are all small and medium-sized targets, remove After the large target detection head with a magnification of 1, the network contains The detection heads with two sampling rates correspond to small target detection and medium target detection respectively.

[0060] Remove large target detection heads to improve detector accuracy and reduce the number of parameters and calculations.

[0061] The present invention also provides a small target detection method based on an improved YOLOv5s algorithm, comprising the following steps:

[0062] Acquire the image of the area to be detected and perform preprocessing;

[0063] Obtain the friction force map of the area to be detected and perform preprocessing;

[0064] Based on the construction method of the present invention, the YOLOv5s model is improved to obtain an optimized YOLOv5s model;

[0065] Input the preprocessed image into the optimized YOLOv5s model, detect small objects in the image, and count the number of small objects in a single area;

[0066] The number of small targets in a single area is compared with the set threshold. When the threshold is exceeded, the area is screened out and marked.

[0067] This method uses the improved YOLOv5s model to detect small targets to determine whether the patient has Sjögren's syndrome and achieve auxiliary diagnosis.

[0068] The present invention also provides a small target detection system based on an improved YOLOv5s algorithm, comprising an image acquisition module, a friction force map acquisition module and a processing module, wherein the image acquisition module is used to acquire an image to be detected, the friction force map acquisition module is used to acquire a friction force map of an area to be detected, the output ends of the image acquisition module and the friction force map acquisition module are respectively connected to the input end of the processing module, and the processing module executes the small target detection method of the present invention, detects small targets in an image, determines whether the number of small targets in a single area exceeds a threshold, and diagnoses whether a patient suffers from Sjögren's syndrome.

[0069] Using this system, by collecting and analyzing images, small target detection results of the images are obtained to diagnose whether the patient has Sjögren's syndrome.

[0070] Furthermore, the friction force diagram acquisition module is a friction tester.

[0071] The equipment is easily accessible and easy to use. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 It is a structural schematic diagram of the method for constructing an improved YOLOv5s model of the present invention;

[0073] Figure 2 It is a schematic diagram of SIOU angle loss calculation of the method for building the improved YOLOv5s model of the present invention;

[0074] Figure 3 It is a detailed schematic diagram of the self-attention layer of the multi-head self-attention module of the method for constructing the improved YOLOv5s model of the present invention;

[0075] Figure 4 It is a structural schematic diagram of the Shuffle attention module of the method for building an improved YOLOv5s model in the present invention. DETAILED DESCRIPTION

[0076] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0077] In the description of the present invention, it is necessary to understand that the terms "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0078] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal connection between two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.

[0079] The present invention discloses a method for constructing an improved YOLOv5s (YOLOv5 is a single-stage target detection algorithm) model for detecting lymphocyte-infiltrated lesions in pathological images and assisting pathological diagnosis. Since lymphocytes are small and difficult to distinguish, the present invention makes four improvements to YOLOv5 to improve the accuracy of lymphocyte detection tasks. Figure 1 As shown, the construction method includes the following steps:

[0080] Use Focal-SIOU loss function to replace CIOU loss function of YOLOv5s model, speed up network convergence and improve model accuracy;

[0081] The Multi-Head Self-Attention Module (MHSA) is introduced into the backbone of the YOLOv5s model to help the network capture more long-term dependencies and cope with the challenges of complex backgrounds.

[0082] In the neck part of the YOLOv5s model, the shuffle attention (SA) attention module is introduced to enhance the model's ability to fuse spatial and channel dimension features;

[0083] A cross-modal image segmentation module is introduced after the shuffle attention module, wherein the cross-modal image segmentation module includes an image feature extraction module, a friction feature extraction module, and a feature fusion module;

[0084] The large object detection head of the YOLOv5s model is removed to obtain the optimized YOLOv5s model. Removing the corresponding detection head improves accuracy while reducing the number of parameters and model complexity.

[0085] In a preferred embodiment of the present invention, the method of replacing the CIOU loss function of the YOLOv5s model with the Focal-SIOU loss function is as follows:

[0086] In computer vision tasks, the efficiency of object detection is highly dependent on the definition of the loss function. Conventional object detection loss functions focus on several indicators of bounding box regression, including distance, overlapping area, and aspect ratio. However, conventional IOU strategies do not take into account the directional information of the true box and the predicted box. This causes the predicted box to wander around during training, resulting in slower model training and poorer fitting, which ultimately affects the detection performance of the model. SIOU takes angle loss into account and solves the above problems. SIOU's loss function mainly consists of four parts: angle loss, distance loss, shape loss, and IOU loss.

[0087] like Figure 2 As shown, assuming that α is the angle between the predicted box and the real box coordinate center (less than or equal to 45 degrees), let C h , C w Respectively represent the horizontal distance and vertical distance between the coordinate centers of the predicted box and the real box, then the straight-line distance σ and the angle α between the coordinate centers of the two are:

[0088]

[0089]

[0090] According to the angle α, the angle loss Λ is defined as:

[0091]

[0092] The distance loss Δ of SIOU is:

[0093]

[0094]

[0095]

[0096] γ=2-Λ

[0097] in, Represent the center point coordinates of the real box and the predicted box respectively; ρ x , ρ y It represents the distance loss factor between the real box and the predicted box in both width and height directions; γ is the angle loss factor.

[0098] The angle loss is integrated into the distance loss function. When the angle α is closer to 45 degrees, the contribution of the angle loss is greater. When the angle α is closer to 0, the contribution of the angle loss is smaller and degenerates into the distance loss. At this time, C w With C h It represents the maximum distance between the predicted box and the true box, not the distance between their center points.

[0099] The shape loss Ω of SIOU is:

[0100]

[0101]

[0102]

[0103] Wherein, θ represents the attention paid to shape loss. To achieve more balanced training, its value needs to be adjusted according to the specific data set. The value range is [2, 6]. In this embodiment, its value is set to 4. The smaller the value, the higher the attention paid to shape loss, so that the model is more inclined to adjust the shape of the prediction box during training, thereby suppressing the feedback of other losses on training. w, h, w gt ,h gt Represents the width and height of the predicted box and the real box respectively; ω w ,ω h Represents the shape loss variable in the width and height directions.

[0104] After integrating the three indicator losses, SIOU's regression box loss function L SIOU for:

[0105]

[0106] Among them, IOU is the intersection-union ratio between the real box and the predicted box.

[0107] When regressing the bounding box of a predicted object, the process suffers from the problem of imbalanced training samples. In an image, there are fewer high-quality anchor boxes with small regression errors than low-quality anchor boxes with large errors. Poor-quality anchor boxes produce excessively large gradients, which negatively affects the training process. To address this issue, Focal-loss is integrated with SIOU to distinguish between high-quality and low-quality anchor boxes. This helps improve the accuracy of regression. The Focal-SIOU loss function is shown below:

[0108] L Focal-SIOU =IOU γ L SIOU

[0109] Among them, γ represents the attention to IOU, and its value range is greater than 0. The larger the value of γ is, the more attention the loss function pays to IOU; the closer the value is to 0, the less attention the loss function pays to IOU and gradually degenerates to L SIOU IOU represents the intersection-over-union ratio between the real box and the predicted box; L SIOU Represents the loss function of SIOU.

[0110] In a preferred embodiment of the present invention, the method of introducing the multi-head self-attention module into the skeleton part of the YOLOv5s model is:

[0111] like Figure 3 As shown in Figure 1, the multi-head self-attention module is a simple and powerful self-attention module that is suitable for a variety of machine vision tasks including image classification, object detection, and instance segmentation.

[0112] The multi-head self-attention layer adds position encoding to make it sensitive to position; while paying attention to feature information, the network is also sensitive to the relative positions between features, which plays a role in efficiently combining information.

[0113] Define the number of heads in the multi-head self-attention layer (for example, the number of heads is 4, etc., which can be adjusted according to the specific application scenario). The input first generates the query vector q, the key vector k and the value vector v through point convolution, and R h , R w Each represents a positional code extracted from the height and width;

[0114] After the position encoding performs the corresponding element addition operation, the position vector r is generated. r and q are matrix multiplied to generate the corresponding content-position vector qr T , and q and k perform matrix multiplication to generate the corresponding content-content vector qk T ;

[0115] qr T With qk T The corresponding elements are added, passed through the softmax layer, and then matrix multiplication is performed with v to finally obtain the output feature z (the output feature extracted after the multi-head self-attention mechanism).

[0116] In a preferred embodiment of the present invention, the method of introducing the shuffleattention module into the neck part of the YOLOv5s model is as follows:

[0117] Attention mechanism has become a key component to improve model detection performance. There are two types of attention mechanisms widely used in machine vision research, spatial attention mechanism and channel attention mechanism, which focus on information in spatial and channel dimensions respectively.

[0118] The channel attention mechanism helps the model confirm the feature information of the detection target, and the spatial attention mechanism helps the model obtain the location information of the detection target. Although the fusion of channel attention and spatial attention will improve performance, it will also increase the number of parameters and computational consumption.

[0119] like Figure 4 As shown in the figure, Shuffle attention reduces the number of parameters and computational cost of the attention mechanism while incorporating feature information from both channel and spatial dimensions, thereby improving the detection accuracy of the detector.

[0120] The input feature map dimension is c*h*w. The input is split into g groups along the channel dimension c, and the dimension of each group is c / g*h*w.

[0121] Each group will be split into two branches along the channel dimension again, and the dimension of each branch becomes c / 2g*h*w;

[0122] The two branches generate their own feature maps through the spatial attention module and the channel attention module respectively, helping the model to focus on the detection target and its location information;

[0123] After extracting information, the two feature maps are concatenated and the dimension becomes c / g*h*w. After all g groups have extracted features, they are concatenated again to get the output. The dimension of the output is still c*h*w, which remains the same as the dimension of the input.

[0124] The output is reordered through the channel reorganization function to ensure the information flow between different groups;

[0125] The spatial attention mechanism and channel attention mechanism used in Shuffle attention are simple to build. Compared with SE attention mechanism and CBAM attention mechanism, they have fewer parameters and lower computational cost while improving accuracy. The channel attention mechanism in the Shuffle attention module first averages the input to obtain a set of channel-related statistics. This set of statistics is linearly transformed and activated by the sigmoid function, and then multiplied with the corresponding elements of the original input to obtain the output result of the collected position information;

[0126] The spatial attention mechanism used in the Shuffle attention module first group-normalizes the input to obtain spatially related statistics. This set of statistics is linearly transformed and passed through the sigmoid activation function, and then multiplied with the corresponding elements of the original input to obtain the output result of the collected target information.

[0127] In a preferred embodiment of the present invention, the cross-modal image segmentation module obtains a color image of the area to be detected and a force map formed by the friction force in the area to be detected, and the force map uses grayscale values ​​to represent the magnitude of the friction force at different positions in the area to be detected, and adjusts the two images to the same resolution, and has Sjögren's disease segmentation label information in each image, and divides the image data pairs into a training set, a validation set, and a test set;

[0128] The image feature extraction module extracts features from the color image of the area to be detected to obtain single-modal color image features;

[0129] The friction force feature extraction module extracts features from the force map of the detection area to obtain single-mode friction force features;

[0130] The feature fusion module includes a first gating module (Relu function), a second gating module and a fusion network. The first gating module obtains color image features and processes the output part greater than the image threshold. The second gating module obtains friction features and processes the output part greater than the friction threshold. The fusion network fuses the features output by the first gating module and the second gating module, and uses the friction features as a channel of the output image.

[0131] In a preferred embodiment of the present invention, the method for removing the large target detection head of the YOLOv5s model is:

[0132] The classic yolov5 model includes The three sampling rate detection heads correspond to small target detection, medium target detection, and large target detection. Since the sample targets are all small and medium-sized targets, remove After the large target detection head with a magnification of 1, the network contains The detection heads with two sampling rates correspond to small target detection and medium target detection respectively.

[0133] Remove large target detection heads to improve detector accuracy and reduce the number of parameters and calculations.

[0134] The present invention also provides a small target detection method based on an improved YOLOv5s algorithm, comprising the following steps:

[0135] The images of the area to be detected were obtained and preprocessed; the images to be detected were divided into pathological block images of size 640*640 at the highest resolution, and 300 images were selected as the experimental data set. The data set was divided into training set, validation set and test set in a ratio of 8:1:1.

[0136] Obtain the friction force map of the area to be detected and perform preprocessing;

[0137] Based on the construction method of the present invention, the YOLOv5s model is improved to obtain an optimized YOLOv5s model;

[0138] Input the preprocessed image into the optimized YOLOv5s model, detect small objects in the image, and count the number of small objects in a single area;

[0139] The number of small targets in a single area is compared with the set threshold. When the threshold is exceeded, this area is screened out and marked, which can be used to diagnose whether the patient has Sjögren's syndrome.

[0140] For example, the experiment sets the training rounds to 100 epochs; batchsize to 5; input image size to 640*640; initial learning rate to 0.01, Adam is used as the optimization algorithm, the decay coefficient is 0.005, and the momentum parameter is 0.937. The labial gland biopsy pathology images are collected, and the labial gland biopsy WSI is divided into pathology block images of size 640*640 at the highest resolution. 300 images are selected as the experimental data set of this article, and the lymphocytes are manually annotated using labelimg based on the lymphocyte discrimination criteria. The data set is divided into training set, validation set, and test set in a ratio of 8:1:1.

[0141] This paper uses target detection metrics to evaluate the performance of the improved YOLOv5s in the lymphocyte detection task. The main metric of interest is mAP 0.5 Since the only target to be detected is lymphocyte, mAP 0.5 It can be expressed as:

[0142]

[0143] Among them, P and R represent precision and recall respectively, satisfying:

[0144]

[0145] Among them, P represents precision, R represents recall, TP refers to instances correctly predicted as positive examples, TN refers to instances incorrectly predicted as negative examples, FP refers to instances incorrectly predicted as positive examples, and FN refers to instances not incorrectly predicted as negative examples.

[0146] Improve the YOLOv5s target detection model, and make improvements in the loss function, feature extraction, and attention mechanism of the original model. In order to evaluate the impact of improvements in different modules and the combination of modules on the performance of the detection model, an ablation experiment is designed on the dataset in this article, using mAP 0.5 As evaluation indicators, the experimental results are shown in Table 1.

[0147] Table 1 Ablation experiment results

[0148]

[0149] The detection accuracy of the original YOLOv5s model on the dataset in this paper is 87.7%. After the multi-head self-attention module is installed on Backbone, the model accuracy is improved by 1.4%, and the number of parameters and GFLOPs are slightly reduced. After the Shuffle attention module is installed on the neck of the model, the model accuracy is improved by 1.5%, and the number of parameters and GFLOPs are slightly increased. After removing the large target detection head, the model accuracy is improved by 1.2%, the number of parameters is greatly reduced, and the GFLOPs also decreases significantly. After replacing the original CIOU with Focal-SIOU, the model accuracy is improved by 0.5% without changing the number of parameters and GFLOPs. After integrating the above improvement strategies, it can be seen that the final detection accuracy of the model reaches 91.1%. Compared with the original network, the accuracy is improved by 3.4%, the number of parameters is reduced by 29.6%, and the GFLOPs is reduced by 10.8%, which proves that the improvement strategy adopted in this paper has a significant improvement effect on lymphocyte detection.

[0150] By comparing the network of the present invention with other networks of similar size, it can be seen that the network of the present invention has certain advantages in both accuracy and parameter quantity, as shown in Table 2.

[0151] Table 2 Model comparison results

[0152] Network Model <![CDATA[mAP 0.5 / %]]> Parameter quantity GFLOPs Yolov7-tiny 85.2 <![CDATA[6.01×10 6 ]]> 13.0 Yolov7 85.5 <![CDATA[9.32×10 6 ]]> 26.7 RetinaNet 69.6 <![CDATA[19.8×10 6 ]]> 61.5 Yolov3-SPP 89.9 <![CDATA[4.12×10 6 ]]> 12.0 Yolov6n 88.3 <![CDATA[4.23×10 6 ]]> 11.8 RT-DETR 88.3 <![CDATA[20×10 6 ]]> 60 This article network 91.1 <![CDATA[4.94×10 6 ]]> 14.1

[0153] The improved YOLOv5s model can fully extract background information and effectively identify interfering cells with similar colors and shapes, such as epithelial cells, to ensure the accuracy of detection. Based on the improved YOLOv5s model, the WSI of lip gland biopsy is detected in blocks, the number of lymphocytes in a single area is counted, and the blocks with lymphocyte numbers greater than the set threshold are marked with color and regarded as suspicious lesions for reference by doctors, which serves the purpose of assisting the diagnosis of Sjögren's syndrome.

[0154] The present invention also provides a small target detection system based on an improved YOLOv5s algorithm, comprising an image acquisition module, a friction force map acquisition module and a processing module, wherein the image acquisition module is used to acquire an image to be detected, the friction force map acquisition module is used to acquire a friction force map of the area to be detected, the output ends of the image acquisition module and the friction force map acquisition module are respectively connected to the input end of the processing module, and the processing module executes the small target detection method of the present invention, detects small targets in an image, determines whether the number of small targets in a single area exceeds a threshold, and diagnoses whether a patient suffers from Sjögren's syndrome. Preferably, the friction force map acquisition module is a friction tester.

[0155] Using this system, by collecting and analyzing images, small target detection results of the images are obtained to diagnose whether the patient has Sjögren's syndrome.

[0156] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0157] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for building an improved YOLOv5s model, It is characterized in that The steps include: Use the Focal-SIOU loss function to replace the CIOU loss function of the YOLOv5s model; introduce the multi-head self-attention module into the skeleton network part of the YOLOv5s model; In the neck part of the YOLOv5s model, the shuffle attention module is introduced; A cross-modal image segmentation module is introduced after the shuffle attention module, wherein the cross-modal image segmentation module includes an image feature extraction module, a friction feature extraction module, and a feature fusion module; Remove the large object detection head of the YOLOv5s model to obtain the optimized YOLOv5s model; The cross-modal image segmentation module obtains a color image of the area to be detected and a force map formed by the friction force of the area to be detected, wherein the force map uses grayscale values ​​to represent the magnitude of the friction force at different positions of the area to be detected, adjusts the two images to the same resolution, and has Sjögren's disease segmentation label information in each image, and divides the image data pairs into a training set, a validation set, and a test set; The image feature extraction module extracts features from the color image of the area to be detected to obtain single-modal color image features; The friction force feature extraction module extracts features from the force map of the detection area to obtain single-mode friction force features; The feature fusion module includes a first gating module, a second gating module and a fusion network. The first gating module obtains color image features and processes the output part greater than the image threshold. The second gating module obtains friction features and processes the output part greater than the friction threshold. The fusion network fuses the features output by the first gating module and the second gating module, and uses the friction features as a channel of the output image.

2. The method for constructing the improved YOLOv5s model as claimed in claim 1, It is characterized in that Using the Focal-SIOU loss function, the method to replace the CIOU loss function of the YOLOv5s model is as follows: Assume that α is the angle between the predicted box and the real box coordinate center, let C h , C w Respectively represent the horizontal distance and vertical distance between the coordinate centers of the predicted box and the real box, then the straight-line distance σ and the angle α between the coordinate centers of the two are: According to the angle α, the angle loss Λ is defined as: The distance loss Δ of SIOU is: γ=2-Λ in, Represent the center point coordinates of the real box and the predicted box respectively; ρ x , ρ y Represents the distance loss factor between the real box and the predicted box in both width and height directions; γ is the angle loss factor; The angle loss is integrated into the distance loss function. When the angle α is closer to 45 degrees, the contribution of the angle loss is greater. When the angle α is closer to 0, the contribution of the angle loss is smaller and degenerates into the distance loss. At this time, C w With C h It represents the maximum distance between the predicted box and the true box, not the distance between their center points. The shape loss Ω of SIOU is: Among them, θ represents the attention to shape loss, and its value needs to be adjusted according to the specific data set; w, h, w gt ,h gt Represents the width and height of the predicted box and the real box respectively; ω w ,ω h Represents the shape loss factor in the width direction and height direction; After integrating the three indicators, the loss function L of SIOU SIOU for: Among them, IOU is the intersection-union ratio between the real box and the predicted box; Focal-loss is integrated with SIOU to distinguish high-quality and low-quality anchor boxes, which helps to improve the accuracy of regression. The Focal-SIOU loss function is: L Focal-SIOU =IOU γ L SIOU Among them, γ represents the attention to IOU, and its value range is greater than 0. The larger the value of γ is, the more attention the loss function pays to IOU; the closer the value is to 0, the less attention the loss function pays to IOU and gradually degenerates to L SIOU ; IOU represents the intersection-over-union ratio between the real box and the predicted box; L SIOU Represents the loss function of SIOU.

3. The method for constructing the improved YOLOv5s model as claimed in claim 1, It is characterized in that The method of introducing the multi-head self-attention module into the skeleton network part of the YOLOv5s model is: The C3 module of the original YOLOv5s network is restructured and integrated with a multi-head self-attention layer; Add position encoding to the multi-head self-attention layer to make it position sensitive; The number of heads in the multi-head self-attention layer is defined to balance accuracy and computation. The input is first generated by point convolution to generate query vector q, key vector k and value vector v, while R h , R w Each represents a positional code extracted from the height and width; After the position encoding performs the corresponding element addition operation, the position vector r is generated. r and q are matrix multiplied to generate the corresponding content-position vector qr T , and q and k perform matrix multiplication to generate the corresponding content-content vector qk T ; qr T With qk T The corresponding elements are added, passed through the softmax layer, and then matrix multiplication is performed with v to finally obtain the output feature z of the multi-head self-attention layer.

4. The method for constructing the improved YOLOv5s model as claimed in claim 1, It is characterized in that In the neck part of the YOLOv5s model, the method of introducing the shuffle attention module is as follows: The input feature map dimension is c*h*w. The input is split into g groups along the channel dimension c, and the dimension of each group is c / g*h*w. Each group will be split into two branches along the channel dimension again, and the dimension of each branch becomes c / 2g*h*w; The two branches generate their own feature maps through the spatial attention module and the channel attention module respectively, helping the model to focus on the detection target and its location information; After extracting information, the two feature maps are concatenated and the dimension becomes c / g*h*w. After all g groups have extracted features, they are concatenated again to get the output. The dimension of the output is still c*h*w, which remains the same as the dimension of the input. The output is reordered through the channel reorganization function to ensure the information flow between different groups; The channel attention mechanism in the shuffle attention module first averages the input to obtain a set of channel-related statistics. After this set of statistics is linearly transformed and activated by the sigmoid function, it is multiplied with the corresponding elements of the original input to obtain the output result of the position information. The spatial attention mechanism used in the Shuffle attention module first group-normalizes the input to obtain spatially related statistics. After linear transformation and passing through the sigmoid activation function, this set of statistics is multiplied with the corresponding elements of the original input to collect the output result of the target information.

5. The method for constructing the improved YOLOv5s model as claimed in claim 1, It is characterized in that The method to remove the large object detection head of the YOLOv5s model is: Since the sample targets are all small and medium-sized targets, remove The large-scale target detection head with a magnification of 100%, the network includes The detection heads with two sampling rates correspond to small target detection and medium target detection respectively.

6. A small target detection method based on improved YOLOv5s algorithm, It is characterized in that The steps include: Acquire the image of the area to be detected and perform preprocessing; Obtain a friction force map of the area to be detected and perform preprocessing; based on the construction method described in any one of claims 1 to 5, improve the YOLOv5s model to obtain an optimized YOLOv5s model; Input the preprocessed image into the optimized YOLOv5s model, detect small objects in the image, and count the number of small objects in a single area; The number of small targets in a single area is compared with the set threshold. When the threshold is exceeded, the area is screened out and marked.

7. A small target detection system based on improved YOLOv5s algorithm, It is characterized in that It includes an image acquisition module, a friction force map acquisition module and a processing module, wherein the image acquisition module is used to acquire an image to be detected, the friction force map acquisition module is used to acquire a friction force map of the area to be detected, the output ends of the image acquisition module and the friction force map acquisition module are respectively connected to the input end of the processing module, and the processing module executes the method described in claim 6 to detect small targets in the image, determine whether the number of small targets in a single area exceeds a threshold, and diagnose whether the patient suffers from Sjögren's syndrome.

8. The small target detection system based on the improved YOLOv5s algorithm as claimed in claim 7, It is characterized in that The friction force diagram acquisition module is a friction tester.