Improved tissue ultrasonic image detection lightweight system and method

Through the improved yolov5n network model and Wise-SIOU loss function, combined with the distillation strategy, the problem of large size, heavy weight and insufficient image processing capabilities of portable ultrasound equipment in resource-constrained areas is solved, and efficient and robust tissue ultrasound image detection is achieved.

CN120339793APending Publication Date: 2025-07-18EAST CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335468.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Portable ultrasonic devices have problems such as large size, heavy weight and insufficient image processing capabilities in applications in resource-limited areas. The prior art lacks robustness in complex or low-quality image processing and complex calculations, making it difficult to meet the demand for immediate detection.

Method used

The improved yolov5n network model is adopted, including the backbone layer, neck layer and head layer, the C3-Res2Block module and SPPF module with Bottle2neck structure are used, combined with the V5_AFPN neck structure network of the ASFF module, and the Wise-SIOU loss function and distillation strategy are introduced to optimize data processing capabilities.

Benefits of technology

It realizes the accuracy and robustness of image detection without increasing the size and weight of the equipment, ensures efficient operation in complex environments, and is suitable for real-time detection in resource-limited areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005321590060000031
    Figure BDA0005321590060000031
  • Figure BDA0005321590060000032
    Figure BDA0005321590060000032
  • Figure BDA0005321590060000041
    Figure BDA0005321590060000041
Patent Text Reader

Abstract

The invention discloses an improved tissue ultrasonic image detection lightweight system, and the system comprises an improved yov5n network model which comprises a backbone layer, a check layer, and a head layer; the detection lightweight system comprises a backbone layer and a check layer, a C3-Res2Block module comprising a Bottle2check structure is used in the backbone layer and the check layer, a short-cut branch is used in an SPPF of the backbone layer, and through a V5AFPN neck structure network comprising an ASFF module, the detection effect of the detection lightweight system is improved. The invention further provides a tissue ultrasound image lightweight detection method which has wide application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image detection, and relates to an improved lightweight system and method for tissue ultrasound image detection. Background Art

[0002] Compared with magnetic resonance imaging (MRI) and computed tomography (CT), ultrasound imaging has the significant advantages of being radiation-free, cost-effective and widely available, and has become an ideal conventional imaging detection method. However, traditional ultrasound devices are usually bulky and lack sufficient mobility, which is particularly problematic in grass-roots and resource-poor areas and often unable to provide convenient detection services. To solve this problem, innovative mobile algorithms and optimized data processing capabilities are needed to empower portable ultrasound devices to achieve instant tissue ultrasound image detection and significantly reduce the volume and weight of the devices.

[0003] In the prior art, portable ultrasound devices have begun to adopt new image processing technologies. For example, in the paper "Automatic thyroid nodule recognition and diagnosis in ultrasound imaging with the YOLOv2 neural network", an automatic image recognition, detection and diagnosis system based on the YOLOv2 neural network was established. Although the artificial intelligence system in this study performed well in thyroid nodule diagnosis, its application and lightweight design on portable devices were not carried out, which limited its application in resource-limited areas. At the same time, this study did not focus on how to improve the performance of the system in complex or low-quality images. In the paper "Evaluation of a deep learning-based computer-aided diagnosis system for distinguishing benign from malignant thyroid nodules in ultrasound images", a computer-aided diagnosis (CAD) system was introduced. The method of this study relied on the fusion of manually designed features (such as HOG, LBP, SIFT) and features extracted by deep learning, and then classified by a support vector machine (SVM). This method required manual feature design and was computationally complex, and was not as efficient and accurate as the end-to-end deep learning method of YOLOv5. In addition, YOLOv5 was also more robust in processing complex or low-quality images. Summary of the Invention

[0004] To address the deficiencies of the existing technology, the objective of the present invention is to provide an improved lightweight system and method for tissue ultrasound image detection with lightweight design and efficient operation. Through an improved mobile ultrasound imaging algorithm, the present invention aims to solve the drawbacks of existing portable ultrasound devices in terms of volume, weight, and real-time image processing capabilities through more efficient data processing techniques and algorithm optimization, thereby enhancing the accuracy and robustness of tissue ultrasound image detection. Meanwhile, a new loss function (Wise-SIO) is designed to further improve the performance of the model in complex environments. This system breaks away from the volume and cost limitations of traditional ultrasound devices, not only enabling efficient operation on mobile devices but also ensuring high detection accuracy.

[0005] By simplifying the structure of the deep learning model and reducing its demand for computing resources, the present invention enables portable ultrasound devices to achieve faster image processing speed and longer battery life without sacrificing image quality, making them more suitable for use in resource-constrained environments. Additionally, the present invention focuses on improving the generalization ability of the model to ensure accurate and reliable imaging results in various clinical situations.

[0006] The present invention provides an improved lightweight system for tissue ultrasound image detection. The detection lightweight system includes: an improved yolov5n network model, and the improved yolov5n network model includes a backbone layer, a neck layer, and a head layer;

[0007] The C3-Res2Block module containing the Bottle2neck structure is used in the backbone layer and the neck layer, a short-cut branch is used in the SPPF of the backbone layer, and the detection effect of the detection lightweight system is enhanced through the V5_AFPN neck structure network containing the ASFF module.

[0008] The backbone layer is used to extract features of the input image and includes a Conv module, a C3-Res2Block module, and an SPPF module;

[0009] The Conv module includes a convolutional layer, a BN layer, and an activation function. The C3-Res2Block module is used to adaptively aggregate the feature maps obtained by Conv, and the SPPF module obtains spatial information through the weighted fusion of global features and local features.

[0010] In a specific embodiment, the convolutional kernel size in the Conv module of the backbone layer is 3*3, and the stride is 2;

[0011] In the present invention, the C3-Res2Block module includes two branches. One branch includes a Conv module and a Bottle2neck module, and the other branch only includes a Conv module. After connecting the outputs of the two branches and then passing through a Conv module, the output of the C3-Res2Block module is realized.

[0012] In a specific embodiment, the convolution kernel size of the Conv module in the C3-Res2Block module is 1×1.

[0013] The Bottle2neck module includes an upsampling convolution layer, a feature map partitioning layer, a fusion convolution layer, a concatenation layer, and a downsampling convolution layer.

[0014] The upsampling convolution layer increases the number of channels of the input feature map, enhances the representation ability of the feature map, and captures more information.

[0015] The feature map partitioning layer equally partitions the upsampled feature map by channels for subsequent per-channel feature fusion and processing.

[0016] The fusion convolution layer integrates information from different channels through feature fusion and convolution operations to generate new high-dimensional features.

[0017] The concatenation layer concatenates multiple convolved feature maps to integrate multi-channel features and enhance the feature expression ability.

[0018] The downsampling convolution layer uses a 1×1 convolution to reduce the number of channels of the concatenated feature map back to the original level, reducing the computational burden and controlling the model complexity.

[0019] In the Bottle2neck module, through residual connection, the input feature map is added to the downsampled feature map per channel to prevent gradient disappearance and promote information flow during the training process.

[0020] In the present invention, the SPPF module includes a convolution layer with multiple convolution parameters. After convolving the feature map respectively, the feature sets obtained from multiple convolutions are connected to the input feature map, and then convolution is used to reduce the dimension to obtain the output feature map.

[0021] In the present invention, the neck layer uses a V5_AFPN neck structure network including an ASFF module, specifically including multiple ASFF_2 and ASFF_3 modules.

[0022] In the improved neck layer, a 1*1 convolution is used to reduce the dimension of the two branches respectively, and then a Concat operation is performed on the two branches. Subsequently, a 1*1 convolution is used to reduce the output dimension by 2.

[0023] The output result is subjected to Softmax operation according to dimensions, and the results after Softmax of the two dimensions are multiplied by input1 and input2 respectively.

[0024] The outputs are added, and the result after addition is then convolved through a 3×3 convolution to obtain the output of the entire module.

[0025] The ASFF_2 module includes: a Dowmsample and Upsample module for adjusting dimensions and sizes, three 1×1 convolution modules, a Concat aggregation module, a Softmax module, and two weight modules;

[0026] The ASFF_3 module processes the pairwise Inputs among the three inputs. The three input inputs are pairwise combined and processed through the ASFF_2 algorithm for each pair of inputs. This algorithm aims to improve the performance and stability of the overall model by optimizing the dependency relationships and interactions among the inputs. The core of this process lies in dynamically adjusting the input data to adapt to the non-linear associations among different inputs.

[0027] The present invention also provides a lightweight detection method for tissue ultrasound images, and the lightweight detection method includes the following steps:

[0028] Step 1: Collect tissue ultrasound images to be detected;

[0029] Step 2: Construct an improved yolov5n network model for image detection and perform training optimization;

[0030] Step 3: Input the tissue ultrasound images to be detected in Step 1 into the constructed improved yolov5n network model to achieve the detection of tissue ultrasound images.

[0031] In Step 2, when training and optimizing the improved yolov5n network model, the optimization of the model is guided and evaluated by introducing the Wise-SIOU loss function; the Wise-SIOU loss function adds three cost functions on the basis of the IOU loss function, including an angle cost Λ, a distance cost Δ, and a shape cost Ω;

[0032] The angle cost Λ is calculated by the following formula:

[0033]

[0034] where,

[0035]

[0036] b gt Cx, b gt Cy is the center coordinate of the ground truth box, b Cx , b Cy is the center coordinate of the predicted box, C h C is the height difference between the centers of the ground truth box and the predicted box, and α is the angle between the line connecting the centers of the ground truth box and the predicted box and its horizontal projection;

[0037] The distance cost Δ is calculated by the following formula:

[0038]

[0039] where,

[0040]

[0041] Δ represents the total loss or error, ρ x and ρ y are the squared normalized errors in the x and y directions, C h is the height difference between the centers of the ground truth box and the predicted box, C w is the width difference between the centers of the ground truth box and the predicted box, and Υ is an adjustment factor depending on Λ;

[0042] The shape cost Ω is calculated by the following formula:

[0043]

[0044] where,

[0045]

[0046] w and h: represent the width and height of the current rectangle, w gt and h gt : represent the width and height of the target rectangle or the reference rectangle, ω w : represents the relative measure of the width difference, ω h : represents the relative measure of the height difference; θ is an adjustable variable used to indicate how much attention the network needs to pay to this cost;

[0047] In a specific embodiment, θ = 4;

[0048] The final calculation of SIOU is shown by the following formula:

[0049]

[0050] By setting the balance scheme of high and low quality samples of Wise - SIOU, the training effect of the object detection model is adjusted and optimized;

[0051] The calculation formula of the loss function is as follows:

[0052]

[0053] Among them,

[0054] L WSIOUvl = R WSIOU *L SIOU ,

[0055]

[0056] Among them, the parameters α and β are user-defined parameters. α is a scaling factor used to adjust the influence of components such as the loss function, and β is a regulation coefficient used to balance the contributions or influences of different components in the loss function or formula; Wg and Hg are the length and width of the minimum bounding rectangle of the two boxes, x gt and y gt are the horizontal and vertical coordinates of the target rectangle, R WSIOU is the attention adjustment factor; L SIOU represents the loss function of SIOU.

[0057] In the present invention, a model distillation strategy that comprehensively considers the output layer and the feature layer is also used;

[0058] During the distillation process of the output layer, the distillation of the output layer is transformed into the distillation of the bounding box, classification, and confidence, and the loss functions are respectively

[0059] Among them,

[0060]

[0061] Among them, L D box is the bounding box loss, CIOU is an index that measures the overlap degree between the predicted box and the true box, is for calculating the predicted bounding box, is the true bounding box;

[0062]

[0063] Among them, N is the batch size, M is the height of the feature map, X is the width of the feature map, Y is the depth or number of classes of the feature map, Λ i,j,k is the weight used to adjust the contribution of each class in the loss calculation, P T i,j is the predicted probability of the k-th class of the teacher network at the position (i, j), P S i,j is the predicted probability of the student network at the same position and class, log(P S i,j ) is the log-likelihood when the class exists;

[0064] The implementation of confidence distillation loss is similar to that of classification distillation loss, both calculating the loss through the prediction results between the teacher network and the student network. The difference is that confidence distillation loss focuses on the confidence of each predicted bounding box (i.e., the probability that the box contains the target), rather than the specific class. Specifically, in object detection, the teacher network and the student network will respectively predict the confidence values of each box, and calculate the difference between the two through binary cross-entropy loss (BCE loss). The calculation method is similar to the classification loss, except that here the predicted confidence values are compared instead of the class probabilities.

[0065] The core objective of confidence distillation loss is to adjust the predicted confidence of the student network to be consistent with that of the teacher network, thereby improving the accuracy of the student network in predicting the confidence of the target bounding box. A weight coefficient λ can be introduced to control the contribution of the confidence loss to the overall loss, so as to balance it with other losses (such as classification loss or localization loss). Ultimately, confidence distillation loss can help the student network better predict the presence or absence of the target bounding box, improving the accuracy and robustness of detection.

[0066] In summary, the implementation of confidence distillation loss is similar to that of classification distillation loss. Both compare the outputs of the teacher and student networks and optimize through BCE loss, but it is for the confidence of the bounding box rather than class information.

[0067] During the output layer distillation process, combining the bounding box distillation loss, classification distillation loss, and confidence distillation loss, the complete output layer distillation loss consists of the weighted sum of these three parts of losses, expressed as follows:

[0068] L output = λ1·L bbox + λ2·L cls + λ3·L obj ,

[0069] where L bbox represents the bounding box distillation loss, L cls represents the classification distillation loss, L obj represents the confidence distillation loss, λ1 represents the bounding box loss weight, λ2 represents the classification loss weight, and λ3 represents the confidence loss weight;

[0070] Specifically, the final loss function of the model is divided into three parts, namely the bounding box loss L box , the classification loss L cls , and the confidence loss L. And these losses are all calculated from the outputs of the three dimensions of the model. Inspired by this, when performing the distillation loss of the output layer, the output can also be decomposed into these three types, and then distillation is performed on each part separately.

[0071] Coefficients are assigned to the feature layer distillation loss and the output layer distillation loss respectively for parameter tuning and optimization.

[0072] During the output layer distillation process, the losses of feature layer distillation and output layer distillation are weighted and summed to obtain the final total loss:

[0073] L total = α·L feature + β·L output

[0074] Where: L feature is the feature layer distillation loss, which measures the difference between the student model and the teacher model in the intermediate feature layer. L output is the output layer distillation loss, including the bounding box distillation loss, classification distillation loss and confidence distillation loss, and is obtained through weighted sum. α and β are adjustment coefficients used to balance their impacts during training.

[0075] By adjusting α and β, the focus of the model can be controlled when learning the intermediate features (L feature ) and the final prediction result (L output ). If α is larger, the model pays more attention to the feature layer; if β is larger, the model pays more attention to the prediction accuracy of the output layer.

[0076] The beneficial effects of the present invention include: aiming at the problem that the accuracy of the selected baseline model is not high enough, the present invention proposes a series of improvement schemes for its network structure, including improving the key modules for feature extraction in its original backbone, improving the original pooling layer structure, and improving the original neck layer for feature fusion, so that the accuracy of the model is improved; aiming at the shortcomings of the original loss function, the present invention proposes an improved loss function scheme called Wise-SIOU to replace the original CIOU loss function, and the accuracy is improved compared with the original loss function; the present invention proposes a distillation strategy that comprehensively considers the feature layer and the output layer. By calculating the KL divergence between the feature layer information of the student network and the teacher network as the feature layer distillation loss; and using the output of the teacher network as the soft label to calculate the distillation localization, confidence and classification losses as the output layer distillation loss, and comprehensively considering these two types of losses to achieve model distillation. The accuracy of model recognition and detection is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0078] Figure 1 It is a schematic diagram of the existing yolov5n network structure.

[0079] Figure 2 It is a structural diagram of the C3 module of the original yolov5.

[0080] Figure 3 It is a structural diagram of the improved Bottle2neck module of the present invention.

[0081] Figure 4 It is a structural diagram of the ASPP-U module of the present invention.

[0082] Figure 5 It is a network structure diagram of V5_AFPN.

[0083] Figure 6 It is a structural diagram of the ASFF_2 module.

[0084] Figure 7 It is a schematic diagram of the letter meanings of the SIOU angle cost calculation.

[0085] Figure 8 It is a schematic diagram of the letter meanings of the SIOU distance cost calculation.

[0086] Figure 9 It is a schematic diagram of the distillation strategy of the present invention.

[0087] Figure 10 It is a schematic diagram of the transverse section annotation of the carotid artery blood vessel of the present invention.

[0088] Figure 11 It is a schematic diagram of the transverse section annotation of the carotid artery plaque of the present invention.

[0089] Figure 12 It is a schematic diagram of the longitudinal section annotation of the intima-media of the carotid artery of the present invention.

[0090] Figure 13 It is a schematic diagram of the longitudinal section annotation of the carotid artery plaque of the present invention.

[0091] Figure 14 It is a heat map of the teacher network and the student network of the present invention. Detailed implementation manners

[0092] Combined with the following specific embodiments and the attached drawings, the present invention will be further described in detail. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and well-known common sense in the art, and the present invention has no particularly restricted content.

[0093] The present invention proposes an improved lightweight method and corresponding system for tissue ultrasound image detection. This method can achieve targeted carotid plaque detection on mobile devices, effectively allowing these devices to replace ultrasound professionals in identifying and diagnosing carotid plaques. This innovation not only enhances portability but also expands access to crucial diagnostic capabilities, potentially transforming stroke prevention strategies. The main contributions of the method of the present invention include significant improvements in model detection performance, the introduction of a new loss function Wise-SIOU to address the deficiencies of existing functions, and a new distillation technique to improve the accuracy of the pruned model, thereby restoring and enhancing model performance.

[0094] Specifically, the present invention improves the model structure and loss function in the lightweight system for tissue ultrasound image detection.

[0095] The benchmark adopted by the lightweight detection system described in the present invention is the yolov5n network model. The structure of the conventional yolov5n network model is as Figure 1 shown. The yolov5n network model is mainly divided into three parts: the backbone layer, the neck layer, and the head layer. Among them, the backbone layer is mainly used for extracting the features of the input image; the neck layer is mainly used for extracting and fusing the feature information generated by the backbone layer and generating multi-scale feature information; the role of the head layer is to convert the multi-scale information output by the neck layer into the detection result.

[0096] In the backbone layer, it performs a dimensionality reduction operation through a convolution with a size of 3*3 and a stride of 2, and its feature extraction is mainly carried out through the C3 module. Finally, a pooled SPPF layer is added to obtain multi-scale information.

[0097] In the neck layer, it adopts the FPN+PAN structure to deeply extract and fuse the features extracted by the backbone. The FPN structure realizes the transfer of high-level features to the low level, enhancing the semantic expression ability of the low-level features. The PAN structure then transmits the low-level features fused by the FPN back to the high level, supplementing the detail information of the high-level features.

[0098] In the head layer, mainly through the way of predefined anchor boxes, three anchor boxes are defined at each of the three scales. Therefore, the output of the model for the aspect ratio of the bounding box can be simplified to the scaling coefficients of these three anchor boxes, reducing the training complexity and difficulty of the model. The output of the final head layer is the output of the model.

[0099] A. Improvement of the model structure

[0100] The schematic diagram of the C3 module structure of the conventional yolov5n is asFigure 2 As shown, after the feature map enters the C3 module, it will be divided into two paths. The left path enters the Conv and Bottleneck modules, and the right path only has one Conv module. After the two paths are concatenated, they enter another Conv module and then are output from the C3 module. The three Conv modules in C3 all have 1×1 convolutional kernels, and their main function is to reduce or increase the dimension, which is not very meaningful for feature extraction. The Bottleneck module uses residual connection, which contains two other Conv modules. The first Conv is used to halve the number of channels, and the second Conv is used to double the number of channels. First reducing the dimension is beneficial for the convolutional kernel to better understand the feature information, and increasing the dimension is beneficial for extracting more detailed features. Finally, a large residual structure is used to avoid gradient loss.

[0101] To improve the model accuracy, the present invention adopts a C3-Res2Block module to replace the C3 module in the original yolov5 model. The main difference between it and the C3 module is that a Bottle2neck structure is proposed to replace the Bottleneck structure in the original C3 module. The structure diagram of Bottle2neck is as Figure 3 shown. The Bottle2neck first performs a dimension-increasing operation through a 1×1 convolution, and then divides the output feature map into four equal parts according to the number of channels. One-fourth of them is directly inherited to the next-layer feature map. Each of the subsequent one-fourth channels first fuses its own feature map after convolution with the previous one-fourth channel, and then performs convolution to obtain the feature map after convolution of this channel number. After all convolutions are completed, these four one-fourth channels are concatenated, and then a dimension-reducing operation is performed through a 1×1 convolution. Finally, it is fused with a shortcut to prevent gradient loss to obtain the output of the entire Bottle2neck.

[0102] For the original SPPF module, multiple parallel dilated convolutions with different sampling rates are used to construct convolutional kernels with different receptive fields through different dilation rates to obtain multi-scale object information. The features extracted for each sampling rate are further processed in separate branches and fused to generate the final result. In the present invention, the global receptive field branch of the original SPPF module is directly replaced with the built-in short-cut branch, reducing the computational amount of the module and improving the convenience of module deployment. The structure diagram of the improved SPPF module (ASPP-U) is as Figure 4As shown in the figure. The improved SPPF module (ASPP-U) first performs 1×1 convolution, 3×3 convolution with a dilation rate of 6, 3×3 convolution with a dilation rate of 12, and 3×3 convolution with a dilation rate of 18 on the input feature map respectively, so as to implement convolution operations with receptive fields of 1×1, 13×13, 25×25, and 37×37 respectively. Then, the feature sets obtained by multiple convolutions and the input feature map are subjected to a Concat operation, and finally a 1×1 convolution is used for dimensionality reduction operation to obtain the output feature map.

[0103] For the original neck network structure of FPN+PAN, the present invention designs a neck structure network of V5_AFPN, and its network structure is as Figure 5 shown Figure 5 shown. The structure inputs P3, P4, and P5, which are feature maps from different layers of the backbone network respectively. Among them, P3 and P4 first pass through a Conv1*1 convolution operation respectively, and then cross-fuse through the ASFF_2 module respectively. The output results are respectively input into the C3-Res2Block module containing the Bottle2neck structure. After P5 passes through a Conv1*1 convolution, the output result is then input into the ASSF_3 module pairwise with the outputs of the above two respectively, and finally output to the C3-Res2Block module containing the Bottle2neck structure. In the neck layer, the interaction of information of feature maps of different sizes is mainly realized through the ASFF module. The schematic diagram of the structure of ASFF_2 is as Figure 6 shown. It first passes through a Dowmsample / Upsample module that jointly adjusts the dimension and size. Which one to use specifically is based on the incoming parameter. If the incoming parameter is 0, input1 is subjected to a Downsample operation to ensure that its dimension and size are consistent with input2. If the incoming parameter is 1, input2 is subjected to an Upsample operation to ensure that its dimension and size are consistent with input1. Then, 1*1 convolution is used for dimensionality reduction operation on the two branches respectively, and then the two branches are subjected to a Concat operation. Then, a 1*1 convolution is used to reduce the output dimension to 2. Then, a Softmax operation is performed according to the dimension to obtain the weights, and the results after Softmax of the two dimensions (i.e., weights weight[0] and weight[1]) are multiplied by input1 and input2 respectively, and then the obtained results are added. The added result is then passed through a 3*3 convolution to obtain the output of the entire module;

[0104] The ASFF_3 module processes the input by pairwise calling ASFF_2 to optimize the dependency relationship and interaction between the inputs.

[0105] In a specific embodiment, the ASFF_3 module receives three input feature maps and uses the ASFF_2 module for fusion in a pairwise input manner. In the ASFF_2 module, first, the sizes of the input feature maps are adjusted through Downsample / Upsample operations to make them match. Subsequently, a dimensionality reduction operation is performed on each pair of inputs, and a 1×1 convolution is used to reduce the dimensionality of the feature maps. Then, the feature maps after dimensionality reduction are merged into a higher-dimensional feature map through a concatenation (Concat) operation. Finally, the merged feature map is further processed through a convolutional layer to obtain the fused features. In this way, the ASFF_3 module fuses three feature maps of different scales into a final multi-scale feature map through two ASFF_2 operations, improving the model's detection ability for targets of different sizes while avoiding information loss.

[0106] After the improvement of the model in the present invention, it has an advantage in improving the recognition accuracy. In the modification of the C3 module part, the model recognition accuracy is improved from 0.551 to 0.574. In the modification of the SPPF module, the model recognition accuracy is improved from 0.551 to 0.609. After integrating these improvements together, the overall model is named plaque-yolov5, and the recognition accuracy of the overall model is improved from 0.551 to 0.639.

[0107] B. Improvement of Loss Function

[0108] The original iou loss function used in yolov5n is CIOU. However, there are two obvious defects in the CIOU loss: one is that the processing of the additional loss term only considers the additional loss of the distance between the center points of the prediction box and the ground truth box and the length and width, and lacks the angular information of the relative positions of the two boxes; the other is that the regression samples with poor quality have a greater impact on the regression loss, while the samples with better regression quality are difficult to further optimize. Therefore, the present invention introduces a loss function called Wise-SIOU during the model training process to solve this problem.

[0109] In SIOU, compared with the most basic IOU, it adds three cost functions, namely the angular cost Λ, the distance cost Δ, and the shape cost Ω. The definitions of their respective costs are as follows:

[0110] The angular cost Λ describes the minimum angle between the connection line of the center points and the x and y axes. When the included angle α is 0 and the connection line of the center points is exactly the x-axis or y-axis, Λ = 0. When the included angle α is 45 degrees, Λ = 1. This term can guide the anchor box to the nearest axis of the target box, and its detailed calculation is shown in the following formula:

[0111]

[0112] The meanings of the letters in the formula are as Figure 7 shown. Specifically, b gt Cx , b gt Cy is the center coordinate of the ground truth box, b Cx , b Cy is the center coordinate of the predicted box, C h is the height difference between the centers of the ground truth box and the predicted box, Figure 7 The C in w is the width difference between the centers of the ground truth box and the predicted box. The included angle α is the angle between the line connecting the centers of the ground truth box and the predicted box and its projection in the horizontal direction.

[0113] The distance cost Δ describes the distance between the centers, and its penalty cost is related to the angle cost. Its detailed calculation is shown in the following formula:

[0114]

[0115] Among them,

[0116]

[0117] The meanings of the letters in the formula are as Figure 8 shown. Specifically, Δ represents the total loss or error, ρ x and ρ y are the squared normalized errors in the x and y directions, C h is the height difference between the centers of the ground truth box and the predicted box, C w is the width difference between the centers of the ground truth box and the predicted box, Υ is an adjustment factor, and its value depends on Λ. Usually, Λ is a fixed adjustment parameter used to adjust the influence of the error. Specifically, taking the x-axis as an example, when the two boxes are almost parallel, the angular distance is almost 0, and at this time Υ is close to 2, then the contribution of the distance between the two boxes to the loss is also very small. When the included angle between the line connecting the two boxes is 45 degrees, at this time Υ is close to 1, and at this time the distance between the two boxes needs to be emphasized and should account for a larger loss.

[0118] The shape cost Ω considers the aspect ratio between the two boxes. It is mainly defined by calculating the ratio of the difference in length (width) between the two boxes to the maximum length (width) of the two boxes. θ is an adjustable variable used to represent how much attention the network needs to pay to this cost. In a specific embodiment of the present invention, the adjustable variable is set to 4, and its calculation method is shown in the following formula:

[0119]

[0120] Among them, w and h: represent the width and height of the current rectangle, w gt and h gt: represents the width and height of the target rectangle or the reference rectangle, ω w : represents the relative measure of the width difference, ω h : represents the relative measure of the height difference;

[0121] Finally, the calculation method of SIOU is shown in the following formula:

[0122]

[0123] Based on SIOU, the present invention proposes a new balance scheme for high and low quality samples of Wise-SIOU, which mainly has three versions: v1, v2, and v3. Version v1 ensures that for anchor boxes of general quality, they can receive sufficient attention, while for anchor boxes of better quality, their attention to the center point distance can be reduced. Version v2 is an intermediate version without specific usage scenarios. Version v3 is the final adopted version.

[0124] The formula for calculating the bounding box loss of version v1 is shown in the following formula:

[0125] L WSIOUvl =R WSIOU *L SIOU ,where

[0126]

[0127] Among them, Wg and Hg are the length and width of the minimum bounding rectangle of the two boxes, x gt and y gt are the horizontal and vertical coordinates of the target rectangle, R WSIOU is the attention adjustment factor, which ensures that for anchor boxes of general quality, they can receive sufficient attention, while for anchor boxes of better quality, their attention to the center point distance can be reduced; L SIOU represents the loss function of SIOU.

[0128] Version v2 is inspired by focalloss [1] and adds a monotonic focusing coefficient on the basis of v1. However, it is found during the training process that its value decreases as L SIOU decreases, resulting in a slower convergence speed of the model in the later stage. Considering this factor, the present invention considers introducing the mean value of L SIOU as the normalization factor. Therefore, the calculation result of version v2 finally obtained is shown in the following formula:

[0129]

[0130] Version v3 defines an outlier degree to describe the quality of the anchor box, and its definition is shown in the following formula:

[0131]

[0132] For this value, if it is small, it indicates a higher quality of the anchor box, and a small gradient gain is assigned to it so that the bounding box regression focuses on the anchor boxes of normal quality; if it is large, a smaller gradient gain is assigned to it to prevent low-quality samples from generating large harmful gradients. Finally, an outlier coefficient is constructed and applied to the v1 version to achieve the design of the v3 version, and its definition is shown in the following formula:

[0133]

[0134] Among them, the parameters α and β are user-defined parameters. α is a scaling factor used to adjust the influence of components such as the components of the loss function, and β is a regulation coefficient used to balance the contributions or influences of different components in the loss function or formula. δ is used to measure the quality (outlier) of the anchor box. It adjusts the gradient gain of the anchor box during training. Anchor boxes with higher quality correspond to smaller δ, while anchor boxes with poorer quality correspond to larger δ, thereby balancing the contributions of anchor boxes of different qualities to the loss function; Wg and Hg are the length and width of the minimum bounding rectangle of the two boxes, x gt and y gt are the horizontal and vertical coordinates of the target rectangle, R WSIOU is the attention adjustment factor; L SIOU represents the loss function of SIOU.

[0135] In the present invention, a model distillation strategy with better effects than the prior art is adopted, comprehensively considering the output layer and the feature layer.

[0136] Traditional model distillation methods mainly detect the input images through the teacher model to obtain corresponding outputs, and then use these outputs as soft labels to train the student model. However, the output results of object detection are different from the logits output that only contains various classifications. It also contains a series of information such as bounding boxes and confidences. Therefore, directly applying the output distillation method for object classification to object detection lacks rationality. In addition, in the network, there are not only output layers, but also a series of intermediate feature layers in the network, which also contain rich image information. Letting the student model directly learn the output of the teacher model without considering these specific features is likely to exceed the learning ability of the student network and lead to a deterioration in the distillation effect.

[0137] Considering the problems existing in the above traditional distillation methods, the present invention designs a distillation strategy that comprehensively considers the output layer and the feature layer. The distillation strategy of the present invention is as Figure 9 shown. Among them, are the output feature maps of the teacher network backbone layer in three dimensions of P3, P4, and P5 respectively, The output feature maps of the student network's backbone layer in the three dimensions of P3, P4, and P5 are used as the input for feature distillation, which is a way of feature layer distillation; The output feature maps of the teacher network's neck layer in the three dimensions of P3, P4, and P5 The output feature maps of the student network's neck layer in the three dimensions of P3, P4, and P5 are used as the input for feature distillation, which is another way of feature layer distillation. Subsequently, one of them will be selected through experiments as the feature distillation layer of the present invention. Output layer distillation uses the three-dimensional results output by the teacher network and the student network on the detection head as the input for distillation. Subsequently, the methods of feature layer distillation and output layer distillation will be introduced in detail.

[0138] A. Feature layer distillation strategy

[0139] The traditional pruning distillation strategy is to use the Softmax function to distill the output result of the teacher network at a determined distillation temperature to obtain soft labels, and then use the soft labels to train the result of the student network.

[0140] In the feature layer of the present invention, the activation map in each channel can be first normalized to make the foreground of the feature map more prominent, and then the KL divergence between the normalized channel activation maps of the teacher and student networks is reduced. In this way, the student network is guided to pay more attention to the regions with large normalized activation values in each channel. The specific steps of distillation should be as follows:

[0141] I. Unify the number of channels of the teacher network and the student network: Generally, the network structure and the number of parameters of the teacher network are larger than those of the student network, and the number of channels in its feature layer is also likely to be larger than that of the student network. Therefore, first, a 1*1 convolution is used to reduce the number of channels of the teacher model to ensure that the feature layers of the student network and the teacher network have the same number of channels.

[0142] II. Perform a normalization transformation on each channel of the feature layers of the student network and the teacher network. Before the transformation, the feature layers of the student network and the teacher network can both be represented as R N*M*X*Y , where N is the number of images in a batch, M is the number of channels, and X*Y is the length and width of the feature map. The transformation method for the activation map of the N*Mth channel is shown in the following formula:

[0143]

[0144] where F i.j is the activation value of the teacher and student networks at the (i, j) position; is the activation value after normalization for the teacher network or student network at the (i, j) - th pixel, and Γ is the distillation temperature at this time.

[0145] III. Calculate the KL - divergence for the activation value maps of the teacher network and the student network after normalization. The final feature - layer loss is the quotient of the sum of the KL - divergence values of each channel and the product of the batch size and the number of channels N * M. The final loss function is shown as the following formula. Through this formula, it can be found that when is larger, that is, when the teacher network attaches importance to this pixel point, its contribution to the loss function is also larger. On the contrary, its contribution to the loss function is smaller.

[0146]

[0147] where F T is the activation value of the teacher network, F s is the activation value of the student network, is the activation value of the teacher network at the (i, j) position. is the activation value of the student network at the (i, j) position, N is the batch size of the feature map (batchsize), M is the number of channels of the feature map (channels), that is, the depth of the output feature map, ∑ N k=0 is to sum over all batches, ∑ M i=1 ∑ M j=0 is to sum over each pixel of the feature map; T represents a parameter related to the temperature (distillation temperature). During the distillation process, the temperature T controls the smoothness of the outputs of the teacher network and the student network. In the setting of knowledge distillation, the temperature T is generally used to control the smoothness of the output probability distribution of the teacher network, thereby affecting the learning process of the student network.

[0148] B. Output - layer distillation strategy

[0149] According to the relevant description of the aforementioned loss function, the final loss function is divided into three parts, namely the bounding - box loss classification loss confidence loss And the acquisition of these losses is calculated from the outputs of three dimensions of the model. Inspired by this, when performing the distillation loss of the output layer, the output can also be decomposed into these three types, and then distillation is performed on each part separately. Therefore, the distillation of the output layer can also be transformed into three parts, namely the distillation for the bounding - box, classification, and confidence. The loss functions of these three types of distillation are respectively The following will introduce these three loss functions respectively:

[0150] I. Bounding Box Distillation Loss: Assume that for the same image input, the outputs at the $i$-th position are and respectively, and the corresponding anchor box is $A$ i . Through these three parameters, the bounding boxes predicted by the teacher and the student are decoded as and respectively. Subsequently, taking the prediction of the teacher network as the soft label, the IOU loss between the two is calculated using CIOU, and its calculation method is shown in the following formula:

[0151]

[0152] where $L$ D box is the bounding box loss, and CIOU is an index measuring the overlap degree between the predicted box and the ground truth box. is the predicted bounding box, and is the ground truth bounding box.

[0153] II. Classification Distillation Loss: Let $K$ be the number of categories, then the dimension of the activation map for classification by the detection head is $R$ N*M*X*Y . Different from the Softmax function used in traditional classification tasks, in object detection, there is the problem of multi-object detection. Therefore, it is more reasonable to use the Sigmoid function to compress the probability information of each classification into the interval $(0, 1)$. After using the Sigmoid function, the predicted classification probabilities of the teacher network and the student network are $P$ T and $P$ S respectively. Thus, $P$ T can be regarded as the label, and $P$ S as the model prediction. The BCE loss is calculated for these two, and in addition, a weight coefficient $\lambda$ can be added to optimize the classification loss. The definitions of the classification loss and the weight coefficient are shown in the following formula:

[0154]

[0155] where $N$ is the batch size, $M$ is the height of the feature map, $X$ is the width of the feature map, $Y$ is the depth or the number of categories of the feature map, that is, the number of predicted categories at each position of the network output. $\Lambda$ i,j,k is the weight used to adjust the contribution of each category in the loss calculation. $P$ T i,j is the predicted probability of the $k$-th category at the position $(i, j)$ of the teacher network. $P$ S i,j is the predicted probability of the student network at the same position and category. $\log(P$ S i,j ) is the log-likelihood when the category exists.

[0156] III. Confidence Distillation Loss: The implementation of confidence distillation loss is similar to that of classification distillation loss, both calculating the loss through the prediction results between the teacher network and the student network. The difference is that confidence distillation loss focuses on the confidence of each predicted bounding box (i.e., the probability that the box contains the target), rather than the specific class. Specifically, in object detection, the teacher network and the student network will respectively predict the confidence values of each box, and calculate the difference between them through the binary cross-entropy loss (BCE loss). The calculation method is similar to the classification loss, except that here the predicted confidence values are compared instead of the class probabilities.

[0157] The core objective of confidence distillation loss is to adjust the predicted confidence of the student network to be consistent with that of the teacher network, thereby improving the accuracy of the student network in predicting the confidence of the target bounding box. A weight coefficient λ can be introduced to control the contribution of the confidence loss to the overall loss, so as to balance it with other losses (such as classification loss or localization loss). Ultimately, confidence distillation loss can help the student network better predict the presence or absence of the target bounding box, improving the accuracy and robustness of detection.

[0158] In summary, the implementation of confidence distillation loss is similar to that of classification distillation loss. Both compare the outputs of the teacher and student networks and are optimized through BCE loss, except that it is for the confidence of the bounding box rather than the class information;

[0159] The final output layer distillation loss function L output is expressed as the weighted sum of three parts:

[0160] L output = λ1·L bbox + λ2·L cls + λ3·L obj

[0161] where,

[0162] L bbox : Bounding box distillation loss, usually using CIOU or GIOU loss, to measure the difference between the predicted bounding box and the ground truth bounding box.

[0163] L cls : Classification distillation loss, measuring the difference between the predicted class and the ground truth class, usually using cross-entropy loss.

[0164] L obj : Confidence distillation loss, measuring the probability that the bounding box contains the target, usually using binary cross-entropy loss.

[0165] where:

[0166] λ1: Bounding box loss weight, affecting the localization accuracy of the bounding box.

[0167] λ2: Classification loss weight, which affects the classification prediction accuracy.

[0168] λ3: Confidence loss weight, which affects the prediction of the target existence probability of the bounding box.

[0169] Coefficients are assigned to the feature layer distillation loss and the output layer distillation loss for parameter tuning and optimization, as shown in the following formula:

[0170] L total = α·L feature + β·L output ,

[0171] where L feature is the feature layer distillation loss, which measures the difference between the student model and the teacher model in the intermediate feature layer; L output is the output layer distillation loss, including the bounding box distillation loss, the classification distillation loss, and the confidence distillation loss; α and β are adjustment coefficients used to balance their impacts during training.

[0172] In the specific implementation manner of the present invention, the dataset used is from the carotid artery ultrasound images of 55 patients collected by a cooperative medical institution. On the premise of consulting relevant medical experts, medical workers were organized to screen and annotate the images. Finally, 2,606 groups of datasets were obtained. Each group of images in the dataset contains an original ultrasound image and an annotated image with annotation information. Among the images in the dataset, they can be divided into two categories: transverse carotid artery data and longitudinal carotid artery data. According to the task requirements of object detection, they can be further divided into four tasks, namely transverse image blood vessel detection, transverse image plaque detection, longitudinal image intima-media detection, and longitudinal image plaque detection. Figure 10 It is the result diagram of carotid artery transverse image blood vessel detection. Figure 11 It is the result diagram of carotid artery transverse image plaque detection. Figure 12 It is the result diagram of carotid artery longitudinal image intima-media detection. Figure 13 It is the result diagram of carotid artery longitudinal image plaque detection.

[0173] Experimental related to model improvement in Example 1

[0174] The results of the experiment are shown in Table 1 below. Compared with the baseline model, the C3-Res2Block module in the improved model has increased the mAP (mean average precision, the larger the value, the stronger the model detection ability) of the model by 2.8 points, reduced the number of parameters by 0.12M, and reduced the computational volume by 0.4 GFLOPS. The present invention believes that the reason is that its Bottle2neck module performs a split operation on the feature map before convolution, reducing the number of input channels of the convolution, thereby reducing the number of parameters and computational volume of the convolution. At the same time, by adding the results of each sliced convolution to the next unsliced slice, feature fusion between different slices is achieved, improving the generalization performance of the model. In the ASPP-U structure, compared with the baseline model, it has improved by 4.5 points and 3.9 points in accuracy and mAP respectively. However, there are also some problems. Its number of parameters has increased a lot compared with the original baseline model. At the same time, it can be found that the increase in its computational volume is much smaller than the increase in the number of parameters. Therefore, it can be speculated that there are many redundant parameters in this module, and these parameters can be processed through model pruning operations. On the premise that the number of parameters and computational volume of V5_AFPN have both decreased significantly compared with the original model, the recall rate is ensured to be stable. In addition, it has improved by 2.5 points and 2.0 points in accuracy and mAP respectively, significantly improving the performance of the model. These three improved networks are comprehensively used and named plaque-yolov5. Compared with the original baseline model, plaque-yolov5 has significantly improved in accuracy, recall rate, and mAP. The increase in its number of parameters and computational volume, as mentioned above, mainly comes from the ASPP-U module, which can be solved through subsequent pruning operations.

[0175] Table 1 Experimental Results Table of Model Structure Improvement

[0176]

[0177] Example 2 Experiments Related to Loss Function Improvement

[0178] Regarding the improvement of the CIOU loss function for YOLOv5 proposed above, in this embodiment, relevant experiments were conducted on the v1, v2, and v3 versions of the proposed Wise-SIOU. Since it has been determined in this article that the improved plaque-YOLOv5 is used as the model for carotid lesion detection, the network used in this experiment is plaque-YOLOv5, and the baseline loss function is the basic CIOU loss function. The experimental results are shown in Table 2 below. It can be seen from the results in the table that compared with the original CIOU loss function, the performance of the v1, v2, and v3 versions of Wise-SIOU has all improved. Especially for the v3 version, when achieving the balance of accuracy and recall rate, the mAP standard has also been significantly improved. Therefore, the v3 version of Wise-SIOU is finally adopted in this invention to replace the original CIOU loss function.

[0179] Table 2 Comparison Experimental Results Table of Loss Function Improvement

[0180]

[0181] Example 3 Relevant Experiments on Model Distillation

[0182] Before model distillation, Pavlo [2] proposed a pruning (GRAD) strategy based on channel weight gradient to prune the original network, achieving a 3-fold acceleration. During distillation, the teacher model selected in this embodiment is YOLOv5m, and the parameters of the teacher model and the student model to be distilled are shown in Table 3 below.

[0183] Table 3 Comparison Table of Student Model and Teacher Model Parameters

[0184]

[0185] When performing feature-level distillation, regarding the selection of the distillation layer, there are two methods: selecting the feature maps of the three dimensions output by the backbone and selecting the feature maps of the three dimensions output by the neck. For these two distillation layer selection strategies, corresponding comparative experiments are also carried out in this embodiment. In the comparative experiment, both output layer distillation and feature layer distillation adopt the distillation strategy proposed above, and the difference only lies in the different selected distillation layers. The results of the experiment are shown in Table 4 below. It can be found from the experimental results that both can significantly improve the accuracy of the model. However, comparatively speaking, the accuracy of distillation effect by selecting the output of the backbone layer is 0.8 percentage points higher than that of distillation by selecting the output of the neck layer. The output of the neck layer is only separated from the final output by one detection head layer. In the model of the present invention, the main function of the head layer during training is to convert the output of the neck layer of each scale into the specified output format, and it does not contain more processing of the feature layer information inside, so there will be more information redundancy between it and the final output distillation layer, thus affecting the final distillation effect. While using the output feature map of the backbone layer as the distillation feature layer will not have this problem. On the one hand, it can be regarded as the distillation of the output result of the backbone layer. On the other hand, it can also be regarded as approximating the input data of the neck layers of the teacher network and the student network. Therefore, combined with the subsequent output distillation, it can also produce a good distillation effect on the neck layer, thus ultimately improving the distillation effect of the entire model.

[0186] Table 4 Comparison of experimental results of different feature distillation layer selections

[0187]

[0188] To show the superiority of the method of the present invention over the traditional method, a comparative experiment is carried out between the method of the present invention and the traditional method. In the experiment, the present invention respectively selects two distillation strategies of using L1-loss and L2-loss for the output layer. The results of the experiment are shown in Table 5 below. The method of the present invention has a relatively significant improvement in accuracy compared with these two traditional model distillation methods.

[0189] Table 5 Comparison of experimental results with the effects of traditional distillation methods

[0190]

[0191] To visualize the distillation effect, the present invention selects the outputs of the neck layers of the teacher network, the student network before distillation, and the student network after distillation, and selects the prediction boxes with the highest confidence for each category among them. And for the prediction, gradient heat maps of the three networks are respectively generated. The generated gradient heat maps are as Figure 14 shown.

[0192] It is not difficult to find from the results in the heat map that for the two categories, a distilled model using a distillation strategy that combines feature-level distillation and output-level distillation. Compared with the model before distillation, it has learned from the teacher network the hot information that the teacher network pays attention to but the student network does not, thus indicating that the distillation strategy used in the present invention can help the student network better learn the knowledge of the teacher network.

[0193] After the above distillation operation, the present invention has achieved improving the performance of the student network by learning the knowledge of the teacher network. To show the advantages of the distilled model compared with traditional lightweight models, the present invention respectively selects several relatively common lightweight backbones to replace the original backbone of yolov5, and compares it with the original yolov5n and the plaque-yolov5 after pruning and distillation through experiments. The results of the experiment are shown in Table 6 below.

[0194] Analyzing the comparison table of the experimental results, it can be found that compared with mobilenetv3-yolov5, which is the best-performing in common lightweight models, the model obtained by pruning and distilling the plaque-yolov5 of the present invention has an increase of 0.154 in mAP50 val an reduction of 0.41M in the number of parameters, and a reduction of 0.2 GFLOPs in the amount of computation. The model proposed by the present invention has achieved better performance than mobiletv3-yolov5 in terms of accuracy, speed and the number of parameters.

[0195] Table 6 Comparison table of experimental results between the distilled network and the traditional lightweight network

[0196]

[0197] In view of the problem that the accuracy of the selected baseline model is not high enough, the present invention proposes a series of improvement schemes for its network structure, including improving the key module of feature extraction in its original backbone, improving the original pooling layer structure and improving the neck layer of the original feature fusion, so that the accuracy of the model is improved; in view of the shortcomings of the original loss function, the present invention proposes a loss function improvement scheme called Wise-SIOU to replace the original CIOU loss function, and its three versions have improved accuracy compared with the original loss function; the present invention proposes a distillation strategy that comprehensively considers the feature layer and the output layer, by calculating the KL divergence between the feature layer information of the student network and the teacher network as the feature layer distillation loss; and taking the output of the teacher network as a soft label, calculating the distillation positioning, confidence and classification loss as the output layer distillation loss, and comprehensively considering these two types of losses to achieve the distillation of the model.

[0198] References

[0199] [1]Lin TY, Goyal P, Girshick R, et al.Focal loss for dense object detection[C] / / Proceedings of the IEEE international conference on computervision.2017:2980-2988.

[0200] [2]Molchanov P,Mallya A,Tyree S,et al.Importance estimation forneural network pruning[C] / / Proceedings of the IEEE / CVF conference on computervision and pattern recognition.2019:11264-11272.

[0201] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the present invention, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.

Claims

1. An improved lightweight system for detecting tissue ultrasound images, characterized in that, The described detection lightweight system includes: an improved yolov5n network model, and the improved yolov5n network model includes a backbone layer, a neck layer, and a head layer; In the backbone layer and the neck layer, a C3-Res2Block module containing a Bottle2neck structure is used. A short-cut branch is used in the SPPF of the backbone layer, and the detection effect of the detection lightweight system is improved through a V5_AFPN neck structure network containing an ASFF module.

2. The detection lightweight system according to claim 1, characterized in that, The backbone layer is used to extract the features of the input image and includes a Conv module, a C3-Res2Block module, and an SPPF module; The Conv module includes a convolutional layer, a BN layer, and an activation function. The C3-Res2Block module is used to adaptively aggregate the feature maps obtained by Conv, and the SPPF module obtains spatial information through the weighted fusion of global features and local features.

3. The detection lightweight system according to claim 2, characterized in that, The C3-Res2Block module contains two branches. One branch contains a Conv module and a Bottle2neck module, and the other branch only contains a Conv module. After the outputs of the two branches are connected and then passed through a Conv module, the output of the C3-Res2Block module is realized; The Bottle2neck module contains an upsampling convolutional layer, a feature map partitioning layer, a fusion convolutional layer, a splicing layer, and a downsampling convolutional layer; The upsampling convolutional layer increases the number of channels of the input feature map, enhances the representation ability of the feature map, and captures more information; The feature map partitioning layer equally partitions the upsampled feature map by channels for subsequent per-channel feature fusion and processing; The fusion convolutional layer integrates the information of different channels through feature fusion and convolutional operations to generate new high-dimensional features; The splicing layer splices multiple convolved feature maps, integrates multi-channel features, and enhances the feature expression ability; The downsampling convolutional layer uses a 1×1 convolution to reduce the number of channels of the spliced feature map back to the original level, reducing the computational burden and controlling the model complexity.

4. The detection lightweight system according to claim 1, characterized in that, In the Bottle2neck module, through residual connection, the input feature map and the downsampled feature map are added channel by channel to prevent gradient disappearance and promote information flow during the training process.

5. The detection lightweight system according to claim 1, wherein, The SPPF module contains convolutional layers with multiple convolutional parameters. After convolving the feature map respectively, the feature sets obtained by multiple convolutions and the input feature map are connected, and then convolution is used for downsampling to obtain the output feature map.

6. The detection lightweight system according to claim 1, wherein The neck layer uses a V5_AFPN neck structure network containing an ASFF module and includes multiple ASFF_2 and ASFF_3 modules; Among them, the ASFF_2 module includes: a Dowmsample and Upsample module for adjusting dimensions and sizes, multiple convolutional modules, a Concat aggregation module, a Softmax module, and a weight module; The ASFF_3 module processes the input by pairwise calling of ASFF_2 to optimize the dependency relationship and interaction between the inputs.

7. A lightweight detection method for tissue ultrasound images, characterized in that, The lightweight detection method includes the following steps: Step 1, collect the tissue ultrasound images to be detected; Step 2, construct an improved yolov5n network model for image detection and perform training optimization; In the backbone layer and neck layer of the improved yolov5n network model, the C3-Res2Block module containing the Bottle2neck structure is used. In the SPPF of the backbone layer, a short-cut branch is used, and the detection effect of the detection lightweight system is improved through the V5_AFPN neck structure network containing the ASFF module; Step 3, input the tissue ultrasound images to be detected in Step 1 into the constructed improved yolov5n network model to realize the detection of tissue ultrasound images.

8. The lightweight detection method according to claim 7, characterized in that, In Step 2, when training the yolov5n network model, the Wise-SIOU loss function is introduced. The Wise-SIOU loss function adds three cost functions on the basis of IOU, including the angle cost Λ, the distance cost Δ, and the shape cost Ω; The calculation of the angle cost Λ is shown as follows: Where, b gt Cx , b gt Cy is the center coordinate of the ground truth box, b Cx , b Cy is the center coordinate of the predicted box, C h is the height difference between the centers of the ground truth box and the predicted box, and α is the angle between the line connecting the centers of the ground truth box and the predicted box and its projection in the horizontal direction; The calculation of the distance cost Δ is shown as follows: Where, Δ represents the total loss or error, ρ x and ρ y are the squared normalized errors in the x and y directions, C h is the height difference between the centers of the ground truth box and the predicted box, C w is the width difference between the centers of the ground truth box and the predicted box, Υ is a numerical adjustment factor depending on Λ; The calculation of the shape cost Ω is shown as follows: Among them, w and h represent the width and height of the current rectangle, w gt and h gt represent the width and height of the target rectangle or the reference rectangle, ω w represents the relative measure of the width difference, ω h represents the relative measure of the height difference, and θ is an adjustable variable; The calculation method of the final SIOU is shown as follows:

9. The lightweight detection method according to claim 8, wherein By setting the balance scheme of high-quality and low-quality samples of Wise-SIOU, the training effect of the improved yolov5n network model is adjusted and optimized; The loss function is expressed as follows: Where, L WSIOUv1 =R WSIOU *L SIOU , Among them, α is the scaling factor and β is the adjustment coefficient; Wg and Hg are the length and width of the minimum bounding rectangle of the two boxes, and x gt and y gt are the horizontal and vertical coordinates of the target rectangle, R WSIOU is the attention adjustment factor; L SIOU represents the loss function of SIOU, and δ is used to measure the quality of the anchor box.

10. The lightweight detection method according to claim 7, wherein Use a distillation strategy that comprehensively considers the feature layer and the output layer to optimize the model performance; During the feature layer distillation process, the activation map in each channel is normalized to make the foreground of the feature map more prominent. Subsequently, the KL divergence between the normalized channel activation maps of the teacher and student networks is reduced to guide the student network to pay more attention to the regions with large normalized activation values in each channel; During the distillation process of the output layer, the distillation of the output layer is transformed into the distillation for bounding boxes, classification, and confidence, and the loss functions are respectively Where, Among them, L D box is the bounding box loss, and CIOU is an index that measures the degree of overlap between the predicted bounding box and the ground truth bounding box. is for calculating the predicted bounding box, and is the ground truth bounding box; Among them, N is the batch size, M is the height of the feature map, X is the width of the feature map, Y is the depth or number of classes of the feature map, and Λ i,j,k is the weight, which is used to adjust the contribution of each class in the loss calculation, and P T i,j is the predicted probability of the k-th class of the teacher network at the position (i, j), and P S i,j is the predicted probability of the student network at the same position and class, and log(P S i,j ) is the log-likelihood when the class exists; The confidence distillation loss For the confidence of the bounding boxes, it is optimized by comparing the outputs of the teacher and student networks and using the BCE loss. During the output layer distillation process, combining the bounding box distillation loss, the classification distillation loss, and the confidence distillation loss, the final output layer distillation loss is obtained, which is expressed as follows: L output = λ1·L bbox + λ2·L cls + λ3·L obj , Among them, L bbox represents the bounding box distillation loss, L cls represents the classification distillation loss, L obj represents the confidence distillation loss, λ1 represents the bounding box loss weight, λ2 represents the classification loss weight, and λ3 represents the confidence loss weight; Coefficients are assigned to the feature layer distillation loss and the output layer distillation loss for parameter tuning and optimization, as shown in the following formula: L total = α·L feature + β·L output , Among them, \(L\) feature is the feature layer distillation loss, which measures the difference between the student model and the teacher model in the intermediate feature layer; \(L\) output is the output layer distillation loss, including the bounding box distillation loss, the classification distillation loss, and the confidence distillation loss; \(\alpha\) and \(\beta\) are adjustment coefficients used to balance their impacts during training.