Lightweight neural network based highway debris detection method

CN122473617BActive Publication Date: 2026-09-11CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610941864.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-11
Estimated Expiration
2046-06-29

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种轻量化神经网络的高速公路抛洒物检测方法,解决现有技术中对微弱小目标特征的误抑制、几何畸变补偿不精准与计算冗余之间的矛盾,以及模型整体计算效率难以满足路侧边缘实时部署要求的问题,实现在不增加计算开销的前提下,对高速场景下微小、模糊、形变抛洒物的微弱特征保留、几何形变自适应补偿与模型轻量化之间的高效协同

Benefits of technology

[0044]1.本发明通过RMS-CBAM中残差通道注意力与动态自适应多尺度空间注意力的协同设计,有效避免了传统CBAM对微弱小目标特征的误抑制,解决了深度网络多次下采样过程中小目标特征被背景噪声淹没的核心难题;实验表明,本发明不仅在自建高速公路抛洒物数据集上的检测精度较YOLOv8n基线有所提升,并且召回率(Recall)稳定达到0.999以上,极大降低了高速公路二次事故诱发的漏检风险。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122473617B_ABST
    Figure CN122473617B_ABST
Patent Text Reader

Abstract

The application discloses a highway litter detection method of a lightweight neural network, and the method comprises the following steps: pre-processing and extracting basic features of a to-be-detected image; performing phased enhancement and fusion on a feature map by embedding a residual multi-scale convolution attention module in a key path of a backbone network; for a minimum scale feature map, a gating signal is generated according to the edge blur degree, so as to dynamically trigger a residual deformable convolution module acting on a bounding box regression branch, and adaptive deformation compensation is realized; a small target focal loss function containing a course IoU regression loss, a small target perception weight distribution, a focal classification loss and a softened distribution focal loss is used to optimize network training; finally, the input image is inferred by using the trained model, and a detection result is output. While controlling the calculation overhead, the application realizes efficient cooperation among the three aspects of weak feature reservation, geometric deformation adaptive compensation and model lightweight of the litter in a high-speed scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and intelligent transportation technology, and in particular relates to a lightweight neural network method for detecting spilled materials on highways. Background Technology

[0002] With the development of intelligent perception and proactive safety early warning technologies for highways, real-time detection of small road debris (such as aluminum cans, plastic bags, and small stones) has become a key research topic. These targets, due to their small size and difficulty in detection, can easily trigger emergency avoidance or collisions when vehicles are traveling at high speeds, leading to secondary accidents and seriously threatening driving safety. Studies show that approximately 15% of traffic accidents globally each year are related to road debris. Therefore, achieving robust and low-latency detection of small debris is of great significance for improving roadside perception capabilities and supporting autonomous driving and proactive hazard avoidance.

[0003] However, detection in this scenario faces significant challenges: the target scale is extremely small, typically less than 32×32 pixels in the image, and its features are weak, easily drowned out by background noise during multi-layer downsampling in deep networks; simultaneously, motion blur and non-rigid deformation caused by high-speed motion can damage the target structure, increasing the difficulty of classification and localization. Currently, detection methods based on mainstream frameworks such as YOLO and Faster R-CNN perform well in general scenarios, but still have significant shortcomings in such extremely small target scenarios. On the one hand, common attention mechanisms (such as CBAM) usually enhance features by suppressing weak response regions, but in small target scenarios, weak features are easily misclassified as noise and suppressed, leading to missed detections. On the other hand, traditional convolutional receptive fields are fixed and difficult to adapt to severe geometric distortions; deformable convolutions can improve deformation modeling capabilities, but global deployment introduces computational redundancy and irrelevant spatial perturbations, affecting classification accuracy. In addition, high-precision models (such as Transformer-based detectors) have high computational costs, making it difficult to meet the low power consumption and real-time response requirements of roadside edge devices; while lightweight models usually have limited performance in small, blurry, high-speed composite degradation scenarios.

[0004] Although existing technologies have improved model performance to some extent by introducing attention mechanisms and deformable convolutions, the following prominent problems still exist: First, static or global attention mechanisms are prone to accidentally suppressing weak features of small targets; second, the global deployment of deformable convolutions introduces irrelevant spatial interference in classification tasks and has low computational efficiency; third, existing architectures struggle to achieve a balance between the efficiency of weak feature preservation, accurate geometric deformation compensation, and real-time inference under edge computing constraints. Summary of the Invention

[0005] The purpose of this invention is to provide a lightweight neural network-based method for detecting debris on highways, which solves the contradiction between false suppression of weak and small target features, inaccurate geometric distortion compensation and computational redundancy in existing technologies, as well as the problem that the overall computational efficiency of the model is difficult to meet the requirements of real-time deployment at roadside edges. It achieves efficient synergy between preserving weak features of small, blurry, and deformed debris in highway scenes, adaptive compensation for geometric distortion and lightweight model, without increasing computational overhead.

[0006] The technical solution adopted in this invention is a lightweight neural network-based method for detecting highway debris, comprising the following steps:

[0007] Step S1: Input the preprocessed image into the backbone network, and perform preliminary downsampling and basic feature mapping through continuous convolutional layers to obtain the basic feature map;

[0008] Step S2: The basic feature map continues to propagate forward in the backbone network. The feature map is enhanced in stages by embedding a residual multi-scale convolutional attention RMS-CBAM module on the critical path of the backbone network to obtain an enhanced feature map.

[0009] Step S3: Perform multi-scale feature fusion on the enhanced feature map; input the smallest scale feature map in the enhanced feature map into the bounding box regression branch, generate a gating signal based on the edge ambiguity of the smallest scale feature map, and dynamically determine whether to activate the conditionally triggered residual deformable convolution module acting on the branch based on the gating signal, and output the deformed feature map.

[0010] Step S4: Calculate the detection head based on the deformed feature map, and optimize the network training process using a small target focus loss function that includes curriculum IoU regression loss, small target perception weight allocation, focus classification loss, and softened distribution focus loss;

[0011] Step S5: Use the trained model to infer the input image and output the detection results.

[0012] Furthermore, in step S2, in the P2 and P3 stages of the backbone network corresponding to the high-resolution feature maps, the first two standard C2f modules in the original network are replaced with C2f-RMS-CBAM modules that integrate RMS-CBAM modules; before the features enter the deep spatial pyramid pooling fast module, an additional global RMS-CBAM module is embedded.

[0013] Furthermore, the RMS-CBAM module includes a residual channel attention submodule and a dynamic adaptive multi-scale spatial attention submodule.

[0014] Furthermore, the residual channel attention submodule specifically includes:

[0015] Global average pooling and global max pooling are performed simultaneously on the input feature map. The two pooling results are then input into a shared multilayer perceptron (MLP) for processing. The outputs of the two MLPs are summed and processed by a sigmoid activation function to generate a channel attention map. The channel attention map is then added to a value of 1 to obtain the scaling factor for each channel. The scaling factor is then multiplied channel by channel with the input feature map to obtain a channel-weighted feature map.

[0016] Furthermore, the dynamic adaptive multi-scale spatial attention submodule specifically includes:

[0017] The channel-weighted feature map output by the residual channel attention submodule is processed in parallel through three convolutional layers with kernel sizes of 3×3, 5×5, and 7×7 to extract features and obtain the feature map. , , ; the feature map , , The concatenated feature map is obtained by stitching along the channel dimension and then input into the scale attention module to generate a scale weight vector. The feature map is then analyzed using the scale weight vector. , , Weighted summation is performed to obtain the fused dynamic features; the dynamic features are processed through a 1×1 convolutional layer and a sigmoid activation function to generate a spatial attention map; the spatial attention map is multiplied element-wise with the channel-weighted feature map to output the final enhanced features;

[0018] The scale attention module specifically includes: performing global average pooling on the stitched feature map, and processing it through a multilayer perceptron and a Softmax function to generate scale weight vectors corresponding to the three scales.

[0019] Further, in step S3, generating a gating signal based on the edge ambiguity of the feature map specifically includes:

[0020] Edge ambiguity is quantified by performing edge detection on the minimum scale feature map and calculating its variance, and the edge ambiguity is used as the input of the gating unit. The input of the gating unit is processed by a multilayer perceptron and then compressed to the [0,1] interval by a sigmoid activation function to generate a gating signal.

[0021] Further, in step S3, dynamically determining whether to activate the conditionally triggered residual deformable convolution module acting on that branch based on the gating signal specifically includes:

[0022] When the gate signal Greater than the preset threshold At that time, the residual deformable convolution module is activated, and features are output according to the following formula. :

[0023]

[0024] in, The input features are those flowing through the bounding box regression branch corresponding to the smallest scale feature map in the enhanced feature map. For input features The output after performing a standard deformable convolution operation;

[0025] When the gate signal Less than or equal to the preset threshold If it is determined that no deformation compensation is needed, the feature is directly output. .

[0026] Furthermore, the residual deformable convolution module performs deformable convolution... The output is formed by adding the input features in the form of residual connections, which is related to the input features. The processing output is ;in, It is a learnable scalar parameter.

[0027] Furthermore, in step S4,

[0028] The total loss of the small target focus loss function The loss is a weighted sum of the course-based IoU regression loss, the focus classification loss after small target perception weight allocation, and the softened distribution focus loss. Its calculation formula is as follows:

[0029]

[0030] in, For course-based IoU regression loss, Focus classification loss after weighting small target perception. To soften the distribution focal loss, , , These are the balance coefficients for regression loss, focus classification loss, and softened distribution focus loss, respectively. Perceive weights for small targets.

[0031] Furthermore, the course-based IoU regression loss The calculation method is as follows:

[0032]

[0033] in, 'b' represents the bounding box predicted by the model, 'b' represents the corresponding ground truth bounding box, and 't' represents the current training epoch. This is the threshold for the preheating round. For generalized intersection and comparison of losses, For complete intersection and union, the loss is compared;

[0034] The small target perception weight The allocation method is as follows:

[0035]

[0036] in, Let be the area of ​​the ground truth bounding box corresponding to the i-th sample. The basic weights of the samples;

[0037] The focus classification loss The calculation formula is:

[0038]

[0039] in, It is the model's predicted probability of the true class. This is a factor used to balance the weights of positive and negative samples; These are adjustable focusing parameters;

[0040] The softened distribution focal loss By introducing a temperature parameter into the softmax function for the distribution focus loss. The softened softmax function is defined as follows:

[0041]

[0042] in, The representation model for the th The original logic value predicted by each discrete coordinate interval; Represents an exponential function; These are the original logic values ​​predicted by the model, corresponding to the discrete intervals of the coordinates.

[0043] The beneficial effects of this invention are:

[0044] 1. This invention effectively avoids the false suppression of weak and small target features by the traditional CBAM through the collaborative design of residual channel attention and dynamic adaptive multi-scale spatial attention in RMS-CBAM, and solves the core problem of small target features being submerged by background noise during multiple downsampling processes in deep networks. Experiments show that this invention not only improves the detection accuracy on the self-built highway spill data set compared with the YOLOv8n baseline, but also achieves a stable recall rate of over 0.999, greatly reducing the risk of missed detections induced by secondary accidents on highways.

[0045] 2. This invention employs a conditionally triggered residual deformable convolution strategy, deploying residual DCN only locally in the P3 regression branch. This accurately compensates for boundary distortion and positioning deviation caused by high-speed motion, while avoiding irrelevant spatial disturbances introduced by the classification branch. Combined with the dynamic adaptive multi-scale spatial attention of RMS-CBAM, the model adaptively adjusts its receptive field, significantly improving the regression accuracy of bounding boxes for non-rigidly deformable projectiles and solving the computational redundancy and semantic interference problems caused by existing global DCN deployments.

[0046] 3. This invention significantly improves performance by embedding optimization modules in stages and using targeted lightweight design, while reducing the number of model parameters by 0.29M and the number of floating-point operations (FLOPs) by 1.8G. This lightweight feature enables it to meet the stringent requirements of roadside edge devices for low power consumption and real-time inference, and reduces the computational overhead by more than 10 times compared to heavy Transformer models (such as RT-DETR), demonstrating excellent potential for engineering implementation.

[0047] 4. This invention is equipped with an STFL loss function, which avoids gradient explosion and overfitting in the early stage of training through four collaborative mechanisms: course-based IoU guidance, small target weight allocation, focus classification and distribution softening, and improves the robustness of the model to fuzzy boundaries and sparse small samples. At the same time, the temperature parameter softening mechanism introduced improves the model's tolerance to boundary uncertainty and enhances the localization robustness in high-speed fuzzy scenes. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a schematic diagram of the system application in an embodiment of the present invention.

[0050] Figure 2 This is a diagram of the overall network architecture of an embodiment of the present invention.

[0051] Figure 3 This is a structural diagram of the residual multi-scale convolutional attention module RMS-CBAM according to an embodiment of the present invention.

[0052] Figure 4 The diagrams show a comparison of attention mechanisms, where (a) is a heatmap of the original YOLOv8n model and (b) is a heatmap of the enhanced YOLOv8n-RSF model. Detailed Implementation

[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0054] like Figures 1-4 As shown in the embodiment, the lightweight neural network method for detecting highway debris includes the following steps:

[0055] Step S1: Image Input and Basic Feature Extraction

[0056] The roadside surveillance images to be detected are first preprocessed to the network's standard input size, 640×640 pixels; subsequently, the image data enters the backbone network of the overall architecture; such as... Figure 2 As shown, the input features are first downsampled and mapped to basic features through two consecutive convolutional layers to extract shallow features such as edges and textures from the image, resulting in a basic feature map. These features form the basis for subsequent deep semantic feature extraction and enhancement.

[0057] Step S2: Dynamic enhancement and preservation of multi-scale features based on RMS-CBAM:

[0058] After the basic feature extraction is completed, the resulting basic feature map continues to propagate forward in the backbone network. In this process, in order to solve the core problem that small target features are lost or suppressed in deep networks due to downsampling and noise interference, this invention embeds a core enhancement module, namely the Residual Multiscale Convolutional Attention (RMS-CBAM) module, in stages on the critical path of the backbone network.

[0059] S21: Module Embedding Strategy and Residual Channel Attention Mechanism

[0060] Based on the propagation characteristics of small target features, this invention makes targeted embeddings at key positions in the backbone network. Specifically, in the P2 and P3 stages corresponding to high-resolution feature maps, the first two standard C2f modules in the original network are replaced with C2f-RMS-CBAM modules that integrate RMS-CBAM modules, aiming to enhance the perception and preservation of small target details in the early stage of feature extraction. Furthermore, before the features enter the deep spatial pyramid pooling fast (SPPF) module, an additional global RMS-CBAM module is embedded to globally enhance the target region at a more abstract semantic level and suppress complex background interference.

[0061] The core of each C2f-RMS-CBAM module is the RMS-CBAM module, whose structure is as follows: Figure 3 As shown; when the input feature map X flows into this module, it first enters the residual channel attention submodule; this submodule performs global average pooling and global max pooling on the input features simultaneously to aggregate spatial information from two different perspectives; the results of these two pooling are fed into a shared multilayer perceptron (MLP) for processing, and after being summarized, the channel attention map is generated by the sigmoid activation function. .

[0062] The calculation process for channel attention is as follows:

[0063]

[0064]

[0065]

[0066] in, represents the Sigmoid activation function; MLP represents a shared multilayer perceptron; AvgPool and MaxPool represent global average pooling and global max pooling operations, respectively; This represents channel-by-channel multiplication; The final scaling factor applied to each channel of the input feature X; This is the feature map after channel attention weighting.

[0067] The key to the above design lies in the introduction of residual connections; a significant drawback of traditional attention mechanisms is that they may compress the weights of weak response channels to near zero, directly leading to the loss of small target features; this invention addresses this by... Adding it to the baseline weight with a value of 1 yields This ensures that the weight gain of any channel in the output feature is not less than 1; therefore, while moderately enhancing the discriminative channel, this mechanism can absolutely preserve the pathway of the original weak features, thereby effectively avoiding false suppression.

[0068] S22: Dynamic Adaptive Multi-Scale Spatial Attention Mechanism

[0069] Features enhanced by channel Enter the dynamic adaptive multi-scale spatial attention submodule; the spatial attention branch of traditional CBAM uses a single 7×7 convolution to generate the attention map, and its fixed receptive field is difficult to adapt to the drastic changes in the scale of projectile targets in high-speed scenes.

[0070] To overcome this limitation, this invention designs a three-branch parallel structure; the three branches use convolutional kernels of three different sizes: 3×3, 5×5, and 7×7, respectively, from the input features... Extracting multi-scale spatial context information yields the corresponding feature maps. , , The calculation for each branch can be uniformly represented as:

[0071]

[0072] in, This represents a feature map extracted by a convolutional kernel of size k×k, which contains spatial contextual information at a specific scale. is the convolution kernel; k is the kernel size.

[0073] To adaptively fuse the aforementioned multi-scale information, the feature map , , First, the data is concatenated along the channel dimension to form a spliced ​​feature map. The concatenated feature map is then fed into a lightweight scale attention module; the purpose of this module is to dynamically evaluate the importance of each scale branch; the specific process is as follows: Global average pooling is performed to compress the spatial dimension and obtain global statistics for each channel. A multilayer perceptron (MLP) is then used to perform a nonlinear transformation and mapping on the pooled features to generate an initial weight vector. Finally, the vector is normalized using a Softmax function to obtain the final scale weight vector. Through this mechanism, the network can automatically adjust the input features based on their scale and fuzziness. , , Assign appropriate fusion weights.

[0074] The process of generating scale weights is as follows:

[0075]

[0076]

[0077] in, This indicates a splicing operation along the channel dimension.

[0078] Based on the obtained scale weight vector Its weight The importance of branches at different scales is correspondingly assigned; the feature maps of each branch are adaptively weighted and fused to obtain the fused dynamic features. The dynamic feature is processed by a 1×1 convolution and a sigmoid activation function to generate the final spatial attention map. The attention map The value for each spatial location indicates the importance of that location in the target task.

[0079] The generation of the spatial attention map and the calculation of the final output are as follows:

[0080]

[0081]

[0082]

[0083] in, This represents the Sigmoid activation function; Represents a 1×1 convolution; ultimately enhances the features. The input feature X is sequentially combined with the channel attention weights. Spatial attention map The result of element-wise multiplication.

[0084] Therefore, this dynamic mechanism enables the network to adjust its dependence on receptive fields of different scales in real time and adaptively according to the actual situation of the input features, thereby significantly improving the model's perception flexibility and feature preservation ability for projectile targets with varying scales and shapes in high-speed scenes. After the feature maps are processed by multiple such modules in the backbone network, the semantic information is continuously deepened. Finally, the SPPF module performs multi-scale feature aggregation and outputs a set of multi-level, enhanced feature maps for use by subsequent network parts.

[0085] Step S3: Multi-scale feature fusion and conditionally triggered geometric deformation compensation:

[0086] The multi-level enhancement features output from the backbone network are fed into the neck network for further fusion; such as Figure 2As shown, the neck network achieves multi-scale feature fusion from top to bottom and bottom to top through operations such as upsampling, concatting, and C2f modules, and finally constructs P3, P4, and P5 feature pyramids with rich semantic and positional information. These feature maps are then fed into the detection head to complete the target classification and bounding box regression tasks in parallel.

[0087] S31: Construction of the gating trigger mechanism:

[0088] To accurately compensate for geometric deformation caused by high-speed motion, while avoiding computational redundancy and semantic interference to the classification branch caused by globally deployed deformable convolutions (DCN), this invention implements a conditional selective residual-DCN strategy in the detection head. This strategy only selectively applies to the bounding box regression branch that processes the smallest scale (P3) feature map, while the classification branch and the detection branches at the P4 and P5 scales remain unchanged by standard convolution. This is based on the fact that the P3 scale feature map contains rich details of small targets, and the regression task is most sensitive to the judgment of geometric deformation.

[0089] When P3 feature map When flowing through the regression branch, the system first generates a gating signal based on the ambiguity estimate of the region's features; where, Represents the set of real numbers. This represents the number of channels in the feature map. Indicates the height of the feature map. The width of the feature map is represented by X. Specifically, the ambiguity is quantified by calculating the edge variance of the input feature X and used as the input to the gating unit. Then, a lightweight multilayer perceptron (MLP) processes the input and compresses it to the [0,1] interval via a sigmoid activation function to generate the final gating signal. The calculation process is as follows:

[0090]

[0091]

[0092] in, This indicates the gating signal input, derived from the input feature map. After edge detection operation Calculation of variance This is used to quantify the edge blurring degree of local regions in the feature map; σ represents the dynamically generated gating value used to assess the likelihood of severe motion deformation in the current region; MLP stands for Lightweight Multilayer Perceptron; σ represents the Sigmoid activation function.

[0093] S32: Design of a residual deformable convolution module:

[0094] Gating signal Used to control whether subsequent deformation compensation operations are activated; only when When the target in the region is significantly deformed, the system determines that the target in the region is significantly deformed and triggers the residual deformable convolution module; otherwise, the path will fall back to the standard convolution operation to avoid introducing unnecessary calculations.

[0095] The preset threshold θ is a hyperparameter used to control the gating trigger sensitivity, and its value range is usually set to 0.3 to 0.7. In this embodiment of the invention, experimental verification on a highway spill dataset shows that setting the threshold θ to 0.5 can achieve a better balance between the accuracy of target detection and the computational efficiency of the model. In actual deployment, this threshold can be adaptively adjusted according to the target movement speed, image blur level and edge device computing power requirements in the actual scene to achieve synergistic optimization of detection robustness and real-time performance.

[0096] The triggered deformable convolutional module, based on standard convolution, predicts a set of learnable spatial sampling offsets through a parallel offset learning branch. This allows the fixed sampling grid of standard convolution to adaptively align with the contour of the deformed target.

[0097] The output of a standard deformable convolution (DCN) is defined as follows:

[0098]

[0099] Where n represents the sampling point index and N represents the total number of sampling points; This indicates the coordinates of the center position of the current convolution window; This represents the fixed relative coordinates of the nth sampling point in the standard convolution. These relative coordinates constitute a predefined, fixed sampling grid. For example, for a 3×3 convolution, N=9; This represents the spatial offset corresponding to the nth sampling point, learned by the network. This represents the convolution weight corresponding to the nth sampling position; This indicates that sampling is performed at a specified location on the feature map X.

[0100] To further stabilize training, this invention adds the output of the deformable convolution to the original input features using residual connections, forming a basic residual deformable convolution module (ResidualDeformConv2d), whose output Y is defined as:

[0101]

[0102] in, Y is a learnable scalar parameter used to dynamically balance the contributions of the original features and the deformable enhancement features; Y represents the output of the residual deformable convolution module.

[0103] S33: Conditional Triggered Output Mechanism:

[0104] By integrating a comprehensive gating mechanism and a residual deformable convolution module, this invention defines a complete conditionally triggered output; in implementation, let This represents the original input features flowing through the P3 regression branch; the output response of the residual deformable convolution module defined in step S32 is... .

[0105] When the gate signal When the system activates the module, it will output the module's output. Fixed parameters in Replace with dynamic gating value This will give you the output of the conditional triggering mechanism: This formula shows that the strength of deformation compensation is related to the degree of ambiguity (from...). (Characteristics) positively correlated.

[0106] when When the system determines that deformation compensation is unnecessary, it directly outputs the original features. In this case, the path is equivalent to a standard convolution (identity mapping), avoiding unnecessary computation.

[0107] Output of conditional triggering mechanism Defined by the following formula:

[0108]

[0109] This design achieves strict on-demand triggering and intensity-controlled geometric deformation compensation.

[0110] The aforementioned conditionally triggered residual deformable convolution strategy ensures effectiveness through a triple mechanism: First, it only operates in regions where high-speed motion leads to severe blurring or positioning errors ( First, it precisely activates geometric deformation compensation; second, it automatically shuts down DCN in clear or large-scale areas, completely avoiding potential interference from irrelevant spatial disturbances to the classification task and unnecessary computational overhead; finally, it uses residual connections and dynamic gating values. Together, they effectively improve the training robustness under unstable offset learning conditions and prevent small target features from degrading during deformation modeling.

[0111] This strategy complements and synergizes with the dynamic adaptive multi-scale attention mechanism of the aforementioned RMS-CBAM module: RMS-CBAM mainly enhances and preserves the features of small targets in the feature extraction stage, while this strategy adaptively compensates for their geometric deformation in the localization and regression stage; the two work together to solve the computational overhead and semantic interference problems caused by traditional global deployment of DCN, significantly improving the localization accuracy of blurred and deformed small targets while keeping the overall floating-point operation volume of the model to a minimum.

[0112] Step S4: End-to-end collaborative optimization based on STFL loss function

[0113] During network training, the predicted results output by the detection head need to be compared with the real labels to optimize network parameters. To address this, this invention designs a Small Object Focus Loss Function (STFL) for end-to-end collaborative supervision. This loss function consists of four collaborative components that work together to address the unique challenges of small object detection in high-speed scenarios.

[0114] S41: Curriculum IoU Loss:

[0115] YOLOv8 baseline models typically use the full intersection-union loss (CIoU) as the bounding box regression loss, which simultaneously considers the overlapping region, center point distance, and aspect ratio. However, for small targets that are severely blurred under high-speed motion, the aspect ratio difference between the predicted box and the ground truth box may be extremely large in the early stages of training. The strong penalty term in CIoU can easily lead to gradient instability, hindering the early convergence of the model.

[0116] To address this issue, the STFL of this invention introduces a course learning strategy; when the training round t is less than a preset warm-up threshold... At this time, a coarse localization guide is adopted using the generalized intersection-union loss (GIoU), which is relatively insensitive to the initial position and is more lenient; when the training round t reaches or exceeds Then, a smooth switch to CIoU loss is applied to achieve more refined regression in later stages. This strategy effectively improves stability during the initial training phase; regression loss... Defined as:

[0117]

[0118] in, represents the bounding box predicted by the model, b represents the corresponding ground truth bounding box, and t represents the current training epoch. The threshold for the preheating rounds used to control loss switching.

[0119] S42: Small-Target-Aware Weighting

[0120] The standard loss function assigns uniform weights to all target samples, failing to consider the sparsity of small targets in the data distribution, resulting in insufficient attention to them during the optimization process. Therefore, STFL introduces small target-aware weights; for the i-th sample, if its target area... If the pixel count is less than 32×32, then a higher weight is assigned to it when calculating the regression loss and classification loss. This significantly increases the contribution of small objectives to the total loss, forcing the optimizer to pay more attention to these difficult samples; the weight allocation formula is as follows:

[0121]

[0122] in, This represents the area of ​​the true bounding box corresponding to the sample. These are the basic weights of the samples.

[0123] S43: Focal Classification Loss

[0124] The original binary cross-entropy loss, in scenarios with extreme imbalance between positive and negative samples, causes training to be dominated by a large number of easily classifiable background samples; STFL replaces this with focus loss, which modulates the loss through a modulation factor. Automatically reduces the loss contribution of easily classified samples, allowing model training to focus on difficult-to-classify samples (such as blurry small targets); focus classification loss. The calculation formula is:

[0125]

[0126] in, It is the model's predicted probability of the true class. This is a factor used to balance the weights of positive and negative samples; These are adjustable focusing parameters.

[0127] S44: Softened Distribution Focal Loss

[0128] YOLOv8 uses Distributed Focus Loss (DFL) to model the probability distribution of the bounding box coordinates through discretization; however, the blurring of small target boundaries often leads to an overly sharp probability distribution in the DFL output, which can easily cause overfitting. Therefore, STFL introduces a temperature parameter into the softmax function of DFL. Softening the model smooths the predicted probability distribution, thereby improving its tolerance to boundary uncertainties and its robustness to localization. The softmax function after softening is defined as follows:

[0129]

[0130] in, The representation model for the th The original logic value predicted by each discrete coordinate interval; Represents an exponential function; The original logits, which are predicted by the model and correspond to the discrete intervals of the coordinates. For temperature parameters, The larger the value, the smoother the output probability distribution; based on this, the softened distribution focus loss is denoted as... .

[0131] S45: Total STFL Loss:

[0132] The total loss of STFL is the weighted sum of the above terms, defined as follows:

[0133]

[0134] in, , , These are the balance coefficients for regression loss, focus classification loss, and softened distribution focus loss, respectively, and their default values ​​are consistent with the YOLOv8 baseline model settings.

[0135] It should be noted that the small target perception weight in this invention This applies only to the bounding box regression and classification losses, not to the softened distribution focus loss. This is because the regression and classification losses directly determine the model's attention to small target samples; by introducing weight enhancement, the gradient contribution of small targets during training can be effectively increased. The softened distribution focus loss, on the other hand, is mainly used for probabilistic modeling of the discrete distribution of the bounding boxes, its core function being to improve the smoothness and uncertainty tolerance of boundary position predictions. Applying additional small target-aware weights to the softened distribution focus loss could easily lead to over-biasing of the coordinate distribution learning, thereby reducing training stability.

[0136] This loss function, along with RMS-CBAM (responsible for feature enhancement and preservation) and conditionally triggered residual deformable convolution (responsible for precise localization), are deeply integrated and work together in the end-to-end training process to drive the model to learn effective representations of small, blurry, and deformable projectiles in high-speed scenes.

[0137] Step S5: Model Inference and Real-Time Detection Output:

[0138] The YOLOv8n-RSF model trained through the above steps can be deployed on roadside edge computing devices. In the inference phase, the input image frame passes through the forward propagation path defined in steps S1 to S3 in sequence: after preprocessing, it is subjected to deep feature enhancement extraction through the backbone network embedded with RMS-CBAM, then multi-scale feature fusion through the neck network, and finally classification and regression prediction are completed in the detection head with embedded conditional triggering mechanism. Thanks to the dynamic lightweight design of RMS-CBAM and the conditional triggering strategy of deformable convolution, the entire network achieves low computational overhead while maintaining high detection accuracy. The original prediction results output by the detection head are post-processed by non-maximum suppression and other methods to obtain the final bounding boxes, categories and confidence scores of all debris in the image, meeting the millisecond-level response requirements of the roadside system for real-time early warning.

[0139] To verify the effectiveness of the proposed YOLOv8n-RSF method and ensure the statistical significance and reproducibility of the evaluation results, this study constructed a self-built dataset for detecting small debris on highways. The original images in this dataset were all collected from real highway roadside monitoring scenarios, fully reproducing the extreme detection conditions under high-speed vehicle driving environments.

[0140] The dataset exhibits the following typical characteristics, placing extremely high demands on the model's robustness and generalization ability: First, the scale of the scattered objects in the dataset is extremely small, typically less than 32×32 pixels in the images; second, due to the high-speed movement of vehicles, the objects often suffer from severe motion blur and non-rigid deformation; third, the objects have diverse appearances, encompassing various common scattered objects such as aluminum cans, plastic bags, shredded paper, and small stones; finally, the image backgrounds are complex, exhibiting low contrast, clutter, and partial occlusion. These combined characteristics significantly increase the detection difficulty of this dataset compared to typical small object subsets in existing mainstream public datasets.

[0141] In the initial phase of the research, this invention conducted experiments based on approximately 800 labeled images. The results showed that on this small dataset, the model was highly susceptible to overfitting, and its generalization performance was difficult to reliably assess. To address the issue of insufficient sample size and improve the model's scene adaptability, this study employed a physically-aware data augmentation strategy. This strategy utilized the OpenCV library to implement a series of augmentation operations, including random rotation, brightness and contrast perturbation, simulated motion blur, and multi-target Mosaic stitching, ultimately expanding the dataset to 3125 images with high quality. All subsequent ablation experiments, module effectiveness comparisons, and performance comparisons with mainstream algorithms were conducted on this augmented dataset, thus ensuring the statistical reliability of the experimental conclusions.

[0142] To intuitively reveal the key impact of data scale on model performance, this study also conducted preliminary comparative experiments on the original small-scale dataset (approximately 800 images). The results are shown in Table 1, further confirming the necessity of data expansion and the basis for the effectiveness of the method of this invention.

[0143] Table 1. Comparison of mainstream object detection algorithms on the original small-scale dataset.

[0144]

[0145] To ensure the robustness of the evaluation results, all experiments in this invention were conducted on an enhanced self-built dataset containing 3,125 images to provide sufficient statistical support.

[0146] To systematically verify the effectiveness of the lightweight neural network-based highway debris detection method (YOLOv8n-RSF) proposed in this invention, this section details the experimental setup, evaluation criteria, and analysis framework.

[0147] The training configuration is as follows: All experiments were implemented using the PyTorch 2.3 deep learning framework, with an NVIDIA RTX 5060 GPU (CUDA 12.8 acceleration). Input images were uniformly scaled to 640×640 pixels, the batch size was set to 16, and the total number of training epochs was 200. The optimization process used the Adam optimizer, with an initial learning rate of 1×10⁻⁶. -3 In the final verification phase, the AdamW optimizer was further employed to achieve optimal convergence.

[0148] The evaluation metrics strictly follow the evaluation standards of general object detection datasets such as MS COCO, including the average accuracy (mAP@0.5) at an IoU threshold of 0.5, the average accuracy (mAP@0.5:0.95) within the IoU threshold range of 0.5 to 0.95 (step size 0.05), and the accuracy for areas smaller than 32. 2 The metrics for small targets at the pixel level include average precision, recall, F1 score, and fitness, an adaptive metric used to comprehensively evaluate model performance. The definitions of these core metrics are as follows:

[0149] Accuracy measures the precision of a model's predictions; that is, the proportion of samples that the model predicts as positive actually being positive.

[0150]

[0151] Recall reflects a model's ability to detect positive samples; that is, the proportion of true positive examples correctly detected by the model.

[0152]

[0153] Wherein, TP represents the number of correctly detected targets; FP represents the background or false alarm targets that are incorrectly detected; and FN represents the real targets that are not detected.

[0154] Mean precision (AP) is a core metric for evaluating the accuracy of an object detection model on a single class. It is the area under the precision-recall curve for that class. The mean mean precision (mAP) is obtained by averaging the APs for all n classes in the dataset, and its calculation formula is as follows:

[0155]

[0156] The F1 score is the harmonic mean of precision and recall, used to comprehensively evaluate the balance between the two:

[0157]

[0158] Fitness is a custom comprehensive evaluation metric used in model training to guide hyperparameter optimization. Its computational weights emphasize a more rigorous mAP@0.5:0.95.

[0159]

[0160] Furthermore, the computational complexity of the model is quantified by the number of parameters and the number of floating-point operations (FLOPs). Models with lower parameter counts and FLOPs require fewer computing resources, making them more suitable for deployment on edge devices with limited computing resources.

[0161] To systematically evaluate the actual contributions of the RMS-CBAM module, the Conditional Selective Residual-DCN strategy, and the various components within the YOLOv8n-RSF framework jointly constructed by the two, this invention employs ablation experiments. All ablation experiments were conducted under a completely uniform experimental configuration: using the original YOLOv8n model as a baseline, and employing its original VFL+BCE classification loss and CIoU regression loss functions, along with the Adam optimizer, on an enhanced dataset of 3,125 self-built highway spill data. This setup aims to eliminate the impact of differences in loss functions or optimizers on module performance evaluation.

[0162] (1) Ablation verification of RMS-CBAM module:

[0163] First, a specific comparative experiment was conducted on the core component RMS-CBAM and the traditional attention mechanism, and the results are shown in Table 2.

[0164] Table 2 Results of RMS-CBAM ablation experiments

[0165]

[0166] Experiments show that, compared to the baseline model YOLOv8n, the static multi-scale fusion RMS-CBAM (static) improves the overall performance index mAP@0.5:0.95 by 1.7%. Further introducing a dynamic adaptive scale attention mechanism, the dynamic RMS-CBAM version further improves the small target detection accuracy (mAP@Small) by 0.8% and the overall performance index Fitness by 0.0032, while reducing computational overhead (FLOPs) by 0.2G. In contrast, directly embedding standard CBAM or C2f-CBAM modules leads to a significant decrease in mAP@0.5:0.95 (decreases of 0.091 and 0.088, respectively). These results demonstrate that the design of introducing residual connections and employing a dynamic adaptive multi-scale spatial attention branch in channel attention effectively prevents the missuppression of weak small target features by traditional attention mechanisms, achieving more robust feature representation while reducing computational cost.

[0167] (2) Verification of the fusion and ablation of the main modules:

[0168] After confirming the effectiveness of RMS-CBAM, the synergistic effect of the Conditional Selective Residual-DCN strategy and the complete RSF fusion framework was further examined, and the results are shown in Table 3.

[0169] Table 3 Ablation Experiment Results of Main Modules

[0170]

[0171] Experiments show that the Conditional Selective Residual-DCN strategy, even when deployed only in the regression branch at the P3 scale, can independently improve mAP@0.5:0.95 by 0.6% and small target mAP by 1.3%, while significantly reducing model complexity (FLOPs decreased by 8.6%, and parameter count decreased by 1.7%). This verifies its effectiveness in compensating for geometric deformation and positioning errors caused by high-speed motion. The RMS-CBAM (Dynamic) module, when used alone, shows the most significant improvement in overall performance, increasing mAP@0.5:0.95 by 2.0%, achieving a maximum Fitness of 0.8083, and a recall of 1.000. When the two are integrated to form a complete YOLOv8n-RSF framework, the model achieves the best overall performance with small target mAP of 0.779 and Fitness of 0.7995, while having the lowest computational cost (FLOPs=6.3G, Params=2.72M). All indicators are better than the application of a single module, which fully verifies that there is a significant complementary and synergistic effect between the feature enhancement provided by dynamic RMS-CBAM and the geometric modeling implemented by conditionally triggered Residual-DCN.

[0172] (3) Ablation verification of loss function and optimizer:

[0173] Under the premise of fixing the YOLOv8n-RSF network structure, a systematic ablation study was conducted on the classification and regression loss functions and the optimizer selection. The results are shown in Table 4.

[0174] Table 4. Ablation experiments of loss function and optimizer (based on YOLOv8n-RSF)

[0175]

[0176] Experimental results show that in this extreme small target scenario, various mainstream IoU loss variants failed to outperform the default YOLOv8 loss function, instead causing mAP@0.5:0.95 to decrease by 1.3% to 3.2%. However, the Small Target Focus Loss Function (STFL) proposed in this invention, under the Adam optimizer, improved mAP@0.5:0.95 by 0.2% and small target mAP by 0.9%. Further application of the AdamW optimizer resulted in optimal model performance, with a Fitness of 0.8042, mAP@0.5:0.95 of 0.783, and a stable recall of 1.000, completely eliminating the missed detection problem for small targets. These data fully validate the effectiveness of the STFL loss function through the four collaborative mechanisms of curriculum-based IoU guidance, small target weight allocation, focus classification, and distribution softening.

[0177] Through the above three sets of ablation experiments, this invention clearly demonstrates the independent and synergistic technical contributions of the dynamic RMS-CBAM module, the ConditionalSelective Residual-DCN strategy, and the STFL loss function, providing sufficient theoretical and experimental basis for the YOLOv8n-RSF model to achieve the optimal balance between high accuracy and high efficiency in the task of detecting small debris on highways.

[0178] To systematically verify the comprehensive performance of the YOLOv8n-RSF method of this invention, a comprehensive comparative analysis was conducted with nine mainstream object detection algorithms covering various architectures under a unified experimental setting and an enhanced self-built dataset (3,125 images). These included classic single-stage detectors (YOLOv5n, YOLOv8n, YOLOv11n), two-stage models (Faster R-CNN), general lightweight architectures (SSD, EfficientDet-D1), Transformer-based frameworks (DETR, RT-DETR), and the latest small-object-specific algorithm published in 2025 (YOLO11n-PRNet). The comparison results are shown in Table 5.

[0179] Table 5 Performance Comparison Results with Mainstream Object Detection Algorithms

[0180]

[0181] The comparative analysis is as follows:

[0182] Performance comparisons with lightweight models demonstrate that the proposed solution exhibits significant advantages in the synergistic optimization of accuracy and efficiency. As shown in Table 5, compared to the baseline model YOLOv8n, YOLOv8n-RSF achieves a 1.6% improvement in mAP@0.5:0.95, a 2.2% improvement in average precision for small targets (mAP@Small), and a recall rate of 1.000, while simultaneously reducing the number of parameters by 0.29M and the number of floating-point operations (FLOPs) by 1.8G. Compared to the latest small-target-specific algorithm YOLO11n-PRNet, the proposed solution achieves a 1.9% improvement in mAP@0.5:0.95, a 4.1% improvement in average precision for small targets, and a 28.1% reduction in computational overhead (FLOPs). These data demonstrate that through the synergistic optimization of dynamic RMS-CBAM, Conditional Selective Residual-DCN, and the STFL loss function, the proposed solution achieves the optimal balance between detection accuracy and computational efficiency within a lightweight framework.

[0183] Comparative analysis with heavy and specialized models further reveals the robustness of this invention in extreme small target scenarios. Experiments show that some heavy models suffer severe performance degradation on this task: the average accuracy of the two-stage detector Faster R-CNN for small targets is only 0.033, and EfficientDet-D1 is even 0.000. Although the Transformer-based models DETR and RT-DETR maintain acceptable recall, their localization accuracy is significantly insufficient, with average accuracy for small targets of only 0.556 and 0.648, respectively. It is worth noting that although the overall evaluation index Fitness of RT-DETR is slightly higher than that of this invention (0.8119 vs 0.8042), its computational cost (FLOPs of 67.91G, parameter count of 42.71M) is about 10 times and 15 times higher than that of this invention, and its average accuracy for small targets (0.648) is much lower than that of this invention (0.787), which cannot meet the stringent requirements of roadside edge devices for low power consumption and real-time response.

[0184] In summary, YOLOv8n-RSF outperforms mainstream methods in all key performance indicators while maintaining extremely low computational complexity (6.3G FLOPs and 2.72M parameters), achieving a balance between high accuracy, high efficiency, and high robustness.

[0185] The interpretability of the model is validated through Gradient-Weighted Class Activation Mapping (Grad-CAM) visualization. For example... Figure 4 As shown in the comparison results, Figure 4 (a) The attention heatmap of the baseline model YOLOv8n is relatively diffuse, with some activation regions distributed in the background; while Figure 4 (b) The attention heatmap of YOLOv8n-RSF in this invention can accurately focus on the geometric boundaries and key texture regions of the projectile, significantly reducing background interference. This visualization result intuitively verifies the effectiveness of the dynamic RMS-CBAM module and the conditionally triggered residual deformable convolution strategy in enhancing the network's feature activation and localization capabilities in small target regions.

[0186] In summary, through the above systematic performance comparison and interpretability analysis, this invention fully verifies the superiority of the YOLOv8n-RSF method in the task of detecting small debris on highways. The proposed solution achieves breakthroughs in core challenges such as feature suppression, geometric deformation compensation, and computational efficiency, providing a complete, efficient, and practical technical solution for roadside intelligent sensing systems while maintaining high recall (stable above 0.999) and real-time performance.

[0187] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0188] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A lightweight neural network-based method for detecting highway debris, characterized in that, Includes the following steps: Step S1: Input the preprocessed image into the backbone network, and perform preliminary downsampling and basic feature mapping through continuous convolutional layers to obtain the basic feature map; Step S2: The basic feature map continues to propagate forward in the backbone network. The feature map is enhanced in stages by embedding a residual multi-scale convolutional attention RMS-CBAM module on the critical path of the backbone network to obtain an enhanced feature map. Step S3: Perform multi-scale feature fusion on the enhanced feature map; The minimum scale feature map in the enhanced feature map is input into the bounding box regression branch. A gating signal is generated based on the edge ambiguity of the minimum scale feature map. Based on the gating signal, it is dynamically determined whether to activate the conditionally triggered residual deformable convolution module acting on the branch, and the deformable feature map after deformation compensation is output. Step S4: Calculate the detection head based on the deformed feature map, and optimize the network training process using a small target focus loss function that includes curriculum IoU regression loss, small target perception weight allocation, focus classification loss, and softened distribution focus loss; Step S5: Use the trained model to infer the input image and output the detection results; The RMS-CBAM module includes a residual channel attention submodule and a dynamic adaptive multi-scale spatial attention submodule. The dynamic adaptive multi-scale spatial attention submodule specifically includes: The channel-weighted feature map output by the residual channel attention submodule is processed in parallel through three convolutional layers with kernel sizes of 3×3, 5×5, and 7×7 to extract features and obtain the feature map. , , ; the feature map , , The concatenated feature map is obtained by stitching along the channel dimension and then input into the scale attention module to generate a scale weight vector. The feature map is then analyzed using the scale weight vector. , , Weighted summation is performed to obtain the fused dynamic features; the dynamic features are processed through a 1×1 convolutional layer and a sigmoid activation function to generate a spatial attention map; the spatial attention map is multiplied element-wise with the channel-weighted feature map to output the final enhanced features; The scale attention module specifically includes: performing global average pooling on the stitched feature map, and processing it through a multilayer perceptron and a Softmax function to generate scale weight vectors corresponding to the three scales.

2. The lightweight neural network-based highway debris detection method according to claim 1, characterized in that, In step S2, In the backbone network, corresponding to the P2 and P3 stages of high-resolution feature maps, the first two standard C2f modules in the original network are replaced with C2f-RMS-CBAM modules that integrate RMS-CBAM modules; before the features enter the deep spatial pyramid pooling fast module, an additional global RMS-CBAM module is embedded.

3. The lightweight neural network-based highway debris detection method according to claim 1, characterized in that, The residual channel attention submodule specifically includes: Global average pooling and global max pooling are performed simultaneously on the input feature map. The two pooling results are then input into a shared multilayer perceptron (MLP) for processing. The outputs of the two MLPs are summed and then processed by a sigmoid activation function to generate a channel attention map. The channel attention map is added to a value of 1 to obtain the scaling factor for each channel. The scaling factor is then multiplied channel by channel with the input feature map to obtain a channel-weighted feature map.

4. The lightweight neural network-based highway debris detection method according to claim 1, characterized in that, In step S3, a gating signal is generated based on the edge ambiguity of the feature map, specifically including: Edge ambiguity is quantified by performing edge detection on the minimum scale feature map and calculating its variance, and the edge ambiguity is used as the input of the gating unit. The input of the gating unit is processed by a multilayer perceptron and then compressed to the [0,1] interval by a sigmoid activation function to generate a gating signal.

5. The method for detecting highway spills using a lightweight neural network according to claim 1, characterized in that, In step S3, the dynamic determination of whether to activate the conditionally triggered residual deformable convolution module acting on the branch based on the gating signal specifically includes: When the gate signal Greater than the preset threshold At that time, the residual deformable convolution module is activated, and features are output according to the following formula. : ; in, The input features are those flowing through the bounding box regression branch corresponding to the smallest scale feature map in the enhanced feature map. For input features The output after performing a standard deformable convolution operation; When the gate signal Less than or equal to the preset threshold If it is determined that no deformation compensation is needed, the feature is directly output. .

6. The method for detecting highway spills using a lightweight neural network according to claim 5, characterized in that, The residual deformable convolution module performs deformable convolution. The output is formed by adding the input features in the form of residual connections, which is related to the input features. The processing output is ;in, It is a learnable scalar parameter.

7. The method for detecting highway spills using a lightweight neural network according to claim 1, characterized in that, In step S4, The total loss of the small target focus loss function The loss is a weighted sum of the course-based IoU regression loss, the focus classification loss after small target perception weight allocation, and the softened distribution focus loss. Its calculation formula is as follows: ; in, For course-based IoU regression loss, Focus classification loss after weighting small target perception. To soften the distribution focal loss, , , These are the balance coefficients for regression loss, focus classification loss, and softened distribution focus loss, respectively. Perceive weights for small targets.

8. The method for detecting highway spills using a lightweight neural network according to claim 7, characterized in that, The course-based IoU regression loss The calculation method is as follows: ; in, 'b' represents the bounding box predicted by the model, 'b' represents the corresponding ground truth bounding box, and 't' represents the current training epoch. This is the threshold for the preheating round. For generalized intersection and comparison of losses, For complete intersection and union, compare the losses; The small target perception weight The allocation method is as follows: ; in, Let be the area of ​​the ground truth bounding box corresponding to the i-th sample. The basic weights of the samples; The focus classification loss The calculation formula is: ; in, It is the model's predicted probability of the true class. This is a factor used to balance the weights of positive and negative samples; These are adjustable focusing parameters; The softened distribution focal loss By introducing a temperature parameter into the softmax function for the distribution focus loss. The softened softmax function is defined as follows: ; in, The representation model for the th The original logic value predicted by each discrete coordinate interval; Represents an exponential function; These are the original logical values ​​predicted by the model, corresponding to the discrete intervals of the coordinates.

Citation Information

Patent Citations

  • Method for detecting and analyzing internal shrinkage cavity defects of die casting based on machine vision

    CN121883472A

  • AI insect condition identification early warning method and system based on multispectral imaging

    CN122116147A