Slag glass detection method based on deep learning
Through deep learning methods, combining multi-scale feature enhancement and dynamically extended depth separation modules, the accuracy and real-time nature of slag glass detection are improved, and the problems of insufficient accuracy and real-time nature in the existing technology are solved, and efficient slag glass detection is achieved.
Patent Information
- Application Number
- CN202510626095.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-15
Smart Images

Figure CN120495633A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a slag glass detection method based on deep learning, belonging to the technical field of industrial slag glass detection. Background Art
[0002] With global population growth and accelerated industrialization, the amount of municipal solid waste generated has increased significantly, making research into recycling technologies particularly important. Slag, a byproduct of waste incineration, contains unburned, recyclable materials such as crushed stone, glass, and ceramics. Its complex composition makes it difficult to efficiently separate these materials using existing technologies. Consequently, large amounts of recyclable glass are forced into landfills, exacerbating resource waste and carbon emissions. To meet the demand for high-purity products, the purity of slag glass must reach a high level of over 99%. Furthermore, to support efficient front-end and back-end production lines, sorting equipment must be capable of processing more than a ton of material per hour. The requirements for high-precision detection and large-scale, real-time processing complicate glass recycling.
[0003] In recent years, deep learning technology has made significant progress in the field of target detection. However, existing target detection algorithms have certain limitations in slag glass detection scenarios. The complex background of slag, and the varying sizes, shapes, and colors of glass targets, especially small glass targets, place high demands on the accuracy and real-time performance of detection algorithms. While two-stage target detection algorithms perform well in terms of detection accuracy, their high computational complexity makes it difficult to meet real-time detection requirements. Single-stage target detection algorithms are currently rarely used in the field of industrial slag glass detection. The model lacks adaptability to complex backgrounds and small targets, making it prone to missed detections and false detections. Summary of the Invention
[0004] In response to the contradiction between high precision and real-time performance in the field of slag glass industrial detection, the present invention provides a slag glass detection method based on deep learning.
[0005] A slag glass detection method based on deep learning of the present invention comprises:
[0006] Collect slag glass images, annotate the slag glass images, and build a data set;
[0007] The slag glass detection model is trained using the dataset. The input of the slag glass detection model is the slag glass image, and the output includes the location and category of the target in the slag glass image.
[0008] The slag glass detection model includes a backbone network, a feature fusion network, and a classification network; the backbone network is used to extract features of different scales from the input image, the feature fusion network fuses the extracted features of different scales, and the fused features of different scales are input into the classification network for position and category detection, and the classification network outputs the position and category of the target;
[0009] The trained slag glass detection model is used to detect slag glass images.
[0010] Preferably, the backbone network includes a convolution module Conv1, a convolution module Conv2, a multi-branch feature fusion structure M_DEDS No. 1, a convolution module Conv3, a multi-branch feature fusion structure M_DEDS No. 2, a convolution module Conv4, a multi-scale feature enhancement module FEM, a convolution module Conv5, a multi-branch feature fusion structure M_DEDS No. 3, a fast spatial pyramid pooling module SPPF and a cross-channel spatial pyramid attention module C2PSA, which are connected in sequence;
[0011] The outputs of the No. 2 multi-branch feature fusion structure M_DEDS, the multi-scale feature enhancement module FEM, and the cross-channel spatial pyramid attention module C2PSA are the extracted features of different scales;
[0012] Each multi-branch feature fusion structure M_DEDS includes convolution module Conv7, convolution module Conv8, convolution module Conv9, dynamic extended depth separable module DEDS No. 1, dynamic extended depth separable module DEDS No. 2, dynamic extended depth separable module DEDS No. 3, dynamic extended depth separable module DEDS No. 4, convolution module Conv10, convolution module Conv9, 11, convolution module Conv12, convolution module Conv13, convolution module Conv14, splicing module No. 1, splicing module No. 2, and splicing module No. 3;
[0013] The input of the convolution module Conv7 is the input of the multi-branch feature fusion structure M_DEDS, and the output of the convolution module Conv7 is simultaneously input to the convolution module Conv8, the convolution module Conv9, and the No. 1 splicing module;
[0014] The output of the convolution module Conv8 is input to the No. 1 dynamically extended depth separable module DEDS, and the output of the No. 1 dynamically extended depth separable module DEDS is input to the No. 2 dynamically extended depth separable module DEDS. The No. 2 splicing module splices the output of the No. 2 dynamically extended depth separable module DEDS and the output of the convolution module Conv9, and the spliced result is input to the convolution module Conv10;
[0015] The output of convolution module Conv10 is simultaneously input to convolution module Conv11, convolution module Conv12, and splicing module No. 1;
[0016] The output of the convolution module Conv11 is input to the dynamic extended depth separable module DEDS No. 3, the output of the dynamic extended depth separable module DEDS No. 3 is input to the dynamic extended depth separable module DEDS No. 4, the splicing module No. 2 splices the output of the dynamic extended depth separable module DEDS No. 4 and the output of the convolution module Conv12, and the spliced result is input to the convolution module Conv13, and the output of the convolution module Conv13 is input to the splicing module No. 1;
[0017] Splicing module No. 1 splices the input, and the spliced result is input to the convolution module Conv14. The output of the convolution module Conv14 is the output of the multi-branch feature fusion structure M_DEDS.
[0018] Preferably, the feature fusion network includes upsampling module No. 1, splicing module No. 4, multi-branch feature fusion structure No. 4 M_DEDS, upsampling module No. 2, splicing module No. 5, multi-branch feature fusion structure No. 5 M_DEDS, convolution module Conv5, splicing module No. 6, multi-branch feature fusion structure No. 6 M_DEDS, convolution module Conv6, splicing module No. 7, multi-branch feature fusion structure No. 7 M_DEDS;
[0019] The output of the cross-channel spatial pyramid attention module C2PSA is input to the upsampling module No. 1, the splicing module No. 4 splices the output of the upsampling module No. 1 and the output of the multi-scale feature enhancement module FEM, and the spliced result is input to the multi-branch feature fusion structure No. 4 M_DEDS, and the output of the multi-branch feature fusion structure M_DEDS No. 4 is input to the upsampling module No. 2, the splicing module No. 5 splices the output of the upsampling module No. 2 and the output of the multi-branch feature fusion structure M_DEDS No. 2, and the spliced result is input to the multi-branch feature fusion structure M_DEDS No. 5, and the multi-branch feature fusion structure M_DEDS No. 5 is input to the multi-branch feature fusion structure M_DEDS No. 5. The output of the feature fusion structure M_DEDS is input to the convolution module Conv5. The No. 6 splicing module splices the output of the convolution module Conv5 and the output of the No. 4 multi-branch feature fusion structure M_DEDS. The spliced result is input to the No. 6 multi-branch feature fusion structure M_DEDS. The output of the No. 6 multi-branch feature fusion structure M_DEDS is input to the convolution module Conv6. The No. 7 splicing module splices the output of the convolution module Conv6 and the output of the cross-channel spatial pyramid attention module C2PSA. The spliced result is input to the No. 7 multi-branch feature fusion structure M_DEDS.
[0020] The outputs of the multi-branch feature fusion structure M_DEDS No. 5, the multi-branch feature fusion structure M_DEDS No. 6, and the multi-branch feature fusion structure M_DEDS No. 7 are fusion features of different scales.
[0021] Preferably, the multi-scale feature enhancement module FEM includes 4 channels, the input features are input into the 4 channels at the same time, and the first channel uses 1×1 standard convolution to generate equivalent features Figure 1 ;
[0022] The second channel uses 1×1 standard convolution, 1×3 standard convolution, 3×1 standard convolution, and 3×3 hole convolution in sequence to generate equivalent features. Figure 2 ;
[0023] The third channel uses 1×1 standard convolution, 3×1 standard convolution, 1×3 standard convolution, and 3×3 hole convolution in sequence to generate equivalent features. Figure 3 ;
[0024] The fourth channel uses 1×1 standard convolution and 3×3 standard convolution in sequence to generate equivalent features. Figure 4 ;
[0025] Equivalent features Figure 1 , equivalent features Figure 2 , equivalent features Figure 3 , equivalent features Figure 4 Through residual connection, the output features of the multi-scale feature enhancement module FEM are obtained.
[0026] Preferably, the dynamically extended depthwise separable module DEDS includes a point-by-point convolution module Conv_pw1, a depthwise separable convolution module Conv_dw, a mobile multi-query attention module Mobile MQA, and a convolution module Conv_pw2 connected in sequence.
[0027] Preferably, when training the slag glass detection model, the training parameters are set, and the cosine annealing learning rate adjustment strategy and data enhancement strategy are adopted. The overall loss function is composed of classification loss, box loss and distribution focus loss. During the training process, the weight ratio of each sub-item in the overall loss function is optimized, and the training is continuously iterated until the model converges.
[0028] Preferably, the method for constructing the data set includes:
[0029] At the industrial site of a waste incineration plant, high-definition linear array industrial cameras and strip LED light sources were used to collect image data of slag glass. The ratio of glass fragments to slag fragments was approximately 4:6, and the particle size was 5mm-30mm. Images were screened to construct a dataset with corresponding features. During annotation, the collected images with a resolution of 1696×200 were stitched together every 5 images, and the resolution was reduced to 1280×1280 to fill in the white edges. The slag and glass were annotated using annotation software.
[0030] The present invention has the beneficial effects of enhancing the feature receptive field by integrating FEM, effectively improving the backbone network's ability to express small target features, and thus enhancing image detection accuracy. Through M_DEDS, the present invention reduces the number of parameters and computational costs, while further optimizing the network's feature extraction capabilities, further improving detection accuracy while ensuring the algorithm's real-time performance. Through data augmentation and reasonable model improvements, the present invention's model is adaptable to different working conditions and exhibits certain generalization and robustness to lighting conditions and glass morphology. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Schematic diagram of the training and testing process of the present invention;
[0032] Figure 2 Schematic diagram of the principle of the slag glass detection model of the present invention;
[0033] Figure 3 This is the principle diagram of the M_DEDS module;
[0034] Figure 4 This is a schematic diagram of the DEDS module;
[0035] Figure 5 This is the principle diagram of the FEM module;
[0036] Figure 6 is the training loss curve, tralin / box_loss represents the box loss during training, tralin / cls_loss represents the classification loss during training, tralin / dfl_loss represents the distribution focus loss during training, val / box_loss represents the box loss during training, val / cls_loss represents the classification loss during training, val / dfl_loss represents the distribution focus loss during training;
[0037] Figure 7 The following are pictures of the samples, where (a) is dirty glass, and (b) is cement and ceramic blocks;
[0038] Figure 8 Detection results under different lighting and blur conditions. DETAILED DESCRIPTION
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0040] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0041] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.
[0042] The slag glass detection method based on deep learning in this embodiment includes:
[0043] Step 1: Collect slag glass images, annotate them, and build a dataset:
[0044] At a waste incineration plant, a high-definition linear array industrial camera and a strip LED light source were used to collect image data of slag glass. The ratio of glass fragments to slag fragments was approximately 4:6, with particle sizes ranging from 5mm to 30mm. 2,909 images were selected to construct a dataset with corresponding features, and after annotation, the total number of samples was 4,398.
[0045] The collected original image needs to be gamma corrected to adjust the brightness and contrast of the image, enhance the brightness and texture details of the dirty glass, and at the same time avoid overexposure of the high-transmittance and high-brightness glass. The gamma value is set to 0.5.
[0046] To improve generalization capabilities, data augmentation operations were performed on the dataset, including random rotation, brightness perturbation, horizontal flipping, and vertical flipping. This increased data diversity and enabled the model to learn more glass objects with different shapes and characteristics.
[0047] Images with a resolution of 1696×200 were stitched together in groups of 5, and the resolution was reduced to 1280×1280 to fill in the white edges. The slag and glass were annotated using annotation software.
[0048] The dataset is divided into training set, test set and validation set according to the ratio of 8:1:1.
[0049] Step 2: Construct a slag glass detection model. The input of the slag glass detection model is the slag glass image, and the output includes the location and category of the target in the slag glass image;
[0050] The slag glass detection model includes a backbone network, a feature fusion network, and a classification network. The backbone network is used to extract features of different scales from the input image. The feature fusion network fuses the extracted features of different scales. The fused features of different scales are input into the classification network for position and category detection. The classification network outputs the location and category of the target.
[0051] In a preferred embodiment, the backbone network of this embodiment includes a convolution module Conv1, a convolution module Conv2, a multi-branch feature fusion structure M_DEDS No. 1, a convolution module Conv3, a multi-branch feature fusion structure M_DEDS No. 2, a convolution module Conv4, a multi-scale feature enhancement module FEM, a convolution module Conv5, a multi-branch feature fusion structure M_DEDS No. 3, a fast spatial pyramid pooling module SPPF and a cross-channel spatial pyramid attention module C2PSA connected in sequence;
[0052] The outputs of the No. 2 multi-branch feature fusion structure M_DEDS, the multi-scale feature enhancement module FEM, and the cross-channel spatial pyramid attention module C2PSA are the extracted features of different scales;
[0053] This implementation introduces a multi-scale feature enhancement module (FEM), which enables the model to simultaneously fuse local details and global context information in the mid-level feature extraction stage, expand the feature receptive field, and thus improve the network's feature representation capability for tiny targets.
[0054] To address the issues of small object feature loss and insufficient semantic information, a multi-scale feature enhancement module (FEM) is introduced to improve the backbone network. The FEM module uses multi-branch standard convolution to extract multi-dimensional discriminative semantic information and extracts rich contextual information through parallel stacking of dilated convolutions.
[0055] The multi-scale feature enhancement module FEM consists of 4 channels. The input features are input into the 4 channels at the same time. The first channel uses 1×1 standard convolution to generate equivalent features. Figure 1 , using the residual structure to generate equivalent feature maps to retain the key information of small targets;
[0056] The second channel uses 1×1 standard convolution, 1×3 standard convolution, 3×1 standard convolution, and 3×3 hole convolution in sequence to generate equivalent features. Figure 2 ;
[0057] The third channel uses 1×1 standard convolution, 3×1 standard convolution, 1×3 standard convolution, and 3×3 hole convolution in sequence to generate equivalent features. Figure 3 ;
[0058] A dilated convolution layer with a dilation rate of 5 is embedded in the two middle channels to scan a wider range of feature areas without adding additional parameters;
[0059] The fourth channel uses 1×1 standard convolution and 3×3 standard convolution in sequence to generate equivalent features. Figure 4 ;
[0060] Equivalent features Figure 1 , equivalent features Figure 2 , equivalent features Figure 3 , equivalent features Figure 4 Through residual connection, the output features of the multi-scale feature enhancement module FEM are obtained.
[0061] In a preferred embodiment, the feature fusion network of this embodiment includes upsampling module No. 1, splicing module No. 4, multi-branch feature fusion structure No. 4 M_DEDS, upsampling module No. 2, splicing module No. 5, multi-branch feature fusion structure No. 5 M_DEDS, convolution module Conv5, splicing module No. 6, multi-branch feature fusion structure No. 6 M_DEDS, convolution module Conv6, splicing module No. 7, multi-branch feature fusion structure No. 7 M_DEDS;
[0062] The output of the cross-channel spatial pyramid attention module C2PSA is input to the upsampling module No. 1, the splicing module No. 4 splices the output of the upsampling module No. 1 and the output of the multi-scale feature enhancement module FEM, and the spliced result is input to the multi-branch feature fusion structure No. 4 M_DEDS, and the output of the multi-branch feature fusion structure M_DEDS No. 4 is input to the upsampling module No. 2, the splicing module No. 5 splices the output of the upsampling module No. 2 and the output of the multi-branch feature fusion structure M_DEDS No. 2, and the spliced result is input to the multi-branch feature fusion structure M_DEDS No. 5, and the multi-branch feature fusion structure M_DEDS No. 5 is input to the multi-branch feature fusion structure M_DEDS No. 5. The output of the feature fusion structure M_DEDS is input to the convolution module Conv5. The No. 6 splicing module splices the output of the convolution module Conv5 and the output of the No. 4 multi-branch feature fusion structure M_DEDS. The spliced result is input to the No. 6 multi-branch feature fusion structure M_DEDS. The output of the No. 6 multi-branch feature fusion structure M_DEDS is input to the convolution module Conv6. The No. 7 splicing module splices the output of the convolution module Conv6 and the output of the cross-channel spatial pyramid attention module C2PSA. The spliced result is input to the No. 7 multi-branch feature fusion structure M_DEDS.
[0063] The outputs of the multi-branch feature fusion structure M_DEDS No. 5, the multi-branch feature fusion structure M_DEDS No. 6, and the multi-branch feature fusion structure M_DEDS No. 7 are fusion features of different scales.
[0064] The M_DEDS structure of this embodiment is composed of multiple convolution Conv layers and DEDS modules. The module is divided into different branches, each branch is first processed by the convolution layer, and then the features are fused between different branches. Each multi-branch feature fusion structure M_DEDS includes convolution module Conv7, convolution module Conv8, convolution module Conv9, dynamic extended depth separable module DEDS No. 1, dynamic extended depth separable module DEDS No. 2, dynamic extended depth separable module DEDS No. 3, dynamic extended depth separable module DEDS No. 4, convolution module Conv10, convolution module Conv9, 11, convolution module Conv12, convolution module Conv13, convolution module Conv14, splicing module No. 1, splicing module No. 2, splicing module No. 3;
[0065] The input of the convolution module Conv7 is the input of the multi-branch feature fusion structure M_DEDS, and the output of the convolution module Conv7 is simultaneously input to the convolution module Conv8, the convolution module Conv9, and the No. 1 splicing module;
[0066] The output of the convolution module Conv8 is input to the No. 1 dynamically extended depth separable module DEDS, and the output of the No. 1 dynamically extended depth separable module DEDS is input to the No. 2 dynamically extended depth separable module DEDS. The No. 2 splicing module splices the output of the No. 2 dynamically extended depth separable module DEDS and the output of the convolution module Conv9, and the spliced result is input to the convolution module Conv10;
[0067] The output of convolution module Conv10 is simultaneously input to convolution module Conv11, convolution module Conv12, and splicing module No. 1;
[0068] The output of the convolution module Conv11 is input to the dynamic extended depth separable module DEDS No. 3, the output of the dynamic extended depth separable module DEDS No. 3 is input to the dynamic extended depth separable module DEDS No. 4, the splicing module No. 2 splices the output of the dynamic extended depth separable module DEDS No. 4 and the output of the convolution module Conv12, and the spliced result is input to the convolution module Conv13, and the output of the convolution module Conv13 is input to the splicing module No. 1;
[0069] Splicing module No. 1 splices the input, and the spliced result is input to the convolution module Conv14. The output of the convolution module Conv14 is the output of the multi-branch feature fusion structure M_DEDS.
[0070] For wide images with a resolution of 1696*200 in the slag glass sorting scenario, this embodiment proposes a dynamically expanded depthwise separable module DEDS. In the depthwise separable convolution, the expansion rate determines the magnification of the input channel to the intermediate hidden layer channel. The dynamic expansion rate dynamically adjusts the expansion rate according to the resolution of the input features to optimize the computational efficiency and accuracy of the model. Since the image width is much larger than the height, the adaptive mechanism can dynamically increase the expansion rate according to the input width, thereby improving the network's ability to capture the fine-grained texture and edge features of slag glass. The depthwise separable convolution can optimize computational efficiency while ensuring accuracy and avoid unnecessary computational overhead.
[0071] In this embodiment, the dynamically extended depthwise separable module DEDS includes a point-by-point convolution module Conv_pw1, a depthwise separable convolution module Conv_dw, a mobile multi-query attention module Mobile MQA, and a convolution module Conv_pw2 connected in sequence.
[0072] Conv_pw1 is a point-by-point convolution used to change the number of channels in the feature map. As the first convolutional layer, it expands the number of channels in the input feature map from inp to hidden_dim. Hidden_dim is calculated based on the input channel number inp and the expansion ratio expand_ratio. Conv_dw is a depthwise separable convolution used to extract spatial information from the feature map. Conv_pw2 is also a point-by-point convolution, compressing the number of channels in the feature map from hidden_dim to oup, and outputting the final feature map.
[0073] Point-by-point convolution module Conv_pw1: used to expand the number of channels of the feature map, using 1x1 convolution and SiLU activation function.
[0074] Depthwise separable convolution Conv_dw: used to extract the spatial information of the feature map, using 3x3 depthwise separable convolution without using an activation function.
[0075] Point-by-point convolution module Conv_pw2: used to compress the number of channels of the feature map, using 1x1 convolution and no activation function.
[0076] This implementation first expands the number of channels through a 1×1 point-by-point convolution module Conv_pw1, with the expansion rate dynamically determined according to the resolution. It then extracts spatial features through a 3×3 depthwise separable convolution Conv_dw, further processes the features through a mobile multi-query attention module Mobile MQA, and finally compresses the channels through a 1×1 point-by-point convolution module Conv_pw2.
[0077] Step 3: Use the dataset to train the slag glass detection model:
[0078] Model training parameter settings: The batch size is set to 32, the momentum is set to 0.937, and the initial learning rate is set to 0.01. A cosine annealing learning rate adjustment strategy and data augmentation strategy are used. As the number of training rounds increases, the learning rate is gradually reduced to ensure the convergence stability of the model.
[0079] The overall loss function consists of classification loss, box loss, and distribution focus loss. The classification loss uses binary cross entropy loss with a weight of 0.5 to measure the accuracy of the model's classification of glass objects. The box loss takes into account factors such as the intersection-over-union ratio and the center point distance, with a weight of 7.5 to ensure that the model can accurately predict the bounding box position of the glass object. The distribution focus loss uses cross entropy to optimize the probability distribution near the label, with a weight of 1.5 to enhance the model's learning ability for difficult samples. The overall loss function is calculated according to the following formula:
[0080] L total =λ cls L cls +λ box L box +λ dfl L dfl
[0081] Among them, the classification loss λ cls , box loss λ box , distribution focusing loss λ obj is the weight coefficient.
[0082] The classification loss in this embodiment adopts binary cross entropy loss, and the calculation formula in binary classification is:
[0083]
[0084] Among them, y i ∈{0,1} is the target category label; p i is the predicted category score; σ is the Sigmoid function.
[0085] The distribution focus loss of this embodiment enables the network to quickly focus on the value near the label, making the probability density at the label as large as possible. The cross entropy function is used to optimize the probabilities of two positions near the label y. The formula is:
[0086] L dfl =-((y r -y)log(p l )+(yy l )log(p r ))
[0087] Among them, p l and p r is the y predicted by the modell and y r The corresponding probability value.
[0088] The positioning loss in this embodiment is used to evaluate the similarity between the predicted box and the true box. The calculation formula is as follows:
[0089]
[0090] Among them, IoU is the intersection-over-union ratio of the predicted box and the real box; ρ 2 (b,b gt ) is the square of the distance between the center of the predicted box and the center of the true box; c is the diagonal length of the minimum enclosing rectangle; v is a parameter that measures the consistency of the aspect ratio.
[0091] Step 4: Use the trained slag glass detection model to detect slag glass images. Through confidence threshold screening and intersection-over-union (IoU) overlap filtering, accurate regression of the detection box is achieved, and the target category information is output. Other models (Halcon-CNN, YOLOv8n, YOLOv8s, and YOLO11s) were trained using the same dataset and training strategy. The parameters of these models are compared in Table 1.
[0092] Table 1 Comparative experiments on test sets
[0093]
[0094] It can be seen that the accuracy and F1 score of the present invention are the highest compared with other common models, reflecting its better detection ability for slag glass.
[0095] Table 2 Real-time detection ablation experiment
[0096]
[0097] After adding the multi-scale feature enhancement module (FEM), the recall rate increased by 1.3%, indicating that the multi-scale feature enhancement module (FEM) can enhance the feature receptive field through multi-branch dilated convolution, thereby improving the feature extraction ability for small targets. The recall rate increased by 3.1%, the overall F1 score increased by 0.0113, and the number of parameters decreased by 25%. This shows that the inverted bottleneck structure reduces the number of network parameters while also increasing the network depth, further improving the feature expression ability. Ultimately, the method of the present invention achieves the highest accuracy and recall rate. In addition, while improving detection performance, a balance between parameter number and computational efficiency is achieved through structural optimization, ensuring the real-time performance of the detection algorithm. The final single frame takes 20.7ms, meeting the 40FPS acquisition frame rate requirement of the line array camera.
[0098] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely illustrative of the principles and applications of the invention. It should be understood that many modifications may be made to the illustrative embodiments, and that other arrangements may be devised, without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in ways other than those described in the original claims. It should also be understood that features described in conjunction with individual embodiments may be employed in conjunction with other described embodiments.
Claims
1. A slag glass detection method based on deep learning, characterized in that: include: Collect slag glass images, annotate the slag glass images, and build a data set; The slag glass detection model is trained using the dataset. The input of the slag glass detection model is the slag glass image, and the output includes the location and category of the target in the slag glass image. The slag glass detection model includes a backbone network, a feature fusion network, and a classification network; the backbone network is used to extract features of different scales from the input image, the feature fusion network fuses the extracted features of different scales, and the fused features of different scales are input into the classification network for position and category detection, and the classification network outputs the position and category of the target; The trained slag glass detection model is used to detect slag glass images.
2. The slag glass detection method based on deep learning according to claim 1, characterized in that: The backbone network includes the sequentially connected convolution module Conv1, convolution module Conv2, multi-branch feature fusion structure M_DEDS No. 1, convolution module Conv3, multi-branch feature fusion structure M_DEDS No. 2, convolution module Conv4, multi-scale feature enhancement module FEM, convolution module Conv5, multi-branch feature fusion structure M_DEDS No. 3, fast spatial pyramid pooling module SPPF and cross-channel spatial pyramid attention module C2PSA; The outputs of the No. 2 multi-branch feature fusion structure M_DEDS, the multi-scale feature enhancement module FEM, and the cross-channel spatial pyramid attention module C2PSA are the extracted features of different scales; Each multi-branch feature fusion structure M_DEDS includes convolution module Conv7, convolution module Conv8, convolution module Conv9, dynamic extended depth separable module DEDS No. 1, dynamic extended depth separable module DEDS No. 2, dynamic extended depth separable module DEDS No. 3, dynamic extended depth separable module DEDS No. 4, convolution module Conv10, convolution module Conv9, 11, convolution module Conv12, convolution module Conv13, convolution module Conv14, splicing module No. 1, splicing module No. 2, and splicing module No. 3; The input of the convolution module Conv7 is the input of the multi-branch feature fusion structure M_DEDS, and the output of the convolution module Conv7 is simultaneously input to the convolution module Conv8, the convolution module Conv9, and the No. 1 splicing module; The output of the convolution module Conv8 is input to the No. 1 dynamically extended depth separable module DEDS, and the output of the No. 1 dynamically extended depth separable module DEDS is input to the No. 2 dynamically extended depth separable module DEDS. The No. 2 splicing module splices the output of the No. 2 dynamically extended depth separable module DEDS and the output of the convolution module Conv9, and the spliced result is input to the convolution module Conv10; The output of convolution module Conv10 is simultaneously input to convolution module Conv11, convolution module Conv12, and splicing module No. 1; The output of the convolution module Conv11 is input to the dynamic extended depth separable module DEDS No. 3, the output of the dynamic extended depth separable module DEDS No. 3 is input to the dynamic extended depth separable module DEDS No. 4, the splicing module No. 2 splices the output of the dynamic extended depth separable module DEDS No. 4 and the output of the convolution module Conv12, and the spliced result is input to the convolution module Conv13, and the output of the convolution module Conv13 is input to the splicing module No. 1; Splicing module No. 1 splices the input, and the spliced result is input to the convolution module Conv14. The output of the convolution module Conv14 is the output of the multi-branch feature fusion structure M_DEDS.
3. The slag glass detection method based on deep learning according to claim 2, characterized in that: The feature fusion network includes upsampling module No. 1, splicing module No. 4, multi-branch feature fusion structure No. 4 M_DEDS, upsampling module No. 2, splicing module No. 5, multi-branch feature fusion structure No. 5 M_DEDS, convolution module Conv5, splicing module No. 6, multi-branch feature fusion structure No. 6 M_DEDS, convolution module Conv6, splicing module No. 7, multi-branch feature fusion structure No. 7 M_DEDS; The output of the cross-channel spatial pyramid attention module C2PSA is input to the upsampling module No. 1, the splicing module No. 4 splices the output of the upsampling module No. 1 and the output of the multi-scale feature enhancement module FEM, and the spliced result is input to the multi-branch feature fusion structure No. 4 M_DEDS, and the output of the multi-branch feature fusion structure M_DEDS No. 4 is input to the upsampling module No. 2, the splicing module No. 5 splices the output of the upsampling module No. 2 and the output of the multi-branch feature fusion structure M_DEDS No. 2, and the spliced result is input to the multi-branch feature fusion structure M_DEDS No. 5, and the multi-branch feature fusion structure M_DEDS No. 5 is input to the multi-branch feature fusion structure M_DEDS No.
5. The output of the feature fusion structure M_DEDS is input to the convolution module Conv5. The No. 6 splicing module splices the output of the convolution module Conv5 and the output of the No. 4 multi-branch feature fusion structure M_DEDS. The spliced result is input to the No. 6 multi-branch feature fusion structure M_DEDS. The output of the No. 6 multi-branch feature fusion structure M_DEDS is input to the convolution module Conv6. The No. 7 splicing module splices the output of the convolution module Conv6 and the output of the cross-channel spatial pyramid attention module C2PSA. The spliced result is input to the No. 7 multi-branch feature fusion structure M_DEDS. The outputs of the multi-branch feature fusion structure M_DEDS No. 5, the multi-branch feature fusion structure M_DEDS No. 6, and the multi-branch feature fusion structure M_DEDS No. 7 are fusion features of different scales.
4. The slag glass detection method based on deep learning according to claim 2, characterized in that: The multi-scale feature enhancement module FEM includes 4 channels, the input features are input into the 4 channels at the same time, and the first channel uses 1×1 standard convolution to generate an equivalent feature map 1; The second channel uses 1×1 standard convolution, 1×3 standard convolution, 3×1 standard convolution, and 3×3 dilated convolution in sequence to generate the equivalent feature map 2; The third channel uses 1×1 standard convolution, 3×1 standard convolution, 1×3 standard convolution, and 3×3 hole convolution in sequence to generate the equivalent feature map 3; The fourth channel uses 1×1 standard convolution and 3×3 standard convolution in sequence to generate the equivalent feature map 4; Equivalent feature map 1, equivalent feature map 2, equivalent feature map 3, and equivalent feature map 4 are connected through residuals to obtain the output features of the multi-scale feature enhancement module FEM.
5. The slag glass detection method based on deep learning according to claim 2, characterized in that: The dynamically extended depthwise separable module DEDS includes a point-by-point convolution module Conv_pw1, a depthwise separable convolution module Conv_dw, a mobile multi-query attention module Mobile MQA, and a convolution module Conv_pw2 connected in sequence.
6. The slag glass detection method based on deep learning according to claim 2, characterized in that When training the slag glass detection model, the training parameters are set, and the cosine annealing learning rate adjustment strategy and data augmentation strategy are adopted. The overall loss function is composed of classification loss, box loss and distribution focus loss. During the training process, the weight ratio of each sub-item in the overall loss function is optimized, and the training is continuously iterated until the model converges.
7. The slag glass detection method based on deep learning according to claim 2, characterized in that: Methods for constructing datasets include: At the industrial site of a waste incineration plant, high-definition linear array industrial cameras and strip LED light sources were used to collect image data of slag glass. The ratio of glass fragments to slag fragments was approximately 4:6, and the particle size was 5mm-30mm. Images were screened to construct a dataset with corresponding features. During annotation, the collected images with a resolution of 1696×200 were stitched together every 5 images, and the resolution was reduced to 1280×1280 to fill in the white edges. The slag and glass were annotated using annotation software.
8. A computer-readable storage device storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the slag glass detection method based on deep learning as described in any one of claims 1 to 7 are implemented.
9. A slag glass detection device based on deep learning, comprising a storage device, a processor, and a computer program stored in the storage device and executable on the processor, characterized in that: The processor executes the computer program to implement the steps of the slag glass detection method based on deep learning as described in any one of claims 1 to 7.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the slag glass detection method based on deep learning as described in any one of claims 1 to 7 are implemented.