Construction waste detection method and storage medium
By using the DE-YOLO network model, combined with the lightweight convolution module, attention mechanism and detection head, the existing construction waste detection methods have been solved, with low accuracy, narrow application scope and large parameters, and efficient and accurate construction waste detection is achieved.
Patent Information
- Application Number
- CN202411960387.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing construction waste detection methods have problems such as low recognition accuracy, narrow application scope and large parameters, and are difficult to effectively deploy in actual industrial environments.
Construction waste detection is performed using the DE-YOLO network model, which includes a backbone network that integrates lightweight convolution modules, a neck network that integrates attention mechanisms, and a detection head that can realize regression prediction and classification.
It improves the accuracy and scope of application of construction waste detection, reduces the amount of calculation and parameter, and enables the model to be deployed and used efficiently in actual industrial environments.
Smart Images

Figure CN120071123A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection in machine vision, and in particular relates to a construction waste detection method and a storage medium. Background Art
[0002] The current solutions for recycling construction waste include crushing the construction waste with a crusher and then using gravity sorting or chemical sorting to process the construction waste into coarse and fine aggregates for use as filler. However, both chemical sorting and gravity sorting require manual initial screening and subsequent screening of the remaining materials. The initial screening stage mainly uses manual sorting, which is extremely inefficient. At the same time, the working environment is generally poor, and long-term sorting work is harmful to the workers' physical and mental health. The problem of automated identification of construction waste needs to be solved urgently.
[0003] In view of the above problems, the construction waste detection method based on deep learning classification in the prior art has achieved a relatively good detection effect. However, due to the accelerating urbanization process, large-scale construction projects, high vehicle transportation costs, limited transportation capacity, or insufficient construction waste treatment facilities, construction waste has been piled up for a long time and stacked seriously. The detection method based on the YOLO model in the prior art adopts the idea of a single-stage detector, and the positioning accuracy of the target is low. For the detection scene where various types of construction waste are mixed and piled up, its applicability is still insufficient. In addition, when the color of the construction waste sample is similar to the background color, it is easy to miss or misdetect. Finally, construction waste detection needs to meet the lightweight conditions. Existing research often does not take into account the complexity of the detection method, so that it cannot be successfully deployed and used in actual industrial environments. Therefore, it is of great significance to study a construction waste detection method with high recognition accuracy, wide application range and low parameter quantity. Summary of the invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and to provide a construction waste detection method and storage medium with high recognition accuracy, wide application range and low parameter quantity.
[0005] To achieve the above object, the present invention is implemented by adopting the following technical solutions:
[0006] In a first aspect, the present invention provides a construction waste detection method, the method comprising:
[0007] Collect images of construction waste to be detected;
[0008] Input the construction waste image into a pre-built and trained DE-YOLO network model, wherein the DE-YOLO network model includes a backbone network fused with a lightweight convolution module, a neck network fused with an attention mechanism, and a detection head capable of achieving regression prediction and classification;
[0009] Abstractly extract the deep features of the construction waste image using the backbone network to generate multi-scale feature maps;
[0010] Sample and fuse the multi-scale feature maps using the neck network, and weight the features fused in each layer through an attention mechanism to generate a multi-scale fused semantic feature map;
[0011] After performing convolutional operations on the multi-scale fused semantic feature map using the detection head, perform regression prediction and classification respectively to obtain the construction waste detection result.
[0012] Combined with the first aspect, further, the training method of the DE-YOLO network model includes:
[0013] Collect construction waste sample images under different distribution scenarios of construction waste;
[0014] Mark the position information and category information of the construction waste in the construction waste sample image to generate a label corresponding to the construction waste image;
[0015] Construct a sample data set using the construction waste sample images marked with the labels;
[0016] Divide the sample data set into a training set, a validation set, and a test set;
[0017] Train the pre-constructed DE-YOLO network model using the training set, validate the trained DE-YOLO network model using the validation set, test the validated DE-YOLO network model using the test set, and use the DE-YOLO network model with the optimal test result as the finally trained DE-YOLO network model.
[0018] Combined with the first aspect, further, the training method of the DE-YOLO network model further includes:
[0019] Construct a total loss function based on the weighted sum of the classification loss, the localization loss, and the confidence loss;
[0020] With the goal of minimizing the value of the total loss function, use the Adam optimizer to iteratively optimize the DE-YOLO network model.
[0021] Combined with the first aspect, further, the backbone network includes a ConvModule module, a Dual_C2f module, and an SPPF module; among them, a Dual_C2f module and a ConvModule module are connected in series to form a downsampling module;
[0022] Abstracting and extracting the deep features of the construction waste image by using the backbone network to generate multi-scale feature maps, including:
[0023] Using two cascaded ConvModule modules to initially capture the low-level feature information of the construction waste image to obtain a low-level feature map;
[0024] Using at least three cascaded downsampling modules to perform downsampling operations on the low-level feature map in sequence to gradually obtain deep feature maps of different scales;
[0025] Performing multi-scale feature extraction on the deep feature map output by the last downsampling module by using the Dual_C2f module and the SPPF module in sequence, and finally outputting feature maps of multiple different scales.
[0026] Combined with the first aspect, further, the ConvModule module includes a two-dimensional convolutional layer Conv2d, a first batch normalization layer, and a SiLu activation function connected in sequence;
[0027] The Dual_C2f module includes a 1x1 convolutional layer Conv1, a split module, a Bottleneck module, a Concat module, and a 1x1 convolutional layer Conv2;
[0028] The Dual_C2f module performs the following processing steps on the input deep feature map:
[0029] Using the 1x1 convolutional layer Conv1 to adjust the number of channels of the input deep feature map to obtain a feature map with adjusted number of channels;
[0030] Using the split module to split the feature map with adjusted number of channels into two feature maps with equal number of channels;
[0031] Inputting the two feature maps obtained by splitting by the split module into multiple cascaded Bottleneck modules for processing in sequence;
[0032] Using the Concat module to concatenate each feature map processed by multiple cascaded Bottleneck modules together to obtain a concatenated feature map;
[0033] Using the 1x1 convolutional layer Conv2 to process the concatenated feature map to obtain the feature map processed by the Dual_C2f module;
[0034] Among them, the Bottleneck module includes a standard convolution module, a second batch normalization module, a first ReLU activation function, and a DualConv module. The DualConv module includes a grouped convolution module, a pointwise convolution module, a fusion module, a third batch normalization module, and a second ReLU activation function connected in sequence;
[0035] The Bottleneck module performs the following processing steps on the input feature map:
[0036] Input the input feature map into the standard convolution module, perform spatial feature extraction on the input feature map through a 3x3 convolution layer, and sequentially input the extracted spatial features into the second batch normalization module and the first ReLU activation function to obtain an enhanced feature map;
[0037] Use the grouped convolution module to divide the number of channels of the enhanced feature map into multiple groups, and perform convolution operations independently on each group to obtain a grouped convolution feature map;
[0038] Use the pointwise convolution module to adjust the number of channels of the grouped convolution feature map to obtain a pointwise convolution feature map;
[0039] Use the fusion module to add the grouped convolution feature map and the pointwise convolution feature map element by element to obtain a first fusion feature map;
[0040] Input the first fusion feature map into the third batch normalization module and the second ReLU activation function for processing and then output to obtain the feature map processed by the Bottleneck module.
[0041] Combined with the first aspect, further, the SPPF module includes a first-layer network structure, a second-layer network structure, a third-layer network structure, and a fourth-layer network structure. Among them, both the first-layer network structure and the fourth-layer network structure are ConvModule modules, the second-layer network structure includes three max-pooling layers, and the third-layer network structure is a Concat module;
[0042] The method for the SPPF module to perform multi-scale feature extraction includes:
[0043] Use the first-layer network structure to process the input feature map, split the processed feature into three paths. One path of the three paths passes through one max-pooling layer in the second-layer network structure for multi-scale feature extraction, another path of the three paths sequentially passes through two max-pooling layers in the second-layer network structure for multi-scale feature extraction, and the remaining path of the three paths sequentially passes through three max-pooling layers in the second-layer network structure for multi-scale feature extraction;
[0044] Using the third-layer network structure, the output features of the three max-pooling layers are concatenated with the output features of the first-layer network structure in the channel dimension to achieve the fusion of multi-scale information;
[0045] The concatenated features are input into the fourth-layer network structure for feature extraction, and a multi-scale feature map is output.
[0046] Combined with the first aspect, further, the neck network introduces an improved FPN-PAN double-tower structure, and upsampling is achieved through the nearest neighbor interpolation algorithm, and downsampling is completed through the ConvModule module;
[0047] Using the neck network to sample and fuse the multi-scale feature map, and weighting the features fused in each layer through an attention mechanism to generate a multi-scale fused semantic feature map, including:
[0048] The multi-scale feature map is upsampled level by level according to the scale and concatenated with the feature map of a larger scale to obtain fused semantic features;
[0049] The fused semantic features are input into the C2f module, and the C2f module introduces an ECA attention mechanism module;
[0050] Using the ECA attention mechanism module to adaptively adjust the channel weights of the fused semantic features to generate an upsampled multi-scale fused semantic feature map;
[0051] The upsampled multi-scale fused semantic feature map is downsampled level by level according to the scale, concatenated with the feature map of a smaller scale, and then input into the C2f module, and the ECA attention mechanism module is used to generate the final multi-scale fused semantic feature map.
[0052] Combined with the first aspect, further, the use of the ECA attention mechanism module to adaptively adjust the channel weights of the fused semantic features includes:
[0053] Capturing the global information in the channel dimension through global average pooling, compressing the fused semantic features into a one-dimensional vector to obtain a global information vector;
[0054] Applying a one-dimensional convolution operation to process the global information vector, using the convolution kernel to capture the dependencies between channels, and adaptively adjusting the weights of each channel;
[0055] Among them, the expression for adaptively adjusting the weights of each channel is:
[0056] ;
[0057] Among them, K is the number of convolution kernels, zis the channel dimension, , are parameters in the relational expression of the convolution kernel and the channel dimension.
[0058] Combined with the first aspect, further, the detection head includes a ConvModule module and a Conv2d module; a set of detection heads are respectively configured for each scale of the multi-scale fusion semantic feature map;
[0059] After performing convolution operations on the multi-scale fusion semantic feature map using the detection head, regression prediction and classification are respectively performed to obtain the construction waste detection result, including:
[0060] The feature maps of each scale in the multi-scale fusion semantic feature map are each split into two paths, one path is processed successively through the ConvModule module and the Conv2d module of the corresponding detection head as the regression prediction result; the other path is processed successively through the ConvModule module and the Conv2d module of the corresponding detection head as the classification result.
[0061] In a second aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of the first aspects are implemented.
[0062] Compared with the prior art, the beneficial effects achieved by the present invention:
[0063] The technical solution provided by the present invention uses a trained DE-YOLO network model to detect construction waste. The DE-YOLO network model includes a backbone network integrating a lightweight convolution module, a neck network integrating an attention mechanism, and a detection head capable of performing regression prediction and classification; among them, the backbone network integrating the lightweight convolution module can reduce the amount of calculation and achieve efficient calculation; combined with the neck network integrating the attention mechanism, it enhances the prominence of useful information in the feature map and at the same time suppresses the interference of redundant information, thereby improving the accuracy of construction waste detection, and further being able to effectively detect and locate construction waste of different sizes. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0065] Figure 1 is a method flow chart of a construction waste detection method provided by an embodiment of the present invention;
[0066] Figure 2 It is a schematic diagram of the network structure of a DE-YOLO network model provided by an embodiment of the present invention;
[0067] Figure 3 is Figure 2 a schematic diagram of the network structure of a Dual_C2f module in
[0068] Figure 4 is Figure 2 a schematic diagram of the network structure of an ECA attention mechanism module in
[0069] Figure 5 It is the actual detection effect diagram of the application example of the present invention. Detailed implementation manners
[0070] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and cannot be used to limit the protection scope of the present invention.
[0071] Embodiment 1:
[0072] This embodiment provides a construction waste detection method. As Figure 1 shown, it is the flow chart of the method provided in this embodiment, which mainly includes the following steps:
[0073] Step S1: Collect construction waste images to be detected;
[0074] Step S2: Input the construction waste images into a pre-constructed and trained DE-YOLO network model, where the DE-YOLO network model includes a backbone network integrating a lightweight convolution module, a neck network integrating an attention mechanism, and a detection head capable of performing regression prediction and classification;
[0075] Step S3: Use the backbone network to abstract and extract the deep features of the construction waste images to generate multi-scale feature maps;
[0076] Step S4: Use the neck network to sample and fuse the multi-scale feature maps, and weight the features fused in each layer through the attention mechanism to facilitate restoring the details lost in the feature maps in Step S3 to generate multi-scale fused semantic feature maps;
[0077] Step S5: After performing convolution operations on the multi-scale fused semantic feature maps using the detection head, perform regression prediction and classification respectively to obtain the construction waste detection results.
[0078] It should be noted that the multi-scale feature maps should be understood as feature maps of multiple different scales. As Figure 2As shown, it is a schematic diagram of the network structure of the DE-YOLO network model provided in this embodiment. In the leftmost box in the figure, a network structure of the backbone network Backbone is provided, which mainly includes a ConvModule module, a Dual_C2f module, and an SPPF module. Among them, a Dual_C2f module and a ConvModule module are connected in series to form a downsampling module. In this embodiment, referring to Figure 2 It can be seen that there are two ConvModule modules, and there are a total of three groups of downsampling modules. Correspondingly, the backbone network finally outputs three feature maps of different scales, which are denoted as feature map C1, feature map C2, and feature map C3 from top to bottom for easy description. Among them, the size of feature map C1 is 80x80, and the number of channels is 256; the size of feature map C2 is 40x40, and the number of channels is 512; the size of feature map C3 is 20x20, and the number of channels is 512.
[0079] In this embodiment, in order to reduce the amount of calculation and the number of parameters and reduce the complexity of the training model, three groups of downsampling modules are selected to gradually obtain multi-scale feature maps, which is convenient for subsequent multi-scale feature fusion operations. In actual applications, the amount of data calculation and data distribution of object detection may be different, and the required image resolution and model performance are also different. When using this method and its model, an appropriate number of downsampling modules can be adopted according to a specific scenario to help the model adapt to different data distributions and improve the performance of the model in a specific scenario.
[0080] Next, in combination with Figure 2 the network structure of the backbone network in, a further detailed description will be given on how the backbone network abstracts and extracts the construction waste image in step S3:
[0081] Two series-connected ConvModule modules are used to initially capture the low-level feature information of the construction waste image to obtain a low-level feature map;
[0082] At least three series-connected downsampling modules are used to perform downsampling operations on the input feature map in sequence to gradually obtain deep feature maps of different scales;
[0083] The deep feature map output by the last downsampling module is then subjected to multi-scale feature extraction using a Dual_C2f module and an SPPF module in sequence, and finally feature maps of multiple different scales are output.
[0084] Specifically, the ConvModule module includes a two-dimensional convolutional layer Conv2d, a first batch normalization layer, and a SiLu activation function connected in sequence. When the construction waste image is input into the DE-YOLO network model, the ConvModule module extracts the spatial information and channel information in the construction waste image through the two-dimensional convolutional layer Conv2d; then normalizes the spatial information and channel information through the first batch normalization layer to accelerate the training process and improve the stability of the model; finally, performs a non-linear transformation on the normalized output information through the SiLu activation function to introduce non-linear features and enhance the expression ability of the model.
[0085] Furthermore, the Dual_C2f module plays a key role in the backbone network. It mainly extracts and transforms the features of the input data through operations such as feature transformation, branch processing, and feature fusion to generate a more representative output. As Figure 3 shown, it is a schematic diagram of the network structure of the Dual_C2f module provided in this embodiment. Among them, the Dual_C2f module includes a 1x1 convolutional layer Conv1, a split module, a Bottleneck module, a Concat module, and a 1x1 convolutional layer Conv2.
[0086] The following combines Figure 3 , and further elaborates in detail on how the Dual_C2f module processes the input deep feature map:
[0087] Use the 1x1 convolutional layer Conv1 to adjust the number of channels of the input deep feature map to obtain a feature map with adjusted number of channels;
[0088] Use the split module to split the feature map with adjusted number of channels into two feature maps with equal number of channels;
[0089] Input the two feature maps obtained by splitting through the split module into multiple cascaded Bottleneck modules in sequence for processing;
[0090] Use the Concat module to splice each path of feature map processed by multiple cascaded Bottleneck modules together to obtain a spliced feature map. The number of channels of the spliced feature map is the sum of the output channels of all paths;
[0091] Use the 1x1 convolutional layer Conv2 to process the spliced feature map to obtain the feature map processed by the Dual_C2f module.
[0092] Specifically, for ease of description, let the height of the deep feature map input to the Dual_C2f module be H 0 , and the width be W 0First, the 1x1 convolution layer Conv1 performs convolution operations on the deep feature map and adjusts the number of channels. The adjusted feature map size becomes (H 1 , W 1 ), the number of channels is adjusted to C 1 ; Secondly, the split module splits the adjusted feature map into two paths, A and B, and the size of each feature map is still (H 1 , W 1 ), but the number of channels becomes D 2 =D 1 / 2, the A-path feature map is sequentially input into multiple serially connected bottleneck Bottleneck modules for processing, and the B-path feature map directly enters the splicing Concat module without passing through the bottleneck Bottleneck module; wherein, the A-path feature map will be divided into multiple groups after entering the bottleneck Bottleneck module, and among these groups, one group will directly enter the splicing Concat module without passing through the next-level serial bottleneck Bottleneck module, and the remaining groups will be processed by the bottleneck Bottleneck module of this level and input into the next-level bottleneck Bottleneck module until the processing is completed; then, the B-path feature map, the feature map output by the last-level bottleneck Bottleneck module in the A-path feature map, and the remaining feature maps that have not been grouped by the next level are spliced and integrated in the channel dimension through the splicing Concat module; finally, the feature map whose channel number is restored after splicing is processed through the 1x1 convolutional layer Conv2 to obtain the feature map processed by the Dual_C2f module.
[0093] It should be noted that after the feature map is split into two feature maps with equal number of channels using the split module, the number of channels of each feature map is D 2 Therefore, in order to facilitate subsequent feature extraction, the number of channels of the deep feature map in the 1x1 convolution layer Conv1 is adjusted to 2*D 2 After being processed by the Dual_C2f module, the model optimizes the capture and integration of complex features to the greatest extent while maintaining computational efficiency, thus improving the performance of the overall model.
[0094] As an embodiment, in order to effectively fuse multi-scale features and enrich the semantic features of feature maps without changing the size of feature maps, an SPPF module is also introduced into the backbone network. The SPPF module includes a four-layer network structure, which are named as the first layer network structure, the second layer network structure, the third layer network structure and the fourth layer network structure for ease of description, wherein the first layer network structure and the fourth layer network structure are both ConvModule modules, the second layer network structure includes three maximum pooling layers, and the third layer network structure is a Concat module.
[0095] When performing multi-scale feature extraction in the SPPF module, first, the input feature map is processed using the first-layer network structure, and the processed features are split into three paths. One of the three paths undergoes multi-scale feature extraction through a max pooling layer in the second-layer network structure. Another of the three paths sequentially undergoes multi-scale feature extraction through two max pooling layers in the second-layer network structure. The remaining one of the three paths sequentially undergoes multi-scale feature extraction through three max pooling layers in the second-layer network structure;
[0096] Secondly, the output features of the three max pooling layers and the output features of the first-layer network structure are concatenated in the channel dimension using the third-layer network structure to achieve the fusion of multi-scale information;
[0097] Finally, the concatenated features are input into the fourth-layer network structure for feature extraction, and a multi-scale feature map is output. The introduction of the SPPF module enables the model to better adapt to inputs of various scales, thereby reducing the computational amount of the model and improving the detection accuracy.
[0098] In some embodiments, as Figure 2 shown, a network structure of the Neck network, an improved FPN-PAN two-tower structure, is provided in the square box in the middle of the DE-YOLO network model. In the improved FPN-PAN two-tower structure, FPN uses upsampling to conduct deep semantic features to the shallow layer and enhance semantic expressions at multiple scales; while PAN, on the contrary, uses downsampling to conduct shallow localization information to the deep layer and enhance the localization ability at multiple scales.
[0099] Next, in combination with Figure 2 the network structure of the Neck network in, a further detailed description will be given of how the Neck network samples and fuses the multi-scale feature map in step S4 and how to weight the features fused in each layer through the attention mechanism:
[0100] In the upsampling path of FPN on the left side of the Neck network, the multi-scale feature map is upsampled step by step according to the scale through the nearest neighbor interpolation algorithm Upsample to restore the spatial resolution of the feature map and obtain an upsampled feature map that is twice as large as the input feature map; then it is concatenated with the feature map of the same scale output by the backbone network respectively to obtain fused semantic features, thereby enhancing semantic expressions at multiple scales;
[0101] Then, the fused semantic features are input into the C2f module introduced with the ECA attention mechanism module to further extract the fused semantic features;
[0102] Then, the channel weights of the fused semantic features are adaptively adjusted using the ECA attention mechanism module, effectively improving the expressive power of the feature map, enabling the network model to better focus on important feature regions, and finally generating an upsampled multi-scale fused semantic feature map.
[0103] Specifically, the Upsample module, Concat module, C2f module, and ECA attention mechanism module are connected in series to form an upsampling structure; during the upsampling process of FPN, two series-connected upsampling structures are passed through. The upsampling process is as follows: when the three feature maps C1, C2, and C3 of different scales finally output by the backbone network are input into the neck network, the feature map C3 is divided into two groups. One group directly enters the downsampling path of PAN to obtain the feature map C3´, with a size of 20x20 and 512 channels; the other group, after passing through the Upsample module, has its length and width doubled while the number of channels remains unchanged. After being concatenated with the feature map C2 of the same scale, it then passes through the C2f module and the ECA attention mechanism module respectively to better extract deep semantic information, and finally outputs the feature map C2´, with a size of 40x40 and 512 channels; the feature map C2´ then passes through an upsampling structure to enlarge the scale, concatenate with the feature map C1, and extract deep semantic information, and then outputs the feature map C1´, with a size of 80x80 and 256 channels; thus, finally generating the upsampled multi-scale fused semantic feature maps C1´, C2´, and C3´.
[0104] In the downsampling path of PAN on the right side of the neck network, the upsampled multi-scale fused semantic feature maps are downsampled level by level through the ConvModule module to enhance the localization expression at multiple scales, obtaining a downsampled feature map that is half the size of the input feature map;
[0105] Finally, after concatenating the downsampled feature map with a feature map of a smaller scale, it is input into the C2f module, and the ECA attention mechanism module is used to generate the final multi-scale fused semantic feature map.
[0106] Further, the ConvModule module, the concatenation Concat module, the C2f module, and the ECA attention mechanism module are connected in series to form a downsampling structure; during the downsampling process of PAN, two series-connected downsampling structures will be passed through. The downsampling process is as follows: when downsampling the upsampled multi-scale fused semantic feature maps C1´, C2´, and C3´, the feature map C1´ is divided into two groups. One group directly enters the detection head Head as the input feature map C1´´, with a size of 80x80 and 256 channels; the other group, after passing through the ConvModule module, has its length and width reduced to half of the original, and the number of channels remains unchanged. After concatenating with the feature map C2´ of the same scale, it then passes through the C2f module and the ECA attention mechanism module respectively to better restore the shallow position information, and finally outputs the feature map C2´´, with a size of 40x40 and 512 channels; the feature map C2´´ then undergoes another downsampling structure to reduce the scale, concatenate with the feature map C3´, and extract the shallow position information, and then outputs the feature map C3´´, with a size of 20x20 and 512 channels; thus generating the final multi-scale fused semantic feature maps C1´´, C2´´, and C3´´ as the input feature maps of the detection head Head with three different scale sizes.
[0107] The improved FPN-PAN two-tower structure is introduced, enabling the DE-YOLO model to generate more accurate multi-scale feature maps and effectively identify and locate targets of different sizes.
[0108] As an alternative embodiment, as Figure 2 shown, a network structure of the detection head Head is provided in the rightmost box of the DE-YOLO network model. The detection head mainly includes a ConvModule module and a Conv2d module; a set of detection heads are respectively configured for the feature maps C1´´, C2´´, and C3´´ corresponding to each scale of the multi-scale fused semantic feature map, with scales of 80x80, 40x40, and 20x20.
[0109] Next, in combination with Figure 2 the network structure of the detection head in, a further detailed description will be given on how the detection head performs convolution operations on the multi-scale fused semantic feature map and how to perform regression prediction and classification in step S5:
[0110] Each scale of the feature map in the multi-scale fused semantic feature map is split into two paths. One path is successively processed by the ConvModule module and the Conv2d module of the corresponding detection head as the regression prediction result; the other path is successively processed by the ConvModule module and the Conv2d module of the corresponding detection head as the classification result.
[0111] It should be noted that the ConvModule module structures in this embodiment are all the same, including a 3*3 two-dimensional convolutional layer Conv2d, a first batch normalization layer, and a SiLu activation function connected in sequence.
[0112] The technical solution provided in this embodiment uses a trained DE-YOLO network model to detect construction waste. The backbone network integrating the lightweight convolutional module can reduce the computational amount and achieve efficient calculation; the neck network integrating the attention mechanism enhances the prominence of useful information in the feature map and suppresses the interference of redundant information at the same time; combined with the detection head that can achieve regression prediction and classification, the DE-YOLO network model thus improves the accuracy of construction waste detection and can effectively detect and locate construction waste of different sizes.
[0113] Embodiment 2:
[0114] This embodiment provides a method for detecting construction waste, which is different from Embodiment 1 in that the Bottleneck module in the Dual_C2f module mainly includes a standard convolutional module, a second batch normalization module, a first ReLU activation function, and a DualConv module. The DualConv module mainly includes a grouped convolution module GroupConvolution, a pointwise convolution module Pointwise Convolution, a fusion module, a third batch normalization module, and a second ReLU activation function connected in sequence.
[0115] The Bottleneck module performs the following processing steps on the input feature map:
[0116] Input the input feature map into the standard convolutional module, perform spatial feature extraction on the input feature map through a 3x3 convolutional layer, and sequentially input the extracted spatial features into the second batch normalization module and the first ReLU activation function to increase the non-linear transformation and further enhance the expression ability of the features to obtain an enhanced feature map;
[0117] Use the grouped convolution module Group Convolution to divide the number of channels of the enhanced feature map into multiple groups through a 3x3 convolutional layer and perform convolutional operations independently on each group. This process deeply explores the local correlation of the features within the group, enhances the model's learning ability for diverse features while reducing the computational amount, thus achieving efficient calculation and retaining rich spatial features to obtain a grouped convolution feature map with the same spatial size;
[0118] Use the pointwise convolution module Pointwise Convolution to adjust the number of channels of the grouped convolution feature map through a 1x1 convolutional layer, capture the dependencies between channels, and obtain a pointwise convolution feature map;
[0119] The grouped convolution feature map and the pointwise convolution feature map are added element by element using a fusion module, thereby fusing the features of the grouped convolution and the pointwise convolution to obtain a first fused feature map. Without changing the size, the first fused feature map achieves information fusion between channels and recombination of features, combining the advantage of grouped convolution in reducing computational complexity and the ability of pointwise convolution to adjust channels.
[0120] The first fused feature map is input into the third batch normalization module to improve training stability, accelerate network convergence, and reduce the impact of internal covariate shift. Immediately afterwards, it is input into the second ReLU activation function to increase non-linear transformation and enhance the expressive power of the feature map, enabling the network to learn more complex features. Finally, the feature map processed by the bottleneck Bottleneck module is obtained as the output.
[0121] As an embodiment, the ECA attention mechanism described in Embodiment 1 can enhance the attention to important features and suppress redundant information while maintaining computational efficiency. Combining Figure 4 With the network structure schematic diagram of the ECA attention mechanism module provided in this embodiment, the method for adaptively adjusting the channel weights of the fused semantic features will be further described in detail:
[0122] GAP, namely global average pooling, captures global information in the channel dimension through global average pooling, compresses the fused semantic features into a one-dimensional vector to obtain a global information vector;
[0123] Apply the one-dimensional convolution Conv1d operation to process the global information vector, use the convolution kernel to capture the dependencies between channels, and adaptively adjust the weights of each channel;
[0124] Among them, the expression for adaptively adjusting the weights of each channel is:
[0125] ;
[0126] Among them, K is the number of convolution kernels, z is the channel dimension, , are the parameters in the relationship formula between the convolution kernel and the channel dimension, defaulting to 2 and 1 respectively, odd refers to rounding up to the nearest odd number.
[0127] Embodiment 3:
[0128] This embodiment provides a construction waste detection method, and will further describe in detail the method for training the DE-YOLO network model in steps S1 and S2 of Embodiment 1:
[0129] Collect the construction waste sample images under different distribution scenarios of construction waste;
[0130] Annotate the position information and category information of the construction waste in the construction waste sample image to generate the label corresponding to the construction waste image;
[0131] Construct a sample data set using the construction waste sample images marked with the said labels;
[0132] Divide the said sample data set into a training set, a validation set and a test set;
[0133] Use the said training set to train the pre-constructed DE-YOLO network model, use the said validation set to validate the trained DE-YOLO network model, use the said test set to test the validated DE-YOLO network model, and take the DE-YOLO network model with the optimal test result as the finally trained DE-YOLO network model.
[0134] It should be noted that when collecting the construction waste sample images, it is necessary to collect images with representativeness, high recycling value, complex scenarios and various types; specifically, the complex scenarios include scenarios such as scattered collection, slight aliasing, and severe aliasing, and the types of construction waste include steel bars, concrete blocks, stones, foams, bricks, etc. In addition, after normalizing the abscissa and ordinate of the center of the bounding box, the width and height of the bounding box in the construction waste sample image to the range of [0, 1], they are used to annotate the position information of the construction waste; the category information is annotated by the type of construction waste to generate the corresponding category label, and the label value uses the natural number sequence starting from 0.
[0135] As an optional embodiment, the training method of the said DE-YOLO network model further includes:
[0136] Construct a total loss function according to the weighted sum of the classification loss, the localization loss and the confidence loss; among them, the classification loss can measure the difference between the predicted category and the true category, the localization loss can measure the gap between the bounding box coordinates of the predicted box and the true bounding box coordinates, and the confidence loss can measure the difference between the confidence of the predicted box and the true confidence;
[0137] With the goal of minimizing the value of the said total loss function, use the Adam optimizer to iteratively optimize the said DE-YOLO network model; among them, the initial learning rate is set to 0.001, and the total number of training times is set to 100; during the training process, perform a validation every 10 training times, test the accuracy of the validation set of the model, and save the set of model weights with the highest validation set accuracy as the final DE-YOLO network model weights.
[0138] Specifically, the expression of the classification loss is:
[0139] ,
[0140] Among them, c is the number of sample categories, is the indicator function of the true label. If the sample belongs to category i, then otherwise . is the probability predicted as category i.
[0141] The localization loss expression is:
[0142] ,
[0143] Among them, DFL is a loss function in object detection, aiming to solve the uncertainty problem in bounding box regression, and is generally expressed as:
[0144] ,
[0145] Among them, n is the number of discrete intervals, is the true distribution, is the predicted distribution;
[0146] CIOU is a method for measuring the similarity of bounding boxes improved on the basis of traditional IOU , and its calculation formula is:
[0147] ,
[0148] Among them, is the intersection over union of two bounding boxes, is the predicted bounding box, is the true bounding box, is the square of the Euclidean distance between the centers of the predicted bounding box and the true bounding box, is the diagonal length of the smallest closed region that can contain both the predicted bounding box and the true bounding box, is used to measure the similarity of the aspect ratios of the predicted bounding box and the true bounding box.
[0149] The confidence loss expression is:
[0150]
[0151] Among them, and are respectively the indicator functions indicating whether the bounding box contains an object, and are the predicted confidences.
[0152] The expression of the total loss function is as follows:
[0153]
[0154] Among them, is the classification loss, is the localization loss, is the confidence loss, and the weight factors , and are used to adjust the importance of each loss part in the total loss.
[0155] In order to verify the detection effect of the embodiments of the present invention on construction waste, a comparative experiment was carried out. The experimental results are shown in Table 1, where the bold indicates the best index.
[0156]
[0157] Table 1 Comparative experimental results of different test models
[0158] As can be seen from Table 1, the model volumes and floating-point operation amounts of the second-order object detection network Faster R-CNN, RT-DETR-resnet50 network, Swin-S network, and YOLOX-s network are relatively large, and the operations are complex, which is not conducive to deployment and use in the actual industrial environment. The SSD network has the highest precision rate, which is 96.75%, but its recall rate is the lowest, only 68.68%, and the model volume and floating-point operation amount are more than ten times that of the DE-YOLO network. Although the recall rates of the YOLOv5s network and the YOLOv8n network are the same as that of the DE-YOLO network, both are 97%, but the detection accuracy of the YOLOv5s network is 87.5%, and the detection accuracy of the YOLOv8n network is 89.8%, both of which are less than 92.5% of the DE-YOLO network. Except for the precision rate, the other performances of the DE-YOLO network are all optimal, and compared with the YOLO series networks, its precision rate is only about 2% lower than that of the YOLOX-s network. From the above experimental results, it can be seen that the DE-YOLO network has good robustness and low model complexity, and all performances reach a satisfactory level.
[0159] Figure 5 is the actual detection effect diagram of the application example of the present invention. As Figure 5 shown, the model constructed in this embodiment still shows high accuracy and stability in detecting various materials in complex scenarios. Among them, from Figure 5(d) It can be seen that the detection confidence of the foam material is as high as 0.95, indicating that the model has outstanding performance in identifying lightweight materials. For small samples such as stones that are prone to overlap or repetition in a single scene, the model can still detect them separately and assign relatively high confidence levels, such as 0.72, 0.76, and 0.78, effectively solving the problems of misdetection and missed detection that occur in traditional models in complex scenarios.
[0160] Combining the four figures (a), (b), (c), and (d), it can be seen that the detection confidence of concrete basically reaches above 0.77, demonstrating the reliability of the model in distinguishing structural materials. The detection confidence of bricks all reaches above 0.65, indicating that even in the case of low edge clarity, the model can still effectively identify them. In addition, the detection confidence of rebars is also relatively significant. Since such objects are slender in shape and easily blend with the background, traditional models often tend to have missed detections or repeated detections. However, the present model can still accurately locate rebars in complex backgrounds, further proving its detection robustness for long and slender structure samples.
[0161] Compared with traditional detection models, the present model can still exhibit excellent detection performance in scenarios with severe sample overlap, complex backgrounds, diverse categories, and large shape differences, effectively improving the accuracy and stability of detection, and having strong generalization ability and engineering application potential.
[0162] In summary, the DE-YOLO network model constructed in the construction waste detection method proposed in the embodiment of the present invention still has high detection accuracy even in different complex scenarios or various backgrounds, and its lightweight model can meet the applications of terminals.
[0163] It should be noted that Figure 1 only the logical order of the method described in this embodiment is shown. On the premise of not conflicting with each other, in other possible embodiments of the present application, the steps shown or described can be completed in a different Figure 1 order than that shown. The method provided in this embodiment can be applied to terminals and can be executed by a construction waste detection device. This device can execute the method provided in any one of the first to third embodiments, has the corresponding functional modules and beneficial effects for executing the method, and this device can be implemented in software and / or hardware. This device can be integrated in a terminal, such as: any smart phone, tablet computer, or computer device with communication functions.
[0164] Embodiment 4:
[0165] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in any one of Embodiments 1 to 3 are implemented.
[0166] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of these features.
[0167] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0168] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0170] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.
Claims
1. A construction waste detection method, characterized in that: The method comprises: Collect images of construction waste to be detected; Input the construction waste image into a pre-built and trained DE-YOLO network model, wherein the DE-YOLO network model includes a backbone network fused with a lightweight convolution module, a neck network fused with an attention mechanism, and a detection head capable of achieving regression prediction and classification; Using the backbone network to abstractly extract the deep features of the construction waste image to generate a multi-scale feature map; The multi-scale feature map is sampled and fused using the neck network, and the fused features of each layer are weighted by an attention mechanism to generate a multi-scale fused semantic feature map; After the detection head performs convolution operation on the multi-scale fusion semantic feature map, regression prediction and classification are performed respectively to obtain the construction waste detection result.
2. The construction waste detection method according to claim 1, characterized in that: The training method of the DE-YOLO network model includes: Collect sample images of construction waste in different distribution scenarios; Marking the location information and category information of the construction waste in the construction waste sample image to generate a label corresponding to the construction waste image; Using sample images of construction waste marked with the labels to construct a sample data set; Dividing the sample data set into a training set, a validation set and a test set; The training set is used to train the pre-built DE-YOLO network model, the validation set is used to validate the trained DE-YOLO network model, the test set is used to test the validated DE-YOLO network model, and the DE-YOLO network model with the optimal value obtained in the test result is used as the final trained DE-YOLO network model.
3. The construction waste detection method according to claim 2, characterized in that: The training method of the DE-YOLO network model also includes: Construct a total loss function based on the weighted sum of classification loss, localization loss, and confidence loss; With the goal of minimizing the value of the total loss function, the Adam optimizer is used to iteratively optimize the DE-YOLO network model.
4. The construction waste detection method according to claim 1, characterized in that: The backbone network includes a ConvModule module, a Dual_C2f module and a SPPF module; Among them, a Dual_C2f module and a ConvModule module are connected in series to form a downsampling module; The method of using the backbone network to abstractly extract the deep features of the construction waste image and generate a multi-scale feature map includes: Two serially connected ConvModule modules are used to preliminarily capture low-level feature information of the construction waste image to obtain a low-level feature map; Using at least three serially connected downsampling modules to sequentially downsample the low-level feature maps to gradually obtain deep-level feature maps of different scales; The deep feature map output by the last downsampling module is then used in turn by the Dual_C2f module and the SPPF module for multi-scale feature extraction, and finally feature maps of multiple different scales are output.
5. The construction waste detection method according to claim 4, characterized in that: The ConvModule module includes a two-dimensional convolutional layer Conv2d, a first batch of normalization layers and a SiLu activation function connected in sequence; The Dual_C2f module includes a 1x1 convolutional layer Conv1, a split module, a bottleneck module, a concatenation module and a 1x1 convolutional layer Conv2; The Dual_C2f module performs the following processing steps on the input deep feature map: Use the 1x1 convolution layer Conv1 to adjust the number of channels of the input deep feature map to obtain the feature map after the number of channels is adjusted; Use the split module to split the feature map after the number of channels is adjusted into two feature maps with equal number of channels; The two feature maps obtained by splitting the split module are sequentially input into multiple series-connected bottleneck modules for processing; Use the Concat module to concatenate each feature map processed by multiple series-connected bottleneck modules to obtain a concatenated feature map. The concatenated feature map is processed using the 1x1 convolutional layer Conv2 to obtain the feature map processed by the Dual_C2f module; The bottleneck module includes a standard convolution module, a second batch normalization module, a first ReLU activation function and a DualConv module, and the DualConv module includes a grouped convolution module, a point-by-point convolution module, a fusion module, a third batch normalization module and a second ReLU activation function connected in sequence; The bottleneck module performs the following processing steps on the input feature map: The input feature map is input into the standard convolution module, spatial features are extracted from the input feature map through a 3x3 convolution layer, and the extracted spatial features are sequentially input into the second batch normalization module and the first ReLU activation function to obtain an enhanced feature map; Using a grouped convolution module to divide the number of channels of the enhanced feature map into multiple groups, and independently performing a convolution operation on each group to obtain a grouped convolution feature map; Use the point-by-point convolution module to adjust the number of channels of the grouped convolution feature map to obtain the point-by-point convolution feature map; Using a fusion module, the grouped convolution feature map and the point-by-point convolution feature map are added element by element to obtain a first fused feature map; The first fused feature map is sequentially input into the third batch normalization module and the second ReLU activation function for processing and output to obtain the feature map processed by the bottleneck module.
6. The construction waste detection method according to claim 4 or 5, characterized in that: The SPPF module includes a first-layer network structure, a second-layer network structure, a third-layer network structure and a fourth-layer network structure, wherein the first-layer network structure and the fourth-layer network structure are both ConvModule modules, the second-layer network structure includes three maximum pooling layers, and the third-layer network structure is a Concat module; The method for extracting multi-scale features by the SPPF module includes: The input feature map is processed by using the first-layer network structure, and the features obtained by the processing are split into three paths, one of the three paths passes through a maximum pooling layer in the second-layer network structure for multi-scale feature extraction, another of the three paths passes through two maximum pooling layers in the second-layer network structure in turn for multi-scale feature extraction, and the remaining one of the three paths passes through three maximum pooling layers in the second-layer network structure in turn for multi-scale feature extraction; The output features of the three maximum pooling layers are concatenated with the output features of the first layer of network structure in the channel dimension by using the third layer of network structure to achieve fusion of multi-scale information; The concatenated features are input into the fourth-layer network structure for feature extraction, and a multi-scale feature map is output.
7. The construction waste detection method according to claim 1, characterized in that: The neck network introduces an improved FPN-PAN dual-tower structure, achieves upsampling through the nearest neighbor interpolation algorithm, and completes downsampling through the ConvModule module; The method of sampling and fusing the multi-scale feature map by using the neck network and weighting the fused features of each layer by an attention mechanism to generate a multi-scale fused semantic feature map includes: The multi-scale feature maps are upsampled step by step according to the scale, and spliced with the feature maps of larger scales to obtain fused semantic features; Input the fused semantic features into the C2f module, and the C2f module introduces an ECA attention mechanism module; Adaptively adjust the channel weights of the fused semantic features using the ECA attention mechanism module to generate an upsampled multi-scale fused semantic feature map; The upsampled multi-scale fusion semantic feature map is downsampled step by step according to the scale, spliced with the feature map of a smaller scale, and input into the C2f module, and the final multi-scale fusion semantic feature map is generated using the ECA attention mechanism module.
8. The construction waste detection method according to claim 7, characterized in that: The method of adaptively adjusting the channel weight of the fused semantic feature by using the ECA attention mechanism module includes: Capturing global information of the channel dimension through global average pooling, compressing the fused semantic features into a one-dimensional vector to obtain a global information vector; Applying a one-dimensional convolution operation to process the global information vector, using a convolution kernel to capture dependencies between channels, and adaptively adjusting the weights of each channel; The expression for adaptively adjusting the weight of each channel is: ; in, K is the number of convolution kernels, z is the channel dimension, , is the parameter in the relationship between the convolution kernel and the channel dimension.
9. The construction waste detection method according to claim 1, characterized in that: The detection head includes a ConvModule module and a Conv2d module; a set of detection heads is configured for each scale of the multi-scale fusion semantic feature map; After the detection head performs convolution operation on the multi-scale fusion semantic feature map, regression prediction and classification are performed respectively to obtain the construction waste detection result, including: The feature map of each scale in the multi-scale fusion semantic feature map is split into two paths, one of which is processed by the ConvModule module and the Conv2d module of the corresponding detection head in turn as a regression prediction result; the other is processed by the ConvModule module and the Conv2d module of the corresponding detection head in turn as a classification result.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 9 are implemented.
Citation Information
Cited By
Construction waste intelligent classification method based on deep learning
CN120807948A
Shield segment detection method and system using image features and feature enhancement
CN120997596A
A shield segment detection method and system using image features and feature enhancement
CN120997596B
River bank illegal building identification method based on multi-scale fusion
CN121121499A
Lightweight small target detection method based on multi-domain modeling and semantic embedding enhancement
CN122090229A