Construction waste intelligent detection and identification method based on image feature enhancement

By introducing defuzzy modules and other feature enhancement technologies into the intelligent sorting detection and identification method of construction waste, the detection problems caused by blurred construction waste in traditional methods are solved, and efficient detection and identification of small-target construction waste is achieved.

CN120147706APending Publication Date: 2025-06-13BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510208462.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional construction waste treatment methods are inefficient, and under dust environment and vibration of conveyor belts, the construction waste image is blurred, resulting in the loss of partial features of small target objects, making it difficult to effectively detect and identify.

Method used

Using the YOLOv10-based intelligent sorting detection and identification method for construction waste, the model's feature extraction capability of small targets is enhanced by introducing defuzzy modules, spatial depth conversion convolution, SimAM attention mechanism and hollow space convolution pooling pyramids in the main layer, and designing lightweight modules in the main layer and neck feature fusion part.

Benefits of technology

It effectively solves the problem of blurred image of construction waste caused by conveyor belt vibration and dust environment, improves the detection and recognition accuracy of small-target construction waste, and enhances the detection speed and generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147706A_ABST
    Figure CN120147706A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of building waste intelligent sorting image data processing, and particularly relates to a building waste intelligent detection and recognition method based on image feature enhancement. The method comprises the following steps: 1, collecting a construction waste image in a construction waste sorting field as an original image sample set, and carrying out the motion blur and noise processing of the image through Python; 2, a detection and recognition model for intelligent sorting of the YOLOv10 construction waste is established; 3, building waste data set images are sent into the YOLOv10 building waste intelligent sorting detection and recognition model established in the step 2 to be trained and verified, label smoothing processing is used in the training process to improve the robustness and generalization ability of the model, and the optimal weight is obtained; and 4, after the optimal weight is obtained, testing is carried out, the optimal weight is loaded, and construction waste data set images of a test set are sent into the YOLOv10 construction waste intelligent sorting detection and recognition model for testing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image data processing for intelligent sorting of construction waste, and in particular relates to an intelligent detection and identification method for construction waste based on image feature enhancement. Background Art

[0002] With the acceleration of urbanization, the construction industry has flourished, and a large amount of construction waste has been generated. How to effectively deal with this waste has become an urgent issue. Construction waste is mainly composed of bricks, concrete, glass, reinforced concrete, plastic pipes, gravel, wood, fiber and metal. Faced with such a huge amount of construction waste, an efficient and reasonable treatment solution is needed. At present, my country's treatment method for construction waste is relatively single, mainly through preliminary screening by processing equipment, separation of light materials by magnetic force and airflow, and finally relying on manual classification. However, the traditional treatment method is inefficient. The use of manual methods not only consumes a lot of human resources, but also poses a threat to personnel health due to the harsh working environment (such as high dust and noise). Therefore, research on aggregate recycling technologies such as reduction, harmlessness and resource utilization of construction waste has become an important issue that society urgently needs to solve. In particular, the effective separation of waste metal, fiber, plastic, wood, etc. in construction waste is a key technical link to improve the reuse rate of building materials. In order to meet the above challenges, it is urgent to develop an intelligent construction waste sorting detection and identification technology to replace the traditional manual operation method and improve the processing efficiency and safety.

[0003] Computer vision technology is increasingly used in intelligent sorting of construction waste, but its accuracy and efficiency are still affected by many factors. In actual operation, due to the different sizes of construction waste, the vibration of the conveyor belt during transportation and the influence of the dusty environment often lead to blurred images, resulting in the loss of some image features of small target objects, which brings difficulties to detection and recognition. Therefore, measures should be taken to solve the problems of motion blur and dust, and enhance the model's ability to detect and identify small target construction waste, thereby improving the overall accuracy and efficiency of intelligent sorting detection.

[0004] Object detection methods based on computer vision and deep learning can be roughly divided into two categories: one is the two-stage detection algorithm represented by Fast R-CNN and Faster R-CNN, and the other is the single-stage detection algorithm represented by YOLO (You Only Look Once) and SSD (Single Shot Detector). The two-stage detection algorithm first determines the regions that may contain objects by generating candidate boxes, and then processes these candidate boxes with a complex architecture to achieve final detection and classification. In contrast, the single-stage detection algorithm directly detects and classifies all spatial regions, focuses on the proposals of all spatial regions, and completes the object detection task at once through a relatively simple architecture. This type of algorithm can achieve fast object detection while maintaining high accuracy, and has the advantages of fast detection speed, easy implementation of the algorithm, and end-to-end optimization. Therefore, in the application scenario of construction waste detection and recognition, single-stage detection algorithms are usually preferred. This method can effectively address the problem of image blurring caused by conveyor belt vibration and dust environment, especially the loss of some features of small target objects, thereby improving the accuracy and stability of recognition. Summary of the Invention

[0005] The purpose of the present invention is to provide an intelligent detection and recognition method for construction waste based on image feature enhancement, which is used for intelligent detection and recognition of construction waste and sorting method to replace manual operation. This method can effectively solve the problem that in the intelligent sorting process, due to conveyor belt vibration and image feature blurring in a dust environment, it is difficult to detect and recognize some features of small target objects, enhance the feature extraction ability for small targets, improve the detection accuracy, and further improve the recycling rate of construction waste.

[0006] In order to achieve the above purpose, the present invention provides the following technical solutions:

[0007] An intelligent detection and recognition method for construction waste based on image feature enhancement, characterized in that: the method includes the following steps:

[0008] Step 1: Collect construction waste images at the construction waste sorting site as the original image sample set, and use Python to perform motion blur and noise processing on the images;

[0009] Step 2: Establish a detection and recognition model for intelligent sorting of construction waste using YOLOv10;

[0010] Introduce a deblurring module, spatial depth conversion convolution, SimAM attention mechanism, and atrous spatial pyramid pooling in the feature extraction part of the backbone layer of the model; design a lightweight module in the feature extraction part of the backbone layer and the feature fusion part of the neck;

[0011] The detection and recognition model for intelligent sorting of construction waste by YOLOv10 includes a backbone feature extraction part, a neck feature fusion part, and a head detection module part;

[0012] The backbone feature extraction part has a total of 14 layers: The first layer is the input image layer; the second layer is the deblurring module; the third layer is the spatial-depth conversion convolution; the fourth layer is the lightweight module, and the feature output by the fourth layer is P0; the fifth layer is the spatial-depth conversion convolution; the sixth layer is the lightweight module; the seventh layer is the SimAM attention mechanism module, and the feature output by the seventh layer is P1; the eighth layer is downsampling; the ninth layer is the C2fCIB module; the tenth layer is the SimAM attention mechanism module, and the feature output by the tenth layer is P2; the eleventh layer is downsampling; the twelfth layer is the C2fCIB module; the thirteenth layer is the atrous spatial pyramid pooling layer; the fourteenth layer is the SimAM attention mechanism module, and the feature output by the fourteenth layer is P3;

[0013] At the beginning of the model backbone layer, a deblurring module in the Deblur-GAN adversarial network is introduced to achieve the deblurring process of moving images;

[0014] The deblurring module has a total of 48 layers. Among them, the first layer is the convolutional layer; the second layer is the instance normalization layer; the third layer is the ReLU activation function; the fourth layer is the convolutional layer; the fifth layer is the instance normalization layer; the sixth layer is the ReLU activation function; the seventh layer is the convolutional layer; the eighth layer is the instance normalization layer; the ninth layer is the ReLU activation function; the tenth to thirty-ninth layers are 6 residual modules, where each residual module is sequentially a convolutional layer, an instance normalization layer, a ReLU activation function, a convolutional layer, and an instance normalization layer; the fortieth layer is the depthwise separable convolutional layer; the forty-first layer is the transposed convolutional layer; the forty-second layer is the instance normalization layer; the forty-third layer is the ReLU activation function; the forty-fourth layer is the transposed convolutional layer; the forty-fifth layer is the instance normalization layer; the forty-sixth layer is the ReLU activation function; the forty-seventh layer is the convolutional layer; the forty-eighth layer is the Tanh activation function;

[0015] The deblurring module consists of two branches. The first branch sends the input feature to a convolutional layer, then to a normalization layer and a ReLU activation function, and then sequentially passes through convolutional layers, normalization layers, ReLU activation functions, convolutional layers, normalization layers, and ReLU activation functions. Then it passes through six residual modules, and then sequentially through a depthwise separable convolutional layer, a transposed convolutional layer, a normalization layer, a ReLU activation function, a transposed convolutional layer, a normalization layer, a ReLU activation function, a convolutional layer, and a Tanh activation function. Finally, it is concatenated with the second branch that only contains the input feature and then outputs the feature. Each residual module consists of two branches. The first branch sequentially sends the input feature to a convolutional layer, a normalization layer, a ReLU activation function, a convolutional layer, and a normalization layer. Then the output of the first branch is concatenated with the second branch that only contains the input feature.

[0016] The spatial-depth conversion convolution includes a spatial depth layer and a progressive convolution layer. During the downsampling process, all information in the channel dimension is retained, improving the model's recognition ability for small targets and low-resolution images.

[0017] After the feature map is input, the spatial-depth conversion convolution module first undergoes downsampling, generates four separate channels through depth convolution, and then the four channels are added together and output after pointwise convolution.

[0018] The SimAM attention mechanism module generates attention weights based on the similarity between pixels, adopts a simple calculation method to significantly reduce the computational complexity, improves the model's detection ability for small targets, and at the same time makes the model training more stable and efficient.

[0019] The structure of the SimAM attention mechanism is that the input feature sequentially passes through a convolutional layer, a normalization layer, and a ReLU activation function, and then through global average pooling to obtain the feature to be mapped in the space. This feature passes through the Sigmoid function to obtain the weight map, and then after element-wise multiplication with the feature map and convolution, it is added to the input feature and then output.

[0020] The Atrous Spatial Pyramid Pooling (ASPP) can effectively process multi-scale objects in images by combining convolutional operations with different dilation rates, obtain global context without losing location information, enhance the multi-scale feature fusion ability, and improve the perception ability for small targets.

[0021] The Atrous Spatial Pyramid Pooling (ASPP) module includes five branches. The first, second, third, and fourth branches are all composed of convolutional layers, and the fifth branch is composed of a pooling layer, a convolutional layer, and an upsampling module. The five branches are concatenated, and finally output after convolutional processing.

[0022] In the neck feature fusion part, the output feature P3 is upsampled and then concatenated with the output feature P2. After passing through the C2fCIB module, the output feature layer P4 is obtained, and then it is upsampled. After that, it is concatenated with the output feature P1. After passing through the C2fCIB module, the output feature P5 is formed. Then it continues to be upsampled, concatenated with the output feature P0, and after passing through the C2fCIB module, the output feature P6 is formed. Subsequently, it enters the small object detection layer of the detection module part. At the same time, the output feature P6 is convolved and then concatenated with the output feature P5. After passing through the C2fCIB module, the output feature P7 is formed. The output feature P7 enters the smaller object detection layer of the detection module part. P7 is downsampled and then concatenated with the output feature P4. After passing through the lightweight module, the output feature P8 is formed. The output feature P8 enters the medium object detection layer of the detection module part. Finally, the output feature P8 is downsampled and then concatenated with the output feature P3. After passing through the lightweight module, the output feature P9 is formed. The output feature P9 enters the large object detection layer of the detection module part;

[0023] In order to improve the detection and recognition rate of the model, lightweight modules are adopted in both the backbone feature extraction part and the neck feature fusion part;

[0024] The lightweight module includes 4 branches. The first branch consists of a convolutional module, which has the same composition as the convolution in the neck feature fusion part; the second branch consists of a convolutional module and a slicing segmentation module; the third branch consists of a convolutional module, a slicing segmentation module, and a UIB module; the fourth branch consists of a convolutional module, a slicing segmentation module, and two identical UIB modules; the 4 branches are concatenated, and finally processed by a convolutional module for output;

[0025] For the YOLOv10 intelligent sorting detection and recognition model of construction waste, the convolution in the neck feature fusion part consists of ordinary convolution, BN batch normalization layer, and SiLU activation function;

[0026] In order to achieve accurate and effective bounding box regression, further improve the detection accuracy and generalization ability of the model, the Focal-EIoU Loss function is used to train the model;

[0027] In order to further improve the detection and recognition rate of the model, the Sophia optimizer is selected to approximate the Hessian matrix and use this information to update the parameters, optimizing the model parameters in the object detection task, including convolutional layer weights, normalization layer parameters, fully connected layer weights, and bias terms;

[0028] Step 3: Feed the construction waste dataset images into the detection and recognition model of intelligent sorting of construction waste by YOLOv10 established in Step 2 for training and validation, and use label smoothing during the training process to improve the robustness and generalization ability of the model, and obtain the optimal weights;

[0029] Step 4: After obtaining the optimal weights, conduct testing, load the optimal weights, and feed the construction waste dataset images of the test set into the detection and recognition model of intelligent sorting of construction waste by the YOLOv10 for testing.

[0030] Among them, in Step 1, the value of the motion blur convolution kernel is set to 15 to control the degree of blur, and the mean value of noise processing is set to 0 and the standard deviation is set to 10 to generate noise, simulating the problem of blurred features of some construction waste images caused by the dust environment and the vibration of the conveyor belt, and making labels for the processed dataset and dividing it into a training set, a validation set and a test set according to 8:1:1.

[0031] Among them, in Step 3, the number of training rounds is set to 300 rounds and the batch size is set to 2 during the training process.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] First, by introducing the deblurring module in the improved Deblur-GAN adversarial network into the backbone layer to preprocess the construction waste dataset images, the super-resolution of the image dataset samples is improved, and the problems of blurred construction waste images caused by conveyor belt vibration and in a dust environment are solved;

[0034] Then, introduce the spatial depth conversion convolution, SimAM attention mechanism module, and atrous spatial pyramid pooling layer in the backbone layer feature extraction part of the model, design lightweight modules in the backbone layer feature extraction part and the neck feature fusion part, and add a small target detection layer for small targets in the head prediction layer part of the model. Adopt the Focal-EIoU Loss effective and accurate bounding box regression loss function, adopt the Sophia optimizer, improve the YOLOv10 object detection model, enhance the feature extraction ability of the backbone layer feature extraction part for small target construction waste, improve the fusion ability of the neck feature fusion part, improve the loss function and optimizer of the model, and improve the detection accuracy, speed and generalization ability of the model;

[0035] Finally, the construction waste dataset images are sent into the improved YOLOv10 intelligent construction waste detection and recognition model for training, validation, and testing. During the training process, label smoothing is used to obtain the optimal weights. Then, for testing, the optimal weights are loaded, and the construction waste dataset images in the test set are sent into the improved YOLOv10 intelligent construction waste detection model for testing and detection. This can effectively solve the problem that due to the vibration of the conveyor belt and the blurring of construction waste images caused by the dust environment, the detection and recognition accuracy of some image features of small target construction waste is poor, enhance the detection ability of small target objects, and achieve accurate intelligent detection and recognition of construction waste.

[0036] The technical solution proposed by the present invention ensures a high detection speed while improving the accuracy, and can provide an effective method for the intelligent detection and recognition of construction waste. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is the specific step diagram of the present invention;

[0038] Figure 2 is the structure diagram of the deblurring module;

[0039] Figure 3 is the structure diagram of the spatial depth conversion convolution;

[0040] Figure 4 is the structure diagram of the SimAM attention mechanism;

[0041] Figure 5 is the structure diagram of the atrous spatial pyramid pooling;

[0042] Figure 6 is the structure diagram of the general inverted bottleneck module (UIB);

[0043] Figure 7 is the designed lightweight structure diagram;

[0044] Figure 8 is the structure diagram of the improved Yolov10 of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0045] The following further describes the specific embodiments of the present invention with reference to the accompanying drawings.

[0046] Referring to Figure 1 , a method for intelligent detection and recognition of construction waste based on image feature enhancement includes the following steps:

[0047] Step 1: Collect construction waste images at the construction waste sorting yard as the original image sample set, and use Python to perform motion blur and noise processing on the images.

[0048] The numerical value of the motion blur convolution kernel is set to 15 to control the degree of blur. The mean of the noise processing is set to 0 and the standard deviation is set to 10 to generate noise, simulating the problem of blurred features of some construction waste images caused by the dust environment and conveyor belt vibration, and making labels for the processed data set and dividing it into a training set, a validation set, and a test set according to 8:1:1;

[0049] In this embodiment, 6 samples of red bricks, stones, wood, plastics, steel bars, and fibers are selected as targets, and a total of 6198 construction waste data set images are used as the original image sample set. Motion blur and noise processing are adopted, and then the processed data set sample set is labeled using LabelImg to make a data set in VOC format, and divided into a training set, a validation set, and a test set according to the ratio of 8:1:1.

[0050] Step 2: Establish a detection and recognition model for intelligent sorting of construction waste by YOLOv10.

[0051] To solve the problem of blurred construction waste images caused by the vibration of the conveyor belt and the dust environment, resulting in the loss of some image features of construction waste, a detection and recognition model for intelligent sorting of construction waste by YOLOv10 is established. A deblurring module, a spatial depth conversion convolution, a SimAM attention mechanism, and an atrous spatial pyramid pooling are introduced in the feature extraction part of the backbone layer of the model; a lightweight module is designed in the feature extraction part of the backbone layer and the feature fusion part of the neck.

[0052] As Figure 8 shown, the detection and recognition model for intelligent sorting of construction waste by YOLOv10 includes a backbone feature extraction part, a neck feature fusion part, and a head detection module part.

[0053] The backbone feature extraction part has 14 layers, including: the first layer is the input image layer; the second layer is the deblurring module; the third layer is the spatial depth conversion convolution; the fourth layer is the lightweight module, and the output feature of the fourth layer is P0; the fifth layer is the spatial depth conversion convolution; the sixth layer is the lightweight module; the seventh layer is the SimAM attention mechanism module, and the output feature of the seventh layer is P1; the eighth layer is downsampling; the ninth layer is the C2fCIB module; the tenth layer is the SimAM attention mechanism module, and the output feature of the tenth layer is P2; the eleventh layer is downsampling; the twelfth layer is the C2fCIB module; the thirteenth layer is the atrous spatial pyramid pooling layer; the fourteenth layer is the SimAM attention mechanism module, and the output feature of the fourteenth layer is P3.

[0054] Deblurring module:

[0055] To solve the problem of image blurring of construction waste caused by dust environment and conveyor belt vibration, a deblurring module in the Deblur-GAN adversarial network is introduced at the beginning of the model backbone layer to realize the deblurring process of moving images.

[0056] As Figure 2 shown, the deblurring module has a total of 48 layers. Among them, the first layer is a convolutional layer; the second layer is an instance normalization layer; the third layer is a ReLU activation function; the fourth layer is a convolutional layer; the fifth layer is an instance normalization layer; the sixth layer is a ReLU activation function; the seventh layer is a convolutional layer; the eighth layer is an instance normalization layer; the ninth layer is a ReLU activation function; the 10th to 39th layers are 6 residual modules. Among them, each residual module is, in turn, a convolutional layer, an instance normalization layer, a ReLU activation function, a convolutional layer, and an instance normalization layer; the 40th layer is a depthwise separable convolutional layer; the 41st layer is a transposed convolutional layer; the 42nd layer is an instance normalization layer; the 43rd layer is a ReLU activation function; the 44th layer is a transposed convolutional layer; the 45th layer is an instance normalization layer; the 46th layer is a ReLU activation function; the 47th layer is a convolutional layer; the 48th layer is a Tanh activation function.

[0057] The deblurring module is composed of 2 branches. The first branch sends the input features to the convolutional layer, and then sends them to the normalization layer and the ReLU activation function. Then, it continues to pass through the convolutional layer, normalization layer, ReLU activation function, convolutional layer, normalization layer and ReLU activation function in turn. Then, it passes through 6 residual modules, and then passes through the depthwise separable convolutional layer, transposed convolutional layer, normalization layer, ReLU activation function, transposed convolutional layer, normalization layer, ReLU activation function, convolutional layer and Tanh activation function in turn. Finally, it is connected to the second branch Concat that only contains the input features and then outputs the features. Each residual module is composed of 2 branches. The first branch sends the input features to the convolutional layer, normalization layer, ReLU activation function, convolutional layer and normalization layer in turn. Then, the output of the first branch is connected to the second branch Concat that only contains the input features.

[0058] Spatial depth conversion convolution:

[0059] The spatial depth conversion convolution includes a spatial depth layer and a progressive convolutional layer. During the downsampling process, all information in the channel dimension is retained, improving the model's recognition ability for small targets and low-resolution images.

[0060] As Figure 3 shown, the spatial depth conversion convolution module first undergoes downsampling after the feature map is input, generates four separate channels through depth convolution, and then the four channels are added and output after pointwise convolution.

[0061] SimAM attention mechanism module:

[0062] The SimAM attention mechanism module generates attention weights based on the similarity between pixels, adopts a simple calculation method to significantly reduce the computational complexity, improves the model's detection ability for small targets, and makes the model training more stable and efficient.

[0063] As Figure 4 shown in the structure of the SimAM attention mechanism, the input features pass through a convolutional layer, a normalization layer, and a ReLU activation function in sequence, and then through global average pooling to obtain the features to be mapped in space. These features pass through the Sigmoid function to obtain a weight map, which is then element-wise multiplied with the feature map and convolved, and then added to the input features for output.

[0064] Atrous Spatial Pyramid Pooling (ASPP):

[0065] Atrous Spatial Pyramid Pooling can effectively process multi-scale objects in images by combining convolutional operations with different dilation rates, obtain global context without losing location information, enhance the multi-scale feature fusion ability, and improve the perception ability for small targets.

[0066] As Figure 5 shown, the Atrous Spatial Pyramid Pooling module includes 5 branches. The 1st, 2nd, 3rd, and 4th branches are all composed of convolutional layers, and the 5th branch is composed of a pooling layer, a convolutional layer, and an upsampling module. The 5 branches are concatenated, and finally, the output is obtained after convolutional processing.

[0067] In the neck feature fusion part, the output feature P3 is upsampled and then concatenated with the output feature P2. After passing through the C2fCIB module, the output feature layer P4 is obtained, and then it is upsampled and concatenated with the output feature P1. After passing through the C2fCIB module, the output feature P5 is formed. Then it is upsampled and concatenated with the output feature P0. After passing through the C2fCIB module, the output feature P6 is formed, and then it enters the small target detection layer in the detection module part. At the same time, the output feature P6 is convolved and concatenated with the output feature P5. After passing through the C2fCIB module, the output feature P7 is formed. The output feature P7 enters the smaller target detection layer in the detection module part. P7 is downsampled and concatenated with the output feature P4. After passing through the lightweight module, the output feature P8 is formed. The output feature P8 enters the medium target detection layer in the detection module part. Finally, the output feature P8 is downsampled and concatenated with the output feature P3. After passing through the lightweight module, the output feature P9 is formed. The output feature P9 enters the large target detection layer in the detection module part.

[0068] Lightweight module:

[0069] To improve the detection and recognition rate of the model, the present invention adopts lightweight modules in both the backbone feature extraction part and the neck feature fusion part, and replaces all C2f modules in the original YOLOv10 model with lightweight C2f modules. This structure can use optional depthwise convolutions to make a temporary trade-off between spatial and channel mixing, can expand the receptive field of the network according to actual needs, improves the flexibility of the network, and enhances the learning ability of the network. After introducing the lightweight C2f module, the number of parameters and the amount of computation are reduced, effectively improving the problem of redundant feature maps of the model, enabling the model to be deployed on mobile devices, and making it easier to implement the detection and recognition of intelligent sorting of construction waste.

[0070] The lightweight module structure consists of Figure 6 the general inverted bottleneck module (UIB). After the input feature map, the general inverted bottleneck module expands the dimension through a convolutional layer, then enters the depthwise separable convolutional layer, including depthwise convolution and pointwise convolution, and finally compresses and reduces the dimension through an activation function and a convolutional layer, and finally outputs the feature map.

[0071] The lightweight module is as Figure 7 shown. It includes 4 branches. The first branch consists of a convolutional module, which is the same as the convolution in the neck feature fusion part. The second branch consists of a convolutional module and a slicing and segmentation module. The third branch consists of a convolutional module, a slicing and segmentation module, and a UIB module. The fourth branch consists of a convolutional module, a slicing and segmentation module, and two identical UIB modules. The 4 branches are connected by Concat, and finally processed and output by a convolutional module.

[0072] For the YOLOv10 intelligent sorting detection and recognition model of construction waste, the convolution in the neck feature fusion part consists of ordinary convolution (Conv2d), BN batch normalization layer, and SiLU activation function.

[0073] To achieve accurate and effective bounding box regression, further improve the detection accuracy and generalization ability of the model, the Focal-EIoU Loss function is used to replace the CIoU Loss function in the original model to train the model. By combining the EIoU and Focal mechanisms, it effectively solves the limitations of the traditional IoU loss function in bounding box regression, especially improves the accuracy of small object detection and the processing ability of difficult example samples, thus bringing a significant performance improvement to the YOLOv10 model.

[0074] To further improve the detection and recognition rate of the model, the Sophia optimizer is selected to replace the traditional SGD or Adam optimizer. The Sophia optimizer can optimize the model parameters in the object detection task, including convolutional layer weights, normalization layer parameters, fully connected layer weights, and bias terms, by approximating the Hessian matrix and using this information for parameter updates. This not only improves the convergence speed of the model but also enhances the generalization ability and detection accuracy of the model. Especially when dealing with large-scale datasets, the Sophia optimizer can converge faster and achieve higher detection accuracy.

[0075] Step 3: Feed the construction waste dataset images into the YOLOv10 construction waste intelligent sorting detection and recognition model established in Step 2 for training and validation, and use label smoothing during the training process to improve the robustness and generalization ability of the model to obtain the optimal weights. Set the number of training epochs to 300 and the batch size to 2 during the training process.

[0076] Step 4: After obtaining the optimal weights, conduct testing. Load the optimal weights and feed the construction waste dataset images in the test set into the YOLOv10 construction waste intelligent sorting detection and recognition model for testing.

[0077] Example:

[0078] 6198 original image samples were taken at the construction waste sorting site of Beijing Urban Green Source Environmental Protection Technology Co., Ltd. Use Python to perform motion blur and noise processing, where the motion blur convolution kernel is set to 15, and the mean of the noise processing is 0, and the standard deviation is 10 to simulate the problem of blurred construction waste image features caused by dust environment and conveyor belt vibration. After processing, the dataset is labeled using LabelImg, including 6 label samples of red bricks, stones, wood, plastic, steel bars, and fibers, forming a VOC format dataset, and divided into a training set, a validation set, and a test set in a ratio of 8:1:1 to complete the production of the dataset.

[0079] To address the problem of blurred construction waste images caused by conveyor belt vibration and dust environment, which in turn leads to the loss of image features, a deblurring module, spatial-depth conversion convolution, SimAM attention mechanism, and atrous spatial pyramid pooling are introduced in the backbone layer feature extraction part of the original YOLOv10 model, and a lightweight module is designed in both the backbone layer feature extraction part and the neck feature fusion part. The designed lightweight module consists of a Universal Inverted Bottleneck module (UIB) to achieve the lightweight effect by compressing the dimensions. A small target detection layer for small targets is added to the head part, and at the same time, the Focal-EIoU Loss loss function and the Sophia optimizer are used to improve the YOLOv10 object detection model.

[0080] This embodiment was completed in an experimental environment with the Windows 11 operating system, NVIDIA GeForce RTX 4060Ti GPU, CUDA 11.6, cuDNN 8.4.1.50, and Python 3.9. The prepared construction waste dataset was input into the improved YOLOv10 intelligent sorting detection and recognition model for construction waste for training and validation. During the training process, label smoothing was adopted, and the optimal weights were obtained through multiple iterations, followed by testing. Specifically, the training was carried out for 300 rounds with a batch size of 2. This embodiment uses frames per second (FPS), F1 score, and mean average precision (mAP) to evaluate the improved YOLOv10 intelligent detection and recognition model for construction waste. The experimental results show that the FPS tested on the test set reached 52.76, the F1 value of red bricks was 0.97, the F1 value of wood was 0.91, the F1 value of stones was 0.97, the F1 value of plastics was 0.95, the F1 value of fibers was 0.85, and the F1 value of steel bars was 0.83, and the mAP value reached 92.14%.

[0081] Subsequently, a verification experiment was carried out to test the grasping effect of construction waste. The experimental equipment includes a camera, a robotic arm, a computer, a conveyor belt, etc. The trained YOLOv10 model was deployed on the computer, and the robotic arm was driven to operate. The experiment simulated the actual working environment, used diverse construction waste samples, and was tested multiple times. Finally, the grasping accuracy of the experiment reached 89.2%, verifying the effectiveness of the proposed method. This method can effectively address the problem of image blurring caused by conveyor belt vibration and dust environment, thereby improving the difficulty of small target detection, enhancing the detection ability of small target objects, and realizing the intelligent detection and recognition of construction waste. The technical solution proposed by the present invention not only improves the accuracy but also maintains a high detection speed, providing an effective solution for the intelligent detection and recognition of construction waste.

Claims

1. A construction waste intelligent detection and identification method based on image feature enhancement, characterized in that: The method comprises the following steps: Step 1: Collect construction waste images at the construction waste sorting site as the original image sample set, and use Python to perform motion blur and noise processing on the images; Step 2: Establish a detection and recognition model for intelligent sorting of construction waste using YOLOv10; In the backbone layer feature extraction part of the model, a deblurring module, spatial depth conversion convolution, SimAM attention mechanism and dilated spatial convolution pooling pyramid are introduced; lightweight modules are designed in the backbone layer feature extraction part and the neck feature fusion part; The YOLOv10 construction waste intelligent sorting detection and recognition model includes a trunk feature extraction part, a neck feature fusion part and a head detection module part; The backbone feature extraction part has 14 layers: the first layer is the input image layer; the second layer is the deblurring module; the third layer is the spatial depth conversion convolution; the fourth layer is the lightweight module, and the output feature of the fourth layer is P0; the fifth layer is the spatial depth conversion convolution; the sixth layer is the lightweight module; the seventh layer is the SimAM attention mechanism module, and the output feature of the seventh layer is P1; the eighth layer is downsampling; the ninth layer is the C2fCIB module; the tenth layer is the SimAM attention mechanism module, and the output feature of the tenth layer is P2; the eleventh layer is downsampling; the twelfth layer is the C2fCIB module; the thirteenth layer is the void space convolution pooling pyramid layer; the fourteenth layer is the SimAM attention mechanism module, and the output feature of the fourteenth layer is P3; The deblurring module in the Deblur-GAN adversarial network is introduced at the beginning of the model backbone layer to achieve deblurring of moving images. The deblurring module has a total of 48 layers, among which the first layer is a convolutional layer; the second layer is an instance normalization layer; the third layer is a ReLU activation function; the fourth layer is a convolutional layer; the fifth layer is an instance normalization layer; the sixth layer is a ReLU activation function; the seventh layer is a convolutional layer; the eighth layer is an instance normalization layer; the ninth layer is a ReLU activation function; the tenth to thirty-nine layers are six residual modules, among which each residual module is a convolutional layer, an instance normalization layer, a ReLU activation function, a convolutional layer, and an instance normalization layer in sequence; the 40th layer is a depthwise separable convolutional layer; the 41st layer is a deconvolutional layer; the 42nd layer is an instance normalization layer; the 43rd layer is a ReLU activation function; the 44th layer is a deconvolutional layer; the 45th layer is an instance normalization layer; the 46th layer is a ReLU activation function; the 47th layer is a convolutional layer; and the 48th layer is a Tanh activation function; The deblurring module is composed of two branches, the first branch sends the input features to the convolution layer, and then to the normalization layer and the ReLU activation function, and then passes through the convolution layer, the normalization layer, the ReLU activation function, the convolution layer, the normalization layer and the ReLU activation function in sequence, and then passes through 6 residual modules, and then passes through the depth-separable convolution layer, the deconvolution layer, the normalization layer, the ReLU activation function, the deconvolution layer, the normalization layer, the ReLU activation function, the convolution layer and the Tanh activation function in sequence, and finally Concat connects with the second branch containing only the input features and outputs the features; wherein each residual module is composed of two branches, the first branch sends the input features to the convolution layer, the normalization layer, the ReLU activation function, the convolution layer and the normalization layer in sequence, and then the output of the first branch is Concat connected with the second branch containing only the input features; The spatial depth conversion convolution includes a spatial depth layer and a step-wise convolution layer. During the downsampling process, all information in the channel dimension is retained, which improves the model's ability to recognize small objects and low-resolution images. The spatial depth conversion convolution module first performs downsampling after the feature map is input, generates four separate channels through deep convolution, and then adds the four channels and performs point-by-point convolution to output; The SimAM attention mechanism module generates attention weights based on the similarity between pixels. It uses a simple calculation method to significantly reduce the computational complexity, improve the model's ability to detect small targets, and make the model training more stable and efficient. The structure of the SimAM attention mechanism is that the input feature passes through the convolution layer, normalization layer, ReLU activation function in sequence, and then undergoes global average pooling to obtain the feature to be mapped in space. The feature passes through the Sigmoid function to obtain a weight map, which is then multiplied by the feature map element by element and then convolved, and then added to the input feature and output; The atrous spatial convolutional pooling pyramid can effectively process multi-scale objects in images by combining convolution operations with different atrous rates, obtain global context without losing position information, enhance the ability to fuse multi-scale features, and improve the perception of small objects. The dilated spatial convolutional pooling pyramid module includes 5 branches, the 1st, 2nd, 3rd and 4th branches are all composed of convolutional layers, the 5th branch is composed of a pooling layer, a convolutional layer and an upsampling module, the 5 branches are concat-connected and finally output after convolution processing; In the neck feature fusion part, the output feature P3 is concat-connected with the output feature P2 after upsampling, and continues to pass through the C2fCIB module to output feature layer P4, and then upsampling is performed, and then concat-connected with the output feature P1, and after passing through the C2fCIB module, the output feature P5 is formed, and then upsampling is continued, and concat-connected with the output feature P0 after passing through the C2fCIB module to form the output feature P6, and then enters the small target detection layer of the detection module part. At the same time, the output feature P6 is concat-connected with the output feature P5 after convolution, and after passing through the C2fCIB module, the output feature P7 enters the smaller target detection layer of the detection module part, and P7 is concat-connected with the output feature P4 after downsampling, and after passing through the lightweight module, the output feature P8 enters the medium target detection layer of the detection module part. Finally, the output feature P8 is concat-connected with the output feature P3 after downsampling, and after passing through the lightweight module, the output feature P9 enters the large target detection layer of the detection module part; In order to improve the detection and recognition rate of the model, lightweight modules are used in the backbone layer feature extraction part and the neck feature fusion part; The lightweight module includes 4 branches, the first branch is composed of a convolution module, which is the same as the convolution module of the neck feature fusion part; the second branch is composed of a convolution module and a slice segmentation module; the third branch is composed of a convolution module, a slice segmentation module and a UIB module; the fourth branch is composed of a convolution module, a slice segmentation module and two identical UIB modules; the four branches are concat-connected, and finally the convolution module processes the output; In the YOLOv10 construction waste intelligent sorting detection and recognition model, the convolution of the neck feature fusion part is composed of ordinary convolution, BN batch normalization layer and SiLU activation function; In order to achieve accurate and effective bounding box regression and further improve the detection accuracy and generalization ability of the model, the Focal-EIoU Loss function is used to train the model; In order to further improve the detection and recognition rate of the model, the Sophia optimizer is selected to approximate the Hessian matrix and use this information to update the parameters. In the target detection task, the model parameters are optimized, including the convolutional layer weights, normalization layer parameters, fully connected layer weights, and bias terms. Step 3: Send the images of the construction waste dataset to the YOLOv10 construction waste intelligent sorting detection and recognition model established in step 2 for training and verification, and use label smoothing processing during the training process to improve the robustness and generalization ability of the model and obtain the optimal weight; Step 4: After obtaining the optimal weights, perform a test, load the optimal weights, and send the images of the test set of the construction waste dataset into the YOLOv10 construction waste intelligent sorting detection and recognition model for testing.

2. The method for intelligent detection and identification of construction waste based on image feature enhancement according to claim 1, characterized in that: In step 1, the value of the motion blur convolution kernel is set to 15 to control the degree of blur, the mean of the noise processing is set to 0 and the standard deviation is set to 10 to generate noise, simulating the dust environment and the problem of blurred image features of some construction waste caused by conveyor belt vibration. The processed data set is labeled and divided into training set, validation set and test set according to the ratio of 8:1:

1.

3. The intelligent detection and identification method for construction waste based on image feature enhancement according to claim 1 or 2, characterized in that: In step 3, the number of training rounds is set to 300 and the batch size is set to 2 during the training process.