An AI-based intelligent sorting system for construction waste
By introducing a lightweight YOLO model and feature fusion strategy, the real-time performance and detection accuracy issues of the construction waste sorting system were resolved, improving the system's real-time response capability and detection accuracy, reducing the missed detection rate, and achieving efficient intelligent sorting of construction waste.
Patent Information
- Application Number
- CN202510616159.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Existing intelligent sorting systems for construction waste suffer from large detection models and slow inference speeds, making it difficult to meet real-time and high throughput requirements. Furthermore, they exhibit low detection accuracy and robustness in dense, multi-category, and small-target environments, resulting in high rates of missed and false detections, hindering efficient identification and rapid response.
We adopt a lightweight YOLO model based on the MobileViTv3 architecture, combining an efficient feature extraction network, dynamic convolution mechanism and feature reuse technology. We introduce a spatially aware dynamic feature fusion strategy and a redundant feature efficient reconstruction mechanism to improve the accuracy of feature extraction and detection and reduce the false negative and false positive rates.
It has improved the real-time response capability and processing efficiency of the intelligent sorting system for construction waste, significantly improved the detection accuracy and system robustness in dense, multi-category, and small-target environments, and reduced the rate of missed detections and false detections.
Smart Images

Figure CN120532770B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent waste sorting, specifically to an intelligent construction waste sorting system based on artificial intelligence. Background Technology
[0002] Currently, construction waste, as an important byproduct of urban construction, has made its intelligent recycling and reuse a key issue for sustainable urban development. Existing intelligent sorting systems for construction waste mainly rely on detection methods based on convolutional neural networks, which suffer from problems such as large traditional detection models and slow inference speeds, making it difficult to meet the real-time and high-throughput requirements of construction waste sorting lines. Secondly, when facing complex environments with dense, multi-category, and small targets, existing systems suffer from low detection accuracy and robustness, resulting in missed detections and false detections. This makes it difficult to balance efficient identification, accurate sorting, and rapid response in practical applications, severely restricting the improvement of the resource utilization level of construction waste. Summary of the Invention
[0003] To address the above issues and overcome the shortcomings of existing technologies, this invention provides an AI-based intelligent construction waste sorting system. Addressing the problems of large size and slow inference speed in traditional detection models, this invention introduces a lightweight YOLO model based on the MobileViTv3 architecture, combined with an efficient feature extraction network, dynamic convolution mechanism, and feature reuse technology. This achieves efficient feature extraction and rapid inference of construction waste, improving the real-time response capability and processing efficiency of the intelligent construction waste sorting system. To address the issues of low detection accuracy and robustness, this invention employs a spatially aware dynamic feature fusion strategy during feature extraction and multi-scale feature fusion, and introduces a redundant feature efficient reconstruction mechanism in the detection head design. This significantly improves the detection accuracy and system robustness of construction waste in dense, multi-category, small-target complex environments, while reducing the false negative and false positive rates.
[0004] The technical solution adopted in this invention is as follows: This invention provides an intelligent construction waste sorting system based on artificial intelligence, comprising a data acquisition module, an intelligent identification module, a sorting decision module, a sorting execution module, and an intelligent feedback optimization module, specifically including the following:
[0005] The data acquisition module integrates a high-resolution industrial camera, lidar, infrared sensor, and weight detection device to collect construction waste data, including images and videos of the construction waste;
[0006] The intelligent recognition module introduces a lightweight YOLO model based on the MobileViTv3 architecture as a waste detection sub-model. It uses an efficient feature extraction network, dynamic convolution mechanism and feature reuse technology to extract features and detect multiple targets in construction waste data, and obtain the recognition results of construction waste, including waste type, quantity and spatial distribution information.
[0007] The sorting decision module uses a sorting optimization algorithm that combines target particle swarm optimization and reinforcement learning to construct a multi-constraint optimization model based on the identification results to generate the optimal sorting strategy.
[0008] The sorting execution module includes a robotic arm, a conveyor belt, sorting rollers, and a pick sorter, which accurately handles and sorts construction waste according to the optimal sorting strategy.
[0009] The intelligent feedback optimization module collects and identifies confidence level, sorting success rate, and sorting error, and evaluates and provides feedback in real time on the running effect of the lightweight YOLO model based on the MobileViTv3 architecture.
[0010] Furthermore, in the data acquisition module, the construction waste data includes the following types: concrete blocks, brick fragments, metal components, plastic pipes, wooden formwork, glass fragments, ceramic products, and insulation materials.
[0011] Furthermore, the intelligent recognition module specifically includes the following steps:
[0012] Step S1: Input preprocessing, collect historical construction waste data, denoise, scale and enhance the construction waste data to obtain processed construction waste data;
[0013] Step S2: Construct a preliminary coding feature extractor. Based on the MobileViTv3 model architecture, introduce inter-layer skip connections and separable convolutional kernel design. Input the processed construction waste data into the lightweight feature extraction network to obtain the preliminary coding feature map, as shown below:
[0014] ;
[0015] in, This represents the initial encoded feature map. , For spatial resolution, This refers to the number of spatial channels;
[0016] Step S3: Dynamic perceptual feature fusion. The initial encoded feature map is subjected to spatial perceptual dynamic feature fusion processing to obtain the fused multi-scale feature map.
[0017] Step S4: Model building. Using YOLOv5 as the detection framework, embedding the MobileViTv3 feature extractor to build a lightweight YOLO model based on the MobileViTv3 architecture.
[0018] Step S5: Construction waste identification. A lightweight YOLO model based on the MobileViTv3 architecture is used to identify construction waste in the processed construction waste data to obtain the identification results.
[0019] Step S6: Result processing. Non-maximum suppression is applied to the identification results from step S5, and the identification results with high confidence are selected to obtain the identification results of construction waste.
[0020] Furthermore, step S3 specifically includes the following steps:
[0021] Step S31: Multi-branch response feature generation. Based on the preliminary encoded feature map, for each spatial location... Generate through multi-branch lightweight convolution operations Given 1 candidate response feature, generate a candidate response feature set, represented as follows:
[0022] ;
[0023] in, and These are the row and column indices for the spatial location, respectively. and These represent the height and width of the initial encoded feature map, respectively. The number of feature channels, Indicates the first The spatial location of candidate response features eigenvectors;
[0024] Step S32: Calculate the fusion weights. Introduce a spatial location-aware fusion weight mechanism. For each spatial location, calculate the fusion score through a lightweight attention encoding network and normalize it using a softmax operation to obtain the fusion weight of each candidate response feature at the corresponding spatial location. The formula used is as follows:
[0025] ;
[0026] in, For the first The spatial location of candidate response features Normalized fusion weights on The score indicates the spatial location;
[0027] Step S33: Multi-scale feature map generation. The candidate response features of each spatial location are weighted and combined according to the fusion weights to obtain the fused response features of all spatial locations, forming a multi-scale feature map, as shown below:
[0028] ;
[0029] in, These are the response characteristics after fusion.
[0030] Furthermore, step S4 specifically includes the following:
[0031] Infrastructure: A lightweight YOLO model based on the MobileViTv3 architecture is built using YOLOv5 as the basic detection framework;
[0032] Backbone Feature Extraction Network: The initial encoded feature extractor is used as the backbone feature extraction network, and multi-scale fused feature maps are extracted through lightweight convolution design;
[0033] Lightweight Neck: In the Neck part, an efficient separable convolution and lightweight feature fusion design are introduced. The efficient separable convolution adopts a depthwise separable convolution and a channel rearrangement mechanism to extract and reorganize information across channels. The lightweight feature fusion obtains the fused lightweight feature map by aggregating and compressing information from the multi-scale fused feature map.
[0034] Lightweight Head: The Head part introduces redundant features for efficient reconstruction design. Based on Ghost, a structural reparameterization mechanism is introduced. Combined with redundant convolution branches, the fused lightweight feature map is efficiently reconstructed and inference is accelerated to obtain the final detection output.
[0035] The beneficial effects achieved by the present invention using the above solution are as follows:
[0036] (1) In view of the problem that traditional detection models are large in size and slow inference speed, this invention introduces a lightweight YOLO model based on the MobileViTv3 architecture, combined with an efficient feature extraction network, dynamic convolution mechanism and feature reuse technology, to achieve efficient extraction and fast inference of construction waste features, thereby improving the real-time response capability and processing efficiency of the intelligent construction waste sorting system.
[0037] (2) To address the issues of low detection accuracy and robustness, this invention employs a spatially perceptive dynamic feature fusion strategy during feature extraction and multi-scale feature fusion, and introduces a redundant feature efficient reconstruction mechanism in the design of the detection head. This significantly improves the detection accuracy and system robustness of construction waste in complex environments with dense, multi-category, and small targets, and reduces the rate of missed detections and false detections. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of a module of an intelligent construction waste sorting system based on artificial intelligence proposed in this invention;
[0039] Figure 2 This is a flowchart of the intelligent recognition module proposed in this invention.
[0040] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation
[0041] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0042] Example 1, see Figure 1 This invention provides an intelligent construction waste sorting system based on artificial intelligence, comprising a data acquisition module, an intelligent identification module, a sorting decision module, a sorting execution module, and an intelligent feedback optimization module, specifically including the following:
[0043] The data acquisition module integrates a high-resolution industrial camera, lidar, infrared sensor, and weight detection device to collect construction waste data, including images and videos of the construction waste;
[0044] The intelligent recognition module introduces a lightweight YOLO model based on the MobileViTv3 architecture as a waste detection sub-model. It uses an efficient feature extraction network, dynamic convolution mechanism and feature reuse technology to extract features and detect multiple targets in construction waste data, and obtain the recognition results of construction waste, including waste type, quantity and spatial distribution information.
[0045] The sorting decision module uses a sorting optimization algorithm that combines target particle swarm optimization and reinforcement learning to construct a multi-constraint optimization model based on the identification results to generate the optimal sorting strategy.
[0046] The sorting execution module includes a robotic arm, a conveyor belt, sorting rollers, and a pick sorter, which accurately handles and sorts construction waste according to the optimal sorting strategy.
[0047] The intelligent feedback optimization module collects and identifies confidence level, sorting success rate, and sorting error, and evaluates and provides feedback in real time on the running effect of the lightweight YOLO model based on the MobileViTv3 architecture.
[0048] Example 2, based on the above examples, includes the following types of construction waste data in the data acquisition module: concrete blocks, brick fragments, metal components, plastic pipes, wooden formwork, glass fragments, ceramic products, and insulation materials.
[0049] Example 3, see Figure 2 This embodiment is based on the above embodiment, and the intelligent recognition module specifically includes the following steps:
[0050] Step S1: Input preprocessing, collect historical construction waste data, denoise, scale and enhance the construction waste data to obtain processed construction waste data;
[0051] Step S2: Construct a preliminary coding feature extractor. Based on the MobileViTv3 model architecture, introduce inter-layer skip connections and separable convolutional kernel design. Input the processed construction waste data into the lightweight feature extraction network to obtain the preliminary coding feature map, as shown below:
[0052] ;
[0053] in, This represents the initial encoded feature map. , For spatial resolution, This refers to the number of spatial channels;
[0054] Step S3: Dynamic perceptual feature fusion. The initial encoded feature map is subjected to spatial perceptual dynamic feature fusion processing to obtain the fused multi-scale feature map.
[0055] Step S4: Model building. Using YOLOv5 as the detection framework, embedding the MobileViTv3 feature extractor to build a lightweight YOLO model based on the MobileViTv3 architecture.
[0056] Step S5: Construction waste identification. A lightweight YOLO model based on the MobileViTv3 architecture is used to identify construction waste in the processed construction waste data to obtain the identification results.
[0057] Step S6: Result processing. Non-maximum suppression is applied to the identification results from step S5, and the identification results with high confidence are selected to obtain the identification results of construction waste.
[0058] Example 4, based on the above examples, specifically includes the following steps in step S3:
[0059] Step S31: Multi-branch response feature generation. Based on the preliminary encoded feature map, for each spatial location... Generate through multi-branch lightweight convolution operations Given 1 candidate response feature, generate a candidate response feature set, represented as follows:
[0060] ;
[0061] in, and These are the row and column indices for the spatial location, respectively. and These represent the height and width of the initial encoded feature map, respectively. The number of feature channels, Indicates the first The spatial location of candidate response features eigenvectors;
[0062] Step S32: Calculate the fusion weights. Introduce a spatial location-aware fusion weight mechanism. For each spatial location, calculate the fusion score through a lightweight attention encoding network and normalize it using a softmax operation to obtain the fusion weight of each candidate response feature at the corresponding spatial location. The formula used is as follows:
[0063] ;
[0064] in, For the first The spatial location of candidate response features Normalized fusion weights on The score indicates the spatial location;
[0065] Step S33: Multi-scale feature map generation. The candidate response features of each spatial location are weighted and combined according to the fusion weights to obtain the fused response features of all spatial locations, forming a multi-scale feature map, as shown below:
[0066] ;
[0067] in, These are the response characteristics after fusion.
[0068] In this embodiment, the code used is as follows:
[0069] import torch
[0070] import torch.nn as nn
[0071] import torch.nn.functional as F
[0072] class MultiBranchFeatureFusion(nn.Module):
[0073] def __init__(self, in_channels, out_channels, n_branches):
[0074] super(MultiBranchFeatureFusion, self).__init__()
[0075] self.n_branches = n_branches
[0076] # Multi-branch lightweight convolution operation
[0077] self.conv_branches = nn.ModuleList([nn.Conv2d(in_channels,out_channels, kernel_size=3, padding=1) for _ in range(n_branches)])
[0078] # Lightweight attention encoding network (for calculating fusion weights)
[0079] self.attention_network = nn.Conv2d(out_channels, 1, kernel_size=1)
[0080] def forward(self, x):
[0081] # Step S31: Generation of Multi-Branch Response Features
[0082] branches = []
[0083] for conv in self.conv_branches:
[0084] branches.append(conv(x)) # Each convolutional branch
[0085] branches = torch.stack(branches, dim=1) # (batch_size, n_branches, C, H, W)
[0086] # Step S32: Calculate the fusion weights
[0087] # Calculate the fusion score for each response feature
[0088] fusion_scores = []
[0089] for i in range(self.n_branches):
[0090] score = self.attention_network(branches[:, i, :, :, :]) # Calculate the attention score for each branch
[0091] fusion_scores.append(score)
[0092] fusion_scores = torch.stack(fusion_scores, dim=1) # (batch_size, n_branches, 1, H, W)
[0093] # Use softmax for normalization
[0094] fusion_weights = F.softmax(fusion_scores, dim=1) # (batch_size, n_branches, 1, H, W)
[0095] # Step S33: Multi-scale feature map generation
[0096] # The candidate response features at each spatial location are weighted and combined according to the fusion weight.
[0097] fused_features = torch.sum(branches * fusion_weights, dim=1)# (batch_size, C, H, W)
[0098] return fused_features
[0099] # Assume the input feature map size is (batch_size, C, H, W)
[0100] batch_size, C, H, W = 8, 64, 32, 32
[0101] input_tensor = torch.randn(batch_size, C, H, W)
[0102] # Initialize and apply the network
[0103] model = MultiBranchFeatureFusion(in_channels=C, out_channels=128, n_branches=4)
[0104] output = model(input_tensor)
[0105] print(output.shape) # Size of the output feature map.
[0106] Example 5, based on the above examples, specifically includes the following in step S4:
[0107] Infrastructure: A lightweight YOLO model based on the MobileViTv3 architecture is built using YOLOv5 as the basic detection framework;
[0108] Backbone Feature Extraction Network: The initial encoded feature extractor is used as the backbone feature extraction network, and multi-scale fused feature maps are extracted through lightweight convolution design;
[0109] Lightweight Neck: In the Neck part, an efficient separable convolution and lightweight feature fusion design are introduced. The efficient separable convolution adopts a depthwise separable convolution and a channel rearrangement mechanism to extract and reorganize information across channels. The lightweight feature fusion obtains the fused lightweight feature map by aggregating and compressing information from the multi-scale fused feature map.
[0110] Lightweight Head: The Head part introduces redundant features for efficient reconstruction design. Based on Ghost, a structural reparameterization mechanism is introduced. Combined with redundant convolution branches, the fused lightweight feature map is efficiently reconstructed and inference is accelerated to obtain the final detection output.
[0111] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0112] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
[0113] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.
Claims
1. An intelligent sorting system for construction waste based on artificial intelligence, characterized in that: It includes a data acquisition module, an intelligent recognition module, a sorting decision module, a sorting execution module, and an intelligent feedback optimization module, specifically including the following: The data acquisition module integrates a high-resolution industrial camera, lidar, infrared sensor, and weight detection device to collect construction waste data, including images and videos of the construction waste; The intelligent recognition module introduces a lightweight YOLO model based on the MobileViTv3 architecture as a waste detection sub-model. It uses an efficient feature extraction network, dynamic convolution mechanism and feature reuse technology to extract features and detect multiple targets in construction waste data, and obtain the recognition results of construction waste. The sorting decision module uses a sorting optimization algorithm that combines target particle swarm optimization and reinforcement learning to construct a multi-constraint optimization model based on the identification results to generate the optimal sorting strategy. The sorting execution module accurately transports and sorts construction waste according to the optimal sorting strategy; The intelligent feedback optimization module collects and identifies confidence level, sorting success rate and sorting error, and evaluates and provides feedback on the running effect of the lightweight YOLO model based on the MobileViTv3 architecture in real time. The intelligent recognition module specifically includes the following steps: Step S1: Input preprocessing, collect historical construction waste data, denoise, scale and enhance the construction waste data to obtain processed construction waste data; Step S2: Construct a preliminary coding feature extractor. Based on the MobileViTv3 model architecture, introduce inter-layer skip connections and separable convolutional kernel design. Input the processed construction waste data into the lightweight feature extraction network to obtain the preliminary coding feature map, as shown below: ; in, This represents the initial encoded feature map. , For spatial resolution, This refers to the number of spatial channels; Step S3: Dynamic perceptual feature fusion. The initial encoded feature map is subjected to spatial perceptual dynamic feature fusion processing to obtain the fused multi-scale feature map. Step S4: Model building. Using YOLOv5 as the detection framework, embedding the MobileViTv3 feature extractor to build a lightweight YOLO model based on the MobileViTv3 architecture. Step S5: Construction waste identification. A lightweight YOLO model based on the MobileViTv3 architecture is used to identify construction waste in the processed construction waste data to obtain the identification results. Step S6: Result processing. Non-maximum suppression is applied to the identification results from step S5, and the identification results with high confidence are selected to obtain the identification results of construction waste.
2. The intelligent construction waste sorting system based on artificial intelligence according to claim 1, characterized in that: Step S3 specifically includes the following steps: Step S31: Multi-branch response feature generation. Based on the preliminary encoded feature map, for each spatial location... Generate through multi-branch lightweight convolution operations Given 1 candidate response feature, generate a candidate response feature set, represented as follows: ; in, and These are the row and column indices for the spatial location, respectively. and These represent the height and width of the initial encoded feature map, respectively. The number of feature channels, Indicates the first The spatial location of candidate response features eigenvectors; Step S32: Calculate the fusion weights. Introduce a spatial location-aware fusion weight mechanism. For each spatial location, calculate the fusion score through a lightweight attention encoding network and normalize it using a softmax operation to obtain the fusion weight of each candidate response feature at the corresponding spatial location. The formula used is as follows: ; in, For the first The spatial location of candidate response features Normalized fusion weights on The score indicates the spatial location; Step S33: Multi-scale feature map generation. The candidate response features of each spatial location are weighted and combined according to the fusion weights to obtain the fused response features of all spatial locations, forming a multi-scale feature map, as shown below: ; in, These are the response characteristics after fusion.
Citation Information
Patent Citations
Hot rolled steel strip defect detection algorithm based on YOLO algorithm
CN117593258A
Winter jujube detection and positioning and mechanical arm picking sequence planning method based on YOLO-MLG and YAGR methods
CN118636150A