A method and system for intelligent sorting of construction waste using multi-source information fusion

Through multi-source information fusion technology, multi-sensors are used to obtain multimodal data of construction waste, combined with Kalman filtering and attention mechanism, and classification convolutional neural networks are used for waste identification and sorting, which solves the problem of efficient and accurate sorting of construction waste in existing technologies and realizes intelligent sorting and resource processing.

CN120451717BActive Publication Date: 2025-10-03GUANGXI QINGHUI ENVIRONMENTAL PROTECTION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510497561.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-10-03
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

Existing construction waste sorting technology relies on manual sorting or a single sensor, which makes it difficult to efficiently and accurately identify multiple categories of waste, resulting in low sorting efficiency and accuracy, and lack of adaptability. It cannot remain stable in diverse environments and cannot achieve high-precision real-time sorting, which limits the recycling rate and sorting efficiency of construction waste.

Method used

A multi-source information fusion method is adopted to arrange multiple sensors to obtain visible light images, three-dimensional information, spectral data, weight and vibration signals in real time. Kalman filtering and attention mechanism are combined for data processing, and classification convolutional neural network is used to identify waste types, and the sorting actuator is controlled for precise delivery.

Benefits of technology

It has greatly improved the automation efficiency and accuracy of construction waste sorting, can operate stably in complex environments, realize the intelligent sorting of waste such as concrete blocks, metals, and plastics, reduce labor costs and environmental burdens, and improve resource processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451717B_ABST
    Figure CN120451717B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for intelligently sorting construction waste using multi-source information fusion. The system deploys multiple sensors of varying types on a conveyor belt to acquire real-time multi-source data, including visible light images, three-dimensional information, spectral data, weight, and vibration signals. Kalman filtering is then used to perform denoising and feature extraction. A multimodal fusion algorithm, combined with an attention mechanism, weightedly fuses the features of different modalities to produce a fused feature vector. This fused feature vector is then input into a classification convolutional neural network, which outputs the type and confidence level of the construction waste. The confidence level is then used to control the sorting actuator for precise sorting. This system can efficiently identify and sort different types of waste, such as concrete blocks, metal, wood, and plastic, in a variety of complex environments, significantly improving automatic sorting efficiency and accuracy, and possessing high industrial application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-source information processing, and specifically to a method and system for intelligent sorting of multi-source information fusion of construction waste. Background Art

[0002] With the accelerating pace of urbanization and industrialization, the amount of construction waste generated by building construction and demolition activities has increased rapidly. Construction waste is complex and includes concrete blocks, bricks and tiles, steel bars, metal components, glass, plastics, wood, and composite materials. These materials often exhibit significant differences in shape, density, texture, and spectral characteristics. Without efficient recycling and reuse methods, this waste not only occupies significant land and landfill resources, but its degradation or disposal processes may also cause secondary pollution to the environment, increasing the socioeconomic burden.

[0003] Existing construction waste sorting and processing technologies often rely on manual sorting or data collection based on a single sensor. Although manual sorting has a certain degree of flexibility, the efficiency and accuracy of manual operations are difficult to guarantee due to factors such as the huge amount of waste, the diverse composition, and the harsh working environment. At the same time, workers are easily exposed to harsh working conditions such as dust and noise for a long time, which can easily affect their health. Automatic sorting methods based on a single sensor, such as simple image processing or magnetic separation processes, are difficult to take into account the complex sorting needs of multiple categories. On the one hand, a single image or spectral information cannot fully identify the subtle differences between different materials; on the other hand, if it is based solely on metal magnetic separation, it cannot effectively distinguish non-metallic waste, resulting in low sorting efficiency and accuracy.

[0004] In recent years, with the development of artificial intelligence and multi-sensor fusion technologies, the application of multi-source data to construction waste sorting has become a research hotspot. For example, visible light imagery can be used to identify the appearance of waste, depth imagery can be used to analyze its geometry, and spectral sensors can be used to detect material characteristics. This, combined with supplementary information such as weight or vibration, allows for the characterization and determination of waste properties from multiple perspectives. This multi-source information collection and fusion approach helps improve accuracy and robustness in the identification process. However, the following bottlenecks are still faced in practical applications: First, the real-time registration and fusion algorithms of multi-source data are relatively complex, and traditional data fusion methods often lack adaptability and cannot remain stable in diverse environments; second, how to highlight key modalities or key feature information in the fusion process to more accurately classify different types of waste is still a difficulty; third, in the sorting process, automated sorting operations are required based on the recognition results, such as precise control of robotic arms or pneumatic valves, which need to be linked with the classification algorithm to achieve efficient and stable waste transportation and delivery; and the existing multi-source data processing does not combine different data sources according to the importance of the data and the advantageous processing performance of the algorithm adopted, resulting in low recognition accuracy, and cannot achieve high-precision recognition and real-time sorting of various types of construction waste, and leads to obstacles in the widespread and large-scale promotion of its application in demolition sites or centralized waste treatment centers, and cannot significantly improve the recycling rate and sorting efficiency of construction waste, thereby reducing labor costs and environmental burdens. Summary of the Invention

[0005] To address the aforementioned issues in the prior art, the present invention provides a method and system for intelligently sorting construction waste using multi-source information fusion. This method involves deploying multiple sensors of varying types on a conveyor belt to acquire real-time multi-source data, including visible light images, three-dimensional information, spectral data, weight, and vibration signals. Kalman filtering is then used to perform feature extraction and denoising. A multimodal fusion algorithm, combined with an attention mechanism, weightedly fuses the features of different modalities to produce a fused feature vector. This fused feature vector is then input into a classification convolutional neural network, which outputs the type and confidence level of the construction waste. The confidence level is then used to control the sorting actuator for precise sorting. This method can efficiently identify and sort different types of waste, such as concrete blocks, metal, wood, and plastic, in a variety of complex environments, significantly improving the efficiency and accuracy of automatic sorting and possessing high industrial application value.

[0006] This application provides a method for intelligent sorting of construction waste by multi-source information fusion, comprising the following steps:

[0007] S1: Use multiple different types of sensors to monitor construction waste on the conveyor belt in real time to obtain multi-source data;

[0008] S2: Preprocess multi-source data;

[0009] S3: Extract features from multi-source data and fuse them using a multimodal fusion algorithm. The fusion algorithm introduces an attention mechanism to weight the contributions of multi-source data of different modalities. Based on the relevance of each modality, weights are adaptively assigned to obtain a fused feature vector.

[0010] S4: Input the fused feature vector into the trained classification convolutional neural network model to output different types of construction waste and their corresponding sorting priorities, and provide the confidence level of the classification results;

[0011] S5: According to the type of construction waste and the classification confidence, the corresponding sorting actuators are controlled to send different types of waste to the corresponding recycling and processing areas. The sorting actuators include robotic arms, pneumatic sorting valves, and mobile trash cans.

[0012] Preferably, the multi-source data includes: an RGB camera to obtain visible light images, a depth camera to obtain three-dimensional information, a spectral sensor to obtain material spectral data, a weight sensor to obtain material weight information, and a vibration sensor to obtain vibration signals; the RGB camera, the depth camera, the spectral sensor, and the weight sensor are respectively installed at different positions and / or different angles of the conveyor belt to achieve multi-angle collection of construction waste; wherein, the RGB camera is used to obtain visible light images and perform surface texture recognition, the depth camera is used to extract the three-dimensional geometric structure of the material, the spectral sensor is used to obtain the spectral reflection characteristics of the material contained in the waste and distinguish different materials, the weight sensor obtains material weight information by detecting the local pressure of the conveyor belt, and the vibration sensor monitors the vibration signal during the transportation process to detect internal defects or loose structures.

[0013] Preferably, the preprocessing includes image denoising on the RGB image and three-dimensional information, and filtering the spectral data, weight data and vibration data using a Kalman filter algorithm.

[0014] Preferably, step S3 includes:

[0015] S3-1: performing feature extraction on the multi-source data respectively to obtain multimodal features including visible light image features, three-dimensional geometric structure features, spectral features, weight features, and vibration features;

[0016] S3-2: The extracted features are input into the attention mechanism module, which calculates the attention weight of each modality feature in the fusion process based on the correlation of each modality feature;

[0017] S3-3: Dynamically adjust the attention weights and perform weighted operations on different modal features to form a fused feature vector.

[0018] Preferably, the multimodal feature acquisition in step S3-1 includes: inputting the visible light image into a feature extraction convolutional neural network, extracting surface texture features as visible light image features through multi-layer convolution, pooling and activation operations; converting the three-dimensional data acquired by the depth camera into a point cloud or depth map, and inputting the converted data into a three-dimensional convolutional neural network or a point cloud network to extract spatial structure features as three-dimensional geometric structure features; performing a wavelet frequency domain transform on the spectral data to obtain representative spectral components as spectral features; performing feature extraction on the weight data and the vibration data respectively to obtain mean, variance, kurtosis and bandpass energy statistical features, thereby obtaining weight features and vibration features;

[0019] Step S3-2 includes: constructing and inputting an attention mechanism module, using the extracted multimodal features to construct a key, a query, and a value, and inputting them into an attention mechanism module including a multi-head attention unit. In each attention head, the similarity between the key and the query is calculated based on cosine similarity, and the results are normalized to obtain the attention weight of each modality;

[0020] The step S3-3 includes: weighting each modal feature according to the attention weight and then splicing them to obtain a fused feature vector.

[0021] Preferably, S5-1: comparing the classification confidence with a preset confidence threshold; if the classification confidence is higher than or equal to the confidence threshold, adopting the classification result; if the classification confidence is lower than the confidence threshold, marking the construction waste as "abnormal" and sending it to a corresponding manual processing area;

[0022] S5-2: Matching the adopted classification results with sorting equipment, and selecting at least one sorting actuator from among a robotic arm, a pneumatic sorting valve, or a mobile trash can according to the type of construction waste and its volume, shape, or density characteristics;

[0023] S5-3: Use the selected sorting actuator to grab, push or suction the construction waste to be sorted and guide it to the corresponding recycling and processing area.

[0024] Preferably, the activation function f1(x) used in the classification convolutional neural network model is expressed as follows:

[0025]

[0026] Wherein, x represents the input value of the current neuron, e is the base of the natural logarithm; a1 represents the attention weight corresponding to the visible light image feature calculated by step S3-2, and a2 represents the attention weight corresponding to the three-dimensional geometric structure feature calculated by step S3-2.

[0027] This application also provides a construction waste multi-source information fusion intelligent sorting system, including:

[0028] The acquisition module uses multiple different types of sensors to monitor construction waste on the conveyor belt in real time and obtain multi-source data;

[0029] Preprocessing module, preprocesses multi-source data;

[0030] The fusion feature vector acquisition module extracts features from multi-source data and fuses the multi-source data using a multimodal fusion algorithm. The fusion algorithm introduces an attention mechanism to weight the contributions of multi-source data of different modalities. Based on the correlation of the data of each modality, the weights are adaptively assigned to obtain a fused feature vector.

[0031] The calculation module inputs the fused feature vector into the trained classification convolutional neural network model, outputs different types of construction waste and their corresponding sorting priorities, and gives the confidence level of the classification results;

[0032] The actuator operation module controls the corresponding sorting actuator according to the type of construction waste and the classification confidence, and sends different types of waste to the corresponding recycling and processing areas. The sorting actuator includes a robotic arm, a pneumatic sorting valve, and a mobile trash can.

[0033] Preferably, the multi-source data includes: an RGB camera to obtain visible light images, a depth camera to obtain three-dimensional information, a spectral sensor to obtain material spectral data, a weight sensor to obtain material weight information, and a vibration sensor to obtain vibration signals; the RGB camera, the depth camera, the spectral sensor, and the weight sensor are respectively installed at different positions and / or different angles of the conveyor belt to achieve multi-angle collection of construction waste; wherein, the RGB camera is used to obtain visible light images and perform surface texture recognition, the depth camera is used to extract the three-dimensional geometric structure of the material, the spectral sensor is used to obtain the spectral reflection characteristics of the material contained in the waste and distinguish different materials, the weight sensor obtains material weight information by detecting the local pressure of the conveyor belt, and the vibration sensor monitors the vibration signal during the transportation process to detect internal defects or loose structures.

[0034] Preferably, the preprocessing includes image denoising on the RGB image and three-dimensional information, and filtering the spectral data, weight data and vibration data using a Kalman filter algorithm.

[0035] The present invention provides a method and system for intelligent sorting of construction waste by integrating multi-source information, which can achieve the following beneficial technical effects:

[0036] 1. This invention simultaneously collects multimodal data, including visible light images, three-dimensional structures, spectral data, weight information, and vibration signals. During the fusion process, an attention mechanism is introduced to weight key modalities. This allows accurate identification of the material and appearance characteristics of construction waste of various types and shapes, reducing blind spots caused by reliance on a single sensor and significantly improving sorting accuracy. By assessing the confidence of the classification results and combining them with sorting priorities, the invention can control robotic arms, pneumatic sorting valves, or mobile trash cans to rapidly sort and release various types of waste. This enables intelligent sorting of various types of construction waste, including concrete blocks, metals, plastics, and wood, shortening processing cycles and reducing labor costs and operational risks.

[0037] 2. This invention utilizes Kalman filtering and multimodal feature extraction algorithms in the preprocessing stage, effectively suppressing noise and filtering out interference. It also uses an attention mechanism to adaptively adjust the weights of each modality during the fusion process, enabling stable operation under complex conditions such as variable ambient lighting, diverse waste sizes, and high noise levels. This invention can be easily integrated with existing conveyor systems, robotic grippers, and pneumatic sorting devices, and deployed in diverse building demolition or centralized waste processing scenarios, encompassing a wider range of applications and helping to achieve efficient recycling and resource processing of construction waste.

[0038] 3. By incorporating attention weights corresponding to visible light image modalities into the activation function calculation of a convolutional neural network, this invention further amplifies the influence of visible light image features during the critical forward propagation phase of the model. This enables the network to more sensitively capture detailed textures and lighting differences when distinguishing waste with high appearance similarity, significantly improving classification accuracy. In traditional convolutional network structures, the activation function is insensitive to the scale or modal information of the input signal. However, with the introduction of attention weights, the differences between the visible light modality and other modalities are more directly quantified in the activation output. This not only enables the network to dynamically "amplify" important modal features, but also provides greater interpretability for subsequent visual analysis and diagnosis. Because attention weights are dynamically calculated based on the correlation and data quality of each modal data, the network's activation output can be adjusted in real time based on factors such as external ambient lighting, material type, and waste shape. This adaptive mechanism enables the sorting system to maintain high robustness and accuracy even in complex and changing scenarios. By incorporating attention weights into the activation function, the neural network can converge more quickly on the key range of visible light image features during training, thereby reducing over-learning of redundant or noisy features and helping the model achieve good results even with limited training samples. Furthermore, the additional "amplification" effect of visible light features also provides a more stable gradient update signal for backpropagation. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0040] Figure 1 This is a flowchart of the steps of a method for intelligent sorting of construction waste by multi-source information fusion according to the present invention;

[0041] Figure 2 This is a schematic diagram of a multi-source information fusion intelligent sorting system for construction waste according to the present invention. DETAILED DESCRIPTION

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0043] Example 1:

[0044] In order to solve the above-mentioned problems in the prior art, the following Figure 1 As shown: This application provides a method for intelligent sorting of construction waste by multi-source information fusion, including the following steps:

[0045] S1: Use multiple different types of sensors to monitor construction waste on the conveyor belt in real time to obtain multi-source data.

[0046] In one embodiment, a conveyor belt system is installed at a construction waste processing center to transport various types of demolition and construction waste. To collect multi-source information about the construction waste on the conveyor belt, this embodiment deploys RGB cameras, depth cameras, spectral sensors, weight sensors, and vibration sensors at various locations and / or angles along the conveyor belt. The RGB cameras are mounted approximately 0.5 to 1 meter above the conveyor belt to capture visible light images, ensuring that the lens covers the width of the conveyor belt and captures material surface information.

[0047] During the acquisition process, the RGB camera captures visible light images at a rate of approximately 20 frames per second, identifies the texture features and color characteristics of the target waste, and outputs image data to the control system at the end of the conveyor belt or after multi-segment light source calibration. Depth camera installation and three-dimensional information acquisition: A depth camera is installed at a certain angle to the RGB camera (such as a 30° to 45° tilt) to capture the three-dimensional structure of the waste from different perspectives. The camera acquires three-dimensional information of the waste in real time, such as point cloud or depth image data, and transmits it to the data fusion unit for subsequent identification of the volume, shape, and stacking of the material.

[0048] Installation of spectral sensor and collection of spectral data: The spectral sensor is installed along the side of the conveyor belt, maintaining an appropriate distance from the waste (about 10 to 30 cm) to ensure that the spectral reflection information of the material can be captured when the material passes through the sensor's field of view. Through multi-channel acquisition in a specific wavelength range (such as visible light to near-infrared band), the sensor distinguishes the spectral reflection differences of different materials (such as concrete, metal, plastic, wood, etc.) in the characteristic band, and synchronously transmits the spectral data to the data processing module. Installation of weight sensor and collection of weight data: The weight sensor is placed on the support structure under the conveyor belt to monitor the load-bearing changes in a short period of time. When the material passes through the sensing area of ​​the weight sensor, the sensor can obtain the local pressure or total pressure value, quantify the material weight signal and output it to the control system, so as to combine the image and spectral information to further determine the density or material characteristics of the waste.

[0049] Vibration sensor installation and signal acquisition: Vibration sensors can be mounted on the conveyor belt or fixed to the sensor frame to monitor the vibration characteristics of the conveyor belt during operation or between the material and the belt surface. When waste with internal cavities, loose structures, or a high degree of fragmentation passes through, the vibration waveform exhibits dynamic characteristics different from those of normal solid materials, which can subsequently assist the sorting algorithm in identifying potentially abnormal or brittle materials. With this sensor arrangement, each sensor communicates with the data acquisition controller via industrial Ethernet or a high-speed serial port. A synchronized triggering mechanism ensures that all sensors collect data from the same waste batch at the same time. The acquired RGB images, 3D information, spectral data, weight data, and vibration signals are uniformly time-stamped and processed for the next level of data preprocessing and fusion. This multi-angle, multi-modal, high-resolution real-time monitoring solution enables the system to more accurately capture material characteristics in complex construction waste sorting scenarios, ensuring the stability and accuracy of subsequent identification and sorting steps.

[0050] S2: Preprocessing the multi-source data; In some embodiments, to ensure the accuracy of subsequent fusion and classification, the collected multi-source data needs to be preprocessed first, including denoising the visible light image and three-dimensional information, and filtering the spectral data, weight data, and vibration data using the Kalman filter algorithm, as follows: Denoising the visible light image and three-dimensional information, input the image frame acquired by the RGB camera into the image denoising module. Preferably, for different noise types (such as random noise, speckle noise, or high ISO noise under low light conditions), median filtering, bilateral filtering, or image denoising algorithms based on convolutional neural networks can be used for denoising. Distortion correction and hole filling are performed on the three-dimensional information captured by the depth camera (such as a depth map or point cloud). For example, an interpolation algorithm is used to fill in the missing pixels in the depth map; or a radius search threshold is used to remove isolated noise points in the point cloud data to obtain a more continuous and realistic three-dimensional shape.

[0051] Kalman filtering of spectral data: Time-series spectral data collected by the spectral sensor is input into the Kalman filter. In the initial stage, the initial state vector and measurement noise variance matrix are set to ensure a reasonable estimate of the sensor's measurement uncertainty. The state estimate and error covariance are iteratively updated at each moment to dynamically correct for random noise and fluctuations in the spectral curve. This smooths out more accurate and representative spectral features, facilitating subsequent material identification. Kalman filtering of weight and vibration data: Weight data is typically represented by continuous mechanical readings, which can be subject to transient interference during conveyor belt operation or material collisions.

[0052] Through Kalman filtering, short-term impulse noise can be suppressed to a low level, resulting in a more robust weight estimate. Kalman filtering is also performed on the time series signal from the vibration sensor to track in real time changes in the vibration waveform that may be caused by the degree of waste fragmentation or loose structure, thereby reducing interference from environmental vibration or mechanical resonance. Through the above-mentioned multi-source preprocessing steps, the present invention can significantly improve the quality and stability of raw data in high-noise, dynamically changing scenarios, providing a more reliable input basis for the subsequent training and inference of multimodal fusion and classification convolutional neural networks.

[0053] S3: Feature extraction is performed on the multi-source data. The multi-source data is fused using a multimodal fusion algorithm. The fusion algorithm introduces an attention mechanism to weight the contributions of the multi-source data of different modalities. Based on the correlation of the data of each modality, the weights are adaptively assigned to obtain a fused feature vector. In some embodiments, this embodiment extracts features from the visible light image, three-dimensional geometric information, spectral data, weight data, and vibration data obtained from the preprocessed multi-source data, and introduces an attention mechanism to perform multimodal weighted fusion to obtain a fused feature vector for subsequent classification and sorting decisions. Multimodal feature extraction (step S3-1) is as follows:

[0054] 1) Visible light image feature extraction: The noise-processed RGB image is fed into a feature extraction convolutional neural network (CNN). The CNN consists of multiple convolutional layers, pooling layers, and nonlinear activation layers (such as Reluctant Units (ReLUs). In each convolutional layer, a 3×3 filter is applied to extract local texture information. The pooling layer downsamples the feature map to reduce computational effort and enhance feature robustness. Finally, the global pooling output of the last convolutional layer or the output of the Flatten layer is used as the visible light image feature vector, effectively characterizing key information such as surface texture, edges, and color of the construction waste.

[0055] 2) 3D geometric structure feature extraction: Convert the 3D information acquired by the depth camera into point cloud data or depth map format. For point cloud data, a point cloud network (such as PointNet or PointNet++) can be used to learn the coordinates of each point and its local geometric relationship. For depth maps, a 3D convolutional neural network (3D-CNN) can be input for voxelization or feature extraction based on 2D slices. Taking PointNet as an example, a multi-layer perceptron (MLP) is used to map the coordinates (x, y, z) of each point in the point cloud to obtain global and local geometric features. The maximum pooling layer is then used to extract the overall material spatial contour features, ultimately outputting a 3D geometric structure feature vector.

[0056] 3) Spectral feature extraction: Spectral data collected by the spectral sensor is subjected to a wavelet transform (e.g., wavelet packet decomposition) to separate noise components while retaining information from key spectral bands. The energy or characteristic components of each level of coefficients obtained from the wavelet frequency domain transform are extracted, and frequency bands with minimal impact on construction waste identification are removed. The most recognizable characteristic components are retained and concatenated into a spectral feature vector, which is used to distinguish reflectance differences between different materials in key bands.

[0057] 4) Weight and vibration characteristics: A simple sliding window capture is performed on the time series data from the weight sensor to statistically determine parameters such as mean, variance, and maximum value. These parameters are combined with the denoised values ​​from the preprocessing stage to form a weight feature vector. Frequency domain or envelope analysis is performed on the time series signals from the vibration sensor to extract metrics such as kurtosis, bandpass energy, RMS value, and frequency domain characteristics to characterize the vibration pattern and structural stability of the waste during operation. These statistics are then assembled into a vibration feature vector.

[0058] In some embodiments, multimodal features have been extracted in advance and visible light image features, three-dimensional geometric features, spectral features, weight features, and vibration features have been obtained respectively. In order to fully integrate these features and highlight key modalities, we use the following steps to construct and input features in the attention mechanism module: Preparation of mapping of Key, Query, and Value. First, in order for the attention mechanism to be able to effectively compare and extract information between multi-source data, it is necessary to map the feature vectors of different modalities accordingly. The system applies a set of trainable linear transformations to each modal feature vector and outputs three sets of vectors, representing "key", "query", and "value". These vectors will respectively assume the functions of comparison, matching, and information carrying in subsequent processes to achieve the measurement of similarity and importance between multiple modalities.

[0059] Parallel calculation of multi-head attention units. In this embodiment, the attention mechanism is designed as a "multi-head" structure, which is intended to allow the system to compare the features of each modality from different subspaces or perspectives in parallel. Specifically, each attention "head" independently compares the "query" vector and the "key" vector, and evaluates the degree of match between them. If the features of a certain modality show a high degree of match in the current subspace, the modality will receive higher attention in the "head"; if the degree of match is low, the attention will be reduced accordingly. This parallel design allows the system to cover feature information from multiple angles and levels at the same time, enhancing the overall fusion accuracy.

[0060] Similarity evaluation and attention weight calculation: To measure the correlation or similarity between features of different modalities, each attention "head" compares the "query" with the "key" and generates a "score" or "attention score" based on this. A high score indicates that the current query is more inclined to absorb more information from this modality. These attention scores are converted into attention weights after a unified normalization step. The larger the attention weight, the more feature information corresponding to the modality should be retained during the fusion process; if the weight is smaller, it means that the importance or credibility of the modality in the current scenario is relatively low. Weighted synthesis of numerical information: Each attention "head" will perform a weighted synthesis of the aforementioned weights with the numerical vectors of each modality to obtain several weighted output information. The results output by different "heads" are usually spliced ​​or superimposed to form the "multi-head" fusion feature required by this embodiment. Through this weighted method, the attention mechanism module can maintain the richness of multi-source data while highlighting the modal features that contribute more to the recognition effect and reducing the interference of noise or secondary information on the system.

[0061] Output the fused feature vector, and finally, merge the weighted results of all attention "heads" to form an aggregated fused feature vector, which comprehensively reflects the key information of construction waste in visible light, depth, spectrum, weight, vibration and other modalities. This fused feature vector can be further used by subsequent classification models or other discriminant logics, so as to achieve higher accuracy and stability in the recognition and sorting links. In summary, through the construction of the attention mechanism module and the feature input process described in this embodiment, the system can flexibly allocate attention among different modal data, automatically highlight modal information with more discriminant value, and realize efficient weighted fusion of multi-source data. This solution has good robustness and adaptability in complex and changeable construction waste sorting scenarios, and provides better feature support for subsequent deep learning classification.

[0062] S4: The fused feature vector is input into a trained classification convolutional neural network model, which outputs different types of construction waste and their corresponding sorting priorities, along with confidence scores for the classification results. In some embodiments, after completing multimodal feature extraction and fusion with the attention mechanism, a fused feature vector is obtained, containing information such as visible light images, 3D geometric information, spectral data, weight, and vibration. This fused feature vector is then input into a pre-trained convolutional neural network (CNN) model to distinguish the types of construction waste and assign corresponding sorting priorities and confidence scores. The following is an example of a specific network structure and workflow: Overall network structure design: Input layer: Receives the fused feature vector output from the previous stage. This vector is typically one-dimensional and can be treated as a "pseudo-image" of a channel, or directly fed into a fully connected layer, depending on design requirements. Convolution and pooling layers: In this embodiment, to ensure that the CNN can still extract local correlations and discriminative information in the high-dimensional space after feature fusion, two to three convolutional layers are used after the input layer. Each convolutional layer uses a small convolution kernel (e.g., 3×3) to extract new local patterns or subtle differences. Pooling layers are inserted between convolutional layers to reduce dimensionality and enhance feature robustness. Feature Transformation Layer / Fully Connected Layer: After convolution and pooling, one or two fully connected layers aggregate the local features extracted by convolution into a more compact representation, preparing for the final classification and priority output. Output Layer: One or more neurons are used to output the predicted construction waste type, and additional neurons or output channels can be used to represent sorting priorities. The confidence level of the classification result is typically given by a Softmax function or similar probabilistic output mechanism.

[0063] 1) Input fused features: The fused feature vector is treated as the initial input to the network, replacing the usual image input. Because the multimodal data has been effectively integrated and reduced in dimension in the early stages, the CNN no longer needs to process the features of different modalities separately, but instead treats them as integrated high-dimensional information.

[0064] 2) Convolution extracts latent patterns. In the first convolution layer, setting a certain number of convolution kernels (e.g., 32 or 64) allows the network to identify local differences or subtle features that may be present in the fused features. Subsequently, the pooling layer downsamples the convolution results to remove noise and reduce computational effort. Successive layers of convolution and pooling help the network capture features at different depths and scales, highlighting differences in the material, shape, or physical properties of construction waste.

[0065] 3) Full connection and activation: The cleverly designed fully connected layer can map the output of the convolutional layer to a more compact feature space, and further amplify key information or suppress redundant information through activation functions (such as ReLU, LeakyReLU, or activation functions embedded with attention weights in the improved solution of the present invention).

[0066] 4) Classification output and priority assessment: The final layer of the network can use activation functions such as Softmax or Sigmoid to generate probability distributions for different categories (e.g., concrete blocks, metals, plastics, wood, glass, etc.), with the category with the highest probability being used as the classification result for construction waste. Furthermore, a parallel neural network channel can be designed in the network's output layer to comprehensively assess sorting priorities based on indicators such as waste resource value, processing difficulty, or riskiness. Larger values ​​indicate higher recycling or priority treatment value, while smaller values ​​indicate lower priority.

[0067] 5) Confidence Calculation: Using Softmax or other probabilistic output mechanisms, the network can assign corresponding probabilities to each category. The highest probability value can be considered the confidence level of the classification result. If the confidence value falls below a pre-set threshold, it is considered "uncertain" or "abnormal," requiring subsequent manual intervention or secondary testing. Training and Application Scenarios: Offline Training: In this embodiment, the classification CNN can first be trained offline using large amounts of multi-source construction waste data to ensure that the network fully learns the characteristic differences between different types of waste. Online Inference: After the system obtains real-time fused feature vectors, it inputs them into the trained CNN for inference. The network outputs the waste type, sorting priority, and corresponding confidence level in real time, which is then used by the sorting mechanism to perform sorting operations. Compared to traditional CNNs with single-modal input, the classification CNN described in this embodiment can process comprehensive features derived from multimodal fusion, making it more discriminative and adaptable, reducing the probability of missed detections and false positives. By setting a sorting priority output channel, high-value or hazardous waste can be quickly identified and prioritized for sorting or isolation, significantly improving resource utilization efficiency and safety. In summary, this embodiment describes in detail the application process and working principle of the fusion feature vector in the classification convolutional neural network. It can not only complete the waste classification efficiently and accurately based on multi-source information, but also flexibly decide the processing order in conjunction with the priority judgment mechanism, thereby having higher application value and actual efficiency in the automated production line of construction waste sorting.

[0068] S5: Based on the type of construction waste and the confidence level of the classification, the corresponding sorting actuators are controlled to deliver different types of waste to the corresponding recycling and processing areas. The sorting actuators include a robotic arm, a pneumatic sorting valve, and a mobile trash can. In some embodiments, after the system completes the classification and confidence assessment of the construction waste, it will control the sorting actuators to deliver different types of waste to the corresponding recycling and processing areas based on the identification output and the set priority. The following is a detailed description of the application of the three actuators: the robotic arm, the pneumatic sorting valve, and the mobile trash can: When the sorting execution instruction is generated, when the classification convolutional neural network outputs the type of target waste (such as metal, concrete, wood, etc.) and gives a confidence value, the system will first compare it with the preset confidence threshold: 1) If the confidence level is higher than (or equal to) the set threshold, the classification result is adopted and the automatic sorting process begins; 2) If the confidence level is lower than the threshold, the batch of materials is marked as "abnormal" or "pending manual processing" and will be inspected by humans or a backup system. Subsequently, the system queries the appropriate sorting actuator type and location from the database or configuration table according to the "recycling and processing area" code corresponding to the waste. For the robotic arm sorting solution, if the system determines that the waste to be sorted is large in volume and requires a more precise grasping position (such as scattered steel bars or large concrete blocks), the robotic arm will be selected for sorting. After obtaining the execution instruction, the robotic arm will perform route planning and posture control of the gripper or adsorption end based on the real-time position detection on the conveyor belt. The robotic arm accurately grasps the target waste and moves it to the designated recycling bin or hopper. Compared with other methods, the robotic arm has higher flexibility and can handle waste of various specifications and shapes.

[0069] The pneumatic sorting valve solution can be used to quickly divert waste when the system identifies certain types of waste (such as lightweight plastics or small debris) at high conveying speeds. The moment the system detects that the material is about to reach the sorting area, it triggers the pneumatic valve to open or close the corresponding channel; using strong airflow or mechanical baffles, the material is quickly blown (or moved) to the corresponding sorting channel. Pneumatic sorting valves are suitable for high-speed processing of waste with large flow rates, relatively small size, and easy to be pushed by airflow. They are highly efficient, but their sorting effect on large or hard materials is limited.

[0070] Mobile waste bins are used at certain disposal sites where installing fixed mechanical equipment at the sorting location is difficult, or when multiple waste types need to be collected and moved to a processing point away from the main conveyor belt. The system automatically controls a trolley-type waste bin to the conveyor belt's receiving port based on waste type and priority. Once the waste arrives at the designated location, the trolley receives the material through a flap or telescopic mechanism and returns it to the designated recycling area along a pre-planned route. This method is suitable for operations requiring high flexibility or dispersed sites, but its frequency and processing speed are generally lower than those of robotic arms and pneumatic sorting valves. Monitoring and exception handling: During the sorting process, the system monitors the execution of waste grabbing or pushing actions, as well as the remaining material on the conveyor belt, in real time. If errors or debris remain in the sorting port after sorting, the system records the anomaly and triggers an alarm or error correction mechanism. If the sorting target repeatedly deviates from the expected classification results, the system marks and archives the corresponding data for subsequent classification model correction or self-learning training. Through the sorting execution scheme of this embodiment, various types of construction waste can be placed in appropriate recycling and processing areas according to the categories and confidence levels output by the classification model. This can not only improve the sorting accuracy, but also flexibly adapt to various scenarios and different waste characteristics, greatly improving the overall operational efficiency.

[0071] Specifically, the multi-source data includes: RGB camera to obtain visible light images, depth camera to obtain three-dimensional information, spectral sensor to obtain material spectral data, weight sensor to obtain material weight information, and vibration sensor to obtain vibration signals; RGB camera, depth camera, spectral sensor, and weight sensor are respectively installed at different positions and / or different angles of the conveyor belt to achieve multi-angle collection of construction waste; among them, RGB camera is used to obtain visible light images and perform surface texture recognition, depth camera is used to extract the three-dimensional geometric structure of materials, spectral sensor is used to obtain spectral reflection characteristics of materials contained in the waste and distinguish different materials, weight sensor obtains material weight information by detecting local pressure of the conveyor belt, and vibration sensor monitors vibration signals during transportation to detect internal defects or loose structures.

[0072] Specifically, the preprocessing includes image denoising on the RGB image and three-dimensional information, and filtering the spectral data, weight data and vibration data using a Kalman filter algorithm.

[0073] Specifically, step S3 includes:

[0074] S3-1: performing feature extraction on the multi-source data respectively to obtain multimodal features including visible light image features, three-dimensional geometric structure features, spectral features, weight features, and vibration features;

[0075] S3-2: The extracted features are input into the attention mechanism module, which calculates the attention weight of each modality feature in the fusion process based on the correlation of each modality feature;

[0076] S3-3: Dynamically adjust the attention weights and perform weighted operations on different modal features to form a fused feature vector.

[0077] Specifically, the multimodal feature acquisition in step S3-1 includes: inputting the visible light image into a feature extraction convolutional neural network, extracting surface texture features as visible light image features through multi-layer convolution, pooling and activation operations; converting the three-dimensional data acquired by the depth camera into a point cloud or depth map, and inputting the three-dimensional convolutional neural network or point cloud network to extract spatial structure features as three-dimensional geometric structure features; performing a wavelet frequency domain transform on the spectral data to obtain representative spectral components as spectral features; performing feature extraction on the weight data and vibration data respectively to obtain mean, variance, kurtosis and bandpass energy statistical features, and obtaining weight features and vibration features;

[0078] Step S3-2 includes: constructing and inputting an attention mechanism module, using the extracted multimodal features to construct a key, a query, and a value, and inputting them into an attention mechanism module including a multi-head attention unit. In each attention head, the similarity between the key and the query is calculated based on cosine similarity, and the results are normalized to obtain the attention weight of each modality;

[0079] The step S3-3 includes: weighting each modal feature according to the attention weight and then splicing them to obtain a fused feature vector.

[0080] Specifically, S5-1: comparing the classification confidence with a preset confidence threshold; if the classification confidence is higher than or equal to the confidence threshold, adopting the classification result; if the classification confidence is lower than the confidence threshold, marking the construction waste as "abnormal" and sending it to the corresponding manual processing area;

[0081] S5-2: Matching the adopted classification results with sorting equipment, and selecting at least one sorting actuator from among a robotic arm, a pneumatic sorting valve, or a mobile trash can according to the type of construction waste and its volume, shape, or density characteristics;

[0082] S5-3: Use the selected sorting actuator to grab, push or suction the construction waste to be sorted and guide it to the corresponding recycling and processing area.

[0083] Specifically, the activation function f1(x) used in the classification convolutional neural network model is expressed as follows:

[0084]

[0085] Wherein, x represents the input value of the current neuron, e is the base of the natural logarithm; a1 represents the attention weight corresponding to the visible light image feature calculated by step S3-2, and a2 represents the attention weight corresponding to the three-dimensional geometric structure feature calculated by step S3-2. In some embodiments, the main workflow of the classification convolutional neural network includes convolution operation, pooling dimensionality reduction, feature fusion and final classification output. In order to better highlight the key modal information and suppress interference, this embodiment has made targeted improvements and optimizations in the activation function link of the network, as follows: The application of the activation function in the hidden layer, and the nonlinear mapping of features between each convolution layer and the fully connected layer in the network. Traditional activation functions such as ReLU or LeakyReLU usually use a unified threshold and mapping method to treat all input features equally.

[0086] In this embodiment, when certain modalities (such as visible light image features) are judged to be more critical, the system will amplify the contribution of the modal features in the forward propagation by adjusting the threshold or output range of the activation function. This can maintain continuous attention to key features in the deep network and avoid the attenuation of key features due to the superposition operation of subsequent layers. In combination with the attention weight for adaptive amplification, during the feature mapping process of the network, the attention weight of each modality will be obtained through the previous stage or parallel attention mechanism module. If the attention weight of a modality is higher, it means that the modality has greater value in distinguishing the types of waste; if the weight is low, the role of the modality in the current judgment scenario is relatively weak.

[0087] To fully utilize this information, this embodiment adaptively scales the input or intermediate output during the activation phase. If the attention weight is high, the intermediate output is significantly retained or even amplified; if the attention weight is low, the intermediate output signal is moderately reduced. This approach can be understood as "dynamically assigning a gain coefficient to the activation function based on the modality's importance," making the network more sensitive to key modalities. Implementation and training process: After the convolutional or fully connected layer, the attention weight corresponding to the current modality is first read, then the eigenvalues ​​are scaled, and the scaled result is passed to the activation function (such as ReLU or Sigmoid). This allows the amplification mechanism for important modalities to be retained in addition to the traditional activation function calculation process, without requiring significant changes to the overall network structure. Training process: During offline training, the network automatically learns the distribution of attention weights that best improves the accuracy of construction waste recognition and adjusts the corresponding input distribution of the activation function. Because the gradient can be traced back to the attention weights during backpropagation, the system can continuously adjust which modalities should be emphasized and which should be suppressed. The final classification output and confidence score are generated at the end of the network using common normalization methods (such as Softmax) to generate a probability distribution for each construction waste category, with the maximum probability value used as the confidence score for the classification result. The difference is that the upper-level features on which each category output relies now take into account the amplifying effect of attention weights on the activation function, thereby strengthening the influence of key modalities on the classification process. If this modality truly provides more effective discrimination, the model's final output will demonstrate higher accuracy and stability. Highlighting Key Modalities: Incorporating attention weights during the activation phase allows the deep network to more sensitively capture important modal features such as visible light images. Enhanced Adaptability: When the contribution of a modality decreases in a specific scenario, the network automatically reduces its activation weight, thereby improving overall noise immunity and robustness. Efficient Training: This activation function strategy does not require significant changes to the network structure. By simply inserting the reading, writing, and scaling of attention weights into the existing activation function calculation process, it significantly improves feature representation and is compatible with mainstream deep learning frameworks. In summary, this embodiment incorporates the idea of ​​multimodal attention weights into the activation function of the classification convolutional neural network, taking into account the high efficiency of the traditional activation function and the differentiated attention advantages of the attention mechanism, which can significantly improve the network's recognition accuracy and stability for various types of construction waste, and demonstrates good practical value in sorting operations under high noise and complex working conditions.

[0088] This application also provides a multi-source information fusion intelligent sorting system for construction waste, such as Figure 2The system primarily comprises an acquisition module, a preprocessing module, a fused feature vector acquisition module, a calculation module, and an actuator operation module. This embodiment provides a multi-source information fusion intelligent sorting system applicable to centralized construction waste treatment facilities or demolition sites. The modules are interconnected via a high-speed data bus or industrial Ethernet to enable real-time data acquisition, transmission, processing, and execution. A specific example is as follows: The acquisition module hardware consists of: 1) an RGB camera for capturing visible light images; 2) a depth camera for capturing the three-dimensional geometric structure of construction waste; 3) a spectral sensor for detecting the spectral reflectance characteristics of the material in different wavelengths; 4) a weight sensor, installed under the conveyor belt or on the support frame, to monitor local pressure and obtain material weight information; and 5) a vibration sensor for capturing vibration signals generated during conveying or by loose material. Connection method: Each sensor is connected to the system's main control server or embedded acquisition terminal via industrial Ethernet (such as EtherCAT) or serial communication (such as RS485). Acquisition modules are typically located at different locations on or around the conveyor belt to enable multi-angle and multi-modal acquisition of waste. Through the combined collection of the sensors, more comprehensive data support is provided for subsequent identification and sorting, thereby improving identification accuracy.

[0089] The preprocessing module hardware consists of: 1) an industrial personal computer (IPC) or high-performance embedded processor for running preprocessing algorithms; 2) a high-speed storage device for caching sensor data and processing results. Connection method: The IPC receives multi-source data from the acquisition module and writes data at high bandwidth through local caching or shared memory mechanisms. Function: De-noises and corrects distortion of images captured by the RGB and depth cameras; performs filtering on spectral, weight, and vibration signals, such as Kalman filtering or other noise reduction algorithms; and transmits the denoised data in a unified format or with a timestamp to the subsequent fusion feature vector acquisition module.

[0090] The hardware components of the fusion feature vector acquisition module are as follows: 1) GPU server or embedded GPU card, used to accelerate multimodal feature extraction in convolutional neural networks; 2) high-bandwidth system bus, used to quickly transmit feature data between multi-core CPUs and GPUs. Connection method: This module is usually on the same industrial computer or server platform as the pre-processing module, interconnected through the motherboard bus (such as PCIe), and share memory or video memory for data exchange. De-noised RGB images, depth information, spectral data, weight data and vibration data are obtained from the pre-processing module; convolutional neural networks or 3D-CNN / point cloud networks are used to extract features from visible light images and three-dimensional data, and feature dimensionality reduction processing such as wavelet transform is performed on spectral data, and the statistical characteristics of weight and vibration signals are recorded; each modal feature is input into the attention mechanism module, and the attention weight is calculated based on the correlation between different modal features, and weighted fusion is performed to finally generate a fusion feature vector.

[0091] The computing module hardware consists of the following components: 1) a GPU or dedicated AI accelerator chip (e.g., NPU, TPU, etc.) for deep learning inference; and 2) a multi-threaded CPU environment for comprehensive management of data pipelines and model calls. The connection method and the fusion feature vector acquisition module are deployed together on the same workstation or server, obtaining the fusion feature vector through internal high-speed communication or shared memory. The fusion feature vector is input into a previously trained offline classification convolutional neural network, which outputs the corresponding construction waste category (e.g., concrete blocks, metal, wood, etc.) and sorting priority in real time. Softmax or other probabilistic mechanisms are used to determine the confidence level of the classification results. The classification results and confidence information are stored or directly transmitted to the actuator operation module to ensure that subsequent sorting instructions are based on the latest recognition results.

[0092] The hardware components of the actuator operation module include: 1) a programmable logic controller (PLC) or motion controller for real-time control of the sorting actuators; 2) a drive system for the robotic arm, pneumatic sorting valve, and mobile trash can; and 3) additional sensor modules (such as collision sensors and position encoders) to monitor actuator status. Connection method: The PLC or motion controller is typically connected to each actuator via a fieldbus (such as PROFINET, EtherCAT, or CANopen), communicating with the upper-level computing module to obtain sorting instructions and output action feedback.

[0093] Receive and analyze construction waste categories, confidence levels, and sorting priority information; when the confidence level is higher than a preset threshold, automatically control the robotic arm or pneumatic valve to perform grabbing, pushing, or suction diversion operations; waste with low confidence or unknown categories is directed to manual processing or abnormal processing areas; if the waste is determined to require independent recycling or bulk collection, the mobile trash can is moved to the conveyor belt under PLC scheduling to receive such waste and transport it to the corresponding recycling and processing area; monitor the completion and accuracy of the execution action in real time, and if sorting deviations or interference occur, the information is fed back to the upper system through an alarm or fallback mechanism for correction.

[0094] The overall connection method and data flow are as follows:

[0095] 1) Construction waste is transported to the multi-sensor coverage area via a conveyor belt;

[0096] 2) The acquisition module sends raw data to the pre-processing module via industrial Ethernet during real-time monitoring;

[0097] 3) After the pre-processing module performs denoising and filtering, the data is transmitted to the fusion feature vector acquisition module;

[0098] 4) The fusion feature vector acquisition module generates a fusion feature vector in a GPU-accelerated environment through feature extraction and attention weighted fusion;

[0099] 5) The computing module uses a deep learning classification algorithm to perform inference and output categories, sorting priorities, and confidence levels;

[0100] 6) The actuator operation module controls the robotic arm, pneumatic sorting valve or mobile trash can based on the recognition results and threshold judgment to place the waste into the corresponding recycling area;

[0101] 7) If the sorting completion does not meet expectations or the confidence level is too low, manual review or exception handling will be triggered.

[0102] Through the above-mentioned hardware components and connection relationships, this system can effectively achieve accurate classification and high-speed sorting of construction waste under the condition of high integration of multi-source information, significantly improving resource recycling efficiency and automation level.

[0103] The acquisition module uses multiple different types of sensors to monitor construction waste on the conveyor belt in real time and obtain multi-source data;

[0104] Preprocessing module, preprocesses multi-source data;

[0105] The fusion feature vector acquisition module extracts features from multi-source data and fuses the multi-source data using a multimodal fusion algorithm. The fusion algorithm introduces an attention mechanism to weight the contributions of multi-source data of different modalities. Based on the correlation of the data of each modality, the weights are adaptively assigned to obtain a fused feature vector.

[0106] The calculation module inputs the fused feature vector into the trained classification convolutional neural network model, outputs different types of construction waste and their corresponding sorting priorities, and gives the confidence level of the classification results;

[0107] The actuator operation module controls the corresponding sorting actuator according to the type of construction waste and the classification confidence, and sends different types of waste to the corresponding recycling and processing areas. The sorting actuator includes a robotic arm, a pneumatic sorting valve, and a mobile trash can.

[0108] Specifically, the multi-source data includes: RGB camera to obtain visible light images, depth camera to obtain three-dimensional information, spectral sensor to obtain material spectral data, weight sensor to obtain material weight information, and vibration sensor to obtain vibration signals; RGB camera, depth camera, spectral sensor, and weight sensor are respectively installed at different positions and / or different angles of the conveyor belt to achieve multi-angle collection of construction waste; among them, RGB camera is used to obtain visible light images and perform surface texture recognition, depth camera is used to extract the three-dimensional geometric structure of materials, spectral sensor is used to obtain spectral reflection characteristics of materials contained in the waste and distinguish different materials, weight sensor obtains material weight information by detecting local pressure of the conveyor belt, and vibration sensor monitors vibration signals during transportation to detect internal defects or loose structures.

[0109] Specifically, the preprocessing includes image denoising on the RGB image and three-dimensional information, and filtering the spectral data, weight data and vibration data using a Kalman filter algorithm.

[0110] The present invention provides a method and system for intelligent sorting of construction waste by integrating multi-source information, which can achieve the following beneficial technical effects:

[0111] 1. This invention simultaneously collects multimodal data, including visible light images, three-dimensional structures, spectral data, weight information, and vibration signals. During the fusion process, an attention mechanism is introduced to weight key modalities. This allows accurate identification of the material and appearance characteristics of construction waste of various types and shapes, reducing blind spots caused by reliance on a single sensor and significantly improving sorting accuracy. By assessing the confidence of the classification results and combining them with sorting priorities, the invention can control robotic arms, pneumatic sorting valves, or mobile trash cans to rapidly sort and release various types of waste. This enables intelligent sorting of various types of construction waste, including concrete blocks, metals, plastics, and wood, shortening processing cycles and reducing labor costs and operational risks.

[0112] 2. This invention utilizes Kalman filtering and multimodal feature extraction algorithms in the preprocessing stage, effectively suppressing noise and filtering out interference. It also uses an attention mechanism to adaptively adjust the weights of each modality during the fusion process, enabling stable operation under complex conditions such as variable ambient lighting, diverse waste sizes, and high noise levels. This invention can be easily integrated with existing conveyor systems, robotic grippers, and pneumatic sorting devices, and deployed in diverse building demolition or centralized waste processing scenarios, encompassing a wider range of applications and helping to achieve efficient recycling and resource processing of construction waste.

[0113] 3. By incorporating attention weights corresponding to visible light image modalities into the activation function calculation of a convolutional neural network, this invention further amplifies the influence of visible light image features during the critical forward propagation phase of the model. This enables the network to more sensitively capture detailed textures and lighting differences when distinguishing waste with high appearance similarity, significantly improving classification accuracy. In traditional convolutional network structures, the activation function is insensitive to the scale or modal information of the input signal. However, with the introduction of attention weights, the differences between the visible light modality and other modalities are more directly quantified in the activation output. This not only enables the network to dynamically "amplify" important modal features, but also provides greater interpretability for subsequent visual analysis and diagnosis. Because attention weights are dynamically calculated based on the correlation and data quality of each modal data, the network's activation output can be adjusted in real time based on factors such as external ambient lighting, material type, and waste shape. This adaptive mechanism enables the sorting system to maintain high robustness and accuracy even in complex and changing scenarios. By incorporating attention weights into the activation function, the neural network can converge more quickly on the key range of visible light image features during training, thereby reducing over-learning of redundant or noisy features and helping the model achieve good results even with limited training samples. Furthermore, the additional "amplification" effect of visible light features also provides a more stable gradient update signal for backpropagation.

[0114] The above is a detailed introduction to a method and system for intelligent sorting of multi-source information fusion of construction waste. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the core idea of ​​the present invention. At the same time, for general technical personnel in this field, according to the ideas and methods of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A method for intelligent sorting of construction waste by multi-source information fusion, characterized in that: Including steps: S1: Use multiple different types of sensors to monitor construction waste on the conveyor belt in real time to obtain multi-source data; S2: Preprocess multi-source data; S3: Extract features from multi-source data and fuse them using a multimodal fusion algorithm. The fusion algorithm introduces an attention mechanism to weight the contributions of multi-source data of different modalities. Based on the relevance of the data of each modality, the weights are adaptively assigned to obtain a fused feature vector. S4: Input the fused feature vector into the trained classification convolutional neural network model to output different types of construction waste and their corresponding sorting priorities, and provide the confidence level of the classification results; S5: According to the type of construction waste and the confidence level of classification, the corresponding sorting actuators are controlled to send different types of waste to the corresponding recycling and processing areas. The sorting actuators include robotic arms, pneumatic sorting valves, and mobile trash cans. The step S3 comprises: S3-1: performing feature extraction on the multi-source data respectively to obtain multimodal features including visible light image features, three-dimensional geometric structure features, spectral features, weight features, and vibration features; S3-2: The extracted features are input into the attention mechanism module, which calculates the attention weight of each modality feature in the fusion process based on the correlation of each modality feature; S3-3: Dynamically adjust the attention weights and perform weighted operations on features of different modalities to form a fused feature vector. Step S3-2 includes: constructing and inputting an attention mechanism module, using the extracted multimodal features to construct a key, a query, and a value, which are then input into an attention mechanism module comprising a multi-head attention unit. In each attention head, the similarity between the key and the query is calculated based on cosine similarity, and the result is normalized to obtain the attention weight of each modality; The step S3-3 includes: weighting each modal feature according to the attention weight and then splicing them to obtain a fusion feature vector; The activation function used in the classification convolutional neural network model is It is expressed as follows: ; in, Represents the input value of the current neuron, and e is the base of the natural logarithm; represents the attention weight corresponding to the visible light image feature calculated in step S3-2, Represents the attention weight corresponding to the three-dimensional geometric structure feature calculated by step S3-2.

2. The method for intelligent sorting of construction waste by multi-source information fusion according to claim 1, characterized in that: The multi-source data includes: an RGB camera to obtain visible light images, a depth camera to obtain three-dimensional information, a spectral sensor to obtain material spectral data, a weight sensor to obtain material weight information, and a vibration sensor to obtain vibration signals; the RGB camera, the depth camera, the spectral sensor, and the weight sensor are respectively installed at different positions and / or different angles of the conveyor belt to achieve multi-angle collection of construction waste; among them, the RGB camera is used to obtain visible light images and perform surface texture recognition, the depth camera is used to extract the three-dimensional geometric structure of the material, the spectral sensor is used to obtain the spectral reflection characteristics of the material contained in the waste and distinguish different materials, the weight sensor obtains material weight information by detecting the local pressure of the conveyor belt, and the vibration sensor monitors the vibration signal during the conveying process to detect internal defects or loose structures.

3. The method for intelligent sorting of construction waste by multi-source information fusion according to claim 1, characterized in that: The preprocessing includes image denoising on the RGB image and three-dimensional information, and filtering the spectrum data, weight data and vibration data using a Kalman filter algorithm.

4. The method for intelligent sorting of construction waste by multi-source information fusion according to claim 1, characterized in that: The multimodal feature acquisition step S3-1 includes inputting the visible light image into a feature extraction convolutional neural network, extracting surface texture features as visible light image features through multi-layer convolution, pooling, and activation operations; converting the three-dimensional data acquired by the depth camera into a point cloud or depth map, and then inputting the data into a three-dimensional convolutional neural network or point cloud network to extract spatial structure features as three-dimensional geometric structure features; Perform wavelet frequency domain transformation on the spectral data to obtain representative spectral components as spectral features; Feature extraction is performed on the weight data and vibration data respectively to obtain statistical features of mean, variance, kurtosis and bandpass energy, and weight features and vibration features are obtained.

5. The method for intelligent sorting of construction waste by multi-source information fusion according to claim 1, characterized in that: S5-1: Compare the classification confidence with a preset confidence threshold. If the classification confidence is higher than or equal to the confidence threshold, adopt the classification result. If the classification confidence is lower than the confidence threshold, mark the construction waste as "abnormal" and send it to the corresponding manual processing area. S5-2: Matching the adopted classification results with sorting equipment, and selecting at least one sorting actuator from among a robotic arm, a pneumatic sorting valve, or a mobile trash can according to the type of construction waste and its volume, shape, or density characteristics; S5-3: Use the selected sorting actuator to grab, push or suction the construction waste to be sorted and guide it to the corresponding recycling and processing area.

6. A multi-source information fusion intelligent sorting system for construction waste, characterized by: include: The acquisition module uses multiple different types of sensors to monitor construction waste on the conveyor belt in real time and obtain multi-source data; Preprocessing module, preprocesses multi-source data; The fusion feature vector acquisition module extracts features from multi-source data and fuses the multi-source data using a multimodal fusion algorithm. The fusion algorithm introduces an attention mechanism to weight the contributions of multi-source data of different modalities. Based on the correlation of the data of each modality, the weights are adaptively assigned to obtain a fused feature vector. The calculation module inputs the fused feature vector into the trained classification convolutional neural network model, outputs different types of construction waste and their corresponding sorting priorities, and gives the confidence level of the classification results; The actuator operation module controls the corresponding sorting actuators according to the type of construction waste and the classification confidence level, and sends different types of waste to the corresponding recycling and processing areas. The sorting actuators include robotic arms, pneumatic sorting valves, and mobile trash cans. The multi-source information fusion intelligent sorting method for construction waste as described in any one of claims 1 to 5 is used to implement the multi-source information fusion intelligent sorting system for construction waste.

7. The multi-source information fusion intelligent sorting system for construction waste according to claim 6, characterized in that: The multi-source data includes: an RGB camera to obtain visible light images, a depth camera to obtain three-dimensional information, a spectral sensor to obtain material spectral data, a weight sensor to obtain material weight information, and a vibration sensor to obtain vibration signals; the RGB camera, the depth camera, the spectral sensor, and the weight sensor are respectively installed at different positions and / or different angles of the conveyor belt to achieve multi-angle collection of construction waste; among them, the RGB camera is used to obtain visible light images and perform surface texture recognition, the depth camera is used to extract the three-dimensional geometric structure of the material, the spectral sensor is used to obtain the spectral reflection characteristics of the material contained in the waste and distinguish different materials, the weight sensor obtains material weight information by detecting the local pressure of the conveyor belt, and the vibration sensor monitors the vibration signal during the conveying process to detect internal defects or loose structures.

8. The multi-source information fusion intelligent sorting system for construction waste according to claim 6, characterized in that: The preprocessing includes image denoising on the RGB image and three-dimensional information, and filtering the spectrum data, weight data and vibration data using a Kalman filter algorithm.

Citation Information

Patent Citations

  • Action recognition method and system based on combination of electrostatic induction and image detection

    CN118747304A

  • On-site intelligent sorting method for construction waste

    WO2025000272A1