Composite material garbage sorting method based on hyper-spectrum technology and AI classification model

By deploying hyperspectral and RGB cameras on the conveyor belt, combining multimodal neural networks and SegFormer deep segmentation networks, the problem of difficult to sort composite waste is solved, and efficient and accurate garbage sorting and resource recycling are achieved.

CN120279340AActive Publication Date: 2025-07-08FUJIAN ZENGZHI ENVIRONMENTAL PROTECTION TECH CO LTD +1

Patent Information

Application Number
CN202510759162.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

Traditional manual sorting or RGB vision-based methods are difficult to accurately distinguish composite waste, resulting in low sorting efficiency, waste of resources and environmental pollution.

Method used

Using a method based on hyperspectral technology and AI classification model, the ultraspectral and RGB cameras are deployed along the conveyor belt to perform multimodal data acquisition, preprocessing, and fusion, and classification is performed using multimodal neural networks. Combined with SegFormer deep segmentation network and target tracking algorithm, precise identification and sorting of garbage materials are achieved.

Benefits of technology

It realizes efficient and precise sorting of composite waste, improves sorting efficiency and resource recycling rate, supports fine-grained identification of multi-material mixed objects, and ensures the accuracy and timeliness of sorting actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279340A_ABST
    Figure CN120279340A_ABST
Patent Text Reader

Abstract

The invention relates to a composite material garbage sorting method based on a hyper-spectrum technology and an AI classification model. The method comprises the following steps that S1, multi-modal data collection is conducted; s2, the collected multi-modal data are preprocessed and fused, and combined features of the multi-modal data are obtained; s3, constructing a classification model based on the multi-modal neural network model, and performing classification according to the joint features; s4, based on a classification result, using a SegFormer deep segmentation network to identify junk components at a pixel level, and realizing accurate segmentation of a plurality of material composition objects; s5, target tracking of the garbage trajectory is kept through a target tracking algorithm; and S6, on the basis of the target tracking result in the S5, the position of the garbage on the conveying belt is predicted in real time, and the sorting action is precisely executed according to the classification and segmentation result. According to the invention, the complex composite material garbage can be efficiently and accurately treated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent sorting, and particularly to a method for sorting composite material waste based on hyperspectral technology and an AI classification model. Background Art

[0002] Composite material waste is composed of multiple materials, usually including metals, plastics, fibers, etc. Traditional manual sorting or RGB vision-based sorting methods are difficult to accurately distinguish these materials, easily resulting in low sorting efficiency, resource waste, and environmental pollution. Traditional sorting technologies (based on manual, mechanical, or simple image recognition) have technical problems of low resolution accuracy and insufficient efficiency when facing complex composite material waste. Summary of the Invention

[0003] In order to solve the above problems, the purpose of the present invention is to provide a method for sorting composite material waste based on hyperspectral technology and an AI classification model, which can efficiently and accurately process complex composite material waste.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions: A method for sorting composite material waste based on hyperspectral technology and an AI classification model, comprising the following steps: S1: Deploy a number of hyperspectral and RGB cameras at different nodes along the conveyor belt, and each node focuses on collecting multi-modal data of the waste in a section; S2: Preprocess the collected multi-modal data, and fuse the RGB and hyperspectral data according to the timestamp and spatial position to obtain the joint features of the multi-modal data; S3: Build a classification model based on the multi-modal neural network model, and classify according to the joint features; S4: Based on the classification result, use the SegFormer deep segmentation network to identify the composition of the waste at the pixel level and realize the segmentation of objects composed of multiple materials; S5: Keep target tracking of the waste trajectory through a target tracking algorithm; S6: Based on the result of the target tracking in S5, predict the position of the waste on the conveyor belt in real time, and perform sorting actions according to the classification and segmentation results.

[0005] Further, S1 is specifically: Pre-measure the section length L and width W at different positions along the conveyor belt; According to the waste flow rate v, combined with the density R of the waste density , determine the sampling frequency f of each node sampling ;

[0006] where d min is the minimum distance between adjacent pieces of garbage; According to the sampling field of view F view =(θ H , θ V ), where θ H is the field of view angle in the horizontal direction and θ V is the field of view angle in the vertical direction, determine the overlapping area to avoid blind spots in garbage detection, and the deployment distance between adjacent nodes S nodes is: ; Each node includes an RGB camera, a hyperspectral imager, a triggering device, and a controller; Assume that the garbage on the conveyor belt moves at a speed v, and the sensor dynamically captures multi-modal information of the object: Calculate the length L capture and width W capture covered by each sampling based on the resolution and FOV of the sensor:

[0007] where h is the vertical height from the sensor to the conveyor belt, where: L capture ≥ L, W capture ≥ W; The displacement Δ capture of the garbage within two sampling time intervals T x is calculated according to the following formula: ; Ensure that Δx ≤ d min to ensure continuous sampling and no missed detection.

[0008] Furthermore, S2 is specifically: Assume that the collected RGB data is R t (x, y, c), corresponding to the timestamp t R ; the hyperspectral data is H t (x, y, λ), corresponding to the timestamp t H ; where, x, y represent the spatial dimensions, λ represents the light wave wavelength in the hyperspectral image; c is the number of channels; For each frame of RGB data R t , find the hyperspectral data H t with the closest timestamp to R t :

[0009] where, is the time difference; It means finding the t that minimizes the time difference in the set of candidate time points value; H value; After time alignment, synchronized multimodal data pairs (R t , H t ) are obtained; Spatial alignment is completed through the transformation matrix T obtained by calibration align , and the formula is:

[0010] where [u, v] Hyper are the pixel coordinates of the hyperspectral data; [u, v] RGB are the pixel coordinates of the RGB data; Align the pixel coordinates of the hyperspectral data with the pixel coordinates of the RGB data to obtain spatially consistent multimodal data; After time and spatial alignment, for each pixel point, directly splice the RGB data and the hyperspectral data to form a joint feature F t (x, y).

[0011] Furthermore, the multimodal neural network model includes a hyperspectral branch, an RGB branch, a multi-layer fusion strategy layer, and a classification head, which are specifically as follows: The hyperspectral branch inputs the preprocessed hyperspectral data , extracts spectral features, and captures material information at different wavelengths: Extract local spectral-spatial features through a 3D convolutional layer : ; where kernel is the convolution kernel size; Channels is the number of convolution output channels; Con3D represents the 3D convolution operation; The spectral Transformer module captures global spectral dependencies : ; where Q, K, and V are the query vector, key vector, and value vector respectively; W Q , W K , W V are the projection parameters of the query vector, key vector, and value vector respectively; d k is the dimension of the attention space; T represents transpose; Attention represents the attention mechanism; LayerNorm is the normalization operation; Finally, through the global average pooling operation, a spectral feature vector f H : ; Among them, GlobalAvgPool represents the global average pooling operation; The RGB branch inputs the preprocessed RGB data R′ and extracts shape and texture features: Through the ResNet-50 backbone network, multi-level spatial features are extracted: ; Among them, ResNet gradually extracts multi-level features through 4 residual blocks ResNet-Block1, ResNet-Block2, ResNet-Block3, and ResNet-Block4, including , , , ; The spatial attention module enhances the features of key regions: ; Among them, Conv2D represents the convolution operation; ⊙ is the element-wise weighted operation; σ is the activation function; w R (x,y) is the attention weight; is the output of the spatial attention module; Global average pooling generates a spatial feature vector: ; The multi-layer fusion strategy layer fuses hyperspectral and RGB features at multiple levels to enhance modal complementarity: At the intermediate layer, the feature maps of the two modalities are directly concatenated to obtain the concatenated feature : ; After concatenation, it is fused through a convolutional layer to obtain the fused feature : ; The classification head outputs the classification result by passing the fused feature through a fully connected layer: ; Among them, is the class probability; Softmax is the activation function; W c and b c are the weights and biases of the fully connected layer.

[0012] Furthermore, for the few-shot composite materials, the classification model introduces meta-learning to optimize the classification model, specifically as follows: Regarding different material classification tasks as the task set τ of meta-learningi ; Each task contains a support set S=(X,Y) and a query set Q’=(X′,Y′); Model training process: Update the model parameters on the support set of task set τ i : ; Among them, is the model parameter, is the learning rate; is the loss function corresponding to the meta-task τ i ; is the updated model parameter on the support set of task set τ i ; represents the gradient of the loss function with respect to the parameter θ; Calculate the loss of the query set, and find the total gradient on all tasks to optimize the global parameters: ; Among them, is the meta-learning rate; When new category of garbage data is input, quickly update through the support set: .

[0013] Furthermore, based on the classification results, use the SegFormer deep segmentation network to identify the composition of garbage at the pixel level and achieve the segmentation of objects composed of multiple materials, as follows: Use the class probability output by the classification model as the prior condition of the segmentation network, and input the joint features and class probabilities of multi-modal data ; Through conditional concatenation, broadcast the classification label to each pixel point to form conditional features : ; Among them, is the extended conditional information; The SegFormer deep segmentation network includes an encoder and a decoder, The encoder uses a pre-trained SegFormer-B5 model, with the input being Fcond(x,y) and the output being multi-scale features : ; Among them, SegFormerEncoder is the encoder of the SegFormer-B5 model; The decoder fuses multi-scale features and generates a segmentation mask M seg : ; Among them, MLP represents the MLP decoder; The cross-entropy loss is adopted to process the pixel-level output: ; Among them, is the One-Hot representation of the pixel true material label; is the predicted segmentation mask; k is the material category index.

[0014] Furthermore, S5 is specifically: Obtain the position, category, and segmentation mask information of the target from the classification and segmentation modules: Generate the minimum bounding rectangle b from the pixel-level mask of the a-th item output by the segmentation module boxa : b boxa =[x min ,y min ,x max ,y max ; Among them, [x min ,y min are the coordinates of the upper left corner of the rectangle frame [x max ,y max are the coordinates of the lower right corner of the rectangle frame; The classification module outputs the category probability vector of the i-th item ; The pixel-level material segmentation result M i (x,y); Use the Kalman filter to predict the position of the target in the next frame; and associate the detection box with the tracking target through the Hungarian algorithm; Process the appearance of new targets, disappearance of old targets, or occlusion according to the following rules: Appearance of new targets: Unmatched detection boxes are initialized as new tracking targets, assigned unique IDs, and the Kalman filter state is initialized; Disappearance of old targets: If the tracking target fails to match the detection box for Nmiss consecutive frames, its trajectory is terminated; Occlusion handling: For short-term occlusion, the tracking is maintained by predicting the position through the Kalman filter; for long-term occlusion, re-identification is performed by combining the material information of the segmentation mask.

[0015] Furthermore, use the Kalman filter to predict the position of the target in the next frame, specifically as follows: Assume that the conveyor belt movement is a linear uniform motion model, and the current state vector x t includes the position and speed of the target: ; where x and y are the coordinates of the target center; v x , v y is the velocity of the target in the conveyor belt coordinate system; the superscript T represents the transpose; Kalman filter prediction equation: ; where x t-1 is the state vector at the previous moment; F is the state transition matrix; is the process noise covariance matrix; is the predicted state of the current frame; w t is the weight coefficient; is the predicted covariance matrix; P t-1 is the error covariance matrix at the previous moment; State transition matrix F: ; where Δt is the time interval between two frames; When the detection box matches the tracking target, update the state: ; where H is the observation matrix; z t is the observation value; R is the observation noise covariance; K t is the Kalman gain; is the predicted covariance matrix; is the updated covariance matrix; x t is the updated state, and I is the identity matrix.

[0016] Furthermore, the detection box and the tracking target are associated through the Hungarian algorithm, as follows: Match the detection box of the current frame with the existing tracking targets, and construct a cost matrix using the Mahalanobis distance and the appearance feature distance : ; where λ1, λ 2, λ3 are the weighting coefficients; is the target and the target 's motion distance; is the target and the target 's appearance distance; is the target and the target 's class distance; Mahalanobis distance: ; where is the covariance matrix of the Kalman filter; is the target predicted state; is the target observed value; If DeepSORT is used, extract the depth features of the detection box and calculate the cosine distance; Calculate the category matching penalty term according to the category label output by the classification module; Minimize the total cost through the Hungarian algorithm to obtain the optimal matching between the detection box and the tracking target.

[0017] Furthermore, based on the motion state of S5 target tracking, predict the spatio-temporal coordinates of the garbage reaching the sorting port, generate sorting instructions according to the classification and segmentation results, and trigger the sorting device to act at the correct time through time-space mapping and delay compensation.

[0018] The present invention has the following beneficial effects: 1. By combining the garbage motion tracking technology in the conveyor belt environment with the real-time classification system, the present invention effectively solves the problems of mixed material identification and automatic sorting, not only improving the sorting efficiency, but also significantly increasing the resource recycling rate, providing important technical support for realizing green circular economy and sustainable development; 2. The garbage classification model based on the multi-modal neural network of the present invention can efficiently process complex multi-material garbage, accurately classify the garbage through joint features, and use the SegFormer deep segmentation network to accurately segment the material composition of the garbage at the pixel level, supporting the fine-grained identification of multi-material mixed objects (such as metal wrapped in plastic, composite materials, etc.); 3. Through the target tracking algorithm, the present invention can continuously track the garbage trajectory, ensure that the target is not lost in the dynamic environment on the conveyor belt, and based on the target tracking results, predict the position of the garbage on the conveyor belt in real time, providing accurate position information for the robotic arm or sorting equipment to ensure the accuracy and timeliness of the sorting action. Description of the Drawings

[0019] Figure 1 is the flowchart of the method of the present invention. Detailed Embodiment

[0020] The following further describes the present invention in detail with reference to the drawings and specific embodiments: Refer to Figure 1 , in this embodiment, a method for sorting composite material garbage based on hyperspectral technology and an AI classification model is provided, including the following steps: S1: Deploy several hyperspectral and RGB cameras at different nodes along the conveyor belt, and each node focuses on the multi-modal data collection of the garbage in one section; S2: Preprocess the collected multi-modal data, and fuse the RGB and hyperspectral data according to the timestamp and spatial position to obtain the joint features of the multi-modal data; S3: Build a classification model based on the multi-modal neural network model and classify according to the joint features; S4: Based on the classification results, use the SegFormer deep segmentation network to identify the composition of the garbage at the pixel level and achieve precise segmentation of objects composed of multiple materials; S5: Keep the target tracking of the garbage trajectory through the target tracking algorithm; S6: Based on the results of the target tracking in S5, predict the position of the garbage on the conveyor belt in real time, and accurately execute the sorting action according to the classification and segmentation results.

[0021] In this embodiment, S1 is specifically as follows: Measure the section length L and width W at different positions along the conveyor belt in advance; According to the garbage flow velocity v (the speed of the conveyor belt), combined with the density R of the garbage density (the number of garbage detected per unit time), determine the sampling frequency f of each node sampling ;

[0022] where d min is the minimum distance between adjacent garbage; According to the sampling field of view (Field of View, FOV) F view =(θ H , θ V ), where θ H is the horizontal field of view angle and θ V is the vertical field of view angle, determine the overlapping area to avoid blind spots in garbage detection, and the deployment spacing between adjacent nodes S nodes is: ; Each node includes an RGB camera, a hyperspectral imager, a trigger device and a controller; Suppose the garbage on the conveyor belt moves at a speed of v, and the sensor dynamically captures the multi-modal information of the object: calculate the length L capture and width W capture covered by each sampling based on the resolution and FOV of the sensor:

[0023] where h is the vertical height from the sensor to the conveyor belt, and: L capture ≥L, W capture ≥W; The displacement Δ of the garbage within two sampling time intervals T capture is calculated according to the following formula: x ; ; Ensure that Δx ≤ d min to ensure continuous sampling and no missed detection.

[0024] In this embodiment, S2 is specifically: Let the collected RGB data be R t (x, y, c), corresponding to the timestamp t R ; the hyperspectral data be H t (x, y, λ), corresponding to the timestamp t H ; where x, y represent the spatial dimensions, λ represents the light wave wavelength in the hyperspectral image; c is the number of channels; For each frame of RGB data R t , find the hyperspectral data H t with the closest timestamp to R t :

[0025] where, is the time difference; represents finding the t value that minimizes the time difference H in the set of candidate time points; After time alignment, the synchronized multimodal data pair (R t , H t ) is obtained; Spatial alignment is completed through the transformation matrix T align obtained by calibration, and the formula is:

[0026] where, [u, v] Hyper is the pixel coordinate of the hyperspectral data; [u, v] RGB is the pixel coordinate of the RGB data; Align the pixel coordinates of the hyperspectral data with the pixel coordinates of the RGB data to obtain spatially consistent multimodal data; After time and spatial alignment, for each pixel point, directly splice the RGB data and the hyperspectral data to form the joint feature Ft (x, y).

[0027] In this embodiment, the multi-modal neural network model includes a hyperspectral branch, an RGB branch, a multi-layer fusion strategy layer, and a classification head, specifically as follows: The hyperspectral branch inputs the preprocessed hyperspectral data , extracts spectral features, and captures material information at different wavelengths: Extracts local spectral-spatial features through a 3D convolutional layer : ; Among them, kernel is the convolutional kernel size; Channels is the number of convolutional output channels; Con3D represents the 3D convolution operation; The spectral Transformer module captures global spectral dependencies : ; Among them, Q, K, and V are the query vector, key vector, and value vector respectively; W Q , W K , W V are the projection parameters of the query vector, key vector, and value vector respectively; d k is the dimension of the attention space; T represents the transpose; Attention represents the attention mechanism; LayerNorm is the normalization operation; Finally, through the global average pooling operation, a spectral feature vector is generated f H : ; Among them, GlobalAvgPool represents the global average pooling operation; The RGB branch inputs the preprocessed RGB data R′ and extracts shape and texture features: Extracts multi-level spatial features through the ResNet-50 backbone network: ; Among them, ResNet gradually extracts multi-level features through 4 residual blocks ResNet-Block1, ResNet-Block2, ResNet-Block3, and ResNet-Block4, including , , , ; The spatial attention module enhances the features of key regions: ; Among them, Conv2D represents a convolutional operation; ⊙ is an element-wise weighted operation; σ is an activation function; w R (x, y) is the attention weight; is the output of the spatial attention module; Global average pooling to generate a spatial feature vector: ; The multi-level fusion strategy layer fuses hyperspectral and RGB features at multiple levels to enhance modal complementarity: Directly splice the feature maps of the two modalities at the intermediate layer to obtain the spliced features : ; Fuse the spliced features through a convolutional layer to obtain the fused features : ; The classification head outputs the classification result by passing the fused features through a fully connected layer: ; Among them, is the class probability; Softmax is the activation function; W c and b c are the weights and biases of the fully connected layer.

[0028] In this embodiment, for the few-shot composite materials (such as glass fiber composite and carbon fiber composite), the classification model introduces meta-learning to optimize the classification model, specifically as follows: Regarding different material classification tasks (such as glass fiber composite vs. carbon fiber composite) as the task set τ i ; Each task contains a support set Q' and a query set S; Model training process: Update the model parameters on the support set of the task set τ i : ; Among them, is the model parameter, is the learning rate; is the loss function corresponding to the meta-task τ i ; is the updated model parameter on the support set of the task τ i ; represents the gradient of the loss function with respect to the parameter θ; Calculate the loss of the query set and find the total gradient on all tasks to optimize the global parameters: ; wherein, is the meta learning rate; When new category of garbage data (such as hybrid composite materials) is input, it is quickly updated through the support set: .

[0029] In this embodiment, based on the classification result, the SegFormer deep segmentation network is used to identify the garbage composition components at the pixel level and achieve precise segmentation of objects composed of multiple materials, specifically as follows: The class probability output by the classification model is used as the prior condition of the segmentation network, and the joint feature and class probability of the multi-modal data are input ; Through conditional concatenation, the classification label is broadcast to each pixel point to form conditional features : ; wherein, is the extended conditional information; The SegFormer deep segmentation network includes an encoder and a decoder, The encoder adopts a pre-trained SegFormer-B5 model, with the input being Fcond(x,y), and outputs multi-scale features : ; wherein, SegFormerEncoder is the encoder of the SegFormer-B5 model; The decoder (Decoder) fuses the multi-scale features and generates a segmentation mask M seg : ; wherein, MLP represents the MLP decoder; The cross-entropy loss is adopted to process the pixel-level output: ; wherein, is the One-Hot representation of the pixel true material label; is the predicted segmentation mask; k is the material category index.

[0030] In this embodiment, S5 is specifically: Obtain the location, category, and segmentation mask information of the target from the classification and segmentation modules: The minimum bounding rectangle b is generated from the pixel-level mask of the a-th item output by the segmentation module boxa : b boxa=[x min ,y min ,x max ,y max ; Among them, [x min ,y min is the coordinate of the upper left corner of the rectangular box, and [x max ,y max is the coordinate of the lower right corner of the rectangular box; the classification module outputs the category probability vector of the i-th item ; the pixel-level material segmentation result M i (x,y) output by SegFormer; Use the Kalman filter to predict the position of the target in the next frame; and associate the detection box with the tracking target through the Hungarian algorithm; Process the appearance of new targets, the disappearance of old targets, or occlusion according to the following rules: Appearance of new targets: Unmatched detection boxes are initialized as new tracking targets, assigned unique IDs, and the Kalman filter state is initialized; Disappearance of old targets: If the tracking target has not been matched with the detection box for Nmiss consecutive frames, its trajectory is terminated; Occlusion handling: For short-term occlusion, maintain tracking by predicting the position through the Kalman filter; for long-term occlusion, perform re-identification by combining the material information of the segmentation mask.

[0031] In this embodiment, the Kalman filter is used to predict the position of the target in the next frame, specifically as follows: Assume that the conveyor belt movement is a linear uniform motion model, and the current state vector x t includes the position and speed of the target: ; Among them, x and y are the center coordinates of the target; v x ,v y is the speed of the target in the conveyor belt coordinate system; the superscript T represents the transpose; Kalman filter prediction equation: ; Among them, x t-1 is the state vector at the previous moment; F is the state transition matrix; is the process noise covariance matrix; is the predicted state of the current frame; w t is the weight coefficient; is the predicted covariance matrix; P t-1 is the error covariance matrix at the previous moment; State transition matrix F: ; where, Δt is the time interval between two frames; After the detection box matches the tracking target, update the state: ; where, H is the observation matrix; z t is the observation value; R is the observation noise covariance; K t is the Kalman gain; is the predicted covariance matrix; is the updated covariance matrix; x t is the updated state, and I is the identity matrix.

[0032] In this embodiment, the detection box and the tracking target are associated by the Hungarian algorithm, specifically as follows: Match the detection box of the current frame with the existing tracking targets, and use the Mahalanobis distance (motion similarity) and the appearance feature distance (category and segmentation consistency) to construct a cost matrix : ; where, λ1, λ 2, λ3 are the weighting coefficients; is the and motion distance of the target; is the and appearance distance of the target; is the and category distance of the target; Mahalanobis distance (motion similarity): ; where, is the covariance matrix of the Kalman filter; is the predicted state of the target; is the observed value of the target; If DeepSORT is used, extract the depth features of the detection box (such as the embedding vector of ResNet-50) and calculate the cosine distance; Calculate the category matching penalty term according to the category label output by the classification module; Minimize the total cost through the Hungarian algorithm to obtain the optimal matching between the detection box and the tracking target.

[0033] In this embodiment, S6 is specifically as follows: Based on the motion state of the target tracking in S5, predict the spatio-temporal coordinates of the garbage reaching the sorting port, generate a sorting instruction according to the classification and segmentation results (material, category), and trigger the sorting device (such as a pneumatic valve, robotic arm) to act at the correct time through time-space mapping and delay compensation.

[0034] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0035] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0036] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0037] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0038] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in any other form. Any person skilled in the relevant art may make changes or modifications using the technical content disclosed above to obtain equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for sorting composite waste based on hyperspectral technology and an AI classification model, characterized in that, It includes the following steps: S1: Deploy several hyperspectral and RGB cameras at different nodes along the conveyor belt, and each node focuses on collecting multimodal data of the garbage in one section; S2: Preprocess the collected multimodal data, and fuse the RGB and hyperspectral data according to the timestamp and spatial position to obtain the joint features of the multimodal data; S3: Construct a classification model based on the multimodal neural network model and classify according to the joint features; S4: Based on the classification results, use the SegFormer deep segmentation network to identify the composition of the garbage at the pixel level and achieve the segmentation of objects composed of multiple materials; S5: Keep the target tracking of the garbage trajectory through the target tracking algorithm; S6: Based on the results of the target tracking in S5, predict the position of the garbage on the conveyor belt in real time, and perform sorting actions according to the classification and segmentation results.

2. The composite material waste sorting method based on hyperspectral technology and AI classification model according to claim 1, wherein The specific content of S1 is as follows: Pre-measure the section length L and width W at different positions along the conveyor belt; According to the garbage flow velocity v, combined with the density R of the garbage density , determine the sampling frequency f of each node sampling ; ; where d min is the minimum distance between adjacent pieces of refuse; According to the sampling field of view F view =(θ H ,θ V ), where θ H is the field of view angle in the horizontal direction, θ V is the field of view angle in the vertical direction, determine the overlapping area, avoid the blind area of garbage detection, and the deployment spacing of adjacent nodes S nodes is: ; Each node includes an RGB camera, a hyperspectral imager, a trigger device and a controller; Suppose the garbage on the conveyor belt moves at a speed of v, and the sensor dynamically captures the multi-modal information of the object: calculate the length of the conveyor belt covered by each sampling based on the resolution and FOV of the sensor L capture and width W capture : ; where h is the vertical height from the sensor to the conveyor belt, and: L capture ≥ L, W capture ≥ W; The displacement Δ of the garbage within two sampling time intervals T capture is calculated according to the following formula: x as follows: ; Ensure that Δx ≤ d min , to ensure continuous sampling and no missed detection.

3. The composite material waste sorting method based on hyperspectral technology and AI classification model according to claim 2, wherein The specific content of S2 is: Let the collected RGB data be R t (x, y, c), corresponding to the time stamp t R ; and the hyperspectral data be H t (x, y, λ), corresponding to the time stamp t H ; where x and y represent the spatial dimensions, λ represents the light wave wavelength in the hyperspectral image; c is the number of channels; For each frame of RGB data R t , find the hyperspectral data H t whose timestamp is closest to that of R t : ; Among them, is the time difference; represents finding the t value that minimizes the time difference H in the set of candidate time points; After time alignment, a synchronized multi-modal data pair (R t , H t ) is obtained; The transformation matrix T obtained through calibration for spatial alignment align is completed, and the formula is: ; where [u, v] Hyper is the pixel coordinate of the hyperspectral data; [u, v] RGB is the pixel coordinate of the RGB data; Align the pixel coordinates of the hyperspectral data with the pixel coordinates of the RGB data to obtain spatially consistent multimodal data; After time and space alignment, for each pixel, the RGB data and hyperspectral data are directly concatenated to form the joint feature F t (x,y).

4. The composite material waste sorting method based on hyperspectral technology and AI classification model according to claim 1, characterized in that The multimodal neural network model includes a hyperspectral branch, an RGB branch, a multi-layer fusion strategy layer and a classification head, which are specifically as follows: The hyperspectral branch inputs the preprocessed hyperspectral data , extracts spectral features, and captures material information at different wavelengths: Extract local spectral-spatial features through 3D convolutional layers : ; where, kernel is the convolution kernel size; Channels is the number of convolution output channels; Con3D represents the 3D convolution operation; Spectral Transformer module, capturing global spectral dependencies : ; Among them, Q, K, and V are the query vector, key vector, and value vector respectively; W Q , W K , W V are the projection parameters of the query vector, key vector, and value vector respectively; d k is the dimension of the attention space; T represents transpose; Attention represents the attention mechanism; LayerNorm is the normalization operation; Finally, through the global average pooling operation, a spectral feature vector is generated. f H : ; where, GlobalAvgPool represents the global average pooling operation; The RGB branch inputs the preprocessed RGB data R′ and extracts shape and texture features: Extract multi-level spatial features through the ResNet-50 backbone network: ; Among them, ResNet gradually extracts multi-level features through 4 residual blocks, namely ResNet-Block1, ResNet-Block2, ResNet-Block3, and ResNet-Block4, including , , , ; Spatial attention module to enhance the features of key regions: ; where, Conv2D represents the convolution operation; ⊙ is the element-wise weighted operation; σ is the activation function; w R (x, y) is the attention weight; is the output of the spatial attention module; Global average pooling to generate a spatial feature vector: ; Multi-layer fusion strategy layer to fuse hyperspectral and RGB features at multiple levels and enhance modal complementarity: Directly splice the feature maps of the two modalities in the intermediate layer to obtain the spliced features : ; After splicing, it is fused through a convolutional layer to obtain the fused features : ; The classification head outputs the classification results of the fused features through the fully connected layer: ; Among them, is the class probability; Softmax is the activation function; W c and b c are the weights and biases of the fully connected layer.

5. The composite material waste sorting method based on hyperspectral technology and AI classification model according to claim 4, characterized in that For the classification model of few-shot composite materials, meta-learning is introduced to optimize the classification model, which is specifically as follows: Take the task of classifying different materials as the task set τ of meta-learning i ; Each task contains a support set S = (X, Y) and a query set Q’ = (X′, Y′); Model training process: Update the model parameters on the support set of task set τ i : ; Among them, is the model parameter, is the learning rate; is the meta-task τ i corresponding loss function; is the model parameter updated on the support set of the task set τ i ; represents the gradient of the loss function with respect to the parameter θ; Calculate the loss of the query set and calculate the total gradient on all tasks to optimize the global parameters: ; Among them, is the meta learning rate; When inputting new category garbage data, quickly update through the support set: 。 6. The composite material waste sorting method based on hyperspectral technology and AI classification model according to claim 3, characterized in that, Based on the classification results, use the SegFormer deep segmentation network to identify the composition of the garbage at the pixel level and achieve the segmentation of objects composed of multiple materials, which is specifically as follows: The class probabilities output by the classification model As a prior condition for the segmentation network, input the joint features and class probabilities of the multimodal data ; Broadcast the classification label to each pixel through conditional splicing to form conditional features : ; Among them, is the expanded conditional information; The SegFormer deep segmentation network includes an encoder and a decoder, The encoder uses a pre-trained SegFormer-B5 model, with the input being Fcond(x,y) and the output being multi-scale features : ; where, SegFormerEncoder is the encoder of the SegFormer-B5 model; The decoder fuses multi-scale features and generates the segmentation mask M seg : ; where, MLP represents the MLP decoder; Using cross-entropy loss Processing pixel-level output: ; where, is the One-Hot representation of the true material label of the pixel; is the predicted segmentation mask; k is the material category index.

7. The composite material waste sorting method based on hyperspectral technology and AI classification model according to claim 1, characterized in that, The specific content of S5 is: Obtain the position, category, and segmentation mask information of the target from the classification and segmentation module: Generate the minimum bounding rectangle b from the pixel-level mask of the a-th item output by the segmentation module boxa : b boxa =[x min ,y min ,x max ,y max ; Among them, [x min , y min is the coordinate of the upper left corner of the rectangular box, and [x max , y max is the coordinate of the lower right corner of the rectangular box; the classification module outputs the category probability vector of the i-th item ; the pixel-level material segmentation result M i (x, y); Use Kalman filtering to predict the position of the target in the next frame; and use the Hungarian algorithm to associate the detection boxes with the tracked targets; Process the appearance of new targets, disappearance of old targets, or occlusions according to the following rules: Appearance of new targets: Unmatched detection boxes are initialized as new tracked targets, assigned unique IDs, and the Kalman filtering state is initialized; Disappearance of old targets: If a tracked target fails to match a detection box for Nmiss consecutive frames, its trajectory is terminated; Occlusion handling: For short-term occlusions, maintain tracking by predicting the position using Kalman filtering; for long-term occlusions, perform re-identification by combining the material information of the segmentation mask.

8. The composite material waste sorting method based on hyperspectral technology and AI classification model according to claim 7, characterized in that The use of Kalman filtering to predict the position of the target in the next frame is specifically as follows: Assume that the conveyor belt movement is a linear uniform motion model, and the current state vector x t includes the position and velocity of the target: ; where x and y are the target center coordinates; v x , v y is the velocity of the target in the conveyor belt coordinate system; the superscript T represents the transpose; Kalman filtering prediction equation: ; Among them, x t-1 is the state vector at the previous moment; F is the state transition matrix; is the process noise covariance matrix; is the predicted state of the current frame; w t is the weight coefficient; is the predicted covariance matrix; P t-1 is the error covariance matrix at the previous moment; State transition matrix F: ; where Δt is the time interval between two frames; After the detection box matches the tracked target, update the state: ; Among them, H is the observation matrix; z t is the observed value; R is the observation noise covariance; K t is the Kalman gain; is the predicted covariance matrix; is the updated covariance matrix; x t is the updated state, and I is the identity matrix.

9. The composite material waste sorting method based on hyperspectral technology and AI classification model according to claim 7, wherein The association of the detection box with the tracked target by the Hungarian algorithm is specifically as follows: Match the detection boxes of the current frame with the existing tracking targets, and construct a cost matrix using Mahalanobis distance and appearance feature distance : ; Among them, λ1, λ 2, λ3 are weighting coefficients; is the and the moving distance of the target; is the and the appearance distance of the target; is the and the category distance of the target; Mahalanobis distance: ; Among them, is the covariance matrix of the Kalman filter; is the predicted state of the target ; is the observed value of the target ; If DeepSORT is used, extract the depth features of the detection box and calculate the cosine distance; According to the class label output by the classification module, calculate the class matching penalty term; Minimize the total cost through the Hungarian algorithm to obtain the optimal matching between the detection box and the tracked target.

10. The composite material waste sorting method based on hyperspectral technology and AI classification model according to claim 1, characterized in that, Specifically, S6 is: Based on the motion state of the target tracking in S5, predict the spatio-temporal coordinates of the garbage reaching the sorting outlet, generate sorting instructions according to the classification and segmentation results, and trigger the sorting device to act at the correct time through time-space mapping and delay compensation.

Citation Information

Patent Citations

  • Multi-dimensional garbage identification and classification system

    CN113102266A

  • Intelligent household garbage sorting system and method based on machine vision

    CN120001667A

  • Solid waste identification and segregation system

    US20160078414A1

  • Hyperspectral remote sensing image classification method based on self-attention context network

    US20230260279A1

Cited By

  • Garbage sorting method based on artificial intelligence

    CN120644394A

  • Multi-dimensional intelligent food material analysis system based on artificial intelligence

    CN120976917A

  • Coal mining machine roller tracking system and method

    CN121010931A

  • Small sample electric aircraft composite material combined structure dynamics modeling method based on attention element learning

    CN121072017A

  • A method for dynamic modeling of composite material structures in electric aircraft based on attention meta-learning

    CN121072017B