Intelligent garbage sorting method and system based on multi-modal recognition
By combining multimodal recognition technology with RGB images, depth sensors and near-infrared spectral sensors, the existing garbage sorting system is solved, and the problem of light sensitivity and difficulty in identifying complex garbage is achieved, efficient garbage recognition and sorting is achieved, and the stability and efficiency of the system are improved.
Patent Information
- Application Number
- CN202510696230.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing garbage sorting system relies on a single RGB visual information, and has problems such as being sensitive to ambient lighting conditions, being difficult to identify garbage with similar materials but different compositions, and being difficult to accurately judge the spatial position relationship of objects in complex stacking scenarios.
Using an intelligent garbage sorting method based on multimodal recognition, combining the data of RGB images, depth sensors and near-infrared spectral sensors, a unified bounding box and classification results are generated through the HybridNet model, and the three-dimensional point cloud data of the object is used to calculate the number of stacked layers at the location of the object, priority is given to the surface object, initial grabbing paths are generated and optimized through DQN depth reinforcement learning.
It realizes comprehensive and efficient identification and sorting of different types of garbage, solves the problem of robots crawling and optimization in dynamic environments, and improves the overall crawling efficiency and the stability of system operation.
Smart Images

Figure CN120206543A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of garbage sorting, and in particular to an intelligent garbage sorting method and system based on multimodal recognition. Background Art
[0002] With the acceleration of urbanization and the improvement of people's living standards, garbage disposal has become an environmental problem that needs to be solved urgently. The traditional manual garbage sorting method has problems such as low efficiency, high cost, and poor working environment. At the same time, the garbage is of various types, different shapes, and complex stacking, which poses a severe challenge to the automated sorting system. Therefore, the development of an efficient and accurate intelligent garbage sorting system is of great practical significance.
[0003] The current mainstream garbage sorting system mainly relies on single RGB visual information for identification and classification. However, this method has the following problems: it is sensitive to ambient lighting conditions and is easily disturbed by shadows, reflections, etc.; it is difficult to accurately identify garbage with similar materials but different components (such as different types of plastics); in complex stacking scenes, single visual information is difficult to accurately determine the spatial position relationship of objects. Summary of the invention
[0004] In order to solve the above problems, the purpose of the present invention is to provide an intelligent garbage sorting method and system based on multimodal recognition, which can realize comprehensive and efficient identification and sorting of different types of garbage.
[0005] To achieve the above object, the present invention adopts the following technical solutions: An intelligent garbage sorting method based on multimodal recognition comprises the following steps: S1: Use the embedded camera to collect the RGB image of the conveyor belt, and use the depth sensor and near-infrared spectrum sensor to capture the three-dimensional shape information and material information of the objects on the conveyor belt; S2: preprocessing the collected image data and sensor data to remove noise and align the modalities; S3: Based on the preprocessed image data and sensor data, a unified bounding box and classification results are generated based on the HybridNet model; S4: Based on the bounding box and classification results output by the HybridNet model, the discrete PID algorithm is used to calculate the motion instructions, and the depth sensor is used to calculate the three-dimensional point cloud data of the object, determine the number of stacking layers at the location of the object, give priority to grasping the surface objects, and generate the initial grasping path based on the position error and stacking priority; S5: Based on the initial grasping path, use DQN deep reinforcement learning to optimize the initial grasping path to the optimal grasping path.
[0006] Furthermore, the noise removal is as follows: The RGB image data is processed by bilateral filtering to remove noise: ; where I(x, y) is the pixel value of the RGB image at point (x, y); is a spatial-domain Gaussian kernel function related to the distance to control the neighborhood range; is a Gaussian kernel function for pixel difference to control color similarity; is a normalization coefficient; is the pixel value of the image after noise removal by bilateral filtering at point (x, y); k is the radius of the filtering window; ; ; where , are typical parameter values, which are used to control the filtering range and pixel similarity respectively; And local brightness normalization is performed on the RGB image to make the brightness gradient uniform: ; where is the mean value within the sliding window; is the standard deviation within the sliding window; The three-dimensional shape information data obtained by the depth sensor uses RANSAC to fit a plane model to remove outliers; The material information obtained by the near-infrared spectral sensor uses reflectance calibration and Savitzky-Golay smoothing.
[0007] Furthermore, the alignment modality is to strictly align the RGB image with the depth point cloud and spectral information by combining the pose calibration algorithm and the multi-modal data alignment process, specifically as follows: Unify the coordinate systems of each sensor so that each sensor works in the same reference coordinate system. The embedded camera uses the calibration board method to estimate the internal parameter matrix K: ; where f x and f y are the focal lengths, and c x and c y are the principal point optical centers; The external parameter calibration uses the embedded camera as the reference coordinate system and completes the rigid body transformation between the depth sensor and the camera with the rotation matrix R and the translation vector T: ; where P RGB and Pdepth respectively represent the pixel coordinates of the RGB image and the depth point cloud; Map the depth data onto the pixel grid of the RGB image: ; where (X, Y, Z) are the three-dimensional point coordinates in the depth data; (x, y) are the 2D coordinates on the image plane; Due to the characteristics of the single-line scanning of the near-infrared spectral sensor, align the spectral points based on the calibrated geometric transformation to the RGB image: ; where, is the homography matrix projected from the near-infrared spectral sensor to the RGB image.
[0008] Furthermore, the HybridNet model consists of an input layer, a multi-modal feature extraction layer, a feature fusion module, and a joint detection head, specifically as follows: The input layer inputs the preprocessed image data and sensor data. The multi-modal feature extraction layer processes the RGB image, depth map, and near-infrared spectral data respectively, and converts them into the representative features F RGB , F Depth and F NIR ; The feature fusion module fuses multi-modal features based on the attention mechanism and feature concatenation; ; where, α RGB , α Depth and α NIR are the attention weights calculated for each modal feature: F fusion is the fused multi-modal feature; Dynamically allocate the weights of each modality α = {α RGB , α Depth , α NIR} through the channel attention mechanism: α = σ(Conv3×3(FRGB⊕FDepth⊕FNIR)) where, ⊕ is element-wise addition; σ is the Sigmoid activation function; α represents the dependence of each position on RGB / Depth / NIR; Conv3×3 represents a 3×3 convolutional network; The joint detection head uses a shared fully convolutional network to extract the fused features for position detection and classification, and generates bounding box predictions and classification probabilities; Generate the following two outputs according to the HybridNet model: Bounding box B i ={xi , y i , w i , h i}, where (x i , y i ) are the coordinates of the i-th bounding box; w i , h i are the width and height of the i-th bounding box respectively; the detection position and size of the object in the image space; Classification result P class (i), the probability distribution of the target class, select the class label C corresponding to the maximum probability i = argmax ; According to the target sorting requirements, filter out the objects that do not belong to the target class, and only retain the classes of interest as the control input. Suppose there are N targets in the image, then the output target set O is: .
[0009] Further, the joint detection head includes a bounding box regression branch and a classification branch, specifically as follows: The bounding box regression branch presets p scales and q aspect ratios, with a total of p*q anchor boxes. The center position of the anchor box (x a , y a ) is calculated according to the feature map grid and is associated with the size of the input image: The regression branch outputs the offset of each anchor box ΔB=(Δx,Δy,Δw,Δh), where (Δx,Δy) are the coordinate offsets of the bounding box; Δw,Δh are the width and height offsets of the bounding box respectively, and calculate the final bounding box: ; Among them, and are the width and height of the center of the anchor box respectively; are the coordinates of the final bounding box; w, h are the width and height of the final bounding box respectively; Use Smooth L1 Loss to measure the error between the bounding box predicted by the model and the true label: ; Among them, is the number of positive samples; ΔB i is the offset value of the i-th bounding box predicted by the model; is the offset value of the true bounding box; smooth L1 is the Smooth L1 loss formula; The classification branch uses the Softmax function to predict the probability distribution of the target class P class: ; Among them, FC represents the fully connected layer; The cross-entropy loss is adopted to measure the difference between the classification probability predicted by the model and the true classification label: ; Among them, is the true classification label; is the predicted probability; Considering the dual objectives of regression and classification in object detection, the joint loss function L total is as follows: ; Among them, , and are weight coefficients; is the fusion weight regularization term.
[0010] Furthermore, based on the bounding boxes and classification results output by the HybridNet model, a discrete PID algorithm is used to calculate the motion instructions, specifically as follows: The image bounding boxes are located in the real-world coordinates through the depth sensor point cloud, and the target sorting area position required by the sorting task is set. The three-dimensional physical position of the target is extracted using the depth sensor point cloud data, and the centroid of the object is used as the actual coordinate of the object P actual,u : ; Among them, ([[]] ) is the actual coordinate of object u, z actual,u is the actual depth value of the object; Let the target position of the sorting area be P target =(x target , y target , z target ), where (x target , y target , z target ) is the actual coordinate of the target position, and the error e between the target position and the actual position is calculated u : ; The motion instructions of the robotic arm are generated to the target object through the discrete PID control algorithm, and the discrete PID control dynamically adjusts the motion instructions according to the position error: ; Among them, is the motion increment of the current control instruction in the x direction; is the current time step error; is the sampling period; K p 、K i 、K d are the proportional, integral, and derivative coefficients of the PID control respectively; is the cumulative error; The same formula is used to calculate other directions (u y , u z ), and finally integrated into the motion instruction of the robotic arm: .
[0011] Furthermore, the depth sensor is used to calculate the three-dimensional point cloud data of the object, determine the stacking layer where the object is located, and preferentially grasp the surface object, specifically as follows: Through the three-dimensional point cloud of the depth sensor, the layer number is estimated in the z-axis direction: Analyze the highest point z of the target object in the i-th bounding box max,i : ; Among them, is the layer number estimation of the i-th bounding box; is the lowest reference height; is the layer division step size; If the point cloud projection of the upper object overlaps with the lower object, it is considered that the target is blocked and not preferentially grasped. The occlusion rate is defined as: ; Among them, is the occlusion rate, is the overlapping area, is the total area of the target bounding box; Allocate priorities according to the layer where the object is located and the occlusion situation : .
[0012] Furthermore, an initial grasping path based on the position error and stacking priority is generated, specifically as follows: According to the target path error e u and the stacking priority, plan the grasping path, sort it according to the target priority, and select the target with the highest priority as the current grasping task P athsorted : P athsorted =Sort(O, Priority i ) Based on the actual position P of the targetactual,u and the target point P in the sorting area target Generate the robotic arm path and plan a straight-line path P from the current position of the robotic arm end to the target object ath (t): ; where P start is the starting point; If the closest distance dobs between the path and the obstacle point cloud is less than , is the obstacle threshold, and a detour point P is inserted into the path avoid : ; where n is the normal vector of the obstacle surface; is the modulus of the normal vector of the obstacle surface; Insert an intermediate point P at a preset safety height above the object P actual,u ; actual,usafe ; Finally, obtain the initial grasping path P ath as: P ath ={P start ,P avoid ,P actual,usafe, ,P actual,u ,P target};
[0013] Furthermore, based on the initial grasping path, use DQN deep reinforcement learning to optimize the initial grasping path into the optimal grasping path, as follows: Input the initial grasping path s0 and initialize the DQN deep reinforcement learning network: s0={P start ,P target ,P obs ,StackInfo}; where StackInfo is information about the number of stacked objects and occlusion; P start is the three-dimensional coordinate of the start end of the robotic arm; P obs is the set of currently scanned obstacles; starting state; Action set A = {±Δa, ±Δb, ±Δc, grasping action}; where ±Δa, ±Δb, ±Δc represent fine-tuning actions in the three-dimensional space along the x, y, and z directions; State update and action selection: At the t-th step, use the greedy policy to select an action: ; Among them, is the state s t and the action a t The Q value of is predicted by DQN; the parameter θ represents the network weight of DQN; Random exploration avoids falling into local optima, and optimizes the use of existing Q values to guide path movement; In the current state s t Execute the action a t to obtain a new state s t+1 and the reward r t ; Update the coordinate positions in the path; Calculate the target Q value: ; Among them, y t is the target Q value; γ is the discount factor; θ − are the fixed parameters of the target network; represents the target network Q target 's prediction of the Q value of the action a' at the next moment state s t+1 ; Update the current Q network parameters; When the reward value r t converges, or when the path length and the number of obstacle avoidance times meet the optimization requirements, output the optimized grasping path.
[0014] An automatic sorting system for recycled resources combined with intelligent recognition, including a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the above-mentioned intelligent garbage sorting method based on multi-modal recognition.
[0015] The present invention has the following beneficial effects: 1. The present invention realizes comprehensive and efficient recognition and sorting of different types of garbage, effectively solving the grasping optimization problem of robots in dynamic environments (such as stacking and obstacle interference); 2. The present invention realizes the common constraint of multi-modal features, classification, and regression through a multi-task loss function. Based on the HybridNet model, it can efficiently complete the target bounding box prediction and classification tasks, and dynamically adjust the degree of help of multi-modal for detection to achieve the optimal performance; 3. The present invention uses a depth sensor to generate point cloud data, calculates the number of layers of the stack where the object is located, can dynamically determine the level of each object in the garbage stacking scenario, preferentially grabs the surface objects, reduces the interference to the stacking structure, improves the overall grasping efficiency and the stability of the system operation, and uses discrete PID to preliminarily plan the motion path. Through the calculation based on the position error and the grasping priority, an initial grasping path is generated. Finally, based on DQN deep reinforcement learning, path optimization can be achieved in a dynamic multi-object and multi-obstacle environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following further describes the present invention in detail with reference to the drawings and specific embodiments: Referring to Figure 1 , in this embodiment, an intelligent garbage sorting method based on multi-modal recognition is provided, including the following steps: S1: Use an embedded camera to collect the RGB image of the conveyor belt, and use a depth sensor and a near-infrared spectroscopy sensor to capture the three-dimensional shape information and material information of the objects on the conveyor belt; S2: Preprocess the collected image data and sensor data to remove noise and align the modalities; S3: Based on the preprocessed image data and sensor data, generate a unified bounding box and classification result based on the HybridNet model; S4: Based on the bounding box and classification result output by the HybridNet model, use the discrete PID algorithm to calculate the motion instruction, and use the depth sensor to calculate the three-dimensional point cloud data of the object, judge the stacking layer number of the object's location, preferentially grab the surface objects, and generate an initial grasping path based on the position error and the stacking priority; S5: Based on the initial grasping path, use DQN deep reinforcement learning to optimize the initial grasping path into an optimal grasping path.
[0018] In this embodiment, the noise removal process is as follows: The RGB image data is processed using bilateral filtering to remove noise: ; where I(x,y) is the pixel value of the RGB image at the point (x,y); is a spatial domain Gaussian kernel function related to the distance , controlling the neighborhood range; is the Gaussian kernel function of pixel difference, controlling the color similarity; is the normalization coefficient; The pixel value of the image after noise removal by bilateral filtering at the point (x, y); k is the radius of the filtering window; ; ; Among them, 、 are typical parameter values, which are used to control the filtering range and pixel similarity respectively; And perform local brightness normalization on the RGB image to make the brightness gradient uniform: ; Among them, is the mean value within the sliding window; is the standard deviation within the sliding window; The three-dimensional shape information data obtained by the depth sensor uses RANSAC to fit a plane model to remove outliers; The material information obtained by the near-infrared spectral sensor adopts reflectance calibration and Savitzky-Golay smoothing.
[0019] In this embodiment, the alignment mode is to strictly align the RGB image with the depth point cloud and spectral information by combining the attitude calibration algorithm and the multi-modal data alignment process, specifically as follows: Unify the coordinate systems of each sensor so that each sensor works in the same reference coordinate system. The embedded camera uses the calibration board method to estimate the internal parameter matrix K: ; Among them, f x and f y are the focal lengths, and c x and c y are the principal point optical centers; The external parameter calibration takes the embedded camera as the reference coordinate system and completes the rigid body transformation of the depth sensor and the camera with the rotation matrix R and the translation vector T: ; Among them, P RGB and P depth respectively represent the pixel point coordinates of the RGB image and the depth point cloud; Map the depth data to the pixel grid of the RGB image: ; Among them, (X, Y, Z) are the three-dimensional point coordinates in the depth data; (x, y) are the 2D coordinates on the image plane; Due to the single-line scanning characteristic of the near-infrared spectral sensor, align the spectral points to the RGB image based on the calibrated geometric transformation: ; Among them, is the homography matrix projected from the near-infrared spectral sensor to the RGB image.
[0020] In this embodiment, the HybridNet model consists of an input layer, a multi-modal feature extraction layer, a feature fusion module, and a joint detection head, specifically as follows: The input layer inputs the preprocessed image data and sensor data. The multi-modal feature extraction layer processes the RGB image, depth map, and near-infrared spectral data respectively, and converts them into characteristic features F RGB , F Depth and F NIR ; The feature fusion module fuses multi-modal features based on the attention mechanism and feature splicing; ; Among them, α RGB , α Depth and α NIR are the attention weights calculated for each modal feature: F fusion is the fused multi-modal feature; Dynamically allocate the weights of each modality through the channel attention mechanism α={α RGB , α Depth , α NIR}: α = σ(Conv3×3(F RGB ⊕ F Depth ⊕ F NIR )) Among them, ⊕ is element-wise addition; σ is the Sigmoid activation function; α represents the dependence degree of each position on RGB / Depth / NIR; Conv3×3 represents a 3×3 convolutional network; The joint detection head uses a shared fully convolutional network to extract the fused features for position detection and classification, and generates bounding box predictions and classification probabilities; Generate the following two outputs according to the HybridNet model: Bounding box B i ={x i , y i , w i , h i}, where (x i , y i ) are the coordinates of the i-th bounding box; w i , h i are the width and height of the i-th bounding box respectively; the detection position and size of the object in the picture space; Classification result Pclass (i) The probability distribution of the target category, and select the category label C corresponding to the maximum probability i = argmax ; According to the target sorting requirements, filter out the objects that do not belong to the target category, and only retain the categories of interest as the control input. Suppose there are N targets in the image, then the output target set O is: .
[0021] In this embodiment, the joint detection head includes a bounding box regression branch and a classification branch, specifically as follows: For the bounding box regression branch, p scales and q aspect ratios are preset, with a total of p*q anchor boxes. The center position (x a , y a ) of the anchor box is calculated according to the feature map grid and is associated with the size of the input image: The regression branch outputs the offset ΔB=(Δx,Δy,Δw,Δh) of each anchor box, where (Δx,Δy) is the coordinate offset of the bounding box; Δw and Δh are the width and height offsets of the bounding box respectively. Calculate the final bounding box: ; Among them, and are the width and height of the center of the anchor box respectively; are the coordinates of the final bounding box; w and h are the width and height of the final bounding box respectively; Adopt Smooth L1 Loss to measure the error between the bounding box predicted by the model and the true label: ; Among them, is the number of positive samples (anchor boxes with corresponding targets); ΔB i is the offset value of the i-th bounding box predicted by the model; is the offset value of the true bounding box; smooth L1 is the Smooth L1 loss formula; The classification branch uses the Softmax function to predict the probability distribution of the target category P class : ; Among them, FC represents the fully connected layer; Adopt cross-entropy loss to measure the difference between the classification probability predicted by the model and the true classification label: ; Among them, is the true classification label (one-hot encoded, with a value of 1 indicating that class i is the true label and 0 otherwise); is the predicted probability; Taking into account the dual objectives of regression (position prediction) and classification (class prediction) in object detection, the joint loss function L total is as follows: ; where , and are weight coefficients; is the fusion weight regularization term.
[0022] In this embodiment, based on the bounding boxes and classification results output by the HybridNet model, a discrete PID algorithm is used to calculate the motion instructions, specifically as follows: The image bounding boxes are located in the real-world coordinates through the depth sensor point cloud, and the target sorting area position required by the sorting task is set. The three-dimensional physical position of the target is extracted using the depth sensor point cloud data, and the centroid of the object is used as the actual coordinate of the object P actual,u : ; where ([[]]END]] ) is the actual coordinate of object u, and z actual,u is the actual depth value of the object; Let the target position of the sorting area be P target =(x target ,y target ,z target ), where (x target ,y target ,z target ) is the actual coordinate of the target position, and the error e u between the target position and the actual position is calculated: ; The motion instructions of the robotic arm are generated to the target object through the discrete PID control algorithm, and the discrete PID control dynamically adjusts the motion instructions according to the position error: ; where is the motion increment of the current control instruction in the x direction; is the current time step of the error; is the sampling period; K p 、K i 、K d are the proportional, integral, and differential coefficients of the PID control respectively; is the cumulative error; Use the same formula to calculate in other directions (u y ,u z ), and finally integrate them into the motion command of the robotic arm: .
[0023] In this embodiment, the depth sensor is used to calculate the three-dimensional point cloud data of the object, judge the stacking layer number of the object's location, and preferentially grasp the surface object, specifically as follows: Estimate the number of layers in the z-axis direction through the three-dimensional point cloud of the depth sensor: Analyze the highest point z of the target object in the i-th bounding box max,i : ; where, is the layer number estimation of the i-th bounding box; is the lowest reference height; is the layer division step size; If the point cloud projection of the upper object overlaps with the lower object, it is considered that the target is occluded and not preferentially grasped. The occlusion rate is defined as: ; where, is the occlusion rate, is the overlapping area, is the total area of the target bounding box; Allocate priorities according to the layer number and occlusion situation of the object : .
[0024] In this embodiment, an initial grasping path based on position error and stacking priority is generated, specifically as follows: According to the target path error e u and stacking priority, plan the grasping path, sort according to the target priority, and select the target with the highest priority as the current grasping task P athsorted : P athsorted =Sort(O,Priority i ) Based on the actual position P of the target actual,u and the target point P in the sorting area target Generate the robotic arm path, and plan the straight-line path P ath (t) from the current position of the robotic arm end to the target object: ; where, P startis the starting point; If the minimum distance dobs between the path and the obstacle point cloud is < , is the obstacle threshold, insert a detour point P avoid : ; where n is the normal vector of the obstacle surface; is the modulus of the normal vector of the obstacle surface; Above the object P actual,u , insert an intermediate point P actual,u at a preset safe height (e.g., z = z actual,usafe + 0.1); Finally, obtain the initial grasping path P ath as: P ath ={P start , P avoid , P actual,usafe , P actual,u , P target};
[0025] In this embodiment, based on the initial grasping path, the DQN deep reinforcement learning is used to optimize the initial grasping path into the optimal grasping path, specifically as follows: Input the initial grasping path s0, and initialize the DQN deep reinforcement learning network: s0={P start , P target , P obs , StackInfo}; where StackInfo is information about the number of stacked objects and occlusion; P start is the three-dimensional coordinate of the start and end of the robotic arm; P obs is the set of currently scanned obstacles; starting state; The action set A = {±Δa, ±Δb, ±Δc, grasping action}; where ±Δa, ±Δb, ±Δc represent fine-tuning actions in the three-dimensional space along the x, y, and z directions; State update and action selection: At the t-th step, use the greedy policy to select an action: ; where, is the Q value of state s t and action a t , predicted by DQN; the parameter θ represents the network weights of DQN; Random exploration avoids getting stuck in local optima and optimally utilizes existing Q-values to guide path movement; In the current state s t Execute action a t to obtain a new state s t+1 and a reward r t ; Update the coordinate positions in the path; Calculate the target Q-value: ; where y t is the target Q-value; γ is the discount factor; θ − are the fixed parameters of the target network; represents the Q of the target network target for predicting the Q-value of the action a′ at the next moment state s t+1 ; Update the current Q-network parameters; When the reward value r t converges, or when the path length and the number of obstacle avoidance times meet the optimization requirements, output the optimized grasping path.
[0026] An automatic sorting system for recycled resources combined with intelligent recognition, including a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the above-mentioned intelligent garbage sorting method based on multi-modal recognition Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0027] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0028] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the function specified in one or more processes and / or blocks Figure 1 in one process or more processes and / or blocks Figure 1 in one block or more blocks.
[0029] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more processes and / or blocks Figure 1 in one process or more processes and / or blocks Figure 1 in one block or more blocks.
[0030] As mentioned above, it is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in any other form. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. An intelligent garbage sorting method based on multimodal recognition, characterized in that, It includes the following steps: S1: Use an embedded camera to collect RGB images of the conveyor belt, and use a depth sensor and a near-infrared spectroscopy sensor to capture the three-dimensional shape information and material information of the objects on the conveyor belt; S2: Preprocess the collected image data and sensor data to remove noise and align modalities; S3: Based on the preprocessed image data and sensor data, generate unified bounding boxes and classification results based on the HybridNet model; S4: Based on the bounding boxes and classification results output by the HybridNet model, use the discrete PID algorithm to calculate the motion instructions, and use the depth sensor to calculate the three-dimensional point cloud data of the object, judge the stacking layer number of the object's location, give priority to grasping the surface object, and generate an initial grasping path based on the position error and stacking priority; S5: Based on the initial grasping path, use DQN deep reinforcement learning to optimize the initial grasping path into an optimal grasping path.
2. The intelligent garbage sorting method based on multimodal recognition according to claim 1, characterized in that The noise removal is specifically as follows: The RGB image data is processed using bilateral filtering to remove noise: ; Where, I(x,y) is the pixel value of the RGB image at point (x,y); is a spatial domain Gaussian kernel function related to the distance which controls the neighborhood range; is the Gaussian kernel function of pixel difference, which controls the color similarity; is the normalization coefficient; is the pixel value of the image after bilateral filtering for noise removal at point (x,y); k is the radius of the filtering window; ; ; Among them, and are typical parameter values, which are used to control the filtering range and pixel similarity respectively; And perform local brightness normalization on the RGB image to make the brightness gradients consistent: ; Among them, is the mean value within the sliding window; is the standard deviation within the sliding window; The three-dimensional shape information data obtained by the depth sensor uses the RANSAC fitting plane model to remove outliers; The material information obtained by the near-infrared spectroscopy sensor uses reflectance calibration and Savitzky-Golay smoothing.
3. The intelligent garbage sorting method based on multimodal recognition according to claim 2, characterized in that, The alignment of modalities is to strictly align the RGB image with the depth point cloud and spectral information by combining the pose calibration algorithm and the multi-modal data alignment process, specifically as follows: Unify the coordinate systems of each sensor so that each sensor works in the same reference coordinate system. The embedded camera uses the calibration board method to estimate the internal parameter matrix K: ; where f x and f y are the focal lengths, and c x and c y are the principal points; The external parameter calibration takes the embedded camera as the reference coordinate system, and completes the rigid body transformation of the depth sensor and the camera with the rotation matrix R and the translation vector T: ; where P RGB and P depth represent the pixel coordinates of the RGB image and the depth point cloud, respectively; Map the depth data to the pixel grid of the RGB image: ; Among them, (X, Y, Z) are the three-dimensional point coordinates in the depth data; (x, y) are the 2D coordinates on the image plane; Due to the single-line scanning feature of the near-infrared spectroscopy sensor, the spectral points are aligned to the RGB image based on the calibrated geometric transformation: ; Among them, is the homography matrix projected from the near-infrared spectral sensor to the RGB image.
4. An intelligent garbage sorting method based on multi-modal recognition according to claim 1, characterized in that, The HybridNet model consists of an input layer, a multi-modal feature extraction layer, a feature fusion module, and a joint detection head, specifically as follows: The input layer inputs the preprocessed image data and sensor data, and the multi-modal feature extraction layer processes RGB images, depth maps, and near-infrared spectral data respectively, and converts them into representative features F RGB , F Depth , and F NIR ; The feature fusion module fuses multi-modal features based on the attention mechanism and feature splicing; ; Among them, α RGB , α Depth and α NIR are attention weights calculated for each modal feature: F fusion is the fused multi-modal feature; Dynamically allocate the weights of each modality through the channel attention mechanism α = {α RGB , α Depth , α NIR}: α = σ(Conv3×3(FRGB⊕FDepth⊕FNIR)) Among them, ⊕ is element-wise addition; σ is the Sigmoid activation function; α represents the dependence degree of each position on RGB / Depth / NIR; Conv3×3 represents a 3×3 convolutional network; The joint detection head uses a shared fully convolutional network to extract fused features for position detection and classification, and generates bounding box predictions and classification probabilities; Generate the following two outputs according to the HybridNet model: Bounding box B i ={x i ,y i ,w i ,h i}, where (x i ,y i ) are the coordinates of the i-th bounding box; w i ,h i are the width and height of the i-th bounding box respectively; the detection position and size of the object in the image space; Classification result P class (i), Probability distribution of the target class, select the class label C corresponding to the maximum probability i =argmax ; According to the target sorting requirements, filter out the objects that do not belong to the target category, and only retain the categories of interest as the control input. Suppose there are N targets in the image, then the output target set O is: 。 5. The intelligent garbage sorting method based on multimodal recognition according to claim 4, wherein The joint detection head includes a bounding box regression branch and a classification branch, specifically as follows: The bounding box regression branch presets p scales and q aspect ratios, with a total of p * q anchor boxes. The center position (x a , y a ) of the anchor box is calculated based on the feature map grid and is associated with the size of the input image: The regression branch outputs the offsets ΔB=(Δx,Δy,Δw,Δh) of each anchor box, where (Δx,Δy) are the coordinate offsets of the bounding box; Δw and Δh are the width and height offsets of the bounding box respectively, and the final bounding box is calculated: ; Among them, and are the width and height of the center of the anchor box, respectively; are the coordinates of the final bounding box; w and h are the width and height of the final bounding box, respectively. Use Smooth L1 Loss to measure the error between the bounding boxes predicted by the model and the ground truth labels: ; Among them, is the number of positive samples; ΔB i is the offset value of the i-th bounding box predicted by the model; is the offset value of the ground truth bounding box; smooth L1 is the Smooth L1 loss formula; The classification branch uses the Softmax function to predict the probability distribution of the target classes P class : ; where FC represents the fully connected layer; Use cross-entropy loss to measure the difference between the classification probabilities predicted by the model and the true classification labels: ; Among them, is the true classification label; is the predicted probability; Considering the dual objectives of regression and classification in object detection, the combined loss function L total is as follows: ; Among them, , and are weight coefficients; is the fusion weight regularization term.
6. The intelligent garbage sorting method based on multimodal recognition according to claim 4, characterized in that, Based on the bounding box and classification result output by the HybridNet model, a discrete PID algorithm is used to calculate the motion instruction, specifically as follows: Locate the image bounding box in the real-world coordinates through the depth sensor point cloud, and set the position of the target sorting area required by the sorting task. Use the depth sensor point cloud data to extract the three-dimensional physical position of the target, and use the centroid of the object as the actual coordinates of the object P actual,u : ; Among them, ( ) is the actual coordinate of object u, and z actual,u is the actual depth value of the object; Let the target position in the sorting area be P target =(x target ,y target ,z target ), where (x target ,y target ,z target ) are the actual coordinates of the target position, and calculate the error e between the target position and the actual position u : ; Generate the motion instruction of the robotic arm to the target object through the discrete PID control algorithm, and the discrete PID control dynamically adjusts the motion instruction according to the position error: ; Among them, is the motion increment of the current control instruction in the x direction; is the current time step error; is the sampling period; K p 、K i 、K d are the proportional, integral, and differential coefficients of the PID control respectively; is the cumulative error; Use the same formula to calculate in other directions (u y , u z ), and finally integrate them into the motion commands of the robotic arm: 。 7. The intelligent garbage sorting method based on multi-modal recognition according to claim 6, wherein, Calculate the three-dimensional point cloud data of the object using the depth sensor, judge the stacking layer number of the position where the object is located, and preferentially grasp the surface object, specifically as follows: Estimate the number of layers in the z-axis direction through the three-dimensional point cloud of the depth sensor: Analyze the highest point z of the target object within the i-th bounding box max,i : ; wherein, is the layer estimation of the i-th bounding box; is the lowest reference height; is the layer division step size; If the point cloud projection of the upper object overlaps with the lower object, it is considered that the target is occluded and not preferentially grasped. The occlusion rate is defined as: ; Among them, is the occlusion rate, is the area of the overlapping region, is the total area of the target bounding box; Assign priorities based on the layer where the object is located and the occlusion situation : 。 8. An intelligent garbage sorting method based on multi-modal recognition according to claim 7, characterized in that, Generate the initial grasping path based on the position error and stacking priority, specifically as follows: According to the target path error e u and the stacking priority, plan the grasping path, sort it according to the target priority, and select the target with the highest priority as the current grasping task P athsorted : P athsorted =Sort(O,Priority i ) Based on the actual position P of the target actual,u and the target point P in the sorting area target Generate the robot arm path, and plan the straight-line path P from the current position of the robot arm end to the target object ath (t): ; where P start is the starting point; If the closest distance dobs between the path and the obstacle point cloud is < , is the obstacle threshold, insert a detour point P into the path avoid : ; where n is the normal vector of the obstacle surface; is the modulus of the normal vector of the obstacle surface; Insert the midpoint P at a preset safe height above the object P actual,u ; actual,usafe ; Finally obtain the initial grasping path P ath which is P ath ={P start ,P avoid ,P actual,usafe, ,P actual,u ,P target}}。 9. The intelligent garbage sorting method based on multi-modal recognition according to claim 8, characterized in that Based on the initial grasping path, use DQN deep reinforcement learning to optimize the initial grasping path into the optimal grasping path, specifically as follows: Input the initial grasping path s0 and initialize the DQN deep reinforcement learning network: s0 = {P start , P target , P obs , StackInfo}; Among them, StackInfo is about the number of stacked objects and occlusion information; P start is the three-dimensional coordinate of the start and end of the robotic arm; P obs is the set of obstacles scanned currently; Starting state; Action set A={±Δa,±Δb,±Δc, grasping action}; where ±Δa,±Δb,±Δc represent fine-tuning actions in the x, y, and z directions in three-dimensional space; State update and action selection: At step t, use the greedy policy to select an action: ; Among them, is the state s t and the action a t The Q value of is predicted by DQN; the parameter θ represents the network weight of DQN; Random exploration to avoid falling into local optimum, and optimize the use of existing Q values to guide path movement; In the current state s t Execute action a t , to obtain a new state s t+1 and a reward r t ; Update the coordinate positions in the path; Calculate the target Q value: ; Among them, y t is the target Q value; γ is the discount factor; θ − are the fixed parameters of the target network; represents the target network Q target for predicting the Q value of the action a′ at the next moment state s t+1 ; Update the current Q network parameters; When the reward value r t converges, or when the path length and the number of obstacle avoidance times meet the optimization requirements, output the optimized grasping path.
10. An intelligent garbage sorting system based on multimodal recognition, characterized in that, It includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in an intelligent garbage sorting method based on multi-modal recognition according to any one of claims 1-9.
Citation Information
Patent Citations
Visual 3D taking and placing method and system based on cooperative robot
CN112476434A
Object recognition and pose estimation method for grabbing robot
CN117036470A
Control method for double mechanical arms
CN117798924A
Imitation bud thinning claw for kiwi fruits and control method thereof
CN118216332A
Automatic solid waste sorting method and device applying machine vision
CN119445493A
Cited By
Renewable resource sorting management system based on image analysis
CN120411657A
Automatic scheduling and efficient sorting system for crystalline silicon leftover materials based on AGV (Automatic Guided Vehicle)
CN120815739A
Material state self-adaptive material grabbing method and automatic material grabbing system
CN121269174A
Robot self-adaptive grabbing control method and system suitable for complex curved surface workpiece
CN121374571A
Robot adaptive grasping control method and system suitable for complex curved surface workpiece
CN121374571B