Garbage classification and recognition system for smart city

Through the garbage classification and recognition system that integrates image acquisition and repair and multimodal features, the accuracy and robustness of the smart city garbage classification system in complex environments is solved, and the automatic linkage of high-precision garbage identification and recycling equipment is realized.

CN120259783AActive Publication Date: 2025-07-04SHANGHAI TIANQI INTELLIGENT BUILDING CO LTD

Patent Information

Application Number
CN202510724593.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing smart city garbage classification system lacks classification accuracy and system robustness in complex environments, making it difficult to achieve high-precision, real-time classification and recycling equipment linkage control.

Method used

The image acquisition and repair module is used to obtain RGB-D images, combine the multimodal feature fusion module and dynamic classification decision module, and use parallel dual-branch convolutional neural network for garbage material and purpose recognition, and trigger the recycling equipment control signal through the classification result output module.

Benefits of technology

Realize high-precision garbage feature extraction under low light and occlusion conditions, improving the intelligence and automation capabilities of the garbage classification system, and ensuring seamless linkage between garbage identification results and recycling equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259783A_ABST
    Figure CN120259783A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of garbage treatment, in particular to a smart city-oriented garbage classification and recognition system, which comprises an image acquisition and restoration module, a multi-modal feature fusion module, a dynamic classification decision module and a classification result output module, wherein the image acquisition and restoration module is used for acquiring and restoring an image of a garbage throwing point and outputting a restored three-dimensional image; the multi-modal feature fusion module is used for fusing the extracted features into multi-modal fusion features through a space-time alignment algorithm; the dynamic classification decision module is used for classifying the parallel double-branch convolutional neural network integrated with the attention mechanism; and the classification result output module is used for triggering a control signal of the garbage recycling equipment. According to the intelligent urban garbage classification system, through collaborative design of multi-modal feature fusion and a parallel classification structure, accurate recognition of garbage materials and purposes in a complex environment is achieved, the recycling equipment can be automatically linked, and the overall intelligent level of the intelligent urban garbage classification system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of garbage disposal, and in particular to a garbage classification and recognition system for smart cities. Background Art

[0002] With the continuous growth of the total amount of urban domestic garbage, the traditional manual sorting method is difficult to meet the needs of modern urban management due to its slow speed, high error rate and high labor intensity; the automatic classification system based on a single two-dimensional image is prone to missed detection and misjudgment in case of insufficient light or occlusion, and lacks the comprehensive utilization of garbage weight and volume information, and cannot accurately reflect the physical characteristics of the objects put in.

[0003] Currently, most garbage classification solutions for smart cities still stay in a single dimension of visual recognition or weight detection, and it is difficult to achieve high-precision, real-time classification and linkage control of recycling equipment in the true sense. Therefore, a garbage classification and recognition system for smart cities is needed to overcome the problems of insufficient classification accuracy and system robustness in complex environments in the prior art. Summary of the Invention

[0004] Based on the above purpose, the present invention provides a garbage classification and recognition system for smart cities.

[0005] A garbage classification and recognition system for smart cities includes an image acquisition and repair module, a multi-modal feature fusion module, a dynamic classification decision module, and a classification result output module; wherein: The image acquisition and repair module: is used to acquire the RGB-D image of the garbage disposal point, perform light compensation on the low-light area and pixel-level repair on the occluded area, and output the repaired three-dimensional image; The multi-modal feature fusion module: is used to receive the repaired three-dimensional image, extract its surface texture features and three-dimensional geometric features, and at the same time access the real-time data of the weight sensor and the volume sensor, and fuse the extracted features into multi-modal fusion features through a spatio-temporal alignment algorithm; The dynamic classification decision module: based on the multi-modal fusion features, uses a parallel dual-branch convolutional neural network with an integrated attention mechanism for classification, the first branch identifies the garbage material type, the second branch identifies the garbage use type, and combines the output results of the two branches to generate a garbage classification label and its confidence score; The classification result output module: is used to encapsulate the garbage classification label and the confidence score into a structured data format, output it to the smart city management platform, and trigger the control signal of the corresponding garbage recycling equipment.

[0006] Optionally, the image acquisition and repair module includes an image acquisition unit, a light compensation unit, a pixel-level repair unit, and an image output unit; wherein: Image acquisition unit: It is used to synchronously acquire RGB-D data of the garbage disposal point by using a structured light depth camera and a wide-angle RGB camera. Among them, the structured light depth camera acquires depth maps at a rate of 30 frames per second, with a field of view angle of 70°, and the resolution of the RGB camera is 1920×1080 pixels; Light compensation unit: It is used to perform multi-scale Retinex processing on the RGB image output by the image acquisition unit. The multi-scale Retinex processing is to perform logarithmic domain transformation, color restoration, and gamma correction in sequence, and compensate the pixel brightness in the low-light area to the gray range of 50-200; Pixel-level repair unit: It is used to perform texture matching repair on the RGB pixel blocks corresponding to the missing areas caused by occlusion in the depth map based on the PatchMatch algorithm. By selecting the pixel blocks with the highest texture similarity to the surrounding of the missing area for filling, and using a bilateral filter to smooth the repair boundary; Image output unit: It is used to generate an RGB-D three-dimensional image by pixel-level mapping of the compensated RGB image and the repaired depth map, and output the three-dimensional image to the multi-modal feature fusion module.

[0007] Optionally, the pixel-level repair unit includes: Occlusion detection sub-unit: It is used to identify the pixel points in the input depth map with pixel values of zero or depth differences from adjacent pixels exceeding the set threshold as the boundaries of the occlusion area; Candidate initialization sub-unit: It is used to randomly select multiple candidate blocks with the same size as the pixel block to be repaired in the non-occluded area. Let the target repair block be , and the candidate block be ; Similarity calculation sub-unit, which is used to calculate the texture similarity for each candidate block and the target block ; Iterative update sub-unit: It is used to traverse the matching results in the adjacent areas of the target image to guide the iteration of the candidate blocks, and retain the matching position corresponding to the minimum texture similarity; Repair filling sub-unit: It is used to copy and cover the pixel information of the candidate block with the minimum similarity to the target occlusion area, and use the bilateral filtering algorithm to smooth the repair boundary.

[0008] Optionally, the multi-modal feature fusion module includes a surface texture feature extraction unit, a three-dimensional geometric feature extraction unit, a physical sensor data acquisition unit, and a spatio-temporal alignment and fusion unit; among them: Surface texture feature extraction unit: It is used to perform local binary pattern encoding on the input repaired RGB image, extract the detailed texture information of the garbage object surface, and form a surface texture feature vector; 3D Geometric Feature Extraction Unit: It is used to construct a spatial point cloud from the input repaired depth map, estimate the surface normal vector, calculate the curvature, and analyze the shape histogram of the point cloud, extract the 3D shape features of the garbage object, and form a 3D geometric feature vector; Physical Sensor Data Acquisition Unit: It is used to collect the data generated by the weight sensor and volume sensor during garbage disposal in real time. After eliminating outliers and interference information through data preprocessing, a standardized physical feature vector is formed; Spatio-Temporal Alignment and Fusion Unit: It is used to establish a spatio-temporal alignment mapping relationship based on the timestamp information of the received 3D image features and physical features, and use interpolation and data synchronization methods to ensure that the image features and physical features are strictly aligned under the same spatio-temporal reference, and splice the aligned feature vectors into a single multi-modal fusion feature vector.

[0009] Optionally, the 3D Geometric Feature Extraction Unit includes: Point Cloud Generation Sub-Unit: Based on the internal parameter matrix of the depth image, it converts the two-dimensional pixel coordinates and depth information into 3D spatial coordinate point clouds, and outputs the spatial coordinate point cloud data; Normal Vector Calculation Sub-Unit: It is used to estimate the normal vector of each spatial point in the point cloud data by using the neighborhood covariance analysis method, and obtain the normal vector of each point; Curvature Calculation Sub-Unit: It is used to calculate the curvature of each spatial point in the point cloud according to the normal vector obtained by the normal vector calculation sub-unit, and construct a local curvature feature description of the point cloud; Shape Histogram Construction Sub-Unit: It is used to construct a shape histogram based on the direction distribution and curvature value statistics of the normal vector of the point cloud data, and normalize the statistical results and output them as a 3D geometric feature vector.

[0010] Optionally, the Spatio-Temporal Alignment and Fusion Unit includes: Timestamp Matching Sub-Unit: It is used to extract the timestamp of the 3D image feature data and the timestamp of the physical sensor data , and calculate the time difference between the two ; Time Alignment Interpolation Sub-Unit: It is used to set a time alignment threshold . When the time difference , it is determined that the current image feature and physical feature are on the same spatio-temporal reference, otherwise linear interpolation is used to interpolate the physical feature data; Feature Splicing Sub-Unit: It is used to concatenate and splice the image feature vector after time alignment and the interpolated physical feature vector in a fixed order to generate a multi-modal fusion feature vector.

[0011] Optionally, the dynamic classification decision module includes a feature input unit, an attention enhancement unit, a material recognition branch unit, a use recognition branch unit, and a label fusion unit; specifically: Feature input unit: It is used to receive the fused feature vector output by the multi-modal feature fusion module, perform dimension verification and normalization processing on it, and use it as the unified input of the dual-branch convolutional neural network. Attention enhancement unit: It is used to impose an integrated mechanism of channel attention and spatial attention on the fused feature vector at the network input stage. Material recognition branch unit: It is constructed as the first branch convolutional neural network structure, used to receive the attention-enhanced feature vector, and through continuous convolutional layers, batch normalization layers and activation functions, extract discriminant features reflecting the material attributes of the garbage, and output the corresponding material classification result vector. Use recognition branch unit: It is constructed as the second branch convolutional neural network structure, used to process the same input feature vector in parallel, focus on extracting the functional use features of the garbage items, and output the use classification result vector. Label fusion unit: It is used to synthesize the output results of the material recognition branch and the use recognition branch, map the output indexes of the two branches to a unique garbage classification label based on a predefined classification mapping table, and calculate the joint confidence score using a probability combination rule.

[0012] Optionally, the attention enhancement unit includes: Channel attention sub-unit: It is used to perform global average pooling and global max pooling operations on the input fused feature vector, respectively obtain the corresponding channel-level statistical information, and learn the weight factors of each feature channel through a shared multi-layer perceptron network to obtain the attention weight vector in the channel dimension. Channel enhancement sub-unit: It is used to apply the channel attention weight vector output by the channel attention sub-unit to the original fused feature vector in a channel-wise multiplication manner to obtain the channel-enhanced feature. Spatial attention sub-unit: Based on the channel-enhanced feature output by the channel enhancement sub-unit, perform max pooling and average pooling processing along the channel dimension respectively to generate the spatial attention mapping feature, and generate a two-dimensional spatial attention mask through convolution operations. Spatial enhancement sub-unit: It is used to apply the spatial attention mask output by the spatial attention sub-unit to the channel-enhanced feature element-wise, enhance the local region features with significant discriminative power for the garbage material and use classification tasks in the spatial dimension, form a fused feature vector enhanced by both channel and spatial dimensions, and output it to the material recognition branch unit and the use recognition branch unit.

[0013] Optionally, the label fusion unit includes: Index mapping subunit: It is used to receive the material category index output by the material recognition branch unit and the usage category index output by the usage recognition branch unit , and based on a predefined two-dimensional classification mapping table , map the two indexes to a unique waste classification label , and the mapping relationship is: ; Probability combination subunit: It is used to receive the material classification probability value output by the material recognition branch unit and the usage classification probability value output by the usage recognition branch unit, and calculate the joint confidence score based on the multiplication fusion rule ; Label output subunit: It is used to output the mapped waste classification label and the joint confidence score to the classification result output module

[0014] Optionally, the classification result output module includes a structured encapsulation unit, a data communication unit, a device matching unit, and a control signal triggering unit; where: Structured encapsulation unit: It is used to receive the waste classification label and the joint confidence score, and use a predefined JSON data structure to encapsulate the label and the score in the form of key-value pairs respectively to form a standard structured data message Data communication unit: It is used to send the structured data message generated by the structured encapsulation unit to the data interface of the smart city management platform through the network interface using the MQTT protocol, and perform packet integrity verification and security encryption processing before sending Device matching unit: It is used to call the association table of the waste classification label and the waste recycling device predefined according to the received waste classification label to determine the device identification code of the corresponding waste recycling device and form a target device control index Control signal triggering unit: Generate a control instruction message that conforms to the communication protocol of the waste recycling device based on the target device control index, and send a control signal to the corresponding waste recycling device through the wireless communication module to perform actual recycling operations such as starting and stopping the device, opening or closing the waste classification door

[0015] Advantages of the present invention: In the present invention, by constructing an image acquisition and repair module, a multi-modal feature fusion module, a dynamic classification decision module, and a classification result output module, it is possible to realize the acquisition and repair of RGB-D images of waste disposal points in low-light and occluded scenarios, and combine physical sensor data such as weight and volume to fuse multi-source feature information under a unified spatio-temporal reference, effectively improving the integrity and accuracy of waste feature extraction

[0016] The present invention realizes the parallel recognition of garbage materials and uses through a dual-branch convolutional neural network integrating an attention mechanism, outputs a unique classification label and a score based on a label fusion and confidence calculation mechanism, and then automatically triggers corresponding recycling equipment control instructions in combination with an equipment mapping table, realizing the seamless linkage between the garbage recognition result and the urban recycling hardware system, and enhancing the intelligence and automation capabilities of the overall garbage classification system. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 Schematic diagram of the garbage classification recognition system according to an embodiment of the present invention; Figure 2 Schematic diagram of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The present invention will be described in detail below with reference to the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; and the drawings are only for more specifically describing the embodiments, and are not intended to specifically limit the present invention.

[0020] It should be noted that in the specification, references to "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment includes such specific features, structures, or characteristics. Additionally, when combining embodiments to describe specific features, structures, or characteristics, implementing such features, structures, or characteristics in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0021] Generally, terms can be understood, at least in part, from their use in the context. For example, at least in part depending on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather, at least in part depending on the context, can allow for the existence of other factors that may not be explicitly described.

[0022] Such as Figure 1 - Figure 2As shown in the figure, a garbage classification recognition system for a smart city includes an image acquisition and restoration module, a multi-modal feature fusion module, a dynamic classification decision-making module, and a classification result output module; among them: Image acquisition and restoration module: It is used to obtain the RGB-D image (i.e., Red-Green-Blue-Depth image) of the garbage disposal point, perform light compensation on the low-light area, perform pixel-level restoration on the occluded area, and output the restored three-dimensional image; Multi-modal feature fusion module: It is used to receive the restored three-dimensional image, extract its surface texture features and three-dimensional geometric features, and at the same time access the real-time data of the weight sensor and the volume sensor, and fuse the extracted features into multi-modal fusion features through the spatio-temporal alignment algorithm; Dynamic classification decision-making module: Based on the multi-modal fusion features, use a parallel dual-branch convolutional neural network with an integrated attention mechanism for classification. The first branch identifies the garbage material type, and the second branch identifies the garbage usage type, and generates a garbage classification label and its confidence score by combining the output results of the two branches; Classification result output module: It is used to encapsulate the garbage classification label and the confidence score into a structured data format, output it to the smart city management platform, and trigger the control signal of the corresponding garbage recycling equipment.

[0023] The image acquisition and restoration module includes an image acquisition unit, a light compensation unit, a pixel-level restoration unit, and an image output unit; among them: Image acquisition unit: It is used to synchronously acquire the RGB-D data of the garbage disposal point by using a structured light depth camera and a wide-angle RGB camera. Among them, the structured light depth camera acquires depth maps at a rate of 30 frames per second, the field of view angle is 70°, and the resolution of the RGB camera is 1920×1080 pixels; Light compensation unit: It is used to perform multi-scale Retinex processing on the RGB image output by the image acquisition unit. The multi-scale Retinex processing is to perform logarithmic domain transformation, color restoration, and gamma correction in sequence, and compensate the pixel brightness of the low-light area to the gray level range of 50-200; Pixel-level restoration unit: It is used to perform texture matching restoration on the RGB pixel blocks corresponding to the missing areas caused by occlusion in the depth map based on the PatchMatch algorithm, fill in the pixel blocks with the highest texture similarity to the surrounding areas of the missing area, and use a bilateral filter to smooth the restored boundary; Image output unit: It is used to generate an RGB-D three-dimensional image by mapping the compensated RGB image and the restored depth map at the pixel level, and output the three-dimensional image to the multi-modal feature fusion module; Based on the collaborative cooperation of the above units, the image acquisition and restoration module can generate a high-precision and complete RGB-D three-dimensional image under low-light and occlusion conditions, providing accurate and reliable input for the multi-modal feature fusion module, and improving the classification accuracy of the entire garbage classification recognition system.

[0024] The pixel-level restoration unit includes: Occlusion detection sub-unit: It is used to identify the pixel points with a pixel value of zero or a depth difference from adjacent pixels exceeding the set threshold in the input depth map as the boundary of the occlusion area; Candidate initialization sub-unit: It is used to randomly select multiple candidate blocks with the same size as the pixel block to be restored in the non-occluded area. Let the target restoration block be , and the candidate block be , where is the number of candidate blocks; Similarity calculation sub-unit, which is used to calculate the texture similarity between each candidate block and the target block . The formula is: , where represents the pixel L2 distance between the target block and the candidate block ; represents the relative coordinate within the block; represents the RGB three-channel pixel vector of the target block at the position ; represents the RGB pixel vector of the candidate block at the same position; represents the Euclidean square norm; Iterative update sub-unit: It is used to traverse the matching results in the adjacent areas of the target image to guide the iteration of the candidate blocks and retain the matching position corresponding to the minimum texture similarity; Restoration filling sub-unit: It is used to copy and cover the pixel information of the candidate block with the minimum similarity to the target occlusion area, and use the bilateral filtering algorithm to smooth the restoration boundary to ensure the continuity of color and texture between the restored area and the surrounding environment; The above pixel-level restoration unit can quickly identify and copy the most similar candidate image blocks in the occlusion area by introducing the PatchMatch texture matching mechanism with clear parameter definitions, realize high-fidelity image reconstruction, effectively improve the image restoration quality, and provide complete and stable input data for subsequent image feature extraction.

[0025] The multi-modal feature fusion module includes a surface texture feature extraction unit, a three-dimensional geometric feature extraction unit, a physical sensor data acquisition unit, and a spatio-temporal alignment and fusion unit; among which: The surface texture feature extraction unit: is used to perform local binary pattern encoding on the input repaired RGB image, extract the detailed texture information on the surface of the garbage object, and form a surface texture feature vector; The three-dimensional geometric feature extraction unit: is used to construct a spatial point cloud from the input repaired depth map, and perform surface normal vector estimation, curvature calculation, and shape histogram analysis on the point cloud, extract the three-dimensional shape features of the garbage object, and form a three-dimensional geometric feature vector; The physical sensor data acquisition unit: is used to collect the data generated by the weight sensor and the volume sensor during garbage disposal in real time, and form a standardized physical feature vector after eliminating outliers and interference information through data preprocessing; The spatio-temporal alignment and fusion unit: is used to establish a spatio-temporal alignment mapping relationship according to the received three-dimensional image features (including the surface texture feature vector and the three-dimensional geometric feature vector) and the timestamp information of the physical features, and use interpolation and data synchronization methods to ensure that the image features and the physical features are strictly aligned under the same spatio-temporal reference, and splice the aligned feature vectors into a single multi-modal fusion feature vector and output it to the dynamic classification decision module.

[0026] The three-dimensional geometric feature extraction unit includes: The point cloud generation sub-unit: converts the two-dimensional pixel coordinates and depth information into three-dimensional space coordinate point clouds based on the internal parameter matrix of the depth image, and outputs the spatial coordinate point cloud data; The normal vector calculation sub-unit: is used to estimate the normal vector of each spatial point in the point cloud data by using the neighborhood covariance analysis method, and obtain the normal vector of each point; The normal vector calculation formula is: , where, is the covariance matrix of the neighborhood points, with a dimension of ; is the three-dimensional coordinate vector of the th point in the neighborhood; is the average three-dimensional coordinate vector of all points in the neighborhood; is the number of points in the neighborhood; the eigenvector corresponding to the smallest eigenvalue obtained through eigenvalue decomposition is the normal vector of the current point; The curvature calculation sub-unit: is used to calculate the curvature of each spatial point in the point cloud according to the normal vector obtained by the normal vector calculation sub-unit, and construct a local curvature feature description of the point cloud based on this; The curvature calculation formula is: , where, is the curvature eigenvalue of the current spatial point; is the minimum eigenvalue of the covariance matrix; is the sum of all eigenvalues of the covariance matrix; Shape histogram construction subunit: used to construct a shape histogram based on the directional distribution of normal vectors and curvature values of point cloud data, and output the statistical results as a three-dimensional geometric feature vector after normalization.

[0027] The specific steps to construct the shape histogram are as follows: Step 1: Convert the normal vector of each point in the point cloud to the spherical coordinate system, and represent the normal vector as the polar angle and azimuth angle. The formulas are respectively: ; ; where, is the polar angle of the normal vector; is the azimuth angle of the normal vector; is the direction component of the normal vector in three-dimensional space; Step 2: Perform normalization processing on the curvature values calculated for each point in the point cloud to make the curvature eigenvalue fall within the interval [0, 1]. The formula is: , where, is the normalized curvature value; is the original curvature value of the current point; respectively represent the maximum and minimum values of the curvature values of all points in the point cloud; Step 3: Uniformly divide the feature space composed of the normal vector direction and the normalized curvature value into intervals. Among them, the polar angle range is divided into intervals; the azimuth angle range is divided into intervals; the curvature value range is divided into intervals; Step 4: For each point in the point cloud data, determine its interval index in the above feature space according to the polar angle, azimuth angle and normalized curvature value of its normal vector. Among them, the polar angle index is , the azimuth angle index is , and the curvature index is ; the interval index is defined as: ; ; , where, is the polar angle quantization step size; is the azimuth angle quantization step size; is the curvature quantization step size; is the floor operation; then count the corresponding positions of each interval index. The formula is; ; Step Five: Normalize the statistically obtained shape histogram to form the final three-dimensional geometric feature vector, with the expression: , where in the formula, are the components of each dimension of the final three-dimensional geometric feature vector; the denominator is the total count of all intervals of the shape histogram, which serves the purpose of normalization.

[0028] The spatio-temporal alignment and fusion unit includes: Timestamp matching sub-unit: used to extract the timestamps of the three-dimensional image feature data and the timestamps of the physical sensor data , and calculate the time difference between the two , with the formula: ; Time alignment interpolation sub-unit: used to set the time alignment threshold . When the time difference , it is determined that the current image feature and physical feature are on the same spatio-temporal basis; otherwise, linear interpolation is used to interpolate the physical feature data, with the formula: , where: represents the interpolated physical feature data; respectively represent the two physical feature data points before and after closest to the timestamp of the image feature; respectively represent the timestamps of the two physical feature data points before and after closest to the timestamp of the image feature; Feature concatenation sub-unit: used to concatenate the image feature vector after time alignment and the interpolated physical feature vector in a fixed order to generate a multi-modal fusion feature vector, with the calculation formula: , where is the finally output fusion feature vector; the symbol [;] represents the feature vector concatenation operation; through the precise matching and interpolation operations of the above spatio-temporal alignment and fusion unit, the precise time synchronization of the image features and physical sensor data is achieved, effectively reducing the spatio-temporal error in the multi-modal data fusion process and significantly improving the accuracy and robustness of waste classification decision-making.

[0029] The dynamic classification decision module includes a feature input unit, an attention enhancement unit, a material recognition branch unit, a usage recognition branch unit, and a label fusion unit; where: Feature input unit: used to receive the fusion feature vector output by the multi-modal feature fusion module, and perform dimension verification and normalization processing on it as the unified input of the dual-branch convolutional neural network; Attention Enhancement Unit: An integrated mechanism used to impose channel attention and spatial attention on the fused feature vector during the network input stage, to enhance the response ability to significant features of materials and uses, suppress irrelevant feature dimensions, and improve classification accuracy; Material Recognition Branch Unit: Constructed as the first branch convolutional neural network structure, used to receive the attention-enhanced feature vector, and through consecutive convolutional layers, batch normalization layers, and activation functions, extract discriminative features reflecting the attributes of waste materials (such as metal, plastic, glass, etc.), and output the corresponding material classification result vector; Use Recognition Branch Unit: Constructed as the second branch convolutional neural network structure, used to process the same input feature vector in parallel, focusing on extracting the functional use features of waste items (such as kitchen waste, hazardous, recyclable, etc.), and output the use classification result vector; Label Fusion Unit: Used to synthesize the output results of the material recognition branch and the use recognition branch, map the output indices of the two branches to a unique garbage classification label based on a predefined classification mapping table, and calculate the joint confidence score using a probability combination rule, which is used to measure the credibility of the final classification result.

[0030] The attention enhancement unit includes: Channel Attention Sub-Unit: Used to perform global average pooling and global max pooling operations on the input fused feature vector, respectively obtain the corresponding channel-level statistical information, and learn the weight factors of each feature channel through a shared multi-layer perceptron network to obtain the attention weight vector in the channel dimension; Channel Enhancement Sub-Unit: Used to apply the channel attention weight vector output by the channel attention sub-unit to the original fused feature vector in a channel-wise multiplication manner, enhance the response intensity of the channels of significant features for garbage classification, and obtain the channel-enhanced features; Spatial Attention Sub-Unit: Based on the channel-enhanced features output by the channel enhancement sub-unit, perform max pooling and average pooling operations along the channel dimension respectively, generate spatial attention mapping features, and generate a two-dimensional spatial attention mask through convolution operations; Spatial Enhancement Sub-Unit: Used to apply the spatial attention mask output by the spatial attention sub-unit to the channel-enhanced features element-wise, enhance the local region features with significant discriminability for the garbage material and use classification tasks in the spatial dimension, form a fused feature vector that is enhanced both in channel and space, and output it to the material recognition branch unit and the use recognition branch unit; Through the above explicit integration of channel and spatial attention mechanisms, the attention enhancement unit effectively improves the network's attention to important information in the fused features, suppresses the interference of irrelevant information, and significantly improves the classification performance of the material and use recognition branches and the classification accuracy of the overall system.

[0031] The label fusion unit includes: Index mapping subunit: used to receive the material category index output by the material recognition branch unit and the usage category index output by the usage recognition branch unit , and based on a predefined two-dimensional classification mapping table , map the two indexes to a unique garbage classification label , and the mapping relationship is: , where is a predefined two-dimensional label mapping table, used to represent the garbage classification label corresponding to any combination of material category and usage category; is the material category index output by the material recognition branch; is the usage category index output by the usage recognition branch; is the garbage classification label obtained after mapping; Table 1 Example of predefined two-dimensional label mapping In Table 1 above, the rows represent the category indexes output by the material recognition branch , and the columns represent the category indexes output by the usage recognition branch ; the value at each intersection is the final output unique garbage classification label; for some combinations that are not applicable logically (such as "glass + kitchen waste"), it can be set as an invalid item "L–" or the default output "other garbage".

[0032] The specific application example of Table 1 is as follows: Assume: the output of the material recognition branch is metal, that is, the index is ; the output of the usage recognition branch is recyclable, that is, the index is ; then: .

[0033] Probability combination subunit: used to receive the material classification probability value output by the material recognition branch unit and the usage classification probability value output by the usage recognition branch unit, and calculate the joint confidence score based on the multiplication fusion rule , and the formula is as follows: , where is the predicted probability value of the corresponding index of the material recognition branch; is the predicted probability value of the corresponding index of the usage recognition branch; is the joint confidence score, indicating the credibility of the final garbage classification label; Label output subunit: used to output the mapped garbage classification label and the joint confidence score Output to the classification result output module; through the classification label index mapping and probability fusion rules of the above-mentioned label fusion unit, the accurate comprehensive decision-making of material and usage information is effectively realized, the confidence and stability of the garbage classification result are improved, and the overall decision-making quality of the smart city garbage classification system is further guaranteed.

[0034] The classification result output module includes a structured encapsulation unit, a data communication unit, a device matching unit, and a control signal triggering unit; among them: The structured encapsulation unit is used to receive the garbage classification label and the combined confidence score, and use a predefined JSON data structure to encapsulate the label and the score in the form of key-value pairs respectively to form a standard structured data message. The data communication unit: is used to send the structured data message generated by the structured encapsulation unit to the data interface of the smart city management platform through the network interface according to the MQTT protocol, and perform packet integrity verification and security encryption processing before sending. The device matching unit: is used to call the predefined association table of garbage classification labels and garbage recycling devices according to the received garbage classification labels, determine the device identification code of the corresponding garbage recycling device, and form a target device control index. Table 2 Example of the association between garbage classification labels and garbage recycling device identification codes In the above Table 2, the garbage classification label number (L) is the final label output by the dynamic classification decision module, such as "L1", "L5", etc.; the corresponding Chinese label name of the label is convenient for management and display; the device number (Device_ID) represents the corresponding garbage recycling device number in the system and is used as the target device of the control instruction; the control instruction type represents the specific control signal generated by the system through this field (such as "open the recyclable door", "close other doors", etc.).

[0035] The control signal triggering unit: generates a control instruction message that conforms to the communication protocol of the garbage recycling device based on the target device control index, and sends a control signal to the corresponding garbage recycling device through the wireless communication module to perform actual recycling operations such as starting and stopping the device, opening or closing the garbage classification door; through the above units' structured encapsulation and secure transmission of the classification label and confidence score by the classification result output module, and the generation of automatic control instructions for accurately matching the garbage recycling device, the seamless cooperation between the system and the smart city platform and recycling device is realized, effectively improving the automation level and management efficiency of urban garbage classification and recycling.

[0036] The present invention covers any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0037] The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A garbage classification recognition system for a smart city, characterized in that, It includes an image acquisition and restoration module, a multi-modal feature fusion module, a dynamic classification decision module, and a classification result output module; among which: The image acquisition and restoration module: It is used to obtain the RGB-D image of the garbage disposal point, perform light compensation on the low-light area, perform pixel-level restoration on the occluded area, and output the restored three-dimensional image; The multi-modal feature fusion module: It is used to receive the restored three-dimensional image, extract its surface texture features and three-dimensional geometric features, and at the same time access the real-time data of the weight sensor and the volume sensor, and fuse the extracted features into multi-modal fusion features through the spatio-temporal alignment algorithm; The dynamic classification decision module: Based on the multi-modal fusion features, it uses a parallel dual-branch convolutional neural network with an integrated attention mechanism for classification. The first branch identifies the garbage material type, and the second branch identifies the garbage usage type, and combines the output results of the two branches to generate the garbage classification label and its confidence score; The classification result output module: It is used to encapsulate the garbage classification label and the confidence score into a structured data format, output it to the smart city management platform, and trigger the control signal of the corresponding garbage recycling equipment.

2. The garbage classification recognition system for a smart city according to claim 1, wherein The image acquisition and restoration module includes an image acquisition unit, a light compensation unit, a pixel-level restoration unit, and an image output unit; among which: The image acquisition unit: It is used to synchronously acquire the RGB-D data of the garbage disposal point by using a structured light depth camera and a wide-angle RGB camera. Among them, the structured light depth camera acquires depth maps at a rate of 30 frames per second, the field of view angle is 70°, and the resolution of the RGB camera is 1920×1080 pixels; The light compensation unit: It is used to perform multi-scale Retinex processing on the RGB image output by the image acquisition unit. The multi-scale Retinex processing is to perform logarithmic domain transformation, color restoration, and gamma correction in sequence, and compensate the pixel brightness of the low-light area to the gray range of 50-200; The pixel-level restoration unit: It is used to perform texture matching restoration on the RGB pixel blocks corresponding to the missing areas caused by occlusion in the depth map based on the PatchMatch algorithm. By selecting the pixel block with the highest texture similarity to the surrounding area of the missing area for filling, and using a bilateral filter to smooth the restored boundary; The image output unit: It is used to generate an RGB-D three-dimensional image by pixel-level mapping of the compensated RGB image and the restored depth map, and output the three-dimensional image to the multi-modal feature fusion module.

3. The waste classification recognition system for a smart city according to claim 2, characterized in that, The pixel-level restoration unit includes: The occlusion detection sub-unit: It is used to identify the pixel points with a pixel value of zero or a depth difference from adjacent pixels exceeding the set threshold in the input depth map as the occlusion area boundary; Candidate initialization subunit: used to randomly select multiple candidate blocks with the same size as the pixel block to be repaired in the non-occluded area. Let the target repair block be , and the candidate block be ; A similarity calculation sub-unit, which is used to calculate the texture similarity for each candidate block and the target block respectively The iterative update sub-unit: It is used to traverse the matching results of adjacent areas in the target image to guide the candidate block iteration, and retain the matching position corresponding to the minimum texture similarity; The restoration filling sub-unit: It is used to copy and cover the pixel information of the candidate block with the smallest similarity to the target occlusion area, and use the bilateral filtering algorithm to smooth the restored boundary.

4. The garbage classification recognition system for a smart city according to claim 1, characterized in that, The multimodal feature fusion module includes a surface texture feature extraction unit, a three-dimensional geometric feature extraction unit, a physical sensor data acquisition unit, and a spatio-temporal alignment and fusion unit; where: The surface texture feature extraction unit: is used to perform local binary pattern encoding on the input repaired RGB image, extract the detailed texture information on the surface of the garbage object, and form a surface texture feature vector; The three-dimensional geometric feature extraction unit: is used to construct a spatial point cloud from the input repaired depth map, and perform surface normal vector estimation, curvature calculation, and shape histogram analysis on the point cloud to extract the three-dimensional shape features of the garbage object and form a three-dimensional geometric feature vector; The physical sensor data acquisition unit: is used to collect the data generated by the weight sensor and the volume sensor during garbage disposal in real time, and form a standardized physical feature vector after eliminating outliers and interference information through data preprocessing; The spatio-temporal alignment and fusion unit: is used to establish a spatio-temporal alignment mapping relationship according to the timestamp information of the received three-dimensional image features and physical features, and use interpolation and data synchronization methods to ensure that the image features and physical features are strictly aligned under the same spatio-temporal reference, and splice the aligned feature vectors into a single multimodal fusion feature vector.

5. The garbage classification recognition system for a smart city according to claim 4, characterized in that, The three-dimensional geometric feature extraction unit includes: The point cloud generation sub-unit: based on the internal parameter matrix of the depth image, converts the two-dimensional pixel coordinates and depth information into three-dimensional space coordinate point clouds, and outputs the spatial coordinate point cloud data; The normal vector calculation sub-unit: is used to estimate the normal vector of each spatial point in the point cloud data by using the neighborhood covariance analysis method to obtain the normal vector of each point; The curvature calculation sub-unit: is used to calculate the curvature of each spatial point in the point cloud according to the normal vector obtained by the normal vector calculation sub-unit, and construct a local curvature feature description of the point cloud based on this; The shape histogram construction sub-unit: is used to construct a shape histogram by statistically analyzing the direction distribution and curvature values of the point cloud data according to the normal vector, and output the statistical result as a three-dimensional geometric feature vector after normalization.

6. The garbage classification recognition system for a smart city according to claim 4, wherein, The spatio-temporal alignment and fusion unit includes: Timestamp matching subunit: used to extract the timestamps of the three-dimensional image feature data and the timestamps of the physical sensor data , and calculate the time difference between the two ; Time alignment interpolation subunit: used to set the time alignment threshold , when the time difference is met, it is determined that the current image feature and the physical feature are on the same spatio-temporal basis; otherwise, linear interpolation is used to interpolate the physical feature data. The feature splicing sub-unit: is used to concatenate the image feature vector after time alignment and the interpolated physical feature vector in a fixed order to generate a multimodal fusion feature vector.

7. The waste classification recognition system for a smart city according to claim 1, characterized in that, The dynamic classification decision module includes a feature input unit, an attention enhancement unit, a material recognition branch unit, a use recognition branch unit, and a label fusion unit; where: The feature input unit: is used to receive the fusion feature vector output by the multimodal feature fusion module, and perform dimension verification and normalization processing on it as the unified input of the two-branch convolutional neural network; The attention enhancement unit: is used to apply an integrated mechanism of channel attention and spatial attention to the fusion feature vector at the network input stage; The material recognition branch unit: is constructed as the first branch convolutional neural network structure, is used to receive the attention-enhanced feature vector, extracts discriminative features reflecting the material attributes of the garbage through continuous convolutional layers, batch normalization layers, and activation functions, and outputs the corresponding material classification result vector; Usage recognition branch unit: Constructed as a second-branch convolutional neural network structure for parallel processing of the same input feature vector, focusing on extracting the functional usage features of waste items and outputting a usage classification result vector; Label fusion unit: Used to synthesize the output results of the material recognition branch and the usage recognition branch, map the output indices of the two branches to a unique waste classification label based on a predefined classification mapping table, and calculate the combined confidence score using a probability combination rule.

8. The garbage classification recognition system for a smart city according to claim 7, characterized in that, The attention enhancement unit includes: Channel attention sub-unit: Used to perform global average pooling and global max pooling operations on the input fused feature vector to obtain corresponding channel-level statistical information, and learn the weight factors of each feature channel through a shared multi-layer perceptron network to obtain an attention weight vector in the channel dimension; Channel enhancement sub-unit: Used to apply the channel attention weight vector output by the channel attention sub-unit to the original fused feature vector in a channel-wise multiplication manner to obtain channel-enhanced features; Spatial attention sub-unit: Based on the channel-enhanced features output by the channel enhancement sub-unit, perform max pooling and average pooling operations along the channel dimension respectively to generate spatial attention mapping features, and generate a two-dimensional spatial attention mask through convolution operations; Spatial enhancement sub-unit: Used to apply the spatial attention mask output by the spatial attention sub-unit to the channel-enhanced features element-wise, enhance the local region features with significant discriminative power for the waste material and usage classification tasks in the spatial dimension, form a fused feature vector enhanced by both channel and spatial dimensions, and output it to the material recognition branch unit and the usage recognition branch unit.

9. The garbage classification recognition system for a smart city according to claim 7, wherein The label fusion unit includes: Index mapping sub-unit: used to receive the material category index output by the material recognition branch unit and the usage category index output by the usage recognition branch unit , based on a predefined two-dimensional classification mapping table , map the two indexes to a unique waste classification label , the mapping relationship is: ; Probability combination subunit: used to receive the material classification probability value output by the material recognition branch unit and the use classification probability value output by the use recognition branch unit, and calculate the joint confidence score based on the multiplication fusion rule ; Label output subunit: used to output the mapped waste classification labels and the combined confidence score to the classification result output module.

10. A garbage classification recognition system for a smart city according to claim 1, characterized in that, The classification result output module includes a structured encapsulation unit, a data communication unit, a device matching unit, and a control signal triggering unit; among them: Structured encapsulation unit: Used to receive the waste classification label and the combined confidence score, and use a predefined JSON data structure to encapsulate the label and the score in the form of key-value pairs respectively to form a standard structured data message; Data communication unit: Used to send the structured data message generated by the structured encapsulation unit to the data interface of the smart city management platform through a network interface using the MQTT protocol, and perform packet integrity verification and security encryption processing before sending; Device matching unit: Used to call the predefined association table of waste classification labels and waste recycling devices according to the received waste classification label to determine the device identification code of the corresponding waste recycling device, and form a target device control index; Control signal triggering unit: Generate a control instruction message that conforms to the waste recycling device communication protocol based on the target device control index, and send a control signal to the corresponding waste recycling device through a wireless communication module to perform actual recycling operations such as starting and stopping the device, opening or closing the waste classification door.

Citation Information

Patent Citations

  • Community garbage disposal method, community server and computer readable storage medium

    CN110969267A

  • Garbage classification method based on classification and detection joint judgment

    CN113657143A

  • Point cloud registration method based on low-dimensional point cloud local feature descriptors

    CN114972459A

  • Control method of intelligent garbage sorting system

    CN116921247A

  • Method and equipment of separate collection for trash automatic identification

    KR1020090000245A

Cited By

  • Background garbage classification resource optimization decision-making system based on big data

    CN120851271A

  • Kitchen garbage intelligent classification method and system based on AI image recognition

    CN122066997A

  • Intelligent classification method and system for kitchen waste based on AI image recognition

    CN122066997B