Garbage classification supervision system based on multi-sensor fusion

By adopting multi-sensor fusion and deep learning algorithms in the garbage classification system, the problems of insufficient space coverage, lack of closed-loop behavior and low recognition accuracy in the existing technology are solved, and efficient and real-time garbage classification and overflow detection are achieved, achieving the goal of smart garbage classification.

CN120047902AInactive Publication Date: 2025-05-27SUZHOU LECHUANG ENVIRONMENTAL PROTECTION TECH CO LTD

Patent Information

Application Number
CN202510518803.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing garbage classification detection technology has problems such as insufficient space coverage, lack of closed-loop behavior, real-time bottlenecks, and low recognition accuracy, especially in complex scenarios and abnormal situations.

Method used

The garbage classification supervision system based on multi-sensor fusion is adopted, combining multi-view cameras, ultrasonic sensors, edge computing layer and execution layer, and garbage classification detection and overflow detection are carried out using the improved visual Transformer algorithm and the front and rear timing comparison algorithm to achieve automated control and early warning prompts.

Benefits of technology

It improves the identification accuracy and real-time nature of garbage classification, enhances the detection ability of garbage overflow, ensures the system's environmental robustness and low-latency inference optimization, and achieves the smart garbage classification goal of "multi-person disposal and zero manual real-time supervision".

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047902A_ABST
    Figure CN120047902A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of garbage classification detection, and discloses a garbage classification supervision system based on multi-sensor fusion, which comprises a multi-view camera, an ultrasonic sensor, an edge calculation layer and an execution layer, the edge calculation layer comprises a data acquisition module, a human body feature detection module, a garbage classification detection module and a garbage overflow detection module; the garbage classification detection module performs garbage classification detection and identification on video image data in target data based on an improved visual Transform algorithm and a front-back time sequence comparison algorithm, so that model updating is automatically started under low-confidence alert, disastrous forgetting is effectively prevented, and meanwhile, the real-time performance and high precision of the system are kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of garbage classification detection, and more specifically to a garbage classification supervision system based on multi-sensor fusion. Background Art

[0002] With the acceleration of the urbanization process, the amount of domestic garbage generated has increased sharply. Accurately and quickly classifying garbage can improve the efficiency of domestic garbage treatment. The traditional manual sorting mode has low efficiency, high cost and is easily affected by subjective factors. Therefore, combining artificial intelligence technology with garbage classification and using intelligent algorithms to detect and classify domestic garbage effectively improves the intelligent level and accuracy of garbage classification.

[0003] However, there are problems such as insufficient space coverage, lack of behavior closed-loop, and real-time bottlenecks in the existing detection and monitoring of garbage classification. A single camera has monitoring blind spots and cannot detect the behavior of illegally discarding garbage outside the garbage pavilion. There is no linkage decision-making in links such as door opening and closing control, throwing detection, and mixed throwing recognition. It is impossible to complete the analysis of multiple video streams within 200 ms, resulting in response delays. Moreover, most use traditional CNN for garbage recognition and classification. Traditional CNN is difficult to capture long-range dependence relationships, and the recognition accuracy drops significantly in complex scenarios. It has insufficient processing ability in abnormal situations such as occlusion and deformation, and insufficient environmental robustness. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a garbage classification supervision system based on multi-sensor fusion to solve the problems existing in the above-mentioned background art.

[0005] The present invention provides the following technical solutions: A garbage classification supervision system based on multi-sensor fusion includes a multi-view camera, an ultrasonic sensor, an edge computing layer, and an execution layer; the edge computing layer includes a data acquisition module, a human feature detection module, a garbage bag and garbage throwing area detection module, and a garbage overflow detection module; The multi-view camera is used to collect video image data of the throwing process in the area outside the garbage pavilion and the area inside the garbage pavilion; the ultrasonic sensor is used to obtain the filling height data of the garbage in the trash can; The data acquisition module is used to collect the target data collected by the multi-view camera and the ultrasonic sensor; The human feature detection module is used to detect human features in the video image data in the target data. If human information is captured, a driving instruction is sent to the execution layer, and at the same time, the data is transmitted to the garbage classification detection module for detection; The garbage classification detection module performs garbage classification detection and recognition on the video image data in the target data based on an improved Vision Transformer algorithm and a front-back time series comparison algorithm, and sends a driving instruction or a warning instruction to the execution layer based on the recognition result; The overflow detection module detects the garbage overflow situation based on the filling height data of the garbage in the target data, and sends a warning instruction to the execution layer based on the detection result; The execution layer includes an automatic control module and an interaction and reminder module. The automatic control module includes a push rod motor and an electric push rod installed in the trash can. One end of the electric push rod is connected to the output end of the push rod motor, and the other end is connected to the lid of the trash can. The push rod motor is used to receive the driving instruction issued by the edge computing layer, and the interaction and reminder module is used to receive the warning instruction issued by the edge computing layer.

[0006] As a specific implementation manner, the improved Vision Transformer algorithm includes: 1) Obtain the video image data collected by the multi-view camera as the input image; 2) Build a Transformer improved model. Based on the Transformer model, introduce the Vision Transformer structure. Add a CNN before the Transformer backbone network to extract local texture features, and then fuse them with the output of Self-Attention. Use residual connection to avoid gradient disappearance caused by deep networks; 3) Perform adaptive Patch segmentation on the input image through the Transformer improved model to obtain a Patch image and extract features, build a serialized Token, and enter the multi-scale attention fusion network for interactive learning to obtain the final fusion features; 4) Input the fusion features into the automatic online learning module for learning and perform low-latency inference optimization.

[0007] As a specific implementation manner, the adaptive Patch segmentation adopts an improved dynamic block strategy, and the formula is expressed as: ; where P adaptive represents the adaptive segmentation mechanism function; X represents the input image tensor, which is a three-channel RGB image matrix; N represents the number of parallel blocks; s i represents the size of the i-th type of block; w i represents the weight parameter of the i-th type of block method; Split(X, s i ) represents the block function, which splits the input image X according to the size s i for segmentation.

[0008] As a specific implementation, the multi-scale attention fusion network proposes an innovative attention calculation formula to achieve the in-depth fusion of spatial and scale information. The innovative attention calculation formula is as follows: ; where M pos represents the relative position encoding matrix, M scale represents the scale adaptation matrix, d k represents the dimension of the attention mechanism, Q represents the query matrix, K represents the key matrix, V represents the value matrix, which are the basic components of the attention mechanism. K T represents the transpose matrix of K, and QK T is used to calculate the attention score, representing the correlation between the query and the key; Softmax() represents converting the attention score into a probability distribution.

[0009] As a specific implementation, the formula for the multi-scale attention fusion network to perform multi-scale feature fusion is as follows: ; where represents the scale weight, which is optimized through backpropagation; MSA c is the multi-head self-attention layer for multi-scale fusion of features, FFN is the feed-forward neural network, c represents the scale level of the features, and C represents the total number of scale levels of the features.

[0010] As a specific implementation, the low-latency inference optimization utilizes edge computing and GPU acceleration, combined with a lightweight network structure design. The automatic online learning module sets the recognition confidence threshold YU. When the confidence of an input sample is lower than YU, the sample is automatically marked and online incremental learning is performed. The update formula of its objective function is expressed as: ; where L total represents the comprehensive loss, including the classification loss and the knowledge distillation loss, L auto represents measuring the annotation error generated by the automatic learning module, which is composed of the deviation between the automatic annotation and the manual confirmation. λ auto represents the automatic learning weight coefficient, η represents the adaptive learning rate, represents the model parameters at time t + 1, that is, the updated model parameters, represents the model parameters at time t, that is, the model parameters before the update, represents the update of the model parameters for transfer learning.

[0011] As a specific implementation, the formula for the comprehensive loss is expressed as: ; where, L CE represents the classification loss, L KD represents the knowledge distillation loss, L EWC represents the elastic weight consolidation loss, λ 1 and λ 2 are the corresponding weight coefficients respectively. λ 1 is used to balance the proportion of the classification loss and the knowledge distillation loss and is a fixed value; λ 2 is used to control the amplitude of parameter change and is a fixed value.

[0012] As a specific implementation manner, after the garbage classification detection module classifies and detects the garbage, if the identified garbage category is the same as the garbage category corresponding to the placed trash can, the recognition result is correct; if the identified garbage category is different from the garbage category corresponding to the placed trash can, the recognition result is incorrect; if the recognition result is incorrect, a warning instruction is sent to the execution layer.

[0013] As a specific implementation manner, the specific way for the overflow detection module to detect the garbage overflow situation is as follows: Obtain the distance from the surface of the garbage in the trash can: ; where, d is the distance from the surface of the garbage in the trash can, v represents the propagation speed of ultrasonic waves in the air, which is related to the ambient temperature, v = 331.4 + 0.6T; where, T is the ambient temperature and t is the echo time; Obtain the overflow rate of the trash can: ; where, MY is the overflow rate and H is the total internal height of the trash can, which is calibrated according to different trash can models; Judge the overflow state of the trash can: When the overflow rate exceeds the preset threshold UI or d ≤ d safe , the detection result is overflow, otherwise the detection result is non - overflow. d safe represents the safe distance from the garbage in the trash can to the surface of the trash can.

[0014] As a specific implementation manner, the specific process of the front - back time - series comparison algorithm is as follows: Step S1, obtain the key frames After the previous resident finishes placing and leaves, the system determines the video image at this moment or time period as the "key frame of placement end" and caches or archives it; When the next resident approaches and before the trash can lid is opened, the system marks the video image at this moment or time period as the "key frame before placement", and during the placement process, continuously records the segment for real - time detection; Step S2, differential detection Find the newly added or changed area ΔM in the picture through frame difference method or differential calculation based on depth features: ΔM = D(F before , F after ), where F before is the previous "key frame at the end of delivery"; F after is the next "key frame before delivery", and D is the differential operation; If an obvious outline of a garbage bag appears in ΔM, it is determined that the garbage bag belongs to this delivery; if the object shown in ΔM already existed in the camera view after the previous delivery, it is judged that the garbage bag is "old garbage" or a leftover; Step S3, Abnormality judgment and temporal fusion After the system detects newly added garbage, it combines the improved Vision Transformer model to re-identify the target area in ΔM to confirm the type, bag-breaking situation, and whether there is mixed delivery; If ΔM cannot match the position of the garbage bag on-site or there are structural differences, it means that there may be "illegal discarding" or "multiple people littering" behaviors, and further temporal fusion analysis is required. Temporal fusion refers to performing action detection on the short video sequences before and after the delivery of the garbage bag, and finally making an accurate attribution judgment by identifying the position changes of the personnel's delivery actions and the changes in the garbage landing points. Specifically: If the action trajectory highly overlaps with the position of the newly added garbage bag in ΔM, it is confirmed that the garbage is done by the current deliverer; If there is an obvious action break or new garbage appears after the deliverer leaves, it is determined as "subsequent personnel" or "re-delivery", and the system records the new time point.

[0015] The technical effects and advantages of the present invention: By designing the edge computing layer, the present invention is conducive to introducing an automatic online learning module and optimizing parameters for low-latency inference, classifying, detecting, and identifying garbage based on the improved Vision Transformer model and the front and rear temporal comparison algorithms, and detecting the garbage overflow situation based on sensor data. It adopts a more intelligent image chunking method. The original path was fixed, while the present invention adopts dynamic chunking adjustment; it has stronger feature extraction ability, and improves the recognition accuracy through the fusion of multi-scale features; a confidence threshold is set. When the forward inference of the model is lower than this confidence level, the data will enter the pending review list, be manually labeled and then enter the training data set, enhancing the accuracy of the data. The improved Vision Transformer model mentioned in the present invention introduces an automatic online learning module and optimizes parameters for low-latency inference, utilizes edge computing and GPU acceleration, and combines a lightweight network structure design to ensure that the overall average inference latency is controlled within 850 ms and the maximum detection latency does not exceed 1 s. Description of the Drawings

[0016] Figure 1 This is the structural diagram of the garbage classification supervision system based on multi-sensor fusion of the present invention; Figure 2 This is the structural diagram of the edge computing layer of the present invention; Figure 3 This is the schematic diagram of the garbage kiosk of the present invention Figure 1 ; Figure 4 This is the schematic diagram of the garbage kiosk of the present invention Figure 2 ; Figure 5 This is the schematic diagram of the automatic opening mechanism of the present invention; Figure 6 This is the schematic diagram of the connection of the functional modules of the garbage classification supervision system based on multi-sensor fusion of the present invention; Reference numerals: 1, overflow sensor; 2, panoramic camera; 3, garbage kiosk; 4, bucket cover pull rod; 5, hook; 6, electric push rod; 7, channel; 8, front camera; 9, top camera; 10, display screen. Detailed implementation manners

[0017] Next, the technical solutions in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. In addition, the forms of the respective structures described in the following embodiments are merely examples, and a garbage classification supervision system based on multi-sensor fusion involved in the present invention is not limited to the respective structures described in the following embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0018] As Figure 1 , 6 shown, the present invention provides a garbage classification supervision system based on multi-sensor fusion, including multi-view cameras, ultrasonic sensors, an edge computing layer, and an execution layer.

[0019] The multi-view camera is used to collect video data of the garbage disposal process in the area outside and inside the garbage pavilion; the ultrasonic sensor is used to obtain the overflow data of the garbage in the trash can; the garbage classification detection module performs garbage classification detection and recognition on the video image data in the target data based on the improved vision Transformer algorithm and the front and back time series comparison algorithm, and sends a driving instruction or a warning instruction to the execution layer based on the recognition result; the execution layer is used to drive warning prompts and perform driving control on the trash cans in the garbage pavilion; the execution layer includes an automatic control module and an interaction and reminder module. The automatic control module includes a push rod motor and an electric push rod installed in the trash can. One end of the electric push rod is connected to the output end of the push rod motor, and the other end is connected to the lid of the trash can. The push rod motor is used to receive the driving instruction sent by the edge computing layer, and the interaction and reminder module is used to receive the warning instruction sent by the edge computing layer.

[0020] As Figure 2 shown, the edge computing layer includes a data acquisition module, a human feature detection module, a garbage classification detection module, and a garbage overflow detection module; The data acquisition module is used to obtain the target data collected by the multi-view camera and the ultrasonic sensor. The target data includes the video data of the target area and the ultrasonic sensor data. The target area is the outside and inside of the garbage pavilion, and the range of the outside area can be set by those skilled in the art according to the actual situation; The human feature detection module is used to detect human features in the video data of the target data. If human information is captured, a driving instruction is sent to the execution layer, and at the same time, the data is transmitted to the garbage classification detection module for detection; The garbage classification detection module introduces an automatic online learning module and low-latency inference optimization parameters, performs garbage classification detection and recognition on the video image data in the target data based on the improved vision Transformer model and the front and back time series comparison algorithm, and sends a warning instruction to the execution layer based on the recognition result; The overflow detection module detects the garbage overflow situation based on the sensor data and sends a warning instruction to the execution layer based on the detection result.

[0021] See Figure 3 、 4As shown in FIGS. 5, the automatic control module includes a trash can disposed in the trash pavilion 3. A multi-view camera and an overflow sensor 1 are provided on the top cover of the trash pavilion 3. The multi-view camera includes a panoramic camera 2, a front camera 8 and a top camera 9 installed on the trash pavilion 3. The panoramic camera 2 and the front camera 8 are mainly used to collect video data of the throwing process in the area outside the trash pavilion 3, and the top camera 9 is mainly used to collect video data of the throwing process in the area inside the trash pavilion 3. The overflow sensors 1 correspond one by one to the trash cans in terms of quantity and position. A lid pull rod 4 is installed at the upper end of the trash can. Symmetric arc-shaped channels 7 are opened on the two side plates of the trash pavilion 3. The two ends of the lid pull rod 4 are respectively inserted into the channels 7. The lid pull rod 4 is connected to the handle on the trash can lid through a hook 5. An electric push rod 6 is provided on the side of the trash pavilion 3. One end of the electric push rod 6 is connected to the output end of the push rod motor, and the other end is connected to the lid pull rod 4. When the push rod motor receives a driving command, it drives the electric push rod 6 to drive the lid pull rod 4 to slide along the channel 7 to open the trash can lid.

[0022] In this embodiment, it should be specifically noted that the human feature detection module reads continuous frame data from the video data of the target data, adjusts the resolution and frame rate to meet the actual processing requirements, and preprocesses the continuous frame images, including grayscale conversion, noise reduction and enhancement, and dynamic range compression, etc. The grayscale conversion converts the color frame into a grayscale image to reduce the computational complexity. The noise reduction and enhancement use Gaussian filtering to eliminate noise and improve the contrast through histogram equalization. The dynamic range compression is aimed at the uneven illumination scene and uses adaptive histogram equalization to optimize the image quality.

[0023] A deep learning model is used for human feature detection. The deep learning model includes, but is not limited to, single-stage detectors, multi-task cascaded networks, and two-stage detectors. The single-stage detectors such as YOLO and SSD directly generate human bounding boxes and confidence levels through convolutional neural networks to achieve end-to-end efficient detection. The multi-task cascaded networks such as MTCNNN can quickly generate candidate human regions, finely screen the candidate boxes and perform bounding box regression, and output the final human position and key point coordinates. The two-stage detectors such as Faster R-CNN first generate candidate regions and then fine-tune the bounding boxes through classification and regression networks, which are suitable for high-precision scenarios.

[0024] In this embodiment, it should be specifically noted that the specific way for the human feature detection module to send a driving instruction to the execution layer is as follows: If human information is recognized, a driving instruction is sent to the execution layer to make the automatic lid-opening mechanism execute the instruction.

[0025] In this embodiment, the improved Vision Transformer algorithm includes: 1) Obtain the video image data collected by the multi-view camera as the input image; 2) Construct an improved Transformer model. Based on the Transformer model, introduce the Vision Transformer structure. Add a CNN before the Transformer backbone network to extract local texture features, and then fuse them with the output of Self-Attention. Use residual connections to avoid gradient disappearance caused by deep networks; 3) Perform adaptive Patch segmentation on the input image through the improved Transformer model to obtain Patch images and extract features, construct serialized Tokens, and enter the multi-scale attention fusion network for interactive learning to obtain the final fused features; 4) Input the fused features into the automatic online learning module for learning and perform low-latency inference optimization.

[0026] Here, an adaptive Patch segmentation mechanism, a multi-scale attention fusion network, an automatic online learning module, and low-latency inference optimization are adopted. The purpose is to adopt a more intelligent image chunking method. The original path was fixed, and in this embodiment, dynamic chunking adjustment is used; it has stronger feature extraction capabilities. By fusing features at multiple scales, the recognition accuracy is improved; a confidence threshold is set. When the forward inference of the model is lower than this confidence level, the data will enter the pending review list and be manually annotated and then enter the training dataset, enhancing the accuracy of the data.

[0027] The adaptive Patch segmentation mechanism adopts an improved dynamic chunking strategy to improve the feature expression accuracy. The formula is expressed as: ; where P adaptive represents the adaptive segmentation mechanism function, X represents the input image tensor, usually a three-channel RGB image matrix; N represents the number of parallel chunks. In this embodiment, N = 3 is selected, indicating that 3 different chunk sizes are used simultaneously; s i represents the size of the i-th chunk, and the range of s i is between 8×8 and 32×32; w i represents the weight parameter of the i-th chunking method, which is a self-adaptive learning parameter and reflects the importance of different chunk sizes; Split(X, s i ) represents the chunking function, that is, the input image X is segmented according to the size s i ; DynamicEmbed( ) represents the function name of the entire dynamic embedding process, which weights and fuses the chunking results at different scales.

[0028] Chunking effect comparison: Table 1

[0029] As can be seen from Table 1, by adopting the adaptive Patch segmentation mechanism and the improved dynamic chunking strategy, the retention rate of features can be increased, and the feature expression accuracy can be improved.

[0030] The multi-scale attention fusion network proposes an innovative attention calculation formula to achieve the in-depth fusion of spatial and scale information. The innovative attention calculation formula is: ; where M pos represents the relative position encoding matrix, which is used to encode the positional relationship between features and helps the model understand spatial position information. M scale represents the scale adaptation matrix, which is used to process features of different scales and enhance the model's adaptability to targets of different sizes. d k represents the dimension of the attention mechanism. In this embodiment, d k = 64. Dividing by here is to scale the attention scores and prevent gradient disappearance; Q represents the query matrix, K represents the key matrix, and V represents the value matrix, which are the basic components of the attention mechanism. K T represents the transpose matrix of K. QK T is used to calculate the attention scores, indicating the correlation between the query and the key; Softmax() is used to convert the attention scores into a probability distribution, ensuring that the sum of all attention weights is 1, highlighting important features and suppressing unimportant features; this formula mainly obtains the features of the image from multiple scales, and finally performs feature weighting according to the results. During the process of extracting the features of garbage, features from local to global can be extracted; The formula for the multi-scale attention fusion network to perform multi-scale feature fusion is: ; where represents the scale weight, which is optimized through backpropagation; MSA c is the multi-head self-attention layer, which is used for multi-scale fusion of features. FFN is the feed-forward neural network. c represents the scale level of the features, and C represents the total number of scale levels of the features; The automatic online learning module and low-latency inference optimization parameters can ensure the long-term high performance and adaptability of the model. The low-latency inference optimization utilizes edge computing and GPU acceleration, combined with a lightweight network structure design, to ensure that the overall average inference latency is controlled within 850 ms and the maximum detection latency does not exceed 1 second. The automatic online learning sets the recognition confidence threshold YU. In this embodiment, YU = 0.85 is selected. When the confidence of an input sample is lower than YU, the sample is automatically marked and online incremental learning is performed. Through expert manual confirmation and decision-making, the update formula of its objective function is expressed as: ; where L total represents the comprehensive loss, including classification loss and knowledge distillation loss, which is used to ensure basic recognition accuracy. L auto represents the annotation error generated by the automatic learning module, which is composed of the deviation between automatic annotation and manual confirmation. λ auto represents the automatic learning weight system, which can be automatically adjusted to balance the importance of basic tasks and automatic learning. η represents the adaptive learning rate, which is used to control the step size of parameter update. represents the model parameters at time t+1, that is, the updated model parameters. represents the model parameters at time t, that is, the model parameters before update. represents the model parameter update of transfer learning; this mechanism realizes that the model update can be automatically started under low-confidence alert, can effectively prevent catastrophic forgetting, and at the same time maintain the real-time performance and high accuracy of the system. ; where L CE represents the classification loss, L KD represents the knowledge distillation loss, L EWC represents the elastic weight consolidation loss, and λ 1 and λ 2 are the corresponding weight coefficients respectively. λ 1 is used to balance the proportion of classification loss and knowledge distillation loss, and control the degree to which the new model learns the knowledge of the old model. It is usually a fixed value, such as 0.5 or 0.7. The larger the value of λ 1 , the more attention is paid to inheriting the knowledge of the old model. The smaller the value of λ 1 , the more inclined to learn new knowledge by itself. λ 2 is used to control the amplitude of parameter change to prevent sudden large changes in model parameters. It is usually a fixed value, such as 0.4 or 0.6, and its purpose is to maintain model stability. The larger the value of λ 2 , the more attention is paid to maintaining the original knowledge. If the data changes little, a larger value of λ 2 can be used. If the scenario changes frequently, a smaller value of λ 2 can be used.

[0031] This mechanism realizes the automatic activation of model updates under low-confidence alerts, effectively preventing catastrophic forgetting while maintaining the real-time performance and high precision of the system.

[0032] Specifically, the core parameter configuration in this example is as follows: Table 2

[0033] Performance Evaluation Based on the comparative experiment data of a 100,000-sample test set (considering both detection and online learning latency), the detection results are as follows: Table 3

[0034] As can be seen from Table 3, while ensuring a detection latency < 1 s, the accuracy, recall rate, and F1 score of the present invention have all increased. Through the automatic online learning module, long-term adaptive updates of the system and rapid response to abnormal situations are achieved.

[0035] In this embodiment, after the garbage classification detection module classifies and detects the garbage, if the identified garbage category is the same as the garbage category corresponding to the trash can into which it is placed, the recognition result is correct; if the identified garbage category is different from the garbage category corresponding to the trash can into which it is placed, the recognition result is incorrect; if the recognition result is incorrect, a warning instruction is sent to the execution layer.

[0036] In this embodiment, it should be specifically noted that the specific method for the overflow detection module to detect the garbage overflow situation is as follows: Obtain the distance from the surface of the garbage in the trash can: ; where d is the distance from the surface of the garbage in the trash can, v represents the propagation speed of ultrasonic waves in the air, which is related to the ambient temperature, v = 331.4 + 0.6T; where T is the ambient temperature and t is the echo time; Obtain the overflow rate of the trash can: ; where MY is the overflow rate and H is the total internal height of the trash can, which is calibrated according to different trash can models; Determine the overflow state of the trash can: When the overflow rate exceeds the preset threshold UI or d ≤ d safe , the detection result is overflow, otherwise the detection result is non-overflow; the preset threshold UI can be specifically set by those skilled in the art according to the specific situation of the trash can, and this embodiment does not specifically limit this specific value. In this embodiment, YU ≥ 90% is selected; d safe is the preset minimum safety distance, that is, the safety threshold of the distance from the surface of the garbage in the trash can. When d ≤ dsafe When it indicates that the distance from the surface of the garbage is too small, it will cause overflow.

[0037] In this embodiment, the propagation speed of the ultrasonic wave in the air is affected by factors such as temperature and humidity, and a correction formula is used for correction. The correction formula is expressed as: ; where v corrected represents the corrected propagation speed of the ultrasonic wave in the air, α represents the humidity compensation coefficient, taking 0.1% to 0.3% per degree Celsius, and can be automatically adjusted through the built-in algorithm of the sensor; ΔT represents the deviation between the current temperature and the calibrated temperature; β represents the humidity compensation coefficient; ΔE represents the deviation between the current humidity and the calibrated humidity.

[0038] In this embodiment, the execution layer receives the warning instruction and gives a warning prompt. The warning prompt includes voice broadcast and a full - overflow warning on the display screen. The voice broadcast gives a voice reminder for wrong garbage - throwing behaviors and automatically ends after the user leaves; the full - overflow warning on the display screen automatically generates a work order and notifies the cleaning unit to clean the overflowing garbage, and automatically settles the account after cleaning. Here, the display screen 10 is set on the top plate of the garbage kiosk 3.

[0039] The drive control of the trash can in the garbage kiosk includes controlling the automatic opening mechanism to execute the instruction and automatically open the trash can lid.

[0040] In addition, a pre - and post - placement comparison algorithm is introduced in the garbage classification detection module. The pre - and post - placement comparison algorithm is used to identify before and after garbage placement. In the garbage classification placement scenario, simply detecting garbage through the current frame sometimes makes it difficult to promptly distinguish whether this is newly placed garbage in this time or old garbage that has not been recognized previously. Therefore, in this embodiment, the pre - and post - placement comparison algorithm is added to improve the accuracy and traceability of recognition; The process of the pre - and post - placement comparison algorithm is specifically as follows: Step S1, obtaining key frames: After a resident finishes placing and leaves, the system will mark the video image at this moment or time period as the "key frame at the end of placement" and archive the "key frame at the end of placement"; when the next resident approaches and before the trash can lid is opened, the video image at this moment or time period is marked as the "key frame before placement", and during the placement process, continuous recording of segments is performed for real - time detection; Step S2, differential detection: After placement ends and before the next placement starts, on the premise that the environmental backgrounds of the two key frames are the same, the newly added areas and changed areas in the key frames are obtained through differential detection. The differential detection is expressed by the formula: ΔM = D(F before , F after ); where ΔM represents the newly added areas and changed areas, F beforeis the previous "key frame at the end of delivery"; F after is the next "key frame before delivery", and D is the differential operation; If an obvious outline of a garbage bag appears in ΔM, it can be determined that the garbage belongs to this delivery behavior. If the object shown in ΔM already existed in the previous post-delivery camera view, then it can be judged as "old garbage" or a leftover, reducing the interference caused by misjudgment to the current delivery user; the garbage bag refers to the entire delivered garbage. Since garbage is always packaged in garbage bags or trash bins, the entire garbage formed after packaging is called a garbage bag; Step S3, anomaly judgment and timing integration: After detecting new garbage, the area in ΔM is recognized again using an improved Vision Transformer model to confirm the type, bag-breaking situation, and whether there is mixed delivery, etc.; if ΔM cannot match the position of the on-site garbage bag or there are structural differences, it means that there may be illegal dumping or multiple people littering, and further timing fusion analysis is required; The timing fusion analysis means performing action detection on the short video sequences before and after this delivery, and finally making an accurate attribution judgment by identifying the position changes of the personnel's delivery actions and the changes in the garbage landing points, etc.; if the action trajectory highly overlaps with the position of the new garbage in ΔM, it is confirmed that the garbage is formed by the current deliverer; if there is an obvious action break or garbage appears after the deliverer leaves, it may be a "subsequent person" or "re-delivery" situation, and at this time, a new time point needs to be recorded.

[0041] The present invention combines a multi-view camera and an ultrasonic sensor, and integrates an edge computing unit inside the device to run multiple deep learning algorithms (including an improved Vision Transformer and a front-back timing comparison algorithm, etc.) to control the entire process of garbage delivery. Through automatic control and behavior analysis, the intelligent garbage classification goal of "multiple people delivering, zero manual real-time supervision" is achieved.

[0042] Finally: The above description is only a preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0043] The above description is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or replacements, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A garbage classification supervision system based on multi-sensor fusion, characterized by: It includes a multi-view camera, an ultrasonic sensor, an edge computing layer and an execution layer; the edge computing layer includes a data acquisition module, a human feature detection module, a garbage bag and garbage delivery area detection module, and a garbage overflow detection module; The multi-view camera is used to collect video image data of the process of placing garbage outside the garbage kiosk and inside the garbage kiosk; The ultrasonic sensor is used to obtain the filling height data of the garbage in the garbage bin; The data acquisition module is used to collect target data collected by the multi-view camera and the ultrasonic sensor; The human feature detection module is used to perform human feature detection on the video image data in the target data. If human information is captured, a driving instruction is sent to the execution layer, and the data is transmitted to the garbage classification detection module for detection; The garbage classification detection module performs garbage classification detection and recognition on the video image data in the target data based on the improved visual Transformer algorithm and the front-to-back time sequence comparison algorithm, and sends a driving instruction or a warning instruction to the execution layer based on the recognition result; The overflow detection module detects the garbage overflow situation based on the filling height data of the garbage in the target data, and sends a warning instruction to the execution layer based on the detection result; The execution layer includes an automation control module and an interaction and reminder module. The automation control module includes a push rod motor and an electric push rod installed in the trash can. One end of the electric push rod is connected to the output end of the push rod motor, and the other end is connected to the lid of the trash can. The push rod motor is used to receive the driving instructions issued by the edge computing layer, and the interaction and reminder module is used to receive the early warning instructions issued by the edge computing layer.

2. According to claim 1, a garbage classification supervision system based on multi-sensor fusion is characterized by: The improved visual Transformer algorithm includes: 1) Obtain video image data collected by a multi-view camera as input image; 2) Build an improved Transformer model. Based on the Transformer model, introduce the Vision Transformer structure, add CNN before the Transformer backbone network to extract local texture features, and then fuse them with the Self-Attention output. Use residual connections to avoid gradient disappearance caused by deep networks. 3) Adaptively segment the input image into patches through the Transformer improved model to obtain patch images and extract features, build serialized tokens, enter the multi-scale attention fusion network for interactive learning, and obtain the final fusion features; 4) Input the fused features into the automatic online learning module for learning and low-latency inference optimization.

3. According to claim 2, a garbage classification supervision system based on multi-sensor fusion is characterized by: The adaptive Patch segmentation adopts an improved dynamic block strategy, which is expressed as follows: ; Among them, P adaptive represents the adaptive segmentation mechanism function; X represents the input image tensor, which is a three-channel RGB image matrix; N represents the number of parallel blocks; s i represents the size of the i-th block; w i Represents the weight parameter of the i-th block splitting method; Split(X,s i ) represents the block function, which divides the input image X into blocks of size s i to split.

4. According to claim 2, a garbage classification supervision system based on multi-sensor fusion is characterized by: The multi-scale attention fusion network proposes an innovative attention calculation formula to achieve deep fusion of space and scale information. The innovative attention calculation formula is: ; Among them, M pos Represents the relative position encoding matrix, M scale represents the scale adaptation matrix, d k represents the dimension of the attention mechanism, Q represents the query matrix, K represents the key matrix, and V represents the value matrix, which is the basic component of the attention mechanism. T represents the transposed matrix of K, QK T Used to calculate the attention score, which represents the correlation between the query and the key; Softmax() represents the conversion of the attention score into a probability distribution.

5. According to claim 4, a garbage classification supervision system based on multi-sensor fusion is characterized by: The multi-scale attention fusion network performs multi-scale feature fusion using the formula: ; in, Represents the scale weight, optimized by back-propagation; MSA c It is a multi-head self-attention layer used for multi-scale fusion of features. FFN is a feed-forward neural network. c represents the scale level of the feature, and C represents the total number of scale levels of the feature.

6. The garbage classification supervision system based on multi-sensor fusion according to claim 2 is characterized by: The low-latency inference optimization utilizes edge computing and GPU acceleration, combined with lightweight network structure design. The automatic online learning module sets the recognition confidence threshold YU. When there is an input sample with a confidence lower than YU, the sample is automatically marked and online incremental learning is performed. The objective function update formula is expressed as: ; Among them, L total represents the comprehensive loss, including classification loss and knowledge distillation loss, L auto represents the labeling error produced by the automatic learning module, which is composed of the deviation between automatic labeling and manual confirmation, λ auto represents the automatic learning weight coefficient, η represents the adaptive learning rate, represents the model parameters at time t+1, that is, the updated model parameters, represents the model parameters at time t, that is, the model parameters before updating, Represents the model parameter update for transfer learning.

7. The garbage classification supervision system based on multi-sensor fusion according to claim 6 is characterized by: The formula for the comprehensive loss is expressed as: Among them, L CE represents the classification loss, L KD represents the knowledge distillation loss, L EWC Represents the elastic weight consolidation loss, λ1 and λ2 are the corresponding weight coefficients respectively. λ1 is used to balance the proportion of classification loss and knowledge distillation loss, which is a fixed value; λ2 is used to control the amplitude of parameter change, which is a fixed value.

8. The garbage classification supervision system based on multi-sensor fusion according to claim 1 is characterized by: After the garbage classification detection module performs classification detection on the garbage, if the identified garbage category is the same as the garbage category corresponding to the garbage bin where the garbage is placed, the recognition result is correct; If the identified garbage category is different from the garbage category corresponding to the garbage bin, the identification result is wrong; If the recognition result is an error, a warning instruction is sent to the execution layer.

9. The garbage classification supervision system based on multi-sensor fusion according to claim 1 is characterized by: The specific method in which the overflow detection module detects the garbage overflow situation is as follows: Get the distance of the garbage surface in the trash can: ; Where d is the distance from the garbage surface in the garbage bin, v represents the propagation speed of ultrasound in the air, which is related to the ambient temperature, v = 331.4 + 0.6T; where T is the ambient temperature, and t is the echo time; Get the trash bin overflow rate: ; Among them, MY is the overflow rate, H is the total height inside the trash can, and it is calibrated according to different trash can models; Determine the overflowing state of the trash can: When the overflow rate exceeds the preset threshold UI or d≤d safe When , the test result is overflow, otherwise the test result is non-overflow, d safe Indicates the safe distance between the garbage in the trash can and the surface of the trash can.

10. The garbage classification supervision system based on multi-sensor fusion according to claim 1 is characterized by: The specific process of the front-to-back timing comparison algorithm is as follows: Step S1, get key frame When a resident has finished delivering and left, the system will identify the video image at that moment or time as the "delivery end key frame" and cache or archive it; When the next resident approaches, before the lid of the trash can is opened, the system marks the video image of that moment or period as a "pre-delivery key frame". During the delivery process, the clips are continuously recorded for real-time detection; Step S2, differential detection By using the frame difference method or differential calculation based on depth features, find the newly added or changed area ΔM in the picture: ΔM=D(F before ,F after ) Among them, F before The previous "delivery end key frame"; F after is the next "pre-delivery key frame", and D is the differential operation; If there is an obvious outline of a garbage bag in ΔM, it is determined that the garbage bag belongs to the current delivery; if the object shown in ΔM already exists in the lens after the previous delivery, it is determined that the garbage bag is "old garbage" or leftovers; Step S3: abnormality judgment and time series fusion After detecting the newly added garbage, the system uses the improved visual Transformer model to re-identify the target area in ΔM to confirm the type, bag breakage, and whether it is mixed. If ΔM cannot match the location of the garbage bag on site or there are structural differences, it means that there may be "illegal dumping" or "multiple littering" behavior, and further time series fusion analysis is required. Time series fusion refers to motion detection of the short video sequence before and after the garbage bag is dropped. By identifying the changes in the position of the person's dropping action and the changes in the garbage landing point, an accurate attribution judgment is finally made. Specifically: If the action trajectory highly overlaps with the location of the newly added garbage bag in ΔM, it is confirmed that the garbage was thrown by the current person; If there is an obvious action gap or new garbage appears after the person who puts the garbage away leaves, it will be judged as "subsequent personnel" or "re-putting", and the system will record the new time point.

Citation Information

Patent Citations

  • Garbage classification method based on classification and detection joint judgment

    CN113657143A

  • Intelligent garbage throwing and transferring method, system and device

    CN118968375A

Cited By

  • Background garbage classification resource optimization decision-making system based on big data

    CN120851271A

  • Garbage throwing behavior dynamic identification system based on Internet of Things data

    CN120909125A

  • Artificial intelligence edge computing terminal for fault diagnosis of industrial equipment

    CN122241361A