Intelligent garbage classification box management and control method
By standardizing and eliminating outliers in urban planning, demographic, and meteorological monitoring data, and combining it with a random forest model to predict trash bin saturation, we solved the data compatibility and quality issues in the intelligent trash sorting bin management and control system, and achieved efficient trash bin layout optimization and transportation scheduling.
Patent Information
- Application Number
- CN202510691457.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing intelligent garbage sorting bin management and control system has compatibility issues in data integration and processing, which makes it difficult to ensure data quality and affects the optimization of garbage bin layout and transportation efficiency.
By acquiring data from urban planning, demographic and meteorological monitoring databases, standardization, integrity testing, outlier removal and logical verification are performed to generate a structured feature set. The random forest model is used to predict the saturation distribution of garbage bins and dynamically generate collection path parameters.
It realizes the intelligent management of garbage bin distribution and transportation, improves garbage collection efficiency, reduces environmental pollution risks, and provides data support for urban management decisions.
Smart Images

Figure CN120611307A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a method for managing and controlling intelligent garbage sorting bins. Background Art
[0002] Problem Background: Intelligent waste sorting bin management and control, a crucial area for urban environmental sanitation and resource recycling, plays a crucial role in improving urban management efficiency and residents' quality of life. With the acceleration of urbanization, the need for waste sorting and disposal is becoming increasingly urgent, making intelligent management and control a key approach to addressing the waste problem. However, many current solutions lack significant shortcomings in data integration and application. They often focus solely on managing a single device and lack the comprehensive utilization and in-depth analysis of multi-source data. This leads to frequent problems such as irrational waste bin layout and inefficient collection and transportation. Against this backdrop, this field faces numerous challenges, particularly in data import and processing. First, due to the complex and inconsistent data sources and formats for various aspects such as urban planning, population distribution, and weather changes, systems often face compatibility issues when integrating this external data. This data heterogeneity directly leads to another, more complex issue: data quality assurance. Errors and omissions are prone to occur during the import process, which in turn affects the accuracy of subsequent analysis and decision-making. Furthermore, due to the lack of effective data cleansing and validation mechanisms, the system is unable to detect and correct these issues in a timely manner, ultimately hindering the effective implementation of functions such as waste bin layout optimization and collection and transportation scheduling. Therefore, how to build an efficient data import mechanism to ensure the accuracy and consistency of multi-source data, and at the same time improve data quality through cleaning and verification methods, has become a key issue in the field of smart garbage sorting box management and control. Summary of the Invention
[0003] The present invention provides a method for controlling an intelligent garbage classification box, which mainly includes:
[0004] Obtaining original data sets from the urban planning database, demographic database, and meteorological monitoring database, performing field mapping processing on the original data sets using a preset standardized template, and generating a standardized basic data set with a unified format;
[0005] Perform integrity check on the standardized basic dataset. If it is detected that the proportion of missing values in a field exceeds a preset threshold, call the historical dataset of the corresponding area for interpolation and filling to generate a complete dataset.
[0006] The standard deviation threshold method is used to perform anomaly detection on the completed data set. If the data point deviates from the mean by more than the preset standard deviation range, it is marked as abnormal data and removed to generate a cleaned data set;
[0007] Performing a logical check on the cleaned dataset and the trash bin layout benchmark information, and triggering a rule base correction mechanism if a data inconsistency is detected to generate an integrated dataset;
[0008] performing a redundant field deletion operation on the integrated data set, and fusing the meteorological time series data with the garbage bin distribution data to generate a structured feature set;
[0009] Performing time series analysis on the structured feature set, marking a high-impact period if a meteorological feature value exceeds a preset threshold, and generating a labeled feature set;
[0010] The labeled feature set is input into the random forest model, and the saturation distribution of garbage bins in each area is predicted based on historical removal records;
[0011] The scheduling optimization model is called according to the saturation distribution. If a high-saturation area is detected, the cleaning path parameters are dynamically generated and the final scheduling plan is output.
[0012] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:
[0013] The present invention discloses a method for intelligent garbage sorting bin management and control. By integrating multi-source data such as urban planning, demographics, and meteorological monitoring, the data is standardized, integrity tested, outliers eliminated, and logically checked to generate a structured feature set. Combined with meteorological time series data, the feature set is subjected to time series analysis to mark high-impact periods. The processed data is input into a random forest model to predict the saturation distribution of garbage bins in each region. Based on the saturation distribution, the collection and transportation path parameters are dynamically generated, and the optimal scheduling plan is output. The present invention realizes the intelligent management of the distribution and collection of urban garbage bins, improves the efficiency of garbage collection, reduces the risk of environmental pollution, and provides data support for urban management decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a flow chart of a smart garbage sorting bin management and control method of the present invention.
[0015] Figure 2 This is a schematic diagram of an intelligent garbage sorting box management and control method of the present invention. DETAILED DESCRIPTION
[0016] The following will describe the technical solutions in the embodiments of the present invention in detail with reference to the accompanying drawings. The described embodiments are only a part of the embodiments of the present invention.
[0017] like Figure 1-2 In this embodiment, a smart garbage classification box management and control method may specifically include:
[0018] Step S101 : obtaining original data sets from the urban planning database, the demographic database, and the meteorological monitoring database, performing field mapping processing on the original data sets using a preset standardized template, and generating a standardized basic data set with a unified format.
[0019] The process involves obtaining raw datasets from urban planning, demographic, and meteorological monitoring databases and performing preliminary cleansing using automated tools to generate a cleaned initial dataset. Based on the cleaned initial dataset, field mapping is performed on the dataset using a pre-defined standardized template to determine whether the fields conform to the unified format requirements. If field content is missing or the format is inconsistent, the data is supplemented or converted using pre-defined rules to determine a mapped dataset with a unified format. Within the mapped dataset, correlation information is obtained between the fields. The distribution characteristics of urban planning, demographic, and meteorological monitoring data are compared to determine whether data redundancy or anomalies exist. If anomalies are detected, the data is filtered using a pre-defined threshold to generate a filtered standard dataset. Based on the filtered standard dataset, a logistic regression model is used to analyze the potential relationships between the data. Feature extraction is performed to identify the cross-influences between urban planning, demographics, and meteorological monitoring, thereby determining a key feature dataset. Using this key feature dataset, a priority order for data integration is determined. Cross-domain data fusion is performed on the highest-priority fields to determine whether the fused data meets completeness requirements. If not, historical data is used to supplement the data to generate a comprehensive, integrated dataset. Based on the integrated comprehensive data set, cluster analysis is used to group the data and classify it based on the distribution patterns of urban planning, demographics, and meteorological monitoring to obtain a classified grouped data set. From this classified grouped data set, the core feature information of each data set is obtained. Multi-dimensional verification of this feature information is performed to determine whether the data meets the business logic requirements. If not, adjustments are made according to preset rules to determine the final business data set.
[0020] For example, in scenarios involving the integration of urban planning, demographics, and meteorological monitoring data, acquiring and cleaning raw datasets is fundamental. Consider obtaining road construction data for a city in 2023 from an urban planning database. This data contains fields such as road length and construction time, but some records are missing the construction time. During cleaning, automated tools can be used to mark missing values as "unknown" and remove completely irrelevant fields, such as notes, to ensure initial usability of the data. This step effectively reduces noise in subsequent processing.
[0021] In one possible implementation, the field mapping operation may perform standardization on the cleansed data.
[0022] For example, the "age range" field in demographic data may exist in two formats: "20-30" or "20 to 30." Predefined rules can be used to standardize the format to "20-30," and missing values can be filled in with "unknown." This unified formatting ensures consistency in subsequent analysis and improves data processing efficiency.
[0023] Specifically, data redundancy and outliers can be detected by comparing distribution characteristics. For example, suppose that 99% of the daily average temperature records at a certain station in meteorological monitoring data are between -5 and 35 degrees Celsius, but a few records are as high as 100 degrees Celsius, a clear anomaly. By filtering with a preset threshold, values outside the reasonable range are removed, ensuring the reliability of the standard dataset. This prevents outliers from interfering with subsequent modeling.
[0024] For example, logistic regression model analysis can be used to explore the relationship between urban planning and demographics. If an analysis finds a positive correlation between road density and population density in a certain area, "road density" can be extracted as a key feature. This helps identify the core influencing factors in urban development and provides a basis for planning decisions.
[0025] In one possible implementation, data fusion priority may be based on feature importance ranking.
[0026] For example, population density and rainfall data from meteorological data are prioritized for integration to analyze their impact on traffic congestion. If data for some areas is missing after integration, historical data from the past five years can be used to supplement the missing data to ensure completeness. This improves the comprehensiveness and practicality of the data.
[0027] Specifically, cluster analysis can group data to reveal distribution patterns.
[0028] For example, by dividing cities into high-density residential areas and low-density industrial areas, and analyzing the differences in meteorological impacts between the two areas, this can help formulate targeted urban management strategies and improve resource allocation efficiency.
[0029] For example, multi-dimensional validation can be used to perform business logic analysis on grouped data. For example, if the population data for a high-density area doesn't match the planned data, for example, if the population far exceeds the planned capacity, preset rules can be used to adjust the data to a reasonable value or flag it as requiring special attention. This ensures that the final business dataset meets actual needs and supports scientific decision-making.
[0030] Step S102 : performing integrity check on the standardized basic dataset. If it is detected that the ratio of missing values in a field exceeds a preset threshold, the historical dataset of the corresponding area is called to perform interpolation filling to generate a complete dataset.
[0031] By performing integrity checks on the basic dataset, the distribution of missing values in the fields is obtained, and it is determined whether the proportion of missing values exceeds the preset threshold. If the proportion of missing values in the field exceeds the preset threshold, the historical dataset is obtained from the corresponding region to determine the data range that can be used for interpolation filling. Based on the distribution characteristics of the historical dataset, the missing fields are filled using linear interpolation to obtain a preliminary completed dataset. For the preliminary completed dataset, data standardization is performed to obtain consistency verification results to determine whether the data distribution meets the preset standards. If the consistency verification results do not meet the standards, additional relevant records are extracted from the historical dataset through the regional matching mechanism to determine candidate values for supplementary filling. Based on the distribution characteristics of the candidate values, a secondary verification method is used to obtain the final completed dataset to determine whether the data integrity meets business requirements. By performing multi-dimensional testing on the final completed dataset, data quality assessment results are obtained to determine the availability of the dataset.
[0032] For example, in the field of data processing related to urban planning, demographics, and meteorological monitoring, the integrity of basic data sets can be checked by analyzing the distribution of missing values in fields to judge data quality. Suppose that in a planning data set for an urban area, certain key fields such as population density or building coverage are missing, and the missing ratio reaches 20%, while the preset threshold is 10%. At this point, the system will trigger the process of obtaining supplementary data from historical data sets. Historical data may include records for the same area over the past five years. By analyzing its distribution trend, such as the pattern of increasing population density year by year in a specific area, the scope of interpolation filling can be preliminarily determined.
[0033] Specifically, linear interpolation is a common method for filling missing fields. For example, if the population density data for a region in 2020 and 2022 were 2,000 and 2,500 people per square kilometer, respectively, and the data for 2021 were missing, linear interpolation could be used to estimate the 2021 figure to be 2,250. This method, based on the assumption that data changes linearly over time, closely matches actual trends and ensures the rationality of data completion.
[0034] In one embodiment, after the data set is initially completed, it needs to be standardized to ensure consistency. For example, the population data units from different sources are unified into units per square kilometer, or the temperature units in meteorological data are unified into degrees Celsius. If the consistency check finds that some data distribution is abnormal, for example, the temperature data in a certain area is 10 degrees higher than that in other areas, it may be necessary to use the regional matching mechanism to extract more relevant records from the historical data as candidate values for supplementation. The selection of candidate values can be based on temporal or spatial proximity, such as giving priority to data from adjacent years or adjacent areas.
[0035] For example, the secondary verification step can further select the optimal fill value by comparing the distribution characteristics of candidate values with existing data. For example, suppose a meteorological monitoring point is missing rainfall data. There are three candidate values: 50 mm, 60 mm, and 100 mm. By analyzing the average rainfall of 55 mm at surrounding monitoring points, 50 mm is ultimately selected as the fill value. This approach can improve data accuracy and consistency.
[0036] In one embodiment, the multi-dimensional detection of the final completed data set may include quality assessment of the temporal dimension and the spatial dimension.
[0037] For example, the temporal dimension checks whether the data conforms to seasonal variations, and the spatial dimension checks whether data from adjacent areas are continuous. If winter temperatures in one area are significantly higher than those in surrounding areas, this may indicate a data anomaly and require further verification. This multi-dimensional testing helps ensure the usability of the dataset and provides a reliable basis for subsequent urban planning decisions.
[0038] For example, once data integrity meets business requirements, high-quality datasets can support more accurate urban resource allocation analysis, such as optimizing public facility layout through complete population density data or predicting the impact of extreme weather on urban traffic through meteorological data. These benefits can significantly improve the efficiency and accuracy of data-driven decision-making.
[0039] In step S103, the standard deviation threshold method is used to perform anomaly detection on the completed data set. If the data point deviates from the mean by more than a preset standard deviation range, it is marked as abnormal data and eliminated to generate a cleaned data set.
[0040] Step 1: Obtain supplementary data from the original dataset. Calculate the mean and standard deviation of the supplementary data. Determine the data range and threshold range through statistical analysis to obtain a preliminary anomaly detection benchmark. Step 2: Based on the preliminary anomaly detection benchmark, calculate the degree of deviation of each data point from the mean. If the deviation exceeds the preset threshold, mark it as an outlier, generating a marked dataset. Step 3: Perform a pruning operation on the marked dataset to remove data points marked as outliers from the dataset, generating a preliminary cleaned dataset. Step 4: Based on the preliminary cleaned dataset, recalculate the mean and standard deviation. By comparing the changes in the data range before and after, determine whether any potential anomalies remain unmarked, generating an updated detection benchmark. Step 5: Based on the updated detection benchmark, if the data range changes beyond the preset fluctuation range, perform the anomaly marking and pruning operations again to generate a secondary cleaned dataset. Step 6: Based on the secondary cleaned dataset, use the K-means clustering algorithm to group the data points. Analyze the distribution characteristics of the data within each group to determine whether the final cleaned dataset meets the preset stability criteria. Step 7: For the final cleaned data set, record the abnormal data removal log and data distribution characteristics, and complete the data cleaning process by storing the processed data and related statistical indicators.
[0041] For example, during the cleansing process of processing basic data sets, detailed analysis can be performed from multiple perspectives to detect and address data integrity and outliers, ensuring that data quality meets business requirements. The following provides specific examples and explanations for each technical topic, focusing on application scenarios in the data cleansing field.
[0042] For example, when calculating the mean and standard deviation of data to determine the anomaly detection benchmark, we can first understand in principle that the mean reflects the central trend of the data, while the standard deviation measures the degree of dispersion of the data.
[0043] In one possible implementation, suppose a sales dataset for a region contains 1,000 data points, with a calculated mean of 500 and a standard deviation of 50. A threshold range can be initially set to twice the standard deviation above and below the mean, or between 400 and 600. Data points outside this range are considered potential anomalies. This approach facilitates rapid identification of extreme values in the data, providing a foundation for subsequent cleaning.
[0044] For example, calculating the degree of deviation is crucial when labeling outliers. For example, if a data point has a value of 750, significantly exceeding the threshold range of 400 to 600 and deviating from the mean by 250, it would be labeled an outlier. This labeling approach helps identify outliers that may be caused by data entry errors or special events, ensuring the reliability of the dataset.
[0045] For example, in the case of outlier data removal, suppose 50 outliers are identified among 1,000 data points. After removal, the remaining 950 data points form a preliminary cleaned data set. This process directly reduces the impact of data noise and lays the foundation for subsequent analysis. The benefit of this removal operation is that it improves data consistency and prevents outliers from interfering with subsequent statistical results.
[0046] For example, when recalculating the mean and standard deviation to identify potential anomalies, suppose that after removing data, the mean becomes 495 and the standard deviation is reduced to 45. The change in data range is small, indicating that the initial cleaning is effective. If the range changes significantly, there may be unmarked anomalies that require further testing. This repeated verification method helps gradually approach the ideal data distribution.
[0047] For example, during the second round of cleaning, if the data range still fluctuates beyond the preset value, such as if the standard deviation changes by more than 10%, the abnormal data will be marked and removed again until the data stabilizes. This iterative approach ensures the gradual optimization of the data and reduces the possibility of misjudgment.
[0048] For example, when using the K-means clustering algorithm to group data, the data can be divided into three groups. The distribution characteristics within each group can be analyzed to determine whether stability conditions are met. For example, if the data distribution in one group is relatively concentrated and the standard deviation is only 20, it indicates that the data quality is high, while another group with a standard deviation of 60 may require further processing. This type of group analysis helps to examine data characteristics from different dimensions and improve the comprehensiveness of data cleaning.
[0049] For example, when recording abnormal data removal and distribution characteristics, you can record in detail the number of outliers removed each time, along with timestamps and statistical indicators of the data distribution, such as the changing trends of the mean and standard deviation. These records provide important evidence for subsequent data tracing and quality assessment, ensuring the auditability of the process and providing data support for business decisions.
[0050] Step S104 : performing a logic check on the cleaned data set and the garbage bin layout benchmark information. If a data inconsistency is detected, a rule base correction mechanism is triggered to generate an integrated data set.
[0051] By obtaining the original records from the cleaned dataset and the trash bin layout, and using data preprocessing tools to unify the format, a preliminary organized dataset is obtained. If there are field mismatches between the preliminary organized dataset and the benchmark information, the data comparison process is triggered, the inconsistent fields are marked, and the range of abnormal data is determined. Based on the marked abnormal data range, the content of the pre-established rule library is obtained, and the data contradictions are compared item by item to determine whether they meet the correction conditions. If the data contradictions meet the correction conditions, the correction mechanism is activated, and the corresponding strategy in the rule library is used to adjust the abnormal data to obtain the corrected data records. The corrected data records are subjected to a secondary logical check with the benchmark information, and the support vector machine algorithm is used to evaluate the data consistency and determine the verification results. Based on the verification results, the data records that meet the consistency requirements are integrated to generate the final integrated dataset, completing the data processing process.
[0052] For example, when processing clean datasets and raw records of waste bin layouts, relevant fields can be extracted from the data source, such as the bin number, location coordinates, and capacity information. Data preprocessing tools can then be used to unify records in different formats into a standardized table format. If some of the raw records contain text data and some data in numeric form, the preprocessing tool will convert all fields into a unified format, such as converting capacity to liters, to ensure consistency in subsequent processing.
[0053] Specifically, when comparing the initially collated dataset with the baseline data, if a field mismatch is detected—for example, if the baseline data requires the trash bin capacity field to be an integer value, but the actual data contains decimals or null values—the data comparison process is triggered. For mismatched fields, the system automatically marks the abnormal data range, for example, marking data in the capacity field that is less than 0 or greater than 1000 liters as abnormal. This marking method helps quickly locate problematic data and provides a basis for subsequent corrections.
[0054] For example, when invoking a rule base based on a range of flagged abnormal data, suppose the rule base stipulates that abnormal capacity values can be corrected using historical averages. If a trash bin's capacity record is -5 liters, which is clearly inconsistent, the system will search the rule base for the bin's average capacity over the past 30 days. If the average is 300 liters, for example, the abnormal data will be replaced with this value. This item-by-item comparison and correction method effectively reduces data inconsistencies and ensures the rationality of data records.
[0055] Specifically, after activating the correction mechanism, adjustments to abnormal data can combine multiple strategies.
[0056] For example, for a garbage bin record with unusual coordinates, if the coordinates are outside the city limits, the rule library might provide an interpolation correction method based on the locations of neighboring garbage bins. For example, assuming the coordinates of the neighboring garbage bin are 120.5 degrees east longitude and 30.2 degrees north latitude, the unusual coordinates are adjusted based on this. This approach maintains spatial consistency of the data.
[0057] For example, when performing a secondary logical check on corrected data records, a support vector machine algorithm is used to assess consistency. This analysis can determine the rationality of the correction results by analyzing the data distribution characteristics. For example, if the corrected capacity data is concentrated between 200 and 500 liters, but a single record still contains 800 liters, the algorithm will flag it as a potential inconsistency and prompt further verification. This assessment method helps improve the overall reliability of the data.
[0058] Specifically, during the generation of the final integrated dataset, the system selects data records that meet consistency requirements. For example, it consolidates all verified bin numbers, locations, and capacity information into a complete table. This integration facilitates subsequent analysis and application. By recording data changes before and after corrections, it creates a traceable processing log, supporting data management.
[0059] Step S105 , performing a redundant field deletion operation on the integrated data set, and fusing the meteorological time series data and the garbage bin distribution data to generate a structured feature set.
[0060] A preliminary scan of the integrated dataset identifies redundant fields, and field correlation analysis is used to determine the redundant fields. Based on the redundant field determination results, deletion operations are performed, and the remaining fields are standardized to obtain a streamlined dataset. Meteorological time series information is extracted from the streamlined dataset, and time series decomposition methods are used to determine the changing trends and cyclical characteristics of the meteorological data. Based on the changing trends and cyclical characteristics of the meteorological data, information related to garbage distribution is integrated, and a preliminary spatiotemporal feature combination is constructed using spatial mapping technology. Structural features are extracted from this preliminary spatiotemporal feature combination, and principal component analysis is used to determine the independence and importance distribution between features. Based on the independence and importance distribution between features, the feature construction process is optimized to obtain the final structured feature set. The structured feature set is stored and formatted to generate standardized data output suitable for subsequent analysis.
[0061] For example, during the initial scan of a consolidated dataset, a field duplication detection tool can be used to identify fields, such as bin numbers, that appear repeatedly across multiple tables. Imagine a dataset contains 1,000 records, 50 of which have identical numbers. These redundant fields increase the data storage burden. Field correlation analysis can further confirm that these repeated fields have low relevance to core business operations, such as bin location distribution. Consequently, these fields are identified as redundant and deleted, streamlining the data and improving subsequent processing efficiency.
[0062] For example, when standardizing the remaining fields, we can standardize the units of the trash bin capacity field to liters. Assuming that some records in the original data are in cubic meters and some in liters, we can use the conversion rules to adjust all data to a consistent format, resulting in a streamlined dataset. This standardization helps reduce data ambiguity and lays the foundation for subsequent analysis.
[0063] For example, when extracting meteorological time series information, we can filter out the past year's rainfall and temperature data from the dataset and use time series decomposition methods to separate the data into trend and cyclical components. Suppose we find that rainfall shows a clear upward trend from June to August each year. This cyclical characteristic can provide a reference for predicting waste distribution. Such analysis helps us understand the potential impact of meteorological factors on waste generation.
[0064] For example, when integrating information about garbage distribution and constructing spatiotemporal feature combinations, spatial mapping techniques can be used to correlate garbage bin locations with meteorological data. For example, if there are 10 garbage bins in an area, and 5 of them are located in areas with high rainfall, it can be inferred that garbage collection frequency in these locations should be increased. This mapping approach facilitates more precise resource allocation planning.
[0065] For example, when extracting structural features and applying principal component analysis, we can determine the independence of features such as trash bin usage rate and location density. If the analysis results show that usage rate contributes 60% to the overall data, while location density contributes only 10%, then usage rate is prioritized as the core feature. This optimization method effectively reduces data dimensionality and improves analysis efficiency.
[0066] For example, when optimizing feature construction, feature weights can be adjusted based on importance distribution. For example, setting the weight of the usage feature to 0.7 and the location density to 0.3 ultimately generates a structured feature set. This adjustment can better reflect business needs and improve data representation capabilities.
[0067] For example, when storing and formatting structured feature sets, the data can be output in a common table format, assuming that each record contains key fields such as trash bin number, location, and usage rate, and ensuring that the data field names are consistent. This standardized output facilitates direct access by subsequent analysis tools, reducing data conversion costs and improving the convenience of data sharing.
[0068] Step S106 , performing time series analysis on the structured feature set, and marking a high-impact period if a meteorological feature value is detected to exceed a preset threshold, thereby generating a labeled feature set.
[0069] By processing the data of structured features, an initial dataset is obtained, and a preliminary time series analysis is performed on it to obtain time series feature data. Based on the time series feature data, a sliding window method is used to extract the dynamic trend of meteorological features and determine the fluctuation range of meteorological features. If the fluctuation range of meteorological features exceeds the preset threshold, the corresponding period is marked as a high-impact period using threshold judgment logic, resulting in marked period data. For this marked period data, an intermediate result of annotated features is generated by combining feature extraction methods to determine the feature distribution of high-impact periods. Key structured feature information is extracted from the feature distribution, and the annotated features are classified using a support vector machine algorithm to obtain a classified feature set. Based on the classified feature set, feature set analysis operations are performed, integrating the contextual information of the time series and judging the completeness of the feature set to obtain the final annotated feature set. By processing the data of the final annotated feature set, the correlation information between meteorological features and high-impact periods is integrated to determine the output data of the business analysis.
[0070] For example, when processing structured feature data to obtain an initial dataset, you can first perform a preliminary time series organization on the data. The core of time series organization is to arrange the data in chronological order to ensure the continuity of subsequent analysis. For example, suppose you are interested in the correlation between a city's meteorological data and the distribution of garbage bins. The initial dataset may contain information such as daily temperature and humidity. This data will be sorted by date to form a complete time series feature data. This organization method helps capture the patterns of meteorological characteristics changing over time.
[0071] For example, when using a sliding window approach to extract dynamic trends in meteorological characteristics, a time window of, say, seven days can be set, sliding daily to observe changes in temperature or humidity. For example, if the temperature gradually rises from 20°C to 28°C over a given week, a sliding window can help identify this rising trend and further determine the range of fluctuations. If the preset threshold is 5°C and the actual fluctuation exceeds 8°C, the threshold judgment logic will mark that period as a high-impact period. This marking method facilitates subsequent focus on key time periods for in-depth analysis.
[0072] For example, when combining labeled time period data with feature extraction methods to generate intermediate results, we can extract the specific meteorological characteristics of high-impact periods. For example, if humidity consistently exceeds 80% during a high-impact period, this characteristic will be recorded and analyzed for its potential impact on waste bin management. For example, excessive humidity may accelerate waste decomposition, affecting the frequency of collection. Determining the characteristic distribution provides the data foundation for subsequent classification.
[0073] For example, when using a support vector machine algorithm to classify labeled features, features can be divided into high-impact and low-impact categories. For example, if humidity and temperature features are classified as high-impact, while wind speed features are classified as low-impact, this classification helps identify the most critical feature sets for business analysis. This classified feature set provides a clear structure for subsequent analysis.
[0074] For example, when integrating time series context to determine feature set completeness, you can check whether the classified feature set covers all key time periods. For example, if a month has three high-impact time periods, but the feature set only contains data for two, you need to supplement the missing information to ensure the completeness of the final annotated feature set. This completeness check helps improve the reliability of the analysis.
[0075] For example, when integrating meteorological characteristics with information related to high-impact periods, we can correlate temperature and humidity data during these periods with the density of trash bins. If a region experiences high temperatures and humidity during these periods, along with a dense distribution of trash bins, increased cleaning frequency may be necessary. This integration of correlated information provides direct decision-making support for business analysis and helps optimize resource allocation.
[0076] For example, the final business analysis output data can be used to generate a report based on the analysis, including the specific characteristics of high-impact periods and waste management recommendations for the corresponding areas. This output data provides guidance for actual operations and significantly improves management efficiency.
[0077] Step S107 , inputting the labeled feature set into a random forest model, and combining it with historical removal records to predict the saturation distribution of garbage bins in each area.
[0078] By extracting feature data from historical collection and transportation records, an initial data set is constructed for subsequent processing. Based on the initial data set, a random forest model is used for training to obtain the saturation value prediction data of the garbage bins in each area. For the saturation value in the predicted data, if the saturation value of a certain area exceeds the preset threshold, the garbage bin distribution in that area is prioritized and the high-priority area range is determined. The historical collection and transportation records within the high-priority area are obtained, and combined with the current saturation value, the correlation between the regional division and the saturation distribution is analyzed to obtain the basis for distribution adjustment. Based on the distribution adjustment basis, the saturation distribution data of each area is recalculated to determine whether there is an uneven distribution phenomenon. If an uneven distribution phenomenon exists, the regional division is dynamically adjusted to generate a new regional division scheme to determine the optimized distribution. Based on the optimized distribution, the saturation value and regional division information in the feature data are updated to obtain the final saturation distribution prediction result.
[0079] For example, when processing historical collection records to build an initial dataset, key information can be extracted from the collection data for the past year, such as the frequency of collection for each garbage bin, when it is full, and the collection cycle in the area. This data can help form a table containing fields such as time, location, and saturation, which serves as the basis for subsequent analysis.
[0080] It should be noted that the construction of the initial dataset needs to ensure the integrity and accuracy of the data, such as reasonably filling in missing data and ensuring the consistency of timestamps, so that subsequent model training will not cause inaccurate predictions due to data bias.
[0081] For example, using a random forest model to predict saturation can be understood as using features from historical data, such as collection frequency and regional foot traffic, to train a model to predict the saturation level of trash bins within a certain timeframe. If, for example, a certain area's trash bins have been full an average of twice a day for the past month, and foot traffic has recently increased by 20%, the model might predict a saturation level of 85%, providing a reference for subsequent processing. This approach offers the advantage of identifying potential high-saturation areas in advance, providing data support for collection and dispatch scheduling.
[0082] For example, when prioritizing areas where saturation exceeds a preset threshold, let's say the threshold is set at 80%. When a region's predicted saturation reaches 90%, the system automatically marks it as a high-priority area. Further analysis of the distribution of waste bins in that area may reveal that some bins are overly concentrated, leading to localized oversaturation. Marking high-priority areas helps allocate more resources to these areas, alleviating pressure.
[0083] For example, when analyzing the correlation between historical clearance records and current saturation values in high-priority areas, by comparing clearance data from the past three months, we can find that the rapid increase in saturation after each clearance in a certain area may be due to the frequent commercial activities in the surrounding area. This correlation analysis provides a basis for subsequent regional division adjustments and optimizes resource allocation efficiency.
[0084] For example, when recalculating saturation distribution data based on the distribution adjustment criteria, if it is found that the distribution of trash bins in a certain area is uneven, such as a high concentration in the north and a sparse distribution in the south, a dynamic adjustment plan can be implemented to reallocate some trash bins to the south to balance the overall saturation. This adjustment can effectively reduce local overload and improve collection efficiency.
[0085] For example, after generating a new zoning plan, the rationality of the optimized distribution can be verified through simulation testing. For example, after adjustment, the saturation level in the northern region dropped from 90% to 75%, while the saturation level in the southern region increased from 50% to 70%, resulting in a more balanced overall distribution. This optimization also reduces wasted travel by transportation vehicles and improves operational efficiency.
[0086] For example, when updating the saturation values and regional division information in the feature data, the adjusted data can be re-entered into the system to form the final saturation distribution prediction results. This update ensures the real-time and accuracy of the data, provides a reliable basis for the formulation of subsequent removal plans, and reduces resource waste.
[0087] Step S108: calling a scheduling optimization model based on the saturation distribution, and dynamically generating cleaning path parameters if a high saturation area is detected, and outputting a final scheduling plan.
[0088] By acquiring real-time saturation distribution data from the monitoring system and comparing it with a pre-established database, the system conducts preliminary screening to determine whether any areas exceed a preset threshold, thereby initially identifying high-saturation areas. For these initially identified high-saturation areas, the scheduling optimization model conducts in-depth analysis, combining regional identification data with historical traffic information. If saturation levels consistently exceed the threshold, the system identifies key areas requiring adjustment. Based on these identified key areas, a dynamic calculation method is used to simulate the removal routes. Combining real-time traffic data and regional saturation distribution, multiple alternative route planning schemes are generated. These alternative route planning schemes are evaluated using a pre-set optimization model and parameter generation rules. If a particular scheme demonstrates higher overall efficiency than the others, it is designated as the preferred removal route. For these prioritized removal routes, real-time resource allocation data is acquired, and route parameters are adjusted using data analysis methods. If resource allocation is insufficient, the parameters are recalculated and new dispatch instructions are generated. Based on the adjusted route parameters and dispatch instructions, the final dispatch plan is generated and automatically distributed to the execution terminal through the system, completing the closed-loop optimization process.
[0089] For example, in actual waste collection and transportation management, obtaining real-time saturation distribution data through monitoring systems is a critical step. For example, suppose there are 10 waste bins in an area of a city. The system updates saturation data hourly and finds that three of them have a saturation level of 85%, exceeding the preset threshold of 80%. After comparing the data with the database, the area containing these three bins is initially identified as a high-saturation area. This method can quickly identify problem areas and provide a foundation for subsequent analysis.
[0090] For example, when invoking the scheduling optimization model for in-depth analysis of initially identified high-saturation areas, historical traffic information can be incorporated. For example, data from the past week shows that this area experiences higher traffic and higher waste generation rates during weekday lunchtime and evening hours. If the current saturation level remains above a threshold, for example, exceeding 80% for three consecutive hours, the area is identified as a priority for adjustment. This analysis helps pinpoint areas requiring urgent attention.
[0091] For example, after defining the scope of a key area, dynamic calculation methods can be used to simulate removal routes, incorporating real-time traffic data. Assuming congestion around the area, the system will simulate three alternative routes based on traffic conditions: Path A takes 40 minutes, Path B takes 50 minutes, and Path C takes 35 minutes but requires a longer detour. Using saturation distribution data, Path C, which covers bins with high saturation, is prioritized. This simulation method is effective in complex environments.
[0092] For example, when evaluating the preferred removal route from multiple alternative routes, the optimization model comprehensively considers parameters such as time, distance, and resource consumption. If Route C has the highest overall efficiency score of 90, while the other routes only score 70 or 75, Route C will be selected as the preferred option. This evaluation mechanism ensures maximum resource utilization.
[0093] For example, when obtaining resource allocation data for a priority removal route, if it is found that there are currently insufficient removal vehicles (for example, only one vehicle is available, while route C requires two vehicles), the route parameters can be adjusted through data analysis to reallocate vehicles or shorten the route coverage. This flexible adjustment method can avoid delays caused by resource shortages.
[0094] For example, when the final dispatch plan is generated based on the adjusted route parameters and sent to the execution terminal, the system automatically sends instructions to the removal vehicles, ensuring that the instructions are clear and executable. For example, the instructions include specific starting and ending points and time requirements. This closed-loop processing method improves execution efficiency and ensures that the removal tasks are completed on time.
[0095] For example, the application of real-time data and dynamic adjustments are central to the entire process. By continuously updating saturation data and traffic information, the system can respond promptly to changes, such as quickly adjusting routing if a particular bin suddenly reaches 100% saturation. This mechanism significantly improves management flexibility and responsiveness.
[0096] For example, as an expansion plan, temporary trash bins can be added to key areas or the frequency of collection can be adjusted. For example, two temporary trash bins could be added to a high-saturation area to alleviate pressure. This approach provides more possibilities for long-term optimization while ensuring the stability of daily management.
[0097] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for controlling an intelligent garbage classification box, characterized in that: The method comprises: Obtaining original data sets from the urban planning database, demographic database, and meteorological monitoring database, performing field mapping processing on the original data sets using a preset standardized template, and generating a standardized basic data set with a unified format; Perform integrity check on the standardized basic dataset. If it is detected that the proportion of missing values in a field exceeds a preset threshold, call the historical dataset of the corresponding area for interpolation and filling to generate a complete dataset. The standard deviation threshold method is used to perform anomaly detection on the completed data set. If the data point deviates from the mean by more than the preset standard deviation range, it is marked as abnormal data and removed to generate a cleaned data set; Performing a logical check on the cleaned dataset and the trash bin layout benchmark information, and triggering a rule base correction mechanism if a data inconsistency is detected to generate an integrated dataset; performing a redundant field deletion operation on the integrated data set, and fusing the meteorological time series data with the garbage bin distribution data to generate a structured feature set; Performing time series analysis on the structured feature set, marking a high-impact period if a meteorological feature value exceeds a preset threshold, and generating a labeled feature set; The labeled feature set is input into the random forest model, and the saturation distribution of garbage bins in each area is predicted based on historical removal records; The scheduling optimization model is called according to the saturation distribution. If a high-saturation area is detected, the cleaning path parameters are dynamically generated and the final scheduling plan is output.
2. The intelligent garbage classification box management and control method according to claim 1 is characterized in that: The method of obtaining the original data set from the urban planning database, the demographic database, and the meteorological monitoring database, and performing field mapping processing on the original data set using a preset standardized template to generate a standardized basic data set with a unified format includes: By obtaining original data sets from urban planning, demographic statistics and meteorological monitoring related databases, the data sets are preliminarily cleaned using automated tools to obtain a cleaned initial data set; Based on the cleaned initial dataset, a field mapping operation is performed on the dataset using a preset standardized template to determine whether the fields meet the unified format requirements. If the field content is missing or the format is inconsistent, it is supplemented or converted according to predefined rules to determine a mapped dataset with a unified format; For the mapped dataset, we obtain the correlation information between each field and compare the distribution characteristics of urban planning, demographics, and meteorological monitoring data to determine whether there is data redundancy or anomalies. If an outlier is detected, we filter it using a preset threshold to obtain a filtered standard dataset. Based on the filtered standard data set, a logistic regression model was used to analyze the potential relationships between the data, extract features based on the cross-influence between urban planning, demographics, and meteorological monitoring, and determine the key feature data set; Through the key feature data set, the priority order of data integration is obtained, and cross-domain data fusion processing is performed on the fields with higher priority to determine whether the fused data meets the integrity requirements. If the integrity is insufficient, it is supplemented with historical data to obtain an integrated comprehensive data set; Based on the integrated comprehensive data set, cluster analysis method is used to group the data, classify the distribution patterns of urban planning, demographics and meteorological monitoring, and obtain the classified grouped data set; Through the classified grouped data sets, the core feature information of each group of data is obtained, and multi-dimensional verification is performed on the feature information to determine whether the data meets the business logic requirements. If not, adjustments are made according to preset rules to determine the final business data set.
3. The intelligent garbage classification box management and control method according to claim 1 is characterized in that: The completeness check is performed on the standardized basic dataset. If it is detected that the proportion of missing values in a field exceeds a preset threshold, the historical dataset of the corresponding area is called for interpolation and filling to generate a complete dataset, including: By performing integrity checks on the basic data set, we can obtain the distribution of missing values in the field and determine whether the proportion of missing values exceeds the preset threshold. If the proportion of missing values in a field exceeds the preset threshold, historical datasets are obtained from the corresponding area to determine the data range that can be used for interpolation filling; According to the distribution characteristics of the historical data set, the missing fields are filled using the linear interpolation method to obtain a preliminary completed data set; For the initially completed data set, perform data standardization, obtain consistency verification results, and determine whether the data distribution meets the preset standards; If the consistency check result does not meet the requirements, the region matching mechanism is used to extract additional relevant records from the historical data set to determine the candidate values for supplementary filling; Based on the distribution characteristics of the candidate values, a secondary verification method is used to obtain the final completed data set to determine whether the data integrity meets business requirements; By performing multi-dimensional testing on the final completed dataset, data quality assessment results are obtained to determine the availability of the dataset.
4. The intelligent garbage classification box management and control method according to claim 1, characterized in that: The standard deviation threshold method is used to perform anomaly detection on the completed data set. If the data point deviates from the mean by more than a preset standard deviation range, it is marked as abnormal data and eliminated to generate a cleaned data set, including: Step 1: Obtain supplementary data from the original data set, calculate the data mean and standard deviation for the supplementary data, determine the data range and threshold range through statistical analysis, and obtain a preliminary anomaly detection benchmark; Step 2: Based on the preliminary anomaly detection benchmark, the degree of deviation of each data point from the data mean is calculated. If the deviation exceeds the preset threshold range, it is marked as anomaly data to obtain the marked data set; Step 3: Perform a culling operation on the labeled data set to remove data points marked as abnormal from the data set to generate a preliminary cleaned data set; Step 4: Based on the preliminary cleaned data set, recalculate the data mean and standard deviation. By comparing the changes in the data range before and after, determine whether there is potential abnormal data that has not been marked, and obtain an updated detection benchmark; Step 5: For the updated detection benchmark, if the data range changes beyond the preset fluctuation range, perform anomaly marking and elimination operations again to generate a secondary cleansing data set; Step 6: Based on the secondary cleaned data set, use the K-means clustering algorithm to group the data points. By analyzing the distribution characteristics of the data in each group, determine whether the final cleaned data set meets the preset stability conditions; Step 7: For the final cleaned data set, record the abnormal data removal log and data distribution characteristics, and complete the data cleaning process by storing the processed data and related statistical indicators.
5. The intelligent garbage classification box management and control method according to claim 1 is characterized in that: The cleaning data set is logically checked against the trash bin layout benchmark information, and if a data inconsistency is detected, a rule base correction mechanism is triggered to generate an integrated data set, including: By obtaining the original records from the cleaned dataset and the bin layout, the data preprocessing tools are used to unify the format and obtain a preliminary organized dataset; If there are any fields that do not match between the initially collated data set and the benchmark information, the data comparison process is triggered to mark the inconsistent fields and determine the scope of the abnormal data; Based on the range of marked abnormal data, the pre-established rule base is obtained, and data inconsistencies are compared item by item to determine whether they meet the correction conditions; If the data contradiction meets the correction conditions, the correction mechanism is activated and the corresponding strategy in the rule base is used to adjust the abnormal data to obtain the corrected data record; Perform a secondary logic check on the corrected data records and the benchmark information, use the support vector machine algorithm to evaluate the data consistency, and determine the verification result; Based on the verification results, the data records that meet the consistency requirements are integrated to generate the final integrated data set, completing the data processing process.
6. The intelligent garbage classification box management and control method according to claim 1, characterized in that: The redundant field deletion operation is performed on the integrated data set, and the meteorological time series data and the garbage bin distribution data are integrated to generate a structured feature set, including: By performing a preliminary scan of the integrated data set, redundant fields are identified, and the field correlation analysis method is used to obtain the redundant field determination results; Based on the results of the redundant fields, the deletion operation is performed, and the remaining fields are standardized to obtain a streamlined data set; Extract meteorological time series information from the streamlined data set and use time series decomposition methods to determine the changing trends and periodic characteristics of meteorological data; Based on the changing trends and periodic characteristics of meteorological data, information related to garbage distribution is integrated and spatial mapping technology is used to construct a preliminary spatiotemporal feature combination; From the preliminary spatiotemporal feature combination, structural features are extracted and principal component analysis is used to determine the independence and importance distribution of features. According to the independence and importance distribution between features, the feature construction process is optimized to obtain the final structured feature set; By storing and formatting structured feature sets, standardized data output suitable for subsequent analysis is generated.
7. The intelligent garbage classification box management and control method according to claim 1, characterized in that: The time series analysis is performed on the structured feature set. If a meteorological feature value is detected to exceed a preset threshold, it is marked as a high-impact period, and a marked feature set is generated, including: By processing the data of structured features, the initial data set is obtained, and the time series feature data is obtained by preliminary sorting of the time series data; Based on the time series characteristic data, the sliding window method is used to extract the dynamic change trend of meteorological characteristics and determine the fluctuation range of meteorological characteristics; If the fluctuation range of meteorological characteristics exceeds the preset threshold, the corresponding period is marked as a high-impact period through the threshold judgment logic, and the marked period data is obtained; For the marked time period data, combined with feature extraction methods, intermediate results of the labeled features are generated to determine the feature distribution of the high-impact time period; Obtain key structural feature information from feature distribution, use support vector machine algorithm to classify the labeled features, and obtain the classified feature set; Based on the classified feature set, perform feature set analysis operations, integrate the context information of the time series, judge the integrity of the feature set, and obtain the final labeled feature set; By processing the data of the final labeled feature set, the correlation information between meteorological characteristics and high-impact periods is integrated to determine the output data of business analysis.
8. The intelligent garbage classification box management and control method according to claim 1, characterized in that: The labeled feature set is input into the random forest model, and the historical removal records are combined to predict the saturation distribution of garbage bins in each area, including: By extracting feature data from historical removal records, an initial data set is constructed for subsequent processing; Based on the initial data set, a random forest model is used for training to obtain the saturation value prediction data of the garbage bins in each area; For the saturation value in the predicted data, if the saturation value of a certain area exceeds the preset threshold, the garbage bin distribution in the area is prioritized and the range of the high-priority area is determined; Obtain historical removal records within high-priority areas, combine them with current saturation values, analyze the correlation between regional divisions and saturation distribution, and obtain the basis for distribution adjustment; Based on the distribution adjustment basis, recalculate the saturation distribution data of each region to determine whether there is uneven distribution; If uneven distribution exists, the regional division is dynamically adjusted to generate a new regional division scheme and determine the optimized distribution; Based on the optimized distribution, the saturation value and region division information in the feature data are updated to obtain the final saturation distribution prediction result.
9. The intelligent garbage classification box management and control method according to claim 1, characterized in that: The scheduling optimization model is called according to the saturation distribution. If a high saturation area is detected, the cleaning path parameters are dynamically generated and the final scheduling plan is output, including: By obtaining real-time saturation distribution data from the monitoring system and comparing and preliminarily screening it with a pre-established database, it is determined whether there are areas exceeding the preset threshold, and preliminary identification results of high-saturation areas are obtained; For the initially identified high-saturation areas, the dispatch optimization model is called for in-depth analysis. Combining regional identification data with historical traffic information, if the saturation is detected to be consistently above the threshold, the key areas requiring adjustment are determined. Based on the identified key areas, a dynamic calculation method is used to simulate the removal routes, combining real-time traffic data and regional saturation distribution to generate multiple alternative route planning schemes; From multiple alternative route planning schemes, a preset optimization model is used to evaluate them. Combined with the parameter generation rules, if a scheme has a higher overall efficiency than other schemes, then this scheme is determined to be the priority removal route; For the priority removal routes, obtain real-time resource allocation data and adjust route parameters through data analysis methods. If resource allocation is insufficient, recalculate parameters and generate new dispatch instructions. Based on the adjusted path parameters and scheduling instructions, the final scheduling plan is generated and automatically sent to the execution terminal through the system, completing the closed-loop processing of the entire optimization process.
Citation Information
Cited By
Intelligent farm waste treatment and energy conversion method and system
CN120912374A
Retrieval enhancement method and system based on unstructured data
CN120994813A
Garbage throwing optimization management method and system based on classification data mining
CN121525993A