Rubbish bin intelligent simulation method and equipment based on reinforcement learning

By dividing the waste storage area into grids and mining historical data, combined with reinforcement learning models, intelligent decision-making and operation instructions are generated, which solves the problem of lack of refinement and intelligence in waste storage management and improves the quality of waste fermentation and the efficiency of incineration power generation.

CN122065639APending Publication Date: 2026-05-19北京朝阳环境集团有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
北京朝阳环境集团有限公司
Filing Date
2025-12-24
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing waste storage management technologies lack refined division and precise management methods, making it difficult to achieve grid-level accurate control over waste distribution and status. They also fail to effectively combine historical operational data with real-time monitoring information for systematic analysis and intelligent decision-making, thus hindering the improvement of waste fermentation quality and incineration power generation efficiency.

Method used

By spatially dividing the target waste bin into grids and mining the waste distribution patterns and status characteristics of each grid unit using historical operational data, cross-unit association is achieved through feature transfer, the total calorific value data of the waste is determined, and reinforcement learning objectives are determined based on the hierarchical decomposition and quantitative evaluation of expected management needs. Multi-source simulation information is generated through dynamic identification of partitions and encoding of real-time operational data, and finally, intelligent decision-making operation instructions are generated by training the reinforcement learning model.

Benefits of technology

It has enabled intelligent and refined upgrades to the management of waste storage facilities, improved the quality of waste fermentation and the efficiency of incineration power generation, adapted to dynamic changes such as fluctuations in waste supply, and provided systematic technical support for the efficient and stable operation of waste storage facilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122065639A_ABST
    Figure CN122065639A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer simulation, and provides a garbage bin intelligent simulation method and equipment based on reinforcement learning, which are used for improving garbage fermentation quality and incineration power generation efficiency. The method comprises the following steps: in response to an intelligent simulation task of a target garbage bin, carrying out space gridding division processing on the target garbage bin to obtain a plurality of grid units; mining a garbage distribution rule and state feature information of each grid unit based on historical operation data of the target garbage bin; cross-unit association is achieved through feature transmission, and total garbage calorific value data are determined; expected management demand features are collected and subjected to hierarchical disassembly, and a reinforcement learning target is determined in combination with total garbage calorific value data; generating aligned multi-source simulation information through partition dynamic identification and real-time operation data coding; and training the initial reinforcement learning model by using the aligned multi-source simulation information, inputting the aligned multi-source actual measurement information into the trained model, and obtaining an intelligent decision operation instruction under the preset calorific value promotion label.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer simulation technology, specifically relating to an intelligent simulation method and device for waste storage based on reinforcement learning. Background Technology

[0002] Waste-to-energy incineration is an important method for reducing and recycling municipal solid waste, and the management level of the waste storage facility directly affects the efficiency of incineration and the effectiveness of waste treatment. Traditional waste storage facility management relies heavily on the experience and judgment of on-site operators. Operators manually control the waste crane to handle, stack, and feed waste by observing the accumulation and fermentation status of the waste within the facility. With the development of waste-to-energy incineration technology, the requirements for refined and intelligent waste storage facility management are increasing. Existing technologies are gradually incorporating data monitoring and analysis methods, such as using sensors to collect environmental parameters like temperature and humidity within the waste storage facility, and combining this data with data on the amount of waste input and incinerated to perform simple statistical analysis, assisting operators in decision-making. Meanwhile, some research attempts to apply simulation technology to waste storage facility management, simulating the accumulation and fermentation process of waste by establishing physical models of the waste storage facility. However, existing simulation models are mostly based on empirical formulas and lack the ability to deeply integrate and dynamically adjust to actual operational data.

[0003] It is evident that existing technologies lack refined division and precise management methods for waste storage space, making it difficult to achieve grid-level precise control over waste distribution and status. Furthermore, they fail to effectively combine historical operational data with real-time monitoring information for systematic analysis and intelligent decision-making, making it difficult to adapt to dynamic changes and complex scenario requirements during waste storage operation, and consequently, making it difficult to improve waste fermentation quality and incineration power generation efficiency. Summary of the Invention

[0004] This application provides a reinforcement learning-based intelligent simulation method and device for waste storage facilities, which can realize intelligent and refined waste storage facility management and improve the quality of waste fermentation and the efficiency of incineration power generation.

[0005] This application provides a reinforcement learning-based intelligent simulation method for waste bins, applied to computer equipment. The method includes: In response to the intelligent simulation task for the target waste bin, the target waste bin is spatially gridded to obtain several grid cells; Based on the historical operational data of the target waste bin, the distribution patterns of each grid cell are mined and the waste accumulation status is traced to obtain the waste distribution patterns and waste status characteristics of each grid cell. By performing cross-cell association processing on the waste distribution pattern and waste status feature information corresponding to each grid cell through feature transfer, inter-grid association features are generated. Combining the waste distribution pattern, the waste status feature information and the inter-grid association features, the total waste calorific value data is determined. The expected management needs of the target waste bin are collected and decomposed hierarchically to obtain hierarchical features of the needs. The hierarchical features of the needs and the total calorific value data of the waste are converted into evaluation indicators with the same dimension or without units. Based on the quantitative relationship between the converted demand evaluation indicators and the calorific value evaluation indicators, the reinforcement learning objective of the intelligent simulation task is determined. Based on the reinforcement learning objective, the dynamic trajectory simulation information of waste in each grid cell is determined by partition dynamic identification. Simultaneously, the real-time simulation operation data of the target waste bin is processed by state encoding to generate state simulation encoding information, and multi-source simulation information after the state simulation encoding information and the dynamic trajectory simulation information are aligned is generated. The aligned multi-source simulation information is used as the input to the initial reinforcement learning model. A preset policy optimization algorithm is called to train the initial reinforcement learning model to generate a post-trained reinforcement learning model. The aligned multi-source measured information is then input into the post-trained reinforcement learning model to obtain intelligent decision-making operation instructions for the target waste bin under a preset heat value enhancement label.

[0006] This application provides a computer device including a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the above-described method.

[0007] This application provides a computer-readable storage medium including a computer program that, when run on a computer device, causes the computer device to perform the steps of the above-described method.

[0008] In this application, the target waste bin is spatially gridded, and historical operational data is used to mine the waste distribution patterns and state characteristics of each grid unit. Feature transfer is used to achieve cross-unit correlation and determine the total calorific value of the waste. Then, based on the hierarchical decomposition and quantitative evaluation of expected management needs, reinforcement learning objectives are determined. Aligned multi-source simulation information is generated through dynamic partitioning and real-time operational data encoding. Finally, intelligent decision-making operation instructions are generated using reinforcement learning model training and experimental information input. This application embodiment comprehensively achieves an intelligent and refined upgrade of waste bin management, breaking through the limitations of traditional waste bin management. The management system, which relies on experience-based judgment and lacks systematic quantitative analysis, is limited. By dividing the data into grids, the management precision is improved from the regional level to the grid level. Combined with historical data mining and cross-unit feature correlation, a comprehensive and accurate understanding of waste distribution, status, and calorific value is achieved. Through the scientific setting of reinforcement learning objectives and model training, the system can automatically explore the optimal operating strategy in complex waste storage operation scenarios. The generated intelligent decision-making operation instructions can effectively guide the actual operation of the waste storage, improve the quality of waste fermentation and the efficiency of incineration power generation, and adapt to dynamic changes such as fluctuations in waste supply. This provides systematic technical support for the efficient and stable operation of the waste storage. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating a reinforcement learning-based intelligent simulation method for waste bins provided in an embodiment of this application.

[0010] Figure 2 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application.

[0011] Figure 3 This is a functional block diagram of a computer device provided in an embodiment of this application. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0013] See Figure 1 This is a reinforcement learning-based intelligent simulation method for waste bins provided in the embodiments of this application. This method can be applied to computer equipment, and the specific process is as follows: steps 110-160.

[0014] Step 110: In response to the intelligent simulation task for the target waste bin, the target waste bin is spatially gridded to obtain several grid cells.

[0015] In this embodiment, the target waste storage facility is a large, enclosed facility in a co-processing base for municipal solid waste and industrial waste. The application scenario needs to cover core requirements such as refined spatial management, achieving grid-level precision beyond traditional "district-level" management, and considering vertical fermentation differences. After receiving the intelligent simulation task for the waste storage facility, the computer equipment first retrieves the three-dimensional spatial parameters of the target waste storage facility, including its horizontal length and width dimensions and vertical height dimensions, while simultaneously acquiring the layout data of the fixed facilities inside the waste storage facility. Specifically, the computer equipment performs spatial gridding processing according to preset grid division rules: horizontally, the bottom surface of the waste storage facility is divided into several square areas based on preset grid side lengths; vertically, it is layered according to preset layer spacing, the layer spacing being set with reference to research data on vertical fermentation differences in waste, ensuring that each vertical layer reflects the fermentation environment differences at different heights. Furthermore, the target waste storage facility is divided into several uniquely identified grid units, each grid unit's identifier including a horizontal position code and a vertical layer code.

[0016] Step 120: Based on the historical operation data of the target waste bin, perform distribution pattern mining and waste accumulation status tracing for each grid cell to obtain the waste distribution pattern and waste status characteristic information corresponding to each grid cell.

[0017] In this embodiment, the computer device first collects historical operational data of the target waste storage area over a continuous period of time. This data includes waste input records, waste transfer records, and waste fermentation monitoring data. Furthermore, for each grid cell, the computer device extracts waste distribution data from the historical operational data within that grid cell. By statistically analyzing the changes in waste distribution density and range over different time periods, it uncovers waste distribution patterns. Simultaneously, the computer device traces the waste accumulation process within the grid cell, including changes in waste accumulation height, accumulation duration, and waste composition, thereby obtaining waste status characteristic information. For example, the waste distribution pattern of a certain grid cell may show that its waste distribution density is higher on weekdays than on weekends, and the waste status characteristic information may include the average accumulation height of waste within that grid cell and the proportion of major waste components.

[0018] Step 130: Perform cross-cell correlation processing on the waste distribution pattern and waste status feature information corresponding to each grid cell through feature transfer to generate inter-grid correlation features. Combine the waste distribution pattern, waste status feature information and inter-grid correlation features to determine the total waste calorific value data.

[0019] In this embodiment of the application, the computer device first performs preliminary sorting of the garbage distribution pattern and garbage status characteristics information of each grid cell to prepare for cross-cell association processing.

[0020] Step 131: Extract the distribution density and distribution time sequence features from the waste distribution pattern of each grid cell, and extract the waste component features and waste accumulation height features from the waste state feature information of each grid cell.

[0021] In this embodiment, the computer device extracts the distribution density and distribution time sequence characteristics of the waste in each grid cell based on its waste distribution pattern: the distribution density characteristic is obtained by calculating the average volume percentage of waste within the grid cell per unit time, and the distribution time sequence characteristic is obtained by analyzing the changing trend of waste distribution density over time. Simultaneously, the computer device extracts waste composition characteristics and waste accumulation height characteristics from the waste state characteristic information of each grid cell: the waste composition characteristic reflects the proportion of different components in the waste within the grid cell, such as the proportion of organic matter, inorganic matter, and recyclables; the waste accumulation height characteristic is the average accumulation height data of the waste within the grid cell.

[0022] Step 132: Establish a grid cell association topology with adjacent grid cells as the association objects, transfer the distribution density characteristics and distribution time series characteristics of each grid cell to the adjacent grid cells, calculate the feature attenuation coefficient during the transfer process, adjust the feature values ​​after transfer based on the feature attenuation coefficient, and generate the inter-grid association features corresponding to each grid cell.

[0023] In this embodiment, the computer device first determines the adjacent grid cells of each grid cell. Adjacent grid cells include those adjacent in the horizontal direction and those adjacent in the vertical direction. Then, a grid cell association topology is established using these adjacent grid cells as the association objects. Furthermore, the computer device transmits the distribution density characteristics and distribution time sequence characteristics of each grid cell to its adjacent grid cells. During this transmission, the computer device calculates a feature attenuation coefficient based on the distance between adjacent grid cells and obstacles to waste transfer. The greater the distance and the more obstacles, the smaller the feature attenuation coefficient. Specifically, the computer device adjusts the transmitted feature values ​​based on the feature attenuation coefficient to obtain the inter-grid association characteristics corresponding to each grid cell. For example, after the distribution density characteristics of a grid cell are transmitted to adjacent grid cells, its feature values ​​are adjusted accordingly based on the attenuation coefficient. The adjusted feature values ​​are a part of the inter-grid association characteristics of that adjacent grid cell.

[0024] Step 133: Standardize or normalize the waste distribution pattern, waste status characteristics and corresponding inter-grid correlation characteristics of each grid cell, and splice or fuse the processed features to generate a single-grid multi-dimensional feature vector for each grid cell. Summarize the single-grid multi-dimensional feature vectors of all grid cells to obtain a multi-dimensional waste feature set.

[0025] In this embodiment, the computer device first standardizes the waste distribution pattern, waste status characteristics, and inter-grid correlation characteristics for each grid cell. The standardization method converts the feature values ​​into values ​​with a mean of 0 and a standard deviation of 1, ensuring comparability between different features. After processing, the computer device concatenates the standardized features to generate a single-grid multi-dimensional feature vector for each grid cell. This vector contains comprehensive information about the grid cell's waste distribution, waste status, and inter-grid correlations. Then, the computer device aggregates the single-grid multi-dimensional feature vectors of all grid cells to obtain a multi-dimensional waste feature set.

[0026] Step 134: Extract the waste component features and distribution density features from the multi-dimensional waste feature set, match the waste component features with the preset component calorific value mapping relationship, and obtain the initial calorific value data corresponding to each grid cell.

[0027] In this embodiment, the computer device extracts the waste component characteristics and distribution density characteristics of each grid cell from a multi-dimensional waste feature set. Furthermore, the computer device invokes a preset component calorific value mapping relationship, which is established based on a large amount of waste calorific value experimental data; different combinations of waste components correspond to different calorific value values. The computer device matches the waste component characteristics of each grid cell with this mapping relationship to obtain the initial calorific value data corresponding to each grid cell. For example, if the proportion of organic matter in the waste component characteristics of a certain grid cell is high, according to the component calorific value mapping relationship, its corresponding initial calorific value data will be relatively high.

[0028] Step 135: Calculate the heat value correction coefficient for each grid cell based on the distribution density characteristics and inter-grid correlation characteristics. The heat value correction coefficient is a dimensionless quantity.

[0029] In this embodiment, the computer device first analyzes the relationship between distribution density characteristics, inter-grid correlation characteristics, and calorific value: the distribution density characteristics reflect the density of waste within a grid cell; a higher distribution density may affect the fermentation efficiency of the waste, thus affecting the calorific value; the inter-grid correlation characteristics reflect the correlation between the grid cell and other grid cells; closely correlated grid cells may have mutual influence on their calorific values. Specifically, the computer device calculates the calorific value correction coefficient for each grid cell using a preset algorithm. The algorithm's input is the standardized value of the distribution density characteristics and the inter-grid correlation characteristics, and the output is a unitless calorific value correction coefficient. For example, a grid cell with a high distribution density characteristic value and a large inter-grid correlation characteristic value may have a calorific value correction coefficient greater than 1 to positively correct the initial calorific value data.

[0030] Step 136: Use the calorific value correction coefficient to weight or scale the initial calorific value data to generate the corrected calorific value data for each grid cell; combine the corrected calorific value data of all grid cells to generate the full calorific value data of the waste.

[0031] In this embodiment, the computer device multiplies the initial calorific value data of each grid cell by the corresponding calorific value correction coefficient to obtain the corrected calorific value data. Specifically, the computer device aggregates the corrected calorific value data of all grid cells to generate the total calorific value data of the target waste bin, which includes the corrected calorific value data of each grid cell and the total calorific value data of the entire waste bin.

[0032] Step 140: Collect the expected management requirements of the target waste bin and decompose them hierarchically to obtain the requirements hierarchical features. Convert the requirements hierarchical features and the total waste calorific value data into evaluation indicators with the same dimension or without units. Based on the quantitative relationship between the converted requirements evaluation indicators and the calorific value evaluation indicators, determine the reinforcement learning objective of the intelligent simulation task.

[0033] In this embodiment, the computer device first collects the expected management requirements of the target waste storage area, including requirements for waste treatment efficiency, waste storage capacity, and waste fermentation quality. Then, the computer device hierarchically decomposes these requirements, breaking down waste treatment efficiency into sub-requirements such as waste transfer efficiency and waste input efficiency; waste storage capacity into sub-requirements such as grid cell capacity utilization and total waste storage area capacity utilization; and waste fermentation quality into sub-requirements such as waste calorific value enhancement and waste fermentation uniformity. This yields a hierarchical set of requirements. Specifically, the computer device converts these hierarchical requirements and the total waste calorific value data into dimensionless or unitless evaluation indicators. For example, waste transfer efficiency is converted into a standardized indicator of waste transfer time, and the total waste calorific value data is converted into a standardized indicator of calorific value. Then, the computer device analyzes the quantitative relationship between the converted requirement evaluation indicators and the calorific value evaluation indicators. Combined with the objective of the intelligent simulation task, it determines the reinforcement learning objective. For example, the reinforcement learning objective might be to maximize the increase in waste calorific value while meeting the waste storage capacity requirement.

[0034] Step 141: Collect the expected management requirements of the target waste storage facility. The expected management requirements include waste treatment efficiency requirements, waste storage capacity requirements, and waste transfer coordination requirements.

[0035] In this embodiment, the computer device interacts with the management system of the target waste storage facility to collect the expected management requirements. Waste treatment efficiency requirements include the average time limit from waste input to fermentation completion, and the daily throughput requirement for waste treatment; waste storage capacity requirements include the maximum allowable stacking height of each grid unit and the total capacity limit of the waste storage facility; waste transfer coordination requirements include the scheduling coordination requirements for waste transfer equipment and the time coordination requirements between waste transfer and waste input.

[0036] Step 142: Decompose the expected management requirements into hierarchical features. Decompose the waste treatment efficiency requirements into treatment timeliness features and treatment quality features. Decompose the waste storage capacity requirements into capacity utilization rate features and capacity dynamic adjustment features. Decompose the waste transfer coordination requirements into transfer sequence features and transfer connection features, thus obtaining multi-dimensional hierarchical requirements.

[0037] In this embodiment, the computer equipment addresses the waste treatment efficiency requirements by breaking them down into treatment timeliness characteristics and treatment quality characteristics: treatment timeliness characteristics include the upper limit of waste treatment time and the time stability requirements for waste treatment; treatment quality characteristics include the calorific value compliance rate after waste fermentation and the uniformity requirements for waste fermentation. Addressing the waste storage capacity requirements, the equipment breaks them down into capacity utilization rate characteristics and capacity dynamic adjustment characteristics: capacity utilization rate characteristics include the capacity utilization rate targets for each grid unit and the total capacity utilization rate target for the waste storage area; capacity dynamic adjustment characteristics include the response speed requirements for adjusting the grid unit capacity according to changes in waste volume and the limits on the magnitude of capacity adjustment. Addressing the waste transfer coordination requirements, the equipment breaks them down into transfer timing characteristics and transfer connection characteristics: transfer timing characteristics include the time interval requirements for waste transfer and the time difference limit between waste transfer and waste input; transfer connection characteristics include the connection time requirements between waste transfer equipment and the smoothness requirements for the connection between waste transfer and waste treatment processes. Through the above breakdown, multi-dimensional hierarchical characteristics of the requirements are obtained.

[0038] Step 143: Perform feature quantification on the demand-level features to generate demand quantification indicators corresponding to each demand-level feature, and perform partitioned heat value statistical processing on the full amount of waste heat value data to obtain the partitioned heat value statistical results of each grid cell.

[0039] In this embodiment, the computer device quantifies each demand level feature. For example, it quantifies the upper limit of waste treatment time in the processing timeliness feature into a specific time value, and quantifies the calorific value compliance rate of fermented waste in the processing quality feature into a percentage value, thereby generating demand quantification indicators corresponding to each demand level feature. Simultaneously, the computer device performs partitioned calorific value statistical processing on the full amount of waste calorific value data. For each grid unit, it statistically analyzes its average calorific value, maximum calorific value, minimum calorific value, and other data within a preset time period, obtaining the partitioned calorific value statistical results for each grid unit.

[0040] Step 144: According to the preset evaluation rules, the processing timeliness characteristics and processing quality characteristics are quantified into the first demand evaluation index; the capacity utilization rate characteristics and capacity dynamic adjustment characteristics are quantified into the second demand evaluation index; the transfer time sequence characteristics and transfer connection characteristics are quantified into the third demand evaluation index; the overall calorific value level of the total waste calorific value data, the zonal calorific value statistical results of each grid unit, and the time sequence change characteristics are quantified into the corresponding first calorific value evaluation index, second calorific value evaluation index, and third calorific value evaluation index; establish the quantitative correlation between the first demand evaluation index and the first calorific value evaluation index, the second demand evaluation index and the second calorific value evaluation index, and the third demand evaluation index and the third calorific value evaluation index, and generate the demand calorific value matching relationship.

[0041] In this embodiment, the computer equipment quantifies different types of features according to preset evaluation rules: processing timeliness and processing quality features are weighted according to preset weights to obtain a first demand evaluation index; capacity utilization and capacity dynamic adjustment features are weighted according to preset weights to obtain a second demand evaluation index; and transfer timing and transfer connection features are weighted according to preset weights to obtain a third demand evaluation index. For the total calorific value data of waste, the overall calorific value level is quantified as the first calorific value evaluation index, the zonal calorific value statistical results of each grid unit are quantified as the second calorific value evaluation index, and the temporal variation characteristics of calorific value are quantified as the third calorific value evaluation index. Specifically, the computer equipment analyzes the correlation between different evaluation indicators to establish a quantitative correlation between demand evaluation indicators and calorific value evaluation indicators, generating a demand-calorific value matching relationship.

[0042] Step 145: Extract demand heat value association pairs that meet the preset conditions based on the demand heat value matching relationship, transform the demand heat value association pairs into objective function constraints for reinforcement learning, and determine the reinforcement learning objective by combining the task setting keywords of the intelligent simulation task.

[0043] In this embodiment, the computer device first sets a preset condition for the matching degree, such as a correlation coefficient greater than a preset threshold. Then, the computer device extracts demand heat value association pairs that meet the preset condition from the demand heat value matching relationship. Furthermore, the computer device transforms the above demand heat value association pairs into objective function constraints for reinforcement learning, for example, transforming "the first demand evaluation index is positively correlated with the first heat value evaluation index" into a constraint term in the objective function. Then, the computer device combines the task setting keywords of the intelligent simulation task, such as "maximize heat value" and "optimize transportation efficiency," to determine the reinforcement learning objective. This objective clarifies the direction and constraints that need to be optimized in the intelligent simulation task.

[0044] Step 150: Based on the reinforcement learning objective, determine the dynamic trajectory simulation information of waste in each grid cell through partitioned dynamic identification, and simultaneously perform state encoding processing on the real-time simulation operation data of the target waste bin to generate state simulation encoding information, and generate multi-source simulation information after aligning the state simulation encoding information with the dynamic trajectory simulation information.

[0045] In this embodiment, the computer device first sets relevant parameters for dynamic identification of waste zones based on the requirements for dynamic waste tracking in the reinforcement learning objective, including identification area division rules and identification frequency. Then, the computer device tracks and simulates the movement trajectory of waste in each grid cell using dynamic identification technology to obtain dynamic trajectory simulation information. Simultaneously, the computer device performs state encoding processing on the real-time simulation operation data of the target waste bin, converting different types of operation data into standardized encoded information to generate state simulation encoded information. Then, the computer device aligns the dynamic trajectory simulation information and the state simulation encoded information based on time to generate aligned multi-source simulation information.

[0046] Step 151: Based on the accuracy requirements of dynamic tracking of waste in the reinforcement learning objective, set the dynamic recognition range of the partition, divide the grid unit of the target waste bin into a core recognition area and an extended recognition area, and the recognition frequency of the core recognition area is higher than that of the extended recognition area.

[0047] In this embodiment, the computer device first extracts the accuracy requirements for dynamic waste tracking from the reinforcement learning objective, including the allowable range of positional error for trajectory tracking and the required time interval for trajectory updates. Then, based on these accuracy requirements, the computer device sets the dynamic recognition range for different zones. Grid cells with more frequent dynamic changes in waste and a greater impact on the reinforcement learning objective are designated as core recognition regions, while grid cells with relatively less dynamic changes and a smaller impact on the reinforcement learning objective are designated as extended recognition regions. Simultaneously, the computer device sets the recognition frequency for the core recognition regions to be higher than that for the extended recognition regions, thereby improving recognition efficiency while ensuring tracking accuracy.

[0048] Step 152: Set the iterative identification termination condition to the overlap of the waste trajectory between two adjacent identifications reaching a preset standard. During the first identification process, collect the waste movement process of each grid unit in the core identification area in real time to generate an initial trajectory point set.

[0049] In this embodiment, the computer device first sets an iterative identification termination condition, namely, the overlap of the waste trajectories obtained from two adjacent identifications reaches a preset percentage standard. During the initial identification process, the computer device monitors the movement of waste in each grid unit within the core identification area in real time through monitoring equipment inside the waste bin, collects the location information of waste at different time points, and generates an initial set of trajectory points, each trajectory point containing location information and time information.

[0050] Step 153: Perform trajectory fitting processing based on the initial trajectory point set to generate an initial garbage trajectory. Extend the initial garbage trajectory to the extended recognition area and collect trajectory points in the extended recognition area to supplement the initial garbage trajectory, thus obtaining the initial complete garbage trajectory.

[0051] In this embodiment, the computer device uses a preset trajectory fitting algorithm to process the initial set of trajectory points and fit it to generate an initial waste trajectory. Then, the computer device extends the initial waste trajectory to an extended identification area, and collects trajectory points within the extended identification area through monitoring equipment to supplement the initial waste trajectory, ensuring that the waste trajectory covers the relevant area of ​​the entire target waste bin, thus obtaining an initial complete waste trajectory.

[0052] Step 154: Proceed to the next iteration of identification, re-collect trajectory points in the core identification area and the extended identification area, generate a new complete garbage trajectory, and calculate the overlap between the new complete garbage trajectory and the initial complete garbage trajectory.

[0053] In this embodiment, the computer device enters the next iteration of the identification process according to a set identification frequency. During this process, the computer device re-collects trajectory points in the core identification area and the extended identification area, and uses the same trajectory fitting and supplementation method as the first identification to generate a new complete waste trajectory. Specifically, the computer device calculates the overlap degree between the new complete waste trajectory and the initial complete waste trajectory. The overlap degree calculation is based on the positional matching and temporal matching of the trajectory points.

[0054] Step 155: If the overlap does not reach the termination condition of iterative identification, the new complete waste trajectory is used as the initial complete waste trajectory to jump to the steps of trajectory acquisition, supplementation and overlap calculation until the overlap reaches the termination condition of iterative identification. The final complete waste trajectory is then used as dynamic trajectory simulation information.

[0055] In this embodiment, the computer device compares the calculated overlap with a preset iterative identification termination condition. If the overlap does not reach the termination condition, the new complete waste trajectory is used as the initial complete waste trajectory, and the steps of trajectory acquisition, supplementation, and overlap calculation are repeated. If the overlap reaches the termination condition, the final complete waste trajectory is determined as dynamic trajectory simulation information, which includes the complete movement trajectory and time information of the waste within the target waste bin.

[0056] Step 156: During the trajectory tracking process, real-time simulation operation data of the target waste bin is collected synchronously. The real-time simulation operation data includes the real-time accumulation amount of waste, real-time temperature of waste, and real-time humidity of waste in each grid unit.

[0057] In this embodiment, while tracking the waste trajectory, the computer device simultaneously collects real-time simulation data via sensor devices within the waste bin. This data includes the real-time accumulation of waste in each grid cell (i.e., the volume or weight of waste within the grid cell per unit time); the real-time temperature of the waste (i.e., the average temperature of the waste within the grid cell); and the real-time humidity of the waste (i.e., the moisture content of the waste within the grid cell). The collected data can be transmitted to the computer device in real time.

[0058] Step 157: Perform state coding processing on the real-time simulation operation data, converting the real-time waste accumulation amount, real-time waste temperature, and real-time waste humidity into standardized coding vectors to generate state simulation coding information.

[0059] In this embodiment, the computer device first preprocesses the collected real-time simulation data, including data cleaning and standardization. Then, using preset encoding rules, the computer device converts the real-time waste accumulation amount, real-time waste temperature, and real-time waste humidity into standardized encoding vectors, with each data type corresponding to a specific dimension in the encoding vector. For example, the real-time waste accumulation amount corresponds to the first dimension of the encoding vector, the real-time waste temperature corresponds to the second dimension, and the real-time waste humidity corresponds to the third dimension. Further, state simulation encoding information is generated and stored in the form of encoding vectors.

[0060] Step 158: Extract the trajectory acquisition timestamp from the dynamic trajectory simulation information and the data acquisition timestamp from the state simulation coding information. Use the timestamp as a reference to perform time-series alignment processing on the dynamic trajectory simulation information and the state simulation coding information, delete information with mismatched timestamps, and obtain aligned multi-source simulation information.

[0061] In this embodiment, the computer device first extracts the timestamp of each trajectory point from the dynamic trajectory simulation information and the timestamp of each encoding vector from the state simulation encoding information. Then, using the timestamps as a reference, the computer device performs time-series alignment processing on the dynamic trajectory simulation information and the state simulation encoding information, associating the trajectory information and encoding information corresponding to the same timestamp. For information with mismatched timestamps, the computer device deletes it to ensure that in the final aligned multi-source simulation information, each piece of information has a corresponding timestamp, and the timestamps are consistent.

[0062] Step 160: Use the aligned multi-source simulation information as input to the initial reinforcement learning model, call the preset policy optimization algorithm to train the initial reinforcement learning model, generate the trained reinforcement learning model, input the aligned multi-source measured information into the trained reinforcement learning model, and obtain the intelligent decision-making operation instructions for the target waste bin under the preset heat value enhancement label.

[0063] In this embodiment, the computer device first constructs an initial reinforcement learning model, which includes a state input layer, a policy network layer, and a value network layer. Then, the computer device uses aligned multi-source simulation information as input to the initial reinforcement learning model and calls a preset policy optimization algorithm, such as the deep deterministic policy gradient algorithm, to train the initial reinforcement learning model. During training, the computer device continuously adjusts the model parameters to optimize the model's policy output and value evaluation capabilities. When the model training reaches a preset stopping condition, a post-trained reinforcement learning model is generated. Then, the computer device inputs aligned multi-source measured information into the post-trained reinforcement learning model. Based on the input information and preset heat value enhancement labels, the model outputs intelligent decision-making operation instructions for the target waste bin, which include decision content related to waste transfer and waste accumulation adjustment.

[0064] Step 161: Perform feature standardization processing on the aligned multi-source simulation information to eliminate the dimensional differences between dynamic trajectory simulation information and state simulation coding information, generate standardized multi-source simulation information, and divide the standardized multi-source simulation information into training dataset and validation dataset according to a preset ratio. The training dataset is used for iterative updating of model parameters, and the validation dataset is used for evaluating the model training effect.

[0065] In this embodiment, the computer device first performs feature standardization processing on the dynamic trajectory simulation information and state simulation coding information in the aligned multi-source simulation information. The standardization method used is to convert the feature values ​​into values ​​with a mean of 0 and a standard deviation of 1 to eliminate the dimensional differences between different information. After processing, standardized multi-source simulation information is generated. Furthermore, the computer device divides the standardized multi-source simulation information into a training dataset and a validation dataset according to a preset ratio, such as 7:3. The training dataset is used for iterative updates of the model's parameters, and the validation dataset is used to evaluate the model's training effect during the model training process.

[0066] Step 162: Initialize the network parameters of the initial reinforcement learning model. The initial reinforcement learning model includes a feature extraction layer, a policy decision layer, and a value evaluation layer. The feature extraction layer is used to extract temporal correlation features and state correlation features from standardized multi-source simulation information. The policy decision layer is used to output the model training policy based on the current input features. The value evaluation layer is used to calculate the value estimate corresponding to the training policy.

[0067] In this embodiment, the computer device first constructs the network structure of an initial reinforcement learning model. The feature extraction layer uses a convolutional neural network structure to extract temporal correlation features and state correlation features from standardized multi-source simulation information. The policy decision layer uses a fully connected neural network structure to output the model training policy based on the extracted features. The value evaluation layer also uses a fully connected neural network structure to calculate the value estimate corresponding to the training policy. Then, the computer device initializes the network parameters of the initial reinforcement learning model using a random initialization method to ensure that the model has a certain degree of randomness at the beginning of training.

[0068] Step 163: Call the preset policy optimization algorithm, set the reward function of the algorithm based on the reinforcement learning objective, the reward value of the reward function is positively correlated with the degree of fit of the model output to meet the preset heat value improvement requirements, incorporate the constraint conditions corresponding to the inter-grid correlation features into the penalty term of the reward function, and trigger the penalty mechanism when feature overflow or parameters deviate from the preset range during model training.

[0069] In this embodiment, the computer device invokes a preset policy optimization algorithm, which is a deep reinforcement learning algorithm. Then, the computer device sets a reward function based on the reinforcement learning objective. The input to the reward function is the degree of fit between the model output and the preset heat value improvement requirement; the higher the fit, the larger the reward value. Simultaneously, the computer device incorporates the constraints corresponding to the inter-grid correlation features into the penalty term of the reward function. When feature overflow occurs during model training—that is, when feature values ​​exceed a preset range or parameters deviate from a preset range—a penalty mechanism is triggered, reducing the reward value.

[0070] Step 164: Input the training dataset into the feature extraction layer of the initial reinforcement learning model, extract the convolutional features from the standardized multi-source simulation information through convolution operation, generate the simulation feature vector, input the simulation feature vector into the policy decision layer, so that the policy decision layer outputs the initial training policy based on the preset policy distribution, and obtain the initial value evaluation result by performing value estimation on the initial training policy through the value evaluation layer.

[0071] In this embodiment, the computer device inputs the training dataset into the feature extraction layer of the initial reinforcement learning model. The feature extraction layer processes the standardized multi-source simulation information through convolution operations to extract convolutional features and generate simulation feature vectors. Then, the computer device inputs the simulation feature vectors into the policy decision layer. Based on a preset policy distribution, such as a Gaussian distribution, the policy decision layer outputs an initial training policy. Finally, the computer device performs a value estimation on the initial training policy through a value evaluation layer to obtain an initial value evaluation result, which reflects the expected value of the initial training policy.

[0072] Step 165: Calculate the reward value corresponding to the initial training policy based on the reward function, update the policy decision layer parameters of the initial reinforcement learning model using the policy gradient descent method in combination with the initial value evaluation results, and simultaneously adjust the convolution kernel parameters of the feature extraction layer to optimize the convolution feature extraction effect, thus completing one iteration of model training.

[0073] In this embodiment, the computer device first calculates the reward value corresponding to the initial training policy based on the reward function. Then, combining the initial value evaluation results, the computer device updates the policy decision layer parameters of the initial reinforcement learning model using the policy gradient descent method, with the parameter update direction maximizing the reward value. Simultaneously, the computer device adjusts the convolution kernel parameters of the feature extraction layer to optimize the convolutional feature extraction effect. After the parameter update is completed, one iteration of model training is finished.

[0074] Step 166: Repeatedly execute the steps of training data input, feature extraction, policy output, value evaluation and parameter update. After each preset number of iterations of training, input the validation dataset into the reinforcement learning model obtained in the current iteration, and calculate the error value between the model output result and the corresponding labeled result of the validation dataset.

[0075] In this embodiment, the computer device iteratively executes the steps of training data input, feature extraction, policy output, value evaluation, and parameter update, using a different subset of training data in each iteration. After completing a preset number of iterations, the computer device inputs the validation dataset into the reinforcement learning model obtained in the current iteration, and the model outputs results. Then, the computer device calculates the error value between the model output and the labeled results corresponding to the validation dataset, using the mean squared error method.

[0076] Step 167: Determine whether the error value is less than the preset error threshold. If the error value is greater than or equal to the preset error threshold, continue to execute the iterative training steps and dynamically adjust the learning rate of the policy gradient descent. If the error value is less than the preset error threshold, stop the iterative training, determine the reinforcement learning model obtained in the current iteration as the post-training reinforcement learning model, and save the network parameters of the post-training reinforcement learning model and the optimal policy parameters during the training process.

[0077] In this embodiment, the computer device compares the calculated error value with a preset error threshold. If the error value is greater than or equal to the preset error threshold, the iterative training steps continue, and the learning rate of the policy gradient descent is dynamically adjusted. The learning rate adjustment adopts an adaptive learning rate adjustment method, so that the learning rate changes as the training process progresses. If the error value is less than the preset error threshold, the iterative training stops, and the reinforcement learning model obtained in the current iteration is determined as the post-trained reinforcement learning model. At the same time, the computer device saves the network parameters of the post-trained reinforcement learning model and the optimal policy parameters during the training process. The optimal policy parameters are the policy parameters that maximize the reward value during the training process.

[0078] Step 168: Input the aligned multi-source measured information into the trained reinforcement learning model to obtain intelligent decision-making operation instructions for the target waste bin under the preset heat value enhancement label.

[0079] In this embodiment, the computer device first acquires aligned multi-source measured information, which is multi-source data from the actual operation of the target waste storage facility, obtained through the same processing flow as the simulation information. Then, the computer device inputs the aligned multi-source measured information into a trained reinforcement learning model. Based on the input information and preset heat values ​​to enhance labels, the model outputs intelligent decision-making operation instructions for the target waste storage facility. These instructions include the direction of waste transfer, the method of adjusting waste accumulation, and other related information.

[0080] Step 1681: Simultaneously input the measured dynamic trajectory information and the measured state encoding information into the feature extraction layer of the reinforcement learning model after training. The feature extraction layer performs feature fusion processing on the measured dynamic trajectory information and the measured state encoding information based on the convolution kernel parameters optimized during training. It also strengthens the key temporal node features in the measured dynamic trajectory information through a temporal attention mechanism to generate a measured fusion feature vector.

[0081] In this embodiment, the computer device synchronously inputs the measured dynamic trajectory information and the measured state encoding information into the feature extraction layer of the trained reinforcement learning model. Based on the optimized convolutional kernel parameters during training, the feature extraction layer performs convolution operations on the measured dynamic trajectory information and the measured state encoding information to extract features. Then, the feature extraction layer fuses the extracted features using feature concatenation. Furthermore, the feature extraction layer strengthens the key temporal node features in the measured dynamic trajectory information through a temporal attention mechanism, assigning higher weights to these key temporal nodes. Finally, a measured fused feature vector is generated, which contains the fused features of the measured dynamic trajectory information and the measured state encoding information, as well as the strengthened key temporal node features.

[0082] Step 1682: Call the policy decision layer of the post-trained reinforcement learning model to perform feature parsing on the measured fused feature vector and extract target features associated with the preset heat value enhancement label. The target features include the garbage dynamic accumulation rate feature of the grid cell, the heat value potential feature corresponding to the real-time state code, and the garbage migration association feature between grids.

[0083] In this embodiment, the computer device invokes the policy decision layer of the trained reinforcement learning model to perform feature parsing on the measured fused feature vector. The policy decision layer processes the measured fused feature vector through fully connected operations to extract target features associated with preset heat value enhancement labels. Target features include the dynamic accumulation rate of waste in the grid cell, i.e., the rate of change of the amount of waste accumulation in the grid cell per unit time; the heat value potential feature corresponding to the real-time state code, i.e., the heat value enhancement potential of waste predicted based on the real-time state code; and the waste migration correlation feature between grids, i.e., the degree of correlation of waste migration between different grid cells.

[0084] Step 1683: Generate a feature matching matrix based on the matching rules between the target features and the preset heat value enhancement labels. Calculate the matching degree between the target features and the preset heat value enhancement labels using the feature matching matrix, and select the target feature combination with the highest matching degree as the decision basis.

[0085] In this embodiment, the computer device first sets matching rules between target features and preset heat value enhancement labels. These matching rules are based on the correlation between the target features and heat value enhancement. Then, the computer device generates a feature matching matrix according to the matching rules, where each element represents the degree of matching between the target features and the preset heat value enhancement labels. Furthermore, the computer device calculates the matching degree between the target features and the preset heat value enhancement labels using the feature matching matrix. The matching degree is calculated based on a weighted summation of the matrix elements. Finally, the computer device selects the target feature combination with the highest matching degree as the decision-making basis.

[0086] Step 1684: Combine the decision-making basis with the optimal strategy parameters saved during the training process to generate an initial decision scheme. The initial decision scheme includes suggestions for waste transfer priority, waste accumulation adjustment parameters, and real-time monitoring frequency adjustment for each grid unit.

[0087] In this embodiment, the computer device, based on the decision-making criteria, calls upon the optimal strategy parameters saved during the training process. These optimal strategy parameters include parameters related to waste transfer, waste accumulation adjustment, and real-time monitoring frequency adjustment. Then, the computer device generates an initial decision plan based on the decision-making criteria and the optimal strategy parameters. The initial decision plan includes: waste transfer priority for each grid unit (grid units with higher priority are prioritized for waste transfer); waste accumulation adjustment parameters, including adjustment values ​​for waste accumulation height and density; and real-time monitoring frequency adjustment suggestions, including the adjustment range and timing for monitoring frequencies in different grid units.

[0088] Step 16841: Perform multi-dimensional structured analysis on the target feature combination corresponding to the decision basis, decompose the feature dimension of the garbage dynamic accumulation rate of the grid cell, the feature dimension of the heat value potential corresponding to the real-time status code, and the feature dimension of garbage migration correlation between grids, extract the feature weight distribution, feature time series change law and feature quantization value under each feature dimension, and integrate them to obtain the feature analysis matrix of the decision basis.

[0089] In this embodiment, the computer device performs multi-dimensional structured analysis on the target feature combination corresponding to the decision-making basis. First, it decomposes the feature dimension of the grid cell's dynamic garbage accumulation rate, the feature dimension of the heat potential corresponding to the real-time state code, and the feature dimension of garbage migration correlation between grids. Then, for each feature dimension, it extracts the feature weight distribution (i.e., the weight of each sub-feature under that feature dimension), the feature temporal change pattern (i.e., how the feature changes over time), and the feature quantization value (i.e., the specific numerical value of the feature). Finally, it integrates the extracted information to obtain the decision-making basis feature analysis matrix, where the rows of the matrix represent feature dimensions and the columns represent feature attributes.

[0090] Step 16842: Retrieve the optimal policy parameters saved during training, determine the parameter type, applicable range, and adjustment threshold of the optimal policy parameters, and generate an optimal policy parameter index library containing parameter indexes, scene labels, and adjustment rules. The scene labels and the feature dimensions and feature weight distribution of the decision basis form a one-to-one correspondence.

[0091] In this embodiment, the computer device retrieves the optimal strategy parameters saved during the training process. Then, it determines the parameter type of the optimal strategy parameters, such as waste transfer parameters or waste accumulation adjustment parameters; the applicable scope of the parameters, i.e., which grid cells or scenarios the parameters are applicable to; and the parameter adjustment threshold, i.e., the adjustment range of the parameters. Furthermore, the computer device generates an optimal strategy parameter index library, which includes parameter indexes, scenario labels, and adjustment rules. The scenario labels and the feature dimensions and feature weight distributions of the decision-making basis form a one-to-one correspondence, facilitating the rapid retrieval of the corresponding optimal strategy parameters based on the decision-making basis.

[0092] Step 16843: Using the feature dimensions, feature weight distribution, and feature temporal change patterns in the decision-making feature analysis matrix as retrieval conditions, perform a layer-by-layer matching query in the optimal strategy parameter index library, filter out the optimal strategy parameters with a matching degree higher than the feature matching degree threshold, and generate a subset of optimal strategy parameters that are adapted to the current decision-making scenario.

[0093] In this embodiment, the computer device uses the feature dimensions, feature weight distribution, and temporal variation patterns of the decision-making feature analysis matrix as retrieval conditions, and performs a layer-by-layer matching query in the optimal strategy parameter index. First, a preliminary matching is performed based on the feature dimensions to filter out the optimal strategy parameters that match the feature dimensions. Then, a secondary matching is performed based on the feature weight distribution to filter out the optimal strategy parameters that match the feature weight distribution. Finally, a tertiary matching is performed based on the temporal variation patterns of the features to filter out the optimal strategy parameters that match the temporal variation patterns. During the filtering process, a feature matching degree threshold is set; only optimal strategy parameters with a matching degree higher than this threshold are retained, generating a subset of optimal strategy parameters adapted to the current decision-making scenario.

[0094] Step 16844: Construct a decision parameter calculation model based on the subset of optimal strategy parameters; normalize the quantified values ​​of each feature in the feature analysis matrix of decision basis to obtain dimensionless feature scores; determine the mapping calculation relationship between each optimal strategy parameter and the normalized feature scores; input the normalized feature scores into the decision parameter calculation model, and obtain the initial values ​​of waste transfer volume, stacking height adjustment, and real-time monitoring frequency corresponding to each grid unit through weighted calculation.

[0095] In this embodiment, the computer device constructs a decision parameter calculation model based on a subset of optimal strategy parameters. The model's input is a feature score, and its output is the initial value of the decision parameters. Then, the computer device normalizes the quantized values ​​of each feature in the decision basis feature analysis matrix by converting them into values ​​between 0 and 1, resulting in a dimensionless feature score. Furthermore, the computer device determines the mapping relationship between each optimal strategy parameter and the normalized feature score. This mapping relationship is based on the parameter type of the optimal strategy parameters and the meaning of the feature score. Finally, the computer device inputs the normalized feature score into the decision parameter calculation model and, through weighted calculation, obtains the initial values ​​of waste transfer volume, stacking height adjustment, and real-time monitoring frequency for each grid unit.

[0096] Step 16845: Combining the grid cell topology of the target waste bin with the parameter coordination patterns between grids in historical operation, the initial values ​​of decision parameters for each grid cell are coordinated and adapted. When there is a conflict between the initial values ​​of decision parameters of adjacent grid cells, the initial values ​​are recalculated and adjusted based on the control rules of the optimal strategy parameter subset and the waste migration correlation characteristics between grids, so as to achieve coordinated rationalization of the initial values ​​of decision parameters for each grid cell.

[0097] In this embodiment, the computer device combines the grid cell topology of the target waste storage area, i.e., the positional and connectivity relationships between grid cells, and the historical operational patterns of inter-grid parameter coordination, to perform collaborative adaptation processing on the initial values ​​of decision parameters for each grid cell. When there are conflicts in the initial values ​​of decision parameters between adjacent grid cells, such as conflicts in waste transfer directions, the computer device recalculates and adjusts the initial values ​​of decision parameters based on the control rules of the optimal strategy parameter subset and the inter-grid waste migration correlation characteristics. The principle of adjustment is to make the initial values ​​of decision parameters of each grid cell collaborative and reasonable, avoiding conflicts.

[0098] Step 16846: Based on the decision types of waste transfer, accumulation adjustment, and real-time monitoring, classify and integrate the collaboratively adapted decision parameters to generate a waste transfer priority ranking table, a waste accumulation adjustment parameter set, and a real-time monitoring frequency adjustment list. The waste transfer priority ranking table is determined based on the dynamic accumulation rate characteristics and calorific value potential characteristics of each grid unit. The waste accumulation adjustment parameter set is associated with the accumulation height limit threshold of each grid unit. The real-time monitoring frequency adjustment list is associated with the waste status change frequency of each grid unit.

[0099] In this embodiment, the computer device categorizes and integrates the collaboratively adapted decision parameters according to the decision type. For the waste transfer decision type, a waste transfer priority ranking table is generated. The ranking table is determined based on the dynamic accumulation rate characteristics and calorific value potential characteristics of each grid unit. Grid units with high dynamic accumulation rates and high calorific value potential have higher priority. For the accumulation adjustment decision type, a waste accumulation adjustment parameter set is generated. This set is associated with the accumulation height limit threshold of each grid unit to ensure that the height after accumulation adjustment does not exceed the limit threshold. For the real-time monitoring decision type, a real-time monitoring frequency adjustment list is generated. This list is associated with the waste status change frequency of each grid unit. Grid units with high waste status change frequencies have higher monitoring frequencies.

[0100] Step 16847: Construct a decision instruction tag set, which includes decision target tags, decision object tags, decision parameter tags, and time sequence execution tags. Fill the corresponding tags with the classified and integrated waste transfer priority ranking table, waste accumulation adjustment parameter set, and real-time monitoring frequency adjustment list. Supplement the execution trigger conditions, execution duration, and correlation logic between tags for each decision parameter to generate an initial decision scheme.

[0101] In this embodiment, the computer device constructs a decision instruction tag set, which includes decision target tags, decision object tags, decision parameter tags, and timing execution tags. Then, the classified and integrated waste transfer priority ranking table, waste accumulation adjustment parameter set, and real-time monitoring frequency adjustment list are filled into the corresponding tags. Furthermore, the computer device supplements the execution trigger conditions for each decision parameter, such as triggering waste transfer when the waste accumulation height reaches a certain value; the execution duration, such as the duration of waste transfer; and the association logic between tags, such as the association relationship between decision target tags and decision parameter tags. Then, an initial decision scheme is generated, which contains all decision-related information.

[0102] Step 1685: Call the value evaluation layer of the post-trained reinforcement learning model to verify the value of the initial decision scheme, calculate the probability that the total calorific value data of the target waste bin reaches the preset calorific value improvement standard after the implementation of the initial decision scheme, and determine whether the probability is greater than the preset probability threshold.

[0103] In this embodiment, the computer device invokes the value evaluation layer of the trained reinforcement learning model to verify the value of the initial decision-making scheme. Based on the initial decision-making scheme and relevant data from the target waste bin, the value evaluation layer calculates the probability that the total calorific value of the waste in the target waste bin will reach a preset calorific value improvement standard after the implementation of the initial decision-making scheme. Then, the computer device compares the calculated probability with a preset probability threshold to determine whether the probability is greater than the preset probability threshold.

[0104] Step 1686: If the probability is less than or equal to the preset probability threshold, the decision parameters are adjusted by the strategy decision layer based on the value assessment results, the decision scheme is regenerated and the value is verified again until the success probability of the generated decision scheme is greater than the preset probability threshold.

[0105] In this embodiment, if the calculated probability is less than or equal to a preset probability threshold, the computer device adjusts the decision parameters based on the value assessment results through the strategy decision layer. The direction of the decision parameter adjustment is to increase the probability of achieving the target. After the adjustment is completed, the computer device regenerates the decision scheme and performs value verification again. This process is repeated until the probability of achieving the target corresponding to the generated decision scheme is greater than the preset probability threshold.

[0106] Step 1687: If the probability is greater than the preset probability threshold, the current decision scheme is transformed into a standardized intelligent decision operation instruction. The intelligent decision operation instruction includes the instruction execution object, execution parameters, execution sequence and execution priority. The execution object is the specific grid unit of the target waste bin, the execution parameters are quantitative indicators including at least the waste transfer volume and the stacking height adjustment value, and the execution sequence is determined based on the key nodes of the dynamic trajectory measured information.

[0107] In this embodiment, if the calculated probability is greater than a preset probability threshold, the computer device converts the current decision into a standardized intelligent decision-making operation instruction. The intelligent decision-making operation instruction includes the instruction execution object, i.e., the specific grid unit of the target waste bin; execution parameters, including quantitative indicators such as waste transfer volume and accumulation height adjustment value; execution timing, determined based on key nodes in the dynamic trajectory measurement information, such as the time point when waste accumulation reaches a certain height; and execution priority, determined according to the importance of the decision. The standardized intelligent decision-making operation instruction facilitates execution and understanding by the target waste bin's management system.

[0108] Optionally, the method further includes: Step 170: Extract multi-dimensional target features corresponding to intelligent decision-making operation instructions. These features include calorific value enhancement, operational efficiency, and resource consumption. Establish a correlation mapping relationship between these multi-dimensional target features and identify constraint association rules between different target features. Based on the correlation mapping relationship and constraint association rules, perform multi-objective collaborative trade-off processing on the intelligent decision-making operation instructions to generate a weight allocation scheme for each target feature. Combining the weight allocation scheme, conduct a multi-objective comprehensive evaluation and optimization of the predicted effects of the calorific value enhancement, operational efficiency, and resource consumption targets corresponding to the intelligent decision-making operation instructions. Adjust the parameters of the generated decision instructions based on the comprehensive evaluation results to obtain multi-objective optimized decision instructions. Match the achievement degree of each target feature corresponding to the multi-objective optimized decision instructions with preset target thresholds to select decision instructions whose achievement degree of all target features meets the preset target thresholds. Perform a compatibility analysis between the selected decision instructions and the grid cell topology features of the target waste bins, output multi-objective optimized decision instructions that meet the compatibility requirements, and simultaneously generate multi-objective trade-off feature records for model optimization.

[0109] In this embodiment, the computer device first extracts multi-dimensional target features corresponding to the intelligent decision-making operation instructions. These features cover calorific value improvement target features, operational efficiency target features, and resource consumption target features. Then, the computer device establishes an association mapping relationship between the multi-dimensional target features and identifies constraint association rules by analyzing the mutual influence between different target features. Furthermore, based on the association mapping relationship and constraint association rules, the intelligent decision-making operation instructions undergo multi-objective collaborative trade-off processing to generate a weight allocation scheme corresponding to each target feature. The weight allocation scheme is determined according to the importance of the target. Specifically, the computer device, in conjunction with the weight allocation scheme, performs multi-objective comprehensive evaluation and optimization of the effect prediction values ​​of each target corresponding to the intelligent decision-making operation instructions. Evaluation indicators include the achievement degree of each target and the degree of synergy between targets. Based on the comprehensive evaluation results, the parameters of the decision instructions are adjusted to obtain multi-objective optimized decision instructions. Afterward, the computer device matches the achievement degree of each target feature corresponding to the multi-objective optimized decision instructions with preset target thresholds, and filters out decision instructions whose achievement degree of all target features meets the preset target thresholds. Then, the computer equipment performs a compatibility analysis between the filtered decision instructions and the topological characteristics of the grid cells of the target waste bin. The analysis includes whether the decision instructions are suitable for the layout of the grid cells and whether they will affect the coordination between the grid cells. The computer outputs multi-objective optimization decision instructions that meet the compatibility requirements and generates multi-objective trade-off feature records for model optimization.

[0110] Optionally, the method further includes: Step 180: Obtain multi-source feedback features during the execution of intelligent decision-making operation instructions. These features include dynamic changes in the waste status of each grid cell, real-time fluctuations in heat value, and response characteristics to the execution of operation instructions. The acquisition sequence of these multi-source feedback features is synchronized with the execution sequence of the intelligent decision-making operation instructions. The multi-source feedback features are then compared with preset decision-making effect benchmark features to extract deviation correlation information between the feedback features and the benchmark features. This deviation correlation information includes the feature deviation dimension, deviation temporal distribution, and deviation correlation strength. Based on the deviation correlation information, the input features of the trained reinforcement learning model are dynamically completed. The completed features include deviation correction features. The system incorporates feedback-related features to ensure that the completed input features match the feature input dimensions of the trained reinforcement learning model. The completed input features are then fed into the trained reinforcement learning model, triggering an iterative learning process that updates the model's policy decision layer and value evaluation layer parameters, generating an iterative learning model adapted to the current feedback state. This iterative learning model is then used to dynamically optimize intelligent decision-making operation instructions, generating optimized instructions that inherit the execution object relationships from the original instructions. Only the execution parameters and timing parameters are adjusted to match the deviation correction requirements of the multi-source feedback features.

[0111] In this embodiment, the computer device first acquires multi-source feedback features during the execution of intelligent decision-making operation instructions. These features include dynamic changes in the waste status of each grid cell, real-time fluctuations in heat value, and operation instruction execution response features. The acquisition timing is synchronized with the execution timing to ensure that the feedback features can reflect the instruction execution effect in a timely manner. Then, the computer device compares the multi-source feedback features with preset decision effect benchmark features. The benchmark features are feature values ​​under ideal conditions. During the comparison, deviation correlation information between the feedback features and the benchmark features is extracted, including feature deviation dimension, deviation temporal distribution, and deviation correlation strength. Furthermore, the computer device dynamically completes the input features of the trained reinforcement learning model based on the deviation correlation information. The completed features include deviation correction features and feedback correlation features, ensuring that the completed input features match the feature input dimension of the model. Specifically, the computer device inputs the completed input features into the trained reinforcement learning model, triggering the model's iterative learning process, updating the model's policy decision layer parameters and value evaluation layer parameters, and generating an iterative learning model adapted to the current feedback state. Then, the computer equipment uses an iterative learning model to dynamically optimize the intelligent decision-making operation instructions. The optimized instructions inherit the execution object association of the original instructions and only adjust the execution parameters and timing parameters, so that the adjusted parameters match the deviation correction requirements of the multi-source feedback features, thereby improving the execution effect of the instructions.

[0112] In a specific application scenario, a municipal solid waste incineration power plant faces long-standing challenges, including insufficient quantitative accuracy, difficulty in refined management due to insufficient waste supply, and a lack of optimization methods. The plant's waste storage area is divided into functional zones such as a stockpiling zone, fermentation zone, feeding zone, and industrial waste co-firing zone. Operators rely on experience to move waste using a garbage crane. Within the same zone, the fermentation time varies significantly between different locations; waste at the bottom may have fermented for 6 days, while waste at the top may have only fermented for 3 days. It is impossible to accurately grasp the fermentation status at each spatial coordinate, and when waste supply is insufficient, it is difficult to ensure sufficient fermentation of the limited amount of waste entering the furnace, resulting in unstable calorific value and severely impacting power generation efficiency. To address these issues, the plant introduced a reinforcement learning-based intelligent optimization system for the waste storage area, achieving refined management through a two-stage process of offline learning and online application.

[0113] During the offline learning phase, the system (which can be understood as the aforementioned computer equipment) first performs spatial gridding on the waste storage area, dividing the entire waste storage area into several grid units of 1m×1m×1m. Each grid unit records information such as filling rate, average age of waste, calorific value, and waste type, achieving a breakthrough from traditional "district-level" management to "grid-level" precision, while also considering the fermentation differences at different vertical heights. Furthermore, the system learns unloading, feeding, and state transition models from historical data of the waste storage area, analyzing patterns such as the unloading time, weight, unloading gate usage, and waste type of sanitation vehicles, mastering patterns such as feeding frequency, weight requirements, and feeding port usage, and clarifying the state change logic of waste transfer between grids. Finally, the system obtains waste calorific value information through a patented method.

[0114] The second part of offline learning is reinforcement learning training. The system defines a state space containing information such as waste distribution status, unloading queues, feeding queues, garbage crane location and load, waste sufficiency, and the functions of each zone. Reward functions related to feeding contribution, stockpiling contribution, and dumping contribution are designed. Feeding contribution considers calorific value, stability, and timeliness; stockpiling contribution focuses on cleaning timeliness and space utilization rationality; and dumping contribution emphasizes optimizing the dumping time interval. The system uses the SoftActor-Critic algorithm for policy optimization, approximating the optimal policy through alternating policy evaluation and improvement. In policy evaluation, the state-action value function and state value function are updated; in policy improvement, the policy that maximizes the expected value is sought. Entropy regularization is used to maintain exploratory nature, ensuring that the optimal operation can be found even in complex scenarios such as insufficient waste supply.

[0115] Upon entering the online application phase, the system first performs dynamic zone identification of waste trajectories, counting the number of waste placed and retrieved in each area over the past 24 hours. Zones with significantly more placements than retrievals are identified as stacking areas; those with significantly more retrievals are identified as feeding areas; and those with very few placements and retrievals are identified as fermentation areas. This adaptive zoning identification supports dynamic zoning and can automatically identify various functional zones, including those for co-firing. Next, the system encodes real-time data, including real-time waste distribution, unloading lists, feeding demand lists, real-time location and load of the waste crane, real-time waste supply rate, and real-time functional status of each area, into system decision states. Then, based on a trained strategy, the system generates optimal actions, determining the grabbing and releasing positions of the waste crane and the weight of waste to be grabbed and released. Finally, the system decodes these optimal actions into operational instructions, such as grabbing waste from the unloading gate area and stacking it in a designated area, grabbing waste from the feeding area and dropping it into the feeding port, or performing a dumping operation, which are then given to the operator or directly integrated with the waste crane control system.

[0116] In practical applications, this system is specifically optimized for situations with insufficient waste supply, maximizing fermentation quality with limited waste volume. When waste supply is low, the system uses refined space management to prioritize the full fermentation of waste in the fermentation zone and rationally adjust waste stacking in the composting zone to avoid excessive accumulation that could negatively impact fermentation. Simultaneously, the system can distinguish between municipal solid waste requiring fermentation and industrial waste that does not, automatically identifying the blending zone and optimizing the blending ratio. This allows for the appropriate blending of industrial waste with fermented municipal solid waste, improving power generation efficiency. For example, if the system detects that municipal solid waste in a certain grid cell has fermented for 7 days, reaching a high calorific value, while the industrial waste in the adjacent blending zone does not require fermentation, it generates an operation command to control the waste crane to grab an appropriate amount of waste from the municipal solid waste grid cell and feed it into the incinerator along with the industrial waste from the blending zone in an optimized ratio. This ensures both the calorific value of the waste entering the furnace and improves waste utilization.

[0117] Based on this, the problems of insufficient quantitative accuracy and lack of optimization methods in waste storage management have been effectively solved. In the context of insufficient waste supply, the calorific value of waste has been improved through refined management, the fluctuation of calorific value has been reduced, and the power generation capacity per ton of waste has been increased, providing strong support for the efficient operation of waste incineration power plants.

[0118] In summary, by spatially dividing the target waste storage area into grids and mining the waste distribution patterns and state characteristics of each grid unit using historical operational data, cross-unit correlation is achieved through feature transfer to determine the total calorific value of the waste. Then, reinforcement learning objectives are determined based on the hierarchical decomposition and quantitative evaluation of expected management needs. Aligned multi-source simulation information is generated through dynamic partitioning and real-time operational data encoding. Finally, intelligent decision-making operation instructions are generated using reinforcement learning model training and experimental data input. This embodiment of the application comprehensively achieves an intelligent and refined upgrade of waste storage area management, breaking through the limitations of traditional waste storage area management. Overcoming the limitations of relying on experience-based judgment and lacking systematic quantitative analysis, this approach improves management precision from the regional level to the grid level through grid-based division. By combining historical data mining and cross-unit feature correlation, it achieves a comprehensive and accurate grasp of waste distribution, status, and calorific value. Through the scientific setting of reinforcement learning objectives and model training, it can automatically explore optimal operating strategies in complex waste storage operation scenarios. The generated intelligent decision-making operation instructions can effectively guide the actual operation of waste storage facilities, improve the quality of waste fermentation and the efficiency of incineration power generation, and adapt to dynamic changes such as fluctuations in waste supply, providing systematic technical support for the efficient and stable operation of waste storage facilities.

[0119] Based on the same inventive concept, embodiments of this application also provide a computer device. See also... Figure 2 As shown, it is a schematic diagram of the structure of a possible computer device provided in an embodiment of this application. Figure 2In the computer device 200, there are a processor 210 and a memory 220. The processor 210 and the memory 220 are connected to each other via a communication bus. The memory 220 stores computer programs that can be executed by the processor 210. By executing the instructions stored in the memory 220, the processor 210 can perform the steps of the above-mentioned intelligent simulation method for waste bins based on reinforcement learning.

[0120] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium including a computer program. When the computer program is run on a computer device, it causes the computer device to perform the steps of the aforementioned reinforcement learning-based intelligent simulation method for waste bins. In some possible implementations, various aspects of the reinforcement learning-based intelligent simulation method for waste bins provided in this application can also be implemented as a program product including a computer program. When the program product is run on a computer device, it causes the computer device to perform the steps in the aforementioned reinforcement learning-based intelligent simulation method for waste bins. For example, the computer device can perform actions such as... Figure 1 The steps are shown in the diagram. The computer-readable storage medium includes volatile or non-volatile or a combination thereof, and may be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium.

[0121] like Figure 3 The diagram shown is a functional block diagram of the computer device provided in this embodiment. The computer device includes a reinforcement learning-based intelligent simulation device for a waste bin. The reinforcement learning-based intelligent simulation device for a waste bin includes: The spatial grid division module is used to respond to the intelligent simulation task for the target waste bin by performing spatial grid division processing on the target waste bin to obtain several grid cells; The grid cell processing module is used to mine the distribution pattern and trace the garbage accumulation status of each grid cell based on the historical operation data of the target garbage bin, so as to obtain the garbage distribution pattern and garbage status characteristic information corresponding to each grid cell; The full calorific value determination module is used to perform cross-cell correlation processing on the waste distribution pattern and waste status feature information corresponding to each grid cell through feature transmission, generate inter-grid correlation features, and determine the full calorific value data of waste by combining the waste distribution pattern, the waste status feature information and the inter-grid correlation features. The reinforcement learning setting module is used to collect the expected management demand features of the target waste bin and decompose them hierarchically to obtain demand hierarchical features. The demand hierarchical features and the total waste calorific value data are converted into evaluation indicators with the same dimension or without units. Based on the quantitative relationship between the converted demand evaluation indicators and the calorific value evaluation indicators, the reinforcement learning objective of the intelligent simulation task is determined. The multi-source simulation alignment module is used to determine the dynamic trajectory simulation information of the waste in each grid cell by partitioning dynamic identification based on the reinforcement learning objective, and simultaneously perform state encoding processing on the real-time simulation operation data of the target waste bin to generate state simulation encoding information and generate multi-source simulation information after the state simulation encoding information and the dynamic trajectory simulation information are aligned. The reinforcement learning decision module is used to take the aligned multi-source simulation information as input to the initial reinforcement learning model, call a preset policy optimization algorithm to train the initial reinforcement learning model, generate a trained reinforcement learning model, input the aligned multi-source measured information into the trained reinforcement learning model, and obtain intelligent decision operation instructions for the target waste bin under the preset heat value enhancement label.

[0122] Accordingly, this application also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above method embodiments.

[0123] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0124] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0125] Finally, it should be noted that the above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for intelligent simulation of waste bins based on reinforcement learning, characterized in that, The method includes: In response to the intelligent simulation task for the target waste bin, the target waste bin is spatially gridded to obtain several grid cells; Based on the historical operational data of the target waste bin, the distribution patterns of each grid cell are mined and the waste accumulation status is traced to obtain the waste distribution patterns and waste status characteristics of each grid cell. By performing cross-cell association processing on the waste distribution pattern and waste status feature information corresponding to each grid cell through feature transfer, inter-grid association features are generated. Combining the waste distribution pattern, the waste status feature information and the inter-grid association features, the total waste calorific value data is determined. The expected management needs of the target waste bin are collected and decomposed hierarchically to obtain hierarchical features of the needs. The hierarchical features of the needs and the total calorific value data of the waste are converted into evaluation indicators with the same dimension or without units. Based on the quantitative relationship between the converted demand evaluation indicators and the calorific value evaluation indicators, the reinforcement learning objective of the intelligent simulation task is determined. Based on the reinforcement learning objective, the dynamic trajectory simulation information of waste in each grid cell is determined by partition dynamic identification. Simultaneously, the real-time simulation operation data of the target waste bin is processed by state encoding to generate state simulation encoding information, and multi-source simulation information after the state simulation encoding information and the dynamic trajectory simulation information are aligned is generated. The aligned multi-source simulation information is used as the input to the initial reinforcement learning model. A preset policy optimization algorithm is called to train the initial reinforcement learning model to generate a post-trained reinforcement learning model. The aligned multi-source measured information is then input into the post-trained reinforcement learning model to obtain intelligent decision-making operation instructions for the target waste bin under a preset heat value enhancement label.

2. The method as described in claim 1, characterized in that, The process involves cross-cell correlation processing of the waste distribution patterns and waste status characteristics corresponding to each grid cell through feature transfer, generating inter-grid correlation features. Combining the waste distribution patterns, waste status characteristics, and inter-grid correlation features, the total waste calorific value data is determined, including: Extract the distribution density and distribution time sequence features from the waste distribution pattern of each grid cell, and extract the waste component features and waste accumulation height features from the waste state feature information of each grid cell; A grid cell association topology is established with adjacent grid cells as the association objects. The distribution density characteristics and distribution time sequence characteristics of each grid cell are transferred to the adjacent grid cells. The feature attenuation coefficient during the transfer process is calculated. Based on the feature attenuation coefficient, the feature values ​​after transfer are adjusted to generate the inter-grid association features corresponding to each grid cell. The waste distribution pattern, waste status characteristics and corresponding inter-grid correlation characteristics of each grid cell are standardized or normalized respectively. The processed features are spliced ​​or fused to generate a single-grid multi-dimensional feature vector for each grid cell. The single-grid multi-dimensional feature vectors of all grid cells are summarized to obtain a multi-dimensional waste feature set. Extract the waste component features and distribution density features from the multi-dimensional waste feature set, match the waste component features with the preset component calorific value mapping relationship, and obtain the initial calorific value data corresponding to each grid cell; The heat value correction coefficient of each grid cell is calculated based on the distribution density characteristics and inter-grid correlation characteristics, and the heat value correction coefficient is a unitless quantity. The initial calorific value data is weighted or scaled using the calorific value correction coefficient to generate corrected calorific value data for each grid cell; the corrected calorific value data of all grid cells are combined to generate the total calorific value data of the waste.

3. The method as described in claim 1, characterized in that, The process involves collecting the expected management needs characteristics of the target waste bin and decomposing them hierarchically to obtain hierarchical needs characteristics. These hierarchical needs characteristics and the total waste calorific value data are then converted into evaluation indicators of the same dimension or without units. Based on the quantitative relationship between the converted needs evaluation indicators and the calorific value evaluation indicators, the reinforcement learning objective of the intelligent simulation task is determined, including: Collect the expected management requirements of the target waste storage facility, which include waste treatment efficiency requirements, waste storage capacity requirements, and waste transfer coordination requirements. The expected management requirements are decomposed hierarchically, with the waste treatment efficiency requirement being decomposed into treatment timeliness and treatment quality characteristics, the waste storage capacity requirement being decomposed into capacity utilization and capacity dynamic adjustment characteristics, and the waste transfer coordination requirement being decomposed into transfer timing and transfer connection characteristics, resulting in multi-dimensional hierarchical requirements. The demand-level features are subjected to feature quantification processing to generate demand quantification indicators corresponding to each demand-level feature, and the full amount of waste heat value data is subjected to partition heat value statistical processing to obtain the partition heat value statistical results of each grid unit. Based on the preset evaluation rules, the processing timeliness and processing quality characteristics are quantified as the first demand evaluation indicators; the capacity utilization rate and capacity dynamic adjustment characteristics are quantified as the second demand evaluation indicators; the transfer time sequence and transfer connection characteristics are quantified as the third demand evaluation indicators; the overall calorific value level of the total waste calorific value data, the zonal calorific value statistical results of each grid unit, and the time sequence change characteristics are quantified as the corresponding first, second, and third calorific value evaluation indicators; a quantitative correlation relationship is established between the first demand evaluation indicators and the first calorific value evaluation indicators, between the second demand evaluation indicators and the second calorific value evaluation indicators, and between the third demand evaluation indicators and the third calorific value evaluation indicators, generating demand calorific value matching relationships; Based on the demand heat value matching relationship, demand heat value association pairs that meet the preset conditions are extracted. The demand heat value association pairs are transformed into objective function constraints for reinforcement learning. The reinforcement learning objective is determined by combining the task setting keywords of the intelligent simulation task.

4. The method as described in claim 1, characterized in that, Based on the reinforcement learning objective, the dynamic trajectory simulation information of waste in each grid cell is determined through partitioned dynamic identification. Simultaneously, state encoding processing is performed on the real-time simulation data of the target waste bin to generate state simulation encoding information. Then, multi-source simulation information aligned with the state simulation encoding information and the dynamic trajectory simulation information is generated, including: Based on the required accuracy of dynamic tracking of waste in the reinforcement learning objective, the dynamic identification range of the partition is set, and the grid unit of the target waste bin is divided into a core identification area and an extended identification area. The identification frequency of the core identification area is higher than that of the extended identification area. The iterative identification termination condition is set when the overlap of the waste trajectory between two adjacent identifications reaches a preset standard. During the first identification process, the waste movement process of each grid unit in the core identification area is collected in real time to generate an initial trajectory point set. Based on the initial set of trajectory points, trajectory fitting is performed to generate an initial waste trajectory. The initial waste trajectory is then extended to the extended recognition area, and trajectory points within the extended recognition area are collected to supplement the initial waste trajectory, resulting in an initial complete waste trajectory. In the next iteration of identification, the trajectory points of the core identification area and the extended identification area are re-collected to generate a new complete garbage trajectory, and the overlap between the new complete garbage trajectory and the initial complete garbage trajectory is calculated; If the overlap does not reach the termination condition of iterative identification, the new complete waste trajectory is used as the initial complete waste trajectory to jump to the steps of trajectory acquisition, supplementation and overlap calculation until the overlap reaches the termination condition of iterative identification. The final complete waste trajectory is then used as dynamic trajectory simulation information. During the trajectory tracking process, real-time simulation operation data of the target waste bin is collected simultaneously. The real-time simulation operation data includes the real-time accumulation amount of waste, real-time temperature of waste, and real-time humidity of waste in each grid unit. The real-time simulation data is processed by state coding, which converts the real-time accumulation amount of waste, the real-time temperature of waste, and the real-time humidity of waste into standardized coding vectors to generate state simulation coding information. The trajectory acquisition timestamp in the dynamic trajectory simulation information and the data acquisition timestamp in the state simulation coding information are extracted. The dynamic trajectory simulation information and the state simulation coding information are time-series aligned based on the timestamps. Information with mismatched timestamps is deleted to obtain aligned multi-source simulation information.

5. The method as described in claim 4, characterized in that, The step of setting a dynamic identification range for the partition based on the accuracy requirements for dynamic tracking of waste in the reinforcement learning objective, and dividing the grid cells of the target waste bin into a core identification region and an extended identification region, includes: Extract the garbage dynamic tracking accuracy requirements from the reinforcement learning objective, and determine the allowable range of position error and time error for trajectory tracking; Analyze the historical garbage accumulation frequency of the grid cells of the target garbage bin, mark the grid cells with historical garbage accumulation frequency higher than a preset frequency threshold as the first accumulation grid cells, and mark the grid cells with historical garbage accumulation frequency lower than or equal to the preset frequency threshold as the second accumulation grid cells; Based on the allowable range of positional error, the spatial accuracy requirement of the core identification area is determined; with the first stacked grid unit as the center, according to the spatial accuracy requirement and referring to the historical stacking frequency distribution, the boundary of the core identification area is delineated, and the boundary of the core identification area covers the first stacked grid unit and an adjacent preset number of second stacked grid units. The grid cells outside the core identification area within the target waste bin are defined as the extended identification area, which covers the remaining second stacking grid cells. The recognition frequency is set based on the allowable time error range. The recognition frequency of the core recognition area is set to the highest frequency that meets the allowable time error range, and the recognition frequency of the extended recognition area is set to a preset ratio of the recognition frequency of the core recognition area. The recognition frequency of the core recognition area is higher than the recognition frequency of the extended recognition area. The core recognition area range, the extended recognition area range, and the corresponding recognition frequency are integrated into the partitioned dynamic recognition parameters.

6. The method according to any one of claims 1-4, characterized in that, The step of using the aligned multi-source simulation information as input to the initial reinforcement learning model, and calling a preset policy optimization algorithm to train the initial reinforcement learning model to generate a trained reinforcement learning model includes: The aligned multi-source simulation information is subjected to feature standardization processing to eliminate the dimensional differences between dynamic trajectory simulation information and state simulation coding information, generating standardized multi-source simulation information. The standardized multi-source simulation information is divided into training dataset and validation dataset according to a preset ratio. The training dataset is used for iterative updating of model parameters, and the validation dataset is used for evaluating the model training effect. The network parameters of the initial reinforcement learning model are initialized. The initial reinforcement learning model includes a feature extraction layer, a policy decision layer, and a value evaluation layer. The feature extraction layer is used to extract temporal correlation features and state correlation features from standardized multi-source simulation information. The policy decision layer is used to output a model training policy based on the current input features. The value evaluation layer is used to calculate the value estimate corresponding to the training policy. The preset strategy optimization algorithm is invoked, and the reward function of the algorithm is set based on the reinforcement learning objective. The reward value of the reward function is positively correlated with the degree of fit of the model output to meet the preset heat value improvement requirements. The constraint conditions corresponding to the inter-grid correlation features are incorporated into the penalty term of the reward function. When feature overflow or parameters deviate from the preset range during model training, the penalty mechanism is triggered. The training dataset is input into the feature extraction layer of the initial reinforcement learning model. Convolutional features in the standardized multi-source simulation information are extracted through convolutional operations to generate simulation feature vectors. The simulation feature vectors are input into the policy decision layer, which outputs an initial training policy based on a preset policy distribution. The initial value evaluation result is obtained by estimating the value of the initial training policy through the value evaluation layer. The reward value corresponding to the initial training strategy is calculated based on the reward function. The policy decision layer parameters of the initial reinforcement learning model are updated using the policy gradient descent method in combination with the initial value evaluation result. The convolution kernel parameters of the feature extraction layer are adjusted synchronously to optimize the convolution feature extraction effect, and one iteration of model training is completed. The training data input, feature extraction, policy output, value evaluation and parameter update steps are executed in a loop. After each preset number of iterations of training, the validation dataset is input into the reinforcement learning model obtained in the current iteration, and the error value between the model output result and the label result corresponding to the validation dataset is calculated. Determine whether the error value is less than a preset error threshold. If the error value is greater than or equal to the preset error threshold, continue the iterative training step and dynamically adjust the learning rate of the policy gradient descent. If the error value is less than the preset error threshold, stop the iterative training, determine the reinforcement learning model obtained in the current iteration as the post-trained reinforcement learning model, and save the network parameters of the post-trained reinforcement learning model and the optimal policy parameters during the training process.

7. The method according to any one of claims 1-4, characterized in that, The aligned multi-source measured information includes dynamic trajectory measured information and state measured encoded information. The step of inputting the aligned multi-source measured information into the trained reinforcement learning model to obtain intelligent decision-making operation instructions for the target waste bin under a preset heat value enhancement label includes: The measured dynamic trajectory information and the measured state encoding information are synchronously input into the feature extraction layer of the reinforcement learning model after training. The feature extraction layer performs feature fusion processing on the measured dynamic trajectory information and the measured state encoding information based on the convolution kernel parameters optimized during training. The key temporal node features in the measured dynamic trajectory information are strengthened through a temporal attention mechanism to generate a measured fusion feature vector. The policy decision layer of the post-trained reinforcement learning model is invoked to perform feature parsing on the measured fused feature vector, and target features associated with the preset heat value enhancement label are extracted. The target features include the garbage dynamic accumulation rate feature of the grid cell, the heat value potential feature corresponding to the real-time state code, and the garbage migration association feature between grids. A feature matching matrix is ​​generated based on the matching rules between the target features and the preset heat value enhancement labels. The matching degree between the target features and the preset heat value enhancement labels is calculated through the feature matching matrix, and the target feature combination with the highest matching degree is selected as the decision basis. Combining the aforementioned decision criteria with the optimal strategy parameters saved during training, an initial decision scheme is generated. The initial decision scheme includes suggestions for waste transfer priority, waste accumulation adjustment parameters, and real-time monitoring frequency adjustment for each grid unit. The value evaluation layer of the trained reinforcement learning model is invoked to verify the value of the initial decision scheme. The probability that the total calorific value data of the target waste bin reaches the preset calorific value improvement standard after the implementation of the initial decision scheme is calculated, and it is determined whether the probability is greater than the preset probability threshold. If the probability is less than or equal to the preset probability threshold, the strategy decision layer adjusts the decision parameters based on the value assessment results, regenerates the decision scheme, and performs value verification again until the success probability of the generated decision scheme is greater than the preset probability threshold. If the probability is greater than the preset probability threshold, the current decision scheme is transformed into a standardized intelligent decision operation instruction. The intelligent decision operation instruction includes the instruction execution object, execution parameters, execution sequence, and execution priority. The execution object is a specific grid unit of the target waste bin. The execution parameters are quantitative indicators that include at least the waste transfer volume and the stacking height adjustment value. The execution sequence is determined based on the key nodes of the dynamic trajectory measured information.

8. The method as described in claim 7, characterized in that, The step of combining the decision criteria with the optimal policy parameters saved during training to generate an initial decision scheme includes: The target feature combination corresponding to the decision basis is subjected to multi-dimensional structured analysis, and the feature dimension of the garbage dynamic accumulation rate of the grid cell, the feature dimension of the heat value potential corresponding to the real-time status code, and the feature dimension of garbage migration association between grids are decomposed. The feature weight distribution, feature time series change law and feature quantization value under each feature dimension are extracted and integrated to obtain the decision basis feature analysis matrix. Retrieve the optimal strategy parameters saved during training, determine the parameter type, applicable range and adjustment threshold of the optimal strategy parameters, and generate an optimal strategy parameter index library containing parameter index, scene label and adjustment rules. The scene label and the feature dimension and feature weight distribution of the decision basis form a one-to-one correspondence. Using the feature dimensions, feature weight distribution, and feature temporal change patterns in the decision-making feature analysis matrix as search conditions, a layer-by-layer matching query is performed in the optimal strategy parameter index library to filter out the optimal strategy parameters with a matching degree higher than the feature matching degree threshold, and to generate a subset of optimal strategy parameters that are adapted to the current decision-making scenario. A decision parameter calculation model is constructed based on the subset of optimal strategy parameters; the quantified values ​​of each feature in the decision basis feature analysis matrix are normalized to obtain dimensionless feature scores; the mapping calculation relationship between each optimal strategy parameter and the normalized feature scores is determined; the normalized feature scores are input into the decision parameter calculation model, and the initial values ​​of waste transfer volume, stacking height adjustment, and real-time monitoring frequency corresponding to each grid unit are obtained through weighted calculation. Based on the grid cell topology of the target waste bin and the parameter coordination rules between grids in historical operation, the initial values ​​of decision parameters of each grid cell are coordinated and adapted. When there is a conflict between the initial values ​​of decision parameters of adjacent grid cells, the initial values ​​are recalculated and adjusted based on the control rules of the optimal strategy parameter subset and the waste migration correlation characteristics between grids, so as to achieve the coordinated rationalization of the initial values ​​of decision parameters of each grid cell. Based on the decision types of waste transfer, accumulation adjustment, and real-time monitoring, the decision parameters after collaborative adaptation are classified and integrated to generate a waste transfer priority ranking table, a waste accumulation adjustment parameter set, and a real-time monitoring frequency adjustment list. The waste transfer priority ranking table is determined based on the dynamic accumulation rate characteristics and calorific value potential characteristics of each grid unit, and the waste accumulation adjustment parameter set is associated with the accumulation height limit threshold of each grid unit. A decision instruction tag set is constructed, which includes decision target tags, decision object tags, decision parameter tags, and time sequence execution tags. The classified and integrated waste transfer priority ranking table, waste accumulation adjustment parameter set, and real-time monitoring frequency adjustment list are filled into the corresponding tags. The execution trigger conditions, execution duration, and correlation logic between tags for each decision parameter are supplemented to generate the initial decision scheme.

9. A computer device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of the reinforcement learning-based intelligent simulation method for waste bins as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program that, when run on a computer device, causes the computer device to perform the steps of the reinforcement learning-based intelligent simulation method for waste bins as described in any one of claims 1 to 8.