A method and system for intelligent optimization of co-pyrolysis feed ratio

By constructing a causal network and using a causal constraint reinforcement learning algorithm to optimize the ratio of raw materials for co-pyrolysis, the problem of strategy instability caused by the lack of a kinetic response mechanism in the existing technology is solved, and a more efficient and reliable optimization of the co-pyrolysis process is achieved.

CN121601097BActive Publication Date: 2026-05-26GREEN HARVEST ENERGY (BEIJING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GREEN HARVEST ENERGY (BEIJING) TECHNOLOGY CO LTD
Filing Date
2026-01-28
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing methods for optimizing the proportion of co-pyrolysis feedstocks lack systematic modeling of the evolution of reaction behavior and kinetic response mechanisms. This results in optimization results that are sensitive to changes in operating conditions and lack strategy stability, making it difficult to reliably extend to complex multi-condition environments.

Method used

By constructing a causal network and combining multi-scale reaction features and kinetic feature datasets, a causal constraint reinforcement learning algorithm is used to optimize the raw material ratio. A temporal difference reinforcement learning algorithm is introduced for policy evaluation to ensure that the optimization strategy conforms to the physical mechanism and causal law.

Benefits of technology

It improves the accuracy and stability of raw material ratio optimization, enhances the interpretability of the model and the efficiency of strategy learning, and provides a reliable intelligent optimization control scheme.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121601097B_ABST
    Figure CN121601097B_ABST
Patent Text Reader

Abstract

This invention provides an intelligent optimization method and system for co-pyrolysis feedstock ratio, relating to the field of data processing technology. The method includes: acquiring feedstock data; constructing a co-pyrolysis reaction characterization feature dataset under different ratio conditions based on the feedstock data, wherein the co-pyrolysis reaction characterization features include multi-scale reaction features and reaction trajectory features; constructing a kinetic feature dataset based on the reaction trajectory features; constructing a causal network for the co-pyrolysis feedstock ratio based on the co-pyrolysis reaction feature dataset and the kinetic feature dataset; constructing a causal constraint reinforcement learning environment based on the causal network; determining a co-pyrolysis feedstock ratio optimization strategy using a temporal difference reinforcement learning algorithm based on the causal constraint reinforcement learning environment; evaluating the reversibility of the co-pyrolysis feedstock ratio optimization strategy to determine the final strategy; and executing the final strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for intelligent optimization of the proportion of co-pyrolysis raw materials. Background Technology

[0002] The significance of intelligent optimization of the ratio of co-pyrolysis raw materials lies in the fact that by rationally controlling the ratio of various raw materials in the co-pyrolysis process, the reaction process can be highly efficient and coordinated, the product distribution can be optimized, and the energy and resource utilization efficiency can be improved.

[0003] Different raw materials differ in composition, pyrolysis characteristics, and reactivity. Their ratio directly affects the reaction initiation temperature, reaction stage evolution, energy conversion efficiency, and final product quality during co-pyrolysis. Therefore, intelligent optimization of raw material ratio is a key technical link to improve the overall performance and operational stability of co-pyrolysis systems, and has important application value in fields such as biomass resource utilization, solid waste treatment, and clean energy conversion.

[0004] However, while there are some research methods for optimizing the feed ratio of co-pyrolysis, most of them focus on model analysis based on experimental experience, statistical regression, or correlation. They usually only depict the superficial relationship between changes in the ratio and products or performance indicators, lacking systematic modeling of the evolution of reaction behavior and kinetic response mechanisms. Even when intelligent algorithms are introduced, they often fail to explicitly incorporate reaction kinetics and causal relationships into the decision-making process. This can easily lead to the optimization results being sensitive to changes in operating conditions and insufficient strategy stability, making it difficult to reliably promote them in complex, multi-condition co-pyrolysis applications. Summary of the Invention

[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide an intelligent optimization method for co-pyrolysis feedstock ratio. This method can solve the problem that although there are some research methods for optimizing co-pyrolysis feedstock ratio, they are mostly based on experimental experience, statistical regression, or correlation-driven model analysis. They usually only describe the superficial relationship between ratio changes and products or performance indicators, lacking systematic modeling of reaction behavior evolution and kinetic response mechanisms. Even when intelligent algorithms are introduced, they often do not explicitly incorporate reaction kinetics and causal relationships into the decision-making process. This easily leads to the optimization results being sensitive to changes in operating conditions and insufficient strategy stability, making it difficult to reliably promote in complex, multi-condition co-pyrolysis applications.

[0006] A first aspect of this invention provides a method for intelligent optimization of the proportion of co-pyrolysis raw materials, comprising:

[0007] S1: Obtain raw material data;

[0008] S2: Based on the raw material data, construct a characterization dataset of co-pyrolysis reaction under different ratio conditions, wherein the characterization dataset of co-pyrolysis reaction includes multi-scale reaction features and reaction trajectory features;

[0009] S3: Construct a kinetic feature dataset based on the described reaction trajectory characteristics;

[0010] S4: Based on the pyrolysis reaction characterization dataset and the kinetic characteristic dataset, construct a causal network for the pyrolysis feedstock ratio;

[0011] S5: Based on the causal network, construct a causal constraint reinforcement learning environment;

[0012] S6: Based on the causal constraint reinforcement learning environment, determine the optimization strategy for the ratio of co-pyrolysis raw materials through the temporal difference reinforcement learning algorithm;

[0013] S7: Perform a reversibility assessment on the co-pyrolysis raw material ratio optimization strategy and determine the final strategy;

[0014] S8: Execute the final strategy.

[0015] A second aspect of this invention provides an intelligent optimization system for the proportioning of co-pyrolysis raw materials, comprising: a processor and a memory;

[0016] The memory stores programs or instructions that can run on the processor, which, when executed by the processor, implement the steps of the intelligent optimization method for co-pyrolysis raw material ratio as described in the first aspect.

[0017] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0018] In this embodiment of the invention, multi-scale reaction features, reaction trajectory features, and corresponding kinetic parameters are extracted based on raw material data. This allows for a comprehensive quantitative characterization of the impact of changes in raw material ratios on the evolution of reaction behavior and kinetic response. By constructing a causal network, the true causal paths between ratio variables, reaction features, and kinetic features are identified, avoiding spurious associations introduced by traditional correlation methods. This improves the interpretability and stability of the model structure. The causal structure is used as a priori constraint for the reinforcement learning environment, ensuring that state evolution, action selection, and reward allocation all conform to the physical mechanisms and causal laws of the co-pyrolysis process. This significantly improves the efficiency and convergence of policy learning. By introducing a policy reversibility evaluation mechanism, the reversibility and correctability of the raw material ratio optimization strategy during execution are assessed, thereby constraining and avoiding irreversible or high-risk ratio adjustment behaviors. Overall, this embodiment of the invention effectively improves the accuracy, stability, and interpretability of raw material ratio optimization, providing an implementable, reliable, and engineering-practical solution for the intelligent optimization control of co-pyrolysis processes. Attached Figure Description

[0019] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0020] Figure 1 This is a schematic flowchart of a method for intelligent optimization of co-pyrolysis raw material ratio provided in an embodiment of the present invention.

[0021] Figure 2 This is a schematic diagram of the structure of an intelligent optimization system for co-pyrolysis raw material ratio provided in an embodiment of the present invention. Detailed Implementation

[0022] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0023] The intelligent optimization method for co-pyrolysis raw material ratio provided by the present invention will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0024] Reference manual attached Figure 1 The diagram shows a flowchart of an intelligent optimization method for co-pyrolysis raw material ratio provided by an embodiment of the present invention.

[0025] This invention provides a method for intelligent optimization of co-pyrolysis raw material ratio, which may include the following steps:

[0026] S1: Obtain raw material data.

[0027] It should be noted that raw material data refers to the set of fundamental data used to characterize the compositional characteristics, physicochemical properties, and basic thermal reaction features of various raw materials participating in the co-pyrolysis process. Raw material data includes basic attribute data such as raw material category identification, composition information, elemental analysis data, industrial analysis data, and moisture content. It also includes basic thermal reaction data such as thermogravimetric or differential thermogravimetric data obtained for individual raw materials under preset process conditions, and characteristic reaction temperature ranges.

[0028] S2: Based on the raw material data, construct a characterization dataset of co-pyrolysis reaction under different ratio conditions. The characterization dataset of co-pyrolysis reaction includes multi-scale reaction features and reaction trajectory features.

[0029] Among them, the characterization features of co-pyrolysis reactions are a set of characteristic information used to describe the reaction behavior, product evolution, and changes in thermal effects of multiple raw materials during co-pyrolysis. Multi-scale reaction features refer to features extracted from different time scales, temperature scales, or reaction stage scales, used to describe the behavioral differences of co-pyrolysis reactions at macroscopic and local stages, such as mass loss rates and exothermic / endothermic characteristics. Reaction trajectory features refer to features reflecting the continuous evolution path of the reaction state over time or temperature during co-pyrolysis, used to describe the dynamic changes in the reaction process.

[0030] In this embodiment of the invention, by constructing a co-pyrolysis reaction characterization feature dataset containing multi-scale reaction features and reaction trajectory features under different raw material ratios, it is possible to systematically and comprehensively characterize the influence of changes in raw material ratios on co-pyrolysis reaction behavior, so that the dynamic evolution characteristics and stage differences of the reaction process can be fully expressed.

[0031] Specifically, since different raw materials have significant differences in composition, thermal stability, volatile release characteristics and reactivity, and changes in the raw material ratio will cause nonlinear changes in the reaction initiation temperature, reaction intensity and reaction stage evolution path during co-pyrolysis, this embodiment of the invention constructs a co-pyrolysis reaction characterization feature dataset.

[0032] The construction process is as follows: Based on raw material data, the types of raw materials participating in co-pyrolysis and their proportions are determined. Co-pyrolysis process data under different raw material proportions are acquired under preset heating rates and reaction atmospheres. Feature extraction is performed on the co-pyrolysis process data to construct multi-scale reaction features that characterize the behavior of the co-pyrolysis reaction at different scales within a unified temperature range and time scale. Simultaneously, reaction trajectory features reflecting the continuous evolution of reaction intensity with temperature or time are extracted. Finally, the multi-scale reaction features and reaction trajectory features are uniformly organized and labeled according to the raw material proportions to construct a dataset of co-pyrolysis reaction characterization features under different proportions, used to characterize the impact of changes in raw material proportions on the evolution of co-pyrolysis reaction behavior.

[0033] S3: Construct a kinetic feature dataset based on the reaction trajectory characteristics.

[0034] In this embodiment of the invention, by constructing a kinetic feature dataset based on reaction trajectory characteristics, it is possible to establish a correspondence between the reaction rate change and stage transformation behavior during co-pyrolysis and kinetic parameters such as apparent activation energy and pre-exponential factor, thereby quantifying the intrinsic kinetic nature of the reaction process.

[0035] In one possible implementation, S3 specifically includes:

[0036] S301: Based on the characteristics of the reaction trajectory, the co-pyrolysis process is divided into stages to determine multiple reaction stages.

[0037] S302: Constructing the Coats–Redfern kinetic model for each reaction stage:

[0038]

[0039] Where ln represents the logarithmic function, Represents the integral mechanism function, The conversion rate represents the proportion of the sample's reaction progress during the pyrolysis process. This indicates the absolute temperature, that is, the instantaneous temperature of the sample during the pyrolysis process. A Indicates pre-exponential factor, R Represents the gas constant. E Indicates the apparent activation energy. This indicates the rate of temperature increase.

[0040] S303: Input the reaction trajectory characteristics into the Coats–Redfern kinetic model to obtain the corresponding kinetic parameters for each reaction stage.

[0041] Specifically, within each reaction stage, the relationship between the conversion rate α and temperature T of the sample in that stage is determined based on the characteristics of the reaction trajectory, and an integral mechanism function is selected. Under the premise of this, the conversion rate and temperature data are substituted into the Coats–Redfern kinetic model, and the conversion rate and temperature data are calculated by adjusting ln(g(α) / T). 2 By performing linear regression fitting between 1 / T and the reaction stage, the apparent activation energy E and pre-exponential factor A corresponding to that reaction stage are derived from the fitting slope and intercept, respectively, thereby obtaining the kinetic parameters characterizing the kinetic properties of each reaction stage.

[0042] S304: Construct a dynamic feature dataset based on dynamic parameters.

[0043] Specifically, the apparent activation energy, pre-exponential factor, and other kinetic parameters obtained by inversion from the Coats–Redfern kinetic model in each reaction stage are uniformly organized and structurally represented according to the raw material ratio, reaction stage, and corresponding temperature range. The kinetic parameters under different ratio conditions are labeled and normalized to construct a kinetic feature dataset that can reflect the differences in the kinetic response of the co-pyrolysis reaction under different raw material ratios and reaction stages. This dataset is used to characterize the influence of changes in raw material ratio on the kinetic behavior of the co-pyrolysis reaction.

[0044] S4: Based on the characterization dataset and kinetic dataset of the co-pyrolysis reaction, construct a causal network for the ratio of co-pyrolysis feedstocks.

[0045] In this embodiment of the invention, by jointly modeling the co-pyrolysis reaction feature dataset and the kinetic feature dataset and using them for causal network construction, it is possible to simultaneously introduce reaction behavior evolution information and kinetic mechanism constraints, thereby overcoming the limitation that single feature or correlation analysis is difficult to reveal the true influence mechanism of raw material ratio. This not only helps to identify the direct and indirect causal paths of raw material ratio changes on the reaction process and kinetic response, but also effectively suppresses the interference of spurious correlations.

[0046] Specifically, due to the multi-scale, strongly coupled, and significantly nonlinear characteristics of the influence of co-pyrolysis feedstock ratio on the reaction process, a single type of feature is insufficient to fully characterize the combined effect of ratio changes on the evolution of reaction behavior and its kinetic response. Therefore, it is necessary to jointly model co-pyrolysis reaction characterization features and kinetic features. To this end, multi-scale reaction features and reaction trajectory features in the co-pyrolysis reaction characterization feature dataset are used as information describing the evolution of reaction behavior caused by ratio changes, while kinetic parameters such as apparent activation energy and pre-exponential factor in the kinetic feature dataset are used as physical constraints reflecting the kinetic response characteristics of the reaction. Under a unified feedstock ratio working condition index, a multivariate time series feature representation is constructed, and based on the feature representation, a causal modeling method is used to identify the directed dependencies between feedstock ratio, reaction features, and kinetic features. This constructs a co-pyrolysis feedstock ratio causal network that can characterize the causal transmission path of "feedstock ratio change—reaction behavior evolution—kinetic response characteristics".

[0047] In one possible implementation, S4 specifically includes:

[0048] S401: Based on the pyrolysis reaction characterization dataset and kinetic characteristic dataset, set the set of nodes used to construct the causal network.

[0049] In this embodiment of the invention, by uniformly setting the node set of the causal network based on the pyrolysis reaction characteristic dataset and the kinetic characteristic dataset, the raw material ratio state, reaction behavior characteristics and kinetic response parameters can be incorporated into the same causal analysis framework, so that various features can be consistently expressed at the structural level.

[0050] Specifically, considering the influence of the feedstock ratio on the co-pyrolysis reaction process, the nodes of the causal network are uniformly defined as key feature variables that characterize the ratio-driven effect, through the evolution of reaction behavior and its corresponding kinetic response. These include: ratio variable nodes characterizing different feedstock ratio states; multi-scale reaction feature nodes and reaction trajectory feature nodes used to depict the behavior of the co-pyrolysis reaction at different temperature ranges, reaction stages, and characteristic scales; and kinetic parameter nodes estimated using the Coats–Redfern kinetic model. Each node in the node set corresponds one-to-one with a feature dimension in both the co-pyrolysis reaction characterization dataset and the kinetic feature dataset.

[0051] S402: The characteristic dataset of the co-pyrolysis reaction and the kinetic characteristic dataset are concatenated to obtain a multivariate time series characteristic set.

[0052] S403: Perform rank transformation on the feature set of multivariate time series to obtain the rank vector.

[0053] In this embodiment of the invention, by performing rank transformation on the multivariate time series feature set and constructing a rank vector, the influence of different feature dimensions, distribution patterns and extreme values ​​on the modeling results can be effectively weakened, thereby improving the robustness and generalization ability of the model under multiple raw material ratio conditions.

[0054] Specifically, the feature vectors at each time point in the multivariate time series feature set are sorted according to the feature dimension, the original values ​​are mapped to the corresponding ranks, and the rank results are used to construct the rank vector. In this way, while maintaining the relative order of each feature over time, the influence of different feature dimensions, distribution differences and outliers on the modeling process is eliminated.

[0055] S404: Construct a Rank-VAR model based on the rank vector.

[0056] The Rank-VAR model is specifically as follows:

[0057]

[0058] in, Represents a rank vector. This indicates that the characteristics of a multivariate time series are concentrated at time 1. t eigenvectors, In the Rank-VAR model, the first... The regression coefficient matrix with lag order, L This represents the maximum lag order of the Rank-VAR model. This indicates that the characteristics of a multivariate time series are concentrated at time 1. eigenvectors, express t The random perturbation term of the Rank-VAR model at time step.

[0059] S405: Construct a directed causal candidate edge set using the Rank-VAR model.

[0060] In this embodiment of the invention, the Rank-VAR model is used to model the rank vector time series, which can identify the potential directed influence relationship between different nodes based on the consideration of time lag effect, so that the time-series dependency structure between raw material ratio, reaction characteristics and kinetic characteristics can be systematically characterized.

[0061] In one possible implementation, S405 specifically includes:

[0062] S4051: Perform multivariate least squares estimation on the Rank-VAR model to obtain a set of estimated coefficient matrices.

[0063] S4052: Construct a single-coefficient significance test statistic based on the set of estimated coefficient matrices.

[0064] Specifically, for each lag order in the Rank-VAR model regression coefficients (represents a node) j Lag For nodes i The impact of different raw material ratios or segmented reaction trajectories was estimated separately, and the estimation results of the same single coefficient under multiple conditions were statistically summarized. A single coefficient significance test statistic based on cross-condition stability was constructed, which is defined as:

[0065]

[0066] in, This indicates the regression coefficients The constructed single-coefficient significance test statistic, In the Rank-VAR model, the first... The ()th order of the lag regression coefficient matrix i , j ) elements, Regression coefficients Estimated mean values ​​under different raw material ratios and operating conditions Indicates standard deviation, This represents the regularization constant used to prevent the denominator from being zero.

[0067] S4053: Based on the single coefficient significance test statistic, set the lag-level directed edge indicator.

[0068] Among them, the lag-level directed edge indicator is a binary identifier variable used to characterize whether there is an effective directed influence relationship between nodes under a specific lag level. Its value is used to indicate whether a node has a statistically significant and stable influence on another node under a given time lag condition.

[0069] Specifically, regarding the Rank-VAR model, the first... Regression coefficients under lag Based on its corresponding single-coefficient significance test statistic For regression coefficients The strength, stability, and consistency of the effect direction under different raw material ratios or segmented reaction trajectory conditions are comprehensively judged. When the statistical quantity meets the preset significance condition and maintains a consistent direction of influence under multiple conditions, it is considered a node. j In the lag pair of nodes i The influence relationship is statistically significant, thus it is marked as a valid lag-level directed edge; otherwise, it is not retained. The corresponding lag-level directed edge indicator is defined as:

[0070]

[0071]

[0072] in, This indicates the directed edge indicator of the lag stage, when At, it indicates the time. Under hysteresis, node j For nodes i The influence relationship is deemed valid when At, it indicates the time. Under hysteresis, node j For nodes i The influence relationship was deemed invalid. This indicates an indicator function; it takes a value of 1 when the statistic is not less than a threshold, and a value of 0 otherwise. The threshold for significance statistic, Indicates a metric for directional consistency. K This indicates the number of raw material ratio conditions during the co-pyrolysis experiment. Indicates the first k Under each raw material ratio condition, the regression coefficients are estimated using the Rank-VAR model, where sign() represents the sign function. This represents the directional consistency threshold.

[0073] S4054: Determine the node-level directed edge indication matrix based on the lag-level directed edge indication.

[0074] Specifically, the directed edge indicators of each node pair obtained under different time lag orders are integrated across lag orders. When the directed edge indicator of a certain node pair is determined to be valid under at least one time lag condition, it is considered that the node pair has a directed influence relationship at the node level, and its corresponding node-level directed edge indicator is set to be valid. Thus, the node-level directed edge indicator matrix is ​​constructed using the node-level directed edge indicators of each node pair as matrix elements, which is used to uniformly characterize the candidate directed dependency structure between nodes.

[0075] S4055: Construct a set of directed causal candidate edges based on the node-level directed edge indicator matrix.

[0076] Specifically, the matrix elements in the node-level directed edge indicator matrix are traversed, and the node pairs marked as valid are extracted as directed edges. The direction of the directed edges is determined according to the directional relationship corresponding to the node pairs. Thus, all directed edges that meet the conditions are summarized to form a directed causal candidate edge set. The directed causal candidate edge set is used to characterize the directed influence relationship between nodes that has statistical significance and stability under at least one time lag condition.

[0077] S406: Perform sparse inverse covariance estimation on the rank vector to determine the sparse precision matrix and the set of conditionally dependent candidate edges.

[0078] Specifically, an extended rank vector variable set containing the current term and multiple lag terms is constructed for the multivariate time series after rank transformation. The covariance matrix is ​​estimated based on the extended rank vector samples. To address the potential issues of non-invertibility or unstable estimation of the covariance matrix under short time series conditions, the LoGo sparse inverse covariance estimation method is used to inversely derive the covariance matrix, resulting in a precision matrix with a sparse structure. J , Representing variables and Direct dependencies still exist even when other variables are controlled. This indicates that the conditions are independent.

[0079] Furthermore, the direct dependencies between variables that still exist under the condition of controlling the other variables are characterized by the non-zero off-diagonal elements in the precision matrix, thereby determining the set of candidate edges for conditional dependencies.

[0080] S407: Define the conditional transition entropy of the directed causal candidate edge set based on the sparse precision matrix.

[0081] In this embodiment of the invention, the conditional transfer entropy is defined for directed causal candidate edges based on a sparse precision matrix. This enables quantitative evaluation of the information transmission strength between nodes under the condition of controlling the influence of relevant variables. This effectively distinguishes between direct causal effects and indirect or spurious correlations, improves the accuracy and reliability of causal edge screening results, and provides an important basis for constructing highly reliable causal networks.

[0082] In one possible implementation, S407 specifically includes:

[0083] S4071: Define the lag set and the condition set.

[0084] Specifically, directed edges are used to characterize the causal relationship between nodes. For any directed edge, the node corresponding to its starting point is defined as the source node, which represents the source of information that has a potential causal effect on other nodes, and the node corresponding to its ending point is defined as the target node, which represents the result variable that is affected by the historical state of the source node at the current moment.

[0085] For each candidate directed edge in the set of directed causal candidate edges, taking the corresponding source node and target node as the analysis objects, the state variables of the source node at historical moments are expanded according to a preset time lag order to form the lag set of the source node, which is used to characterize the historical information of the source node's possible influence on the target node. At the same time, the historical lag terms of the target node itself, as well as other nodes in the sparse precision matrix that have a direct conditional dependency on the current state of the target node and their corresponding historical lag terms, together form a condition set, which is used to characterize the minimum conditional constraint environment after controlling for the influence of relevant variables, thereby providing a unified information structure foundation for the subsequent calculation of conditional transition entropy.

[0086] S4072: Define the conditional transition entropy based on the lag set, the condition set, and the sparse precision matrix.

[0087] Specifically, given that the lag set and the condition set are already determined, the current item of the target node, the lag item of the source node, and the variables in the condition set are uniformly mapped to the variable index space corresponding to the sparse precision matrix. The local sub-blocks in the sparse precision matrix related to the current item of the target node and the lag item of the source node are used to characterize the degree of change in the uncertainty of the target node before and after introducing the historical information of the source node under the given condition set constraints. Thus, the conditional transition entropy is defined through the algebraic operation of the local sub-blocks of the precision matrix.

[0088] Let sparse precision matrix J The four sub-matrices obtained after reordering the variables according to their roles in the causal relationship are as follows: , , as well as The conditional transition entropy is specifically:

[0089]

[0090] in, T ( ) represents the conditional transition entropy. Indicates the source node, Z3 represents the target node, Z3 represents the condition set, and log represents the logarithmic function. This represents the precision sub-block corresponding to the current state of the target node, used to characterize the original uncertainty of the target node without introducing any historical or conditional variables. This represents a joint precision sub-block consisting of historical lag terms from the source node and variables from the condition set, used to characterize the condition dependency structure among all known condition variables. and These represent the cross-precision sub-blocks between the current state of the target node and its history and condition variables, respectively. Represents a determinant. -1 This represents the inverse of a matrix.

[0091] S408: Based on the conditional transition entropy, filter the set of directed causal candidate edges to determine the set of significant edges for TE.

[0092] Specifically, the conditional transition entropy value is compared with a preset significance criterion to screen out directed edges that have a statistically significant information transmission effect. The directed edges that meet the significance condition are then summarized to form the TE significant edge set. The TE significant edge set is used to characterize the directed relationships between nodes that still have a direct causal effect after excluding spurious correlations and indirect effects.

[0093] The significance criteria include at least one of the following: (1) The conditional transfer entropy value is not less than a preset threshold. (2) The significance test result constructed based on the conditional transfer entropy meets the preset significance level. (3) Under different raw material ratio conditions or reaction trajectory segmentation conditions, the conditional transfer entropy meets the significance requirements in most conditions.

[0094] S409: The causal network is obtained by fusing the set of directed causal candidate edges, the set of conditional dependency candidate edges, and the set of TE significant edges, and combining them with the set of nodes.

[0095] Specifically, using the node set defined by S401 as the node basis of the causal network, the directed causal candidate edge set obtained based on the Rank-VAR model, the conditional dependency candidate edge set determined based on the sparse precision matrix, and the TE significant edge set obtained based on conditional transfer entropy screening are jointly constrained and fused. Only directed edges that simultaneously satisfy the requirements of directional significance, conditional dependency existence, and significant information transmission effect are retained, and the directed edges are mapped to the corresponding node pairs in the node set as the edge set of the causal network, thereby constructing the causal network.

[0096] S5: Construct a causal constraint reinforcement learning environment based on causal networks.

[0097] In this embodiment of the invention, by constructing a causal constraint reinforcement learning environment based on a causal network, the causal relationship between changes in raw material ratio and co-pyrolysis reaction behavior and kinetic response can be explicitly embedded into the reinforcement learning process. This allows policy learning to occur only within the scope of state transitions and actions that conform to the identified causal mechanism, thereby effectively reducing ineffective exploration and unreasonable decisions, and improving the convergence efficiency, stability, and physical interpretability of the ratio optimization strategy.

[0098] Specifically, the causal network constructed in the above steps can be used to explicitly constrain the action space and state transition path of subsequent raw material ratio optimization strategies, avoiding invalid explorations that violate the co-pyrolysis reaction mechanism during strategy learning. This embodiment of the invention uses the causal network as a structural prior for the reinforcement learning environment, uniformly constraining and modeling the environment's state space, action space, state transition relationships, and reward feedback mechanism. Specifically, the state value of each node in the causal network at the current moment is used as the state variable of the reinforcement learning environment, representing the joint state of the raw material ratio state, co-pyrolysis reaction behavior characteristics, and kinetic response characteristics. The raw material ratio adjustment amount or direction is defined as the action of the reinforcement learning environment, and based on the directed causal edges from the ratio variable nodes to the reaction characteristic nodes and kinetic parameter nodes in the causal network, causal consistency constraints are applied to the selectable range of actions and their influence paths. Meanwhile, during the interaction with the environment, state transitions are only allowed to occur along the causal transmission paths determined in the causal network. When constructing rewards, the reaction performance index and kinetic advantage index are assigned according to the effective paths in the causal network. This results in a causal constraint reinforcement learning environment that is constrained by the causal network in terms of state evolution, action, and reward feedback. This environment ensures that the subsequent policy learning process conforms to the identified causal mechanism of "change in raw material ratio - evolution of reaction behavior - kinetic response characteristics".

[0099] S6: Based on the causal constraint reinforcement learning environment, the optimization strategy for the ratio of co-pyrolysis raw materials is determined by the temporal difference reinforcement learning algorithm.

[0100] In this embodiment of the invention, by employing a temporal difference reinforcement learning algorithm for policy learning in a causal constraint reinforcement learning environment, it is possible to fully utilize sample interaction information while taking into account both long-term benefits and immediate feedback. This allows the raw material ratio optimization strategy to gradually converge to a better-performing adjustment scheme under the premise of satisfying causal consistency constraints, thereby improving the stability, efficiency, and practical feasibility of the optimization results.

[0101] In one possible implementation, S6 specifically includes:

[0102] S601: In a causal constraint reinforcement learning environment, initialize the behavior policy Q function, the value estimation Q function, and the learning parameters, including the learning rate parameter, the discount factor, and the policy temperature parameter.

[0103] The behavioral policy Q function (actQ) characterizes the behavioral tendency of a reinforcement learning agent to choose various possible actions given a state in the current learning phase. Its value reflects the immediate evaluation result of taking the corresponding action under the state-action pair, and is mainly used to guide sample collection and actual action selection. The value estimation Q function (valQ) estimates the expected long-term reward that can be obtained by executing the current or target policy under a given state-action pair. Its value is used to evaluate the long-term value of the action and serves as an important basis for policy updates and convergence determination.

[0104] S602: Generate the Boltzmann policy based on the initialized behavior policy Q function.

[0105] It should be noted that the Boltzmann strategy, also known as the soft maximum strategy, is a stochastic strategy that transforms the value of the Q-function of the behavioral policy into a probability distribution of action selection through an exponential mapping. Its characteristic is that while prioritizing high-value actions, it still retains a certain probability of exploring other actions, thereby achieving a balance between exploration and exploitation.

[0106] Specifically, after initializing the behavioral policy Q-function, the values ​​of the behavioral policy Q-function under each state-action pair are used as the basis for action value evaluation. An exponential mapping and normalization process is used to construct the probability distribution of action selection, thereby generating the corresponding Boltzmann policy. The Boltzmann policy assigns higher selection probabilities to high-value actions while reserving a certain selection probability for lower-value actions. This ensures policy stability while maintaining the ability to explore the space for adjusting raw material ratios, achieving a balance between exploration and utilization. Its policy form can be expressed as:

[0107]

[0108] in, This represents the Boltzmann policy, i.e., at time... In a state When selecting raw material ratio adjustment action The probability, e Represents an exponential function. Indicates the strategy temperature parameter. Indicates time The behavioral policy Q function in state-action pairs The value at that location, b Represents the state The candidate action index below, Indicates the current state. Indicates at time t The behavioral policy Q function in the state With action The value that the corresponding state-action pair can take.

[0109] S603: Under the Boltzmann strategy, sample collection is performed to obtain a sample set.

[0110] In this embodiment of the invention, by collecting samples under the Boltzmann strategy, it is possible to ensure that the adjustment of the proportion of high-value raw materials is explored first, while avoiding premature convergence of the strategy due to over-reliance on a single proportion scheme. This results in a more comprehensive coverage of the co-pyrolysis reaction state space in the statistical sense of the samples, thereby improving the stability and reliability of subsequent strategy evaluation and optimization results.

[0111] Specifically, under the Boltzmann policy constraint, the raw material ratio adjustment action is selected probabilistically based on the current state, and the action is executed in a causal constraint reinforcement learning environment to obtain the corresponding immediate reward and state transition result. During the sample collection process, the state, sampling action, immediate reward, and next state at each time step are recorded, and the above interaction samples are continuously accumulated within a sample batch period to form a sample set used to characterize the interaction between the current strategy and the co-pyrolysis reaction process.

[0112] In one possible implementation, S603 specifically includes:

[0113] S6031: Sample the actions according to the Boltzmann strategy to obtain sampled actions.

[0114] S6032: Perform the sampling action to obtain an immediate reward and the next state.

[0115] It should be noted that the immediate reward refers to the immediate feedback quantity generated by the change in the reaction process caused by the adjustment of the co-pyrolysis feedstock ratio after performing the sampling action in the current state. It is used to characterize the immediate impact of this ratio adjustment on the performance and kinetics of the co-pyrolysis reaction. The immediate reward can consist of factors such as changes in co-pyrolysis product yield, energy efficiency improvement, reaction stability indicators, or risk penalties, and is comprehensively reflected in scalar form to reflect the immediate contribution of the current action to the optimization objective. The next state refers to the new state evolved by the system under the constraints of the co-pyrolysis reaction mechanism and operating conditions after performing the sampling action in the current state. It is used to describe the updated results of the co-pyrolysis reaction characteristics, kinetic parameters, and related state variables after the feedstock ratio adjustment. The next state consists of the actual evolution results of the reaction process or a set of state variables obtained based on sample statistics and model predictions, and serves as the state input for subsequent action selection and value updates.

[0116] For example, in the co-pyrolysis feedstock ratio optimization process of this invention embodiment, the immediate reward can be represented as the performance feedback directly generated by the co-pyrolysis reaction after the current feedstock ratio adjustment action is executed. This includes, for example, the increase in the yield of the target product (such as combustible gas or oil), the decrease in unit energy consumption, the improvement in reaction stability, or the penalty applied when entering unfavorable operating conditions such as high tar or high coking. The next state represents the new state evolved by the system under the constraints of the pyrolysis reaction mechanism and operating conditions after the ratio adjustment action is executed. It can be composed of the updated feedstock ratio, the thermal reaction characteristics within the reaction temperature range, the estimated kinetic parameters, and the values ​​of relevant nodes in the causal network. This state is used to fully characterize the impact of this action on the co-pyrolysis reaction process and the system's operating state, and serves as the state input for subsequent strategy decisions and value updates.

[0117] S6033: Combine the sampling action, immediate reward, next state, and current state to obtain the sample set.

[0118] Specifically, in the sample collection process of this invention, at each time step, the sampling action selected according to the Boltzmann strategy under the current state is recorded, and the corresponding immediate reward and next state are obtained after the sampling action is executed. The current state, sampling action, immediate reward, and next state together constitute a state-action-reward-state transition sample, and the above samples are continuously recorded and accumulated within a sample batch cycle to form a sample set. The sample set is used to characterize the actual interaction results between the raw material ratio adjustment behavior and the co-pyrolysis reaction process under the current fixed strategy conditions.

[0119] S604: Update the initialized value estimation Q function based on the sample set:

[0120]

[0121] in, express t The value estimation Q-function at time +1, i.e., the updated value estimation Q-function. express t The Q-function for estimating the value at time t, express t The state at any given moment, Representing state The sampling action below, This represents the learning rate parameter. Indicates the discount factor. express t Instant rewards for each moment Indicates by t In the probability distribution of the policy determined by the behavior policy at a given time, in the state... Select action b The probability, express t+ The state at time 1 b Indicates the sampling action index.

[0122] In this embodiment of the invention, by updating the initialized value estimation Q function based on the sample set, the state-action-reward-state transition information collected in batches can be fully utilized under fixed policy conditions to smooth the random fluctuations caused by a single sampling, thereby obtaining a more stable and reliable value estimation result. This provides an accurate value reference for subsequent policy evaluation and behavioral policy Q function updates, and effectively improves the convergence and robustness of the overall optimization process.

[0123] S605: Calculate the average reward and average transfer for the sample set:

[0124]

[0125] in, Represents state-action pairs in a sample set. Average instant reward, This represents the state-action pair within a sample batch. The corresponding cumulative reward value, This represents the normalization factor, used to avoid division by zero, and is generally taken as... , This indicates that in the sample set, from state Through action Transition to state The average transition probability, This indicates that within a sample batch, the triplet The number of times it appears, Indicates the next state.

[0126] S606: Construct sample-batch time-series difference error based on average reward, average transfer, and the updated value estimation Q function.

[0127] In one possible implementation, S606 specifically includes:

[0128] S6061: Construct the next-state maximum value term for the sample set based on the average transition and the updated value estimation Q function.

[0129]

[0130] in, This represents the maximum value item in the next state of the sample set. max This indicates maximization.

[0131] S6062: Construct sample-batch temporal difference error based on the maximum value term of the next state:

[0132]

[0133] in, Expression – Action Pair Sample-batch time-series difference error, The Q-function represents the behavioral policy.

[0134] S607: Update the initialized behavior strategy Q function based on the sample-batch time-series difference error.

[0135] Specifically, while maintaining the original behavioral strategy Q function, the sample-batch temporal difference error is gradually accumulated into the corresponding behavioral strategy Q value according to the preset learning rate. This enables the forward raw material ratio adjustment action to effectively reflect the comprehensive impact on subsequent states and immediate rewards under the statistical significance of the samples, thereby prompting the behavioral strategy Q function to gradually converge towards the direction of raw material ratio adjustment with higher returns and controllable risks.

[0136] It should be noted that those skilled in the art can set the preset learning rate according to actual needs, and this invention does not limit this.

[0137] S608: Update the Boltzmann policy based on the updated behavior policy Q function.

[0138] Specifically, after updating the Q-function of the behavioral policy, the updated Q-value is used as the new basis for evaluating the value of actions. The raw material ratio adjustment actions in each state are re-weighted, and a new action selection probability distribution is constructed through exponential mapping and normalization, thereby updating the original Boltzmann policy. This update process ensures that ratio adjustment actions with higher value in the Q-function are assigned higher selection probabilities in subsequent decisions, while retaining necessary exploration probabilities for actions with lower value. This allows the policy to dynamically evolve with the learning process, gradually converging towards a more profitable and risk-controlled raw material ratio optimization direction.

[0139] S609: Determine the optimization strategy for the co-pyrolysis feedstock ratio based on the updated Boltzmann strategy.

[0140] Specifically, after the behavior policy Q function is updated and the corresponding Boltzmann policy is generated, the Boltzmann policy is used as the raw material ratio decision rule. Under a given co-pyrolysis state, the selection probability of each candidate raw material ratio adjustment action is calculated according to the policy, and the final raw material ratio adjustment scheme to be executed is determined based on the selection probability.

[0141] For example, the Boltzmann strategy can be represented as an exponential probability mapping function with the behavioral policy Q function as input. By exponentially weighting and normalizing the Q values ​​corresponding to different actions, it ensures that the actions of adjusting the raw material ratio with higher value are selected with a higher probability, while the actions of lower value still retain a certain probability of being selected. This ensures the optimization effect while avoiding the strategy from converging to a local optimum too early, and ultimately forms a co-pyrolysis raw material ratio optimization strategy.

[0142] S7: Conduct a reversibility assessment of the co-pyrolysis feedstock ratio optimization strategy and determine the final strategy.

[0143] In this embodiment of the invention, by evaluating the reversibility of the co-pyrolysis raw material ratio optimization strategy, it is possible to obtain better optimization benefits while constraining the reversibility and correctability of the strategy in the actual execution process. This effectively avoids the uncontrollable state of the system due to strategy errors or uncertain disturbances, thereby significantly improving the safety, stability and engineering feasibility of the strategy implementation process while ensuring the optimization effect of co-pyrolysis reaction performance.

[0144] In one possible implementation, S7 specifically includes:

[0145] S701: Calculate the reversibility index of the co-pyrolysis raw material ratio optimization strategy based on the sample average transition probability and backoff accessibility constraints.

[0146] In this embodiment of the invention, since the co-pyrolysis raw material ratio optimization strategy drives the evolution of the system state through continuous raw material ratio adjustment during actual execution, and the state evolution process has uncertainty and path dependence characteristics in the statistical sense of the sample, when the strategy guides the system into an unfavorable reaction range or a high-risk operating condition, the lack of effective backtracking ability will make the optimization process difficult to correct or even uncontrollable. Therefore, it is necessary to quantitatively evaluate the backtrackability and correctability of the optimization strategy during execution.

[0147] To this end, embodiments of the present invention introduce a fallback reachability constraint. The fallback reachability constraint is used to limit whether, after performing any forward raw material ratio adjustment action, under the premise of satisfying causal constraints and operating condition boundary conditions, there still exists a feasible control path that can guide the system state from the evolved state back to the original state or its allowed neighborhood. The fallback control path consists of a set of fallback actions. The set of fallback actions is a set of feasible correction actions determined after performing a forward action, based on the current state, the adjustable range of raw material ratio, and causal constraints, through rule mapping, sample transfer statistical analysis, or constraint projection.

[0148] Based on this, firstly, using the average transition probability obtained from the sample set statistics, the probability distribution characteristics of the system state evolving to each possible next state after the optimization strategy performs the forward raw material ratio adjustment action are characterized. Then, combining the backoff reachability constraints and the corresponding backoff action set, for each possible next state reached after the forward action, the reachability and backoff accuracy of the system to return to the original state or its allowed neighborhood through allowed backoff actions are evaluated, under the premise of satisfying causal constraints and operating boundary conditions. Finally, the action selection probabilities of the optimization strategy in different states are weighted and summarized to construct a reversibility index to characterize the overall backoffability, correctability, and execution controllability of the co-pyrolysis raw material ratio optimization strategy in the statistical sense of the samples.

[0149]

[0150] in, This indicates the optimization strategy for co-pyrolysis feedstock ratio. The reversibility index, Indicates the state The expectation operator, Indicates the state Below, optimization strategy for co-pyrolysis feedstock ratio Select sampling action The probability, Indicates a back motion. Indicates sampling action In state The corresponding set of rollback actions, Indicates from state Execute rollback action Back to normal The sample average transition probability, Indicates the state Execute rollback action The expected state that may be achieved later. Indicates the backoff residual distance. This represents the backoff precision penalty coefficient. e This represents an exponential function.

[0151] S702: Determine if the reversibility index is greater than or equal to the threshold. If yes, determine the co-pyrolysis raw material ratio optimization strategy as the final strategy. Otherwise, adjust the learning parameters of the temporal differential reinforcement learning algorithm and return to step S6 until the reversibility index is greater than or equal to the threshold.

[0152] Specifically, when the reversibility index is greater than or equal to the threshold, the strategy is considered to have sufficient backtracking capability and controllability in the statistical sense of the samples, and thus the co-pyrolysis raw material ratio optimization strategy is determined as the final strategy. When the reversibility index is less than the threshold, it indicates that the current strategy has insufficient backtracking capability or risk accumulation problems during execution. In this case, the learning parameters of the temporal difference reinforcement learning algorithm are adjusted, and the process returns to step S6 to relearn and update the strategy until a co-pyrolysis raw material ratio optimization strategy with a reversibility index that meets the threshold requirement is obtained.

[0153] S8: Execute the final policy.

[0154] In summary, this invention constructs a causal constraint reinforcement learning framework to model the co-pyrolysis raw material ratio adjustment process as a learnable and feedback-enabled decision problem. Temporal difference reinforcement learning is used to gradually form a raw material ratio optimization strategy through sample interactions. Furthermore, a reversibility evaluation mechanism based on the average transition probability of samples and backoff reachability constraints is introduced to perform closed-loop correction of the strategy's execution controllability. This effectively suppresses the risk of irreversibility while ensuring optimization benefits, achieving a synergistic balance between performance improvement, safety, and stability in the co-pyrolysis raw material ratio optimization strategy. This significantly enhances the reliability and applicability of the method in practical engineering applications.

[0155] Reference manual attached Figure 2 The diagram shows a schematic of the structure of an intelligent optimization system for co-pyrolysis raw material ratio provided in an embodiment of the present invention.

[0156] This invention provides an intelligent optimization system 20 for co-pyrolysis raw material ratio, comprising: a processor 201 and a memory 202;

[0157] The memory 202 stores programs or instructions that can run on the processor 201. When the program or instructions are executed by the processor 201, they implement the steps of the above-mentioned intelligent optimization method for co-pyrolysis raw material ratio and achieve the same technical effect. To avoid repetition, the present invention will not elaborate further.

[0158] It should be understood that the processor 201 in this embodiment of the invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0159] It should also be understood that the memory 202 in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DR RAM).

[0160] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0161] It should be understood that, in various embodiments of the present invention, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0162] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0163] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0164] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0165] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0166] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0167] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0168] This invention provides a readable storage medium comprising: storing a program or instructions on the readable storage medium, wherein when the program or instructions are executed by a processor, the program or instructions implement the steps of the above-described intelligent optimization method for co-pyrolysis raw material ratio, and can achieve the same technical effect. To avoid repetition, this invention will not elaborate further.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the protection scope of the present invention.

Claims

1. A method for intelligent optimization of the ratio of co-pyrolysis raw materials, characterized in that, include: S1: Obtain raw material data; S2: Based on the raw material data, construct a characterization dataset of co-pyrolysis reaction under different ratio conditions, wherein the characterization dataset of co-pyrolysis reaction includes multi-scale reaction features and reaction trajectory features; S3: Construct a kinetic feature dataset based on the described reaction trajectory characteristics; S4: Based on the pyrolysis reaction characterization dataset and the kinetic characteristic dataset, construct a causal network for the pyrolysis feedstock ratio; Specifically, S4 includes: S401: Based on the pyrolysis reaction characterization dataset and the kinetic characteristic dataset, set a set of nodes for constructing a causal network; S402: The pyrolysis reaction characterization dataset and the kinetic characteristic dataset are concatenated to obtain a multivariate time series feature set; S403: Perform a rank transformation on the multivariate time series feature set to obtain a rank vector; S404: Construct a Rank-VAR model based on the rank vector; Specifically, the Rank-VAR model is as follows: ; in, Represents a rank vector. This indicates that the characteristics of a multivariate time series are concentrated at time 1. t eigenvectors, In the Rank-VAR model, the first... l The regression coefficient matrix with lag order, L This represents the maximum lag order of the Rank-VAR model. This indicates that the characteristics of a multivariate time series are concentrated at time 1. t-l eigenvectors, express t The random perturbation term of the Rank-VAR model at time step; S405: Construct a directed causal candidate edge set using the Rank-VAR model; S406: Perform sparse inverse covariance estimation on the rank vector to determine the sparse precision matrix and the set of conditionally dependent candidate edges; S407: Define the conditional transition entropy of the directed causal candidate edge set based on the sparse precision matrix; S408: Based on the conditional transition entropy, the set of directed causal candidate edges is filtered to determine the set of significant TE edges; S409: The directed causal candidate edge set, the conditional dependency candidate edge set, and the TE significant edge set are fused together with the node set to obtain the causal network; S5: Based on the causal network, construct a causal constraint reinforcement learning environment; S6: Based on the causal constraint reinforcement learning environment, determine the optimization strategy for the ratio of co-pyrolysis raw materials through the temporal difference reinforcement learning algorithm; S7: Perform a reversibility assessment on the co-pyrolysis raw material ratio optimization strategy and determine the final strategy; S8: Execute the final strategy.

2. The intelligent optimization method for co-pyrolysis raw material ratio according to claim 1, characterized in that, S3 specifically includes: S301: Based on the aforementioned reaction trajectory characteristics, the co-pyrolysis process is divided into stages to determine multiple reaction stages; S302: Within each of the aforementioned reaction stages, a Coats–Redfern kinetic model is constructed based on the characteristics of the reaction trajectory; S303: Input the reaction trajectory features into the Coats–Redfern kinetic model to obtain the kinetic parameters corresponding to each of the reaction stages; S304: Construct the dynamic feature dataset based on the dynamic parameters.

3. The intelligent optimization method for co-pyrolysis raw material ratio according to claim 1, characterized in that, Specifically, S405 includes: S4051: Perform multivariate least squares estimation on the Rank-VAR model to obtain a set of estimated coefficient matrices; S4052: Construct a single-coefficient significance test statistic based on the set of estimated coefficient matrices; S4053: Based on the single-coefficient significance test statistic, set the lag-level directed edge indicator; S4054: Determine the node-level directed edge indication matrix based on the hysteresis-level directed edge indication quantity; S4055: Construct the set of directed causal candidate edges based on the node-level directed edge indication matrix.

4. The intelligent optimization method for co-pyrolysis raw material ratio according to claim 1, characterized in that, Specifically, S407 includes: S4071: Define the lag set and the condition set; S4072: Define the conditional transition entropy based on the hysteresis set, the condition set, and the sparse precision matrix.

5. The intelligent optimization method for co-pyrolysis raw material ratio according to claim 1, characterized in that, S6 specifically includes: S601: In the causal constraint reinforcement learning environment, initialize the behavior policy Q function, the value estimation Q function, and the learning parameters, wherein the learning parameters include: learning rate parameter, discount factor, and policy temperature parameter; S602: Generate the Boltzmann policy based on the initialized behavior policy Q function; S603: Under the Boltzmann strategy, sample collection is performed to obtain a sample set; S604: Update the initialized value estimation Q function based on the sample set; S605: Calculate the average reward and average transfer of the sample set; S606: Construct sample-batch temporal difference error based on the average reward, the average transfer, and the updated value estimation Q function; S607: Update the initialized behavior strategy Q function based on the sample-batch time-series difference error; S608: Update the Boltzmann policy according to the updated behavior policy Q function; S609: Determine the optimization strategy for the co-pyrolysis feedstock ratio based on the updated Boltzmann strategy.

6. The intelligent optimization method for co-pyrolysis raw material ratio according to claim 5, characterized in that, Specifically, S603 includes: S6031: Sample the actions according to the Boltzmann strategy to obtain sampled actions; S6032: Perform the sampling action to obtain an immediate reward and the next state; S6033: Combine the sampling action, the immediate reward, the next state, and the current state to obtain the sample set.

7. The intelligent optimization method for co-pyrolysis raw material ratio according to claim 5, characterized in that, Specifically, S606 includes: S6061: Based on the average transition and the updated value estimation Q function, construct the next state maximum value term of the sample set; S6062: Construct the sample-batch time-series difference error based on the maximum value term of the next state.

8. The intelligent optimization method for co-pyrolysis raw material ratio according to claim 1, characterized in that, Specifically, S7 includes: S701: Calculate the reversibility index of the co-pyrolysis raw material ratio optimization strategy based on the sample average transition probability and backoff reachability constraints; S702: Determine whether the reversibility index is greater than or equal to the threshold; if so, determine the co-pyrolysis raw material ratio optimization strategy as the final strategy; otherwise, adjust the learning parameters of the temporal differential reinforcement learning algorithm and return to step S6 until the reversibility index is greater than or equal to the threshold.

9. A smart optimization system for the proportioning of co-pyrolysis raw materials, characterized in that, include: Processor and memory; The memory stores programs or instructions that can run on the processor, and when the program or instructions are executed by the processor, they implement the steps of the intelligent optimization method for co-pyrolysis raw material ratio as described in any one of claims 1 to 8.