A complex system control decision method and system fusing wrt and spt

By integrating WRT and SPT into a complex system control decision-making method, a pre-decision table is constructed and real-time control decisions are implemented. This solves the problems of cognitive limitations and poor interpretability of deterministic causal driving models in complex system research, and achieves efficient and reliable decision support.

CN120373906BActive Publication Date: 2026-01-27NANJING NARI GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510497823.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-04-23
Filing Date
2025-04-21
Publication Date
2026-01-27
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

In existing technologies, models that rely solely on deterministic causal driving forces suffer from limitations in cognition, difficulty in quickly and accurately identifying model parameters, mismatch between model predictions and actual results, poor interpretability, difficulty in ensuring learning quality under small sample sizes, and low credibility of analysis results, making the models unsuitable for large-scale problems.

Method used

By integrating Whole Reduction Theory (WRT) and Symbol String Pre-training (SPT), a pre-decision table containing multiple sets of mappings is constructed by pre-training the symbol string information of the target system and its objective environment. The self-attention mechanism is used to optimize control decisions and implement control decisions under real-time changes.

Benefits of technology

It enables fast, reliable, and adaptive decision recommendations in high-dimensional complex systems, ensuring the interpretability and reliability of causal analysis. By optimizing pre-training through reinforcement learning, it approximates the states and patterns of the real world, thereby improving the performance and accuracy of analysis and decision-making models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373906B_ABST
    Figure CN120373906B_ABST
Patent Text Reader

Abstract

A complex system control decision method and system fusing WRT and SPT, the method comprising two-layer architecture: the upper layer classifies the rough granulation of uncertain factors according to the structural information and language and other non-structural information of each related field inside and outside the system, and provides the information required for generating training samples to the causal algorithm of the lower layer, including the corresponding prediction working condition, disturbance event and control measures available at the time of each learning sample example, and efficiently organizes the training process of deep reinforcement learning in high-dimensional decision space; the lower layer adopts a causal-driven deterministic analysis algorithm based on WRT to complete the mapping from system state to control decision; the two layers cooperate to realize the pre-training of decision, and periodically update the pre-decision table according to the results. Once the actual disturbance is detected and identified in the objective system, it is matched with the pre-decision table to quickly complete the control decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information science and decision support, specifically relating to a complex system control decision-making method that integrates holistic reductionism (WRT) and symbol string pre-training (SPT) technology. Background Technology

[0002] Research on complex systems such as energy, electricity, transportation, and the environment needs to consider multiple dimensions of information, physics, and society. The research objects are characterized by strong time-varying, strong nonlinearity, and strong uncertainty. Therefore, it is urgent to achieve accurate and rapid quantitative research through the deep integration of digital and intelligent technologies with domain knowledge.

[0003] Currently, research methods that rely solely on deterministic causal driving models, while possessing strong interpretability, are prone to limitations in the input, output, and reasoning aspects of the models, leading to situations where they are detached from reality and ignore the overall picture. Mismatches between model predictions and actual results are common. Model parameters are sometimes difficult to identify quickly and accurately, and it is also difficult to quickly correct models after they are falsified, resulting in analytical conclusions that cannot keep up with the rapid evolution of the system. Furthermore, the models only possess mechanistic interpretability at low dimensions, and this characteristic cannot be maintained at high dimensions.

[0004] In summary, research methods that rely solely on deterministic causal driving models face difficulties in cognition, parameterization, and mechanistic analysis, which limit the scope of applicable problems and cannot guarantee solutions for arbitrary problem sizes.

[0005] In recent years, purely data-driven models, especially machine learning models, have made rapid progress. Generative large-scale model technologies such as ChatGPT and Sora have shown breakthrough development, and the dawn of Artificial General Intelligence (AGI) is emerging. AI for Science (AI4S) is gradually becoming a new research paradigm, and the development of digital intelligence technology is ushering in new opportunities. However, such models usually have poor interpretability, difficulty in ensuring learning quality under small sample sizes, and low reliability of analysis results. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a control decision-making method and system for complex systems that integrates WRT and SPT. This invention pre-trains the symbol string information of the target system and its objective environment to construct a pre-decision table containing multiple sets of mappings. During transfer execution, it finds and implements the corresponding control decision based on the real-time changing feature vector of the target system.

[0007] The first aspect of the present invention provides 1. a control decision method for complex systems integrating WRT and SPT, characterized by comprising the following steps:

[0008] In response to the received symbol string information of the target system and its objective environment, the decision model integrating WRT and SPT is pre-trained:

[0009] The pre-training includes: mapping the uncertainties contained in the symbol string to coarse-grained tiers, prioritizing the analysis queue according to the self-attention mechanism for the representative states in each tier, optimizing the control decision of each item in the queue based on the deterministic analysis of the overall restoration, completing the one-to-one mapping from the predicted quantity of the observable quantity to the control decision, and constructing a pre-decision table containing multiple sets of mappings.

[0010] When the decision model is transferred and executed, it responds to the real-time changes in the feature vector of the target system, and finds and implements the corresponding control decision by matching it with the condition feature vector in the pre-decision table.

[0011] A second aspect of the present invention provides a control decision-making method for complex systems that integrates WRT and SPT, comprising the following steps:

[0012] Collect raw corpora of the target system and its objective environment, and extract symbol string information;

[0013] Construct a decision model that integrates WRT and SPT. The SPT layer is the upper layer of the decision model and is used for uncertainty analysis, while the WRT layer is the lower layer of the decision model and is used for deterministic analysis.

[0014] Pre-training the SPT layer includes: mapping the uncertainty factors contained in the symbol string to coarse-grained classification, constructing an event table containing knowledge of the uncertain environment, and generating an event queue that needs to be pre-analyzed.

[0015] Pre-training the WRT layer includes: the upper SPT layer pushes the event table to the WRT layer, optimizes the corresponding control decisions based on causal deterministic analysis for representative states in each segment, and generates a pre-decision table containing quantitative knowledge of deterministic domain.

[0016] The transfer execution of the decision model responds to the real-time changes in the target system's feature vectors. By matching these feature vectors with the conditional feature vectors in the pre-decision table, the corresponding control decisions are found and implemented.

[0017] Preferably, the target system includes: a power system, an energy system, an economic system, or an environmental system;

[0018] The objective environment includes: a collection of other systems that interact with the target system;

[0019] Symbol strings include sequences of bytes, numbers, characters, or words extracted from the original corpus; the original corpus includes text, sound, images, pictures, numbers, or a combination of these in the form of a document library or database.

[0020] Preferably, the pre-training of the SPT layer specifically includes:

[0021] By mapping the uncertainties contained in the symbol string to coarse-grained subdivisions;

[0022] The representative state in each tier is determined based on the control decision type to be trained in the WRT layer deterministic analysis.

[0023] Based on the obtained hierarchical representation of observables, an event table containing knowledge of uncertain environmental conditions is constructed.

[0024] For each representative state in each tier, the events in the event table are prioritized according to the self-attention mechanism to generate an event queue.

[0025] Preferably, the step of mapping the uncertainty factors contained in the symbol string to coarse-grained subdivisions includes:

[0026] For qualitative descriptions of observables in a symbol string, quantitative tiered representations are obtained by querying a semantic database;

[0027] For the uncertain quantitative representation of observable quantities in a symbol string, a probability density function is used;

[0028] For uncertain quantitative representations involving multiple observables changing simultaneously in a symbol string, a joint probability density function is used. For the representation in the form of probability density function, a classification operation should be performed with reference to the semantic library settings or the knowledge learned from the function itself. A one-to-one or one-to-many mapping from an observable to one or more observables is defined through the semantic library.

[0029] Preferably, determining the representative state in each tier based on the control decision type to be trained in the WRT layer deterministic analysis includes:

[0030] For preventive control decisions that need to be applied before a risk event occurs, representativeness refers to the expected value of the observable quantity for all samples within that tier.

[0031] For emergency control decisions that need to be applied after a risk event occurs, the representative state refers to the extreme value of all samples in the most unfavorable direction within the observable range.

[0032] The representative state value method should be specified in advance through the semantic library.

[0033] Preferably, the step of prioritizing events in the event table according to a self-attention mechanism and generating an event queue for representative states in each tier includes:

[0034] Estimate the pre-control risk value based on the representative state value of each event in the coarse-grained classification value in the event table;

[0035] Sort the event table in descending order according to the pre-control risk value to generate a queue of events that need to be pre-analyzed;

[0036] We will prioritize the events that are ranked first, i.e. the high-risk events, and conduct the aforementioned deterministic analysis based on causality one by one.

[0037] Even when parallel computing capabilities are available, it is still necessary to conduct the aforementioned deterministic analysis based on causal drive for each batch of high-risk events.

[0038] Preferably, the pre-training of the WRT layer specifically includes:

[0039] Simulation simulation: used to simulate the operation of a system under a given system state and decision, within a given time range, based on a simulation model that reflects the objective laws of the target system;

[0040] Knowledge extraction: Used to extract quantitative knowledge from simulation parameter trajectories that is beneficial to achieving research objectives;

[0041] Decision support: For specific decision problems, under a given system state, it explores the values ​​of relevant decision parameters in a simulation model, then calls back the simulation and knowledge extraction to generate quantitative analysis results of the target system characteristics under the trial decision; through repeated trials and callbacks, it generates multiple rounds of analysis results; based on the results of a single or multiple rounds of analysis, it generates a new next or next set of trial directions; it repeats the above process until the trial no longer produces better results; the best-performing decision, and the feature vector of the initiation condition of that decision, are stored as a record in the pre-decision table.

[0042] Preferably, the pre-training of the WRT layer further includes:

[0043] To enhance the decision-making model with intelligence, which can be used for the falsification and correction of the decision-making model.

[0044] Preferably, the simulation specifically includes:

[0045] In response to the received event table, the simulation model required for deterministic analysis is obtained from the SPT layer;

[0046] In real-time mode, data reflecting the current state of the target system is substituted into the simulation model; in non-real-time mode, data reflecting the target system at past or assumed future times are substituted into the simulation model.

[0047] For each event record to be trained, the hierarchical state of each observable related to the event record is substituted and mapped to the corresponding simulation model parameters to complete the modification of the simulation model.

[0048] Based on the modified model data, causal simulations are conducted to generate simulation parameter trajectories within a given time range.

[0049] Preferably, the knowledge extraction specifically includes:

[0050] A series of planar orthogonal pattern libraries are built based on domain expert knowledge, and each pattern contains only 2 parameters;

[0051] Based on the settings of each mode, the simulation model is reduced to order 2 by treating other parameters not involved in that mode as constants (hereinafter referred to as parameters). By traversing all modes, an underdetermined entropy-preserving order reduction mapping matrix of size M*1 is generated, where M is the number of modes.

[0052] For an element of an underdetermined entropy-preserving reduced-order matrix, expand it into a single row;

[0053] Combine all the rows of the patterns to generate a well-defined entropy-preserving reduced-order mapping matrix of size M*N, where M is the number of patterns and N is the number of time steps;

[0054] The linear analytical method is used to analyze each element system of the well-defined entropy-preserving reduced-order matrix, so as to achieve a quantitative expression of a class of characteristics of each element system;

[0055] It aggregates the characteristic expressions of all metasystems, restores and outputs the overall characteristic expression of the target system.

[0056] Preferably, the intelligent enhancement specifically includes:

[0057] Obtain sample results of each observable from the SPT layer;

[0058] Simulation results of observables to be verified are extracted from the results of multiple deterministic cases and recorded to construct a verification sample library;

[0059] When the simulation results of the same observable quantity generally show significant deviations from the observed sample results in the validation sample library, and it is certain that this is not an observation sampling problem, the parameter identification function is activated based on the quantitative analysis knowledge of variable correlation accumulated in the simulation. This guides the generation of new model correction trial directions, which are then fed back to the SPT layer. Through the adjustment of the causal simulation model parameters, the simulation model is corrected, so that the model output continuously approaches the actual behavior of the target system.

[0060] Preferably, during the migration execution, the SPT layer pushes the actual feature vector of the target system to the WRT layer, and the WRT layer queries the pre-decision table for matching control measures and applies them to the target system;

[0061] In response to the reinforcement learning task performed in the SPT layer, the WRT layer also provides deterministic cases to the SPT layer to update or continuously improve the training sample library; or, it will use deterministic cases to support the falsification and correction tasks of the decision model and enrich the validation sample library.

[0062] Preferably, the migration execution specifically includes:

[0063] When the actual feature vector of the target system pushed by the SPT layer is received, it is determined whether the changes in the target system meet the migration execution conditions. If so, the vector is compared with the feature vector stored in each record of the pre-decision table. When the matching degree of the two exceeds the threshold, it is considered that a matching record has been found. The control measures are then extracted from the record and applied to the target system immediately. If no matching record is found, the fallback decision to deal with the problem is searched and then applied to the target system immediately.

[0064] The specific decisions trained to deal with each possible system state are aggregated into a set. Reinforcement learning is then performed based on the data in the set to obtain a general decision that can deal with all possible states.

[0065] Preferably, the reinforcement learning specifically includes:

[0066] In response to reinforcement learning tasks performed at the SPT layer, the WRT layer must be able to serve as a deterministic case inference tool for the SPT layer, simulating environmental responses for each decision trial, generating a sample library that supports reinforcement learning training needs, and feeding it back to the SPT layer.

[0067] In response to the model correction task performed in the SPT layer, the WRT layer must be able to serve as a deterministic case inference tool for the SPT layer, to obtain the corresponding simulation sample for each observable observation sample, and to feed it back to the SPT layer.

[0068] In response to the above reinforcement learning or model correction tasks, the samples returned directly after the simulation inference step are the time response trajectories of observable quantities; the samples returned directly after the quantification analysis step are the quantitative index expressions of the target system under the specified state; and the samples returned directly after the decision support step are the control measures of the target system under the specified state.

[0069] A third aspect of the present invention provides a complex system control decision system that integrates WRT and SPT, and runs a complex system control decision method that integrates WRT and SPT according to the first aspect, comprising: a decision model;

[0070] The decision model includes an SPT layer for uncertainty analysis and a WRT layer for deterministic analysis; the SPT layer is the upper layer of the decision model, and the WRT layer is the lower layer of the decision model; after the decision model is pre-trained, the SPT layer pushes the actual feature vector of the target system to the WRT layer, and the WRT layer queries the pre-decision table for matching control measures and applies them to the target system.

[0071] A fourth aspect of the present invention provides a control decision method for complex systems that integrates WRT and SPT, based on A-WARMAP, to implement adaptive power grid wide-area monitoring, analysis, protection and control. The method includes the following steps:

[0072] Collect raw corpora of A-WARMAP and its objective environment, construct scenario and working condition corpora, and extract symbol string information;

[0073] A decision model integrating WRT and SPT is constructed. The decision model includes an SPT layer and a WRT layer. The SPT layer, as the upper layer of the decision model, is used for uncertainty analysis and pre-training of power system safety margin control decisions. Each example is sorted in the pre-fault table according to the risk level value when no control is applied. The WRT layer, as the lower layer of the decision model, is used for deterministic analysis.

[0074] Pre-training the SPT layer includes: mapping the uncertainty factors contained in the symbol string to coarse-grained classification, constructing an event table containing knowledge of the uncertain environment, and generating an event queue that needs to be pre-analyzed.

[0075] The WRT layer is pre-trained to generate the optimal control decision for each fault in the pre-fault queue, and the results are summarized into a pre-decision table.

[0076] The migration execution WRT queries the pre-decision table based on the actual operating conditions and faults, finds the matching decision, and executes it.

[0077] Preferably, the process of collecting the original corpus of A-WARMAP and its objective environment, constructing a scenario and operational corpus, and extracting symbol string information includes:

[0078] Collect corpora reflecting changes in power system operating conditions and scenarios from any of the multiple information sources, including the Internet, dedicated data networks, or buffer networks connecting the two.

[0079] Preprocessing of the corpus includes: for structured meteorological data, data cleaning, statistics, transformation and integration are completed according to information sources and classifications; for unstructured early warning corpus, it is directly added to the unstructured storage container in the form of strings.

[0080] Preferably, the pre-training of the SPT layer specifically includes:

[0081] By mapping the uncertainties contained in the symbol string to coarse-grained classifications, including: failure probability classification, loss classification, and risk classification;

[0082] Determine the representative status in each sub-category;

[0083] Based on the obtained tiered representation of observables, an event table incorporating knowledge of uncertain environmental conditions is constructed, including: constructing and sorting a pre-fault table; from the calculation results of the above steps, forming a record in the format of "time, line number, fault probability, pre-control loss, pre-control risk, fault type"; merging all records into the pre-fault table, i.e., the event table; sorting the pre-fault table according to pre-control risk from largest to smallest to form a pre-fault queue that needs to be pre-analyzed.

[0084] For each representative state in each tier, the events in the event table are prioritized according to the self-attention mechanism to generate an event queue, forming a pre-fault queue that needs to be pre-analyzed.

[0085] Preferably, the pre-training of the WRT layer specifically includes:

[0086] Simulation and deduction: Under given power system states and decisions, based on a simulation model that reflects the target power system, deduce the system stability behavior within a given time range;

[0087] Knowledge extraction: Extract quantitative knowledge from simulation parameter trajectories that is beneficial to achieving the goal of stability quantitative analysis;

[0088] Decision support: For prevention / emergency / corrective emergency control decision problems, explore the optimal values ​​of relevant decision parameters in the simulation model;

[0089] For the worst possible operating conditions and faults of the target power system, the decisions made at this time are recorded; the preventive / emergency / corrective control decisions trained to deal with each possible system state are summarized into a set, and reinforcement learning is performed based on the data in the set to obtain a general emergency control decision that can deal with all possible states.

[0090] Synchronization of decision tables: Tables in the pre-fault queue that have not yet completed the control strategy update are designated as "tables in progress"; after the update is completed, the corresponding records should be promptly sent to the "online decision table", and finally, through periodic updates, the decision table group on the dispatch side is synchronized to the decision table group on the plant side via the communication system.

[0091] Preferably, the pre-training of the WRT layer further includes:

[0092] Intelligent Enhancement: The SPT layer reads the sample results of observables from WAMS / SCADA; it extracts the simulation inference results of observables to be verified from the inference results of the same scenario setting and records them to build a verification sample library;

[0093] When the simulation results of the power angle swing curve of the same unit generally show significant deviations from the observed sample results in the verification sample library, and it is confirmed that the problem does not belong to the observation sampling problem, the parameter identification function is activated based on the accumulated quantitative analysis knowledge of variable correlation. This guides the generation of new model correction and exploration directions, which are then fed back to the SPT layer. Through the adjustment of the causal simulation model parameters, the simulation model is corrected, so that the model output continuously approaches the actual behavior of the target system.

[0094] Preferably, the decision support specifically includes:

[0095] Enter a new iteration round, query the quantitative analysis results in knowledge extraction, if the system is unstable, find the unstable metasystem, extract its clustering pattern, and find the leading group; otherwise, end the optimization.

[0096] For each leading unit in the group, the stability margin calculation results after the control are tested based on the EEAC algorithm; the sensitivity is calculated based on the ratio of the stability margin difference before and after control to the unit's power generation.

[0097] If no measure can improve the system stability margin, the optimization ends; otherwise, all trial measures are sorted in descending order according to sensitivity, the highest-ranked switching measure is applied, and the next round begins.

[0098] Based on the optimized switching decision and the activation condition of the decision, the activation condition is converted into a feature vector, forming a record in the pre-decision table.

[0099] Preferably, the migration execution WRT specifically includes:

[0100] The system collects operating conditions and fault information of each node in the actual power system, and also collects actual fault information.

[0101] A new online power system model is constructed based on the collected information; if the online model changes, a new round of pre-training of WRT is started simultaneously to update the pre-decision table;

[0102] When a fault message is received and the fault is a situation requiring emergency control, immediately merge online operating conditions and fault information, and calculate the feature vector;

[0103] The vector is compared with the feature vector stored in each record of the pre-decision table. When the matching degree between the two exceeds the threshold, it is considered that a matching record has been found. Then, the control measures are extracted from the record and the emergency control strategy is immediately implemented.

[0104] If no matching record is found, an emergency control fallback decision is immediately executed.

[0105] A fifth aspect of the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when loaded onto the processor, implements a complex system control decision method integrating WRT and SPT as described in the first and third aspects.

[0106] A sixth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a complex system control decision method integrating WRT and SPT as described in the first and third aspects.

[0107] Compared with the prior art, the beneficial effects of the present invention include at least the following:

[0108] This invention provides a complex system control decision-making method integrating WRT and SPT. The method comprises a two-layer architecture: the upper layer coarsely classifies uncertainties based on structured information from relevant domains both inside and outside the system, as well as unstructured information such as language, and provides the lower-layer causal algorithm with the information needed to generate training samples, including the predicted operating conditions, disturbance events, and available control measures corresponding to each training sample example, efficiently organizing the deep reinforcement learning training process in a high-dimensional decision space; the lower layer employs a WRT-based causal-driven deterministic analysis algorithm to complete the mapping from system state to control decisions. The two layers collaboratively perform pre-training of decisions and periodically update the pre-decision table based on the results. Once an actual disturbance is detected and identified in the objective system, it is matched with the pre-decision table, quickly completing the control decision.

[0109] In summary, this invention proposes a novel research method and implementation platform for complex systems by integrating WRT and SPT. This method and system provide decision-makers with a real-time analysis and decision-making model based on digital intelligence technology. While ensuring the interpretability and reliability of rigorous causal analysis, it can extract various uncertain information contained in speech / sound / image / number symbol strings, construct a pre-decision table containing multiple sets of mappings, and provide reliable, rapid, and adaptive decision suggestions for preventing and controlling various risk events in complex systems. In scenarios with high real-time requirements, it can also provide direct closed-loop control output. Furthermore, by continuously optimizing the organization of pre-training through reinforcement learning and constantly approximating the states and patterns of the real world through model correction, the performance and accuracy of the analysis and decision-making model are further improved.

[0110] This invention provides online (or offline) decision-makers with a digitally intelligent real-time analysis and decision-making platform. While ensuring the reliability and interpretability of rigorous causal analysis, it incorporates speech / audio / video information reflecting uncertainties, rapidly providing reliable and adaptive decision suggestions for preventing various risk events in complex systems, and even directly achieving closed-loop control. Therefore, this invention has broad application prospects and significant practical value in the fields of information science and decision support, particularly in the research of complex systems. Attached Figure Description

[0111] Figure 1 This is a control block diagram of a complex system control decision method that integrates WRT and SPT, provided according to an embodiment of the present invention.

[0112] Figure 2 This is a control block diagram of the A-WARMAP system provided according to an embodiment of the present invention. Detailed Implementation

[0113] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this invention.

[0114] like Figure 1 As shown, Embodiment 1 of the present invention provides a complex system control decision method that integrates WRT and SPT.

[0115] Overall, the key features of this invention are: first, the integration of causal rules and conditions into the decision-making process to enhance the descriptive output of causal relationships, expand the transparency of the model, and achieve a balance between performance and interpretability; second, the introduction of digital twin technology to generate samples through simulation and inference to compensate for insufficient sample size; and third, the introduction of mechanisms, result verification, error correction, and policy fallback mechanisms to improve the credibility of the generated results.

[0116] Regarding deterministic causal driving models, firstly, addressing the analysis and control problems of complex power systems, the inventors' team developed the only theoretically proven and engineering-application-compliant quantitative analysis method for power system transient stability—the Extended Equal Area Criteria (EEAC). The FASTEST software, based on EEAC, for quantitative analysis and optimization decision-making of power system safety and stability, has been exported to 29 users in countries such as the US and Canada, placing the team in a leading international position in terms of both software integrators and user numbers. Next, they developed the Wide Area Monitoring Analysis Protection-control system (WARMAP), which has achieved widespread engineering applications. Furthermore, they constructed the Cyber-Physical-Social System in Energy (CPSSE) framework, supporting the study of the power-energy-environment system as a whole. Then, based on the summary and refinement of the common ideas in the research of the above complex systems, the Whole Reductionism Thinking (WRT) was proposed. By reducing the high-dimensional model of the complex system to a series of 2-dimensional metasystem models, the entropy-preserving mapping from the high-dimensional whole to the low-dimensional group was realized, and a bridge between holism and reductionism was built, creating conditions for the quantitative mechanism research of high-dimensional time-varying nonlinear complex systems.

[0117] This invention aims to further develop the WRT method in handling uncertain and unstructured knowledge, combining it with machine learning techniques to solve the technical problems of poor interpretability of purely data-driven models, difficulty in ensuring learning quality under small sample sizes, and low reliability of analysis results.

[0118] The core concept of this invention includes at least the following: collecting raw corpora of the target system and its objective environment, and extracting symbol string information; in response to the received symbol string information of the target system and its objective environment, pre-training a decision model integrating WRT and SPT, and constructing an event table by mapping the uncertainty factors contained in the symbol strings into coarse-grained tiers; prioritizing the event queue to be analyzed according to the self-attention mechanism for representative states in each tier, and then optimizing the control decision of each item in the event queue based on the deterministic analysis of the overall reconstruction, completing the one-to-one mapping from the predicted quantity of observable quantities to the control decision, and constructing a pre-decision table containing multiple sets of mappings; during transfer execution, in response to the real-time changes of the feature vector of the target system, finding the corresponding control decision and implementing it by matching it with the conditional feature vector in the pre-decision table.

[0119] In a further embodiment, the complex system control decision-making method integrating WRT and SPT includes:

[0120] Step 1: Collect the original corpus of the target system and its objective environment, and extract symbol string information.

[0121] Preferably, but not limited to, the target system includes, but is not limited to, power systems, energy systems, economic systems, and environmental systems; the objective environment refers to the collection of other systems that interact with the target system, including, but not limited to, technical, economic, environmental, and social systems; a symbol string refers to a sequence of bytes, numbers, characters, or words extracted from the original corpus; the original corpus includes, but is not limited to, a document library or database in the form of text, sound, pictures, images, numbers, or a mixture of these, including, but not limited to, meteorological disaster warnings, equipment failure alarms, emergency reports, government / market announcements, expert consultation results, social survey / sensing / forecasting / statistics / planning / decision-making data, etc.

[0122] Artificial intelligence is used to extract symbol strings from raw corpora, and a semantic database is used to identify qualitative or quantitative information related to observable quantities in the symbol strings, which is then organized into formatted strings. The artificial intelligence technologies mentioned include, but are not limited to, expert systems, large language models, and other heuristic algorithms and machine learning methods.

[0123] Observable quantities are state parameters of the target system and its objective environment, including but not limited to physical, economic, and environmental parameters; the predictable quantities of the observable quantities are quantitative predictions of the observable quantities, including but not limited to the time, location, probability, intensity of changes in the observable quantities, and the resulting losses and risks.

[0124] Step 2 involves constructing a decision model integrating WRT and SPT. Specifically, the decision model includes an SPT (Symbol-strings Pre-Training) layer and a WRT (Whole Reductionism Thinking) layer. The SPT layer is the upper layer of the decision model, used for uncertainty analysis, while the WRT layer is the lower layer, used for deterministic analysis. Subsequent steps involve pre-training the decision model, including optimizing the internal data structures and parameters of the SPT and WRT layers.

[0125] Step 3, pre-training the SPT layer, includes: mapping the uncertainty factors contained in the symbol string to coarse-grained classification, constructing an event table containing knowledge of the uncertain environment, and generating an event queue that needs to be pre-analyzed.

[0126] Preferably, but not limitingly, step 3 specifically includes:

[0127] Step 3.1 involves mapping the uncertainties contained in the symbol string to coarse-grained segments.

[0128] Further, but not restrictively, the tiered representation of observables can be implemented using the following method:

[0129] For qualitative descriptions of observables in a symbol string, quantitative tiered representations are obtained by querying a semantic database;

[0130] For the uncertain quantitative representation of observable quantities in a symbol string, a probability density function is used;

[0131] For uncertain quantitative representations involving multiple observables changing simultaneously in a symbol string, a joint probability density function is used. For the representation in the form of probability density function, the classification operation should be performed with reference to the semantic library settings or the knowledge learned from the function itself. It is allowed to define a one-to-many mapping from one observable to multiple observables through the semantic library.

[0132] The semantic library includes, but is not limited to, the following: a mapping table from key features expressed in the form of symbol strings in the original corpus to corresponding observables; a mapping table from qualitative descriptions of observables to quantitative classifications; a one-to-many mapping table from one observable classification to multiple observable classifications; and a method for determining the values ​​of representative states in each observable classification.

[0133] Step 3.2: Determine the representative state in each segment, wherein the representative state in each segment is determined in conjunction with the control decision type to be trained in the WRT layer deterministic analysis.

[0134] Further preferred but not restrictive, for preventive control decisions to be applied before a risk event occurs, the representative state refers to the expected value of all samples within the tier of the observable quantity; for emergency control decisions to be applied after a risk event occurs, the representative state refers to the extreme value of all samples within the tier of the observable quantity in the most unfavorable direction, including but not limited to the maximum and minimum values ​​of all samples; the method for determining the representative state should be specified in advance through a semantic library.

[0135] Step 3.3: Based on the obtained tiered representation of observables, construct an event table containing knowledge of uncertain environmental conditions, where an event refers to a process in which one or more observables change due to an uncertain factor.

[0136] Further preferred, but not limiting, step 3.3 specifically includes:

[0137] An event is constructed starting from a single observable, multiple independent observables, or multiple observables that change simultaneously; each event must select a subdivision from each of the relevant observables;

[0138] The probability of occurrence is determined based on the source of the event: a single observable is taken from the result of the probability density function or the semantic library setting value; multiple independent observables are taken from the product of the probability values ​​of each member observable; multiple observables that change simultaneously are taken from the result of the joint probability density function or the semantic library setting value; each possible combination of categories should be constructed as an event, and duplicate events are not allowed to be constructed.

[0139] Multiple event records are summarized into an event table.

[0140] Step 3.4: For representative states in each segment, prioritize the events in the event table according to the self-attention mechanism to generate an event queue.

[0141] Further preferred, but not restrictive, the steps of the self-attention mechanism are as follows: first, estimate the pre-control risk value of each event in the event table based on the representative state value of each event in the coarse-grained classification value; then, sort the event table in descending order according to the pre-control risk value to generate a queue of events to be pre-analyzed; the deterministic analysis based on causality must be carried out one by one from the events ranked at the top, that is, the risk head events; when parallel computing capability is available, the deterministic analysis based on causality must still be carried out batch by batch for the risk head events.

[0142] Step 4, pre-training the WRT layer, includes: the upper SPT layer pushing the event table to the WRT layer to start the pre-training of the WRT layer; for representative states in each tier, optimizing the corresponding control decisions based on causal-driven deterministic analysis to generate a pre-decision table containing quantitative knowledge of the deterministic domain; when the number of records in the pre-decision table constructed by the WRT layer reaches a threshold, the decision model reaches the transfer condition; when the construction of the WRT layer pre-decision table is completed, the training of the decision model ends.

[0143] In a preferred embodiment, the process of constructing a pre-decision table containing multiple sets of mappings through WRT training, and optimizing the control decision of each item in the event queue based on deterministic analysis of overall restoration includes: simulation deduction, knowledge extraction, decision support, and intelligence enhancement.

[0144] Preferably, but not limitingly, step 4 specifically includes:

[0145] Step 4.1: Simulation and deduction, responsible for deduce the operation process of the target system under the given system state and decision, within a given time range, based on the simulation model that reflects the objective laws of the target system.

[0146] Further preferred, but not limiting, step 4.1 specifically includes:

[0147] In response to the received event table, the simulation model required for deterministic analysis is obtained from the SPT layer, i.e., the symbol string pre-training layer.

[0148] In real-time mode, data reflecting the current state of the target system is substituted into the model; in non-real-time mode, such as, but not limited to, non-real-time mode for research purposes, data reflecting the target system at past or hypothetical future times are substituted into the model; the user decides which state of data to substitute into the model.

[0149] For each event record to be trained, the hierarchical state of each observable related to the event record is substituted and mapped to the corresponding simulation model parameters to complete the modification of the simulation model.

[0150] Based on the modified model data, causal simulations are conducted to generate simulation parameter trajectories within a given time range.

[0151] Step 4.2: Knowledge Extraction, responsible for extracting quantitative knowledge from the simulation parameter trajectory that is conducive to achieving the research objectives.

[0152] Further preferred, but not limiting, step 4.2 specifically includes:

[0153] A series of planar orthogonal pattern libraries are built based on domain expert knowledge, and each pattern contains only 2 parameters;

[0154] Based on the settings of each mode, the simulation model is reduced to order 2 by treating other parameters not involved in that mode as constants (hereinafter referred to as parameters). By traversing all modes, an underdetermined entropy-preserving order reduction mapping matrix of size M*1 is generated, where M is the number of modes.

[0155] For an element of the underdetermined entropy-preserving reduction matrix, it is expanded into a row. The method is as follows: at each simulation time step, extract the corresponding parameter trajectory from the simulation parameter trajectory generated in step 4.1, and substitute the data of each time step on the trajectory into each parameter of the element model; by traversing all time steps, construct a sequence of size 1*N, where N is the number of time steps. Each element in the sequence is called a meta-system, which is the reduced-order model of the simulation results of each time step with the parameters substituted; combine the rows of all modes to generate a well-determined entropy-preserving reduction mapping matrix of size M*N, where M is the number of modes and N is the number of time steps;

[0156] Linear analytical methods are used to analyze each element system of the well-defined entropy-preserving reduced-order matrix, so as to achieve a quantitative expression of a class of characteristics of each element system. The linear analytical methods include, but are not limited to: eigenvalue analysis and equal area criterion. The selected element system characteristics should be able to support the overall research objectives of the target system. The element system characteristics include, but are not limited to: stability, sufficiency, and environmental friendliness.

[0157] Aggregate the characteristic expressions of all metasystems to restore and output the overall characteristic expression of the target system. Aggregation methods include, but are not limited to: taking the minimum value, taking the maximum value, and taking the expected value.

[0158] Step 4.3, Decision Support, is responsible for probing the values ​​of relevant decision parameters in the simulation model for a specific decision problem under a given system state. Then, it calls back to steps 4.1 and 4.2 to generate quantitative analysis results of the target system characteristics under the probing decision. Through repeated probing and callbacks, multiple rounds of analysis results are generated. Based on the results of a single or multiple rounds of analysis, a new next or next set of probing directions is generated. The above process is repeated until the probing no longer produces results more favorable to the research objective. The best-performing decision and the feature vector of the initiation condition of that decision are stored as a record in the pre-decision table.

[0159] As a further preferred but non-limiting implementation, step 4 further includes:

[0160] Step 4.4: Improve the model with intelligence for falsification and correction.

[0161] Further preferred, but not limiting, step 4.4 specifically includes:

[0162] Obtain sample results of each observable from the SPT layer;

[0163] Simulation results of observables to be verified are extracted from the results of multiple deterministic cases and recorded to construct a verification sample library;

[0164] When the simulation results of the same observable quantity generally show significant deviations from the observed sample results in the validation sample library, and it is certain that the problem does not belong to the observation sampling problem, the parameter identification function is activated based on the quantitative analysis knowledge of variable correlation accumulated in step 4.1. This guides the generation of new model correction trial directions, which are then fed back to the SPT layer. Through the adjustment of the causal simulation model parameters, the simulation model is corrected, so that the model output continuously approaches the actual behavior of the target system.

[0165] Step 5: After the decision model meets the migration conditions and the WRT layer pre-decision table is constructed, the migration is executed. In response to the real-time changes in the target system feature vector, the corresponding control decision is found and implemented by matching it with the condition feature vector in the pre-decision table.

[0166] Preferably, but not limitingly, step 5 specifically includes:

[0167] During the transfer learning process, the SPT layer pushes the actual feature vector of the target system to the WRT layer. The WRT layer then queries the pre-decision table for matching control measures and applies them to the target system. In response to the reinforcement learning task executed in the SPT layer, the WRT layer also provides the SPT layer with deterministic case updates or continuously improves the training sample library. In addition, these cases can also support the model's falsification and correction tasks, enriching the validation sample library.

[0168] During migration execution, in response to the real-time changes in the target system's feature vector, the corresponding control decision is found and implemented by matching it with the conditional feature vectors in the pre-decision table.

[0169] Further, but not restrictively, the specific process of migration execution includes:

[0170] The detailed process of migration execution is as follows: When the actual feature vector of the target system pushed by the SPT layer is received, it is first determined whether the changes in the target system meet the migration execution conditions. If so, the vector is compared with the feature vector stored in each record of the pre-decision table. When the matching degree of the two exceeds the threshold, it is considered that a matching record has been found. Then, the control measures are extracted from the record and applied to the target system immediately. If no matching record is found, the fallback decision to deal with the problem is searched and then applied to the target system immediately. The fallback decision includes, but is not limited to: for a specific decision problem, the decision generated based on the worst possible state of the target system; and the specific decisions trained to deal with each possible system state are summarized into a set, and reinforcement learning is performed based on the data in the set to obtain a general decision that can deal with various possible states.

[0171] Further optimization, but not limitation, of the specific processes of reinforcement learning includes:

[0172] In response to reinforcement learning tasks performed at the SPT layer, the WRT layer must be able to serve as a deterministic case inference tool for the SPT layer, simulating environmental responses for each decision trial, generating a sample library that supports reinforcement learning training needs, and feeding it back to the SPT layer.

[0173] In response to the model correction task performed in the SPT layer, the WRT layer must be able to serve as a deterministic case inference tool for the SPT layer, to obtain the corresponding simulation sample for each observable observation sample, and to feed it back to the SPT layer.

[0174] In response to the above reinforcement learning or model correction tasks, the samples returned directly after the simulation inference step are the time response trajectories of observable quantities; the samples returned directly after the quantification analysis step are the quantitative index expressions of the target system under the specified state; and the samples returned directly after the decision support step are the control measures of the target system under the specified state.

[0175] Embodiment 2 of the present invention provides a system control system integrating WRT and SPT, comprising: the decision model described in Embodiment 1, wherein the decision model includes: an SPT layer for uncertainty analysis and a WRT layer for deterministic analysis; the SPT layer is the upper layer of the decision model, and the WRT layer is the lower layer of the decision model; after the decision model is pre-trained as described in Embodiment 1, the SPT layer pushes the actual feature vector of the target system to the WRT layer, and the WRT layer queries the pre-decision table for matching control measures and applies them to the target system.

[0176] like Figure 2 As shown, Embodiment 3 of the present invention provides a complex system control decision-making method integrating WRT and SPT. Based on the Adaptive-Wide Area Monitoring Analysis Protection-control system (A-WARMAP) framework, it implements environmentally adaptive power grid wide-area monitoring analysis protection control. The method includes the following steps:

[0177] Step 1: Collect the original corpus of A-WARMAP and its objective environment, construct a scenario and working condition corpus, and extract symbol string information.

[0178] Preferably, but not limitingly, step 1 specifically includes:

[0179] Step 1.1: Collect corpus reflecting changes in power system operating conditions and scenarios from any of the multiple information sources, including the Internet, dedicated data networks, or buffer networks connecting the two.

[0180] Step 1.2: Complete the preprocessing of the corpus: For structured meteorological data, conventional methods can be used to complete data cleaning / statistics / conversion / integration according to information sources and classifications; for unstructured meteorological disaster early warning corpus, it is directly added to the unstructured storage container in the form of strings.

[0181] Further preferred, but not limiting, step 1.2 specifically includes:

[0182] A semantic database is constructed to support the parsing of unstructured meteorological corpora. The construction steps include: offline extraction of text pattern features based on domain experience to form a string pattern database of key factors. Each table in the database stores a vocabulary of relevant factors, including: time, administrative division name, geographical location, disaster warning type, warning level, meteorological element type, predicted value, trend type, degree of change value, etc.; based on the user's requirements for coarseness, a mapping table is formed between meteorological feature strings and classification levels. A mapping in the table might look like "8-level gale → [1, 4]", where 1 represents the disaster-causing factor of gale and 4 represents the classification number of a 7-8 level gale. In addition, the semantic database can also record common knowledge, rules, and conceptual systems in the field of meteorological disasters. The above information can be stored in either structured tables or unstructured formats.

[0183] The process of constructing formatted strings based on a large language model includes: First, extracting a string from the corpus according to natural language delimiters (such as commas, periods, and carriage returns); second, segmenting the corpus into words, extracting keywords, and matching them with a pattern vocabulary stored in a semantic database, and performing pattern encoding based on the matching results; third, encoding meteorological features according to a meteorological feature grading mapping table in the semantic database; fourth, forming a symbol string model of the background features of the routes within the research scope, including the following parameters: route segment name, route name, starting tower number, ending tower number, geographical features, channel features, ontological features, etc.; finally, querying the matching relationship between spatial information and spatial pattern encoding in the route background feature symbol string model, and combining it with meteorological feature encoding information to achieve meteorological feature encoding for each route.

[0184] Step 2, construct a decision model that integrates WRT and SPT, such as Figure 2 As shown, the A-WARMAP system consists of a two-layer architecture. Specifically, the decision model includes: the SPT (Symbol-strings Pre-Training) layer and the WRT (Whole Reductionism Thinking) layer. The SPT layer, as the upper layer of the decision model, is used for uncertainty analysis and pre-training for power system safety and sufficiency control decisions. It can sort each example in the pre-fault table according to the risk level value when no control is applied. The following example, "generating a pre-fault table from meteorological disaster early warning corpus", illustrates the specific implementation steps of the SPT layer.

[0185] Step 3, pre-training the SPT layer, includes: mapping the uncertainty factors contained in the symbol string to coarse-grained classification, constructing an event table containing knowledge of the uncertain environment, and generating an event queue that needs to be pre-analyzed.

[0186] Preferably, but not limitingly, step 3 specifically includes:

[0187] Step 3.1 involves mapping the uncertainties contained in the symbol string to coarse-grained classifications, including: failure probability classification, loss classification, and risk classification.

[0188] The fault probability classification includes: First, based on domain knowledge, a method is constructed that comprehensively considers one or more meteorological characteristic factors, real-time operating status, maintenance status, and other information for each line to assess the fault rate of the line under a given fault type; Second, it is determined whether the number of classifications for the same meteorological characteristic on the object is unique. If so, the extreme value of each meteorological characteristic factor in the most unfavorable direction within its classification is extracted, and the line fault rate is calculated; if not, all permutations and combinations of the classifications involved in each factor are enumerated, and the extreme values ​​of each meteorological characteristic factor in the most unfavorable direction within the enumerated classifications are extracted in turn, and the line fault rate is calculated; The classification mapping table in the semantic database is read, and based on the distribution of all fault rate calculation results, all classifications of the line's fault probability under a given fault type are marked and recorded.

[0189] Loss classification includes: First, based on domain knowledge, constructing a method to estimate line pre-control losses by comprehensively considering the importance of each line in the entire power grid topology and real-time power status; Second, completing the assessment and recording of line losses.

[0190] Risk classification includes: First, extracting the fault probability classification and pre-control loss classification information for each line in sequence, and taking the extreme value in the most unfavorable direction within the classification as the result; Second, by multiplying the line loss by the fault probability result, the pre-control risk of the line is assessed and recorded.

[0191] Step 3.2: Determine the representative status in each sub-segment.

[0192] Step 3.3: Based on the obtained tiered representation of observables, construct an event table that includes knowledge of uncertain environmental conditions.

[0193] Further preferred but not restrictive, step 3.3 specifically includes: constructing and sorting a pre-fault table. First, from the calculation results of the above steps, a record is formed in the format of "time, line number, fault probability, pre-control loss, pre-control risk, fault type". Second, all records are merged into the pre-fault table, i.e., the event table. The pre-fault table is sorted from largest to smallest pre-control risk to form a pre-fault queue that needs to be pre-analyzed.

[0194] Step 3.4: For representative states in each tier, prioritize events in the event table according to the self-attention mechanism to generate an event queue. This involves sorting the pre-fault table from highest to lowest pre-control risk, forming a pre-fault queue requiring pre-analysis. Preferably, considering limited computing power, high-risk faults are extracted from the pre-fault queue, and deterministic optimization analysis based on overall reconstruction is performed batch by batch using parallel computing tools.

[0195] Step 4: Pre-train the WRT layer. The pre-training of WRT is mainly based on deterministic analysis of the overall reconstruction to generate the optimal control decision for each fault in the pre-fault queue, which is then summarized into a pre-decision table.

[0196] Preferably, but not restrictively, in step 4, the analysis objectives include: short-term prevention / emergency control and medium-to-long-term prevention / correction control of the power system; the analysis content includes: power flow interruption, transient security and stability quantitative analysis (EEAC) based on the extended equal area criterion, energy change rate, and trajectory characteristic roots; the analysis method includes: optimal control strategy search based on WRT.

[0197] Preferably, but not limitingly, step 4 specifically includes:

[0198] Step 4.1, Simulation and Deduction, is used to deduce the system stability behavior within a given time range based on a simulation model that reflects the target power system under given power system state and decision conditions.

[0199] Further preferred, but not limiting, step 4.1 specifically includes:

[0200] In response to the received pre-operating condition and pre-fault tables, the electromechanical / electromagnetic transient simulation model required for stability analysis is obtained from the SPT layer;

[0201] In the online operating state, data reflecting the current state of the target power system are substituted into the model;

[0202] For each fault record to be trained, the tiered state of the observables related to the fault record, such as, but not limited to, the power flow tiered value of a certain line, is substituted into and mapped to the corresponding simulation model parameters, such as, but not limited to, the power flow value of a certain line, to complete the modification of the simulation model.

[0203] Based on the modified model data, transient simulations of the power system are conducted to generate simulation parameter trajectories within the required time range, such as, but not limited to, unit power angle swing curves.

[0204] Step 4.2, knowledge extraction, is used to extract quantitative knowledge from the simulation parameter trajectory, such as, but not limited to, the unit power angle swing curve, which is conducive to achieving the goal of stability quantitative analysis.

[0205] Further preferred, but not limiting, step 4.2 specifically includes:

[0206] Based on the EEAC method, for each time slice in the trajectory, the power angle swing curves of N units are divided into two groups based on each trajectory gap in that slice. Within each group, an inertia weighting formula is applied to aggregate the multi-machine trajectories into a single-machine trajectory, thereby achieving the positioning of the inertia centerline. Each gap corresponds to a grouping method, and each grouping method is a planar orthogonal pattern. By summarizing all time slices and all gaps in each slice and removing duplicates, a planar orthogonal pattern library is constructed, with each pattern containing only two groups.

[0207] Based on trajectory aggregation in each mode, the simulation model is reduced to order 2; by traversing all modes, an underdetermined entropy-preserving order reduction mapping matrix of size M*1 is generated, where M is the number of modes.

[0208] For an element of the underdetermined entropy-preserving reduction matrix, it is expanded into a row. The method is as follows: at each simulation time step, the corresponding parameter trajectory is extracted from the simulation parameter trajectory generated in step 4.1.4, and the data of each time step on the trajectory is substituted into each parameter of the element model; by traversing all time steps, a sequence of size 1*N is constructed, where N is the number of time steps. Each element in the sequence is called a meta-system, which is the reduced-order model of the simulation results of each time step with the parameters substituted; the rows of all modes are combined to generate a well-determined entropy-preserving reduction mapping matrix of size M*N, where M is the number of modes and N is the number of time steps;

[0209] In each element system of a well-defined entropy-preserving reduced-order matrix, the stability margin of the element system is calculated using the EEAC analytical method.

[0210] The characteristics of all the subsystems are aggregated, the minimum stability margin of all the subsystems is taken, and the result is restored to the whole system as an expression of the overall stability of the target power system and output.

[0211] Step 4.3: Decision support, used to explore the optimal values ​​of relevant decision parameters in the simulation model for prevention / emergency / corrective control decision problems, such as, but not limited to, the total cut-off amount of all units at the first moment of the fault and the optimal value of the cut-off action for each unit.

[0212] Further preferred, but not limiting, step 4.3 specifically includes:

[0213] Step 4.3.1: Enter a new iteration round, query the quantitative analysis results in Step 4.2. If the system is unstable, find the unstable metasystem, extract its clustering pattern, and find the leading group (generally composed of units that have lost synchronization stability); otherwise, end the optimization and proceed to Step 4.3.5.

[0214] Step 4.3.2: For each leading unit in the group, calculate the stability margin after the control is cut off based on the EEAC algorithm; calculate the sensitivity based on the ratio of the stability margin difference before and after control to the unit's power generation.

[0215] Step 4.3.4: If no measure can improve the system stability margin, then end the optimization and proceed to step 4.3.5; otherwise, sort all the trial measures in descending order according to sensitivity, apply the highest-ranked switching measure, and proceed to the next round, returning to step 4.3.1.

[0216] Step 4.3.5: Based on the optimized machine switching decision and its initiation conditions (such as corresponding operating conditions and faults), the initiation conditions are converted into feature vectors to form a pre-decision table. Figure 2 A record from the China Index Online Decision Table.

[0217] Step 4.4: Generate and store fallback decisions.

[0218] Further preferred, but not limiting, step 4.4 specifically includes:

[0219] The generation method is as follows: for the worst operating conditions and faults that may occur in the target power system, record the decisions made at this time; summarize the emergency control decisions trained to deal with each possible system state into a set, and perform reinforcement learning based on the data in the set to obtain a general emergency control decision that can deal with various possible states;

[0220] The storage method is as follows: the decisions generated based on the above method are stored in the scheduling side as an offline decision table, and stored in the plant side as a fallback table.

[0221] Step 4.5: Synchronization of decision table groups. Considering the limited computing power, tables in the pre-fault queue that have not yet completed the control strategy update are designated as "tables in progress". After the update is completed, the corresponding records should be promptly sent to the "online decision table". Finally, through periodic updates, the decision table group on the scheduling side is synchronized to the decision table group on the plant side via the communication system.

[0222] Step 4.6: Intelligent Enhancement, including: The SPT layer reads sample results of observable quantities from the Wide Area Measurement System (WAMS) or Supervisory Control and Data Acquisition (SCADA) system, preferably but not limited to, sample results of the unit's power angle swing curve after disturbance; extracts and records simulation results of observable quantities to be verified from the simulation results of the same scenario setting, such as, but not limited to, simulation results of the unit's power angle swing curve after disturbance, and constructs a verification sample library; when the simulation results of the power angle swing curve of the same unit generally have significant deviations from the observed sample results in the verification sample library, and it is confirmed that it is not an observation sampling problem, based on the quantitative analysis knowledge of variable correlation accumulated in Step 4.1, the parameter identification function is activated to guide the generation of new model correction trial directions, which are fed back to the SPT layer. Through the adjustment of the causal simulation model parameters, the simulation model is corrected, so that the model output continuously approaches the actual behavior of the target system.

[0223] Step 5: Migration execution of WRT. The migration execution of WRT is responsible for querying the pre-decision table based on the actual operating conditions and faults, finding the matching decision and executing it.

[0224] Preferably, but not limitingly, step 5 specifically includes:

[0225] Step 5.1: Collect injection quantities, topology quantities, line power flow and other operating condition information of each node from the actual new power system, and collect actual fault information at the same time;

[0226] Step 5.2: Construct a new online power system model based on the collected information; if the online model changes, simultaneously start a new round of pre-training of WRT and update the pre-decision table;

[0227] Step 5.3: When a fault message is received and the fault is a situation requiring emergency control, immediately merge the online operating conditions and fault information, and calculate the feature vector;

[0228] Step 5.4: Compare the vector with the feature vector stored in each record of the pre-decision table. When the matching degree between the two exceeds the threshold, it is considered that a matching record has been found. Then, the control measures are extracted from the record and the emergency control strategy is executed immediately.

[0229] Step 5.5: If no matching record can be found, immediately execute the emergency control fallback decision.

[0230] Embodiment 4 of the present invention provides a computer-readable storage medium for storing one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to perform the complex system control decision method integrating WRT and SPT according to Embodiment 1.

[0231] Embodiment 5 of the present invention provides a computing device including one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a complex system control decision method integrating WRT and SPT according to Embodiment 1.

[0232] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.

[0233] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A control decision-making method for complex systems integrating WRT and SPT, characterized in that, Includes the following steps: Collect raw corpora of the target system and its objective environment, and extract symbol string information; Construct a decision model that integrates WRT and SPT. The SPT layer is the upper layer of the decision model and is used for uncertainty analysis, while the WRT layer is the lower layer of the decision model and is used for deterministic analysis. Pre-training the SPT layer includes: mapping the uncertainty factors contained in the symbol string to coarse-grained classification, constructing an event table containing knowledge of the uncertain environment, and generating an event queue that needs to be pre-analyzed. Pre-training the WRT layer includes: the upper SPT layer pushes the event table to the WRT layer, optimizes the corresponding control decisions based on causal deterministic analysis for representative states in each segment, and generates a pre-decision table containing quantitative knowledge of deterministic domain. The transfer execution of the decision model is implemented in response to the real-time changes in the feature vector of the target system. By matching it with the condition feature vector in the pre-decision table, the corresponding control decision is found and implemented. Among them, WRT stands for holistic reduction thinking, and SPT stands for symbol string pre-training.

2. The complex system control decision-making method integrating WRT and SPT according to claim 1, characterized in that: The target system includes: power system, energy system, economic system, or environmental system; The objective environment includes: a collection of other systems that interact with the target system; Symbol strings include sequences of bytes, numbers, characters, or words extracted from the original corpus; the original corpus includes text, sound, images, pictures, numbers, or a combination of these in the form of a document library or database.

3. A complex system control decision-making method integrating WRT and SPT according to claim 1 or 2, characterized in that: The pre-training of the SPT layer specifically includes: By mapping the uncertainties contained in the symbol string to coarse-grained subdivisions; The representative state in each tier is determined based on the control decision type to be trained in the WRT layer deterministic analysis. Based on the obtained hierarchical representation of observables, an event table containing knowledge of uncertain environmental conditions is constructed. For each representative state in each tier, the events in the event table are prioritized according to the self-attention mechanism to generate an event queue.

4. The complex system control decision-making method integrating WRT and SPT according to claim 3, characterized in that: The process of mapping the uncertainty factors contained in the symbol string to coarse-grained subdivisions includes: For qualitative descriptions of observables in a symbol string, quantitative tiered representations are obtained by querying a semantic database; For the uncertain quantitative representation of observable quantities in a symbol string, a probability density function is used; For uncertain quantitative representations involving multiple observables changing simultaneously in a symbol string, a joint probability density function is used. For the representation in the form of probability density function, a classification operation should be performed with reference to the semantic library settings or the knowledge learned from the function itself. A one-to-one or one-to-many mapping from an observable to one or more observables is defined through the semantic library.

5. The complex system control decision-making method integrating WRT and SPT according to claim 3, characterized in that: The determination of representative states in each tier based on the control decision type to be trained in the WRT layer deterministic analysis includes: For preventive control decisions that need to be applied before a risk event occurs, representativeness refers to the expected value of the observable quantity for all samples within that tier. For emergency control decisions that need to be applied after a risk event occurs, the representative state refers to the extreme value of all samples in the most unfavorable direction within the observable range. The representative state value method should be specified in advance through the semantic library.

6. The complex system control decision-making method integrating WRT and SPT according to claim 3, characterized in that: The process of prioritizing events in the event table according to the self-attention mechanism and generating an event queue for representative states in each tier includes: Estimate the pre-control risk value based on the representative state value of each event in the coarse-grained classification value in the event table; Sort the event table in descending order according to the pre-control risk value to generate a queue of events that need to be pre-analyzed; We will prioritize the events that are ranked first, i.e. the high-risk events, and conduct the aforementioned deterministic analysis based on causality one by one. Even when parallel computing capabilities are available, it is still necessary to conduct the aforementioned deterministic analysis based on causal drive for each batch of high-risk events.

7. A complex system control decision-making method integrating WRT and SPT according to claim 1 or 2, characterized in that: The pre-training of the WRT layer specifically includes: Simulation simulation: used to simulate the operation of a system under a given system state and decision, within a given time range, based on a simulation model that reflects the objective laws of the target system; Knowledge extraction: Used to extract quantitative knowledge from simulation parameter trajectories that is beneficial to achieving research objectives; Decision support: For specific decision problems, under a given system state, it explores the values ​​of relevant decision parameters in a simulation model, then calls back the simulation and knowledge extraction to generate quantitative analysis results of the target system characteristics under the trial decision; through repeated trials and callbacks, it generates multiple rounds of analysis results; based on the results of a single or multiple rounds of analysis, it generates a new next or next set of trial directions; it repeats the above process until the trial no longer produces better results; the best-performing decision, and the feature vector of the initiation condition of that decision, are stored as a record in the pre-decision table.

8. The complex system control decision-making method integrating WRT and SPT according to claim 7, characterized in that: The pre-training of the WRT layer also includes: To enhance the decision-making model with intelligence, which can be used for the falsification and correction of the decision-making model.

9. A complex system control decision-making method integrating WRT and SPT according to claim 7, characterized in that: The simulation specifically includes: In response to the received event table, the simulation model required for deterministic analysis is obtained from the SPT layer; In real-time mode, data reflecting the current state of the target system is substituted into the simulation model; in non-real-time mode, data reflecting the target system at past or assumed future times are substituted into the simulation model. For each event record to be trained, the hierarchical state of each observable related to the event record is substituted and mapped to the corresponding simulation model parameters to complete the modification of the simulation model. Based on the modified model data, causal simulations are conducted to generate simulation parameter trajectories within a given time range.

10. A complex system control decision-making method integrating WRT and SPT according to claim 7, characterized in that: The knowledge extraction specifically includes: A series of planar orthogonal pattern libraries are built based on domain expert knowledge, and each pattern contains only 2 parameters; Based on the settings of each mode, the simulation model is reduced to order 2 by treating other parameters not involved in that mode as constants (hereinafter referred to as parameters). By traversing all modes, an underdetermined entropy-preserving order reduction mapping matrix of size M*1 is generated, where M is the number of modes. For an element of an underdetermined entropy-preserving reduced-order matrix, expand it into a single row; Combine all the rows of the patterns to generate a well-defined entropy-preserving reduced-order mapping matrix of size M*N, where M is the number of patterns and N is the number of time steps; The linear analytical method is used to analyze each element system of the well-defined entropy-preserving reduced-order matrix, so as to achieve a quantitative expression of a class of characteristics of each element system; It aggregates the characteristic expressions of all metasystems, restores and outputs the overall characteristic expression of the target system.

11. A complex system control decision-making method integrating WRT and SPT according to claim 8, characterized in that: The aforementioned intelligent enhancement specifically includes: Obtain sample results of each observable from the SPT layer; Simulation results of observables to be verified are extracted from the results of multiple deterministic cases and recorded to construct a verification sample library; When the simulation results of the same observable quantity generally show significant deviations from the observed sample results in the validation sample library, and it is certain that this is not an observation sampling problem, the parameter identification function is activated based on the quantitative analysis knowledge of variable correlation accumulated in the simulation. This guides the generation of new model correction trial directions, which are then fed back to the SPT layer. Through the adjustment of the causal simulation model parameters, the simulation model is corrected, so that the model output continuously approaches the actual behavior of the target system.

12. A complex system control decision-making method integrating WRT and SPT according to claim 1 or 2, characterized in that: During the migration execution, the SPT layer pushes the actual feature vector of the target system to the WRT layer, and the WRT layer queries the pre-decision table for matching control measures and applies them to the target system. In response to the reinforcement learning task performed in the SPT layer, the WRT layer also provides deterministic cases to the SPT layer to update or continuously improve the training sample library; or, it will use deterministic cases to support the falsification and correction tasks of the decision model and enrich the validation sample library.

13. A complex system control decision-making method integrating WRT and SPT according to claim 1 or 2, characterized in that: The migration execution specifically includes: When the actual feature vector of the target system pushed by the SPT layer is received, it is determined whether the changes in the target system meet the migration execution conditions. If so, the vector is compared with the feature vector stored in each record of the pre-decision table. When the matching degree of the two exceeds the threshold, it is considered that a matching record has been found. The control measures are then extracted from the record and applied to the target system immediately. If no matching record is found, the fallback decision to deal with the problem is searched and then applied to the target system immediately. The specific decisions trained to deal with each possible system state are aggregated into a set. Reinforcement learning is then performed based on the data in the set to obtain a general decision that can deal with all possible states.

14. The complex system control decision-making method integrating WRT and SPT according to claim 13, characterized in that: The reinforcement learning specifically includes: In response to reinforcement learning tasks performed at the SPT layer, the WRT layer must be able to serve as a deterministic case inference tool for the SPT layer, simulating environmental responses for each decision trial, generating a sample library that supports reinforcement learning training needs, and feeding it back to the SPT layer. In response to the model correction task performed in the SPT layer, the WRT layer must be able to serve as a deterministic case inference tool for the SPT layer, to obtain the corresponding simulation samples for each observable observation sample, and to feed them back to the SPT layer. In response to the above reinforcement learning or model correction tasks, the samples returned directly after the simulation inference step are the time response trajectories of observable quantities; the samples returned directly after the quantification analysis step are the quantitative index expressions of the target system under the specified state; and the samples returned directly after the decision support step are the control measures of the target system under the specified state.

15. A complex system control decision system integrating WRT and SPT, running a complex system control decision method integrating WRT and SPT according to any one of claims 1 to 14, characterized in that, include: Decision-making models; The decision model includes: an SPT layer for uncertainty analysis and a WRT layer for deterministic analysis; the SPT layer is the upper layer of the decision model, and the WRT layer is the lower layer of the decision model; after the decision model is pre-trained, the SPT layer pushes the actual feature vector of the target system to the WRT layer, and the WRT layer queries the pre-decision table for matching control measures and applies them to the target system; Among them, WRT stands for holistic reduction thinking, and SPT stands for symbol string pre-training.

16. A control decision-making method for complex systems integrating WRT and SPT, characterized in that, Based on A-WARMAP, an adaptive power grid wide-area monitoring, analysis, protection, and control method is implemented, comprising the following steps: Collect raw corpora of A-WARMAP and its objective environment, construct scenario and working condition corpora, and extract symbol string information; A decision model integrating WRT and SPT is constructed. The decision model includes an SPT layer and a WRT layer. The SPT layer, as the upper layer of the decision model, is used for uncertainty analysis and pre-training of power system safety margin control decisions. Each example is sorted in the pre-fault table according to the risk level value when no control is applied. The WRT layer, as the lower layer of the decision model, is used for deterministic analysis. Pre-training the SPT layer includes: mapping the uncertainty factors contained in the symbol string to coarse-grained classification, constructing an event table containing knowledge of the uncertain environment, and generating an event queue that needs to be pre-analyzed. The WRT layer is pre-trained to generate the optimal control decision for each fault in the pre-fault queue, and the results are summarized into a pre-decision table. The migration execution WRT queries the pre-decision table based on the actual operating conditions and faults, finds the matching decision, and executes it. Among them, WRT stands for holistic reduction thinking, and SPT stands for symbol string pre-training.

17. A complex system control decision-making method integrating WRT and SPT according to claim 16, characterized in that: The process of collecting raw corpora of A-WARMAP and its objective environment, constructing scenario and operational condition corpora, and extracting symbol string information includes: Collect corpora reflecting changes in power system operating conditions and scenarios from any of the multiple information sources, including the Internet, dedicated data networks, or buffer networks connecting the two. Preprocessing of the corpus includes: for structured meteorological data, data cleaning, statistics, transformation and integration are completed according to information sources and classifications; for unstructured early warning corpus, it is directly added to the unstructured storage container in the form of strings.

18. A complex system control decision-making method integrating WRT and SPT according to claim 16, characterized in that: The pre-training of the SPT layer specifically includes: By mapping the uncertainties contained in the symbol string to coarse-grained classifications, including: failure probability classification, loss classification, and risk classification; Determine the representative status in each sub-category; Based on the obtained tiered representation of observables, an event table containing knowledge of uncertain environmental conditions is constructed, including: constructing and sorting a pre-fault table, forming a record from the calculation results of the above steps in the format of "time, line number, fault probability, pre-control loss, pre-control risk, fault type"; merging all records into the pre-fault table, i.e., the event table; sorting the pre-fault table from largest to smallest pre-control risk to form a pre-fault queue that needs to be pre-analyzed; For each representative state in each tier, the events in the event table are prioritized according to the self-attention mechanism to generate an event queue, forming a pre-fault queue that needs to be pre-analyzed.

19. A complex system control decision-making method integrating WRT and SPT according to claim 16, characterized in that: The pre-training of the WRT layer specifically includes: Simulation and deduction: Under given power system states and decisions, based on a simulation model that reflects the target power system, deduce the system stability behavior within a given time range; Knowledge extraction: Extract quantitative knowledge from simulation parameter trajectories that is beneficial to achieving the goal of stability quantitative analysis; Decision support: For prevention / emergency / corrective emergency control decision problems, explore the optimal values ​​of relevant decision parameters in the simulation model; For the worst operating conditions and faults that may occur in the target power system, record the decisions made at that time; summarize the preventive / emergency / corrective control decisions trained to deal with each possible system state into a set, and perform reinforcement learning based on the data in the set to obtain a general emergency control decision that can deal with all possible states; Synchronization of decision tables: Tables in the pre-fault queue that have not yet completed the control strategy update are designated as "tables in progress"; after the update is completed, the corresponding records should be promptly sent to the "online decision table", and finally, through periodic updates, the decision tables on the dispatch side are synchronized to the decision tables on the plant side via the communication system.

20. A complex system control decision-making method integrating WRT and SPT according to claim 16, characterized in that: The pre-training of the WRT layer also includes: Intelligent Enhancement: The SPT layer reads the sample results of observables from WAMS / SCADA; it extracts the simulation inference results of observables to be verified from the inference results of the same scenario setting and records them to build a verification sample library; When the simulation results of the power angle swing curve of the same unit generally show significant deviations from the observed sample results in the verification sample library, and it is confirmed that the problem does not belong to the observation sampling problem, the parameter identification function is activated based on the accumulated quantitative analysis knowledge of variable correlation. This guides the generation of new model correction and exploration directions, which are then fed back to the SPT layer. Through the adjustment of the causal simulation model parameters, the simulation model is corrected, so that the model output continuously approaches the actual behavior of the target system.

21. A complex system control decision-making method integrating WRT and SPT according to claim 19, characterized in that: The decision support specifically includes: Enter a new iteration round, query the quantitative analysis results in knowledge extraction, if the system is unstable, find the unstable metasystem, extract its clustering pattern, and find the leading group; otherwise, end the optimization. For each leading unit in the group, the stability margin calculation results after the control are tested based on the EEAC algorithm; the sensitivity is calculated based on the ratio of the stability margin difference before and after control to the unit's power generation. If no measure can improve the system stability margin, the optimization ends; otherwise, all trial measures are sorted in descending order according to sensitivity, the highest-ranked switching measure is applied, and the next round begins. Based on the optimized switching decision and the activation condition of the decision, the activation condition is converted into a feature vector, forming a record in the pre-decision table.

22. A complex system control decision-making method integrating WRT and SPT according to claim 16, characterized in that: The migration execution WRT specifically includes: The system collects operating conditions and fault information of each node in the actual power system, and also collects actual fault information. A new online power system model is constructed based on the collected information; if the online model changes, a new round of pre-training of WRT is started simultaneously to update the pre-decision table; When a fault message is received and the fault is a situation requiring emergency control, the online operating conditions and fault information are immediately merged, and a feature vector is calculated. The vector is compared with the feature vector stored in each record of the pre-decision table. When the matching degree between the two exceeds the threshold, it is considered that a matching record has been found. Then, the control measures are extracted from the record and the emergency control strategy is immediately implemented. If no matching record is found, an emergency control fallback decision is immediately executed.

23. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements a complex system control decision-making method that integrates WRT and SPT according to any one of claims 1 to 14 and 16 to 22; wherein WRT is holistic reduction thinking and SPT is symbol string pre-training.

24. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a complex system control decision-making method that integrates WRT and SPT according to any one of claims 1 to 14 and 16 to 22; wherein WRT is holistic reduction thinking and SPT is symbol string pre-training.

Citation Information

Patent Citations

  • Two-phase method for real time process control

    CN1080409A

  • Emergency generator tripping decision-making method based on knowledge fusion and deep reinforcement learning

    CN115566665A