Complex system control decision method and system fusing WRT and SPT

By integrating WRT and SPT methods, a pre-decision table is constructed, which solves the explanatoryness of the causal-driven model and the credibility of the data-driven model in complex system research, and achieves efficient and reliable real-time decision support, which is suitable for complex systems such as electricity, energy, and environment.

CN120373906AActive Publication Date: 2025-07-25NANJING NARI GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510497823.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-23
Filing Date
2025-04-21
Publication Date
2025-07-25
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

In the research of complex systems, the problem of poor interpretability, model parameters are difficult to quickly and accurately identify and correct, and it is difficult to adapt to high-dimensional problems, and the pure data-driven model has poor interpretability and low credibility in analysis results.

Method used

Combining holistic reduction theory (WRT) and symbol string pre-training technology (SPT), by pre-training the symbol string information of the target system and its objective environment, a pre-decision table containing multiple sets of mappings is constructed, combining self-attention mechanism and causal-driven analysis, coarse-grained classification and deterministic analysis of uncertain factors are realized, pre-decision tables are generated, and control decisions are quickly matched in real-time changes.

Benefits of technology

It provides a real-time analysis and decision-making platform that can extract uncertain information while ensuring the interpretability and reliability of causal analysis, quickly provides trusted and adaptable decision-making suggestions, and improves the accuracy and performance of complex system analysis and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373906A_ABST
    Figure CN120373906A_ABST
Patent Text Reader

Abstract

The invention discloses a complex system control decision method and system fusing WRT and SPT, and the method comprises two layers of architectures: an upper layer carries out the coarse graining classification of uncertain factors according to the structural information and language and other non-structural information of the related fields inside and outside the system, and provides the information needed by the generation of a training sample for a causal algorithm of a lower layer; comprising prediction working conditions corresponding to each learning sample example, disturbance events, control measures which can be called at that time and the like, and a training process of deep reinforcement learning in a high-dimensional decision space is efficiently organized; a lower layer adopts a cause and effect driven deterministic analysis algorithm based on WRT to complete mapping from a system state to a control decision; the two layers cooperate to realize pre-training of decision, and a pre-decision table is periodically updated according to the result. Once the actual disturbance is detected and recognized in the objective system, the actual disturbance is matched with the pre-decision table, and the control decision is quickly completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information science and decision support, and specifically relates to a complex system control decision method integrating holistic reduction theory (WRT) and symbolic string pre-training technology (SPT). Background Art

[0002] Research on complex systems such as energy, electricity, transportation, and the environment needs to consider multi-dimensional factors such as information, physics, and society. The research objects have strong time-varying, strong nonlinear, and strong uncertainty characteristics, and there is an urgent need to achieve accurate and rapid quantitative research through the deep integration of digital intelligence technology and domain knowledge.

[0003] At present, although the research method that solely relies on deterministic causal-driven models has strong interpretability, the input, output, reasoning and other links of the model are easily limited by cognition, and are out of touch with reality and ignore the overall situation. Mismatches between model predictions and actual results are common; model parameters are sometimes difficult to identify quickly and accurately, and it is difficult to quickly correct the model after it is falsified, resulting in the inability of the analysis conclusions to track the rapid evolution of the system in a timely manner; the model is only mechanistically analyzable in low dimensions, and this feature cannot be maintained in high dimensions.

[0004] In summary, the difficulties in cognition, parameter assignment, and mechanism analysis of research methods that solely rely on deterministic causal-driven models limit the expansion of the scale of problems to which such models and methods can be applied, and cannot guarantee solutions to problems of any scale.

[0005] In recent years, pure data-driven models, especially machine learning models, have made rapid progress. Generative large model technologies such as ChatGPT and Sora have shown breakthrough developments. Artificial General Intelligence (AGI) is beginning to emerge. AI-driven scientific research (AI for Science, AI4S) has gradually become a new scientific research paradigm, and the development of digital intelligence technology has ushered in new opportunities. However, such models are usually poorly interpretable, it is difficult to ensure the quality of learning under small samples, and the credibility of the analysis results is low. Summary of the invention

[0006] In order to solve the deficiencies in the prior art, the present invention provides a complex system control decision method and system integrating WRT and SPT. The present invention pre-trains the symbol string information of the target system and its objective environment, constructs a pre-decision table containing multiple groups of mappings, and finds and implements the corresponding control decision according to the real-time changing target system feature vector during migration execution.

[0007] The first aspect of the present invention provides 1. A complex system control decision method integrating WRT and SPT, characterized in that it comprises the following steps: In response to the received symbolic string information of the target system and its objective environment, pre-train the decision-making model that fuses WRT and SPT: The pre-training includes: mapping the uncertainty factors contained in the symbolic string into coarse-grained bins, and for the representative states in each bin, sorting the analysis queue according to the self-attention mechanism, and then optimizing the control decision of each item in the queue based on the overall restored deterministic analysis, completing the one-to-one mapping from the predicted value of the observable quantity to the control decision, and constructing a pre-decision table containing multiple groups of mappings; When the decision-making model is migrated and executed, in response to the real-time change of the feature vector of the target system, find the corresponding control decision and implement it by matching with the conditional feature vector in the pre-decision table.

[0008] The second aspect of the present invention provides a complex system control decision-making method that fuses WRT and SPT, including the following steps: Collect the original corpus of the target system and its objective environment, and extract symbolic string information; Construct a decision-making model that fuses WRT and SPT. The SPT layer is the upper layer of the decision-making model for uncertainty analysis, and the WRT layer is the lower layer of the decision-making model for deterministic analysis; Pre-train the SPT layer, including: mapping the uncertainty factors contained in the symbolic string into coarse-grained bins, constructing an event table containing uncertainty environment knowledge, and generating an event queue to be pre-analyzed; Pre-train the WRT layer, including: the upper SPT layer pushes the event table to the WRT layer, and for the representative states in each bin, optimize the corresponding control decision based on the causality-driven deterministic analysis, and generate a pre-decision table containing deterministic domain quantization knowledge; Implement the migration execution of the decision-making model. In response to the real-time change of the feature vector of the target system, find the corresponding control decision and implement it by matching with the conditional feature vector in the pre-decision table.

[0009] Preferably, the target system includes: a power system, an energy system, an economic system, or an environmental system; The objective environment includes: a set of other systems that interact with the target system; The symbolic string includes: a sequence composed of bytes, numbers, characters, or tokens extracted from the original corpus; the original corpus includes a document library or database form composed of text, sound, pictures, images, numbers, or a mixture of multiple of them.

[0010] Preferably, the pre-training of the SPT layer specifically includes: Mapping the uncertainty factors contained in the symbolic string into coarse-grained bins; Determine the representative state in each bin based on the type of control decision to be trained in the WRT layer deterministic analysis; Construct an event table containing uncertain environment knowledge based on the obtained binned representation of observables; For the representative state in each bin, sort the events in the event table according to the self-attention mechanism to generate an event queue.

[0011] Preferably, the mapping of the uncertainty factors contained in the symbol string to coarse-grained bins includes: For the qualitative description of observables in the symbol string, obtain the quantitative binned representation by querying the semantic library; For the uncertain quantification expression of observables in the symbol string, represent it with a probability density function; For the uncertain quantification expression involving the simultaneous change of multiple observables in the symbol string, represent it with a joint probability density function; for the representation in the form of a probability density function, the binning operation should refer to the settings of the semantic library or the knowledge learned from the function itself; define a one-to-one or one-to-many mapping from one observable to one or more observables through the semantic library.

[0012] Preferably, the determination of the representative state in each bin based on the type of control decision to be trained in the WRT layer deterministic analysis includes: For the preventive control decision to be applied before the occurrence of a risk event, the representative state refers to the expected value of the observables for all samples within the bin; For the emergency control decision to be applied after the occurrence of a risk event, the representative state refers to the extreme value of the observables for all samples within the bin in the most unfavorable direction; The method for obtaining the representative state value should be specified in advance through the semantic library.

[0013] Preferably, the sorting of the events in the event table according to the self-attention mechanism for the representative state in each bin to generate an event queue includes: Estimate the pre-control risk value according to the representative state value of each event in the event table in the coarse-grained bin value; Sort the event table in descending order according to the pre-control risk value to generate an event queue to be pre-analyzed; First, perform the causal-driven deterministic analysis one by one from the events with higher rankings, that is, the risk head events; When parallel computing capabilities are available, the causal-driven deterministic analysis must still be carried out batch by batch for the risk head events.

[0014] Preferably, the pre-training of the WRT layer specifically includes: Simulation deduction: used to deduce the operation process of a target system within a given time range based on a simulation model that reflects the objective laws of the target system under a given system state and decision; Knowledge extraction: used to extract quantitative knowledge beneficial to the achievement of research goals from the simulation parameter trajectory; Decision support: used for a specific decision problem, under a given system state, to explore the values of relevant decision parameters in the simulation deduction model, and then call back simulation deduction and knowledge extraction to generate quantitative analysis results of the characteristics of the target system under the exploratory decision; through repeated exploration and callback, multiple rounds of analysis results are generated; based on the single-round or multi-round analysis results, new next or next group of exploratory directions are generated; repeat the above process until no better results can be obtained through exploration; save the decision with the best performance and the eigenvector of the start condition of this decision as a record in the pre-decision table.

[0015] Preferably, the pre-training of the WRT layer further includes: Intelligence improvement of the decision model, used for the falsification and correction of the decision model.

[0016] Preferably, the simulation deduction specifically includes: In response to the received event table, obtain the simulation model required for deterministic analysis from the SPT layer; In the real-time state, substitute the data reflecting the current state of the target system into this simulation model; in the non-real-time state, substitute the data reflecting the target system at past time or future assumed time into this simulation model; For each event record to be trained, substitute the binned states of each observable quantity related to this event record and map them to the corresponding simulation model parameters to complete the modification of the simulation model; Based on the modified model data, carry out causal simulation deduction to generate a simulation parameter trajectory within a given time range.

[0017] Preferably, the knowledge extraction specifically includes: Construct a series of plane orthogonal mode libraries based on domain expert knowledge, and each mode only contains 2 parameters; Based on the setting of each mode, by regarding the other parameters not involved in this mode as constants, and the other parameters are hereinafter referred to as parameter variables, reduce the simulation model to the second order; by traversing all modes, generate an underdetermined entropy-preserving reduced-order mapping matrix, with a size of M*1, where M is the number of modes; For an element of the underdetermined entropy-preserving reduced-order matrix, expand it into a row; Combine the rows of all modes to generate a well-determined entropy-preserving reduced-order mapping matrix, with a size of M*N, where M is the number of modes and N is the number of time steps; In each element system of the well-posed entropy-preserving reduced-order matrix, a linear analytical method is used for expansion analysis to achieve the quantitative expression of a class of characteristics of each element system; Aggregate the characteristic expressions of all element systems, restore and output the overall characteristics of the target system.

[0018] Preferably, the wisdom improvement specifically includes: Obtain the sample results of each observable quantity collected from the SPT layer; Extract and record the simulation deduction results of the observable quantities to be verified from the deduction results of multiple deterministic cases, and construct a verification sample library; When there are significant deviations between the simulation deduction results and the observed sample results of the same observable quantity in the verification sample library, and it is certain that it does not belong to the problem of observation sampling, based on the quantitative analysis knowledge of variable correlation accumulated by simulation deduction, start the parameter identification function, guide the generation of a new model correction trial direction, feedback it to the SPT layer, and complete the simulation model correction by adjusting the causal simulation deduction model parameters, so that the model output continuously approaches the actual behavior of the target system.

[0019] Preferably, when the migration is executed, the SPT layer pushes the actual feature vector of the target system to the WRT layer, and the WRT layer queries and matches the control measures in the pre-decision table and applies them to the target system; In response to the reinforcement learning task executed in the SPT layer, the WRT layer also provides the SPT layer with the update of deterministic cases or continuously improves the training sample library; or, uses the deterministic cases to support the falsification and correction tasks of the decision-making model to enrich the verification sample library.

[0020] Preferably, the migration execution specifically includes: When receiving the actual feature vector of the target system pushed by the SPT layer, judge whether the change of the target system reaches the migration execution condition; if so, by comparing this vector with the feature vectors saved in each record of the pre-decision table, when the matching degree of the two exceeds the threshold, it is considered that a matching record is found, then extract the control measure from the record and immediately apply it to the target system; if no matching record is found, search for a fallback decision to deal with this problem, and then immediately apply it to the target system; Summarize the specific decisions trained for each possible system state into a set, and perform reinforcement learning based on the data in the set to obtain a general decision that can handle various possible states.

[0021] Preferably, the reinforcement learning specifically includes: In response to the reinforcement learning task executed at the SPT layer, the WRT layer must be able to serve as a deterministic case deduction tool for the SPT layer, simulate the environmental response for each decision trial, generate a sample library to support the reinforcement learning training requirements, and feed it back to the SPT layer; In response to the model correction task executed at the SPT layer, the WRT layer must be able to serve as a deterministic case deduction tool for the SPT layer, obtain the corresponding simulation samples for the observation samples of each observable quantity, and feed it back to the SPT layer.

[0022] In response to the above reinforcement learning or model correction task, the sample directly returned after the simulation deduction step is the time response trajectory of the observable quantity; the sample directly returned after the quantitative analysis step is the quantitative index expression of the target system under the specified state; the sample directly returned after the decision support step is the control measure of the target system under the specified state.

[0023] The third aspect of the present invention provides a complex system control decision-making system integrating WRT and SPT, which operates according to a complex system control decision-making method described in the first aspect, including: a decision-making model; The decision-making model includes: an SPT layer for uncertainty analysis and a WRT layer for deterministic analysis; the SPT layer is the upper layer of the decision-making model, and the WRT layer is the lower layer of the decision-making model; after pre-training the decision-making model, the SPT layer pushes the actual feature vector of the target system to the WRT layer, and the WRT layer queries and matches the control measures in the pre-decision table and applies them to the target system.

[0024] The fourth aspect of the present invention provides a complex system control decision-making method integrating WRT and SPT, which implements adaptive power grid wide-area monitoring, analysis, protection and control based on A-WARMAP. The method includes the following steps: Collect the original corpus of A-WARMAP and its objective environment, construct a scenario and working condition corpus, and extract string information; Construct a decision-making model integrating WRT and SPT. The decision-making model includes: an SPT layer and a WRT layer. The SPT layer is the upper layer of the decision-making model and is used for uncertainty analysis. It implements pre-training for power system security adequacy control decision-making, and sorts each case according to the risk grading value when no control is applied in the pre-fault table; the WRT layer is the lower layer of the decision-making model and is used for deterministic analysis; Pre-train the SPT layer, including: mapping the uncertainty factors contained in the string to coarse-grained grading, constructing an event table containing uncertainty environment knowledge, and generating an event queue to be pre-analyzed; Pre-train the WRT layer, generate the optimal control decision for each fault in the pre-fault queue, and summarize it into a pre-decision table; Migrate and execute WRT. According to the actual working conditions and faults, query the pre-decision table, find the matching decision and execute it.

[0025] Preferably, the collecting of the original corpus of A-WARMAP and its objective environment, building the scene and working condition corpus, and extracting the symbol string information comprises: Collect data reflecting power system operating conditions and scenario changes from any number of information sources on the Internet, dedicated data networks, or buffer networks connecting the two; Preprocessing corpus includes: for structured meteorological data, completing data cleaning, statistics, conversion and integration according to information source and classification; for unstructured warning corpus, directly adding it to the unstructured storage container in the form of a string.

[0026] Preferably, the pre-training of the SPT layer specifically includes: By mapping the uncertainty factors contained in the symbol string into coarse-grained classifications, including: failure probability classification, loss classification and risk classification; Determine and describe the representative status in each bin; Based on the obtained graded representation of observable quantities, an event table containing uncertain environmental knowledge is constructed, including: constructing and sorting a pre-fault table, from the calculation results of the above steps, forming a record in the format of "time, line number, fault probability, pre-control loss, pre-control risk, fault type"; merging all records into the pre-fault table, that is, the event table; sorting the pre-fault table from large to small according to the pre-control risk, forming a pre-fault queue that needs to be pre-analyzed; For the representative states in each bin, the events in the event table are prioritized according to the self-attention mechanism to generate an event queue and form a pre-fault queue that needs to be pre-analyzed.

[0027] Preferably, the pre-training of the WRT layer specifically includes: Simulation: Under given power system status and decision-making, based on the simulation model reflecting the target power system, the system stability behavior within a given time range is simulated; Knowledge extraction: Extract quantitative knowledge that is beneficial to achieving the goal of quantitative stability analysis from simulation parameter trajectories; Decision support: For preventive / emergency / corrective emergency control decision-making problems, explore the optimal values of relevant decision parameters in the simulation model; For the worst possible operating conditions and faults in the target power system, the decisions made at this time are recorded; the preventive / emergency / corrective control decisions obtained from training for each possible system state are summarized into a set, and reinforcement learning is performed based on the data in the set to obtain general emergency control decisions that can cope with various possible states.

[0028] Synchronization of the decision table group: The tables in the pre-fault queue that have not completed the control strategy update are regarded as "tables being updated"; after the update is completed, the corresponding records should be sent to the "online decision table" in a timely manner, and finally, through periodic updates, the decision table group on the dispatching side is synchronized to the decision table group on the substation side via the communication system.

[0029] Preferably, the pre-training of the WRT layer further includes: Wisdom improvement: The SPT layer reads the sample results of observable quantities from WAMS / SCADA; extracts and records the simulation deduction results of observable quantities to be verified from the deduction results of the same scenario setting, and constructs a verification sample library. When there are significant deviations in the simulation deduction results of the power angle swing curves of the same unit and the observed sample results in the verification sample library, and it is certain that it does not belong to the observation sampling problem, based on the accumulated quantitative analysis knowledge of variable correlation, the parameter identification function is activated to guide the generation of a new model correction exploration direction, which is fed back to the SPT layer. Through the adjustment of the causal simulation deduction model parameters, the simulation model is corrected to make the model output continuously approach the actual behavior of the target system.

[0030] Preferably, the decision support specifically includes: Enter a new iteration round, query the quantitative analysis results in knowledge extraction. If the system is unstable, find the meta-system that has lost stability, extract its clustering pattern, and find the leading group; otherwise, end the optimization. For each unit in the leading group, calculate the stability margin calculation results after tentative removal based on the EEAC algorithm; calculate the sensitivity according to the ratio of the difference in stability margins before and after control to the generating power of the unit. If none of the measures can improve the system stability margin, end the optimization; otherwise, sort all the tentative measures in descending order according to the sensitivity, apply the top-ranked generator tripping measure, and enter the next round. Based on the optimized generator tripping decision and the activation condition of this decision, convert the activation condition into a feature vector to form a record in the pre-decision table.

[0031] Preferably, the migration execution of WRT specifically includes: Collect the operating conditions and fault information of each node from the actual power system, and at the same time collect the actual fault information. Construct a new online model of the power system based on the collected information; if there are changes in the online model, synchronously start a new round of pre-training of WRT and update the pre-decision table. When receiving fault information and the fault belongs to the situation that requires emergency control, immediately merge the online operating conditions: fault and other information, and calculate the feature vector. Compare this vector with the feature vectors stored in each record of the pre - decision table. When the matching degree between the two exceeds the threshold, it is determined that a matching record is found. Then, extract the control measures from the record and immediately execute the emergency control strategy; If no matching record is found, immediately execute the fallback decision for emergency control.

[0032] The fifth aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements a complex system control decision - making method integrating WRT and SPT according to the first and third aspects.

[0033] The sixth aspect of the present invention provides a computer - readable storage medium. The computer - readable storage medium stores a computer program. When the computer program is executed by a processor, it implements a complex system control decision - making method integrating WRT and SPT according to the first and third aspects.

[0034] Compared with the prior art, the beneficial effects of the present invention at least include: The present invention provides a complex system control decision - making method integrating WRT and SPT. This method includes a two - layer architecture: the upper layer coarsely classifies uncertain factors according to the structural information and non - structural information such as language in various relevant fields inside and outside the system, and provides the information required to generate training samples for the causal algorithm in the lower layer, including the predicted working conditions, disturbance events, and available control measures corresponding to each learning sample case, etc., efficiently organizing the training process of deep reinforcement learning in a high - dimensional decision space; the lower layer uses a causal - driven deterministic analysis algorithm based on WRT to complete the mapping from the system state to the control decision; the two layers cooperate to achieve pre - training of the decision, and periodically update the pre - decision table according to the results. Once an actual disturbance is detected and identified in the objective system, it is matched with this pre - decision table to quickly complete the control decision.

[0035] In summary, by integrating WRT and SPT, the present invention proposes a brand - new complex system research method and its implementation platform. This method and system provide a real - time analysis and decision - making model based on digital intelligence technology for decision - makers. While ensuring the interpretability and reliability of strict causal analysis, it can extract various uncertain information contained in language / audio / video / digital symbol strings, construct a pre - decision table containing multiple groups of mappings, provide credible, fast, and adaptive decision - making suggestions for preventing various risk events in complex systems, and can also provide direct closed - loop control output in scenarios with high real - time requirements. At the same time, by continuously optimizing the organization of pre - training through reinforcement learning and continuously approaching the state and laws of the real world through model correction, the performance and accuracy of the analysis and decision - making model are further improved.

[0036] The present invention provides a digital intelligent real-time analysis and decision-making platform for online (or offline) decision-makers. While ensuring the reliability and interpretability of strict causal analysis, it incorporates language / voice / image information reflecting uncertain factors to quickly provide trustworthy and adaptive decision-making suggestions for preventing and controlling various risk events in complex systems, and even directly achieve closed-loop control. Therefore, the present invention has broad application prospects and important practical value in the research of complex systems in the field of information science and decision support. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a control block diagram of a complex system control decision-making method integrating WRT and SPT according to an embodiment of the present invention; Figure 2 is a control block diagram of an A-WARMAP system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0039] As Figure 1 shown, Embodiment 1 of the present invention provides a complex system control decision-making method integrating WRT and SPT.

[0040] Generally speaking, as the prominent substantial features of the present invention, one is to incorporate causal rules and conditions into the decision-making process, strengthen the descriptive output of the causal relationship in the decision-making process, expand the transparency of the model, and achieve a balance between the performance and interpretability of the model; the second is to introduce the digital twin technology, generate samples through simulation deduction to make up for the shortage of the number of samples; the third is to introduce links such as mechanism analysis, result verification, error correction, and strategy backup to improve the credibility of the generated results.

[0041] In terms of deterministic causality-driven models, first, regarding the analysis and control of complex power systems, the team where the inventors are located has invented the Extended Equal Area Criteria (EEAC), the only method for quantitatively analyzing the transient stability of power systems that has been theoretically proven and applied in engineering internationally to date. The power system security and stability quantitative analysis and optimization decision-making software FASTEST, with EEAC as the core, has been exported to 29 users in countries such as the United States and Canada, leading the world in terms of both software integrators and the number of users. Subsequently, the Wide Area Monitoring Analysis Protection-control system (WARMAP) was developed and widely applied in engineering. Further, a Cyber-Physical-Social Systemin Energy (CPSSE) framework was constructed to support the study of the power-energy-environment system as a whole. Then, based on the summary and refinement of the common ideas in the research of the above complex systems, the Whole ReductionismThinking (WRT) was proposed. By reducing the high-dimensional model of the complex system to a series of two-dimensional meta-system models, an entropy-preserving mapping from the high-dimensional whole to the low-dimensional population was achieved, bridging the gap between holism and reductionism and creating conditions for the quantitative mechanism research of high-dimensional time-varying nonlinear complex systems.

[0042] This invention is committed to further developing the WRT method in dealing with uncertain and unstructured knowledge, combining it with machine learning techniques to solve the technical problems of poor interpretability of pure data-driven models, difficulty in ensuring the learning quality under small samples, and low credibility of analysis results.

[0043] The core concept of this invention at least includes: collecting the original corpus of the target system and its objective environment, and extracting symbolic string information; in response to the received symbolic string information of the target system and its objective environment, pre-training the decision-making model integrating WRT and SPT, constructing an event table by mapping the uncertainty factors contained in the symbolic string into coarse-grained bins; for the representative states in each bin, sorting the event queue to be analyzed according to the self-attention mechanism, and then optimizing the control decision of each item in the event queue based on the deterministic analysis of whole reduction, completing the one-to-one mapping from the predicted value of the observable quantity to the control decision, and constructing a pre-decision table containing multiple groups of mappings; during migration execution, in response to the real-time change of the feature vector of the target system, finding the corresponding control decision by matching with the conditional feature vector in the pre-decision table and implementing it.

[0044] In a further embodiment, the complex system control decision-making method integrating WRT and SPT includes: Step 1, collect the original corpus of the target system and its objective environment, and extract symbol string information.

[0045] Preferably but not limited to, the target system includes but is not limited to power systems, energy systems, economic systems, environmental systems; the objective environment refers to the set of other systems that interact with the target system, including but not limited to technical, economic, environmental, and social systems; the symbol string refers to a sequence composed of bytes, numbers, characters, or tokens extracted from the original corpus; the original corpus includes but is not limited to documents or databases in the form of text, sound, pictures, images, numbers, or a combination of multiple of them, including but not limited to meteorological disaster warnings, equipment failure alarms, emergency reports, government / market announcements, expert consultation results, social survey / sensing / prediction / statistics / planning / decision-making data, etc.

[0046] Extract symbol strings from the original corpus based on artificial intelligence, identify qualitative or quantitative information related to observable quantities in the symbol strings in combination with the semantic library, and organize them into formatted strings. The artificial intelligence technologies include but are not limited to expert systems, large language models, and other heuristic algorithms and machine learning methods.

[0047] The observable quantity is a state parameter of the target system and its objective environment, including but not limited to physical, economic, and environmental parameters; the predicted quantity of the observable quantity is the quantitative prediction result of the observable quantity, including but not limited to the time, location, probability, intensity of the change of the observable quantity, as well as the losses and risks caused.

[0048] Step 2, construct a decision-making model integrating WRT and SPT. Specifically, the decision-making model includes: the SPT (Symbol-strings Pre-Training) layer and the WRT (Whole Reductionism Thinking) layer. The SPT layer is the upper layer of the decision-making model and is used for uncertainty analysis. The WRT layer is the lower layer of the decision-making model and is used for deterministic analysis. The subsequent steps of pre-training the decision-making model include optimizing the internal data structure and parameters of the SPT layer and the WRT layer.

[0049] Step 3, pre-train the SPT layer, including: map the uncertainty factors contained in the symbol strings into coarse-grained bins, construct an event table containing uncertainty environment knowledge, and generate an event queue to be pre-analyzed.

[0050] Preferably but not limited to, Step 3 specifically includes: Step 3.1, map the uncertainty factors contained in the symbol strings into coarse-grained bins.

[0051] Further preferably but not limited thereto, the hierarchical representation of observables is achieved by the following method: For the qualitative description of observables in the symbol string, obtain the quantitative hierarchical representation by querying the semantic library; For the uncertain quantification expressions of observables in the symbol string, represent them with probability density functions; For the uncertain quantification expressions in the symbol string involving the simultaneous change of multiple observables, represent them with joint probability density functions; for the representation in the form of probability density functions, perform hierarchical operations with reference to the settings of the semantic library or the knowledge learned from the function itself; allow defining a one-to-many mapping from one observable to multiple observables through the semantic library.

[0052] The semantic library includes but is not limited to the following: a mapping table from the key features expressed in the form of symbol strings in the original corpus to the corresponding observables, a mapping table from the qualitative description of observables to the quantitative hierarchical representation, a one-to-many mapping table from the hierarchical classification of one observable to the hierarchical classification of multiple observables, and a value-taking method for the representative states in each hierarchical classification of observables.

[0053] Step 3.2, determine the representative states in each hierarchical classification, and the representative states in each hierarchical classification are determined in combination with the control decision types to be trained in the WRT layer deterministic analysis.

[0054] Further preferably but not limited thereto, for the preventive control decisions to be applied before the occurrence of risk events, the representative state refers to the expected value of the observable in all samples within this hierarchical classification; for the emergency control decisions to be applied after the occurrence of risk events, the representative state refers to the extreme values of the observable in all samples in the most unfavorable direction, including but not limited to the maximum value and the minimum value of all samples; the value-taking method for the representative states should be specified in advance through the semantic library.

[0055] Step 3.3, based on the obtained hierarchical representation of observables, construct an event table containing uncertainty environment knowledge, where the event refers to the process in which one or more observables change due to one uncertainty factor.

[0056] Further preferably but not limited thereto, step 3.3 specifically includes: Start constructing an event from a single observable, multiple independent observables, or multiple simultaneously changing observables; each event must select a hierarchical classification from each relevant observable; Determine the occurrence probability according to the construction source of the event: For a single observable, take the result of the probability density function or the set value in the semantic library; for multiple independent observables, take the product of the probability values of each member observable; for multiple observables that change simultaneously, take the result of the joint probability density function or the set value in the semantic library; each possible binning combination should be constructed as an event, and duplicate events are not allowed to be constructed; Summarize multiple event records into an event table.

[0057] Step 3.4, for the representative states in each binning, sort the events in the event table according to the self-attention mechanism to generate an event queue.

[0058] Further preferably but not restrictively, the steps of the self-attention mechanism are as follows: First, estimate the pre-control risk value according to the representative state value of each event in the event table in the coarse-grained binning value; then, sort the event table in descending order according to the pre-control risk value to generate an event queue to be pre-analyzed; it is necessary to preferentially conduct the causality-driven deterministic analysis one by one from the events with higher rankings, that is, the risk head events; when there is parallel computing ability, it is still necessary to conduct the causality-driven deterministic analysis batch by batch for the risk head events.

[0059] Step 4, pre-train the WRT layer, including: the upper SPT layer pushes the event table to the WRT layer to start the pre-training of the WRT layer; for the representative states in each binning, optimize the corresponding control decisions based on the causality-driven deterministic analysis to generate a pre-decision table containing quantitative knowledge in the deterministic domain; when the number of records in the pre-decision table constructed by the WRT layer reaches the threshold, the decision model reaches the migration condition; when the pre-decision table of the WRT layer is constructed, the training of the decision model ends.

[0060] In a preferred embodiment, a pre-decision table containing multiple sets of mappings is constructed through WRT training, and the process of optimizing the control decisions of each item in the event queue based on the overall restored deterministic analysis includes: simulation deduction, knowledge extraction, decision support, and wisdom improvement.

[0061] Preferably but not restrictively, step 4 specifically includes: Step 4.1: Simulation deduction, which is responsible for deducing the operation process based on the simulation model reflecting the objective laws of the target system within a given time range under the given system state and decision.

[0062] Further preferably but not restrictively, step 4.1 specifically includes: In response to the received event table, obtain the simulation model required for deterministic analysis from the SPT layer, that is, the symbol string pre-training layer; In the real-time state, data reflecting the current state of the target system is substituted into the model; in the non-real-time state, for example but not limited to, for research purposes in the non-real-time state, data reflecting the target system at past time or future assumed time is substituted into the model; it is determined by the user what state of data to substitute into the model; For each event record to be trained, the binned states of each observable related to the event record are substituted and mapped to the corresponding simulation model parameters, completing the modification of the simulation model; Based on the modified model data, causal simulation deduction is carried out to generate the simulation parameter trajectory within a given time range.

[0063] Step 4.2: Knowledge extraction, responsible for extracting quantitative knowledge beneficial to achieving the research goal from the simulation parameter trajectory.

[0064] Further preferably but not restrictively, step 4.2 specifically includes: Construct a series of plane orthogonal mode libraries based on domain expert knowledge, and each mode only contains 2 parameters; Based on the setting of each mode, by treating other parameters not involved in this mode as constants, and other parameters are hereinafter referred to as parametric variables, the simulation model is reduced to the second order; by traversing all modes, an underdetermined entropy-preserving reduced mapping matrix is generated, with a size of M*1, where M is the number of modes; For an element of the underdetermined entropy-preserving reduced matrix, expand it into a row. The method is: at each simulation time step, extract the corresponding parametric variable trajectory from the simulation parameter trajectory generated in step 4.1, and substitute the data of each time step on the trajectory into each parametric variable of the element model; by traversing all time steps, construct a sequence with a size of 1*N, where N is the number of time steps, and each element in the sequence is called a meta-system, and the meta-system is the reduced model of the parametric variable substituted into the simulation results of each time step; combine the rows of all modes to generate a well-determined entropy-preserving reduced mapping matrix with a size of M*N, where M is the number of modes and N is the number of time steps; Use linear analysis methods to expand and analyze in each meta-system of the well-determined entropy-preserving reduced matrix to achieve the quantitative expression of a class of characteristics of each meta-system. The linear analysis methods include but are not limited to: eigenvalue analysis, equal area criterion. The selected meta-system characteristics should be able to support the overall research goal of the target system. The meta-system characteristics include but are not limited to: stability, adequacy, environmental friendliness; Aggregate the characteristic expressions of all meta-systems to achieve the restoration and output of the overall characteristic expression of the target system. The aggregation methods include but are not limited to: taking the minimum value, taking the maximum value, taking the expected value.

[0065] Step 4.3, Decision Support: Responsible for, in response to a specific decision-making problem and given the system state, exploring the values of relevant decision-making parameters in the simulation and deduction model, then calling back Steps 4.1 and 4.2 to generate quantitative analysis results of the characteristics of the target system under the tentative decision; through repeated exploration and callback, generating multiple rounds of analysis results; based on the single-round or multi-round analysis results, generating a new next or next set of exploration directions; repeating the above process until exploration no longer produces results more favorable to the research objective; storing the best-performing decision and the eigenvector of the activation conditions of this decision as a record in the pre-decision table.

[0066] As a further preferred but non-limiting implementation, Step 4 further includes: Step 4.4: Intelligence enhancement of the model, used for falsification and correction of the model.

[0067] Further preferably but non-limitingly, Step 4.4 specifically includes: Obtaining the sample results of each observable quantity collected from the SPT layer; Extracting and recording the simulation and deduction results of the observable quantities to be verified from the deduction results of multiple deterministic cases, and constructing them into a verification sample library; When there are generally significant deviations between the simulation and deduction results and the observed sample results of the same observable quantity in the verification sample library, and it is certain that it does not belong to the observed sampling problem, based on the quantitative analysis knowledge of variable correlation accumulated in Step 4.1, activate the parameter identification function, guide the generation of new model correction exploration directions, feedback to the SPT layer, and complete the simulation model correction through the adjustment of the causal simulation and deduction model parameters, so that the model output continuously approaches the actual behavior of the target system.

[0068] Step 5, after the decision model meets the migration conditions and the pre-decision table in the WRT layer is constructed, implement migration execution. In response to the real-time changes of the target system eigenvector, by matching with the conditional eigenvector in the pre-decision table, find the corresponding control decision and implement it.

[0069] Preferably but non-limitingly, Step 5 specifically includes: During migration execution, the SPT layer pushes the actual eigenvector of the target system to the WRT layer, and the WRT layer queries and matches the control measures in the pre-decision table and applies them to the target system; in response to the reinforcement learning task executed in the SPT layer, the WRT layer also provides the SPT layer with updates of deterministic cases or continuously improves the training sample library; in addition, these cases can also support the falsification and correction tasks of the model and enrich the verification sample library.

[0070] During migration execution, in response to the real-time changes of the target system eigenvector, by matching with the conditional eigenvector in the pre-decision table, find the corresponding control decision and implement it.

[0071] Further preferably but not limited thereto, the specific process of migration execution includes: The detailed process of migration execution is as follows: when the actual feature vector of the target system pushed by the SPT layer is received, first determine whether the change of the target system reaches the migration execution condition; if so, by comparing this vector with the feature vectors saved in each record of the pre-decision table, when the matching degree between the two exceeds the threshold, it is determined that a matching record is found, then the control measures are extracted from the record and immediately applied to the target system; if no matching record is found, search for a fallback decision to address this problem, and then immediately apply it to the target system; the fallback decision includes but is not limited to: for a specific decision problem, a decision generated according to the worst state that may occur in the target system; summarizing the specific decisions obtained by training for each possible system state into a set, and performing reinforcement learning based on the data in the set, so as to obtain a general decision that can handle various possible states.

[0072] Further preferably but not limited thereto, the specific process of reinforcement learning includes: In response to the reinforcement learning task executed in the SPT layer, the WRT layer must be able to serve as a deterministic case deduction tool for the SPT layer, simulate the environmental response for each decision trial, generate a sample library to support the reinforcement learning training requirements, and feedback it to the SPT layer; In response to the model correction task executed in the SPT layer, the WRT layer must be able to serve as a deterministic case deduction tool for the SPT layer, obtain the corresponding simulation samples for the observation samples of each observable quantity, and feedback it to the SPT layer.

[0073] In response to the above reinforcement learning or model correction task, the sample directly returned after the simulation deduction step is the time response trajectory of the observable quantity; the sample directly returned after the quantitative analysis step is the quantitative index expression of the target system in the specified state; the sample directly returned after the decision support step is the control measure of the target system in the specified state.

[0074] Embodiment 2 of the present invention provides a system control system integrating WRT and SPT, including: the decision model described in Embodiment 1, the decision model includes: an SPT layer for uncertainty analysis and a WRT layer for deterministic analysis; the SPT layer is the upper layer of the decision model, and the WRT layer is the lower layer of the decision model; after the decision model is pre-trained as described in Embodiment 1, the SPT layer pushes the actual feature vector of the target system to the WRT layer, and the WRT layer queries the matching control measures in the pre-decision table and applies them to the target system.

[0075] As Figure 2As shown in the figure, Embodiment 3 of the present invention provides a complex system control decision-making method that integrates WRT and SPT. Based on the Adaptive-Wide Area Monitoring Analysis Protection-control system (A-WARMAP) framework, wide-area monitoring, analysis, protection, and control of the power grid with environmental adaptability are implemented. The method includes the following steps: Step 1: Collect the original corpus of A-WARMAP and its objective environment, construct a scenario and operating condition corpus, and extract string information.

[0076] Preferably but not restrictively, Step 1 specifically includes: Step 1.1: Collect the corpus reflecting the changes in the operating conditions and scenarios of the power system from any number of information sources on the Internet, a dedicated data network, or a buffer network connecting the two.

[0077] Step 1.2: Complete the preprocessing of the corpus: For structured meteorological data, conventional methods can be used to complete data cleaning / statistics / transformation / integration according to the information source and classification; for unstructured meteorological disaster warning corpus, it is directly added to the unstructured storage container in the form of a string.

[0078] More preferably but not restrictively, Step 1.2 specifically includes: Construct a semantic library for supporting the parsing of unstructured meteorological corpus. The construction steps include: According to domain experience, offline refine the text pattern features to form a string pattern library of a series of key factors. Each table in the library stores the vocabulary of relevant factors, and these factors include: time, administrative division name, geographical orientation, disaster warning type, warning level, meteorological element type, predicted value, trend type, change degree value...; According to the user's requirements for the degree of coarsening, form a mapping table from meteorological feature strings to bins. A certain mapping in the table is in the form of "gale of level 8 → [1, 4]", where 1 represents the disaster-causing factor of the gale and 4 represents the bin number where the wind of level 7-8 is located. In addition, the semantic library can also record the common sense, rules, and concept system in the field of meteorological disasters. The above information can be in the form of a structured table or in unstructured storage.

[0079] Construct a formatted string based on a large language model, including: First, extract a string from a corpus according to natural language delimiters (such as commas, periods, carriage returns); Second, segment the corpus, extract keywords from it, and match them with the pattern vocabulary stored in the semantic library, and perform pattern encoding according to the matching results; Then, according to the meteorological feature classification mapping table in the semantic library, perform meteorological feature encoding; Again, form a symbolic string model of the line background features within the research scope, including the following parameters: line segment name, line name to which it belongs, starting tower number, ending tower number, geographical features, corridor features, body features,...; Finally, query the matching relationship between the spatial information and the spatial pattern encoding in the line background feature symbolic string model, and combine the meteorological feature encoding information to achieve the meteorological feature encoding of each line.

[0080] Step 2, construct a decision-making model that integrates WRT and SPT, as Figure 2 shown, the A-WARMAP system consists of a two-layer architecture. Specifically, the decision-making model includes: the SPT (Symbol-strings Pre-Training) layer and the WRT (Whole Reductionism Thinking) layer. The SPT layer, as the upper layer of the decision-making model, is used for uncertainty analysis and implements pre-training for power system security adequacy control decisions. It can sort each case according to the risk classification value without control in the pre-fault table. The following takes "generating a pre-fault table from meteorological disaster warning corpus" as an example to illustrate the specific implementation steps of the SPT layer.

[0081] Step 3, perform pre-training on the SPT layer, including: mapping the uncertainty factors contained in the symbol string into coarse-grained classifications, constructing an event table containing uncertainty environment knowledge, and generating an event queue to be pre-analyzed.

[0082] Preferably but not restrictively, Step 3 specifically includes: Step 3.1, by mapping the uncertainty factors contained in the symbol string into coarse-grained classifications, including: fault probability classification, loss classification, and risk classification.

[0083] The failure probability grading includes: First, based on domain knowledge, construct a method for evaluating the failure rate of a line under a given failure type by comprehensively considering one or more meteorological characteristic factors, real-time operating status, maintenance status, etc. of each line; Second, determine whether the number of grades for the same meteorological characteristic on the object is unique. If so, take the extreme values of each meteorological characteristic factor in the most unfavorable direction within the grade where it is located, and calculate the line failure rate; If not, enumerate all permutations and combinations of the grades involved by each factor, take the extreme values of each meteorological characteristic factor in the enumerated grades in the most unfavorable direction in turn, and calculate the line failure rate; Read the grading mapping table in the semantic library, and mark and record all the grades of the failure probability of the line under the given failure type according to the distribution of all the calculated results of the failure rate.

[0084] The loss grading includes: First, based on domain knowledge, construct a method for estimating the pre-control loss of a line by comprehensively considering the importance of each line in the entire power grid topology, real-time power status, etc.; Second, complete the evaluation and record of the line loss.

[0085] The risk grading includes: First, take out the failure probability grading and pre-control loss grading information of each line in turn, and take the extreme value in the most unfavorable direction within the grade as the result; Second, complete the evaluation and record of the pre-control risk of the line by multiplying the loss of the line by the failure probability result.

[0086] Step 3.2, determine the representative state in each of the above-mentioned grades.

[0087] Step 3.3, based on the obtained graded representation of the observable quantities, construct an event table containing knowledge of the uncertain environment.

[0088] Further preferably but not restrictively, Step 3.3 specifically includes: Construct and sort the pre-fault table. First, form a record in the format of "time, line number, failure probability, pre-control loss, pre-control risk, failure type" from the calculation results of the above steps; Second, import all the records into the pre-fault table, that is, the event table; Sort the pre-fault table in descending order of pre-control risk to form a pre-fault queue to be pre-analyzed.

[0089] Step 3.4, for the representative state in each grade, sort the priorities of each event in the event table according to the self-attention mechanism to generate an event queue, that is, sort the pre-fault table in descending order of pre-control risk to form a pre-fault queue to be pre-analyzed. Preferably, considering the limited computing power, take out the head-of-risk faults from the pre-fault queue to be pre-analyzed, and use parallel computing tools to carry out deterministic optimization analysis based on overall restoration batch by batch.

[0090] Step 4: Pre-train the WRT layer. The pre-training of WRT is mainly based on the deterministic analysis of overall restoration, generating the optimal control decisions for each fault in the pre-fault queue and summarizing them into a pre-decision table.

[0091] Preferably but not restrictively, in Step 4, the analysis purposes include: short-term preventive / emergency control of power systems, medium- and long-term preventive / corrective control; the analysis contents include: interrupted power flow, transient security and stability quantitative analysis (EEAC) based on the extended equal area criterion, energy change rate, trajectory characteristic roots; the analysis methods include: searching for the best control strategy based on WRT.

[0092] Preferably but not restrictively, Step 4 specifically includes: Step 4.1: Simulation and deduction, which is used to deduce the system stability behavior within a given time range based on the simulation model reflecting the target power system under the given power system state and decision.

[0093] Further preferably but not restrictively, Step 4.1 specifically includes: In response to the received pre-operation condition and pre-fault table, obtain the electromechanical / electromagnetic transient simulation model required for stability analysis from the SPT layer; Under the online operating state, substitute the data reflecting the current state of the target power system into this model; For each fault record to be trained, substitute the binned state of the observable quantity related to this fault record, such as but not limited to, the power flow binned value of a certain line, and map it to the corresponding simulation model parameter, such as but not limited to the power flow value of a certain line, to complete the modification of the simulation model; Based on the modified model data, conduct power system transient simulation and deduction to generate the simulation parameter trajectory within the required time range, such as but not limited to the rotor angle swing curve of the unit.

[0094] Step 4.2: Knowledge extraction, which is used to extract the quantitative knowledge conducive to the achievement of the stability quantitative analysis goal from the simulation parameter trajectory, such as but not limited to the rotor angle swing curve of the unit.

[0095] Further preferably but not restrictively, Step 4.2 specifically includes: Based on the EEAC method, for each time slice in the trajectory, based on each trajectory gap in this slice, divide the rotor angle swing curves of N units into 2 groups, and within each group, use the inertia weighted formula to aggregate the multi-machine trajectory into a single-machine trajectory to achieve the positioning of the inertia center line; each gap corresponds to a clustering method, and each clustering method is a plane orthogonal mode; summarize all time slices and all gaps on each slice, remove duplicates, and thus construct a plane orthogonal mode library, and each mode only contains 2 groups; Based on the trajectory aggregation under each mode, the simulation model is reduced to the second order; by traversing all modes, an underdetermined entropy-preserving reduced mapping matrix is generated, with a size of M*1, where M is the number of modes; For an element of the underdetermined entropy-preserving reduced matrix, expand it into a row in the following way: at each simulation time step, extract the corresponding parameter variable trajectory from the simulation parameter trajectories generated in Step 4.1.4, and substitute the data of each time step on the trajectory into each parameter variable of the element model; by traversing all time steps, construct a sequence with a size of 1*N, where N is the number of time steps, and each element in the sequence is called a meta-system, and the meta-system is the reduced model of the simulation results with the parameter variables substituted into each time step; combine the rows of all modes to generate a well-determined entropy-preserving reduced mapping matrix with a size of M*N, where M is the number of modes and N is the number of time steps; In each meta-system of the well-determined entropy-preserving reduced matrix, use the EEAC analysis method to calculate the stability margin of the meta-system; Aggregate the characteristic expressions of all meta-systems, take the minimum stability margin among all meta-systems, restore it to the whole, and output it as the expression of the overall stability of the target power system.

[0096] Step 4.3: Decision support, which is used to explore the optimal values of relevant decision parameters in the simulation deduction model for preventive / emergency / corrective control decision problems, such as but not limited to, the total shedding amount of all units at the first moment of a fault and the optimal values of the shedding actions for each specific unit.

[0097] Further preferably but not restrictively, Step 4.3 specifically includes: Step 4.3.1: Enter a new iteration round, query the quantitative analysis results in Step 4.2. If the system is unstable, find the meta-system that loses stability, extract its clustering mode, and find the leading group (generally composed of units that lose synchronous stability); otherwise, end the optimization and enter Step 4.3.5; Step 4.3.2: For each unit in the leading group, based on the EEAC algorithm, tentatively calculate the stability margin after shedding; calculate the sensitivity according to the ratio of the difference in stability margin before and after control to the generating power of the unit; Step 4.3.4: If none of the measures can improve the system stability margin, end the optimization and enter Step 4.3.5; otherwise, sort all tentative measures in descending order according to the sensitivity, apply the top-ranked generator tripping measure, and enter the next round, returning to Step 4.3.1; Step 4.3.5: Based on the optimized generator tripping decision and information such as the starting conditions (such as the corresponding operating conditions and faults) of the decision, convert the starting conditions into feature vectors to form a record in the pre-decision table ( Figure 2 referred to as the on-line decision table).

[0098] Step 4.4: Generate and store fallback decisions.

[0099] Further preferably but not restrictively, Step 4.4 specifically includes: The generation method is as follows: For the most severe working conditions and faults that may occur in the target power system, record the decisions generated at this time; Summarize the emergency control decisions trained for each possible system state into a set, and perform reinforcement learning based on the data in the set to obtain general emergency control decisions that can handle various possible states; The storage method is as follows: Store the decisions generated based on the above method in the dispatching side as an offline decision table and in the substation side as a fallback table.

[0100] Step 4.5: Synchronization of the decision table group. Considering the limited computing power, the tables in the pre-fault queue that have not completed the control strategy update are regarded as "tables being updated"; After the update is completed, the corresponding records should be sent to the "online decision table" in a timely manner. Finally, through periodic updates, the decision table group on the dispatching side is synchronized to the decision table group on the substation side via the communication system.

[0101] Step 4.6: Intelligence improvement, including: The SPT layer reads the sample results of observable quantities from the Wide Area Measurement System (WAMS) or the Supervisory Control and Data Acquisition (SCADA) system. Preferably but not restrictively, such as the sample results of the disturbed generator rotor angle swing curve; Extract and record the simulation deduction results of observable quantities to be verified from the deduction results of the same scenario setting, such as but not limited to the simulation deduction results of the disturbed generator rotor angle swing curve, and construct a verification sample library; When there are significant deviations between the simulation deduction results and the observed sample results of the rotor angle swing curve of the same generator in the verification sample library and it is certain that it does not belong to the problem of observation sampling, based on the quantitative analysis knowledge of variable correlation accumulated in Step 4.1, start the parameter identification function, guide the generation of a new model correction exploration direction, feedback it to the SPT layer, and complete the simulation model correction by adjusting the causal simulation deduction model parameters, so that the model output continuously approaches the actual behavior of the target system.

[0102] Step 5, Migrate and execute WRT. The migration and execution of WRT are responsible for querying the pre-decision table according to the actual working conditions and faults that occur, finding the matching decisions and executing them.

[0103] Preferably but not restrictively, Step 5 specifically includes: Step 5.1: Collect the injection quantities of each node in the actual new power system: Topological quantities: Working condition information such as line power flow, and at the same time collect actual fault information; Step 5.2: Construct an online model of the new power system based on the collected information; if there are changes in the online model, start a new round of pre-training of WRT and update the pre-decision table synchronously. Step 5.3: When a fault message is received and the fault belongs to a situation that requires emergency control, immediately merge the online operating conditions: fault and other information, and calculate the feature vector. Step 5.4: Compare the vector with the feature vectors saved in each record of the pre-decision table. When the matching degree between the two exceeds the threshold, it is determined that a matching record is found. Then, extract the control measures from the record and immediately execute the emergency control strategy. Step 5.5: If no matching record is found, immediately execute the fallback decision for emergency control.

[0104] Embodiment 4 of the present invention provides a computer-readable storage medium storing one or more programs. The one or more programs include instructions that, when executed by a computing device, cause the computing device to execute the complex system control decision method integrating WRT and SPT according to Embodiment 1.

[0105] Embodiment 5 of the present invention provides a computing device, including one or more processors, a memory, and one or more programs. One or more programs are stored in the memory and configured to be executed by the one or more processors. The one or more programs include instructions for executing the complex system control decision method integrating WRT and SPT according to Embodiment 1.

[0106] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present disclosure.

[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A complex system control decision-making method integrating WRT and SPT, characterized in that, It includes the following steps: In response to the received symbolic string information of the target system and its objective environment, pre-train the decision-making model that fuses WRT and SPT: The pre-training includes: mapping the uncertainty factors contained in the symbolic string into coarse-grained bins, sorting the analysis queue by priority according to the self-attention mechanism for the representative states in each bin, and then optimizing the control decisions of each item in the queue based on the globally restored deterministic analysis to complete the one-to-one mapping from the predicted quantity of observable variables to the control decisions, and constructing a pre-decision table containing multiple groups of mappings; When the decision-making model migrates and executes, in response to the real-time change of the target system feature vector, find the corresponding control decision and implement it by matching with the conditional feature vector in the pre-decision table.

2. A complex system control decision-making method integrating WRT and SPT, characterized in that, It includes the following steps: Collect the original corpus of the target system and its objective environment, and extract the symbolic string information; Construct a decision-making model that fuses WRT and SPT. The SPT layer is the upper layer of the decision-making model for uncertainty analysis, and the WRT layer is the lower layer of the decision-making model for deterministic analysis; Pre-train the SPT layer, including: mapping the uncertainty factors contained in the symbolic string into coarse-grained bins, constructing an event table containing uncertainty environment knowledge, and generating an event queue to be pre-analyzed; Pre-train the WRT layer, including: the upper SPT layer pushes the event table to the WRT layer, and based on the causal-driven deterministic analysis for the representative states in each bin, optimize the corresponding control decisions to generate a pre-decision table containing deterministic domain quantization knowledge; Implement the migration and execution of the decision-making model. In response to the real-time change of the target system feature vector, find the corresponding control decision and implement it by matching with the conditional feature vector in the pre-decision table.

3. The complex system control decision-making method that fuses WRT and SPT according to claim 2, characterized in that: The target system includes: a power system, an energy system, an economic system or an environmental system; The objective environment includes: a set of other systems that interact with the target system; The symbolic string includes: a sequence composed of bytes, numbers, characters or tokens extracted from the original corpus; the original corpus includes a document library or database form composed of multiple types of text, sound, pictures, images, numbers or mixtures thereof.

4. The complex system control decision-making method that fuses WRT and SPT according to claim 2 or 3, characterized in that: The pre-training of the SPT layer specifically includes: Mapping the uncertainty factors contained in the symbolic string into coarse-grained bins; Determining the representative states in each bin based on the control decision types to be trained in the deterministic analysis of the WRT layer; Based on the obtained binned representation of the observable variables, constructing an event table containing uncertainty environment knowledge; For the representative states in each bin, sorting the events in the event table by priority according to the self-attention mechanism to generate an event queue.

5. The complex system control decision-making method that fuses WRT and SPT according to claim 4, characterized in that: The mapping of the uncertainty factors contained in the symbolic string into coarse-grained bins includes: For the qualitative description of observable quantities in the symbolic string, the quantitative graded representation is obtained by querying the semantic library; The uncertain quantitative expression of observable quantities in the symbol string is expressed as a probability density function; For uncertain quantitative expressions involving simultaneous changes in multiple observables in a symbol string, a joint probability density function is used to represent them. For the representation in the form of a probability density function, the grading operation should be performed with reference to the settings of the semantic library or the knowledge learned from the function itself. A one-to-one or one-to-many mapping from one observable to one or more observables is defined through the semantic library.

6. The complex system control decision method integrating WRT and SPT according to claim 4 is characterized in that: The determining of the representative state in each bin based on the control decision type to be trained in the WRT layer deterministic analysis includes: For preventive control decisions that need to be applied before a risk event occurs, the representative state refers to the expected value of all samples of the observable quantity in that bin; For emergency control decisions that need to be applied after a risk event occurs, the representative state refers to the extreme value of the observable quantity in the most unfavorable direction for all samples in the bin; The method for obtaining representative state values should be specified in advance through the semantic library.

7. The complex system control decision method integrating WRT and SPT according to claim 4 is characterized in that: For the representative state in each bin, the events in the event table are prioritized according to the self-attention mechanism to generate an event queue, including: Estimate the pre-control risk value based on the representative state value of each event in the coarse-grained classification value in the event table; Sort the event table in descending order according to the pre-control risk value to generate a queue of events that require preliminary analysis; Prioritize the events with the highest ranking, that is, the top risk events, and carry out the causal-driven deterministic analysis one by one; When parallel computing capabilities are available, the causal-driven deterministic analysis described above must still be carried out batch by batch for risk head events.

8. A complex system control decision method integrating WRT and SPT according to claim 2 or 3, characterized in that: The pre-training of the WRT layer specifically includes: Simulation deduction: used to deduce the operation process of the target system based on the simulation model that reflects the objective laws of the target system under given system status and decision-making within a given time range; Knowledge extraction: used to extract quantitative knowledge from simulation parameter trajectories that is beneficial to achieving research goals; Decision support: It is used to test the values of relevant decision parameters in the simulation deduction model under a given system state for specific decision problems, and then call back the simulation deduction and knowledge extraction to generate quantitative analysis results of the target system characteristics under the trial decision; through repeated trials and callbacks, multiple rounds of analysis results are generated; based on single-round or multi-round analysis results, a new next or next group of trial directions is generated; the above process is repeated until the trial no longer produces better results; the decision with the best performance and the feature vector of the starting conditions of the decision are stored as a record in the pre-decision table.

9. The complex system control decision method integrating WRT and SPT according to claim 8, characterized in that: The pre-training of the WRT layer further includes: Wisely improve the decision-making model for the falsification and correction of the decision-making model.

10. A complex system control decision-making method integrating WRT and SPT according to claim 8, characterized in that: The simulation deduction specifically includes: In response to the received event table, obtain the simulation model required for deterministic analysis from the SPT layer; In the real-time state, substitute the data reflecting the current state of the target system into the simulation model; in the non-real-time state, substitute the data reflecting the target system at a past time or a future assumed time into the simulation model; For each event record to be trained, substitute the binned states of each observable quantity related to the event record into and map them to the corresponding simulation model parameters to complete the modification of the simulation model; Based on the modified model data, carry out causal simulation deduction to generate the simulation parameter trajectory within a given time range.

11. A complex system control decision-making method integrating WRT and SPT according to claim 8, characterized in that: The knowledge extraction specifically includes: Construct a series of planar orthogonal mode libraries based on domain expert knowledge, and each mode only contains 2 parameters; Based on the setting of each mode, by treating other parameters not involved in the mode as constants (hereinafter referred to as parameter variables), reduce the simulation model to the second order; by traversing all modes, generate an underdetermined entropy-preserving reduced-order mapping matrix with a size of M*1, where M is the number of modes; For an element of the underdetermined entropy-preserving reduced-order matrix, expand it into a row; Combine the rows of all modes to generate a well-determined entropy-preserving reduced-order mapping matrix with a size of M*N, where M is the number of modes and N is the number of time steps; Use linear analysis methods to expand and analyze each element system in the well-determined entropy-preserving reduced-order matrix to achieve the quantitative expression of a class of characteristics of each element system; Aggregate the characteristic expressions of all element systems, restore and output the overall characteristics of the target system.

12. A complex system control decision-making method integrating WRT and SPT according to claim 9, characterized in that: The wise improvement specifically includes: Obtain the sample results of each observable quantity collected from the SPT layer; Extract and record the simulation deduction results of the observables to be verified from the deduction results of multiple deterministic cases, and construct a verification sample library; When there are significant deviations in general between the simulation deduction results and the observed sample results of the same observable quantity in the verification sample library, and it is certain that it does not belong to the observed sampling problem, based on the quantitative analysis knowledge of variable correlation accumulated in the simulation deduction, start the parameter identification function, guide the generation of a new model correction trial direction, feedback it to the SPT layer, and complete the correction of the simulation model by adjusting the causal simulation deduction model parameters, so that the model output continuously approaches the actual behavior of the target system.

13. A complex system control decision-making method integrating WRT and SPT according to claim 2 or 3, characterized in that: When migrating and executing, the SPT layer pushes the actual feature vector of the target system to the WRT layer, and the WRT layer queries and matches the control measures in the pre-decision table and applies them to the target system; In response to the reinforcement learning task executed at the SPT layer, the WRT layer also provides deterministic case updates or continuously improves the training sample library for the SPT layer simultaneously; or, it will use deterministic cases to support the falsification and correction tasks of the decision-making model to enrich the verification sample library.

14. A complex system control decision-making method integrating WRT and SPT according to claim 2 or 3, characterized in that: The migration execution specifically includes: When receiving the actual feature vector of the target system pushed by the SPT layer, determine whether the change of the target system reaches the migration execution condition; if so, by comparing this vector with the feature vectors saved in each record of the pre-decision table, when the matching degree between the two exceeds the threshold, it is considered that a matching record is found, then extract the control measures from the record and immediately apply them to the target system; if no matching record is found, search for a fallback decision to handle this problem, and then immediately apply it to the target system; Summarize the specific decisions obtained by training for each possible system state into a set, and perform reinforcement learning based on the data in the set to obtain a general decision that can handle various possible states.

15. A complex system control decision-making method integrating WRT and SPT according to claim 14, characterized in that: The specific content of the reinforcement learning includes: In response to the reinforcement learning task executed at the SPT layer, the WRT layer must be able to act as a deterministic case deduction tool for the SPT layer, simulate the environmental response for each decision-making trial, generate a sample library to support the reinforcement learning training requirements, and feedback it to the SPT layer; In response to the model correction task executed at the SPT layer, the WRT layer must be able to act as a deterministic case deduction tool for the SPT layer, obtain the corresponding simulation samples for the observation samples of each observable quantity, and feedback it to the SPT layer. In response to the above reinforcement learning or model correction task, the samples directly returned after the simulation deduction step are the time response trajectories of the observable quantities; the samples directly returned after the quantitative analysis step are the quantitative index expressions of the target system in the specified state; the samples directly returned after the decision support step are the control measures of the target system in the specified state.

16. A complex system control decision-making system integrating WRT and SPT, operating according to a complex system control decision-making method as claimed in any one of claims 1 to 15, characterized in that, Including: A decision-making model; The decision-making model includes: an SPT layer for uncertainty analysis and a WRT layer for deterministic analysis; the SPT layer is the upper layer of the decision-making model, and the WRT layer is the lower layer of the decision-making model; after pre-training the decision-making model, the SPT layer pushes the actual feature vector of the target system to the WRT layer, and the WRT layer queries the matching control measures in the pre-decision table and applies them to the target system.

17. A complex system control decision-making method integrating WRT and SPT, characterized in that Based on A-WARMAP, implement adaptive wide-area monitoring, analysis, protection, and control of the power grid. The method includes the following steps: Collect the original corpus of A-WARMAP and its objective environment, construct a scenario and working condition corpus, and extract string information; A decision model integrating WRT and SPT is constructed. The decision model includes: SPT layer and WRT layer. The SPT layer is the upper layer of the decision model, which is used for uncertainty analysis, implements pre-training of power system safety and sufficiency control decisions, and sorts each case in the pre-fault table according to the risk classification value when no control is added; the WRT layer is the lower layer of the decision model, which is used for deterministic analysis; Pre-training the SPT layer includes: mapping the uncertainty factors contained in the symbol string into coarse-grained bins, constructing an event table containing uncertainty environment knowledge, and generating an event queue that needs to be pre-analyzed; Pre-train the WRT layer to generate the optimal control decision for each fault in the pre-fault queue and summarize it into a pre-decision table; Migrate and execute WRT. According to the actual working conditions and faults, query the pre-decision table, find the matching decision and execute it.

18. The complex system control decision method integrating WRT and SPT according to claim 17, characterized in that: The collecting of the original corpus of A-WARMAP and its objective environment, the building of the scene and working condition corpus, and the extraction of symbol string information include: Collect data reflecting power system operating conditions and scenario changes from any number of information sources on the Internet, dedicated data networks, or buffer networks connecting the two; Preprocessing corpus includes: for structured meteorological data, completing data cleaning, statistics, conversion and integration according to information source and classification; for unstructured warning corpus, directly adding it to the unstructured storage container in the form of a string.

19. The complex system control decision method integrating WRT and SPT according to claim 17, characterized in that: The pre-training of the SPT layer specifically includes: By mapping the uncertainty factors contained in the symbol string into coarse-grained classifications, including: failure probability classification, loss classification and risk classification; Determine and describe the representative status in each bin; Based on the obtained graded representation of observable quantities, an event table containing uncertain environmental knowledge is constructed, including: constructing and sorting a pre-fault table, from the calculation results of the above steps, forming a record in the format of "time, line number, fault probability, pre-control loss, pre-control risk, fault type"; merging all records into the pre-fault table, that is, the event table; sorting the pre-fault table from large to small according to the pre-control risk, forming a pre-fault queue that needs to be pre-analyzed; For the representative states in each bin, the events in the event table are prioritized according to the self-attention mechanism to generate an event queue and form a pre-fault queue that needs to be pre-analyzed.

20. The complex system control decision method integrating WRT and SPT according to claim 17, characterized in that: The pre-training of the WRT layer specifically includes: Simulation: Under given power system status and decision-making, based on the simulation model reflecting the target power system, the system stability behavior within a given time range is simulated; Knowledge extraction: Extract quantitative knowledge that is beneficial to achieving the goal of quantitative stability analysis from simulation parameter trajectories; Decision support: For preventive / emergency / corrective emergency control decision-making problems, explore the optimal values of relevant decision parameters in the simulation model; Record the decisions made at this time for the most severe working conditions and faults that may occur in the target power system; summarize the preventive / emergency / corrective control decisions obtained through training for each possible system state into a set, and perform reinforcement learning based on the data in the set to obtain general emergency control decisions that can handle various possible states. Synchronization of decision table groups: The tables in the pre-fault queue that have not completed the control strategy update are regarded as "tables being updated"; after the update is completed, the corresponding records should be sent to the "online decision table" in a timely manner, and finally, through periodic updates, the decision table group on the dispatching side is synchronized to the decision table group on the substation side via the communication system.

21. A complex system control decision method integrating WRT and SPT according to claim 17, characterized in that: The pre-training of the WRT layer further includes: Wisdom improvement: The SPT layer reads the sample results of observable quantities from WAMS / SCADA; extracts and records the simulation deduction results of observable quantities to be verified from the deduction results of the same scenario setting, and constructs a verification sample library. When there are significant deviations in the simulation deduction results of the power angle swing curves of the same unit and the observed sample results in the verification sample library, and it is certain that it does not belong to the problem of observation sampling, based on the accumulated quantitative analysis knowledge of variable correlation, start the parameter identification function, guide the generation of a new model correction exploration direction, feedback it to the SPT layer, and complete the simulation model correction by adjusting the causal simulation deduction model parameters, so that the model output continuously approaches the actual behavior of the target system.

22. A complex system control decision method integrating WRT and SPT according to claim 17, characterized in that: The decision support specifically includes: Enter a new iteration round, query the quantitative analysis results in knowledge extraction. If the system is unstable, find the meta-system that has lost stability, extract its clustering pattern, and find the leading group; otherwise, end the optimization. For each unit in the leading group, calculate the stability margin calculation results after tentative removal based on the EEAC algorithm; calculate the sensitivity according to the ratio of the difference in stability margins before and after control to the generating power of the unit. If none of the measures can improve the system stability margin, end the optimization; otherwise, sort all the tentative measures in descending order according to the sensitivity, apply the top-ranked generator tripping measure, and enter the next round. Based on the optimized generator tripping decision and the activation condition of this decision, convert the activation condition into a feature vector to form a record in the pre-decision table.

23. A complex system control decision method integrating WRT and SPT according to claim 17, characterized in that: The migration and execution of WRT specifically includes: Collect the working conditions and fault information of each node from the actual power system, and at the same time collect the actual fault information. Construct an online model of the new power system according to the collected information; if there are changes in the online model, simultaneously start a new round of pre-training of WRT and update the pre-decision table. When receiving fault information and the fault belongs to the situation that requires emergency control, immediately merge the online working conditions: fault and other information, and calculate the feature vector. Compare this vector with the feature vectors saved in each record of the pre-decision table. When the matching degree between the two exceeds the threshold, it is determined that a matching record has been found. Then, extract the control measures from the record and immediately execute the emergency control strategy; If no matching record is found, immediately execute the fallback decision for emergency control.

24. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements a complex system control decision method that fuses WRT and SPT according to any one of claims 1 to 15, 17 to 23.

25. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a complex system control decision method that fuses WRT and SPT according to any one of claims 1 to 15, 17 to 23.

Citation Information

Patent Citations

  • Two-phase method for real time process control

    CN1080409A

  • Emergency generator tripping decision-making method based on knowledge fusion and deep reinforcement learning

    CN115566665A

  • New energy delivery scale prediction method based on system dynamics model

    CN117217488A

  • Vehicle decision control model training method, device and equipment and vehicle decision control method, device and equipment

    CN117755341A

  • System controller for controlling a control network having an open communication protocol via proprietary communication

    US20030074433A1