Microgrid safety response method for M3M-ToT self-evolution reasoning
Through the M3M-ToT self-evolutionary reasoning method, combined with the large language reasoning model and the MITRE framework, a microgrid security response strategy is constructed, which solves the problem that traditional strategies are difficult to adapt to rapid change attacks, and realizes the adaptive and efficient defense of the microgrid security strategy.
Patent Information
- Application Number
- CN202510529924.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional microgrid security strategies are difficult to adapt to the rapidly changing network environment and evolving attack methods. Especially when facing large-scale and diversified security incidents, how to dynamically adjust security strategies has become a challenge.
The M3M-ToT self-evolutionary reasoning method is adopted to build a closed-loop knowledge graph of "attack-defense-counter" through a large language inference model and MITRE Shield, ATT&CK and Engage framework. The mapping relationship is iteratively updated using offense and defense data to generate a variety of potential defense strategies, and optimize the defense strategy through the self-evolution mechanism.
It has achieved flexible adjustment and self-evolution of the microgrid security defense system, can effectively respond to complex threats, improve the accuracy and efficiency of defense strategies, avoid resource conflicts, and ensure network security and stability.
Smart Images

Figure CN120455051A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of microgrid security protection, and specifically to a microgrid security response method based on M3M-ToT self-evolutionary reasoning. Background Art
[0002] With the rapid development of information technology, microgrid cybersecurity issues are becoming increasingly prominent. Traditional security policies often rely on predefined rules and signatures, making them difficult to adapt to rapidly changing network environments and evolving attack vectors. Against this backdrop, the MITRE Engage and MITRE Shield frameworks have emerged, providing new approaches and methods for cybersecurity protection. MITRE Engage emphasizes planning and executing adversary engagement strategies to help organizations more proactively address cyber threats and enhance their defense capabilities. MITRE Engage and MITRE Shield not only focus on comprehensive monitoring and analysis of attack behavior but also help organizations promptly identify and respond to various security threats through attack lifecycle management, threat intelligence sharing, and automated response. They are closely integrated with the MITRE ATT&CK framework, providing defenders with a wealth of resources and tools to help them better understand and counter adversary tactics, techniques, and procedures. However, even with frameworks like MITRE Engage, dynamic adaptation of security policies remains a challenge, especially in the face of large-scale and diverse security incidents. Enforcing appropriate security policies in response to evolving threats remains a pressing issue. Summary of the Invention
[0003] This paper proposes a microgrid security response method based on M3M-ToT self-evolutionary reasoning. This method uses a mind tree to perform reasoning through a large language reasoning model, and uses attack and defense data to iteratively update the mapping relationship. It provides multiple potential defense strategies that match the current attack type, providing a comprehensive set of alternative solutions for subsequent screening of the optimal strategy. The self-evolutionary mechanism further enhances the dynamic adaptability of security strategies, providing strong support for network security protection. The technical solutions provided by this paper are as follows:
[0004] A microgrid security response method based on M3M-ToT self-evolutionary reasoning includes the following steps:
[0005] S1, when the microgrid is attacked, through network situation awareness, the defense strategy library is searched to identify whether there is a defense strategy that matches the current attack pattern. If a matching attack pattern is retrieved, the corresponding defense strategy is executed after expert decision-making; if no matching attack pattern is retrieved, a candidate strategy set is generated and entered into the subsequent process;
[0006] S2: Collect and integrate mapping relationships in the MITRE Shield, ATT&CK, and Engage frameworks, build an "attack-defense countermeasure" closed-loop knowledge graph, and establish a three-dimensional mapping equation for the M3M model.
[0007] S3, inputs attack-related data into the large language reasoning model to extract key information, inputs the key information into the thinking tree TOT reasoning framework built based on the M3M model, and generates defense strategies by reasoning;
[0008] S4, the strategies generated by the expert evaluation thinking tree reasoning framework are selected as new defense strategies that can be effectively implemented and achieve the expected defense effect as new defense items;
[0009] S5, the new defense items generated are used to study and summarize the entire defense process through the thinking tree framework, and the actual effect of the defense strategy is analyzed using the association matrix adaptation and conflict detection optimization algorithm.
[0010] S6, adds the new defense items learned and summarized to the historical items of the defense strategy library, so that the strategy library is updated.
[0011] Preferably, the multi-dimensional information collected in the S1 search phase includes: timestamp, event type, event source IP address, event source MAC address, asset number, unique identifier of the response strategy, indicative threat information, network traffic, audit logs, policy change records, event urgency, device type, problem type, device status, geographic location code, associated threat pattern ID, trigger condition, response action, causal chain identifier, policy scope and machine learning model ID.
[0012] Preferably, the three-dimensional mapping equation of the M3M model is as follows:
[0013]
[0014] in, is the ATT&CK attack vector, Y=σ(W xy X) is the Shield defense strategy, and Z is the final output;
[0015] Defense explanation part is the defense-countermeasure correlation matrix, β is a parameter, represents tensor product, ⊙ represents element-wise multiplication;
[0016] Active counter part W xz X is a linear transformation where W xy Is the correlation matrix, which linearly combines the input attack vector X; γX °2 is a nonlinear adversarial term, where γ is the coefficient, X°2 It represents a square operation on each element of the attack vector X, and ReLU is a nonlinear activation function.
[0017] Preferably, in the TOT reasoning framework, each node represents a semantic unit, and the nodes are connected by edges to form an association network. When an input sequence is received, the most relevant nodes are found in the TOT according to the sequence, and output is generated based on these nodes. The quality of each node is evaluated by the value function V(s). For the node state s, its value can be expressed as:
[0018] V(s)=O·P LM (s)+(1-O)·H(s)
[0019] Among them, P LM (s) is the probability of the language model generating state s, H(s) is the heuristic function; O is the weight coefficient. For the current node s t , the formula for generating candidate branches is expressed as:
[0020]
[0021] Among them, C(s t ) represents the current node s t A set of candidate branches, each element in the set is from s t A possible subsequent state or branch generated, s t Indicates the current node or state, Represents the i-th candidate branch or subsequent state generated from the current node, and the subscript t+1 indicates that this is relative to the current node s t The next state of , and the superscript i is used to distinguish different candidate branches.
[0022] Preferably, the mind tree reasoning framework iterates and optimizes the mapping relationship between attack features and defense actions using a defense confidence update formula based on the reasoning results and the actual attack and defense situation. The defense confidence update formula is as follows:
[0023]
[0024] Among them, β (t) is the parameter value at the tth iteration, which is obtained by the parameter value β of the previous iteration (t-1) Update to get, β (t-1) is the parameter value at the t-1th iteration, and is the parameter value before the update; It is a learning rate-related term, which represents a time decay factor or the total number of time steps. It is used to control the size of the learning rate. As the number of iterations increases, the learning rate gradually decreases, which helps to gradually converge during the optimization process. is the gradient of the objective function with respect to the parameter β, is the objective function, which is the weighted sum of multiple terms, and the gradient Represents the rate of change of the objective function to the parameter β, which is used to guide the direction of parameter update; geo w eight i Represents the weight of the i-th one, which is used to weight different contributions in the objective function.
[0025] Preferably, the adaptive formula of the correlation matrix is as follows:
[0026]
[0027] Among them, W xy is the correlation matrix, which is used to store the correlation relationship between the input vector X and the output vector Y. t is the time variable, which represents the change of the correlation matrix over time. λ is the urgency driving parameter, which is used to control the learning rate and determines how quickly the system adjusts the correlation matrix in an emergency. A larger λ value means that the system responds faster to emergencies. X is the attack vector and Y is the defense strategy. is the conflict penalty term. When the system detects a conflict, this term will increase, thereby slowing down the update speed of the association matrix.
[0028] Preferably, the conflict detection algorithm is specifically as follows: in the detect_conflict function, first, the masks of the two strategies s1 and s2 are bitwise ANDed, that is, s1.mask & s2.mask. If the result is not 0, it means that the ranges of the two strategies overlap and there may be a conflict; then, the actions of the two strategies are bitwise ANDed, s1.action & s2.action, and it is determined whether the result is equal to 0b11. If it is equal to 0b11, it means that the actions of the two strategies are 1 at the same time in some bits, and there is a conflicting action combination; only when the above two conditions are met at the same time, it is determined that the two strategies are in conflict, and the function returns true, otherwise it returns false.
[0029] Preferably, the method further comprises the following steps:
[0030] S7, regularly optimize and organize the strategy library, remove outdated or inefficient strategies, and establish a mathematical model combining dynamic threshold screening and time decay factors:
[0031]
[0032] Retention conditions: J i (q)≥θ
[0033] Among them, the scoring function J i(q) represents the comprehensive utility score of strategy i at time t, which is composed of the original utility and efficiency index modified by timeliness. The utility parameter F i (q) represents the baseline utility value of strategy i at time q, calculated through historical data or real-time feedback; time decay factor Used to control the natural decay of the strategy utility over time, K represents the decay rate parameter (K>0), q0 represents the initial introduction time of the strategy; efficiency ratio Indicates the input-output efficiency of the quantitative strategy, η i (q) represents the profit index of the current time window, Ψ i (q) represents the corresponding resource consumption index, the weight coefficient L, Used to balance the importance of utility continuity and immediate efficiency, satisfying The screening threshold θ is used to dynamically adjust the retention critical value, which is automatically adjusted according to the scale of the strategy library:
[0034] θ(q)=μ J (q)-kσ J (q)
[0035] Among them, μ J (q) is the current average score, σ J (q) is the standard deviation, k is the adjustment coefficient;
[0036] Regularly recalculate the scores of all strategies, calculate the mean and standard deviation of the current score distribution, and automatically eliminate inefficient strategies based on the retention criteria:
[0037]
[0038] where ε q represents the strategy set of the qth generation, ε q+1 represents the set of strategies of the q+1th generation that remain after elimination, ε i represents the i-th strategy, J i (q) represents the score of the i-th strategy in the q-th generation; θ(q) represents the elimination threshold, and only when the score of the strategy J i (q) is greater than or equal to the elimination threshold θ(q), the strategy will be retained to the next generation. By comparing the score of each strategy with the elimination threshold, the strategies with lower scores are automatically eliminated.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] First, within the microgrid's security defense system, neural network technology is introduced to optimize the defense strategy update mechanism. The core of this process is to calculate gradients by analyzing current attack and defense data. Parameters are then updated based on defense confidence, allowing for flexible adjustments to the defense strategy to address emerging threats. Based on new attack and defense data, the system calculates the gradient of the current defense strategy—indicating the direction and degree of strategy adjustment. Combined with a preset learning rate, the strategy parameters are gradually adjusted, optimizing the defense strategy to better address new and complex threats. During the training and inference process of the large language model, model parameters are continuously adjusted to improve performance. By analyzing current attack and defense data to calculate gradients and updating and adjusting the defense strategy based on defense confidence, the system improves its ability to respond to new and complex threats.
[0041] Second, a dynamic mapping model was constructed, comprehensively considering multiple factors, including attack characteristics, defense strategies, and the microgrid, to map attack feature vectors to corresponding defense actions. This mapping equation is continuously optimized through iterative updates of attack and defense data. This mapping equation is presented as a three-dimensional mapping equation, the core function of which is to map the input attack vector to the output. It consists of two modules: "Defense Interpretation" and "Active Countermeasures." Through mathematical operations such as linear transformations, nonlinear activation functions, tensor products, and element-by-element multiplication, the model can flexibly adjust defense strategies based on complex attack scenarios and generate effective countermeasures, thereby enhancing the security of the microgrid.
[0042] Third, by collecting feedback data after defense strategy implementation, using algorithms to analyze and optimize the defense strategy library, and introducing conflict penalties to avoid resource conflicts, a closed-loop adaptive defense system is formed to continuously improve the accuracy and efficiency of defense strategies. In microgrid defense scenarios, defense strategies may compete for limited resources, affecting overall defense effectiveness. By introducing conflict penalties during the defense strategy optimization process, when resource competition arises between different defense strategies, the system will suppress such conflicts by adding conflict penalties. This not only promotes more scientific and reasonable resource allocation in the microgrid, but also ensures that various defense strategies are executed in a coordinated and consistent manner, further improving overall defense effectiveness and self-evolution. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0044] Figure 1 It is a flow chart of the solution of the present invention;
[0045] Figure 2 It is a functional module diagram of the present invention. DETAILED DESCRIPTION
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0047] In order to make the above-mentioned objects, features and effects of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0048] Example 1:
[0049] The scenario discussed in this method is the microgrid intranet scenario, not the external Internet business. It focuses on the network security protection within the microgrid and responds to the network attacks that may be suffered within the microgrid. This method uses a large language reasoning model to perform reasoning, obtains the iterative update mapping relationship of attack and defense data, and provides a variety of potential defense strategies that match the current attack type, providing a comprehensive alternative plan for the subsequent screening of the optimal strategy, while achieving a self-reinforcement effect. In the security protection of microgrids, this method can flexibly adjust the defense strategy according to the real-time attack situation, and effectively respond to the ever-changing network threat environment. Figure 1 、 2 As shown in FIG, a microgrid security response method based on M3M-ToT self-evolutionary reasoning is specifically provided, which includes the following steps:
[0050] S1, in the search phase of microgrid network situational awareness, the defense strategy library is retrieved to identify whether there is a defense strategy that matches the current attack pattern. If a matching attack pattern is retrieved, the corresponding defense strategy is executed after expert decision-making to effectively deal with the current network threat; if a matching attack pattern is not retrieved, a candidate strategy set is generated to enter the subsequent process so as to enter the subsequent analysis and decision-making process and further formulate appropriate defense measures to ensure the security of the microgrid network. The multi-dimensional information collected in the search phase includes: timestamp, event type, event source IP address, event source MAC address, asset number, unique identifier of the response strategy, indicative threat information, network traffic, audit log, policy change record, event urgency, device type, problem type, device status, geographic location code, associated threat pattern ID, trigger condition, response action, causal chain identifier, policy scope and machine learning model ID. The following is a detailed explanation based on actual conditions:
[0051] Matching an attack pattern: Suppose the system detects that the microgrid network is under a DDoS attack, characterized by a large number of abnormal traffic requests within a short period of time. The system searches the defense policy library and finds defense policies that match the DDoS attack, such as traffic cleaning and traffic blocking. After expert decision-making, considering business continuity and the impact on users, the system ultimately chooses to implement the traffic cleaning policy. Using specialized cleaning equipment, the attack traffic is cleaned, preserving normal business traffic, effectively countering the current DDoS attack and ensuring the normal operation of the microgrid network.
[0052] Situation where no attack pattern is matched: Suppose the system detects a new type of malware attack within the microgrid network. Its attack pattern differs from previously known patterns, being more stealthy and destructive. The system fails to find a matching defense strategy in the defense strategy library. In this case, the system constructs a candidate strategy set, which may include isolating infected devices, updating the virus database, and restricting network access. After entering the subsequent analysis and decision-making process, the system utilizes a mindset tree and a large language model, and expert decision-making ultimately develops appropriate defense measures. These measures may include isolating infected devices to prevent further spread of the malware, while also urgently updating the virus database to identify and remove the malware and ensure the security of the microgrid network.
[0053] S2 builds the M3M model, a three-dimensional mapping model based on the MITRE framework for active defense coordination in microgrid security response. By integrating mapping relationships from the MITRE Shield, ATT&CK, and Engage frameworks, a comprehensive "attack-defense-countermeasure" closed-loop knowledge graph is constructed. The core of M3M lies in its three-dimensional mapping equation, which can map attack vectors to defense strategies and countermeasures, thereby achieving precise defense and active countermeasures against complex network attacks. The specific steps are:
[0054] S2.1, collect and integrate the mapping relationships in MITRE Shield, ATT&CK and Engage frameworks to build a closed-loop knowledge graph of "attack-defense-counterattack". This is similar to building a bridge between different knowledge systems, associating and integrating key information such as attack patterns, defense strategies and countermeasures scattered in various frameworks, and laying a solid foundation for subsequent model construction. The knowledge graph is not a simple list, but treats attack methods, defense mechanisms and countermeasures as interrelated nodes and edges, and through systematic combing and presentation, reveals the intricate internal connections between them. For example, a specific phishing attack (T1566 in the ATT&CK framework) is closely linked to the corresponding email filtering defense strategy (a measure in MITRE Shield) and the corresponding tracing countermeasures to form an organic knowledge system, which provides strong support for in-depth analysis and decision-making.
[0055] Some attack-defense mapping relationships are shown in Table 1.
[0056]
[0057]
[0058] S2.2, construct a three-dimensional mapping equation to map the input attack vector X to the output Z, achieving precise defense and proactive countermeasures against complex network attacks. This equation consists of two parts: "defense interpretation" and "active countermeasures." It combines multiple operations such as linear transformations, nonlinear activation functions, tensor products, and element-by-element multiplication to flexibly adjust defense strategies and generate effective countermeasures. This three-dimensional mapping equation is as follows:
[0059]
[0060] in, is the ATT&CK attack vector (dimension = 852), Y = σ(W xy X) is the Shield defense strategy (including 230 non-zero elements), is the defense-countermeasure correlation matrix (interpretability core), γX °2 is a nonlinear adversarial term (enhanced APT defense).
[0061] Y=σ(W xy X) This part is the Shield defense strategy, where W xy It is an association matrix that performs a linear transformation on the input attack vector X to generate an intermediate representation. Then, the intermediate representation is nonlinearly transformed through the activation function σ to obtain the defense strategy Y. Here, σ is a Sigmoid function or other type of activation function used to map the result of the linear transformation to a specific range, such as (0, 1). It is a defense-countermeasure association matrix, which can be mathematically regarded as a linear transformation matrix. It is used to map the defense strategy Y to a higher-dimensional space or perform some specific combination operation. This matrix is the core of interpretability, which means it plays a key role in the interpretability of the model. represents the tensor product or Kronecker product, a matrix operation used to combine two matrices into a larger matrix. It is used to perform a tensor product operation on the defense strategy Y and the parameter β, generating a more complex matrix structure. ⊙ represents element-by-element multiplication, which multiplies the corresponding elements of the two matrices. It is used to multiply the correlation matrix α by the tensor product result obtained previously, element-by-element, to generate the output of the defense explanation component.
[0062] In the active countermeasure part, W xz X is a linear transformation where W xy Is an association matrix that linearly combines the input attack vector X to generate a preliminary countermeasure signal. °2 is a nonlinear countermeasure term used to enhance the defense capability against advanced persistent threats (APT), where γ is a coefficient and X °2 ReLU represents the square operation of each element of the attack vector X. This is a mathematically nonlinear transformation used to capture nonlinear features in the input data. ReLU is a nonlinear activation function, also known as a rectified linear unit. ReLU introduces nonlinearity, enabling the model to learn more complex patterns while alleviating the vanishing gradient problem. It performs nonlinear activation on the integrated countermeasure signal, ultimately generating the output of the active countermeasure component.
[0063] An ATT&CK attack vector X encompasses 852 different attack signatures or behavioral patterns. A defense strategy Y is a 230-dimensional vector, and Z is the final output, representing a defensive action that combines defense strategies and countermeasures. This is determined jointly by the defense interpretation component and the active countermeasure component. Each attack signature (a dimension of X) is associated with 230 defense strategies (a dimension of Y) via a corresponding weight in the weight matrix. The mapping from Y to Z occurs in the defense interpretation component, using the defense-countermeasure association matrix α and the parameter β to map defense strategy Y to a more complex defense interpretation output. There are 852 possible mappings from X to Z, and this entire formula maps attack vector X to the final output Z. This process combines the mapping from X to Y and the mapping from Y to Z, as well as the direct mapping of X by the active countermeasure component.
[0064] S3 leverages the reasoning capabilities of a large language inference model and employs the MindTree language model reasoning framework to iteratively update mapping relationships, ensuring continuous adaptation to the ever-changing threat landscape. During this process, defense confidence update formulas are used to optimize and update defense strategies, resulting in more effective new defense strategies and achieving self-evolution.
[0065] S3.1, the detected microgrid attack related data is input into the large language reasoning model for reasoning.
[0066] For example, if a microgrid detects a denial of service (DoS) attack, it will collect a series of attack-related data, such as the attack source IP address, attack time, attack intensity, network traffic, and the status of the attacked device. This data is fed into a large language inference model, which uses natural language processing technology to deeply analyze and understand it, extracting key information and patterns. For example, it may detect unusually frequent connection requests from the attack source IP address or unusual peaks in network traffic, providing a foundation for subsequent defense strategy development.
[0067] S2 inputs the inference results of the large language model into the Thinking Tree reasoning framework built on the M3M model. The Thinking Tree (ToT) is a hierarchical reasoning and decision-making framework built on the MITRE 3D model for microgrid security response. When a microgrid detects an attack, it generates auxiliary response recommendations for manual evaluation and fine-tuning. By learning from the microgrid's historical data and contextual information, it forms inference logic. Its tree-like structure defines fine-tuning strategies for different attack scenarios. Combined with the reasoning capabilities of the large language model, it enables efficient response to complex security threats and improves the accuracy and adaptability of microgrid defense measures.
[0068] In the ToT framework, each node represents a semantic unit, such as a concept, a topic, or a keyword. These nodes are connected by edges, forming a complex association network. When the model receives an input sequence, it searches for the most relevant nodes in ToT based on this sequence and generates output based on these nodes. The quality of each node (partial solution) is evaluated by the value function V(s). For example, for the node state s, its value can be expressed as:
[0069] V(s)=O·P LM (s)+(1-O)·H(s)
[0070] Among them, P LM (s) is the probability of the language model generating state s, H(s) is the heuristic function (such as logical consistency, domain knowledge score); O is the weight coefficient. For the current node s t , the formula for generating candidate branches can be expressed as:
[0071]
[0072] Among them, C(s t ) represents the current node s t The candidate branch set, which is a set containing multiple elements, each element is from s t A possible subsequent state or branch generated, s t Indicates the current node or state, Represents the i-th candidate branch or subsequent state generated from the current node, and the subscript t+1 indicates that this is relative to the current node s t The next state of , and the superscript i is used to distinguish different candidate branches.
[0073] Continuing with the example of a DoS attack, the large language model infers key information such as anomalies in the attack source IP address and network traffic, and then feeds these results into a mind-tree reasoning framework. Each node in this framework represents a semantic unit, such as concepts, topics, or keywords like "attack source," "network traffic," and "microgrid device status." Nodes are connected by edges to form a complex network of associations. After receiving the input sequence, the model searches for the most relevant nodes in the mind-tree. For example, nodes related to "anomaly in attack source IP" might include "malicious IP identification" and "IP blacklist management," while nodes related to "anomaly in network traffic" might include "traffic monitoring" and "bandwidth restriction strategy." Based on these nodes, the model generates outputs, such as recommendations to block the anomalous IP address or adjust network bandwidth allocation. The advantage of the mind-tree reasoning framework lies in its ability to consider multiple thought paths simultaneously. For example, when facing a complex attack, it can consider defense from the perspective of the attack source, as well as from multiple perspectives such as network traffic and device protection. It also dynamically adjusts the generated output strategy based on the different semantic information in the input sequence. For example, as the attack intensity increases, the strength of the defense measures can be adjusted accordingly, making the generated defense strategy more flexible and diverse, adapting to different attack scenarios and defense needs.
[0074] S3.3, based on the inference results and the actual attack and defense situation, adjust and optimize the mapping relationship between attack features and defense actions so that the mapping relationship can adapt to the ever-changing threat environment and improve the flexibility and effectiveness of the defense strategy.
[0075] After implementing actual defense strategies generated by the mind tree reasoning framework, it is discovered that some defense actions may not be effective. For example, after banning abnormal IP addresses, attacks can still continue through other channels, or the bandwidth restriction strategy is too strict, affecting normal business. At this time, the mapping relationship between attack characteristics and defense actions is adjusted and optimized based on the actual defense situation. For example, the monitoring and analysis of other IP addresses in the area where the attack source IP address is located can be increased to determine whether there is a possibility of coordinated attacks, thereby expanding the identification range of attack characteristics; or the bandwidth restriction strategy can be adjusted to set a more reasonable bandwidth threshold to balance the defense effect and normal business needs. This allows the mapping relationship to adapt to the ever-changing threat environment, improve the flexibility and effectiveness of the defense strategy, ensure that defense measures can promptly respond to emerging attack methods, such as more subtle distributed denial of service attacks (DDoS), and maintain the safe operation of the microgrid.
[0076] S3.4, using the defense confidence update formula to optimize defense strategy updates, by analyzing the current attack and defense data to calculate the gradient, according to the defense confidence update, adjust the defense strategy to achieve self-evolution. The defense confidence update formula is as follows:
[0077]
[0078] Among them, β (t) is the parameter value at the tth iteration, which is obtained by the parameter value β of the previous iteration (t-1) Update to get, β (t-1) is the parameter value at the t-1th iteration, and is the parameter value before the update; It is a learning rate-related term, which represents a time decay factor or the total number of time steps. It is used to control the size of the learning rate. As the number of iterations increases, the learning rate gradually decreases, which helps to gradually converge during the optimization process. is the gradient of the objective function with respect to the parameter β, is the objective function, which is the weighted sum of multiple terms, and the gradient Represents the rate of change of the objective function to the parameter β, which is used to guide the direction of parameter update; geo w eight i Represents the weight of the ith one, which is used to weight different contributions in the objective function. In this scheme, if the effect of the defense strategy is good, it is weighted, and vice versa.
[0079] For example, in the process of DoS attack defense, after several strategy adjustments, the defense strategy needs to be further optimized. Assume that the parameter of the current defense strategy is β (t), calculates the gradient by analyzing the current attack and defense data, where the objective function is the weighted sum of multiple items, such as defense effect items, resource consumption items, business impact items, etc. The weight of each item is determined according to the actual situation. If a certain defense action is effective, such as successfully blocking most attack traffic and having little impact on normal business, then increase its corresponding weight, otherwise reduce the weight. Then, according to the defense confidence update formula, it gradually decreases as the number of iterations increases. Controlling the learning rate ensures gradual convergence during the optimization process. This allows the defense strategy to be adjusted based on updated defense confidence, allowing it to continuously evolve based on actual defense effectiveness. For example, parameters such as IP blocking and bandwidth management strategies can be gradually optimized to achieve better defense results, ensuring that the microgrid maintains efficient and reliable defense capabilities in the face of constantly changing attack threats.
[0080] S4, the strategies generated by the expert evaluation thinking tree reasoning framework are used as new defense strategies that can be effectively implemented and achieve the expected defense effects as new defense items.
[0081] During the evaluation process, experts will comprehensively consider multiple key factors, including the practical feasibility of the strategy (i.e., whether it can be implemented under the current microgrid's technical and resource conditions), the potential impact on network performance (such as whether it will reduce network speed or stability), resource consumption (such as whether it will occupy excessive computing or storage resources), and possible side effects (such as whether it will interfere with normal business operations), etc., to ensure that the selected strategy is technically feasible, affordable in terms of resources, and has no significant negative impact on business. For example, when faced with a new malware attack, experts will exclude strategies from the thinking tree reasoning that require special hardware that the microgrid does not have, as well as strategies that may over-consume computing resources and affect the operation of important businesses. They will retain feasible strategies and further evaluate their impact on network performance, ultimately screening out feasible strategies with minimal impact on network performance.
[0082] After completing a comprehensive evaluation of the defense strategy, the experts will only decide to implement the selected defense strategy if they ensure that it can be effectively implemented in the actual operating environment and achieve the desired defensive effect. That is, it can effectively enhance the security of the microgrid and resist the current attack threats, while not adversely affecting the normal operation and business development of the microgrid. This will ensure accurate and effective security protection for the microgrid in the event of an attack. For example, if the experts determine that a certain strategy can effectively prevent an attack and will not interfere with the normal operation of the microgrid after implementation, they will implement the strategy, adjust access control and encrypted communications, successfully prevent the attack, and ensure the safe operation of the microgrid.
[0083] After this new strategy is implemented and successfully defends against the attack, it is added to the historical defense entries as a newly generated historical defense entry. When the microgrid faces a similar attack again in the future, it can directly retrieve this historical defense entry from the defense strategy library, enabling a rapid response and ensuring the safe and stable operation of the microgrid.
[0084] In step S5, the newly generated defense items are studied and summarized throughout the entire defense process using a mind-tree framework. The effectiveness of the defense strategy is analyzed, including whether it successfully prevented the attack, its impact on the microgrid's network performance, and whether it caused new problems. Using a correlation matrix adaptation and conflict detection optimization algorithm, a new defense strategy is learned and summarized to adapt to similar attack patterns, achieving self-evolution.
[0085] S5.1. During the learning process, the correlation matrix adaptive formula is used to adjust the correlation between attack and defense to achieve self-evolution. The correlation matrix adaptive formula is as follows:
[0086]
[0087] Among them, W xy is the correlation matrix, which is used to store the correlation relationship between the input vector X and the output vector Y. t is the time variable, which represents the change of the correlation matrix over time. λ is the urgency driving parameter, which is used to control the learning rate and determines how quickly the system adjusts the correlation matrix in an emergency. A larger λ value means that the system responds faster to emergencies. X is the attack vector and Y is the defense strategy. is the conflict penalty term. When the system detects a conflict, this term will increase, thereby slowing down the update speed of the association matrix.
[0088] This mechanism optimizes the weights of defense strategies within large language models. Based on feedback from defense effectiveness and response time, the model dynamically adjusts defense strategy weights, suppressing conflicts, avoiding resource contention, and improving defense effectiveness. For example, strategies with significant effectiveness in actual defense are weighted higher, while strategies with poor performance or negative impacts are weighted lower. This allows the defense strategy library to continuously evolve, improving its accuracy and efficiency, and continuously optimizing the defense strategy library.
[0089] Assume that the main attack vectors X facing a microgrid include malicious text injection and data exfiltration attempts, while defense strategies Y include text content filtering and encrypted communications. The association matrix stores the relationships between these attack vectors and defense strategies, with its elements representing the weights between the input attack vectors and the output defense strategies. Over time t, the system continuously adjusts the association matrix based on actual conditions. For example, when a new type of malicious text injection attack is detected, the urgency-driven parameter λ increases, accelerating the adjustment of the association matrix and increasing the weight of the defense strategy against this new attack, thereby more effectively responding to emergencies. Furthermore, if the system detects that certain defense strategies conflict with other strategies during execution—for example, if two defense strategies simultaneously process the same text, resulting in a degradation of system performance—the conflict penalty term increases, slowing down the update of the association matrix and preventing further escalation of the conflict.
[0090] In S5.2, when resource contention arises between different defense strategies, conflict penalties are added to suppress such conflicts. Conflict detection optimization uses a binary mask mechanism to encode the strategy range into a form that can be operated on by bits. Bitwise operations in the conflict detection algorithm are then used to quickly determine conflicts between strategies. Finally, the parallel computing capabilities of SIMD instructions are leveraged to significantly improve conflict detection efficiency.
[0091] Conflict detection algorithm, bitwise AND operation to determine mask overlap: In the detect_conflict function, first perform a bitwise AND operation on the masks of the two strategies s1 and s2, that is, s1.mask & s2.mask. If the result is not 0, it means that the ranges of the two strategies overlap and there may be a conflict. Then, perform a bitwise AND operation on the actions of the two strategies, that is, s1.action & s2.action, and determine whether the result is equal to 0b11. If it is equal to 0b11, it means that the actions of the two strategies are 1 at the same time in some bits, that is, there is a conflicting action combination. Only when the above two conditions are met at the same time, the two strategies are considered to be in conflict, and the function returns true, otherwise it returns false.
[0092] Consider multiple defense strategies, such as real-time monitoring of critical equipment and encrypted transmission of power lines. These strategies may compete for limited network bandwidth and computing resources. The conflict detection optimization algorithm uses a binary mask mechanism to encode the policy ranges in a form that can be operated on by bits. For example, the monitoring strategy for all network devices is encoded as 0b00, with a mask of 0b1111; the monitoring strategy for subnet devices is encoded as 0b01, with a mask of 0b0011; and the monitoring strategy for a single critical device is encoded as 0b10, with a mask of 0b0001. In the detect_conflict function, a bitwise AND operation is performed on the masks of two strategies, s1 and s2, such as s1.mask & s2.mask. If the result is not 0, the scopes of the two strategies overlap, indicating a possible resource conflict. The actions of the two strategies are then bitwise ANDed to determine whether there are conflicting action combinations. If there is a conflict, it is suppressed by adding a conflict penalty term, which makes the microgrid more reasonable in resource allocation, ensures that each defense strategy can be executed in a coordinated manner, and improves the overall defense effectiveness and self-evolution.
[0093] S6, adds the new defense items learned and summarized to the historical items of the defense strategy library, so that the strategy library is updated.
[0094] Suppose the Thinking Tree has learned and summarized a new defense strategy for "phishing email attacks", such as "detecting potential malicious attachments through email attachment scanning technology", and integrated this new strategy into the historical entries of the defense strategy library, so that the content of the strategy library is updated and enriched, providing more methods and basis for subsequent defense work.
[0095] Over time, we continue to learn new attack patterns and corresponding defense methods. For example, we've learned defense strategies for various attacks, such as malware dissemination and phishing links. The number of strategies in the defense strategy library has grown from dozens to hundreds or even more. The richness of these strategies has gradually improved the overall defense capabilities of the defense strategy library, enabling it to more effectively defend against various network threats.
[0096] S7: Regularly optimize and organize the policy library to remove outdated or inefficient policies. During the optimization process, it was found that some early defense policies for specific old viruses were no longer applicable as viruses disappeared and new technologies were applied. They took up space and affected retrieval efficiency. Therefore, these policies were identified and removed. At the same time, some inefficient policies were replaced or improved. By removing poor policies, we ensured that the policy library could provide efficient and practical support for network security defense and avoid redundant or outdated content that affects system performance. We identified and removed outdated or inefficient policies, combined dynamic threshold screening with time decay factors to establish a mathematical model. The following formula describes the evolution of the policy retention probability:
[0097]
[0098] Retention conditions: J i (q)≥θ
[0099] Among them, the scoring function J i (q) represents the comprehensive utility score of strategy i at time t, which is composed of the original utility and efficiency index modified by timeliness. The utility parameter F i (q) represents the baseline utility value of strategy i at time q (normalized from 0 to 1), calculated using historical data or real-time feedback; time decay factor Used to control the natural decay of the strategy utility over time, K represents the decay rate parameter (K>0), and q0 represents the initial introduction time of the strategy. Indicates the input-output efficiency of the quantitative strategy, η i (q) represents the profit index of the current time window, Ψ i (q) represents the corresponding resource consumption index, the weight coefficient L, Used to balance the importance of utility continuity and immediate efficiency, satisfying The screening threshold θ is used to dynamically adjust the retention critical value (the recommended initial value is 0.6), which can be automatically adjusted according to the scale of the strategy library:
[0100] θ(q)=μ J (q)-kσ J (q)
[0101] Among them, μ J (q) is the current average score, σ J (q) is the standard deviation, and k is the adjustment coefficient.
[0102] Periodically (e.g., at ΔT intervals), recalculate the scores of all strategies, calculate the mean and standard deviation of the current score distribution, and automatically eliminate inefficient strategies based on the retention criteria:
[0103]
[0104] where ε q represents the strategy set of the qth generation, ε q+1 represents the set of strategies of the q+1th generation that remain after elimination, ε i represents the i-th strategy, J i (q) represents the score of the i-th strategy in the q-th generation; θ(q) represents the elimination threshold, which is used to decide which strategies are retained. i (q) is greater than or equal to the elimination threshold θ(q), the strategy will be retained to the next generation. By comparing the score of each strategy with the elimination threshold, the strategies with lower scores are automatically eliminated.
[0105] Example 2:
[0106] The computer-readable storage medium of this embodiment stores a computer program thereon, which, when executed by a processor, implements the steps of the microgrid security response method of M3M-ToT self-evolutionary reasoning in embodiment 1.
[0107] The computer-readable storage medium of this embodiment may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal; the computer-readable storage medium of this embodiment may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card, a secure digital card, a flash memory card, etc. equipped on the terminal; further, the computer-readable storage medium may also include both an internal storage unit of the terminal and an external storage device.
[0108] The computer-readable storage medium of this embodiment is used to store computer programs and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.
[0109] Example 3:
[0110] The computer device of this embodiment includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the microgrid security response method of M3M-ToT self-evolutionary reasoning in Example 1 are implemented.
[0111] In this embodiment, the processor can be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The memory can include read-only memory and random access memory, and provide instructions and data to the processor. A part of the memory can also include non-volatile random access memory. For example, the memory can also store information about the device type.
[0112] Those skilled in the art will clearly understand that each embodiment can be implemented by means of software plus a necessary general-purpose hardware platform, or of course, by means of hardware. Based on this understanding, the essence of the above technical solution or the portion that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0113] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A microgrid security response method based on M3M-ToT self-evolutionary reasoning, characterized by: The following steps are involved: S1, when the microgrid is attacked, through network situation awareness, the defense strategy library is searched to identify whether there is a defense strategy that matches the current attack pattern. If a matching attack pattern is retrieved, the corresponding defense strategy is executed after expert decision-making; if no matching attack pattern is retrieved, a candidate strategy set is generated and entered into the subsequent process; S2: Collect and integrate mapping relationships in the MITRE Shield, ATT&CK, and Engage frameworks, build an "attack-defense-counter" closed-loop knowledge graph, and establish a three-dimensional mapping equation for the M3M model. S3, inputs attack-related data into the large language reasoning model to extract key information, inputs the key information into the thinking tree TOT reasoning framework built based on the M3M model, and generates defense strategies by reasoning; S4, the strategies generated by the expert evaluation thinking tree reasoning framework are selected as new defense strategies that can be effectively implemented and achieve the expected defense effect as new defense items; S5: Use the mind tree framework to study and summarize the entire defense process for the newly generated defense items, and use the correlation matrix adaptation and conflict detection optimization algorithm to analyze the actual effect of the defense strategy; S6, adds the new defense items learned and summarized to the historical items of the defense strategy library, so that the strategy library is updated.
2. The microgrid security response method based on M3M-ToT self-evolutionary reasoning according to claim 1 is characterized in that: The multi-dimensional information collected during the S1 search phase includes: timestamp, event type, event source IP address, event source MAC address, asset number, unique identifier of the response policy, indicative threat information, network traffic, audit logs, policy change records, event urgency, device type, problem type, device status, geographic location code, associated threat pattern ID, trigger condition, response action, causal chain identifier, policy scope, and machine learning model ID.
3. The microgrid security response method based on M3M-ToT self-evolutionary reasoning according to claim 1 is characterized in that: The three-dimensional mapping equation of the M3M model is as follows: in, is the ATT&CK attack vector, Y = σ(W xy X) is the Shield defense strategy, and Z is the final output; Defense explanation part is the defense-countermeasure correlation matrix, β is a parameter, represents tensor product, ⊙ represents element-wise multiplication; Active counter part W xz X is a linear transformation where W xy Is the correlation matrix, which linearly combines the input attack vector X; γX °2 is a nonlinear adversarial term, where γ is the coefficient, X °2 It represents a square operation on each element of the attack vector X, and ReLU is a nonlinear activation function.
4. The microgrid security response method based on M3M-ToT self-evolutionary reasoning according to claim 3 is characterized in that: In the TOT reasoning framework, each node represents a semantic unit. Nodes are connected by edges to form an association network. When an input sequence is received, the most relevant nodes are found in the TOT based on the sequence, and output is generated based on these nodes. The quality of each node is evaluated by the value function V(s). For a node state s, its value can be expressed as: V(s)=O·P LM (s)+(1-O)·H(s) Among them, P LM (s) is the probability of the language model generating state s, H(s) is the heuristic function; O is the weight coefficient, for the current node s t , the formula for generating candidate branches is expressed as: Among them, C(s t ) represents the current node s t A set of candidate branches, each element in the set is from s t A possible subsequent state or branch generated, s t Indicates the current node or state, Represents the i-th candidate branch or subsequent state generated from the current node, and the subscript t+1 indicates that this is relative to the current node s t The next state of , and the superscript i is used to distinguish different candidate branches.
5. The microgrid security response method based on M3M-ToT self-evolutionary reasoning according to claim 4 is characterized in that: The MindTree reasoning framework iterates and optimizes the mapping relationship between attack features and defense actions using the defense confidence update formula based on the reasoning results and the actual attack and defense situation. The defense confidence update formula is as follows: Among them, β (t) is the parameter value at the tth iteration, which is obtained by the parameter value β of the previous iteration (t-1) Update to get, β (t-1) is the parameter value at the t-1th iteration, and is the parameter value before the update; It is a learning rate-related term, which represents a time decay factor or the total number of time steps. It is used to control the size of the learning rate. As the number of iterations increases, the learning rate gradually decreases, which helps to gradually converge during the optimization process. is the gradient of the objective function with respect to the parameter β, is the objective function, which is the weighted sum of multiple terms, and the gradient Represents the rate of change of the objective function to the parameter β, which is used to guide the direction of parameter update; geo w eight i Represents the weight of the i-th one, which is used to weight different contributions in the objective function.
6. The microgrid security response method based on M3M-ToT self-evolutionary reasoning according to claim 1 is characterized in that: The adaptive formula of the correlation matrix is as follows: Among them, W xy is the correlation matrix, which is used to store the correlation relationship between the input vector X and the output vector Y. t is the time variable, which represents the change of the correlation matrix over time. λ is the urgency driving parameter, which is used to control the learning rate and determines the speed at which the system adjusts the correlation matrix in an emergency. A larger λ value means that the system responds faster to emergencies. X is the attack vector, and Y is the defense strategy. is the conflict penalty term. When the system detects a conflict, this term will increase, thereby slowing down the update speed of the association matrix.
7. The microgrid security response method based on M3M-ToT self-evolutionary reasoning according to claim 6 is characterized in that: The conflict detection algorithm is specifically as follows: in the detect_conflict function, first, the masks of the two strategies s1 and s2 are bitwise ANDed, that is, s1.mask & s2.mask. If the result is not 0, it means that the ranges of the two strategies overlap and there may be a conflict; then, the actions of the two strategies are bitwise ANDed, s1.action & s2.action, and it is determined whether the result is equal to 0b11. If it is equal to 0b11, it means that the actions of the two strategies are 1 at the same time in some bits, and there is a conflicting action combination; only when the above two conditions are met at the same time, it is determined that the two strategies conflict, and the function returns true, otherwise it returns false.
8. The microgrid security response method based on M3M-ToT self-evolutionary reasoning according to claim 1 is characterized in that: The following steps are also included: S7, regularly optimize and organize the strategy library, remove outdated or inefficient strategies, and establish a mathematical model combining dynamic threshold screening and time decay factors: Retention conditions: J i (q)≥θ Among them, the scoring function J i (q) represents the comprehensive utility score of strategy i at time t, which is composed of the original utility and efficiency index modified by timeliness. The utility parameter F i (q) represents the baseline utility value of strategy i at time q, calculated through historical data or real-time feedback; time decay factor Used to control the natural decay of the strategy utility over time, K represents the decay rate parameter (K>0), q0 represents the initial introduction time of the strategy; efficiency ratio Indicates the input-output efficiency of the quantitative strategy, η i (q) represents the profit index of the current time window, Ψ i (q) represents the corresponding resource consumption index, the weight coefficient L, Used to balance the importance of utility continuity and immediate efficiency, satisfying The screening threshold θ is used to dynamically adjust the retention critical value, which is automatically adjusted according to the scale of the strategy library: θ(q)=μ J (q)-kσ J (q) Among them, μ J (q) is the current average score, σ J (q) is the standard deviation, k is the adjustment coefficient; Regularly recalculate the scores of all strategies, calculate the mean and standard deviation of the current score distribution, and automatically eliminate inefficient strategies based on the retention criteria: where ε q represents the strategy set of the qth generation, ε q+1 represents the set of strategies of the q+1th generation that remain after elimination, ε i represents the i-th strategy, J i (q) represents the score of the i-th strategy in the q-th generation; θ(q) represents the elimination threshold, and only when the score of the strategy J i (q) is greater than or equal to the elimination threshold θ(q), the strategy will be retained to the next generation. By comparing the score of each strategy with the elimination threshold, the strategies with lower scores are automatically eliminated.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the microgrid security response method of M3M-ToT self-evolutionary reasoning as described in any one of claims 1 to 8 are implemented.
10. A computer device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps in the microgrid security response method of M3M-ToT self-evolutionary reasoning according to any one of claims 1-8 are implemented.
Citation Information
Cited By
Terminal equipment control method and system based on Internet of Things
CN120710796A
Combination optimization solving system and method based on general attack and defense framework
CN120975358A