Multi-agent based distributed coordination method and system for power system

By adopting a distributed collaborative approach based on multi-agent systems, the problems of data processing latency and low reliability of centralized control in large power systems are solved, and efficient collaborative operation of power nodes and accurate power allocation are achieved, thereby improving the overall performance and stability of the system.

CN122418897BActive Publication Date: 2026-08-25CHENGDU BIKONG SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610857349.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-08-25
Estimated Expiration
2046-06-15

AI Technical Summary

Technical Problem

Traditional centralized control methods face problems such as data processing delays, low reliability, and insufficient coordination among multiple power supply nodes in large power systems, making it difficult to achieve efficient and stable operation.

Method used

A distributed collaborative approach based on multi-agents is adopted. By acquiring global situational information, a distributed collaborative decision-making model is used to identify the negotiation intentions and commitments of multi-agents, generate a multi-cluster task decomposition structure, and realize the collaborative grouping of power nodes and the allocation of power output tasks.

Benefits of technology

It improves the overall performance and stability of the power system, reduces system operating costs, and ensures the consistency of goals and the rationality of task allocation among the intelligent agents in the collaborative process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122418897B_ABST
    Figure CN122418897B_ABST
Patent Text Reader

Abstract

The application provides a kind of power supply system distributed collaborative method and system based on multi-agent, it is related to power supply system control technical field, first acquisition power supply system current running time global situation information set, including node state parameters and power transmission topological connection data, then it is input distributed collaborative decision model and is identified to multi-agent negotiation intention, obtain multi-agent intention distribution set, then according to intention characteristic vector similarity cross-agent negotiation commitment matching is carried out, and the multi-agent commitment protocol set is generated, then the power supply node carrying the same commitment protocol identification is subjected to task collaborative grouping, and the multi-cluster task decomposition structure is obtained, finally based on commitment protocol identification, cluster level resource scheduler is called to carry out power output task allocation, and the distributed collaborative scheduling instruction set is generated.The application improves the power supply system collaborative efficiency and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system control technology, and more specifically, to a distributed collaborative method and system for power systems based on multiple agents. Background Technology

[0002] In the field of power system operation and management, the traditional centralized control method has always been the mainstream approach. Centralized control relies on a central controller to collect all information from the entire power system, including the status parameters of each power node and the power transmission topology connection data between them, and then the central controller performs unified decision-making and scheduling. However, as the scale of power systems continues to expand, the number of power nodes increases, and the structure becomes increasingly complex, centralized control methods face many challenges.

[0003] On the one hand, the central controller needs to process massive amounts of data, which places extremely high demands on its computing power and communication bandwidth, easily leading to data processing delays and an inability to respond promptly to dynamic changes in the system. On the other hand, centralized control systems have low reliability; once the central controller fails, the entire power system may be paralyzed, severely affecting the stable operation of the system. Furthermore, while existing distributed control methods have solved some of the problems of centralized control to a certain extent, they still have shortcomings in multi-power node coordination, lacking effective agent negotiation mechanisms and task coordination grouping strategies, making it difficult to achieve efficient and stable operation of the power system. Summary of the Invention

[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a distributed cooperative method for a power system based on multiple agents, the method comprising: Obtain a global situation information set of the power system at the current operating moment. The global situation information set includes node status parameters uploaded by each power node and power transmission topology connection data between each power node. The global situation information set is input into a preset distributed collaborative decision-making model for multi-agent negotiation intent recognition. Each agent in the distributed collaborative decision-making model generates a power adjustment negotiation intent for adjacent power nodes based on its corresponding node state parameters and the power transmission topology connection data, resulting in a multi-agent intent distribution set containing the intent feature vectors output by each agent. Based on the vector space similarity calculation results between each intention feature vector in the multi-agent intention distribution set, cross-agent negotiation commitment matching is performed, and a corresponding target commitment protocol identifier is assigned to each intention feature vector to generate a multi-agent commitment protocol set containing the correspondence between commitment protocol identifiers and intention feature vectors. Task collaborative grouping is performed on the power nodes corresponding to the intent feature vectors carrying the same commitment protocol identifier in the multi-agent commitment protocol set. Power nodes with the same commitment protocol identifier are divided into the same collaborative task cluster, resulting in a multi-cluster task decomposition structure that includes the correspondence between collaborative task cluster identifier and power node. Based on the commitment protocol identifier corresponding to each collaborative task cluster in the multi-cluster task decomposition structure, the cluster-level resource scheduler in the distributed collaborative decision-making model is invoked to allocate power output tasks to the power nodes within the collaborative task cluster, generating a set of distributed collaborative scheduling instructions containing the target power value and power output timing of each power node.

[0005] Furthermore, embodiments of the present invention also provide a distributed cooperative system for power systems based on multi-agent systems, comprising: A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the above-described distributed cooperative method for a multi-agent power system by executing the machine-executable instructions.

[0006] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions stored in a computer-readable storage medium, a processor of a multi-agent-based distributed power system cooperative system reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the multi-agent-based distributed power system cooperative system to execute the above-described multi-agent-based distributed power system cooperative method.

[0007] Based on the above, by acquiring the global situational information set of the power system at the current operating moment, this global situational information set is input into a preset distributed collaborative decision-making model for multi-agent negotiation intent recognition. This enables each agent to generate power adjustment negotiation intents for adjacent power nodes based on its corresponding node state parameters and power transmission topology connection data, forming a multi-agent intent distribution set containing the intent feature vectors of each agent. This achieves preliminary information interaction and intent expression among the multiple agents. Cross-agent negotiation commitment matching is performed based on the vector space similarity calculation results between the intent feature vectors in the multi-agent intent distribution set. A corresponding target commitment protocol identifier is assigned to each intent feature vector, generating a multi-agent commitment protocol set. This effectively coordinates the actions among the multiple agents and ensures the consistency of goals among the agents during the collaborative process. The power nodes corresponding to intent feature vectors carrying the same commitment protocol identifier in the multi-agent commitment protocol set are grouped into task collaborative groups, resulting in a multi-cluster task decomposition structure. Power nodes with similar intents and commitments are grouped into the same collaborative task cluster, improving the rationality of task allocation and collaborative efficiency. Based on the commitment protocol identifier corresponding to each collaborative task cluster in the multi-cluster task decomposition structure, the cluster-level resource scheduler in the distributed collaborative decision-making model is invoked to allocate power output tasks to the power nodes within the collaborative task cluster, generating a set of distributed collaborative scheduling instructions. This achieves accurate power allocation and orderly collaborative operation of each power node within the power system, effectively improving the overall performance and stability of the power system and reducing system operating costs. Attached Figure Description

[0008] Figure 1 This is a schematic diagram of the execution flow of the distributed collaborative method for power systems based on multi-agent systems provided in an embodiment of the present invention.

[0009] Figure 2 This is a schematic diagram of exemplary hardware and software components of a distributed collaborative power system based on multi-agent systems provided in an embodiment of the present invention. Detailed Implementation

[0010] Figure 1 This is a flowchart illustrating a distributed collaborative method for a power system based on multiple agents, provided in one embodiment of the present invention. A detailed description follows.

[0011] Step S110: Obtain the global situation information set of the power system at the current operating moment. The global situation information set includes the node status parameters uploaded by each power node and the power transmission topology connection data between each power node.

[0012] In this embodiment, a large-scale regional data center power supply system is taken as an example. This large-scale regional data center power supply system consists of multiple distributed power nodes, such as mains power access nodes, diesel generator set nodes, energy storage battery array nodes, and power conversion nodes that supply power to IT loads in different areas.

[0013] First, at the start of each preset control cycle (e.g., every 100 milliseconds), a global situational awareness data collection and aggregation operation is performed. In this embodiment, a situational awareness reporting request is simultaneously broadcast to all registered power nodes via a network communication interface. Upon receiving the request, each power node packages the real-time operating data monitored by its local controller into a data packet of a preset format and reports it. The reported data packet first passes through a pre-processing data cleaning and format unification module, which is responsible for removing obviously abnormal noise data and converting heterogeneous data from different manufacturers and models of equipment into an internally unified data structure.

[0014] Specifically, for mains-connected nodes, the reported node status parameters include the instantaneous output power value, denoted as P_U_instant (in kilowatts); the instantaneous input power value, which is usually associated with the output power since mains power is the energy input, but can also be monitored independently, denoted as P_U_in; the state of charge percentage of the internal energy storage unit, which can be considered inapplicable or constant at 100% for mains-connected nodes; and the real-time operating efficiency percentage of the power conversion module, denoted as η_U. For energy storage battery array nodes, the reported node status parameters include the instantaneous output power value P_B_instant (positive for discharging), the instantaneous input power value P_B_in (positive for charging), the state of charge percentage SOC_B of the internal energy storage unit, and the operating efficiency percentage η_B of the bidirectional power conversion module (charging and discharging efficiencies may differ, but the model will select the corresponding efficiency value based on the power flow direction).

[0015] Simultaneously, the physical connection relationships and real-time power flow data between each power node are obtained from network switches or smart grid gateways, i.e., power transmission topology connection data. This data exists in the form of an adjacency list. For example, for energy storage node B, its topology connection data includes a list of adjacent node identifiers, such as [mains node U, IT load node L1], and a corresponding list of power transmission line parameters, including the transmission line resistance value R_BU between it and mains node U (in ohms); the line reactance value X_BU (in ohms); and the real-time transmission power value P_trans_UB currently flowing from node U to node B on this line (in kilowatts). Finally, all the collected node status parameters and topology connection data are encapsulated according to a globally unique timestamp (e.g., a Unix timestamp, denoted as T_current) to form a global situational information set S_global.

[0016] Step S120: Input the global situation information set into a preset distributed collaborative decision-making model to identify the multi-agent negotiation intention. Each agent in the distributed collaborative decision-making model generates a power adjustment negotiation intention for adjacent power nodes based on its corresponding node state parameters and the power transmission topology connection data, thereby obtaining a multi-agent intention distribution set containing the intention feature vectors output by each agent.

[0017] Next, the global situational information set S_global generated in step S110 above is used as input and passed to a pre-trained distributed cooperative decision-making model. This distributed cooperative decision-making model is a deep reinforcement learning framework based on a multi-agent system. After receiving S_global, the distributed cooperative decision-making model can distribute it to agents instantiated for each power node. Each agent is an independent computational entity running the same neural network structure, but its parameters may be updated independently.

[0018] Step S121: The node status parameters uploaded by each power node in the global situation information set are classified by data index according to the power node identifier, and a single node status parameter record uniquely corresponding to each power node is generated. The single node status parameter record includes the instantaneous output power value, instantaneous input power value, state of charge percentage of the internal energy storage unit, and real-time working efficiency percentage of the power conversion module at the current operating moment of the power node.

[0019] Specifically, the distributed collaborative decision-making model first performs data distribution and indexing operations. A data distributor in the model constructs a single-node state parameter record R_state_B for each power node, such as the energy storage node identified as ID_B, based on the data in S_global. This single-node state parameter record is a structure whose internal fields are precisely filled: the instantaneous output power value field corresponding to ID_B in the global information is assigned to the P_out field of R_state_B; the instantaneous input power value field is assigned to the P_in field; the state of charge percentage field is assigned to the SOC field; and the operating efficiency percentage field is assigned to the η field. This generates an independent data record for each node containing all its own state information.

[0020] Step S122: Input the power transmission topology connection data between each power node in the global situation information set into the topology relationship parsing module of the distributed collaborative decision-making model to deconstruct the connection relationship, generate a list of adjacent node identifiers with each power node as the center node and power transmission line parameters with each power node as the center node. The power transmission line parameters include the transmission line resistance value, transmission line reactance value and real-time transmission power value on the current transmission line between the power node and each adjacent node.

[0021] Meanwhile, the power transmission topology connection data in the global situational information set S_global is fed into the model's built-in topology relation parsing module. This module traverses the entire system's topology graph, generating a local topology view centered on each node. For example, for the energy storage node ID_B, the module parses the adjacency list in S_global, finds all nodes directly connected to ID_B, and generates a list of neighboring node identifiers L_neighbor_B, for example, [ID_U, ID_L1]. Then, for each neighboring node in the list, the module extracts the corresponding line parameters from S_global, generating a parallel power transmission line parameter list L_line_B, where each element is a triplet. For the line with ID_U, the triplet is (R_BU, X_BU, P_trans_UB); for the line with ID_L1, the triplet is (R_BL1, X_BL1, P_trans_BL1).

[0022] Step S123: Deploy a corresponding agent instance for each power node. Simultaneously input the individual node state parameter record, the neighbor node identifier list centered on the power node, and the power transmission line parameters centered on the power node into the intent encoder module of the agent instance corresponding to the power node for feature encoding, and generate an original state feature vector containing the power node's own state information and the connection relationship information with neighboring nodes.

[0023] In this embodiment, a corresponding agent instance is then activated for each power node, for example, Agent_B is activated for the energy storage node ID_B. The intent encoder module of Agent_B is a multi-input, single-hidden-layer fully connected neural network. This intent encoder module receives three inputs: the individual node state parameter record R_state_B, whose values ​​are concatenated into a one-dimensional vector V_raw_B, for example, with four dimensions, corresponding to P_out, P_in, SOC, and η respectively; the neighbor node identifier list L_neighbor_B; and the power transmission line parameter list L_line_B. The intent encoder first flattens the values ​​(R, X, P_trans) in L_line_B and concatenates them with V_raw_B to form a new vector V_combined_B. Then, it passes through a linear transformation layer: V_encoded_B = f(W_enc·V_combined_B + b_enc), where f is a non-linear activation function, for example, using a rectified linear unit. W_enc is the encoder's weight matrix, and b_enc is the bias vector. This transformation generates a higher-dimensional original state feature vector F_raw_B that contains the node's own state and local topology information. This original state feature vector captures the comprehensive characteristics of the node's current operating condition.

[0024] Step S124: Input the original state feature vector generated by the intent encoder module of the agent instance corresponding to each power node into the attention calculation module of the neighboring nodes of the agent instance for attention weight allocation. Based on the feature similarity calculation results between the original state feature vector of the power node and the original state feature vectors of each neighboring node of the power node, generate the attention weight coefficient for each neighboring node, and obtain the inter-node attention distribution matrix containing the attention weight coefficients of the power node and all neighboring nodes.

[0025] After generating the raw state feature vector for each node, such as F_raw_B for Agent_B, these vectors are fed into a neighbor node attention calculation module shared by all agents. This module aims to allow each agent to assess the importance of negotiating with its neighbors.

[0026] Step S1241: Simultaneously send the original state feature vector generated by the intent encoder module of the agent instance corresponding to each power node to the neighboring node attention calculation module of the agent instance and the neighboring node attention calculation module of the agent instance corresponding to each neighboring node of the power node, so that the neighboring node attention calculation module of each agent instance simultaneously receives the original state feature vector of its own power node and the original state feature vectors of all neighboring nodes.

[0027] To perform attention computation, a vector broadcast is executed. The attention computation module of Agent_B's neighboring nodes not only receives its own F_raw_B, but also receives the original state feature vectors of all its neighboring nodes (such as Agent_U and Agent_L1) through the network, namely F_raw_U and F_raw_L1. In this way, Agent_B's local module has its own feature vectors and those of all its neighbors.

[0028] Step S1242: In the neighboring node attention calculation module of each agent instance, the original state feature vector of the power node where it is located is multiplied by the original state feature vector of each neighboring node to obtain the initial attention score between the node where it is located and each neighboring node. The initial attention score is used to characterize the degree of mutual attention between the two power nodes in the current operating state for power regulation negotiation.

[0029] Within the attention calculation module of Agent_B, dot product attention calculation is performed. First, the attention score with neighboring node U is calculated: e_BU = F_raw_B · F_raw_U, which is the sum of the corresponding element-wise multiplications of the two vectors, resulting in a scalar value e_BU. The larger this scalar value, the more relevant node B's current state is to node U's state, and the more necessary negotiation is. Similarly, the score with neighboring node L1 is calculated: e_BL1 = F_raw_B · F_raw_L1.

[0030] Step S1243: In the neighboring node attention calculation module of each agent instance, the initial attention score between the node itself and all neighboring nodes is numerically normalized. By dividing each initial attention score by the sum of all initial attention scores, a normalized attention weight coefficient corresponding to each initial attention score is generated. The magnitude of the normalized attention weight coefficient is positively correlated with the magnitude of the initial attention score.

[0031] To ensure the attention weights are comparable and summed to one, the module normalizes the calculated scores. The sum of all neighbor scores is calculated: sum_e_B = e_BU + e_BL1. Then, the normalized attention weight for node U is calculated: α_BU = e_BU / sum_e_B. Similarly, the weight for node L1 is calculated: α_BL1 = e_BL1 / sum_e_B. These two coefficients, α_BU and α_BL1, represent how much attention (i.e., negotiation weight) Agent_B should allocate to its neighbors U and L1 in this decision-making process.

[0032] Step S1244: In the neighboring node attention calculation module of each agent instance, the normalized attention weight coefficients corresponding to the node where it is located and each neighboring node are arranged in the order of the neighboring node's identifier to construct a one-dimensional row vector. Each element position of this row vector corresponds one-to-one with the identifier of each neighboring node, and the value of each element position is the normalized attention weight coefficient corresponding to that neighboring node.

[0033] Next, Agent_B arranges the calculated normalized weight coefficients according to the order of its neighbor node identifier list L_neighbor_B (e.g., U first, then L1), forming a row vector V_att_B=[α_BU, α_BL1]. This vector is the inter-node attention distribution vector of node B.

[0034] Step S1245: Output the row vector containing normalized attention weight coefficients constructed by the neighboring node attention calculation module of each agent instance as the node attention distribution vector corresponding to the power node. Stack all the node attention distribution vectors corresponding to the power nodes vertically according to the order of the power node identifiers to form a two-dimensional node attention distribution matrix. The row index of the node attention distribution matrix corresponds to the power node identifier that initiated the attention calculation, and the column index of the node attention distribution matrix corresponds to the neighboring node identifier that is being paid attention to. The value of each element in the node attention distribution matrix is ​​the normalized attention weight coefficient from the power node corresponding to the row index to the neighboring node corresponding to the column index.

[0035] In this embodiment, the attention distribution vectors generated by all agents are aggregated. For example, the attention distribution vector of the mains node U is V_att_U=[α_UL1, α_UB] (assuming its neighbors are L1 and B), and the attention distribution vector of the load node L1 is V_att_L1=[α_L1U, α_L1B] (assuming its neighbors are U and B). In this embodiment, the above vectors are stacked vertically in the order of node identification (e.g., B, U, L1) to form a three-row, two-column matrix M_attention. The rows of the matrix represent source nodes, and the columns represent target nodes (i.e., neighbors). The elements in M_attention, such as the first column of the second row, represent the attention weights from node U to its first neighbor (which may be L1).

[0036] Step S125: Input the original state feature vector of the agent instance corresponding to each power node and the inter-node attention distribution matrix of the agent instance into the intent generation module of the agent instance for intent feature mapping. The input original state feature vector is nonlinearly transformed by the fully connected network layer in the intent generation module to generate a preliminary intent feature vector.

[0037] Returning to Agent_B, its intent generation module receives two inputs: one is the original state feature vector F_raw_B generated in step S123, and the other is the row vector corresponding to Agent_B extracted from the global attention matrix M_attention, namely V_att_B. The intent generation module first processes F_raw_B through a fully connected network layer. This fully connected network layer performs a linear transformation: V_pre_intent_B = W_intent1·F_raw_B + b_intent1, followed by a non-linear activation function, such as the hyperbolic tangent function, to generate a preliminary intent feature vector P_intent_B. This vector P_intent_B contains the node's original desire to adjust based on its own state, but it has not yet considered the negotiation relationship with its neighbors.

[0038] Step S126: The intention weight fusion layer in the intention generation module multiplies each attention weight coefficient in the inter-node attention distribution matrix with the corresponding feature dimension in the preliminary intention feature vector element by element to generate an intention feature vector enhanced by the attention mechanism.

[0039] Next, in the intent-weighted fusion layer, Agent_B fuses its attention distribution vector V_att_B with the initial intent feature vector P_intent_B. Specifically, since P_intent_B is a multi-dimensional vector, while V_att_B typically has a smaller dimension (equal to the number of neighbors), they cannot be directly multiplied element-wise. Therefore, this module first maps V_att_B through a linear mapping layer to the same dimension as P_intent_B, obtaining the expanded attention vector V_att_expanded_B. Then, an element-wise multiplication (Hadamard product) operation is performed: F_intent_B = P_intent_B ⊙ V_att_expanded_B. Here, ⊙ represents element-wise multiplication at corresponding positions. This operation means that, based on the degree of attention to each neighbor, different feature dimensions of the initial intent are enhanced or suppressed. For example, if Agent_B is highly concerned about node U (α_BU is large), then the dimension values ​​related to power upward adjustment in the mapped V_att_expanded_B will be larger, thus highlighting the intention of upward adjustment in F_intent_B. The final F_intent_B is the intention feature vector output by Agent_B, which contains the negotiation relationship.

[0040] Step S127: Collect and aggregate the intent feature vectors output by the intent generation module of the agent instance corresponding to each power node according to the power node identifier. Combine the intent feature vectors corresponding to all power nodes into a multi-dimensional vector set, and output the multi-dimensional vector set as the multi-agent intent distribution set. Each intent feature vector in the multi-agent intent distribution set carries its corresponding power node identifier and the power adjustment target direction identifier to which the intent feature vector is targeted. The power adjustment target direction identifier is used to indicate whether the power adjustment expectation expressed by the intent feature vector is to increase or decrease the output power.

[0041] After all agents have completed intent generation, the intent aggregation phase begins.

[0042] Step S1271: After the intent generation module of the agent instance corresponding to each power node completes the intention feature vector generation operation, the intention feature vector, its corresponding power node identifier, and the current power adjustment target direction identifier of the power node are data encapsulated to generate an intent data encapsulation unit containing vector data and identifier data. The vector data stores all the numerical elements of the intention feature vector, and the identifier data stores the power node identifier and the power adjustment target direction identifier.

[0043] Each agent, such as Agent_B, after generating the intent feature vector F_intent_B, will determine the main adjustment direction based on its internal policy (e.g., based on its SOC value). If the SOC is too high, the intent may be to reduce charging or increase discharging, with the direction labeled "increase output power"; if the SOC is too low, the intent may be to increase charging or reduce discharging, with the direction labeled "decrease output power". Agent_B encapsulates F_intent_B, its node identifier ID_B, and the direction label (e.g., using the enumeration values ​​DIR_INC to represent increase and DIR_DEC to represent decrease) into a data unit U_intent_B.

[0044] Step S1272: Send the intent data encapsulation units corresponding to all power nodes to the preset central intent collector for data aggregation. The central intent collector temporarily stores the received intent data encapsulation units in a temporary data buffer according to the time sequence.

[0045] These encapsulation units are asynchronously sent to a central intent collector service. The central intent collector service temporarily stores all received units, such as U_intent_B, U_intent_U, and U_intent_L1, in a high-speed memory cache according to the order of receipt, awaiting further processing.

[0046] Step S1273: After the central intent collector completes the receiving operation of the intent data encapsulation units corresponding to all power nodes, the intent data aggregation process is started. Each intent data encapsulation unit is read sequentially from the temporary data buffer, and the power node identifier and power adjustment target direction identifier are extracted from the identifier data of each intent data encapsulation unit.

[0047] After confirming that it has received units from all online nodes (by checking against a pre-defined node list), the central intent collector begins processing the data in the buffer. It reads each unit one by one, first parsing out the identification data within.

[0048] Step S1274: Based on the extracted power node identifier, perform the first-level classification of the read intent data encapsulation units, and group the intent data encapsulation units with the same power node identifier into the intent data subset corresponding to the same power node. The intent data subset corresponding to each power node contains all intent feature vectors generated by that power node in this negotiation intent recognition process.

[0049] After resolving the identifiers, the collector groups them according to the node IDs. For example, all units identified as ID_B are grouped into a list called Set_B. Since each node may generate multiple intents per cycle (e.g., for negotiation), Set_B may contain multiple F_intent_B vectors.

[0050] Step S1275: After completing the first-level classification, perform a second-level classification on the intent data encapsulation units within the intent data subset corresponding to each power node. Based on the power adjustment target direction identifier in the identifier data of each intent data encapsulation unit, further divide the intent data encapsulation units within the intent data subset into groups for increasing power intent data and groups for decreasing power intent data.

[0051] Next, within each node's data subset, the collector performs a secondary partition based on the direction identifier. For example, in Set_B, all cells with the direction identifier DIR_INC are placed into the subgroup Set_B_inc, and all cells with the direction identifier DIR_DEC are placed into the subgroup Set_B_dec.

[0052] Step S1276: Select an intention feature vector from the power increase intention data group corresponding to each power node as the representative power increase intention feature vector of that power node, and select an intention feature vector from the power decrease intention data group corresponding to each power node as the representative power decrease intention feature vector of that power node. Combine the representative power increase intention feature vector and the representative power decrease intention feature vector of each power node to form the final intention feature vector pair of that power node.

[0053] To simplify the subsequent matching process, each node retains only the most representative increase / decrease intent. The central intent collector, in Set_B_inc, selects the highest-scoring vector as the representative increase intent feature vector F_inc_B for node B by calculating the magnitude of each vector or through a pre-trained "intent representativeness evaluation network." Similarly, a representative decrease intent feature vector F_dec_B is selected from Set_B_dec. The final intent of node B is a vector pair (F_inc_B, F_dec_B).

[0054] Step S1277: Arrange the final intent feature vector pairs of all power nodes in the order of power node identifiers to form a data structure with the power node identifier as the index and the final intent feature vector pair of each power node as the index value. Output this data structure as the multi-agent intent distribution set. In the multi-agent intent distribution set, the final intent feature vector pair corresponding to each power node carries an increase output power identifier and the final intent feature vector carries a decrease output power identifier.

[0055] Finally, the central intent collector organizes the representative intent pairs from all nodes into an ordered mapping structure, Map_intent, according to node ID order. For example, Map_intent[ID_B] = (F_inc_B, F_dec_B). This Map_intent is the final multi-agent intent distribution set, which will serve as the input for the next step, where each intent vector clearly indicates its desired power adjustment direction.

[0056] Step S130: Perform cross-agent negotiation commitment matching based on the vector space similarity calculation results between each intent feature vector in the multi-agent intent distribution set, assign a corresponding target commitment protocol identifier to each intent feature vector, and generate a multi-agent commitment protocol set containing the correspondence between commitment protocol identifiers and intent feature vectors.

[0057] After obtaining the multi-agent intent distribution set Map_intent, cross-agent commitment matching begins, with the aim of pairing nodes that want to increase power with nodes that want to decrease power to form a preliminary agreement.

[0058] Step S131: Extract all representative increase intention feature vectors carrying the increase output power identifier from the multi-agent intention distribution set to form an increase intention feature vector set, and extract all representative decrease intention feature vectors carrying the decrease output power identifier from the multi-agent intention distribution set to form a decrease intention feature vector set.

[0059] In this embodiment, Map_intent is first traversed, and all feature vectors representing increasing intent, such as F_inc_B and F_inc_U, are extracted and placed into a set V_set_inc. At the same time, all feature vectors representing decreasing intent, such as F_dec_B and F_dec_L1, are extracted and placed into another set V_set_dec.

[0060] Step S132: Perform pairwise vector space similarity calculation on each representative increasing intention feature vector in the set of increasing intention feature vectors and each representative decreasing intention feature vector in the set of decreasing intention feature vectors. By calculating the cosine similarity value between each pair of representative increasing intention feature vectors and representative decreasing intention feature vectors, generate an initial intention matching similarity matrix containing all possible pairings.

[0061] Next, a pairing calculator is launched. This calculator calculates the cosine similarity for each vector in V_set_inc (e.g., F_inc_B) and each vector in V_set_dec (e.g., F_dec_L1). For vectors F_inc_B and F_dec_L1, the cosine similarity sim_incB_decL1 is calculated by taking the dot product of the two vectors and then dividing by the product of their magnitudes. That is, sim = (F_inc_B·F_dec_L1) / (||F_inc_B||*||F_dec_L1||). The calculator performs this operation on all possible combinations and fills the results into a matrix M_sim. The row indices of the matrix correspond to vectors in V_set_inc (representing adding nodes), and the column indices correspond to vectors in V_set_dec (representing removing nodes). Each element in M_sim is a similarity value between -1 and 1; the larger the value, the more matched the intentions of the two nodes.

[0062] Step S133: Based on the magnitude of each similarity value in the initial intent matching similarity matrix, select candidate intent pairings from all possible pairings whose similarity values ​​exceed a preset matching threshold. Each candidate intent pairing includes a feature vector representing an increase in intent and its corresponding power node identifier, and a feature vector representing a decrease in intent and its corresponding power node identifier.

[0063] In this embodiment, a matching threshold is preset, for example, denoted by θ_match. Then, the M_sim matrix is ​​traversed, and for each element sim_incB_decL1 greater than θ_match, its corresponding pairing (adding node B, removing node L1) is recorded as a candidate pairing combination Candidate_B_L1. This candidate pairing combination includes the identifiers ID_B and ID_L1 of the two nodes, as well as their corresponding intent vectors F_inc_B and F_dec_L1.

[0064] Step S134: Perform conflict detection and resolution on all selected candidate intent pairings, identify candidate intent pairings that contain the same feature vector representing an increase intent being matched by multiple feature vectors representing a decrease intent, and identify candidate intent pairings that contain the same feature vector representing a decrease intent being matched by multiple feature vectors representing an increase intent. Mark the above candidate intent pairings with matching conflicts as intent pairings to be negotiated.

[0065] Next, conflict detection is performed on the candidate pairing list. The conflict detector counts the number of times each added node appears in the candidate list. For example, if added node B has candidate pairings with both removed node L1 and removed node U (i.e., Cand_B_L1 and Cand_B_U exist simultaneously), it indicates that node B has been matched multiple times, resulting in a conflict. Similarly, if removed node L1 is paired with both added node B and added node U, it also constitutes a conflict. In this embodiment, all candidate combinations involving such "one-to-many" or "many-to-one" situations, such as Cand_B_L1 and Cand_B_U, are marked as "pairing combinations to be negotiated" and placed in a negotiation list L_negotiate.

[0066] Step S135: Invoke the preset conflict resolution negotiation protocol to conduct multiple rounds of negotiation interaction for each intention pairing combination to be negotiated. In each round of negotiation interaction, the agent instance representing the power node corresponding to the intention feature vector that has a matching conflict and the agent instance representing the power node corresponding to the intention feature vector that has a matching conflict exchange virtual negotiation information. Adjust the similarity calculation weight between their respective intention feature vectors according to the exchanged virtual negotiation information, and recalculate the adjusted similarity value until there is no longer a matching conflict in all intention pairing combinations to be negotiated, and generate a set of conflict-free final intention pairing combinations.

[0067] In order to resolve the conflict, a multi-round negotiation agreement was initiated.

[0068] Step S1351: Mark the agent instance representing the power node corresponding to the intention feature vector of the increase in each intention pairing combination to be negotiated as the increasing agent, and mark the agent instance representing the power node corresponding to the intention feature vector of the decrease in each intention pairing combination to be negotiated as the decreasing agent.

[0069] When handling conflicts, the roles are first clearly defined. For each combination in the list to be negotiated, such as Cand_B_L1, the agent that adds node B is marked as Agent_B_inc, and the agent that removes node L1 is marked as Agent_L1_dec.

[0070] Step S1352: Initialize the first round of negotiation interaction. Each adding agent sends its current representative adding intention feature vector and the corresponding expected power adjustment magnitude value to all adding agents with matching conflicts. Each reducing agent sends its current reducing intention feature vector and the corresponding expected power adjustment magnitude value to all adding agents with matching conflicts.

[0071] Negotiation begins. All labeled amplifying agents, such as Agent_B_inc, not only decode their intent feature vector F_inc_B, but also a specific desired power adjustment magnitude, such as ΔP_B_inc (in kilowatts), and broadcast it through a virtual negotiation channel to all conflicting reducing agents (such as Agent_L1_dec and Agent_U_dec). Similarly, reducing agents also send their vector F_dec_L1 and desired reduction magnitude ΔP_L1_dec to all conflicting amplifying agents.

[0072] Step S1353: After receiving the reduction intention feature vector and the power adjustment expectation magnitude value from multiple reduction agents, each increasing agent calculates the first round negotiation weight adjustment coefficient corresponding to each reducing agent based on the degree of matching between the power adjustment expectation magnitude value of each reducing agent and its own increasing power adjustment expectation magnitude value. The first round negotiation weight adjustment coefficient is inversely proportional to the absolute value of the difference between the two expectation magnitude values.

[0073] After receiving ΔP_L1_dec from Agent_L1_dec and ΔP_U_dec from Agent_U_dec, Agent_B_inc begins calculating the adjustment coefficients. It calculates the absolute value of the magnitude difference with L1: diff_L1 = |ΔP_B_inc - ΔP_L1_dec|. Then, it calculates a temporary candidate adjustment coefficient: w_temp_L1 = 1 / (diff_L1 + ε), where ε is a very small positive number to prevent division by zero. Similarly, w_temp_U is calculated. Then, the temporary coefficients are normalized to obtain the final adjustment coefficients. For example, w_L1 = w_temp_L1 / (w_temp_L1 + w_temp_U). This w_L1 represents the additional weight that B should give to L1 in terms of magnitude matching. The closer the magnitude match (the smaller the diff), the greater the weight. Here, w_temp_U represents the intermediate result obtained by the adding agent when calculating the temporary adjustment coefficients of the subtracting agent U. Similar to w_temp_L1, it is calculated as w_temp_U = 1 / (diff_U + ε), where diff_U is the absolute value of the difference between the expected magnitude of the adding agent and the expected magnitude of the subtracting agent U, and ε is a very small positive number used to prevent division by zero. w_temp_U is then used for normalization to calculate the final weight adjustment coefficient w_U.

[0074] Step S1354: After receiving the addition intention feature vector and the power adjustment expectation magnitude value from multiple addition agents, each reducing agent calculates the first round negotiation weight adjustment coefficient corresponding to each adding agent based on the matching degree between the power adjustment expectation magnitude value of each adding agent and its own reduction power adjustment expectation magnitude value. The first round negotiation weight adjustment coefficient is inversely proportional to the absolute value of the difference between the two expectation magnitude values.

[0075] Similarly, Agent_L1_dec will also calculate its adjustment coefficients for the increasing agents B and U based on the received ΔP_B_inc and ΔP_U_inc, for example, w_B (from L1's perspective). ΔP_U_inc represents the expected magnitude of the increased power received by the decreasing agent (such as L1) from the other increasing agent U in the negotiation. w_B and w_U represent the temporary negotiation weight adjustment coefficients for increasing agents B and U, calculated from the perspective of the decreasing agent L1, based on the degree of matching between the received expected magnitude values ​​(ΔP_B_inc and ΔP_U_inc) of increasing agents B and U and its own expected reduction magnitude. Their calculation logic is the same as that of the increasing agent in calculating w_temp_L1 and w_temp_U, that is, inversely proportional to the absolute value of the difference in expected magnitudes, obtained after normalization.

[0076] Step S1355: Each adding agent uses the calculated first-round negotiated weight adjustment coefficient for each reducing agent to perform a weighted correction on the initial intent matching similarity between itself and each reducing agent, generating a first-round corrected similarity value.

[0077] Now, Agent_B_inc uses the weights w_L1 and w_U it just calculated to correct its similarity with L1 and U. The corrected similarity sim'_B_L1 = sim_incB_decL1 * w_L1, sim'_B_U = sim_incB_decU * w_U. Similarly, Agent_L1_dec also uses its calculated weights w_B and w_U to correct the original similarity, obtaining sim'_L1_B and sim'_L1_U. The above corrected similarity reflects the new matching degree after one round based on the adjustment magnitude. sim_incB_decU represents the initial cosine similarity value between the increasing intent feature vector F_inc_B of the increasing party node B and the decreasing intent feature vector F_dec_U of the decreasing party node U in the initial intent matching similarity matrix M_sim. sim'_L1_B and sim'_L1_U represent, from the perspective of the reducing agent L1, the corrected similarity values ​​between the intentions of the increasing agents B and U after the first round of negotiation. They are obtained by the reducing agent L1 using its own calculated weight adjustment coefficients w_B and w_U to weight and correct the initial similarities sim_incB_decL1 and sim_incU_decL1, respectively.

[0078] Step S1356: Based on the similarity values ​​corrected in the first round, reconstruct the matching relationship between the pairs of intents to be negotiated, identify whether there are still pairs of intents to be negotiated that have matching conflicts after the first round of negotiation interaction, and if there are still pairs of intents to be negotiated that have matching conflicts, then start the second round of negotiation interaction.

[0079] In this embodiment, all corrected similarities are collected, and steps S133 and S134 are executed again, i.e., re-screening and detecting conflicts. For example, if, after correction, the similarity between B and L1 becomes very high, while the similarity with U becomes very low, then "B and L1" may become the unique match, and the conflict is resolved. However, if new or old conflicts still exist, the next round of negotiation begins.

[0080] Step S1357: In the second round of negotiation interaction, the agent that increases the number of matching conflicts and the agent that decreases the number of matching conflicts further exchange the negotiation weight adjustment coefficients calculated in the first round of negotiation interaction, and adjust their first-round corrected similarity values ​​against each other again according to the negotiation weight adjustment coefficients provided by the other party, generate the second-round corrected similarity values, and reconstruct the matching relationship again according to the second-round corrected similarity values.

[0081] In the second round, the agents no longer exchange the original magnitude values, but rather the weight adjustment coefficients calculated in the previous round. For example, Agent_B_inc sends its calculated w_L1 and w_U to Agent_L1_dec and Agent_U_dec. Upon receiving this, Agent_L1_dec merges it with its own weight w_B calculated in the first round (e.g., by averaging) to generate a new joint adjustment coefficient w'_B_L1, which is then used to further adjust the similarity. This represents a deeper level of negotiation compromise. w'_B_L1 represents a new, joint negotiated weight adjustment coefficient generated in the second round of negotiation interaction by the reducing agent L1 after receiving the first-round weight adjustment coefficient w_L1 from the increasing agent B, merging it with its own weight w_B calculated for B in the first round (e.g., by averaging). This coefficient will be used to further adjust the similarity sim'_L1_B after the first round of correction, thus generating the second-round corrected similarity.

[0082] Step S1358: Repeat the negotiation interaction multiple times. In each round of negotiation interaction, agents with matching conflicts exchange the negotiation weight adjustment coefficients calculated in the previous round, and iteratively correct the similarity values ​​based on the exchanged information, until there is no situation in all the intention pairings to be negotiated where any representative increase intention feature vector is matched by multiple representative decrease intention feature vectors, and there is no situation in any representative decrease intention feature vector being matched by multiple representative increase intention feature vectors.

[0083] This negotiation process will proceed iteratively until a stable state is reached, where each added node forms a strong match with at most one removed node, and vice versa. This is similar to a distributed, intent-based matching market clearing process.

[0084] Step S1359: Sort all the pairing combinations that have been formed after multiple rounds of negotiation and interaction and that do not have any matching conflicts in order of the final corrected similarity value from high to low, and select a preset number of pairing combinations at the top of the sorting as the set of conflict-free final intention pairing combinations for output. Each final intention pairing combination in the set of conflict-free final intention pairing combinations contains a final consensus representing an increase in intention feature vector and its corresponding power node identifier, and a final consensus representing a decrease in intention feature vector and its corresponding power node identifier.

[0085] After negotiation, a conflict-free pairing list is obtained. This embodiment then sorts these pairs based on the final corrected similarity, and may select the top K pairs with the highest similarity as the final protocol based on the overall system adjustment requirements. For example, the final selected pairs are (ID_B, ID_L1) and (ID_U, ID_L2), which constitute the final intention pairing combination set.

[0086] Step S136: Assign a globally unique commitment protocol identifier to each final intent pairing in the set of conflict-free final intent pairing combinations, and associate and bind the commitment protocol identifier with the feature vector representing increasing intent and the feature vector representing decreasing intent contained in the final intent pairing combination, to generate a commitment protocol record containing the commitment protocol identifier and the feature vectors representing increasing intent and decreasing intent bound to the commitment protocol identifier.

[0087] Finally, a protocol is generated for each final pairing. A globally unique commitment protocol identifier, such as PID_001, is generated for each pairing (ID_B, ID_L1). Then, a commitment protocol record R_protocol_001 is created. This commitment protocol record contains three core parts: the protocol identifier PID_001, the bound adder intent vector F_inc_B and its node ID_B, and the bound deleter intent vector F_dec_L1 and its node ID_L1.

[0088] Step S137: Gather all commitment protocol records and arrange them in the order of commitment protocol identifiers to form a data set with the commitment protocol identifier as the index and each commitment protocol record as the index value. Output this data set as the multi-agent commitment protocol set. Each commitment protocol record in the multi-agent commitment protocol set also includes a power node identifier corresponding to the increase intention feature vector bound to the commitment protocol identifier and a power node identifier corresponding to the decrease intention feature vector bound to the commitment protocol identifier.

[0089] All generated protocol records, such as R_protocol_001, R_protocol_002, etc., are collected, sorted by PID, and form a mapping Map_protocol. Map_protocol[PID_001] = R_protocol_001. This Map_protocol is the final set of multi-agent commitment protocols, which clarifies which nodes have reached a preliminary consensus on power regulation.

[0090] Step S140: Perform task collaborative grouping on the power nodes corresponding to the intent feature vectors carrying the same commitment protocol identifier in the multi-agent commitment protocol set, and divide the power nodes with the same commitment protocol identifier into the same collaborative task cluster to obtain a multi-cluster task decomposition structure containing the correspondence between collaborative task cluster identifier and power node.

[0091] Next, according to the agreement, the nodes will be grouped to form a cluster to execute tasks.

[0092] Step S141: Parse each commitment protocol record in the multi-agent commitment protocol set, and extract the commitment protocol identifier contained in each commitment protocol record, as well as the power node identifier corresponding to the feature vector representing the increase intention and the power node identifier corresponding to the feature vector representing the decrease intention bound to the commitment protocol identifier.

[0093] In this embodiment, each record in Map_protocol is traversed. For record R_protocol_001, its commitment protocol identifier PID_001, as well as the bound adder node ID_B and remover node ID_L1, are parsed out.

[0094] Step S142: Group and aggregate all extracted power node identifiers according to the extracted commitment protocol identifier, and group all power node identifiers with the same commitment protocol identifier into the same temporary power node group to form a temporary power node group corresponding to a specific commitment protocol identifier.

[0095] In this embodiment, a temporary group mapping Temp_group is created. For PID_001, ID_B and ID_L1 are both placed in a list with PID_001 as the key, i.e., Temp_group[PID_001]=[ID_B, ID_L1]. Similarly, for PID_002, Temp_group[PID_002]=[ID_U, ID_L2].

[0096] Step S143: Deduplicate the power node identifiers within each temporary power node group to ensure that the same power node identifier appears only once in the same temporary power node group, and generate a deduplicated list of power node identifiers corresponding to each commitment protocol identifier.

[0097] Although it is unlikely that a node will appear twice in the same protocol, deduplication is performed on each group to prevent anomalies. For example, if a list is [ID_B, ID_B], it will be corrected to [ID_B]. The resulting deduplicated list L_group_PID001 = [ID_B, ID_L1].

[0098] Step S144: Assign a globally unique collaborative task cluster identifier to each deduplicated power node identifier list corresponding to the commitment protocol identifier, and associate and bind the collaborative task cluster identifier with the corresponding commitment protocol identifier to form a correspondence record between cluster identifier and protocol identifier.

[0099] In this embodiment, a new identifier is assigned to each cluster to be formed, for example, cluster identifier GID_101 corresponds to protocol PID_001. An association record R_link is created, with the content of (GID_101, PID_001).

[0100] Step S145: Combine the data of each collaborative task cluster identifier and its corresponding deduplicated power node identifier list into a cluster node correspondence data unit. The collaborative task cluster node correspondence data unit includes a collaborative task cluster identifier part and a power node identifier list part. The power node identifier list part is arranged in the original number order of the power node identifiers.

[0101] Create a cluster node mapping data unit U_cluster_GID101. This cluster node mapping data unit contains two parts: the header is the cluster identifier GID_101, and the body is a list of node identifiers L_group_PID001, in which the node IDs are sorted by their original numerical values.

[0102] Step S146: Traverse and scan the power node identifier list in the corresponding relational data unit of each cluster node to identify whether there are duplicate power node identifiers in the same power node identifier list. If there are duplicate power node identifiers, the power node identifier list is deduplicated and the duplicate power node identifiers are deleted, and a unique power node identifier record is retained.

[0103] Double-check to ensure there are no duplicates in the list.

[0104] Step S147: Sort the power node identifier list in the corresponding relational data unit of each cluster node, and sort the power node identifiers in the list in ascending order according to the numerical value of the power node identifiers to generate a sorted power node identifier list.

[0105] Sort the list to ensure data consistency and readability. The resulting sorted list is L_group_PID001_sorted.

[0106] Step S148: Encapsulate the collaborative task cluster identifier and the sorted power node identifier list in each cluster node correspondence data unit to generate a standard format cluster node correspondence encapsulation unit. The cluster node correspondence encapsulation unit includes a fixed data header and a variable data body. The data header stores the collaborative task cluster identifier, and the data body stores the sorted power node identifier list.

[0107] GID_101 and L_group_PID001_sorted are encapsulated into a standard unit U_std_GID101.

[0108] Step S149: Sort all standard format cluster node correspondence encapsulation units according to the generation time order of the collaborative task cluster identifier stored in its data header, and generate an ordered encapsulation unit sequence.

[0109] Sort U_std_GID101, U_std_GID102, etc., according to the order in which the GIDs were generated, to form a sequence Seq_cluster_units.

[0110] Step S1410: Package the ordered encapsulation unit sequence as a whole to generate a composite data structure containing the total number of encapsulation units and the specific content of each encapsulation unit, and output the composite data structure as the multi-cluster task decomposition structure.

[0111] The `Seq_cluster_units` are packaged into a data structure `Struct_task_decomp`, which contains the sequence length and the specific content of each unit. `Struct_task_decomp` is the final multi-cluster task decomposition structure.

[0112] Step S1411: While outputting the multi-cluster task decomposition structure, a power node participation cluster status index table is generated. The power node participation cluster status index table uses each power node identifier as an index item. Under each index item, it records which collaborative task cluster identifiers the power node corresponding to that power node identifier appears in. The power node participation cluster status index table is used to quickly query all collaborative task clusters that each power node participates in at the current time.

[0113] In addition, to facilitate subsequent scheduling and querying, an inverted index table, Map_node_to_clusters, is constructed. For example, for node ID_B, which appears in GID_101, then Map_node_to_clusters[ID_B] = [GID_101]. If node B also participates in another cluster, a new entry is added to the list.

[0114] Step S150: Based on the commitment protocol identifier corresponding to each collaborative task cluster in the multi-cluster task decomposition structure, the cluster-level resource scheduler in the distributed collaborative decision-making model is invoked to allocate power output tasks to the power nodes in the collaborative task cluster, generating a set of distributed collaborative scheduling instructions containing the target power value and power output timing of each power node.

[0115] Finally, based on the formed cluster, fine-grained power scheduling is performed.

[0116] Step S151: Extract each collaborative task cluster identifier from the multi-cluster task decomposition structure, and search for the commitment protocol record corresponding to the commitment protocol identifier associated with and bound to the collaborative task cluster identifier from the multi-agent commitment protocol set according to each collaborative task cluster identifier.

[0117] In this embodiment, the cluster identifier, such as GID_101, is first obtained from Struct_task_decomp. Then, the protocol identifier PID_001 corresponding to GID_101 is found through the previously created association record R_link. Finally, the commitment protocol record R_protocol_001 is retrieved from Map_protocol.

[0118] Step S152: Parse each found commitment agreement record, extract the expected power adjustment magnitude corresponding to the feature vector representing the increase intention and the expected power adjustment magnitude corresponding to the feature vector representing the decrease intention bound in the commitment agreement record, and use them as the cluster-level power adjustment target reference value corresponding to the collaborative task cluster.

[0119] Parse R_protocol_001 to extract the expected power adjustment magnitude agreed upon in the protocol. For example, decode ΔP_B_inc (expected increase) from the increasing party's intention vector F_inc_B, and decode ΔP_L1_dec (expected decrease) from the decreasing party's intention vector F_dec_L1. In this embodiment, the smaller of the two values, or the average of the two, may be used as the cluster-level power adjustment target reference value ΔP_cluster_ref for this cluster GID_101.

[0120] Step S153: Input the cluster identifier of each collaborative task cluster and its corresponding cluster-level power adjustment target reference value into the cluster-level resource scheduler in the distributed collaborative decision-making model. The cluster-level resource scheduler obtains the instantaneous output power value and the state of charge percentage of the internal energy storage unit in the node state parameters of each power node at the current running time according to the power node identifier list in each collaborative task cluster.

[0121] Input (GID_101, ΔP_cluster_ref) into the cluster-level resource scheduler. The scheduler finds its node list L_group_PID001_sorted=[ID_B, ID_L1] based on GID_101. Then, the scheduler obtains the current instantaneous output power values ​​P_B_instant and P_L1_instant of node B and node L1, as well as their state of charge percentages SOC_B and SOC_L1 (for load nodes, SOC may represent their adjustable load capacity percentage) from the latest global situation information set S_global.

[0122] Step S154: In the cluster-level resource scheduler, for each collaborative task cluster, based on the expected increase and decrease of power magnitude values ​​in the cluster-level power adjustment target reference value corresponding to the collaborative task cluster, and combined with the instantaneous output power value of each power node in the collaborative task cluster and the state of charge percentage of the internal energy storage unit, a power balance adjustment equation set for the collaborative task cluster is established. The power balance adjustment equation set uses the power output adjustment amount of each power node as the unknown and the total power adjustment amount of the cluster equal to the cluster-level power adjustment target reference value as the equality constraint condition.

[0123] The scheduler establishes an optimization problem for cluster GID_101. Let the power adjustment of node B be ΔP_B_adj (which can be positive or negative, with positive indicating increased output), and the adjustment of node L1 be ΔP_L1_adj. The total power adjustment within the cluster must equal the target value: ΔP_B_adj + ΔP_L1_adj = ΔP_cluster_ref. Furthermore, the adjustment is also constrained by the node's own capabilities. For example, for energy storage node B, its maximum increase in output power is limited by its SOC_B and maximum discharge power P_B_max_discharge, i.e., ΔP_B_adj ≤ min(P_B_max_discharge - P_B_instant, f(SOC_B)), where f is a function mapping SOC to available discharge capacity. Similarly, for load node L1, its ability to reduce power (i.e., perform load reduction) is limited by its current load and minimum load requirement.

[0124] Step S155: Call the linear programming algorithm built into the cluster-level resource scheduler to solve the power balance adjustment equations of each collaborative task cluster, calculate the power output adjustment amount that each power node in the collaborative task cluster needs to adjust, and algebraically add the current instantaneous output power value of each power node to the calculated power output adjustment amount to obtain the target power value of each power node.

[0125] The scheduler uses its built-in linear programming solver to find the optimal ΔP_B_adj and ΔP_L1_adj, satisfying the aforementioned equality and inequality constraints. The objective can be to minimize the total adjustment cost or achieve SOC equilibrium, etc. After obtaining the solution, the target power value for node B is calculated: P_B_target = P_B_instant + ΔP_B_adj; the target power value for node L1 is calculated: P_L1_target = P_L1_instant + ΔP_L1_adj.

[0126] Step S156: Based on the power regulation time response requirement parameters in the cluster-level power regulation target reference value corresponding to each collaborative task cluster, and combined with the magnitude of the power output adjustment of each power node, assign a power output timing identifier to each power node. The power output timing identifier is used to indicate the time sequence and rate of change requirements that the power node should follow in the process of reaching its target power value.

[0127] The scheduler also parses time response requirements from the protocol log, such as requiring adjustment to be completed within three seconds. For nodes with larger adjustment amounts, earlier start or faster rate changes may be required. The scheduler assigns a timing identifier T_seq_B to each node based on its adjustment amount ΔP_adj and its power change rate limit (e.g., the energy storage node can change ΔP_rate_B per second). T_seq_B can be an encoding, such as "start immediately at maximum rate," "start at half rate after a one-second delay," etc.

[0128] Step S157: Associate and bind the identifier of each power node in each collaborative task cluster with its corresponding target power value and power output timing identifier to generate a single-point scheduling instruction unit for a single power node, and package all single-point scheduling instruction units in the same collaborative task cluster to generate a cluster-level scheduling instruction package.

[0129] Generate instruction unit C_B for node B: its content is (ID_B, P_B_target, T_seq_B). Similarly, generate C_L1. Then, package C_B and C_L1 into a cluster instruction package Pkg_GID101.

[0130] Step S158: Gather all cluster-level scheduling instruction packages corresponding to all collaborative task clusters to form a distributed collaborative scheduling instruction set with the collaborative task cluster identifier as the organization unit, and send the distributed collaborative scheduling instruction set to the local controller of each power node through a preset instruction delivery channel to trigger each power node to perform a power output adjustment operation according to the target power value and power output timing identifier in its corresponding single-point scheduling instruction unit.

[0131] Finally, all cluster instruction packets, such as Pkg_GID101 and Pkg_GID102, are aggregated to form the final distributed collaborative scheduling instruction set, Set_commands. In this embodiment, a secure and encrypted communication protocol is used to distribute these instructions in parallel to the local controllers of each power node. For example, after receiving instruction C_B, the local controller of node B parses P_B_target and T_seq_B, and according to timing requirements, controls its power conversion module to adjust the output power, ultimately achieving the goal of coordinating with node L1 to complete overall power regulation.

[0132] The process of acquiring the global situation information set of the power system at the current operating moment, which includes node status parameters uploaded by each power node and power transmission topology connection data between the power nodes, and before inputting the global situation information set into a preset distributed collaborative decision-making model for multi-agent negotiation intent recognition, further includes performing dynamic credit assessment and dynamic granting of negotiation qualifications for the power nodes based on the replay of negotiation history experience. Specifically, this includes: Step S210: After obtaining the global situation information set of the power system at the current operating moment, and before inputting the global situation information set into the preset distributed collaborative decision-making model for multi-agent negotiation intention recognition, the step further includes performing dynamic credit assessment and dynamic granting of negotiation qualifications for the power nodes based on the playback of negotiation history experience on the global situation information set.

[0133] To ensure the reliability and security of the system and prevent unreliable nodes from interfering with the negotiation process, a pre-emptive credit assessment and qualification screening module is introduced before executing the core negotiation steps.

[0134] Step S220: Retrieve a complete set of historical negotiation records of each power node participating in multi-agent negotiation within multiple consecutive historical negotiation cycles before the current running time from the historical operation log database of the preset distributed collaborative decision-making model.

[0135] In this embodiment, a high-performance time-series database specifically for storing historical operational data is first accessed. The complete historical negotiation records of all nodes over a past period, such as the past one hundred negotiation cycles, are retrieved from this database. These records are encapsulated into a large dataset H_record_set. For node B, its historical record includes: the historical intent feature vector F_intent_B_t output for each historical cycle t; the historical commitment protocol identifier PID_B_t finally reached in each cycle t; and the feedback data of the node's actual power execution result for the protocol, collected after each cycle t, such as the actual power adjustment amount ΔP_actual_B_t and the actual power change direction Dir_actual_B_t.

[0136] Step S230: Perform intent fulfillment consistency analysis on the historical intent feature vector corresponding to each power node in the complete historical negotiation record set.

[0137] The credit evaluator begins behavioral analysis for each node. For node B, the evaluator iterates through all its historical records. For each historical period t, it compares the expected power adjustment direction (e.g., DIR_INC or DIR_DEC) decoded from the intent feature vector F_intent_B_t of node B's output in that period with the actual power change direction Dir_actual_B_t at the end of that period. If they match, for example, if the intent increases and the output also actually increases, then the period is marked as a consistency period. The total number of consistency periods over N consecutive historical periods is counted to generate an intent fulfillment consistency period count C_consistency_B.

[0138] Step S240: Perform protocol fulfillment integrity analysis on the historical commitment protocol identifier corresponding to each power node in the complete historical negotiation record set.

[0139] Simultaneously, the evaluator performs a completeness analysis. For each historical period t of node B, the required power adjustment amplitude ΔP_agreed_B_t is extracted from the protocol PID_B_t it achieved. Then, the ratio of the actual power adjustment amplitude ΔP_actual_B_t to the required amplitude is calculated: r_fulfill_B_t = ΔP_actual_B_t / ΔP_agreed_B_t. If this ratio exceeds a preset threshold, for example, if the threshold θ_fulfill is set to 95%, i.e., r_fulfill_B_t > 0.95, then the period is marked as a protocol fulfillment qualified period. The total number of qualified periods within consecutive historical periods is counted to generate a protocol fulfillment qualified period count C_fulfill_B.

[0140] Step S250: Input the intent fulfillment consistency cycle count and protocol fulfillment qualification cycle count corresponding to each power node into the preset power node credit score accumulation function for weighted summation calculation. By assigning different weight coefficients to the intent fulfillment consistency cycle count and the protocol fulfillment qualification cycle count, the total historical negotiation credit score accumulated by each power node before the current running time is generated.

[0141] Next, the credit score is calculated by combining these two indicators. A pre-defined credit score accumulation function is used, for example: Credit_B = w1 * C_consistency_B + w2 * C_fulfill_B, where w1 and w2 are pre-set weighting coefficients, for example, w1 is 0.4 and w2 is 0.6, representing the completeness of fulfilling the agreement compared to... Figure 1 Consistency is more important. This function calculates the total historical negotiated credit score of node B, Credit_B.

[0142] Step S260: Compare the total historical negotiation credit score of each power node with the preset negotiation qualification granting score threshold, identify power nodes whose total historical negotiation credit score is lower than the negotiation qualification granting score threshold, mark the above power nodes as power nodes with insufficient credit to be observed at the current running time, and generate a list of power nodes with insufficient credit to be observed containing all power nodes with insufficient credit to be observed.

[0143] In this embodiment, a credit score threshold θ_credit is established. The credit score of each node is compared with θ_credit. For example, if Credit_B < θ_credit, then node B is determined to have insufficient credit. The identifier ID_B of all such nodes, such as node B, is added to a list of nodes with insufficient credit, L_unreliable.

[0144] Step S270: Extract the node status parameters corresponding to each under-credit power node to be observed in the list of under-credit nodes from the global situation information set, input the above node status parameters into a preset temporary negotiation observation period data channel for caching, and temporarily remove the node status parameters corresponding to each under-credit power node to be observed in the list of under-credit nodes from the global situation information set, and temporarily disconnect the power transmission line connection relationship description data associated with each under-credit power node to be observed in the power transmission topology connection data, and generate a temporary negotiation participation global situation information set containing only information on credit qualified power nodes.

[0145] In this embodiment, the original global situation information set S_global is "filtered" based on L_unreliable. First, the state parameters of node B are copied from S_global and stored in a dedicated temporary observation cache Cache_observe. Then, all state parameters of node B are removed from S_global, and the topology connection data is traversed to remove all connection relationship descriptions related to node B (e.g., removing node B's information from the adjacency table and removing ID_B from other nodes' neighbor lists). After this operation, S_global is simplified to a temporary negotiation-participation global situation information set S_global_clean containing only creditworthy nodes (such as U, L1, etc.).

[0146] Step S280: The temporary negotiation participation global situation information set is used as the input data for the multi-agent negotiation intention recognition processing of the global situation information set into the preset distributed collaborative decision-making model. At the same time, the list of nodes with insufficient credit and their corresponding node status parameters are sent to the preset credit observation period independent negotiation simulator for bypass virtual negotiation simulation. After the total historical negotiation credit score of the power node with insufficient credit exceeds the negotiation qualification granting score threshold in multiple consecutive subsequent running times, the power node with insufficient credit is re-included into the formal global situation information set to participate in subsequent multi-agent negotiations.

[0147] The formal negotiation process will be initiated based on S_global_clean (i.e., executing step S120). Simultaneously, a bypass credit observation period independent negotiation simulator will be launched to isolate and evaluate nodes with insufficient credit.

[0148] For example, in step S281: create an independent virtual negotiation agent instance for each power supply node with insufficient credit in the list of power supply nodes with insufficient credit, and input the node state parameters corresponding to the power supply node with insufficient credit into the virtual state encoder module of the corresponding virtual negotiation agent instance for virtual state feature encoding to generate the virtual original state feature vector corresponding to the power supply node with insufficient credit.

[0149] The simulator creates an independent virtual agent Agent_B_virtual for node B. The state parameters of node B in the cache are input into its virtual state encoder (the structure is the same as the encoder in step S123) to generate a virtual raw state feature vector F_raw_B_virtual.

[0150] Step S282: Input the virtual original state feature vector corresponding to each under-credit power node to be observed into the virtual topology connection construction module of the preset credit observation period independent negotiation simulator. Based on the power transmission line connection relationship description data originally associated with the under-credit power node in the power transmission topology connection data, construct virtual connection relationships between each virtual negotiation agent instance and the real agent instances corresponding to other credit qualified power nodes, forming a credit observation period virtual negotiation topology structure containing virtual connection relationships.

[0151] The simulator retrieves historical topology data for node B from the original S_global and uses this data to establish virtual connections between its virtual agent Agent_B_virtual and real agents (such as Agent_U_real and Agent_L1_real). Although B is excluded in real negotiations, these connections exist in the simulator, forming a virtual topology.

[0152] Step S283: During the formal multi-agent negotiation process at each current running moment, each virtual negotiation agent instance in the virtual negotiation topology of the credit observation period is synchronously activated with the real agent instance of the corresponding credit-qualified power node in the temporary negotiation participation global situation information set, so that each virtual negotiation agent instance can receive the real-time intent feature vector output by the real agent instance with which it has a virtual connection relationship during the formal negotiation process, and generate the virtual response intent feature vector corresponding to the virtual negotiation agent instance based on the received real-time intent feature vector.

[0153] When the actual negotiation is underway (in progress of step S120), the virtual agent Agent_B_virtual is also activated. It is able to "listen" to the real-time intent feature vector F_intent_U_real generated by its virtual neighbor (such as the real agent Agent_U_real) during the formal negotiation. Based on F_intent_U_real and its own state F_raw_B_virtual, Agent_B_virtual simulates the virtual response intent feature vector F_intent_B_virtual that it could generate if it participated in the negotiation.

[0154] Step S284: Input the virtual response intent feature vector generated by each virtual negotiation agent instance into the virtual negotiation commitment matching module of the credit observation period independent negotiation simulator, and perform virtual negotiation commitment matching and deduction with the real-time intent feature vectors of other virtual negotiation agent instances or real agent instances in the virtual space to generate the virtual commitment agreement identifier of the under-credit power node to be observed in this negotiation cycle.

[0155] The virtual commitment matching module in the simulator uses the same logic as in step S130 to match the virtual intent F_intent_B_virtual with the real intent that has been monitored, and deduce the virtual commitment protocol identifier PID_B_virtual that node B should have reached in this negotiation.

[0156] Step S285: After the current negotiation cycle ends, compare the virtual power adjustment expectation direction expressed by the virtual response intent feature vector corresponding to each under-credit power node with the actual power change direction in actual operation of the under-credit power node to perform a virtual intent fulfillment consistency comparison. Then, compare the virtual protocol power adjustment magnitude corresponding to the virtual commitment protocol identifier of the under-credit power node with the actual power adjustment magnitude in actual operation of the under-credit power node to perform a virtual protocol fulfillment integrity comparison. Update the total historical negotiation credit score of the under-credit power node based on the comparison results.

[0157] After the negotiation period ends, the actual power change direction Dir_actual_B and actual adjustment magnitude ΔP_actual_B of Node B during that period are obtained. Then, the simulator compares the desired direction in the virtual intent F_intent_B_virtual with Dir_actual_B, and the required magnitude in the virtual protocol PID_B_virtual with ΔP_actual_B. Based on these two comparison results (similar to steps S230 and S240), a "virtual credit score increase value" is calculated and added to Node B's total credit score Credit_B.

[0158] Step S286: Repeat the above virtual negotiation simulation and credit score update operation at each subsequent running time, and after each update, determine whether the total historical negotiation credit score of the under-credit power node has accumulated to exceed the negotiation qualification granting score threshold. Once it is detected that the total historical negotiation credit score of an under-credit power node exceeds the negotiation qualification granting score threshold, immediately remove the under-credit power node from the under-credit node list.

[0159] In this embodiment, steps S281 to S285 are repeated in each control cycle to continuously update the credit score of node B. After one cycle ends, it is checked again whether Credit_B>=θ_credit. Once the condition is met, node B is removed from the L_unreliable list.

[0160] Step S287: Add the node status parameters corresponding to the power node that has completed the credit score accumulation and reached the negotiation qualification granting score threshold back to the global situation information set at subsequent running times. At the same time, reconnect the power transmission line connection relationship description data associated with the power node in the power transmission topology connection data, so that the power node regains full qualification to participate in the formal multi-agent negotiation process.

[0161] After the data acquisition step in the next cycle (step S110), it is detected that node B is no longer in L_unreliable. Therefore, when constructing the global situational information set, the state parameters of node B and its topological connectivity can be included again to generate a complete S_global containing node B. At this point, after the observation period and credit repair, node B is re-admitted to the formal negotiation process, realizing dynamic, historical behavior-based admission control.

[0162] In one exemplary embodiment, a multi-agent-based distributed collaborative power system is provided. This multi-agent-based distributed collaborative power system can be a terminal, server, etc., and its internal structure diagram can be as follows: Figure 2 As shown, this multi-agent-based distributed collaborative power system includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, near-field communication, or other technologies. When the computer program is executed by the processor, it implements a multi-agent-based distributed collaborative power system method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device can be a touch layer covering the display screen, or a button, trackball, or touchpad set on the shell of a multi-agent power system distributed collaborative system, or an external keyboard, touchpad, or mouse, etc.

[0163] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

Claims

1. A distributed cooperative method for a power system based on multi-agent systems, characterized in that, The method includes: Obtain a global situation information set of the power system at the current operating moment. The global situation information set includes node status parameters uploaded by each power node and power transmission topology connection data between each power node. The global situation information set is input into a preset distributed collaborative decision-making model for multi-agent negotiation intent recognition. Each agent in the distributed collaborative decision-making model generates a power adjustment negotiation intent for adjacent power nodes based on its corresponding node state parameters and the power transmission topology connection data, resulting in a multi-agent intent distribution set containing the intent feature vectors output by each agent. Based on the vector space similarity calculation results between each intention feature vector in the multi-agent intention distribution set, cross-agent negotiation commitment matching is performed, and a corresponding target commitment protocol identifier is assigned to each intention feature vector to generate a multi-agent commitment protocol set containing the correspondence between commitment protocol identifiers and intention feature vectors. Task collaborative grouping is performed on the power nodes corresponding to the intent feature vectors carrying the same commitment protocol identifier in the multi-agent commitment protocol set. Power nodes with the same commitment protocol identifier are divided into the same collaborative task cluster, resulting in a multi-cluster task decomposition structure that includes the correspondence between collaborative task cluster identifier and power node. Based on the commitment protocol identifier corresponding to each collaborative task cluster in the multi-cluster task decomposition structure, the cluster-level resource scheduler in the distributed collaborative decision-making model is invoked to allocate power output tasks to the power nodes within the collaborative task cluster, generating a set of distributed collaborative scheduling instructions containing the target power value and power output timing of each power node.

2. The distributed cooperative method for power systems based on multi-agent systems according to claim 1, characterized in that, The process involves inputting the global situational information set into a preset distributed collaborative decision-making model for multi-agent negotiation intent recognition. Each agent in the distributed collaborative decision-making model generates a power adjustment negotiation intent for adjacent power nodes based on its corresponding node state parameters and the power transmission topology connection data. This results in a multi-agent intent distribution set containing the intent feature vectors output by each agent, including: The node status parameters uploaded by each power node in the global situation information set are indexed and classified according to the power node identifier, and a unique individual node status parameter record is generated for each power node. The individual node status parameter record includes the instantaneous output power value, instantaneous input power value, state of charge percentage of the internal energy storage unit, and real-time working efficiency percentage of the power conversion module at the current operating moment of the power node. The power transmission topology connection data between each power node in the global situation information set is input into the topology relationship parsing module of the distributed collaborative decision-making model to deconstruct the connection relationship, generating a list of adjacent node identifiers with each power node as the central node and power transmission line parameters with each power node as the central node. The power transmission line parameters include the transmission line resistance value, transmission line reactance value and real-time transmission power value on the current transmission line between the power node and each adjacent node. For each power node, a corresponding agent instance is deployed. The state parameters of the individual node corresponding to each power node, the list of neighboring node identifiers centered on the power node, and the power transmission line parameters centered on the power node are simultaneously input into the intent encoder module of the agent instance corresponding to the power node for feature encoding, generating an original state feature vector containing the power node's own state information and its connection relationship with neighboring nodes. The original state feature vector generated by the intent encoder module of the agent instance corresponding to each power node is input into the attention calculation module of the neighboring nodes of the agent instance for attention weight allocation. Based on the feature similarity calculation results between the original state feature vector of the power node and the original state feature vector of each neighboring node of the power node, attention weight coefficients for each neighboring node are generated, resulting in a node attention distribution matrix containing the attention weight coefficients of the power node and all neighboring nodes. The original state feature vector of the agent instance corresponding to each power node and the inter-node attention distribution matrix of the agent instance are simultaneously input into the intent generation module of the agent instance for intent feature mapping. The input original state feature vector is nonlinearly transformed by the fully connected network layer in the intent generation module to generate a preliminary intent feature vector. The intent weighted fusion layer in the intent generation module multiplies each attention weight coefficient in the inter-node attention distribution matrix with the corresponding feature dimension in the preliminary intent feature vector element by element to generate an intent feature vector enhanced by the attention mechanism. The intent feature vectors output by the intent generation module of the agent instance corresponding to each power node are collected and aggregated according to the power node identifier. The intent feature vectors corresponding to all power nodes are combined into a multi-dimensional vector set, and the multi-dimensional vector set is output as the multi-agent intent distribution set. Each intent feature vector in the multi-agent intent distribution set carries its corresponding power node identifier and the power adjustment target direction identifier to which the intent feature vector is targeted. The power adjustment target direction identifier is used to indicate whether the power adjustment expectation expressed by the intent feature vector is to increase or decrease the output power.

3. The distributed cooperative method for power systems based on multi-agent systems according to claim 2, characterized in that, The original state feature vector generated by the intent encoder module of the agent instance corresponding to each power node is input into the neighboring node attention calculation module of the agent instance for attention weight allocation. Based on the feature similarity calculation results between the original state feature vector of the power node and the original state feature vectors corresponding to each neighboring node of the power node, attention weight coefficients for each neighboring node are generated, resulting in an inter-node attention distribution matrix containing the attention weight coefficients of the power node and all neighboring nodes, including: The original state feature vector generated by the intent encoder module of the agent instance corresponding to each power node is simultaneously sent to the neighboring node attention calculation module of the agent instance and the neighboring node attention calculation module of the agent instance corresponding to each neighboring node of the power node, so that the neighboring node attention calculation module of each agent instance simultaneously receives the original state feature vector of its own power node and the original state feature vectors of all neighboring nodes. In the neighboring node attention calculation module of each agent instance, the original state feature vector of the power node where it is located is multiplied by the original state feature vector of each neighboring node to obtain the initial attention score between the power node where it is located and each neighboring node. The initial attention score is used to characterize the degree of mutual attention between the two power nodes in the current operating state for power regulation negotiation. In the neighboring node attention calculation module of each agent instance, the initial attention score between the power node where it is located and all neighboring nodes is normalized. By dividing each initial attention score by the sum of all initial attention scores, a normalized attention weight coefficient corresponding to each initial attention score is generated. The magnitude of the normalized attention weight coefficient is positively correlated with the magnitude of the initial attention score. In the neighboring node attention calculation module of each agent instance, the normalized attention weight coefficients corresponding to the power node where the agent is located and each neighboring node are arranged in the order of the neighboring node's identifier to construct a one-dimensional row vector. Each element position of this row vector corresponds one-to-one with the identifier of each neighboring node, and the value of each element position is the normalized attention weight coefficient corresponding to that neighboring node. The row vector containing normalized attention weight coefficients constructed by the neighboring node attention calculation module of each agent instance is output as the inter-node attention distribution vector corresponding to the power node. The inter-node attention distribution vectors corresponding to all power nodes are stacked vertically in the order of power node identifiers to form a two-dimensional inter-node attention distribution matrix. The row index of the inter-node attention distribution matrix corresponds to the power node identifier that initiated the attention calculation, and the column index of the inter-node attention distribution matrix corresponds to the neighboring node identifier that is being paid attention to. The value of each element in the inter-node attention distribution matrix is ​​the normalized attention weight coefficient from the power node corresponding to the row index to the neighboring node corresponding to the column index.

4. The distributed cooperative method for power systems based on multi-agent systems according to claim 2, characterized in that, The step of collecting and aggregating the intent feature vectors output by the intent generation module for each agent instance corresponding to each power node according to the power node identifier, combining the intent feature vectors corresponding to all power nodes into a multi-dimensional vector set, and outputting the multi-dimensional vector set as the multi-agent intent distribution set includes: After the intent generation module of the agent instance corresponding to each power node completes the intention feature vector generation operation, it encapsulates the intention feature vector with its corresponding power node identifier and the current power adjustment target direction identifier of the power node to generate an intent data encapsulation unit containing vector data and identifier data. The vector data stores all the numerical elements of the intention feature vector, and the identifier data stores the power node identifier and the power adjustment target direction identifier. All intent data encapsulation units corresponding to power nodes are sent to a preset central intent collector for data aggregation. The central intent collector temporarily stores the received intent data encapsulation units in a temporary data buffer according to the time sequence of the received intent data encapsulation units. After the central intent collector completes the receiving operation of the intent data encapsulation units corresponding to all power nodes, the intent data aggregation process is started. Each intent data encapsulation unit is read sequentially from the temporary data buffer, and the power node identifier and power adjustment target direction identifier are extracted from the identifier data of each intent data encapsulation unit. The intent data encapsulation units read are classified at the first level based on the extracted power node identifier. Intent data encapsulation units with the same power node identifier are grouped into the intent data subset corresponding to the same power node. Each intent data subset corresponding to a power node contains all intent feature vectors generated by that power node during this negotiation intent recognition process. After completing the first level of classification, the second level of classification is performed on the intent data encapsulation units within the intent data subset corresponding to each power node. Based on the power adjustment target direction identifier in the identifier data of each intent data encapsulation unit, the intent data encapsulation units within the intent data subset are further divided into power increase intent data groups and power decrease intent data groups. Select one intention feature vector from the power increase intention data group corresponding to each power node as the representative power increase intention feature vector of that power node, and select one intention feature vector from the power decrease intention data group corresponding to each power node as the representative power decrease intention feature vector of that power node. Combine the representative power increase intention feature vector and the representative power decrease intention feature vector of each power node to form the final intention feature vector pair of that power node. The final intent feature vector pairs of all power nodes are arranged in the order of power node identifiers to form a data structure with the power node identifier as the index and the final intent feature vector pair of each power node as the index value. This data structure is output as the multi-agent intent distribution set. In the multi-agent intent distribution set, the final intent feature vector pair corresponding to each power node carries an increase output power identifier in the increase intent feature vector and a decrease intent feature vector carries a decrease output power identifier.

5. The distributed cooperative method for power systems based on multi-agent systems according to claim 1, characterized in that, The step of performing cross-agent negotiation commitment matching based on the vector space similarity calculation results between each intent feature vector in the multi-agent intent distribution set, assigning a corresponding target commitment protocol identifier to each intent feature vector, and generating a multi-agent commitment protocol set containing the correspondence between commitment protocol identifiers and intent feature vectors includes: Extract all representative increase intention feature vectors carrying the increase output power identifier from the multi-agent intention distribution set to form an increase intention feature vector set, and extract all representative decrease intention feature vectors carrying the decrease output power identifier from the multi-agent intention distribution set to form a decrease intention feature vector set. For each representative feature vector of increasing intent in the set of increasing intent feature vectors and each representative feature vector of decreasing intent in the set of decreasing intent feature vectors, calculate the vector space similarity between each pair of representative feature vectors of increasing intent and representative feature vectors of decreasing intent. By calculating the cosine similarity value between each pair of representative feature vectors of increasing intent and representative feature vectors of decreasing intent, generate an initial intent matching similarity matrix containing all possible pairings. Based on the magnitude of each similarity value in the initial intent matching similarity matrix, candidate intent pairings with similarity values ​​exceeding a preset matching threshold are selected from all possible pairings. Each candidate intent pairing contains a feature vector representing an increase in intent and its corresponding power node identifier, and a feature vector representing a decrease in intent and its corresponding power node identifier. Conflict detection and resolution are performed on all selected candidate intent pairs. Candidate intent pairs that contain the same feature vector representing an increase intent being matched by multiple feature vectors representing a decrease intent, and candidate intent pairs that contain the same feature vector representing a decrease intent being matched by multiple feature vectors representing an increase intent, are identified. The candidate intent pairs with matching conflicts are marked as intent pairs to be negotiated. The pre-defined conflict resolution negotiation protocol is invoked to conduct multiple rounds of negotiation interaction for each intention pairing combination to be negotiated. In each round of negotiation interaction, the agent instance representing the power node corresponding to the intention feature vector that has a matching conflict and the agent instance representing the power node corresponding to the intention feature vector that has a matching conflict exchange virtual negotiation information. The similarity calculation weight between their respective intention feature vectors is adjusted according to the exchanged virtual negotiation information, and the adjusted similarity value is recalculated until there is no longer a matching conflict in all intention pairing combinations to be negotiated, generating a set of conflict-free final intention pairing combinations. Assign a globally unique commitment protocol identifier to each final intent pairing combination in the set of conflict-free final intent pairing combinations, and associate and bind the commitment protocol identifier with the feature vector representing increasing intent and the feature vector representing decreasing intent contained in the final intent pairing combination, to generate a commitment protocol record containing the commitment protocol identifier and the feature vectors representing increasing intent and the feature vectors representing decreasing intent bound to the commitment protocol identifier; All commitment protocol records are collected and arranged in the order of commitment protocol identifiers to form a data set with the commitment protocol identifier as the index and each commitment protocol record as the index value. This data set is output as the multi-agent commitment protocol set. Each commitment protocol record in the multi-agent commitment protocol set also includes a power node identifier corresponding to the increase intention feature vector bound to the commitment protocol identifier and a power node identifier corresponding to the decrease intention feature vector bound to the commitment protocol identifier.

6. The distributed cooperative method for power systems based on multi-agent systems according to claim 5, characterized in that, The process involves invoking a preset conflict resolution negotiation protocol to conduct multiple rounds of negotiation interactions for each intent pairing to be negotiated. In each round of negotiation, agent instances representing the addition of power nodes corresponding to intent feature vectors and agent instances representing the reduction of power nodes corresponding to intent feature vectors exchange virtual negotiation information. The similarity calculation weights between their respective intent feature vectors are adjusted based on the exchanged virtual negotiation information, and the adjusted similarity values ​​are recalculated until no further matching conflicts exist among all intent pairings to be negotiated, generating a final set of conflict-free intent pairings, including: In each pair of intents to be negotiated, the agent instance representing the power node corresponding to the intention to increase is marked as the increasing agent, and the agent instance representing the power node corresponding to the intention to decrease is marked as the decreasing agent. Initiate the first round of negotiation interaction. Each adding agent sends its current representative adding intention feature vector and the corresponding expected power adjustment magnitude value to all adding agents that have matching conflicts with it. Each reducing agent sends its current reducing intention feature vector and the corresponding expected power adjustment magnitude value to all adding agents that have matching conflicts with it. After receiving the reduction intention feature vector and the power adjustment expectation magnitude value from multiple reduction agents, each adding agent calculates the first round negotiation weight adjustment coefficient for each reducing agent based on the degree of matching between the power adjustment expectation magnitude value of each reducing agent and its own added power adjustment expectation magnitude value. The first round negotiation weight adjustment coefficient is inversely proportional to the absolute value of the difference between the two expectation magnitude values. After receiving the addition intention feature vector and the power adjustment expectation magnitude value from multiple addition agents, each reducing agent calculates the first round negotiation weight adjustment coefficient for each adding agent based on the degree of matching between the power adjustment expectation magnitude value of each adding agent and its own reduction power adjustment expectation magnitude value. The first round negotiation weight adjustment coefficient is inversely proportional to the absolute value of the difference between the two expectation magnitude values. Each adding agent uses the calculated first-round negotiated weight adjustment coefficient for each reducing agent to perform a weighted correction on the initial intent matching similarity between itself and each reducing agent, generating a first-round corrected similarity value. Based on the similarity values ​​corrected in the first round, the matching relationship between the pairs of intents to be negotiated is reconstructed. It is then identified whether there are still pairs of intents to be negotiated that have matching conflicts after the first round of negotiation interaction. If there are still pairs of intents to be negotiated that have matching conflicts, the second round of negotiation interaction is initiated. In the second round of negotiation, the agents that still have matching conflicts, the adding agent and the reducing agent, further exchange the negotiation weight adjustment coefficients calculated in the first round of negotiation. They then adjust their first-round corrected similarity values ​​against each other based on the negotiation weight adjustment coefficients provided by the other party, generate the second-round corrected similarity values, and reconstruct the matching relationship based on the second-round corrected similarity values. Repeatedly execute multiple rounds of negotiation interaction. In each round of negotiation interaction, agents with matching conflicts exchange the negotiation weight adjustment coefficients calculated in the previous round, and iteratively correct the similarity values ​​based on the exchanged information, until there is no situation in all the intention pairings to be negotiated where any feature vector representing the intention to increase is matched by multiple feature vectors representing the intention to decrease, and there is no situation in any feature vector representing the intention to decrease is matched by multiple feature vectors representing the intention to increase. All pairings that have been formed after multiple rounds of negotiation and interaction and have no matching conflicts are sorted in descending order of their final corrected similarity scores. A preset number of pairings at the top of the sorted list are selected as the set of conflict-free final intent pairings for output. Each final intent pairing in the set of conflict-free final intent pairings includes a final agreed-upon feature vector representing an increase in intent and its corresponding power node identifier, and a feature vector representing a decrease in intent and its corresponding power node identifier.

7. The distributed cooperative method for power systems based on multi-agent systems according to claim 1, characterized in that, The step of task collaborative grouping of power nodes corresponding to intent feature vectors carrying the same commitment protocol identifier in the multi-agent commitment protocol set, and grouping power nodes with the same commitment protocol identifier into the same collaborative task cluster, yields a multi-cluster task decomposition structure containing the correspondence between collaborative task cluster identifiers and power nodes, including: Parse each commitment protocol record in the multi-agent commitment protocol set, and extract the commitment protocol identifier contained in each commitment protocol record, as well as the power node identifier corresponding to the feature vector representing the increase intention and the power node identifier corresponding to the feature vector representing the decrease intention bound to the commitment protocol identifier; Based on the extracted commitment protocol identifier, all extracted power node identifiers are grouped and aggregated. All power node identifiers with the same commitment protocol identifier are grouped into the same temporary power node group, forming a temporary power node group corresponding to a specific commitment protocol identifier. Deduplicat the power node identifiers within each temporary power node group to ensure that the same power node identifier appears only once in the same temporary power node group, and generate a deduplicated list of power node identifiers corresponding to each commitment protocol identifier. Assign a globally unique collaborative task cluster identifier to the deduplicated list of power node identifiers corresponding to each commitment protocol identifier, and associate and bind the collaborative task cluster identifier with the corresponding commitment protocol identifier to form a record of the correspondence between cluster identifier and protocol identifier; Each collaborative task cluster identifier and its corresponding deduplicated power node identifier list are combined into a cluster node correspondence data unit. This collaborative task cluster node correspondence data unit includes a collaborative task cluster identifier part and a power node identifier list part. The power node identifier list part is arranged according to the original number order of the power node identifiers. The power node identifier list in the corresponding data unit of each cluster node is traversed and scanned to identify whether there are duplicate power node identifiers in the same power node identifier list. If duplicate power node identifiers exist, the power node identifier list is deduplicated and the duplicate power node identifiers are deleted, and a unique power node identifier record is retained. Sort the power node identifier list in the data unit corresponding to each cluster node by sorting the list elements, and sort the power node identifiers in the list in ascending order according to the numerical value of the power node identifiers to generate a sorted power node identifier list. The collaborative task cluster identifier and the sorted power node identifier list in each cluster node correspondence data unit are encapsulated to generate a standard format cluster node correspondence encapsulation unit. This encapsulation unit contains a fixed data header and a variable data body. The data header stores the collaborative task cluster identifier, and the data body stores the sorted power node identifier list. All standard format cluster node correspondence encapsulation units are sorted according to the generation time order of the collaborative task cluster identifier stored in their data header, generating an ordered encapsulation unit sequence. The ordered sequence of encapsulation units is packaged as a whole to generate a composite data structure containing information on the total number of encapsulation units and the specific content of each encapsulation unit. This composite data structure is then output as the multi-cluster task decomposition structure. While outputting the multi-cluster task decomposition structure, a power node participation cluster index table is generated. This power node participation cluster index table uses each power node identifier as an index item. Under each index item, it records which collaborative task cluster identifiers the power node corresponding to that power node identifier appears in. The power node participation cluster index table is used to quickly query all collaborative task clusters that each power node participates in at the current time.

8. The distributed cooperative method for power systems based on multi-agent systems according to claim 1, characterized in that, Based on the commitment protocol identifier corresponding to each collaborative task cluster in the multi-cluster task decomposition structure, the cluster-level resource scheduler in the distributed collaborative decision-making model is invoked to allocate power output tasks to the power nodes within the collaborative task cluster, generating a set of distributed collaborative scheduling instructions containing the target power value and power output timing of each power node, including: Extract each collaborative task cluster identifier from the multi-cluster task decomposition structure, and search for the commitment protocol record corresponding to the commitment protocol identifier associated with and bound to the collaborative task cluster identifier from the multi-agent commitment protocol set based on each collaborative task cluster identifier. Analyze each found commitment agreement record, extract the expected power adjustment magnitude corresponding to the feature vector representing the increase intention and the expected power adjustment magnitude corresponding to the feature vector representing the decrease intention bound in the commitment agreement record, and use them as the cluster-level power adjustment target reference value for the collaborative task cluster. Each collaborative task cluster identifier and its corresponding cluster-level power adjustment target reference value are input into the cluster-level resource scheduler in the distributed collaborative decision-making model. The cluster-level resource scheduler obtains the instantaneous output power value and the state of charge percentage of the internal energy storage unit in the node state parameters of each power node at the current running time based on the power node identifier list in each collaborative task cluster. In the cluster-level resource scheduler, for each collaborative task cluster, based on the expected increase and decrease of power magnitude values ​​in the cluster-level power adjustment target reference value corresponding to the collaborative task cluster, and combined with the instantaneous output power value of each power node in the collaborative task cluster and the state of charge percentage of the internal energy storage unit, a power balance adjustment equation set for the collaborative task cluster is established. The power balance adjustment equation set uses the power output adjustment amount of each power node as the unknown and the total power adjustment amount of the cluster equal to the cluster-level power adjustment target reference value as the equality constraint condition. The linear programming algorithm built into the cluster-level resource scheduler is invoked to solve the power balance adjustment equations of each collaborative task cluster, calculate the power output adjustment amount that each power node in the collaborative task cluster needs to adjust, and algebraically add the current instantaneous output power value of each power node to the calculated power output adjustment amount to obtain the target power value of each power node. Based on the power regulation time response requirement parameters in the cluster-level power regulation target reference value corresponding to each collaborative task cluster, and combined with the magnitude of the power output adjustment of each power node, a power output timing identifier is assigned to each power node. The power output timing identifier is used to indicate the time sequence and rate of change requirements that the power node should follow in the process of reaching its target power value. Associate and bind each power node identifier in each collaborative task cluster with its corresponding target power value and power output timing identifier to generate a single-point scheduling instruction unit for a single power node, and package all single-point scheduling instruction units in the same collaborative task cluster to generate a cluster-level scheduling instruction package. All cluster-level scheduling instruction packages corresponding to the collaborative task clusters are aggregated to form a distributed collaborative scheduling instruction set organized by the collaborative task cluster identifier. This distributed collaborative scheduling instruction set is then sent to the local controller of each power node through a preset instruction delivery channel to trigger each power node to perform a power output adjustment operation based on the target power value and power output timing identifier in its corresponding single-point scheduling instruction unit.

9. A distributed cooperative system for power supply systems based on multi-agent systems, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the distributed cooperative method for a multi-agent power system according to any one of claims 1 to 8 by executing the machine-executable instructions.

10. A computer program product, characterized in that, The computer program product includes machine-executable instructions stored in a computer-readable storage medium. A processor of the multi-agent-based distributed power system cooperative system reads the machine-executable instructions from the computer-readable storage medium and executes the machine-executable instructions, causing the multi-agent-based distributed power system cooperative system to perform the multi-agent-based distributed power system cooperative method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • WebUI self-adaptive test system and method based on multi-agent cooperation and multi-mode perception

    CN121434099A

  • Computing power center multi-agent energy consumption collaboration method and system

    CN121635654A