A cross-system data security sharing method for a green power supply chain

By optimizing the participant set through clustering and screening methods based on network latency and resource information in the green power grid supply chain, the overfitting problem of the federated learning model is solved, high-quality cross-system data security sharing is achieved, and the traceability capability of the green supply chain is guaranteed.

CN121770898BActive Publication Date: 2026-05-22TECH TRAINING CENT OF STATE GRID HUBEI ELECTRIC POWER CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TECH TRAINING CENT OF STATE GRID HUBEI ELECTRIC POWER CO LTD
Filing Date
2026-02-28
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing federated learning node selection methods based on optimizing node performance and communication costs fail to effectively consider the strong correlation between multiple participants in power grid data scenarios, leading to overfitting of the federated learning model. This affects the accuracy and stability of cross-system data sharing and cannot meet the traceability requirements of green supply chains in new power systems.

Method used

By clustering based on network latency, a participant structure graph is constructed, aggregation coefficients are determined, candidate participants are screened out, and the participant set is optimized for training the federated learning model.

Benefits of technology

It improves the accuracy and stability of federated learning models, ensures the reliability of green supply chain data sharing, and realizes the visibility, manageability and traceability of carbon flows, meeting the green traceability requirements of new power systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121770898B_ABST
    Figure CN121770898B_ABST
Patent Text Reader

Abstract

The application discloses a cross-system data security sharing method for a green power grid supply chain, and relates to the technical field of resource allocation. The method comprises the following steps: in response to obtaining a first candidate participant in the green power grid supply chain, clustering each first candidate participant based on the network delay between the first candidate participants to obtain at least one first participant cluster; determining a first aggregation coefficient of the first participant cluster based on the resource information of each first candidate participant in the first participant cluster; the first aggregation coefficient is used to represent the aggregation degree of the first participant cluster; based on the first aggregation coefficient, the first candidate participant in the first participant cluster is screened out until the first aggregation coefficient meets a preset stop condition, and a plurality of target participants are obtained; the target participants are used to share data to train a federated learning model. The application can realize cross-system data security sharing for the green power grid supply chain with high quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of resource allocation technology, and more specifically to a cross-system data security sharing method for a green power grid supply chain. Background Technology

[0002] Currently, the construction of new power systems places high demands on the green traceability capabilities of the supply chain. It is necessary to accurately track and record the green electricity attributes and carbon emission data of electricity throughout the entire process "from generation to consumption". This relies on connecting data from multiple links such as power sources and power grids to achieve visible, manageable and traceable carbon flow.

[0003] Currently, when implementing secure cross-system data sharing based on federated learning, the selection of federated learning nodes is mainly based on optimizing node performance and related communication costs, thereby training the federated learning model and ensuring the execution of privacy-preserving computation tasks.

[0004] However, in power grid data scenarios, there may be strong interrelationships among multiple participants. Existing federated learning node selection methods, which optimize node performance and communication costs, do not take this into account. This can easily lead to overfitting during federated learning model training, hindering high-quality cross-system data security sharing for the green power grid supply chain. Summary of the Invention

[0005] This application provides a method for cross-system data security sharing in a green power grid supply chain, aiming to achieve high-quality cross-system data security sharing in a green power grid supply chain.

[0006] A first aspect of this application provides a cross-system data security sharing method for a green power grid supply chain, comprising:

[0007] In response to the acquisition of the first candidate participants in the green power grid supply chain, the first candidate participants are clustered based on the network latency between them to obtain at least one first participant cluster.

[0008] Based on the resource information of each first candidate participant in the first participant cluster, the first aggregation coefficient of the first participant cluster is determined; the first aggregation coefficient is used to characterize the degree of aggregation of the first participant cluster.

[0009] Based on the first aggregation coefficient, the first candidate participants in the first participant cluster are screened out until the first aggregation coefficient meets the preset stopping condition, resulting in multiple target participants; the target participants are used to share data to train the federated learning model.

[0010] Furthermore, this application also proposes determining the first aggregation coefficient of the first participant cluster based on the resource information of each first candidate participant in the first participant cluster, including:

[0011] Using the first candidate participant as the node and the network topology between the first candidate participants as the edge, a participant structure graph is constructed.

[0012] Based on the resource information of each first candidate participant in the first participant cluster and the associated candidate participants, the center evaluation value of each first candidate participant is determined; the associated candidate participants are located in the same first participant cluster as the first candidate participants and are connected to the first candidate participants in the participant structure diagram; the center evaluation value is used to characterize the probability that the first candidate participant is the cluster center node of the first participant cluster.

[0013] The first aggregation coefficient of the first participant cluster is determined based on the central evaluation value of each first candidate participant in the first participant cluster.

[0014] Furthermore, this application also proposes that the resource information includes communication performance;

[0015] Based on the resource information of each first candidate participant in the first participant cluster and the associated candidate participants, the central evaluation value of each first candidate participant is determined, including:

[0016] The average communication performance of each first candidate participant in the first participant cluster is obtained by averaging the communication performance of the first participant cluster.

[0017] The average network latency between each associated candidate participant and the first candidate participant is calculated by averaging the network latency of the associated candidate participants.

[0018] The central evaluation value of the first candidate participant is determined based on the number of participants, average network latency, average communication performance of the associated candidate participants, and the communication performance of the first candidate participant.

[0019] Furthermore, this application also proposes determining the first aggregation coefficient of the first participant cluster based on the central evaluation value of each first candidate participant in the first participant cluster, including:

[0020] The first candidate participant with the highest center evaluation value in the first participant cluster is determined as the cluster center node of the first participant cluster;

[0021] The mean value of the difference between the central evaluation value of each first candidate participant and the central node of the first participant cluster is averaged to obtain the first average value.

[0022] Based on the first average value and the center evaluation value of the cluster center node, the first aggregation coefficient of the first participating party cluster is determined.

[0023] Furthermore, this application also proposes that, in response to obtaining the first candidate participants in the green power grid supply chain, before clustering the first candidate participants based on the network latency between them to obtain at least one cluster of first participants, the process further includes:

[0024] Based on the resource information of each participant in the green power grid supply chain, the busyness of each participant is determined.

[0025] Based on the busyness of each participant, the first candidate participant with a busyness level lower than the preset busyness threshold is selected from all participants.

[0026] Furthermore, this application also proposes that resource information includes real-time load and computing performance;

[0027] Based on resource information from each participant in the green power grid supply chain, the busyness level of each participant is determined, including:

[0028] Divide the real-time load of the participants by the computing performance to obtain the busy rating;

[0029] The busyness rating is normalized to obtain the busyness level of the participants.

[0030] Furthermore, this application also proposes that, before filtering out the first candidate participants in the first participant cluster based on the first aggregation coefficient until the first aggregation coefficient meets a preset stopping condition and multiple target participants are obtained, the process further includes:

[0031] Based on the resource information of each first candidate participant, the clusters of each first participant are updated to obtain the clusters of the second participants corresponding to the clusters of the first participants.

[0032] The second participating party cluster is compared with the first participating party cluster to determine the second aggregation coefficient of the second participating party cluster; the second aggregation coefficient is used to characterize the degree of aggregation of the second participating party cluster.

[0033] Based on the first aggregation coefficient, the first candidate participants in the first participant cluster are screened out until the first aggregation coefficient meets the preset stopping condition, resulting in multiple target participants, including:

[0034] Based on the second aggregation coefficient, the second candidate participants in the second participant cluster are screened out until the second aggregation coefficient meets the preset stopping condition, resulting in multiple target participants.

[0035] Furthermore, this application also proposes that the resource information includes computing performance and communication performance;

[0036] Based on the resource information of each first candidate participant, the clusters of each first participant are updated to obtain the second participant clusters corresponding to the first participant clusters, including:

[0037] The optimal weight of each first candidate participant is determined based on the busyness, computing performance, and communication performance of each first candidate participant.

[0038] Based on the preferred weights of each first candidate participant, the clusters of each first participant are updated to obtain the second participant clusters corresponding to the first participant clusters.

[0039] Furthermore, this application also proposes to compare the second participant cluster with the first participant cluster to determine the second aggregation coefficient of the second participant cluster, including:

[0040] Obtain the ratio of the number of second participants in the second participant cluster to the number of first participants in the first participant cluster;

[0041] The average value is obtained by averaging the ratio between the center evaluation value of each second candidate participant in the second participant cluster and the center evaluation value of the cluster center node of the first participant cluster.

[0042] The second aggregation coefficient of the second participant cluster is determined based on the ratio of the number of participants, the second average value, and the first aggregation coefficient of the first participant cluster.

[0043] Furthermore, this application also proposes to screen out second candidate participants in the second participant cluster based on a second aggregation coefficient until the second aggregation coefficient meets a preset stopping condition, thereby obtaining multiple target participants, including:

[0044] The third average value corresponding to the second aggregation coefficient of each second participant cluster is normalized to obtain the aggregation evaluation degree.

[0045] In response to a aggregation evaluation score greater than a preset aggregation threshold, the second candidate participant with the smallest weight in the target second participant cluster with the largest second aggregation coefficient is screened out until the aggregation evaluation score is no greater than the preset aggregation threshold, thus obtaining multiple target participants.

[0046] The present invention has the following beneficial effects:

[0047] The cross-system data security sharing method for a green power grid supply chain provided in this application firstly, in response to obtaining the first candidate participants in the green power grid supply chain, clustering is performed based on the network latency between the first candidate participants to obtain at least one first participant cluster. Network latency reflects the communication status between participants, and clustering can group participants with similar communication status into one category. Next, based on the resource information of each first candidate participant in the first participant cluster, a first aggregation coefficient of the first participant cluster is determined. This first aggregation coefficient can characterize the degree of aggregation of the first participant cluster. Finally, based on the first aggregation coefficient, the first candidate participants in the first participant cluster are screened out until the first aggregation coefficient meets a preset stopping condition, resulting in multiple target participants. In this way, the correlation between participants is fully considered, and the excessively high local aggregation is avoided through clustering and adjustment of the first aggregation coefficient. Therefore, the target participants can share data more effectively to train the federated learning model, ensuring the accuracy and stability of the federated learning model, thereby achieving high-quality cross-system data security sharing for a green power grid supply chain. Attached Figure Description

[0048] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a flowchart illustrating a cross-system data security sharing method for a green power grid supply chain, as provided in one embodiment of the present invention.

[0050] Figure 2 This is a schematic diagram illustrating the process of determining the first aggregation coefficient of a first participating party cluster according to an embodiment of the present invention;

[0051] Figure 3 This is a schematic diagram illustrating the process for determining a first candidate participant according to an embodiment of the present invention;

[0052] Figure 4 This is a schematic diagram illustrating the process of determining the second aggregation coefficient of a second participating party cluster, as provided in an embodiment of the present invention. Detailed Implementation

[0053] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a cross-system data security sharing method for a green power grid supply chain proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0055] In traditional power grid green supply chain data sharing processes, federated learning technology is used to achieve secure cross-system data sharing. However, the node selection mechanism makes decisions solely based on optimizing node performance and communication costs. This mechanism fails to consider the strong correlations among multiple participants in power grid data scenarios, leading to overfitting in the federated learning model during the training phase. These strong correlations manifest as high coupling between participants in terms of business logic, geographical location, or data characteristics, resulting in decreased model generalization ability and consequently affecting the reliability and accuracy of cross-system data sharing. Consequently, the tracking of green electricity attributes and carbon emission data cannot meet the stringent requirements of new power systems for green supply chain traceability.

[0056] For example, in the actual operation of a regional power grid's green supply chain, multiple wind farms, photovoltaic power plants, and power grid companies act as data-sharing participants, exhibiting a densely connected network topology. Furthermore, the real-time power generation and load data of each participant show significant time-series correlations. Moreover, when selecting federated learning nodes based on communication latency and computational performance, the system only focuses on the state of individual resources, failing to identify strong correlations between participants. Consequently, during the federated learning model training phase, the high similarity of participant data is erroneously amplified, model parameters overfit local data distributions, leading to cross-system data sharing results deviating from the true carbon flow state, systematic biases in green electricity attribute records, and direct damage to the traceability of carbon emission data.

[0057] If the above problems are not addressed, the overfitting phenomenon in federated learning models will continue to worsen, making it impossible to achieve visibility, manageability, and traceability of carbon flows in the cross-system data sharing process. Specifically, data consistency among various links in the power grid green supply chain will be disrupted, the matching of green electricity attributes between the power generation side and the grid side will fail, ultimately leading to the loss of green traceability capabilities in the supply chain and making it difficult to achieve the goals of building a new power system.

[0058] In this regard, such as Figure 1As shown, this application provides a flowchart illustrating a cross-system data security sharing method for a green power grid supply chain. This cross-system data security sharing method for a green power grid supply chain can be applied to electronic devices and may include the following steps S110 to S130:

[0059] S110, in response to obtaining the first candidate participants in the green power grid supply chain, cluster the first candidate participants based on the network latency between them to obtain at least one first participant cluster;

[0060] S120, Based on the resource information of each first candidate participant in the first participant cluster, determine the first aggregation coefficient of the first participant cluster; the first aggregation coefficient is used to characterize the degree of aggregation of the first participant cluster;

[0061] S130, based on the first aggregation coefficient, the first candidate participants in the first participant cluster are screened out until the first aggregation coefficient meets the preset stopping condition, resulting in multiple target participants; the target participants are used to share data to train the federated learning model.

[0062] For ease of understanding, the following explains some key terms in this embodiment:

[0063] A green power grid supply chain refers to a collaborative network formed by various parties involved in the production, transmission, distribution, and consumption of electricity, including equipment suppliers, grid operators, and electricity users, with the goal of achieving a green and low-carbon energy transition. Its core lies in tracking and managing the green attributes and carbon emission data throughout the entire lifecycle of electricity.

[0064] First-line candidate participants refer to entities or nodes in the green power grid supply chain that have the potential to participate in federated learning data sharing, such as power plants, transmission companies, distribution companies, and large electricity consumers. These participants possess the data that needs to be shared and may participate in the training of the federated learning model.

[0065] Network latency refers to the time required for a data packet to travel from its source to its destination. In distributed systems, network latency is a key factor affecting data transmission efficiency and system performance.

[0066] Clustering is the process of grouping a set of objects into several groups (clusters) based on their similarity. In data processing, clustering can help identify the inherent structure or patterns in data.

[0067] First participant cluster: refers to a set of first candidate participants with similar network latency characteristics, obtained by clustering the first candidate participants based on network latency.

[0068] Resource information: refers to various types of data possessed by the first candidate participant that can be used to assess its ability and contribution to federated learning, such as computing power, storage capacity, communication bandwidth, and data volume.

[0069] First aggregation coefficient: This refers to an indicator used to quantitatively characterize the degree of aggregation within a cluster of participants with the first aggregation factor. This first aggregation coefficient can reflect the closeness or similarity between participants within the cluster.

[0070] Preset stopping conditions: These are predetermined criteria used to determine whether the screening process should be terminated during the screening of the first candidate participants. For example, the screening process stops when the first aggregation coefficient reaches a certain threshold or the number of participants within the cluster reaches its minimum value.

[0071] Target participants: These are the participants selected after clustering and screening processes to share data and train the federated learning model. These participants are considered the most suitable set of nodes for participating in federated learning.

[0072] Federated learning models refer to a distributed machine learning paradigm that allows multiple participants to collaboratively train a global model without directly sharing the original data. Each participant trains its local model using its private data and then sends the model updates (rather than the original data) to an aggregation server for aggregation, thus protecting data privacy.

[0073] For ease of understanding, the specific implementation steps in this embodiment are explained below:

[0074] First, in response to the acquisition of primary candidate participants in the green power grid supply chain, these primary candidate participants are clustered based on the network latency between them, resulting in at least one primary participant cluster. In the green power grid supply chain, it is first necessary to acquire a group of potential primary candidate participants capable of participating in data sharing. These primary candidate participants can be different entities within the power grid, such as power plants, transmission stations, distribution stations, and large industrial users. These participants can be acquired through manual registration, automatic system discovery, or pre-configuration. For example, system administrators can manually enter all power grid entities that meet the basic criteria as primary candidate participants. After acquiring these primary candidate participants, it is necessary to assess their network connectivity. One implementation method is to acquire network latency data by sending probe packets between the primary candidate participants and measuring their round-trip time (RTT). For example, network connectivity tests can be periodically performed between the primary candidate participants, recording and calculating the average network latency. Subsequently, based on these average network latency, the primary candidate participants are clustered. A variety of clustering algorithms can be chosen. For example, the K-Means algorithm can be used, which uses network latency as a distance metric to group participants with similar network latency into one category.

[0075] Secondly, based on the resource information of each first candidate participant in the first participant cluster, a first aggregation coefficient is determined for the first participant cluster; this first aggregation coefficient is used to characterize the degree of aggregation of the first participant cluster. After obtaining the first participant clusters, it is necessary to further evaluate the internal aggregation degree of each first participant cluster. This can be achieved by analyzing the resource information of each first candidate participant within the cluster. Resource information may include, but is not limited to, computing power, storage capacity, and communication bandwidth. For example, data such as the number of CPU cores, memory size, hard disk space, and network interface speed of each first candidate participant can be collected. There are several ways to determine the first aggregation coefficient. One implementation is that, for each first participant cluster, the average or median of the resource information of all first candidate participants within the cluster can be calculated, and this can be used as a preliminary measure of the aggregation degree of the cluster. For example, the average computing power and average communication bandwidth of all first candidate participants within the cluster can be calculated. Furthermore, these average resource information can be weighted and summed to obtain a comprehensive value, which serves as the first aggregation coefficient of the first participant cluster. The higher the first aggregation coefficient, the more similar or superior the participants in the cluster are in terms of resource information, thus characterizing the higher the degree of aggregation of the cluster.

[0076] Finally, based on the first aggregation coefficient, the first candidate participants in the first participant cluster are screened out until the first aggregation coefficient meets a preset stopping condition, resulting in multiple target participants. These target participants are used to share data to train the federated learning model. After determining the first aggregation coefficient for each first participant cluster, the first candidate participants within the cluster need to be screened out based on these coefficients to optimize the set of nodes participating in federated learning. The purpose of screening is to reduce the degree of aggregation within the cluster. The screening process can be iterative. For example, a preset stopping condition can be set, such as stopping screening when the first aggregation coefficient falls below a certain threshold or when the number of remaining participants in the cluster reaches a preset minimum. In each iteration, the first candidate participant that contributes the most to the first aggregation coefficient of the current cluster or has the worst resource information can be identified. For example, the participant with the lowest computing power or the narrowest communication bandwidth in the cluster can be identified. Subsequently, the identified first candidate participant is removed from the cluster. After removal, the first aggregation coefficient of the cluster needs to be recalculated, and the preset stopping condition needs to be checked again. This process is repeated until the preset stopping condition is met. Ultimately, the remaining first-choice participants in each cluster are identified as the target participants. These target participants will be used for subsequent federated learning model training; they are considered to have good synergy in terms of network latency and resource information, enabling them to support high-quality federated learning tasks.

[0077] The following example will provide a more detailed explanation of the above technical solution:

[0078] Suppose a green power grid supply chain scenario includes multiple entities: power plant A, power plant B, transmission company C, distribution company D, and electricity consumer companies E and F. These entities all wish to participate in a federated learning task to jointly train a federated learning model for predicting green electricity consumption or carbon emissions.

[0079] First, the system responds by acquiring these entities as first-line candidate participants. For example, through a pre-defined registration mechanism, the system identifies power plants A and B, transmission company C, distribution company D, and electricity consumers E and F as first-line candidate participants. Next, the system clusters these first-line candidate participants based on the network latency between them. For instance, the system continuously probes the network, measuring and recording the pairwise network latency between A, B, C, D, E, and F. The system can use the K-Means algorithm to cluster these network latency data as a distance metric. Assume that after clustering, two first-line participant clusters are obtained: Cluster 1: containing power plants A and B; Cluster 2: containing transmission company C, distribution company D, large electricity consumer E, and large electricity consumer F. In this way, participants with similar network latency are grouped together.

[0080] Subsequently, the system determines the first aggregation coefficient of each cluster based on the resource information of the first candidate participants in each first participant cluster. For example, for cluster 1 (A, B), the system collects the computing power (e.g., number of CPU cores) and communication bandwidth of A and B. Assume that A has 8 computing cores and a bandwidth of 1Gbps; B has 6 computing cores and a bandwidth of 800Mbps. The system can simply use the sum of the computing power of all participants within the cluster as the first aggregation coefficient of the cluster. Therefore, the first aggregation coefficient of cluster 1 is 8 + 6 = 14. For cluster 2 (C, D, E, F), the system also collects their resource information and calculates the first aggregation coefficient of cluster 2 in the same way. This first aggregation coefficient is used to characterize the degree of aggregation of participants within the cluster in terms of resource information.

[0081] Finally, based on the first aggregation coefficient, the system filters out the first candidate participants in the first participant cluster until the first aggregation coefficient meets a preset stopping condition, resulting in multiple target participants. For example, the system sets the preset stopping condition as follows: filtering stops when the number of participants in a cluster is less than 2, or filtering stops when the first aggregation coefficient is lower than a certain threshold. For cluster 1, its first aggregation coefficient is 14. Assuming the system determines that the first aggregation coefficient is high, but the number of participants in the cluster is 2, reaching the preset minimum value, A and B in cluster 1 are retained as target participants. For cluster 2, assuming its initial first aggregation coefficient is a certain value, the system can iteratively filter out the participant with the worst resource information in the cluster. For example, if the computing power and communication bandwidth of electricity company F are the lowest in cluster 2, the system will filter it out first. After filtering out F, the first aggregation coefficient of cluster 2 will be recalculated. If the first aggregation coefficient is still high at this time, and the number of participants in the cluster is still greater than the preset minimum value, the system may continue to filter out the next participant with the worst resource information, such as electricity company E. This process continues until the first aggregation coefficient of cluster 2 reaches a preset stopping condition, or the number of remaining participants in the cluster reaches a preset minimum value. Ultimately, only transmission company C and distribution company D may remain as target participants in cluster 2. Through this filtering mechanism, an optimized set of target participants (e.g., A, B, C, D) is obtained, enabling more efficient and high-quality data sharing for training the federated learning model.

[0082] In this embodiment, firstly, in response to obtaining the first candidate participants in the green power grid supply chain, clustering is performed based on the network latency among the first candidate participants to obtain at least one first participant cluster. Network latency reflects the communication status between participants, and clustering can group participants with similar communication status into one category. Next, based on the resource information of each first candidate participant in the first participant cluster, a first aggregation coefficient is determined for the first participant cluster. This first aggregation coefficient characterizes the degree of aggregation of the first participant cluster. Finally, based on the first aggregation coefficient, the first candidate participants in the first participant cluster are screened out until the first aggregation coefficient meets a preset stopping condition, resulting in multiple target participants. In this way, the correlation between participants is fully considered, and the excessively high local aggregation is avoided through clustering and adjustment of the first aggregation coefficient. Therefore, target participants can more effectively share data to train the federated learning model, ensuring the accuracy and stability of the federated learning model, thereby achieving high-quality cross-system data security sharing for the green power grid supply chain.

[0083] In some embodiments described above, a first aggregation coefficient is proposed to be determined based on the resource information of each first candidate participant in the first participant cluster, in order to characterize the degree of aggregation of the cluster. However, in practical applications, relying solely on resource information may not be sufficient to capture the complex network relationships between members within the cluster, resulting in an inaccurate assessment of the degree of aggregation of the cluster and affecting the subsequent screening effect of participants.

[0084] In this regard, such as Figure 2 As shown, this application further proposes that S120 includes the following S121 to S123:

[0085] S121, using the first candidate participant as the node and the network topology between the first candidate participants as the edge, construct the participant structure graph;

[0086] S122, Based on the resource information of each first candidate participant in the first participant cluster and the associated candidate participants, determine the center evaluation value of each first candidate participant; the associated candidate participants are located in the same first participant cluster as the first candidate participants and are connected to the first candidate participants in the participant structure diagram; the center evaluation value is used to characterize the possibility of the first candidate participant being the cluster center node of the first participant cluster.

[0087] S123, Based on the central evaluation value of each first candidate participant in the first participant cluster, determine the first aggregation coefficient of the first participant cluster.

[0088] In this embodiment, a participant structure graph is constructed by taking the first candidate participant as a node and the network topology between the first candidate participants as edges. This means that the first candidate participants in the green power grid supply chain are abstracted as nodes in the graph, and edges are established between these nodes according to their actual network connections or logical communication relationships, thereby forming a graph structure that can reflect the interconnection between the participants.

[0089] The centrality evaluation value of each first-candidate participant is determined based on its resource information and associated candidate participants within the first-candidate participant cluster. This involves comprehensively considering the first-candidate participant's own resource status and its connections with other cluster members in the network structure to quantify its importance or central position within the cluster. Associated candidate participants specifically refer to participants located within the same first-candidate participant cluster and directly connected to it in the constructed participant structure graph. A higher centrality evaluation value indicates a greater likelihood that the first-candidate participant will become the cluster center node of its respective cluster. For example, the centrality evaluation value can be obtained by weighting a first-candidate participant's resource information (such as computing power and communication bandwidth) with the number of associated candidate participants and the network latency or connection strength between them. Better resource information, more associated candidate participants, and lower network latency result in a higher centrality evaluation value.

[0090] The first aggregation coefficient of a cluster is determined based on the central evaluation values ​​of all first candidate participants within the first participant cluster. This means that after calculating the central evaluation values ​​of all first candidate participants within the cluster, a certain aggregation method is used to obtain an index that can reflect the overall aggregation degree of the cluster. For example, the central evaluation values ​​of all first candidate participants in the first participant cluster can be statistically processed, such as calculating the average, median, or weighted average, and used as the first aggregation coefficient of the cluster.

[0091] This application's solution first constructs a participant structure diagram, clarifying the network topology relationships among the first-candidate participants. Based on this, when determining the central evaluation value of each first-candidate participant, it considers not only its own resource information but also the network connectivity relationships between it and related candidate participants within the same cluster. This approach allows the central evaluation value to more comprehensively reflect a participant's actual influence within the cluster and its potential as a cluster center node. Finally, based on these more representative central evaluation values, the first aggregation coefficient is determined, enabling this coefficient to more accurately characterize the aggregation degree of the first-candidate participant's cluster, thus overcoming the evaluation bias that may result from relying solely on resource information. This method integrates network structure information into the calculation of the aggregation coefficient, making the evaluation of the cluster aggregation degree more refined and reasonable, and providing a more reliable basis for subsequent participant selection.

[0092] The following is a concrete example. Assume a cluster of first-participant participants contains three first-participant candidates, P1, P2, and P3. First, construct a participant structure graph based on the network connections between P1, P2, and P3. For example, if P1 has direct network connections to both P2 and P3, but P2 and P3 do not, then there is an edge between P1 and P2, and another edge between P1 and P3. Next, determine the center evaluation value for each first-participant candidate. For example, the center evaluation value for P1 can be calculated based on its own communication performance, computing performance, and other resource information, as well as factors such as network latency and connection stability between P1 and its associated candidate participants P2 and P3. The center evaluation values ​​for P2 and P3 are calculated in a similar manner. Finally, based on the center evaluation values ​​of P1, P2, and P3, the first aggregation coefficient of the cluster can be calculated. For example, the average of all center evaluation values ​​can be taken as the first aggregation coefficient.

[0093] By employing the aforementioned technical solution, the determination of the first aggregation coefficient for the first participating party's cluster can fully consider the network topology and interrelationships among the first candidate participating parties. This allows the determined first aggregation coefficient to more accurately and comprehensively reflect the degree of aggregation of the cluster. This helps to more effectively identify participating parties that not only possess superior resources but also exhibit good connectivity and a central position in the network structure during subsequent screening processes. Consequently, this improves the efficiency of data security sharing and the quality of federated learning model training in the power grid's green supply chain.

[0094] In some embodiments described above in this application, a method is proposed to determine the central evaluation value of each first candidate participant based on the resource information of each first candidate participant in the first participant cluster and the associated candidate participants, thereby determining the first aggregation coefficient. However, in practical applications, if the "resource information" is not specifically defined and the method for comprehensively considering the performance of the participant and its connectivity within the cluster is not explained in detail, the calculation of the central evaluation value may be inaccurate, thus affecting the accuracy of the first aggregation coefficient. This, in turn, affects the selection effect of the target participant, resulting in poor efficiency and stability of the finally selected participant in data sharing and federated learning model training.

[0095] In this regard, this application further proposes resource information including communication performance; S122 includes:

[0096] The average communication performance of each first candidate participant in the first participant cluster is obtained by averaging the communication performance of the first participant cluster.

[0097] The average network latency between each associated candidate participant and the first candidate participant is calculated by averaging the network latency of the associated candidate participants.

[0098] The central evaluation value of the first candidate participant is determined based on the number of participants, average network latency, average communication performance of the associated candidate participants, and the communication performance of the first candidate participant.

[0099] In this embodiment, communication performance refers to the ability of participants to transmit and interact within a network environment, characterizing the efficiency and reliability of data sharing. Communication performance may include, but is not limited to, metrics such as network bandwidth, data transmission rate, network throughput, and packet loss rate. For example, communication performance may refer to the maximum amount of data a participant can transmit within a specific time period, or the time required for data to be transmitted from one participant to another. By incorporating communication performance into resource information, the potential of participants in data sharing can be assessed more accurately.

[0100] The communication performance of each first-candidate participant in the first-participant cluster is averaged to obtain the average communication performance of the cluster, aiming to obtain the overall level of communication performance of all first-candidate participants within the cluster. This method can eliminate the impact of extreme fluctuations in the communication performance of individual participants on the overall evaluation, providing a more representative indicator of communication capability. The averaging process can employ various methods such as arithmetic mean, weighted mean, or geometric mean. For example, the arithmetic mean of the communication performance of all first-candidate participants can be calculated to reflect the average data transmission capability of the cluster.

[0101] The average network latency between each associated candidate participant and the first candidate participant is calculated by averaging the network latency of the associated candidate participants. This average network latency aims to quantify the network connection quality between a specific first candidate participant and its directly connected associated candidate participants. Network latency is the time required for data to travel from the source to the destination and is a key indicator of network response speed. By averaging the network latency among the associated candidate participants connected to a specific first candidate participant, the average communication response speed of that first candidate participant in the local network environment can be obtained. Averaging can be performed using methods such as the arithmetic mean or the median to reflect the average communication efficiency between the first candidate participant and its neighboring nodes.

[0102] The determination of the centrality evaluation value of the first candidate participant, based on the number of participants in the associated candidate participants, average network latency, average communication performance, and the communication performance of the first candidate participant, is a comprehensive evaluation process aimed at assessing the suitability of the first candidate participant as the cluster central node. Specifically, the number of participants in the associated candidate participants reflects the connection breadth or local network density of the first candidate participant; average network latency reflects its average communication efficiency with neighboring nodes; average communication performance reflects the overall communication capability of the entire cluster; and the communication performance of the first candidate participant itself reflects its individual capability as a potential central node. By integrating these factors, a multi-dimensional evaluation model can be constructed to ensure that the evaluation value can comprehensively and accurately characterize the likelihood of the first candidate participant serving as the cluster central node.

[0103] This application's scheme, when determining the first aggregation coefficient of the first participant's cluster, introduces communication performance as resource information and refines the calculation method of the center evaluation value, making the assessment of the probability of the first candidate participant as the cluster center node more accurate. Specifically, after constructing the participant structure graph, in order to more accurately evaluate the center evaluation value of each first candidate participant, this scheme first clarifies that the resource information includes communication performance. Subsequently, by averaging the communication performance of all first candidate participants within the first participant's cluster, the average communication performance of the cluster can be obtained, which provides a benchmark for evaluating the performance of an individual participant in the overall environment. At the same time, for each first candidate participant, the network latency between it and its directly connected associated candidate participants is averaged to obtain the average network latency of the first candidate participant, which directly reflects the quality of its local connectivity. Finally, these refined indicators, namely the number of participants in the associated candidate participants, the average network latency, the average communication performance of the cluster, and the communication performance of the first candidate participant itself, are combined to determine the center evaluation value of the first candidate participant. This comprehensive approach considers not only the communication capabilities of each participant but also their connectivity within the cluster and the average level of the entire cluster, making the calculation of the center evaluation value more comprehensive and objective. In this way, participants with superior communication capabilities and network connectivity, who are more suitable as cluster center nodes, can be identified more accurately. This allows the subsequent first aggregation coefficient determined based on the center evaluation value to more accurately reflect the degree of cluster aggregation, thereby optimizing the selection process for target participants and ensuring that the ultimately selected participants have higher efficiency and stability in data sharing and federated learning model training.

[0104] Specifically, to overcome potential dimension issues in calculating the central evaluation value of the first candidate participant, a standardization process is introduced during data processing. Specifically, for indicators with different dimensions involved in the central evaluation value calculation, such as the number of participants, network latency, and communication performance, these indicators are standardized and converted into dimensionless values ​​before calculation. For example, the Z-score standardization method can be used to convert the value of each indicator into a value with a mean of 0 and a standard deviation of 1.

[0105] Therefore, the central evaluation value of the first candidate participant can be determined by the following formula 1:

[0106] Formula 1

[0107] In formula 1, The center evaluation value is used to characterize the t-th participant in the m-th first participant cluster. The result after standardization of the number of participants used to characterize the associated candidate participants connected to the t-th participant in the m-th first participant cluster. The result of normalizing the average network latency of the associated candidate participants connected to the t-th first candidate participant in the m-th first participant cluster. The result after normalization is used to characterize the average communication performance of the m-th first-participant cluster. The result after normalization is used to characterize the communication performance of the t-th first candidate participant in the m-th first participant cluster.

[0108] It should be noted that, in order to ensure that the calculation results are meaningful, when performing fractional operations in this embodiment of the invention, if the denominator is 0, a parameter adjustment factor greater than 0 needs to be added to the denominator before summing to prevent the denominator from being 0. The value of the parameter adjustment factor can be set by the implementer according to the actual situation, and this invention does not impose any special restrictions.

[0109] in, A larger value indicates that the t-th first candidate participant has more associated candidate participants and a smaller average distance, meaning that the t-th first candidate participant in the m-th first participant cluster is more closely connected to other nodes in the m-th first participant cluster. The strength of the performance of the t-th first candidate participant in the cluster of the m-th first participant is reflected in the cluster. The larger the scale of information exchange that the central node needs to handle, the stronger its communication performance.

[0110] As a specific implementation, when determining the center evaluation value of the first candidate participant, the communication performance can be set as the average network bandwidth of the participants, and the network latency as the round-trip time of data packets between the participants. First, the system can collect the average network bandwidth data of all first candidate participants in the first participant cluster and calculate its arithmetic mean as the average communication performance of the cluster. For example, if there are three participants A, B, and C in the cluster, with bandwidths of 100Mbps, 80Mbps, and 120Mbps respectively, then the average communication performance is 100Mbps. Next, for a specific first candidate participant, such as participant A, the system can measure the network latency between it and all associated candidate participants (e.g., B and C, which are directly connected to A), and calculate the arithmetic mean of these latency as the average network latency of the associated candidate participants. For example, if the latency from A to B is 20ms and the latency from A to C is 30ms, then the average network latency of A is 25ms. Simultaneously, the number of participants in the associated candidate participants is 2. Based on this, the center evaluation value of participant A can be determined using the above formula 2. In this way, a quantified central evaluation value can be obtained, which can be used for the subsequent calculation of the first aggregation coefficient.

[0111] The above technical solution clarifies that communication performance is included in resource information and refines the calculation process of the center evaluation value, making the assessment of the likelihood of the first candidate participant as a cluster center node more accurate and comprehensive. Specifically, by comprehensively considering the communication performance of the first candidate participant itself, the average network latency between it and associated candidate participants, the number of associated candidate participants, and the average communication performance of the entire cluster, the actual performance and connection quality of the participants in the network environment can be more accurately reflected. This helps to more effectively identify potential center nodes with good communication capabilities and stable network connections, so that the first aggregation coefficient determined based on these center evaluation values ​​can more realistically represent the degree of aggregation of the cluster. Finally, in the process of selecting target participants, participants more suitable for data sharing and federated learning model training can be selected, effectively improving the training efficiency of the federated learning model and the stability of data sharing, and reducing the risk of training interruption or data transmission failure due to poor network performance or unstable connections.

[0112] In some embodiments described above in this application, a first aggregation coefficient for a first participant cluster is determined based on the central evaluation value of each first candidate participant in the first participant cluster. However, after obtaining only the central evaluation value of each first candidate participant, there may still be operational ambiguity in how to systematically and accurately transform these discrete central evaluation values ​​into a "first aggregation coefficient" that can comprehensively reflect the aggregation degree of the entire cluster, thus affecting the accurate evaluation of the cluster quality.

[0113] In this regard, this application further proposes that S123 includes:

[0114] The first candidate participant with the highest center evaluation value in the first participant cluster is determined as the cluster center node of the first participant cluster;

[0115] The mean value of the difference between the central evaluation value of each first candidate participant and the central node of the first participant cluster is averaged to obtain the first average value.

[0116] Based on the first average value and the center evaluation value of the cluster center node, the first aggregation coefficient of the first participating party cluster is determined.

[0117] In this embodiment, determining the first candidate participant with the highest center evaluation value in the first participant cluster as the cluster center node means that within the formed cluster, by comparing the center evaluation values ​​of all first candidate participants, the participant with the highest center evaluation value is selected as the representative of that cluster. This cluster center node is the most representative and core participant in the cluster, and its center evaluation value reflects its probability of becoming the cluster center. For example, the cluster center node can be determined by traversing all first candidate participants in the cluster, comparing their center evaluation values ​​one by one, and recording the maximum value and its corresponding participant. Alternatively, all first candidate participants within the cluster can be pre-sorted in descending order of their center evaluation values, and then the first participant in the sorted order can be directly selected as the cluster center node.

[0118] Subsequently, the mean values ​​of the centrality evaluation values ​​between each first candidate participant and the cluster center node in the first participant cluster are averaged to obtain the first average value. This step aims to quantify the average level of the "distance" or "similarity" between other participants within the cluster and the selected cluster center node. The centrality evaluation value difference reflects the degree of deviation of each participant from the cluster center node in terms of "centrality." By averaging these differences, an overall index reflecting the degree of dispersion within the cluster can be obtained. For example, the absolute difference between the centrality evaluation value of each first candidate participant (including the cluster center node itself, whose difference is zero) and the centrality evaluation value of the cluster center node can be calculated. Then, all these absolute differences are summed and divided by the total number of first candidate participants in the cluster to obtain the first average value.

[0119] Finally, based on the first average value and the center evaluation value of the cluster center node, the first aggregation coefficient of the first participating party cluster is determined. This step combines the aforementioned two indicators to form a single, comprehensive quantitative indicator that can characterize the degree of cluster aggregation. Specifically, the first aggregation coefficient of the first participating party cluster can be determined using the following formula 2:

[0120] Formula 2

[0121] In formula 2, The first aggregation coefficient is used to characterize the m-th first participant cluster. The center evaluation value used to characterize the cluster center node of the m-th first participant cluster. The center evaluation value is used to characterize the t-th participant in the m-th first participant cluster. Used to characterize the number of first candidate participants in the m-th first participant cluster. Used to characterize the normalization function.

[0122] in, This indicates the prominence of the central evaluation value of the cluster center node within the entire cluster. The larger the value, the more prominent its central evaluation, and the greater the degree of cluster aggregation. Conversely, the smaller the value, the more similar the distribution of central evaluation values ​​among the nodes in the cluster may be, and the less the degree of cluster aggregation.

[0123] Through the above technical solution, this application provides a clear and quantifiable method for calculating the first aggregation coefficient. This makes the assessment of the aggregation degree of the first participant cluster no longer dependent on vague judgments, but based on objective data and clear calculation logic. This precise quantification capability significantly improves the accuracy and efficiency of subsequent screening of the first participant cluster, ensuring that the finally selected target participants can form a highly aggregated, high-performance federated learning model training group, thereby optimizing the data security sharing effect in the power grid green supply chain.

[0124] In some of the embodiments described above in this application, in the cross-system data security sharing method for a green power grid supply chain, directly clustering all potential participants based on network latency may result in the selection of nodes with excessively high resource loads or insufficient performance among the first candidate participants. These nodes may not be able to effectively participate in the subsequent federated learning model training, thereby affecting the efficiency and stability of model training, and even causing training interruptions or model performance degradation.

[0125] In this regard, such as Figure 3 As shown, this application further proposes that, prior to S110, it also includes the following S310 to S320:

[0126] S310 determines the busyness of each participant based on resource information of each participant in the green power grid supply chain.

[0127] S320: Based on the busyness of each participant, select the first candidate participant whose busyness is less than the preset busyness threshold from among the participants.

[0128] In this embodiment, the resource information of each participant in the green power grid supply chain refers to the computing, storage, and networking capabilities possessed by each participant that can be used for federated learning tasks. This information may include, but is not limited to, the participant's CPU utilization, memory usage, disk I / O, network bandwidth, and GPU availability. By acquiring this resource information, the current operating status and processing capabilities of the participants can be comprehensively assessed.

[0129] Determining the workload level of each participant is an indicator of their current resource utilization and workload. Workload level can be calculated by comprehensively considering various resource information from each participant; for example, it can be achieved through a weighted average of CPU utilization, memory usage, and network bandwidth utilization, or by mapping these metrics using a specific function. Higher workload level indicates that the participant is currently handling more tasks, and therefore has fewer resources available for federated learning.

[0130] The preset busyness threshold is a pre-defined value used to determine whether a participant is suitable to participate in the federated learning task. This threshold can be set in various ways, such as based on the resource requirements of the federated learning task, the overall resource status of the green power grid supply chain, historical data analysis, or expert experience. For example, it can be set as a percentage cap on resource utilization or an upper limit on a comprehensive busyness score. For instance, the preset busyness threshold could be 0.5.

[0131] Selecting the first candidate participants whose busyness is less than a preset busyness threshold means that the system selects nodes with low current busyness and sufficient resources to participate in federated learning from all potential participants based on the preset busyness threshold. In practice, the system can iterate through all participants, calculate their busyness, and then mark participants with busyness lower than or equal to the preset threshold as qualified "first candidate participants", thus forming a higher-quality participant pool.

[0132] Before clustering the first candidate participants, this application's solution first assesses the resource information of each participant in the power grid green supply chain. Specifically, the system acquires the resource information of each participant and calculates their workload based on this information. Workload, as a key indicator of a participant's current workload and available resources, reflects their suitability for participating in the federated learning task. Subsequently, the system compares the calculated workload with a preset workload threshold, selecting only those participants with workloads below the threshold as "first candidate participants" for subsequent clustering and federated learning. This pre-screening mechanism ensures that participants entering the clustering stage have sufficient resources to support the federated learning task, avoiding the inclusion of resource-constrained or underperforming nodes in the federated learning process, thus providing a more reliable foundation for subsequent network latency-based clustering. This mechanism allows subsequent federated learning model training to be conducted within a group of more capable and stable participants, significantly improving the efficiency of federated learning and the model's convergence speed.

[0133] As a specific implementation method, in the green power grid supply chain, each participant can periodically report its resource information, such as CPU utilization, memory usage, and network bandwidth utilization, to the central coordination node. Upon receiving this information, the central coordination node can calculate the busyness of each participant using a weighted average method. For example, CPU utilization can be weighted at 0.5, memory usage at 0.3, and network bandwidth utilization at 0.2, and these weighted values ​​can be summed to obtain a comprehensive busyness score. Simultaneously, a busyness threshold can be preset, for example, setting the comprehensive busyness score to not exceed 0.7. When selecting the first candidate participant, the central coordination node will iterate through all participants, calculate their busyness, and filter out those with busyness scores below 0.7. These filtered participants, due to their lower current resource load, are considered more suitable for participating in the federated learning task, thus forming a high-quality set of first candidate participants, laying a good foundation for subsequent federated learning model training.

[0134] Using the above technical solution, before clustering the first candidate participants, the busyness of each participant is determined based on their resource information, and participants with busyness levels below a preset busyness threshold are selected as first candidate participants. This pre-screening process effectively avoids including nodes with excessively high resource loads or insufficient performance in the federated learning task, thereby significantly improving the efficiency and stability of subsequent federated learning model training. Since all participants have sufficient resources, data sharing and model updates can be performed more smoothly, reducing training interruptions or delays caused by resource bottlenecks and ensuring the reliability of the federated learning process and the quality of model convergence.

[0135] In some embodiments described above, this application proposes determining the busyness of each participant based on their resource information and selecting first-candidate participants according to their busyness, thereby optimizing subsequent clustering and federated learning model training processes. However, in practice, accurately and effectively quantifying the busyness of each participant to ensure that the selected first-candidate participants are truly in a suitable state to participate in federated learning is a problem that needs to be solved. If the busyness assessment is inaccurate, it may lead to the system misjudging the availability of participants, thus affecting the efficiency of federated learning and the quality of model training.

[0136] In this regard, this application further proposes resource information including real-time load and computing performance; S310 includes:

[0137] Divide the real-time load of the participants by the computing performance to obtain the busy rating;

[0138] The busyness rating is normalized to obtain the busyness level of the participants.

[0139] In this embodiment, real-time load refers to the amount of tasks or resources a participant is processing at a given moment. This may include, but is not limited to, CPU utilization, memory usage, disk I / O rate, and network bandwidth usage. Real-time load is a key indicator for measuring a participant's current working status, reflecting its remaining capacity to handle new tasks. Computing performance refers to the hardware or software processing capabilities of a participant. This may include, but is not limited to, CPU clock speed, number of cores, memory size, GPU computing power, and storage capacity. Computing performance is a fundamental indicator for measuring a participant's potential processing capabilities, reflecting the maximum workload it can handle under ideal conditions.

[0140] The step of dividing a participant's real-time load by its computing performance to obtain a busy rating aims to quantify its current workload by comparing its real-time workload with its maximum processing capacity. For example, this could be achieved by dividing CPU utilization (real-time load) by maximum CPU processing capacity (computing performance), or by dividing current memory usage by total memory capacity. This ratio intuitively reflects the proportion of a participant's resources being utilized; a higher value indicates a busier participant.

[0141] The step of normalizing the busyness evaluation values ​​to obtain the busyness level of each participant aims to transform the busyness evaluation values ​​of different participants or different types of participants onto a unified and comparable scale, such as the [0, 1] interval. This helps to eliminate the impact of differences in hardware configuration among different participants, making the busyness level universal and comparable. Common normalization methods include min-max normalization or Z-score normalization. Through normalization, the busyness level of each participant can be assessed more fairly and accurately, providing a reliable basis for subsequent screening.

[0142] This application proposes a method for quantifying busyness by introducing real-time load and computing performance as key indicators for measuring the resource information of participating parties. Specifically, it first acquires real-time load and computing performance data for each participant in the power grid green supply chain. Real-time load reflects the resources consumed by a participant in its current tasks, while computing performance represents the maximum processing capacity a participant can provide. By comparing real-time load and computing performance, a preliminary busyness assessment value is obtained, which intuitively represents the proportion of resources occupied by each participant. To ensure the comparability of busyness assessment values ​​among different participants and to eliminate evaluation biases that may be caused by differences in hardware configuration, this application further normalizes the busyness assessment value to obtain a uniform scale of busyness. This busyness accurately reflects the current available resource status of each participant, providing a reliable quantitative basis for subsequent selection of first-candidate participants based on busyness. In this way, it can be ensured that only those participants with relatively free resources and who can effectively participate in the federated learning task are selected as the first candidate participants, thereby improving the efficiency and stability of federated learning model training and avoiding training interruptions or performance degradation caused by insufficient participant resources.

[0143] Specifically, the busyness level of the participants can be determined using the following formula 3:

[0144] Formula 3

[0145] In formula 3, Used to characterize the busyness of the i-th participant. Used to characterize the real-time load of the i-th participant. Used to characterize the computational performance of the i-th participant; Used to characterize the Sigmoid function, here it is used for proportional normalization.

[0146] The following is a concrete example. For each participant in the green power grid supply chain, the system can periodically collect its current CPU utilization as real-time load and obtain its maximum CPU processing capacity (e.g., obtained through benchmarking or hardware specifications) as computing performance. For example, a participant's current CPU utilization is 80%, and its maximum processing capacity is 100%. In this case, 80% can be divided by 100% to obtain a busy rating of 0.8. To convert this rating into a uniform busy level, a minimum-maximum normalization method can be used. Assuming that the busy ratings of all participants range from 0.1 to 0.9, then 0.8 can be subtracted from the minimum value of 0.1, and then divided by the difference between the maximum and minimum values ​​(0.9-0.1), i.e., (0.8-0.1) / (0.9-0.1) = 0.7 / 0.8 = 0.875. Thus, the busy level of this participant is determined to be 0.875. In this way, even if the CPU models and performance of different participants differ, their workload can be compared on a uniform scale, thus providing a fair basis for subsequent selection.

[0147] Through the aforementioned technical solution, this application can more accurately assess the workload of each participant in the green power grid supply chain. By combining real-time load with computing performance and performing normalization, the workload assessment considers not only the current working status of the participants but also their inherent processing capabilities, thus avoiding the bias that may arise from a single indicator. This accurate workload assessment enables more effective identification of nodes with sufficient resources and the ability to stably participate in the federated learning task when selecting first-line candidate participants, thereby improving the accuracy of pre-screening. This helps reduce the risk of federated learning task interruptions or performance degradation due to the selection of participants with insufficient resources, thereby improving the overall efficiency and reliability of federated learning model training.

[0148] In some embodiments of this application described above, the first participant cluster is screened out using a first aggregation coefficient to obtain the target participants. However, relying solely on the initial first aggregation coefficient for screening may result in poor accuracy, leading to poor performance or efficiency of the screened target participants in subsequent federated learning model training, or failure to fully utilize existing resources, thus affecting the final data sharing and model training results.

[0149] In this regard, such as Figure 4 As shown, this application further proposes that, prior to S130, it also includes the following S410 to S420:

[0150] S410, based on the resource information of each first candidate participant, update the clusters of each first participant to obtain the clusters of the second participants corresponding to the clusters of the first participants;

[0151] S420, compare the second participating party cluster with the first participating party cluster to determine the second aggregation coefficient of the second participating party cluster; the second aggregation coefficient is used to characterize the degree of aggregation of the second participating party cluster;

[0152] S130 includes:

[0153] Based on the second aggregation coefficient, the second candidate participants in the second participant cluster are screened out until the second aggregation coefficient meets the preset stopping condition, resulting in multiple target participants.

[0154] In this embodiment, the scheme aims to adjust and optimize the existing clusters of first-candidate participants based on more comprehensive resource information of the first-candidate participants, in order to generate more representative or optimized clusters of second-candidate participants. Resource information may include, but is not limited to, computing power, storage capacity, communication bandwidth, data quality, data volume, and historical contributions to federated learning. The update process may involve re-evaluating and adjusting the members of existing clusters. For example, recalculating the similarity or distance between participants based on resource information, or moving certain first-candidate participants from one cluster to another according to a preset update strategy, or even forming new clusters or merging existing clusters. Another implementation is to assign a weight or priority based on resource information to each first-candidate participant while maintaining the original clustering structure, thereby performing preliminary screening of the first-candidate participants in the first-candidate participant clusters.

[0155] After updating the clusters, the degree of aggregation of the new second-participant clusters needs to be evaluated. The second aggregation coefficient is an indicator that measures the tightness or homogeneity among members within the second-participant cluster. It can be determined in several ways. For example, it can be calculated based on the mean, variance, or weighted average of the resource information (such as computing power, communication performance, etc.) of each second-participant candidate participant in the second-participant cluster to reflect the concentration of resources within the cluster. Another way to determine it is by comparing the differences between the second-participant clusters and the first-participant clusters in terms of structure, membership composition, or resource distribution, and combining these differences to quantify the degree of aggregation of the second-participant clusters. For example, if the updated clusters are more balanced or concentrated in resource distribution, their second aggregation coefficient may be higher.

[0156] This scheme refines and improves the original screening process. It demonstrates that by introducing an updated second-party cluster and a second aggregation coefficient, the final target participant screening process will be based on this new second aggregation coefficient. This means that the screening decision no longer depends solely on the initial first aggregation coefficient, but takes into account the updated clustering state and aggregation degree. The screening process may include iteratively removing the second candidate participant from the second-party cluster that contributes the most to the aggregation degree or has the least resource mismatch, until the second aggregation coefficient reaches a preset stopping condition, such as reaching a certain threshold, or the number of members in the cluster reaching a preset range.

[0157] This application's scheme, based on preliminary clustering and calculation of a first aggregation coefficient, further introduces a dynamic update mechanism for the first participant clusters. Specifically, before initial screening based on the first aggregation coefficient, the system uses the resource information of each first candidate participant to re-evaluate and adjust the existing first participant clusters, thereby obtaining second participant clusters that are more optimized and reflect the current resource status. Subsequently, the method compares the clusters before and after the update and, based on the characteristics of the second participant clusters, determines a second aggregation coefficient that better characterizes their aggregation degree. By introducing the second aggregation coefficient, this method can more precisely evaluate the quality and stability of the updated clusters. Finally, the screening process no longer relies solely on the initial first aggregation coefficient, but instead iteratively screens the second candidate participants in the second participant clusters based on this second aggregation coefficient, which has been updated and re-evaluated with resource information, until the second aggregation coefficient meets a preset stopping condition. This mechanism makes the selection process of target participants more dynamic and adaptive, better able to adapt to the volatility of resource information of participants in the green power grid supply chain, and ensures that the final selected target participants not only have good aggregation in terms of network latency, but also more reasonable resource allocation, thereby providing a more stable and efficient data sharing foundation for federated learning model training.

[0158] Through the aforementioned technical solutions, this method, based on preliminary clustering and aggregation coefficient evaluation, introduces a dynamic update mechanism for participant clusters and a refined screening mechanism based on the updated aggregation coefficients. This allows the selection process of target participants to more fully consider the dynamic changes in participant resource information within the green power grid supply chain, avoiding the problems of insufficient resource utilization or poor screening results that may result from relying solely on initial evaluation. By updating the clusters and re-evaluating their aggregation degree, this method can more accurately identify participants that not only have good aggregation properties in terms of network latency but are also more matched and optimized in terms of computing, communication, and other resources. This significantly improves the overall quality and stability of the final selected target participants, thereby providing a more reliable and efficient data sharing foundation for federated learning model training, effectively ensuring the smooth progress of the federated learning task and the improvement of model performance.

[0159] In some embodiments described above, this application proposes updating the first participant cluster based on the resource information of the first candidate participant to obtain the second participant cluster. However, in practice, relying solely on resource information for updating may not fully reflect the actual contribution potential of each participant in the federated learning task, resulting in inefficient clusters in subsequent screening processes and difficulty in effectively identifying the optimal target participant.

[0160] In this regard, this application further proposes that the resource information includes computing performance and communication performance; S410 includes:

[0161] The optimal weight of each first candidate participant is determined based on the busyness, computing performance, and communication performance of each first candidate participant.

[0162] Based on the preferred weights of each of the first candidate participants, the clusters of each of the first participants are updated to obtain the second participant clusters corresponding to the first participant clusters.

[0163] In this embodiment, computational performance refers to the ability of a participant to perform computational tasks, such as processing data and training models. It can be quantified as floating-point operations per second (FLOPS), the number of CPU cores, GPU computing power, and memory size. For example, computational performance can be evaluated based on the CPU model, GPU model, and quantity provided by the participant; or it can be determined by running benchmark programs to measure the completion time of a specific computational task. Communication performance refers to the efficiency and reliability of data transmission between participants. It can be quantified as network bandwidth, network latency, and packet loss rate. For example, communication performance can be evaluated by measuring the data transmission rate between a participant and the central server or other participants; or it can be determined by periodically sending probe packets and calculating the average round-trip time (RTT) and packet loss rate.

[0164] Based on this, the preferred weights of each first-candidate participant are determined according to their workload, computational performance, and communication performance. The preferred weight is a comprehensive indicator used to quantify the importance or priority of each first-candidate participant in updating clusters. This weight comprehensively considers the participant's workload, computational performance, and communication performance, aiming to more comprehensively evaluate its suitability as a participant in federated learning. There are several ways to determine the preferred weights. For example, the preferred weights can be determined using the following formula 4:

[0165] Formula 4

[0166] In formula 4, Used to characterize the preferred weight of the i-th participant. Used to characterize the busyness of the i-th participant. Used to characterize the computational performance of the i-th participant. Used to characterize the communication performance of the i-th participant.

[0167] Alternatively, multi-attribute decision analysis methods, such as the Analytic Hierarchy Process (AHP) or the TOPSIS method, can be used to determine the optimal weights of each participant by constructing a judgment matrix or calculating the distance to the ideal solution, with busyness, computing performance, and communication performance as decision attributes.

[0168] Subsequently, based on the preferred weights of each first candidate participant, the clusters of first participants are updated to obtain the second participant clusters corresponding to the first participant clusters. Updating the clusters refers to adjusting and optimizing the original first participant clusters according to these preferred weights after determining them, to form new, better second participant clusters. Specifically, the first candidate participants in all first participant clusters are sorted in descending order of their preferred weights; those ranked higher indicate better performance and are more suitable for training the federated learning model. Then, the bottom 20% of first candidate participants are removed from the first participant clusters to obtain the second participant clusters corresponding to the first participant clusters.

[0169] This embodiment specifically defines resource information as computational and communication performance, and combines this with the busyness of participants to more comprehensively and precisely evaluate the overall capabilities and suitability of each first candidate participant in the federated learning task, thereby determining their preferred weights. Based on these preferred weights, the first participant clusters are updated, ensuring that the resulting second participant clusters more accurately reflect the actual contribution potential of each participant. This avoids the limitations of updating clusters based solely on single or general resource information, significantly improving the quality and effectiveness of the clustering results. Therefore, when subsequently filtering participants based on the second aggregation coefficient, high-quality and efficient target participants can be identified more efficiently, thereby optimizing the training process of the federated learning model, improving the model's convergence speed and performance, and reducing overall communication and computational overhead.

[0170] In some embodiments described above in this application, a method for screening participants is proposed by updating clusters and determining their aggregation coefficients. However, when determining the aggregation degree of the updated clusters, if the aggregation coefficients are recalculated solely based on the updated members, it may not adequately reflect the dynamic changes in the structure and internal consistency of the clusters during the update process. This could affect the accuracy and continuity of the cluster aggregation degree assessment, and consequently, the selected target participants may not be optimal.

[0171] In this regard, this application further proposes that S420 includes:

[0172] Obtain the ratio of the number of second participants in the second participant cluster to the number of first participants in the first participant cluster;

[0173] The average value is obtained by averaging the ratio between the center evaluation value of each second candidate participant in the second participant cluster and the center evaluation value of the cluster center node of the first participant cluster.

[0174] The second aggregation coefficient of the second participant cluster is determined based on the ratio of the number of participants, the second average value, and the first aggregation coefficient of the first participant cluster.

[0175] In this embodiment, the ratio of the number of second participants in the second-participant cluster to the number of first participants in the first-participant cluster is obtained to quantify the change in cluster size before and after the update. This ratio reflects the direct impact of the update operation on the cluster size and is an important structural indicator for evaluating cluster stability and aggregation. This ratio can be obtained by directly counting the number of cluster members before and after the update and then performing a division operation. For example, before the update, the system can record the member list of the first-participant cluster and count its number; after the update, it can record the member list of the second-participant cluster and count its number, and then calculate the ratio. Alternatively, the system can dynamically maintain a cluster member count counter during each update operation (e.g., filtering out participants) to obtain and calculate the ratio in real time.

[0176] The second average is obtained by averaging the ratios between the center evaluation values ​​of each second candidate participant in the second participant cluster and the center evaluation value of the cluster center node in the first participant cluster. This step aims to quantify the degree of aggregation or similarity between each member in the updated second participant cluster and the core node of the original first participant cluster. By calculating the ratios and averaging them, the impact of the update operation on the internal consistency of the clusters can be assessed, i.e., whether the updated members still maintain a high correlation with the original core, thus providing a key quality indicator for determining the subsequent aggregation coefficient. This process first obtains the center evaluation value of each second candidate participant in the second participant cluster and the center evaluation value of the cluster center node in the first participant cluster. Then, the ratio between the center evaluation value of each second candidate participant and the center evaluation value of the cluster center node is calculated. Finally, all these ratios are arithmetically averaged to obtain the second average.

[0177] Based on the ratio of participating parties, the second average value, and the first aggregation coefficient of the first participating party cluster, the second aggregation coefficient of the second participating party cluster is determined. This step is crucial for quantifying the aggregation degree of the updated clusters by comprehensively considering multiple factors. It combines changes in cluster size (participating party ratio), the consistency between cluster members and the original core (second average value), and the aggregation degree of the previous round (first aggregation coefficient), thus providing a more comprehensive and dynamic assessment of aggregation degree. Specifically, the second aggregation coefficient of the second participating party cluster can be determined using the following formula 5:

[0178] Formula 5

[0179] In formula 5, The second aggregation coefficient is used to characterize the m-th second-participant cluster. The number of second participants used to characterize the m-th second participant cluster. Used to characterize the number of first participants in the m-th first participant cluster. The center evaluation value is used to characterize the cluster center node i of the m-th first participant cluster. The center evaluation value is used to characterize the i-th second candidate participant in the m-th second participant cluster. The first aggregation coefficient is used to characterize the m-th first participant cluster.

[0180] It should be noted that, in order to ensure that the calculation results are meaningful, when performing fractional operations in this embodiment of the invention, if the denominator is 0, a parameter adjustment factor greater than 0 needs to be added to the denominator before summing to prevent the denominator from being 0. The value of the parameter adjustment factor can be set by the implementer according to the actual situation, and this invention does not impose any special restrictions.

[0181] The proposed method determines the updated second aggregation coefficient by comprehensively considering the changes in the number of participants in the cluster before and after the update, the ratio of the updated members to the original cluster center node's central evaluation value, and the aggregation coefficient before the update. This method not only quantifies the structural changes in the cluster but also assesses the consistency between its internal members and the core, and inherits the evaluation results from the previous round, thus providing a more comprehensive, dynamic, and continuous aggregation degree assessment.

[0182] Through the above technical solution, this application can more accurately and dynamically evaluate the aggregation degree of updated clusters. This comprehensive evaluation method avoids the information loss that may result from simply recalculating the aggregation coefficients, making the perception of changes in cluster structure and internal consistency more sensitive. Therefore, in the subsequent participant screening process, decisions can be made based on more accurate aggregation coefficients, which helps to screen out target participants with higher aggregation degree and better internal consistency, thereby improving the efficiency and reliability of federated learning model training.

[0183] In some embodiments described above in this application, a method for filtering out second candidate participants in a second participant cluster is proposed based on a second aggregation coefficient. However, in its implementation, relying solely on a single second aggregation coefficient for filtering may fail to adequately reflect the overall aggregation quality of each second participant cluster, resulting in an insufficiently refined filtering process. Alternatively, when the aggregation level is still high, it may be difficult to effectively identify and remove participants with low contributions to the training of the federated learning model, thereby affecting the quality of the final target participants and the training effect of the federated learning model.

[0184] In this regard, this application further proposes that S130 includes:

[0185] The third average value corresponding to the second aggregation coefficient of each second participant cluster is normalized to obtain the aggregation evaluation degree.

[0186] In response to a aggregation evaluation score greater than a preset aggregation threshold, the second candidate participant with the smallest weight in the target second participant cluster with the largest second aggregation coefficient is screened out until the aggregation evaluation score is no greater than the preset aggregation threshold, thus obtaining multiple target participants.

[0187] In this embodiment, the third average value corresponding to the second aggregation coefficient of each second participant cluster is normalized to obtain the aggregation evaluation score. The third average value is the result of calculating the arithmetic mean of the second aggregation coefficients of each second participant cluster. The normalization process aims to transform the third average value to a uniform scale for easier comparison and evaluation. For example, min-max normalization can be used to linearly transform the data to the [0,1] interval, or Z-score normalization can be used to transform the data into a distribution with a mean of 0 and a standard deviation of 1. Through normalization, the aggregation evaluation score can more fairly and accurately reflect the aggregation quality of each second participant cluster.

[0188] In response to a convergence evaluation score exceeding a preset convergence threshold, the candidate second participant with the smallest weight in the target second participant cluster with the largest second convergence coefficient is screened out until the convergence evaluation score is no greater than the preset convergence threshold, resulting in multiple target participants. The preset convergence threshold is a pre-defined critical value used to determine whether the convergence evaluation score of the current second participant cluster reaches an acceptable level. This threshold can be empirically set or determined through experimental optimization based on factors such as the needs of the actual application scenario, the performance requirements of the federated learning model, and system resource limitations. For example, the preset convergence threshold can be set to 0.3.

[0189] When the aggregation evaluation score exceeds this threshold, it indicates that the aggregation degree of the current second-participant cluster is still too high, or that there are still participants within it that need further optimization, thus requiring a screening operation. The specific strategy for this screening operation is as follows: First, identify the target second-participant cluster with the largest second aggregation coefficient. This target second-participant cluster typically represents the cluster with the highest current aggregation degree or the one most in need of optimization. Then, within this target second-participant cluster, further identify the second candidate participant with the smallest preferred weight. The preferred weight can comprehensively consider resource information such as the participant's workload, computational performance, and communication performance, reflecting the participant's potential contribution or suitability to the federated learning model training. The participant with the smallest preferred weight is usually considered to be the participant with the lowest contribution to the federated learning model training or the worst resource conditions in the current cluster. By prioritizing the screening of such participants, the overall quality of the clusters can be effectively improved. The screening process is iterative. After each participant is screened out, the relevant aggregation evaluation score is recalculated and compared with the preset aggregation threshold again until the aggregation evaluation score is no longer greater than the preset aggregation threshold. At this point, the remaining second candidate participant is determined as the final target participant.

[0190] Through the above technical solution, this application can quantitatively evaluate the aggregation quality of the second participant clusters and perform dynamic, iterative screening based on this. This avoids the problems of incomplete or excessive screening that may result from simply relying on a single aggregation coefficient. By prioritizing the elimination of participants with the lowest weights from the clusters with the highest aggregation degree, the participant set can be optimized more effectively, ensuring that the final selected target participants not only meet the aggregation requirements but also have higher overall quality and greater contribution to the training of the federated learning model, thereby improving the efficiency of federated learning and model performance.

[0191] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0192] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A cross-system data security sharing method for a green power grid supply chain, characterized in that, The method includes: In response to the acquisition of the first candidate participants in the green power grid supply chain, the first candidate participants are clustered based on the network latency between them to obtain at least one first participant cluster. Based on the resource information of each of the first candidate participants in the first participant cluster, a first aggregation coefficient of the first participant cluster is determined; the first aggregation coefficient is used to characterize the degree of aggregation of the first participant cluster. Based on the first aggregation coefficient, the first candidate participants in the first participant cluster are screened out until the first aggregation coefficient meets a preset stopping condition, resulting in multiple target participants; the target participants are used to share data to train the federated learning model. Determining the first aggregation coefficient of the first participating party cluster includes: Using the first candidate participant as the node and the network topology between the first candidate participants as the edge, a participant structure graph is constructed. Based on the resource information of each first candidate participant in the first participant cluster and the associated candidate participants, the center evaluation value of each first candidate participant is determined; the associated candidate participants are located in the same first participant cluster as the first candidate participants and are connected to the first candidate participants in the participant structure diagram; the center evaluation value is used to characterize the probability that the first candidate participant is the cluster center node of the first participant cluster. Based on the central evaluation value of each of the first candidate participants in the first participant cluster, the first aggregation coefficient of the first participant cluster is determined.

2. The cross-system data security sharing method for a green power grid supply chain according to claim 1, characterized in that, The resource information includes communication performance; The step of determining the central evaluation value of each first candidate participant based on the resource information of each first candidate participant in the first participant cluster and the associated candidate participants includes: The average communication performance of each of the first candidate participants in the first participant cluster is obtained by averaging the communication performance of the first participant cluster. The average network latency between each of the associated candidate participants and the first candidate participant is averaged to obtain the average network latency of the associated candidate participants. Based on the number of participants in the associated candidate participants, the average network latency, the average communication performance, and the communication performance of the first candidate participant, the central evaluation value of the first candidate participant is determined.

3. The cross-system data security sharing method for a green power grid supply chain according to claim 1, characterized in that, The step of determining the first aggregation coefficient of the first participant cluster based on the central evaluation value of each of the first candidate participants in the first participant cluster includes: The first candidate participant with the largest center evaluation value in the first participant cluster is determined as the cluster center node of the first participant cluster; The mean value of the difference between the central evaluation value of each first candidate participant in the first participant cluster and the cluster center node is averaged to obtain the first average value. Based on the first average value and the center evaluation value of the cluster center node, the first aggregation coefficient of the first participating party cluster is determined.

4. The cross-system data security sharing method for a green power grid supply chain according to claim 1, characterized in that, Before the step of clustering the first candidate participants based on the network latency between them to obtain at least one cluster of first participant groups in response to obtaining the first candidate participants in the green power grid supply chain, the method further includes: Based on the resource information of each participant in the green power grid supply chain, the busyness of each participant is determined. Based on the busyness of each participant, the first candidate participant whose busyness is less than a preset busyness threshold is selected from the participants.

5. The cross-system data security sharing method for a green power grid supply chain according to claim 4, characterized in that, The resource information includes real-time load and computing performance; The determination of the busyness of each participant based on resource information of each participant in the green power grid supply chain includes: Divide the real-time load of the participating party by the computing performance to obtain the busy evaluation value; The busyness rating is normalized to obtain the busyness level of the participating party.

6. The cross-system data security sharing method for a green power grid supply chain according to claim 1, characterized in that, Before the step of filtering out the first candidate participants in the first participant cluster based on the first aggregation coefficient until the first aggregation coefficient meets a preset stopping condition and multiple target participants are obtained, the method further includes: Based on the resource information of each of the first candidate participants, the clusters of each of the first participants are updated to obtain the second participant clusters corresponding to the first participant clusters. The second participating party cluster is compared with the first participating party cluster to determine the second aggregation coefficient of the second participating party cluster; the second aggregation coefficient is used to characterize the degree of aggregation of the second participating party cluster. Based on the first aggregation coefficient, the first candidate participants in the first participant cluster are screened out until the first aggregation coefficient meets a preset stopping condition, resulting in multiple target participants, including: Based on the second aggregation coefficient, the second candidate participants in the second participant cluster are screened out until the second aggregation coefficient meets the preset stopping condition, resulting in multiple target participants.

7. The cross-system data security sharing method for a green power grid supply chain according to claim 6, characterized in that, The resource information includes computing performance and communication performance; The step of updating the clusters of each first participant based on the resource information of each first candidate participant to obtain a second participant cluster corresponding to the first participant cluster includes: Based on the busyness of each first candidate participant, the computing performance, and the communication performance, the preferred weight of each first candidate participant is determined. Based on the preferred weights of each of the first candidate participants, the clusters of each of the first participants are updated to obtain the second participant clusters corresponding to the first participant clusters.

8. The cross-system data security sharing method for a green power grid supply chain according to claim 6, characterized in that, The step of comparing the second participating party cluster with the first participating party cluster to determine the second aggregation coefficient of the second participating party cluster includes: Obtain the ratio of the number of second participants in the second participant cluster to the number of first participants in the first participant cluster; The average value is obtained by averaging the ratio between the center evaluation value of each of the second candidate participants in the second participant cluster and the center evaluation value of the cluster center node of the first participant cluster. Based on the ratio of the number of participants, the second average value, and the first aggregation coefficient of the first participant cluster, the second aggregation coefficient of the second participant cluster is determined.

9. The cross-system data security sharing method for a green power grid supply chain according to claim 6, characterized in that, Based on the second aggregation coefficient, second candidate participants in the second participant cluster are screened out until the second aggregation coefficient meets a preset stopping condition, resulting in multiple target participants, including: The third average value corresponding to the second aggregation coefficient of each second participant cluster is normalized to obtain the aggregation evaluation degree. In response to the aggregation evaluation degree being greater than a preset aggregation threshold, the second candidate participant with the smallest weight in the target second participant cluster with the largest second aggregation coefficient is screened out until the aggregation evaluation degree is not greater than the preset aggregation threshold, thus obtaining multiple target participants.