Coal mine cross-business big data federated analysis method

Through the central server collecting and dividing node collections, differential privacy technology and encryption mechanism are used to solve the problems of privacy protection and data accuracy in the federal analysis of cross-business big data of coal mines, and achieve safe and efficient data sharing and analysis.

CN120296784AActive Publication Date: 2025-07-11CHINA UNIV OF MINING & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510372629.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The existing technology is difficult to achieve secure and efficient data sharing in coal mine cross-business big data federal analysis, especially in terms of privacy protection and data accuracy. It is mainly due to the lack of integrity in the graph structure and the risk of privacy leakage due to the data being distributed on different local clients.

Method used

The subgraph data of local clients is collected privately through the central server, and the differential privacy technology and node division mechanism are used to build a global graph and allocate node collections to the local client, ensuring that the query results of each set are collected only once, reducing duplicate calculations, and protecting data privacy through encryption and perturbation.

Benefits of technology

It realizes the protection of user privacy in federal analysis, reduces duplicate calculations, improves data processing efficiency, ensures data security while ensuring the accuracy of global graph analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296784A_ABST
    Figure CN120296784A_ABST
Patent Text Reader

Abstract

The invention discloses a coal mine cross-business big data federated analysis method. The method comprises the following steps: S10, local clients construct sub-graphs according to a business data relationship, and a central server privately collects coal mine business sub-graph data from a plurality of local clients and aggregates a global graph meeting differential privacy; s20, dividing all nodes into a plurality of mutually disjoint node sets by the central server based on each piece of service node degree information in the global graph, distributing each node set to a corresponding local client, and ensuring that a related query result of each set is only collected once; and S30, based on the divided nodes and global graph information, a plurality of local clients cooperatively complete a coal mine graph data analysis task in a federated scene. According to the method, the problem of data islands under cross-scene and cross-business analysis scenes of the coal mine can be effectively solved, the data security is well protected, and powerful technical support is provided for constructing a safe, efficient and shared coal mine big data analysis platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data security, and specifically, to a cross-business big data federated analysis method for coal mines. Background Art

[0002] As the main energy source in China, coal resources will still play an important role in a quite long time in the future. As a unit integrating coal production, washing, processing, sales, transportation and other operations, coal mines have made great contributions to ensuring the stable energy supply in China. With the digital and intelligent transformation and upgrading of coal mines, the demand for cross-business and cross-scenario comprehensive big data analysis in each coal mine is gradually increasing. On the one hand, the existing data silos in business operations pose obstacles to comprehensive big data analysis, and at the same time, the data security of coal mines has become an important part of China's energy security. Due to the complexity of coal mine operations and scenarios, the generated data and the relationships between data are also complex. The proposal of graph structure provides a basis for the modeling and representation of complex coal mine data. However, with the increasingly serious privacy issues and the constraints of regulations such as the Personal Information Protection Law of the People's Republic of China, the Data Security Law of the People's Republic of China, and the General Data Protection Regulation, the analysis methods based on graph-structured data are facing more and more challenges. For example, in coal mine business scenarios, the communication information for collaborative operations between various departments in different coal mine enterprises can be modeled into different graph data, which contain various sensitive information such as various plan information, production capacity information, and national energy policy information of sensitive units, and cannot be directly and unconstrainedly applied to big data analysis. Therefore, collaborative analysis of coal mine big data from different enterprises while protecting sensitive information is the key to realizing a safe, efficient, and shared coal mine intelligent analysis platform.

[0003] Based on the above analysis, differential privacy is an effective technical means to achieve coal mine data collaborative analysis and sensitive information protection. The existing differential privacy graph analysis methods are mainly divided into centralized models and local models. In the centralized scenario, a trusted server holds the global graph composed of multiple nodes and edges, but this model is prone to problems such as privacy leakage and data intrusion. In the local scenario, each client only has the information of a single user and its first-order neighbors, and does not trust the server. Instead, it protects privacy by directly perturbing local sensitive data. However, this approach will introduce a large amount of noise, thus affecting the accuracy of the results.

[0004] Although differential privacy graph analysis has been widely studied, existing methods are still difficult to apply to the federated analysis of cross-business big data in coal mines. The main reasons include: under the federated analysis framework, the data of each coal mine business system is distributed among different local clients, and each client only has a part of the global data, resulting in a lack of integrity in the constructed graph structure, making it difficult for a single client to identify complete business association relationships, thus increasing the difficulty of computing statistics; in addition, the data of different business systems may be cross or overlapping. Although a single client can provide privacy protection for edges, reporting the same information multiple times may increase the risk of being identified, thus causing privacy leakage of edges. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for federated analysis of cross-business big data in coal mines to solve the problems proposed in the above background technology.

[0006] According to one aspect of the present application, a method for federated analysis of cross-business big data in coal mines includes the following steps:

[0007] Local clients construct subgraphs according to business data relationships, and the central server privately collects coal mine business subgraph data from multiple local clients and aggregates a global graph that satisfies differential privacy.

[0008] The central server divides all nodes into multiple non-overlapping node sets based on the business node degree information in the global graph, and assigns each node set to the corresponding local client to ensure that the relevant query results of each set are only collected once.

[0009] Based on the divided nodes and global graph information, multiple local clients cooperate to complete the coal mine graph data analysis task in the federated scenario.

[0010] Preferably, the method for the central server to privately collect business data from multiple local clients and construct subgraph information according to business data relationships includes:

[0011] The first client initializes the flag vector of edge information, perturbs each element in the flag vector through the random response mechanism, and encrypts the perturbed flag vector using the public key.

[0012] Starting from the second client, the elements in the flag vector are updated in turn according to its subgraph data, and the elements are perturbed again through the random response mechanism. If the edge obtained after perturbation exists, it is encrypted using the public key and replaces the corresponding element.

[0013] All clients cooperate to decrypt the flag vector using the private key. If the element in the flag vector indicates the existence of an edge, the corresponding edge is incorporated into the global graph.

[0014] Preferably, the initialization of the flag vector of node information by the first client includes the first client initializing a flag vector Y. For each element of Y, if the edge exists, it is marked as 1, otherwise it is marked as 0;

[0015] From the second client onwards, updating the elements in the flag vector in sequence according to the new graph data includes updating each element in sequence from the second client. If the edge exists, it is marked as 1, otherwise it is marked as 0. Subsequently, 1 or 0 is perturbed through the random response mechanism. If the perturbation results in 1, the client encrypts it using the public key and then replaces the corresponding encrypted element.

[0016] Preferably, the method of dividing all nodes into multiple non - overlapping node sets based on the degree information and assigning each node set to the corresponding local client includes:

[0017] Each client calculates the degree information of each node, perturbs the degree information of each node using the Laplace mechanism, generates the perturbed degree information and sends it to the central server;

[0018] The central server compares the multiple degree values of the same node based on the perturbed degree information of the nodes collected from each local client, assigns the node to the client reporting the maximum degree value, generates the corresponding node set, and determines the node partition information;

[0019] The central server returns the node partition information to each local client.

[0020] According to another aspect of the present application, an electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the coal mine cross - business big data federated analysis method according to any one of claims 1 - 4.

[0021] According to another aspect of the present application, a computer - readable storage medium stores a computer program thereon, and the program is executed by a processor to be used to implement the coal mine cross - business big data federated analysis method according to any one of claims 1 - 4.

[0022] According to another aspect of the present application, a computer program product, the computer program is executed to be used to implement the coal mine cross - business big data federated analysis method according to any one of claims 1 - 4.

[0023] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention protects user privacy through differential privacy technology to prevent the leakage of sensitive information; reasonably divides and allocates nodes to reduce duplicate calculations and improve data processing efficiency. At the same time, the data is kept on the local client to avoid the privacy risks brought by centralized storage. By sharing computing tasks, while ensuring data security, the accuracy of global graph analysis is guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a flowchart of a method for federated analysis of cross-business big data in coal mines according to an embodiment of the present invention;

[0025] Figure 2 It is a flowchart of a method for the central server to collect subgraph information from multiple local clients in a privacy-protected manner according to an embodiment of the present invention;

[0026] Figure 3 It is an example diagram of a method for the central server to collect subgraph information from multiple local clients in a privacy-protected manner according to an embodiment of the present invention;

[0027] Figure 4 It is a block diagram of a method for dividing all nodes into multiple non-overlapping node sets based on the degree information according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0029] Please refer to Figure 1 , according to an embodiment of the present invention, a method for federated analysis of cross-business big data in coal mines is provided, including the following steps:

[0030] The method for federated analysis of cross-business big data in coal mines includes the following steps:

[0031] S10: The local client constructs a subgraph according to the business data relationship, and the central server privately collects the coal mine business subgraph data from multiple local clients and aggregates a global graph that satisfies differential privacy.

[0032] Collect sub - graph information from multiple local clients through a central server and aggregate it into a global graph while ensuring privacy protection. Since each client only has information about local users and their first - order neighbors, the central server uses a privacy - protection mechanism, such as differential privacy technology, to collect and aggregate this information. By adding noise to each local client, it ensures that no sensitive data is leaked in the collected sub - graph.

[0033] For example, the central server will add perturbations to the local data so that even if an attacker accesses the global graph, they cannot restore the specific information of individuals.

[0034] S20: The central server divides all nodes into multiple non - overlapping node sets based on the business node degree information in the global graph, and assigns each node set to the corresponding local client, ensuring that the relevant query results of each set are only collected once.

[0035] After the central server aggregates the global graph structure, it first obtains the degree information of all nodes, and then based on this degree information, divides all nodes into multiple non - overlapping node sets, ensuring that the relevant query results of each node set are only collected once, avoiding repeated calculations and information leakage.

[0036] S30: Based on the divided nodes and the global graph information, multiple local clients cooperate to complete the coal mine graph data analysis task in the federated scenario.

[0037] After completing the node division, the central server can perform federated graph analysis based on the division information of the node sets and the generated global graph structure. The central server executes the graph analysis task through the data provided by each local client, and each client only needs to process the node set assigned to it and submit the calculation results. In this way, privacy is protected while ensuring the accuracy and effectiveness of graph analysis.

[0038] In one embodiment, as Figure 2 shown, the method by which the central server collects sub - graph information from multiple local clients in a privacy - protected manner includes:

[0039] S110: The first client initializes the flag vector of the node information, then perturbs each element in the flag vector through the random response mechanism, and then encrypts the perturbed flag vector using the public key.

[0040] The first client initializes a flag vector to represent node information. Each element represents the connection between a node and other nodes. For example, whether there is an edge connection. Then, each element is perturbed through a random response mechanism to ensure the privacy of the client's local data is not leaked. The perturbed flag vector will be encrypted using the public key to prevent data from being tampered with or leaked during transmission, ensuring privacy protection.

[0041] S120: Starting from the second client, update the elements in the flag vector in sequence according to the new graph data, and perturb the elements again through the random response mechanism. If an edge exists after perturbation, encrypt it using the public key and replace the corresponding element.

[0042] Starting from the second client, update the elements in the flag vector according to its local graph data. Each client perturbs the elements in the flag vector to represent the privacy of its local data. If the perturbed element represents the existence of an edge, such as an edge between two nodes, the client encrypts this information using the public key and replaces the original element, ensuring that each client protects the privacy of its subgraph information and transmits it, and at the same time preventing information from being leaked or tampered with during transmission through encryption.

[0043] S130: All clients cooperate to decrypt the flag vector using the private key. If the element in the flag vector represents the existence of an edge, incorporate the corresponding edge into the global graph.

[0044] In one embodiment, further, the initialization of the flag vector for edge information by the first client includes the first client initializing a flag vector Y. For each element of Y, if the edge exists, it is marked as 1, otherwise it is marked as 0. For example, under the premise of satisfying differential privacy, the process of multiple clients collaborating to calculate the union of edge information is as Figure 3 shown: Suppose there are three clients C1, C2, and C3. Among them, C1 has a private edge set E1 = {e1, e2}, C2 has a private edge set E2 = {e2, e3}, C3 has a private edge set E3 = {e5}, and the node set is V = {v1, v2, v3, v4}. Therefore, E1, E2, and E3 are all subsets of the maximum possible edge set E, where E = {e1, e2, e3, e4, e5, e6}, and the specific definition of each edge is as follows: e1 = (v1, v2), e2 = (v1, v3), e3 = (v1, v4), e4 = (v2, v3), e5 = (v2, v4), e6 = (v3, v4). The goal of this method is to find the union E = E1 ∪ E2 ∪ E3 of the private edge sets of clients C1, C2, and C3.

[0045] The first client initializes the flag vector Y and marks the corresponding elements according to the existence of edges in the graph. If the edge exists, it is marked as 1; if the edge does not exist, it is marked as 0. At the same time, ensure that the status of each edge can be reflected by the flag vector. Initialize the flag vector to ensure that the existence or non-existence of each edge can be accurately represented in binary form, preparing for subsequent perturbation and encryption processes.

[0046] Starting from the second client, update the elements in the flag vector successively according to the new graph data, including each element in the successive update starting from the second client. If the edge exists, it is marked as 1; otherwise, it is marked as 0. Subsequently, perturb 1 or 0 through the random response mechanism. If the perturbation results in 1, the client uses the public key to encrypt and then replaces the corresponding encrypted element.

[0047] Each client updates the flag vector according to its local graph data to converge the information of the entire graph, further protect privacy and prepare for the next perturbation and encryption. Through perturbation and encryption, protect the local data privacy of each client, and at the same time ensure that the encrypted data can be synthesized into a global graph at the central server to avoid leakage of sensitive information.

[0048] Collect graph data through the cooperation of distributed clients, and at the same time use differential privacy technology to protect the data privacy of each client. When the central server decrypts and synthesizes the data collected from all clients, a global graph can be obtained without leaking any sensitive information of the clients.

[0049] Furthermore, in order to make full use of the real local subgraph information and further improve the accuracy of statistical results, additional communication between the server and the clients is allowed. Each client calculates the intermediate answer based on its local real subgraph data and the perturbed global graph data.

[0050] In one embodiment, the method of dividing all nodes into multiple non-overlapping node sets based on degree information and assigning each node set to the corresponding local client includes:

[0051] Each client calculates the degree information of each node, perturbs the degree information of each node using the Laplace mechanism, generates the perturbed degree information and sends it to the central server.

[0052] Each client calculates the degree information of the nodes and uses the Laplace mechanism for perturbation to ensure the privacy protection of the degree information of each node during the transmission process, preventing the leakage of sensitive data. By perturbing the degree information, data is collected under the premise of privacy protection to ensure that the final degree information is unrecognizable to other clients.

[0053] Based on the degree information of the perturbed nodes collected from each local client, the central server compares multiple degree values of the same node, assigns the node to the client reporting the maximum degree value, generates the corresponding node set, and determines the node partition information.

[0054] After the central server aggregates the perturbed degree information transmitted by each client, by comparing the degree information reported by different clients, it determines the "maximum degree value" of each node and assigns the node to the client reporting this maximum degree value. This process ensures that each node is assigned to the most suitable client, avoiding duplicate reporting and calculation. By comparing the degree information, the nodes are reasonably assigned to the corresponding clients, ensuring the rationality of the partition and the calculation efficiency.

[0055] The central server returns the node partition information to each local client.

[0056] The central server returns the determined node partition information to each local client to guide the local client on how to process and analyze the corresponding node set, ensuring that each client clearly knows which node sets it is responsible for processing, facilitating subsequent graph analysis tasks.

[0057] For example, in one embodiment, the process of the central server collecting degree information and partitioning the node set under the premise of satisfying differential privacy is given. As Figure 4 shown, assume there are three local clients C1, C2, and C3. Among them, the degree information of C1 is D1 = {3, 3, 0, 2, 1, 1}, the degree information of C2 is D2 = {0, 3, 3, 1, 1, 2}, and the degree information of C3 is D3 = {3, 0, 3, 1, 2, 1}. The specific steps are as follows:

[0058] In the first step, the three clients calculate the degree information of each node, perturb it using the Laplace mechanism to obtain D′1, D′2, and D′3, and then send them to the central server.

[0059] In the second step, for each node, the central server collects multiple degree information from local clients. By comparing the magnitudes of the degrees, the node will be assigned to the client with the maximum degree, generating the corresponding node set. For example, for node v3, the second client has the maximum degree of 3. Therefore, v3 is partitioned to C2.

[0060] In the third step, the central server returns the node partition information to the local clients, that is, U1 = {v1, v2, v4}, U2 = {v3, v6}, U3 = {v5}.

[0061] According to another aspect of the present application, an electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the method in any possible implementation manner of the first aspect and the second aspect.

[0062] Specifically, the function can be implemented in a modular or software form, and when used as an independent application or embedded in a computer device, it can be stored in a computer-readable storage medium. Based on this, the technical solution of the present application or the part that contributes to the prior art can be stored in a storage medium and used by a computer device (such as a personal computer, a server, a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes, but is not limited to: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, optical disks, and other media capable of storing program codes.

[0063] According to another aspect of the present application, a computer-readable storage medium stores a computer program, and the program is executed by a processor for the instructions in any possible implementation manner of the first aspect or the second aspect.

[0064] Among them, the computer-readable storage medium can include various forms of media, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, optical disks, etc., which can store program codes to enable a computer device (including a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method.

[0065] According to another aspect of the present application, a computer program product includes a computer program, and the computer program is executed for the method in any possible implementation manner of the first aspect or the second aspect.

[0066] Parts not involved in the present invention are the same as the prior art or can be implemented using the prior art. Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A cross-business big data federated analysis method for coal mines, characterized in that It includes the following steps: The local client constructs a subgraph according to the business data relationship. The central server privately collects the coal mine business subgraph data from multiple local clients and aggregates a global graph that satisfies differential privacy; Based on the business node degree information in the global graph, the central server divides all nodes into multiple non-overlapping node sets, assigns each node set to the corresponding local client, and ensures that the relevant query results of each set are only collected once; Based on the divided nodes and the global graph information, multiple local clients collaborate to complete the coal mine cross-business big data analysis task in the federated scenario.

2. The coal mine cross-business big data federated analysis method according to claim 1, wherein, The method for the central server to collect business data from multiple local clients in a privacy-preserving manner and construct subgraph information according to the business data relationship includes: The first client initializes the flag vector of the edge information, perturbs each element in the flag vector through the random response mechanism, and encrypts the perturbed flag vector using the public key; Starting from the second client, the elements in the flag vector are updated in turn according to its subgraph data, and the elements are perturbed again through the random response mechanism. If the perturbed edge exists, it is encrypted using the public key and replaces the corresponding element; All clients collaborate to decrypt the flag vector using the private key. If the element in the flag vector indicates that the edge exists, the corresponding edge is incorporated into the global graph.

3. The coal mine cross-business big data federated analysis method according to claim 2, wherein, The first client initializing the flag vector of the node information includes the first client initializing a flag vector Y. For each element of Y, if the edge exists, it is marked as 1, otherwise it is marked as 0; Starting from the second client, updating the elements in the flag vector in turn according to the new graph data includes starting from the second client and updating each element in turn. If the edge exists, it is marked as 1, otherwise it is marked as 0. Subsequently, 1 or 0 is perturbed through the random response mechanism. If the perturbed result is 1, the client encrypts it using the public key and then replaces the corresponding encrypted element.

4. The coal mine cross-business big data federated analysis method according to claim 1, wherein, The method for dividing all nodes into multiple non-overlapping node sets based on the degree information and assigning each node set to the corresponding local client includes: Each client calculates the degree information of each node, perturbs the degree information of each node using the Laplace mechanism, generates the perturbed degree information and sends it to the central server; Based on the degree information of the perturbed nodes collected from each local client, the central server compares the multiple degree values of the same node, and assigns the node to the client reporting the maximum degree value, generates the corresponding node set, and determines the node partition information; The central server returns the node partition information to each local client.

5. An electronic device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the coal mine cross-business big data federated analysis method according to any one of claims 1-4.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to be used to implement the coal mine cross-business big data federated analysis method according to any one of claims 1-4.

7. A computer program product, comprising a computer program, characterized in that, The computer program is executed to be used to implement the coal mine cross-business big data federated analysis method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Energy data fusion calculation method in federated learning scene based on secure sharing

    CN113449329A

  • Differential privacy and denoising data protection method under vertical federated framework

    CN115470520A

  • Distributed training method and device, terminal equipment and computer readable medium

    CN116450889A

  • Federal learning privacy protection method and system

    CN118972171A

  • Hardware protection for differential privacy

    US20190147188A1