Model generation method, device and product based on zero-knowledge proof and federated learning

By adopting a method based on zero-knowledge proof and federated learning in the training of regulatory analysis models in the gas industry, blockchain technology is used to ensure data privacy and model training accuracy, solving the problems of low training accuracy and data breach risk in traditional mode.

CN120197731AActive Publication Date: 2025-06-24SHENZHEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510689286.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-24
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Under the traditional model, the training accuracy of the gas industry regulatory analysis model is low and there is a risk of data leakage, which is difficult to meet the regulatory authorities' demand for refined risk control.

Method used

Using a method based on zero-knowledge proof and federated learning, the first gradient and proof parameters of each participant are obtained through the blockchain node device, multiple second gradients are determined, and the global model is updated based on these second gradients.

Benefits of technology

Improve the accuracy of model training, ensure data privacy, avoid the risk of commercial secret leakage and public panic, and enhance the training accuracy of the global model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197731A_ABST
    Figure CN120197731A_ABST
Patent Text Reader

Abstract

The embodiment of the invention is suitable for the technical field of data processing, and provides a model generation method, device and product based on zero-knowledge proof and federated learning, and the method comprises the steps: obtaining a first gradient and a proof parameter sent by each participant; the first gradient comprises a gradient generated when a participant carries out model training based on local data and a federated learning architecture, and the proof parameters comprise parameters used for verifying that the local data meet preset constraint conditions; the federal learning architecture comprises initial model parameters for model training; determining a plurality of second gradients; the second gradient is the first gradient corresponding to an effective parameter in the proof parameters; and updating the initial model parameters in the global model based on the plurality of second gradients to obtain an updated global model. By adopting the method, the model training accuracy can be improved on the basis of ensuring data privacy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data processing, and particularly relates to a model generation method, device, and product based on zero-knowledge proof and federated learning. Background Art

[0002] In the field of gas industry supervision, regulatory agencies usually need to build an analysis model that can reflect the safety situation of the entire industry in real time and accurately, in order to achieve purposes such as early warning of gas accident risks and optimization of emergency response strategies. Usually, the training of this analysis model depends on the massive historical data accumulated by gas operators, including privacy data such as the spatio-temporal distribution of safety events, types of equipment failures, and effects of disposal measures.

[0003] However, the above privacy data not only concerns the enterprise's own competitive advantages (such as trade secrets such as equipment failure handling capabilities and risk prevention and control strategies), but may also cause public panic about gas safety due to leakage, leading to a social trust crisis.

[0004] Therefore, the requirements for data privacy and protection of trade secrets make it difficult for enterprises to directly share raw data. In the traditional mode, only a small amount of data uploaded by enterprises can be obtained for training, resulting in low accuracy of model training and difficulty in meeting the needs of regulatory authorities for refined risk control. For example, it is difficult to generate a risk assessment report with high accuracy. Summary of the Invention

[0005] The embodiments of this application provide a model generation method, device, and product based on zero-knowledge proof and federated learning, which can solve the problems of low accuracy of the model trained in the traditional mode and the risk of data leakage.

[0006] In a first aspect, the embodiments of this application provide a model generation method based on zero-knowledge proof and federated learning, which is applied to a node device of a blockchain. The method includes: Obtain the first gradients and proof parameters sent by each participating party; the first gradients include the gradients generated by the participating party during model training based on local data and the federated learning architecture, and the proof parameters include the parameters used to verify that the local data meets the preset constraint conditions; the federated learning architecture includes the initial model parameters for model training; Determine a plurality of second gradients; the second gradients are the first gradients corresponding to the valid parameters in the proof parameters; Update the initial model parameters in the global model based on the plurality of second gradients to obtain an updated global model.

[0007] In an embodiment, updating the initial model parameters in the global model based on the plurality of second gradients to obtain an updated global model includes: Aggregate the plurality of second gradients to obtain an aggregated gradient; Calculate the product of the aggregated gradient and the preset learning rate; Determine the difference between the initial model parameters and the product as the target model parameters; Replace the initial model parameters with the target model parameters to obtain the updated global model.

[0008] In one embodiment, aggregating multiple second gradients to obtain an aggregated gradient includes: Determine the historical contribution degree corresponding to each participant; the historical contribution degree is used to quantify the improvement effect of the first gradient sent by the participant at the historical moment on the global model; Determine the aggregation weight corresponding to each second gradient based on the historical contribution degree; the aggregation weight with a high historical contribution degree is greater than or equal to the aggregation weight with a low historical contribution degree; Respectively perform weighted summation of each second gradient and the corresponding aggregation weight to obtain the aggregated gradient.

[0009] In one embodiment, the method further includes: Count the total number of obtained signatures; the signature is a proof for the participant to request the aggregated gradient; If the total number is greater than or equal to the preset number, publish the aggregated gradient.

[0010] In a second aspect, an embodiment of the present application provides a model generation device based on zero-knowledge proof and federated learning, which is applied to a node device of a blockchain. The device includes: An acquisition module, configured to acquire the first gradients and proof parameters sent by each participant; the first gradient includes the gradient generated by the participant during model training based on local data and the federated learning architecture, and the proof parameters include the parameters for verifying that the local data meets the preset constraint conditions; the federated learning architecture includes the initial model parameters for model training; A determination module, configured to determine multiple second gradients; the second gradient is the first gradient corresponding to the valid parameter in the proof parameters; An update module, configured to update the initial model parameters in the global model based on the multiple second gradients to obtain the updated global model.

[0011] In a third aspect, an embodiment of the present application provides another model generation method based on zero-knowledge proof and federated learning, which is applied to an electronic device. The method includes: Perform model training based on local data and the federated learning architecture to obtain the first gradient generated during the training process; the federated learning architecture includes the initial model parameters for model training; Generate proof parameters for verifying that the local data meets the preset constraint conditions; Send the first gradient and the proof parameters to the node device of the blockchain; the node device is used to determine a plurality of second gradients, and update the initial model parameters of the global model based on the plurality of second gradients to obtain an updated global model; the second gradient is the first gradient corresponding to the valid parameter in the proof parameters.

[0012] In one embodiment, model training is performed based on local data and a federated learning architecture to obtain the first gradient generated during the training process, including: Obtain the macroscopic metrics and the federated learning architecture sent by the node device; Obtain the training data corresponding to the macroscopic metrics from the local data; Perform model training based on the training data and the federated learning architecture to obtain the first gradient.

[0013] In one embodiment, model training is performed based on local data and a federated learning architecture to obtain the first gradient generated during the training process, including: When performing model training on the initial model parameters in the model based on local data, obtain the third gradient generated during the training process; the third gradient includes at least one; Perform noise processing on each third gradient respectively to obtain the corresponding noise gradient; Aggregate each noise gradient to generate the first gradient.

[0014] In one embodiment, performing noise processing on each third gradient respectively to obtain the corresponding noise gradient includes: For any one of the third gradients, determine the sensitivity category of the local data used correspondingly when generating the third gradient; Determine the noise corresponding to the third gradient based on the sensitivity category; the noise corresponding to the high-sensitivity category is greater than the noise corresponding to the low-sensitivity category; Process the third gradient based on the noise to obtain the noise gradient.

[0015] Fourthly, an embodiment of the present application provides another model generation device based on zero-knowledge proof and federated learning, which is applied to an electronic device. The device includes: A training module, configured to perform model training based on local data and a federated learning architecture to obtain the first gradient generated during the training process; the federated learning architecture includes the initial model parameters for model training; A generation module, configured to generate proof parameters for verifying that the local data meets the preset constraint conditions; A sending module, configured to send the first gradient and the proof parameters to the node device of the blockchain; the node device is used to determine a plurality of second gradients, and update the initial model parameters of the global model based on the plurality of second gradients to obtain an updated global model; the second gradient is the first gradient corresponding to the valid parameter in the proof parameters.

[0016] In a fifth aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method according to the first aspect or the third aspect above is implemented.

[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method according to the first aspect or the third aspect above is implemented.

[0018] In a fifth aspect, an embodiment of the present application provides a computer program product, which when running on a computer device, causes the computer device to execute the method according to the first aspect or the third aspect above.

[0019] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: After obtaining the first gradient and the proof parameter generated by the participating party based on the local data and the federated learning architecture, the second gradient can be determined from multiple first gradients based on the validity of the proof parameter. That is, the second gradient can be recognized as "trusted", and the model update information carried by it is obtained by training based on the local data that meets the preset constraint conditions, rather than being obtained by training with false local data, improving the accuracy of subsequent global model updates based on the second gradient. Moreover, the first gradient is generated and uploaded by the participating party when training the initial model parameters in the federated learning architecture based on the local data. Therefore, it can be considered that the participating party only uploads gradient information rather than the original local data. Furthermore, it can ensure that the local sensitive data remains local, avoiding the risk of commercial secret leakage or public panic caused by direct data sharing. Finally, the initial model parameters in the global model are updated based on multiple second gradients to obtain the updated global model. Based on this, since multiple second gradients are generated by each participating party based on its own local data, they essentially integrate the characteristics of a large amount of data. Furthermore, updating the initial model parameters based on multiple second gradients can enable the global model to capture deep patterns that cannot be covered by the data of a single participating party, improving the training accuracy of the updated global model. Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 It is a flowchart of the implementation of a model generation method based on zero-knowledge proof and federated learning provided by an embodiment of the present application; Figure 2 It is a schematic diagram of an implementation manner for generating a first gradient in a model generation method based on zero - knowledge proof and federated learning provided by an embodiment of the present application; Figure 3 It is a schematic diagram of an implementation manner for generating a first gradient in a model generation method based on zero - knowledge proof and federated learning provided by another embodiment of the present application; Figure 4 It is a schematic diagram of an implementation manner for generating a noise gradient in a model generation method based on zero - knowledge proof and federated learning provided by an embodiment of the present application; Figure 5 It is a schematic diagram of an implementation manner for updating a global model in a model generation method based on zero - knowledge proof and federated learning provided by an embodiment of the present application; Figure 6 It is a schematic diagram of an implementation manner for generating an aggregated gradient in a model generation method based on zero - knowledge proof and federated learning provided by an embodiment of the present application; Figure 7 It is a flowchart of an implementation of a model generation method based on zero - knowledge proof and federated learning provided by another embodiment of the present application; Figure 8 It is a schematic diagram of an application scenario in a model generation method based on zero - knowledge proof and federated learning provided by another embodiment of the present application; Figure 9 It is a schematic diagram of the structure of a model generation device based on zero - knowledge proof and federated learning provided by an embodiment of the present application; Figure 10 It is a schematic diagram of the structure of a model generation device based on zero - knowledge proof and federated learning provided by another embodiment of the present application; Figure 11 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0022] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well - known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0023] It should be understood that when used in the description of the present application specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0024] It should be noted that the information collection process (such as the face image collection process, fingerprint information collection process, etc.) / feature extraction process involved in the present application is executed with the user's knowledge and permission, that is, the information collection process / feature extraction process complies with the requirements of laws and regulations and does not belong to acts that harm the public interest.

[0025] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0026] In the field of gas industry supervision, regulatory agencies usually need to build an analysis model that can reflect the safety situation of the entire industry in real time and accurately, in order to achieve purposes such as early warning of gas accident risks and optimization of emergency response strategies. Usually, the training of this analysis model relies on a large amount of historical data accumulated by gas operators, including privacy data such as the spatio-temporal distribution of safety events, equipment failure types, and the effects of disposal measures.

[0027] However, the above privacy data not only concerns the enterprise's own competitive advantages (such as trade secrets such as equipment failure handling capabilities and risk prevention and control strategies), but may also cause public panic about gas safety due to leakage, leading to a social trust crisis.

[0028] Therefore, the requirements for data privacy and trade secret protection make it difficult for enterprises to directly share raw data. In the traditional mode, only a small amount of data uploaded by enterprises can be obtained for training, resulting in low training accuracy of the model and difficulty in meeting the needs of regulatory authorities for refined risk control.

[0029] Based on this, in order to improve the training accuracy of the model, an embodiment of the present application provides a model generation method based on zero-knowledge proof and federated learning, which can be applied to the node devices of the blockchain.

[0030] Among them, the blockchain is a distributed ledger technology. By packing data into "blocks" in chronological order and linking them into a chain structure in a cryptographic manner, it realizes the immutability, traceability, and decentralized storage of data. Its core features include: Decentralization: There is no centralized management agency, and the data is distributed on multiple node devices in the network, avoiding single-point failures.

[0031] Tamper-proof: Each block contains the hash value of the previous block. Once the data is written, it is difficult to tamper with, ensuring integrity.

[0032] Transparency: Network participants can view the data on the chain (subject to permissions), enhancing trust.

[0033] Security: Ensure network security through encryption algorithms (such as SHA-256) and consensus mechanisms (such as PoW, PoS).

[0034] In the context of federated learning, blockchain is mainly used for: Storing and validating gradient data: Nodes receive gradients (such as the first gradient, the second gradient) and proof parameters sent by each participant, and verify the validity of the data (such as whether local data complies with privacy constraints). Ensuring the trustworthy circulation of data: Automatically execute rules through smart contracts (for example, only aggregate valid gradients) to avoid malicious data contaminating the global model. Tracing and auditing: All operations are recorded on the chain, facilitating subsequent verification of the behavior of participants (for example, who sent gradient data and when).

[0035] A node device is a participating entity in the blockchain network and can be a hardware device such as a server, computer, or mobile phone. Each node device needs to run blockchain client software to participate in network communication, data verification, and maintain the blockchain ledger.

[0036] Please refer to Figure 1 , Figure 1 which shows the implementation flowchart of a model generation method based on zero-knowledge proof and federated learning provided by an embodiment of the present application. The method includes the following steps: S101. Obtain the first gradients and proof parameters sent by each participant.

[0037] In one embodiment, the above-mentioned first gradient includes the gradient generated by a participant during model training based on local data and a federated learning architecture, and the federated learning architecture includes initial model parameters for model training.

[0038] The above-mentioned participant refers to an independent entity participating in the federated learning task (for example, electronic devices such as mobile phones and tablets). Generally, each participant has its own local data (for example, various types of data such as accident location, time, device model, and treatment measures). Among them, the local data can be data in the gas field, medical field, etc., and no limitation is made here. In this embodiment, the data in the gas field is taken as an example for the following explanation.

[0039] It should be noted that since privacy data not only concerns the enterprise's own competitive advantage but may also cause public panic about gas safety due to leakage, leading to a social trust crisis. Therefore, it is necessary to collaboratively train the global model while protecting data privacy.

[0040] Among them, the first gradient is the direction and magnitude of the model parameter update calculated by the participating party when training the initial model based on local data and the federated learning architecture. Among them, the gradient is the core signal for model optimization, indicating the direction in which the model parameters should be adjusted to minimize the loss function. In the equipment fault prediction task in the gas field, the first gradient may include the parameter update amount of the gas equipment operation parameters to improve the recognition ability of safety features such as leakage and pressure anomalies.

[0041] The above-mentioned federated learning architecture is a distributed machine learning framework that allows participating parties to collaboratively train a global model without sharing the original data. Generally, the federated learning architecture can include various defined information such as initial model parameters, communication protocol security mechanisms, etc.

[0042] Exemplarily, the initial model parameters are the initial state of the global model (for example, randomly initialized neural network weights), which can be randomly set by the node device or actively set by the regulatory party, and then published to each participating party.

[0043] In addition, the communication protocol can define information such as gradient transmission and model aggregation rules between the participating party and the node device.

[0044] In addition, the security mechanism can include information such as data encryption and privacy protection algorithms (for example, differential privacy, secure multi-party computation).

[0045] As an example, the electronic device corresponding to the participating party can generate the first gradient according to the steps S201-S203 as shown in Figure 2 The details are as follows: S201. Obtain the macroscopic indicators and the federated learning architecture sent by the node device.

[0046] S202. Obtain the training data corresponding to the macroscopic indicators from the local data.

[0047] In one embodiment, the above-mentioned node device has been explained and will not be elaborated here. Among them, the macroscopic indicator refers to the statistical characteristic value extracted from the local data such as gas equipment operation, user behavior, and safety events, which is used to generally reflect the overall safety situation, equipment operation efficiency, or user behavior pattern of the industry or region.

[0048] In one embodiment, the macroscopic indicator can be generated by aggregating the statistics of the original data (such as count, mean, ratio, etc.), only retaining the global features without including the details of individual data, which is the product of realizing "data can be used but not seen" in federated learning.

[0049] It should be noted that the corresponding to the macroscopic indicator is the detail indicator, which is used to describe the specific details of the data.

[0050] As an example, the macro indicators can be safety situation indicators, which are used to reflect the safety risk level and accident distribution characteristics of the gas system. The corresponding detailed indicators can be the number of gas accidents per unit area (such as per square kilometer); the proportion of high-risk areas: the proportion of the number of geographical areas where the number of accidents exceeds the threshold; the repeated accident rate: the proportion of the same equipment / area having two or more accidents within a certain period, and other indicators.

[0051] It can be understood that when the electronic device obtains the corresponding training data from the local data based on the macro indicators, the obtained training data may not need to include detailed data such as the specific number of accidents, accident locations, user information, etc.

[0052] As an example, the macro indicators can include but are not limited to accident frequency indicators, illegal operation indicators, and accident severity indicators. Correspondingly, the detailed indicators corresponding to the accident frequency indicators can be the number of gas leakage accidents, the number of gas explosion accidents, the number of gas poisoning accidents, the number of daily accidents, the number of weekly accidents, the number of monthly accidents, and other indicators. The detailed indicators corresponding to the illegal operation indicators can be illegal equipment, illegal time, illegal location, illegal type, and other detailed indicators. And the detailed indicators corresponding to the accident severity indicators can be casualties, property losses, and other detailed indicators.

[0053] Exemplarily, taking the accident frequency indicator as an example, the electronic device only needs to count the total number of accidents occurring in a certain week as the training data, rather than counting the number of accidents corresponding to each type within a week as the training data. Furthermore, during the training process, the node device does not need to intervene in the specific accident details, further providing privacy protection for the local data.

[0054] It can be understood that the training data corresponding to the macro indicators is obtained through operations (such as summation, averaging), eliminating individual data identifiers (such as the number of gas leakage accidents). Even if the training data corresponding to the macro indicators is leaked, it is impossible to reverse-engineer the data corresponding to the detailed indicators in the macro indicators. For example, "the number of gas accidents in a certain area within a week is 3 times" does not involve any specific accident types and the number of accidents of each type.

[0055] In one embodiment, the above-mentioned macro indicators can be pre-defined by the regulatory party in an indicator system (such as the "Macro Indicator Specification for Gas Industry Federated Learning"), clarifying each macro indicator. The electronic device can automatically traverse the local data to obtain the training data corresponding to each macro indicator.

[0056] In another embodiment, in order to enable the regulatory party (node device) to verify the authenticity of the local data (such as the training data corresponding to the macro indicators) through a smart contract, the electronic device can generate proof parameters for verifying that the local data meets the preset constraint conditions.

[0057] Among them, the preset constraint conditions can be set according to the actual situation, and there is no limitation on this. Exemplarily, the preset constraint conditions can be that the number of accidents in the past 3 months ≤ 5 times.

[0058] It should be noted that the preset constraint conditions are rules predefined to ensure data compliance. This rule can be used to achieve data distribution alignment and reflect local data characteristics.

[0059] Exemplarily, the preset constraint conditions can be that the proportion of training data corresponding to "leakage accidents" needs to be within the range of [20%, 40%]. Furthermore, it can make the distribution of training data used for training consistent with the training objectives of the global model (for example, in the accident data of each participating party, the proportion of "leakage accidents" is within the range of [20%, 40%].

[0060] In one embodiment, the electronic device can generate proof parameters of statistical metrics based on two zero - knowledge proof technologies, zk - SNARKs and zk - STARKs. Or, generate proof parameters based on the homomorphic encryption algorithm. In this embodiment, there is no limitation on the way of generating proof parameters.

[0061] Taking the generation of proof parameters by zk - SNARKs as an example, zk - SNARKs is a technology for efficient proof of fixed statistical metrics, applicable to statistical metrics within a fixed time window (for example, "the number of accidents in the past 3 months ≤ 5"). The electronic device can pre - design a circuit structure for verifying "the number of accidents ≤ 5", and the data is fixed when generating the proof. Then, the electronic device can generate public parameters (such as proof keys and verification keys), and generate proof parameters corresponding to local data based on the circuit structure. Finally, the electronic device can upload the public parameters and proof parameters to the node device. At this time, the node device only needs to verify the validity of the proof parameters based on the public parameters.

[0062] And, for the zk - STARKs technology, it is applicable to dynamically updated data (such as "the real - time number of accidents ≤ threshold"), without the need for a trusted setup, and the proof can be incrementally updated. The way of generating proof parameters by the zk - STARKs technology is similar to that of zk - SNARKs, and will not be elaborated here.

[0063] It should be noted that for incremental update of the proof, when a new accident occurs in the local data record, the electronic device can generate proof parameters corresponding to the incremental part of the training data based on the zk - SNARKs technology: there is no need to generate proof parameters corresponding to all training data. Furthermore, it can reduce the computing resources of the electronic device.

[0064] S203. Perform model training based on the training data and the federated learning architecture to obtain the first gradient.

[0065] In one embodiment, when performing model training, the electronic device can be pre-set with a neural network model, and the initial model parameters in the neural network model can be the initial model parameters in the global model set in the node device to achieve federated learning among multiple electronic devices.

[0066] In one embodiment, the set neural network model can be a lightweight neural network model. For example, it can be a long short-term memory network, a gated recurrent unit network, etc. The present invention is not limited thereto.

[0067] As an example, taking the long short-term memory network as an example, during the model training process, the input layer is used to input training data, and the hidden layer can extract spatial features and capture the temporal dependencies between features. The output layer can perform classification based on the extracted features (for example, classification such as normal or faulty). For example, output the probability corresponding to normal and the probability corresponding to faulty. Then, the long short-term memory network can calculate the training loss using a loss function. Finally, based on the training loss, backpropagation is performed to calculate the first gradient. For example, calculate the gradients of the weights and bias parameters of each layer. Among them, the generated first gradient is used to iteratively update the model parameters of the current long short-term memory network.

[0068] Among them, the above content is only an example of obtaining the first gradient. In this embodiment, the process of obtaining the first gradient through model training is not limited.

[0069] It can be understood that by obtaining the macroscopic metrics and the federated learning architecture sent by the node device, corresponding training data can be obtained from the local data for model training, making the training more targeted and helping to improve the model training efficiency and convergence speed. Moreover, relying on the federated learning architecture for model training can ensure that the data is trained locally, avoid the cross-node flow of the original local data, and enhance the security of private data.

[0070] It should be added that in addition to the above basic process, the following optimization scenarios can also be considered in practical applications. Exemplarily, when training a model based on training data, data desensitization processing can also be performed on the training data to seek a balance between data utilization and privacy security. By stripping sensitive information, both the needs of model training are met and regulatory requirements are complied with, avoiding the risk of data abuse.

[0071] Exemplarily, the electronic device can generate desensitized features of the training data through desensitization technology and delete the training data that meets the prohibited features.

[0072] Among them, the desensitized features refer to the features that are retained after processing and can be used for model training, but the sensitive information has been removed or blurred, and only the key dimensions at the business level are retained.

[0073] Exemplarily, the desensitized features can be features such as accident types (e.g., "Pipeline Leakage_Level2"), response duration intervals (e.g., "1 - 2 hours"), etc.

[0074] Specifically, abstract the specific accident type into a standardized code. "Pipeline Leakage" represents the accident category, and "Level2" represents the severity level. At this time, the corresponding training data retains the business attributes of the accident type, facilitating the model to learn the laws of different types / levels of accidents, while avoiding the leakage of sensitive information such as the location and responsible party of specific accidents.

[0075] Also, divide the actual response duration into an interval range instead of being accurate to specific minutes. This can retain the key indicators of response efficiency for model optimization of emergency response strategies, while avoiding exposing the time-sensitive details of specific events.

[0076] Also, the prohibited features can include features such as specific coordinates, device serial numbers, operator IDs, etc., which can further avoid the leakage of privacy data.

[0077] In another embodiment, it should be noted that during the training process of the deep learning model, to achieve model convergence, multiple iterative optimizations are usually required. Since in each iteration, the model calculates and generates a gradient based on the current batch of training data, if the gradient generated in each iteration is directly uploaded to the node device as the first gradient, it will cause consumption of network bandwidth, affecting the transmission efficiency and training cost. In addition, the unprocessed gradient usually contains a large amount of model training information. If the gradient is directly uploaded, there may be a risk of privacy leakage. An attacker may reverse-engineer the training data through the gradient and thus obtain sensitive data.

[0078] Based on this, in order to reduce the network resources required for uploading the first gradient and ensure the security of privacy data, the electronic device can generate the first gradient according to the Figure 3 S301 - S303 steps shown below. Details are as follows: S301. When training the initial model parameters in the model based on local data, obtain the third gradient generated during the training process of the model.

[0079] In one embodiment, the local data has been explained above and will not be elaborated here. It should be noted that when performing step S301, the local data can be processed based on the above S201 - S202 steps first to obtain the training data. Then, based on the training data, train the initial model parameters in the model and obtain the third gradient generated during the training process of the model.

[0080] Among them, the model can be the lightweight neural network model described above and will not be elaborated here.

[0081] Moreover, the above-mentioned third gradient includes at least one. Since each iteration described in the above S203 step will generate a gradient, when only one iteration is performed, a corresponding third gradient will be obtained. Moreover, when multiple iterations are performed, multiple corresponding third gradients will be obtained. Among them, the third gradient can be considered as the original density that has not been processed.

[0082] S302. Perform noise processing on each third gradient respectively to obtain the corresponding noise gradient.

[0083] In one embodiment, the electronic device can perform noise processing on the third gradient to obtain a noise gradient based on methods such as Laplace Noise and Gaussian Noise, and there is no limitation on this.

[0084] As an example, the electronic device can perform noise processing on the third gradient to obtain a noise gradient according to the S401 - S403 steps shown as follows: Figure 4 Details are as follows: S401. For any third gradient, determine the sensitivity category of the local data corresponding to when generating the third gradient.

[0085] In one embodiment, during the model training process, due to the large scale of local data, usually only part of the local data is used for training each time during the model training process. It can be understood that each local data usually corresponds to a sensitivity category. For example, categories such as high sensitivity, general sensitivity, and low sensitivity. Among them, the local data corresponding to high sensitivity usually has a higher privacy level, and the local data corresponding to low sensitivity usually has a lower privacy level. In this embodiment, there is no limitation on the types of sensitivity categories. Among them, the corresponding sensitivity category of each local data can be set in advance.

[0086] Exemplarily, when performing model training for the first time, low - sensitivity local data can be used for training to obtain the first third gradient. Moreover, when performing model training for the second time, high - sensitivity local data can be used for training to obtain the second third gradient.

[0087] S402. Determine the noise corresponding to the third gradient based on the sensitivity category.

[0088] In one embodiment, the noise corresponding to the high - sensitivity category is greater than the noise corresponding to the low - sensitivity category. Exemplarily, there can be preset noises corresponding to various sensitivity categories in the electronic device to determine the noise corresponding to the third gradient based on this corresponding relationship.

[0089] Specifically, the electronic device can associate the local data used in each model training process with the corresponding generated third gradient. Furthermore, after determining the sensitivity category corresponding to the local data, the noise corresponding to the local data can be determined based on the above corresponding relationship.

[0090] In another embodiment, the electronic device can also preset an initial noise for each third gradient. After that, the initial noise is increased or decreased based on the sensitivity category to obtain the noise corresponding to the third gradient.

[0091] Exemplarily, when the local data corresponding to the third gradient is highly sensitive data, the initial noise can be added to a preset value to obtain the noise corresponding to the third gradient; and when the local data corresponding to the third gradient is moderately sensitive data, the initial noise can be determined as the noise corresponding to the third gradient; and when the local data corresponding to the third gradient is low-sensitivity data, the initial noise can be subtracted from a preset value to obtain the noise corresponding to the third gradient. In this embodiment, the method for determining the noise is not limited.

[0092] As an example, the electronic device can determine the noise according to the following formula. Details are as follows: ; Wherein, represents the noise corresponding to the third gradient, represents the preset gradient sensitivity, represents the sensitivity value corresponding to the sensitivity category. Among them, the higher the sensitivity category, the lower the corresponding sensitivity value. Furthermore, the noise calculated by the above formula will be higher, so that the noise added to the third gradient subsequently will also be larger, and the privacy protection degree of the local data will also be improved.

[0093] S403. Process the third gradient based on the noise to obtain a noise gradient.

[0094] In one embodiment, the electronic device can perform an operation (such as addition or multiplication) on the noise and the third gradient to obtain a noise gradient; or, multiply the noise by a corresponding weight and then add it to the third gradient to obtain the above noise gradient. In this embodiment, the method for obtaining the noise gradient is not limited.

[0095] It should be noted that during the data processing, first determine the sensitivity category of the local data corresponding to the generation of the third gradient, and then, based on the principle that the noise corresponding to the high-sensitivity category is greater than the noise corresponding to the low-sensitivity category, match the noise for the sensitivity category for the gradient. Furthermore, when obtaining the noise gradient by processing the third gradient based on the determined noise, differential noise addition can be used to focus on protecting high-sensitivity data to prevent privacy leakage, and moderately process low-sensitivity data to avoid excessive interference with the original data features, thereby avoiding excessive perturbation of the data. From the perspective of compliance management, operating according to the data sensitivity classification is in line with the requirements of taking differential protection measures for data at different levels in many industry specifications, and follows the hierarchical management specification. Ultimately, hierarchical enhanced privacy protection is achieved, while protecting data privacy, ensuring the training effect of the model, and improving the system security and the trust of the participating parties.

[0096] S303. Aggregate each noise gradient to generate a first gradient.

[0097] In one embodiment, if the electronic device undergoes multiple iterations during the process of training the model, a large number of noise gradients will be generated correspondingly. At this time, if all the noise gradients are directly uploaded to the node device, it will consume a large amount of network resources and reduce the security when uploading the first gradient. Based on this, the electronic device can aggregate each noise gradient to generate a first gradient and upload it.

[0098] Exemplarily, the electronic device can sum multiple noise gradients with weights to generate a first gradient for uploading, thereby reducing the amount of data required to be uploaded and reducing network resources. And since the first gradient is aggregated based on the noise gradients, generating the first gradient based on the noise gradients after adding noise can also reduce the possible risk of privacy leakage.

[0099] It should be noted that since each noise gradient covers the weights and bias parameters of each layer in the model, during the aggregation process, it is necessary to aggregate the weights of the same layer and the bias parameters of the same layer respectively, and finally obtain the first gradient composed of the aggregated weights and bias parameters of each layer.

[0100] To sum up the above description, it is an example of each participating party generating the first gradient and the proof parameters locally. Based on the above example, the participating party can participate in the global model training without sharing local data, thereby protecting local data privacy while improving the training accuracy of the updated global model.

[0101] Among them, after generating the first gradient and the proof parameters, in order to further ensure data security, the electronic device can also encrypt the first gradient and the proof parameters first, and then upload the encrypted first gradient and proof parameters.

[0102] S102. Determine multiple second gradients.

[0103] In one embodiment, the above-mentioned second gradient is the first gradient corresponding to the valid parameter among the proof parameters. The methods for verifying that the proof parameter is a valid parameter include, but are not limited to, cryptographic proof matching, homomorphic encryption verification, etc., and are not limited thereto.

[0104] Exemplarily, if the electronic device generates proof parameters based on a zero-knowledge proof (ZKP) circuit, the node device can use the verification key (including parameters such as the circuit's constraints and hash function) to perform the following steps for verification.

[0105] Specifically, the node device can parse the structure of the proof parameter through the verification key and check whether it conforms to the preset circuit logic (such as arithmetic constraints, threshold constraints, etc.). Then, using the hash function in the public parameters, verify whether the proof parameter correctly "binds" the calculation process of the local data to ensure that there is no forgery or tampering. Finally, through a random challenge (such as the challenge value c in ZKP) and a response mechanism, verify whether the proof parameter can correctly answer the authenticity questions about the local data to ensure that the data is not maliciously constructed.

[0106] Alternatively, when generating proof parameters based on the homomorphic encryption algorithm, the public parameters may include the encryption public key and the homomorphic operation rules. At this time, the node device can use the public key in the public parameters to verify the encryption legality of the proof parameter (such as signature verification). Finally, according to the homomorphic operation rules, check whether the calculation results (such as gradient aggregation, statistical values) in the proof parameter conform to the expected mathematical transformation (such as additive homomorphism, multiplicative homomorphism) to indirectly verify the validity of the local data.

[0107] In this embodiment, the methods for verifying that the proof parameter is a valid parameter are not limited.

[0108] It should be noted that for the first gradient of the proof parameter being an invalid parameter, it can be considered that this first gradient may be generated based on "contaminated" local data during the training process and thus has no reference value. Therefore, subsequent processing can be carried out only for the second gradient.

[0109] S103. Update the initial model parameters in the global model based on multiple second gradients to obtain an updated global model.

[0110] In one embodiment, the node device can aggregate multiple second gradients to obtain an aggregated gradient. Then, update the initial model parameters with the aggregated gradient (for example, add or subtract) to obtain updated model parameters. At this time, the global model containing the updated model parameters is the updated global model.

[0111] As an example, the electronic device can according to Figure 5The steps S501 - S504 shown update the global model. Details are as follows: S501. Aggregate multiple second - order gradients to obtain an aggregated gradient.

[0112] In one embodiment, the manner of aggregating the second - order gradients to obtain the aggregated gradient can be similar to the manner of aggregating the noise gradients to obtain the first - order gradient above. Exemplarily, since each second - order gradient respectively covers the weights and bias parameters of each layer in the model, during the aggregation process, it is necessary to aggregate the weights of the same layer and the bias parameters of the same layer respectively, and finally obtain an aggregated gradient composed of the aggregated weights and bias parameters of each layer.

[0113] It should be noted that when aggregating the first - order gradients of each participant, if a simple average weighting or data - volume - weighted manner is used for aggregation, there are significant limitations.

[0114] On the one hand, this method does not consider the actual contribution differences of the historical data and historical first - order gradients of the participants to the optimization of the global model, which may lead to the underestimation of the value of high - quality data, while low - quality or "polluted" data interferes with the model convergence. For example, due to the lack of equipment monitoring data or sensor failures in some gas enterprises, the gradient update not only fails to improve the model performance, but may even reduce the prediction accuracy.

[0115] On the other hand, the lack of an incentive mechanism may cause the participants to lack the motivation to improve data quality and optimize gradients, easily forming a vicious cycle of "bad money drives out good money", seriously restricting the application effect of federated learning in the field of gas safety and the collaborative development of the industry.

[0116] Based on this, in order to be able to quantify the historical contributions of the participants, optimize the aggregated gradient after aggregation, and improve the performance of the updated global model, the electronic device can generate the aggregated gradient according to the steps S601 - S603 shown as follows: Figure 6 Details are as follows: S601. Determine the historical contribution degree corresponding to each participant.

[0117] In one embodiment, the above - mentioned historical contribution degree is used to quantify the improvement effect of the first - order gradient sent by the participant at the historical moment on the global model. Among them, the first - order gradient at the historical moment can be 1 or multiple, and there is no limit to this. In one embodiment, the above - mentioned historical contribution degree can be set by the supervisor, or the node device can determine the historical contribution degree of each participant from dimensions such as model performance improvement, gradient effectiveness, and data quality.

[0118] Among them, the improvement of model performance is used to measure the improvement of the prediction metrics of the global model by the gradients of the participating parties; the gradient effectiveness is used to evaluate the consistency between the gradient direction of the participating parties and the global optimal update direction, reflecting the effectiveness of the gradient in optimizing the model; the data quality score is used to evaluate the data quality of the participating parties from aspects of data integrity, timeliness, and compliance.

[0119] As an example, the node device can record the prediction metrics of the global model before each iteration (for example, the fault prediction accuracy metric). After the gradients uploaded by each participating party are aggregated, the updated prediction metrics can be recorded, and the improvement value corresponding to the prediction metrics can be calculated. For example, the above improvement value is obtained by subtracting the prediction metrics before the update from the prediction metrics after the update. Then, the improvement value is normalized to the [0, 1] interval to avoid the influence of different metric magnitudes on the results.

[0120] Example: In the 5th iteration, Gas Company A increased the accuracy of the global model from 82% to 85%, with an improvement value of 3%; Company B only increased the accuracy from 82% to 82.5%, with an improvement value of 0.5%. After normalization: The score (historical contribution) of Company A: 0.83, the score (historical contribution) of Company B: 0.

[0121] And, as another example, the node device can calculate the similarity between the second gradient of the participating party and the aggregated gradient (for example, by methods such as cosine similarity and Euclidean distance similarity); then, if the similarity is lower than a preset threshold (for example, 0.5), the first gradient can be regarded as an invalid gradient and the score is set to 0; otherwise, the similarity is retained as the original value of the historical contribution.

[0122] Exemplarily, the cosine similarity between the second gradient of Company C and the aggregated gradient is 0.8, and the similarity of Company D is only 0.2 due to data anomalies. Therefore, the historical contribution corresponding to Company C can be determined to be 0.8, and the historical contribution corresponding to Company D is 0.

[0123] In one embodiment, the node device can determine the historical contribution based on any of the above methods, or can first determine the contribution by the above multiple methods respectively, and then, perform a weighted sum of the contributions determined by the multiple methods respectively to obtain the historical contribution corresponding to each participating party.

[0124] S602. Determine the aggregation weight corresponding to each second gradient based on the historical contribution.

[0125] S603. Respectively perform a weighted sum of each second gradient and the corresponding aggregation weight to obtain the aggregated gradient.

[0126] In one embodiment, the aggregation weight with a high historical contribution is greater than or equal to the aggregation weight with a low historical contribution.

[0127] As an example, the node device can preset the aggregation weights corresponding to multiple contribution degree ranges respectively. Then, based on the contribution degree range where the historical contribution degree is located, the aggregation weight corresponding to the second gradient is determined.

[0128] Alternatively, the node device can also set the weight adjustment values corresponding to multiple contribution degree ranges respectively, and the preset weights corresponding to each participant. Then, based on the contribution degree range where the historical contribution degree is located, the weight adjustment value corresponding to the second gradient is determined, so as to adjust the preset weight based on the weight adjustment value to obtain the above-mentioned aggregation weight. In this embodiment, the method for determining the aggregation weight is not limited.

[0129] In one embodiment, after obtaining each second gradient, each second gradient can be weighted and summed with the corresponding aggregation weight to obtain the above-mentioned aggregation gradient.

[0130] In this embodiment, quantifying the improvement effect of the second gradient of each participant on the global model based on the historical contribution degree can accurately identify high-quality participants with high data quality and strong gradient effectiveness, so that participants with high historical contribution degree can obtain higher weights during the aggregation of the second gradient, motivating them to continuously provide high-quality data; for participants with low historical contribution degree, the gradient aggregation weight is reduced, effectively suppressing the interference of invalid gradients caused by data loss, anomalies or malicious attacks on the model. Through the differential weight allocation mechanism, the risk that malicious nodes damage the model training by uploading "polluted" gradients can be significantly reduced, and the misleading of the model optimization direction by low-quality data can be avoided, thereby accelerating the convergence speed of the global model and improving the model prediction accuracy and generalization ability.

[0131] S502. Calculate the product of the aggregation gradient and the preset learning rate.

[0132] S503. Determine the difference between the initial model parameter and the product as the target model parameter.

[0133] S504. Replace the initial model parameter with the target model parameter to obtain the updated global model.

[0134] In one embodiment, the above-mentioned preset learning rate can determine the size of the adjustment of the model parameters (such as weights, biases, etc.) according to the aggregation gradient during each iteration. Among them, the preset learning rate can be set by the supervisor according to the actual situation, and this is not limited.

[0135] It can be understood that if the learning rate is too large, the parameter update step may exceed the optimal solution position, resulting in model oscillation or even divergence (non-convergence); if the learning rate is too small, the parameter update is slow, the model convergence speed is greatly reduced, and it may fall into a local optimal solution or fail to converge for a long time.

[0136] Among them, the node device can determine the difference between the initial model parameters and the product as the updated target model parameters.

[0137] As an example, the node device can determine the target model parameters according to the following formula. Details are as follows: ; Among them, represents the target model parameters, represents the initial model parameters at the t-th global model update, represents the preset learning rate, represents the aggregated gradient.

[0138] It can be understood that when performing the (t + 1)-th global model update, the initial model parameters in the global model at this time can be considered as .

[0139] In one embodiment, the node device can update the global model according to the above method; or, after obtaining the target model parameters, the node device can send the target model parameters to the electronic device corresponding to the supervisor, so that the electronic device corresponding to the supervisor can update the global model based on the target model parameters. In this embodiment, the method for updating the global model is not limited.

[0140] Among them, after the electronic device corresponding to the supervisor updates the global model based on the model parameters, the electronic device can generate a risk assessment report based on the global model, and display the corresponding macro indicators and the risk assessment report on the dashboard.

[0141] It should be noted that in the gradient aggregation stage, since the second gradient comes from the local training of each participant, after integration, it can fuse the data features of multiple parties, avoid the limitations of a single data source, and thus improve the generalization ability of the model. Moreover, by introducing the preset learning rate, the parameter update step size can be accurately controlled to ensure the stability and accuracy of model optimization. Finally, by determining the difference between the initial model parameters and the product as the target model parameters and performing updates, the model can gradually adjust the parameters based on the direction and magnitude of the aggregated gradient to reduce the loss function value, obtain a global model with better performance and stronger adaptability, and improve the prediction accuracy and reliability of the model in actual application scenarios.

[0142] In this embodiment, after obtaining the first gradients and proof parameters generated by the participating parties based on their local data and the federated learning architecture, the second gradients can be determined from multiple first gradients based on the validity of the proof parameters. That is, the second gradients can be recognized as "trusted", and the model update information they carry is obtained by training based on local data that meets the preset constraints, rather than training with false local data, which improves the accuracy of subsequent global model updates based on the second gradients. Moreover, the first gradients are generated and uploaded by the participating parties when training the initial model parameters in the federated learning architecture based on their local data. Therefore, it can be considered that the participating parties only upload gradient information rather than the original local data. Furthermore, it can ensure that local sensitive data remains local, avoiding the risk of commercial secret leakage or public panic caused by direct data sharing. Finally, the initial model parameters in the global model are updated based on multiple second gradients to obtain the updated global model. Based on this, since multiple second gradients are generated by each participating party based on its own local data, they essentially integrate the characteristics of a large amount of data. Furthermore, updating the initial model parameters based on multiple second gradients can enable the global model to capture deep patterns that cannot be covered by the data of a single participating party, improving the training accuracy of the updated global model.

[0143] In another embodiment, to ensure traceability and security, the blockchain usually stores the aggregated gradients. Under this mechanism, the participating parties need to rely on the aggregated gradients to update their local models, but randomly sending the aggregated gradients to each participating party may trigger a trust crisis.

[0144] Based on this, to ensure the legality and credibility of the aggregated gradients and avoid blindly spreading the aggregated gradients and reducing the trust risk generated by the participating parties, the node device can count the total number of signatures obtained; the signature is the proof for the participating party to request the aggregated gradients. Then, when the total number is greater than or equal to the preset number, the aggregated gradients are released. Otherwise, when the total number is less than the preset number, the release of the aggregated gradients is prohibited.

[0145] Among them, the preset number can be set according to the actual situation and is not limited in this regard. Moreover, the signature can be considered as a proof for the authenticity and validity of the participating party's initiative to request the aggregated gradients.

[0146] It can be understood that by counting the total number of signatures of the participating parties requesting the aggregated gradients and setting the preset number as the release threshold, the credibility of the process can be effectively strengthened: only when the total number of signatures reaches or exceeds the preset number, the release operation of the aggregated gradients is triggered. Furthermore, through the signature mechanism, the participating parties' consensus recognition of the aggregation result can be achieved, ensuring the legality and credibility of the aggregated gradients, avoiding blindly spreading the aggregated gradients, and guaranteeing the trust risk generated by the participating parties.

[0147] Refer to Figure 7, which shows a schematic diagram of the steps for generating another global model provided by the embodiments of the present application, may specifically include the following steps: S701. Train a model based on local data and a federated learning architecture to obtain a first gradient generated during the training process; the federated learning architecture includes initial model parameters for model training.

[0148] Among them, the execution subject in this embodiment may be an electronic device. The electronic device may obtain the macro indicators and the federated learning architecture sent by the node device; obtain the training data corresponding to the macro indicators from the local data; and train the model based on the training data and the federated learning architecture to obtain the first gradient.

[0149] Alternatively, when training the initial model parameters in the model based on local data, obtain a third gradient generated during the training process of the model; there is at least one third gradient; then, perform noise processing on each third gradient to obtain a corresponding noise gradient; and aggregate each noise gradient to generate the first gradient.

[0150] In addition, when obtaining the noise gradient, for any third gradient, determine the sensitivity category of the local data corresponding to the generation of the third gradient; determine the noise corresponding to the third gradient based on the sensitivity category; the noise corresponding to the high-sensitivity category is greater than the noise corresponding to the low-sensitivity category; and process the third gradient based on the noise to obtain the noise gradient.

[0151] S702. Generate proof parameters for verifying that the local data meets the preset constraint conditions.

[0152] S703. Send the first gradient and the proof parameters to the node device of the blockchain; the node device is used to determine a plurality of second gradients and update the initial model parameters of the global model based on the plurality of second gradients to obtain an updated global model; the second gradient is the first gradient corresponding to the valid parameter in the proof parameters.

[0153] In one embodiment, the various terms and processing procedures corresponding to S701-S703 above have been explained in the above Figures 2 to 4 corresponding embodiments, and will not be described herein again.

[0154] In this embodiment, the electronic device can train based on local data and initial model parameters to obtain the first gradient, which not only makes full use of the data value of the participants but also avoids the direct exposure of the original data, ensuring data privacy and security. Then, proof parameters are generated to verify that the local data meets the preset constraint conditions, effectively ensuring data quality, compliance, and integrity, and preventing "contaminated" data from interfering with model training. Finally, the first gradient and the proof parameters are sent to the blockchain node device, and the immutable property of the blockchain can be used for evidence storage. The second gradient corresponding to the valid proof parameters is screened through node verification, filtering out invalid or malicious gradients. Based on this, when updating the global model based on the screened second gradient, the training accuracy of the updated global model can be improved, while avoiding data leakage and security risks and promoting multi-party security.

[0155] To illustrate the solution in this application more clearly, specific embodiments are used below to elaborate on the solution in this application. For details, refer to Figure 8 , Figure 8 which is a schematic diagram of an application scenario in a model generation method based on zero-knowledge proof and federated learning provided in another embodiment of this application.

[0156] Taking the example of a provincial gas safety management center jointly building a device failure prediction model (global model) with multiple gas enterprises based on federated learning and blockchain technology. The participants include: node devices: verification nodes in the blockchain network (deployed by the management center), electronic devices (local servers of each gas enterprise, holding user gas device data), and preset constraint conditions: desensitized device operation data that complies with the "Urban Gas Data Security Specification".

[0157] First, the node device can broadcast the federated learning architecture (including initial model parameters) and macro indicators (for example, abnormal device pressure data in the recent 3 months).

[0158] Gas enterprise A can screen the corresponding training data from the local database. Enterprise A trains the model based on the local training data to generate multiple layers of the third gradient. Then, differential noise is added to the third gradient of different sensitive layers based on the sensitivity category of the local data corresponding to the generation of the third gradient to obtain the noise gradient. Finally, the noise gradients are aggregated to obtain the first gradient.

[0159] At the same time, enterprise A can use zero-knowledge proof (ZKP) technology to generate proof parameters corresponding to the training data and send the proof parameters and the first gradient to the node device.

[0160] The node device can verify the validity of the proof parameters uploaded by each enterprise respectively and determine the second gradients corresponding to the valid proof parameters. After that, based on the historical contribution degree corresponding to each participant, the weight corresponding to each second gradient can be determined to perform weighted summation on the second gradients to obtain the aggregated gradient. For example, determine the historical contribution degree corresponding to enterprise A to determine the weight corresponding to the second gradient uploaded by enterprise A and participate in subsequent processing.

[0161] Finally, the node device can calculate the product of the aggregated gradient and the preset learning rate and determine the difference between the initial model parameters and the product as the target model parameters to replace the initial model parameters with the target model parameters to obtain the updated global model.

[0162] Refer to Figure 9 , Figure 9 FIG. shows a schematic diagram of a model generation device based on zero-knowledge proof and federated learning provided by an embodiment of the present application. The device can be applied to a node device of a blockchain. The model generation device 900 based on zero-knowledge proof and federated learning may include an acquisition module 910, a determination module 920, and an update module 930, where: The acquisition module 910 is configured to acquire the first gradients and proof parameters sent by each participant; the first gradients include the gradients generated by the participant during model training based on local data and the federated learning architecture, and the proof parameters include the parameters used to verify that the local data meets the preset constraint conditions; the federated learning architecture includes the initial model parameters for model training.

[0163] The determination module 920 is configured to determine a plurality of second gradients; the second gradients are the first gradients corresponding to the valid parameters in the proof parameters.

[0164] The update module 930 is configured to update the initial model parameters in the global model based on the plurality of second gradients to obtain the updated global model.

[0165] In one embodiment, the first update module 930 is further configured to: Aggregate a plurality of second gradients to obtain an aggregated gradient; calculate the product of the aggregated gradient and the preset learning rate; determine the difference between the initial model parameters and the product as the target model parameters; replace the initial model parameters with the target model parameters to obtain the updated global model.

[0166] In one embodiment, the first update module 930 is further configured to: Determine the historical contribution degree corresponding to each participating party; the historical contribution degree is used to quantify the improvement effect of the first gradient sent by the participating party at the historical moment on the global model; determine the aggregation weight corresponding to each second gradient based on the historical contribution degree; the aggregation weight with a high historical contribution degree is greater than or equal to the aggregation weight with a low historical contribution degree; respectively perform weighted summation of each second gradient and the corresponding aggregation weight to obtain the aggregated gradient.

[0167] In one embodiment, the model generation device 900 based on zero-knowledge proof and federated learning further includes: A statistics module, configured to count the total number of signatures obtained; the signature is a proof for a participating party to request to obtain the aggregated gradient.

[0168] A publishing module, configured to publish the aggregated gradient if the total number is greater than or equal to a preset number.

[0169] It should be understood that Figure 9 In the structural schematic diagram of the model generation device based on zero-knowledge proof and federated learning shown, each module is used to execute Figure 1 、 Figure 5 and Figure 6 the respective steps in the corresponding embodiments, and for Figure 1 、 Figure 5 and Figure 6 the respective steps in the corresponding embodiments have been explained in detail in the above embodiments. For details, please refer to Figure 1 、 Figure 5 and Figure 6 as well as Figure 1 、 Figure 5 and Figure 6 the relevant descriptions in the corresponding embodiments, which will not be elaborated here.

[0170] Referring to Figure 10 , Figure 10 FIG. shows a schematic diagram of another model generation device based on zero-knowledge proof and federated learning provided by an embodiment of the present application. This device can be applied to an electronic device. The model generation device 1000 based on zero-knowledge proof and federated learning may include a training module 1010, a generation module 1020, and a sending module 1030, where: The training module 1010 is configured to perform model training based on local data and a federated learning architecture to obtain the first gradient generated during the training process; the federated learning architecture includes initial model parameters for model training.

[0171] The generation module 1020 is configured to generate proof parameters for verifying that the local data meets the preset constraint conditions.

[0172] A sending module 1030 is configured to send a first gradient and a proof parameter to a node device of a blockchain; the node device is configured to determine a plurality of second gradients and update initial model parameters of a global model based on the plurality of second gradients to obtain an updated global model; the second gradient is the first gradient corresponding to a valid parameter in the proof parameter.

[0173] In one embodiment, the training module 1010 is further configured to: Obtain a macroscopic metric and a federated learning architecture sent by a node device; obtain training data corresponding to the macroscopic metric from local data; perform model training based on the training data and the federated learning architecture to obtain a first gradient.

[0174] In one embodiment, the training module 1010 is further configured to: When performing model training on initial model parameters in a model based on local data, obtain a third gradient generated during the training process of the model; the third gradient includes at least one; perform noise processing on each third gradient respectively to obtain a corresponding noise gradient; aggregate each noise gradient to generate a first gradient.

[0175] In one embodiment, the training module 1010 is further configured to: For any third gradient, determine a sensitivity category of local data corresponding to when generating the third gradient; determine noise corresponding to the third gradient based on the sensitivity category; the noise corresponding to a high-sensitivity category is greater than the noise corresponding to a low-sensitivity category; perform processing on the third gradient based on the noise to obtain a noise gradient.

[0176] It should be understood that Figure 10 In the structural schematic diagram of the model generation device based on zero-knowledge proof and federated learning shown, each module is configured to execute Figures 2 - 4 and Figure 7 the respective steps in the corresponding embodiments, and for Figures 2 - 4 and Figure 7 the respective steps in the corresponding embodiments have been explained in detail in the above embodiments. For details, please refer to Figures 2 - 4 and Figure 7 and Figures 2 - 4 and Figure 7 the relevant descriptions in the corresponding embodiments, which will not be elaborated here.

[0177] Figure 11 is a structural schematic diagram of a computer device provided in an embodiment of the present application. As Figure 11As shown, the computer device 1100 of this embodiment includes: a processor 1110, a memory 1120, and a computer program 1130 stored in the memory 1120 and executable on the processor 1110, such as a program for the model generation method based on zero-knowledge proof and federated learning. When the processor 1110 executes the computer program 1130, it implements the steps in each of the above embodiments of the model generation method based on zero-knowledge proof and federated learning, such as Figure 1 S101 to S103 shown in Figure 7 or S701 - S703 shown in Figure 9 or Figure 10 When the processor 1110 executes the computer program 1130, it implements the functions of each module in the corresponding Figure 9 or Figure 10 embodiment. For example, Figure 9 or Figure 10 the functions of each module shown in. For specific details, please refer to Figure 9 or Figure 10 the relevant descriptions in the corresponding embodiments.

[0178] Exemplarily, the computer program 1130 can be divided into one or more modules. One or more modules are stored in the memory 1120 and executed by the processor 1110 to implement the model generation method based on zero-knowledge proof and federated learning provided by the embodiments of the present application. One or more modules can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program 1130 in the computer device 1100. For example, the computer program 1130 can implement the model generation method based on zero-knowledge proof and federated learning provided by the embodiments of the present application.

[0179] The computer device 1100 may include, but is not limited to, a processor 1110 and a memory 1120. Those skilled in the art can understand that Figure 10 this is only an example of the computer device 1100 and does not constitute a limitation on the computer device 1100. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.

[0180] The so-called processor 1110 may be a central processing unit, or may also be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0181] The memory 1120 may be an internal storage unit of the computer device 1100, such as the hard disk or memory of the computer device 1100. The memory 1120 may also be an external storage device of the computer device 1100, such as a plug-in hard disk, a smart memory card, a flash memory card, etc. equipped on the computer device 1100. Further, the memory 1120 may also include both the internal storage unit of the computer device 1100 and the external storage device.

[0182] An embodiment of the present application provides a computer-readable storage medium, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the model generation method based on zero-knowledge proof and federated learning in the above-mentioned various embodiments.

[0183] An embodiment of the present application provides a computer program product. When the computer program product runs on a computer device, it causes the computer device to execute the model generation method based on zero-knowledge proof and federated learning in the above-mentioned various embodiments.

[0184] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A model generation method based on zero - knowledge proof and federated learning, characterized in that, A node device applied to a blockchain, the method comprising: Obtain first gradients and proof parameters sent by each participating party; the first gradients include gradients generated by the participating party during model training based on local data and a federated learning architecture, and the proof parameters include parameters for verifying that the local data meets preset constraint conditions; the federated learning architecture includes initial model parameters for model training; Determine a plurality of second gradients; the second gradients are the first gradients corresponding to the valid parameters in the proof parameters; Update the initial model parameters in the global model based on the plurality of second gradients to obtain the updated global model.

2. The method according to claim 1, characterized in that The updating the initial model parameters in the global model based on the plurality of second gradients to obtain the updated global model includes: Aggregate the plurality of second gradients to obtain an aggregated gradient; Calculate the product of the aggregated gradient and a preset learning rate; Determine the difference between the initial model parameters and the product as the target model parameters; Replace the initial model parameters with the target model parameters to obtain the updated global model.

3. The method according to claim 2, wherein The aggregating the plurality of second gradients to obtain an aggregated gradient includes: Determine the historical contribution degree corresponding to each participating party; the historical contribution degree is used to quantify the improvement effect of the first gradients sent by the participating party at historical moments on the global model; Determine the aggregation weight corresponding to each second gradient based on the historical contribution degree; the aggregation weight with a higher historical contribution degree is greater than or equal to the aggregation weight with a lower historical contribution degree; Respectively perform weighted summation of each second gradient and the corresponding aggregation weight to obtain the aggregated gradient.

4. The method according to claim 2, characterized in that, The method further includes: Count the total number of signatures obtained; the signature is the proof for the participating party to request the aggregated gradient; If the total number is greater than or equal to a preset number, publish the aggregated gradient.

5. A model generation method based on zero-knowledge proof and federated learning, characterized in that, Applied to an electronic device, the method includes: Perform model training based on local data and a federated learning architecture to obtain first gradients generated during the training process; the federated learning architecture includes initial model parameters for model training; Generate proof parameters for verifying that the local data meets preset constraint conditions; Send the first gradients and the proof parameters to a node device of the blockchain; the node device is used to determine a plurality of second gradients and update the initial model parameters of the global model based on the plurality of second gradients to obtain the updated global model; the second gradients are the first gradients corresponding to the valid parameters in the proof parameters.

6. The method according to claim 5, wherein The performing model training based on local data and a federated learning architecture to obtain first gradients generated during the training process includes: Obtain the macroscopic metrics sent by the node device and the federated learning architecture; Obtain training data corresponding to the macroscopic metrics from the local data; Perform model training based on the training data and the federated learning architecture to obtain the first gradients.

7. The method according to claim 5 or 6, characterized in that, The performing model training based on local data and a federated learning architecture to obtain first gradients generated during the training process includes: When performing model training on the initial model parameters in the model based on the local data, obtain third gradients generated by the model during the training process; the third gradients include at least one; Perform noise processing on each of the third gradients respectively to obtain corresponding noise gradients; Aggregate each of the noise gradients to generate the first gradient.

8. The method according to claim 7, characterized in that The performing noise processing on each of the third gradients respectively to obtain corresponding noise gradients includes: For any one of the third gradients, determine the sensitivity category of the local data correspondingly used when generating the third gradient; Determine the noise corresponding to the third gradient based on the sensitivity category; the noise corresponding to the high sensitivity category is greater than the noise corresponding to the low sensitivity category; Process the third gradient based on the noise to obtain the noise gradient.

9. A computer device, characterized in that, Comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the computer device implements the method according to any one of claims 1-4 or 5-8.

10. A computer program product, characterized in that, Comprising a computer program, when the computer program is run, the method according to any one of claims 1-4 or 5-8 is executed.

Citation Information

Patent Citations

  • Decentralized federated learning method

    CN115549922A

  • Verifiable privacy protection federal learning method based on block chain

    CN118400087A

  • Gradient aggregation federal learning method based on combination of zero knowledge proof and block chain technology

    CN119420489A