Trusted data space privacy protection method and system based on zero knowledge proof
Through the Internet of Things data collection and the combination of zero-knowledge proof and blockchain technology, the problem of insufficient privacy protection in the multi-source data sharing process is solved, and efficient and secure data sharing and verification are achieved.
Patent Information
- Application Number
- CN202510410702.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The lack of dynamic privacy protection mechanisms for multi-source data sharing in the prior art, resulting in the easy leakage of sensitive data and difficulty in realizing trustworthy verification.
Real-time running data from multiple data sources is collected through the Internet of Things, real-time proof complexity is determined based on comprehensive analysis of security perception information, data privacy, data anomalies and data volatility, and zero-knowledge proof is used for encrypted storage and sharing.
It realizes efficient, trustworthy and secure sharing and verification of multi-source heterogeneous data without leaking the original data content, improving the security and compliance of the data space.
Smart Images

Figure CN120257329A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a privacy protection method and system for a trusted data space based on zero-knowledge proof. Background Art
[0002] A large amount of heterogeneous data is collected in real time and widely used in scenarios such as intelligent manufacturing and smart cities. How to achieve trusted sharing while ensuring data security and privacy has become the core challenge in the construction of the current trusted data space system.
[0003] Existing data privacy protection methods mostly rely on static encryption or unified-level privacy policies, and it is difficult to achieve dynamic and differentiated processing according to data sensitivity, abnormal situations, and network risks. Especially in the process of data sharing and verification, there is a lack of a zero-knowledge proof strategy that is flexible, controllable, efficient, and secure. Therefore, there is an urgent need to construct a dynamic privacy protection mechanism based on multi-dimensional complexity evaluation to achieve accurate matching and trusted proof for different data scenarios. Summary of the Invention
[0004] This application provides a privacy protection method and system for a trusted data space based on zero-knowledge proof, aiming to solve the technical problems in the prior art that there is a lack of a dynamic privacy protection mechanism for multi-source data during the sharing process, resulting in easy leakage of sensitive data and difficult implementation of trusted verification.
[0005] In view of the above problems, this application provides a privacy protection method and system for a trusted data space based on zero-knowledge proof.
[0006] In the first aspect disclosed in this application, a privacy protection method for a trusted data space based on zero-knowledge proof is provided. The method includes collecting real-time operation data of multiple data sources through the Internet of Things to obtain a real-time operation data set and a data source information set; based on the data source information set, performing zero-knowledge proof on the real-time operation data set according to the real-time proof complexity to generate a proof data set, where the real-time proof complexity is determined by comprehensive analysis of security perception information, data privacy level, data abnormality level, and data volatility level; using blockchain technology to encrypt and store the proof data set, and performing data sharing and data verification according to the encrypted storage data set.
[0007] Another aspect disclosed in this application provides a trusted data space privacy protection system based on zero-knowledge proof. The system includes an operation data collection module for collecting real-time operation data of multiple data sources through the Internet of Things to obtain a real-time operation data set and a data source information set; a proof data set generation module for performing zero-knowledge proof on the real-time operation data set according to the real-time proof complexity based on the data source information set to generate a proof data set, where the real-time proof complexity is determined by comprehensive analysis of security awareness information, data privacy, data anomaly, and data volatility; and a data encryption storage module for using blockchain technology to perform data sharing and data verification on the proof data set according to the encrypted storage data set.
[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0009] Due to the adoption of a data privacy protection technical solution that collects multi-source data through the Internet of Things, performs zero-knowledge proof based on real-time proof complexity, and uses blockchain for encrypted storage and verification, it solves the technical problems in the prior art that there is a lack of dynamic privacy protection mechanism in the process of sharing and verifying multi-source heterogeneous data, which is prone to lead to leakage of sensitive information and difficulty in trusted verification, and achieves the technical effect of realizing efficient, trusted, and secure sharing while ensuring data privacy.
[0010] The above description is only an overview of the technical solutions of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the following specifically illustrates the specific embodiments of this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a schematic flowchart of a trusted data space privacy protection method based on zero-knowledge proof provided by an embodiment of this application;
[0012] Figure 2 It is a schematic flowchart of performing zero-knowledge mapping proof on a real-time operation data set in a trusted data space privacy protection method based on zero-knowledge proof provided by an embodiment of this application;
[0013] Figure 3 It is a schematic structural diagram of a trusted data space privacy protection system based on zero-knowledge proof provided by an embodiment of this application.
[0014] Description of the reference numerals: operation data collection module 11, proof data set generation module 12, data encryption storage module 13. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The overall idea of the technical solution provided in this application is as follows:
[0016] The embodiments of the present application provide a privacy protection method and system for a trusted data space based on zero-knowledge proof. First, various data source information is collected and identified. Subsequently, the real-time proof complexity is evaluated based on data sensitivity and security posture, and a matching zero-knowledge proof method is used to perform privacy calculations on the data. Finally, the proof results are encrypted and stored in the blockchain, realizing secure sharing and trusted verification of data without revealing the original content.
[0017] After introducing the basic principle of the present application, the various non-limiting implementation manners of the present application will be specifically introduced below with reference to the accompanying drawings of the specification.
[0018] Embodiment 1
[0019] As Figure 1 shown, the embodiments of the present application provide a privacy protection method for a trusted data space based on zero-knowledge proof, and the method includes:
[0020] Step S100: Through the Internet of Things, collect the real-time operation data of multiple data sources, and obtain a real-time operation data set and a data source information set.
[0021] Specifically, the Internet of Things refers to a network system that interconnects various physical devices through sensors, communication devices, and network protocols to realize the perception, collection, transmission, and processing of information. A data source refers to a device or system that can provide real-time data. A real-time operation data set refers to a data set that describes the operation state of a device or environmental information collected from each data source at a specific time node, such as temperature, current, output, etc. A data source information set refers to the metadata corresponding to each piece of data, including the ID, type, location, and collection timestamp of the data source.
[0022] First, various types of Internet of Things sensors will be deployed in an industrial scenario, such as temperature and humidity sensors, current detectors, vibration monitoring modules, etc. Each device is assigned a unique ID for data identification. The Internet of Things gateway or edge computing node is responsible for periodically waking up these sensors, collecting their respective real-time data at the set monitoring time node, and sending this data to the data processing center through a wireless communication protocol (such as MQTT, LoRa, NB-IoT, etc.).
[0023] During the collection, the system will also synchronously record the relevant information of each data source, including the device ID, data type, collection location, and timestamp, to form a data source information set. In addition, to prevent data from being tampered with or leaked during transmission, the collection end usually performs preliminary encryption processing through a lightweight encryption algorithm.
[0024] Through this step, not only the unified and accurate data collection of multi-source devices is achieved, but also a solid data foundation is laid for subsequent data trusted computing and privacy protection. Especially on the premise that the data source is traceable and the type is clear, the system can dynamically adjust the privacy protection strategy according to the sensitivity and abnormal characteristics of different devices, thus significantly improving the intelligence and security of the data space.
[0025] Step S200: Based on the data source information set, perform zero-knowledge proof on the real-time operation data set according to the real-time proof complexity, and generate a proof data set, where the real-time proof complexity is determined by comprehensive analysis of security perception information, data privacy level, data abnormality level, and data volatility.
[0026] Specifically, zero-knowledge proof (ZKP) is a cryptographic technology used to prove to a third party that a certain data attribute or state holds without exposing the original data. Common forms include zk-SNARK, zk-STARK, etc. The proof data set refers to the data structure generated after ZKP processing, which contains corresponding proof information without revealing the original data and is used for subsequent verification, evidence storage, or sharing. The real-time proof complexity refers to the complexity level dynamically calculated based on factors such as the current security state of the data, privacy sensitivity level, and abnormal risk, which determines the strength and strategy of ZKP. Security perception information refers to the monitoring information of the current network or system security state, such as network attack frequency, virus scan results, access anomalies, etc. Data privacy level is used to describe the sensitivity level of data, such as ordinary data, business secrets, user privacy information, etc. Data abnormality level is used to reflect the degree to which data deviates from the normal range. The higher the proportion of abnormal data, the greater the abnormality level. Data volatility is used to measure the stability of data over a period of time, often expressed in the form of "standard deviation / mean".
[0027] After collecting the "real-time operation data set" and the corresponding "data source information set", first perform comprehensive security and privacy analysis on each data source. Specifically, start from four dimensions: the current network attack frequency, the sensitivity level of the data, the abnormality level and volatility generated by comparison with historical data. This information can be obtained through historical data models, threshold rule libraries, and behavior analysis algorithms (such as Z-score analysis, sliding window).
[0028] Next, after quantifying these four indicators, a value is output through weighted calculation or a neural network model, which is the "real-time proof complexity". This complexity value will be used as input and fed into a "proof method decision maker" to match a suitable proof method from a pre-configured ZKP policy library. For example, when the complexity is high, zk-STARK can be selected; when the complexity is low, a lightweight hash-based ZKP can be chosen. After selecting the proof method, zero-knowledge mapping processing is performed on the real-time data. For example, a zero-knowledge proof is generated for the fact that "the temperature is within the specified range" instead of exposing the specific temperature value. This proof information is added to the "proof data set" for subsequent uploading, verification, or blockchain operations. The entire process is jointly completed by data analysis (such as SparkStreaming), cryptographic libraries (such as libsnark, ZoKrates), and a security engine, and is deployed on edge nodes or cloud platforms.
[0029] This step realizes differential privacy protection and dynamic allocation of computing resources: providing stronger protection for high-risk and high-privacy data to prevent the leakage of key data; adopting a lightweight solution for low-risk data to improve the system processing efficiency. In addition, the real-time complexity dynamic evaluation mechanism also has a certain security adaptation ability, which can adjust the ZKP policy in a timely manner according to the external attack situation and data anomalies, greatly enhancing the credibility, anti-attack ability, and compliance of the data space.
[0030] Step S300: Use blockchain technology to encrypt and store the proof data set, and perform data sharing and data verification based on the encrypted storage data set.
[0031] Specifically, encrypted storage means that before writing data into the blockchain, the data is encrypted to ensure that the data on the chain cannot be recognized even if it is read. Data sharing means providing the proof data set to other users or systems within the authorized scope for them to verify the authenticity and compliance of the data without accessing the original data content. Data verification means reading the encrypted proof data on the chain and combining zero-knowledge verification algorithms to perform legality verification on the shared data to ensure that the data has not been tampered with and meets the business logic requirements.
[0032] After completing the zero-knowledge proof of the real-time data, the generated proof data set is encrypted. Specifically, it is achieved through symmetric encryption algorithms or asymmetric encryption algorithms, or a hash function can also be used to generate a digest value for evidence storage. After encryption, the data is packaged into a block structure and written into the blockchain through a consensus mechanism (such as PoW, PoS, or PBFT) to ensure its immutability and timestamp authenticity.
[0033] After storage is completed, the certified dataset can be shared on the blockchain. Any verifier with access rights can obtain the corresponding encrypted certified data from the blockchain and, in combination with the original zero-knowledge verification algorithm (such as a zk-SNARK verifier), complete the verification of the authenticity and compliance of the data without decrypting the data content. The data verification logic can be automatically executed through a smart contract to achieve data access control and permission management. At the same time, all operations are recorded on the blockchain, with traceability, ensuring the transparency and compliance of the sharing process.
[0034] Through this step, secure sharing and trusted verification are achieved without revealing the original data. The encryption process combined with the zero-knowledge mechanism effectively blocks the leakage of sensitive information; the blockchain ensures the immutability and traceability of the data.
[0035] Furthermore, through the Internet of Things, real-time operation data from multiple data sources is collected to obtain a real-time operation dataset and a data source information set, including: through the Internet of Things, at a predetermined monitoring time node, data is collected from multiple predetermined data sources on an industrial automated production line, and the data is preliminarily encrypted to obtain a real-time operation dataset; through the Internet of Things, the data source ID, data type, location information, and timestamp of each predetermined data source are synchronously collected to obtain a data source information set.
[0036] Specifically, an industrial automated production line refers to a production line in manufacturing that achieves continuous operation through automated equipment (such as sensors, PLC controllers, robotic arms, etc.), usually with highly digital and intelligent features. A predetermined data source refers to a device or sensor node that has been configured and designated in the system to be monitored and data collected, such as a temperature sensor, current monitor, production counter, etc. The predetermined monitoring time node is used to ensure the timeliness and synchronization of the data. The system triggers the collection operation at a preset time interval or time point, such as collecting every 10 seconds or at the whole hour every day. The data source ID refers to the number that uniquely identifies each data source, such as "TEMP-001" representing the 1st temperature sensor. The location information is used to represent the geographical or logical location of the data source on the industrial production line, such as "North Area Packaging Line" or "Workshop Position 3".
[0037] On the industrial automated production line, multiple types of sensors and intelligent devices are deployed as predetermined data sources, and the collection interval or trigger time point of each device is set in the configuration. For example, the temperature sensor collects data every 30 seconds, while the vibration sensor collects data only when an anomaly is detected. The system sends data collection instructions to these sensors synchronously at each predetermined monitoring time node through an Internet of Things gateway or embedded controller deployed at the edge. Each data source responds by collecting the corresponding real-time operation data, such as temperature values, current values, etc., and immediately performs preliminary encryption processing locally to ensure the security of the data during transmission.
[0038] At the same time, the source information of each data is collected synchronously, including the ID, type, location information and timestamp of the data source, which constitutes the "data source information set". This information can be provided by the sensor itself or generated by the system and uploaded to the data processing platform or cloud as metadata together with the encrypted data.
[0039] The entire process coordinates device communication and data upload through IoT platform software (such as MQTT middleware, OPC UA server, etc.), and implements edge encryption through lightweight encryption libraries (such as wolfSSL, mbedTLS). Ultimately, two key data structures are formed: real-time running data set and data source information set.
[0040] Through this step, the unified collection, identification and encryption protection of multi-type and heterogeneous industrial data are realized, ensuring the integrity and consistency of the data and providing high-quality basic data for subsequent privacy computing and trusted verification.
[0041] Furthermore, based on the data source information set, zero-knowledge proof is performed on the real-time running data set according to the real-time proof complexity to generate a proof data set, including: matching and calling the data identification indicator set and the data judgment threshold set based on the data source information set; based on the data identification indicator set and the data judgment threshold set, zero-knowledge mapping proof is performed on the real-time running data set according to the real-time proof complexity to generate the proof data set.
[0042] Specifically, the data identification indicator set refers to a set of quantitative indicators used to judge and analyze data status, characteristics or behaviors, such as maximum value, minimum value, mean value, growth rate, change trend, abnormal pattern recognition value, etc. The data judgment threshold set refers to the judgment limit that matches the identification indicator, which is used to define whether the data is in an abnormal state, a safe state, or a critical condition that requires encryption processing. For example, the safety threshold of the temperature sensor is 80°C, and the current overload threshold is 20A. Zero-knowledge mapping proof refers to the process of converting raw data into zero-knowledge proof that does not expose its specific content through mathematical mapping, such as proving that "the temperature is within a safe range" without disclosing the actual temperature value.
[0043] After obtaining the data source information set, these metadata (such as data source ID, type, location, etc.) are used to match the corresponding "data identification metric set" and "data judgment threshold set". These metrics and thresholds are pre-stored in the policy library and customized for different types of data sources (such as temperature, current, images, etc.). For example, temperature data will correspond to metrics such as average value, extreme value, rise rate, etc., and risk thresholds such as 70°C / 90°C. After the matching is completed, for each piece of data in the "real-time operation data set", based on the allocated real-time proof complexity, a zero-knowledge mapping proof strategy of corresponding strength is adopted for processing. This process first maps the original data to some kind of property assertion (such as "the value is less than the threshold", "the change rate is within the normal range", etc.), and then uses ZKP tools to generate proof information that does not expose the original value. For example, the system does not directly expose "current = 18.3A", but generates a proof through ZKP to show that the proposition "the current value is less than 20A" holds. This way abstracts the original sensitive data into a mathematical proposition and encrypts the mapping, thus effectively protecting privacy.
[0044] Through this step, refined control and dynamic adaptation of data privacy protection are achieved. Each data type corresponds to different identification metrics and thresholds, making the ZKP process more accurate and context-aware; at the same time, the proof method is automatically matched according to the real-time evaluated complexity, achieving a balance between security and computational efficiency.
[0045] Furthermore, as Figure 2 shown, according to the data identification metric set and the data judgment threshold set, zero-knowledge mapping proof is performed on the real-time operation data set according to the real-time proof complexity, including: randomly selecting a first predetermined data source, and obtaining the first real-time operation data, the first data identification metric, and the first data judgment threshold of the first predetermined data source, as well as the first monitoring data sequence within a preset historical time range; calculating the first real-time proof complexity according to the first monitoring data sequence; matching and obtaining the first proof method in the proof method decision maker according to the first real-time proof complexity; performing zero-knowledge proof on the first real-time operation data according to the first data identification metric and the first data judgment threshold according to the first proof method, generating the first proof data, and adding it to the proof data set.
[0046] Specifically, the first predetermined data source refers to an object randomly selected from multiple configured sensors or devices for performing zero-knowledge proof, such as a certain temperature sensor, voltage probe, etc. The first real-time operation data refers to the real-time data values obtained by the data source during the current acquisition cycle, such as the temperature being 85°C, the current being 16.5 A, etc. The first data identification index refers to an index extracted from a predefined index set for identifying data behavior characteristics, such as mean, maximum value, slope, stability coefficient, etc. The first data judgment threshold refers to the judgment boundary corresponding to the identification index, used to determine whether the data is in the normal, abnormal, or high-risk range. For example, the temperature upper limit is 80°C, and the current overload is 20 A. The first monitoring data sequence refers to the data sequence collected from the data source's history within a certain time range, used for behavior analysis, pattern recognition, and dynamic evaluation. The first real-time proof complexity refers to the proof strength level calculated through comprehensive analysis of historical data behavior, security posture, and privacy sensitivity, used to select an appropriate ZKP method. The proof method decision maker refers to a module used to match the corresponding zero-knowledge proof strategy according to the real-time complexity level. The first proof method refers to the specific ZKP technical solution selected by the decision maker and adapted to the current complexity. The first proof data is the result output by the ZKP process, used to prove the legality and compliance of the data without disclosing its original value.
[0047] When performing privacy-preserving calculations, a predetermined data source is randomly selected as the processing object for the current ZKP task, such as selecting "Sensor-A (temperature)". First, the latest real-time operation data is obtained from the device, and the identification indexes (such as volatility, change trend) and judgment thresholds associated with its type are synchronously retrieved.
[0048] Next, the data history module is called to obtain the monitoring data sequence collected from the sensor within a preset time window (such as the past 30 minutes, 1 hour), used to calculate the current real-time proof complexity of the data source. This complexity evaluation is based on: historical fluctuation conditions, abnormal ratio, sensor privacy level, current network security posture, etc., and is obtained through a mathematical model or scoring function.
[0049] After obtaining the complexity value, the system inputs this value into the "proof method decision maker", and this module matches the appropriate first proof method according to the complexity interval (such as low, medium, high). For example, low complexity corresponds to hash-based ZKP, medium complexity uses zk-SNARK, and high complexity selects zk-STARK.
[0050] Finally, the system constructs a zero-knowledge proposition based on the obtained "identification index" and "judgment threshold", such as "the current temperature value is less than 90°C", and then calls the selected ZKP engine (such as ZoKrates, libsnark, STARKy) to generate a zero-knowledge proof. The proof is encapsulated as the first proof data and added to the proof data set for subsequent storage, sharing or verification.
[0051] This step implements a dynamic and adaptive execution mechanism for data privacy protection. Through the integrated judgment of historical behavior, indicator thresholds and real-time security status, the system can flexibly determine the protection strength of each piece of data to avoid a "one-size-fits-all" approach. At the same time, verifiable proofs are generated through zero-knowledge mapping, which not only protects privacy but also ensures the verifiability and credibility of data. In addition, this method supports intelligent selection between diverse ZKP technologies, improves the system's performance adaptability, resource utilization efficiency and security policy accuracy in different data environments, and has wide practicality and feasibility in scenarios such as industrial data space, smart cities, and Internet of Vehicles.
[0052] Furthermore, a first real-time proof complexity is calculated based on the first monitoring data sequence, including: obtaining a first data privacy level of the first predetermined data source, and real-time security perception information of the industrial automatic production line, wherein the real-time security perception information is the frequency of network attacks within a preset time interval; performing data feature analysis based on the first monitoring data sequence to obtain a first data anomaly coefficient and a first data fluctuation coefficient; and calculating the first real-time proof complexity based on the first data privacy level, the network attack frequency, the first data anomaly coefficient and the first data fluctuation coefficient.
[0053] Specifically, the first data privacy level refers to the privacy sensitivity assessment result of the data source, which can be low, medium or high, and is defined according to the commercial secrets, user privacy or core technical information involved in the data content. Real-time security perception information refers to the security threat information monitored within a set time period (such as the past 15 minutes), such as network attack frequency, number of malicious requests, intrusion detection alarms, etc. Network attack frequency refers to the number of attack behaviors occurring in the network per unit time, which is used to quantify the network security risk level. The first data anomaly coefficient is used to indicate the proportion of data in the monitoring data sequence that exceeds the normal range, reflecting whether the data source has potential failures or the possibility of being tampered with. The first data fluctuation coefficient is used to reflect the stability of the data, usually calculated as standard deviation ÷ mean value. The larger the value, the more drastic the data fluctuation. The first real-time proof complexity refers to a dynamic complexity level calculated based on the above multiple factors, which is used to guide the selection and configuration of the zero-knowledge proof process.
[0054] Before protecting the privacy of a predetermined data source, we first evaluate whether the data requires a high-intensity zero-knowledge proof. To this end, we call the monitoring data sequence collected by the data source within the preset historical time interval to perform data behavior analysis.
[0055] Next, extract and calculate the following four factors to obtain data privacy level and security perception information, including: data privacy level, set through data tags, device policies or manual annotations, such as "temperature data is medium privacy" and "product images are high privacy"; network attack frequency, through security monitoring platforms (such as IDS, SIEM) or local gateway security log analysis, obtain the number of abnormal network behaviors within a certain period of time, such as scanning, injection, DDOS, etc. Calculate the anomaly coefficient and fluctuation coefficient, including: anomaly coefficient calculation, determine how much of the data does not meet the definition of the "judgment threshold set" (such as the number of times the temperature exceeds 80°C ÷ the total number of times); fluctuation coefficient calculation, perform mean and standard deviation analysis on historical sequences, and calculate "standard deviation ÷ mean" to obtain the fluctuation coefficient. The larger the value, the more unstable it is.
[0056] The above four indicators will be normalized and input into a weighted evaluation model (such as a linear model, fuzzy logic system, decision tree, etc.) to output a real-time proof complexity value. Among them, the weighted evaluation model comprehensively scores the real-time proof complexity by setting the weights of each influencing factor. First, the four core indicators - data privacy level, network attack frequency, data anomaly coefficient and data volatility coefficient are normalized to ensure dimensional consistency. Then, the weight coefficient of each indicator is set according to the actual scenario, for example, the privacy level and security situation are given a higher weight, and the total score is calculated by a linear weighted formula. Finally, the score is mapped to a complexity level range to guide the selection of ZKP strategies and realize a flexible and secure data privacy protection mechanism.
[0057] This step realizes a multi-dimensional comprehensive assessment of the characteristics, security risks and privacy attributes of the data itself. It no longer relies on static rules, but builds a real-time, dynamic complexity assessment system to achieve refined privacy protection. Its advantages include: automatically tightening the proof strength when the data fluctuates violently or the network environment is bad; using lightweight proofs to improve processing efficiency when the data is stable and the privacy level is low; avoiding resource waste, while strengthening the ability to transmit data trustably under the threat of attacks.
[0058] Further, perform data feature analysis on the first monitoring data sequence to obtain a first data anomaly coefficient and a first data fluctuation coefficient, including: calculating the proportion of abnormal data in the first monitoring data sequence to obtain the first data anomaly coefficient, where abnormal data are data that do not meet the first data judgment threshold; calculating a first data mean value and a first data standard deviation based on the first monitoring data sequence, and setting the ratio of the first data mean value to the first data standard deviation as the first data fluctuation coefficient.
[0059] Specifically, the first data mean value refers to the average value of the monitoring data sequence and is used to measure the general level of the data. The first data standard deviation refers to the degree to which each data value in the data sequence deviates from the mean value and is used to measure the volatility or stability of the data.
[0060] First, obtain the monitoring data sequence within its historical time period from a certain predetermined data source, such as the temperature values collected every minute in the past 30 minutes, a total of 30 pieces. Then, check these data item by item according to the preset judgment threshold, and count the number of data that do not meet the conditions, which are abnormal data.
[0061] Divide the number of abnormal data by the total number of samples to obtain the current anomaly coefficient of the data source. Subsequently, calculate the mean value and standard deviation of these 30 pieces of data. This process can be completed through data analysis. The edge computing node can use the NumPy or Pandas library in Python, and the embedded device can adopt the sliding window calculation logic implemented in C language, both of which support real-time processing. The anomaly coefficient and the fluctuation coefficient will be used as two key input factors for subsequent evaluation of the "real-time proof complexity" and directly affect the selected zero-knowledge proof strength.
[0062] Taking the current sensor of a certain device as an example, collect the current data within the past 10 minutes (a total of 20 pieces): Judgment threshold: The current should be between 5A and 15A; among the actual data, 4 pieces exceed The mean value of all data is 10A, Based on these two coefficients, it is initially judged that there are certain anomalies in the device data and it has medium volatility. Therefore, medium or high zero-knowledge proof complexity is assigned in the subsequent steps.
[0063] This step realizes the behavior-level analysis of the data state through feature extraction of historical monitoring data, and has strong practicality especially in data anomaly recognition and fluctuation sensitivity evaluation. Specifically, it can accurately identify potential anomaly sources and support early fault detection; dynamically quantify data stability and determine the selection of ZKP strength.
[0064] Further, the first real-time proof complexity is calculated based on the first data privacy level, network attack frequency, first data anomaly coefficient, and first data fluctuation coefficient, including: calculating the ratios of the first data privacy level to the historical maximum data privacy level, the network attack frequency to the historical maximum network attack frequency, the first data anomaly coefficient to the historical maximum first data anomaly coefficient, and the first data fluctuation coefficient to the historical maximum first data fluctuation coefficient, respectively, and calculating the first data proof scale by weighted calculation; multiplying the first data proof scale by the historical maximum first proof complexity to obtain the first real-time proof complexity.
[0065] Specifically, the historical maximum value (such as the maximum privacy level, maximum attack frequency, etc.) refers to the maximum reference value of the corresponding index that has appeared in the historical records and is used to standardize the current data state. The historical maximum first proof complexity refers to the maximum ZKP proof strength generated in history and can be regarded as the system capacity or the upper limit of the maximum computational complexity.
[0066] After analyzing the privacy level, network security situation, anomaly situation, and fluctuation situation of a certain data source, each index is normalized by taking the ratio with its corresponding historical maximum value. For example, if the current data privacy level is 0.8 and the historical highest is 1.0, then the ratio of this item is 0.8; if the current attack frequency is 5 times per minute and the historical highest is 10 times per minute, then the ratio is 0.5, and so on. According to the pre-configured weights (for example: privacy level 40%, attack frequency 30%, anomaly coefficient 20%, fluctuation coefficient 10%), these ratios are weighted and summed to obtain a comprehensive value between 0 and 1, which is called the first data proof scale. Then, this scale value is multiplied by the historical maximum first proof complexity recorded in the system (such as set to 1.0 or a specific complexity coefficient) to calculate the first real-time proof complexity. This value can be used as a decision input to select an appropriate zero-knowledge proof method.
[0067] Suppose the historical maximum values recorded by the system are as follows: maximum privacy level = 1.0; maximum network attack frequency = 10 times / minute; maximum anomaly coefficient = 0.6; maximum fluctuation coefficient = 0.3; maximum historical proof complexity = 1.0 (normalized upper limit). The current situation of a certain temperature data source is: current privacy level = 0.9; attack frequency = 6 times / minute; anomaly coefficient = 0.3; fluctuation coefficient = 0.15; then the four ratios are: privacy level ratio = 0.9 / 1.0 = 0.9; attack frequency ratio = 6 / 10 = 0.6; anomaly coefficient ratio = 0.3 / 0.6 = 0.5; fluctuation coefficient ratio = 0.15 / 0.3 = 0.5; Suppose the weights are: privacy (0.4), attack frequency (0.3), anomaly degree (0.2), fluctuation degree (0.1), then: the first data proof scale = 0.4×0.9 + 0.3×0.6 + 0.2×0.5 + 0.1×0.5 = 0 = 0.69. Finally, the first real-time proof complexity = 0.69×1.0 = 0.69, and this value is mapped to the "medium-high" level, and a medium-high strength zero-knowledge proof scheme such as zk-SNARK or enhanced STARK will be selected accordingly.
[0068] This step realizes the quantification, normalization and unified fusion evaluation of multi-source indicators, taking into account privacy, risk and data behavior characteristics, and improving the scientificity and controllability of the system's ZKP complexity configuration.
[0069] Furthermore, according to the first real-time proof complexity, a first proof method is matched and obtained in the proof method decision maker, including: configuring a proof method database, where the proof method database can be dynamically updated regularly; determining a proof complexity threshold based on historical monitoring data analysis, and dividing and determining several complexity intervals according to a predetermined complexity step size; constructing a proof method decision maker according to the proof method database and the several complexity intervals, and inputting the first real-time proof complexity into the proof method decision maker to match and obtain the first proof method.
[0070] Specifically, the proof method database refers to a set of available zero-knowledge proof technical solutions and their metadata (such as algorithm name, calculation overhead, applicable scope, etc.) maintained in the system. The complexity threshold is used to divide the boundary values of the proof complexity levels. For example, "below 0.3 is low complexity", "0.3 - 0.7 is medium", "above 0.7 is high complexity". The complexity step size is used to define the minimum unit that divides the entire complexity space (0 to 1) into several intervals, such as 0.1, 0.2, etc., which is convenient for subsequent mapping and matching. The complexity interval refers to several continuous sections divided by the threshold or step size. For example: [0, 0.3), [0.3, 0.7), [0.7, 1.0], which is used to establish the "complexity - proof method" mapping relationship.
[0071] First, a proof method database is built and recorded, which records all supported zero-knowledge proof methods and related information, including name, applicable complexity range, operating performance, algorithm type, encryption strength, computing resource consumption, etc. In order to maintain advancement and flexibility, the database supports dynamic updates and can be regularly expanded or replaced according to the release of new algorithms or changes in security requirements. Specifically,
[0072] According to the currently supported zero-knowledge proof protocol types (such as zk-SNARK, zk-STARK, Bulletproof, Plonk, etc.), its core parameter information is collected, including computational complexity, encryption strength, verification efficiency, applicable scenarios, etc. Secondly, set adaptation conditions for each proof method, such as the complexity range to which it is adapted, the recommended data type or the resource consumption level. Subsequently, a unified data table structure is constructed to store various attributes in a standardized manner, and a dynamic update mechanism is supported, so that new ZKP schemes can be regularly accessed or weight strategies can be adjusted. Finally, a "complexity range-proof method" mapping table is established in combination with the system's historical operation data and security policies for real-time call by the proof method decision maker. The entire database can be deployed on the cloud or edge nodes, and combined with the API interface to provide flexible access methods to achieve intelligent and controllable scheduling of proof methods. The flexible expansion, fine adaptation and rapid response of the ZKP strategy are guaranteed, which helps to improve the overall privacy protection efficiency of the system.
[0073] Next, based on historical operation data and security policies, the system sets several complexity thresholds for the "real-time proof complexity" or divides the complexity intervals into equal steps. Then, the "first real-time proof complexity" calculated in real time is used as input to the "proof method decision maker", which matches according to the interval mapping relationship and returns the corresponding "first proof method".
[0074] This step enables the zero-knowledge proof process to move from static configuration to dynamic scheduling and intelligent matching. The core advantages include: intelligently selecting the most suitable data protection method based on real-time complexity; dynamically expanding the ZKP solution to improve the system's adaptability to new technologies; avoiding resource waste caused by "over-protection" and privacy leakage risks caused by "insufficient protection"; achieving a balance between high security, low resource consumption and rapid response capabilities; supporting refined decision-making logic (such as adaptation by data type and scenario) to provide a basis for large-scale applications.
[0075] Further, a proof method decision maker is constructed based on the proof method database and the several complexity intervals, including: evaluating the complexity of multiple proof methods in the proof method database to determine multiple proof complexities; dividing the multiple proof complexities according to the several complexity intervals, and mapping to determine several proof method sets; based on a decision tree, constructing the proof method decision maker according to the mapping relationship between the complexity intervals and the proof method sets and the several complexity intervals and the several proof method sets.
[0076] Specifically, the multiple proof methods refer to various ZKP algorithm schemes stored in the database, such as zk-SNARK, zk-STARK, Bulletproof, Plonk, Sigma protocol, etc. The proof complexity refers to a comprehensive performance index used to measure the computing resources, time overhead, proof data size, etc. required for each proof method during execution, usually represented in numerical form for easy comparison. The complexity interval is used to divide the real-time proof complexity value into several hierarchical intervals, and the proof method set refers to a set of optional proof methods corresponding to each complexity interval, that is, a subset of the "multiple proof methods".
[0077] First, a quantitative evaluation needs to be carried out on each zero-knowledge proof method in the existing proof method database. The evaluation contents include: generation time, verification time, proof size, computing resource consumption, security level, etc. After standardization processing, they are summarized into a proof complexity value, which is used to represent the overhead and capabilities of the method itself. Next, according to the preset complexity interval division rules (such as dividing into 5 segments with a step size of 0.2), the complexity values of each proof method are divided into the corresponding intervals to establish the mapping relationship of complexity interval → proof method set. For example: the interval [0.0–0.2): includes lightweight Hash-based ZKP; the interval [0.2–0.5): includes Bulletproof; the interval [0.5–0.8): includes zk-SNARK; the interval [0.8–1.0]: includes zk-STARK and Plonk. Then, based on the above intervals and the corresponding proof method sets, a decision tree model is constructed: the branch nodes of the tree represent "whether the real-time proof complexity falls within a certain interval", and the leaf nodes represent the specific "recommended proof method set". Whenever a new real-time proof complexity is used as input, judge its interval from the top layer of the decision tree in turn, and quickly output the matching ZKP method.
[0078] The "proof method decision maker" constructed by this method achieves the following key goals: structuring and automating the ZKP selection process, reducing human intervention; realizing the efficient mapping and quick response of "real-time complexity → best ZKP scheme"; supporting the flexible expansion and dynamic update of proof methods to adapt to the development of new technologies;
[0079] In summary, the privacy protection method for a trusted data space based on zero-knowledge proof provided by the embodiments of the present application has the following technical effects:
[0080] 1. By constructing a privacy protection and data trusted sharing mechanism through zero-knowledge proof and blockchain technology, it ensures the secure verification of sensitive data without revealing the content, enhancing the security, compliance, and verifiability of the data space.
[0081] 2. Based on the accurate identification and evaluation of data characteristics from the background information of the data source, and dynamically selecting the zero-knowledge proof method in combination with the real-time complexity, it effectively improves the matching degree and execution efficiency of the privacy protection strategy.
[0082] 3. Introducing a multi-dimensional feature analysis mechanism to comprehensively evaluate data privacy, security risks, anomaly levels, and volatility characteristics, effectively realizing the dynamic calculation of ZKP complexity and risk-sensitive response, and improving the privacy protection accuracy.
[0083] Embodiment 2
[0084] Based on the same inventive concept as the privacy protection method for a trusted data space based on zero-knowledge proof in the foregoing embodiments, as Figure 3 shown, the embodiments of the present application provide a privacy protection system for a trusted data space based on zero-knowledge proof, and the system includes:
[0085] An operating data acquisition module 11, configured to collect real-time operating data of multiple data sources through the Internet of Things, and obtain a real-time operating data set and a data source information set; a proof data set generation module 12, configured to perform zero-knowledge proof on the real-time operating data set according to the real-time proof complexity based on the data source information set, and generate a proof data set, where the real-time proof complexity is determined by comprehensive analysis of security perception information, data privacy degree, data anomaly degree, and data volatility degree; a data encryption storage module 13, configured to use blockchain technology to perform, and perform data sharing and data verification according to the encrypted storage data set on the proof data set.
[0086] Further, the operating data acquisition module 11 is further configured to perform the following steps: through the Internet of Things, at a predetermined monitoring time node, collect data from multiple predetermined data sources of an industrial automatic production line, and perform preliminary encryption processing on the data to obtain a real-time operating data set; through the Internet of Things, synchronously collect the data source ID, data type, location information, and time stamp of each predetermined data source to obtain a data source information set.
[0087] Further, the proof data set generation module 12 is further configured to perform the following steps: match and call a data identification metric set and a data judgment threshold set based on the data source information set; perform a zero-knowledge mapping proof on the real-time operation data set according to the data identification metric set and the data judgment threshold set according to the real-time proof complexity, and generate the proof data set.
[0088] Further, the proof data set generation module 12 is further configured to perform the following steps: randomly select a first predetermined data source, and obtain first real-time operation data, a first data identification metric, and a first data judgment threshold of the first predetermined data source, as well as a first monitoring data sequence within a preset historical time range; calculate a first real-time proof complexity according to the first monitoring data sequence; match and obtain a first proof method in the proof method decision maker according to the first real-time proof complexity; perform a zero-knowledge proof on the first real-time operation data according to the first data identification metric and the first data judgment threshold according to the first proof method, generate first proof data, and add it to the proof data set.
[0089] Further, the proof data set generation module 12 is further configured to perform the following steps: obtain a first data privacy level of the first predetermined data source and real-time security perception information of the industrial automatic production line, where the real-time security perception information is the network attack frequency within a preset time interval; perform data feature analysis according to the first monitoring data sequence to obtain a first data anomaly coefficient and a first data fluctuation coefficient; calculate a first real-time proof complexity according to the first data privacy level, the network attack frequency, the first data anomaly coefficient, and the first data fluctuation coefficient.
[0090] Further, the proof data set generation module 12 is further configured to perform the following steps: calculate the proportion of abnormal data in the first monitoring data sequence to obtain a first data anomaly coefficient, where the abnormal data is data that does not meet the first data judgment threshold; calculate a first data mean and a first data standard deviation according to the first monitoring data sequence, and set the ratio of the first data mean to the first data standard deviation as the first data fluctuation coefficient.
[0091] Further, the proof data set generation module 12 is further configured to perform the following steps: calculate the ratios of the first data privacy level to the historical maximum data privacy level, the network attack frequency to the historical maximum network attack frequency, the first data anomaly coefficient to the historical maximum first data anomaly coefficient, and the first data fluctuation coefficient to the historical maximum first data fluctuation coefficient respectively, and calculate a first data proof scale by weighting; multiply the first data proof scale by the historical maximum first proof complexity to obtain the first real-time proof complexity.
[0092] Further, the proof data set generation module 12 is further configured to perform the following steps: configure a proof method database, where the proof method database can be dynamically updated regularly; determine a proof complexity threshold based on historical monitoring data analysis, and divide and determine several complexity intervals according to a predetermined complexity step size; construct a proof method decision maker according to the proof method database and the several complexity intervals, and input the first real-time proof complexity into the proof method decision maker to match and obtain a first proof method.
[0093] Further, the proof data set generation module 12 is further configured to perform the following steps: evaluate the complexity of multiple proof methods in the proof method database to determine multiple proof complexities; divide the multiple proof complexities according to the several complexity intervals, and map and determine several proof method sets; based on a decision tree, construct the proof method decision maker according to the mapping relationship between the complexity intervals and the proof method sets and the several complexity intervals and the several proof method sets.
[0094] Any step of the method described above can be stored as a computer instruction or program in an unrestricted computer memory and can be called and recognized by an unrestricted computer processor to implement any method in the embodiments of the present application, and no redundant restrictions are imposed here.
[0095] Further, the first or second mentioned above does not only represent an order relationship, but also represents a specific concept, and / or means that multiple elements can be selected individually or in whole. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these modifications and variations.
Claims
1. A privacy protection method for a trusted data space based on zero-knowledge proof, characterized in that, The method includes: Collecting real-time operation data of multiple data sources through the Internet of Things to obtain a real-time operation data set and a data source information set; Based on the data source information set, performing zero-knowledge proof on the real-time operation data set according to the real-time proof complexity to generate a proof data set, where the real-time proof complexity is determined by comprehensive analysis of security perception information, data privacy level, data anomaly degree, and data fluctuation degree; Using blockchain technology to encrypt and store the proof data set, and performing data sharing and data verification based on the encrypted storage data set.
2. The privacy protection method for a trusted data space based on zero-knowledge proof according to claim 1, wherein, Collecting real-time operation data of multiple data sources through the Internet of Things to obtain a real-time operation data set and a data source information set, including: Through the Internet of Things, at a predetermined monitoring time node, collecting data from multiple predetermined data sources of an industrial automatic production line and performing preliminary encryption processing on the data to obtain a real-time operation data set; Through the Internet of Things, synchronously collecting the data source ID, data type, location information, and time stamp of each predetermined data source to obtain a data source information set.
3. The privacy protection method for a trusted data space based on zero-knowledge proof according to claim 1, wherein Based on the data source information set, performing zero-knowledge proof on the real-time operation data set according to the real-time proof complexity to generate a proof data set, including: Based on the data source information set, matching and invoking a data recognition index set and a data judgment threshold set; According to the data recognition index set and the data judgment threshold set, performing zero-knowledge mapping proof on the real-time operation data set according to the real-time proof complexity to generate the proof data set.
4. The privacy protection method for a trusted data space based on zero-knowledge proof according to claim 3, wherein Performing zero-knowledge mapping proof on the real-time operation data set according to the data recognition index set and the data judgment threshold set according to the real-time proof complexity, including: Randomly selecting a first predetermined data source, and obtaining the first real-time operation data, the first data recognition index, and the first data judgment threshold of the first predetermined data source, as well as the first monitoring data sequence within a preset historical time range; Calculating a first real-time proof complexity based on the first monitoring data sequence; Matching and obtaining a first proof method in a proof method decision maker according to the first real-time proof complexity; According to the first data recognition index and the first data judgment threshold, performing zero-knowledge proof on the first real-time operation data according to the first proof method to generate a first proof data, and adding it to the proof data set.
5. The privacy protection method for a trusted data space based on zero-knowledge proof according to claim 4, wherein Calculating a first real-time proof complexity based on the first monitoring data sequence, including: Obtaining the first data privacy level of the first predetermined data source and the real-time security perception information of the industrial automatic production line, where the real-time security perception information is the network attack frequency within a preset time interval; Performing data feature analysis on the first monitoring data sequence to obtain a first data anomaly coefficient and a first data fluctuation coefficient; Calculating a first real-time proof complexity based on the first data privacy level, the network attack frequency, the first data anomaly coefficient, and the first data fluctuation coefficient.
6. The privacy protection method for a trusted data space based on zero-knowledge proof according to claim 5, characterized in that Performing data feature analysis on the first monitoring data sequence to obtain a first data anomaly coefficient and a first data fluctuation coefficient, including: Calculate the proportion of abnormal data in the first monitoring data sequence to obtain the first data anomaly coefficient, where the abnormal data is data that does not meet the first data judgment threshold. Calculate the first data mean and the first data standard deviation based on the first monitoring data sequence, and set the ratio of the first data mean to the first data standard deviation as the first data fluctuation coefficient.
7. The privacy protection method for a trusted data space based on zero-knowledge proof according to claim 5, characterized in that Calculate the first real-time proof complexity based on the first data privacy level, the network attack frequency, the first data anomaly coefficient, and the first data fluctuation coefficient, including: Calculate the ratios of the first data privacy level to the historical maximum data privacy level, the network attack frequency to the historical maximum network attack frequency, the first data anomaly coefficient to the historical maximum first data anomaly coefficient, and the first data fluctuation coefficient to the historical maximum first data fluctuation coefficient respectively, and calculate the first data proof scale by weighted calculation. Multiply the first data proof scale by the historical maximum first proof complexity to obtain the first real-time proof complexity.
8. The method for protecting the privacy of a trusted data space based on zero-knowledge proof according to claim 4, wherein Match and obtain the first proof method in the proof method decision maker according to the first real-time proof complexity, including: Configure the proof method database, where the proof method database can be dynamically updated regularly. Determine the proof complexity threshold based on historical monitoring data analysis, and divide and determine several complexity intervals according to a predetermined complexity step size. Construct a proof method decision maker according to the proof method database and the several complexity intervals, and input the first real-time proof complexity into the proof method decision maker to match and obtain the first proof method.
9. The privacy protection method for a trusted data space based on zero-knowledge proof according to claim 8, wherein Construct a proof method decision maker according to the proof method database and the several complexity intervals, including: Evaluate the complexity of multiple proof methods in the proof method database to determine multiple proof complexities. Divide the multiple proof complexities according to the several complexity intervals, and map and determine several proof method sets. Based on the decision tree, construct the proof method decision maker according to the several complexity intervals and the several proof method sets according to the mapping relationship between the complexity intervals and the proof method sets.
10. A trusted data space privacy protection system based on zero-knowledge proof, characterized in that, For implementing the zero-knowledge proof-based trusted data space privacy protection method according to any one of claims 1 to 9, the system includes: An operation data collection module, configured to collect real-time operation data of multiple data sources through the Internet of Things to obtain a real-time operation data set and a data source information set. A proof data set generation module, configured to perform zero-knowledge proof on the real-time operation data set according to the real-time proof complexity based on the data source information set to generate a proof data set, where the real-time proof complexity is determined by comprehensive analysis of security perception information, data privacy, data anomaly, and data fluctuation. A data encryption storage module, configured to use blockchain technology to perform data sharing and data verification on the encrypted storage data set according to the proof data set.
Citation Information
Patent Citations
Block chain privacy data sharing method based on zero knowledge proof
CN114499900A
Electric power information security storage supervision system based on block chain
CN119150347A
Private data protection method and system based on homomorphic encryption and federated learning
CN119513919A
Zero knowledge prover
US20230267195A1
Cited By
Internet of Things data intelligent sharing method and system
CN121126475A
Data processing method and system for anonymous person attribute abundance optimization
CN121211510A
A data processing method and system for anonymous person attribute abundance optimization
CN121211510B
Synchronous governance method for credible version of carbon factor
CN122390766A