Trusted data space privacy protection method and system based on zero-knowledge proof

By collecting data through the Internet of Things and using zero-knowledge proof and blockchain technology for encrypted storage, the problem of sensitive information leakage in multi-source heterogeneous data sharing is solved, and efficient and reliable data sharing and verification are achieved.

CN120257329BActive Publication Date: 2025-10-17LINGSHU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510410702.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-10-17
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

Existing technologies lack dynamic privacy protection mechanisms in the process of multi-source heterogeneous data sharing and verification, which makes sensitive information easy to leak and makes trusted verification difficult.

Method used

Multi-source data is collected through the Internet of Things, zero-knowledge proof is performed based on real-time proof complexity, and blockchain is used for encrypted storage and verification to generate a zero-knowledge proof data set to achieve dynamic privacy protection.

Benefits of technology

Under the premise of ensuring data privacy, efficient, reliable and secure data sharing and verification are achieved, which improves the security and credibility of the data space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257329B_ABST
    Figure CN120257329B_ABST
Patent Text Reader

Abstract

The application discloses a trusted data space privacy protection method and system based on zero-knowledge proof, and belongs to the technical field of data processing. The method comprises the following steps: collecting real-time running data of multiple data sources through an Internet of Things, and obtaining a real-time running data set and a data source information set; performing zero-knowledge proof on the real-time running data set according to real-time proof complexity based on the data source information set, and generating a proof data set, wherein the real-time proof complexity is determined based on comprehensive analysis of security perception information, data privacy degree, data anomaly degree and data fluctuation degree; and using a block chain technology to encrypt and store the proof data set, and performing data sharing and data verification according to the encrypted and stored data set. The application solves the technical problem that there is a lack of a dynamic privacy protection mechanism for multi-source data in a sharing process in the prior art, so that sensitive data is prone to be leaked, and trusted verification is difficult to implement.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of data processing, and particularly relates to a trusted data space privacy protection method and system based on zero-knowledge proof. BACKGROUND

[0002] Massive heterogeneous data is collected and widely applied in intelligent manufacturing, smart city and other scenarios in real time, and how to realize trusted sharing under the premise of guaranteeing data security and privacy has become a core challenge in the construction of a current data space trusted system.

[0003] Existing data privacy protection methods mostly rely on static encryption or unified level privacy policies, and it is difficult to realize dynamic and differentiated processing according to data sensitivity, abnormal conditions and network risks, especially in the data sharing and verification process, and there is a lack of flexible and controllable zero-knowledge proof strategies with efficiency and security. Therefore, it is urgent to build a dynamic privacy protection mechanism based on multi-dimensional complexity evaluation to realize accurate matching and trusted proof of different data scenarios. SUMMARY

[0004] The application provides a trusted data space privacy protection method and system based on zero-knowledge proof, aiming to solve the technical problems in the prior art that there is a lack of dynamic privacy protection mechanism for multi-source data in the sharing process, which leads to easy leakage of sensitive data and difficulty in realizing trusted verification.

[0005] In view of the above problems, the application provides a trusted data space privacy protection method and system based on zero-knowledge proof.

[0006] The first aspect of the application discloses a trusted data space privacy protection method based on zero-knowledge proof, which comprises the following steps: collecting real-time running data of a plurality of data sources through an Internet of Things, obtaining a real-time running data set and a data source information set; based on the data source information set, performing zero-knowledge proof on the real-time running data set according to real-time proof complexity to generate a proof data set, wherein the real-time proof complexity is determined based on comprehensive analysis of security perception information, data privacy degree, data abnormality degree and data fluctuation degree; and using a block chain technology to encrypt and store the proof data set, and performing data sharing and data verification according to the encrypted and stored data set.

[0007] In another aspect of the present application, a trusted data space privacy protection system based on zero-knowledge proof is provided, which comprises a running data collection module for collecting real-time running data of multiple data sources through an Internet of Things, obtaining a real-time running data set and a data source information set; a proof data set generation module for performing zero-knowledge proof on the real-time running data set according to real-time proof complexity based on the data source information set, generating a proof data set, wherein the real-time proof complexity is determined based on comprehensive analysis of security perception information, data privacy degree, data anomaly degree and data fluctuation degree; and a data encryption storage module for performing data sharing and data verification according to an encrypted storage data set by using blockchain technology on the proof data set.

[0008] The one or more technical solutions provided in the present application have at least the following technical effects or advantages:

[0009] Due to the technical solution of collecting multi-source data through an Internet of Things, performing zero-knowledge proof based on real-time proof complexity, and performing encryption storage and verification by using blockchain, the technical problem of lack of dynamic privacy protection mechanism in the sharing and verification process of multi-source heterogeneous data in the prior art is solved, thereby sensitive information is prevented from being leaked and trusted verification is facilitated, and the technical effect of efficient, trusted and secure sharing under the premise of protecting data privacy is achieved.

[0010] The above description is only a summary of the technical solutions of the present application, in order to more clearly understand the technical means of the present application, the specific embodiments of the present application can be implemented according to the content of the description, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 A flowchart of a trusted data space privacy protection method based on zero-knowledge proof is provided for the embodiments of the present application;

[0012] Figure 2 A flowchart of zero-knowledge mapping proof of a real-time running data set in a trusted data space privacy protection method based on zero-knowledge proof is provided for the embodiments of the present application;

[0013] Figure 3 A structural diagram of a trusted data space privacy protection system based on zero-knowledge proof is provided for the embodiments of the present application.

[0014] Explanation of reference signs: running data collection module 11, proof data set generation module 12, data encryption storage module 13. DETAILED DESCRIPTION

[0015] The general idea of the technical solutions provided in the present application is as follows:

[0016] The embodiment of the application provides a trusted data space privacy protection method and system based on zero-knowledge proof. First, various types of data source information are collected and identified, then the real-time proof complexity is proved according to the data sensitivity and the security situation, the data is calculated by the matched zero-knowledge proof method, and finally the proof result is encrypted and stored in the block chain, so that the data is safely shared and trustedly verified without leaking the original content.

[0017] After introducing the basic principle of the application, the various non-limiting embodiments of the application will be specifically introduced in combination with the drawings of the specification.

[0018] Embodiment one

[0019] As shown in the figure, the embodiment of the application provides a trusted data space privacy protection method based on zero-knowledge proof, which comprises: Figure 1

[0020] Step S100: Collecting real-time running data of multiple data sources through the Internet of Things, obtaining a real-time running data set and a data source information set.

[0021] Specifically, the Internet of Things refers to a network system that connects various physical devices through sensors, communication devices and network protocols to realize the perception, collection, transmission and processing of information. The data source refers to a device or system that can provide real-time data. The real-time running data set refers to a data set collected from each data source at a specific time node, describing the running state of the device or environmental information, such as temperature, current, yield, etc. The data source information set refers to the metadata corresponding to each data, including the ID, type, location and collection timestamp of the data source.

[0022] First, a variety of types of Internet of Things sensors will be laid out in the industrial scene, such as temperature and humidity sensors, current detectors, vibration monitoring modules, etc. Each device is assigned a unique ID for data identification. The Internet of Things gateway or edge computing node is responsible for waking up these sensors regularly, collecting real-time data at the set monitoring time node, and sending these data to the data processing center through wireless communication protocols (such as MQTT, LoRa, NB-IoT, etc.).

[0023] At the same time of collection, the system will also record the relevant information of each data source synchronously, including the device ID, data type, collection location and timestamp, to form the data source information set. In addition, to prevent data from being tampered with or leaked during transmission, the collection end usually performs preliminary encryption processing through a lightweight encryption algorithm.

[0024] ​Through this step, not only is the unified and accurate data collection of multi-source devices achieved, but also a solid data foundation is laid for subsequent data trusted computing and privacy protection. Especially under the premise of traceable data sources and clear types, the system can dynamically adjust the privacy protection strategy according to the sensitivity and abnormal characteristics of different devices, thereby significantly improving the intelligence and security of the data space.

[0025] Step S200: Based on the data source information set, a zero-knowledge proof is performed on the real-time running data set according to a real-time proof complexity, and a proof data set is generated, wherein the real-time proof complexity is determined based on comprehensive analysis of security perception information, data privacy degree, data abnormality degree and data fluctuation degree.

[0026] Specifically, zero-knowledge proof (Zero-Knowledge Proof, ZKP) is a cryptographic technique used to prove to a third party that a certain data property or state is true without exposing the original data. Common forms include zk-SNARK, zk-STARK, etc. The proof data set refers to the data structure generated after ZKP processing, which contains corresponding proof information without revealing the original data, and is used for subsequent verification, evidence storage or sharing. Real-time proof complexity refers to the complexity level dynamically calculated according to the current security state, privacy sensitivity and abnormal risk of the data, which determines the strength and strategy of ZKP. Security perception information refers to the monitoring information of the current network or system security state, such as network attack frequency, virus scan results, access anomalies, etc. Data privacy degree is used to describe the sensitivity level of data, such as ordinary data, business secrets, user privacy information, etc. Data abnormality degree is used to reflect the degree of data deviation from the normal interval, and the higher the proportion of abnormal data, the greater the abnormality degree. Data fluctuation degree is used to measure the stability of data within a period of time, often expressed in the form of "standard deviation / mean".

[0027] After collecting the "real-time running data set" and the corresponding "data source information set", first, comprehensive security and privacy analysis is performed on each data source. Specifically, starting from four dimensions: current network attack frequency, data sensitivity level, abnormality degree and fluctuation degree compared with historical data. These information can be obtained through historical data model, threshold rule library and behavior analysis algorithm (such as Z-score analysis, sliding window).

[0028] Then, after quantifying the four indicators, a numerical value, namely "real-time proof complexity", is output through weighted calculation or a neural network model. This complexity value will be input into a "proof method decision maker" to match the appropriate proof method from a pre-configured ZKP strategy library. For example, when the complexity is high, zk-STARK can be selected; when the complexity is low, a lightweight hash-based ZKP can be selected. After selecting the proof method, real-time data is processed for zero-knowledge mapping, such as generating a zero-knowledge proof for the fact that "the temperature is within the specified range" rather than exposing the specific temperature value. This proof information is added to the "proof data set" for subsequent uploading, verification, or on-chain operation. The entire process is completed by data analysis (such as SparkStreaming), encryption libraries (such as libsnark, ZoKrates), and security engines in collaboration, deployed on edge nodes or cloud platforms.

[0029] This step realizes differentiated privacy protection and dynamic allocation of computing resources: higher protection is provided for high-risk, high-privacy data to prevent critical data leakage; lightweight solutions are used for low-risk data to improve system processing efficiency. In addition, the real-time complexity dynamic evaluation mechanism also has certain security adaptive ability, which can adjust the ZKP strategy in time according to external attack situation and data abnormality, greatly enhancing the credibility, attack resistance, and compliance of the data space.

[0030] Step S300: Using blockchain technology, the proof data set is encrypted and stored, and data sharing and data verification are performed based on the encrypted storage data set.

[0031] Specifically, encrypted storage refers to encrypting data before writing it to the blockchain to ensure that even if the data on the chain is read, it cannot be identified. Data sharing refers to providing the proof data set to other users or systems within the authorized range for them to verify the authenticity and compliance of the data without accessing the original data content. Data verification refers to verifying the legality of shared data by reading the encrypted proof data on the chain and combining it with the zero-knowledge verification algorithm to ensure that the data has not been tampered with and meets the business logic requirements.

[0032] After completing the zero-knowledge proof of real-time data, the generated proof data set is encrypted. Specifically, symmetric encryption algorithms or asymmetric encryption algorithms can be used, or a hash function can be used to generate a digest value for evidence. After encryption, the data is packaged into a block structure and written to the blockchain through a consensus mechanism (such as PoW, PoS, or PBFT) to ensure its tamper resistance and timestamp authenticity.

[0033] After storage is completed, the proof data set can realize data sharing on the chain. Any verification party with access permission can obtain the corresponding encrypted proof data from the chain, and combine the original zero-knowledge verification algorithm (such as zk-SNARK verifier) to complete the authenticity and compliance verification of the data without decrypting the data content. Data verification logic can be automatically executed through a smart contract to realize data access control and permission management. At the same time, all operations are recorded on the chain, which is traceable, ensuring the transparency and compliance of the sharing process.

[0034] Through this step, secure sharing and trusted verification are realized without revealing the original data. The encryption processing and zero-knowledge mechanism effectively block the leakage of sensitive information; the blockchain ensures the data's immutability and traceability.

[0035] Further, through the Internet of Things, real-time running data of multiple data sources is collected to obtain a real-time running data set and a data source information set, including: through the Internet of Things, data of multiple predetermined data sources of an industrial automatic production line is collected at a predetermined monitoring time node, and the data is preliminarily encrypted to obtain a real-time running data set; through the Internet of Things, the data source ID, data type, location information, and timestamp of each predetermined data source are synchronously collected to obtain a data source information set.

[0036] Specifically, the industrial automatic production line refers to a production line that realizes continuous operation through automated equipment (such as sensors, PLC controllers, mechanical arms, etc.) in manufacturing, usually with high data and intelligent characteristics. The predetermined data source refers to a device or sensor node in the system that has been configured and specified to be monitored and data collected, such as a temperature sensor, a current monitor, a production counter, etc. The predetermined monitoring time node is used to ensure the timeliness and synchronization of the data, and the system triggers the collection operation at a preset time interval or time point, such as every 10 seconds or every whole hour. The data source ID refers to a number that uniquely identifies each data source, such as "TEMP-001" representing the 1st temperature sensor. The location information is used to represent the geographical or logical position of the data source in the industrial production line, such as "North packaging line" or "Workshop No. 3".

[0037] Multiple types of sensors and intelligent devices are deployed on the industrial automatic production line as predetermined data sources, and the collection interval or trigger time point of each device is set in the configuration. For example, a temperature sensor collects data every 30 seconds, while a vibration sensor collects data only when an anomaly is detected. The system sends data collection instructions to these sensors synchronously at each predetermined monitoring time node through the Internet of Things gateway or embedded controller deployed at the edge. Each data source responds to collect corresponding real-time running data such as temperature values and current values, and immediately performs preliminary encryption processing locally to ensure the security of the data during transmission.

[0038] At the same time, the source information of each data point is collected simultaneously, including the data source ID, type, location, and timestamp. This information constitutes the "data source information set." This information can be provided by the sensor itself or generated by the system. It is then uploaded to the data processing platform or cloud as metadata along with the encrypted data.

[0039] The entire process coordinates device communication and data uploads through IoT platform software (such as MQTT middleware and OPC UA servers), and implements edge encryption through lightweight encryption libraries (such as wolfSSL and mbedTLS). This ultimately results in two key data structures: a real-time operational dataset and a data source information set.

[0040] This step enables the unified collection, identification, and encryption of multi-type, heterogeneous industrial data. This ensures data integrity and consistency, providing high-quality basic data for subsequent privacy-focused computing and trusted verification.

[0041] Furthermore, based on the data source information set, zero-knowledge proof is performed on the real-time running data set according to the real-time proof complexity to generate a proof data set, including: matching and calling the data identification indicator set and the data judgment threshold set based on the data source information set; based on the data identification indicator set and the data judgment threshold set, zero-knowledge mapping proof is performed on the real-time running data set according to the real-time proof complexity to generate the proof data set.

[0042] Specifically, a data identification indicator set refers to a collection of quantitative metrics used to judge and analyze data status, characteristics, or behavior, such as maximum, minimum, mean, growth rate, trend, and abnormal pattern recognition values. A data judgment threshold set refers to the judgment limits associated with the identification indicators, used to define whether data is in an abnormal state, a safe state, or a critical condition requiring encryption. For example, the safety threshold of a temperature sensor is 80°C, and the current overload threshold is 20A. Zero-knowledge mapping proof refers to the process of converting raw data into a zero-knowledge proof through mathematical mapping without revealing its specific content, such as proving that "the temperature is within a safe range" without revealing the actual temperature value.

[0043] After obtaining the data source information set, the metadata such as data source ID, type, location, etc. are used to match the corresponding "data identification index set" and "data judgment threshold set". These indexes and thresholds are pre-stored in the policy library and customized for different types of data sources (such as temperature, current, image, etc.). For example, temperature data corresponds to indexes such as average value, extreme value, and rising rate, and risk thresholds such as 70°C / 90°C. After matching, for each piece of data in the "real-time running data set", based on the assigned real-time proof complexity, the corresponding intensity of zero-knowledge mapping proof strategy is used for processing. This process first maps the original data into some attribute assertion (such as "numerical value less than threshold value", "change rate within normal range", etc.), and then uses ZKP tools to generate proof information that does not expose the original numerical value. For example, the system does not directly expose "current = 18.3A", but generates a proof that "the current value is less than 20A" through ZKP. This way abstracts the original sensitive data into mathematical propositions and encrypts the mapping, effectively protecting privacy.

[0044] Through this step, fine-grained control and dynamic adaptation of data privacy protection are realized. Different identification indexes and thresholds correspond to different data types, making the ZKP process more accurate and context-aware; at the same time, according to the complexity evaluated in real time, the proof method is automatically matched, realizing the balance between security and computational efficiency.

[0045] Further, as shown in Figure 2 According to the data identification index set and the data judgment threshold set, the real-time running data set is subjected to zero-knowledge mapping proof according to the real-time proof complexity, including: randomly selecting a first predetermined data source, and obtaining first real-time running data of the first predetermined data source, a first data identification index, and a first data judgment threshold, as well as a first monitoring data sequence in a preset historical time range; calculating a first real-time proof complexity according to the first monitoring data sequence; matching to obtain a first proof method in a proof method decision maker according to the first real-time proof complexity; performing zero-knowledge proof on the first real-time running data according to the first proof method according to the first data identification index and the first data judgment threshold, to generate first proof data and add it to the proof data set.

[0046] In particular, the first predetermined data source refers to a randomly selected object from a plurality of configured sensors or devices for performing zero-knowledge proof, such as a certain temperature sensor, voltage probe, etc. The first real-time running data refers to the real-time data value obtained by the data source in the current collection period, such as a temperature of 85°C, a current of 16.5A, etc. The first data identification index refers to an index extracted from a pre-defined index set for identifying data behavior characteristics, such as mean, maximum, slope, stability coefficient, etc. The first data judgment threshold refers to a judgment limit corresponding to the identification index, used to determine whether the data is in a normal, abnormal or high-risk interval. For example, the upper limit of temperature is 80°C, and the current overload is 20A. The first monitoring data sequence refers to the data sequence collected from the data source in a certain time range, used for behavior analysis, pattern recognition and dynamic evaluation. The first real-time proof complexity refers to the proof strength level calculated based on comprehensive analysis of historical data behavior, security situation and privacy sensitivity, used to select a suitable ZKP method. The proof method decision maker refers to a module for matching the corresponding zero-knowledge proof strategy according to the real-time complexity level. The first proof method refers to the specific ZKP technical solution selected by the decision maker to adapt to the current complexity. The first proof data is the result output by the ZKP process, used to prove the legitimacy and compliance of the data without revealing its original value.

[0047] In performing privacy protection calculation, a predetermined data source is randomly selected as the processing object of the current ZKP task, such as "Sensor-A (temperature)". First, the latest real-time running data of the device is obtained, and the identification index (such as volatility, change trend) and judgment threshold associated with its type are synchronously called.

[0048] Next, the data history module is called to obtain the monitoring data sequence collected from the sensor in a preset time window (such as the past 30 minutes, 1 hour), which is used to calculate the real-time proof complexity of the data source. This complexity evaluation is based on historical volatility, abnormality ratio, sensor privacy level, current network security situation, etc., and is obtained through a mathematical model or scoring function.

[0049] After obtaining the complexity value, the system inputs the value into the "proof method decision maker", which matches the appropriate first proof method according to the complexity interval (such as low, medium, high). For example, low complexity corresponds to hash-based ZKP, medium complexity uses zk-SNARK, and high complexity selects zk-STARK.

[0050] Finally, the system constructs a zero-knowledge proposition based on the obtained "identification indicators" and "judgment thresholds", such as "the current temperature value is less than 90℃", and then calls the selected ZKP engine (such as ZoKrates, libsnark, STARKy) to generate a zero-knowledge proof. This proof is encapsulated as the first proof data and added to the proof data set for subsequent storage, sharing or verification.

[0051] This step realizes a dynamic and adaptive execution mechanism for data privacy protection. By combining historical behavior, indicator thresholds, and real-time security state, the system can flexibly determine the protection strength of each piece of data, avoiding a "one-size-fits-all" approach. At the same time, by generating verifiable proofs through zero-knowledge mapping, privacy is ensured while ensuring data verifiability and credibility. In addition, this method supports intelligent selection between various ZKP technologies, improving the performance adaptability, resource utilization efficiency, and security policy accuracy of the system in different data environments. It has wide applicability and feasibility in industrial data space, smart cities, and Internet of Vehicles scenarios.

[0052] Further, the first real-time proof complexity is calculated according to the first monitoring data sequence, including: obtaining the first data privacy level of the first predetermined data source, and real-time security perception information of the industrial automatic production line, wherein the real-time security perception information is the network attack frequency in a preset time interval; performing data feature analysis on the first monitoring data sequence to obtain a first data anomaly coefficient and a first data fluctuation coefficient; and calculating the first real-time proof complexity according to the first data privacy level, the network attack frequency, the first data anomaly coefficient, and the first data fluctuation coefficient.

[0053] Specifically, the first data privacy level refers to the privacy sensitivity evaluation result of the data source, which can be low, medium, or high level, and is defined based on the business secrets, user privacy, or core technology information involved in the data content. Real-time security perception information refers to the security threat information monitored within a set time period (such as the past 15 minutes), such as network attack frequency, malicious request number, intrusion detection alarm, etc. Network attack frequency refers to the number of attack behaviors in the network per unit time, which quantifies the network security risk level. The first data anomaly coefficient is used to represent the proportion of data in the monitoring data sequence that exceeds the normal range, reflecting the possibility of potential failure or tampering of the data source. The first data fluctuation coefficient is used to reflect the stability of the data, and the calculation method is usually standard deviation ÷ average value. The larger the value, the more intense the data fluctuation. The first real-time proof complexity refers to a dynamic complexity level calculated based on the above factors, which is used to guide the selection and configuration of the zero-knowledge proof process.

[0054] Before privacy protection of a certain predetermined data source, first assess whether the data needs high-intensity zero-knowledge proof. For this purpose, the monitoring data sequence collected by the data source in the preset historical time interval is called to perform data behavior analysis.

[0055] Next, the following four types of factors are extracted and calculated to obtain data privacy level and security perception information, including: data privacy level, set by data tags, device policies, or manual annotation, such as "temperature data is medium privacy" and "product image is high privacy"; network attack frequency, obtained by analyzing security monitoring platform (such as IDS, SIEM) or local gateway security logs for a certain period of time, such as scanning, injection, DDOS, etc. Calculate the abnormal coefficient and fluctuation coefficient, including: abnormal coefficient calculation, judge how many proportions in the data do not meet the definition of the "judgment threshold set" (such as the number of times the temperature exceeds 80℃ ÷ total times); fluctuation coefficient calculation, mean and standard deviation analysis on the historical sequence, calculate "standard deviation ÷ mean" to get the fluctuation coefficient, the larger the more unstable.

[0056] The above four indicators will be normalized and input into a weighted evaluation model (such as linear model, fuzzy logic system, decision tree, etc.), and an output of a real-time proof complexity value will be obtained. Among them, the weighted evaluation model sets the weight of each influencing factor to comprehensively score the real-time proof complexity. First, normalize the four core indicators - data privacy level, network attack frequency, data abnormal coefficient, and data fluctuation coefficient to ensure dimensional consistency. Then, set the weight coefficient of each indicator according to the actual scene, for example, give higher weight to privacy level and security situation, and calculate the total score through linear weighting formula. Finally, the score is mapped to the complexity level interval to guide the ZKP strategy selection and realize a flexible and secure data privacy protection mechanism.

[0057] This step realizes multi-dimensional comprehensive evaluation of data itself characteristics, security risks and privacy attributes, no longer relying on static rules, but building a real-time and dynamic complexity evaluation system, thereby realizing fine privacy protection. Its advantages include: automatically increasing proof strength when data fluctuation is severe or network environment is poor; using lightweight proof when data is stable and privacy level is low to improve processing efficiency; avoiding resource waste while strengthening data transmission capability under attack threat.

[0058] Further, the data feature analysis is performed according to the first monitoring data sequence to obtain a first data anomaly coefficient and a first data fluctuation coefficient, including: calculating a proportion of abnormal data in the first monitoring data sequence to obtain the first data anomaly coefficient, wherein the abnormal data is data that does not satisfy a first data judgment threshold; calculating a first data mean and a first data standard deviation from the first monitoring data sequence, and setting a ratio of the first data mean and the first data standard deviation as the first data fluctuation coefficient.

[0059] Specifically, the first data mean refers to the average value of the monitoring data sequence, which is used to measure the general level of the data. The first data standard deviation refers to the degree of deviation of each data value in the data sequence from the mean, which is used to measure the volatility or stability of the data.

[0060] First, a monitoring data sequence in a historical time period is obtained from a predetermined data source, such as temperature values collected every minute for the past 30 minutes, totaling 30. Then, according to a preset judgment threshold, each data is checked, and the number of data that does not meet the condition is counted, which is the abnormal data.

[0061] The number of abnormal data is divided by the total number of samples to obtain the current anomaly coefficient of the data source. Then, the mean and standard deviation of the 30 data are calculated. This process can be completed by data analysis, and the edge computing node can use the NumPy or Pandas library of Python, and the embedded device can use the sliding window calculation logic implemented in C language, which supports real-time processing. The anomaly coefficient and the fluctuation coefficient will be two key input factors for subsequent evaluation of "real-time proof complexity", which directly affects the strength of the zero-knowledge proof selected.

[0062] Taking the current of a device as an example, the current data in the past 10 minutes (20 data) is collected: the judgment threshold is that the current should be between 5A and 15A; in the actual data, there are 4 data exceeding The mean of all data is 10A, According to the two coefficients, it is preliminarily judged that the device data has certain abnormality and has moderate volatility, so moderate or high zero-knowledge proof complexity is allocated in the subsequent steps.

[0063] This step realizes the behavior-level analysis of the data state through feature extraction of the historical monitoring data, especially in data anomaly identification and fluctuation sensitivity evaluation. Specifically, it can accurately identify potential abnormal sources and support early fault detection; dynamically quantify data stability to determine ZKP strength selection.

[0064] Further, the first real-time proof complexity is calculated according to the first data privacy level, network attack frequency, first data anomaly coefficient and first data fluctuation coefficient, comprising: calculating the ratio of the first data privacy level to the historical maximum data privacy level, the network attack frequency to the historical maximum network attack frequency, the first data anomaly coefficient to the historical maximum first data anomaly coefficient, and the first data fluctuation coefficient to the historical maximum first data fluctuation coefficient, respectively, and calculating the first data proof scale by weighted calculation; multiplying the first data proof scale by the historical maximum first proof complexity to obtain the first real-time proof complexity.

[0065] Specifically, the historical maximum value (such as the maximum privacy level, the maximum attack frequency, etc.) refers to the maximum reference value of the corresponding indicator that has appeared in the historical record, which is used to standardize the current data state. The historical maximum first proof complexity refers to the maximum ZKP proof strength that has been generated in history, which can be regarded as the system capacity or the upper limit of the maximum calculation complexity.

[0066] When the privacy level, network security situation, abnormal situation and fluctuation situation of a certain data source are analyzed, each indicator is standardized by comparing it with the corresponding historical maximum value. For example, if the current data privacy level is 0.8 and the historical maximum is 1.0, the ratio is 0.8; if the current attack frequency is 5 times per minute and the historical maximum is 10 times per minute, the ratio is 0.5, and so on. According to the pre-configured weight (for example: privacy level 40%, attack frequency 30%, anomaly coefficient 20%, fluctuation coefficient 10%), the weighted sum of these ratios is obtained, which is a comprehensive value between 0 and 1, called the first data proof scale. Then, multiplying this scale value by the historical maximum first proof complexity recorded by the system (such as 1.0 or a certain specific complexity coefficient), the first real-time proof complexity is calculated. This value can be used as a decision input to select the appropriate zero-knowledge proof method.

[0067] Assume the system records the historical maximum values as follows: maximum privacy level = 1.0; maximum network attack frequency = 10 times / minute; maximum anomaly coefficient = 0.6; maximum fluctuation coefficient = 0.3; maximum historical proof complexity = 1.0 (normalized upper limit). The current situation of a certain temperature data source is: current privacy level = 0.9; attack frequency = 6 times / minute; anomaly coefficient = 0.3; fluctuation coefficient = 0.15; then the four ratio values are: privacy level ratio = 0.9 / 1.0 = 0.9; attack frequency ratio = 6 / 10 = 0.6; anomaly coefficient ratio = 0.3 / 0.6 = 0.5; fluctuation coefficient ratio = 0.15 / 0.3 = 0.5; set the weights as follows: privacy (0.4), attack frequency (0.3), anomaly degree (0.2), fluctuation degree (0.1), then: first data proof scale = 0.4 x 0.9 + 0.3 x 0.6 + 0.2 x 0.5 + 0.1 x 0.5 = 0 = 0.69 Finally, the first real-time proof complexity = 0.69 x 1.0 = 0.69, which is mapped to the "moderately high" level, and a medium-high intensity zero-knowledge proof scheme such as zk-SNARK or enhanced STARK will be selected accordingly.

[0068] This step realizes the quantification, normalization and unified fusion evaluation of multi-source indicators, taking into account privacy, risk and data behavior characteristics, and improves the scientificity and controllability of the system for ZKP complexity configuration.

[0069] Further, according to the first real-time proof complexity, a first proof mode is matched and obtained in the proof mode decision maker, including: configuring a proof mode database, wherein the proof mode database can be dynamically updated regularly; determining a proof complexity threshold based on historical monitoring data analysis, and dividing a plurality of complexity intervals according to a predetermined complexity step size; constructing a proof mode decision maker according to the proof mode database and the plurality of complexity intervals, and inputting the first real-time proof complexity into the proof mode decision maker to match and obtain the first proof mode.

[0070] Specifically, the proof mode database refers to a set of available zero-knowledge proof technology schemes and their metadata (such as algorithm name, calculation overhead, applicable scope, etc.) maintained in the system. The complexity threshold is used to divide the boundary values of the proof complexity level, for example, "less than 0.3 is low complexity", "0.3-0.7 is medium", "more than 0.7 is high complexity". The complexity step size is used to define the minimum unit for equally dividing the entire complexity space (0 to 1) into a plurality of intervals, such as 0.1, 0.2, etc., for subsequent mapping and matching. The complexity interval refers to a plurality of continuous segments divided by the threshold or step size, for example: [0, 0.3), [0.3, 0.7), [0.7, 1.0], which is used to establish a "complexity-proof mode" mapping relationship.

[0071] Firstly, a database of proof methods is constructed and proven, which records all supported zero-knowledge proof methods and their related information, including name, applicable complexity range, running performance, algorithm type, encryption strength, and computational resource consumption. To maintain advancement and flexibility, this database supports dynamic updates and can be periodically expanded or replaced according to new algorithm releases or changes in security requirements. Specifically,

[0072] According to the current supported zero-knowledge proof protocol types (such as zk-SNARK, zk-STARK, Bulletproof, Plonk, etc.), the core parameter information is collected, including computational complexity, encryption strength, verification efficiency, and applicable scenarios. Secondly, for each proof method, the adaptation conditions are set, such as the complexity interval, recommended data type, or resource consumption level. Then, a unified data table structure is constructed to standardize the storage of various attributes and support dynamic update mechanisms that can regularly access new ZKP schemes or adjust weight strategies. Finally, combined with system historical running data and security policies, a "complexity interval-proof method" mapping table is established for real-time calling by the proof method decision maker. The entire database can be deployed on the cloud or edge nodes and combined with API interfaces to provide flexible access methods for intelligent and controllable scheduling of proof methods. This ensures the flexible expansion, fine adaptation, and rapid response of ZKP strategies, which helps to improve the overall privacy protection efficiency of the system.

[0073] Next, the system sets several complexity threshold values or uses equal-step division to divide the complexity interval based on historical running data and security policies. Then, the first real-time proof complexity calculated in real time is input into the "proof method decision maker" for matching according to the interval mapping relationship, and the corresponding "first proof method" is returned.

[0074] This step makes the zero-knowledge proof process from static configuration to dynamic scheduling and intelligent matching, with core advantages including: intelligent selection of the most suitable data protection method based on real-time complexity; dynamic expansion of ZKP schemes to improve the system's ability to adapt to new technologies; avoiding resource waste caused by "overprotection" and privacy leakage risks caused by "insufficient protection"; achieving a balance between high security, low resource consumption, and fast response capability; supporting fine-grained decision logic (such as self-adaptation according to data type and scenario) to provide a foundation for large-scale applications.

[0075] Further, the proof mode database and the plurality of complexity intervals are used to build a proof mode decision maker, including: complexity evaluation of a plurality of proof modes in the proof mode database, to determine a plurality of proof complexity; dividing the plurality of proof complexity according to the plurality of complexity intervals, to map and determine a plurality of proof mode sets; and building the proof mode decision maker according to the plurality of complexity intervals and the plurality of proof mode sets based on a decision tree and according to the mapping relationship between the complexity intervals and the proof mode sets.

[0076] Specifically, the plurality of proof modes refers to various ZKP algorithm schemes stored in the database, such as zk-SNARK, zk-STARK, Bulletproof, Plonk, Sigma protocol, etc. The proof complexity refers to a comprehensive performance indicator for measuring the required computing resources, time overhead, and proof data size of each proof mode during execution, usually represented in numerical form for easy comparison. The complexity interval is used to divide the real-time proof complexity value into several levels, and the proof mode set refers to a group of optional proof mode sets corresponding to each complexity interval, i.e., a subset of the plurality of proof modes.

[0077] First, each zero-knowledge proof method in the existing proof mode database needs to be quantitatively evaluated. The evaluation content includes: generation time, verification time, proof size, computing resource consumption, security level, etc. After standardization processing, it is summarized as a proof complexity value, which is used to represent the overhead and capability of the method itself. Next, according to the preset complexity interval division rule (such as dividing into 5 segments with a step of 0.2), the complexity value of each proof mode is divided into the corresponding interval, and the mapping relationship between the complexity interval and the proof mode set is established. For example: interval [0.0-0.2): contains lightweight Hash-based ZKP; interval [0.2-0.5): contains Bulletproof; interval [0.5-0.8): contains zk-SNARK; interval [0.8-1.0]: contains zk-STARK and Plonk. Then, based on the above interval and the corresponding proof mode set, a decision tree model is constructed: the branch node of the tree represents whether the real-time proof complexity falls within a certain interval, and the leaf node represents the specific "recommended proof mode set". Whenever there is a new real-time proof complexity as input, the decision tree top layer is used to determine the interval it belongs to, and the matching ZKP mode is quickly output.

[0078] The "proof mode decision maker" constructed by this method achieves the following key goals: structuring and automating the ZKP selection process, reducing human intervention; achieving efficient mapping and fast response from "real-time complexity to best ZKP scheme"; supporting flexible expansion and dynamic update of proof modes, adapting to new technology development;

[0079] In summary, the trusted data space privacy protection method based on zero-knowledge proof provided by the embodiments of the application has the following technical effects:

[0080] 1. A privacy protection and data trusted sharing mechanism is constructed by zero-knowledge proof and blockchain technology, which ensures that sensitive data is verified securely without revealing the content, thereby improving the security, compliance and verifiability of the data space.

[0081] 2. The data features are accurately identified and evaluated based on the data source background information, and the zero-knowledge proof mode is dynamically selected in combination with real-time complexity, thereby effectively improving the matching degree and execution efficiency of the privacy protection strategy.

[0082] 3. A multi-dimensional feature analysis mechanism is introduced to comprehensively evaluate data privacy, security risk, abnormality degree and volatility characteristics, thereby effectively realizing dynamic calculation of ZKP complexity and risk-sensitive response, and improving the privacy protection accuracy.

[0083] Embodiment Two

[0084] Based on the same inventive concept as the trusted data space privacy protection method based on zero-knowledge proof in the foregoing embodiments, as shown in Figure 3 The embodiments of the application provide a trusted data space privacy protection system based on zero-knowledge proof, which comprises:

[0085] The running data acquisition module 11 is configured to acquire real-time running data of a plurality of data sources through the Internet of Things, obtain a real-time running data set and a data source information set; the proof data set generation module 12 is configured to perform zero-knowledge proof on the real-time running data set based on the data source information set according to real-time proof complexity, and generate a proof data set, wherein the real-time proof complexity is determined based on comprehensive analysis of security perception information, data privacy degree, data abnormality degree and data volatility degree; and the data encryption storage module 13 is configured to perform data sharing and data verification according to the encrypted storage data set by using blockchain technology on the proof data set.

[0086] Further, the running data acquisition module 11 is further configured to perform the following steps: through the Internet of Things, data of a plurality of predetermined data sources of an industrial automatic production line is acquired at a predetermined monitoring time node, and the data is preliminarily encrypted to obtain a real-time running data set; through the Internet of Things, the data source ID, data type, location information and timestamp of each predetermined data source are synchronously acquired to obtain a data source information set.

[0087] Furthermore, the proof data set generation module 12 is also used to perform the following steps: matching and calling the data identification indicator set and the data judgment threshold set based on the data source information set; performing zero-knowledge mapping proof on the real-time running data set according to the real-time proof complexity according to the data identification indicator set and the data judgment threshold set to generate the proof data set.

[0088] Furthermore, the proof data set generation module 12 is also used to perform the following steps: randomly select a first predetermined data source, and obtain the first real-time operation data, the first data identification index and the first data judgment threshold of the first predetermined data source, and the first monitoring data sequence within a preset historical time range; calculate the first real-time proof complexity according to the first monitoring data sequence; according to the first real-time proof complexity, match and obtain the first proof method in the proof method decider; according to the first data identification index and the first data judgment threshold, perform zero-knowledge proof on the first real-time operation data according to the first proof method, generate first proof data, and add it to the proof data set.

[0089] Furthermore, the proof data set generation module 12 is also used to perform the following steps: obtaining the first data privacy level of the first predetermined data source and the real-time security perception information of the industrial automatic production line, wherein the real-time security perception information is the frequency of network attacks within a preset time interval; performing data feature analysis based on the first monitoring data sequence to obtain a first data anomaly coefficient and a first data fluctuation coefficient; and calculating a first real-time proof complexity based on the first data privacy level, the network attack frequency, the first data anomaly coefficient and the first data fluctuation coefficient.

[0090] Furthermore, the proof data set generation module 12 is also used to perform the following steps: calculate the proportion of abnormal data in the first monitoring data sequence to obtain a first data anomaly coefficient, wherein the abnormal data is data that does not meet the first data judgment threshold; calculate the first data mean and the first data standard deviation based on the first monitoring data sequence, and set the ratio of the first data mean to the first data standard deviation as the first data fluctuation coefficient.

[0091] Furthermore, the proof data set generation module 12 is also used to perform the following steps: respectively calculate the ratios of the first data privacy level and the historical maximum data privacy level, the network attack frequency and the historical maximum network attack frequency, the first data anomaly coefficient and the historical maximum first data anomaly coefficient, and the first data fluctuation coefficient and the historical maximum first data fluctuation coefficient, and perform weighted calculations to obtain the first data proof scale; multiply the first data proof scale by the historical maximum first proof complexity to obtain the first real-time proof complexity.

[0092] Further, the proof data set generation module 12 is further configured to perform the following steps: configuring a proof mode database, wherein the proof mode database can be dynamically updated periodically; determining a proof complexity threshold based on historical monitoring data analysis, and dividing a plurality of complexity intervals according to a predetermined complexity step size; constructing a proof mode decision maker according to the proof mode database and the plurality of complexity intervals, and inputting the first real-time proof complexity into the proof mode decision maker to match and obtain a first proof mode.

[0093] Further, the proof data set generation module 12 is further configured to perform the following steps: performing complexity evaluation on a plurality of proof modes in the proof mode database to determine a plurality of proof complexities; dividing the plurality of proof complexities according to the plurality of complexity intervals to map and determine a plurality of proof mode sets; constructing the proof mode decision maker according to the plurality of complexity intervals and the plurality of proof mode sets based on a decision tree and according to a mapping relationship between the complexity intervals and the proof mode sets.

[0094] Any step of the above method can be stored as computer instructions or programs in an unrestricted computer memory and can be called and recognized by an unrestricted computer processor to implement any method in the embodiments of the present application. No redundant limitation is made herein.

[0095] Further, the above-mentioned first or second order relationship represents a specific concept, and / or refers to a plurality of elements which can be selected individually or collectively. Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the present application and its equivalent technologies, the present application intends to include these modifications and variations.

Claims

1. A trusted data space privacy protection method based on zero-knowledge proof, characterized by: Methods include: Through the Internet of Things, real-time operation data from multiple data sources are collected to obtain real-time operation data sets and data source information sets; Based on the data source information set, performing a zero-knowledge proof on the real-time running data set according to a real-time proof complexity to generate a proof data set, wherein the real-time proof complexity is determined based on a comprehensive analysis of security perception information, data privacy, data anomaly, and data volatility; Using blockchain technology, the proof data set is encrypted and stored, and data sharing and data verification are performed based on the encrypted stored data set; The method of performing zero-knowledge proof on the real-time running data set based on the data source information set according to the real-time proof complexity to generate a proof data set includes: Based on the data source information set, the data identification indicator set and the data judgment threshold set are matched and called; According to the data identification indicator set and the data judgment threshold set, performing zero-knowledge mapping proof on the real-time running data set according to the real-time proof complexity to generate the proof data set; According to the data identification indicator set and the data judgment threshold set, a zero-knowledge mapping proof is performed on the real-time running data set according to the real-time proof complexity, including: Randomly selecting a first predetermined data source, and obtaining first real-time operating data, a first data identification indicator, and a first data judgment threshold of the first predetermined data source, as well as a first monitoring data sequence within a preset historical time range, wherein the first predetermined data source refers to an object randomly selected from a plurality of configured sensors or devices for performing zero-knowledge proof, the first data identification indicator refers to an indicator extracted from a predefined indicator set and used to identify data behavior characteristics, and the first data judgment threshold refers to a judgment limit corresponding to the identification indicator, used to determine whether the data is in a normal, abnormal, or high-risk range; Calculating a first real-time proof complexity according to the first monitoring data sequence; According to the first real-time proof complexity, matching and obtaining a first proof method in a proof method decision device; Performing zero-knowledge proof on the first real-time operation data according to the first data identification index and the first data judgment threshold in accordance with the first proof method to generate first proof data, and adding the first proof data to the proof data set; The first real-time proof complexity is calculated based on the first monitoring data sequence, including: Obtaining a first data privacy level of the first predetermined data source and real-time security perception information of the industrial automatic production line, wherein the real-time security perception information is a frequency of network attacks within a preset time interval; Performing data feature analysis on the first monitoring data sequence to obtain a first data anomaly coefficient and a first data fluctuation coefficient; A first real-time proof complexity is calculated based on the first data privacy level, the network attack frequency, the first data anomaly coefficient and the first data fluctuation coefficient.

2. The trusted data space privacy protection method based on zero-knowledge proof according to claim 1 is characterized in that: Through the Internet of Things, real-time operation data from multiple data sources is collected to obtain real-time operation data sets and data source information sets, including: Through the Internet of Things, data is collected from multiple predetermined data sources of industrial automatic production lines at predetermined monitoring time nodes, and the data is initially encrypted to obtain real-time operation data sets; Through the Internet of Things, the data source ID, data type, location information and timestamp of each predetermined data source are synchronously collected to obtain a data source information set.

3. The trusted data space privacy protection method based on zero-knowledge proof according to claim 1 is characterized in that: Performing data feature analysis based on the first monitoring data sequence to obtain a first data anomaly coefficient and a first data fluctuation coefficient includes: Calculating the proportion of abnormal data in the first monitoring data sequence to obtain a first data abnormality coefficient, wherein the abnormal data is data that does not meet the first data judgment threshold; A first data mean and a first data standard deviation are calculated based on the first monitoring data sequence, and a ratio of the first data mean to the first data standard deviation is set as a first data fluctuation coefficient.

4. The trusted data space privacy protection method based on zero-knowledge proof according to claim 1 is characterized in that: The first real-time proof complexity is calculated according to the first data privacy level, the network attack frequency, the first data anomaly coefficient, and the first data fluctuation coefficient, including: Calculate the ratios of the first data privacy level to the historical maximum data privacy level, the network attack frequency to the historical maximum network attack frequency, the first data anomaly coefficient to the historical maximum first data anomaly coefficient, and the first data fluctuation coefficient to the historical maximum first data fluctuation coefficient, and perform a weighted calculation to obtain a first data proof scale; The first real-time proof complexity is obtained by multiplying the first data proof scale by the historical maximum first proof complexity.

5. The trusted data space privacy protection method based on zero-knowledge proof according to claim 1 is characterized in that: According to the first real-time proof complexity, matching and obtaining a first proof method in a proof method decision unit includes: Configuring a certification method database, wherein the certification method database can be dynamically updated regularly; Determine the proof complexity threshold based on historical monitoring data analysis, and divide it into several complexity intervals according to the predetermined complexity step size; A proof method decision maker is constructed according to the proof method database and the plurality of complexity intervals, and the first real-time proof complexity is input into the proof method decision maker to obtain a first proof method through matching.

6. The trusted data space privacy protection method based on zero-knowledge proof according to claim 5 is characterized in that: Constructing a proof method decision maker according to the proof method database and the plurality of complexity intervals, including: Performing complexity evaluation on multiple proof methods in the proof method database to determine multiple proof complexities; Dividing the plurality of proof complexities according to the plurality of complexity intervals, and mapping and determining a plurality of proof method sets; Based on the decision tree and in accordance with the mapping relationship between complexity intervals and proof method sets, the proof method decider is constructed according to the plurality of complexity intervals and the plurality of proof method sets.

7. A trusted data space privacy protection system based on zero-knowledge proof, characterized by: A system for executing the trusted data space privacy protection method based on zero-knowledge proof according to any one of claims 1 to 6, comprising: The operation data collection module is used to collect real-time operation data from multiple data sources through the Internet of Things, and obtain real-time operation data sets and data source information sets; a proof data set generation module, configured to perform zero-knowledge proof on the real-time running data set based on the data source information set and generate a proof data set according to a real-time proof complexity, wherein the real-time proof complexity is determined based on a comprehensive analysis of security perception information, data privacy, data anomaly, and data volatility; The data encryption storage module is used to utilize blockchain technology to perform data sharing and data verification on the certification data set based on the encrypted storage data set.

Citation Information

Patent Citations

  • Block chain privacy data sharing method based on zero knowledge proof

    CN114499900A

  • Electric power information security storage supervision system based on block chain

    CN119150347A