Industrial data desensitization method, system and equipment and medium

By constructing a six-dimensional binding relationship model and a tiered collaborative desensitization mechanism, the problem of insufficient adaptability and effectiveness of existing industrial data desensitization solutions in industrial cluster network environments has been solved. This has enabled refined data desensitization and dynamic access control, and improved the security and controllability of data sharing and circulation.

CN121859362AActive Publication Date: 2026-04-14江西冠英智能科技股份有限公司 +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing industrial data anonymization solutions lack adaptability and effectiveness in industrial cluster network environments, struggle to handle complex security challenges across levels and entities, and lack a time dimension as a core dynamic adjustment factor for anonymization strategies, resulting in data distortion or insufficient protection. Furthermore, their access control is coarse-grained, failing to meet the requirements of refined data governance and compliance auditing.

Method used

A six-dimensional binding relationship model is constructed, integrating the time dimension and business rules. By binding industrial data variable types, time dimensions, node levels, and cluster roles in multiple dimensions, a tiered collaborative desensitization mechanism is designed to achieve fine-grained, collaborative, and traceable desensitization processing. Furthermore, targeted decryption and accuracy restoration are performed in conjunction with full-link traceability logs.

Benefits of technology

It achieves refined desensitization of industrial variables such as temperature and rotation speed, dynamically adjusts the desensitization intensity, builds a defense-in-depth system, supports extremely fine-grained dynamic access control, meets the needs of refined access control and compliance auditing in complex industrial networks, and improves the security and controllability of data sharing and flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121859362A_ABST
    Figure CN121859362A_ABST
Patent Text Reader

Abstract

The invention relates to a desensitization method, system, equipment and medium for industrial data, and the method comprises the steps: building and dynamically maintaining an industrial field standard data variable rule table, and defining a desensitization strategy and a sensitivity level matched with specific industrial variables such as temperature, rotating speed and the like in physical meanings, time scenes and node levels of the specific industrial variables; the method comprises the following steps: adding a multi-dimensional identifier to original industrial data, establishing a six-dimensional binding relation model of data variable type-time dimension-node hierarchy-cluster role-sensitive level-rule entry, driving the data to execute a time-coordinated stepped desensitization process in four node hierarchies of a workshop, an enterprise, a park and a cluster, according to the method, refined access control and data decryption fusing time, space and role dimensions are realized according to the model and a full-link tracing log, and the problems that rules and industrial scenes are disjointed, dynamic time adjustment is lacked, cross-node collaboration is insufficient and authority control coarse granularity is caused in the prior art are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial data security processing technology, specifically to a method, system, device, and medium for desensitizing industrial data. Background Technology

[0002] Industrial data anonymization is one of the core technical means in the intersection of the Industrial Internet and data security. Its core objective is to prevent the leakage of sensitive information during data sharing, circulation, and analysis by means of technical transformation, replacement, or masking, while ensuring data availability. In the typical scenario of industrial clusters, industrial data usually flows and converges along the vertical hierarchy of "workshop-enterprise-industry-cluster" and the horizontal network across enterprises. This means that data anonymization not only needs to address the sensitivity of individual data points, but also needs to cope with the complex security challenges brought about by cross-level and cross-entity collaboration.

[0003] However, existing general data anonymization solutions still face a series of limitations in adaptability and effectiveness when dealing with industrial cluster network environments. First, most solutions have relatively generic anonymization rules, lacking fine-grained modeling of industrial variables with specific physical meanings and reasonable value ranges, such as temperature, pressure, and rotational speed. This can easily lead to data distortion or insufficient protection after anonymization. Second, industrial production has significant time-series characteristics, such as real-time monitoring, shift-based statistics, and periodic analysis. However, existing solutions typically do not use the time dimension as a core dynamic adjustment factor in their anonymization strategies, making it difficult to automatically adjust protection strength according to different time scenarios such as real-time, shift periods, and statistical cycles. Third, in multi-level network architectures, existing methods often employ centralized or endpoint-based anonymization, lacking tiered and collaborative anonymization mechanisms across workshops, enterprises, and industrial parks, making it difficult to defend against reverse engineering attacks that link multi-source data. Finally, access control models often focus on user identity and data classification, failing to deeply integrate the multi-dimensional context of data, such as time, space (node ​​location), and business roles. This results in coarse-grained access control, making it difficult to meet the requirements of refined data governance and compliance auditing. Summary of the Invention

[0004] Based on this, the purpose of this invention is to provide a method, system, device, and medium for desensitizing industrial data that can be deeply integrated with the time dimension and business rules for multi-level architectures of industrial clusters, and can achieve fine-grained, collaborative, and traceable data.

[0005] The objective of this invention is achieved through the following solution:

[0006] In a first aspect, the present invention provides a method for desensitizing industrial data, comprising the following steps:

[0007] S1: Define and organize the types, time dimensions, desensitization algorithms and compliance criteria of industrial data variables in a rule-based manner. Configure the corresponding desensitization algorithms, sensitivity levels and access rules for each type of industrial variable under different time scenarios and node levels, and generate a versioned standard data variable rule table for the industrial field.

[0008] S2: Timestamp and feature-analyze the raw industrial data collected at the workshop node. Based on the standard data variable rule table in the industrial field, add the data variable type, the time period attribute and statistical period corresponding to the collection time, and the initial sensitivity level determined based on the rule table to each piece of raw industrial data, and generate industrial data to be desensitized with multi-dimensional data identifiers.

[0009] S3: Model the relationship between the standard data variable rule table in the industrial field and the dimensions involved in the industrial data to be desensitized. Associate and bind the data variable type, time dimension, node level, cluster role, data sensitivity level with the specific entries in the rule table in multiple dimensions to build a six-dimensional binding relationship model.

[0010] S4: Based on the industrial data to be desensitized and the six-dimensional binding relationship model, in the four-level nodes of the industrial cluster network of workshop, enterprise, park and cluster, the collaborative desensitization calculation is performed in sequence from fine-grained to coarse-grained and from local to global. The desensitization processing results of each level are used as the input of the next level. After the desensitization processing is completed at the cluster node, the processing metadata of each level is integrated to generate a full-link traceability log.

[0011] S5: Based on the user's data access request and current access time, it matches access permissions and decryption rules through a six-dimensional binding relationship model, and performs targeted decryption and precision restoration based on the full-link traceability log to generate result data that matches the current user's permissions and access time.

[0012] Secondly, the present invention provides an industrial data desensitization system, which is configured with the following modules:

[0013] The data rule definition module is used to define and organize the types, time dimensions, desensitization algorithms and compliance criteria of industrial data variables in a rule-based manner. It configures the corresponding desensitization algorithms, sensitivity levels and access rules for each type of industrial variable under different time scenarios and node levels, and generates a versioned standard data variable rule table for the industrial field.

[0014] The data to be anonymized generation module is used to timestamp and analyze the features of the raw industrial data collected at the workshop node. Based on the standard data variable rule table in the industrial field, it adds the data variable type, the time period attribute and statistical period corresponding to the collection time, and the initial sensitivity level determined based on the rule table to each piece of raw industrial data, generating industrial data to be anonymized with multi-dimensional data identifiers.

[0015] The six-dimensional relationship modeling module is used to model the relationships between the standard data variable rule table in the industrial field and the dimensions involved in the industrial data to be desensitized. It associates and binds the data variable type, time dimension, node level, cluster role, data sensitivity level with the specific entries in the rule table in multiple dimensions to build a six-dimensional binding relationship model.

[0016] The collaborative desensitization module is used to perform collaborative desensitization calculations in the four-level nodes of the industrial cluster network—workshop, enterprise, park, and cluster—based on the industrial data to be desensitized and the six-dimensional binding relationship model. It performs calculations from fine-grained to coarse-grained and from local to global. The desensitization results of each level are used as inputs for the next level. After the desensitization is completed at the cluster node, the processing metadata of each level is integrated to generate a full-link traceability log.

[0017] The data access decryption module is used to match access permissions and decryption rules based on the user's data access request and the current access time through a six-dimensional binding relationship model, and to perform targeted decryption and accurate restoration based on the full-link traceability log, generating result data that matches the current user's permissions and access time.

[0018] Thirdly, this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any of the above-mentioned methods for desensitizing industrial data.

[0019] Fourthly, this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements any of the above-mentioned methods for desensitizing industrial data.

[0020] In summary, the industrial data anonymization method provided in this application, by constructing a standardized rule system deeply adapted to industrial scenarios, enables the definition of refined and interpretable anonymization strategies for industrial variables with specific physical meanings, such as temperature and rotational speed, thereby solving the problem of the disconnect between general rules and industrial semantics. By integrating multiple time dimensions, such as real-time, time period, and cycle, into the entire data processing chain as core constraints, the system can dynamically adjust the anonymization intensity and calculation strategy according to the time context of the data, effectively addressing the lack of time adaptability in static anonymization schemes. By designing a four-level node-tiered collaborative anonymization mechanism of "workshop-enterprise-park-cluster," single-point sensitive information, variable association risks, and cross-enterprise characteristic differences are eliminated sequentially in the step-by-step processing, constructing a defense-in-depth system to solve the security risks of reverse derivation caused by cross-node data splicing. By establishing a six-dimensional binding relationship model integrating data variables, time, nodes, roles, and sensitivity levels, the system can support extremely fine-grained dynamic permission control of "who, when, where, and what level of data precision is accessed," meeting the needs of refined access control and compliance auditing in complex industrial networks. Ultimately, by combining end-to-end traceability logs, this method forms a complete security closed loop that ensures traceable rules, verifiable processing, and auditable access. While guaranteeing data availability, it significantly improves the security and controllability of industrial data sharing and circulation in industrial cluster environments.

[0021] To better understand and implement this invention, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description

[0022] Figure 1 A flowchart illustrating a method for desensitizing industrial data provided in this application embodiment;

[0023] Figure 2 A schematic diagram illustrating the process of generating end-to-end traceability logs provided in an embodiment of this application;

[0024] Figure 3 This is a schematic diagram of the structure of an industrial data desensitization system provided in another embodiment of this application. Detailed Implementation

[0025] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Preferred embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.

[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0027] In one embodiment, such as Figure 1 As shown, a method for de-identifying industrial data is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0028] S1: Define and organize the types, time dimensions, desensitization algorithms, and compliance criteria of industrial data variables in a rule-based manner. Configure corresponding desensitization algorithms, sensitivity levels, and access rules for each type of industrial variable under different time scenarios and node levels, and generate a versioned standard data variable rule table for the industrial field.

[0029] Specifically, the system defines industrial data variable types in a rule-based manner based on relevant laws and standards for industrial data security management. The classification of industrial data variable types combines the physical attributes and business applications of data in industrial production scenarios to form a unified classification system, with each type corresponding to a unique identifier. The system also classifies the time dimension in a rule-based manner, distinguishing different time scenarios based on the temporal characteristics of industrial production, clarifying the definition standards and applicable scope of each time scenario, and providing a foundation for configuring scenario-based de-identification strategies. The system pre-sets multiple de-identification algorithms, clarifying the applicable scenarios and operational logic of each algorithm. The selection of algorithms is related to the industrial data variable type, time scenario, and business availability requirements.

[0030] Furthermore, the system establishes a sensitivity level classification system, clearly defining the judgment criteria for data at each level. Different sensitivity levels correspond to different protection strengths and access control requirements. The system defines access rules, and the configuration of access rules combines sensitivity levels, cluster roles, and time scenarios to clarify data access permissions and operational restrictions under different combinations of conditions. The system establishes a link between compliance basis and rule entries, ensuring that the formulation of each rule is supported by corresponding laws or standards. These elements are systematically organized to form a standard data variable rule table for the industrial field, and the rule table adopts a version management mechanism. When elements in the rule table are added, modified, or deleted, the system automatically triggers a version update, generates a new version identifier, and records relevant information about the version change, including the change time, change content, and change subject. All version information and change records are persistently stored by the system to ensure the traceability of the rules.

[0031] S2: Timestamp and analyze the raw industrial data collected at the workshop node. Based on the standard data variable rule table in the industrial field, add the data variable type, the time period attribute and statistical period corresponding to the collection time, and the initial sensitivity level determined based on the rule table to each raw industrial data to generate industrial data to be desensitized with multi-dimensional data identifiers.

[0032] Specifically, the system receives raw industrial data collected from workshop nodes and generates a timestamp for each piece of raw industrial data. The timestamp is generated synchronously with the data acquisition action, ensuring that the timestamp accurately reflects the timing information of the data acquisition. The system performs feature analysis on the raw industrial data, identifying the variable type corresponding to the data through the identification information in the data transmission protocol, ensuring that the variable type identification results are consistent with the variable type system in the rule table. Based on the time dimension division standards in the rule table, the system analyzes the time period attribute and statistical period corresponding to each data timestamp. The determination of the time period attribute and statistical period strictly follows the relevant definitions in the rule table. Based on the data variable type and acquisition scenario, combined with the sensitivity level division standards in the rule table, the system determines the initial sensitivity level of each data point. The initial sensitivity level determination process completely follows preset rules without additional manual intervention.

[0033] Furthermore, the system integrates the raw data, timestamps, parsed variable types, time period attributes, statistical periods, initial sensitivity levels, and relevant information from the data acquisition devices, encapsulating them into a standardized data structure to form industrial data to be anonymized, carrying multi-dimensional data identifiers. The system performs integrity verification on the encapsulated industrial data to be anonymized. The verification process is implemented through a specific algorithm to ensure that the data is not lost or tampered with during transmission and parsing. Data that fails verification is temporarily stored by the system, and a retransmission request is sent to the relevant nodes. The retransmitted data undergoes the parsing and encapsulation process again. Data that passes verification is transmitted to subsequent processing stages to ensure the continuity of the data processing flow.

[0034] S3: Model the relationship between the standard data variable rule table in the industrial field and the dimensions involved in the industrial data to be desensitized. Associate and bind the data variable type, time dimension, node level, cluster role, data sensitivity level with the specific entries in the rule table in multiple dimensions to build a six-dimensional binding relationship model.

[0035] Specifically, the system reads relevant fields from the standard data variable rule table in the industrial field, extracting dimensional information such as data variable type, time dimension, node level, applicable role, and sensitivity level. Simultaneously, it extracts the corresponding dimensional information from the industrial data to be de-identified, ensuring consistency in the dimensional systems of the two types of data. The system performs association processing on the extracted dimensional information, establishing a correspondence between data variable type, time dimension, node level, cluster role, data sensitivity level, and specific entries in the rule table, forming a six-dimensional association set. Preferably, the system constructs a dedicated storage structure to hold the six-dimensional association set. This storage structure supports high-concurrency queries and dynamic updates, ensuring rapid retrieval of corresponding rule entries during subsequent data processing. The system generates mapping relationships based on the six-dimensional association set. Each set of dimensions corresponds to a unique rule table entry. The generation process of the mapping relationship strictly follows preset logic to ensure the accuracy of the mapping results.

[0036] Furthermore, the system sets priority rules for mapping and matching. When a query request cannot find a completely matching dimension combination, it performs degenerate matching according to a preset priority to ensure that each query request obtains the corresponding rule entry. The system provides a standardized query interface that receives six-dimensional information as input and outputs the corresponding rule table entries and related details. The system periodically verifies the mapping relationship, comparing the association information in the rule table with that in the storage structure. When inconsistencies are found, a repair process is automatically triggered to ensure that the mapping relationship remains synchronized with the rule table. When the rule table is updated, the system automatically updates the six-dimensional association set and mapping relationship in the storage structure, deletes the association information corresponding to the old version, and generates the association content corresponding to the new version. A specific mechanism is used during the update process to ensure that the query service is not interrupted.

[0037] S4: Based on the industrial data to be de-identified and the six-dimensional binding relationship model, collaborative de-identification calculations are performed sequentially from fine-grained to coarse-grained and from local to global in the four-level nodes of the industrial cluster network: workshop, enterprise, park, and cluster. The de-identification processing results of each level are used as the input of the next level. After the de-identification processing is completed at the cluster node, the metadata of each level is integrated to generate a full-link traceability log.

[0038] Specifically, the system transmits the industrial data to be anonymized to the workshop node. System components in the workshop node obtain the corresponding anonymization rules based on a six-dimensional binding relationship model. According to the configuration in the rule table, they apply the corresponding anonymization algorithm to process the data. After processing, the system adds relevant metadata to the data, containing key information from the processing. The system then transmits the anonymization results from the workshop node to the enterprise node. System components in the enterprise node receive the data, match the anonymization rules at the enterprise level based on the six-dimensional binding relationship model, and simultaneously perform correlation analysis to identify related data sets within the same statistical period, executing collaborative anonymization processing.

[0039] After processing, the system appends enterprise-level processing metadata and continues to transmit the data to the park node. Upon receiving cross-enterprise data, the system components in the park node first unify the time-period configurations of each enterprise and perform cross-enterprise collaborative de-identification processing according to the requirements of the rule table. During processing, the system integrates the associated data from different enterprises, applies the corresponding de-identification algorithms to complete the processing, and then appends park-level processing metadata before transmitting the data to the cluster node. The system components in the cluster node receive the data transmitted from the park node and perform final de-identification processing based on compliance rules. After processing, the system integrates the processing metadata from the workshop, enterprise, park, and cluster levels to generate a full-link traceability log. The system associates the final de-identified data with the full-link traceability log and stores it in a preset storage area. The log contains processing information at each level, rule application information, and data flow information, ensuring the traceability of the data throughout its entire lifecycle. The system guarantees the security and integrity of the stored data.

[0040] S5: Based on the user's data access request and current access time, it matches access permissions and decryption rules through a six-dimensional binding relationship model, and performs targeted decryption and precision restoration based on the full-link traceability log to generate result data that matches the current user's permissions and access time.

[0041] Specifically, the system receives data access requests from users, including user identifiers, data query conditions, and access time information. The system verifies the user identifier in the access request based on a preset authentication mechanism. Simultaneously, the system extracts the access time information and, combined with the time period configuration in the rule table, verifies the user's time-period access permissions. After successful verification, the system parses the data query conditions in the request, extracting key dimensions such as data variable types and node levels. Based on a six-dimensional binding relationship model, it matches the corresponding access permissions and decryption rules. According to the matched decryption rules, the system requests the corresponding key from the key management module. Key acquisition is based on the correspondence between user permissions and data sensitivity levels, ensuring the security of key transmission and usage.

[0042] Furthermore, the system locates the processing trajectory of target data through end-to-end traceability logs, performs targeted decryption according to decryption rules, and the decryption process corresponds in reverse to the anonymization processing flow at each level to achieve accurate data restoration. During the restoration process, the system controls the degree of data restoration based on user permissions to ensure that the output data matches the user's permissions and access time. The system performs compliance verification on the restored data, based on the compliance requirements and access rules in the rule table. After successful verification, the system encapsulates the result data in a preset format, including the data itself and relevant access compliance information. The system feeds back the encapsulated result data to the user and records relevant information about this access operation, including the accessing user, access time, data range, decryption rules, and output results. This record is synchronized to the end-to-end traceability log to ensure the auditability of the access operation. Throughout the entire process, the system strictly adheres to data security management standards to ensure the legality and security of data access.

[0043] In summary, the industrial data anonymization method provided in this application, by constructing a standardized rule system deeply adapted to industrial scenarios, enables the definition of refined and interpretable anonymization strategies for industrial variables with specific physical meanings, such as temperature and rotational speed, thereby solving the problem of the disconnect between general rules and industrial semantics. By integrating multiple time dimensions, such as real-time, time period, and cycle, into the entire data processing chain as core constraints, the system can dynamically adjust the anonymization intensity and calculation strategy according to the time context of the data, effectively addressing the lack of time adaptability in static anonymization schemes. By designing a four-level node-tiered collaborative anonymization mechanism of "workshop-enterprise-park-cluster," single-point sensitive information, variable association risks, and cross-enterprise characteristic differences are eliminated sequentially in the step-by-step processing, constructing a defense-in-depth system to solve the security risks of reverse derivation caused by cross-node data splicing. By establishing a six-dimensional binding relationship model integrating data variables, time, nodes, roles, and sensitivity levels, the system can support extremely fine-grained dynamic permission control of "who, when, where, and what level of data precision is accessed," meeting the needs of refined access control and compliance auditing in complex industrial networks. Ultimately, by combining end-to-end traceability logs, this method forms a complete security closed loop that ensures traceable rules, verifiable processing, and auditable access. While guaranteeing data availability, it significantly improves the security and controllability of industrial data sharing and circulation in industrial cluster environments.

[0044] In one embodiment, S1 of the desensitization method for industrial data provided by the present invention specifically includes the following steps:

[0045] S11: Create a structured rule table on the cluster node server, generate a unique rule identifier for each rule entry in the structured rule table, set a data variable type field in the structured rule table to define the physical type and unit of industrial data, and set a time dimension type field to define the time dimension type.

[0046] Specifically, the system performs the creation of structured rule tables on the cluster node servers. The creation process follows database table structure design specifications to ensure the table structure has scalability and query efficiency. The system configures basic storage attributes for the structured rule tables, determining the storage engine, field data types, and indexing strategies. The indexing strategies are optimized for frequently queried fields. The system automatically generates a unique rule identifier for each rule entry in the structured rule table. The rule identifier is generated using preset encoding rules, and the encoding includes rule creation time information and a random sequence to ensure the uniqueness and traceability of the identifier. The system sets a data variable type field in the structured rule table. This field defines the physical type and unit of industrial data. The definition of physical type covers various data forms in industrial production, and the unit and physical type form a fixed association, ensuring standardized description of data variable types. The system also sets a time dimension type field, which explicitly defines the specific type of time dimension. The division of time dimension types is based on the temporal characteristics of industrial data generation, flow, and analysis, covering different time scenario requirements.

[0047] S12: Configure time period adaptation rules for each rule entry in the structured rule table. Define the start and end times of working hours, non-working hours, and overtime hours in the time period adaptation rule field of the structured rule table in key-value pair format, and configure conflict resolution strategies.

[0048] Specifically, the system configures time-period adaptation rules for each rule entry in the structured rule table, storing the configuration through the time-period adaptation rule field. During configuration, the system defines the start and end times of working hours, non-working hours, and overtime hours in key-value pair format. The key of the key-value pair identifies the time-period type, and the value records the corresponding start and end time information. This storage format ensures that the system can quickly parse and compare the data. The system also configures conflict resolution strategies for each rule entry. These strategies are set based on the actual operational needs of industrial production, addressing potential overlaps and gaps in conflict scenarios across different time periods, and clarifying the priority logic for conflict handling. The system stores the specific content of the conflict resolution strategies through rule table fields. The strategy content includes conflict judgment conditions and processing methods, ensuring that when conflicts arise between the time-period configurations of different rules, the system can automatically execute processing according to the preset strategy. During configuration, the system performs correlation checks on the time-period adaptation rules and conflict resolution strategies to ensure that the strategies effectively cover potential conflict scenarios.

[0049] S13: Configure the desensitization algorithm type and algorithm parameters for each rule entry in the structured rule table. Define the algorithm type in the desensitization algorithm type field of the structured rule table, and store the algorithm parameters in key-value pair format in the algorithm parameter field of the structured rule table.

[0050] Specifically, based on the physical characteristics and data flow scenarios of industrial data, the system configures a corresponding de-identification algorithm type for each rule entry in the structured rule table. The specific category of the algorithm is clearly defined in the de-identification algorithm type field of the structured rule table. The coverage of algorithm types includes the core forms required for industrial data de-identification, meeting the de-identification processing needs of different data variable types and different sensitivity levels. The system sets an algorithm parameter field in the structured rule table. This field stores the parameter information of the corresponding de-identification algorithm in a key-value pair format. The key name corresponds one-to-one with the specific parameter item of the algorithm, ensuring accurate matching when parameters are called. The parameter configuration content corresponding to the key-value pair strictly follows the algorithm's execution logic requirements, providing necessary parameter support for the correct operation of the algorithm during the de-identification calculation process. The configuration process of algorithm type and algorithm parameter strictly follows the combination logic of data variable type and time dimension type. That is, the same data variable type corresponds to different de-identification algorithm types and parameter configurations under different time dimension types. The algorithm type and algorithm parameter corresponding to each rule entry form a fixed association relationship. The association result is stored in the structured rule table. Subsequent de-identification calculation processes can directly retrieve and call the algorithm through the rule identifier, ensuring efficient execution of the de-identification algorithm and consistency of the de-identification effect.

[0051] S14: Set the data sensitivity level field for each rule entry in the structured rule table to define the data sensitivity level, set the node adaptation rule field to define the applicable node level, and set the applicable role set field to define the applicable role set.

[0052] Specifically, the system sets a data sensitivity level field for each rule entry in the structured rule table. This field explicitly defines the industrial data sensitivity level corresponding to each rule. The sensitivity level classification is based on the importance of industrial data in production and operation activities and the potential impact of data leakage, forming a standardized sensitivity level system. This system covers different protection levels required for industrial data security management, ensuring that all industrial data can find a corresponding sensitivity level classification. The system also sets a node adaptation rule field in the structured rule table. This field defines the node level to which the rule entry applies. The node level fully corresponds to the four-layer architecture of the industrial cluster network: workshop, enterprise, park, and cluster. During configuration, the system specifies one or more applicable node levels for each rule entry based on the application scope of the rule. The specification method is achieved through standardized node level identifiers, ensuring that the rule can be accurately invoked when data flows to the corresponding node level.

[0053] Furthermore, the system also sets an applicable role set field. This field defines the cluster role scope to which rule entries apply. The role set is divided based on the responsibilities and permissions of different positions in the industrial cluster, covering various core roles such as industrial production management, equipment maintenance, and data auditing. During configuration, the system clearly defines the applicable roles corresponding to each rule entry, forming an applicable role set through the combination of role identifiers, ensuring accurate matching between access permissions and roles. After completing the configuration of the above three fields, the system initiates a correlation verification process. During the verification process, the system checks the matching logic between data sensitivity level and node adaptation rules, ensuring that rules for high-sensitivity data only adapt to high-privilege node levels. It also verifies the correspondence between data sensitivity level and applicable role set, ensuring that access rules for high-sensitivity data are only open to high-privilege roles. In addition, it verifies the correlation between node adaptation rules and applicable role set, ensuring that the permission scope corresponding to the role is consistent with the responsibility scope of the node level. After successful verification, the system writes the configuration information into the corresponding fields of the rule table, completing the configuration of sensitivity level, node adaptation rules, and applicable role set. The configuration result provides a foundation for subsequent access control and rule matching, ensuring that access and processing of industrial data comply with security management logic.

[0054] S15: Associate compliance basis with each rule entry in the structured rule table, reference external laws and regulations or internal normative clauses in the compliance basis field of the structured rule table, set a version number field to record the rule version, and generate and record a version change log for each rule change, generating a versioned industrial standard data variable rule table.

[0055] Specifically, the system initiates a compliance basis association process for each rule entry in the structured rule table. This association is achieved through the compliance basis field of the structured rule table. This field references relevant clauses from external laws and regulations or internal standards. The reference method uses standardized clause identifiers, whose design integrates the unique number of the law or standard with the clause number, ensuring quick location of the corresponding clause content. The association of compliance basis is based on the security requirements and processing logic of the rule entry. The system retrieves clauses related to the rule entry from its built-in compliance clause library. This library contains various external laws and regulations and internal standards related to industrial data security, and it supports regular updates to keep pace with the latest compliance requirements. During the association process, the system verifies the compatibility between the rule entry and the compliance basis, ensuring that the content of the rule entry conforms to the requirements of the corresponding clause. After successful verification, the standardized clause identifier is written into the compliance basis field, ensuring that each rule entry has clear compliance support. The system sets a version number field in the structured rule table. The version number is generated using a preset encoding rule, which includes a major version number, a minor version number, and a revision number. The major version number corresponds to major changes in the rule table structure, the minor version number corresponds to the addition or adjustment of rule entry types, and the revision number corresponds to the modification or deletion of a single rule entry. The version number can clearly distinguish the different iteration stages of the rule table.

[0056] Preferably, the system performs version management operations for each rule change. When a rule entry in the rule table is created, modified, or deleted, the system automatically triggers the generation process of a version change log. The log includes information such as the time the change was triggered, the specific content of the change, the entity initiating the change, the rule status before the change, and the rule status after the change. The system associates the version change log with the rule table version number for storage. The storage adopts a distributed storage architecture to ensure the security and integrity of the log data. Each version of the rule table is associated with the corresponding change log, and it supports querying the corresponding change record by version number. Through the above field configuration and version management process, the system generates a versioned industrial standard data variable rule table. Version management supports backtracking operations on historical versions. When it is necessary to view or use a historical version of the rule table, the system can quickly retrieve the corresponding rule table and change log by version number. At the same time, version management provides complete data support for the auditing of rule changes, ensuring that the entire lifecycle of the rules is under supervision.

[0057] In one embodiment, step S2 of the desensitization method for industrial data provided by the present invention specifically includes the following steps:

[0058] S21: Standardize and timestamp the raw signals collected by sensors or programmable logic controllers at workshop nodes, convert the physical signals into structured industrial data points and attach the collection timestamps to generate raw industrial data with timestamps.

[0059] Specifically, the system establishes stable data transmission connections with sensors and programmable logic controllers (PLCs) at workshop nodes. These connections adhere to common industrial communication protocol specifications, ensuring the stability and reliability of raw signal transmission. The system receives raw signals collected by sensors and PLCs through pre-defined signal receiving interfaces. These raw signals include various physical signals reflecting the industrial production status. The system initiates standardization processing on the received raw signals. The parsing logic is based on the communication protocol type between the sensors and PLCs. The system has a built-in multi-protocol parsing module, capable of adapting to the signal output formats of different types of equipment, converting non-standardized raw signals into unified format digital signals.

[0060] Furthermore, based on the inherent correspondence between signals and industrial data, the system converts abstract electrical and optical signals into data expressions with industrial production significance. The conversion process strictly follows the mapping logic between signals and physical quantities, ensuring that the converted data reflects the actual production status. During signal processing, the system generates a real-time acquisition timestamp for each data point. Timestamp generation is synchronized with data acquisition, employing a globally unified time standard to ensure consistency and comparability of time information across data acquired from different devices at different times. The system integrates the standardized digital signals with the acquisition timestamps to construct structured industrial data points. Each data point includes the core data content after signal conversion, the acquisition timestamp, and basic identification information of the acquisition device.

[0061] S22: Based on the versioned industrial standard data variable rule table, feature extraction and rule mapping are performed on the original industrial data with timestamps. The data variable type is parsed, and the data is determined to be in working hours, non-working hours or overtime hours according to the time period definition in the industrial standard data variable rule table based on the collection timestamp, generating parsed data with time dimension attributes.

[0062] Specifically, the system invokes a versioned industrial standard data variable rule table to initiate the feature extraction process for the original industrial data with timestamps. During feature extraction, the system parses key feature information in the original industrial data, including the physical attribute features, signal output features, and data acquisition device association features. The extraction of these features is based on the system's built-in feature recognition logic, which is correlated with the data variable type definitions in the rule table. The system accurately parses the data variable types by comparing the feature information with the data variable type fields in the rule table. The parsing process uses a rule table index query mechanism, quickly locating the corresponding variable type definition through feature keyword matching, avoiding processing delays caused by full table scans.

[0063] After variable type parsing is completed, the system extracts the collection timestamp from the original industrial data and precisely compares the timestamp information with the time period adaptation rule field in the rule table. The comparison process follows the key-value pair format of the time period definition in the rule table, using a time range comparison algorithm to determine the time period type of the collection timestamp, i.e., working hours, off-hours, or overtime. During time period determination, if a timestamp simultaneously matches multiple time period definitions, the system will automatically handle the conflict resolution strategy configured in the rule table to ensure the uniqueness of the time period determination result. After completing variable type parsing and time period determination, the system initiates a rule mapping verification process to verify whether the parsed variable type and time period information match the corresponding entries in the rule table. It also verifies the consistency of the data format with the requirements of the rule table to avoid information mismatch caused by data transmission errors or parsing logic deviations. After successful verification, the system integrates the parsed variable type and time period information with the core data content in the original industrial data to generate parsed data with time dimension attributes. This data not only retains the core information of the original industrial data but also adds the parsed feature attributes and time dimension attributes.

[0064] S23: Perform statistical cycle classification and sensitivity level labeling on the parsed data with time dimension attributes. According to the preset daily, weekly and monthly cycle division rules, classify it into the corresponding statistical cycle, query the standard data variable rule table in the industrial field to assign it an initial sensitivity level, integrate equipment identification, variable type, time attribute and sensitivity level information to generate industrial data to be desensitized with multi-dimensional data identification.

[0065] Specifically, the system loads preset statistical period division rules, which are logically consistent with the time dimension type field in the standard data variable rule table for the industrial field, ensuring that period classification and the time configuration of the rule table are coordinated. The system performs statistical period classification processing on the parsed data with time dimension attributes. During the classification process, the system extracts the collection timestamp of the parsed data and compares it one by one with the preset daily, weekly, and monthly period division rules based on the date information corresponding to the timestamp. The system accurately classifies the data into the corresponding statistical period category through time-based attribution logic. After the statistical period classification is completed, the system again calls the versioned standard data variable rule table for the industrial field. By querying the corresponding rule entries in the rule table through the association information of the data variable type and collection scenario in the parsed data, the system assigns an initial sensitivity level to the parsed data based on the data sensitivity level field of the rule entries. The assignment of the initial sensitivity level strictly follows the definition standards in the rule table, ensuring that the sensitivity level labeling results are consistent for data of the same type and under the same scenario.

[0066] Furthermore, the system integrates and analyzes core information from the data, such as equipment identification, data variable type, time dimension attribute, and initial sensitivity level. It also supplements this with related information, including the workshop node identification to which the collected equipment belongs, data collection sequence identification, and data integrity verification results, constructing a multi-dimensional data identification system covering all data attributes. During the integration process, the system performs information correlation verification, checking the accuracy of the correspondence between various identification information and the data subject, the completeness of data fields, and whether the matching logic between sensitivity level, data variable type, and collection scenario conforms to the rule table requirements. After successful verification, the system generates industrial data to be de-identified, carrying multi-dimensional data identification. This data format conforms to the standardized requirements for cross-node transmission and can be directly recognized by subsequent enterprise node processing flows. The system stores the industrial data to be de-identified in a preset cache area of ​​the workshop node and simultaneously initiates a data transmission readiness check. After the check passes, the data is transmitted to the enterprise node according to the preset transmission priority, ensuring the orderly progress of the four-layer node collaborative de-identification calculation process.

[0067] In one embodiment, step S3 of the desensitization method for industrial data provided by the present invention specifically includes the following steps:

[0068] S31: Perform multi-dimensional deconstruction and correlation analysis on all entries in the versioned industrial standard data variable rule table, extract the core constraint dimensions of each rule entry such as data variable type, time dimension type, applicable node level, applicable role set and data sensitivity level, and generate dimension deconstruction results.

[0069] Specifically, the system reads each rule entry in the rule table sequentially and performs multi-dimensional deconstruction operations on each entry based on the field structure of the rule table. During the deconstruction process, the system accurately extracts the core constraint dimensions corresponding to each rule entry, including data variable type, time dimension type, applicable node level, applicable role set, and data sensitivity level. The extraction logic strictly follows the definition specifications of each field in the rule table to ensure that the extracted dimension information is completely consistent with the content of the rule entry. For the data variable type dimension, the system extracts the complete definition information of the data variable type field of the rule entry; for the time dimension type dimension, it extracts the explicit time scenario classification in the time dimension type field; for the applicable node level dimension, it extracts all applicable levels defined in the node adaptation rule field; for the applicable role set dimension, it extracts all role identifiers contained in the applicable role set field; and for the data sensitivity level dimension, it extracts the sensitivity level set in the data sensitivity level field.

[0070] After extracting the dimensions of a single rule entry, the system performs analysis on the relationships between the dimensions of that entry. The analysis includes the logical adaptability of each dimension and the consistency of its applicable scope, ensuring that the correspondence between data variable types and applicable node levels, applicable role sets, and data sensitivity levels conforms to the logic of industrial data security management and avoids dimension conflicts. The system integrates the core constraint dimension extraction results of each rule entry with the correlation analysis conclusions to form a dimension deconstruction record for that single rule entry. The deconstruction records of all rule entries are then summarized to generate the dimension deconstruction result.

[0071] S32: Based on the dimensional deconstruction results and the dimensional space involved in the industrial data to be desensitized carrying multi-dimensional data identifiers, perform batch generation of six-dimensional binding records. For each rule entry, under all applicable node levels and role combinations, create a binding record that associates the specific dimension value with the unique identifier of the rule entry, and generate the original binding relationship set.

[0072] Specifically, the system retrieves dimensional deconstruction results from a pre-set database and simultaneously acquires the dimensional space information involved in the industrial data to be de-identified, which carries multi-dimensional data identifiers. This dimensional space corresponds to the core constraint dimensions in the deconstruction results, ensuring that the generated binding records can cover the dimensional combinations required for actual data processing. Based on the dimensional deconstruction results, for each rule entry, the system extracts all elements from its applicable node level and applicable role set, and initiates the combination operation of applicable node level and role. The operation logic forms an independent combination between each applicable node level and each applicable role, ensuring that no valid level and role matching relationship is missed.

[0073] After the combination operation is completed, the system creates a binding record for each combination. Each binding record contains the specific dimension value corresponding to the combination and a unique identifier for the rule entry. The specific dimension value covers the data variable type, time dimension type, applicable node level, applicable role, and data sensitivity level. The unique identifier for the rule entry is associated with the corresponding rule table entry, ensuring a precise mapping between the binding record and the rule entry. The system performs batch processing on all rule entries according to the above logic. A parallel processing mechanism is used during batch generation to improve the efficiency of binding record generation and avoid processing delays caused by a large number of rule entries. After all binding records are generated, the system integrates them into an original binding relationship set. The set is stored in a database table structure to ensure the orderliness and queryability of data storage.

[0074] S33: Perform index optimization and function encapsulation on the original binding relationship set. Create a composite database index based on key fields such as data variable type, time dimension type, node level, role identifier, and data sensitivity level. Define a query function that receives these five dimension parameters, quickly retrieves and returns the most matching rule entries through the index, and completes the six-dimensional binding relationship model.

[0075] Specifically, the system initiates the creation process of a composite database index for key fields in the original set of binding relationships. These key fields include data variable type, time dimension type, node level, role identifier, and data sensitivity level. During index creation, the system combines and sorts these five key fields according to preset index building rules, forming the core structure of the composite index. The combination and sorting logic is based on the frequency of use and matching priority of each field in subsequent query scenarios, ensuring that the index maximizes query efficiency. The system uses the database's built-in index optimization tools to perform performance optimization operations on the composite index. Optimization directions include index structure adjustment and query path planning, reducing resource consumption and response time during index queries, ensuring that the index can quickly adapt to high-concurrency query scenarios.

[0076] After the composite index is created, the system receives five parameters: data variable type, time dimension type, node level, role identifier, and data sensitivity level. Based on these parameters, it quickly retrieves the original binding relationship set and returns the unique identifier of the most matching rule entry. During function definition, the system clearly defines the parameter receiving format and data type to ensure the accuracy of parameter transmission. The retrieval logic employs a composite index calling mechanism. After receiving parameters, the function automatically converts them into index query conditions, locates matching records in the original binding relationship set through the index, and follows a preset matching priority rule. When multiple matching records exist, the optimal result is returned according to priority. The returned result includes the unique identifier of the most matching rule entry and the associated core rule information, ensuring that subsequent processes can quickly obtain the corresponding de-identification rules.

[0077] Furthermore, the system encapsulates the query function, providing a standardized calling interface to support calls from system components at various nodes. The function also incorporates built-in exception handling logic, returning standardized error messages when incomplete parameters or no query results occur. After completing the creation of the composite index and the encapsulation of the query function, the system integrates the original binding relationship set, the composite index, and the query function to construct a six-dimensional binding relationship model.

[0078] In one embodiment, such as Figure 2 As shown, step S4 of the desensitization method for industrial data provided by this invention specifically includes the following steps:

[0079] S41: Industrial data to be desensitized, carrying multi-dimensional data identifiers, undergoes rapid transformation processing based on real-time scenarios and single-point data. At the workshop node, the corresponding millisecond-level desensitization algorithm is matched according to the data variable type, real-time time attribute, and initial sensitivity level to quickly process the original data to remove some sensitive details and generate a workshop-level desensitized data package with complete metadata.

[0080] Specifically, the system receives industrial data to be de-identified, carrying multi-dimensional data identifiers, through a pre-defined communication interface on the workshop node. The communication interface supports adaptation to various industrial communication protocols. The system's built-in protocol parsing module performs format conversion on the transmitted data, ensuring that data output from different devices can be uniformly identified. The transmission process strictly adheres to industrial data transmission security standards to prevent leakage risks during data transmission. The system performs multi-dimensional identifier information extraction on the received data, extracting data variable types, real-time time attributes, initial sensitivity levels, and device association information. This extracted information serves as the core input parameters for rule matching, providing a complete data foundation for subsequent algorithm matching. The system calls the query function of the six-dimensional binding relationship model, passing the core input parameters to the function in a pre-defined order. The function traverses the original binding relationship set using a composite index to quickly locate rule entries matching the parameters. The matching process strictly follows the constraints of node adaptation rules and time dimension types in the rule table, ensuring that the millisecond-level de-identification algorithm obtained is compatible with the real-time processing scenario of the workshop node, avoiding de-identification failure or data distortion caused by algorithm mismatch with the processing scenario. Preferably, the de-identification algorithm used by the system follows a general modified formula:

[0081]

[0082] in, The original data to be anonymized. The algorithm parameters represent the configuration of the rule entries. This represents a desensitization mapping function defined based on the algorithm type. The algorithm uses this formula to quickly transform sensitive details of single-point industrial data, removing sensitive information while preserving the basic usability and application value of the data in industrial production scenarios.

[0083] Furthermore, the system synchronously collects metadata about the processing process. This metadata includes information such as the desensitization algorithm type, algorithm parameters, processing timestamp, unique identifier for rule entries, rule version number, changes in sensitivity level before and after data processing, workshop node identifier, and equipment identifier, ensuring traceable records for each processing step. The system integrates and encapsulates the desensitized core data with the collected metadata in a standardized format. This standardized format specifies the order of data fields, encoding methods, and storage formats, ensuring the compatibility and integrity of data packets during cross-node transmission.

[0084] After encapsulation, the system performs real-time verification on the generated workshop-level de-identified data packets. The verification includes the consistency between the data transformation results and the rule requirements, the integrity of metadata fields, and the standardization of the data packet format. If data is found to be non-compliant during the verification process, the system marks the data packet as abnormal and re-executes the de-identification process. Data packets that pass the verification are stored in the local cache area of ​​the workshop node. The cache area adopts a data classification storage mechanism and is partitioned and managed according to data variable type and statistical period. At the same time, the system pushes the data packets to the enterprise node according to the preset transmission strategy. The transmission strategy dynamically adjusts the transmission order based on network bandwidth status and data priority.

[0085] S42: Perform multi-variable collaborative desensitization processing on workshop-level desensitized data packages from multiple workshops based on statistical cycles and production line logic. At the enterprise node, perform spatiotemporal alignment and correlation analysis on multi-source data from the same cycle and the same production line, and apply collaborative algorithms to eliminate inference risks between variables, generating enterprise-level desensitized data packages with correlation protection capabilities.

[0086] Specifically, the system receives workshop-level de-identified data packets from multiple workshop nodes in batches through the aggregation interface of the enterprise node. The aggregation interface supports high-concurrency data reception and can handle transmission requests from multiple workshop nodes simultaneously. During the reception process, the system performs integrity verification and source legitimacy verification on each data packet. Integrity verification is achieved by comparing the verification information of the data packet with preset verification standards. Source legitimacy verification compares the information such as the workshop node identifier and rule version number contained in the data packet with the system's filing information. Data packets that pass the verification proceed to the subsequent processing flow, while data packets that fail the verification trigger a retransmission mechanism. The system sends a retransmission request to the corresponding workshop node to ensure that no data is lost.

[0087] The system extracts metadata and core data from the anonymized data packages at the workshop level. Based on the statistical period identifier and production line association information in the metadata, it initiates a data classification and organization process. This process initially groups the data according to the statistical period identifier, and then further groups the data within the same statistical period according to the production line association information. This ensures that multi-source data from the same period and the same production line are grouped into the same processing set, facilitating subsequent collaborative anonymization. The system performs spatiotemporal alignment on the multi-source data within the same processing set. Time alignment is based on the collection timestamps in the data packages, using a time calibration algorithm to adjust for time recording discrepancies between different data. Spatial alignment is based on production line spatial association information, using a spatial logic matching algorithm to determine the corresponding positional relationship of data within the production line, eliminating spatiotemporal discrepancies between multi-source data and ensuring data correlation and consistency. The system then performs correlation analysis on the spatiotemporally aligned multi-source data, using a data correlation identification algorithm to uncover the inherent correlation logic and feature mapping relationships between data. The analysis results serve as an important basis for collaborative anonymization, eliminating the risk of inference between variables.

[0088] Preferably, the system invokes a six-dimensional binding relationship model, inputting parameters such as the statistical period of the processing set, data variable type, and applicable role set into the model to match and obtain the collaborative de-identification algorithm and parameters corresponding to the enterprise node. The system then initiates the collaborative de-identification algorithm execution process, which follows the associated data processing formula:

[0089]

[0090] in, This represents processing data from a single variable within a collection. This represents a set of related data within a processing set. Representative parameters of the collaborative desensitization algorithm, This represents the collaborative de-identification mapping function. The algorithm uses this formula to uniformly process related data sets, eliminating potential association inference risks after single-variable de-identification through collaborative data transformation, ensuring data security in multi-variable scenarios. During processing, the system collects enterprise-level de-identification metadata, including collaborative algorithm type, algorithm parameters, processing timestamp, statistical cycle identifier, production line identifier, related data processing records, unique identifiers for rule entries, rule version numbers, and enterprise node identifiers. The metadata and workshop-level metadata are linked through unique data identifiers to construct a complete processing trajectory record. The system integrates and encapsulates the collaboratively de-identified core data and enterprise-level metadata in a standardized format, generating an enterprise-level de-identified data package with association protection capabilities. After encapsulation, the system performs a secondary verification, checking the consistency of de-identification of related data, the effectiveness of eliminating inference risks, and the integrity of metadata. Verified enterprise-level de-identified data packages are stored in a preset database on the enterprise node. The database uses a partitioned storage strategy to improve data read and write efficiency. Simultaneously, the system pushes the data package to the park node according to a preset transmission protocol. The transmission protocol supports data compression transmission to reduce network bandwidth consumption.

[0091] S43: Perform secondary collaborative processing on enterprise-level de-identified data packages from multiple enterprises based on a unified time period standard and cross-entity data fusion. Unify the time period definition of each enterprise within the jurisdiction at the park node, and aggregate and perform secondary de-identification on cross-enterprise data of the same period to eliminate the characteristic differences between enterprises, and generate park-level de-identified data packages that meet the park's regulatory requirements.

[0092] Specifically, the system receives enterprise-level de-identified data packets from multiple enterprise nodes in batches through the cross-enterprise aggregation gateway at the park node. The cross-enterprise aggregation gateway has multi-channel data reception capabilities and can process transmission data from different enterprise nodes simultaneously. During the reception process, the system verifies the format, integrity, and rule version compatibility of the data packets. Format verification ensures that the data packets conform to standardized transmission specifications, integrity verification is achieved through data checksum comparison, and rule version compatibility verification is achieved by comparing the rule version number in the data packet with the currently effective rule version number at the park node. After passing the verification, the data packets are classified and stored according to enterprise identifier and statistical period. The storage adopts a distributed storage architecture to ensure the reliability and scalability of data storage.

[0093] Furthermore, the system extracts time period configuration information from the metadata of each enterprise-level de-identified data package, including the start and end times of working hours, non-working hours, and overtime hours. Then, it compares the time period adaptation rules and conflict resolution strategies in the standard data variable rule table for the industrial field to uniformly adjust the time period definitions of different enterprises, eliminating time period differences between enterprises. The unified time period standard is stored in the local configuration library of the park node. At the same time, the system synchronizes the unified time period standard to each enterprise node through a data synchronization mechanism to ensure the consistency of the time dimension in subsequent cross-enterprise data processing. The system extracts the core data and metadata from each enterprise-level de-identified data package. Based on the unified time period standard and statistical period identifier, it aggregates cross-enterprise data of the same period. The aggregation logic classifies and summarizes the data according to data variable type, production line association attribute, and statistical period. First, the data is grouped by data variable type, then the data of the same variable type is subdivided according to production line association attribute, and finally the data is integrated according to statistical period to form a cross-enterprise data set.

[0094] Preferably, the system invokes a six-dimensional binding relationship model, inputting dimensional information from the cross-enterprise data set, including data variable types, unified time period attributes, statistical periods, park node levels, and applicable regulatory roles, into the model to obtain a park-level secondary collaborative data masking algorithm and parameters. The system then initiates the execution process of the secondary collaborative data masking algorithm, which follows the cross-enterprise data fusion formula:

[0095]

[0096] in, Represents a cross-enterprise data set, Represents secondary synergistic desensitization parameters. This represents the data fusion and de-identification mapping function. The algorithm uses this formula to process the feature differences of cross-enterprise data sets. Through data fusion and secondary transformation, it eliminates the data feature identification between enterprises, ensuring privacy and security in cross-enterprise data sharing scenarios. During processing, the system collects park-level de-identified metadata, including secondary collaboration algorithm type, algorithm parameters, unified time period records, cross-enterprise data aggregation records, processing timestamps, unique identifiers for rule entries, rule version numbers, park node identifiers, and other information. The metadata and enterprise-level metadata form a full-link association through unique data identifiers.

[0097] Preferably, the system integrates and encapsulates the core data after secondary collaborative de-identification with the park-level metadata in the format required by park supervision, generating a park-level de-identified data package. After encapsulation, the system performs compliance verification, which is based on park supervision rules, compliance clauses, and data de-identification strength requirements. The park-level de-identified data package that passes the verification is stored in the park node distributed storage system. The storage system adopts a data multi-copy backup mechanism to ensure data security. At the same time, the system pushes the data package to the cluster nodes according to a preset strategy. The push strategy is dynamically adjusted according to the importance of the data and the transmission priority.

[0098] S44: Strengthen the processing of campus-level de-identified data packets from multiple campuses based on long-term statistics and global compliance requirements. Perform long-term data aggregation across campuses on cluster nodes and match and apply the final de-identification algorithm that meets the privacy protection strength requirements according to the latest compliance requirements to generate the final de-identified data. Integrate workshop-level de-identified data packets, enterprise-level de-identified data packets and campus-level de-identified data packets to generate full-link traceability logs.

[0099] Specifically, the system receives large-scale, de-identified data packets from multiple park nodes via the cloud access interface of the cluster nodes. The cloud access interface supports elastic scaling to adapt to data transmission needs of parks of varying sizes. During reception, the system employs a distributed reception mechanism, processing transmitted data in parallel through multiple data receiving nodes to improve the aggregation efficiency of data from multiple parks. The system performs integrity checks, source legitimacy checks, and rule version consistency checks on each park-level de-identified data packet. Integrity checks are performed by comparing preset fields with actual fields in the data packet; source legitimacy checks are performed by verifying the registration information of the park node identifier; and rule version consistency checks are performed by comparing the rule version number in the data packet with the currently effective rule version number of the cluster node. Data packets that pass the checks enter the long-cycle data aggregation process. The system extracts core data and metadata from the de-identified data packages at each park level, and initiates cross-park long-term data aggregation processing based on statistical period identifiers. The aggregation logic classifies and integrates data according to data variable type, statistical period level, and regional association attribute. First, the data is initially classified according to data variable type, then the data of the same variable type is hierarchically divided according to statistical period level, and finally the data of the same level is integrated according to regional association attribute to form a long-term data set covering multiple parks.

[0100] Furthermore, the system invokes a versioned industrial standard data variable rule table, extracts the latest compliance clauses and global privacy protection strength requirements through the rule parsing module, and inputs the dimensional information of the long-term data set and compliance requirements into a six-dimensional binding relationship model. The model uses a composite index to retrieve and match the final de-identification algorithm and parameters suitable for the cluster nodes. The matching process focuses on verifying the consistency between data processing and global compliance requirements to ensure that the algorithm strength meets privacy protection standards. The system then initiates the final de-identification algorithm execution process, which follows the compliance enhancement formula:

[0101]

[0102] in, Represents a long-term data set across multiple industrial parks. Represents parameters for enhanced compliance. This represents the compliance-based data masking mapping function. The algorithm uses this formula to perform enhanced data masking on long-term datasets across different industrial parks. The processing logic is based on compliance requirements and algorithm parameters, achieving global privacy protection through data enhancement and transformation while ensuring that the data meets the basic requirements for subsequent statistical analysis and compliance auditing. During processing, the system collects cluster-level data masking metadata, including the final data masking algorithm type, algorithm parameters, long-term data aggregation records, compliance clause identifiers, processing timestamps, unique identifiers for rule entries, rule version numbers, cluster node identifiers, and other information. The system integrates data masking metadata at the workshop, enterprise, park, and cluster levels, tracing the entire data processing trajectory chronologically and by node level. The metadata integration covers algorithm execution records, rule matching records, data transformation records, time period processing records, statistical cycle correlation records, compliance verification records, and other information at each level.

[0103] The system generates a full-chain traceability log from the integrated metadata according to a standardized log format. The log includes core content such as unique data identifiers, processing information at each node, rule application information, compliance basis information, and data sensitivity level change trajectories. The system defines the core data after cluster-level enhanced anonymization as the final anonymized data, establishes a connection with the full-chain traceability log through unique data identifiers, and stores it in a distributed data warehouse on the cluster nodes. The storage adopts a multi-replica backup mechanism to ensure data security. The data warehouse supports query and audit operations by data identifier, time range, node level, compliance basis, and other dimensions, providing support for the security management and compliance traceability of industrial data throughout its entire lifecycle.

[0104] In one embodiment, step S5 of the desensitization method for industrial data provided by the present invention specifically includes the following steps:

[0105] S51: Authentication and parsing of the user's identity credentials and access request context in the data access request, verifying the legitimacy of the user's identity, and parsing out the data variable type requested for access, the expected time dimension, the user's node level, role, and expected data sensitivity level, and generating a legitimate access context.

[0106] Specifically, the system receives data access requests submitted by users, extracts identity credentials and access request context information from the request message. The identity credentials include a user identifier and authentication carrier information. The user identifier uses a system-uniformly assigned encoding format, and the authentication carrier information is the verification information pre-registered by the user in the system. The access request context covers the target data query conditions and a description of the expected data format. The target data query conditions include information such as the data-related equipment and production process, while the description of the expected data format specifies the user's requirements for the data output format. The system performs an identity validity verification process, comparing the extracted identity credentials with a unified identity authentication database stored on the cluster nodes field by field. The verification logic follows a preset identity authentication protocol, which specifies the comparison order and matching criteria for each field of the identity credentials. If the verification result is invalid, the system returns an identity authentication failure message and records an abnormal access log, including the access time, user identifier, request content, and reason for failure.

[0107] When the verification result is valid, the system parses the access request context, matches keywords in the target data query conditions with data variable type identifiers in the rule table to determine the type of data variable requested for access; determines the user's node level based on the organization information associated with the user identifier; extracts the user role based on the binding relationship between user identifiers and roles in the identity authentication database; establishes a mapping between the description of data precision in the user request and the sensitivity level identifier in the rule table to determine the expected data sensitivity level; and parses the expected time dimension, including time period attributes and statistical period types, in conjunction with the time description in the query conditions. The system organizes all parsing results in a unified format to generate a valid access context, with each dimension using standardized encoding to maintain consistency with the dimension definitions of the six-dimensional binding relationship model.

[0108] S52: Based on the legitimate access context and the current system time of the user, fine-grained permission matching is performed through a six-dimensional binding relationship model. Rule entries that simultaneously satisfy constraints of data dimension, time dimension, node level, user role and sensitivity level are searched, and a set of rule matching results containing precise access control and decryption rules is generated.

[0109] Specifically, the system extracts the dimensional codes of the legitimate access context, obtains the current system time through the time synchronization mechanism of the cluster nodes, and converts the current system time into the corresponding time dimension code according to the time period adaptation rules in the rule table, supplementing it into the dimension information for permission matching. The system combines the data variable type identifier, the time dimension code corresponding to the current system time, the user's node level identifier, the user role type identifier, and the expected data sensitivity level identifier in a fixed order to form standardized query parameters, which are then passed to a dedicated query function of the six-dimensional binding relationship model. The query function traverses the original binding relationship set through a composite database index, filtering rule entries that simultaneously satisfy all dimensional constraints. The filtering process follows the principle of exact matching first; when there are no exact matching entries, filtering is performed according to preset degenerate matching rules.

[0110] Furthermore, the system deduplicates the matched rule entries, removing duplicate records, and then sorts them according to the compliance priority and sensitivity level suitability of the corresponding compliance basis. The compliance priority is determined based on the current validity of the legal provisions, and the sensitivity level suitability is determined based on the degree of correspondence between the rule sensitivity level and the expected sensitivity level. After sorting, a rule matching result set containing precise access control rules and decryption rules is generated. The access control rules specify the data level and access permission type that the user is allowed to access, and the decryption rules include the decryption algorithm type, key derivation parameters, and allowed precision level of restoration. The rule matching result set is uniquely associated with the user identifier and the legitimate access context.

[0111] S53: Based on the access control and decryption rules in the rule matching result set and the full-link traceability log, perform targeted data decryption and precision restoration processing. According to the access granularity specified by the rule, find the corresponding level of data form from the log, use the key derived from the rule to decrypt, or when authorized, reverse deduce higher precision data based on the desensitization algorithm of the log record, and generate result data that perfectly matches the current user's permissions and access time.

[0112] Specifically, the system loads the rule matching result set and the full-link traceability log, extracts access control rules and decryption rules from the result set, and clarifies the data levels that users are allowed to access and the corresponding decryption parameters. Based on the access granularity specified by the access control rules, and combined with the data variable type identifiers and time dimension information in the legal access context, the system traverses the full-link traceability log, and locates the corresponding level of de-identified data form and related storage indexes and metadata records through multi-dimensional joint retrieval. The system initiates a targeted decryption process, deriving parameters according to the key in the decryption rules, and generating a key by combining the user identifier and access time. This key and the decryption algorithm type serve as the input to the decryption function, performing decryption operations on the de-identified data located in the traceability log. If the access control rules allow restoration to higher-precision data, the system extracts the corresponding de-identification algorithm type and parameters from the full-link traceability log, calls the inverse function of the de-identification algorithm, performs reverse derivation on the de-identified data, and restores the higher-precision data.

[0113] The system performs a consistency check between the decrypted or reverse-derived data and the access context, verifying the match between the data's variable type, time dimension, sensitivity level, and the user's request. It also checks if the data format meets the user's expectations. If the check passes, the system encapsulates the data in the user-expected format, generating result data. This result data includes the data subject, data source hierarchy identifier, data generation timestamp, and compliance reference information. If the check fails, the system returns a data adaptation failure message and records the processing log. The result data is then sent back to the user through a secure channel, with data encryption performed during transmission.

[0114] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0115] Based on the same inventive concept, this application also provides an industrial data desensitization system for implementing the aforementioned industrial data desensitization method. The solution provided by this system is similar to the implementation described in the above method; therefore, the specific limitations of one or more industrial data desensitization system embodiments provided below can be found in the limitations of the industrial data desensitization method described above, and will not be repeated here.

[0116] Preferably, such as Figure 3 As shown, the present invention provides an industrial data anonymization system 600, which is configured with the following modules:

[0117] The data rule definition module 610 is used to define and organize the types, time dimensions, desensitization algorithms and compliance basis of industrial data variables in a rule-based manner. It configures the corresponding desensitization algorithms, sensitivity levels and access rules for each type of industrial variable under different time scenarios and node levels, and generates a versioned standard data variable rule table for the industrial field.

[0118] The data to be desensitized generation module 620 is used to timestamp and analyze the features of the raw industrial data collected at the workshop node. Based on the standard data variable rule table in the industrial field, it adds the data variable type, the time period attribute and statistical period corresponding to the collection time, and the initial sensitivity level determined based on the rule table to each piece of raw industrial data, generating industrial data to be desensitized carrying multi-dimensional data identifiers.

[0119] The six-dimensional relationship modeling module 630 is used to model the relationship between the standard data variable rule table in the industrial field and the dimensions involved in the industrial data to be desensitized. It associates and binds the data variable type, time dimension, node level, cluster role, data sensitivity level with the specific entries in the rule table in multiple dimensions to build a six-dimensional binding relationship model.

[0120] The collaborative desensitization processing module 640 is used to perform collaborative desensitization calculations in the four-level nodes of the industrial cluster network (workshop, enterprise, park, and cluster) based on the industrial data to be desensitized and the six-dimensional binding relationship model. It performs calculations from fine-grained to coarse-grained and from local to global. The desensitization processing results of each level are used as inputs for the next level. After the cluster node completes the desensitization processing, it integrates the processing metadata of each level to generate a full-link traceability log.

[0121] The data access decryption module 650 is used to match access permissions and decryption rules based on the user's data access request and the current access time through a six-dimensional binding relationship model, and to perform targeted decryption and precision restoration based on the full-link traceability log, generating result data that matches the current user's permissions and access time.

[0122] Preferably, the data rule definition module 610 provided in this application is configured with the following units:

[0123] The rule table structure creation unit is used to create structured rule tables on cluster node servers, generate a unique rule identifier for each rule entry in the structured rule table, set a data variable type field in the structured rule table to define the physical type and unit of industrial data, and set a time dimension type field to define the time dimension type.

[0124] The time period rule configuration unit is used to configure time period adaptation rules for each rule entry in the structured rule table. The start and end times of working hours, non-working hours and overtime hours are defined in key-value pair format in the time period adaptation rule field of the structured rule table, and conflict resolution strategies are configured.

[0125] The desensitization algorithm configuration unit is used to configure the desensitization algorithm type and algorithm parameters for each rule entry in the structured rule table. The algorithm type is defined in the desensitization algorithm type field of the structured rule table, and the algorithm parameters are stored in key-value pair format in the algorithm parameter field of the structured rule table.

[0126] The rule adaptation configuration unit is used to set the data sensitivity level field for each rule entry in the structured rule table to define the data sensitivity level, set the node adaptation rule field to define the applicable node level, and set the applicable role set field to define the applicable role set.

[0127] The rule version management unit is used to associate compliance basis with each rule entry in the structured rule table, reference external laws and regulations or internal normative clauses in the compliance basis field of the structured rule table, set the version number field to record the rule version, and generate and record the version change log for each rule change, generating a versioned industrial standard data variable rule table.

[0128] Preferably, the data generation module 620 to be de-identified provided in this application is configured with the following units:

[0129] The raw data time stamping unit is used to standardize and timestamp the raw signals collected by sensors or programmable logic controllers at workshop nodes, converting physical signals into structured industrial data points and attaching collection timestamps to generate raw industrial data with timestamps.

[0130] The data rule mapping and parsing unit is used to perform feature extraction and rule mapping processing on the original industrial data with timestamps based on the versioned industrial standard data variable rule table, parse the data variable type to which it belongs, and determine whether it is in the working period, non-working period or overtime period according to the time period definition in the industrial standard data variable rule table based on the collection timestamp, and generate parsed data with time dimension attributes.

[0131] The data classification and sensitivity labeling unit is used to classify and label the sensitivity level of analytical data with time dimension attributes according to statistical cycles. It classifies the data into the corresponding statistical cycle according to the preset daily, weekly and monthly cycle division rules, and queries the standard data variable rule table in the industrial field to assign an initial sensitivity level. It integrates equipment identification, variable type, time attribute and sensitivity level information to generate industrial data to be desensitized with multi-dimensional data identification.

[0132] Preferably, the six-dimensional relationship modeling module 630 provided in this application is configured with the following units:

[0133] The rule dimension deconstruction and analysis unit is used to perform multi-dimensional deconstruction and correlation analysis on all entries in the versioned industrial standard data variable rule table. It extracts the core constraint dimensions of each rule entry, such as data variable type, time dimension type, applicable node level, applicable role set, and data sensitivity level, and generates dimension deconstruction results.

[0134] The batch binding record generation unit is used to perform batch generation of six-dimensional binding records based on the dimensional deconstruction results and the dimensional space involved in the industrial data to be desensitized carrying multi-dimensional data identifiers. For each rule entry, under all applicable node levels and role combinations, a binding record is created that associates the specific dimension value with the unique identifier of the rule entry, generating the original binding relationship set.

[0135] The binding model optimization building unit is used to optimize the index and encapsulate the function of the original binding relationship set. It creates a composite database index based on key fields such as data variable type, time dimension type, node level, role identifier and data sensitivity level, and defines a query function that receives these five dimension parameters, quickly retrieves and returns the most matching rule entries through the index, thus completing the six-dimensional binding relationship model.

[0136] Preferably, the collaborative desensitization processing module 640 provided in this application is configured with the following units:

[0137] The workshop-level rapid desensitization processing unit is used to rapidly transform industrial data to be desensitized based on real-time scenarios and single-point data, which carries multi-dimensional data identifiers. At the workshop node, the corresponding millisecond-level desensitization algorithm is matched according to the data variable type, real-time time attribute and initial sensitivity level to quickly process the original data to remove some sensitive details and generate a workshop-level desensitized data package with complete metadata.

[0138] The enterprise-level collaborative desensitization processing unit is used to perform multi-variable collaborative desensitization processing on workshop-level desensitized data packages from multiple workshops based on statistical cycles and production line logic. At the enterprise node, it performs spatiotemporal alignment and correlation analysis on multi-source data from the same cycle and the same production line, and applies collaborative algorithms to eliminate the inference risk between variables, generating enterprise-level desensitized data packages with correlation protection capabilities.

[0139] The park-level secondary collaborative desensitization unit is used to perform secondary collaborative processing on enterprise-level desensitized data packages from multiple enterprises based on a unified time period standard and cross-entity data fusion. It unifies the time period definition of each enterprise within the jurisdiction at the park node, and aggregates and performs secondary desensitization on cross-enterprise data of the same period to eliminate the feature differences between enterprises and generate park-level desensitized data packages that meet the park's regulatory requirements.

[0140] The cluster-level compliance de-identification and log integration unit is used to strengthen the processing of campus-level de-identified data packets from multiple campuses based on long-term statistics and global compliance requirements. It performs cross-campus long-term data aggregation on cluster nodes and matches and applies the final de-identification algorithm that meets the privacy protection strength requirements according to the latest compliance requirements to generate the final de-identified data. It also integrates the processing metadata of workshop-level, enterprise-level, and campus-level de-identified data packets to generate full-link traceability logs.

[0141] Preferably, the data access decryption module 650 provided in this application is configured with the following units:

[0142] The access request authentication and parsing unit is used to authenticate and parse the identity credentials and access request context in the user's data access request, verify the legitimacy of the user's identity, and parse out the data variable type requested for access, the expected time dimension, the user's node level, role, and expected data sensitivity level, and generate a legitimate access context.

[0143] The access permission rule matching unit is used to perform fine-grained permission matching based on the legitimate access context and the current system time of the user, through a six-dimensional binding relationship model. It searches for rule entries that simultaneously satisfy constraints of data dimension, time dimension, node level, user role and sensitivity level, and generates a set of rule matching results containing precise access control and decryption rules.

[0144] The targeted decryption and precision restoration unit is used to perform targeted data decryption and precision restoration based on access control and decryption rules and full-link traceability logs in the rule matching result set. According to the access granularity specified by the rule, it finds the data form of the corresponding level from the log, decrypts it using the key derived from the rule, or, when authorized, reverse-engineers higher precision data based on the desensitization algorithm of the log record, generating result data that perfectly matches the current user's permissions and access time.

[0145] In one embodiment, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method for desensitizing industrial data.

[0146] In one embodiment, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for desensitizing industrial data.

[0147] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0148] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0149] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for desensitizing industrial data, characterized in that, Includes the following steps: S1: Define and organize the types, time dimensions, desensitization algorithms and compliance criteria of industrial data variables in a rule-based manner. Configure the corresponding desensitization algorithms, sensitivity levels and access rules for each type of industrial variable under different time scenarios and node levels, and generate a versioned standard data variable rule table for the industrial field. S2: The raw industrial data collected at the workshop node is timestamped and its features are parsed. Based on the industrial standard data variable rule table, each piece of raw industrial data is appended with its data variable type, the time period attribute and statistical period corresponding to the collection time, and the initial sensitivity level determined based on the rule table, to generate industrial data to be desensitized carrying multi-dimensional data identifiers. S3: Perform relationship modeling on the industrial standard data variable rule table and the dimensions involved in the industrial data to be desensitized, and associate and bind the data variable type, time dimension, node level, cluster role, data sensitivity level with the specific entries in the rule table in multiple dimensions to construct a six-dimensional binding relationship model; S4: Based on the industrial data to be desensitized and the six-dimensional binding relationship model, in the four-level nodes of workshop, enterprise, park and cluster in the industrial cluster network, the collaborative desensitization calculation is performed in sequence from fine-grained to coarse-grained and from local to global. The desensitization processing result of each level is used as the input of the next level. After the desensitization processing is completed at the cluster node, the processing metadata of each level is integrated to generate a full-link traceability log. S5: Based on the user's data access request and current access time, the access permissions and decryption rules are matched through the six-dimensional binding relationship model, and targeted decryption and precision restoration are performed based on the full-link traceability log to generate result data that matches the current user's permissions and access time.

2. The method according to claim 1, characterized in that, S1 includes: S11: Create a structured rule table on the cluster node server, generate a unique rule identifier for each rule entry in the structured rule table, set a data variable type field in the structured rule table to define the physical type and unit of industrial data, and set a time dimension type field to define the time dimension type; S12: Configure time period adaptation rules for each rule entry in the structured rule table. Define the start and end times of working hours, non-working hours, and overtime hours in key-value pair format in the time period adaptation rule field of the structured rule table, and configure conflict resolution strategies. S13: Configure a desensitization algorithm type and algorithm parameters for each rule entry in the structured rule table, define the algorithm type in the desensitization algorithm type field of the structured rule table, and store the algorithm parameters in key-value pair format in the algorithm parameter field of the structured rule table; S14: Set a data sensitivity level field for each rule entry in the structured rule table to define the data sensitivity level, set a node adaptation rule field to define the applicable node level, and set an applicable role set field to define the applicable role set; S15: Associate compliance basis with each rule entry in the structured rule table, reference external laws and regulations or internal normative clauses in the compliance basis field of the structured rule table, set a version number field to record the rule version, and generate and record a version change log for each rule change, thereby generating a versioned industrial standard data variable rule table.

3. The method according to claim 1, characterized in that, S2 includes: S21: Standardize and timestamp the raw signals collected by sensors or programmable logic controllers at workshop nodes, convert the physical signals into structured industrial data points and attach the collection timestamps to generate raw industrial data with timestamps. S22: Based on the versioned industrial standard data variable rule table, feature extraction and rule mapping are performed on the original industrial data with timestamps, the data variable type to which it belongs is parsed, and according to the collection timestamp, the time period definition in the industrial standard data variable rule table is compared to determine whether it is in the working period, non-working period or overtime period, and parsed data with time dimension attributes is generated. S23: Perform statistical cycle classification and sensitivity level labeling on the parsed data with time dimension attributes. According to the preset daily, weekly, and monthly cycle division rules, classify it into the corresponding statistical cycle, and query the industrial field standard data variable rule table to assign it an initial sensitivity level. Integrate equipment identification, variable type, time attribute, and sensitivity level information to generate industrial data to be desensitized carrying multi-dimensional data identification.

4. The method according to claim 1, characterized in that, S3 includes: S31: Perform multi-dimensional deconstruction and correlation analysis on all entries in the versioned industrial standard data variable rule table, extract the core constraint dimensions of each rule entry such as data variable type, time dimension type, applicable node level, applicable role set and data sensitivity level, and generate dimension deconstruction results; S32: Based on the dimensional deconstruction results and the dimensional space involved in the industrial data to be desensitized carrying multi-dimensional data identifiers, perform batch generation of six-dimensional binding records. For each rule entry, under all applicable node levels and role combinations, create a binding record that associates the specific dimension value with the unique identifier of the rule entry, and generate the original binding relationship set. S33: Perform index optimization and function encapsulation on the original binding relationship set. Create a composite database index based on key fields such as data variable type, time dimension type, node level, role identifier, and data sensitivity level. Define a query function that receives these five dimension parameters, quickly retrieves and returns the most matching rule entries through the index, and completes the six-dimensional binding relationship model.

5. The method according to claim 1, characterized in that, S4 includes: S41: The industrial data to be desensitized, carrying multi-dimensional data identifiers, undergoes rapid transformation processing based on real-time scenarios and single-point data. At the workshop node, the corresponding millisecond-level desensitization algorithm is matched according to its data variable type, real-time time attribute and initial sensitivity level to quickly process the original data to remove some sensitive details and generate a workshop-level desensitized data package with complete metadata. S42: Perform multi-variable collaborative desensitization processing on the workshop-level desensitized data packages from multiple workshops based on statistical cycles and production line logic. At the enterprise node, perform spatiotemporal alignment and correlation analysis on multi-source data from the same cycle and the same production line, and apply collaborative algorithms to eliminate the inference risk between variables, generating enterprise-level desensitized data packages with correlation protection capabilities. S43: Perform secondary collaborative processing on the enterprise-level de-identified data packages from multiple enterprises based on a unified time period standard and cross-entity data fusion. Unify the time period definition of each enterprise within the jurisdiction at the park node, and aggregate and perform secondary de-identification on cross-enterprise data of the same period to eliminate the feature differences between enterprises and generate a park-level de-identified data package that meets the park's regulatory requirements. S44: The campus-level de-identified data packets from multiple campuses are subjected to enhanced processing based on long-term statistics and global compliance requirements. Cross-campus long-term data aggregation is performed on the cluster nodes, and the final de-identification algorithm that meets the privacy protection strength requirements is matched and applied according to the latest compliance requirements to generate the final de-identified data. The workshop-level de-identified data packets, the enterprise-level de-identified data packets, and the campus-level de-identified data packets are integrated to generate a full-link traceability log.

6. The method according to any one of claims 1-5, characterized in that, S5 includes: S51: Authentication and parsing of the user's identity credentials and access request context in the user's data access request, verifying the legitimacy of the user's identity, and parsing out the data variable type requested for access, the expected time dimension, the user's node level, role, and expected data sensitivity level, and generating a legitimate access context; S52: Based on the legitimate access context and the current system time of the user, fine-grained permission matching is performed through the six-dimensional binding relationship model to find rule entries that simultaneously satisfy the constraints of data dimension, time dimension, node level, user role and sensitivity level, and generate a set of rule matching results containing precise access control and decryption rules; S53: Based on the access control and decryption rules in the rule matching result set and the full-link traceability log, perform targeted data decryption and precision restoration processing. According to the access granularity specified by the rule, find the corresponding level of data form from the log, use the key derived from the rule to decrypt, or, when authorized, reverse deduce higher precision data based on the desensitization algorithm of the log record to generate result data that perfectly matches the current user's permissions and access time.

7. A data anonymization system for industrial data, characterized in that, The system includes: The data rule definition module is used to define and organize the types, time dimensions, desensitization algorithms and compliance criteria of industrial data variables in a rule-based manner. It configures the corresponding desensitization algorithms, sensitivity levels and access rules for each type of industrial variable under different time scenarios and node levels, and generates a versioned standard data variable rule table for the industrial field. The data to be desensitized generation module is used to timestamp and analyze the features of the raw industrial data collected at the workshop node. Based on the standard data variable rule table of the industrial field, it adds the data variable type, the time period attribute and statistical period corresponding to the collection time, and the initial sensitivity level determined based on the rule table to each piece of raw industrial data, and generates industrial data to be desensitized carrying multi-dimensional data identifiers. The six-dimensional relationship modeling module is used to model the relationship between the standard data variable rule table in the industrial field and the dimensions involved in the industrial data to be desensitized. It associates and binds the data variable type, time dimension, node level, cluster role, data sensitivity level with the specific entries in the rule table in multiple dimensions to construct a six-dimensional binding relationship model. The collaborative desensitization processing module is used to perform collaborative desensitization calculations in the four-level nodes of the industrial cluster network (workshop, enterprise, park, and cluster) based on the industrial data to be desensitized and the six-dimensional binding relationship model. The calculations are performed sequentially from fine-grained to coarse-grained and from local to global. The desensitization processing results of each level are used as the input of the next level. After the desensitization processing is completed at the cluster node, the processing metadata of each level is integrated to generate a full-link traceability log. The data access decryption module is used to match access permissions and decryption rules based on the user's data access request and current access time through the six-dimensional binding relationship model, and to perform targeted decryption and precision restoration based on the full-link traceability log to generate result data that matches the current user's permissions and access time.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Data desensitization method, device and system

    CN106611129A

  • Data security hierarchical management and control method and device

    CN115062345A

  • Method for realizing big data application of thermal power plant production process

    CN115329365A

  • Method and device for carrying out data desensitization display according to user role permission

    CN117373597A

  • Desensitization method and device and server

    CN119628950A

Cited By

  • Sensitive data full-link desensitization method and system

    CN122263176A