Data privacy protection and sharing method in man-machine-object fusion environment

By comprehensively sorting out and classifying data, introducing artificial intelligence and big data analysis technologies for risk assessment and threat modeling, combining encryption algorithms that resistant to quantum computing and distributed storage, the problem of lack of systematicity and accuracy of data privacy protection in human-machine-material fusion environment is solved, and data is efficient, safe and reliable protection is achieved.

CN119945711AInactive Publication Date: 2025-05-06GUANGDONG BAIYUN UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202411848501.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

There are problems of lack of systematicity and accuracy in data privacy protection in human-machine-material fusion environment, including the failure to comprehensively sort out the data sources, the inability to achieve accurate classification, the risk assessment is not accurate enough, the data encryption and storage strategies are not systematic enough, and the user participation and monitoring system is incomplete.

Method used

By comprehensively sorting out and classifying data, introducing artificial intelligence and big data analysis technologies for risk assessment and threat modeling, formulating detailed privacy policies and specifications, using quantum computing-resistant encryption algorithms and distributed storage, establishing accurate identity authentication mechanisms, developing advanced data anonymization and desensitization technologies, formulating detailed data sharing protocols, enhancing user education and awareness, and establishing a comprehensive data privacy monitoring system.

Benefits of technology

It realizes accurate identification and accurate protection of data, improves the accuracy of risk assessment and the depth of threat modeling, ensures the security of data in the era of quantum computing, enhances user participation and privacy protection awareness, promptly discovers and handles security incidents, and ensures the security of the data sharing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945711A_ABST
    Figure CN119945711A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data security, and particularly discloses a data privacy protection and sharing method in a man-machine-object fusion environment, which comprises the following steps: S1, determining a data range and sensitivity, comprehensively sorting and classifying data, pre-judging an affected encrypted data type according to a new technology trend, and planning a strategy in advance; according to the method, data security is ensured by sorting and classifying data, removing and labeling error data, applying strategies such as anti-quantum encryption and the like, then risk assessment and modeling are carried out by virtue of artificial intelligence and big data, identity authentication is carried out by virtue of biological characteristics and behavior analysis, anonymization, desensitization and the like are processed by virtue of an algorithm, and meanwhile, a monitoring range index is determined; the method depends on advanced tool monitoring and rapid response according to abnormity, continuously improves strategies, pays attention to user education, establishes a feedback reward mechanism, formulates a data sharing protocol and actively participates in industrial standard and international cooperation, so as to promote comprehensive development and application of a data privacy protection technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data security technology, and in particular to a method for protecting and sharing data privacy in a human-machine-object fusion environment. Background Art

[0002] Data privacy protection in the human-machine-object fusion environment is of vital importance. In the complex environment of ubiquitous computing formed by the cross-fusion of the ternary space composed of people, machines and objects, each link of data collection, storage, transmission, processing and sharing faces many potential risks. In order to effectively protect the sensitive data of individuals, organizations and other subjects from unauthorized access, leakage, abuse or tampering, safeguard the legitimate rights and interests and privacy security of data subjects, and ensure that the human-machine-object fusion system can operate reliably and achieve sustainable development, a series of comprehensive measures must be taken, including the use of advanced technical means such as encryption, access control, and data desensitization, the formulation of rigorous privacy policies, and the strengthening of employee training to improve their privacy protection awareness. At the same time, strict compliance with relevant laws and regulations and other management and compliance measures, however, there are many problems with data privacy protection in the existing human-machine-object fusion environment;

[0003] First, data processing and security strategies lack systematicity and precision. In the complex environment of human-machine-object integration, existing technologies fail to comprehensively sort out data sources and cannot achieve accurate classification, which makes it difficult to accurately identify data of different sensitivity levels. In addition, there is a lack of a mechanism for regular reclassification, which makes it difficult to adapt to the characteristics of data that change over time. In terms of risk assessment and threat modeling, advanced technologies have not been fully utilized to deeply explore risks, and there is a lack of simulated attack experiments to test protection measures, resulting in inaccurate risk warnings. In the fields of data encryption and storage, anonymization and desensitization, and privacy enhancement technology, there is a lack of systematic strategies and clear quantitative standards, which makes it difficult to effectively respond to the challenges brought by emerging technologies such as quantum computing, which easily leads to data leakage risks. At the same time, data sharing agreements are not detailed and comprehensive enough, there is a lack of strict review of the recipient, and there is no real-time monitoring mechanism, which cannot effectively guarantee the security of the sharing process.

[0004] Second, the user participation and monitoring system is not perfect. In terms of user education and awareness raising, the existing technology is obviously insufficient, the educational activities are relatively simple in form, and the updates are not timely. There is a lack of perfect feedback channels and effective reward mechanisms, which makes users' participation in data privacy protection low. In terms of data privacy monitoring, the monitoring scope is not comprehensive enough, there is a lack of precise indicators and advanced technical tools, the monitoring management platform is not integrated enough, and security incidents cannot be discovered and handled in time. It is difficult to dynamically optimize the monitoring system to adapt to complex and changing security conditions. In summary, data privacy protection in the existing human-machine-object fusion environment faces severe challenges and needs to be improved and perfected from multiple aspects to ensure that data privacy is effectively protected and promote the healthy development of the human-machine-object fusion system. Summary of the invention

[0005] The purpose of the present invention is to provide a method for data privacy protection and sharing in a human-machine-object fusion environment, so as to solve the problems raised in the above background technology that data processing and security strategies lack systematicity and accuracy during the application of existing technologies, as well as imperfect user participation and monitoring systems.

[0006] To achieve the above object, the present invention provides a method for data privacy protection and sharing in a human-machine-object fusion environment, comprising the following steps:

[0007] S1: Clarify the scope and sensitivity of data, comprehensively sort out and classify data, predict the types of encrypted data that will be affected by new technology trends and plan strategies in advance, encrypt particularly sensitive data such as those involving national security and major commercial secrets, establish a dynamic data list for automatic scanning and updating, and update key data-related information in real time to cope with the rapid changes in data in the human-machine-object fusion environment;

[0008] S2: Risk assessment and threat modeling. We introduce artificial intelligence and big data analysis technologies to carry out risk assessment and threat modeling, implement more comprehensive and in-depth risk assessments on data in the human-machine-object fusion environment, specifically analyze network traffic data and device logs in the past week to determine potential risk points, conduct simulated attack experiments of attack scenarios in the threat simulation laboratory every month, and test and improve protection measures;

[0009] S3: Formulate privacy policies and regulations, clarify the legal responsibilities and penalties of enterprises and organizations in data privacy protection, protect user rights by signing legally binding privacy agreements with users, stipulate the scope of fines for policy violations, arrange employee training every year, regularly update policies and regulations to adapt to changes in regulations and technology, strengthen employee training, and enhance their compliance awareness;

[0010] S4: Data encryption and secure storage, research and apply encryption algorithms that are resistant to quantum computing, prepare for data privacy protection in the quantum era, use multi-layer encryption technology, differentiate by sensitivity, adopt distributed storage, set up multiple storage nodes and each node has 3 copies of data backup, set up access control policies, assign different permissions according to user types, establish a distributed, redundant secure storage system, store data in fragments and encrypt each fragment independently;

[0011] S5: Access control and identity authentication. Use biometrics, behavioral analysis and other technologies to establish a precise identity authentication mechanism. The legitimacy of the identity is determined based on the user's typing speed, mouse movement trajectory and other behavioral characteristics. In terms of encryption algorithm selection, data transmission uses a 128-bit symmetric encryption algorithm, and highly sensitive data transmission uses a 256-bit asymmetric encryption algorithm combined with a digital signature. A role-based and attribute-based access control model is used to grant corresponding data access rights according to the user's position level and work requirements.

[0012] S6: Data anonymization and desensitization processing, develop more advanced data anonymization and desensitization technologies, use deep learning-based anonymization technology to improve anonymization and desensitization effects, control the probability of re-identification of anonymized data to less than 1%, use k-anonymity technology for medical data, and l-diversity technology for financial data, establish standard specifications for data anonymization and desensitization, and clarify that at least 90% of personal identity information must be removed when anonymizing medical data, so as to guide relevant entities to correctly handle data and prevent privacy leakage risks;

[0013] S7: Secure data sharing agreement. Develop a detailed data sharing agreement template to clarify all aspects of data sharing. Before signing the contract, require the data recipient to meet a certain ISO27001 security level and conduct a strict assessment and review. Use blockchain technology to build a decentralized sharing platform, monitor data sharing behavior on the platform every second, ensure that the sharing process is transparent, traceable and cannot be tampered with, and monitor and audit sharing behavior in real time to deal with violations in a timely manner.

[0014] S8: Privacy enhancement technology application, increase investment in privacy enhancement technology research and development, cooperate with universities and research institutions to carry out relevant technology research and application pilots, accumulate experience data, establish an evaluation system, determine the range of differential privacy noise addition to balance data privacy and availability, compare different technologies to select technologies that are suitable for the human-machine-object fusion environment, and strengthen user education to improve user awareness and acceptance of privacy enhancement technology;

[0015] S9: User education and awareness raising: Through various forms of user education activities, short videos are released every month to enhance users' knowledge of data privacy and protection awareness. Case analysis and simulated attacks are used to let users intuitively know the dangers of privacy leakage and protection methods. User feedback channels are also established to reward users with points for timely reporting of privacy issues, encourage feedback and handle responses in a timely manner, and reward users' privacy protection behaviors to increase their enthusiasm;

[0016] S10: Continuous monitoring and improvement, establishing a comprehensive data privacy monitoring system, establishing a monitoring system covering networks, devices, applications, etc., monitoring and analyzing network traffic every millisecond, controlling data dynamics and security status in real time, using artificial intelligence and machine learning technology to analyze monitoring data, issue early warnings, and promptly detect potential privacy leakage risks. Comprehensively evaluate data privacy protection measures every six months, formulate improvement plans based on the results, and conduct regular evaluation and improvement work. In addition, actively participate in the formulation of industry standards and international cooperation to promote the development and application of data privacy protection technology;

[0017] S11: Application of data watermarking technology: Use data watermarking technology to embed watermarks containing information such as the owner and authorization scope into the data to be shared. In relevant scenarios, the watermark can be used to trace the source and authorization status. The watermark strength is set moderately to ensure that normal use does not affect the quality. Illegal copying is easy to detect. There are requirements for detection accuracy and success rate. Robustness is also optimized to balance strength and quality in high-precision scenarios.

[0018] S12: Comprehensive technology application and collaboration, combining edge computing with privacy protection, and performing preliminary privacy processing on edge devices. For example, camera data in smart homes is encrypted at the edge gateway and then transmitted to the cloud. Homomorphic encryption hardware acceleration modules are developed to increase speed and control costs. They are applied to scenarios with high real-time requirements, optimize multi-party secure computing protocols, shorten computing time, and reduce vulnerability rates. They are used in scenarios such as data sharing in medical institutions and strictly evaluate and prevent risks.

[0019] Furthermore, the comprehensive sorting and classification of data in S1 includes the following steps:

[0020] 1) Data collection and organization: Determine the data sources in the human-machine-object fusion environment, including intelligent devices, computer systems, etc., investigate and record the information of each data source in detail, use data cleaning tools to remove duplicate and erroneous data, control the removal ratio of duplicate data within 5%, and the accuracy of outlier detection is more than 90%, ensuring that the data is accurate and complete;

[0021] 2) Data classification standard formulation: classify data from multiple perspectives such as sensitivity, owner, and life cycle stage, clarify the definition and characteristics of each classification. For example, highly sensitive data, once leaked, will cause significant losses, laying the foundation for subsequent accurate classification and ensuring that different types of data receive appropriate privacy protection strategies;

[0022] 3) Data classification and labeling: Classify data according to standards, use automated tools combined with manual review to ensure accuracy, label data to facilitate subsequent management and protection, and conduct regular review and reclassification to adapt to changes. This should be done once a quarter, and each time should be within one week to ensure that privacy protection measures remain effective.

[0023] Furthermore, the introduction of artificial intelligence and big data analysis technology in S2 includes the following steps:

[0024] 1) Data collection and preprocessing: Use artificial intelligence collection tools to collect data from multiple data sources in the human-machine-object fusion environment, such as smart factory equipment status data, preprocess the data, use data cleaning algorithms to remove noise and outliers, control the proportion of outliers, determine the normal range of network traffic and remove abnormal data;

[0025] 2) Feature extraction and model training: Use machine learning algorithms to extract features, select appropriate algorithms to train models, use random forest algorithms for threat detection, set accuracy and recall targets, evaluate performance through cross-validation, and continuously adjust parameters to improve model accuracy and generalization capabilities to ensure stability;

[0026] 3) Risk assessment and threat warning: Apply the model to actual data to assess risks, classify them into high, medium and low levels, establish a real-time warning system, respond quickly when risks are high, analyze the assessment results monthly, and update the model and warning rules based on new threats to respond to potential risks in a timely manner and ensure data security.

[0027] Furthermore, the research and application of quantum computing-resistant encryption algorithms in S4 include the following steps:

[0028] 1) Demand analysis: Determine the type of data and security level to be protected, such as setting financial transaction data to the highest level 9. The evaluation system determines the links that need to be quantum-resistant, such as the IoT and cloud transmission channels, requiring the encryption algorithm to have a low probability of being cracked by quantum attacks;

[0029] 2) Algorithm research: In-depth research on various quantum-resistant algorithms, comparing their principles, advantages, limitations, and performance, and selecting appropriate algorithms according to different scenarios. For example, speed and memory should be considered for resource-constrained devices.

[0030] 3) Algorithm selection: Select the most suitable algorithm based on the requirements and research results, taking into account scalability and compatibility. For example, if real-time performance is high, a hash algorithm can be selected, and it must be integrated with the existing system;

[0031] 4) Testing and optimization: Testing the algorithm, including performance and security, and optimizing and adjusting based on the results to improve speed and security, such as increasing encryption speed by more than 10%;

[0032] 5) Deployment and maintenance: Deploy the algorithm to each link to ensure the success rate, establish a maintenance mechanism, regularly evaluate and upgrade, train personnel, fully cover security assessments, and provide no less than 10 hours of training each year.

[0033] Furthermore, the use of biometric recognition, behavior analysis and other technologies in S5 includes the following steps:

[0034] 1) Biometric identification settings: Collect biometric information such as fingerprints, faces, and irises, set accuracy, resolution, and collection time standards for each identification method, and store the collected information in an encrypted database with an encryption strength of AES256 bits or more to ensure information security and prevent illegal access;

[0035] 2) Behavior analysis modeling: collect behavioral data such as typing speed and mouse trajectory, use machine learning algorithms to build models, and use training data sets of no less than 1,000 samples. Control the training time so that the model's behavior recognition accuracy reaches more than 95%;

[0036] 3) Comprehensive identity authentication: Authentication is based on biometrics and behavioral analysis results. If both match, the authentication probability is 100%. If one matches, it drops to 70% and needs to be re-verified. The authentication response is within 2 seconds. If multiple failures occur, the account is automatically locked and the login behavior is recorded for subsequent review.

[0037] Furthermore, the development of more advanced data anonymization and desensitization technology in step S6 specifically includes the following steps:

[0038] 1) Data feature analysis: Comprehensively analyze the data that needs to be anonymized and desensitized, classify the types, clarify the structure and sensitivity level, and count the proportion of each type of data. For example, among millions of data sets, personal identity information accounts for 20%, and highly sensitive information accounts for 30%, laying the foundation for targeted processing;

[0039] Application of anonymization technology: Use algorithms such as k-anonymity and l-diversity, set corresponding parameters according to data types, evaluate the anonymized data, control the probability of re-identification and the success rate of attack tests, and ensure the anonymization effect and security;

[0040] 2) Desensitization technology selection: select the appropriate desensitization method according to the data type and purpose, control the precision and amplitude of the numerical type, and ensure the accuracy of keyword replacement for the text type, etc., evaluate the quality of the desensitized data, ensure the similarity with the original data and meet the analysis application requirements;

[0041] 3) Continuous monitoring and optimization: Establish a monitoring mechanism to conduct a comprehensive monthly check of anonymized and desensitized data, covering multiple aspects. Adjust strategies in a timely manner based on the results, control optimization time, and re-evaluate to ensure data security and meet expected results.

[0042] Furthermore, the detailed data sharing agreement template in S7 includes the following steps:

[0043] 1) Clarify the purpose and scope of sharing: clearly explain the purpose of data sharing, limit the scope of use, and specify the type of data involved, the upper limit of the amount, the time and geographical scope, etc. For example, medical data sharing stipulates the content and amount of data, and cross-border sharing clarifies the authority and legal differences to ensure accurate and compliant sharing;

[0044] 2) Formulate security requirements: stipulate that AES256-bit encryption be used for transmission, control transmission error rate, clarify the security responsibilities of the recipient, set access control, assign permissions based on roles, and strictly verify and approve processes to ensure data transmission and access security;

[0045] 3) Determine the liability for breach of contract: clearly define the liability for breach of contract for violation of the agreement, determine compensation for data leakage based on the sensitivity of the data, establish a dispute resolution mechanism, negotiate first, and then arbitrate or sue if no results are achieved. Suspend sharing during the dispute period to ensure that problems are handled in a timely manner and responsibilities are clearly defined.

[0046] 8. The method for data privacy protection and sharing in a human-machine-object fusion environment according to claim 1, wherein the step of establishing a comprehensive data privacy monitoring system in S10 comprises the following steps:

[0047] 1) Determine the monitoring scope and indicators: clearly define the monitoring scope of each part of the human-machine-object fusion environment, set specific indicators such as data access frequency, such as monitoring of equipment and data centers, set access frequency accuracy, abnormal traffic exceeding threshold alarms, and conduct in-depth investigations when the risk index is high;

[0048] 2) Select monitoring technology and tools: Select advanced technologies such as artificial intelligence algorithms with an accuracy rate of more than 90%, deploy professional tools, build a centralized management platform, and ensure a response time of no more than 5 seconds. Integrate data for unified management and analysis to ensure timely discovery and handling of security incidents;

[0049] 3) Implement monitoring and response mechanism: 24-hour uninterrupted monitoring, dedicated personnel on duty, abnormal level response, regular analysis and summary reports, adjustment of strategies and measures based on reports, optimization of monitoring system, and special investigation and prevention in case of an increase in similar incidents.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] First, in the present invention, comprehensive data management and protection, the technical solution has a comprehensive data management and protection mechanism, which uses data combing and classification to clarify the various data sources in the human-machine-object fusion environment, and formulates detailed classification standards, which can accurately identify data of different sensitivity levels, and provide a basis for subsequent privacy protection measures. Through data collection and organization, duplicate and erroneous data are removed to ensure the accuracy and integrity of the data. At the same time, different types of data are classified and labeled to facilitate management and protection. In terms of encryption and storage, quantum computing-resistant encryption algorithms and distributed storage are used to ensure data security and reliability. For sensitive data, higher-level encryption is used, and the data is divided into multiple fragments for storage. Even if some storage nodes are attacked, the complete data cannot be obtained. In addition, a strict access control strategy is set up to assign different permissions according to user roles, further ensuring data security;

[0052] Secondly, in the present invention, this technical solution uses artificial intelligence and big data analysis technology to perform risk assessment and threat modeling. By collecting a large amount of data for preprocessing and feature extraction, a high-accuracy model is trained, which can timely discover potential risk points and provide real-time warnings. At the same time, biometric recognition and behavioral analysis technology are used for identity authentication, which improves the accuracy and security of authentication. In terms of data anonymization and desensitization, advanced algorithms and strict parameter control are used to reduce the risk of re-identification while ensuring the availability of data. In addition, privacy enhancement technologies such as differential privacy and homomorphic encryption are also applied, and an evaluation system is established to select the most suitable technology. Through data watermarking technology, the source and authorization of data can be traced, which enhances the manageability of data. In terms of comprehensive technical application and collaboration, edge computing is combined for preliminary privacy processing to improve the real-time and security of data.

[0053] Third, in the present invention, this technical solution, by determining the specific monitoring scope and indicators and selecting advanced monitoring technologies and tools, can detect abnormal situations in a timely manner and respond quickly according to the severity. At the same time, a 24-hour uninterrupted monitoring system is implemented to ensure that the dynamics and security status of the data are always under control, and the monitoring results are regularly analyzed and summarized. The monitoring strategies and safety measures are continuously adjusted according to actual conditions to achieve continuous improvement. In addition, the solution also focuses on user education and awareness enhancement. By carrying out various forms of educational activities and establishing feedback channels and reward mechanisms, users' understanding and participation in data privacy are improved. In terms of data sharing, a detailed data sharing agreement template is formulated, the responsibilities and security requirements of all parties are clarified, and the data sharing process is ensured to be safe and controllable. By actively participating in the formulation of industry standards and international cooperation, the development and application of data privacy protection technology are promoted. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 Method flow chart. DETAILED DESCRIPTION

[0055] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0056] See also Figure 1 In an embodiment of the present invention, a method for protecting and sharing data privacy in a human-machine-object fusion environment includes the following steps:

[0057] S1: Clarify the scope and sensitivity of data, comprehensively sort out and classify data, combine the development trend of new technologies such as quantum computing, predict the type of encrypted data that may be affected, and plan response strategies in advance; for particularly sensitive data involving national security and major commercial secrets, the 256-bit Advanced Encryption Standard "AES-256" can be used for encryption; establish a dynamic data list, automatically scan and update the data list once an hour to ensure the accuracy of information, and update the source, purpose, storage location and access rights of the data in real time to adapt to the rapid changes of data in the human-machine-object fusion environment;

[0058] Comprehensive combing and classification of data includes the following steps;

[0059] 1): Data collection and organization. First, identify all possible data sources in the human-machine-object fusion environment, including various smart devices such as smartphones, IoT sensors, smart home devices, computer systems, databases, and manually input data. In a typical smart factory environment, there may be thousands of IoT sensors, and the amount of data generated per second is between 1KB and 10KB;

[0060] Conduct a detailed investigation and record of each data source, including the data type (such as text, images, audio, video, etc.), the frequency of generation, the storage location, and the data format. For IoT sensors, record the type of data collected (such as temperature, humidity, pressure, etc.), the collection frequency (ranging from once every minute to once every 10 minutes), and the server or cloud platform where the data is stored. The data format may include JSON, XML, CSV, etc. The proportion of data in each format may vary in different scenarios;

[0061] Preliminary sorting of the collected data to remove duplicate data and obviously erroneous data can be done using data cleaning tools and algorithms, such as data deduplication algorithms and outlier detection algorithms, to ensure the accuracy and integrity of the data. Through the data deduplication algorithm, the removal ratio of duplicate data can be controlled within 5%, and the detection accuracy of outliers can reach more than 90%;

[0062] 2): Data classification standard formulation: According to the nature and purpose of the data, detailed data classification standards can be formulated, and classification can be carried out from multiple angles;

[0063] Classify by data sensitivity: Data is divided into highly sensitive data (such as personal identity information, financial data, medical data, etc.), moderately sensitive data (such as equipment operation status data, internal business data of the enterprise, etc.) and generally sensitive data (such as public environmental data, general statistical data, etc.). Highly sensitive data usually accounts for between 10% and 20%, moderately sensitive data accounts for about 30% to 40%, and generally sensitive data accounts for 40% to 60%;

[0064] Classify by data owner: clarify whether the data belongs to individual users, corporate organizations, or public institutions. Data of different owners may require different privacy protection strategies. Personal user data may account for 30% to 40%, corporate organization data accounts for 40% to 50%, and public institution data accounts for 10% to 20%;

[0065] Classification by data life cycle stage: Data is divided into data generation stage, storage stage, transmission stage, processing stage and destruction stage, etc. Data in each stage has its own specific privacy risks and protection needs. In the data generation stage, the privacy risk is relatively low, and the protection needs are mainly to ensure the accuracy and integrity of the data. The proportion of data in the data generation stage to the total data volume is usually 100%. "As the data flows, the amount of data in each stage will gradually decrease."

[0066] Develop clear definitions and feature descriptions for each category so that data can be accurately identified and classified in the subsequent classification process. For highly sensitive data, it can be defined as data that may cause significant losses to individuals or organizations once leaked, and specific features are listed, including personal ID numbers, bank account information, etc. The recognition accuracy of highly sensitive data must reach more than 95%;

[0067] 3): Data classification and labeling. According to the established classification standards, each piece of data is classified. Automated data classification tools and algorithms can be used in combination with manual review to ensure the accuracy of classification. For text data, natural language processing technology can be used for keyword extraction and semantic analysis to determine which category it belongs to. The accuracy of automated classification should reach more than 80%, and the overall accuracy after manual review should reach more than 98%;

[0068] Label the classified data so that it can be quickly identified and processed in the subsequent data management and privacy protection process. The labeling can include the data classification code, sensitivity level, owner information, etc. For a highly sensitive personal medical data, it can be labeled as "H-personal medical data-owner name". The time cost of labeling is controlled between 1 second and 5 seconds for each data;

[0069] Data should be reclassified and labeled regularly to adapt to data changes and new privacy protection needs. Over time, the nature and use of data may change, so data needs to be reviewed and reclassified regularly to ensure the effectiveness of privacy protection measures. Data should be comprehensively reviewed and reclassified every quarter, and each review and reclassification should be completed within one week.

[0070] S2: Risk assessment and threat modeling, introducing artificial intelligence and big data analysis technology to conduct a more comprehensive and in-depth risk assessment of data in the human-machine-object fusion environment; analyzing 10,000 network traffic data and device logs in the past week to identify potential risk points; establishing a threat simulation laboratory to conduct a simulated attack experiment once a month, covering 5 different types of attack scenarios, simulating data privacy threats in various extreme situations, such as large-scale network attacks, physical equipment destruction, etc., to test and improve protection measures;

[0071] The introduction of artificial intelligence and big data analysis technologies includes the following steps:

[0072] Data collection and preprocessing: using artificial intelligence data collection tools to collect data from various data sources in the human-machine-object fusion environment, including equipment logs, network traffic data, user behavior data, etc. Specifically, the equipment operation status data is collected from the equipment in the smart factory once an hour, and the data volume is between 10MB and 50MB;

[0073] Pre-process the collected data to remove noise and outliers. Use the data cleaning algorithm in big data analysis technology to control the removal ratio of outliers to less than 3%. For network traffic data, determine the normal range of traffic by analyzing historical data. Data outside this range is considered an outlier and removed.

[0074] 2): Feature extraction and model training: Use artificial intelligence machine learning algorithms to extract features from preprocessed data. For user behavior data, features such as user login time, operation frequency, and visited pages can be extracted;

[0075] Use the big data analysis platform to train the model and select the appropriate machine learning algorithm, such as decision tree, random forest, neural network, etc. For threat detection tasks, use the random forest algorithm for training. The size of the training data set is 100GB, and the ratio of positive samples (threat data) to negative samples (normal data) is 1:5;

[0076] Continuously adjust model parameters to improve model accuracy and generalization ability, set the model accuracy target at more than 90% and the recall rate at more than 85%, evaluate model performance through cross-validation and other methods, and ensure the stability of the model on different data sets;

[0077] 3): Risk assessment and threat warning: Apply the trained model to the actual data to conduct risk assessment. According to the output of the model, determine the potential threats and risk levels faced by the data. The risk levels are divided into three levels: high, medium, and low. When the threat probability predicted by the model is greater than 70%, it is judged as high risk.

[0078] Establish a real-time threat warning system to immediately issue an alarm when a high-risk situation is detected. The alarm response time is controlled within 1 minute to ensure that countermeasures can be taken in a timely manner.

[0079] Regularly analyze and summarize risk assessment results, adjust models and early warning strategies based on actual conditions, conduct a comprehensive analysis of risk assessment results every month, and update models and early warning rules based on emerging threat types and trends;

[0080] S3: Formulate privacy policies and specifications. Formulate stricter and more detailed privacy policies and specifications to clarify the legal responsibilities and penalties of enterprises and organizations in data privacy protection; sign legally binding privacy agreements with users to ensure that the rights and interests of users are fully protected; clearly stipulate that the fine for violating the privacy policy is between 100,000 yuan and 1 million yuan; review the privacy policy once a quarter, provide employees with at least 8 hours of privacy policy training each year, and regularly review and update the privacy policy and specifications to adapt to the ever-changing legal and technological environment; at the same time, strengthen employee training to improve employees' awareness of compliance with the privacy policy;

[0081] S4: Data encryption and secure storage, research and apply encryption algorithms that resist quantum computing, adopt encryption algorithms based on lattice cryptography, etc., to prepare for data privacy protection in the era of quantum computing in advance; at the same time, adopt multi-layer encryption technology, use 128-bit encryption for general sensitive data, and 256-bit or higher encryption for highly sensitive data; for the storage system, adopt distributed storage, with no less than 5 data storage nodes, and the number of data backups for each node is 3; set up access control policies, for general users, access rights are limited to reading some non-sensitive data, and their permission allocation ratio is 30%; for advanced users, they can read and modify some sensitive data, and the permission allocation ratio is 60%; for administrators, they have the highest authority, and the permission allocation ratio is 100%; establish a distributed, redundant secure storage system, divide the data into at least 10 fragments and store them in different physical locations, and each fragment uses an independent encryption key to ensure that the data will not be lost or stolen due to a single point of failure during the storage process; adopt data sharding storage and encryption technology, even if some storage nodes are attacked, the complete data cannot be obtained;

[0082] Researching and applying quantum computing-resistant encryption algorithms involves the following steps:

[0083] 1) Demand analysis, determine the type of data and security level requirements that need to be protected in the human-machine-object fusion environment. For sensitive data involving financial transactions, the security level requirements are extremely high, and encryption algorithms that can resist quantum computing attacks must be used. The security level for financial transaction data is set to the highest level 9, requiring that the probability of successful cracking of the encryption algorithm should be less than 0.0001% when facing possible future quantum computing attacks;

[0084] Evaluate existing data storage and transmission systems to determine which links require quantum computing-resistant encryption protection. The data transmission channel between IoT devices and cloud servers may need to be encrypted to prevent data leakage under quantum computing attacks. For example, after evaluation, it was found that the data transmission risk factor between IoT devices and cloud servers was 7 out of 10, which indicates that this link needs to adopt quantum computing-resistant encryption protection with a security level of at least 7.

[0085] 2) Algorithm research: conduct in-depth research on various quantum computing-resistant encryption algorithms, such as lattice-based algorithms, hash-based algorithms, and coding-based algorithms, and understand the principles, advantages, and limitations of each algorithm. For example, lattice-based algorithms have certain advantages in security and efficiency, but the computational complexity is relatively high. The security strength index of lattice-based algorithms can reach 8 out of 10, the encryption speed is between 5MB and 10MB per second, and the key length is 256 bits to 512 bits. The security strength index of hash-based algorithms is 7, and the encryption speed is between 8MB and 10MB per second. 5MB, key length is 128 to 256 bits, the security strength index of the coding-based algorithm is 6.5, the encryption speed is between 3MB and 8MB per second, and the key length is 192 to 384 bits. Compare the performance of different algorithms in different scenarios, including encryption speed, key length, memory usage, etc. For example, for resource-constrained IoT devices, it is necessary to select quantum computing-resistant encryption algorithms with faster encryption speed and smaller memory usage. For low-power IoT devices, the encryption algorithm is required to occupy no more than 5MB of memory and the encryption speed is not less than 3MB per second.

[0086] 3) Algorithm selection: Based on the results of demand analysis and algorithm research, select the most suitable quantum computing-resistant encryption algorithm for the human-machine-object fusion environment. For example, if the real-time requirements for data transmission are high, a hash-based quantum computing-resistant encryption algorithm can be selected, and its encryption speed is relatively fast. When the real-time requirements for data transmission are within 500 milliseconds, a hash-based encryption algorithm should be selected, and its encryption speed should reach more than 10MB per second.

[0087] Consider the scalability and compatibility of the algorithm, and ensure that the selected algorithm can be integrated with existing data processing and storage systems. For example, the selected encryption algorithm should be compatible with existing database management systems and cloud computing platforms. The compatibility requirement is more than 90%, that is, it can be smoothly integrated with more than 90% of the mainstream database management systems and cloud computing platforms on the market;

[0088] 4) Testing and optimization: testing the selected quantum computing-resistant encryption algorithm in a real environment, including encryption and decryption performance testing, security testing, etc. For example, using attack tools that simulate quantum computers to attack encrypted data to verify the security of the algorithm. In security testing, the algorithm is required to have a success rate of not less than 99% in resisting attacks by simulated quantum computers;

[0089] According to the test results, the algorithm is optimized and adjusted. For example, encryption parameters are adjusted and algorithm implementation is optimized to improve encryption speed and security. The encryption speed target is set to process more than 10MB of data per second, while ensuring that security meets the requirements for resisting quantum computing attacks. The encryption speed after optimization should increase by more than 10%, and the security strength index should increase by at least 0.5.

[0090] 5) Deployment and maintenance: Deploy the optimized anti-quantum computing encryption algorithm to all links in the human-machine-object fusion environment, including data storage devices, communication channels, applications, etc. For example, embed anti-quantum computing encryption modules in IoT devices to ensure that the data collected by the devices is securely protected during transmission and storage. The deployment success rate should reach more than 95%, ensuring that the encryption algorithm can be effectively applied in all key links in the human-machine-object fusion environment;

[0091] Establish a maintenance mechanism for encryption algorithms, and regularly update and upgrade the algorithms to cope with changing security threats and technological developments. For example, conduct a security assessment of encryption algorithms every six months, and make necessary updates and upgrades based on the assessment results. At the same time, provide training to personnel using encryption algorithms to ensure that they can correctly use and maintain encryption systems. The coverage of security assessments should reach 100%, and the training time for personnel using encryption algorithms should be no less than 10 hours per year.

[0092] S5: Access control and identity authentication, using biometrics, behavioral analysis and other technologies to establish a more accurate identity authentication mechanism; by analyzing the user's typing speed between 40 and 80 words per minute, the specific pattern of the mouse movement trajectory and other behavioral characteristics, determine whether the user's identity is legitimate, and the accuracy of identity authentication should reach more than 95%; in the selection of encryption algorithms, for general data transmission, use a 128-bit symmetric encryption algorithm, and the encryption speed should be more than 10MB per second; for highly sensitive data transmission, use a 256-bit asymmetric encryption algorithm combined with a digital signature, and the encryption time is controlled within 500 milliseconds. Use a role-based and attribute-based access control model to grant different data access rights according to the user's position level and work needs, from only being able to view part of the data to being able to modify specific data;

[0093] Using biometrics, behavioral analysis and other technologies includes the following steps:

[0094] 1): Biometric identification settings, collect user's biometric information, such as fingerprints, facial features, irises, etc. For fingerprint identification, the collection resolution should be above 500dpi, the recognition accuracy should be above 99%, the accuracy of facial feature recognition should be above 98%, and the fluctuation range of recognition accuracy should be controlled within 2% under different lighting conditions. The accuracy of iris recognition should reach above 99.5%, and the collection time should be controlled within 1 second;

[0095] The collected biometric information is stored in a secure database and protected by encryption technology. The encryption strength of the database should reach Advanced Encryption Standard "AES" 256 bits or above to ensure that the biometric information cannot be illegally obtained;

[0096] 2): Behavior analysis modeling, collecting user behavior data, including typing speed, mouse movement trajectory, operation habits, etc. For example, the monitoring range of typing speed is set to 30 to 120 words per minute, and the average mouse movement speed is between 5 cm and 20 cm per second;

[0097] Use machine learning algorithms to model and analyze user behavior data and establish a user behavior feature model. The size of the model's training data set is no less than 1,000 samples, and the training time is controlled within 24 hours. After training, the model's recognition accuracy of user behavior must reach more than 95%;

[0098] 3): Comprehensive identity authentication, combining biometrics and behavioral analysis results for comprehensive identity authentication. When the biometrics and behavioral analysis results match at the same time, the probability of authentication passing is 100%. If only one of them matches, the probability of authentication passing is reduced to 70%, and further verification steps are required;

[0099] During the identity authentication process, the response time is controlled within 2 seconds to ensure that users can quickly log in and access the system. At the same time, in the case of multiple consecutive authentication failures, the system should automatically lock the account and automatically unlock it after 30 minutes, or unlock it through manual review. During the lock period, any attempt to log in will be recorded for subsequent security review.

[0100] S6: Data anonymization and desensitization processing, develop more advanced data anonymization and desensitization technologies, adopt anonymization technology based on deep learning, improve the effect of anonymization and desensitization, and reduce the risk of re-identification; the probability of re-identification of data after anonymization should be less than 1%; in terms of anonymization technology parameters, for medical data, use k-anonymity technology, and set the k value to be greater than 5; for financial data, use l-diversity technology, and set the l value to be greater than 3; establish standards and specifications for data anonymization and desensitization, and clearly stipulate that in the anonymization of medical data, at least 90% of personal identity information must be removed, guide enterprises and organizations to correctly process data, and avoid privacy leaks due to improper processing;

[0101] The development of more advanced data anonymization and desensitization technologies includes the following steps:

[0102] 1): Data feature analysis: conduct a comprehensive analysis of the data that needs to be anonymized and desensitized to determine the type, structure and sensitivity of the data. The data is divided into different types such as personal identity information, financial data, and medical data. The sensitivity of each type of data is divided into three levels: high, medium, and low. For highly sensitive data, such as ID card numbers and mobile phone numbers in personal identity information, strict anonymization and desensitization must be performed. The proportion of various types of data is counted in order to formulate targeted processing strategies. In a data set containing 1 million data items, personal identity information data accounts for 20%, of which highly sensitive data accounts for 30% of personal identity information data;

[0103] 2): Application of anonymization technology, using advanced anonymization algorithms such as k-anonymity and l-diversity. For medical data using k-anonymity technology, the k value is set to 8 or more to ensure that there are at least 8 different records in each equivalence class, making it difficult to identify individual data. For financial data, l-diversity technology can be used, and the l value is set to 5 or more to ensure that the sensitive attributes in each equivalence class have at least 5 different values. The anonymized data is evaluated to ensure the effectiveness of anonymization. The probability of re-identification of the anonymized data should be less than 0.5%. By simulating attacks, the anonymized data is attacked using known background knowledge and attack algorithms to verify the security of anonymization. The success rate of the attack test should be less than 1%;

[0104] 3): Desensitization technology selection, according to the type and purpose of the data, select the appropriate desensitization technology. For numerical data, data truncation, randomization and other methods can be used for desensitization. The accuracy of data truncation is controlled to two decimal places, and the amplitude of randomization is controlled within 10% of the original data. For text data, keyword replacement, fuzzification and other methods can be used for desensitization. The accuracy of keyword replacement should reach more than 95%, and the readability of the fuzzified text should be controlled below 30%, that is, the original meaning of the text after fuzzification is difficult to directly understand, but it can still retain certain statistical characteristics. The quality of the desensitized data is evaluated to ensure the availability of the data. The similarity between the desensitized data and the original data should be controlled at more than 70%, and at the same time, the desensitized data should be able to meet the needs of subsequent data analysis and application. In a data analysis task, the deviation between the results of the analysis using the desensitized data and the results of the analysis using the original data should be controlled within 5%;

[0105] 4): Continuous monitoring and optimization, establish a monitoring mechanism for data anonymization and desensitization, regularly check and evaluate the processed data, conduct a comprehensive inspection of the anonymized and desensitized data every month, including the degree of anonymization, desensitization effect, availability, etc., and adjust the anonymization and desensitization strategies in a timely manner according to the monitoring results. If the data is found to be at risk of being re-identified or the availability is insufficient, it should be optimized immediately. The optimization time should be controlled within 24 hours to ensure that the security and availability of the data are always guaranteed. At the same time, the optimization results should be re-evaluated to ensure that the expected effect is achieved;

[0106] S7: Secure data sharing agreement. Develop a detailed data sharing agreement template to clarify the purpose, scope, method, security requirements and liability for breach of contract of data sharing. Before signing a data sharing agreement, require the data recipient's security level to reach a certain level of the international security standard ISO27001, and conduct strict security assessment and review of the data recipient. Use blockchain and other technologies to establish a decentralized data sharing platform, monitor the data sharing behavior on the blockchain once a second, and ensure that the data sharing process is transparent, traceable and cannot be tampered with. At the same time, monitor and audit data sharing behavior in real time to promptly discover and deal with violations.

[0107] The development of a detailed data sharing agreement template includes the following development steps:

[0108] 1): Clarify the purpose and scope of data sharing, and accurately explain the specific purpose of data sharing, such as scientific research cooperation in specific fields, accurate business analysis and decision-making, etc.; for data sharing in scientific research cooperation, strictly stipulate that data can only be used for established research projects and cannot be used for any other commercial purposes without authorization; when determining the scope of data sharing, list the data types involved in detail, such as text data, image data, numerical data, etc., and clearly define the upper limit of the data volume; for example, for medical data sharing, stipulate that only diagnostic data and treatment result data of specific diseases can be shared, and the data volume shall not exceed 100GB; if mixed sharing of multiple types of data is involved, explain the approximate proportion of each type of data, such as 30% for text data, 40% for image data, and 30% for numerical data;

[0109] Clarify the time range of the data, such as only sharing specific data within the past year; at the same time, specify the geographical scope of the data. If cross-border data sharing is involved, clarify the differences in data usage rights and legal constraints in different countries and regions;

[0110] 2): Formulate security requirements and stipulate the encryption method of data transmission. At least 256-bit Advanced Encryption Standard "AES" encryption algorithm must be used to ensure the security of data during transmission; the error rate of data transmission is strictly controlled below 0.1%, and encryption strength detection is performed before each data transmission to ensure that the encryption meets the standards. The detection time shall not exceed 10 minutes;

[0111] Clearly define the security responsibilities of the data recipient and require the recipient to have effective security protection measures. For example, the firewall protection level must meet enterprise-level standards and be able to resist common network attacks and malware intrusions. The response time of the intrusion detection system must not exceed 5 seconds. The effectiveness of the recipient's security protection system should reach more than 95%, and a comprehensive test and evaluation of the security protection system should be conducted once a month to ensure its continued effectiveness.

[0112] Carefully set data access permission control, and accurately assign different access permissions according to different user roles and needs. For example, ordinary users are only granted read-only permissions and can only access some non-sensitive data, which accounts for 60% of the total shared data. Advanced users can have read and write permissions, but operations on sensitive data require additional approval processes, and the approval time is strictly controlled within 24 hours. Each data access is subject to permission verification, and the verification time must not exceed 2 seconds.

[0113] 3): Determine the liability for breach of contract and clarify the specific liability of both parties when they breach the data sharing agreement; in the case of data leakage, stipulate that the breaching party must bear the corresponding legal liability and pay a certain amount of compensation. The amount of compensation is determined according to the sensitivity of the data and the severity of the leakage, generally between 100,000 yuan and 1 million yuan. If it involves highly sensitive data leakage, the compensation can be increased to 2 million yuan;

[0114] Establish an efficient dispute resolution mechanism. When disputes arise between the two parties during the data sharing process, they should first be resolved through friendly negotiation. The negotiation time should not exceed 30 days. If the negotiation fails, it can be resolved through arbitration or litigation. The arbitration or litigation time should be controlled within 90 days to ensure that the problem can be handled in a timely manner. During the dispute resolution period, data sharing will be suspended until the dispute is resolved.

[0115] S8: Application of privacy-enhancing technologies. Increase investment in the research and development of privacy-enhancing technologies such as differential privacy and homomorphic encryption. Invest at least RMB 5 million in the research and development of privacy-enhancing technologies to improve the maturity and availability of the technologies. Cooperate with universities and research institutes to conduct research and application pilots of related technologies and accumulate experience and data. Establish an evaluation system for privacy-enhancing technologies. Through evaluation, determine that the noise addition range of differential privacy is between 0.1 and 0.5 to balance data privacy and availability. Evaluate and compare different technologies and select the privacy-enhancing technologies that are most suitable for the human-machine-object fusion environment. At the same time, strengthen education for users to improve their awareness and acceptance of privacy-enhancing technologies.

[0116] S9: User education and awareness-raising; Carry out various forms of user education activities, using online courses, short videos, comics, etc., produce short videos of 5 to 10 minutes in length, and publish at least 5 new educational contents every month to improve users' understanding of data privacy and protection awareness; Through case analysis, simulated attacks, etc., let users intuitively understand the hazards of data privacy leakage and protection methods; Establish user feedback channels, give 100 to 500 points to users who report privacy issues in a timely manner, encourage users to report data privacy issues in a timely manner, and handle and respond to user feedback in a timely manner; At the same time, reward users for privacy protection behavior, and use points, coupons, etc. to improve user enthusiasm;

[0117] S10: Continuous monitoring and improvement, establish a comprehensive data privacy monitoring system, including network monitoring, equipment monitoring, application monitoring, etc., monitor and analyze network traffic every millisecond, and grasp the dynamics and security status of data in real time; use artificial intelligence and machine learning technology to analyze and warn monitoring data, and promptly discover potential privacy leakage risks; conduct a comprehensive assessment of data privacy protection measures every six months, formulate an improvement plan based on the assessment results, and implement it within 3 months; regularly evaluate and improve data privacy protection measures, and formulate improvement plans and measures based on the assessment results; at the same time, actively participate in industry standard setting and international cooperation to promote the development and application of data privacy protection technology;

[0118] Establishing a comprehensive data privacy monitoring system includes the following steps:

[0119] 1): Determine the monitoring scope and indicators, and make it clear that the monitoring scope covers all data storage locations, transmission channels and processing nodes in the human-machine-object fusion environment, including cloud servers, local databases, IoT devices, network communication links, etc. For example, for a system with 1,000 IoT devices and 5 data centers, ensure that each device and data center is effectively monitored;

[0120] Set specific monitoring indicators, such as data access frequency, abnormal traffic, data leakage risk index, etc. The data access frequency monitoring accuracy is once a minute. For normal business data, the access frequency should be between 10 and 100 times per minute. The detection threshold of abnormal traffic is set to 2 times the normal traffic. Once the threshold is exceeded, an alarm is triggered immediately. The data leakage risk index is calculated based on a variety of factors, ranging from 0 to 10. When the risk index reaches 5 or above, an in-depth investigation is conducted;

[0121] 2): Select monitoring technology and tools. Choose advanced monitoring technology, such as artificial intelligence-driven anomaly detection algorithms and encrypted traffic analysis technology. The accuracy of artificial intelligence anomaly detection algorithms must reach more than 90%, and they can quickly and accurately identify abnormal behaviors in large amounts of data. Encrypted traffic analysis technology can detect potential security threats without decrypting data, and the analysis speed must be no less than 100MB per second;

[0122] Deploy professional monitoring tools, such as network traffic monitoring equipment and data access log analysis software. The monitoring accuracy of network traffic monitoring equipment reaches the data packet level and can monitor all data traffic in the network in real time. The data access log analysis software can process 100GB of log data within 1 hour and extract key information for risk assessment.

[0123] Establish a centralized monitoring management platform to integrate the data of all monitoring tools into one platform for unified management and analysis. The platform's response time should not exceed 5 seconds, ensuring that security incidents can be discovered and handled in a timely manner.

[0124] 3): Implement a monitoring and response mechanism, continuously monitor data privacy, implement a 24-hour uninterrupted monitoring system, and assign dedicated personnel to monitor the platform. Each shift should not exceed 8 hours.

[0125] When an abnormal situation is detected, the response mechanism is immediately activated. According to the severity of the abnormality, it is divided into three levels: low, medium and high. Low-level abnormalities are handled within 1 hour, medium-level abnormalities are handled within 30 minutes, and high-level abnormalities are handled within 15 minutes.

[0126] Analyze and summarize the monitoring results regularly, generate a monitoring report every week, and record in detail the monitored security incidents, handling situations and improvement suggestions. According to the monitoring report, adjust the monitoring strategies and security measures in a timely manner, and continuously optimize the data privacy monitoring system. For example, if an increase in the same type of security incidents is monitored for two consecutive weeks, a special investigation should be conducted immediately and targeted preventive measures should be taken.

[0127] S11: Application of data watermark technology. For data that needs to be shared, data watermark technology is used to embed invisible watermarks, which contain information such as the data owner and the scope of authorized use. In scenarios such as medical imaging data and IoT environmental monitoring data, once the data is illegally disseminated or abused, the data source and authorization status can be traced by detecting the watermark. The strength of the watermark is set at a moderate level, which not only ensures that the data quality is not affected when the data is used normally, but also can be easily detected when it is illegally copied. The detection accuracy of the watermark should reach more than 95%. After 10 data format conversions, the detection success rate of the watermark should still be no less than 80%. The robustness of the data watermark is continuously optimized to ensure that it can still be effectively detected after the data is modified, converted, and other operations. At the same time, in scenarios with extremely high requirements for data accuracy, the watermark strength and data quality are balanced to avoid unacceptable impacts on the data.

[0128] S12: Comprehensive technology application and collaboration. Combine edge computing with privacy protection to perform preliminary privacy processing on edge devices close to the data source, such as data encryption and anonymization. In a smart home system, the video data collected by a smart camera is encrypted through face recognition at the edge gateway, and only the encrypted feature data is sent to the cloud server for storage and sharing. The encryption processing time of the edge device should be controlled within 100 milliseconds to ensure data real-time performance. Develop a homomorphic encryption hardware acceleration module. The hardware acceleration module should increase the computing speed of homomorphic encryption by at least 5 times and control the cost within 100,000 yuan to improve the computing speed of homomorphic encryption so that it can be applied to scenarios with high real-time requirements, such as financial data sharing and cloud computing. At the same time, pay attention to the cost and technology adaptation risks of the hardware acceleration module. Optimize the multi-party secure computing protocol. The optimized multi-party secure computing protocol should shorten the computing time by more than 30%, and the vulnerability discovery rate of the security assessment should be lower than 0.1% to improve computing efficiency and security. In scenarios such as sharing patient data among multiple medical institutions for joint research, adopt more efficient secret sharing schemes and zero-knowledge proof technologies, etc. And conduct strict security assessments on the optimized protocol to prevent new security risks.

[0129] The algorithm formula of the k-anonymity algorithm in this method is:

[0130] Equivalence class judgment: If r1 = (a11, a12,..., a1n) and r2 = (a21, a22,..., a2n), for all i = 1, 2,..., m "m < n", when a1i = a2i, r1 and r2 are in the same equivalence class;

[0131] Equivalence class size requirement: For all i = 1, 2,..., p, |Ci| ≥ k, where {Ci} is the set of equivalence classes;

[0132] The algorithm formula of the l-diversity algorithm in this method is:

[0133] Diversity requirement: Based on k-anonymity, let the sensitive attribute be As. For the equivalence class Ci, let the set of different values of the sensitive attribute As in Ci be Vis. For all i = 1, 2,..., p, |Vis| ≥ l;

[0134] Decision tree - Information gain "ID3 algorithm"

[0135] Entropy of the dataset: Entropy(D) = -∑(i = 1 to k)p(ci)log2p(ci), where p(ci) is the probability of the class ci in the dataset D;

[0136] Information gain: Gain(D,A)=Entropy(D)-∑(i=1 to v)(|Di| / |D|)Entropy(Di), v is the number of different values ​​of attribute A, {Di} is the subset after division;

[0137] Random Forest - Out-of-bag data "OOB" error estimation;

[0138] OOB_error = (1 / n)∑(i=1 to n)I(yi(x)≠y(x)), where n is the total number of OOB samples, I(·) is the indicator function, yi(x) is the predicted category of sample x by the i-th tree, and y(x) is the true category;

[0139] Neural Networks - Single Hidden Layer Forward Propagation;

[0140] Hidden layer output: h = σ(W1x + b1), x = (x1, x2, ..., xn) is the input vector, W1 is the weight matrix "m × n", b1 is the bias vector "m × 1), σ(·) is the activation function "such as σ(z) = 1 / (1 + e^(-z))";

[0141] Output layer output: y = σ(W2h + b2), W2 is the weight matrix "p × m", b2 is the bias vector "p × 1);

[0142] Differential privacy-Laplace mechanism;

[0143] M(f(D))=f(D)+Y, where Y follows the Laplace distribution Lap(Δf / ε), Δf=max(|f(D1)-f(D2)|) "D1, D2 are adjacent data sets", ε is the privacy budget, and f is the function from data set D to the set of real numbers;

[0144] Differential privacy-exponential mechanism;

[0145] Output probability: P(x) is proportional to exp((εu(D,x)) / (2Δu)), Δu = max(|u(D1,r)-u(D2,r)|) "D1, D2 are adjacent data sets, r belongs to the output domain", ε is the privacy budget, and u is the utility function from the data set and the output domain to the real number set;

[0146] The homomorphic encryption algorithm formula in this method is "Additive Homomorphic Encryption Example":

[0147] Enc_pk(m1+m2)=Eval_pk(c1,c2), c1=Enc_pk(m1), c2=Enc_pk(m2), pk is the public key, Enc(·) is the encryption function, and Eval_pk(·,·) is the homomorphic computation function.

[0148] The working principle of the present invention is: determine various data sources in the human-machine-object fusion environment, including intelligent devices, computer systems, databases and manually input data, etc., to lay the foundation for comprehensive data combing, conduct detailed investigation and records of each data source, use data cleaning tools to remove duplicate and erroneous data, ensure data accuracy and completeness, control the removal ratio of duplicate data within 5%, and achieve an outlier detection accuracy of more than 90%. Formulate classification standards from multiple perspectives such as sensitivity, owner and life cycle stage, and divide data into highly sensitive, moderately sensitive and generally sensitive data. Clarify the privacy protection strategies of data of different owners, as well as the risks and protection requirements of data at different stages, formulate clear definitions and feature descriptions for each classification, and achieve an accuracy rate of more than 95% for highly sensitive data identification. Use a combination of automation and manual review to classify, improve accuracy, label the classified data, facilitate subsequent management and protection, and conduct regular review and reclassification to ensure that privacy protection measures are continuously effective.

[0149] In terms of risk assessment and threat modeling, artificial intelligence and big data analysis technologies are introduced to collect data from multiple data sources. Status data is collected from smart factory equipment every hour to ensure the comprehensiveness and real-time nature of the data. Data is pre-processed to remove noise and outliers. Machine learning algorithms are used for feature extraction and model training. Appropriate algorithms are selected to improve model accuracy and generalization capabilities. Threats are classified and a real-time early warning system is established. Risk assessment results are analyzed monthly, and models and early warning rules are updated.

[0150] In terms of data encryption, we will study and apply quantum computing-resistant encryption algorithms, determine the types of protected data and security level requirements, conduct in-depth research on multiple algorithms, select appropriate algorithms based on needs, consider scalability and compatibility, conduct testing and optimization, deploy them to various links, and establish maintenance mechanisms;

[0151] Use biometrics and behavior analysis technology to establish an accurate identity authentication mechanism, collect user biometric information and encrypt and store it, collect behavior data for modeling and analysis, combine the two for comprehensive authentication, control the response time within 2 seconds, and handle continuous authentication failures;

[0152] In terms of data anonymization and desensitization, we analyze the data type and sensitivity, use advanced algorithms such as k-anonymity and l-diversity, evaluate and attack the anonymized data, select appropriate desensitization technology, control data quality and availability, establish a monitoring mechanism, and check and adjust strategies every month.

[0153] In addition, privacy enhancement technology, data watermark technology and comprehensive technology applications work together to provide strong support for data privacy protection, increase R&D investment, cooperate with universities, establish an evaluation system, use data watermark technology to trace data, combine edge computing for privacy processing, develop homomorphic encryption hardware acceleration modules, and optimize multi-party secure computing protocols;

[0154] Establish a comprehensive data privacy monitoring system, clearly define the monitoring scope to cover all data storage locations, transmission channels and processing nodes, set specific monitoring indicators, select advanced monitoring technologies and tools, establish a centralized monitoring management platform, ensure timely detection and handling of security incidents, implement a monitoring and response mechanism, monitor 24 hours a day, handle abnormalities according to severity, regularly analyze and summarize monitoring results, adjust monitoring strategies and security measures, and in terms of data sharing, formulate detailed agreement templates, clarify the purpose, scope, method, security requirements and liability for breach of contract, stipulate encryption methods, security responsibilities of the recipient and access permission control, establish a dispute resolution mechanism, and at the same time, focus on user education and awareness-raising, carry out various educational activities, produce short videos, and use case analysis and simulated attacks to let users understand the dangers of privacy leaks and protection methods, and establish feedback channels.

Claims

1. A method for data privacy protection and sharing in a human-machine-object fusion environment, characterized in that: The steps include: S1: Clarify the scope and sensitivity of data, comprehensively sort out and classify data, predict the types of encrypted data that will be affected by new technology trends and plan strategies in advance, encrypt particularly sensitive data such as those involving national security and major commercial secrets, establish a dynamic data list for automatic scanning and updating, and update key data-related information in real time to cope with the rapid changes in data in the human-machine-object fusion environment; S2: Risk assessment and threat modeling. We introduce artificial intelligence and big data analysis technologies to carry out risk assessment and threat modeling, implement more comprehensive and in-depth risk assessments on data in the human-machine-object fusion environment, specifically analyze network traffic data and device logs in the past week to determine potential risk points, conduct simulated attack experiments of attack scenarios in the threat simulation laboratory every month, and test and improve protection measures; S3: Formulate privacy policies and regulations. Formulate strict and detailed privacy policies and regulations, clarify the legal responsibilities and penalties of enterprises and organizations in data privacy protection, protect user rights by signing legally binding privacy agreements with users, stipulate the scope of fines for policy violations, arrange employee training every year, regularly update policies and regulations to adapt to changes in regulations and technology, strengthen employee training, and enhance their compliance awareness; S4: Data encryption and secure storage, research and apply encryption algorithms that are resistant to quantum computing, prepare for data privacy protection in the quantum era, use multi-layer encryption technology, differentiate by sensitivity, adopt distributed storage, set up multiple storage nodes and each node has 3 copies of data backup, set up access control policies, assign different permissions according to user types, establish a distributed, redundant secure storage system, store data in fragments and encrypt each fragment independently; S5: Access control and identity authentication. Use biometrics, behavioral analysis and other technologies to establish a precise identity authentication mechanism. The legitimacy of the identity is determined based on the user's typing speed, mouse movement trajectory and other behavioral characteristics. In terms of encryption algorithm selection, data transmission uses a 128-bit symmetric encryption algorithm, and highly sensitive data transmission uses a 256-bit asymmetric encryption algorithm combined with a digital signature. A role-based and attribute-based access control model is used to grant corresponding data access rights according to the user's position level and work requirements. S6: Data anonymization and desensitization processing, develop more advanced data anonymization and desensitization technologies, use deep learning-based anonymization technology to improve anonymization and desensitization effects, control the probability of re-identification of anonymized data to less than 1%, use k-anonymity technology for medical data, and l-diversity technology for financial data, establish standard specifications for data anonymization and desensitization, and clarify that at least 90% of personal identity information must be removed when anonymizing medical data, so as to guide relevant entities to correctly handle data and prevent privacy leakage risks; S7: Secure data sharing agreement. Develop a detailed data sharing agreement template to clarify all aspects of data sharing. Before signing the contract, require the data recipient to meet a certain ISO27001 security level and conduct a strict assessment and review. Use blockchain technology to build a decentralized sharing platform, monitor data sharing behavior on the platform every second, ensure that the sharing process is transparent, traceable and cannot be tampered with, and monitor and audit sharing behavior in real time to deal with violations in a timely manner. S8: Privacy enhancement technology application, increase investment in privacy enhancement technology research and development, cooperate with universities and research institutions to carry out relevant technology research and application pilots, accumulate experience data, establish an evaluation system, determine the range of differential privacy noise addition to balance data privacy and availability, compare different technologies to select technologies that are suitable for the human-machine-object fusion environment, and strengthen user education to improve user awareness and acceptance of privacy enhancement technology; S9: User education and awareness raising: Through various forms of user education activities, short videos are released every month to enhance users' knowledge of data privacy and protection awareness. Case analysis and simulated attacks are used to let users intuitively know the dangers of privacy leakage and protection methods. User feedback channels are also established to reward users with points for timely reporting of privacy issues, encourage feedback and handle responses in a timely manner, and reward users' privacy protection behaviors to increase their enthusiasm; S10: Continuous monitoring and improvement, establishing a comprehensive data privacy monitoring system, establishing a monitoring system covering networks, devices, applications, etc., monitoring and analyzing network traffic every millisecond, controlling data dynamics and security status in real time, using artificial intelligence and machine learning technology to analyze monitoring data, issue early warnings, and promptly detect potential privacy leakage risks. Comprehensively evaluate data privacy protection measures every six months, formulate improvement plans based on the results, and conduct regular evaluation and improvement work. In addition, actively participate in the formulation of industry standards and international cooperation to promote the development and application of data privacy protection technology; S11: Application of data watermarking technology: Use data watermarking technology to embed watermarks containing information such as the owner and authorization scope into the data to be shared. In relevant scenarios, the watermark can be used to trace the source and authorization status. The watermark strength is set moderately to ensure that normal use does not affect the quality. Illegal copying is easy to detect. There are requirements for detection accuracy and success rate. Robustness is also optimized to balance strength and quality in high-precision scenarios. S12: Comprehensive technology application and collaboration, combining edge computing with privacy protection, performing preliminary privacy processing on edge devices, developing homomorphic encryption hardware acceleration modules to increase speed and control costs, and applying them to scenarios with high real-time requirements. Optimize multi-party secure computing protocols, shorten computing time, and reduce vulnerability rates. Used in scenarios such as medical institution data sharing and strictly evaluate and prevent risks.

2. The method for data privacy protection and sharing in a human-machine-object fusion environment according to claim 1 is characterized in that: The comprehensive sorting and classification of data in S1 includes the following steps: 1) Data collection and organization: Determine the data sources in the human-machine-object fusion environment, including intelligent devices, computer systems, etc., investigate and record the information of each data source in detail, and use data cleaning tools to remove duplicate and erroneous data to ensure that the data is accurate and complete; 2) Development of data classification standards: classify data from multiple perspectives such as sensitivity, owner, and life cycle stage, clarify the definition and characteristics of each classification, lay the foundation for subsequent accurate classification, and ensure that different types of data receive appropriate privacy protection strategies; 3) Data classification and labeling: Classify data according to standards, use automated tools combined with manual review to ensure accuracy, label data to facilitate subsequent management and protection, and conduct regular review and reclassification to adapt to changes. This should be done once a quarter, and each time should be within one week to ensure that privacy protection measures remain effective.

3. The method for data privacy protection and sharing in a human-machine-object fusion environment according to claim 1, characterized in that: The introduction of artificial intelligence and big data analysis technology in S2 includes the following steps: 1) Data collection and preprocessing: Use artificial intelligence collection tools to collect data from multiple data sources in the human-machine-object fusion environment, preprocess the data, use data cleaning algorithms to remove noise and outliers, control the proportion of outliers, determine the normal range of network traffic and remove abnormal data; 2) Feature extraction and model training: Use machine learning algorithms to extract features, select appropriate algorithms to train models, use random forest algorithms for threat detection, set accuracy and recall targets, evaluate performance through cross-validation, and continuously adjust parameters to improve model accuracy and generalization capabilities to ensure stability; 3) Risk assessment and threat warning: Apply the model to actual data to assess risks, classify them into high, medium and low levels, establish a real-time warning system, respond quickly when risks are high, analyze the assessment results monthly, and update the model and warning rules based on new threats to respond to potential risks in a timely manner and ensure data security.

4. The method for data privacy protection and sharing in a human-machine-object fusion environment according to claim 1, characterized in that: The research and application of quantum computing-resistant encryption algorithms in S4 include the following steps: 1) Demand analysis: determine the type of data and security level to be protected, evaluate the system to determine the quantum-resistant encryption link, and require the encryption algorithm to have a low probability of being cracked by quantum attacks; 2) Algorithm research: conduct in-depth research on various quantum-resistant algorithms, compare their principles, advantages, limitations and performance, and select appropriate algorithms according to different scenarios; 3) Algorithm selection: Select the most suitable algorithm based on the requirements and research results, taking into account scalability and compatibility. A hash algorithm with high real-time requirements can be selected and integrated with the existing system; 4) Testing and optimization: Test the algorithm, including performance and security, and make optimization adjustments based on the results to improve speed and security; 5) Deployment and maintenance: Deploy the algorithm to each link to ensure the success rate, establish a maintenance mechanism, regularly evaluate and upgrade, train personnel, fully cover security assessments, and provide no less than 10 hours of training each year.

5. The method for data privacy protection and sharing in a human-machine-object fusion environment according to claim 1, characterized in that: The use of biometric recognition, behavior analysis and other technologies in S5 includes the following steps: 1) Biometric identification settings: Collect biometric information such as fingerprints, faces, and irises, set accuracy, resolution, and collection time standards for each identification method, and store the collected information in an encrypted database with an encryption strength of AES256 bits or more to ensure information security and prevent illegal access; 2) Behavior analysis modeling: collect behavioral data such as typing speed and mouse trajectory, use machine learning algorithms to build models, and use training data sets of no less than 1,000 samples. Control the training time so that the model's behavior recognition accuracy reaches more than 95%; 3) Comprehensive identity authentication: Authentication is based on biometrics and behavioral analysis results. If both match, the authentication probability is 100%. If one matches, it drops to 70% and needs to be re-verified. The authentication response is within 2 seconds. If multiple failures occur, the account is automatically locked and the login behavior is recorded for subsequent review.

6. The method for data privacy protection and sharing in a human-machine-object fusion environment according to claim 1, characterized in that: The development of more advanced data anonymization and desensitization technology in step S6 specifically includes the following steps: 1) Data feature analysis: Comprehensively analyze the data that needs to be anonymized and desensitized, classify the types, clarify the structure and sensitivity level, and count the proportion of each type of data to lay the foundation for targeted processing; Application of anonymization technology: Use algorithms such as k-anonymity and l-diversity, set corresponding parameters according to data types, evaluate the anonymized data, control the probability of re-identification and the success rate of attack tests, and ensure the anonymization effect and security; 2) Desensitization technology selection: select the appropriate desensitization method according to the data type and purpose, control the precision and amplitude of the numerical type, and ensure the accuracy of keyword replacement for the text type, etc., evaluate the quality of the desensitized data, ensure the similarity with the original data and meet the analysis application requirements; 3) Continuous monitoring and optimization: Establish a monitoring mechanism to conduct a comprehensive monthly check of anonymized and desensitized data, covering multiple aspects. Adjust strategies in a timely manner based on the results, control optimization time, and re-evaluate to ensure data security and meet expected results.

7. The method for data privacy protection and sharing in a human-machine-object fusion environment according to claim 1, characterized in that: The detailed data sharing agreement template in S7 includes the following steps: 1) Clarify the purpose and scope of sharing: clearly explain the purpose of data sharing, limit the scope of use, and specify the type of data involved, the upper limit of the amount, the time and geographical scope, etc. For example, medical data sharing stipulates the data content and amount, and cross-border sharing clarifies the authority and legal differences to ensure accurate and compliant sharing; 2) Establish security requirements: stipulate that AES256-bit encryption be used for transmission, control transmission error rate, clarify the security responsibilities of the recipient, set access control, assign permissions based on roles, and strictly verify and approve processes to ensure data transmission and access security; 3) Determine the liability for breach of contract: clearly define the liability for breach of contract for violation of the agreement, determine compensation for data leakage based on the sensitivity of the data, establish a dispute resolution mechanism, negotiate first, and then arbitrate or sue if no results are achieved. Suspend sharing during the dispute period to ensure that problems are handled in a timely manner and responsibilities are clearly defined.

8. The method for data privacy protection and sharing in a human-machine-object fusion environment according to claim 1, characterized in that: The establishment of a comprehensive data privacy monitoring system in S10 includes the following steps: 1) Determine the monitoring scope and indicators: clearly define the monitoring scope of all parts of the human-machine-object fusion environment, set specific indicators such as data access frequency, such as monitoring of equipment and data centers, set access frequency accuracy, set alarms for abnormal traffic exceeding thresholds, and conduct in-depth investigations when the risk index is high; 2) Select monitoring technology and tools: Select advanced technologies such as artificial intelligence algorithms with an accuracy rate of more than 90%, deploy professional tools, build a centralized management platform, and ensure a response time of no more than 5 seconds. Integrate data for unified management and analysis to ensure timely discovery and handling of security incidents; 3) Implement monitoring and response mechanism: 24-hour uninterrupted monitoring, dedicated personnel on duty, abnormal level response, regular analysis and summary reports, adjustment of strategies and measures based on reports, optimization of monitoring system, and special investigation and prevention in case of an increase in similar incidents.

Citation Information

Cited By

  • Method for quickly starting function application based on short-distance wireless communication technology

    CN120434637A

  • Operating personnel safety portrait data sharing method and system based on block chain technology

    CN120688098A

  • Cross-industry data sharing method and system supporting dynamic policy negotiation

    CN121750666A

  • Method, device and system for automatically identifying sensitive data to carry out de-identification processing

    CN121980617A