Private data processing method and device, equipment and storage medium

By acquiring user level tags and data characteristics, determining privacy protection levels and models, and adopting appropriate privacy protection technologies, the challenge of user privacy protection in big data analysis has been solved, achieving efficient and accurate privacy protection and data analysis.

CN120995490APending Publication Date: 2025-11-21INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510857825.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In the process of big data analysis, how to effectively protect user privacy and prevent the leakage of personal information is a challenge that existing technologies struggle to provide targeted and practical privacy protection.

Method used

By acquiring user level tags, data types, and characteristic attributes of the data to be processed, the privacy protection level and model are determined, and different privacy protection technologies such as data de-identification, homomorphic encryption, and differential privacy are used for targeted processing.

Benefits of technology

It achieves precise privacy protection for different data types and user levels, reduces computational overhead and storage burden, and improves the efficiency of data analysis and the effectiveness of privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995490A_ABST
    Figure CN120995490A_ABST
Patent Text Reader

Abstract

The invention provides a privacy data processing method and device, equipment and a storage medium, and the privacy data processing method comprises the steps: obtaining to-be-processed data and a user level label corresponding to the to-be-processed data, and determining the data type and feature attribute of the to-be-processed data; determining a privacy protection level and a privacy protection model of the to-be-processed data according to the data type, the feature attribute and the user level tag; and performing privacy protection processing on the to-be-processed data according to the privacy protection level and the privacy protection model. According to the method and the device, the sexual privacy protection level and the privacy protection model for the to-be-processed data are determined through the data type, the feature attribute and the user level label, so that effective privacy protection on the to-be-processed data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data technology, and in particular to a method, apparatus, device and storage medium for processing privacy data. Background Technology

[0002] With the development of big data technologies, big data analytics has been widely applied, penetrating various fields and industries. Using big data analytics, shopping websites can recommend products of interest to users, increasing sales revenue. Currently, to conduct big data analytics, a large amount of user-related data, including personal information, preferences, and browsing history, is collected by relevant companies and organizations for analysis. This data is highly sensitive because it contains users' personal information; even slight misuse can lead to privacy leaks. Therefore, how to effectively protect user privacy during big data mining has become a pressing issue. Summary of the Invention

[0003] This invention provides a privacy data processing method, apparatus, device, and storage medium to overcome the deficiencies in the prior art and achieve effective privacy protection of data.

[0004] This invention provides a method for processing privacy data, comprising: Obtain the data to be processed and the user level tags corresponding to the data to be processed, and determine the data type and characteristic attributes of the data to be processed; Based on the data type, the feature attributes, and the user level label, determine the privacy protection level and privacy protection model of the data to be processed; The data to be processed is subjected to privacy protection processing based on the privacy protection level and the privacy protection model.

[0005] According to a privacy data processing method provided by the present invention, determining the data type and characteristic attributes of the data to be processed includes: Determine the data type of the data to be processed; The data to be processed is preprocessed, and a feature vector is constructed based on the preprocessed data to be processed. The feature vector is input into the sensitive feature recognition model to obtain the feature attributes of the data to be processed.

[0006] According to a privacy data processing method provided by the present invention, the step of performing privacy protection processing on the data to be processed based on the privacy protection level and the privacy protection model includes: The desensitization parameters and data sampling parameters corresponding to the privacy protection model are determined based on the privacy protection level; Based on the desensitization parameters, data sampling parameters, and the privacy protection model, the data to be processed is subjected to privacy protection processing.

[0007] According to a privacy data processing method provided by the present invention, determining the privacy protection level and privacy protection model of the data to be processed based on the data type, the feature attributes, and the user level label includes: The risk level of the data to be processed is assessed based on a privacy risk assessment model. Based on the risk level, the data type, the characteristic attributes, and the user level label, determine the privacy protection level and privacy protection model of the data to be processed.

[0008] According to a privacy data processing method provided by the present invention, after performing privacy protection processing on the data to be processed based on the privacy protection level and the privacy protection model, the method further includes: Based on the data to be processed and the data after privacy protection, the desensitization effect is evaluated; The privacy protection model is optimized based on the results of the desensitization effect evaluation.

[0009] According to a privacy data processing method provided by the present invention, before obtaining the data to be processed and the user level tag corresponding to the data to be processed, the method further includes: Obtain user data from multiple users; Based on the user data, multiple users are clustered to obtain multiple user clusters; The user level label for each user cluster is determined based on the user data.

[0010] According to a privacy data processing method provided by the present invention, determining the privacy protection level and privacy protection model of the data to be processed based on the data type, the feature attributes, and the user level label includes: If the data type is numerical, the feature attribute is a non-sensitive feature, the user level label is a first level label, the privacy protection level is determined to be the first protection level, and the privacy protection model is at least one of the data de-identification model, data masking model, and data aggregation model; If the data type is numeric, text, or date, the feature attribute is a non-sensitive feature, the user level label is a second-level label or a third-level label, the privacy protection level is determined to be the second protection level, and the privacy protection model is at least one of homomorphic encryption, data desensitization encryption, or lexicalization. If the feature attribute is a sensitive feature, the privacy protection level is determined to be the third protection level, and the privacy protection model is at least one of the differential privacy model, secure multi-party computation model, and anonymization model.

[0011] The present invention also provides a privacy data processing apparatus, comprising: The first determining module is configured to obtain the data to be processed and the user level tag corresponding to the data to be processed, and to determine the data type and feature attributes of the data to be processed; The second determining module is configured to determine the privacy protection level and privacy protection model of the data to be processed based on the data type, the feature attributes, and the user level label. The privacy protection processing module is configured to perform privacy protection processing on the data to be processed according to the privacy protection level and the privacy protection model.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the privacy data processing method as described above.

[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the privacy data processing method as described above.

[0014] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the privacy data processing method as described above.

[0015] The privacy data processing method, apparatus, device, and storage medium provided by this invention acquire the data to be processed and the corresponding user level tag, determine the data type and characteristic attributes of the data to be processed, determine the privacy protection level and privacy protection model of the data to be processed based on the data type, characteristic attributes, and user level tag, and perform privacy protection processing on the data to be processed according to the privacy protection level and privacy protection model. This invention determines a personalized privacy protection level and privacy protection model for the data to be processed by using data type, characteristic attributes, and user level tag, making privacy protection of the data to be processed more targeted and operable, and enabling effective privacy protection. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating the privacy data processing method provided by the present invention.

[0018] Figure 2 This is a schematic diagram of the privacy data processing device provided by the present invention.

[0019] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0021] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0022] Figure 1 This is a flowchart illustrating a privacy data processing method according to an exemplary embodiment. Figure 1 As shown in an exemplary embodiment, the privacy data processing method includes steps 110 to 130, which are described in detail below.

[0023] Step 110: Obtain the data to be processed and the user level tag corresponding to the data to be processed, and determine the data type and feature attributes of the data to be processed.

[0024] In this embodiment of the invention, the data to be processed and the user level tag corresponding to the data to be processed are obtained, and the data type and feature attributes of the data to be processed are determined.

[0025] Step 120: Determine the privacy protection level and privacy protection model of the data to be processed based on the data type, the feature attributes, and the user level label.

[0026] In this embodiment of the invention, the privacy protection level and privacy protection model of the data to be processed are determined based on the data type, feature attributes, and user level tags.

[0027] Optional privacy protection levels include: Level 1 protection (Low): This includes simple desensitization of non-sensitive data, such as data generalization and data desensitization. The second level of protection (Medium) includes taking stricter privacy protection measures for data that is sensitive to a certain extent, which may require the use of encryption or anonymization technologies; Level 3 Protection (High): This level typically involves the most sensitive data and may require advanced privacy protection technologies such as differential privacy to ensure extremely high data privacy.

[0028] Step 130: Perform privacy protection processing on the data to be processed according to the privacy protection level and the privacy protection model.

[0029] In this embodiment of the invention, privacy protection processing is performed on the data to be processed according to the privacy protection level and privacy protection model in order to protect data privacy.

[0030] In this embodiment of the invention, determining the privacy protection level and model through user level tags, data types, and characteristic attributes helps to accurately select privacy protection operations and avoid introducing excessive noise during the privacy protection process, thereby reducing the impact on data analysis results. Secondly, setting the privacy protection level based on user level tags, data types, and characteristic attributes helps to clarify which data requires what level of protection, thus avoiding unnecessary encryption or protection measures and reducing resource consumption. For example, effective privacy protection operations can alleviate the storage burden caused by data privacy processing to some extent.

[0031] In an exemplary embodiment of the present invention, determining the data type and characteristic attributes of the data to be processed includes: Determine the data type of the data to be processed; The data to be processed is preprocessed, and a feature vector is constructed based on the preprocessed data to be processed. The feature vector is input into the sensitive feature recognition model to obtain the feature attributes of the data to be processed.

[0032] In this embodiment of the invention, the data to be processed by big data users is obtained, and the data type and characteristic attributes of the data are determined.

[0033] Optional data types include numeric, text, and date types.

[0034] Optionally, feature attributes include sensitive features and non-sensitive features.

[0035] Optionally, sensitive features may involve personal privacy, such as ID card numbers, bank accounts, medical records, financial transaction passwords, personal phone numbers, and other sensitive personal data.

[0036] Optionally, non-sensitive features typically refer to general information, such as basic identification information, general social characteristics (e.g., occupation, education level), and public information. This information usually does not contain sensitive, personal privacy, or confidential information. Through the above operations, the system can effectively distinguish between these two feature categories, thereby adopting appropriate privacy protection measures to ensure that sensitive information is properly handled.

[0037] The data to be processed is cleaned to remove noise and redundant information, and then various attributes are extracted as feature vectors. A sensitive feature identification model is then used to determine the characteristic attributes of the data. This model is trained based on machine learning models such as logistic regression and support vector machines. Based on this model, features such as credit card numbers and biometric information can be labeled as sensitive features. On the other hand, non-sensitive features that do not involve personal privacy, such as age, gender, and general purchasing preferences, are identified. These features typically do not involve sensitive user information and can be widely used in the analysis.

[0038] In this embodiment of the invention, data types and feature attributes are identified through appropriate algorithms and models, such as rule-based methods, supervised learning, or unsupervised learning methods, which can accurately identify sensitive and non-sensitive features.

[0039] In an exemplary embodiment of the present invention, the step of performing privacy protection processing on the data to be processed according to the privacy protection level and the privacy protection model includes: The desensitization parameters and data sampling parameters corresponding to the privacy protection model are determined based on the privacy protection level; Based on the desensitization parameters, data sampling parameters, and the privacy protection model, the data to be processed is subjected to privacy protection processing.

[0040] In this embodiment of the invention, the data to be processed is processed according to the privacy protection level and privacy protection model through parameter selection and data sampling and / or hierarchical processing methods.

[0041] Different privacy protection levels have pre-set de-identification parameters, such as the degree of noise addition and the magnitude of data perturbation. The higher the privacy protection level, the larger the corresponding de-identification parameters.

[0042] Data sampling parameters characterize the parameters used to select data from the dataset when protecting the privacy of the data to be processed. When sampling data based on data sampling parameters, it is necessary to sample uniformly from the dataset to be processed to ensure that the selected data is representative and has a certain degree of diversity, thereby reducing computational overhead.

[0043] In this embodiment of the invention, the de-identification algorithm parameter settings are optimized by using de-identification parameters, while data sampling technology is used to reduce the amount of data processed. This achieves the technical effect of reducing computational overhead, improving processing efficiency, maintaining data quality, and reducing the risk of privacy leaks.

[0044] In the hierarchical processing approach, the data to be processed is categorized based on the sensitivity of each data point, such as separating physiological indicators and diagnostic information into two different levels. Then, multi-level processing is performed based on this categorization. For lower-sensitivity data (physiological indicators), encryption algorithms are used for de-identification; for high-sensitivity data (diagnostic information), data deletion or data de-labeling methods are used for de-identification.

[0045] In this embodiment of the invention, corresponding processing strategies are formulated for different types of data attributes. For example, differential privacy methods are used to process time series data. By classifying data according to its sensitivity, processing data at multiple levels, and customizing processing schemes, the precision and flexibility of data anonymization are improved, better meeting the privacy needs of different data types and enhancing the adaptability and privacy protection capabilities of data anonymization.

[0046] In the above implementation, by using desensitization parameters, data sampling parameters, and hierarchical processing methods, the requirements of different privacy protection levels can be achieved more accurately, thereby improving the level of privacy protection. Employing strategies such as desensitization parameters, data sampling parameters, and hierarchical processing methods can reduce the computational load and time overhead during the desensitization process. In data desensitization, by carefully selecting parameters and sampling representative data samples, the impact on data quality can be minimized. The hierarchical processing method employs different processing strategies for data types with different levels of sensitivity, thereby achieving personalized and customized data desensitization processing and better protecting data privacy.

[0047] This invention combines parameter selection, data sampling, and hierarchical processing to achieve more efficient, accurate, and controllable results in the data anonymization process, further improving data privacy protection, reducing computational overhead, and maximizing data quality protection, thus achieving accurate and efficient data anonymization.

[0048] In an exemplary embodiment of the present invention, determining the privacy protection level and privacy protection model of the data to be processed based on the data type, the feature attribute, and the user level tag includes: The risk level of the data to be processed is assessed based on a privacy risk assessment model. Based on the risk level, the data type, the characteristic attributes, and the user level label, determine the privacy protection level and privacy protection model of the data to be processed.

[0049] In this embodiment of the invention, a risk assessment mechanism is introduced to dynamically adjust the privacy protection level and privacy protection model based on the risk situation during data processing, in order to cope with the ever-changing threats and privacy leakage risks.

[0050] Potential privacy risks are assessed using a privacy risk assessment model. Then, based on the risk level, user level tags, data types, and characteristic attributes, the privacy protection level and privacy protection model are determined.

[0051] A privacy risk assessment model is a framework or methodology used to assess and quantify potential privacy risks in the processing of personal data. Such models typically combine multiple factors, including data sensitivity, data access control, data sharing policies, and data storage security, to comprehensively evaluate the degree of risk that data may face at different stages.

[0052] In an exemplary embodiment of the present invention, after performing privacy protection processing on the data to be processed according to the privacy protection level and the privacy protection model, the method further includes: Based on the data to be processed and the data after privacy protection, the desensitization effect is evaluated; The privacy protection model is optimized based on the results of the desensitization effect evaluation.

[0053] In this embodiment of the invention, the desensitization effect is evaluated and the level of privacy protection is quantified based on the original data to be processed and the data after privacy protection.

[0054] Desensitization effectiveness evaluation refers to assessing the effectiveness and quality of data privacy protection during the data processing process by analyzing (e.g., mean, variance) and comparing the data after privacy protection.

[0055] Optionally, when evaluating the effectiveness of data anonymization, it is possible to check whether sensitive data has been effectively obscured or hidden. Ensure that the basic statistical properties and valid characteristics of the data after privacy protection are maintained at a certain level.

[0056] Optionally, metrics such as information entropy and mutual information can be used to detect the similarity or difference between the privacy-protected data and the data to be processed.

[0057] Optionally, quantifying the level of privacy protection includes, but is not limited to: analyzing the degree of data distortion introduced during the privacy protection process and quantifying it using metrics (such as mean squared error). For example: 1) Mutual information quantification: calculating the mutual information between data before and after privacy protection to measure the anonymization effect; 2) KL divergence measurement: using KL divergence to compare the similarity between two probability distributions and quantifying the information loss introduced by privacy protection; 3) Mean squared error (MSE): calculating the mean squared error between data before and after privacy protection to assess data changes; 4) Differential privacy budget: quantifying the strength of privacy protection in differential privacy mechanisms; 5) Data re-identifiability assessment: assessing whether the risk of re-identifiability still exists after data privacy protection. Through these methods, the effectiveness of privacy protection can be comprehensively evaluated from different perspectives, and objective metrics can be provided through quantitative indicators to help decision-makers better understand the impact of the anonymization process and further optimize data processing workflows and privacy protection measures. For example, 0 represents no anonymization, and 1 represents complete anonymity. If the proportion of anonymous identifiers in a certain column of data reaches 80%, the score for that item is 0.8.

[0058] In this embodiment of the invention, a differential privacy loss function is calculated based on the data to be processed and the privacy-protected data to evaluate the anonymization effect and quantify the level of privacy protection. The differential privacy loss function includes, but is not limited to, cross-entropy or mean squared error, such as a privacy loss that increases linearly with the number of queries.

[0059] Using differential privacy loss functions can more accurately assess the effectiveness of anonymization because it can quantify the amount of privacy information leaked between different data sets. Furthermore, differential privacy loss functions can serve as an effective privacy protection engine, helping to monitor changes in privacy risks and privacy breaches during data processing. Finally, quantifying the level of privacy protection provides managers with intuitive and comparable data, aiding in the development of appropriate privacy protection measures and decisions.

[0060] Furthermore, this invention optimizes the differential privacy loss function based on the following information: a. Sensitivity Parameters: Sensitivity parameters for different data attributes can be dynamically defined based on data characteristics and privacy requirements. Adjusting these sensitivity parameters affects the calculation results of the loss function, thereby improving the anonymization effect.

[0061] b. Privacy Budget: A well-allocated privacy budget can balance privacy protection and data availability. By optimizing the allocation of the privacy budget, false alarm rates can be minimized or privacy requirements can be met.

[0062] c. Noise Addition Mechanism: Select the appropriate noise type (such as Laplace noise, Gaussian noise) and parameter settings. Adjust the noise amplitude according to needs to balance data accuracy and privacy protection.

[0063] Optimization strategies and methods include, but are not limited to, one of the following: a) Dynamically update the loss function using real-time data streams to ensure the evaluation results are effective in dynamic scenarios; b) Based on transfer learning: Optimize relevant loss function parameters in advance by leveraging prior knowledge and existing datasets; c) Design different differential privacy loss functions based on potential threat scenarios.

[0064] By optimizing the differential privacy loss function using the methods and strategies described above, the evaluation can be made more accurate and better suited to actual needs, achieving better results in balancing data availability and privacy protection performance.

[0065] In this embodiment of the invention, the data conversion process of the privacy protection model is optimized based on the results of the desensitization effect evaluation.

[0066] In one implementation, the data transformation process of the privacy protection model is optimized based on the results of the desensitization effect evaluation and the adaptive data transformation factor.

[0067] Optionally, the adaptive data transformation factor is typically a numerical value representing a certain attribute or parameter. It can be a number, coefficient, weight, or other form of quantification. This value can represent different meanings or functions depending on the system design, and is used to adjust the specific methods of data processing. Therefore, the adaptive data transformation factor plays a role in dynamically adjusting the data transformation process in the privacy protection model, flexibly adjusting the data processing methods through changes in its value to adapt to different needs and scenarios.

[0068] An adaptive data transformation factor is a variable that can be dynamically adjusted based on real-time privacy requirements and data characteristics. This factor plays a role in regulating the data transformation process within a privacy protection model, making it more flexible and intelligent. It can autonomously adjust the degree or method of anonymization or encryption based on data requirements and privacy levels in different scenarios.

[0069] Optionally, the adaptive data transformation factor can flexibly adjust the data transformation process according to user needs and data characteristics, achieving more personalized privacy protection. The adaptive data transformation factor can better balance the relationship between data utility and privacy protection, improving overall system efficiency. It reduces reliance on manual intervention, lowers operational costs, and enhances the system's self-learning and optimization capabilities.

[0070] In the above implementation, the data transformation process of the privacy protection model is optimized based on the results of the desensitization effect evaluation and the adaptive data transformation factor. This includes adjusting the degree of data desensitization according to the privacy protection level and the adaptive factor. While ensuring data usability, aggregation technology is used to protect individual data to the greatest extent. The adaptive data transformation factor is used to dynamically adjust the algorithm parameters to control the data processing process, ensuring that data analysis goals are achieved while protecting privacy. For example, if the desensitization effect evaluation result is 0.8, combined with an adaptive data transformation factor of 0.2, the anonymization standards and rules are emphasized more during the data transformation process, reducing the possibility of sensitive information leakage.

[0071] Furthermore, by monitoring changes in privacy protection levels, the adaptive data transformation factor is automatically adjusted to improve the model's flexibility and robustness. Based on feedback from actual results, the privacy protection model parameters and data processing strategies are adjusted in real time, continuously iterating and optimizing the model.

[0072] By regularly evaluating the effectiveness of data anonymization and optimizing it with adaptive data transformation factors, the model achieves adaptability and dynamic adjustment. This ensures a high level of protection for sensitive data while retaining as much useful information as possible, thus balancing privacy protection and data value. This approach combines the interaction of weight values, privacy protection levels, and the privacy protection model to maximize data utilization while protecting user privacy. It constitutes a flexible and adaptive privacy protection framework that helps maintain data security and privacy control in data application environments.

[0073] In an exemplary embodiment of the present invention, before obtaining the data to be processed and the user level tag corresponding to the data to be processed, the method further includes: Obtain user data from multiple users; Based on the user data, multiple users are clustered to obtain multiple user clusters; The user level label for each user cluster is determined based on the user data.

[0074] In this embodiment of the invention, big data users are clustered, and based on the clustering results, each big data user is assigned a user level label, which represents the user's sensitivity to the risk of privacy information release / leakage. Big data users refer to individuals or users within a dataset whose behaviors, attributes, or characteristics are collected and analyzed. These users may involve large-scale data collected from fields such as social media, e-commerce, and health, typically including personal information, consumption habits, interaction records, browsing history, and purchase records.

[0075] Clustering is an unsupervised learning method that groups individuals with similar characteristics in a dataset together. Clustering large datasets of users facilitates the automatic identification and grouping of users sharing similar features. The final clustering results clearly define the similarities and differences between user groups. Optional clustering algorithms include, but are not limited to, K-means clustering, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), and hierarchical clustering.

[0076] Optionally, users can be divided into three distinct user clusters through clustering. Each cluster corresponds to a user level label, which includes a first-level label, a second-level label, and a third-level label, with the sensitivity to privacy increasing progressively among the three levels. These labels allow for a better understanding of different users' levels of concern regarding privacy issues, enabling the implementation of appropriate privacy protection measures.

[0077] In one embodiment of the present invention, user level labels are adjusted in real time based on changes in user behavior patterns. Changes in user behavior patterns refer to changes in a user's behavior, preferences, interests, etc., at different times or in different contexts. These changes may be influenced by factors such as the external environment, personal experiences, and social impacts. In privacy protection and data processing, considering changes in user behavior patterns can help to more accurately assess a user's privacy sensitivity, more intelligently allocate user levels, and provide targeted privacy protection measures.

[0078] In an exemplary embodiment of the present invention, determining the privacy protection level and privacy protection model of the data to be processed based on the data type, the feature attribute, and the user level tag includes: If the data type is numerical, the feature attribute is a non-sensitive feature, the user level label is a first level label, the privacy protection level is determined to be the first protection level, and the privacy protection model is at least one of the data de-identification model, data masking model, and data aggregation model; If the data type is numeric, text, or date, the feature attribute is a non-sensitive feature, the user level label is a second-level label or a third-level label, the privacy protection level is determined to be the second protection level, and the privacy protection model is at least one of homomorphic encryption, data desensitization encryption, or lexicalization. If the feature attribute is a sensitive feature, the privacy protection level is determined to be the third protection level, and the privacy protection model is at least one of the differential privacy model, secure multi-party computation model, and anonymization model.

[0079] In this embodiment of the invention, the privacy protection level and privacy protection model are determined based on user level, data type, and feature attributes.

[0080] Specifically, if the data type is numerical, the feature attribute is a non-sensitive feature, the user level label is a first-level label, and the privacy protection level is set to a lower first-level protection level, it may focus on data de-identification processing. Therefore, at least one of the privacy protection technologies such as data de-identification, data masking, or data aggregation can be considered as the privacy protection model. This can effectively protect data privacy, while retaining the basic characteristics of the data, reducing the risk of data association, and providing a certain degree of protection for user privacy.

[0081] If the data type is numeric, text, or date-based, the feature attribute is non-sensitive, and the user level label is a second-level or third-level label, then setting the privacy protection level to the medium second-level protection provides relatively strict protection for the data to be processed. Therefore, at least one of the encryption technologies such as homomorphic encryption, data anonymization encryption, or tokenization can be selected as the privacy protection model to ensure the confidentiality and integrity of the data to be processed and reduce the risk of data leakage. Furthermore, for the data to be processed at the second-level protection level, access permissions can be appropriately controlled to prevent unauthorized access.

[0082] If the feature attributes are sensitive, a high level of privacy protection is required. The privacy protection level should be set to the highest level, the third level, to ensure privacy security. Therefore, the privacy protection model should employ at least one of the advanced protection models, such as differential privacy technology, secure multi-party computation, and anonymization methods, to ensure that the data is not parsed or associated with personal identity, thereby effectively protecting the user's sensitive information, providing strict privacy protection, and ensuring that the data analysis results have reasonable accuracy, preventing the data from being re-identified or inferred.

[0083] The above implementation methods take into account user level tags, data types, and characteristic attributes to determine the corresponding privacy protection levels and privacy protection models, thereby effectively protecting the privacy information of different types of data and improving data privacy security.

[0084] This invention comprehensively considers different user levels, data types, and characteristic attributes, and combines corresponding privacy protection levels and model selection. Each implementation method, by employing specific privacy protection technologies, can effectively protect data privacy, ensuring data security and privacy while maintaining data validity and availability as much as possible.

[0085] In this embodiment of the invention, by dividing user level tags into low, medium, and high levels, the needs of users for data privacy are considered more meticulously, realizing a personalized privacy protection strategy. Targeted privacy protection measures are provided for different data types (numerical, text, and date) to ensure that data is properly protected, thereby improving data security. By combining sensitive and non-sensitive features and comprehensively considering the impact of data attributes on privacy protection, an appropriate privacy protection model is selected, making protection more comprehensive and effective.

[0086] This invention establishes a systematic privacy protection framework for different privacy protection levels and model selection schemes in various scenarios, in order to better address various privacy leakage risks. Through these improvements, the proposed technical solution can better meet the privacy protection needs of different user groups and data types, and improve the flexibility and security level of data privacy management.

[0087] In an exemplary embodiment of the present invention, the data to be processed is processed according to the privacy protection level and the privacy protection model in order to achieve data desensitization; Privacy protection models maintain the usability of data while protecting data privacy by introducing noise and generating appropriate perturbations to protect sensitive information.

[0088] Optional privacy protection models include, but are not limited to: 1) Data Anonymization: This involves removing or hiding sensitive information through data generalization, data perturbation, and other methods to protect data privacy.

[0089] 2) Differential Privacy: A method that provides encrypted data privacy by adding noise to query results to prevent the personal identification of the data.

[0090] 3) Encryption: Using symmetric / asymmetric encryption algorithms to ensure the security of data during transmission and storage.

[0091] 4) Attribute Transformation: Transform or replace sensitive attributes to preserve data integrity while protecting data privacy.

[0092] 5) Uncertainty Data Processing: By increasing the uncertainty of data, the potential sensitivity of the data is reduced.

[0093] Sensitive data is anonymized using personalized processing methods based on different privacy protection levels and models, improving adaptability and flexibility to various situations. Different privacy protection models allow for the selection of data anonymization techniques more suitable for specific scenarios, enhancing both the effectiveness of anonymization and security.

[0094] In one embodiment of the present invention, the privacy protection model combines deep neural networks and differential privacy mechanisms to achieve dynamic perturbation desensitization by learning the representation of sensitive data and introducing a privacy-aware mechanism during the embedding process.

[0095] For example, after generating privacy representation features through a privacy protection model, the obtained privacy representation can be perturbed according to the principle of differential privacy to ensure the protection of sensitive information. For example, differential privacy can be used to inject noise, such as Laplace noise or Gaussian noise, into the representation features.

[0096] The privacy-preserving model includes a privacy-aware embedding layer and differential privacy constraints, which are described below: Privacy-aware embedding layer: Transforms individual data into privacy-preserving representation features, as shown in the following formula: ; Where x represents the input data received by the model, E(x) is the output obtained after performing a certain transformation on the input (x), and f(⋅) is the activation function. and These represent the weights and biases of the embedding layer.

[0097] Differential privacy constraints: Differential privacy protection is achieved by combining embedded representations and employing differential privacy loss. The function introduces new parameters λ and δ, and the formula is as follows: ; represents the differential privacy loss function, used to measure the model's performance while adhering to differential privacy constraints; n represents the total number of samples; and Let represent the original data and the perturbated data of the (i)th data point, respectively; and λ and δ represent the probability distribution of the corresponding data after processing by the activation function (f(⋅)); λ is a parameter that balances privacy and data quality, and δ is used to control the protection level of differential privacy in order to dynamically adjust the strength of differential privacy constraints and improve privacy adaptability.

[0098] The privacy-preserving model considers differential privacy constraints during training. It optimizes the loss function to minimize both prediction error and differential privacy loss, and dynamically adjusts the parameter δ to balance the needs of privacy protection and data usability.

[0099] In this embodiment of the invention, the integrated application of steps such as user clustering, user level assignment, privacy protection processing, evaluation, and quantification avoids the problem of poor model accuracy in the aforementioned federated learning. First, by clustering users and assigning user level labels, not only can privacy protection strategies be better managed, but differentiated processing methods can also be adopted to address the uneven distribution of data among participants and insufficient data volume. Second, determining the privacy protection level and privacy protection model based on user level labels, data characteristics, and feature attributes helps avoid increased communication overhead caused by frequent exchange of raw data and improves data privacy. Furthermore, evaluating the desensitization effect and quantifying the privacy protection level indicators can effectively monitor the effectiveness of data protection and potential risks, providing data support for further improvements. Finally, optimizing the model transformation process based on the privacy protection level helps to improve processing efficiency and promote model accuracy and performance while maintaining data privacy. The process of this proposal can effectively coordinate privacy protection and data processing requirements, helping to address the challenge of declining model accuracy in federated learning.

[0100] In this embodiment of the invention, by clustering users and assigning them user level labels, the level of privacy protection and processing strategies can be effectively adjusted to adapt to the needs and privacy requirements of different user groups. Secondly, determining the privacy protection level and model based on user level, data type, and characteristics helps to accurately select de-identification operations and avoid introducing excessive noise during the privacy protection process, thereby reducing the impact on data analysis results. Furthermore, evaluating the de-identification effect and quantifying the privacy protection level provides a basis for data improvement and monitors the effectiveness of privacy protection measures to ensure data availability and confidentiality. Finally, optimizing the model transformation process based on the privacy protection level, by adopting differentiated processing for different privacy protection needs, ensures data privacy while minimizing introduced disturbances to improve the accuracy of analysis results. In summary, this process comprehensively considers the balance between privacy protection and data analysis accuracy, and effectively avoids the problem of decreased accuracy of data analysis results caused by noise introduced by differential privacy through a systematic approach.

[0101] In this invention, privacy protection operations typically reduce data redundancy and dimensionality, thereby reducing data volume and storage space usage. Effective privacy protection operations can alleviate the storage burden caused by data encryption to some extent. Secondly, setting privacy protection levels based on user level, data type, and characteristic attributes helps clarify which data requires what level of protection, thus avoiding unnecessary encryption or protection measures and reducing storage resource consumption. Finally, evaluating the anonymization effect and quantifying the privacy protection level allows for identifying areas for improvement during data processing, enabling more effective management of storage resources and protection of private data. This invention can effectively control the storage resources required for data processing while protecting user privacy, avoiding excessive pressure on storage requirements.

[0102] In this invention, users are first categorized and assigned privacy levels, which reduces unnecessary information exchange and redundant computation. That is, communication occurs only when necessary, thereby reducing overall communication overhead. Privacy protection levels and models are determined based on data type and characteristics, and data anonymization is effectively performed, reducing unnecessary information transmission. Optimizing data processing efficiency also reduces communication overhead. Finally, the anonymization effect is evaluated and the privacy protection level is quantified. Appropriate protection measures and models are customized for different privacy protection needs, ensuring that communication overhead is minimized while maintaining data privacy. This invention effectively manages big data while protecting user privacy, avoiding unnecessary information exchange and minimizing communication overhead. The entire process systematically considers the balance between privacy protection and communication efficiency, thus helping to solve the problem of increased communication overhead in secure multi-party computation.

[0103] The privacy data processing apparatus provided by the present invention will be described below. The privacy data processing apparatus described below can be referred to in correspondence with the privacy data processing method described above. It should be noted that the apparatus provided in the embodiments below and the method provided in the embodiments above belong to the same concept, and the specific manner in which each module and unit performs its operation has been described in detail in the method embodiments, and will not be repeated here.

[0104] In one exemplary embodiment of the present invention, please refer to Figure 2 , Figure 2 This is a privacy data processing apparatus according to an exemplary embodiment, comprising the following modules.

[0105] The first determining module 210 is configured to acquire the data to be processed and the user level tag corresponding to the data to be processed, and to determine the data type and feature attributes of the data to be processed. The second determining module 220 is configured to determine the privacy protection level and privacy protection model of the data to be processed based on the data type, the feature attributes, and the user level label. The privacy protection processing module 230 is configured to perform privacy protection processing on the data to be processed according to the privacy protection level and the privacy protection model.

[0106] In an exemplary embodiment of the present invention, the first determining module 210 includes: The first determining submodule is configured to determine the data type of the data to be processed. A submodule is configured to preprocess the data to be processed and construct a feature vector based on the preprocessed data to be processed. The input submodule is configured to input the feature vector into the sensitive feature recognition model to obtain the feature attributes of the data to be processed.

[0107] In an exemplary embodiment of the present invention, the privacy protection processing module 230 includes: The second determining submodule is configured to determine the de-identification parameters and data sampling parameters corresponding to the privacy protection model based on the privacy protection level; The privacy protection processing submodule is configured to perform privacy protection processing on the data to be processed based on the desensitization parameters, data sampling parameters, and the privacy protection model.

[0108] In an exemplary embodiment of the present invention, the second determining module 220 includes: The evaluation submodule is configured to assess the risk level of the data to be processed based on a privacy risk assessment model; The third determination submodule is configured to determine the privacy protection level and privacy protection model of the data to be processed based on the risk level, the data type, the feature attributes, and the user level label.

[0109] In an exemplary embodiment of the present invention, the privacy data processing apparatus further includes: The evaluation module is configured to evaluate the desensitization effect based on the data to be processed and the data after privacy protection. The optimization module is configured to optimize the privacy protection model based on the results of the desensitization effect evaluation.

[0110] In an exemplary embodiment of the present invention, the privacy data processing apparatus further includes: The acquisition module is configured to acquire user data from multiple users. The clustering module is configured to perform clustering processing on multiple users based on the user data to obtain multiple user clusters; The third determining module is configured to determine the user level label for each user cluster based on the user data.

[0111] In an exemplary embodiment of the present invention, the privacy protection processing module 230 includes: The fourth determination submodule is configured to determine the privacy protection level as the first protection level if the data type is numerical, the feature attribute is a non-sensitive feature, the user level label is a first level label, and the privacy protection model is at least one of the data de-identification model, data masking model, and data aggregation model. The fifth determination submodule is configured to determine the privacy protection level as the second protection level if the data type is numeric, text, or date, the feature attribute is a non-sensitive feature, the user level label is a second level label or a third level label, and the privacy protection model is at least one of homomorphic encryption, data desensitization encryption, or lexicalization. The sixth determination submodule is configured to determine the privacy protection level as the third protection level if the feature attribute is a sensitive feature, and the privacy protection model is at least one of differential privacy model, secure multi-party computation model, and anonymization model.

[0112] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include a processor 310, a communications interface 320, a memory 330, and a communication bus 340, wherein the processor 310, communications interface 320, and memory 330 communicate with each other via the communication bus 340. The processor 310 can invoke logical instructions in the memory 330 to execute a privacy data processing method, which includes: acquiring data to be processed and a user level tag corresponding to the data to be processed, and determining the data type and characteristic attributes of the data to be processed; Based on the data type, the feature attributes, and the user level label, determine the privacy protection level and privacy protection model of the data to be processed; The data to be processed is subjected to privacy protection processing based on the privacy protection level and the privacy protection model.

[0113] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0114] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, the computer program being executed by a processor, the computer being able to execute the privacy data processing method provided by the above methods, the method including: obtaining data to be processed and user level tags corresponding to the data to be processed, and determining the data type and characteristic attributes of the data to be processed; Based on the data type, the feature attributes, and the user level label, determine the privacy protection level and privacy protection model of the data to be processed; The data to be processed is subjected to privacy protection processing based on the privacy protection level and the privacy protection model.

[0115] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the privacy data processing method provided by the above methods, the method comprising: acquiring data to be processed and a user level tag corresponding to the data to be processed, and determining the data type and characteristic attributes of the data to be processed; Based on the data type, the feature attributes, and the user level label, determine the privacy protection level and privacy protection model of the data to be processed; The data to be processed is subjected to privacy protection processing based on the privacy protection level and the privacy protection model.

[0116] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0117] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing privacy data, characterized in that, include: Obtain the data to be processed and the user level tags corresponding to the data to be processed, and determine the data type and characteristic attributes of the data to be processed; Based on the data type, the feature attributes, and the user level label, determine the privacy protection level and privacy protection model of the data to be processed; The data to be processed is subjected to privacy protection processing based on the privacy protection level and the privacy protection model.

2. The privacy data processing method according to claim 1, characterized in that, Determining the data type and characteristic attributes of the data to be processed includes: Determine the data type of the data to be processed; The data to be processed is preprocessed, and a feature vector is constructed based on the preprocessed data to be processed. The feature vector is input into the sensitive feature recognition model to obtain the feature attributes of the data to be processed.

3. The privacy data processing method according to claim 1, characterized in that, The privacy protection processing of the data to be processed according to the privacy protection level and the privacy protection model includes: The desensitization parameters and data sampling parameters corresponding to the privacy protection model are determined based on the privacy protection level; Based on the desensitization parameters, data sampling parameters, and the privacy protection model, the data to be processed is subjected to privacy protection processing.

4. The privacy data processing method according to claim 1, characterized in that, The step of determining the privacy protection level and privacy protection model of the data to be processed based on the data type, the feature attributes, and the user level label includes: The risk level of the data to be processed is assessed based on a privacy risk assessment model. Based on the risk level, the data type, the characteristic attributes, and the user level label, determine the privacy protection level and privacy protection model of the data to be processed.

5. The privacy data processing method according to claim 1, characterized in that, After performing privacy protection processing on the data to be processed according to the privacy protection level and the privacy protection model, the method further includes: Based on the data to be processed and the data after privacy protection, the desensitization effect is evaluated; The privacy protection model is optimized based on the results of the desensitization effect evaluation.

6. The privacy data processing method according to claim 1, characterized in that, Before obtaining the data to be processed and the user level tag corresponding to the data to be processed, the method further includes: Obtain user data from multiple users; Based on the user data, multiple users are clustered to obtain multiple user clusters; The user level label for each user cluster is determined based on the user data.

7. The privacy data processing method according to any one of claims 1 to 6, characterized in that, The step of determining the privacy protection level and privacy protection model of the data to be processed based on the data type, the feature attributes, and the user level label includes: If the data type is numerical, the feature attribute is a non-sensitive feature, the user level label is a first level label, the privacy protection level is determined to be the first protection level, and the privacy protection model is at least one of the data de-identification model, data masking model, and data aggregation model; If the data type is numeric, text, or date, the feature attribute is a non-sensitive feature, the user level label is a second-level label or a third-level label, the privacy protection level is determined to be the second protection level, and the privacy protection model is at least one of homomorphic encryption, data desensitization encryption, or lexicalization. If the feature attribute is a sensitive feature, the privacy protection level is determined to be the third protection level, and the privacy protection model is at least one of the differential privacy model, secure multi-party computation model, and anonymization model.

8. A privacy data processing device, characterized in that, include: The first determining module is configured to obtain the data to be processed and the user level tag corresponding to the data to be processed, and to determine the data type and feature attributes of the data to be processed; The second determining module is configured to determine the privacy protection level and privacy protection model of the data to be processed based on the data type, the feature attributes, and the user level label. The privacy protection processing module is configured to perform privacy protection processing on the data to be processed according to the privacy protection level and the privacy protection model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the privacy data processing method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the privacy data processing method as described in any one of claims 1 to 7.