Big data privacy processing method and device, computer device and storage medium
By preprocessing big data, setting privacy processing rules, and merging and integrating them, the privacy leakage risk in existing big data privacy processing technologies has been resolved, and the protection and utilization of data security and privacy have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN EWARE INFORMATION TECH CO LTD
- Filing Date
- 2024-11-27
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, big data privacy processing still faces the risk of privacy leakage after encryption technology is decrypted, and de-identification technology may re-obtain the original information through the analysis of multiple related information during data processing, thus failing to effectively protect personal privacy.
By acquiring big data, preprocessing it to extract private data, setting privacy processing rules such as encryption, desensitization, generalization, masking, and anonymization, merging and integrating datasets, and monitoring and managing them, we can ensure that the data meets business needs and privacy protection requirements.
This approach achieves the goal of protecting the security of privacy data while fully leveraging its value, effectively safeguarding both the security and privacy of privacy data.
Smart Images

Figure CN122113147A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data technology, and in particular to a big data privacy processing method, apparatus, computer equipment, and storage medium. Background Technology
[0002] Big data privacy processing refers to a series of operations involving the collection, storage, management, and sharing of personal data with third parties, aimed at protecting personal privacy. With the rapid development of the digital age, the issue of personal data privacy protection has become increasingly prominent, making big data privacy processing a crucial topic in the field of data protection.
[0003] Current technologies for big data privacy processing face several significant shortcomings. For example, the limitations of existing technologies are a major issue. Although data protection technologies such as encryption and anonymization are widely used, many unresolved problems remain. For instance, while encryption effectively protects data security during transmission and storage, privacy remains at risk once data is decrypted. Furthermore, while anonymization can render sensitive information irreversible, the original information can still be retrieved through the analysis and integration of multiple related pieces of information during data processing. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a big data privacy processing method, which employs the following technical solution, including:
[0005] Acquiring big data;
[0006] The big data is preprocessed to extract private data from the preprocessed big data;
[0007] Set privacy processing rules for the privacy data, and perform privacy protection processing on the privacy data according to the privacy processing rules;
[0008] The privacy-protected data is merged and integrated to form a new dataset;
[0009] The merged dataset is monitored and managed to meet business needs and privacy protection requirements.
[0010] Preferably, the step of acquiring big data specifically includes:
[0011] Determine the data acquisition method;
[0012] Data is collected according to the determined data collection method, and the collected data is big data.
[0013] Preferably, the step of preprocessing the big data and extracting privacy data from the preprocessed big data specifically includes:
[0014] The big data is cleaned;
[0015] Feature extraction is performed on the cleaned big data;
[0016] The machine learning algorithm is trained using the feature vectors of the training samples, enabling it to identify personal privacy information.
[0017] The trained machine learning model is used to identify unknown test data and extract the private data contained therein.
[0018] Preferably, the step of setting privacy processing rules for the privacy data and performing privacy protection processing on the privacy data according to the privacy processing rules specifically includes:
[0019] Set encryption, desensitization, generalization, blocking, and anonymization rules for the aforementioned privacy data;
[0020] The privacy data is encrypted, desensitized, generalized, blocked, and anonymized according to the rules of encryption, desensitization, generalization, blocking, and anonymization.
[0021] Preferably, the step of merging and integrating the privacy-protected data to form a new dataset specifically includes:
[0022] Merge data that has undergone privacy protection processing;
[0023] Connect data that has undergone privacy protection processing;
[0024] Integrate, correlate, transform, and analyze data that has undergone privacy protection processing.
[0025] Preferably, the step of monitoring and managing the fused dataset to ensure it meets business needs and privacy protection requirements specifically includes:
[0026] Perform a data consistency check on the merged dataset;
[0027] Perform a data integrity check on the merged dataset;
[0028] A data quality assessment is performed on the merged dataset.
[0029] Preferably, after the step of monitoring and managing the fused dataset to ensure it meets business needs and privacy protection requirements, the method further includes:
[0030] Monitor the timeliness of the merged dataset;
[0031] Enhanced data privacy protection management is implemented for the merged dataset.
[0032] To address the aforementioned technical problems, the present invention also provides a big data privacy processing device, which employs the following technical solution, including:
[0033] The acquisition module is used to acquire big data;
[0034] The extraction module is used to preprocess the big data and extract private data from the preprocessed big data.
[0035] The processing module is used to set privacy processing rules for the privacy data and perform privacy protection processing on the privacy data according to the privacy processing rules;
[0036] The integration module is used to merge and integrate privacy-protected data to form a new dataset;
[0037] The monitoring module is used to monitor and manage the merged dataset to ensure it meets business needs and privacy protection requirements.
[0038] To address the aforementioned technical problems, the present invention also provides a computer device that employs the technical solution described below, comprising a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the aforementioned big data privacy processing method.
[0039] To address the aforementioned technical problems, the present invention also provides a computer-readable storage medium, which employs the technical solution described below. The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the aforementioned big data privacy processing method.
[0040] Compared with existing technologies, the present invention has the following main advantages: First, it acquires big data; then, it preprocesses the big data to extract private data from the preprocessed big data; next, it sets privacy processing rules for the private data and performs privacy protection processing on the private data according to the privacy processing rules; then, it merges and integrates the privacy-protected data to form a new dataset; finally, it monitors and manages the merged dataset to ensure it meets business needs and privacy protection requirements; it can integrate different private data according to actual needs, fully utilize the value of private data, and effectively protect the security and privacy of private data. Attached Figure Description
[0041] To more clearly illustrate the solutions in this invention, the accompanying drawings used in the description of the embodiments of this invention will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0042] Figure 1 This is a flowchart of an embodiment of the big data privacy processing method of the present invention;
[0043] Figure 2 This is a schematic diagram of the structure of one embodiment of the big data privacy processing device of the present invention;
[0044] Figure 3 This is a schematic diagram of the structure of an embodiment of the computer device of the present invention. Detailed Implementation
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein in the specification is for the purpose of describing particular embodiments only and is not intended to limit the invention; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings are used to distinguish different objects and not to describe a particular order.
[0046] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0047] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0048] It should be noted that the big data privacy processing method provided in the embodiments of the present invention is generally executed by a server / terminal device, and correspondingly, the big data privacy processing device is generally set in the server / terminal device.
[0049] It should be understood that the number of terminal devices, networks, and servers is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be used.
[0050] Example 1
[0051] Please refer to Figure 1 The flowchart illustrates an embodiment of the big data privacy processing method of the present invention. The big data privacy processing method includes the following steps:
[0052] Step S1: Obtain big data.
[0053] In this embodiment, the electronic device (e.g., a server / terminal device) on which the big data privacy processing method runs can receive big data privacy processing requests via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAXX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future-developed wireless connection methods.
[0054] In some optional implementations of this embodiment, step S1, obtaining big data, may further include the following steps:
[0055] S11, Determine the data acquisition method.
[0056] For example, by collecting data collection needs and objectives, we can analyze what information or insights can be obtained from the data, such as market trends, user behavior, and operational effectiveness. We need to determine the data sources and types, understanding the data's origin and format, including structured data (such as databases and spreadsheets) and unstructured data (such as social media, text, and images).
[0057] S12, Data collection is carried out according to the determined data collection method, and the collected data is big data.
[0058] By using web crawlers, we can simulate client-side network requests and automatically retrieve network information according to data collection needs and goals. Scripts or programs can be written to extract data from web pages in an automated manner.
[0059] Data can also be collected through open databases, directly obtaining data from the target database. This method is highly accurate and real-time, but requires database access permissions.
[0060] Data can also be collected through software interfaces, enabling interconnectivity between different software data through data interfaces provided by various software vendors. This can be achieved through custom API calls and data parsing.
[0061] Data can also be collected through software robots, which can collect data from both client software and websites, making them suitable for complex data collection scenarios.
[0062] Step S2: Preprocess the big data and extract the privacy data from the preprocessed big data.
[0063] In some optional implementations of this embodiment, step S2, preprocessing the big data and extracting privacy data from the preprocessed big data, may further include the following steps:
[0064] S21, Clean the big data.
[0065] Data cleaning is one of the key steps in big data preprocessing. Its purpose is to remove redundant, duplicate, and abnormal information from the data, thereby improving the quality and usability of the data.
[0066] Data deduplication and redundancy processing removes redundant and duplicate information from the data, avoiding biases in data analysis and decision-making.
[0067] Missing and outlier handling, filling in missing values, removing or correcting outliers, ensuring data integrity and accuracy.
[0068] Data formatting and standardization involves standardizing data from different formats into a consistent format, making it conform to predetermined specifications and facilitating subsequent analysis and processing.
[0069] S22, feature extraction is performed on the cleaned big data.
[0070] Feature extraction is the process of extracting useful feature values from preprocessed data to provide a foundation for subsequent privacy information identification.
[0071] Import training samples. Import known training samples that contain personal privacy information.
[0072] Preprocess the training samples to extract effective feature values. These feature values can be personal privacy information such as names, ID numbers, and phone numbers.
[0073] Feature vector formation involves summarizing the extracted feature values to form a feature vector, which provides input for the training of subsequent machine learning algorithms.
[0074] S23 uses the feature vectors of the training samples to train the machine learning algorithm, enabling it to identify personal privacy information.
[0075] Choosing a machine learning algorithm: Select a suitable machine learning algorithm based on business needs and data characteristics, such as support vector machine, neural network, etc.
[0076] Algorithm training involves inputting the feature vectors of training samples into a machine learning algorithm for training, enabling the algorithm to learn the features of personal privacy information.
[0077] Model evaluation involves assessing the performance of a trained machine learning model using methods such as cross-validation to ensure it can accurately identify personal privacy information.
[0078] Training machine learning algorithms requires the use of various machine learning tools and frameworks, such as TensorFlow and PyTorch. These tools and frameworks can help enterprises and organizations train and evaluate machine learning algorithms, improving the accuracy and performance of models.
[0079] S24. Use a trained machine learning model to identify unknown test data and extract the privacy data contained therein.
[0080] Privacy information identification uses a trained machine learning model to identify unknown test data and determine whether it contains personal privacy information.
[0081] Import test data, including collected unknown test data that may contain personal privacy information.
[0082] Preprocess the test data to extract effective feature values.
[0083] Feature vector formation involves summarizing the extracted feature values to form a feature vector.
[0084] Privacy information identification uses a trained machine learning model to identify and judge feature vectors to determine whether they contain personal privacy information.
[0085] Results are saved to the database for subsequent analysis and processing.
[0086] Step S3: Set privacy processing rules for the privacy data, and perform privacy protection processing on the privacy data according to the privacy processing rules.
[0087] In some optional implementations of this embodiment, step S3, setting privacy processing rules for the privacy data and performing privacy protection processing on the privacy data according to the privacy processing rules, may further include the following steps:
[0088] S31, set encryption, desensitization, generalization, blocking, and anonymization rules for the privacy data.
[0089] Encryption is the process of encoding sensitive data to prevent unauthorized decryption. Encryption algorithms include, but are not limited to, symmetric encryption (such as AES), asymmetric encryption (such as RSA), and hash functions (such as SHA). These algorithms provide strong protection for data, preventing it from being stolen or tampered with during storage and transmission.
[0090] Data anonymization is a process of replacing sensitive information with non-sensitive information. This includes character substitution, such as replacing specific characters or character sequences in sensitive data with other symbols (such as an asterisk "") or specific strings; and hash encryption, which uses hash functions to convert sensitive information into a fixed-length string. Anonymization can protect privacy while still allowing data analysis and processing.
[0091] Generalization protects privacy by reducing the precision of the data. For example, a specific age can be converted into an age range, or a precise location can be converted into a less precise location. This approach can reduce the risk of data leakage while maintaining data availability.
[0092] Masking involves hiding some or all of sensitive data. For example, when displaying a credit card number, the first six and last four digits can be retained, while the middle part is replaced with asterisks. This method preserves the basic structure of the data while hiding key information.
[0093] Anonymization refers to the process of processing personal information so that it cannot identify a specific natural person and cannot be restored. This includes replacing real names with pseudonyms and using technologies such as data exchange and data aggregation to ensure that data cannot be associated with the original data subject. Anonymization can be applied to scenarios such as data publishing and market research to protect user privacy.
[0094] S32, according to the encryption, desensitization, generalization, shielding and anonymization processing rules, the privacy data is encrypted, desensitized, generalized, shielded and anonymized.
[0095] For example, symmetric encryption such as AES and DES uses the same key for both encryption and decryption, making it suitable for rapid processing of large amounts of data. Asymmetric encryption, on the other hand, uses a public key and a private key; the public key encrypts the data, and the private key decrypts it. It is commonly used for digital signatures and key exchange.
[0096] For example, this can be achieved by replacing sensitive data (such as replacing names with "John Doe"), using fake data instead of real data, or using a hash function to convert the data into an irreversible string. Static data masking is completed before data is extracted and copied to a non-production environment, while dynamic data masking is performed in real time during the data query process.
[0097] For example, the date can be randomly replaced with a day within a year. This method ensures that the results of analyzing the perturbed data are consistent with the results of the original data, while protecting privacy.
[0098] For example, this can be achieved by deleting sensitive information, using fake data, or replacing real data with code.
[0099] Step S4 involves merging and integrating the privacy-protected data to form a new dataset.
[0100] In some optional implementations of this embodiment, step S4, merging and integrating the privacy-protected data to form a new dataset, may further include the following steps:
[0101] S41 merges the privacy-protected data.
[0102] In existing technologies, the location of privacy data is usually only passively detected. There is a lack of need to process privacy data from different locations by merging them according to the actual situation, which is not conducive to making full use of privacy data.
[0103] Data merging is the process of combining two or more datasets. In this embodiment, vertical merging and horizontal merging methods can be used to merge privacy-protected data, facilitating targeted analysis and processing of different types of privacy-protected data.
[0104] Vertical merging combines columns from two or more privacy-preserving datasets, such as merging privacy-preserving data from different time points into a single time series. This allows for tracking the movements of the same person at different points in time, enabling analysis and extraction of value from privacy-preserving data.
[0105] Horizontal merging involves concatenating rows from two or more privacy-preserving datasets, such as combining privacy-preserving data from different sources into a single composite dataset. This allows for tracking the movements of the same person at different locations and times, enabling analysis and extraction of value from their privacy-preserving data.
[0106] S42 connects the data that has undergone privacy protection processing.
[0107] Data joining is the process of connecting two datasets based on shared columns. For example, inner joins, left joins, and right joins can be used to join privacy-preserving data.
[0108] Inner joins only retain records where shared columns exist in both datasets.
[0109] A left join retains all records in the left dataset, even if there are no matching records in the right dataset.
[0110] A right join retains all records in the right dataset, even if there are no matching records in the left dataset.
[0111] S43 integrates, correlates, transforms, and analyzes data that has undergone privacy protection processing.
[0112] The goal of step S43 is to integrate data from multiple types and sources and discover correlations, trends, and patterns. The fusion of privacy-preserving data can be categorized into relationship-based fusion, content-based fusion, and structure-based fusion.
[0113] Relationship-based fusion is performed by integrating data based on logical relationships that have undergone privacy protection processing, such as linking data of different family members based on family relationships.
[0114] Content-based fusion involves integrating content from privacy-protected data, such as combining news reports from different sources to obtain comprehensive news information.
[0115] Structure-based fusion integrates data based on its structure after privacy protection processing, such as merging database tables of different formats to generate a unified database.
[0116] Step S5: Monitor and manage the fused dataset to ensure it meets business requirements and privacy protection requirements.
[0117] In some optional implementations of this embodiment, step S5, monitoring and managing the fused dataset to ensure it meets business requirements and privacy protection needs, may further include the following steps:
[0118] S51, Perform a data consistency check on the merged dataset.
[0119] Data consistency checks are a crucial step in ensuring data consistency across different data sources. By comparing identical fields across different data sources, inconsistencies can be identified and corrected accordingly.
[0120] S52, Perform a data integrity check on the merged dataset.
[0121] Data integrity checks are a crucial step in ensuring that all records in a dataset are complete and intact. By checking for missing values, outliers, and other abnormalities, the integrity and accuracy of the data can be guaranteed.
[0122] S53, perform a data quality assessment on the merged dataset.
[0123] Data quality assessment is the process of evaluating the overall quality of a dataset. By calculating metrics such as precision, recall, and F1 score, the quality of the dataset can be assessed, and corresponding optimizations can be made based on the assessment results.
[0124] In some optional implementations of this embodiment, after step S5, the electronic device may further perform the following steps:
[0125] S6, monitor the timeliness of the merged dataset.
[0126] Real-time monitoring updates the frequency of the merged data sources to ensure that the merged dataset reflects the latest business dynamics in a timely manner.
[0127] Data aging analysis involves periodically analyzing the timeliness of merged data to identify and process outdated or invalid data.
[0128] The purpose of step S6 is to ensure that the timeliness of the merged data is crucial to supporting rapidly changing business needs and helps improve the timeliness and effectiveness of decision-making.
[0129] S7, Strengthen data privacy protection management of the merged dataset.
[0130] For example, access control can be implemented through role definition. Different user roles and permission levels can be defined based on business needs and data sensitivity. By allocating permissions, corresponding data access and operation permissions can be assigned to different roles, ensuring the implementation of the principle of least privilege.
[0131] The purpose of access control is to effectively prevent unauthorized access and data leakage, and protect data privacy through strict access control.
[0132] For example, data auditing and monitoring can be performed by recording audit logs to document all access to and operations on the dataset, including information such as time, user, and operation type. By detecting abnormal behavior and utilizing machine learning, potential abnormal access and operation behaviors can be detected and alerted.
[0133] The purpose of data auditing and monitoring is to help to promptly identify and address the risks of data privacy breaches, providing strong technical support for data protection.
[0134] The beneficial effects of implementing this embodiment are as follows: First, big data is acquired; then, the big data is preprocessed to extract private data from the preprocessed big data; then, rules for privacy processing are set for the private data, and privacy protection processing is performed on the private data according to the privacy processing rules; then, the privacy-protected data is merged and integrated to form a new dataset; finally, the merged dataset is monitored and managed to meet business needs and privacy protection requirements; different private data can be integrated according to actual needs, making full use of the value of private data, and effectively protecting the security and privacy of private data.
[0135] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0136] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0137] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0138] Example 2
[0139] Further reference Figure 2 As a response to the above Figure 1 The present invention provides an embodiment of a big data privacy processing device, which is implemented in accordance with the method shown. Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0140] like Figure 2 As shown, the big data privacy processing device 60 described in this embodiment includes: an acquisition module 61, an extraction module 62, a processing module 63, an integration module 64, and a monitoring module 65. Wherein:
[0141] Module 61 is used to acquire big data;
[0142] Extraction module 62 is used to preprocess the big data and extract privacy data from the preprocessed big data;
[0143] Processing module 63 is used to set privacy processing rules for the privacy data and perform privacy protection processing on the privacy data according to the privacy processing rules;
[0144] Integration module 64 is used to merge and integrate privacy-protected data to form a new dataset;
[0145] The monitoring module 65 is used to monitor and manage the fused dataset to ensure it meets business needs and privacy protection requirements.
[0146] The beneficial effects of implementing this embodiment are as follows: First, big data is acquired; then, the big data is preprocessed to extract private data from the preprocessed big data; then, rules for privacy processing are set for the private data, and privacy protection processing is performed on the private data according to the privacy processing rules; then, the privacy-protected data is merged and integrated to form a new dataset; finally, the merged dataset is monitored and managed to meet business needs and privacy protection requirements; different private data can be integrated according to actual needs, making full use of the value of private data, and effectively protecting the security and privacy of private data.
[0147] Example 3
[0148] To address the aforementioned technical problems, embodiments of the present invention also provide a computer device. Please refer to [link / reference needed]. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.
[0149] The aforementioned computer device 7 includes a memory 71, a processor 72, and a network interface 73 that are interconnected via a system bus. It should be noted that the figure only shows a computer device 7 with components 71, 72, and 73; however, it should be understood that it is not required to implement all the shown components, and more or fewer components may be implemented instead. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0150] The aforementioned computer devices can be desktop computers, laptops, handheld computers, and cloud servers, among other computing devices. These devices can facilitate human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0151] The aforementioned memory 71 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the aforementioned memory 71 may be an internal storage unit of the aforementioned computer device 7, such as the hard disk or memory of the computer device 7. In other embodiments, the aforementioned memory 71 may also be an external storage device of the aforementioned computer device 7, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 7. Of course, the aforementioned memory 71 may also include both the internal storage unit and its external storage device of the aforementioned computer device 7. In this embodiment, the aforementioned memory 71 is typically used to store the operating system and various application software installed on the aforementioned computer device 7, such as computer-readable instructions for big data privacy processing methods. In addition, the aforementioned memory 71 can also be used to temporarily store various types of data that have been output or will be output.
[0152] In some embodiments, the processor 72 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 72 is typically used to control the overall operation of the computer device 7. In this embodiment, the processor 72 is used to execute computer-readable instructions stored in the memory 71 or to process data, such as executing computer-readable instructions for the big data privacy processing method described above.
[0153] The network interface 73 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 7 and other electronic devices.
[0154] The beneficial effects of implementing this embodiment are as follows: First, big data is acquired; then, the big data is preprocessed to extract private data from the preprocessed big data; then, rules for privacy processing are set for the private data, and privacy protection processing is performed on the private data according to the privacy processing rules; then, the privacy-protected data is merged and integrated to form a new dataset; finally, the merged dataset is monitored and managed to meet business needs and privacy protection requirements; different private data can be integrated according to actual needs, making full use of the value of private data, and effectively protecting the security and privacy of private data.
[0155] Example 4
[0156] The present invention also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the big data privacy processing method described above.
[0157] The beneficial effects of implementing this embodiment are as follows: First, big data is acquired; then, the big data is preprocessed to extract private data from the preprocessed big data; then, rules for privacy processing are set for the private data, and privacy protection processing is performed on the private data according to the privacy processing rules; then, the privacy-protected data is merged and integrated to form a new dataset; finally, the merged dataset is monitored and managed to meet business needs and privacy protection requirements; different private data can be integrated according to actual needs, making full use of the value of private data, and effectively protecting the security and privacy of private data.
[0158] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0159] Obviously, the embodiments described above are merely some embodiments of the present invention, not all embodiments. The accompanying drawings show preferred embodiments of the present invention, but do not limit the patent scope of the present invention. The present invention can be implemented in many different forms; rather, these embodiments are provided to provide a more thorough and complete understanding of the disclosure of the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the patent protection scope of this invention.
Claims
1. A method for handling big data privacy, characterized in that, Includes the following steps: Acquiring big data; The big data is preprocessed to extract private data from the preprocessed big data; Set privacy processing rules for the privacy data, and perform privacy protection processing on the privacy data according to the privacy processing rules; The privacy-protected data is merged and integrated to form a new dataset; The merged dataset is monitored and managed to meet business needs and privacy protection requirements.
2. The big data privacy processing method according to claim 1, characterized in that, The steps for acquiring big data specifically include: Determine the data acquisition method; Data is collected according to the determined data collection method, and the collected data is big data.
3. The big data privacy processing method according to claim 1, characterized in that, The step of preprocessing the big data and extracting privacy data from the preprocessed big data specifically includes: The big data is cleaned; Feature extraction is performed on the cleaned big data; The machine learning algorithm is trained using the feature vectors of the training samples, enabling it to identify personal privacy information. The trained machine learning model is used to identify unknown test data and extract the private data contained therein.
4. The big data privacy processing method according to claim 1, characterized in that, The steps of setting privacy processing rules for the privacy data and performing privacy protection processing on the privacy data according to the privacy processing rules specifically include: Set encryption, desensitization, generalization, blocking, and anonymization rules for the aforementioned privacy data; The privacy data is encrypted, desensitized, generalized, blocked, and anonymized according to the rules of encryption, desensitization, generalization, blocking, and anonymization.
5. The big data privacy processing method according to claim 1, characterized in that, The steps of merging and integrating the privacy-protected data to form a new dataset specifically include: Merge data that has undergone privacy protection processing; Connect data that has undergone privacy protection processing; Integrate, correlate, transform, and analyze data that has undergone privacy protection processing.
6. The big data privacy processing method according to claim 1, characterized in that, The steps for monitoring and managing the fused dataset to ensure it meets business needs and privacy protection requirements specifically include: Perform a data consistency check on the merged dataset; Perform a data integrity check on the merged dataset; A data quality assessment is performed on the merged dataset.
7. The big data privacy processing method according to claim 1, characterized in that, Following the step of monitoring and managing the fused dataset to ensure it meets business needs and privacy protection requirements, the following is also included: Monitor the timeliness of the merged dataset; Enhanced data privacy protection management is implemented for the merged dataset.
8. A big data privacy processing device, characterized in that, include: The acquisition module is used to acquire big data; The extraction module is used to preprocess the big data and extract private data from the preprocessed big data. The processing module is used to set privacy processing rules for the privacy data and perform privacy protection processing on the privacy data according to the privacy processing rules; The integration module is used to merge and integrate privacy-protected data to form a new dataset; The monitoring module is used to monitor and manage the merged dataset to ensure it meets business needs and privacy protection requirements.
9. A computer device comprising a memory and a processor, the memory storing computer-readable instructions, wherein the processor, when executing the computer-readable instructions, implements the steps of the big data privacy processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the big data privacy processing method as described in any one of claims 1 to 7.