A cloud service privacy data detection method and cloud server

By building an initialization model on a cloud server and training it on a local server, a global model is formed, which solves the possible privacy data leakage problem during the construction of the privacy data detection model in traditional technology, and improves the security level of privacy data protection.

CN119442334BActive Publication Date: 2025-05-13BEIJING TIMES XINWEI INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510044152.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-11
Publication Date
2025-05-13
Estimated Expiration
2045-01-11

AI Technical Summary

Technical Problem

When traditional privacy data detection technology builds a privacy data identification model in a cloud service environment, there are security risks of privacy data leakage.

Method used

By building an initialization model on a cloud server and sending it to multiple local servers for training, a global model is formed, reducing the risk of privacy data leakage.

Benefits of technology

On the basis of ensuring the accuracy of the privacy data identification model, this method significantly improves the security level of privacy data protection and reduces the possible risk of privacy data leakage in traditional technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119442334B_ABST
    Figure CN119442334B_ABST
Patent Text Reader

Abstract

This application provides a cloud service privacy data detection method and a cloud server, relating to the field of electronic digital data processing. The method includes: constructing an initial model for privacy data identification based on publicly available cloud data and a pre-set machine learning algorithm on the cloud server, and distributing it to multiple local servers; training the initial model on the local data stored on each local server to obtain a corresponding local machine learning model; aggregating the model parameters of the multiple local machine learning models onto the initial model according to the proportion of local data on each local server within the cloud server to obtain a global model; and using this global model to identify target cloud data obtained by the user to obtain a privacy data detection result. This method reduces the risk of privacy data leakage that may occur during the process of directly building a dataset and training a privacy data identification model in the cloud service environment, and improves the security level of privacy data protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of electronic digital data processing, and in particular to a cloud service privacy data detection method and a cloud server. Background Art

[0002] In today's digital age, data privacy has become an important global issue. With the rapid development and widespread application of cloud computing technology, a large amount of sensitive data is stored and processed on cloud servers. This increases the risk of data leakage, which requires effective privacy protection and monitoring of this data.

[0003] The privacy data detection technology in related technologies is mainly based on centralized model processing. In this method, the local data of each server needs to be transmitted to the central server for integration, and then the centralized model is trained based on the integrated cloud data for privacy analysis and detection of the subsequently collected data.

[0004] However, the construction process of the model used for privacy data detection and analysis in traditional privacy data detection technology requires sending the collected local data to the central server for information sharing, which leads to the security risk of privacy data leakage during the model training process. Summary of the invention

[0005] The present application provides a cloud service privacy data detection method and a cloud server, which are used to address the problem of privacy data leakage that may exist in the process of building a privacy data identification model in a cloud service environment using traditional technologies.

[0006] In a first aspect, the present application provides a cloud service privacy data detection method, which is applied to a cloud server, and the method includes:

[0007] Building an initialization model for identifying private data based on publicly available cloud data and a preset machine learning algorithm in a cloud server, the cloud server being in communication connection with one or more local servers;

[0008] The initialization model is sent to multiple local servers, and a data set is constructed based on the local data stored in each local server. The initialization model is trained respectively to obtain a local machine learning model corresponding to each local server, and the local server provides data support for the cloud server.

[0009] Obtain the data ratio of the local data corresponding to each local server in the cloud server, and the model parameters of the local machine learning model uploaded by each local server;

[0010] Aggregate the model parameters on the initialization model according to the data proportion to obtain a global model;

[0011] When it is detected that a user obtains target cloud data from the cloud server through the cloud service API interface, the global model is used to perform privacy data identification on the target cloud data to obtain a privacy data detection result.

[0012] Through the above embodiments, the cloud server constructs an initialization model and sends it to multiple local servers for training, so that the model data training process between different local servers is carried out separately. It only needs to aggregate the model training parameters of each local server to obtain a global model for privacy data identification of all cloud service data. This method reduces the risk of privacy data leakage that may be caused by the traditional technology of directly building a data set in the cloud service environment to train the privacy data identification model, and improves the security level of privacy data protection on the basis of ensuring the accuracy of the privacy data identification model.

[0013] In some embodiments, the step of constructing a data set based on the local data stored in each local server, training the initialization model respectively, and obtaining a local machine learning model corresponding to each local server specifically includes:

[0014] Obtain local data through a local server and perform preprocessing to obtain sample data. The preprocessing includes data cleaning, formatting, normalization, standardization, and missing value processing.

[0015] The sample data is divided into privacy data types to obtain label data;

[0016] Building an example data set based on the example data and the label data;

[0017] The sample data set is used to train and test the initialized model to obtain a local machine learning model for identifying and classifying private data.

[0018] Through the above embodiments, the cloud server controls each local server to pre-process the local data and divide the privacy data types, so that the trained local machine learning model can effectively learn the characteristics of the privacy data, and then identify and classify the privacy data, thereby improving the accuracy and efficiency of model data recognition.

[0019] In some embodiments, the step of aggregating the model parameters on the initialization model according to the data proportion to obtain a global model specifically includes:

[0020] Determine the weight of each local server according to the data proportion;

[0021] The weighted average of the model parameters is calculated according to the weight to obtain the corresponding parameters of the global model.

[0022] Through the above embodiment, the cloud server aggregates the model parameters uploaded by each local server according to the data proportion of each local server in the cloud server to build a global model. This method allocates weights by considering the importance and representativeness of each local server data, ensuring that the global model can more fairly and effectively reflect the characteristics of each local data during the integration process, improving the generalization ability of the model, and enabling the global model to maintain a high recognition effect on new data.

[0023] In some embodiments, the step of using the global model to identify the privacy data of the target cloud data to obtain the privacy data detection result specifically includes:

[0024] Determine one or more target local servers corresponding to the target cloud data, wherein the target cloud data is stored in the corresponding target local servers;

[0025] Using a local machine learning model corresponding to the target local server to perform privacy data identification on the target cloud data, to obtain a first identification result;

[0026] Using the global model to perform privacy data identification on the target cloud data, obtaining a second identification result;

[0027] The first recognition result and the second recognition result are integrated to obtain the privacy data detection result.

[0028] Through the above embodiments, the cloud server uses the global model and the local machine learning model to identify the privacy data of the target cloud data, and integrates the identification results of the two. This method combines the professionalism of the local model with the generalization ability of the global model to identify and verify privacy data from different angles, thereby greatly improving the accuracy and reliability of identification.

[0029] In some embodiments, after the step of integrating the first recognition result and the second recognition result to obtain the privacy data detection result, the step further includes:

[0030] Determine target private data where there is a deviation between the first recognition result and the second recognition result;

[0031] Re-label the target private data and send it to the corresponding local server, and update the sample data set of the local server;

[0032] Incrementally train the local machine learning model corresponding to the local server based on the updated example data set.

[0033] Through the above embodiments, the cloud server can continuously optimize the performance of the local machine learning model by identifying and re-labeling the deviation data and then performing incremental training, continuously improve the accuracy and efficiency of the model data processing, and ensure that the privacy protection measures are continuously improved over time.

[0034] In some embodiments, when it is detected that a user obtains target cloud data from the cloud server through the cloud service API interface, after the step of using the global model to perform privacy data identification on the target cloud data to obtain the privacy data detection result, the method further includes:

[0035] Summarize and classify the privacy detection results within a preset time period, and count the number of occurrences and distribution of different types of privacy data;

[0036] Generate detection reports based on the occurrence and distribution of different types of privacy data;

[0037] The risk assessment level corresponding to the test report is determined based on predefined risk assessment standards, and the risk assessment levels include high risk, medium risk and low risk.

[0038] Through the above embodiments, the cloud server's aggregation, classification and risk assessment of privacy data detection results can not only help users understand the distribution of privacy data, but also assess potential privacy leakage risks according to predefined standards, so that managers can timely understand the security status of the system, take corresponding protective measures, and effectively prevent and reduce the occurrence of data leakage incidents.

[0039] In some embodiments, after the step of determining the risk assessment level corresponding to the test report based on the predefined risk assessment criteria, the method further includes:

[0040] When the risk assessment level is detected as high risk, an alarm notification is triggered.

[0041] Through the above embodiment, when the cloud server detects that the risk assessment level is high risk, it will trigger an alarm notification, thereby achieving a rapid response to abnormal situations and ensuring that when data privacy faces high risks, relevant personnel can take measures quickly, thereby minimizing the possibility and impact of data leakage.

[0042] In a second aspect, the present application provides a cloud server, the cloud server comprising: one or more processors and a memory;

[0043] The memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions so that the cloud server can implement a cloud service privacy data detection method provided in the above embodiment, which will not be repeated here.

[0044] In a third aspect, the present application provides a computer-readable storage medium comprising instructions. When the instructions are executed on a cloud server, the cloud server executes a cloud service privacy data detection method provided in the above embodiment, which will not be repeated here.

[0045] In a fourth aspect, the present application provides a computer program product. When the computer program product runs on a cloud server, the cloud server can implement a cloud service privacy data detection method provided in the above embodiment, which will not be repeated here.

[0046] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0047] 1. The cloud server builds an initialization model and distributes it to multiple local servers for independent training. It only needs to aggregate the training parameters of these local servers to form a global model. This method significantly reduces the risk of privacy data leakage that may be caused by directly building and training the privacy data identification model in the cloud service environment. Compared with traditional technologies, this method greatly improves the security level of privacy data protection while ensuring the accuracy of the privacy data identification model.

[0048] 2. The cloud server uses both the global model and the local machine learning model trained on each local server to identify private data on the target cloud data, and integrates the identification results of the two. This dual identification mechanism uses the expertise of the local model and the generalization ability of the global model to verify and identify data from different angles, significantly improving the accuracy and reliability of private data identification. This method optimizes the data processing process and enhances the overall identification effect of the system by effectively combining the advantages of different models.

[0049] 3. The cloud server aggregates, classifies and conducts risk assessment on the privacy data detection results, which not only enables users to clearly understand the distribution of privacy data, but also comprehensively assesses the potential privacy leakage risks according to the preset risk assessment criteria. This analysis and assessment mechanism allows managers to obtain the security status of the system in a timely manner and take appropriate preventive measures, thereby effectively preventing and reducing the occurrence of data leakage incidents. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1It is a flowchart of a cloud service privacy data detection method in an embodiment of the present application;

[0051] Figure 2 It is another flowchart of a cloud service privacy data detection method in an embodiment of the present application;

[0052] Figure 3 It is a schematic diagram of a physical device structure of a cloud server in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to be used as limitations to the present application. As used in the specification and appended claims of the present application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include plural expressions, unless there is a clear indication to the contrary in the context. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations comprising one or more listed items.

[0054] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as suggesting or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, unless otherwise specified, "plurality" means two or more.

[0055] For ease of understanding, the following is a flow chart of the method provided in this implementation. Figure 1 The figure is a flow chart of a cloud service privacy data detection method in an embodiment of the present application.

[0056] S101. Construct an initialization model for privacy data identification based on the public cloud data in the cloud server and a preset machine learning algorithm.

[0057] The cloud server first collects and organizes public cloud data. This data can come from various public channels, such as public data sets, open API interfaces, web crawling, etc. The cloud server preprocesses this data, including but not limited to data cleaning, format conversion, feature extraction, etc., to ensure data quality and consistency.

[0058] Next, the cloud server selects to build an initialization model based on a preset machine learning algorithm. Among them, the optional machine learning algorithms include but are not limited to logistic regression, decision tree, support vector machine, neural network, etc. The selection of specific algorithms requires relevant technical personnel to consider factors such as data characteristics, model complexity, and computing resource limitations, which are not limited here. Of course, the cloud server can compare and evaluate different algorithms and select the algorithm with the best performance as the basis for the initialization model.

[0059] It should be noted that this step builds an initialization model based on public cloud data and a preset machine learning algorithm. This model will serve as the basis for subsequent distributed training and will eventually be aggregated into a global privacy data identification model.

[0060] S102: Send the initialization model to multiple local servers, build a data set based on the local data stored in each local server, train the initialization models respectively, and obtain a local machine learning model corresponding to each local server.

[0061] The cloud server distributes the initialization model constructed in step S101 to multiple local servers. Each local server has its own local data, which may come from different business systems, devices, sensors, etc. The local server collects, stores and manages the local data.

[0062] After obtaining the initialization model, each local server performs preprocessing, feature extraction, labeling and other steps on the local uncle to construct a training data set. Then, each local server uses the training data set it has constructed to train the initialization model. Specifically, the local server can use various optimization techniques, such as gradient descent, regularization, early stopping, etc., to improve the convergence speed and generalization ability of the model. At the same time, the local server can also adjust the hyperparameters of the model, such as learning rate, batch size, number of epochs, etc., to obtain the best performance. After training, each local server finally obtains its own local machine learning model. While inheriting the basic capabilities of the initialization model, these models also integrate the characteristics of local data and business needs. They can better identify and protect the privacy data processed by local servers.

[0063] S103. Obtain the data ratio of the local data corresponding to each local server in the cloud server, and the model parameters of the local machine learning model uploaded by each local server.

[0064] The cloud server obtains the data scale and distribution of each local server. Specifically, the cloud server can require each local server to report its own data volume, data type, data quality and other indicators. By aggregating this information, the cloud server can calculate the proportion of each local server's data in the entire system. Among them, the data proportion reflects the data contribution of each local server to the entire system. Local servers with more data usually have more reliable and representative models trained. Therefore, in the subsequent model aggregation process, the cloud server will give different weights to different local models according to the data proportion.

[0065] In addition, the cloud server obtains the parameters of the machine learning model trained by each local server. Model parameters include the model's architecture, weights, biases, and other information, which determine the model's performance and behavior. The local server uploads the model parameters so that the cloud server can understand and reproduce these local models.

[0066] It should be noted that the local server uploads model parameters instead of original data, which is a privacy protection strategy. In this way, the local server does not need to directly share the original data that may contain sensitive information, but only needs to share the trained model parameters. By aggregating these model parameters, the cloud server can obtain a global and more generalized machine learning model, while also reducing the risk of privacy data leakage. In actual operation, the local server can upload the model parameters to the cloud server through secure communication protocols such as HTTPS, SSL, etc. After receiving these parameters, the cloud server also verifies their integrity and validity to ensure that they have not been tampered with or damaged.

[0067] For example, suppose that the cloud server is connected to 100 local servers, each of which has local data of different sizes. Through data statistics, the cloud server finds that the data volume of 10 servers accounts for 80% of the total data volume, and the data volume of the remaining 90 servers accounts for only 20%. This shows that these 10 servers have a greater contribution in terms of data. At the same time, after completing local training, each server uploads its own model parameters to the cloud server. After verifying the validity of these parameters, the cloud server will weight the model parameters according to the data proportion of the server. The model parameters of servers with a large data proportion will obtain greater weights when aggregated, and have a greater influence on the final global model.

[0068] S104: Aggregate the model parameters on the initialization model according to the data proportion to obtain a global model.

[0069] The cloud server aggregates the model parameters collected from different local servers to generate a global machine learning model.

[0070] Specifically, the cloud server assigns weights to the model parameters of each local server according to the data share of each server. The server with a larger data share has a higher weight for its model parameters, allowing servers with larger data volumes and higher quality to play a greater role in the aggregation process, allowing the global model to learn and inherit more of their features. For example, assuming there are n local servers, and the data share of the i-th server is , the model parameters uploaded are Then, the calculation formula of the parameters of the global model is: .

[0071] In actual operation, cloud servers can use distributed computing frameworks such as Apache Spark and TensorFlowFederated to efficiently complete this aggregation process. These frameworks provide features such as parallel computing, fault tolerance, and communication optimization, which can greatly speed up the aggregation speed and improve the aggregation efficiency. The global model parameters obtained by aggregation will be applied to the initialization model to generate the final global machine learning model.

[0072] S105. When it is detected that the user obtains target cloud data from the cloud server through the cloud service API interface, the global model is used to perform privacy data identification on the target cloud data to obtain a privacy data detection result.

[0073] The cloud server uses the aggregated global machine learning model to identify and protect the privacy data of the cloud data requested by users.

[0074] Specifically, the cloud server continuously monitors and processes data requests from users. When a user requests certain cloud data from the cloud server through a standard cloud service API interface, such as RESTful API, gRPC, etc., the cloud server automatically triggers the privacy data identification process.

[0075] First, the cloud server performs necessary preprocessing on the cloud data requested by the user, such as format conversion, feature extraction, etc., to match the input requirements of the global model. Then, the cloud server inputs the preprocessed data into the global model and uses the model's reasoning ability to identify and classify the privacy of the data.

[0076] Furthermore, the global model conducts a comprehensive analysis and judgment of the input cloud data to identify the privacy information that may be contained therein, such as personal name, ID number, phone number, home address, etc. The model will give a privacy risk score for each data block and the type of privacy that may be involved. These will be returned to the cloud server as the results of privacy data detection.

[0077] After obtaining the privacy data detection results, the cloud server processes the results according to the preset privacy protection strategy. For the identified privacy data, the cloud server can take measures such as desensitization, encryption, and filtering to maximize the protection of user privacy while ensuring data availability. The processed data will be returned to the user through the API interface, completing the entire request-response process.

[0078] In actual applications, the cloud server can also continuously monitor and evaluate the performance of privacy data identification. By recording indicators such as the accuracy, recall rate, and response time of the model and generating performance reports regularly, if the performance of the model is detected to have degraded, the cloud server can trigger the retraining or optimization of the model to ensure the effectiveness and stability of privacy protection.

[0079] In the above embodiment, the cloud server constructs an initialization model and sends it to multiple local servers for training, so that the model data training process between different local servers is carried out separately. It only needs to aggregate the model training parameters of each local server to obtain a global model for privacy data identification of all cloud service data. This method reduces the risk of privacy data leakage that may be caused by the traditional technology of directly building a data set in the cloud service environment to train the privacy data identification model, and improves the security level of privacy data protection on the basis of ensuring the accuracy of the privacy data identification model.

[0080] The following is a more detailed description of the process of the method provided by this embodiment. Figure 2 As shown, it is another flow chart of a cloud service privacy data detection method in an embodiment of the present application.

[0081] S201. Acquire local data through a local server and pre-process it to obtain sample data.

[0082] The cloud server controls each local server through instructions to obtain the local data stored by each server. These local data may come from various business systems, device sensors, user input and other channels, showing different formats, quality and distribution characteristics. Then, by performing systematic data preprocessing (including but not limited to data cleaning, formatting, normalization, standardization and missing value processing, etc.) on the acquired local data, we can obtain clean and complete sample data in a unified format.

[0083] S202: Classify the sample data into privacy data types to obtain label data.

[0084] In actual operation, the division and labeling of privacy data can be carried out by combining rule matching and manual review. Specifically, the local server can write regular expressions or other matching rules based on predefined privacy data patterns, perform automatic preliminary screening of sample data, and mark the parts suspected of containing privacy data. Then, a dedicated security auditor will review the screening results, remove the misjudged parts, and supplement the missing privacy data to obtain the classification label corresponding to the sample data.

[0085] S203: construct an example data set based on the example data and the label data, and train and test the initialization model to obtain a local machine learning model for identifying and classifying private data.

[0086] After completing the preprocessing of the sample data and the labeling of the privacy data, the local server integrates the sample data and the corresponding labeled data to build a complete sample data set for training the privacy data identification model.

[0087] Specifically, the local server organizes and stores the sample data and the corresponding label data in a certain format and specification to form a structured sample data set. The optional data set formats include but are not limited to CSV, JSON, TFRecord, etc. The specific format needs to be selected according to the characteristics of the sample data and the needs of subsequent model training, which is not limited here. In addition, when it comes to sample data sets, the consistency of sample data and label data is required to ensure that each sample data has a corresponding privacy type label. At the same time, the entire data set must be divided into a training set and a test set in a certain ratio (such as 7:3), which are used for model training and verification, respectively.

[0088] After obtaining the sample data set, the local server starts training the initialization model. Specifically, the initialization model is loaded into the memory or GPU of the local server, and the model structure is adjusted and optimized as necessary according to the characteristics of the sample data set, such as adjusting the number of layers and parameters of the neural network. Then, the model is iteratively trained using the training set data. Each time a batch of sample data is randomly selected from the training set, the gap between the model's prediction results on this batch of data and the true label (loss function) is calculated, and the model parameters are updated using an optimization algorithm (such as gradient descent) to make the model's prediction results as close to the true label as possible. This process is repeated until the model's loss function reaches a lower level or the preset number of training rounds (Epoch) is reached.

[0089] In addition, the trained model is used to predict the example data in the test set, and the prediction results are compared with the true labels. The performance of the model can be evaluated by calculating the relationship between evaluation indicators such as accuracy, precision, recall rate and preset indicator thresholds, and the optimal model can be selected for subsequent identification and classification of privacy data.

[0090] In the above embodiment, the cloud server controls each local server to pre-process the local data and divide the privacy data types, so that the trained local machine learning model can effectively learn the characteristics of the privacy data, and then identify and classify the privacy data, thereby improving the accuracy and efficiency of model data recognition.

[0091] S204. Obtain the data ratio of the local data corresponding to each local server in the cloud server, and the model parameters of the local machine learning model uploaded by each local server.

[0092] This step is the same as step S103 and will not be described in detail here.

[0093] S205. Determine the weight of each local server based on the data proportion.

[0094] After obtaining the data share information of each local server, the cloud server further converts it into specific aggregation weights for use in subsequent global model aggregation.

[0095] Specifically, the cloud server constructs a weight distribution function so that servers with a large data share obtain higher weights, while servers with a small data share obtain lower weights, so that the aggregation results can better reflect the characteristics of key data. Optionally, the weight distribution function can be a linear function, that is, the data share is directly used as the weight; it can also be a piecewise function, that is, different weight intervals are divided according to the size of the data share, and different mapping functions are used in each interval. Of course, technical personnel can also construct other weight distribution functions in the cloud server according to actual needs, which is not limited here.

[0096] S206. Calculate the weighted average of the model parameters according to the weights to obtain the corresponding parameters of the global model.

[0097] After determining the aggregate weight of each local server, the cloud server uses a weighted average algorithm to fuse the model parameters uploaded by each local server according to the weight ratio to form the parameters of the global model. Specifically, the cloud server first unifies the model parameters uploaded by each local server into the same format and dimension to form a standardized parameter vector or matrix. Then, for each model parameter uploaded by a local server, the cloud server multiplies it by the aggregate weight of the local server to obtain the weight parameter product. Finally, the weight parameter products of all local servers are summed up to obtain the total weight parameter product. Then, the sum of the weight parameter products is divided by the sum of the weights to obtain the final global model parameters.

[0098] In the above embodiment, the cloud server aggregates the model parameters uploaded by each local server according to the proportion of each local server data in the cloud server to build a global model. This method allocates weights by considering the importance and representativeness of each local server data, ensuring that the global model can more fairly and effectively reflect the characteristics of each local data during the integration process, improving the generalization ability of the model, and enabling the global model to maintain a high recognition effect on new data.

[0099] S207: Determine one or more target local servers corresponding to the target cloud data.

[0100] After receiving a user request to access specific cloud data, the cloud server first determines which local server or servers the data is actually stored on.

[0101] Specifically, the cloud server can maintain a data distribution directory to record the correspondence between the unique identifier of each data block and the local server where it is located. When a user requests a certain data, the cloud server parses the request, extracts the identifier of the target data, and then queries the data distribution directory to find the target local server where the data is stored.

[0102] S208. Use the local machine learning model corresponding to the target local server to identify the privacy data of the target cloud data to obtain a first identification result.

[0103] After determining the local servers where the target cloud data is located, the cloud server sends the data identification task to these target local servers. Each target local server uses its own trained local machine learning model to identify and classify the target data.

[0104] Specifically, the target local server loads the cloud data into memory and performs necessary preprocessing, such as format conversion and feature extraction, to match the input requirements of the local machine learning model. The local server then inputs the preprocessed data into the local machine learning model, uses the model's reasoning ability to analyze the data, identifies the privacy information that may be contained therein, and gives the corresponding privacy type and risk level. These identification results form a privacy data report, which is returned to the cloud server as the first identification result.

[0105] S209: Use the global model to identify the privacy data of the target cloud data to obtain a second identification result.

[0106] After receiving the first recognition result returned by the target local server, the cloud server inputs the target cloud data into the global machine learning model, performs a second round of privacy data recognition, and obtains a second recognition result.

[0107] S210: Integrate the first recognition result and the second recognition result to obtain a privacy data detection result.

[0108] The cloud server aligns and compares the first and second recognition results. For each data segment, check whether the privacy type and risk level given by the local machine learning model and the global model are consistent. If they are completely consistent, it is determined that both models are very sure that the data belongs to a specific privacy type; if there is a difference, it is determined that the two models have different judgments on the data and further analysis is needed.

[0109] Furthermore, after detecting that the two models have different judgments on the cloud data, the cloud server determines the second recognition result output by the global model as the final privacy data detection result. Finally, the cloud server summarizes the final privacy data detection result into a formal privacy data detection report. The report lists all the privacy information fragments in the target cloud data, as well as their privacy type, risk level, and location details.

[0110] In the above embodiment, the cloud server uses the global model and the local machine learning model to identify the privacy data of the target cloud data, and integrates the identification results of the two. This method combines the professionalism of the local model with the generalization ability of the global model to identify and verify privacy data from different angles, thereby greatly improving the accuracy and reliability of identification.

[0111] S211. Determine target privacy data where there is a deviation between the first recognition result and the second recognition result.

[0112] When generating a privacy data detection report, the cloud server may detect that there is a significant deviation between the recognition results of the local machine learning model and the global model on certain privacy data. Some data fragments are marked as high-risk privacy by one model, but are considered normal data by another model.

[0113] Specifically, the cloud server can set a deviation threshold. When the difference between the two models in judging the privacy type or risk level of the same data segment exceeds this threshold, the segment will be marked as deviant data and recorded.

[0114] S212: Re-label the target private data and send it to the corresponding local server, and update the sample data set of the local server.

[0115] Specifically, the cloud server extracts the deviation data (target privacy data) and submits it, along with its contextual information, to a dedicated privacy review team. These reviewers are all experts who have received professional training and have a deep understanding of privacy compliance policies, and can make relatively authoritative judgments on privacy data.

[0116] Furthermore, the audit team will carefully analyze each deviation data, refer to the privacy protection standards of the industry and region, and the internal data security system of the enterprise, reconfirm its privacy attributes, and give a clear privacy type and risk level to obtain a new data label. The cloud server packages the updated data label together with the original data and sends it to the corresponding local server. After receiving this data, the local server merges it into its own sample data set.

[0117] S213. Incrementally train the local machine learning model corresponding to the local server based on the updated example data set.

[0118] The cloud server uses the newly labeled deviation data to optimize and upgrade the local machine learning model through incremental learning. Specifically, after the local server receives the updated data sent by the cloud server, it merges it into the original sample data set to form an enhanced data set containing more high-quality samples. Then, the local server uses this enhanced data set to retrain the original local machine learning model.

[0119] S214: Summarize and classify the privacy detection results within a preset time period, and count the number of occurrences and distribution of different types of privacy data.

[0120] The cloud server conducts comprehensive statistics and analysis on the privacy data detection results accumulated over a period of time (such as a day or a week), and mines the associated features and distribution patterns. Specifically, the cloud server extracts all privacy data detection records within a preset time period from the log database, including detailed information such as detection time, data source, privacy type, risk level, data content, etc. Then, the cloud server deduplicates and cleans these original records, removes some incomplete or abnormal data, and obtains a structured privacy data statistics table.

[0121] Next, the cloud server groups and summarizes these data according to different privacy types. The optional privacy data types include but are not limited to name, address, phone number, ID number, bank card number, etc. For each privacy type, the cloud server counts their total number of occurrences, proportions, and the number of distributions at different risk levels, forming a statistical sub-table of the privacy type. These sub-tables show the distribution characteristics and change trends of different privacy data.

[0122] S215: Generate a detection report based on the number of occurrences and distribution of different types of privacy data.

[0123] After completing the summary statistics of privacy data, the cloud server needs to organize these statistical results into a formal privacy data detection report for reference and review by senior managers or relevant departments.

[0124] Specifically, the cloud server generates the main content of the report based on the statistical sub-table of privacy types obtained in step S214. For each privacy type, the report shall list its total number of occurrences, week-on-week growth rate, maximum value, minimum value, average value, median and other basic statistics, and shall be accompanied by intuitive visualization charts, such as bar charts, pie charts, line charts, etc., to clearly show the scale distribution and change trend of the privacy type data.

[0125] S216. Determine the risk assessment level corresponding to the test report based on predefined risk assessment standards.

[0126] The cloud server extracts a set of key risk assessment indicators from the detection report, such as the total amount of high-risk privacy data, the proportion of high-risk privacy, the average daily growth rate of privacy data, the average handling time of privacy incidents, etc. These indicators reflect the quantity, distribution, change trend and response efficiency of privacy data in the system from different aspects, and can comprehensively measure the privacy protection level of the system.

[0127] Then, the cloud server substitutes the extracted risk assessment indicator values ​​into the predefined risk matrix for matching and judgment. Among them, the risk matrix is ​​a multi-indicator and multi-level risk qualitative standard. The horizontal axis of the matrix is ​​the various risk assessment indicators, and the vertical axis is the level division of risk severity, such as "high risk", "medium risk" and "low risk". Each cell in the matrix corresponds to a risk judgment condition, such as "high-risk privacy ratio>5%" corresponds to "high risk", and "average daily growth rate<1% and average disposal time<12 hours" corresponds to "low risk".

[0128] By determining whether each indicator value triggers the judgment conditions in the risk matrix one by one, the cloud server can determine the privacy risk assessment level corresponding to the detection report.

[0129] In the above embodiment, the cloud server's aggregation, classification and risk assessment of privacy data detection results can not only help users understand the distribution of privacy data, but also assess potential privacy leakage risks according to predefined standards, so that managers can timely understand the security status of the system, take corresponding protective measures, and effectively prevent and reduce the occurrence of data leakage incidents.

[0130] S217: When the risk assessment level is detected to be high risk, an alarm notification is triggered.

[0131] If the cloud server evaluates the risk assessment level corresponding to the current detection report as "high risk" in step S216, the cloud server can pop up a high-risk level warning sign on the monitoring interface, such as a flashing red exclamation mark or a pop-up message, to attract the attention of the on-duty personnel. At the same time, the cloud server can send warning text messages and emails to the managers in the pre-set address book to inform them of important information such as the generation time of the privacy data detection report, core statistical data, and risk level determination results, which are not limited here.

[0132] In the above embodiment, when the cloud server detects that the risk assessment level is high risk, it will trigger an alarm notification, thereby achieving a rapid response to abnormal situations and ensuring that when data privacy faces high risks, relevant personnel can take quick measures to minimize the possibility and impact of data leakage.

[0133] The cloud server in the embodiment of the present invention is an electronic device. Figure 3 A schematic diagram of the architecture of an electronic device suitable for implementing an embodiment of the present invention is shown.

[0134] It should be noted that Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0135] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructions (computer programs), or by controlling related hardware through instructions (computer programs), and the instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. The electronic device of this embodiment includes a storage medium and a processor, wherein a plurality of instructions are stored in the storage medium, and the instructions can be loaded by the processor to execute any step of the method provided in the embodiment of the present invention.

[0136] Specifically, the storage medium and the processor are electrically connected directly or indirectly to realize data transmission or interaction. For example, these elements can be electrically connected to each other through one or more signal lines. The storage medium stores computer execution instructions for implementing the data access control method, including at least one software function module that can be stored in the storage medium in the form of software or firmware. The processor executes various functional applications and data processing by running the software program and module stored in the storage medium. The storage medium can be, but is not limited to, random access storage medium (Random Access Memory, referred to as: RAM), read-only storage medium (Read Only Memory, referred to as: ROM), programmable read-only storage medium (Programmable Read-Only Memory, referred to as: PROM), erasable read-only storage medium (Erasable Programmable Read-Only Memory, referred to as: EPROM), electrically erasable read-only storage medium (Electric Erasable Programmable Read-Only Memory, referred to as: EEPROM), etc. Among them, the storage medium is used to store programs, and the processor executes the program after receiving the execution instruction.

[0137] Furthermore, the software programs and modules in the above-mentioned storage medium may also include an operating system, which may include various software components and / or drivers for managing system tasks (such as memory management, storage device control, power management, etc.), and may communicate with various hardware or software components to provide an operating environment for other software components. The processor may be an integrated circuit chip having signal processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc., which may implement or execute the various methods, steps, and logic flow diagrams disclosed in this embodiment. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0138] Since the instructions stored in the storage medium can execute the steps in any method provided in the embodiments of the present invention, the beneficial effects of any method provided in the embodiments of the present invention can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0139] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A cloud service privacy data detection method, applied to a cloud server, characterized in that: The method comprises: Building an initialization model for identifying private data based on public cloud data and a preset machine learning algorithm in a cloud server, wherein the cloud server is in communication with one or more local servers; The initialization model is sent to multiple local servers, and a data set is constructed based on the local data stored in each local server, and the initialization model is trained respectively to obtain a local machine learning model corresponding to each local server, and the local server provides data support for the cloud server; Obtain the data ratio of the local data corresponding to each of the local servers in the cloud server, and the model parameters of the local machine learning model uploaded by each of the local servers; Aggregating the model parameters on the initialization model according to the data proportion to obtain a global model; When it is detected that the user obtains target cloud data from the cloud server through the cloud service API interface, one or more target local servers corresponding to the target cloud data are determined, and the target cloud data is stored in the corresponding target local servers; Using a local machine learning model corresponding to the target local server to perform privacy data identification on the target cloud data, to obtain a first identification result; Using the global model to perform privacy data identification on the target cloud data, to obtain a second identification result; Integrate the first recognition result and the second recognition result to obtain the privacy data detection result; for each data segment, check whether the privacy type and risk level output by the local machine learning model and the global model are consistent; if they are completely consistent, determine that the target cloud data belongs to a specific privacy type; if there is a difference, determine the second recognition result as the final privacy data detection result; Determine target private data where there is a deviation between the first recognition result and the second recognition result; Re-labeling the target private data and sending it to the corresponding local server, and updating the example data set of the local server; Incrementally train the local machine learning model corresponding to the local server based on the updated example data set.

2. The method according to claim 1, characterized in that The step of constructing a data set based on the local data stored in each of the local servers, training the initialization models respectively, and obtaining a local machine learning model corresponding to each of the local servers specifically includes: Acquire local data through a local server and perform preprocessing to obtain sample data, wherein the preprocessing includes data cleaning, formatting, normalization, standardization and missing value processing; Classify the sample data into privacy data types to obtain label data; Building an example data set based on the example data and the label data; The initialized model is trained and tested using the example data set to obtain a local machine learning model for identifying and classifying private data.

3. The method according to claim 1, characterized in that: The step of aggregating the model parameters on the initialization model according to the data proportion to obtain a global model specifically includes: Determining a weight of each of the local servers according to the data proportion; The weighted average values ​​of the model parameters are calculated according to the weights to obtain the corresponding parameters of the global model.

4. The method according to claim 1, characterized in that: After the step of using the global model to identify the target cloud data for privacy when it is detected that the user obtains the target cloud data from the cloud server through the cloud service API interface to obtain the privacy data detection result, the method further includes: Summarize and classify the privacy data detection results within a preset time period, and count the number of occurrences and distribution of different types of privacy data; Generate detection reports based on the occurrence and distribution of different types of privacy data; The risk assessment level corresponding to the test report is determined based on predefined risk assessment standards, and the risk assessment level includes high risk, medium risk and low risk.

5. The method according to claim 4, characterized in that After the step of determining the risk assessment level corresponding to the test report based on the predefined risk assessment standard, the method further includes: When it is detected that the risk assessment level is high risk, an alarm notification is triggered.

6. A cloud server, characterized in that: The cloud server includes: one or more processors and memory; The memory is coupled to the one or more processors, and the memory is used to store computer program codes, wherein the computer program codes include computer instructions, and the one or more processors call the computer instructions to enable the cloud server to execute the method according to any one of claims 1 to 5.

7. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on the cloud server, the cloud server executes the method as described in any one of claims 1 to 5.

8. A computer program product, characterized in that When the computer program product runs on a cloud server, the cloud server is caused to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Remote sensing image classification method

    CN103345643A

  • Federal learning-based social media user privacy information protection method and system

    CN114003957A

  • Privacy leakage detection method and system based on terminal adaptive federal learning

    CN118761100A