Data processing method and apparatus
By modeling the dataset in both ciphertext and plaintext, generating and comparing the processing results, the contradiction between data privacy and model training effectiveness in joint modeling is resolved, thereby improving the accuracy and security of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
- Filing Date
- 2023-11-27
- Publication Date
- 2026-07-10
AI Technical Summary
In joint modeling, how can we ensure both the privacy and security of the data while maintaining the training effectiveness of the machine learning model, especially by addressing the discrepancy between encrypted computation results and plaintext?
By acquiring the dataset to be processed corresponding to the target business, ciphertext and plaintext modeling are performed to generate ciphertext business processing models and plaintext business processing models. The accuracy of the ciphertext business processing model is evaluated by comparing the differences between the ciphertext and plaintext processing results. Privacy computing techniques such as multi-party secure computation and homomorphic encryption are used to encrypt and associate the data.
This improved the accuracy of the encrypted business processing model, ensuring data privacy and security while enhancing the training effect of the machine learning model.
Smart Images

Figure CN117610044B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of multi-party secure computation technology, and in particular to a data processing method. Background Technology
[0002] The application of privacy-preserving computing technology in the field of AI modeling provides a safe and reliable data fusion method for joint modeling. Joint modeling is a machine learning method involving multiple parties that can fuse multiple datasets to improve the accuracy and generalization ability of the model.
[0003] In joint modeling, due to the privacy data of multiple participants, each dataset needs to be encrypted. Encryption algorithms may lose some precision, resulting in a certain degree of deviation between the encrypted calculation results and the plaintext. This leads to insufficient processing accuracy of the trained machine learning model. Therefore, how to ensure both data privacy and security while ensuring the training effect of the machine learning model has become an important problem that technical personnel urgently need to solve. Summary of the Invention
[0004] In view of the above, embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a data processing method is provided, comprising:
[0006] Obtain at least one dataset to be processed and one test dataset corresponding to the target business;
[0007] Ciphertext modeling is performed based on each dataset to be processed to generate a ciphertext business processing model; plaintext modeling is performed based on each dataset to be processed to generate a plaintext business processing model.
[0008] The test dataset is input into the ciphertext service processing model to obtain the ciphertext processing result, and the test dataset is input into the plaintext service processing model to obtain the plaintext processing result.
[0009] Based on the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, plaintext-ciphertext error information is obtained, and the training result of the ciphertext service processing model is determined based on the plaintext-ciphertext error information.
[0010] According to a second aspect of the embodiments of this specification, a data processing apparatus is provided, comprising:
[0011] The acquisition module is configured to acquire at least one pending dataset and one test dataset corresponding to the target business.
[0012] The generation module is configured to perform ciphertext modeling based on each dataset to be processed, generate a ciphertext business processing model, and perform plaintext modeling based on each dataset to be processed, generate a plaintext business processing model.
[0013] The input module is configured to input the test dataset into the ciphertext service processing model to obtain the ciphertext processing result, and input the test dataset into the plaintext service processing model to obtain the plaintext processing result;
[0014] The determination module is configured to obtain plaintext-ciphertext error information based on the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, and to determine the training result of the ciphertext service processing model based on the plaintext-ciphertext error information.
[0015] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising:
[0016] Memory and processor;
[0017] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above-described data processing method.
[0018] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of the data processing method described above.
[0019] According to a fifth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described data processing method.
[0020] This specification provides a data processing method according to one embodiment, which involves obtaining at least one dataset to be processed and a test dataset corresponding to a target service; performing ciphertext modeling based on each dataset to be processed to generate a ciphertext service processing model, and performing plaintext modeling based on each dataset to be processed to generate a plaintext service processing model; inputting the test dataset into the ciphertext service processing model to obtain a ciphertext processing result, and inputting the test dataset into the plaintext service processing model to obtain a plaintext processing result; obtaining plaintext-ciphertext error information based on the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, and determining the training result of the ciphertext service processing model based on the plaintext-ciphertext error information.
[0021] The method provided in this manual involves generating a ciphertext business processing model using the dataset to be processed, followed by plaintext modeling using the same dataset to generate a plaintext business processing model. A test dataset is then input into both the ciphertext and plaintext business processing models to obtain ciphertext and plaintext processing results. The accuracy of the ciphertext business processing model is evaluated by comparing the differences between the two results. This comparison of ciphertext and plaintext results improves the accuracy of the ciphertext business processing model's processing results. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating a data processing method provided in one embodiment of this specification;
[0023] Figure 2 This is a flowchart of the data processing method for numerical text classification scenarios provided in one embodiment of this specification;
[0024] Figure 3 This is a timing diagram of a data processing method provided in one embodiment of this specification;
[0025] Figure 4 This is a schematic diagram of the structure of a data processing device provided in one embodiment of this specification;
[0026] Figure 5 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0027] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0028] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0029] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0030] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0031] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0032] Privacy-preserving computation, also known as privacy-protecting computation, enables the analysis and computation of data without disclosing the original data, ensuring the secure flow of data in a "usable but invisible" manner.
[0033] Feature engineering: Essentially an engineering activity, the purpose of which is to extract and process features from raw data for use by models or algorithms.
[0034] Joint modeling refers to the process of multiple machine learning models (called Collaborative Modeling Entities, CMEs) working together to build a larger, more complex model. In joint modeling, multiple models are merged into a larger model, which can be used for various tasks such as prediction, classification, and clustering. Joint modeling can help develop more accurate and robust models, allowing different business models to obtain information from different perspectives and complement each other.
[0035] The application of privacy-preserving computing technology in the field of AI modeling provides a safe and reliable data fusion method for joint modeling. Joint modeling is a machine learning method involving multiple parties that can fuse multiple datasets to improve the accuracy and generalization ability of the model.
[0036] In collaborative modeling, due to the involvement of privacy data from multiple participants, each dataset needs to be encrypted. Encryption algorithms may suffer from some precision loss, leading to a discrepancy between the encrypted calculation results and the plaintext. This can result in insufficient processing accuracy of the trained machine learning model. However, in practical applications, a small, acceptable range of deviation is often acceptable. Therefore, how to ensure both data privacy and security while maintaining the training effectiveness of the machine learning model has become a crucial problem that technical personnel urgently need to solve.
[0037] Based on this, a data processing method is provided in this specification. This specification also relates to a data processing apparatus, a computing device, and a computer-readable storage medium, which will be described in detail in the following embodiments.
[0038] See Figure 1 , Figure 1 A flowchart of a data processing method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0039] Step 102: Obtain at least one dataset to be processed and a test dataset corresponding to the target business.
[0040] Specifically, the target business refers to the actual business processed in practical applications, such as text processing. More specifically, taking text-based numerical data processing tasks within text processing as an example, it may also include at least one of the following: binary classification, multi-class classification, regression, and clustering tasks. In the embodiments provided in this specification, the specific form of the target business is not limited, but is subject to actual application. Preferably, in one or more embodiments provided in this specification, the target business corresponds to the business corresponding to text-based numerical data.
[0041] It should be noted that the methods provided in the embodiments of this specification are applied to privacy-preserving computing scenarios, i.e., multiple data providers provide datasets to be processed. For example, with two data providers, there are two datasets to be processed; with four data providers, there are four datasets to be processed. There is a one-to-one correspondence between the datasets to be processed and the data providers. The datasets to be processed are used to subsequently train the corresponding business processing models.
[0042] The test dataset is used to detect whether the trained business processing model meets the training conditions; that is, the test dataset is used to verify whether the business processing model is trained well.
[0043] Step 104: Perform ciphertext modeling based on each dataset to be processed to generate a ciphertext business processing model; perform plaintext modeling based on each dataset to be processed to generate a plaintext business processing model.
[0044] The method provided in this specification requires performing ciphertext modeling and plaintext modeling separately based on the same dataset to be processed, obtaining corresponding ciphertext business processing models and plaintext business processing models. The accuracy of the ciphertext business processing model is then verified by comparing the data processing results of the ciphertext business processing model and the plaintext business processing model.
[0045] Specifically, based on each dataset to be processed, ciphertext modeling is performed to generate a ciphertext business processing model, including:
[0046] Based on privacy computing technology, each dataset to be processed is processed to generate a ciphertext modeling training dataset and a ciphertext modeling test dataset corresponding to each dataset to be processed.
[0047] An initial ciphertext business processing model is generated by training various ciphertext modeling training datasets using a federated learning approach.
[0048] The initial encrypted business processing model was tested based on the encrypted modeling test dataset to obtain the test results of the encrypted business processing model.
[0049] If the test results of the encrypted service processing model meet the test conditions, the initial encrypted service processing model is determined to be an encrypted service processing model.
[0050] Encrypted modeling is performed on each dataset to be processed to generate a encrypted business processing model; specifically, this refers to joint modeling based on privacy-preserving computation. Joint modeling based on privacy-preserving computation is a technique that jointly models multiple datasets to be processed while protecting data privacy.
[0051] In practical privacy computing, data encryption and encrypted data association are key steps in protecting data privacy. Therefore, it is necessary to process the business data in each dataset using privacy computing techniques, enabling the business data in each dataset to be integrated and encrypted and associated. This generates a encrypted modeling dataset and a encrypted modeling test dataset for training the encrypted business processing model.
[0052] Specifically, privacy-preserving computation technologies include multi-party secure computation and homomorphic encryption. Multi-party secure computation enables computation and data association among multiple data providers while protecting data privacy. Specifically, it includes:
[0053] a. Data providers determine the computation tasks and security protocols. Specifically, each data provider needs to first determine the computation tasks to be performed. These computation tasks refer to the tasks corresponding to the target business, including operations such as data association, aggregation, and encryption. Appropriate security protocols, such as multi-party secure computation protocols and homomorphic encryption protocols, are then determined among the data providers.
[0054] b. Data Encryption. Each data provider encrypts the business data in its own dataset to protect its data privacy. Encryption methods can include homomorphic encryption, differential privacy, and other techniques.
[0055] c. Data transmission. Each data provider sends encrypted data to other data providers to facilitate data association and computation among them.
[0056] d. Secure computation. Data providers use secure protocols for data association and computation, such as SMC protocol and homomorphic encryption protocol. These protocols ensure that data providers can obtain computation results without sharing the original dataset to be processed.
[0057] e. Decryption of Calculation Results. After secure computation is completed, each data provider decrypts the calculation results to obtain the original associated results of the dataset to be processed. Decryption methods can use techniques such as homomorphic encryption and differential privacy.
[0058] In practical applications, before processing each dataset based on privacy computing technology, the business data in each dataset can be preprocessed in advance to standardize and normalize the business data in the dataset.
[0059] Data preprocessing refers to the cleaning, integration, transformation, and normalization of business data in the original dataset before encrypted modeling, in order to better support subsequent encrypted modeling and data analysis. The purpose of data preprocessing is to improve the quality and accuracy of business data and eliminate noise, thereby achieving better data analysis and encrypted modeling. Data preprocessing includes, but is not limited to, data splitting, filtering business data based on preset filtering conditions, stratified sampling, weighted sampling, etc.
[0060] Data splitting specifically refers to dividing an existing dataset into training, validation, and test sets. This can be done according to a preset splitting ratio or by random splitting. Taking a preset splitting ratio as an example, the ratio of training, validation, and test sets is determined beforehand. Typically, the dataset is split in a 7:2:1 ratio, with 70% of the data used as the training set, 20% as the validation set, and 10% as the test set. Taking random splitting as an example, a random sampling method is used to randomly split the dataset according to the splitting ratio. When performing random splitting, it is important to ensure the stability of the class distribution in the training, validation, and test sets to avoid class imbalance.
[0061] Filtering business data based on preset filtering conditions specifically refers to filtering out samples that do not meet the requirements based on certain conditions in the business data, thereby improving the accuracy and generalization ability of the model. Specifically, the preset filtering conditions must first be determined, and then the business data in the dataset to be processed is cleaned according to these conditions, deleting or marking any business data that does not meet the conditions from the dataset.
[0062] Stratified sampling specifically refers to ensuring that the proportion of different categories of business data in the sampled dataset remains relatively stable, thereby avoiding class imbalance in the dataset. Specifically:
[0063] 1. Determine the categories: Before performing stratified sampling, it is necessary to clarify the categories in the sample data. For example, for binary classification problems, the samples can be divided into positive samples and negative samples, and for multi-class classification problems, the samples can be divided into multiple categories.
[0064] 2. Calculate the data proportions. Calculate the proportion of each category in the original dataset to be processed. For example, in a binary classification problem, if the ratio between positive and negative samples is 5:1, then the proportion of positive samples is 83.3% and the proportion of negative samples is 16.7%.
[0065] 3. Determine the sampling ratio. Based on the proportion of each category in the original dataset to be processed, determine the proportion of each category in the sampled dataset. For example, to avoid class imbalance, the proportion of each category in the sampled dataset can be set to the proportion of each category in the original data.
[0066] 4. Perform data sampling according to the sampling ratio. Specifically, after determining the proportion of each category of data in the sampled dataset, sample the corresponding proportion of samples from each category. For example, in a binary classification problem, if 100 samples are drawn from the original data, 83 positive samples and 17 negative samples are drawn.
[0067] In addition to stratified sampling, weighted sampling can also be used. Different weights are assigned to each sample based on its importance, thus better reflecting the contribution of each sample to the model. Specifically:
[0068] 1. Determine the weight of each data sample. Depending on the actual situation, a weight can be assigned to each sample to reflect its contribution to the model. For example, in the field of financial risk control, a defaulting customer can be given a higher weight to better reflect its impact on the model building.
[0069] 2. Calculate the total weight of the samples. After determining the weight of each data sample, calculate the sum of the weights of all data samples for subsequent weighted sampling.
[0070] 3. Determine the sampling ratio. Based on the weight of each data sample, determine the proportion of each sample in the sampled dataset. The sampling ratio is calculated by dividing the weight of each data sample by the sum of the sample weights. For example, for a sample set, if the weights of each data sample are 1, 2, 3, 4, and 5, then the sum of the sample weights is 15, and the proportions of each data sample are 1 / 15, 2 / 15, 3 / 15, 4 / 15, and 5 / 15, respectively.
[0071] 4. Perform weighted sampling: Based on the proportion of each data sample in the sampled dataset, use weighted random sampling to extract a corresponding proportion of samples from the original data.
[0072] After data association and preprocessing, a ciphertext modeling training dataset and a ciphertext modeling test dataset are obtained. Then, ciphertext modeling can be performed based on each ciphertext modeling training dataset through federated learning to train and generate an initial ciphertext business processing model.
[0073] Specifically, in the process of encrypted modeling, feature engineering is an important step in model building. Its purpose is to extract features useful for model building from the training data to better reflect the characteristic information and classification relationships of the samples. Feature engineering includes at least one of the following: missing value imputation, feature encoding, feature standardization, feature derivation, outlier handling, discrete encoding, and numerical encoding.
[0074] Missing value imputation refers to filling in missing values in the encrypted modeling training dataset using appropriate methods to improve the model's accuracy and generalization ability. Missing value imputation specifically includes:
[0075] 1. Identify missing values. First, it is necessary to identify the missing values in the encrypted modeling training dataset and determine the number and location of the missing values.
[0076] 2. Analyze the exact cause. For missing values in the encrypted modeling training dataset, it is necessary to analyze the reasons for their occurrence and their impact on model building. For example, are the missing values randomly lost, or are they related to the characteristics of the samples?
[0077] 3. Select a filling method. Based on the type and cause of the missing values, select an appropriate filling method. Commonly used filling methods include mean filling, median filling, mode filling, nearest neighbor filling, regression filling, etc.
[0078] 4. Perform imputation: Fill in the missing values according to the determined imputation method. For example, for numerical features, the mean or median can be used for imputation; for categorical features, the mode can be used for imputation; for time series, interpolation can be used for imputation, and so on.
[0079] Feature encoding is used to transform data in a encrypted modeling training dataset into numerical features. It can use embedded encoding or one-hot encoding, primarily to convert training data in an encrypted modeling training dataset into numerical data for training and prediction. Taking one-hot encoding as an example, it specifically includes:
[0080] 1. Identify discrete features. First, identify which features are discrete, such as occupation, education level, etc.
[0081] 2. Determine the encoding method. Based on the actual situation, choose an appropriate encoding method. Generally, binary encoding or one-hot encoding can be used.
[0082] 3. Map discrete features to the coding space. Map discrete features to the coding space through encoding methods. For example, for binary encoding, male and female can be encoded as 0 or 1. For one-hot encoding, male and female can be encoded as [1,0] and [0,1].
[0083] 4. Incorporate the encoded features into the original data to facilitate model training and prediction.
[0084] Feature standardization aims to eliminate the influence of different indicator units and facilitate comparability between indicators. Differences in units can affect the distance calculation results in some models. Specific feature standardization steps include:
[0085] 1. Identify the features that need to be standardized. In the implementation methods provided in this specification, these are typically numerical features.
[0086] 2. Calculate the mean and standard deviation of the features.
[0087] 3. For each feature that needs to be standardized, standardize it based on the formula, i.e. (eigenvalue - mean) / standard deviation.
[0088] 4. Integrate the standardized features into the original data to facilitate model training and prediction.
[0089] Feature derivation primarily involves combining and transforming original features to generate new features, thereby improving the model's accuracy and generalization ability. Specifically, feature derivation includes:
[0090] 1. Identify the features that need to be derived, that is, identify which original features need to be derived, such as whether there is a correlation between certain features, whether they need to be combined, etc.
[0091] 2. Determine the derivation method. Based on the actual situation, select an appropriate feature derivation method. Commonly used derivation methods include linear combination, polynomial features, cross features, discretization, etc.
[0092] 3. Perform feature derivation: Based on the selected feature derivation method, combine, transform, or perform other operations on the original features to obtain new features. For example, for linear combinations, multiple features can be weighted and summed; for polynomial features, the original features can be exponentially operated on, and so on.
[0093] 4. Integrate the derived features into the original data to facilitate model training and prediction.
[0094] Outliers are unreasonable values present in the encrypted modeling training dataset. It's important to note that unreasonable values deviate from the normal range, not necessarily errors. Identifying and handling outliers in the encrypted modeling training dataset can improve the model's accuracy and generalization ability. The specific outlier handling steps are as follows:
[0095] 1. The presence of outliers in the encrypted modeling training dataset can be determined through methods such as visualization analysis and statistical analysis.
[0096] 2. Analyze the causes of anomalies. For outliers, it is necessary to analyze the reasons for their occurrence and their impact on model building in order to select appropriate handling methods.
[0097] 3. Based on the cause of the anomaly and the data distribution, select an appropriate outlier handling method. Common outlier handling methods include deleting outliers, replacing outliers, smoothing outliers, etc.
[0098] 4. Outlier handling methods can be used to process outliers. For example, to delete outliers, they can be removed from the encrypted modeling training dataset; to replace outliers, the mean, median, etc. can be used; and to smooth outliers, methods such as moving average can be used.
[0099] Discrete coding is primarily used to transform numerical features into discrete features to facilitate model training and prediction. Discrete coding can be WOE coding. Specifically:
[0100] 1. Identify the features that need to be discretely encoded, which are usually numerical features.
[0101] 2. Determine the grouping method. For numerical feature types that need to be encoded using WOE, it is necessary to determine a suitable grouping method. Commonly used grouping methods include equal-interval grouping, equal-frequency grouping, and clustering grouping.
[0102] 3. For each group, calculate its corresponding WOE value based on the preset WOE formula.
[0103] 4. Map the WOE values back to the original data to obtain the discrete features encoded by WOE.
[0104] Numerical encoding is primarily used to transform discrete features into numerical features to facilitate model training and prediction. Numerical encoding can be target encoding, specifically:
[0105] 1. Identify the features that need to be Target encoded, which are usually discrete features.
[0106] 2. Determine the appropriate coding method based on the actual situation. Commonly used coding methods include smooth coding, basic target coding, and mean coding, etc.
[0107] 3. For each feature that needs to be Target encoded, calculate its mean in the target variable.
[0108] 4. Calculate the encoding value of discrete features according to the selected encoding method. For example, for smooth encoding, Bayesian smoothing and other methods can be used to calculate the encoding value.
[0109] 5. Incorporate the encoded features into the original data to facilitate model training and prediction.
[0110] In the process of generating an initial encrypted business processing model through federated learning by training various encrypted modeling training datasets, federated learning can be understood as using privacy-preserving techniques to jointly model and train the data, resulting in a jointly modeled encrypted business processing model. Currently, algorithms supported by privacy-preserving joint modeling training include GBDT binary classification, XGBOOST binary classification, logistic regression binary classification, XGBOOST regression, and XGBOOST multi-class classification.
[0111] Taking GBDT binary classification as an example, an initial decision tree is first trained. Then, the residual of the current model is calculated, and the residual is used to train the next decision tree, and so on, until a preset number of iterations is reached or the loss value is less than a threshold. Finally, GBDT combines multiple decision trees to generate a strong classifier for binary classification tasks. Specifically, the GBDT binary classification algorithm modeling first uses the GradientBoostingClassifier class to define the GBDT classification model and sets appropriate parameters, such as tree depth, learning rate, and number of iterations. The fit() method is used to train the model on the encrypted modeling training dataset, and tools such as the GridSearchCV class are used to adjust the model parameters to improve the model's accuracy and generalization ability.
[0112] Taking XGBoost binary classification as an example, we define an XGBoost classification model using the XGBClassifier class and set the corresponding parameters, such as tree depth, learning rate, number of iterations, etc. We then train the model on the encrypted modeling training dataset using the fit() method and adjust the model parameters using tools such as the GridSearchCV class to improve the model's accuracy and generalization ability.
[0113] Taking logistic regression binary classification as an example, we define a logistic regression classification model using the LogisticRegression class and set the corresponding parameters, such as regularization strength. We then use the fit() method to train the model on the encrypted modeling training dataset and use tools such as the GridSearchCV class to adjust the model parameters to improve the model's accuracy and generalization ability.
[0114] Taking XGBoost regression as an example, XGBoost regression is a commonly used learning algorithm when building regression models. The XGBRegressor class is used to define the XGBoost regression model and set the corresponding parameters, such as tree depth, learning rate, number of iterations, etc. The fit() method is used to train the model on the encrypted modeling training dataset. The model parameters are adjusted through methods such as cross-validation and grid search to improve the accuracy and generalization ability of the model.
[0115] Taking XGBoost multi-class classification as an example, XGBoost multi-class classification is a commonly used machine learning algorithm for building classification models to solve multi-class problems. Specifically, the XGBClassifier class is used to define the XGBoost classification model, and corresponding parameters are set, such as tree depth, learning rate, number of iterations, etc. The fit() method is used to train the model on the encrypted modeling training dataset. Through methods such as cross-validation and grid search, the model parameters are adjusted to improve the model's accuracy and generalization ability.
[0116] After training the initial encrypted business processing model using various encrypted modeling training datasets, it is necessary to use the encrypted modeling test dataset to predict the results. During prediction, it is crucial to ensure that the encrypted modeling test data and the encrypted modeling training data have the same feature dimensions and representations to guarantee correct predictions. Specifically, the trained initial encrypted business processing model is loaded, and the `predict()` method is used to predict the test data in the encrypted modeling test dataset. The model's output test results are then displayed or saved.
[0117] If the test results of the encrypted business processing model meet the preset test conditions, it means that the initial encrypted business processing model has been successfully trained and can be used as the encrypted business processing model. If the test results of the encrypted business processing model do not meet the preset test conditions, the initial encrypted business processing model will continue to be trained until the test results of the encrypted business processing model meet the preset test conditions.
[0118] The methods provided in this specification can also be used to validate and evaluate the encrypted business processing model using test data, assessing its effectiveness and validity. Specifically, the validation and evaluation of the encrypted business processing model includes multi-classification evaluation, regression model evaluation, contribution evaluation, and confusion matrix analysis.
[0119] Multi-class evaluation includes binary or multi-class classification. When evaluating a model, it's necessary to combine the model's performance metrics with actual needs, selecting appropriate metrics for evaluation. Furthermore, methods such as cross-validation are needed to evaluate the model and obtain more accurate performance metrics. Specifically, multi-class evaluation requires calculating model performance metrics, such as accuracy, recall, precision, and F1 score, based on the predicted results and the true classification labels of the test data. ROC curves are plotted based on the model's prediction results and the true classification labels of the test data, and the AUC value is calculated. The thresholds for the model's prediction results are adjusted according to actual needs to obtain more suitable classification results.
[0120] The regression model evaluation calculates performance metrics such as mean squared error and root mean square error based on the predicted results and the true classification labels of the test data. A prediction graph is then plotted and visualized using the predicted results and the true classification labels of the test data. Different regression algorithms are used to train the model, and the performance of different models is compared to select the better model.
[0121] Contribution assessment requires evaluating the importance of features to the overall sample, categorized into the importance of features from local data providers and those from remote data providers. By analyzing the weights, coefficients, and variable importance of each feature in the model, the contribution of each feature to the prediction results is evaluated. Based on the feature importance assessment results, a feature importance graph is plotted and visualized. Based on the feature importance assessment results, features with high importance are selected for model training to improve the model's accuracy and generalization ability.
[0122] A confusion matrix is used to compare a classifier's predictions with the true class labels and to calculate the classifier's performance metrics. Specifically, the predictions are first compared with the true class labels in the test data to construct the confusion matrix. The rows of the confusion matrix represent the classes predicted by the model, and the columns represent the true classes in the test data. Each element of the matrix represents the number of times the model correctly classifies one true class as another. Based on the confusion matrix, performance metrics such as accuracy, recall, precision, and F1 score are calculated. Furthermore, the confusion matrix can be visualized to better understand the model's performance.
[0123] At this point, encrypted business processing models are generated based on each dataset to be processed. At the same time, plaintext models are also generated based on each dataset to be processed.
[0124] In the implementation method provided in this specification, plaintext modeling can be processed with reference to conventional machine learning algorithms. During the plaintext modeling process, the datasets to be processed by each data provider do not need to be encrypted or associated.
[0125] Using the same dataset to be processed, ciphertext modeling and plaintext modeling are performed separately to generate ciphertext business processing models and plaintext business processing models. The purpose is to ensure that the model parameters of the ciphertext business processing model and the plaintext business processing model are the same in the subsequent comparison process, reduce the difference in data processing results caused by the difference in model parameters, and better compare the plaintext and ciphertext processing results.
[0126] Step 106: Input the test dataset into the ciphertext service processing model to obtain the ciphertext processing result, and input the test dataset into the plaintext service processing model to obtain the plaintext processing result.
[0127] After training the encrypted and plaintext business processing models, the same test dataset can be input into the encrypted and plaintext business processing models respectively for processing.
[0128] It is important to note that when using the encrypted and plaintext business processing models to process the test data in the test dataset, the same test data and the same model parameters should be used. For example, the learning rate, tree depth, training feature sampling ratio, training sample sampling ratio, tree data, minimum number of samples per leaf node, and minimum number of samples required for node splitting should be set to the same parameters for both the encrypted and plaintext business processing models.
[0129] After setting the same test data and model parameters, the ciphertext business processing model generates ciphertext processing results based on the test dataset, and the plaintext business processing model generates plaintext processing results based on the test dataset.
[0130] Step 108: Obtain plaintext-ciphertext error information based on the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, and determine the training result of the ciphertext service processing model based on the plaintext-ciphertext error information.
[0131] After obtaining the ciphertext processing result and the plaintext processing result, the plaintext-ciphertext error information can be obtained by comparing the two results. Specifically, the plaintext-ciphertext error information is the difference between the ciphertext processing result and the plaintext processing result. This error information represents the error information between the ciphertext business processing model and the plaintext business processing model. Furthermore, to better obtain the plaintext-ciphertext error information, the test dataset of the ciphertext business processing model, the test dataset of the plaintext business processing model, the ciphertext processing result, and the plaintext processing result can be further compared to obtain the plaintext-ciphertext error information.
[0132] Specifically, based on the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, plaintext-ciphertext error information is obtained, including:
[0133] Determine the target plaintext-ciphertext comparison strategy corresponding to the target service;
[0134] Based on the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result of the target plaintext-ciphertext comparison strategy, the plaintext-ciphertext error information corresponding to the target plaintext-ciphertext comparison strategy is obtained.
[0135] In practical applications, different target services correspond to different plaintext-ciphertext comparison strategies. It is necessary to first determine the corresponding target plaintext-ciphertext comparison strategy, and then compare the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result based on the target plaintext-ciphertext comparison strategy to obtain the plaintext-ciphertext error information corresponding to the target plaintext-ciphertext comparison strategy. The details of each plaintext-ciphertext comparison strategy will be introduced below.
[0136] In one specific embodiment provided in this specification, the target plaintext-ciphertext comparison strategy includes a data association strategy;
[0137] Accordingly, the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result are compared based on the target plaintext-ciphertext comparison strategy, including:
[0138] Compare the number of records in each field, the number of missing values, and at least one of the following in the data distribution graph of the test dataset of the encrypted business processing model and the test dataset of the plaintext business processing model.
[0139] Data association strategy refers to a comparison strategy that compares the number of records, the number of missing values, and the data distribution map of each field in the test dataset of the encrypted business processing model and the test dataset of the plaintext business processing model to see if there are any differences. Data association strategy is used to evaluate the accuracy and security of data association and can discover potential problems and vulnerabilities.
[0140] Specifically, comparing the number of records for each field in the test dataset of the encrypted business processing model can verify whether the encrypted data association retains all records and whether the number of records is consistent with the number in the plaintext processing result; comparing the number of missing values can verify whether the encrypted data association in the encrypted processing result correctly handles missing values and whether the number of missing values is consistent with the data in the test dataset of the plaintext business processing model; comparing the data distribution plot can reveal whether there are differences in data distribution between the test dataset of the encrypted business processing model and the test dataset of the plaintext business processing model, and whether the accuracy and security of the encrypted modeling are affected.
[0141] It is important to note that comparing the test datasets of the encrypted and plaintext business processing models to determine if there are differences in the number of records, the number of missing values, and the data distribution. This requires considering factors such as data type, encryption algorithm, and specific application scenario, and validation on a sufficiently large dataset is necessary to ensure the reliability and accuracy of the results. During the comparison process, it is crucial to maintain consistency in the processing methods and model parameters for both the encrypted and plaintext business processing model test datasets to guarantee the reliability of the comparison results.
[0142] In another specific embodiment provided in this specification, the target plaintext-ciphertext comparison strategy includes a data preprocessing strategy;
[0143] Accordingly, the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result are compared based on the target plaintext-ciphertext comparison strategy, including:
[0144] Determine the target plaintext-ciphertext comparison method corresponding to the data preprocessing strategy;
[0145] The test datasets of the ciphertext service processing model and the plaintext service processing model are compared based on the target plaintext-ciphertext comparison method.
[0146] In the actual model training phase, due to the inherent randomness of data sampling, the sampling results of the test datasets for the encrypted and plaintext business processing models cannot be completely identical. Therefore, when comparing the test datasets for the encrypted and plaintext business processing models, the impact of errors introduced by random data sampling needs to be considered. If the random error is small, it can be mitigated by comparing the mean, quantiles, etc. Specifically, data preprocessing strategies can be used to compare the test datasets for the encrypted and plaintext business processing models.
[0147] In practical applications, data preprocessing strategies also correspond to different plaintext-ciphertext comparison methods. Specifically, the target plaintext-ciphertext comparison method includes at least one of the following: sampling comparison method, mean comparison method, quantile comparison method, and coefficient comparison method; wherein...
[0148] The sampling comparison method includes sampling the test dataset of the encrypted service processing model and the test dataset of the plaintext service processing model a preset number of times, and comparing the data based on the sampling results;
[0149] The mean comparison method includes calculating the ciphertext mean corresponding to the test dataset of the ciphertext service processing model, calculating the plaintext mean corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext mean and the plaintext mean.
[0150] The quantile comparison method includes calculating the ciphertext quantiles corresponding to the test dataset of the ciphertext service processing model, calculating the plaintext quantiles corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext quantiles with the quantiles.
[0151] The coefficient comparison method includes determining the ciphertext coefficient information corresponding to the test dataset of the ciphertext service processing model, determining the plaintext coefficient information corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext coefficient information and the plaintext coefficient information.
[0152] For the sampling comparison method, test datasets for both the encrypted and plaintext business processing models can be sampled, and the sampling results can be compared. Since sampling results contain random errors, multiple samplings can be performed for a preset number of times, and the average result can be taken to reduce errors caused by random sampling.
[0153] For the mean comparison method, we can compare whether there is a significant difference between the means of the test dataset of the encrypted business processing model and the test dataset of the plaintext business processing model. Specifically, statistical methods such as t-tests can be used to make the judgment.
[0154] For the quantile comparison method, we can compare whether there is a significant difference between the quantiles of the test dataset of the encrypted business processing model and the test dataset of the plaintext business processing model. Specifically, non-parametric test methods such as Wilcoxon can be used to make the judgment.
[0155] For the coefficient comparison method, we can compare whether there is a significant difference in the correlation coefficient between the test dataset of the encrypted business processing model and the test dataset of the plaintext business processing model. Statistical methods such as Spearman correlation coefficient can be used to make a judgment.
[0156] In another specific embodiment provided in this specification, the target plaintext-ciphertext comparison strategy includes a feature engineering strategy;
[0157] Accordingly, the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result are compared based on the target plaintext-ciphertext comparison strategy, including:
[0158] Determine the target feature encoding comparison method corresponding to the feature engineering strategy;
[0159] The test datasets of the ciphertext service processing model and the plaintext service processing model are compared based on the target feature encoding comparison method.
[0160] The plaintext-ciphertext comparison strategy also includes feature engineering strategies. Feature encoding is a commonly used feature engineering method. After determining the target feature encoding comparison method corresponding to the feature engineering strategy, the test dataset of the ciphertext business processing model and the test dataset of the plaintext business processing model are further compared based on the target feature encoding comparison method.
[0161] Specifically, the target feature encoding comparison method includes at least one of the following: encoding feature quantity comparison method, encoding feature map comparison method, encoding feature category comparison method, and encoding feature statistical information comparison method.
[0162] The method for comparing the number of coded features includes obtaining the number of coded features and the name of coded features in the test dataset of the coded service processing model, obtaining the number of coded features and the name of coded features in the test dataset of the plaintext service processing model, comparing the number of coded features and the number of coded features, and comparing the name of coded features and the name of coded features.
[0163] The encoding feature map comparison method includes obtaining the ciphertext encoding feature map corresponding to the test dataset of the ciphertext service processing model, obtaining the plaintext encoding feature map corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext encoding feature map and the plaintext encoding feature map.
[0164] The encoding feature category comparison method includes obtaining the ciphertext encoding category corresponding to the test dataset of the ciphertext service processing model, obtaining the plaintext encoding category corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext encoding category and the plaintext encoding category.
[0165] The encoding feature statistical information comparison method includes obtaining the ciphertext encoding mean, ciphertext encoding variance, and ciphertext encoding quantiles corresponding to the test dataset of the ciphertext service processing model; obtaining the plaintext encoding mean, plaintext encoding variance, and plaintext encoding quantiles corresponding to the test dataset of the plaintext service processing model; comparing the ciphertext encoding mean with the plaintext encoding mean; comparing the ciphertext encoding variance with the plaintext encoding variance; and comparing the ciphertext encoding quantiles with the plaintext encoding quantiles.
[0166] For the coding feature quantity comparison method, the number of features and feature names in the test dataset of the ciphertext business processing model and the test dataset of the plaintext business processing model can be compared to verify whether the ciphertext modeling correctly preserves all feature information.
[0167] For the coding feature map comparison method, the feature distribution maps of the test dataset of the ciphertext business processing model and the test dataset of the plaintext business processing model can be compared to see if they are consistent, which is used to verify whether the ciphertext modeling process of feature distribution is accurate.
[0168] For the coding feature category comparison method, we can compare whether the feature types of the test dataset of the ciphertext business processing model and the test dataset of the plaintext business processing model are consistent, including numeric, character, Boolean, etc.
[0169] For the coding feature statistical information comparison method, the coding feature statistical information of the test dataset of the ciphertext business processing model and the test dataset of the plaintext business processing model can be compared, such as the mean, variance, quantile and other statistical quantities after feature encoding. By comparing whether the mean, variance, quantile and other statistical quantities of the encoded ciphertext and plaintext are consistent, it can be verified whether the ciphertext modeling is accurate in processing the data after feature encoding.
[0170] The above-mentioned comparison strategies are all plaintext-ciphertext comparison strategies for modeling. In another specific embodiment provided in this specification, the target plaintext-ciphertext comparison strategy also includes a model performance comparison strategy.
[0171] Accordingly, the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result are compared based on the target plaintext-ciphertext comparison strategy, including:
[0172] Based on the ciphertext processing results, determine the ciphertext model effect information of the ciphertext business processing model;
[0173] Based on the plaintext processing results, determine the plaintext model effect information of the plaintext service processing model;
[0174] Compare the ciphertext model effect information with the plaintext model effect information.
[0175] In the embodiments provided in this specification, in addition to comparing the plaintext and ciphertext models, the ciphertext business processing model and the plaintext business processing model are also compared simultaneously. Specifically, the same dataset is used to build both the ciphertext and plaintext business processing models, and the prediction and evaluation results of the models are compared and analyzed. The model performance comparison strategy can evaluate the accuracy of ciphertext modeling and the differences between ciphertext and plaintext modeling. It can also verify the effectiveness and security of the encryption algorithm. Validation is performed on a sufficient amount of data to ensure the reliability and accuracy of the evaluation results. When comparing the ciphertext and plaintext business processing models using the model performance comparison strategy, it is essential to ensure that the processing methods and model parameters for the ciphertext and plaintext data remain consistent to guarantee the reliability of the comparison results.
[0176] In the comparison process based on the model performance comparison strategy, the ciphertext model performance information of the ciphertext business processing model is determined according to the ciphertext processing results, and the plaintext model performance information of the plaintext business processing model is determined according to the plaintext processing results. Model performance information refers to the evaluation indicators used to measure model performance, such as AUC, KS, F1 score, Gini coefficient, Kappa coefficient, LogLoss, coefficient of determination, residual histogram, etc.
[0177] Because the algorithms used in privacy-preserving computation products may experience some loss of precision, there is a certain degree of deviation between the ciphertext processing results and the plaintext processing results. The acceptable deviation range for each evaluation metric will be provided by technical personnel. Combining this with information such as the privacy protection effectiveness, computational efficiency, and resource consumption of the data provider, selecting appropriate evaluation metrics can help effectively assess the model effectiveness and security of privacy-preserving computation joint modeling.
[0178] In the above steps, based on the test dataset of the ciphertext service processing model and the test dataset of the plaintext service processing model, the ciphertext processing result and the plaintext processing result, plaintext-ciphertext error information is obtained.
[0179] In the embodiments provided in this specification, determining the training result of the ciphertext service processing model based on the plaintext-ciphertext error information includes:
[0180] If the plaintext-ciphertext error information is greater than or equal to a preset error information threshold, the training result of the ciphertext business processing model is determined to be unsuccessful.
[0181] If the plaintext-ciphertext error information is less than a preset error information threshold, the training result of the ciphertext business processing model is determined to be passed.
[0182] The preset error information threshold represents the acceptable error range. If the plaintext-ciphertext error information is greater than or equal to the preset error information threshold, the plaintext-ciphertext comparison test fails, and the training result of the ciphertext business processing model is deemed unsuccessful. Conversely, if the plaintext-ciphertext error information is less than the preset error information threshold, the plaintext-ciphertext comparison test passes, and the training result of the ciphertext business processing model is deemed successful.
[0183] If the training result of the encrypted service processing model fails, the encrypted service processing model needs to be retrained, and the plaintext-ciphertext comparison method provided in the embodiments of this specification should be continued until the plaintext-ciphertext error information is less than the preset error information threshold.
[0184] The method provided in the embodiments of this specification provides a testing method for the accuracy of privacy calculation based on plaintext-ciphertext comparison. It can compare the ciphertext processing results with the plaintext processing results, and the test dataset of the ciphertext business processing model with the test dataset of the plaintext business processing model to evaluate the accuracy of ciphertext modeling, so that the error between the encryption calculation results and the plaintext results is within an acceptable range.
[0185] Meanwhile, during the plaintext-ciphertext accuracy testing process, different comparison strategies were determined for different stages of the ciphertext business processing model. Plaintext-ciphertext comparisons were performed from multiple dimensions, providing multi-faceted information for the final comparison results and making the plaintext-ciphertext error information more comprehensive. This approach ensured both data privacy and security while also guaranteeing the training effectiveness of the ciphertext business processing model.
[0186] The following is in conjunction with the appendix Figure 2 Taking the application of the data processing method provided in this specification in a numerical text classification scenario as an example, the data processing method will be further explained. Among other things, Figure 2 The present specification illustrates a flowchart of a data processing method for numerical text classification scenarios provided by an embodiment of this specification, which specifically includes the following steps.
[0187] Step 202: Receive at least one training dataset and a test dataset to be processed.
[0188] Step 204: Perform ciphertext modeling based on each training dataset to be processed, and generate a ciphertext model.
[0189] Step 206: Input the test dataset into the ciphertext model to obtain the ciphertext processing results.
[0190] Step 208: Perform plaintext modeling based on each training dataset to be processed, and generate a plaintext model.
[0191] Step 210: Input the test dataset into the plaintext model to obtain the plaintext processing results.
[0192] Step 212: Compare the ciphertext processing results with the plaintext processing results to obtain plaintext-ciphertext error information.
[0193] Step 214: Determine whether the plaintext / ciphertext error information is within the error range. If yes, proceed to step 216; otherwise, proceed to step 218.
[0194] Step 216: Confirm that the encrypted model training is complete.
[0195] Step 218: Determine if the encrypted model needs to be retrained.
[0196] The method provided in the embodiments of this specification offers a test method for the accuracy of privacy calculations in plaintext-ciphertext comparison applied to numerical text classification scenarios. It compares the ciphertext processing results with the plaintext processing results to evaluate the accuracy of the ciphertext numerical text processing model, ensuring that the error between the encrypted calculation results and the plaintext results is within an acceptable range.
[0197] The following is in conjunction with the appendix Figure 3 This document explains one of the data processing methods provided. Figure 3 A timing diagram of the data processing method provided in the embodiments of this specification is shown. In this embodiment, taking the deployment of the encrypted service processing model and the plaintext service processing model on the cloud side as an example, the privacy computing terminal product is deployed in the cloud. Users can log in to the cloud, select the services they want to use, and thus assemble the encrypted service processing model that they need.
[0198] Step 302: The user logs into the cloud page, uploads the dataset required for model training, selects the algorithm components of the privacy computing product, and generates the initial encrypted business processing model.
[0199] Step 304: The privacy computing terminal product uses the dataset to train the initial encrypted business processing model, obtains the encrypted business processing model, and obtains the encrypted processing result.
[0200] Step 306: The privacy computing terminal product uploads the encrypted processing result to the plaintext-ciphertext comparison service.
[0201] Step 308: The user uploads the same dataset and model configuration description to the plaintext computing service.
[0202] Step 310: The plaintext computing service uses the dataset to train a plaintext business processing model and obtains the plaintext processing results.
[0203] Step 312: The plaintext calculation service uploads the plaintext processing results to the plaintext-ciphertext comparison service.
[0204] Step 314: The plaintext-ciphertext comparison service compares the ciphertext processing result with the plaintext processing result to obtain the plaintext-ciphertext error.
[0205] Step 316: The plaintext-ciphertext comparison service will feed back the plaintext-ciphertext error to the user so that the user can judge the availability of the privacy computing terminal product based on the plaintext-ciphertext error.
[0206] The method provided in the embodiments of this specification provides a testing method for the accuracy of privacy calculation based on plaintext-ciphertext comparison. It can compare the ciphertext processing results with the plaintext processing results, and the test dataset of the ciphertext business processing model with the test dataset of the plaintext business processing model to evaluate the accuracy of ciphertext modeling, so that the error between the encryption calculation results and the plaintext results is within an acceptable range.
[0207] Corresponding to the above method embodiments, this specification also provides data processing apparatus embodiments. Figure 4 A schematic diagram of the structure of a data processing apparatus according to one embodiment of this specification is shown. Figure 4 As shown, the device includes:
[0208] The acquisition module 402 is configured to acquire at least one pending dataset and one test dataset corresponding to the target business.
[0209] The generation module 404 is configured to perform ciphertext modeling based on each dataset to be processed to generate a ciphertext business processing model, and to perform plaintext modeling based on each dataset to be processed to generate a plaintext business processing model.
[0210] Input module 406 is configured to input the test dataset into the ciphertext service processing model to obtain ciphertext processing results, and input the test dataset into the plaintext service processing model to obtain plaintext processing results;
[0211] The determination module 408 is configured to obtain plaintext-ciphertext error information based on the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, and to determine the training result of the ciphertext service processing model based on the plaintext-ciphertext error information.
[0212] Optionally, the determining module 408 is further configured to:
[0213] Determine the target plaintext-ciphertext comparison strategy corresponding to the target service;
[0214] Based on the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result of the target plaintext-ciphertext comparison strategy, the plaintext-ciphertext error information corresponding to the target plaintext-ciphertext comparison strategy is obtained.
[0215] Optionally, the target plaintext-ciphertext comparison strategy includes a data association strategy;
[0216] The determining module 408 is further configured as follows:
[0217] Compare the number of records in each field, the number of missing values, and at least one of the following in the data distribution graph of the test dataset of the encrypted business processing model and the test dataset of the plaintext business processing model.
[0218] Optionally, the target plaintext-ciphertext comparison strategy includes a data preprocessing strategy;
[0219] The determining module 408 is further configured as follows:
[0220] Determine the target plaintext-ciphertext comparison method corresponding to the data preprocessing strategy;
[0221] The test datasets of the ciphertext service processing model and the plaintext service processing model are compared based on the target plaintext-ciphertext comparison method.
[0222] Optionally, the target plaintext-ciphertext comparison method includes at least one of the following: sampling comparison method, mean comparison method, quantile comparison method, and coefficient comparison method; wherein,
[0223] The sampling comparison method includes sampling the test dataset of the encrypted service processing model and the test dataset of the plaintext service processing model a preset number of times, and comparing the data based on the sampling results;
[0224] The mean comparison method includes calculating the ciphertext mean corresponding to the test dataset of the ciphertext service processing model, calculating the plaintext mean corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext mean and the plaintext mean.
[0225] The quantile comparison method includes calculating the ciphertext quantiles corresponding to the test dataset of the ciphertext service processing model, calculating the plaintext quantiles corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext quantiles with the quantiles.
[0226] The coefficient comparison method includes determining the ciphertext coefficient information corresponding to the test dataset of the ciphertext service processing model, determining the plaintext coefficient information corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext coefficient information and the plaintext coefficient information.
[0227] Optionally, the target plaintext-ciphertext comparison strategy includes a feature engineering strategy;
[0228] The determining module 408 is further configured as follows:
[0229] Determine the target feature encoding comparison method corresponding to the feature engineering strategy;
[0230] The test datasets of the ciphertext service processing model and the plaintext service processing model are compared based on the target feature encoding comparison method.
[0231] Optionally, the target feature encoding comparison method includes at least one of the following: encoding feature quantity comparison method, encoding feature map comparison method, encoding feature category comparison method, and encoding feature statistical information comparison method, wherein,
[0232] The method for comparing the number of coded features includes obtaining the number of coded features and the name of coded features in the test dataset of the coded service processing model, obtaining the number of coded features and the name of coded features in the test dataset of the plaintext service processing model, comparing the number of coded features and the number of coded features, and comparing the name of coded features and the name of coded features.
[0233] The encoding feature map comparison method includes obtaining the ciphertext encoding feature map corresponding to the test dataset of the ciphertext service processing model, obtaining the plaintext encoding feature map corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext encoding feature map and the plaintext encoding feature map.
[0234] The encoding feature category comparison method includes obtaining the ciphertext encoding category corresponding to the test dataset of the ciphertext service processing model, obtaining the plaintext encoding category corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext encoding category and the plaintext encoding category.
[0235] The encoding feature statistical information comparison method includes obtaining the ciphertext encoding mean, ciphertext encoding variance, and ciphertext encoding quantiles corresponding to the test dataset of the ciphertext service processing model; obtaining the plaintext encoding mean, plaintext encoding variance, and plaintext encoding quantiles corresponding to the test dataset of the plaintext service processing model; comparing the ciphertext encoding mean with the plaintext encoding mean; comparing the ciphertext encoding variance with the plaintext encoding variance; and comparing the ciphertext encoding quantiles with the plaintext encoding quantiles.
[0236] Optionally, the target plaintext-ciphertext comparison strategy includes a model performance comparison strategy;
[0237] The determining module 408 is further configured as follows:
[0238] Based on the ciphertext processing results, determine the ciphertext model effect information of the ciphertext business processing model;
[0239] Based on the plaintext processing results, determine the plaintext model effect information of the plaintext service processing model;
[0240] Compare the ciphertext model effect information with the plaintext model effect information.
[0241] Optionally, the determining module 408 is further configured to:
[0242] If the plaintext-ciphertext error information is greater than or equal to a preset error information threshold, the training result of the ciphertext business processing model is determined to be unsuccessful.
[0243] If the plaintext-ciphertext error information is less than a preset error information threshold, the training result of the ciphertext business processing model is determined to be passed.
[0244] Optionally, the generation module 404 is further configured to:
[0245] Based on privacy computing technology, each dataset to be processed is processed to generate a ciphertext modeling training dataset and a ciphertext modeling test dataset corresponding to each dataset to be processed.
[0246] An initial ciphertext business processing model is generated by training various ciphertext modeling training datasets using a federated learning approach.
[0247] The initial encrypted business processing model was tested based on the encrypted modeling test dataset to obtain the test results of the encrypted business processing model.
[0248] If the test results of the encrypted service processing model meet the test conditions, the initial encrypted service processing model is determined to be an encrypted service processing model.
[0249] The apparatus provided in the embodiments of this specification provides a testing device for the accuracy of privacy calculations based on plaintext-ciphertext comparison. It can compare the ciphertext processing results with the plaintext processing results, and the test dataset of the ciphertext business processing model with the test dataset of the plaintext business processing model, to evaluate the accuracy of ciphertext modeling, so that the error between the encryption calculation results and the plaintext results is within an acceptable range.
[0250] Meanwhile, during the plaintext-ciphertext accuracy testing process, different comparison strategies were determined for different stages of the ciphertext business processing model. Plaintext-ciphertext comparisons were performed from multiple dimensions, providing multi-faceted information for the final comparison results and making the plaintext-ciphertext error information more comprehensive. This approach ensured both data privacy and security while also guaranteeing the training effectiveness of the ciphertext business processing model.
[0251] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the data processing apparatus is basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the data processing method embodiments.
[0252] Figure 5 A structural block diagram of a computing device 500 according to one embodiment of this specification is shown. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.
[0253] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0254] In one embodiment of this specification, the above-described components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0255] Computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). Computing device 500 can also be a mobile or stationary server.
[0256] The processor 520 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-described data processing method.
[0257] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the data processing method embodiments.
[0258] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described data processing method.
[0259] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the data processing method embodiments.
[0260] An embodiment of this specification also provides a computer program, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described data processing method.
[0261] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer program embodiments are basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the data processing method embodiments.
[0262] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0263] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0264] It should be noted that the above description describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0265] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0266] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A data processing method, comprising: Obtain at least one dataset to be processed and one test dataset corresponding to the target business; Ciphertext modeling is performed based on each dataset to be processed to generate a ciphertext business processing model; plaintext modeling is performed based on each dataset to be processed to generate a plaintext business processing model. The test dataset is input into the ciphertext service processing model to obtain the ciphertext processing result, and the test dataset is input into the plaintext service processing model to obtain the plaintext processing result. Based on the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, plaintext-ciphertext error information is obtained, and the training result of the ciphertext service processing model is determined based on the plaintext-ciphertext error information.
2. The method as described in claim 1, wherein the plaintext-ciphertext error information is obtained based on the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, comprises: Determine the target plaintext-ciphertext comparison strategy corresponding to the target service; Based on the test dataset of the ciphertext business processing model, the test dataset of the plaintext business processing model, the ciphertext processing result, and the plaintext processing result of the target plaintext-ciphertext comparison strategy, the plaintext-ciphertext error information corresponding to the target plaintext-ciphertext comparison strategy is obtained.
3. The method as described in claim 2, wherein the target plaintext-ciphertext comparison strategy includes a data association strategy; Accordingly, the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, based on the target plaintext-ciphertext comparison strategy, include: Compare the number of records in each field, the number of missing values, and at least one of the following in the data distribution graph of the test dataset of the encrypted business processing model and the test dataset of the plaintext business processing model.
4. The method as described in claim 2, wherein the target plaintext-ciphertext comparison strategy includes a data preprocessing strategy; Accordingly, the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, based on the target plaintext-ciphertext comparison strategy, include: Determine the target plaintext-ciphertext comparison method corresponding to the data preprocessing strategy; The test datasets of the ciphertext service processing model and the plaintext service processing model are compared based on the target plaintext-ciphertext comparison method.
5. The method as described in claim 4, wherein the target plaintext-ciphertext comparison method includes at least one of sampling comparison, mean comparison, quantile comparison, and coefficient comparison; wherein, The sampling comparison method includes sampling the test dataset of the encrypted service processing model and the test dataset of the plaintext service processing model a preset number of times, and comparing the data based on the sampling results; The mean comparison method includes calculating the ciphertext mean corresponding to the test dataset of the ciphertext service processing model, calculating the plaintext mean corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext mean and the plaintext mean. The quantile comparison method includes calculating the ciphertext quantiles corresponding to the test dataset of the ciphertext service processing model, calculating the plaintext quantiles corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext quantiles with the quantiles. The coefficient comparison method includes determining the ciphertext coefficient information corresponding to the test dataset of the ciphertext service processing model, determining the plaintext coefficient information corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext coefficient information and the plaintext coefficient information.
6. The method as described in claim 2, wherein the target plaintext-ciphertext comparison strategy includes a feature engineering strategy; Accordingly, the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, based on the target plaintext-ciphertext comparison strategy, include: Determine the target feature encoding comparison method corresponding to the feature engineering strategy; The test datasets of the ciphertext service processing model and the plaintext service processing model are compared based on the target feature encoding comparison method.
7. The method as described in claim 6, wherein the target feature encoding comparison method includes at least one of the following: encoding feature quantity comparison method, encoding feature map comparison method, encoding feature category comparison method, and encoding feature statistical information comparison method, wherein... The method for comparing the number of coded features includes obtaining the number of coded features and the name of coded features in the test dataset of the coded service processing model, obtaining the number of coded features and the name of coded features in the test dataset of the plaintext service processing model, comparing the number of coded features and the number of coded features, and comparing the name of coded features and the name of coded features. The encoding feature map comparison method includes obtaining the ciphertext encoding feature map corresponding to the test dataset of the ciphertext service processing model, obtaining the plaintext encoding feature map corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext encoding feature map and the plaintext encoding feature map. The encoding feature category comparison method includes obtaining the ciphertext encoding category corresponding to the test dataset of the ciphertext service processing model, obtaining the plaintext encoding category corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext encoding category and the plaintext encoding category. The encoding feature statistical information comparison method includes obtaining the ciphertext encoding mean, ciphertext encoding variance, and ciphertext encoding quantiles corresponding to the test dataset of the ciphertext service processing model; obtaining the plaintext encoding mean, plaintext encoding variance, and plaintext encoding quantiles corresponding to the test dataset of the plaintext service processing model; comparing the ciphertext encoding mean with the plaintext encoding mean; comparing the ciphertext encoding variance with the plaintext encoding variance; and comparing the ciphertext encoding quantiles with the plaintext encoding quantiles.
8. The method as described in claim 2, wherein the target plaintext-ciphertext comparison strategy includes a model performance comparison strategy; Accordingly, the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, based on the target plaintext-ciphertext comparison strategy, include: Based on the ciphertext processing results, determine the ciphertext model effect information of the ciphertext business processing model; Based on the plaintext processing results, determine the plaintext model effect information of the plaintext service processing model; Compare the ciphertext model effect information with the plaintext model effect information.
9. The method as described in claim 1, wherein determining the training result of the ciphertext business processing model based on the plaintext-ciphertext error information comprises: If the plaintext-ciphertext error information is greater than or equal to a preset error information threshold, the training result of the ciphertext business processing model is determined to be unsuccessful. If the plaintext-ciphertext error information is less than a preset error information threshold, the training result of the ciphertext business processing model is determined to be passed.
10. The method as described in claim 1, wherein ciphertext modeling is performed based on each dataset to be processed to generate a ciphertext business processing model, comprising: Based on privacy computing technology, each dataset to be processed is processed to generate a ciphertext modeling training dataset and a ciphertext modeling test dataset corresponding to each dataset to be processed. An initial ciphertext business processing model is generated by training various ciphertext modeling training datasets using a federated learning approach. The initial encrypted business processing model was tested based on the encrypted modeling test dataset to obtain the test results of the encrypted business processing model. If the test results of the encrypted service processing model meet the test conditions, the initial encrypted service processing model is determined to be an encrypted service processing model.
11. A data processing apparatus, comprising: The acquisition module is configured to acquire at least one pending dataset and one test dataset corresponding to the target business. The generation module is configured to perform ciphertext modeling based on each dataset to be processed, generate a ciphertext business processing model, and perform plaintext modeling based on each dataset to be processed, generate a plaintext business processing model. The input module is configured to input the test dataset into the ciphertext service processing model to obtain the ciphertext processing result, and input the test dataset into the plaintext service processing model to obtain the plaintext processing result; The determination module is configured to obtain plaintext-ciphertext error information based on the test dataset of the ciphertext service processing model, the test dataset of the plaintext service processing model, the ciphertext processing result, and the plaintext processing result, and to determine the training result of the ciphertext service processing model based on the plaintext-ciphertext error information.
12. The apparatus of claim 11, wherein the determining module is further configured to: Determine the target plaintext-ciphertext comparison strategy corresponding to the target service; Based on the test dataset of the ciphertext business processing model, the test dataset of the plaintext business processing model, the ciphertext processing result, and the plaintext processing result of the target plaintext-ciphertext comparison strategy, the plaintext-ciphertext error information corresponding to the target plaintext-ciphertext comparison strategy is obtained.
13. The apparatus of claim 12, wherein the target plaintext-ciphertext comparison strategy includes a data association strategy; The determining module is further configured as follows: Compare the number of records in each field, the number of missing values, and at least one of the following in the data distribution graph of the test dataset of the encrypted business processing model and the test dataset of the plaintext business processing model.
14. The apparatus of claim 12, wherein the target plaintext-ciphertext comparison strategy includes a data preprocessing strategy; The determining module is further configured as follows: Determine the target plaintext-ciphertext comparison method corresponding to the data preprocessing strategy; The test datasets of the ciphertext service processing model and the plaintext service processing model are compared based on the target plaintext-ciphertext comparison method.
15. The apparatus of claim 14, wherein the target plaintext-ciphertext comparison method includes at least one of sampling comparison, mean comparison, quantile comparison, and coefficient comparison; wherein, The sampling comparison method includes sampling the test dataset of the encrypted service processing model and the test dataset of the plaintext service processing model a preset number of times, and comparing the data based on the sampling results; The mean comparison method includes calculating the ciphertext mean corresponding to the test dataset of the ciphertext service processing model, calculating the plaintext mean corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext mean and the plaintext mean. The quantile comparison method includes calculating the ciphertext quantiles corresponding to the test dataset of the ciphertext service processing model, calculating the plaintext quantiles corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext quantiles with the quantiles. The coefficient comparison method includes determining the ciphertext coefficient information corresponding to the test dataset of the ciphertext service processing model, determining the plaintext coefficient information corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext coefficient information and the plaintext coefficient information.
16. The apparatus of claim 12, wherein the target plaintext-ciphertext comparison strategy includes a feature engineering strategy; The determining module is further configured as follows: Determine the target feature encoding comparison method corresponding to the feature engineering strategy; The test datasets of the ciphertext service processing model and the plaintext service processing model are compared based on the target feature encoding comparison method.
17. The apparatus of claim 16, wherein the target feature encoding comparison method includes at least one of the following: encoding feature quantity comparison method, encoding feature map comparison method, encoding feature category comparison method, and encoding feature statistical information comparison method, wherein, The method for comparing the number of coded features includes obtaining the number of coded features and the name of coded features in the test dataset of the coded service processing model, obtaining the number of coded features and the name of coded features in the test dataset of the plaintext service processing model, comparing the number of coded features and the number of coded features, and comparing the name of coded features and the name of coded features. The encoding feature map comparison method includes obtaining the ciphertext encoding feature map corresponding to the test dataset of the ciphertext service processing model, obtaining the plaintext encoding feature map corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext encoding feature map and the plaintext encoding feature map. The encoding feature category comparison method includes obtaining the ciphertext encoding category corresponding to the test dataset of the ciphertext service processing model, obtaining the plaintext encoding category corresponding to the test dataset of the plaintext service processing model, and comparing the ciphertext encoding category and the plaintext encoding category. The encoding feature statistical information comparison method includes obtaining the ciphertext encoding mean, ciphertext encoding variance, and ciphertext encoding quantiles corresponding to the test dataset of the ciphertext service processing model; obtaining the plaintext encoding mean, plaintext encoding variance, and plaintext encoding quantiles corresponding to the test dataset of the plaintext service processing model; comparing the ciphertext encoding mean with the plaintext encoding mean; comparing the ciphertext encoding variance with the plaintext encoding variance; and comparing the ciphertext encoding quantiles with the plaintext encoding quantiles.
18. The apparatus of claim 12, wherein the target plaintext-ciphertext comparison strategy includes a model performance comparison strategy; The determining module is further configured as follows: Based on the ciphertext processing results, determine the ciphertext model effect information of the ciphertext business processing model; Based on the plaintext processing results, determine the plaintext model effect information of the plaintext service processing model; Compare the ciphertext model effect information with the plaintext model effect information.
19. The apparatus of claim 11, wherein the determining module is further configured to: If the plaintext-ciphertext error information is greater than or equal to a preset error information threshold, the training result of the ciphertext business processing model is determined to be unsuccessful. If the plaintext-ciphertext error information is less than a preset error information threshold, the training result of the ciphertext business processing model is determined to be passed.
20. The apparatus of claim 11, wherein the generation module is further configured to: Based on privacy computing technology, each dataset to be processed is processed to generate a ciphertext modeling training dataset and a ciphertext modeling test dataset corresponding to each dataset to be processed. An initial ciphertext business processing model is generated by training various ciphertext modeling training datasets using a federated learning approach. The initial encrypted business processing model was tested based on the encrypted modeling test dataset to obtain the test results of the encrypted business processing model. If the test results of the encrypted service processing model meet the test conditions, the initial encrypted service processing model is determined to be an encrypted service processing model.
21. A computing device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 10.
22. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Precision control and accumulated error eliminating method applied to fully homomorphic encryption
CN108809619A
Cryptographic processing device, cryptographic processing method, and cryptographic processing program
JP7228286B1