Data security processing and remote proof model protection system and method for machine learning

By integrating user data through data collection and preprocessing modules, filling missing data with the Pearson correlation coefficient and FP-Growth algorithm, and verifying machine learning models using trusted metrics, the shortcomings of machine learning platforms in data processing and security are addressed, achieving more efficient data security processing and model protection.

CN120316835BActive Publication Date: 2025-09-26BEIJING INFORMATION SCI & TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510402264.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-09-26
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

Existing machine learning platforms have difficulty meeting complex data processing requirements when dealing with large-scale, multi-source, and heterogeneous data, and there are data security and model security issues. In particular, the risks of data leakage, tampering, and unauthorized access are high in multi-user shared platforms.

Method used

The data acquisition module is used to integrate user data in different formats, and the data preprocessing module is used to fill in missing data and balance data. The Pearson correlation coefficient and FP-Growth algorithm are combined to mine association rules to fill in data. The integrity of the machine learning model is calculated using trustworthy metrics, and the integrity of the modeling process and results is verified through remote proof technology.

Benefits of technology

It improves the data processing capabilities of the machine learning platform, ensures the security and reliability of data and models, prevents data tampering and model attacks, and achieves the integrity and credibility of the machine learning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316835B_ABST
    Figure CN120316835B_ABST
Patent Text Reader

Abstract

The present invention discloses a data security processing and remote attestation model protection system and method for machine learning, belonging to the fields of computer program analysis technology, information security, and data services. By combining the Pearson correlation coefficient, FP-Growth algorithm, and cosine similarity to fill missing values ​​in discrete data, SMOTE-based minority group oversampling, and constructing a measurement engine, the system inserts stubs into the machine learning model program code to complete the integrity measurement of the modeling process and modeling results. The present invention adopts the above-mentioned data security processing and remote attestation model protection system and method for machine learning, traversing the control flow graph of the static model program code through a breadth-first search method to calculate the trusted execution process measurement value of each sub-process, and comparing each measurement value with the trusted measurement value to verify the measurement result, thereby ensuring that the code and results during the modeling process have not been tampered with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer program analysis technology, information security and data service fields, and in particular to data security processing and remote proof model protection systems and methods for machine learning. Background Art

[0002] In today's digital age, machine learning and artificial intelligence technologies are driving the development of various industries at an unprecedented pace. Whether in e-commerce, healthcare, financial services, or intelligent manufacturing, machine learning platforms have become indispensable tools for businesses and research institutions. However, with the continuous expansion of data scale and the increasing complexity of models, the data processing solutions in existing machine learning platforms are mostly basic and cannot meet the needs of more advanced and complex data processing. In addition, data security and model security issues are becoming increasingly prominent, becoming a major obstacle to the widespread application of machine learning platforms.

[0003] Existing machine learning platforms typically only provide basic data preprocessing capabilities, such as data cleaning based on medians and means, feature extraction, and data transformation. While these capabilities can meet the needs of some simple tasks, they are insufficient when processing large-scale, multi-source, and heterogeneous data. During data transmission and storage, the risks of data leakage, tampering, and unauthorized access always exist. Data security issues are particularly prominent in environments where multiple users share a platform. Malicious users may compromise the accuracy and reliability of models and data through methods such as reverse engineering, model theft, and model attacks. To address these challenges, it is necessary to develop more advanced and powerful data preprocessing methods and introduce machine learning model protection technologies based on remote attestation to ensure the security and reliability of the machine learning process and model results. Summary of the Invention

[0004] The purpose of the present invention is to provide a data security processing and remote attestation model protection system and method for machine learning. It can propose a data security preprocessing solution for the machine learning platform, and combine it with remote attestation technology to protect the integrity and credibility of the machine learning process and model results, so as to solve the shortcomings of existing machine learning platforms in data processing capabilities and security in the machine learning modeling process.

[0005] To achieve the above objectives, the present invention provides a data security processing and remote attestation model protection system for machine learning, including the following modules:

[0006] Data collection module: collects user data and integrates user data in different formats into a unified data format;

[0007] Data preprocessing module: automatically fill or delete missing data, and fill in the mean, median and discrete data based on correlation coefficient;

[0008] Machine learning process and model security protection processing module: Modeling is performed according to the selected modeling method;

[0009] Local storage module: used to upload and test the name and parameters of the machine learning model.

[0010] The present invention also provides a method for processing machine learning data security and a remote proof model protection system. The data preprocessing module mines association rules in the data and fills in missing data by combining the Pearson correlation coefficient, FP-Growth algorithm, and cosine similarity. The steps are as follows:

[0011] S1. Mining frequent itemsets in non-missing data sources and calculating the correlation between data attributes, and calculating the overall correlation within the mined items;

[0012] S2. Select the filling item based on the similarity between the non-missing predecessor item of the item where the missing data is located and the complete data mining item; if the similarity of the filling items is consistent, the weighted confidence is used to further select the filling rule.

[0013] Preferably, the data acquisition module acquires a data set and analysis data thereof, the analysis data including data type, data distribution and statistical data; the data acquisition module performs data analysis and detects whether there are missing values ​​in the data analysis results.

[0014] Preferably, the row processing tool of the data preprocessing module is used to process unbalanced data. When one category of data is significantly less than another category, the data is divided into test set data and training set data, wherein the training set data undergoes a machine learning process; then, through the synthetic minority sample oversampling technology, the number of samples of different categories is balanced in the training set data; data combinations are generated according to the possible values ​​of each column, and the values ​​of numerical columns are determined by the starting point and step size.

[0015] Preferably, the security protection processing module of the machine learning process and model verifies the measurement results by verifying the signature of the measurement results and comparing each measurement value with the trusted measurement value to ensure that the code and results are not tampered with during the modeling process.

[0016] Preferably, the method for calculating the trust metric value is: traversing the control flow graph by breadth-first search, exploring each node layer by layer, performing offline static analysis on the control flow graph of the model program code, and calculating the trusted execution process metric value of each sub-process, that is, the cumulative hash value.

[0017] Therefore, the present invention adopts the above-mentioned machine learning data security processing and remote proof model protection system and method to propose a data security preprocessing solution for the machine learning platform, and combines it with remote proof technology to protect the integrity and credibility of the machine learning process and model results, so as to solve the shortcomings of existing machine learning platforms in data processing capabilities and security in the machine learning modeling process.

[0018] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a complete flow chart of the data processing method for machine learning and the model protection method based on remote attestation according to the embodiment of the data security processing and remote attestation model protection system and method of machine learning of the present invention;

[0020] Figure 2 2. A diagram of a method for calculating a trust metric value set in accordance with an embodiment of a system and method for secure data processing and remote proof model protection for machine learning according to the present invention;

[0021] Figure 3 Schematic diagram of calculating model static code metrics in accordance with an embodiment of the data security processing and remote proof model protection system and method for machine learning of the present invention;

[0022] Figure 4 This is a detailed flow chart of the data security processing and remote proof model protection system and method embodiment of the machine learning present invention, which uses a trusted measurement engine to calculate dynamic execution process measurement values. DETAILED DESCRIPTION

[0023] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0024] Unless otherwise defined, technical or scientific terms used in the present invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.

[0025] Example 1

[0026] Figure 1 The following is a complete flow chart of the data processing method for machine learning and the model protection method based on remote proof. Figure 1 As shown, Figure 1 This paper describes in detail the data preprocessing methods for the machine learning platform and the process of measuring and verifying the static code, execution process, and execution results of specific programs during the modeling process. The data security processing and remote attestation model protection methods for machine learning include four parts:

[0027] Part I: Data Collection and Preprocessing

[0028] The data collection module collects user data and consolidates it into a unified format. This module acquires a unified data set and analyzes the data. The analysis includes data type, distribution, and statistical data. Parameters obtained from the analysis include the valid value, maximum value, minimum value, median, mode, composite value, variance, length, width, category, and the presence of missing values. This analysis utilizes algorithms from related technologies and automated analysis tools.

[0029] The steps of data preprocessing are as follows:

[0030] S1. Fill in missing data:

[0031] The missing data in user data is handled by automatic filling or deletion, and the average value, median and discrete data based on correlation coefficient can be filled.

[0032] The method of filling discrete data based on correlation coefficient is as follows:

[0033] 1) Divide the source data data into the non-missing data part and the missing data part, calculate the Pearson correlation coefficient, and obtain the coefficient correlation matrix P;

[0034] 2) Apriori + distance metric: Generate frequent 1-item sets and sort them using cosine similarity. Generate frequent item sets of higher dimensions and perform pruning in combination with distance metric.

[0035] 3) FP-Growth + Distance Metric: Build an FP tree and sort the items by distance metric. Recursively mine the FP tree and prune items with low support and large distance.

[0036] 4) Combine rules: Generate association rules and calculate the similarity between the rules and the target itemset. Prune rules with low confidence and large distance. During the FP-tree construction and mining process, prune itemsets with low support and large distance to the target itemset to reduce the size of the FP-tree.

[0037] 5) Filling based on distance similarity: Calculate the cosine similarity between each of the mined frequent itemsets and the missing data item. Calculate the cosine similarity CS using the non-missing portion of the missing data item and the rules mined from the complete dataset. Select the frequent itemset with the highest similarity to fill the missing data. If there are multiple frequent itemsets with high similarity, consider weighted filling or multi-rule filling.

[0038] The cosine similarity They are the complete data rule and the part of the item set of missing records that does not contain missing items;

[0039] 6) Get the padded data set Data.

[0040] S2. Processing of imbalanced data:

[0041] To handle imbalanced data, when one class of data is significantly less than another, the system divides the data into a test set and a training set. Using synthetic minority oversampling, the system balances the number of samples from different classes in the training data. To combine data, the system generates data combinations based on the possible values ​​for each column. For numeric columns, the possible values ​​are determined by the starting point and step size.

[0042] The method for processing unbalanced data is the improved SMOTE algorithm:

[0043] Input: Minority class sample set P in the training set data = {p1,p2,…p n}, the number of samples N required for synthesis of each class; the number of nearest neighbors K, the number of nearest neighbors involved in synthesis D (D <K)

[0044] Output: synthetic sample set S.

[0045] The SMOTE algorithm process is as follows:

[0046] Step 1: Find the K nearest neighbor set

[0047] Step 2: From knn i Randomly select D neighbor samples from

[0048] Step 3: Select D random real numbers in

[0049] Step 4: Calculate the selected neighbors With sample p a Vector difference The calculation method is in, is a randomly selected D nearest neighbor sample knn a The nth item in P a is an item in the minority class sample set P.

[0050] Step 5: Calculate the synthetic sample vector:

[0051] Step 6: Count the generated new samples into the set.

[0052] Part II: Calculation of Credibility Metric Sets

[0053] First, the machine learning model program code is divided into blocks, and a stub is inserted before each code module. The control flow graph is traversed by breadth-first search (BFS), and each node is explored layer by layer. The control flow graph of the model program code is statically analyzed offline, and the legal execution process metric value (i.e., cumulative hash value) of each sub-process is calculated, such as Figure 2 shown.

[0054] The algorithm steps are as follows:

[0055] Input: Contains the control flow graph of the model program and the start and end node information of each sub-process;

[0056] Output: The trusted execution process measurement value (cumulative hash value) corresponding to each sub-process.

[0057] Step 1: Initialization: The trust metric value set Z is initialized to an empty set, and the hash value is initialized to 0, indicating the initial cumulative hash value.

[0058] Step 2: Traverse all child nodes of the same level node 1 and calculate the new cumulative hash value of the current node, h2 = H (ID i ,h1); where h1 represents the hash value obtained by hashing the ID of the first program execution module, h2 represents the hash value obtained by cumulative hashing the ID of the second program execution module and h1, and H represents the hash calculation, ID i Indicates the ID of the i-th program execution module;

[0059] Step 3: Traverse all child nodes of the same level node 2 and calculate the new cumulative hash value for the current node, h3 = H (ID i ,h2); where h3 represents the hash value obtained by performing cumulative hash calculation on the ID of the third program execution module and h2.

[0060] Step 4: If the current node is not the exit node of the sub-process, keep the current cumulative hash value unchanged and continue to traverse all child nodes of the same-level node; if the child node is the exit node of the sub-process, add the current cumulative hash value to the trust metric value set Z.

[0061] Step 5: After all nodes have been traversed, the trust metric value set Z is returned.

[0062] Part III: Model Static Code Metrics Calculation Process

[0063] like Figure 3 As shown, when the machine learning model program starts to be executed, the measurement engine initializes the machine learning model program code state after the instrumentation, and measures the static program code to form a static code measurement value.

[0064] Part 4: Modeling process and verification of modeling process and results

[0065] First, the machine learning model program code is divided into blocks, and stubs are inserted before each code module. Then the modeling process is carried out, and the integrity of the modeling process is measured through the measurement engine. Finally, the measurement results are verified by verifying the signature of the measurement results and comparing each measurement value with the trusted measurement value to ensure that the code and results have not been tampered with during the modeling process.

[0066] Figure 4 This is the process of using a trusted measurement engine to measure the integrity of the machine learning modeling process, and further explains the measurement execution process. Figure 4 As shown:

[0067] 1) The jump instruction of the instrumentation module controls the program module jump, executes the corresponding program module, and triggers the control module path trigger, sending the ID corresponding to the current program module to the measurement engine, and then continues to execute subsequent instructions. After receiving the program module ID, the measurement engine performs a cumulative hash calculation. This cumulative hash value represents the measurement value of the code execution process.

[0068] 2) After the machine learning model program is executed, the execution result (model parameters and name) and the one-time random number (Nonce) are returned to the measurement engine. The measurement engine ends the cumulative hash calculation and forms the execution process measurement value F2. The calculation method of F2 is the same as Figure 2 Each trusted metric value in the trusted metric value set is calculated in the same way, and then the model parameters and name are hashed to form the execution result metric value M; finally, the measurement engine hashes the static code metric value F1, the execution process metric value F2, the execution result M, and the Nonce, and signs the calculation result. The signature, F1, F2, and M form the proof material and are encrypted with the trusted measurement engine private key and sent to the verification end.

[0069] The verification process on the verification end is as follows:

[0070] 1) Signature verification:

[0071] The signature is decrypted using the public key of the trusted measurement engine to obtain a hash value.

[0072] Use the Nonce of this request saved by the prover, the static code measurement value F1 in the proof material, the execution process measurement value F2, and the execution result M to perform a hash calculation to obtain another hash value.

[0073] If the two hash values ​​are identical, the static code metric F1, the execution process metric F2, and the execution result M are authentic and valid. Otherwise, it means the metric itself is incorrect and no further verification is necessary. This prevents attackers from using old metric values ​​to launch replay attacks or impersonate the metric engine to provide forged metric results.

[0074] 2) Verification of static code metrics:

[0075] Calculate the hash value of the instrumented executable file corresponding to the machine learning model. Compare the static code metrics in the proof material with the hash value of the executable file. If they match, the machine learning model's static code has not been tampered with; otherwise, it has been tampered with. This prevents attackers from tampering with the executable file before the model is executed, thereby changing its subsequent behavior and thus tampering with the machine learning training process.

[0076] 3) Validation of execution process metrics:

[0077] The execution process measurement values ​​in the measurement results are compared with the pre-calculated trusted execution measurement value set Z to ensure the integrity of the machine learning model program code execution process.

[0078] If the execution process metric value is the same as any trustworthy metric value in set Z, it indicates that the execution process of the machine learning model program code has not been tampered with; otherwise, it indicates that it has been tampered with. This prevents attackers from exploiting software vulnerabilities to modify control data or key constraint data during model program execution, thereby changing its runtime behavior and achieving data tampering.

[0079] 4) Verification of the execution result measurement value:

[0080] The received execution result M is hashed and compared with the decrypted execution result measurement value to verify the integrity of the machine learning model program execution result.

[0081] If the two are identical, it means the execution result has not been tampered with during the process of the machine learning platform sending the execution result to the verifier; otherwise, it means it has been tampered with. This prevents attackers from tampering with data after the machine learning model program code is executed.

[0082] Therefore, the present invention adopts the above-mentioned machine learning data security processing and remote proof model protection system and method to propose a data security preprocessing solution for the machine learning platform, and combines it with remote proof technology to protect the integrity and credibility of the machine learning process and model results, so as to solve the shortcomings of existing machine learning platforms in data processing capabilities and security in the machine learning modeling process.

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A machine learning data security processing and remote proof model protection system, characterized by: Includes the following modules: Data collection module: collects user data and integrates user data in different formats into a unified data format; Data preprocessing module: automatically fill or delete missing data, and fill in the mean, median and discrete data based on correlation coefficient; Machine learning process and model security protection processing module: Modeling is performed according to the selected modeling method; Local storage module: used to upload and test the name and parameters of the machine learning model; The data security processing and remote proof model protection system for machine learning are as follows: The data preprocessing module mines association rules in the data and fills in missing data by combining the Pearson correlation coefficient, FP-Growth algorithm, and cosine similarity. The steps are as follows: S1. Mining frequent itemsets in non-missing data sources and calculating the correlation between data attributes, and calculating the overall correlation within the mined items; S2. Select the filling item based on the similarity between the non-missing predecessor item of the item where the missing data is located and the complete data mining item; if the similarity of the filling items is consistent, further select the filling rule using the weighted confidence; The method of filling discrete data based on correlation coefficient is as follows: 1) Divide the source data data into the non-missing data part and the missing data part, calculate the Pearson correlation coefficient, and obtain the coefficient correlation matrix P; 2) Apriori + distance metric: Generate frequent 1-item sets and sort them using cosine similarity; generate frequent item sets of higher dimensions and perform pruning in combination with distance metric; 3) FP-Growth + distance metric: Build an FP tree, sort items by distance metric, recursively mine the FP tree, and prune items with low support and large distance; 4) Combining rules: Generate association rules, calculate the similarity between the rules and the target itemset, and prune rules with low confidence and large distance. During the construction and mining of the FP tree, prune itemsets with low support and large distance from the target itemset, thereby reducing the size of the FP tree. 5) Filling based on distance similarity: In the mined frequent itemsets, calculate the cosine similarity between each itemset and the missing data item, and use the non-missing part of the missing data item to calculate the cosine similarity with the rules mined from the complete data set , select the frequent item set with the highest similarity to fill the missing data. If there are multiple frequent item sets with high similarity, perform weighted filling or multi-rule filling; The cosine similarity , 、 They are the complete data rule and the part of the item set of missing records that does not contain missing items; 6) Get the padded data set Data.

2. The machine learning data security processing and remote attestation model protection system according to claim 1, characterized in that: The data acquisition module acquires a data set and its analysis data, and the analysis data includes data type, data distribution and statistical data; the data acquisition module performs data analysis and detects whether there are missing values ​​in the data analysis results.

3. The machine learning data security processing and remote attestation model protection system according to claim 1, characterized in that: The row processing tool of the data preprocessing module is used to process unbalanced data. When one category of data is significantly less than another, the data is divided into test set data and training set data, where the training set data undergoes a machine learning process. Then, through the synthetic minority sample oversampling technology, the number of samples of different categories in the training set data is balanced. Data combinations are generated based on the possible values ​​of each column. The values ​​of numerical columns are determined by the starting point and step size.

4. The machine learning data security processing and remote attestation model protection system according to claim 1, characterized in that: The security protection processing module of the machine learning process and model verifies the measurement results by verifying the signature of the measurement results and comparing each measurement value with the trusted measurement value to ensure that the code and results are not tampered with during the modeling process.

5. The machine learning data security processing and remote attestation model protection system according to claim 4, characterized in that: The method for calculating the trustworthy metric value is as follows: traverse the control flow graph by breadth-first search, explore each node layer by layer, perform offline static analysis on the control flow graph of the model program code, and calculate the trustworthy execution process metric value of each sub-process, that is, the cumulative hash value.

Citation Information

Patent Citations

  • Data credibility verification method and device in data element scene

    CN117113299A

  • Systems and methods for the use of digital agents

    US20250016520A1