Data security protection system and education data security method based on machine learning

By adopting federated learning methods in educational data processing, building a mapping model to optimize parameter transmission and thread number configuration, the problems of high risk of educational data leakage and low training and testing efficiency are solved, and security and efficiency are improved.

CN119961989BActive Publication Date: 2025-08-08PARTY SCHOOL OF THE SICHUAN PROVINCIAL COMMITTEE OF THE COMMUNIST PARTY OF CHINA SICHUAN ADMINISTRATION COLLEGE (SICHUAN LONG MARCH CADRE COLLEGE)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510013525.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-08-08
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

In the process of educational data processing, the prior art has the problem of high risk of educational data leakage and low efficiency in model training and testing.

Method used

Using a federated learning method based on machine learning, model training and testing are distributed to various terminals. By building a mapping model with the number of model parameters transmission and educational data leakage risk level, and a mapping model with the time-consuming and threaded number of model training, the configuration of parameter transmission and threaded number is optimized, the risk of data leakage is reduced and the training and testing efficiency is improved.

Benefits of technology

It effectively reduces the risk of educational data leakage, improves the training and testing efficiency of the global model, and improves security and efficiency through iterative optimization of the number of parameter transmissions and thread configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961989B_ABST
    Figure CN119961989B_ABST
Patent Text Reader

Abstract

The present invention discloses a data security protection system and an education data security method based on machine learning, which relate to the field of data protection. The present invention optimizes the current federated learning model from two aspects: the security risk of education data and the total time consumption in the current federated learning model, thereby improving the security of education data and the training and testing efficiency of the global model; wherein, by establishing a mapping model between the number of model parameters transmitted each time and the risk level of education data leakage, an optimization basis is provided for the subsequent iterative adjustment of the number of model parameters transmitted by each terminal user in the current federated learning model; and then, by establishing a mapping model between the number of transmitted parameters, the number of threads allocated to decryption and model training, and the overall time consumption, and based on this, the mapping model provides an optimization basis for the subsequent iterative adjustment of the efficiency of decryption and training in the current federated learning model and the number of threads allocated to the corresponding number of threads.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data protection processing, and more specifically, relates to a data security protection system and an educational data security method based on machine learning. Background Art

[0002] The Chinese patent with announcement number CN113221161B discloses an information protection method and a readable storage medium in the online education big data scenario, including analyzing the privacy protection capabilities of different teaching business interaction servers, and determining the number of online education course contents to be distributed to each teaching business interaction server based on the number of multiple online education course contents to be output and the privacy protection evaluation results of multiple teaching business interaction servers.

[0003] The processing of educational data generally involves classification or clustering of educational data for analysis; classification or clustering operations are generally performed using corresponding models; and model training requires the support of a large amount of educational data; and in the process of collecting a large amount of educational data, there is a risk of educational data leakage. Summary of the Invention

[0004] In response to the problems in the related technology, the present invention proposes a data security protection (security protection, i.e., security) system and an educational data security method based on machine learning to overcome the above-mentioned technical problems existing in the existing related technology.

[0005] To solve the above technical problems, the present invention is achieved through the following technical solutions:

[0006] The present invention is a method for securing educational data based on machine learning, comprising the following steps:

[0007] S1. Collect the education data corresponding to the current central server and multiple current education data owners and the parameters of the education data training and testing model to obtain the current education data global classification model, the current local education data matrix set and the current final model parameter data matrix;

[0008] S2. Collect data from multiple sets of historical central servers and corresponding multiple education data owners to obtain a historical transmission parameter number matrix, a historical education data leakage risk level data set, a historical decryption thread data set, a historical training and testing thread data set, a total number of transmission model parameters, and a historical time-consuming data set; and construct a final education data leakage risk mapping model and a final model training time-consuming mapping model;

[0009] S3. Optimize the number of model parameters encrypted and transmitted by each current education data owner to the current central server in conjunction with the final education data leakage risk mapping model and the current final model parameter data matrix to obtain the current final transmission parameter number set; collect data on the model parameters encrypted and transmitted by each current education data owner to the current central server to obtain the current total number of transmission parameters, the current number of decryption threads, and the current number of model training threads;

[0010] S4. Optimize the current number of decryption threads and the current number of model training threads in conjunction with the final model training time-consuming mapping model to obtain the current final number of decryption threads and the current final number of model training threads;

[0011] This solution adopts the method of federated learning, which disperses the training and testing process of the model for classifying or clustering educational data on various terminals. Each educational data owner trains and tests the educational data model in the local environment. After the model training and testing are completed, the parameters of the trained model are transmitted to the central model, and the parameters of the central model are updated and iterated. Therefore, there is no need to collect educational data, which greatly reduces the risk of educational data leakage. However, since there is still a risk of parameter leakage in the process of transmitting model parameters, and data thieves will infer the relevant characteristics of the educational data based on the parameters of the model, not all model parameters can be transmitted each time, but only part of the model parameters can be transmitted, and the transmitted model parameters need to be encrypted to ensure the security of the parameters. Based on this, in this solution, by collecting the educational data corresponding to the current central server and multiple current educational data owners and the educational data training The parameters of the training and testing model provide data support for the subsequent optimization of the number of model parameters encrypted and transmitted by each current education data owner to the current central server; by collecting data from multiple groups of historical central servers and corresponding multiple education data owners, data support is provided for the subsequent construction of the final education data leakage risk mapping model and the final model training time mapping model; by constructing the final education data leakage risk mapping model, a model is provided for determining whether the results of the subsequent optimization of the number of model parameters encrypted and transmitted by each current education data owner to the current central server in the current federated learning model meet the requirements; by constructing the final model training time mapping model, a judgment model is provided for the subsequent judgment of whether the central server processing parameter decryption of the current federated learning model and the number of model training threads are set in accordance with the requirements; and then it is determined whether the processing parameter decryption and the number of model training threads need to be optimized.

[0012] Preferably, the S1 comprises the following steps:

[0013] S11. Set the current central server and multiple current education data owners to obtain the current education data owner set. , a i Indicates the setting i Current education data owners, Indicates the total number of current education data owners set; the current central server constructs a model for classifying education data, which is recorded as the current education data global classification model;

[0014] Set multiple education data types to get education data type set , Indicates the setting i Educational data types, Represents the total number of set education data types; according to the education data type set, the education data of each current education data owner in the current education data owner set is set to obtain the current local education data matrix set , b i Indicates the current education data owner i The education data matrix of the current education data owner is as follows:

[0015] ;

[0016] in, b ijk Indicates the current education data owner i The current education data owner j The first k types of educational data, Indicates the current education data owner i The total number of educational datasets owned by the current educational data owner;

[0017] S12. Each current education data owner in the current education data owner set builds a model for classifying its own local education data matrix on the local server to obtain the current education data local classification model set , Indicates the current education data owner i Each current education data owner builds a model on the local server to classify his or her own local education data matrix;

[0018] Set the parameter type of each current education data local classification model in the current education data local classification model set to obtain a model parameter type set , Represents the first local classification model of each current education data i parameter types, c Represents the total number of parameter types of each local classification model of current education data;

[0019] S13. Each current education data owner in the current education data owner set uses its own local education data matrix to train and test its own current education data local classification model on the local server in conjunction with the model parameter type set and the current local education data matrix set; after the training and testing are completed, various types of parameter data of each current education data local classification model are obtained to obtain the current final model parameter data matrix ;as follows,

[0020] ;

[0021] in, Indicates the obtained i The current local classification model of educational data j Types of parameter data;

[0022] Educational data types include structured data, such as students' grades, student status information, and teachers' teaching plans; unstructured data, such as teaching videos, audio, and images; process data, such as classroom interactions, online assignments, and web searches; outcome data, such as grades, levels, and quantities; identity information data, which is data related to personal identity, such as name, date of birth, and contact information; user interaction data, which is data generated when users interact with educational platforms or resources, such as participation indicators, click-through rates, page views, and pop-up rates; inferred content data, which is data that infers the effectiveness of content by analyzing user interactions with content, and system-wide data, such as rosters, grades, disciplinary records, and attendance information; the owners of educational data are the end users who own the educational data, and the subsequent training and parameterization of each local model are performed locally by these end users; the central server is the receiving server to which each end user transmits the parameters of the trained model. After receiving the model parameters transmitted by each end user, the central server calculates the average of these parameters to train the global model and sends the averaged parameters back to each end user; this process is repeated until the global model in the central server converges.

[0023] Preferably, said S2 comprises the following steps:

[0024] S21. Collect multiple groups of historical central servers and corresponding multiple education data owners to obtain a historical central server set and a historical education data owner matrix; set a model for classifying education data for each central server in the historical central server set to obtain a global classification model set for historical education data; set a local classification model for education data for each education data owner in the historical education data owner matrix to obtain a local classification model matrix for historical education data;

[0025] Each education data owner in the historical education data owner matrix cooperates with the historical education data local classification model matrix to train and test its own education data local classification model using its own education data; after the training and testing are completed, the parameter data of each education data local classification model in the historical education data local classification model matrix is obtained according to the model parameter type set to obtain a historical model parameter data set matrix;

[0026] S22, cooperate with the historical model parameter data set matrix to collect the number of parameters encrypted and sent by each education data owner in the historical education data owner matrix to the corresponding historical central server and the corresponding education data leakage risk level data, and obtain the historical transmission parameter number matrix And the corresponding historical education data leakage risk level dataset , Indicates the i The central server of the group history and the corresponding multiple education data owners’ education data leakage risk level data when transmitting model parameters, The total number of central servers representing the collected history and the corresponding multiple educational data owners; as follows,

[0027] ;

[0028] in, Indicates the collected i The central server of the group history and the corresponding multiple education data owners j The number of model parameters encrypted and transmitted by each education data owner to the corresponding central service, d i Indicates the collected i a central server of the group history and a total number of education data owners in the corresponding plurality of education data owners;

[0029] S23, constructing a final education data leakage risk mapping model using the historical transmission parameter number matrix and the historical education data leakage risk level dataset;

[0030] S24. When each education data owner in the historical education data owner matrix encrypts and sends parameters to the central server in the corresponding historical central server set at the same time, the number of threads used for education data decryption in each central server in the historical central server set and the number of threads used for training and testing each education data global classification model in the historical education data global classification model set, the total number of model parameters transmitted, and the total time consumption data for training and testing are collected to obtain a historical decryption thread data set. , historical training and testing thread dataset , the total number of transmission model parameters And historical time-consuming data sets ; 、 、 Respectively represent the collected i The number of threads allocated to education data decryption by the central server of the group history and the corresponding central servers of multiple education data owners, as well as the number of threads used for training and testing of each education data global classification model in the historical education data global classification model set and the total time consumption of training and testing;

[0031] S25, using the historical decryption thread data set, the historical training and testing thread data set, and the historical time-consuming data set to construct a final model training time-consuming mapping model;

[0032] Since the number of parameters transmitted to the central server each time has a great correlation with the possibility of the educational data being inferred, this scheme establishes a mapping model between the number of model parameters transmitted each time and the risk level of educational data leakage, and based on this mapping model, provides an optimization basis for the subsequent iterative adjustment of the number of model parameters transmitted by each terminal user in the current federated learning model; in addition, since the model parameters transmitted by each terminal user are often encrypted, the central server needs to decrypt these model parameters in advance when using them for global model training, and then perform training operations; and the efficiency of decryption and training is highly correlated with the number of threads allocated to the corresponding ones, therefore, this scheme establishes a mapping model between the number of parameters transmitted, the number of threads allocated to decryption and model training, and the overall time consumption, and based on this mapping model, provides an optimization basis for the subsequent iterative adjustment of the efficiency of decryption and training in the current federated learning model and the number of threads allocated to the corresponding ones.

[0033] Preferably, S3 includes the following steps:

[0034] S31, collect the number of model parameters encrypted and transmitted by each current education data owner in the current education data owner set to the current central server, and obtain the current transmission parameter number set , Indicates the current education data owner i The number of model parameters encrypted and transmitted by the current education data owner to the current central server;

[0035] S32, set the leakage risk level threshold; input the current transmission parameter number set into the final education data leakage risk mapping model for mapping, and obtain the current education data leakage risk level data. When the current education data leakage risk level data is greater than or equal to the leakage risk level threshold, the current transmission parameter number set is adjusted until the current education data leakage risk level data is less than the leakage risk level threshold, and the current final transmission parameter number set is obtained. , e i Indicates the adjusted current education data owner concentration i The number of model parameters encrypted and transmitted by the current education data owner to the current central server; otherwise, there is no need to adjust the current set of transmitted parameters, and the current set of transmitted parameters is used as the current final set of transmitted parameters;

[0036] S33. According to the current final transmission parameter number set, the total number of model parameters encrypted and transmitted by each current education data owner in the current education data owner set to the current central server, the number of threads used by the current central server for decryption, and the number of threads used by the current central server for training and testing the global classification model of the current education data are collected to obtain the total number of current transmission parameters. , Current number of decryption threads , Current model training threads ; ;

[0037] By inputting the current set of transmission parameters into the final education data leakage risk mapping model for mapping, the corresponding data leakage risk level data is obtained; by setting the leakage risk level threshold, a quantitative judgment standard is provided for determining whether the current education data leakage risk level data meets the requirements; by collecting the total number of model parameters encrypted and transmitted by each current education data owner to the current central server, the number of threads used by the current central server for decryption, and the number of threads used by the current central server for training and testing the global classification model of the current education data, data support is provided for the subsequent determination of whether the total time consumed by the current central server for model training and testing meets the requirements.

[0038] Preferably, the S4 comprises the following steps:

[0039] S41, input the total number of current transmission parameters, the current number of decryption threads and the current number of model training threads into the final model training time-consuming mapping model for mapping, and obtain the current time-consuming data ;

[0040] S42. Set a current model training test time threshold; when the current time data is greater than or equal to the current model training test time threshold, adjust the current number of decryption threads and the current number of model training threads until the current time data is less than the current model training test time threshold, and obtain the current final number of decryption threads and the current final number of model training threads; otherwise, there is no need to adjust the current number of decryption threads and the current number of model training threads, and use the current number of decryption threads and the current number of model training threads as the current final number of decryption threads and the current final number of model training threads, respectively;

[0041] By inputting the total number of current transmission parameters, the current number of decryption threads, and the current number of model training threads into the final model training time mapping model for mapping, the corresponding time consumption data is obtained, which provides a data basis for subsequent determination of whether the current time consumption data meets the requirements; by setting the current model training test time consumption threshold, a quantitative judgment standard is provided for determining whether the current time consumption data meets the requirements.

[0042] Preferably, adjusting the current number of decryption threads and the current number of model training threads in S42 includes the following steps:

[0043] S421, build thread number to optimize locust population , Indicates the number of threads in the optimized locust population i A locust, Indicates the size of the locust population optimized by the number of threads; sets the maximum number of iterations of the locust population optimized by the number of threads to And the current number of iterations is , respectively recorded as the second maximum number of iterations and the second current number of iterations; the search space dimension of the thread number optimization locust population is 2;

[0044] S422, set the value range of the current number of decryption threads and the current number of model training threads, and obtain the value range of the number of decryption threads [ f 21 , f 22 ] and the training thread number range[ f 31 , f 32 ], f 21 、 f22 Respectively represent the lower limit and upper limit of the current decryption thread number, f 31 、 f 32 Respectively represent the lower limit and upper limit of the number of training threads of the current model;

[0045] The number of threads is set according to the decryption thread number value interval and the training thread number value interval to optimize the initial position of each locust in the locust population, and a second initial position matrix is obtained. ;as follows,

[0046] ;

[0047] in, 、 Respectively represent the number of threads in the locust population i The components of the initial position of the locust in the first and second search space dimensions are calculated as follows:

[0048] ; ;

[0049] Where, rand 2i1 、 rand 2i2 Respectively for 、 Generate a random number between 0 and 1;

[0050] S423, according to the current time-consuming data Setting the number of threads to optimize the fitness function of the locust population ;as follows,

[0051] ;

[0052] S424, start iteration, before iteration, set the second current iteration number Set to 1; in the first round of iteration, the number of threads will be used to optimize the fitness function of the locust population And cooperate with the final model training time-consuming mapping model to calculate the second initial position matrix The fitness value of the initial position of each locust in the third fitness value set is obtained to obtain a third fitness value set; the maximum fitness in the third fitness value set and the corresponding initial position of the locust are respectively used as the third global optimal fitness and the third global optimal position; the second initial position matrix is adjusted according to the third global optimal fitness and the third global optimal position. The initial position of each locust is updated; after the update is completed, the second current iteration number Add 1 and enter the next iteration;

[0053] In each other round of iteration, the number of threads will be used to optimize the fitness function of the locust population The time-consuming mapping model of the final model training is used to calculate the fitness value of the position of each locust in the locust population optimized by the number of threads updated in the previous iteration process to obtain a fourth fitness value set; the maximum fitness in the fourth fitness value set and the corresponding locust position are respectively used as the fourth global optimal fitness and the fourth global optimal position; the position of each locust in the locust population optimized by the number of threads updated in the previous iteration process is updated again according to the fourth global optimal fitness and the fourth global optimal position; after the update is completed, the second current iteration number is set. Add 1 and enter the next iteration;

[0054] S425, when When , stop the iteration and get the second final global optimal position; otherwise, continue to iterate until The second final global optimal position and the total number of current transmission parameters are input into the final model training time-consuming mapping model for mapping to obtain the optimized current time-consuming data;

[0055] When the optimized current time consumption data is less than the current model training test time consumption threshold, the components of the second final global optimal position in the first and second dimensions are used as the current final decryption thread number and the current final model training thread number respectively; otherwise, return to S424 and continue to iterate until the optimized current time consumption data is less than the current model training test time consumption threshold;

[0056] By adopting the locust optimization algorithm, the number of threads used by the current central server for decryption and the number of threads used for training and testing models are iteratively optimized multiple times, and the total time consumed by the entire global model training and testing is used as the fitness function; therefore, as the iterations proceed, the total time consumed by the entire global model training and testing becomes smaller and smaller, thereby improving the efficiency of global model training and testing.

[0057] The present invention also discloses a data security protection system based on machine learning, comprising a current federated model construction module, a first current model data acquisition module, a historical federated model data acquisition module, a first mapping model construction module, a second mapping model construction module, a transmission parameter number optimization module, a second current model data acquisition module, and a thread number optimization module;

[0058] The current federated model construction module is used to set the data and data training and testing models corresponding to the current central server and multiple current data owners, and obtain the current data global classification model, the current local data matrix set, and the current data local classification model set;

[0059] The first current model data acquisition module is used to collect the model parameters after the current data local classification model set trains and tests the data in the current local data matrix set to obtain the current final model parameter data matrix;

[0060] The historical federated model data collection module is used to collect the number of parameters encrypted and sent by multiple groups of historical central servers and corresponding multiple data owners to the central server in the corresponding historical central server each time, the corresponding data leakage risk level data, the number of data decryption threads of each central server, the number of threads for training and testing of each data global classification model, the total number of transmitted model parameters, and the total time consumption data of training and testing, and obtain the historical transmission parameter number matrix, historical data leakage risk level data set, historical decryption thread data set, historical training and testing thread data set, total number of transmission model parameters set, and historical time consumption data set;

[0061] A first mapping model construction module is used to construct a final data leakage risk mapping model using a historical transmission parameter number matrix and a historical data leakage risk level dataset;

[0062] The second mapping model construction module is used to use the historical decryption thread data set, the historical training and testing thread data set, and the historical time-consuming data set to construct the final model training time-consuming mapping model;

[0063] The transmission parameter number optimization module is used to optimize the number of model parameters encrypted and transmitted by each current data owner to the current central server in conjunction with the final data leakage risk mapping model and the current final model parameter data matrix to obtain the current final transmission parameter number set;

[0064] The second current model data acquisition module is used to collect the total number of model parameters encrypted and transmitted by each current data owner to the current central server, the number of threads used by the current central server for decryption, and the number of threads used by the current central server for training and testing the global classification model of the current data, and obtain the total number of currently transmitted parameters, the current number of decryption threads, and the current number of model training threads;

[0065] The thread number optimization module is used to optimize the current number of decryption threads and the current number of model training threads in conjunction with the final model training time-consuming mapping model to obtain the current final number of decryption threads and the current final number of model training threads.

[0066] Preferably, said data comprises educational data.

[0067] The present invention has the following beneficial effects:

[0068] 1. The present invention optimizes the current federated learning model from two aspects: the security risk of educational data and the total time consumption in the current federated learning model. That is, the number of parameters transmitted by each educational data owner to the current central server each time and the number of threads for decryption and training and testing models in the current central server are optimized respectively, thereby improving the security of educational data and the training and testing efficiency of the global model. Among them, by establishing a mapping model between the number of model parameters transmitted each time and the risk level of educational data leakage, an optimization basis is provided for the subsequent iterative adjustment of the number of model parameters transmitted by each terminal user in the current federated learning model. Then, by establishing a mapping model between the number of parameters transmitted, the number of threads allocated to decryption and model training, and the overall time consumption, and based on this, the mapping model provides an optimization basis for the subsequent iterative adjustment of the efficiency of decryption and training in the current federated learning model and the number of threads allocated to the corresponding number of threads.

[0069] 2. In the present invention, by setting the current model training test time threshold, a quantitative judgment standard is provided for determining whether the current time consumption data meets the requirements.

[0070] 3. In the present invention, the number of threads used for decryption and the number of threads used for training and testing the model on the current central server are iteratively optimized multiple times by adopting the locust optimization algorithm, and the total time consumed for the entire global model training and testing is used as the fitness function; therefore, as the iterations proceed, the total time consumed for the entire global model training and testing becomes smaller and smaller, thereby improving the efficiency of global model training and testing.

[0071] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] In order to more clearly illustrate the technical solutions of the embodiments of the invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the invention. For ordinary technicians in this field, they can also obtain drawings based on these drawings without paying any creative work.

[0073] Figure 1 Schematic diagram of the process of the educational data security method based on machine learning of the present invention. DETAILED DESCRIPTION

[0074] The following will clearly and completely describe the technical solutions in the embodiments of the invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0075] In the description of the present invention, it should be understood that the terms "opening", "upper", "lower", "top", "middle", "inside" and the like indicating orientation or positional relationship are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the components or elements referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the invention.

[0076] Example 1

[0077] See also Figure 1 This embodiment is a method for securing educational data based on machine learning, comprising the following steps:

[0078] S1. Collect the education data corresponding to the current central server and multiple current education data owners and the parameters of the education data training and testing model to obtain the current education data global classification model, the current local education data matrix set and the current final model parameter data matrix;

[0079] Said S1 comprises the following steps:

[0080] S11. Set the current central server and multiple current education data owners to obtain the current education data owner set. , a i Indicates the setting i Current education data owners, Indicates the total number of current education data owners set; the current central server constructs a model for classifying education data, which is recorded as the current education data global classification model;

[0081] Set multiple education data types to get education data type set , Indicates the setting i Educational data types, Represents the total number of set education data types; according to the education data type set, the education data of each current education data owner in the current education data owner set is set to obtain the current local education data matrix set , b i Indicates the current education data owner iThe education data matrix of the current education data owner is as follows:

[0082] ;

[0083] in, b ijk Indicates the current education data owner i The current education data owner j The first k types of educational data, Indicates the current education data owner i The total number of educational datasets owned by the current educational data owner;

[0084] S12. Each current education data owner in the current education data owner set builds a model for classifying its own local education data matrix on the local server to obtain the current education data local classification model set , Indicates the current education data owner i Each current education data owner builds a model on the local server to classify his or her own local education data matrix;

[0085] Set the parameter type of each current education data local classification model in the current education data local classification model set to obtain a model parameter type set , Represents the first local classification model of each current education data i parameter types, c Represents the total number of parameter types of each local classification model of current education data;

[0086] S13. Each current education data owner in the current education data owner set uses its own local education data matrix to train and test its own current education data local classification model on the local server in conjunction with the model parameter type set and the current local education data matrix set; after the training and testing are completed, various types of parameter data of each current education data local classification model are obtained to obtain the current final model parameter data matrix ;as follows,

[0087] ;

[0088] in, Indicates the obtained i The current local classification model of educational data j Types of parameter data;

[0089] S2. Collect data from multiple sets of historical central servers and corresponding multiple education data owners to obtain a historical transmission parameter number matrix, a historical education data leakage risk level data set, a historical decryption thread data set, a historical training and testing thread data set, a total number of transmission model parameters, and a historical time-consuming data set; and construct a final education data leakage risk mapping model and a final model training time-consuming mapping model;

[0090] The S2 comprises the following steps:

[0091] S21. Collect multiple groups of historical central servers and corresponding multiple education data owners to obtain a historical central server set and a historical education data owner matrix; set a model for classifying education data for each central server in the historical central server set to obtain a global classification model set for historical education data; set a local classification model for education data for each education data owner in the historical education data owner matrix to obtain a local classification model matrix for historical education data;

[0092] Each education data owner in the historical education data owner matrix cooperates with the historical education data local classification model matrix to train and test its own education data local classification model using its own education data; after the training and testing are completed, the parameter data of each education data local classification model in the historical education data local classification model matrix is obtained according to the model parameter type set to obtain a historical model parameter data set matrix;

[0093] S22, cooperate with the historical model parameter data set matrix to collect the number of parameters encrypted and sent by each education data owner in the historical education data owner matrix to the corresponding historical central server and the corresponding education data leakage risk level data, and obtain the historical transmission parameter number matrix And the corresponding historical education data leakage risk level dataset , Indicates the i The central server of the group history and the corresponding multiple education data owners’ education data leakage risk level data when transmitting model parameters, The total number of central servers representing the collected history and the corresponding multiple educational data owners; as follows,

[0094] ;

[0095] in, Indicates the collected i The central server of the group history and the corresponding multiple education data owners jThe number of model parameters encrypted and transmitted by each education data owner to the corresponding central service, d i Indicates the collected i a central server of the group history and a total number of education data owners in the corresponding plurality of education data owners;

[0096] S23, constructing a final education data leakage risk mapping model using the historical transmission parameter number matrix and the historical education data leakage risk level dataset;

[0097] The S23 includes the following steps:

[0098] S231. Construct a first initial SVM mapping model and set a first training data ratio; divide the historical transmission parameter number matrix and the historical education data leakage risk level data set according to the first training data ratio to obtain a historical transmission parameter number training data matrix, a historical education data leakage risk level training data set, a historical transmission parameter number test data matrix, and a historical education data leakage risk level test data set;

[0099] S232: Setting a first training error threshold; inputting the historical transmission parameter number training data matrix and the historical education data leakage risk level training data set as training data and training labels, respectively, into a first initial SVM mapping model for training; during the training process, when the training error is less than the first training error threshold, stopping the training to obtain a first trained SVM mapping model; otherwise, continuing the training until the training error is less than the first training error threshold;

[0100] S233. Set a first test accuracy threshold; input the historical transmission parameter number test data matrix and the historical education data leakage risk level test data set as test data and test labels respectively into the first trained SVM mapping model for testing; after the test is completed, obtain the first test accuracy; when the first test accuracy is greater than or equal to the first test accuracy threshold, use the first trained SVM mapping model as the final education data leakage risk mapping model; otherwise, return to S232 to continue training the first trained SVM mapping model until the first test accuracy is greater than or equal to the first test accuracy threshold;

[0101] S24. When each education data owner in the historical education data owner matrix encrypts and sends parameters to the central server in the corresponding historical central server set at the same time, the number of threads used for education data decryption in each central server in the historical central server set and the number of threads used for training and testing each education data global classification model in the historical education data global classification model set, the total number of model parameters transmitted, and the total time consumption data for training and testing are collected to obtain a historical decryption thread data set. , historical training and testing thread dataset , the total number of transmission model parameters And historical time-consuming data sets ; 、 、 Respectively represent the collected i The number of threads allocated to education data decryption by the central server of the group history and the corresponding central servers of multiple education data owners, as well as the number of threads used for training and testing of each education data global classification model in the historical education data global classification model set and the total time consumption of training and testing;

[0102] S25, using the historical decryption thread data set, the historical training and testing thread data set, and the historical time-consuming data set to construct a final model training time-consuming mapping model;

[0103] The S25 includes the following steps:

[0104] S251. Construct a second initial SVM mapping model and set a second training data ratio; partition the historical decryption thread data set, the historical training and testing thread data set, the total number of transmission model parameters, and the historical time-consuming data set according to the second training data ratio to obtain a historical decryption thread training data set, a historical training and testing thread training data set, a total number of transmission model parameters training set, a historical time-consuming training data set, a historical decryption thread test data set, a historical training and testing thread test data set, a total number of transmission model parameters test set, and a historical time-consuming test data set;

[0105] S252: Setting a second training error threshold; inputting the historical decryption thread training data set, the historical training test thread training data set, the total number of transmission model parameters training set as training data, and the historical time-consuming training data set as training labels into a second initial SVM mapping model for training; during the training process, if the training error is less than the second training error threshold, stopping the training to obtain a second trained SVM mapping model; otherwise, continuing the training until the training error is less than the second training error threshold;

[0106] S253, setting a second test accuracy threshold; inputting the historical decryption thread test data set, the historical training test thread test data set, the total number of transmission model parameters test set as test data, and the historical time-consuming test data set as test labels into the second trained SVM mapping model for testing; after the test is completed, obtaining a second test accuracy; when the second test accuracy is greater than or equal to the second test accuracy threshold, using the second trained SVM mapping model as the final model training time-consuming mapping model; otherwise, returning to S252 to continue training the second trained SVM mapping model until the second test accuracy is greater than or equal to the second test accuracy threshold;

[0107] S3. Optimize the number of model parameters encrypted and transmitted by each current education data owner to the current central server in conjunction with the final education data leakage risk mapping model and the current final model parameter data matrix to obtain the current final transmission parameter number set; collect data on the model parameters encrypted and transmitted by each current education data owner to the current central server to obtain the current total number of transmission parameters, the current number of decryption threads, and the current number of model training threads;

[0108] S3 includes the following steps:

[0109] S31, collect the number of model parameters encrypted and transmitted by each current education data owner in the current education data owner set to the current central server, and obtain the current transmission parameter number set , Indicates the current education data owner i The number of model parameters encrypted and transmitted by the current education data owner to the current central server;

[0110] S32, set the leakage risk level threshold; input the current transmission parameter number set into the final education data leakage risk mapping model for mapping, and obtain the current education data leakage risk level data. When the current education data leakage risk level data is greater than or equal to the leakage risk level threshold, the current transmission parameter number set is adjusted until the current education data leakage risk level data is less than the leakage risk level threshold, and the current final transmission parameter number set is obtained. , e i Indicates the adjusted current education data owner concentration i The number of model parameters encrypted and transmitted by the current education data owner to the current central server; otherwise, there is no need to adjust the current set of transmitted parameters, and the current set of transmitted parameters is used as the current final set of transmitted parameters;

[0111] In S32, the current set of transmission parameters is adjusted until the current education data leakage risk level data is less than the leakage risk level threshold, and obtaining the current final set of transmission parameters includes the following steps:

[0112] S321. Constructing parameters to optimize locust populations , Indicates the number of parameters to optimize the locust population i A locust, Indicates the size of the locust population optimized by the number of parameters; Set the maximum number of iterations of the parameter number optimization locust population to And the current number of iterations is , respectively recorded as the first maximum number of iterations and the first current number of iterations; the search space dimension of the parameter number optimization locust population is ;

[0113] S322, set the value range of each transmission parameter number in the current transmission parameter number set [ f 11 , f 12 ], f 11 、 f 12 Respectively represent the lower limit and upper limit of the value of each transmission parameter number in the current transmission parameter number set;

[0114] The number of parameters is set according to the value interval of each transmission parameter number in the current transmission parameter number set to optimize the initial position of each locust in the locust population, and a first initial position matrix is obtained. ;as follows,

[0115] ;

[0116] in, Indicates that the number of parameters is optimized in the locust population i The initial position of the locust is j The components in the search space dimension are calculated as follows:

[0117] ;

[0118] Where, ceil represents the rounding function, rand 1ij Indicates that Generate a random number between 0 and 1;

[0119] S323, based on the current education data leakage risk level data Setting the number of parameters to optimize the fitness function of locust population ;as follows,

[0120] ;

[0121] S324, start iteration, before iteration, set the first current iteration number Set to 1; in the first round of iteration, the number of parameters will be used to optimize the fitness function of the locust population And calculate the first initial position matrix in conjunction with the final education data leakage risk mapping model The fitness value of the initial position of each locust in the first fitness value set is obtained to obtain a first fitness value set; the maximum fitness in the first fitness value set and the corresponding initial position of the locust are respectively used as the first global optimal fitness and the first global optimal position; the first initial position matrix is adjusted according to the first global optimal fitness and the first global optimal position. The initial position of each locust is updated; after the update is completed, the first current iteration number Add 1 and enter the next iteration;

[0122] In each other round of iteration, the number of parameters will be used to optimize the fitness function of the locust population And cooperate with the final education data leakage risk mapping model to calculate the fitness value of the position of each locust in the locust population by optimizing the number of parameters updated in the previous iteration process, and obtain a second fitness value set; the maximum fitness in the second fitness value set and the corresponding locust position are respectively used as the second global optimal fitness and the second global optimal position; according to the second global optimal fitness and the second global optimal position, the number of parameters updated in the previous iteration process is optimized and the position of each locust in the locust population is updated again; after the update is completed, the first current number of iterations is Add 1 and enter the next iteration;

[0123] S325, when When , stop the iteration and get the first final global optimal position; otherwise, continue to iterate until until the time; taking the first final global optimal position as the optimized current transmission parameter number set; inputting the optimized current transmission parameter number set into the final education data leakage risk mapping model for mapping, and obtaining the optimized current education data leakage risk level data;

[0124] When the optimized current education data leakage risk level data is less than the leakage risk level threshold, the optimized current transmission parameter number set is used as the current final transmission parameter number set; otherwise, return to S324 and continue to iterate until the optimized current education data leakage risk level data is less than the leakage risk level threshold;

[0125] The locust optimization algorithm can quickly converge to the optimal solution. By simulating the social interactions and movement strategies of locusts, it can effectively explore the search space and reduce the risk of falling into the local optimal solution. By dynamically adjusting the control parameters, it can balance global exploration and local refinement search at different stages of the algorithm, thereby improving the overall performance of the algorithm. Based on the above advantages, this scheme uses the locust optimization algorithm to perform multiple iterative optimizations on the number of model parameters transmitted by each current education data owner to the current central server each time, and uses the education data leakage risk level corresponding to the adjusted number of parameters transmitted each time as the fitness function. Therefore, as the iteration proceeds, the current education data leakage risk level becomes smaller and smaller, thereby ensuring the security of the entire education data processing process.

[0126] S33. According to the current final transmission parameter number set, the total number of model parameters encrypted and transmitted by each current education data owner in the current education data owner set to the current central server, the number of threads used by the current central server for decryption, and the number of threads used by the current central server for training and testing the global classification model of the current education data are collected to obtain the total number of current transmission parameters. , Current number of decryption threads , Current model training threads ; ;

[0127] S4. Optimize the current number of decryption threads and the current number of model training threads in conjunction with the final model training time-consuming mapping model to obtain the current final number of decryption threads and the current final number of model training threads;

[0128] The S4 comprises the following steps:

[0129] S41, input the total number of current transmission parameters, the current number of decryption threads and the current number of model training threads into the final model training time-consuming mapping model for mapping, and obtain the current time-consuming data ;

[0130] S42. Set a current model training test time threshold; when the current time data is greater than or equal to the current model training test time threshold, adjust the current number of decryption threads and the current number of model training threads until the current time data is less than the current model training test time threshold, and obtain the current final number of decryption threads and the current final number of model training threads; otherwise, there is no need to adjust the current number of decryption threads and the current number of model training threads, and use the current number of decryption threads and the current number of model training threads as the current final number of decryption threads and the current final number of model training threads, respectively;

[0131] Adjusting the current number of decryption threads and the current number of model training threads in S42 includes the following steps:

[0132] S421, build thread number to optimize locust population , Indicates the number of threads in the optimized locust population i A locust, Indicates the size of the locust population optimized by the number of threads; sets the maximum number of iterations of the locust population optimized by the number of threads to And the current number of iterations is , respectively recorded as the second maximum number of iterations and the second current number of iterations; the search space dimension of the thread number optimization locust population is 2;

[0133] S422, set the value range of the current number of decryption threads and the current number of model training threads, and obtain the value range of the number of decryption threads [ f 21 , f 22 ] and the training thread number range[ f 31 , f 32 ], f 21 、 f 22 Respectively represent the lower limit and upper limit of the current decryption thread number, f 31 、 f 32 Respectively represent the lower limit and upper limit of the number of training threads of the current model;

[0134] The number of threads is set according to the decryption thread number value interval and the training thread number value interval to optimize the initial position of each locust in the locust population, and a second initial position matrix is obtained. ;as follows,

[0135] ;

[0136] in, 、 Respectively represent the number of threads in the locust population i The components of the initial position of the locust in the first and second search space dimensions are calculated as follows:

[0137] ; ;

[0138] Where, rand 2i1 、 rand 2i2 Respectively for 、 Generate a random number between 0 and 1;

[0139] S423, according to the current time-consuming data Setting the number of threads to optimize the fitness function of the locust population ;as follows,

[0140] ;

[0141] S424, start iteration, before iteration, set the second current iteration number Set to 1; in the first round of iteration, the number of threads will be used to optimize the fitness function of the locust population And cooperate with the final model training time-consuming mapping model to calculate the second initial position matrix The fitness value of the initial position of each locust in the third fitness value set is obtained to obtain a third fitness value set; the maximum fitness in the third fitness value set and the corresponding initial position of the locust are respectively used as the third global optimal fitness and the third global optimal position; the second initial position matrix is adjusted according to the third global optimal fitness and the third global optimal position. The initial position of each locust is updated; after the update is completed, the second current iteration number Add 1 and enter the next iteration;

[0142] In each other round of iteration, the number of threads will be used to optimize the fitness function of the locust population The time-consuming mapping model of the final model training is used to calculate the fitness value of the position of each locust in the locust population optimized by the number of threads updated in the previous iteration process to obtain a fourth fitness value set; the maximum fitness in the fourth fitness value set and the corresponding locust position are respectively used as the fourth global optimal fitness and the fourth global optimal position; the position of each locust in the locust population optimized by the number of threads updated in the previous iteration process is updated again according to the fourth global optimal fitness and the fourth global optimal position; after the update is completed, the second current iteration number is set. Add 1 and enter the next iteration;

[0143] S425, when , stop the iteration and get the second final global best position; otherwise, continue to iterate until The second final global optimal position and the total number of current transmission parameters are input into the final model training time-consuming mapping model for mapping to obtain the optimized current time-consuming data;

[0144] When the optimized current time consumption data is less than the current model training test time consumption threshold, the components of the second final global optimal position in the first and second dimensions are used as the current final decryption thread number and the current final model training thread number respectively; otherwise, return to S424 and continue to iterate until the optimized current time consumption data is less than the current model training test time consumption threshold.

[0145] Example 2

[0146] This embodiment discloses a data security protection system based on machine learning, which can implement the method of the above embodiment, including a current federated model construction module, a first current model data acquisition module, a history education federated model data acquisition module, a first mapping model construction module, a second mapping model construction module, a transmission parameter number optimization module, a second current model data acquisition module, and a thread number optimization module;

[0147] The current federated model building module is used to set the education data and education data training and testing models corresponding to the current central server and multiple current education data owners, and obtain the current education data global classification model, the current local education data matrix set and the current education data local classification model set;

[0148] The first current model data acquisition module is used to collect the model parameters after the current education data local classification model set is trained and tested on the data in the current local education data matrix set to obtain the current final model parameter data matrix;

[0149] The historical education federated model data acquisition module is used to collect multiple groups of historical central servers and the number of parameters encrypted and sent by the corresponding multiple education data owners to the central server concentrated in the corresponding historical central server each time, the corresponding education data leakage risk level data, the number of threads for education data decryption of each central server, the number of threads for training and testing of each education data global classification model, the total number of transmitted model parameters and the total time consumption data for training and testing, and obtain the historical transmission parameter number matrix, the historical education data leakage risk level data set, the historical decryption thread data set, the historical training and testing thread data set, the total number of transmission model parameters set and the historical time consumption data set;

[0150] The first mapping model construction module is used to construct a final education data leakage risk mapping model using the historical transmission parameter number matrix and the historical education data leakage risk level dataset;

[0151] The second mapping model construction module is used to construct a final model training time-consuming mapping model using a historical decryption thread data set, a historical training and testing thread data set, and a historical time-consuming data set;

[0152] The transmission parameter number optimization module is used to optimize the number of model parameters encrypted and transmitted by each current education data owner to the current central server in conjunction with the final education data leakage risk mapping model and the current final model parameter data matrix to obtain the current final transmission parameter number set;

[0153] The second current model data acquisition module is used to collect the total number of model parameters encrypted and transmitted by each current education data owner to the current central server, the number of threads used by the current central server for decryption, and the number of threads used by the current central server for training and testing the global classification model of the current education data, and obtain the total number of currently transmitted parameters, the current number of decryption threads, and the current number of model training threads;

[0154] The thread number optimization module is used to optimize the current number of decryption threads and the current number of model training threads in conjunction with the final model training time-consuming mapping model to obtain the current final number of decryption threads and the current final number of model training threads.

[0155] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0156] The preferred embodiments of the invention disclosed above are intended only to help illustrate the invention. These preferred embodiments do not exhaust all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.

Claims

1. An educational data security method based on machine learning, characterized in that: The following steps are involved: S1. Collect the education data corresponding to the current central server and multiple current education data owners and the parameters of the education data training and testing model to obtain the current education data global classification model, the current local education data matrix set and the current final model parameter data matrix; S2. Collect data from multiple sets of historical central servers and corresponding multiple education data owners to obtain a historical transmission parameter number matrix, a historical education data leakage risk level data set, a historical decryption thread data set, a historical training and testing thread data set, a total number of transmission model parameters, and a historical time-consuming data set; And build the final education data leakage risk mapping model and the final model training time mapping model; S3. Optimize the number of model parameters encrypted and transmitted by each current education data owner to the current central server in conjunction with the final education data leakage risk mapping model and the current final model parameter data matrix to obtain the current final transmission parameter number set; collect data on the model parameters encrypted and transmitted by each current education data owner to the current central server to obtain the current total number of transmission parameters, the current number of decryption threads, and the current number of model training threads; S4. Optimize the current number of decryption threads and the current number of model training threads in conjunction with the final model training time-consuming mapping model to obtain the current final number of decryption threads and the current final number of model training threads; The S2 comprises the following steps: S21, collecting parameter data of the education data local classification model from multiple sets of historical central servers and corresponding multiple historical education data owners to obtain a historical model parameter data set matrix; S22. Collect the number of parameters encrypted and sent by each historical education data owner to the corresponding historical central server and the corresponding education data leakage risk level data in conjunction with the historical model parameter data set matrix, and obtain a historical transmission parameter number matrix and a corresponding historical education data leakage risk level data set; S23, constructing a final education data leakage risk mapping model using the historical transmission parameter number matrix and the historical education data leakage risk level dataset; S24. When each owner of historical education data encrypts and sends parameters to the central server in the corresponding historical central server set at the same time, the number of threads used for education data decryption in each historical central server, the number of threads used for training and testing each education data global classification model in the historical education data global classification model set, the total number of transmitted model parameters, and the total time consumption data for training and testing are collected to obtain a historical decryption thread data set, a historical training and testing thread data set, a total number of transmitted model parameters, and a historical time consumption data set; S25, using the historical decryption thread data set, the historical training and testing thread data set, and the historical time-consuming data set to construct a final model training time-consuming mapping model; The final education data leakage risk mapping model adopts the SVM model; the final model training time-consuming mapping model adopts the SVM model.

2. The educational data security method based on machine learning according to claim 1 is characterized in that: Said S1 comprises the following steps: S11. Set a current central server and multiple current education data owners; set multiple education data types to obtain an education data type set; collect education data of each current education data owner according to the education data type set to obtain a current local education data matrix set; S12. Each current education data owner constructs a model for classifying its own local education data matrix on a local server to obtain a current education data local classification model set; sets a parameter type for each current education data local classification model in the current education data local classification model set to obtain a model parameter type set; S13. Each current education data owner uses his own local education data matrix on the local server to train and test his own current education data local classification model in conjunction with the model parameter type set and the current local education data matrix set; after the training and testing are completed, various types of parameter data of each current education data local classification model are obtained to obtain the current final model parameter data matrix.

3. The educational data security method based on machine learning according to claim 1, characterized in that: S3 includes the following steps: S31. Collect the number of model parameters encrypted and transmitted by each current education data owner to the current central server to obtain a current set of transmitted parameters; S32, setting a leakage risk level threshold; inputting the current transmission parameter number set into the final education data leakage risk mapping model for mapping, to obtain current education data leakage risk level data; When the current education data leakage risk level data is greater than or equal to the leakage risk level threshold, the current transmission parameter number set is adjusted to obtain the current final transmission parameter number set; otherwise, the current transmission parameter number set does not need to be adjusted, and the current transmission parameter number set is used as the current final transmission parameter number set; S33. According to the current final set of transmission parameters, the total number of model parameters encrypted and transmitted by each current education data owner to the current central server, the number of threads used by the current central server for decryption, and the number of threads used by the current central server for training and testing the global classification model of the current education data are collected to obtain the current total number of transmission parameters, the current number of decryption threads, and the current number of model training threads.

4. The method for securing educational data based on machine learning according to claim 3, wherein: In S32, the locust optimization algorithm is used to adjust the current set of transmission parameters.

5. The educational data security method based on machine learning according to claim 1 is characterized in that: The S4 comprises the following steps: S41, inputting the total number of current transmission parameters, the current number of decryption threads, and the current number of model training threads into the final model training time-consuming mapping model for mapping to obtain current time-consuming data; S42. Set the current model training test time threshold; when the current time data is greater than or equal to the current model training test time threshold, adjust the current number of decryption threads and the current number of model training threads until the current time data is less than the current model training test time threshold, and obtain the current final number of decryption threads and the current final number of model training threads; otherwise, there is no need to adjust the current number of decryption threads and the current number of model training threads, and use the current number of decryption threads and the current number of model training threads as the current final number of decryption threads and the current final number of model training threads, respectively.

6. The method for securing educational data based on machine learning according to claim 5, wherein: Adjusting the current number of decryption threads and the current number of model training threads in S42 includes the following steps: S421, constructing a locust population optimized by the number of threads; setting the maximum number of iterations of the locust population optimized by the number of threads to be And the current number of iterations is , respectively recorded as the second maximum number of iterations and the second current number of iterations; S422, setting the value intervals of the current number of decryption threads and the current number of model training threads to obtain a value interval of the number of decryption threads and a value interval of the number of training threads; setting the number of threads to optimize the initial position of each locust in the locust population according to the value interval of the number of decryption threads and the value interval of the number of training threads, to obtain a second initial position matrix; S423, setting the number of threads according to the current time-consuming data to optimize the fitness function of the locust population; S424, start iteration; in each round of iteration, the fitness function of the locust population is optimized by the number of threads, and the time-consuming mapping model of the final model training is used to calculate the fitness value of the position of each locust in the locust population optimized by the number of threads updated in the previous round of iteration, and the position of each locust in the locust population optimized by the number of threads updated in the previous round of iteration is updated again; S425, when When , stop the iteration and get the second final global optimal position; otherwise, continue to iterate until The second final global optimal position and the total number of current transmission parameters are input into the final model training time-consuming mapping model for mapping to obtain the optimized current time-consuming data; When the optimized current time consumption data is less than the current model training test time consumption threshold, the components of the second final global optimal position in the first and second dimensions are used as the current final decryption thread number and the current final model training thread number respectively; otherwise, return to S424 and continue to iterate until the optimized current time consumption data is less than the current model training test time consumption threshold.

7. A data security protection system, characterized by: Used to implement the educational data security method based on machine learning as described in any one of claims 1-6, comprising The current federated model building module is used to obtain the current data global classification model, the current local data matrix set, and the current data local classification model set; A first current model data acquisition module is used to obtain the current final model parameter data matrix; The historical federated model data collection module is used to obtain the historical transmission parameter number matrix, the historical data leakage risk level data set, the historical decryption thread data set, the historical training and testing thread data set, the total number of transmission model parameters, and the historical time consumption data set; A first mapping model construction module is used to construct a final data leakage risk mapping model using a historical transmission parameter number matrix and a historical data leakage risk level dataset; The second mapping model construction module is used to use the historical decryption thread data set, the historical training and testing thread data set, and the historical time-consuming data set to construct the final model training time-consuming mapping model; A transmission parameter number optimization module is used to obtain the current final transmission parameter number set; The second current model data acquisition module is used to obtain the total number of current transmission parameters, the current number of decryption threads and the current number of model training threads; The thread number optimization module is used to obtain the current final number of decryption threads and the current final number of model training threads.

Citation Information

Patent Citations

  • Information protection methods and readable storage media in the context of big data in online education

    CN113221161B

  • Cluster federated learning method and device fusing parameter optimization

    CN116362329A

  • Systems and methods for securing electronic data with embedded security engines

    WO2017156414A2