Data protection model training and data protection method and device, and storage medium
Patent Information
- Application Number
- CN202211089109.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-09-07
AI Technical Summary
[0003]本说明书实施例提供一种数据保护模型训练及数据保护方法、装置以及存储介质,可以解决相关技术中数据保护模型的性能较差的技术问题
[0011]本说明书一些实施例提供的技术方案带来的有益效果至少包括:
Smart Images

Figure CN116150774B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of computer information security technology, and in particular to a data protection model training and data protection method, apparatus and storage medium. Background Technology
[0002] Artificial intelligence technology has developed rapidly in recent years and has been gradually applied to various daily scenarios, such as self-service payment, automatic identity verification, and information collection. When providing automatic user services through various edge devices, it is necessary to collect, transmit, process, and store user information. User information contains a large amount of private data. Therefore, in order to avoid the leakage of users' personal privacy information, it is necessary to manage and protect user information efficiently. Summary of the Invention
[0003] This specification provides a data protection model training method, apparatus, and storage medium, which can solve the technical problem of poor performance of data protection models in related technologies.
[0004] Firstly, embodiments of this specification provide a data protection model training method, the method comprising: Obtain at least two original sample data, input each original sample data into an initial network model, and obtain the sample output data corresponding to each original sample data, wherein the initial network model is constructed based on a preset protection function; Obtain the standard output data after the preset protection function processes the original data of each sample, and obtain the first distillation loss based on the standard output data and the output data of each sample. Calculate the first correlation between each standard output data and the second correlation between each sample output data, and obtain the second distillation loss based on the first correlation and the second correlation. A first loss function is constructed based on the first distillation loss and the second distillation loss. The initial network model is then trained based on the first loss function to obtain a first data protection model.
[0005] Secondly, embodiments of this specification provide a data protection method, which includes: In response to a data encryption request, the target original data is encrypted based on a data encryption model to obtain the target encrypted data corresponding to the target original data; In response to the data decryption request, the target encrypted data is decrypted based on the data decryption model to obtain the target decrypted data corresponding to the original target data; Wherein, the data encryption model or the data decryption model is a data protection model trained by the data protection model training method described in any of the above embodiments of the specification.
[0006] Thirdly, embodiments of this specification provide a data protection model training apparatus, the apparatus comprising: The data acquisition module is used to acquire at least two sample raw data, input each sample raw data into the initial network model, and obtain the sample output data corresponding to each sample raw data. The initial network model is constructed based on a preset protection function. The first loss calculation module is used to obtain the standard output data after the preset protection function processes the original data of each sample, and to obtain the first distillation loss based on the standard output data and the output data of each sample. The second loss calculation module is used to calculate the first correlation between each standard output data and the second correlation between each sample output data, and to obtain the second distillation loss based on the first correlation and the second correlation. The first model training module is used to construct a first loss function based on the first distillation loss and the second distillation loss, and to perform a first training on the initial network model based on the first loss function to obtain a first data protection model.
[0007] Fourthly, embodiments of this specification provide a data protection device, which includes: An encryption module is used to respond to data encryption requests and encrypt the target original data based on a data encryption model to obtain the target encrypted data corresponding to the target original data. The decryption module is used to respond to data decryption requests, decrypt the target encrypted data based on the data decryption model, and obtain the target decrypted data corresponding to the target original data. Wherein, the data encryption model or the data decryption model is a data protection model trained by the data protection model training method described in any of the above embodiments of the specification.
[0008] Fifthly, embodiments of this specification provide a computer program product containing instructions that, when run on a computer or processor, cause the computer or processor to perform the steps of the method described above.
[0009] Sixthly, embodiments of this specification provide a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the steps of the method described above.
[0010] In a seventh aspect, embodiments of this specification provide a terminal including a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being adapted to be loaded by the processor and to execute the steps of the method described above.
[0011] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following: This specification provides a data protection model training method. First, an initial network model is constructed based on a preset protection function. The original sample data is input into the initial network model to obtain the output data for each sample. Then, a first distillation loss is obtained by comparing the standard output data after processing the original sample data with the output data of each sample, based on the preset protection function. A second distillation loss is obtained based on the first correlation between the standard output data and the second correlation between the output data of each sample. Finally, the initial network model is trained using the first and second distillation losses to obtain a first data protection model. Since the first correlation between the standard output data and the original sample data reflects the computational characteristics of the preset protection function after processing, using the first correlation between the standard output data and the second correlation between the output data of each sample to construct the loss calculation for training the data protection model allows the data protection model to more accurately fit the computational characteristics and computational capabilities of the preset protection function from the perspective of the correlation between the output data, resulting in a more accurate data protection model. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 An exemplary system architecture diagram of a data protection model training method provided in the embodiments of this specification; Figure 2 A flowchart illustrating a data protection model training method provided in an embodiment of this specification; Figure 3 A flowchart illustrating a data protection model training method provided in an embodiment of this specification; Figure 4 A flowchart illustrating a data protection model training method provided in an embodiment of this specification; Figure 5 A flowchart illustrating a data protection method provided in an embodiment of this specification; Figure 6 A structural block diagram of a data protection model training device provided in the embodiments of this specification; Figure 7 A structural block diagram of a data protection device provided in the embodiments of this specification; Figure 8 This is a schematic diagram of the structure of a terminal provided in an embodiment of this specification. Detailed Implementation
[0014] To make the features and advantages of the embodiments of this specification more apparent and understandable, the technical solutions of the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the embodiments of this specification.
[0015] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those described in this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the embodiments described in this specification as detailed in the appended claims.
[0016] In recent years, artificial intelligence (AI) technology has developed rapidly, and various related applications have begun to expand on a large scale and gradually become integrated into people's daily lives, such as self-service payment scenarios, intelligent recommendation scenarios, and autonomous driving assistance scenarios. However, since AI algorithms typically require analysis based on large amounts of data, the process of serving users often involves the collection, transmission, processing, and storage of users' private data. This poses a security risk of leakage of users' personal privacy information. Therefore, it is necessary to protect user information data to ensure user information security.
[0017] Data protection typically involves using data protection functions to encrypt and decrypt data during computation. Currently, in the field of data protection, from the perspective of computational data types, data protection strategies can be divided into two categories. The first category involves computation based on the original data. This type of strategy encrypts the original data during transmission and storage. When analysis and computation are required, the encrypted data is first decrypted and restored to obtain the original data, which is then used for computation. This method ensures the accuracy of the computation results, but still faces the security risk of information leakage during the computation process. The second category involves computation based on encrypted data. This type of strategy encrypts the data first, and subsequent data processing, including transmission, storage, and computation, is all performed on top of the encrypted data. This strategy provides stronger protection for information and data security.
[0018] For users' private data, a second type of data protection strategy with better encryption performance is usually adopted to ensure data security. For example, homomorphic encryption can be used. Homomorphic encryption can encrypt data and make the result of the calculation on the encrypted data the same as the result of the calculation on the original data. However, homomorphic encryption algorithms have a large amount of computation and low computational efficiency, and require high computing power from the running devices. In daily life scenarios, there are often some edge devices with low computing power that involve users' private data, such as self-service cash registers, which only have basic network access capabilities and cannot perform large-scale calculations. This makes it impossible for such devices to have homomorphic encryption performance on a large scale, thus threatening user data security.
[0019] Therefore, this specification provides a data protection model training method. First, the original data of each sample is input into an initial network model constructed based on a preset protection function to obtain the output data of each sample. Then, based on the preset protection function, the standard output data after processing the original data of each sample and the output data of each sample are used to obtain a first distillation loss. Based on the first correlation between the standard output data and the second correlation between the output data of each sample, a second distillation loss is obtained. Finally, the initial network model is trained according to the first distillation loss and the second distillation loss to obtain a first data protection model, thereby solving the technical problem of poor performance of the aforementioned data protection model.
[0020] Please see Figure 1 , Figure 1 This is an exemplary system architecture diagram of a data protection model training method provided in the embodiments of this specification.
[0021] like Figure 1 As shown, the system architecture may include a terminal 101, a network 102, and a server 103. The network 102 serves as the medium for providing a communication link between the terminal 101 and the server 103. The network 102 may include various types of wired or wireless communication links, such as wired communication links including fiber optic cables, twisted-pair cables, or coaxial cables, and wireless communication links including Bluetooth communication links, Wireless-Fidelity (Wi-Fi) communication links, or microwave communication links, etc.
[0022] Terminal 101 can interact with server 103 via network 102 to receive messages from or send messages to server 103. Alternatively, terminal 101 can interact with server 103 via network 102 to receive messages or data sent to server 103 by other users. Terminal 101 can be hardware or software. When terminal 101 is hardware, it can be various electronic devices, including but not limited to smartwatches, smartphones, tablets, laptops, and desktop computers. When terminal 101 is software, it can be installed in the aforementioned electronic devices and can be implemented as multiple software programs or software modules (e.g., to provide distributed services) or as a single software program or software module; no specific limitation is made here.
[0023] In the embodiments of this specification, terminal 101 can first construct an initial network model based on a preset protection function, input the original data of each sample into the initial network model, and obtain the output data of each sample; then, terminal 101 obtains a first distillation loss by processing the original data of each sample with the standard output data and the output data of each sample based on the preset protection function; further, terminal 101 obtains a second distillation loss based on the first correlation between the standard output data and the second correlation between the output data of each sample; finally, the initial network model is trained first based on the first distillation loss and the second distillation loss to obtain a first data protection model.
[0024] Server 103 can be an integrated server providing various services. It should be noted that server 103 can be either hardware or software. When server 103 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 103 is software, it can be implemented as multiple software programs or software modules (e.g., used to provide distributed services), or as a single software program or software module; no specific limitations are made here.
[0025] Alternatively, the system architecture may not include server 103. In other words, server 103 may be an optional device in the embodiments of this specification. That is, the method provided in the embodiments of this specification can be applied to a system structure that only includes terminal 101. The embodiments of this specification do not limit this.
[0026] It should be understood that Figure 1 The number of terminals, networks, and servers shown is only illustrative; the number can be any number of terminals, networks, and servers depending on the implementation requirements.
[0027] Please see Figure 2 , Figure 2This is a flowchart illustrating a data protection model training method provided in an embodiment of this specification. The execution entity in this embodiment can be a terminal executing data protection model training, a processor within the terminal executing the data protection model training method, or a data protection model training service within the terminal executing the data protection model training method. For ease of description, the following uses the processor within the terminal as an example to describe the specific execution process of the data protection model training method.
[0028] like Figure 2 As shown, data protection model training methods can include at least: S201. Obtain at least two sample raw data, input each sample raw data into the initial network model, and obtain the sample output data corresponding to each sample raw data. The initial network model is constructed based on a preset protection function.
[0029] Optionally, to meet user needs while protecting user information security, user data needs to be protected through methods such as data encryption and data anonymization. However, data protection strategies with good encryption effects often employ computationally intensive and inefficient data protection functions. Therefore, some edge devices providing data services to users, due to their limited computing power, cannot achieve adequate data encryption. Thus, to deploy data protection functions in low-computing-power electronic devices, it is necessary to reduce the computational load required for data protection processing and improve computational efficiency.
[0030] Furthermore, since neural network models can autonomously learn methods to solve pre-defined tasks, for some mathematical problems, compared to traditional and cumbersome formula calculations, neural network models can learn computational features in the data processing process through continuous iterative training based on initial and result data to fit the computational effect of formulas. Moreover, because neural networks learn directly from computational results, the computational load required to achieve the same computational effect is significantly reduced compared to formulas. This means that trained neural network models can be deployed on low-computing-power devices, enabling efficient computation with lower computational requirements. Specifically, the network structure can be chosen to be a lightweight network suitable for operation on edge devices, such as MobileNetV2.
[0031] Therefore, to deploy efficient data protection strategies, such as homomorphic encryption strategies, in devices, a neural network model can be used to fit a preset protection function corresponding to the efficient data protection strategy, thereby obtaining a data protection model that can fit the computational characteristics of the preset protection function. Before training the initial network model, it is first necessary to construct an initial network model based on the specific preset protection function, which facilitates subsequent training of the initial network model based on sample data. It should be noted that the preset protection function can be a preset encryption function or a preset decryption function, wherein the preset encryption function can be a homomorphic encryption function and the preset decryption function can be a homomorphic decryption function. The choice of the type of preset protection function does not limit the embodiments of this specification and can be selected based on actual needs.
[0032] Optionally, before training the initial network model, it is necessary to first obtain input data, i.e., obtain the original sample data, and then input each original sample data into the initial network model to obtain the sample output data of the initial network model based on the original sample data. When selecting the number of original sample data, considering the training effect of the initial network model, the number of original sample data is generally not one. Furthermore, considering that the sample output data needs to be fitted with the standard output data output by the preset protection function for the same original sample data, and the calculation characteristics of the preset protection function can be reflected in various ways through the standard output data, such as the value characteristics of a single output data, the correlation characteristics between multiple output data, and the value characteristics of the output data within a fixed original data range, at least two original sample data can be obtained so that the initial network model outputs at least two sample output data, which facilitates the training of the initial network model based on multiple sample output data.
[0033] S202. Obtain the standard output data of each sample after the preset protection function processes the original data of each sample, and obtain the first distillation loss based on the standard output data and the output data of each sample.
[0034] Optionally, as described in the above embodiments, the sample output data is the output of the initial network model for the original sample data. The fitting target of the sample output data is the standard output data after the preset protection function processes the same original sample data. After obtaining the sample output data, it is necessary to obtain the standard output data after the preset protection function processes each original sample data. Then, each standard output data can be compared with each sample output data. The initial network model is trained based on the data comparison results so that the initial network model learns the computational features of the preset protection function based on the data comparison results.
[0035] Optionally, when training the initial network model, a loss function is typically constructed based on training requirements. The loss function evaluates the degree to which the predicted values output by the network model differ from the true values. Based on the loss function, the optimization direction of the network model can be guided, making the network model's output closer to the standard values until a preset fitting effect is achieved. In the embodiments of this specification, when constructing the loss function, each standard output data and each sample output data can be used as part of the loss function construction. Since the inputs of both the preset protection function and the initial network model are the original sample data, in the knowledge distillation framework, the preset protection function is used as the teacher network, and the initial network model is used as the student network. The initial network model can learn the computational characteristics of the preset protection function; that is, based on each standard output data and each sample output data, the first distillation loss for training the initial network model can be obtained.
[0036] S203. Calculate the first correlation between each standard output data and the second correlation between each sample output data, and obtain the second distillation loss based on the first correlation and the second correlation.
[0037] Optionally, the first distillation loss obtained from each sample output data and each standard output data is the loss value between the output data of the initial network model and the preset protection function. Considering that the computational characteristics of the preset protection function can be reflected not only in the value characteristics of a single output data, but also in the correlation characteristics between multiple output data, the correlation between multiple output data is the relative distance between the vectors corresponding to the output data. The correlation between multiple output data can also represent the vector layout of the output data and the positional information between each output data. Therefore, it can be further stated that the first correlation between multiple standard output data can express the computational characteristics of the preset protection function, and the second correlation between multiple sample output data can express the computational characteristics of the initial network model. That is, after calculating the first correlation between each standard output data and the second correlation between each sample output data, the second distillation loss can be obtained based on the first correlation and the second correlation.
[0038] Optionally, the number of standard output data points can vary when calculating the first correlation, as long as the first correlation possesses certain characteristics to express the computational features of the preset protection function. For example, the first correlation can express the distance between standard output data points, or it can express the positional distribution between standard output data points. In this case, the first correlation can be the correlation between two standard output data points, the correlation between three standard output data points, or the correlation between any number of standard output data points. This specification does not limit this in the embodiments. It should be noted that the calculation method for the second correlation is the same as that for the first correlation, ensuring the rationality of using the second distillation loss to train the initial network model.
[0039] S203. Construct a first loss function based on the first distillation loss and the second distillation loss, and perform the first training on the initial network model based on the first loss function to obtain the first data protection model.
[0040] Optionally, after obtaining the first distillation loss and the second distillation loss, an initial network model can be trained based on the first distillation loss and the second distillation loss. Specifically, a first loss function is constructed based on the first distillation loss and the second distillation loss, and then the initial network model is trained based on the first loss function to obtain a first data protection model with less computational cost and higher data protection performance than the preset protection function.
[0041] Optionally, when constructing the first loss function, the sum of the first distillation loss and the second distillation loss can be used as the first loss function. Different weight values of the first distillation loss and the second distillation loss can be preset based on the actual training task objective to control the training focus of the initial network model. The embodiments in this specification do not specifically limit the setting of the weight values.
[0042] This specification provides a data protection model training method. First, an initial network model is constructed based on a preset protection function. The original data of each sample is input into the initial network model to obtain the output data of each sample. Then, a first distillation loss is obtained by comparing the standard output data after processing the original data with the output data of each sample based on the preset protection function. A second distillation loss is obtained based on the first correlation between the standard output data and the second correlation between the output data of each sample. Finally, the initial network model is trained using the first and second distillation losses to obtain a first data protection model. Since the first correlation between the standard output data and the original sample data reflects the computational characteristics of the preset protection function after processing, using the first correlation between the standard output data and the second correlation between the output data of each sample to construct the loss calculation for training the data protection model allows the data protection model to more accurately fit the computational characteristics and computational capabilities of the preset protection function from the perspective of the correlation between the output data, resulting in a more accurate data protection model.
[0043] Please see Figure 3 , Figure 3 This is a flowchart illustrating a data protection model training method provided in an embodiment of this specification.
[0044] like Figure 3 As shown, data protection model training methods can include at least: S301. Obtain at least two sample raw data, input each sample raw data into the initial network model, and obtain the sample output data corresponding to each sample raw data. The initial network model is constructed based on a preset protection function.
[0045] For details regarding step S301, please refer to the description in step S201; it will not be repeated here.
[0046] S302. Obtain the standard output data of each sample after the preset protection function processes the original data of each sample.
[0047] Optionally, if the sample output data is the output of the initial network model for the original sample data, then the fitting target of the sample output data is the standard output data after the preset protection function processes the same original sample data. After obtaining the sample output data, it is necessary to obtain the standard output data after the preset protection function processes each original sample data, so that the initial network model can be trained based on each standard output data and each sample output data.
[0048] S303. Calculate the first sub-distillation loss between each standard output data and the corresponding sample output data, and take the sum of the first sub-distillation losses as the first distillation loss.
[0049] Optionally, after obtaining the standard output data and the sample output data, the first distillation loss between the output of the preset protection function and the output of the initial network model can be calculated based on the standard output data and the sample output data. Specifically, when calculating the first distillation loss, the correspondence between the standard output data and the sample output data should be followed. The standard output data and the sample output data are both obtained by processing the same original sample data. Therefore, there is a correspondence between the standard output data corresponding to the same original sample data in the preset protection function and the sample output data in the initial network model.
[0050] It is easy to understand that in order to fit the initial network model, it is necessary to compare and calculate the sample output data of the initial network model for a certain sample of original data with the standard output data of the preset protection function for the same sample of original data. Only in this way can the training effect be achieved. That is, it is necessary to calculate the first sub-distillation loss between each standard output data and the sample output data corresponding to each standard output data, and finally use the sum of each first sub-distillation loss as the first distillation loss.
[0051] For example, when the original sample data has , , , The preset protection function processes the original data of each sample to obtain the standard output data. , , , The initial network model processes the original data of each sample to obtain the sample output data. , , , In this context, the original sample data, standard output data, and sample output data with the same data index have a corresponding relationship. Therefore, the absolute value of the Euclidean distance between each standard output data and its corresponding sample output data is calculated to obtain the first sub-distillation loss. The sum of all first sub-distillation losses is then used as the first distillation loss, i.e., .
[0052] S304. Based on the same preset rule, group each standard output data and each sample output data to obtain at least one set of standard output data and the sample output data set corresponding to each set of standard output data.
[0053] Optionally, as can be seen from the above embodiments, the computational characteristics of the preset protection function can be reflected not only in the value characteristics of a single output data, but also in the correlation characteristics between multiple output data. The correlation between multiple output data is the relative distance between the vectors corresponding to the output data, which can represent the vector layout of the output data and the positional information between each output data. Furthermore, the second distillation loss can be obtained based on the first correlation between each standard output data and the second correlation between each sample output data, so that the loss function of the initial network model can fit the computational characteristics of the preset protection function from more angles and can fit the same data protection performance as the preset protection function more quickly.
[0054] Optionally, the first correlation and the second correlation are the correlations between multiple data. Therefore, the calculation of the first correlation is related to at least two standard output data, and the calculation of the second correlation is related to at least two sample output data. In order to calculate the first correlation and the second correlation, it is necessary to group each standard output data and each sample output data to obtain at least one set of standard output data and at least one set of sample output data, so that the correlation can be calculated based on each data set.
[0055] Furthermore, since the output data of the preset protection function and the initial network model are the same original sample data, there is a correspondence between the standard output data and the sample output data corresponding to the same original sample data. When calculating the second distillation loss based on the first correlation and the second correlation, the loss calculation between the first correlation and the second correlation also needs to follow the correspondence between the standard output data and the sample output data. Therefore, in order to ensure the correspondence between the first correlation and the second correlation, each standard output data and each sample output data can be grouped based on the same preset rule, which ensures that there is still a correspondence between each standard output data group and each sample output data group based on the original sample data.
[0056] For example, the original sample data includes , , , Then the standard output data is , , , The sample output data is , , , At this point, the standard output data is divided into two standard output data groups based on preset rules. , Therefore, based on the same preset rule, the sample output data can be divided into two sample output data groups. , It is easy to understand that the standard output data group With sample output data group There is a corresponding relationship, standard output data group With sample output data group There is a corresponding relationship.
[0057] Optionally, for specific preset rules used for grouping, one feasible implementation is to set a data selection area with a preset size, and group the data in the same data selection area into the same group. In actual settings, in order to reduce the amount of calculation, the preset size of the data selection area can be small. At this time, the data selection area includes less data, which can ensure that less data is grouped into the same group, thereby reducing the amount of data.
[0058] S305. Calculate the first correlation between the standard output data in each standard output data group and the second correlation between the sample output data in each sample output data group.
[0059] Optionally, the cosine distance between the vectors corresponding to two data points is the local correlation. The local correlation between data can represent the relative distance and position information between data. Therefore, when calculating the first correlation of each standard output data group and the second correlation of each sample output data group, the local correlation between each standard output data in each standard output data group can be calculated as the first correlation, and the local correlation between each sample output data in each sample output data group can be calculated as the second correlation.
[0060] For example, when the original sample data has , , , Then the standard output data is , , , The sample output data is , , , After grouping according to the same preset rules, standard output data groups are obtained. , and sample output data set , At this point, the local correlation of each standard output data group is calculated, that is, the first correlation is... , And the local correlation of each sample output data group was calculated, that is, the second correlation is... , Among them, the first correlation With the second correlation Correspondingly, the first correlation With the second correlation correspond.
[0061] S306. Calculate the second sub-distillation loss between each first correlation and the second correlation corresponding to each first correlation, and take the sum of each second sub-distillation loss as the second distillation loss.
[0062] Optionally, after obtaining the first correlation of the standard output data in each set of standard output data and the second correlation of the sample output data in each set of sample output data, the first correlation and the second correlation can be compared and calculated. In the specific calculation process, since the standard output data and the sample output data have a corresponding relationship based on the original sample data, and the standard output data set and the sample output data set still have a corresponding relationship after being grouped according to the same preset rule, the second sub-distillation loss between the first correlation and the second correlation corresponding to the first correlation can be calculated, and the sum of the second sub-distillation losses is finally taken as the second distillation loss.
[0063] For example, when the original sample data has , , , Then the standard output data is , , , The sample output data is , , , After grouping, the standard output data group is obtained. , and sample output data set , And calculate the first similarity to obtain , The second correlation is , Based on the correspondence, the second sub-distillation loss between each first correlation and its corresponding second correlation can be calculated, and the sum of all second sub-distillation losses is taken as the second distillation loss. .
[0064] S307. Construct a first loss function based on the first distillation loss and the second distillation loss, and perform the first training on the initial network model based on the first loss function to obtain the first data protection model.
[0065] Optionally, after obtaining the first distillation loss and the second distillation loss, an initial network model can be trained based on the first distillation loss and the second distillation loss. That is, a first loss function is constructed based on the first distillation loss and the second distillation loss, and then the initial network model is trained on the first loss function until convergence, resulting in a first data protection model with lower computational cost and higher data protection performance than the preset protection function. Specifically, when constructing the first loss function, the sum of the first distillation loss and the second distillation loss can be used as the first loss function.
[0066] For example, when the original sample data has , , , Based on the first distillation loss and second distillation losses The first loss function can be constructed as follows: .
[0067] In the embodiments of this specification, a data protection model training method is provided. The first distillation loss is the sum of the first sub-distillation losses between each standard output data and the corresponding sample output data. The standard output data and sample output data are grouped according to the same preset rule to ensure the correspondence between the standard output data groups and the sample output data groups. Based on this correspondence, the second sub-distillation loss is calculated between each first correlation and the corresponding second correlation. The sum of these second sub-distillation losses is then used as the second distillation loss. The initial network model is trained not only based on the value characteristics of a single output data point but also on the correlations between multiple data points. This allows for fitting the computational characteristics of the preset protection function from multiple data perspectives, ultimately resulting in a first data protection model with lower computational complexity and equally efficient data protection performance compared to the preset protection function.
[0068] Please see Figure 4 , Figure 4 This is a flowchart illustrating a data protection model training method provided in an embodiment of this specification.
[0069] like Figure 4 As shown, user software requirement processing methods may include at least: S401. Obtain at least two sample raw data, input each sample raw data into the initial network model, and obtain the sample output data corresponding to each sample raw data. The initial network model is constructed based on a preset protection function.
[0070] S402. Obtain the standard output data of each sample after the preset protection function processes the original data of each sample, and obtain the first distillation loss based on the standard output data and the output data of each sample.
[0071] S403. Calculate the first correlation between each standard output data and the second correlation between each sample output data, and obtain the second distillation loss based on the first correlation and the second correlation.
[0072] S404. Construct a first loss function based on the first distillation loss and the second distillation loss, and perform the first training on the initial network model based on the first loss function to obtain the first data protection model.
[0073] For details regarding steps S401-S404, please refer to the detailed descriptions in steps S201-S204, which will not be repeated here.
[0074] S405. Obtain the original sample data, input the original sample data into the first data protection model, and obtain the first loss result based on the first loss function.
[0075] Optionally, for the first data protection model obtained from the first training, considering the complexity of the preset protection function and the diversity of the original sample data, the accuracy of the first data protection model may become higher and higher in multiple iterations of training, and a large number of network parameters will be optimized in the process. However, excessively high computational accuracy on the original sample data may lead to overfitting in real-world scenarios. Furthermore, among the large number of network parameters, there are also parameters that are of low computational importance and can be ignored. Therefore, in order to further optimize the first data protection model, a second training can be performed on the first data protection model to perform network pruning, remove some negligible network parameters from the first data protection model, further reduce the computational load of the data protection model, improve the computational efficiency of the model, and enhance the applicability of the model in real-world scenarios.
[0076] Optionally, to ensure that the performance of the trained model does not deviate significantly from that of the pre-trained model during the second training of the first data protection model, the same original sample data from the first training can be used. The first loss function from the first training can be used as part of the loss function in the second training to constrain the optimization direction of the data protection model in the second training. This optimizes computational load and efficiency while maintaining the model's data protection performance. Therefore, the original sample data used in the first training can be obtained, input into the first data protection model, and the first loss result can be obtained based on the first loss function. Subsequently, the first loss result can be used as part of the loss function for the second training to ensure the model's data protection performance.
[0077] S406. Calculate the sparse loss of the first data protection model based on the network parameters in the first data protection model.
[0078] Optionally, in order to prune the network of the first data protection model, remove redundant network parameters, and reduce the size of the first data protection model, it is first necessary to calculate based on the network parameters in the first data protection model to determine the sparsity of the first data protection model. The sparser the network, the more network parameters can be pruned, so the pruned network can be smaller and the computational efficiency will be higher. The first data protection model is trained according to the network sparsity to adjust the sparsity of the network parameters in the model until the sparsity of the model reaches the preset target.
[0079] Furthermore, the sparse constraints of the first data protection model can be calculated based on the L1 norm. The L1 norm refers to the sum of the absolute values of each element in a vector, also known as the "Lasso regularization (Least Absolute Shrinkage and Selection Operator (LASSO)" or the L1 regularization of linear regression. Since the L1 norm of a network model is obtained by adding the absolute values of each network parameter, the parameter value is proportional to the model complexity. Therefore, the more complex the model, the larger its L1 norm, which ultimately leads to a larger loss function related to the L1 norm. This means that the model still needs to be optimized by network pruning.
[0080] In the embodiments of this specification, the sparse loss of the first data protection model can be calculated based on the network parameters in the first data protection model. When the network parameters of the first data protection model are used... Therefore, the sparse loss of the first data protection model can be expressed using the L1 paradigm as follows: .
[0081] S407. Construct a second loss function based on the first loss result and the sparse loss, and perform a second training on the first data protection model based on the second loss function to obtain the second data protection model.
[0082] Optionally, after obtaining the first loss result and the sparse loss, a second loss function can be constructed based on the first loss result and the sparse loss. Then, the first data protection model can be trained a second time based on the second loss function to obtain a second data protection model with smaller computational size and higher computational efficiency.
[0083] Specifically, since the first loss result fits the direction closer to the preset protection function, while the sparse loss fits the direction closer to the pruned network parameters, in order to achieve a balance between data protection performance and computational efficiency in the final second data protection model, the preset first loss result and sparse loss can be set to have different weight values to adjust their proportion in the second loss function, thereby controlling the training direction of the first data protection model. That is, when constructing the second loss function, the first loss weight of the first loss result and the second loss weight of the sparse loss are first obtained, and the second loss function is constructed based on the product of the first loss weight and the first loss result and the product of the second loss weight and the sparse loss.
[0084] For ease of understanding, the first loss weight is represented as... The second loss weight is expressed as The first loss result corresponding to the first loss function of the first data protection model is: The sparse loss of the first data protection model is Then the second loss function is expressed as, .
[0085] Optionally, when training the first data protection model using the second loss function, since the first data protection model continuously adjusts its parameters and network structure during training, the corresponding first loss weights and second loss weights need to be modified based on each network adjustment. This can be done manually, or by using a meta-network for learning loss weights as a preset weight network model. The meta-network can perform unsupervised learning, training based on the second loss result corresponding to each second loss function of the first data protection model, and continuing to output updated first loss weights for the first loss function and second loss weights for sparse loss. That is, based on the second loss result obtained in the previous training process and the preset weight network model, the preset weight network model is trained based on the second loss result to obtain the first loss weights for the first loss function and the second loss weights for sparse loss.
[0086] Optionally, after the first data protection model is connected to the preset weight network model, while performing a second training on the first data protection model based on the second loss function, a third training on the preset weight network model can also be performed based on the second loss function. In this case, the first data protection model and the preset weight network model are trained alternately.
[0087] In the embodiments of this specification, a data protection model training method is provided. A first data protection model is trained a second time. Based on the sparse loss of the first data protection model, the sparsity of the network parameters of the first data protection model is determined. While ensuring the data protection performance of the first data protection model, unnecessary network parameters are pruned according to the sparse loss, reducing the network size of the first data protection model. After training, a second data protection model with excellent data protection performance, high computational efficiency, and smaller network size is obtained, which enhances the deployment capability and adaptability of the second data protection model in real-world scenarios.
[0088] Please see Figure 5 , Figure 5 This is a flowchart illustrating a data protection method provided in an embodiment of this specification.
[0089] like Figure 5 As shown, data protection methods may include at least: S501. Respond to the data encryption request, encrypt the target original data based on the data encryption model, and obtain the target encrypted data corresponding to the target original data.
[0090] Optionally, in practical application scenarios, to protect user privacy data, a data protection model can be deployed in the device. In this case, the device can obtain and process relevant data for protection based on the data protection model. When data encryption is required, the device, having deployed a data encryption model, can first respond to the data encryption request and then encrypt the obtained target raw data based on the deployed data encryption model to obtain the target encrypted data corresponding to the target raw data. The data protection model used is the data protection model from any embodiment of this specification.
[0091] In the embodiments of this specification, when the preset protection function is a homomorphic encryption function, the data protection model can achieve data protection performance equivalent to that of the homomorphic encryption function. In actual scenarios, after the device responds to a data encryption request, it encrypts the target original data based on the data encryption model to obtain the target encrypted data. The target encrypted data can then be directly uploaded to the server so that the server can perform calculations on the target encrypted data. Due to the encryption effect of the data encryption model, the calculation result of the server on the target encrypted data is equivalent to the calculation result of the server on the target original data. This allows the server to directly calculate based on the target encrypted data to meet the user's response without knowing the target original data. This avoids the target original data being exposed during data transmission, storage, and calculation. At this time, the data encryption model reduces the device's computational pressure while achieving a homomorphic, efficient, and secure encryption effect with a small amount of computation, greatly improving the encryption performance of low-computing-power devices and providing more rigorous protection for user information security.
[0092] S502. Respond to the data decryption request, and decrypt the target encrypted data based on the data decryption model to obtain the target decrypted data corresponding to the original target data.
[0093] Similarly, when data needs to be decrypted, the device has a data decryption model deployed in it. At this time, it can first respond to the data decryption request, and then decrypt the target decryption data based on the deployed data decryption model to obtain the corresponding target decryption data. The data protection model used is the data protection model in any embodiment of this specification.
[0094] In the embodiments of this specification, a data protection method is provided. In a practical application scenario, the data protection model of any of the foregoing embodiments is deployed and used. In response to a data encryption request, the target original data is encrypted based on the data encryption model to obtain target encrypted data corresponding to the target original data. In response to a data decryption request, the target encrypted data is decrypted based on the data decryption model to obtain target decrypted data corresponding to the target original data. This reduces the computational load of the device when performing data protection processing, improves computational efficiency, and achieves data security assurance.
[0095] Please see Figure 6 , Figure 6 This is a structural block diagram of a data protection model training device provided in an embodiment of this specification. Figure 6 As shown, the data protection model training device 600 includes: The data acquisition module 610 is used to acquire at least two sample raw data, input each sample raw data into the initial network model, and obtain the sample output data corresponding to each sample raw data. The initial network model is constructed based on a preset protection function. The first loss calculation module 620 is used to obtain the standard output data after the preset protection function processes the original data of each sample, and to obtain the first distillation loss based on the standard output data and the output data of each sample. The second loss calculation module 630 is used to calculate the first correlation between each standard output data and the second correlation between each sample output data, and to obtain the second distillation loss based on the first correlation and the second correlation. The first model training module 630 is used to construct a first loss function based on the first distillation loss and the second distillation loss, and to perform a first training on the initial network model based on the first loss function to obtain a first data protection model.
[0096] Optionally, the first loss calculation module 620 is further used to calculate the first sub-distillation loss between each standard output data and the sample output data corresponding to each standard output data, and to use the sum of each first sub-distillation loss as the first distillation loss.
[0097] Optionally, the second loss calculation module 630 is further configured to group each standard output data and each sample output data according to the same preset rule to obtain at least one set of standard output data and each set of sample output data corresponding to the standard output data; calculate the first correlation between each standard output data in each set of standard output data, and calculate the second correlation between each sample output data in each set of sample output data.
[0098] Optionally, the second loss calculation module 630 is further configured to calculate the second sub-distillation loss between each first correlation and the second correlation corresponding to each first correlation, and to use the sum of each second sub-distillation loss as the second distillation loss.
[0099] Optionally, the data protection model training device 600 further includes: a second model training module, used to acquire sample raw data, input each sample raw data into the first data protection model, and obtain a first loss result based on a first loss function; calculate the sparse loss of the first data protection model based on the network parameters in the first data protection model; construct a second loss function based on the first loss result and the sparse loss, and perform a second training on the first data protection model based on the second loss function to obtain a second data protection model.
[0100] Optionally, the second model training module is also used to obtain the first loss weight of the first loss result and the second loss weight of the sparse loss; and to construct the second loss function based on the product of the first loss weight and the first loss result and the product of the second loss weight and the sparse loss.
[0101] Optionally, the second model training module is also used to obtain the first loss weight of the first loss function and the second loss weight of the sparse loss based on the second loss result obtained in the previous training process and the preset weight network model.
[0102] Optionally, the second model training module is also used to perform a second training on the first data protection model based on the second loss function and a third training on the preset weight network model.
[0103] Optionally, the preset protection function is either a preset encryption function or a preset decryption function.
[0104] In this embodiment, a data protection model training device is provided, comprising: a data acquisition module for constructing an initial network model based on a preset protection function, inputting original sample data into the initial network model to obtain output data for each sample; a first loss calculation module for obtaining a first distillation loss based on the standard output data and the output data of each sample after processing the original sample data using the preset protection function; a second loss calculation module for obtaining a second distillation loss based on a first correlation between the standard output data and a second correlation between the output data of each sample; and a first model training module for performing a first training on the initial network model based on the first and second distillation losses to obtain a first data protection model. Since the first correlation between the standard output data and the original sample data reflects the computational characteristics of the preset protection function after processing, using the first correlation between the standard output data and the second correlation between the output data of each sample to construct the loss calculation for training the data protection model allows the data protection model to more accurately fit the computational characteristics and computational capabilities of the preset protection function from the perspective of the correlation between the output data, resulting in a more accurate data protection model.
[0105] Please see Figure 7 , Figure 7 This is a structural block diagram of a data protection device provided as an embodiment of this specification. Figure 7 As shown, the data protection device 700 includes: The encryption module 710 is used to respond to data encryption requests, encrypt the target original data based on the data encryption model, and obtain the target encrypted data corresponding to the target original data. The decryption module 720 is used to respond to data decryption requests, decrypt the target encrypted data based on the data decryption model, and obtain the target decrypted data corresponding to the original target data. The data encryption model or data decryption model is the data protection model trained by the data protection model training method in any embodiment of this specification.
[0106] In the embodiments of this specification, a data protection device is provided, wherein, in a practical application scenario, the data protection model of any of the foregoing embodiments is deployed and used. An encryption module is used to respond to a data encryption request and encrypt the target original data based on the data encryption model to obtain target encrypted data corresponding to the target original data. A decryption module is used to respond to a data decryption request and decrypt the target encrypted data based on the data decryption model to obtain target decrypted data corresponding to the target original data. This reduces the computational load during data protection processing, improves computational efficiency, and achieves data security assurance.
[0107] This specification provides a computer program product containing instructions that, when run on a computer or processor, cause the computer or processor to perform the steps of any of the methods described in the above embodiments.
[0108] This specification also provides a computer storage medium that can store multiple instructions adapted for loading by a processor and executing the steps of any of the methods described in the above embodiments.
[0109] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a terminal provided in an embodiment of this specification. Figure 8 As shown, terminal 800 may include: at least one terminal processor 801, at least one network interface 803, user interface 803, memory 805, and at least one communication bus 802.
[0110] The communication bus 802 is used to enable communication between these components.
[0111] The user interface 803 may include a display screen and a camera. Optionally, the user interface 803 may also include a standard wired interface and a wireless interface.
[0112] The network interface 803 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0113] The terminal processor 801 may include one or more processing cores. The terminal processor 801 connects to various parts within the terminal 800 using various interfaces and lines. It executes various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 805, and by calling data stored in the memory 805. Optionally, the terminal processor 801 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The terminal processor 801 may integrate one or more of the following: a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed on the screen; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the terminal processor 801.
[0114] The memory 805 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 805 may include a non-transitory computer-readable storage medium. The memory 805 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 805 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 805 may also be at least one storage device located remotely from the aforementioned terminal processor 801. Figure 8 As shown, the memory 805, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a data protection model training program.
[0115] exist Figure 8In the terminal 800 shown, the user interface 803 is mainly used to provide an input interface for the user and to obtain the user's input data; while the terminal processor 801 can be used to call the data protection model training program stored in the memory 805 and specifically perform the following operations: Obtain at least two original sample data, input each original sample data into the initial network model, and obtain the sample output data corresponding to each original sample data. The initial network model is constructed based on a preset protection function. Obtain the standard output data of each sample after processing the original data of each sample by the preset protection function, and obtain the first distillation loss based on the standard output data and the output data of each sample. Calculate the first correlation between each standard output data and the second correlation between each sample output data, and obtain the second distillation loss based on the first and second correlations. A first loss function is constructed based on the first distillation loss and the second distillation loss. The initial network model is then trained based on the first loss function to obtain the first data protection model.
[0116] In some embodiments, when the terminal processor 801 executes the first distillation loss based on each standard output data and each sample output data, it specifically performs the following steps: calculating the first sub-distillation loss between each standard output data and the sample output data corresponding to each standard output data, and taking the sum of each first sub-distillation loss as the first distillation loss.
[0117] In some embodiments, when the terminal processor 801 performs the calculation of the first correlation between each standard output data and the second correlation between each sample output data, it specifically performs the following steps: grouping each standard output data and each sample output data according to the same preset rule to obtain at least one set of standard output data and a sample output data set corresponding to each set of standard output data; calculating the first correlation between each standard output data in each set of standard output data and calculating the second correlation between each sample output data in each set of sample output data.
[0118] In some embodiments, when the terminal processor 801 executes the second distillation loss based on the first correlation and the second correlation, it specifically performs the following steps: calculating the second sub-distillation loss between each first correlation and the second correlation corresponding to each first correlation, and taking the sum of each second sub-distillation loss as the second distillation loss.
[0119] In some embodiments, after obtaining the first data protection model, the terminal processor 801 further performs the following steps: acquiring sample raw data, inputting the sample raw data into the first data protection model, and obtaining a first loss result based on a first loss function; calculating the sparse loss of the first data protection model based on the network parameters in the first data protection model; constructing a second loss function based on the first loss result and the sparse loss, and performing a second training on the first data protection model based on the second loss function to obtain a second data protection model.
[0120] In some embodiments, when the terminal processor 801 constructs a second loss function based on the first loss result and the sparse loss, it specifically performs the following steps: obtaining the first loss weight of the first loss result and the second loss weight of the sparse loss; constructing the second loss function based on the product of the first loss weight and the first loss result and the product of the second loss weight and the sparse loss.
[0121] In some embodiments, when the terminal processor 801 performs the following steps to obtain the first loss weight of the first loss function and the second loss weight of the sparse loss, it specifically performs the following steps: based on the second loss result obtained in the previous training process and the preset weight network model, it obtains the first loss weight of the first loss function and the second loss weight of the sparse loss.
[0122] In some embodiments, when the terminal processor 801 performs a second training of the first data protection model based on a second loss function, it specifically performs the following steps: performing a second training of the first data protection model based on the second loss function and performing a third training of a preset weight network model.
[0123] In some embodiments, the preset protection function is a preset encryption function or a preset decryption function.
[0124] Optionally, in Figure 8 In the terminal 800 shown, the user interface 803 is mainly used to provide an input interface for the user and to obtain the user's input data; while the terminal processor 801 can also be used to call the data protection program stored in the memory 805 and specifically perform the following operations: In response to the data encryption request, the target original data is encrypted based on the data encryption model to obtain the target encrypted data corresponding to the target original data; In response to the data decryption request, the target encrypted data is decrypted based on the data decryption model to obtain the target decrypted data corresponding to the original target data; The data encryption model or data decryption model is a data protection model trained by any of the data protection model training methods included in the above embodiments.
[0125] In the several embodiments provided in this specification, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0126] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0127] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this specification are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)).
[0128] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this specification is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this specification. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this specification.
[0129] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0130] The above is a description of a data protection model training and data protection method, apparatus and storage medium provided in this specification. For those skilled in the art, based on the ideas of the embodiments in this specification, there will be changes in the specific implementation and application scope. Therefore, the content of this specification should not be construed as a limitation of this specification.
Claims
1. A data protection model training method, the method comprising: Obtain at least two original sample data, input each original sample data into an initial network model, and obtain the sample output data corresponding to each original sample data, wherein the initial network model is constructed based on a preset protection function; Obtain the standard output data after the preset protection function processes the original data of each sample, and obtain the first distillation loss based on the standard output data and the output data of each sample. Calculate the first correlation between each standard output data and the second correlation between each sample output data, and obtain the second distillation loss based on the first correlation and the second correlation. A first loss function is constructed based on the first distillation loss and the second distillation loss, and the initial network model is trained based on the first loss function to obtain a first data protection model; The preset protection function is a preset encryption function or a preset decryption function, which is used to perform encryption or decryption calculations on the data.
2. The method according to claim 1, wherein obtaining the first distillation loss based on each standard output data and each sample output data includes: Calculate the first sub-distillation loss between each standard output data and the corresponding sample output data, and use the sum of the first sub-distillation losses as the first distillation loss.
3. The method according to claim 1, wherein calculating the first correlation between standard output data and the second correlation between sample output data comprises: Based on the same preset rule, each standard output data and each sample output data are grouped to obtain at least one set of standard output data and each set of sample output data corresponding to the standard output data; Calculate the first correlation between the standard output data in each standard output data group and the second correlation between the sample output data in each sample output data group.
4. The method according to claim 3, wherein obtaining the second distillation loss based on the first correlation and the second correlation comprises: Calculate the second sub-distillation loss between each first correlation and the second correlation corresponding to each first correlation, and take the sum of the second sub-distillation losses as the second distillation loss.
5. The method according to any one of claims 1 to 3, wherein after obtaining the first data protection model, it further comprises: The original data of the samples are obtained, and the original data of each sample are input into the first data protection model, and a first loss result is obtained based on the first loss function; The sparse loss of the first data protection model is calculated based on the network parameters in the first data protection model. A second loss function is constructed based on the first loss result and the sparse loss. The first data protection model is then trained a second time based on the second loss function to obtain the second data protection model.
6. The method according to claim 5, wherein constructing the second loss function based on the first loss result and the sparse loss comprises: Obtain the first loss weight of the first loss result and the second loss weight of the sparse loss; A second loss function is constructed based on the product of the first loss weight and the first loss result, and the product of the second loss weight and the sparse loss.
7. The method according to claim 6, wherein obtaining the first loss weight of the first loss function and the second loss weight of the sparse loss comprises: Based on the second loss result obtained from the previous training process and the preset weight network model, the first loss weight of the first loss function and the second loss weight of the sparse loss are obtained.
8. The method according to claim 7, wherein the second training of the first data protection model based on the second loss function comprises: The first data protection model is trained a second time based on the second loss function, and the preset weight network model is trained a third time.
9. A data protection method, the method comprising: In response to a data encryption request, the target original data is encrypted based on a data encryption model to obtain the target encrypted data corresponding to the target original data; In response to the data decryption request, the target encrypted data is decrypted based on the data decryption model to obtain the target decrypted data corresponding to the original target data; Wherein, the data encryption model or the data decryption model is a data protection model trained by the data protection model training method according to any one of claims 1 to 8.
10. A data protection model training device, the device comprising: The data acquisition module is used to acquire at least two sample raw data, input each sample raw data into the initial network model, and obtain the sample output data corresponding to each sample raw data. The initial network model is constructed based on a preset protection function. The first loss calculation module is used to obtain the standard output data after the preset protection function processes the original data of each sample, and to obtain the first distillation loss based on the standard output data and the output data of each sample. The second loss calculation module is used to calculate the first correlation between each standard output data and the second correlation between each sample output data, and to obtain the second distillation loss based on the first correlation and the second correlation. The first model training module is used to construct a first loss function based on the first distillation loss and the second distillation loss, and to perform a first training on the initial network model based on the first loss function to obtain a first data protection model; The preset protection function is a preset encryption function or a preset decryption function, which is used to perform encryption or decryption calculations on the data.
11. A data protection device, the device comprising: An encryption module is used to respond to data encryption requests and encrypt the target original data based on a data encryption model to obtain the target encrypted data corresponding to the target original data. The decryption module is used to respond to data decryption requests, decrypt the target encrypted data based on the data decryption model, and obtain the target decrypted data corresponding to the target original data. Wherein, the data encryption model or the data decryption model is a data protection model trained by the data protection model training method according to any one of claims 1 to 8.
12. A computer program product comprising instructions that, when run on a computer or processor, cause the computer or processor to perform the steps of the method as claimed in any one of claims 1 to 8 or 9.
13. A computer storage medium storing a plurality of instructions adapted for loading by a processor and performing the steps of the method as claimed in any one of claims 1 to 8 or 9.
14. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the steps of the method as claimed in any one of claims 1 to 8 or 9.
Citation Information
Patent Citations
Model training method and device and data processing method and device
CN112651511A
Model construction method, device and equipment based on privacy protection
CN113221717A