A data de-sensitization method and system
By using encoder training methods and model adjustments, a portion of sensitive data is extracted and processed as desensitized data, solving the problem of easily recoverable desensitized data and achieving a higher level of privacy protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-07
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies are insufficient to effectively protect the privacy of sensitive data, and the anonymized data provided can easily be restored to the original data, leading to privacy leaks.
The sample data is processed by an encoder training method, and a portion of the data is extracted as de-identified data. The encoder parameters are adjusted by a task processing model and a data reconstruction model to reduce the difference between the task prediction results and the reference standard, and increase the difference between the reconstructed data and the sample data, thereby improving the security of the de-identified data.
This improves the security of de-identified data, reduces the possibility of recovering the original data from the de-identified data, and enhances the privacy protection effect.
Smart Images

Figure CN114357519B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the technical field of data processing, and particularly relates to an encoder training method, a data de-sensitization method and system. BACKGROUND
[0002] In various business applications, sensitive data with sensitive information may be used, such as data of biological characteristics of a person, such as a face or voice, or data with sensitive information such as a certificate number, password, account number, etc. For sensitive data, data de-sensitization is required to hide the sensitive information therein, and the de-sensitized data is used for various business applications to achieve reliable protection of sensitive data. Due to the importance of sensitive data protection, it is desirable to improve the effect of data de-sensitization and obtain de-sensitized data that is not easily restored to the original data to better achieve privacy protection.
[0003] Therefore, there is an urgent need for a data de-sensitization method and system to improve the effect of data de-sensitization and better achieve privacy protection. SUMMARY
[0004] One aspect of the present specification provides an encoder training method, comprising: processing sample data through an encoder to obtain sample encoded data; extracting part of the data from the sample encoded data based on a preset position; determining sample de-sensitized data corresponding to the sample data based on the part of the data; processing the sample de-sensitized data corresponding to each through at least one task processing model to obtain task prediction results corresponding to each task processing model; processing the sample encoded data through a data reconstruction model to obtain reconstructed data; and adjusting at least the model parameters of the encoder to reduce the difference between each task prediction result and the corresponding reference standard, and to increase the difference between the reconstructed data and the sample data.
[0005] Another aspect of the present specification provides an encoder training system, comprising: a sample data encoding module configured to process sample data through an encoder to obtain sample encoded data; a sample encoded data extraction module configured to extract part of the data from the sample encoded data based on a preset position; a sample de-sensitized data determination module configured to determine sample de-sensitized data corresponding to the sample data based on the part of the data; a task prediction module configured to process the sample de-sensitized data corresponding to each through at least one task processing model to obtain task prediction results corresponding to each task processing model; a data reconstruction module configured to process the sample encoded data through a data reconstruction model to obtain reconstructed data; and a parameter adjustment module configured to adjust at least the model parameters of the encoder to reduce the difference between each task prediction result and the corresponding reference standard, and to increase the difference between the reconstructed data and the sample data.
[0006] Another aspect of the present specification provides an encoder training apparatus, comprising at least one storage medium and at least one processor, the at least one storage medium being configured to store computer instructions; and the at least one processor being configured to execute the computer instructions to implement the encoder training method.
[0007] One aspect of the present specification provides a data de-sensitization method, comprising: processing original data by using an encoder to obtain encoded data; extracting partial data from the encoded data based on a preset position; and determining de-sensitization data corresponding to the original data based on the partial data.
[0008] Another aspect of the present specification provides a data de-sensitization system, comprising: a data encoding module configured to process original data by using an encoder to obtain encoded data; an encoded data extraction module configured to extract partial data from the encoded data based on a preset position; and a de-sensitization data determination module configured to determine de-sensitization data corresponding to the original data based on the partial data.
[0009] Another aspect of the present specification provides a data de-sensitization apparatus, comprising at least one storage medium and at least one processor, the at least one storage medium being configured to store computer instructions; and the at least one processor being configured to execute the computer instructions to implement the data de-sensitization method. BRIEF DESCRIPTION OF DRAWINGS
[0010] The present specification will be further described in the manner of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same numbers represent the same structures, wherein:
[0011] Figure 1 is an application scenario diagram of an encoder training system or a data de-sensitization system according to some embodiments of the present specification;
[0012] Figure 2 is a block diagram of an encoder training system according to some embodiments of the present specification;
[0013] Figure 3 is a block diagram of a data de-sensitization system according to some embodiments of the present specification;
[0014] Figure 4 is an exemplary flowchart of an encoder training method according to some embodiments of the present specification;
[0015] Figure 5 is an exemplary schematic diagram of a training architecture of an encoder according to some embodiments of the present specification;
[0016] Figure 6is an exemplary schematic diagram of another training architecture of an encoder according to some embodiments of the present specification;
[0017] Figure 7 is an exemplary flow chart of a data de-identification method according to some embodiments of the present specification. DETAILED DESCRIPTION
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present specification, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some examples or embodiments of the present specification, and for those skilled in the art, the present specification can also be applied to other similar scenarios without creative labor. Unless it is clear from the language context or otherwise stated, the same reference numbers in the drawings represent the same structures or operations.
[0019] It should be understood that the "system", "device", "unit" and / or "module" used in the present specification is a method for distinguishing different components, elements, parts, sections or assemblies at different levels. However, if other words can achieve the same purpose, the words can be replaced by other expressions.
[0020] As shown in the present specification and claims, unless the context clearly indicates otherwise, "one", "a", "an" and / or "the" do not refer to the singular, but can also include the plural. Generally speaking, the terms "comprise" and "include" only indicate the inclusion of the steps and elements explicitly identified, and these steps and elements do not constitute an exclusive list, and the method or device can also include other steps or elements.
[0021] Flowcharts are used in the present specification to illustrate the operations performed by the system according to the embodiments of the present specification. It should be understood that the preceding or subsequent operations are not necessarily performed in sequence. On the contrary, each step can be processed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or more steps of the operation can be removed from these processes.
[0022] Figure 1 is an exemplary schematic diagram of an application scenario of an encoder training system or a data de-identification system according to some embodiments of the present specification.
[0023] The application scenario 100 can involve various business scenarios of data de-identification, such as business data de-identification, database operation de-identification, data export data de-identification, and the like.
[0024] In various business applications, sensitive data with sensitive information may be involved. For example, data about human faces, voices, and other biometric features, or data with sensitive information such as ID numbers, passwords, and account numbers. More information about sensitive data can be found in step 410 and its related description. For sensitive data, it is necessary to deform or modify it to hide or remove sensitive information, so as to achieve data desensitization. The desensitized data obtained is used in various business applications, for example, face image data with hidden face features is used for face recognition, and for example, transaction information without passwords is used for transaction risk prediction, which can achieve reliable protection of sensitive data.
[0025] It is inevitable that the desensitized data provided to the business application may be used to restore the original data. For example, others or business parties can additionally train a data reconstruction network, so that the data reconstruction network can reconstruct the same reconstruction data as the original data based on the desensitized data provided to the business application. If the original data is restored, it may cause the leakage of sensitive data and cannot guarantee privacy security. Therefore, it is desirable to improve the effect of data desensitization to obtain desensitized data that is not easy to be restored, so as to better achieve privacy protection.
[0026] Therefore, some embodiments of the present specification propose an encoder training method and system, and a data desensitization method and system. The encoder training method includes obtaining sample encoding data by processing sample data through an encoder, extracting part of the data from the sample encoding data based on a preset position to determine sample desensitized data corresponding to the sample data, processing the corresponding sample desensitized data through at least one task processing model to obtain corresponding task prediction results, and processing the sample encoding data through a data reconstruction model to obtain reconstruction data. Adjust the model parameters of the encoder based on the loss function, so that the difference between each task prediction result and the corresponding reference standard is reduced, and the difference between the reconstruction data and the sample data is increased. After the encoder is trained, the original data can be processed by the encoder to obtain encoding data, and part of the data can be extracted from the encoding data based on the preset position to determine the desensitized data corresponding to the original data, so as to achieve data desensitization. Through the encoder training method of the present specification, the preset position in the encoding data can become the position corresponding to the desensitized data, and then one or more desensitized data extracted from the preset position can be provided to the corresponding task processing model to implement business tasks and meet the effect requirements of the task prediction results. At the same time, it is difficult for the reconstruction network to restore the original data from these encoding data, which further improves the accuracy of task processing, reduces the possibility of restoring the original data from the desensitized data, and makes the desensitized data more secure.
[0027] As Figure 1As shown, the application scenario 100 of the encoder training system or the data de-identification system can include a server 110, a storage device 120, a network 130, and a user terminal 140.
[0028] The server 110 can be configured to manage resources and process data and / or information such as sample data, raw data from at least one component of the system or an external data source (e.g., a cloud data center), implement an encoder training method and a data de-identification method in a platform or a business field. In some embodiments, the server 110 can be a single server or a server group. The server group can be centralized or distributed (e.g., the server 110 can be a distributed system), and can be dedicated or simultaneously provided by other devices or systems. In some embodiments, the server 110 can be regional or remote. In some embodiments, the server 110 can be implemented on a cloud platform or provided in a virtual manner. For example only, the cloud platform can include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-layer cloud, etc., or any combination thereof.
[0029] The server 110 can include a processor 112. The processor 112 can process data and / or information such as sample data, raw data obtained from other device or system components. The processor 112 can execute program instructions based on these data, information, and / or processing results to perform one or more functions described in the present application. For example, the processor 112 can process sample data through an encoder to obtain encoded data, extract partial data from the sample encoded data based on a preset position to determine sample de-identification data and provide it to a corresponding task processing model, process the sample de-identification data corresponding to each task processing model through at least one task processing model to obtain a task prediction result corresponding to each task processing model, process the sample encoded data through a data reconstruction model to obtain reconstructed data, adjust at least the model parameters of the encoder based on a loss function, and then obtain a trained encoder. The processor 112 can also process raw data through the trained encoder to obtain encoded data, and extract partial data from the encoded data based on a preset position to determine de-identification data corresponding to the raw data. In some embodiments, the processor 112 can include one or more sub-processing devices (e.g., single-core processing devices or multi-core multi-core processing devices). For example only, the processor 112 can include a central processing unit (CPU), an application-specific integrated circuit (ASIC), an application-specific instruction processor (ASIP), a graphics processing unit (GPU), a physical processing unit (PPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a microcontroller unit, a reduced instruction set computer (RISC), a microprocessor, etc., or any combination thereof.
[0030] The storage device 120 can be used to store data and / or instructions. Data refers to a digitized representation of information, which can include various types such as binary data, textual data, image data, video data, etc. Instructions refer to programs that can control a device or instrument to perform a specific function. The storage device 120 can store data and / or information obtained from other devices or system components, such as sample data, raw data, sample de-identified data, de-identified data, trained encoder, etc. The storage device 120 can include one or more storage components, each of which can be a standalone device or a part of other devices. In some embodiments, the storage device 120 can include random access memory (RAM), read-only memory (ROM), mass storage, removable storage, volatile read-write memory, etc., or any combination thereof. Exemplarily, the mass storage can include a magnetic disk, an optical disk, a solid-state disk, etc. In some embodiments, the storage device 120 can be implemented on a cloud platform.
[0031] The network 130 can connect components of the system and / or connect the system with external parts. The network 130 enables communication between components of the system and between the system and external parts, facilitating exchange of data and / or information. In some embodiments, the network 130 can be any one or more of wired or wireless networks. For example, the network 130 can include a cable network, a fiber-optic network, a telecommunication network, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network, a ZigBee network, near-field communication (NFC), an intra-device bus, an intra-device line, a cable connection, etc., or any combination thereof. In some embodiments, network connections between parts of the system can be in one of the above manners or in multiple manners. In some embodiments, the network 130 can be in various topologies or combinations of topologies, such as point-to-point, shared, hub-and-spoke, etc. In some embodiments, the network 130 can include one or more network access points. For example, the network 130 can include wired or wireless network access points, such as base stations and / or network switching points 130-1, 130-2,..., through which one or more components of the scenario 100 can connect to the network 130 to exchange data and / or information.
[0032] Terminal 140 refers to one or more terminal devices or software used by a user or a business party. In some embodiments, the terminal 140 can be used by any user or business party, such as an individual, an enterprise, etc. In some embodiments, the terminal 140 can be one or any combination of a mobile device 140-1, a tablet computer 140-2, a laptop computer 140-3, a desktop computer 140-4, and other devices with input and / or output functions. The above examples are only used to illustrate the broadness of the terminal 140 device range, but not to limit the range thereof.
[0033] In some embodiments, the server 110, the terminal 140, and other possible system components can include a storage device 120.
[0034] The server 110 can communicate with the storage device 120 and the terminal 140 through the network 130 to obtain data and / or information, for example, the server 110 can obtain sample data from the storage device 120 through the network 130, etc. The server 110 can execute program instructions based on the obtained data, information, and / or processing results to implement training of an encoder and to implement data de-sensitization. For example, the server 110 can train the encoder using the sample data. In some embodiments, one or more trained task processing models can be deployed on the terminal 140. The server 110 can send de-sensitized data to the terminal 140 through the network 130, so that the terminal 140 processes the de-sensitized data using the task processing model to obtain a task prediction result. In some embodiments, the terminal 140 can obtain the trained encoder from the server 110 through the network 130 to implement data de-sensitization processing. Further, the terminal 140 can process original data using the trained encoder and transmit the obtained de-sensitized data to a device (not shown in the figure) of another business party to obtain a task prediction result. The information transmission relationship between the above devices is only an example, and the present application is not limited thereto.
[0035] Figure 2 is a block diagram of an encoder training system according to some embodiments of the present specification.
[0036] In some embodiments, the encoder training system 200 can be implemented on the processor 112.
[0037] In some embodiments, the encoder training system 200 can include a sample data encoding module 210, a sample encoded data extraction module 220, a sample de-sensitized data determination module 230, a task prediction module 240, a data reconstruction module 250, and a parameter adjustment module 260.
[0038] In some embodiments, the sample data encoding module 210 can be used to process sample data through an encoder to obtain sample encoded data.
[0039] In some embodiments, the sample encoding data extraction module 220 can be configured to extract partial data from the sample encoding data based on a preset position. In some embodiments, the partial data comprises at least one task-specific data and task-common data.
[0040] In some embodiments, the sample desensitization data determination module 230 can be configured to determine sample desensitization data corresponding to the sample data based on the partial data. In some embodiments, the sample desensitization data determination module 230 can be further configured to combine at least one task-specific data with the task-common data respectively, thereby obtaining at least one sample desensitization data corresponding to the at least one task processing model.
[0041] In some embodiments, the task prediction module 240 can be configured to process the sample desensitization data corresponding to the at least one task processing model respectively by the at least one task processing model, thereby obtaining the task prediction result corresponding to each task processing model.
[0042] In some embodiments, the data reconstruction module 250 can be configured to process the sample encoding data by a data reconstruction model, thereby obtaining reconstructed data. In some embodiments, the data reconstruction model can comprise a generator and a discriminator; the generator is configured to process the sample encoding data, thereby obtaining reconstructed data; the discriminator is configured to process input data, thereby obtaining a corresponding score; the score reflects the probability of the discriminator determining that the processed data is true.
[0043] In some embodiments, the parameter adjustment module 260 can be configured to adjust at least the model parameters of the encoder based on the loss function, so as to reduce the difference between the task prediction result and the corresponding reference standard, and increase the difference between the reconstructed data and the sample data. In some embodiments, the parameter adjustment module 260 can be further configured to adjust at least the model parameters of the encoder based on the loss function, so as to reduce the correlation between the partial data and the remaining part of the encoded data. In some embodiments, the parameter adjustment module 260 can be further configured to adjust at least the model parameters of the encoder, so as to make the modulus of the partial data not less than a preset value, and make the modulus of the remaining part of the encoded data not less than a preset value. In some embodiments, the parameter adjustment module 260 can be further configured to adjust at least the model parameters of the encoder based on the loss function, so as to reduce the correlation between the at least one task-specific data and any two of the remaining part of the encoded data. In some embodiments, the parameter adjustment module 260 can be further configured to adjust at least the model parameters of the encoder, so as to make the modulus of the two or more task-specific data not less than a preset value, and make the modulus of the remaining part of the encoded data not less than a preset value. In some embodiments, the parameter adjustment module 260 can be further configured to adjust the model parameters of the at least one task processing model and the model parameters of the data reconstruction model, so as to reduce the difference between the task prediction result and the corresponding reference standard, and increase the difference between the reconstructed data and the sample data. In some embodiments, the parameter adjustment module 260 can be further configured to adjust the model parameters of the at least one task processing model, so as to reduce the difference between the task prediction result and the corresponding reference standard; adjust the model parameters of the data reconstruction model, so as to reduce the difference between the reconstructed data and the sample data; process the sample de-sensitized data corresponding to the at least one task processing model using the adjusted at least one task processing model, to obtain the secondary task prediction result corresponding to the at least one task processing model; process the sample encoded data using the adjusted data reconstruction model, to obtain the secondary reconstructed data; adjust the model parameters of the encoder, so as to reduce the difference between the secondary task prediction result and the corresponding reference standard, and increase the difference between the secondary reconstructed data and the sample data. In some embodiments, the parameter adjustment module 260 can be further configured to process the reconstructed data using the discriminator to obtain a corresponding score; the score reflects the probability that the processed data is true as judged by the discriminator; and adjust the model parameters of the generator, so as to increase the score. In some embodiments, the parameter adjustment module 260 can be further configured to process the sample data using the discriminator to obtain a corresponding score; and adjust the model parameters of the discriminator, so as to reduce the score corresponding to the reconstructed data, and increase the score corresponding to the sample data.
[0044] Figure 3is a block diagram of a data de-identification system according to some embodiments of the present specification.
[0045] In some embodiments, the data de-identification system 300 can be implemented on the processor 112.
[0046] In some embodiments, the encoder training system 300 can comprise a data encoding module 310, an encoded data extracting module 320, and a de-identified data determining module 330.
[0047] In some embodiments, the data encoding module 310 can be configured to process the original data with an encoder to obtain encoded data.
[0048] In some embodiments, the encoded data extracting module 320 can be configured to extract partial data from the encoded data based on a preset position. In some embodiments, the partial data comprises at least one task-specific data and task-common data.
[0049] In some embodiments, the de-identified data determining module 330 can be configured to determine the de-identified data corresponding to the original data based on the partial data. In some embodiments, the de-identified data determining module 330 can also be configured to directly determine the partial data as the de-identified data. In some embodiments, the de-identified data determining module 330 can also be configured to combine at least one task-specific data with the task-common data respectively to obtain at least one de-identified data.
[0050] In some embodiments, the encoder is trained to obtain the preset position as a corresponding position of the de-identified data in the encoded data during the training process of the encoder. In some embodiments, the training of the encoder can be implemented by the encoder training system 200.
[0051] It should be appreciated that the illustrated system and its modules can be implemented in various ways. For example, in some embodiments, the system and its modules can be implemented in hardware, software, or a combination of software and hardware. The hardware portions can be implemented with specialized logic; the software portions can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-described methods and systems can be implemented using computer-executable instructions and / or in processor control code, for example, provided on a carrier medium such as a disk, CD or DVD-ROM, programmable memory such as read-only memory (firmware), or data carrier such as an optical or electrical signal carrier. The system and its modules of the present specification can be implemented not only in hardware circuitry such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., but also in software, for example, executed by various types of processors, and in a combination of the above-mentioned hardware circuitry and software (e.g., firmware).
[0052] It should be noted that the above description of the system and its modules is for the convenience of description only and does not limit the present specification to the scope of the embodiments. It can be understood that, for those skilled in the art, after understanding the principles of the system, the modules can be combined arbitrarily or connected to other modules to form a subsystem without departing from the principles.
[0053] Figure 4 is an exemplary flowchart of an encoder training method according to some embodiments of the present specification.
[0054] In some embodiments, the method 400 can be performed by the processor 112. In some embodiments, the method 400 can be implemented by the encoder training system 200 deployed on the processor 112. It should be noted that the method 400 can be regarded as a round of model training method, and in some embodiments, multiple rounds of iterative model training can be performed based on the method 400. One round of training can be based on one batch of sample data, and one batch of sample data can be 100, 1000, etc. For the convenience of understanding, the method 400 can also be regarded as a training process based on one sample data.
[0055] As shown in Figure 4 the method 400 can include:
[0056] Step 410: processing the sample data by the encoder to obtain sample encoded data.
[0057] In some embodiments, the step 410 can be performed by the sample data encoding module 210.
[0058] The sample data can be various types of data, such as text, picture, voice, video, and the like. In some embodiments, the sample data can include sensitive data. The sensitive data, also referred to as private data, refers to data including sensitive or important information that should not be disclosed or leaked, such as data about the face, voice, and the like of a person, or data with sensitive information such as a certificate number, bank account number, email, password, medical information, and the like.
[0059] In some embodiments, if sensitive data needs to be used, the sensitive data can be desensitized to obtain desensitized data for various business applications to achieve reliable protection of the sensitive data. Data desensitization, also referred to as data de-privatization or data deformation, refers to deformation or modification of sensitive data to hide or remove sensitive information, such as data processing by an encoder to achieve data deformation or modification to obtain encoded data. The data obtained after data desensitization can be referred to as desensitized data.
[0060] The encoder can refer to any algorithm or device that processes data to achieve data deformation or conversion to obtain encoded data. In some embodiments, the encoder can include any one or a combination of a plurality of the following: a fully connected layer, an encoding network (such as an encoding network in a U-Net network, a Transformer network, and the like), an encoding model (such as a neural network model such as NN, RNN, CNN, LSTM, and the like), and the like. In some embodiments, the encoder can also be part of a machine learning model, such as the first several layers in a DNN model, i.e., the output of a certain hidden layer in a certain machine learning model is used as the encoded data.
[0061] In some embodiments, the processing of the input data by the encoder can include any one or a combination of a plurality of the following: data reorganization, data value replacement, data shuffling, data mean valueization, data offset, data encryption, feature extraction, data compression, data encoding, and the like.
[0062] In this specification, the encoded data obtained by the encoder processing the sample data can be referred to as sample encoded data.
[0063] Step 420: Extracting part of the data from the sample encoded data based on a preset position.
[0064] In some embodiments, this step 420 can be performed by the sample encoded data extraction module 220.
[0065] In some embodiments, the encoded data output by the encoder can be used entirely for the business task, for example, the output encoded data is provided entirely as de-identified data to the task processing model to implement the business task. The task processing model refers to a data processing model used to implement the business task, and more description about the task processing model can be referred to step 430 and its related description.
[0066] In some embodiments, the encoded data output by the encoder can be partially used for the business task. Specifically, the encoded data can include part of data important for the business task and part of data unimportant for the business task. Among them: the part of data important for the business task (which can be referred to as task important data in this specification, and hereinafter referred to as task important data) can provide data required or necessary for implementing the business task. The part of data unimportant for the business task is unnecessary for implementing the business task. Therefore, part of the encoded data can be provided to the task processing model to implement the business task.
[0067] In some embodiments, the encoded data can be divided into two parts based on a preset position, and one of the divided parts can be used as task important data, and the remaining part can be used as part of data unimportant for the business application. The preset position can refer to the range of data, the division position of data, etc. In some embodiments, the preset position can be set according to experience or actual needs. For example, the encoded data can be divided into two equal parts, and the front part corresponds to the preset position, that is, the task important data, and the rear part is part of data unimportant for the business application. For another example, the encoded data can not be divided equally, but the preset position corresponds to a longer segment of the encoded data, so that the part of data corresponding to the preset position can be further divided, and each segment of data obtained by further dividing is equal in length to the remaining part of the encoded data except the preset position.
[0068] In some embodiments, the encoded data can be in the form of a vector, a matrix, or a higher-dimensional tensor.
[0069] As an example, the encoder processes the input data and obtains encoded data in the form of a 256-dimensional vector. The preset position in the encoded data can be the position of half the length of the vector, such as the first 128 bits in the vector, based on which the encoded data can be divided and formed into two vectors with a length of 128 dimensions, one of which is used as task important data, and the other is part of data unimportant for the business task. The preset position in the encoded data can also be other positions in the vector, such as the last 192 bits in the 256-dimensional vector, or the second 64 bits in the vector, etc.
[0070] As another example, the encoder processes the input data to obtain encoded data in the form of a matrix with a size of 256*256. The preset position in the encoded data can be a position of half of the matrix, such as a position corresponding to the first 128 rows of the matrix. Based on the preset position, the encoded data can be divided and formed into two matrices with a size of 128*256, one of which is task important data, and the other is part of data that is not important for the business task. The preset position in the encoded data can also be other, such as a position corresponding to the right 128 columns of the matrix in the previous example, or a position corresponding to the first 192 rows of the matrix in the previous example.
[0071] As another example, the encoder processes the input data to obtain encoded data in the form of a tensor with a size of 64*64*4, such as a 64*64*4 feature map (also referred to as a feature map) composed of 4 channels. The preset position in the encoded data can be a position corresponding to half of the tensor, such as a position corresponding to a 64*64*2 feature map composed of 2 channels. Based on the preset position, the encoded data can be divided and formed into two tensors with a size of 64*64*2, such as two feature maps with a size of 64*64 and a channel number of 2, one of which is task important data, and the other is part of data that is not important for the business task. The preset position in the encoded data can also be other, such as a 64*64*3 tensor composed of the first 3 channels of the tensor in the previous example.
[0072] At step 430, sample de-identification data corresponding to the sample data is determined based on the part of data.
[0073] In some embodiments, the step 430 can be performed by the sample de-identification data determination module 230.
[0074] Based on the foregoing, the data corresponding to the preset position in the encoded data is task important data, and therefore, this part of data can be provided to the business task for corresponding task processing. Since the encoded data has undergone data transformation or deformation, to a large extent, the privacy data can be hidden, and the task important data extracted based on the preset position is only a part of the encoded data, further increasing the difficulty of obtaining the original data by reverse calculation. Therefore, in this specification, the data to be provided to the task processing model or used for task prediction is also referred to as de-identification data. In this specification, the part of data extracted from the sample encoded data obtained by processing the sample data by the encoder to obtain the de-identification data can be referred to as sample de-identification data.
[0075] In some embodiments, different task processing models can have different corresponding de-identification data. In some embodiments, one or more pieces of de-identification data corresponding to the business processing model can be extracted from the encoded data output by the encoder based on the preset position, to be provided to the corresponding task processing model to implement the corresponding business task.
[0076] In some embodiments, the task processing model is one, and the partial data extracted based on the preset position can be directly provided as the de-identification data corresponding to the task processing model. As shown in FIG. 4A, the encoded data is divided into two parts, one of which is the partial data extracted based on the preset position, i.e., the task important data, such as the a1 part in FIG. 4A, which is provided as the de-identification data to the task processing model to implement the business task, and the other is the remaining part of data that is not important to the business task, such as the a2 part in FIG. 4A. Figure 5 Figure 5 Figure 5 Figure 5 An exemplary schematic diagram of the training architecture of the encoder is shown in FIG. 4B, and more details about the same can be found in step 450 and related content. Figure 5
[0077] In some embodiments, the task processing model can be two or more, and the partial data extracted based on the preset position can be further divided to obtain two or more task-specific data and task-common data. The task-specific data refers to data important to one task processing model but relatively unimportant to other task processing models, in other words, the task-specific data is specific to only one task processing model. The task-common data refers to data important to two or more task processing models, in other words, the task-common data is common to two or more task processing models.
[0078] In some embodiments, each task-specific data and the task-common data can be combined to obtain the de-identification data corresponding to each task processing model. As shown in FIG. 4C, the partial encoded data extracted based on the preset position can be divided into three parts, one of which is the task-specific data of the task processing model 1, such as the b1 part in FIG. 4C, one of which is the task-specific data of the task processing model 2, such as the b2 part in FIG. 4C, and one of which is the task-common data of the task processing models 1 and 2, such as the b3 part in FIG. 4C. Figure 6 Figure 6 Figure 6 Figure 6 Figure 6 The b4 part in FIG. 4C is the part of the encoded data other than the preset position. Figure 6 An exemplary schematic diagram of the training architecture of the encoder is shown in FIG. 4D, and more details about the same can be found in step 450 and related content. As shown in FIG. 4D, the task-specific data b1 and the task-common data b3 can be combined to obtain the de-identification data corresponding to the task processing model 1, which is provided to the task processing model 1, and the task processing model 1 can implement the corresponding business task. Similarly, the data obtained by combining the task-specific data b2 and the task-common data b3 can be provided as the de-identification data to the task processing model 2, and the task processing model 2 can implement the corresponding business task. Figure 6 Figure 6
[0079] Step 440, processing the corresponding sample de-identification data of each task processing model respectively by at least one task processing model to obtain the task prediction result corresponding to each task processing model.
[0080] In some embodiments, the step 440 can be performed by the task prediction module 240.
[0081] The task processing model refers to a data processing model used to implement a business task. The de-identification data or sample de-identification data is input into the task processing model, and the model can obtain the corresponding task prediction result. The task processing model can include various data processing models that can be used for business tasks, such as NN, RNN, CNN, GNN, GCN, etc. The business task can be any data processing task, such as face recognition, entity classification, inter-entity relationship prediction, etc.
[0082] The task processing model can include one or more, and the obtained one or more sample de-identification data is respectively input into the corresponding task processing model. Each task processing model can obtain the corresponding task prediction result.
[0083] Step 450, processing the sample encoding data by a data reconstruction model to obtain reconstructed data.
[0084] In some embodiments, the step 450 can be performed by the data reconstruction module 250.
[0085] The data reconstruction model refers to a model used to reconstruct the input encoding data to obtain the original data. The encoding data output by the encoder (which can also be sample encoding data) is input into the data reconstruction model, and the model can obtain the reconstructed data, such as reconstructed original data or reconstructed sample data.
[0086] In some embodiments, the data reconstruction model can include the decoder part of the U-Net network, the generator in the generative adversarial network (GAN), etc.
[0087] In some embodiments, the data reconstruction model can also include the discriminator in the generative adversarial network (GAN).
[0088] The generator can be used to process the input encoded data (which can also be sample encoded data) to obtain reconstructed data such as reconstructed sample data. The discriminator can be used to discriminate the input data such as the reconstructed data output by the generator to obtain a corresponding score to reflect the probability that the input data is real. Wherein, the input data is real means that the input data is real data rather than machine generated reconstructed data, the higher the score output by the discriminator, the greater the probability that the input data is real data, on the contrary, the lower the output score, the greater the probability that the input data is machine generated reconstructed data. The main role of the discriminator is to train jointly with the generator to make the reconstructed data generated by the generator as realistic as possible. In some embodiments, the data reconstruction model can be pre-trained, and a round of training process of the data reconstruction model can be: fixing the model parameters of the discriminator, inputting the reconstructed data output by the generator into the discriminator, adjusting the model parameters of the generator, so that the reconstructed data generated by the generator is processed by the discriminator to obtain a larger score, to "confuse" the discrimination of the discriminator, to "fake as real", at this time, the discriminator can be pre-trained. In this way, multiple rounds of adjustment of the parameters of the generator are made, and then the trained data reconstruction model is obtained. In some embodiments, a round of training process of the data reconstruction model can be: training the generator and the discriminator at the same time, at this time, the discriminator can be used to process multiple reconstructed data and sample data to output corresponding scores of the reconstructed data and the sample data respectively, and the model parameters of the discriminator are adjusted so that the score corresponding to the reconstructed data is as small as possible and the score corresponding to the sample data (i.e. real data) is as large as possible. The reconstructed data output by the generator in the current round is processed by the discriminator with adjusted parameters to obtain a corresponding score, at this time, the model parameters of the discriminator are fixed, and only the model parameters of the generator are adjusted to make the score output by the discriminator as large as possible. Multiple rounds of iterative training are repeated in this way, and the training of the data reconstruction model can be completed.
[0089] In step 460, at least the model parameters of the encoder are adjusted to reduce the difference between the task prediction results and the corresponding reference standard, and to increase the difference between the reconstructed data and the sample data.
[0090] In some embodiments, the step 460 can be performed by the parameter adjustment module 260.
[0091] In some embodiments, the training of the encoder can be implemented based on one or more task processing models and a data reconstruction model.
[0092] Figure 5 is an exemplary schematic diagram of a training architecture of an encoder. As Figure 5As shown, the encoder processes the input data (original data or sample data) to obtain encoded data, and based on the preset position, the encoded data is divided into two parts a1 and a2. The a1 part is the task processing model a corresponding de-sensitization data (which can also be sample de-sensitization data), which is input into the task processing model a to achieve the prediction task. The entire encoded data, i.e., the a1 and a2 parts, can be used as a whole data representation to input the data reconstruction model to obtain the reconstructed data.
[0093] Figure 6 is another exemplary schematic diagram of the training architecture of the encoder, which is suitable for the case of multi-task processing model. Figure 6 As shown, the encoder processes the input data (original data or sample data) to obtain encoded data, and based on the preset position, the encoded data is divided into two parts a1 and a2. The a1 part is the task processing model a corresponding de-sensitization data (which can also be sample de-sensitization data), which is input into the task processing model a to achieve the prediction task. The entire encoded data, i.e., the a1 and a2 parts, can be used as a whole data representation to input the data reconstruction model to obtain the reconstructed data. Figure 6 The part filled with shading in is further divided into three parts b1, b2, and b3. The b4 part is the remaining part of the encoded data other than the preset position. The b1 part is the task processing model 1 corresponding task-specific data, the b2 part is the task processing model 2 corresponding task-specific data, and the b3 part is the task processing model 1 and the task processing model 2 task-shared model. The b1 part and the b2 part are combined to form a data representation as de-sensitization data (which can also be sample de-sensitization data) input into the task processing model 1 to achieve the corresponding prediction task. The b3 part and the b2 part are combined to form a data representation as de-sensitization data (which can also be sample de-sensitization data) input into the task processing model 2 to achieve the corresponding prediction task. The entire encoded data, i.e., the b1, b2, b3, and b4 parts, can form a whole data representation to input the data reconstruction model to obtain the reconstructed data.
[0094] In some embodiments, the loss function can be determined based on the difference between the task prediction result output by each task processing model (such as Figure 5 in, or Figure 6 the task processing model 1 and the task processing model 2 in), and the difference between the reconstructed data obtained by the data reconstruction model and the real original data, and the model parameters of the encoder are adjusted based on the loss function to achieve the training of the encoder.
[0095] The reference standard of the task prediction result can be the label of the sample data. As an example, for the task of face recognition, the reference standard of the task processing model is the identity label of the person corresponding to the sample de-sensitization data input into the task processing model, such as Zhang San.
[0096] To improve the encoding effect of the encoder to obtain the task prediction effect of the task processing model that can meet the task processing model and the desensitization data that is not easy to be restored to the original data, in some embodiments, the target of adjusting the model parameters of the encoder based on the loss function can include reducing the difference between the task prediction effect of each task processing model (for example, Figure 5 a single task processing model in Figure 6 task processing model 1, task processing model 2) and the corresponding reference standard, and increasing the difference between the reconstructed data and the corresponding sample data.
[0097] As an example, taking a single task processing model in Figure 5 as an example, the loss function L F of the task processing model can be obtained based on the difference between the task prediction result and the corresponding reference standard. R The loss function L F of the data reconstruction model can also be obtained based on the difference between the reconstructed data and the real original data. R The loss function L E of the encoder training can be obtained based on L R , L F , where E represents the encoder, F represents the task processing model, R represents the data reconstruction model, and λ is a weight coefficient. λ can be determined according to experience or actual demand. Exemplarily, it can take a number in the range of (0, 1). When the encoder is trained, the model parameters of the encoder can be adjusted to minimize the loss function.
[0098] Taking two task processing models in Figure 6 as an example, similar to the training architecture in Figure 5 , the loss function L F1 of the task processing model 1, the loss function L F2 of the task processing model 2, and the loss function L R of the data reconstruction model can be obtained based on the difference between the task prediction result and the corresponding reference standard. Further, the loss function L E of the encoder training can be obtained as L R + λ (L F1 + L F2 ), where F1 represents the task processing model 1 and F2 represents the task processing model 2. When the encoder is trained, the model parameters of the encoder can be adjusted to minimize the loss function.
[0099] In some embodiments, the loss function for training the encoder can further include a term related to the correlation between the partial data corresponding to the preset position and the remaining partial data in the sample encoded data, so that adjusting the model parameters of the encoder based on the loss function can further include minimizing the correlation between the partial data and the remaining partial data in the sample encoded data. The correlation minimization can refer to minimizing the similarity between the two data or the value of the dot product of the two data. For example, Figure 5 The correlation minimization between a1 and a2 in the above formula can refer to minimizing the value of a1·a2 or the similarity between a1 and a2. During the training of the encoder, the loss function can be represented as L E = (-L R + λ1L F + λ2||1·a2||, and the goal of adjusting the encoder parameters can be to minimize the loss function.
[0100] In some embodiments, the loss function for training the encoder can further include a term related to the modulus of the partial data and the modulus of the remaining partial data in the sample encoded data. At least adjusting the model parameters of the encoder can further increase the modulus, for example, at least adjusting the model parameters of the encoder during the training of the encoding model can make the modulus greater than a preset value so that the modulus can not be close to zero. The preset value can be set as needed, such as 2, 5, etc. In some embodiments, the preset values corresponding to the modulus can be the same or different. During the training of the encoding model, the loss function can be represented as The goal of adjusting the encoder parameters can be to minimize the loss function, and minimizing the loss function can make the partial data take values with larger modulus in the value range, so that the modulus of the partial data is not close to zero or larger. Through this embodiment, while minimizing the correlation between the partial data and the remaining partial data in the sample encoded data, the values of the elements in the partial data and the remaining partial data in the sample encoded data can be prevented from being too small, thereby losing the significance of information representation.
[0101] In some embodiments, when the partial data is further subdivided into two or more task-specific data and task-common data, the loss function for training the encoder can further include a term related to the correlation between any two of the task-specific data and the remaining partial data in the encoded data, so that adjusting the model parameters of the encoder based on the loss function can further include minimizing the correlation between any two of the task-specific data and the remaining partial data in the encoded data. For example, minimizing the correlation between b1 and b4, b2 and b4, and b1 and b2. Figure 6 The correlation minimization between b1 and b4, the correlation minimization between b2 and b4, and the correlation minimization between b1 and b2 in the above formula. During the training of the encoder, the loss function can be represented as wherein, denotes a set of indices of the task-specific data and the algebraic symbol corresponding to the remaining part. The objective of adjusting the model parameters of the encoder can be to minimize the loss function.
[0102] In some embodiments, the loss function of the encoder training can further include terms related to the modulus of each task-specific data and the modulus of the remaining part of the sample encoded data. Adjusting the model parameters of the encoder can further increase the modulus of each task-specific data and the modulus of the remaining part of the encoded data. For example, when training the encoding model, adjusting the model parameters of the encoder can make the modulus of each task-specific data greater than a preset value and the modulus of the remaining part of the encoded data greater than a preset value, so that the modulus can not be close to zero or greater. When training the encoder, the loss function can be represented as The objective of adjusting the model parameters of the encoder is to minimize the loss function.
[0103] Through the present embodiment, when adjusting the model parameters of the encoder, by minimizing the correlation between the part of data corresponding to the preset position and the remaining part of the sample encoded data, or minimizing the correlation between any two of the two or more task-specific data and the remaining part of the encoded data, the training of the encoder to separate the part of data corresponding to the preset position and the remaining part of the data, or each task-specific data and the remaining part of the data more thoroughly can be achieved, and the effect of task prediction of the task processing model can be further improved.
[0104] In the foregoing embodiments, the model parameters of the encoder are mainly adjusted, at this time, it can be considered that the task prediction model and the data reconstruction model have been trained.
[0105] In some embodiments, the training process of the encoding model can further include training the task prediction model, i.e., adjusting the model parameters of at least one task processing model, and can further include training the data reconstruction model, i.e., adjusting the model parameters of the data reconstruction model. At this time, the encoder, the task prediction model, and the data reconstruction model are jointly trained. Specifically, one or more rounds of iterative training can be performed using sample data, wherein one round of iterative training includes:
[0106] In some embodiments, after inputting the sample data into the encoder to obtain the sample encoded data, the sample de-sensitization data obtained based on the task important data can be input into the corresponding task processing model to obtain the task prediction result, and the sample encoded data can be input into the data reconstruction model to obtain the reconstructed data. The loss function L F of the task processing model can be obtained based on the difference between the task prediction result and the corresponding reference standard. RAt this time, the parameters of the encoder can be fixed, and the model parameters of the task prediction model and the model parameters of the data reconstruction model are adjusted respectively based on the loss functions L F , L R , so as to minimize the respective loss functions.
[0107] In some embodiments, after the model parameter adjustment of the respective task processing models and the data reconstruction model, the adjusted at least one task processing model can be used to process the corresponding sample de- sensitized data respectively, to obtain the secondary task prediction results corresponding to the respective task processing models, and the adjusted data reconstruction model can be used to process the sample encoded data, to obtain the secondary reconstruction data. Further, the loss function L E = -L R + λL F can be determined based on the secondary task prediction results and the secondary reconstruction data. At this time, the model parameters of the task prediction model and the data reconstruction model are fixed, and only the model parameters of the encoder are adjusted, so that the loss function L E is minimized. In this way, one round of joint training is completed. In some embodiments, a correlation and a modulus constraint term can also be added to the above three loss functions respectively. Exemplarily, the loss function optimization target during joint training can be represented as
[0108] In some embodiments, when the data reconstruction model includes a generator and a discriminator, the training of the data reconstruction model in one round of joint training can be completed by referring to the one round of iterative training process of the data reconstruction model in step 450. In other words, one round of joint training can nest one round of joint training of the generator and the discriminator in the data reconstruction model.
[0109] Figure 7 is an exemplary flowchart of a data de- sensitization method according to some embodiments of the present specification.
[0110] In some embodiments, the method 700 can be executed by the processor 112. In some embodiments, the method 700 can be implemented by the data de- sensitization system 300 deployed on the processor 112.
[0111] As shown in Figure 7 , the method 700 can include:
[0112] Step 710: processing original data by using an encoder to obtain encoded data.
[0113] In some embodiments, the step 710 can be executed by the data encoding module 310.
[0114] The original data can be various types of data, such as text, pictures, voice, video, and the like. The original data can include sensitive data. In some embodiments, the original data can be input into an encoder, and the encoder processes the original data to obtain encoded data. The encoder can be obtained by training, and the training method of the encoder can refer to Figure 4 and related contents.
[0115] The method of processing the original data by the encoder to obtain the encoded data is similar to the method of processing the sample data by the encoder to obtain the sample encoded data. For more information about sensitive data, encoders, and encoders processing data to obtain encoded data, see step 410 and related descriptions.
[0116] In step 720, part of the data is extracted from the encoded data based on a preset position.
[0117] In some embodiments, step 720 can be performed by an encoded data extraction module 320.
[0118] In some embodiments, the encoded data output by the encoder can be used for business tasks. In some embodiments, the encoded data can be divided into two parts based on a preset position. The preset position can refer to the range of data, the division position of data, and the like. In the training method of the encoder, with the training process of the encoder (see related descriptions of flow 400), the preset position becomes the corresponding position of the de-identified data in the encoded data, that is, the part of data corresponding to the preset position can be provided as de-identified data to the task processing model to implement the corresponding business task. For more information about the preset position and extracting part of the data from the encoded data, see step 420 and related descriptions.
[0119] In step 730, the de-identified data corresponding to the original data is determined based on the part of the data.
[0120] In some embodiments, step 730 can be performed by a de-identified data determination module 330.
[0121] As mentioned earlier, the part of data corresponding to the preset position can be provided as de-identified data to the task processing model to implement the corresponding business task. In some embodiments, different task processing models can have different corresponding de-identified data. In some embodiments, one or more business processing models corresponding to the de-identified data can be extracted from the encoded data output by the encoder based on the preset position, to be provided to the corresponding task processing model respectively, to implement the corresponding business task.
[0122] In some embodiments, when the task processing model is one, the part of data extracted based on the preset position can be directly used as the de-identified data corresponding to the task processing model.
[0123] In some embodiments, the task processing model can be two or more, and the partial data extracted based on the preset position can be further divided into two or more task-specific data and task-common data. In some embodiments, each task-specific data can be combined with the task-common data to obtain the de-identification data corresponding to each task processing model.
[0124] For more information about determining the de-identification data based on the partial data in the encoded data, see step 430 and its related description.
[0125] It should be noted that the above description of the processes and methods is only for example and illustration, and does not limit the scope of the present specification. Those skilled in the art can make various modifications and changes to the processes and methods under the guidance of the present specification. However, these modifications and changes are still within the scope of the present specification. For example, the order of steps in the processes and methods can be changed, steps in different processes and methods can be combined, and so on.
[0126] The embodiments of the present specification also provide an encoder training apparatus, comprising at least one storage medium and at least one processor, the at least one storage medium is used to store computer instructions; the at least one processor is used to execute the computer instructions to implement an encoder training method. The method can comprise: processing sample data through an encoder to obtain sample encoded data; extracting partial data from the sample encoded data based on a preset position; determining sample de-identification data corresponding to the sample data based on the partial data; processing the sample de-identification data corresponding to each task processing model through at least one task processing model to obtain task prediction results corresponding to each task processing model; processing the sample encoded data through a data reconstruction model to obtain reconstructed data; and adjusting at least the model parameters of the encoder to reduce the difference between each task prediction result and the corresponding reference standard, and to increase the difference between the reconstructed data and the sample data.
[0127] The embodiments of the present specification also provide a data de-identification apparatus, comprising at least one storage medium and at least one processor, the at least one storage medium is used to store computer instructions; the at least one processor is used to execute the computer instructions to implement a data de-identification method. The method can comprise: processing original data through an encoder to obtain encoded data; extracting partial data from the encoded data based on a preset position; and determining de-identification data corresponding to the original data based on the partial data.
[0128] The beneficial effects that the embodiments of the present specification can bring include but are not limited to: (1) by training the encoder to divide the encoded data into two parts based on the preset position, and providing the part of data corresponding to the preset position to the business processing model as the de-sensitized data, and providing all the encoded data to the data reconstruction model to obtain the reconstructed data, so as to reduce the difference between the task prediction result and the corresponding reference standard, and increase the difference between the reconstructed data and the sample data, adjust the model parameters of the encoder as the target, realize the improvement of the encoding effect of the encoder to obtain the de-sensitized data that can meet the task prediction effect of the task processing model and is not easy to recover the original data; (2) the part of data corresponding to the preset position can be further subdivided to obtain task-specific data and task-common data, which realizes the cleaning of the de-sensitized data input into each task processing model, removes other data that is not important to the business task, and avoids interference; (3) the trained encoder can respectively gather each part of the task important data corresponding to each task processing model based on the corresponding part divided based on the preset position, so as to separate the task important data corresponding to each task processing model, so that the task important data extracted based on the preset position is only a part of the encoded data, further increasing the difficulty of obtaining the original data, and the part of data corresponding to the preset position can be further subdivided to obtain task-specific data and task-common data, which realizes the cleaning of the de-sensitized data input into each task processing model, removes other data that is not important to the business task, and avoids interference, further improving the task prediction effect of the task processing model. It should be noted that different embodiments can have different beneficial effects, and in different embodiments, the beneficial effects that can be obtained can be any one or a combination of the above, or any other beneficial effects that can be obtained.
[0129] The above has described the basic concepts, and it is obvious that the above detailed disclosure is only as an example and does not constitute a limitation on the present specification. Although it is not explicitly stated here, those skilled in the art can make various modifications, improvements and corrections to the present specification. Such modifications, improvements and corrections are suggested in the present specification, so such modifications, improvements and corrections still belong to the spirit and scope of the exemplary embodiments of the present specification.
[0130] At the same time, specific words are used in the present specification to describe the embodiments of the present specification. As "one embodiment", "an embodiment", and / or "some embodiments" means a certain feature, structure or characteristic related to at least one embodiment of the present specification. Therefore, it should be emphasized and noted that the "an embodiment" or "one embodiment" or "one alternative embodiment" mentioned in different places in the present specification does not necessarily refer to the same embodiment. In addition, certain features, structures or characteristics in one or more embodiments of the present specification can be properly combined.
[0131] Moreover, those skilled in the art will appreciate that the various aspects of the disclosure can be illustrated and described in connection with a number of various kinds of species or circumstances, including any new and useful processes, machines, products, or compositions of matter, or any new and useful improvements thereof, as described and claimed. The various aspects of the disclosure can be implemented entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by combinations of hardware and software. The above hardware or software can be referred to as a "data block", "module", "engine", "unit", "component", or "system". Furthermore, the various aspects of the disclosure can be manifested as computer products in one or more computer-readable media using, for example, computerized equipment, such as general purpose computers, laptop computers, personal digital assistants, cellular telephones, wireless communication devices, etc. These computer-readable media can include a computer-readable storage medium that can be any available medium or a combination of two or more of the following: a volatile memory, a non-volatile memory, a storage device, and a data transmission.
[0132] The computer storage media can include a propagated data signal with the computer program code embodied therein, e.g., in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. Computer storage media can be any media suitable for storing computer readable instructions, including a volatile memory, a non-volatile memory, a storage device, or any suitable combination thereof. The computer storage media can be transported by any suitable medium, including a wireless medium, a wired medium, or any suitable combination thereof.
[0133] The computer program code for carrying out the operations of the aspects of the disclosure can be written in any one or more of a variety of programming languages, including an object-oriented programming language such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, and the like, conventional procedural programming languages, such as the "C" programming language, Visual Basic, Fortran 2003, Perl, COBOL 2002, PHP, ABAP, dynamic programming languages, such as Python, Ruby and Groovy, or other programming languages. The program code can execute entirely on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic program code can be downloaded from an on-demand computing platform, such as Amazon Web Services, Microsoft Azure, or Google Cloud Platform, or a similar on-demand computing platform.
[0134] Furthermore, the order of the processing elements and sequences described in this specification are not intended to be construed as a limitation, unless specifically stated, but are included to provide a complete description of one or more embodiments of the present specification. Regardless of the particular sequence of processing elements and sequences, however, the description herein of a process should be understood to include any and all combinations of one or more elements, and sequences that can be perceived as either open-ended or specific.
[0135] Similarly, it is to be noticed that the term "comprising", used in the description, should not be interpreted as being restricted only to the means listed thereafter. It is to be understood that the term "comprising" means that any additional element, which is not specifically mentioned, is optionally present or can be added. In some embodiments, the description of an embodiment using terms such as "comprising", "having", "including" or "carrying" can also be interpreted using the term "consisting of". In some embodiments, the description of an embodiment using terms such as "comprising", "having", "including" or "carrying" can also be interpreted using the term "consisting of".
[0136] Some embodiments use numerical ranges to describe quantities of components, attributes, etc. It should be understood that the use of a numerical range herein is merely intended to reflect an approximate value, and that in some embodiments, the numerical range is modified by the words "about" or "approximately." Unless otherwise indicated, "about" or "approximately" means ±20% of the indicated value. Accordingly, numerical parameters in the description and claims are approximations, and thus can vary depending upon the requirements of the particular embodiment. In some embodiments, numerical parameters are determined by the use of common rounding techniques. Although the numerical ranges and parameters setting forth the broad scope of the embodiments of the specification are approximations, the numerical values set forth in the specific examples are reported as precisely as practicable. The numerical values set forth in the specific examples are provided to be as precise as reasonably possible. However, some variations may
[0137] Each patent, patent application, patent publication, and other material, articles, books, instructions, documents, that has been identified herein, is hereby incorporated herein by reference in its entirety for all purposes. Except in instances where the instant specification contradicts or contradicts away part of the incorporated material, articles, books, instructions, documents, the incorporated material, articles, books, instructions, documents are hereby incorporated by reference for all purposes. It is specifically intended that the description, definitions, and / or terminology used herein be interpreted as consistent with the description, definitions, and / or terminology used throughout the incorporated material, articles, books, instructions, documents.
[0138] Finally, it should be understood that the embodiments described herein are only given by way of example and that other modifications can occur to persons skilled in the art. Therefore, the scope of the present description is not intended to be limited to the embodiments described herein but is only limited by the claims that follow.
Claims
1. An encoder training method, comprising: The sample data is processed by the encoder to obtain the sample encoded data; Extract a portion of the data from the sample encoded data based on a preset location; Based on the aforementioned partial data, determine the de-identified sample data corresponding to the sample data; The corresponding sample desensitized data is processed by at least one task processing model to obtain the task prediction results corresponding to each task processing model. The sample encoded data is processed by a data reconstruction model to obtain reconstructed data; At least the model parameters of the encoder should be adjusted to reduce the difference between the prediction results of each task and the corresponding reference standard, and to increase the difference between the reconstructed data and the sample data.
2. The method of claim 1, at least adjusting the model parameters of the encoder, further reduces the correlation between the partial data and the remaining part of the encoded data.
3. The method as described in claim 2, wherein at least the model parameters of the encoder are adjusted, and the modulus of the partial data is not less than a preset value, and the modulus of the remaining part of the encoded data is not less than a preset value.
4. The method as described in claim 1, wherein the number of task processing models is two or more; the partial data includes two or more task-specific data and task-shared data; the step of determining the sample anonymized data corresponding to the sample data based on the partial data includes: Two or more task-specific data are combined with the task-shared data to obtain two or more de-identified sample data, which correspond one-to-one with two or more task processing models.
5. The method as described in claim 4, wherein the task-specific data, the task-shared data, and the remaining portion of the encoded data are of equal length.
6. The method of claim 4, wherein at least the model parameters of the encoder are adjusted, and the correlation between the two or more task-specific data and any two of the remaining portions of the encoded data is reduced.
7. The method of claim 6, wherein at least the model parameters of the encoder are adjusted, and the modulus values of the two or more task-specific data are not less than a preset value, and the modulus values of the remaining portion of the encoded data are not less than a preset value.
8. The method of claim 1, wherein adjusting at least the model parameters of the encoder to reduce the difference between the prediction results of each task and the corresponding reference standard, and to increase the difference between the reconstructed data and the sample data, further comprises: Adjust the model parameters of the at least one task processing model and the model parameters of the data reconstruction model.
9. The method of claim 8, wherein adjusting at least the model parameters of the encoder to reduce the difference between the prediction results of each task and the corresponding reference standard, and to increase the difference between the reconstructed data and the sample data, comprises: Adjust the model parameters of the at least one task processing model to reduce the difference between the prediction results of each task and the corresponding reference standard; Adjust the model parameters of the data reconstruction model to reduce the difference between the reconstructed data and the sample data; The corresponding sample desensitized data is processed using at least one adjusted task processing model to obtain the secondary task prediction results corresponding to each task processing model. The sample encoded data is processed using the adjusted data reconstruction model to obtain secondary reconstructed data. The model parameters of the encoder are adjusted to reduce the difference between the prediction results of each secondary task and the corresponding reference standard, and to increase the difference between the secondary reconstructed data and the sample data.
10. The method of claim 9, wherein the data reconstruction model includes a generator and a discriminator; the generator is used to process the sample encoded data to obtain reconstructed data; Adjusting the model parameters of the data reconstruction model to reduce the difference between the reconstructed data and the sample data includes: The reconstructed data is processed by a discriminator to obtain the corresponding score; The score reflects the probability that the discriminator determines the processed data to be true; Adjust the model parameters of the generator to increase the score.
11. The method of claim 10, wherein adjusting the model parameters of the data reconstruction model to reduce the difference between the reconstructed data and the sample data further comprises: The sample data is processed by a discriminator to obtain the corresponding score; The model parameters of the discriminator are adjusted so that the score corresponding to the reconstructed data decreases and the score corresponding to the sample data increases.
12. The method of claim 1, wherein the partial data is of equal length to the remaining portion of the encoded data.
13. The method of claim 1, wherein the adjustment of at least the encoder's model parameters is based on a loss function; the loss function includes a first part, a second part, a correlation constraint term, and a modulus constraint term, wherein, The first part reflects the difference between the prediction results of each task and the corresponding reference standard, and the second part reflects the difference between the reconstructed data and the sample data; the correlation constraint term reflects the correlation between the partial data and the remaining part of the encoded data, and the modulus constraint term reflects the modulus of the partial data and the modulus of the remaining part of the encoded data; or, the correlation constraint term reflects the pairwise correlation between two or more task-specific data and the remaining part of the encoded data, and the modulus constraint term reflects the modulus of each task-specific data and the modulus of the remaining part of the encoded data.
14. An encoder training system, comprising: The sample data encoding module is used to process sample data through the encoder to obtain sample encoded data; The sample encoded data extraction module is used to extract partial data from the sample encoded data based on a preset position. A sample desensitization data determination module is used to determine the sample desensitization data corresponding to the sample data based on the partial data; The task prediction module is used to process the corresponding desensitized sample data through at least one task processing model to obtain the task prediction results corresponding to each task processing model. The data reconstruction module is used to process the sample encoded data through a data reconstruction model to obtain reconstructed data; The parameter adjustment module is used to adjust at least the model parameters of the encoder so as to reduce the difference between the prediction results of each task and the corresponding reference standard, and to increase the difference between the reconstructed data and the sample data.
15. An encoder training apparatus, comprising at least one storage medium and at least one processor, the at least one storage medium being used to store computer instructions; the at least one processor being used to execute the computer instructions to implement the encoder training method as described in any one of claims 1 to 13.
16. A data anonymization method, comprising: The raw data is processed using an encoder to obtain encoded data; Extract a portion of the data from the encoded data based on a preset location; The data includes two or more task-specific data and task-shared data. Based on the aforementioned partial data, the de-identified data corresponding to the original data is determined, which further includes combining two or more task-specific data with the task-shared data respectively to obtain two or more de-identified data.
17. The method of claim 16, wherein the remaining portions of the task-specific data, task-shared data, and coded data, excluding the aforementioned portion of data, are of equal length.
18. The method of claim 16, wherein determining the de-identified data corresponding to the original data based on the partial data comprises: The aforementioned portion of the data is directly identified as the de-identified data.
19. The method of claim 18, wherein the partial data is of equal length to the remaining portion of the encoded data excluding the partial data.
20. The method of claim 16, wherein the encoder is obtained by the training method of any one of claims 1 to 13, and the preset position becomes the position of the desensitized data in the encoded data during the training process of the encoder.
21. A data anonymization system, comprising: The data encoding module is used to process the raw data using the encoder to obtain encoded data. The encoded data extraction module is used to extract a portion of the data from the encoded data based on a preset location; the portion of the data includes two or more task-specific data and task-shared data. The de-identified data determination module is used to determine the de-identified data corresponding to the original data based on the partial data. It further includes combining two or more task-specific data with the task-shared data to obtain two or more de-identified data.
22. A data desensitization apparatus, comprising at least one storage medium and at least one processor, wherein the at least one storage medium is used to store computer instructions; and the at least one processor is used to execute the computer instructions to implement the data desensitization method as described in any one of claims 16 to 20.
Citation Information
Patent Citations
Image processing method and system, image recognition model training method and system and image recognition method and system
CN112966737A
Training method and device for data generation system based on differential privacy
CN113642731A