Learning for Converting Confidential Data Using Variable Distribution Storage
The autoencoder with a parameterized loss function addresses inefficiencies in data transformation by dynamically preserving data distribution and balance, enabling real-time, cost-effective data streaming with optimal utility and privacy.
Patent Information
- Application Number
- JP2023557029
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-24
- Filing Date
- 2022-03-23
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2042-03-23
AI Technical Summary
Current data transformation systems are costly and inefficient, requiring stateful processing and sacrificing data integrity to preserve data distribution, which hinders real-time data streaming and fails to balance data utility and privacy needs effectively.
Utilizing an autoencoder with a parameterized loss function to dynamically transform confidential data in real-time while preserving the distribution of data values, controlled by a policy that balances data utility and privacy needs through adjustable parameter coefficients.
Enables fast, stateless data transformation that maintains data integrity and distribution, allowing for real-time data streaming and optimal balance between data utility and privacy, reducing the need for costly preprocessing.
Smart Images

Figure 0007702204000001 
Figure 0007702204000002 
Figure 0007702204000003
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to systems for data security / privacy, and more particularly to dynamically and in real-time transforming sensitive data of a data asset during data dissemination to a data consumer. Certain transformations to data during an anonymization operation are achieved while attempting to preserve the data distribution of the output transformed values to a desired degree with respect to the input data. This is achieved by using an autoencoder for transformation and a policy-based parameterized loss function for data distribution control.
Background Art
[0002] Data dissemination is the distribution or transmission of data to a data consumer. A data consumer can be, for example, a human; an entity such as a business, company, corporation, organization, institution or agency; a software application; or an online service. Data security is the process of protecting data by adopting a set of policies that identify the relative importance of different data sets, the confidentiality of different data sets, and the regulatory compliance requirements corresponding to different data sets, and then applying the appropriate policies to secure a given data set. Elements of data security can include confidentiality, integrity, and availability. These elements can be used as a guide to keep sensitive data in a state protected from unauthorized access. For example, confidentiality ensures that data is accessed only by authorized users. Integrity ensures that data is accurate. Availability ensures that data is usable and accessible to meet the needs of a data consumer.
[0003] Data privacy is the relationship between data collection and propagation, data privacy expectations, and the regulatory issues surrounding them. Data privacy presents challenges in protecting an individual's confidential data or personally identifiable information while simultaneously using the data. Personally identifiable information is information such as a name, address, phone number, or social security number that corresponds to an identifiable person and can be used for the purpose of identifying that particular person.
Summary of the Invention
[0004] According to an exemplary embodiment, a computer-implemented method is provided for preserving the distribution of data values of a data asset in a data anonymization operation. An autoencoder is used to perform data anonymization on selected rows in a data asset by converting the data values of confidential data in a set of related row cells of a column of interest and placing them in a conversion buffer. A loss function value for data anonymization of the selected rows is generated using a loss function that includes parameter coefficients specified in a policy enforcement decision. The loss function value is compared to a loss function threshold. In response to determining that the loss function value is greater than the loss function threshold based on this comparison, the converted data values in the conversion buffer are posted to an output buffer that is labeled as output using a forward mapping to actual row cell values suitable for a particular user. For the output of the data asset, the output buffer labeled as output is transferred to the next row. According to another exemplary embodiment, a computer system and a computer program product are provided for preserving the distribution of data values of a data asset in a data anonymization operation.
Brief Description of the Drawings
[0005]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7A
Figure 7B
Figure 7C
[0006] The present invention may be a system, method, or computer program product, or a combination thereof, at an integrated technical detail level. This computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0007] The computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy (registered trademark) disk, mechanically encoded devices such as punch cards or raised structures in grooves having instructions recorded thereon, and suitable combinations thereof. The computer-readable storage medium as used herein should not be construed to be a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0008] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a respective computing / processing device, or can be downloaded from an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. This network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers those computer-readable program instructions for storage on a computer-readable storage medium within the respective corresponding computing / processing device.
[0009] Computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or configuration data for an integrated circuit, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk® and C++, and procedural programming languages such as the "C" programming language or similar programming languages. These computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, executed as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on a remote computer or remote server. In the last scenario above, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or this connection may be implemented to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to perform aspects of the present invention, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may utilize the state information of these computer-readable program instructions to personalize the electronic circuit and thereby execute these computer-readable program instructions.
[0010] In this specification, aspects of the present invention are described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It is understood that each block of these flowcharts and / or block diagrams and combinations of blocks in these flowcharts and / or block diagrams can be implemented by computer-readable program instructions.
[0011] These computer-readable program instructions can be provided to a processor of a computer or to a processor of other programmable data processing apparatus that forms a machine, in such a way that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, cause means to be generated for performing the functions / operations specified in the blocks of these flowcharts and / or block diagrams and / or both. These computer-readable program instructions can further be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, or other device or combinations thereof to function in a particular manner, such that the computer-readable storage medium storing the instructions includes a product that includes instructions for performing the aspects of the functions / operations specified in the blocks of these flowcharts and / or block diagrams and / or both.
[0012] These computer-readable program instructions can further be loaded onto a computer, other programmable apparatus, or other device such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a process implemented by the computer, in such a way that the instructions, when executed on the computer, other programmable apparatus, or other device, perform the functions / operations specified in the blocks of these flowcharts and / or block diagrams and / or both.
[0013] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in those flowcharts or block diagrams may represent a module, segment, or portion of instructions that include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may be executed in an order different than that shown in the figures. For example, two blocks shown in succession may in fact be executed as one step, executed simultaneously, substantially simultaneously, or partially or fully overlapping in time, or the blocks may sometimes be executed in the reverse order depending on the functionality involved. It should also be noted that each block of those block diagrams or flowcharts, or combinations of blocks in those block diagrams or flowcharts, or both, can be implemented by a dedicated system based on hardware that executes the specified functions or operations or a combination of dedicated hardware and computer instructions.
[0014] Next, referring to the figures, and in particular FIGS. 1 and 2, there is shown a diagram of a data processing environment in which an exemplary embodiment can be implemented. FIGS. 1 and 2 are intended to be merely examples, and it should be understood that FIGS. 1 and 2 are not intended to state or imply any limitation as to the environments in which different embodiments can be implemented. Many modifications can be made to the environments shown.
[0015] FIG. 1 shows a diagram of a network of a data processing system capable of implementing an exemplary embodiment. Network data processing system 100 is a network of computers, data processing systems, and other devices capable of implementing an exemplary embodiment. Network data processing system 100 includes network 102, and network 102 is a medium used to provide communication links between computers, data processing systems, and other devices connected together within network data processing system 100. Network 102 can include connections such as, for example, wired communication links, wireless communication links, fiber optic cables, and the like.
[0016] In the example shown, server 104 and server 106 are connected to network 102 along with storage 108. Server 104 and server 106 may be, for example, server computers having a high-speed connection to network 102. Further, server 104 and server 106 each provide one or more data privacy services by physically converting confidential data of a data asset (e.g., a rectangular data set consisting of columns and rows) requested by a user into anonymized data values while storing the distribution of the values of the data asset to a degree specified based on a policy using an autoencoder including a loss function. It should also be noted that server 104 and server 106 may each represent a group of servers within one or more data centers. Alternatively, server 104 and server 106 may each represent a number of computing nodes within one or more cloud environments.
[0017] The network 102 is also connected to clients 110, 112, and 114. Clients 110, 112, and 114 are clients of servers 104 and 106. In this example, clients 110, 112, and 114 are shown as desktop or personal computers having a wired communication link to network 102. However, clients 110, 112, and 114 are merely examples, and it should be noted that clients 110, 112, and 114 may represent other types of data processing systems such as network computers, laptop computers, handheld computers, smartphones, smart televisions, etc., having, for example, a wired or wireless communication link to network 102. The users (i.e., data consumers) corresponding to clients 110, 112, and 114 can utilize clients 110, 112, and 114 to access data assets hosted or protected by servers 104 and 106. The data assets hosted or protected by servers 104 and 106 can be any type of data set (such as transaction data, marketing data, financial data, or health management data, etc.) including confidential data (such as names, addresses, phone numbers, social security numbers, credit card numbers, etc.) that can identify an individual as an individual and cannot be accessed without the explicit consent of that individual.
[0018] Storage 108 is a network storage device that can store any type of data in a structured format or an unstructured format. Further, storage 108 may represent a plurality of network storage devices. Further, storage 108 can store a plurality of different data assets protected by servers 104 and 106. Further, storage 108 can store other types of data, such as authentication or credential data that may include, for example, a username, password, and biometric template associated with a client device user.
[0019] Further, it should be noted that network data processing system 100 may include any number of additional servers, clients, storage devices, and other devices not shown. Program code located within network data processing system 100 may be stored on a computer-readable storage medium or a set of computer-readable storage media and may be downloaded to a computer or other data processing device for use. For example, program code may be stored on a computer-readable storage medium on server 104 and may be downloaded to client 110 via network 102 for use on client 110.
[0020] In the illustrated example, network data processing system 100 can be implemented as several different types of communication networks, such as, for example, the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a telecommunications network, or any combination thereof. FIG. 1 is intended to be merely an example, and it is not intended that FIG. 1 limit the architecture of different exemplary embodiments.
[0021] As used herein, when used with respect to a plurality of items, "some" means one or more of the plurality of items. For example, "some different types of communication networks" means one or more different types of communication networks. Similarly, when used with respect to a plurality of items, "a set of" means one or more of those items.
[0022] Furthermore, when used with a list of items, the term "at least one of" means that different combinations of one or more of the listed items can be used, and it may be that only one of each item in the list is required. In other words, "at least one of" can use any combination and any number of items from the list, but not all of the items in the list are required. An item can be a particular object, thing, or category.
[0023] For example, without limitation, "at least one of item A, item B, or item C" can include item A, item A and item B, or item B. This example can also include item A, item B, as well as item C, or item B and item C. Of course, any combination of these items can exist. In some exemplary examples, "at least one of" can be, for example, without limitation, 2 item As, 1 item B, and 10 item Cs, 4 item Bs and 7 item Cs, or other suitable combinations.
[0024] Next, referring to FIG. 2, a diagram of a data processing system according to an exemplary embodiment is shown. The data processing system 200 is an example of a computer, such as the server 104 of FIG. 1, into which computer-readable program code or instructions for implementing the data privacy process of the exemplary embodiment can be placed. In this example, the data processing system 200 includes a communication fabric 202 that provides communication between a processor unit 204, a memory 206, a persistent storage 208, a communication unit 210, an input / output (I / O) unit 212, and a display 214.
[0025] The processor unit 204 serves to execute instructions regarding software applications and programs that may be loaded in the memory 206. Depending on the particular implementation, the processor unit 204 may be a set of one or more hardware processor devices, or may be a multi-core processor.
[0026] Memory 206 and persistent storage 208 are examples of storage devices 216. As used herein, a computer-readable storage device or computer-readable storage medium is any hardware that can store, temporarily or permanently, for example, but not limited to, data, computer-readable program code in a functional form, or other suitable information, or a combination thereof. Further, a computer-readable storage device or computer-readable storage medium excludes propagation media such as transient signals. Further, a computer-readable storage device or computer-readable storage medium may represent a set of computer-readable storage devices or a set of computer-readable storage media. In these examples, memory 206 may be, for example, random access memory (RAM), or other suitable volatile or non-volatile storage device such as flash memory. Depending on the particular implementation, persistent storage 208 can take various forms. For example, persistent storage 208 may include one or more devices. For example, persistent storage 208 may be a disk drive, a solid state drive, a writable optical disk, a rewritable magnetic tape, or some combination thereof. The medium used by persistent storage 208 may be removable. For example, a removable hard drive may be used for persistent storage 208.
[0027] In this example, the persistent storage 208 stores a data distribution preserver 218. However, although the data distribution preserver 218 is shown as being within the persistent storage 208, it should be noted that in an exemplary alternative embodiment, the data distribution preserver 218 may be a separate component of the data processing system 200. For example, the data distribution preserver 218 may be a hardware component coupled to the communication fabric 202, or may be a combination of a hardware component and a software component. In another exemplary alternative embodiment, a first set of components of the data distribution preserver 218 may be placed on the data processing system 200, and a second set of components of the data distribution preserver 218 may be placed on a second data processing system such as the server 106 of FIG. 1, for example.
[0028] The data distribution preserver 218 controls a process of dynamically converting in real time the confidential data of a data asset during the distribution of data to the requesting data consumer. The data distribution preserver 218 achieves a specific transformation on the requested data that preserves the transformed values of the output data distribution to the desired degree with respect to the input data during the anonymization operation. The data distribution preserver 218 declaratively controls this anonymization operation based on a policy using an autoencoder that includes a parameterized loss function.
[0029] As a result, the data processing system 200 operates as a dedicated computer system that enables the data distribution preserver 218 within the data processing system 200 to convert the confidential data in the data asset while preserving the distribution of the original data asset. In particular, the data distribution preserver 218 transforms the data processing system 200 into a dedicated computer system as compared to currently available general-purpose computer systems that do not have a data distribution preserver 218.
[0030] In this example, communication unit 210 provides communication with other computers, data processing systems, and devices via a network such as network 102 of FIG. 1. Communication unit 210 can provide communication by using both physical communication links and wireless communication links. The physical communication link can establish a physical communication link for data processing system 200 by using, for example, wires, cables, universal serial bus, or other physical technologies. The wireless communication link can establish a wireless communication link for data processing system 200 by using, for example, shortwave, high frequency, ultra-high frequency, microwave, wireless fidelity (Wi-Fi (registered trademark)), Bluetooth (registered trademark) technology, global system for mobile communications (GSM), code division multiple access (CDMA), second generation (2G), third generation (3G), fourth generation (4G), 4G long term evolution (LTE), LTE advanced, fifth generation (5G), or other wireless communication technologies or standards.
[0031] Input / output unit 212 enables the input and output of data with other devices that may be connected to data processing system 200. For example, input / output unit 212 can provide connections for user input via a keypad, keyboard, mouse, microphone, or other suitable input device or combinations thereof. Display 214 provides a mechanism for displaying information to the user. Display 214 may include a touch screen function to enable, for example, performing on-screen selections or entering data through a user interface.
[0032] Instructions for an operating system, an application, or a program, or a combination thereof, may be disposed within storage device 216 that communicates with processor unit 204 through communication fabric 202. In this example for illustration, these instructions are in a functional form on persistent storage 208. These instructions can be loaded into memory 206 for execution by processor unit 204. Processor unit 204 can execute processes of different embodiments using computer-executable instructions that may be located in a memory such as memory 206. These program instructions are referred to as program code, computer-usable program code, or computer-readable program code, and may be read and executed by a processor within processor unit 204. In different embodiments, these program instructions may be implemented on different computer-readable physical storage devices such as memory 206 or persistent storage 208.
[0033] Program code 220 is selectively disposed in a functional form on removable computer-readable medium 222, and program code 220 may be loaded into data processing system 200 or transferred to data processing system 200 for execution by processor unit 204. Program code 220 and computer-readable medium 222 form computer program product 224. In one example, computer-readable medium 222 may be computer-readable storage medium 226 or computer-readable signal medium 228.
[0034] In these exemplary examples, the computer-readable storage medium 226 is not a medium that propagates or transmits the program code 220, but rather a physical or tangible storage device used to store the program code 220. The computer-readable storage medium 226 may include, for example, an optical or magnetic disk inserted or disposed within a drive or other device that is part of the persistent storage 208 for transfer to a storage device such as a hard drive that is part of the persistent storage 208. The computer-readable storage medium 226 may take the form of persistent storage such as a hard drive, thumb drive, or flash memory connected to the data processing system 200.
[0035] Alternatively, the program code 220 can also be transferred to the data processing system 200 using the computer-readable signal medium 228. The computer-readable signal medium 228 may be, for example, a propagated data signal that includes the program code 220. For example, the computer-readable signal medium 228 may be an electromagnetic signal, an optical signal, or other suitable type of signal. These signals may be transmitted via a communication link such as a wireless communication link, an optical fiber cable, a coaxial cable, a wire, or other suitable type of communication link.
[0036] Furthermore, as used herein, "computer-readable medium 222" can be singular or plural. For example, program code 220 can be placed on a computer-readable medium 222 in the form of a single storage device or system. In another example, program code 220 can be placed on computer-readable media 222 that are distributed within a number of data processing systems. In other words, some of the instructions in program code 220 can be placed within one data processing system, and other instructions in program code 220 can be placed within one or more other data processing systems. For example, a portion of program code 220 can be placed within a computer-readable medium 222 within a server computer, and another portion of program code 220 can be placed within a computer-readable medium 222 located within a set of client computers.
[0037] It is not intended to limit the architecture of the different components shown of data processing system 200 to the manner in which different embodiments can be implemented. In some illustrative examples, one or more of the components can be incorporated into another component, or alternatively, one or more of the components can form a part of another component. For example, in some illustrative examples, memory 206 or a portion thereof can be incorporated into processor unit 204. Those different illustrative embodiments can be implemented within a data processing system that includes components in addition to, or instead of, the components described with respect to data processing system 200. The other components shown in FIG. 2 can be different from those shown for purposes of illustration. Those different embodiments can be implemented using any hardware device or system capable of executing program code 220.
[0038] In another example, a bus system may be used to implement the communication fabric 202, which may be composed of one or more buses such as a system bus or an input / output bus. Of course, this bus system can be implemented using any suitable type of architecture that provides for the transfer of data between different components or devices connected to the bus system.
[0039] In dynamic data distribution where data access is controlled by a policy, it is clear that each data consumer accessing the data through the distribution layer (e.g., the user in a user context) may encounter different applicable transformations of the value of that data that are visible to that particular data consumer. With respect to storage and computational resource costs, it is prohibitively expensive to preprocess every known data asset a priori for each different data consumer. As a result, there is a need for a system that, in each instance of data access, executes the necessary data transformations as quickly as possible, with the data asset in the form of raw data assets. This system will require a fast transformation method that uses a stateless approach to the rows of data and can therefore use streaming. Further, in most data consumption use cases, there should be a parametric approach to the loss function to determine the trade-off between data utility needs and data privacy needs. This parametric approach to the loss function requires an additional mechanism that needs to exist in addition to the actual operation of data transformation. However, this approach does not currently exist in data transformation systems.
[0040] Current systems that perform data transformation rely on costly classical deterministic algorithmic transformation methods such as tokenization, redaction, obfuscation, etc. Each of these methods requires the implementation of a specific transformation type. Some of these transformation types are highly stateful and require the complete processing of the source data asset before attempting to generate the first output row (i.e., tuple) of data, which hinders data streaming use cases.
[0041] There are few systems that attempt to preserve the distribution of data transformation outputs. Furthermore, these systems are known to sacrifice the integrity of row data and be single-column based. Exemplary embodiments utilize a trained autoencoder that generates the necessary physical transformation of data using an adjustable parameter coefficient of a loss function to achieve the required variability in data value distribution. An autoencoder is a type of artificial neural network used in unsupervised learning. In other words, an autoencoder does not require labeled input data to enable learning. Typically, an autoencoder has an input layer, one or more internal hidden layers that perform data processing, and an output layer. The training of an autoencoder is performed by backpropagation.
[0042] For each row of data in the data asset to be transformed, an exemplary embodiment selects columns of interest (e.g., attributes) in the data asset that require transformation as inputs to an autoencoder that includes the loss function of the exemplary embodiment. The exemplary embodiment selects those columns based on one or more policies. The policies include, for example, columns of interest such as names, salaries, addresses, social security numbers, credit card numbers, or other types of confidential information that need to be transformed before a particular data consumer (e.g., a user) accesses such information in a particular context such as the geographical location when accessing the request, the time of the access request, the role when accessing the request. A particular example of a policy is, for example, "if the data asset is confidential, the data asset includes a name, and the data asset includes a salary, the transformation ((pseudonymization of name, salary)) preserves the distribution 0.75". This particular example shows a sample set of policies that can drive a policy enforcement decision and subsequent transformation of the requested data asset to preserve the distribution of the values of the requested data asset to a defined degree based on the policy at the time of distribution.
[0043] The policy may further include a best fitting statistical distribution, such as a normal distribution, a lognormal distribution, or a beta distribution, for the columns of interest, along with the parameters of the best fitting statistical distribution, such as the mean, minimum value, maximum value, etc. Further, the policy may include parameter coefficients of a loss function for balancing the data utility needs of data consumers and the data privacy needs of individuals, such as "ρ" (rho), which evaluates the amount of variance between the data asset input and the data asset output within the columns of interest; "φ" (phi), which evaluates the amount of variance within one row among the multiple columns of interest between the data asset input and the data asset output; and "τ" (tau), which evaluates the average amount of variance of the columns of interest across the entire data asset between the data asset input and the data asset output. High data utility for a data consumer means data having values close to the original data values. High data privacy for an individual means data that does not expose any confidential information.
[0044] It should be noted that in an exemplary embodiment, an autoencoder can be selected from an autoencoder library for a specific data asset, a specific column of interest including related data classes within the data asset, a group of columns of interest within the data asset, or a specific row within the columns of interest. The exemplary embodiment can utilize a two-dimensional arrangement of the autoencoder library. For example, the first dimension can be based on the columns of interest including related data classes, and the second dimension can be based on the data utility needs versus the data privacy needs. The second dimension can be, for example, a coarse value of a loss function corresponding to the selection of the depth of value distribution preservation to a predetermined degree, such as X% of value distribution preservation for the data asset output indicated by the policy.
[0045] Exemplary embodiments can train an autoencoder from any data asset input, whether it is all available data assets or a subset of data assets, based on the desired type or category of the data asset. Further, it should be noted that the loss function parameter coefficient rho is on a sliding scale. For example, when the value of rho is zero, the autoencoder is incentivized not to consider how close the output of the data asset is to the input of the data asset, so complete data privacy exists and data utility does not exist because random values are generated for the row cells of the columns of interest. Conversely, when the value of rho is as large as possible, there is no data privacy and no complete data utility because the data asset output is the same as the data asset input. Further, the loss function parameter coefficient phi is also on a sliding scale. For example, the value of phi controls the amount of variance between two columns of interest in the data asset input relative to the amount of variance between the same two columns of interest in the data asset output. This comparison captures the dependency between the columns of interest.
[0046] To control the balance between the needs for data privacy and the needs for data utility, exemplary embodiments observe different data asset outputs using classical or standard deterministic transformation methods by varying the loss function parameter coefficients rho and phi. This classical deterministic transformation method provides ground truth when training the autoencoder. Active autoencoder learning is based on the backpropagation of the calculated loss function value as reinforcement or penalty when compared to a specified loss function threshold. For example, a calculated loss function value greater than the specified threshold is treated as reinforcement, and a calculated loss function value less than the specified threshold is treated as a penalty.
[0047] Furthermore, exemplary embodiments minimize the exposure of sensitive data (e.g., unlinkability) by calculating the entropy of the input data asset. Entropy quantifies the amount of uncertainty contained in the value of the input data asset. Exemplary embodiments further determine an entropy threshold for the input data asset. Additionally, exemplary embodiments can define a policy to further include the entropy threshold. Exemplary embodiments can generate a policy enforcement decision based on the entropy of the data asset. Additionally, exemplary embodiments can define a policy to include the sensitivity of the data asset with respect to a transformation change. Exemplary embodiments can generate a policy enforcement decision based on the sensitivity of the data asset with respect to a transformation change.
[0048] Exemplary embodiments generate anonymized data for the row cells of the columns of interest while maintaining the original distribution of the values based on a policy that includes the entropy threshold and the sensitivity of the data asset with respect to the transformation. Additionally, when the data asset has a certain level of sensitivity, exemplary embodiments utilize a Laplacian noise function to add Laplacian noise while performing a distribution-preserving transformation to provide an increase in data privacy. By paying attention to the change in the distribution caused by the addition of Laplacian noise, a data consumer or user can still utilize the transformed data. Adding noise from the Laplacian distribution function to the data asset output (i.e., the anonymized data values of the columns of interest across multiple rows) provides a differential privacy adjustment to the data asset output to prevent distribution inference. Differential privacy enables sharing information about a data asset by describing patterns within the data asset while withholding confidential information corresponding to individuals within the data asset.
[0049] Exemplary embodiments utilize a rectangular data set known herein as a data asset. Exemplary embodiments can assign a particular data asset to a particular transformation based on one or more policies and a particular data consumer attempting to access the data asset. Exemplary embodiments profile each received data asset to detect a data class corresponding to each data asset. A policy provides an ordered set of transformations required for a particular access request. These transformations are based on the relevant data classes within a particular data asset and the current policies within the system. The depth of distributed preservation is achieved by parameterizing the coefficient tau of the loss function. Exemplary embodiments perform only pseudonymization transformations. However, exemplary alternative embodiments can perform other transformation types using a single column autoencoder without using a group column autoencoder.
[0050] Accordingly, exemplary embodiments provide one or more technical solutions that solve a technical problem by transforming confidential data of a requested data asset while preserving the distribution of values of the data asset to a defined extent. As a result, these one or more technical solutions provide a technical effect and a practical use in the field of data privacy.
[0051] Next, referring to FIG. 3, there is shown a diagram illustrating an example of a data discovery, data classification, and autoencoder-based training and library maintenance process according to an exemplary embodiment. The data discovery, data classification, and autoencoder-based training and library maintenance process 300 can be implemented on a computer such as the server 104 of FIG. 1 or the data processing system 200 of FIG. 2.
[0052] In this example, the data discovery, data classification, and autoencoder-based training and library maintenance process 300 includes the raw, uncurated input data asset 302, data profiler 304, data asset catalog 306, data asset profile and data class best fit hyperplane storage 308, physical data storage 310, and autoencoder library 312. However, it should be noted that the data discovery, data classification, and autoencoder-based training and library maintenance process 300 is intended to be merely an example and is not intended to be a limitation on the exemplary embodiments. In other words, the data discovery, data classification, and autoencoder-based training and library maintenance process 300 can include more or fewer components than those illustrated. For example, one component can be split into two or more components, two or more components can be combined into one component, or components not shown can be added.
[0053] The raw, uncurated input data asset 302 is a rectangular (e.g., relational) data set consisting of columns and rows. Further, the raw, uncurated input data asset 302 can represent any type of data set that can identify an individual as an individual and includes sensitive data such as, for example, name, address, phone number, social security number, salary, etc. The raw, uncurated input data asset 302 is input into the data profiler 304. The data profiler 304 can represent any type of data profiler that can detect data classes of interest that include sensitive information within the raw, uncurated input data asset 302.
[0054] A data distribution storage device, such as the data distribution storage device 218 in FIG. 2, registers the input data asset 302 in the data asset catalog 306 and determines that the input data asset 302 is a new data asset. The data distribution storage device further stores all data asset profiles and their corresponding data class best-fit hyperplanes in the data asset profile and data class best-fit hyperplane storage 308. Further, at 314, the data distribution storage device determines whether a set of autoencoders, consisting of one or more autoencoders for the new data asset, exists in the autoencoder library 312.
[0055] If the data distribution saver determines that there is currently no set of autoencoders for the new data asset in the autoencoder library 312, then at 316, the data distribution saver generates a new set of autoencoders for the new data asset. Further, at 318, while profiling the new data asset (i.e., the raw, unmanaged input data asset 302), the data distribution saver determines a set of best-fit hyperplanes consisting of one or more best-fit hyperplanes for the detected data classes of interest. Further, at 320, the data distribution saver executes a row-by-row data reading process for the new data asset from the actual data storage 310. Further, at 322, using a fixed adaptive distribution preservation threshold or a configured distribution preservation threshold at the data asset level of the new data asset, the data distribution saver uses the determined set of best-fit hyperplanes for the data classes of interest, the row-by-row data reading, and the stored data class-based history transformation 324 as inputs to simulate the implementation. If the data distribution saver determines that there is currently a set of autoencoders for the new data asset in the autoencoder library 312, the data distribution saver uses the row-by-row data reading and the stored data class-based history transformation 324 as inputs, using a fixed adaptive distribution preservation threshold or a configured distribution preservation threshold at the data asset level of the new data asset, to simulate the implementation.
[0056] At 326, the data distribution saver performs labeling using the normal or classical deterministic transformation, value mapping, and row cell value embeddings of the new data asset. At 328, the data distribution saver further performs an actual row cell level inverse mapping using data class embeddings. The data distribution saver stores this actual row cell level inverse mapping in the map store 330.
[0057] At 332, the data distribution saver performs at least one of base training of a new set of autoencoders for its new data asset or additional training of one or more existing autoencoders to form an autoencoder training set 334. The data distribution saver uses the autoencoder training set 334 to train the autoencoders in the autoencoder library 312. The autoencoder library 312 includes a plurality of different autoencoders. For example, the autoencoder library 312 can include a group of autoencoders for one data asset, one autoencoder for one data asset, one autoencoder for one data class of interest within a data asset, one autoencoder for a particular row or transformation type within a data asset, and the like.
[0058] Next, referring to FIG. 4, there is shown a diagram illustrating an example of a dynamic data distribution and active autoencoder training process using policy enforcement, according to an exemplary embodiment. The dynamic data distribution and active autoencoder training process 400 using policy enforcement can be implemented on a computer such as the server 104 of FIG. 1 or the data processing system 200 of FIG. 2.
[0059] In this example, the dynamic data distribution and active auto-encoder training process 400 using policy enforcement includes user 402, data distribution / access layer 404, previously profiled and curated data asset 406, policy enforcement point 408, policy decision point 410, physical data storage 412, and auto-encoder library 414. However, it should be noted that the dynamic data distribution and active auto-encoder training process 400 using policy enforcement is intended to be merely an example and is not intended as a limitation to exemplary embodiments. In other words, the dynamic data distribution and active auto-encoder training process 400 using policy enforcement can include more or fewer components than the illustrated components. For example, one component can be divided into two or more components, two or more components can be combined into one component, or components not shown can be added.
[0060] User 402 is a data consumer. User 402 can be, for example, a human, a process, an application, a service, or a system. User 402 submits a data distribution request 416 for input data asset 418 to data distribution / access layer 404. User 402 can submit a data distribution request 416 with a specific user context. The user context can be, for example, the location of the originator who submitted the data distribution request 416, the day of the week and time when user 402 submitted the data distribution request 416, and the like. Data distribution / access layer 404 sends data distribution request 416 to policy enforcement point 408.
[0061] Policy enforcement point 408 sends a policy enforcement decision request corresponding to data distribution request 416, the user, and the user context of data distribution request 416 to policy decision point 410. Policy decision point 410 selects a set of policies corresponding to data distribution request 416, the user, and the user context of data distribution request 416. Policy decision point 410 generates a data class-based policy enforcement decision based on the selected policies. Policy decision point 410 sends the data class-based policy enforcement decision to policy enforcement point 408. At 420, policy enforcement point 408 saves the data class-based policy enforcement decision.
[0062] At 422, a data distribution storage, such as data distribution storage 218 in FIG. 2 for example, selects a set of autoencoders for input data asset 418 from autoencoder library 414. At 422, while selecting a set of autoencoders, the data distribution storage uses input data asset 418 as a reference to retrieve actual data from actual data storage 412. At 424, the data distribution storage uses the selected set of autoencoders for row-centric processing of the rows in the row buffer. At 426, the data distribution storage uses a loss function to perform loss function value calculation for the rows in the row buffer. Further, at 428, the data distribution storage compares the loss function value with a distribution storage threshold based on the data class-based policy enforcement decision. This distribution storage threshold is an adaptable threshold or a configured threshold.
[0063] At 430, the data distribution saver performs a determination as to whether the calculated loss function value is greater than the distribution save threshold. If the data distribution saver determines that the calculated loss function value is greater than the distribution save threshold, then at 432, the data distribution saver processes the rows in the row buffer using the actual row cell value mapping with data class value embedding. The data distribution saver retrieves the actual row cell value mapping from the map store 434. The data distribution saver uses the processing of the rows in the row buffer at 432 as the data distribution response to the data distribution request 416. The data distribution saver sends this data distribution response to the data distribution / access layer 404. The data distribution / access layer 404 then sends this data distribution response to the user 402. Alternatively, the data distribution saver can optionally also generate an output data asset 436.
[0064] If the data distribution saver determines that the calculated loss function value is less than or equal to the distribution save threshold, then at 438, the data distribution saver utilizes a sampler for random row samples from the row buffer. Further, at 440, the data distribution saver saves the random row samples from the row buffer to an under-threshold penalty buffer for autoencoder regularization to prevent overfitting of the autoencoder. If the calculated loss function value was greater than the distribution save threshold, then at 442, the data distribution saver saves the random row samples from the row buffer to an over-threshold enhancement buffer.
[0065] The data distribution preservator performs labeling using a normal or classical deterministic transformation with reverse-direction row cell value mapping embedding from the row buffer, at 444, using a below-threshold penalty buffer and an above-threshold penalty buffer. At 446, the data distribution preservator generates a small active learning autoencoder training set using labeling with a normal deterministic transformation. The smaller the training set, the shorter the time required to train the autoencoder. The data distribution preservator trains the autoencoders in the autoencoder library 414 using the active learning autoencoder training set.
[0066] Referring now to FIG. 5, there is shown a diagram depicting a particular example of an autoencoder that includes a loss function, according to an exemplary embodiment. The autoencoder 500 that includes a loss function can be, for example, the autoencoder 411 of FIG. 4. The autoencoder 500 that includes a loss function includes an autoencoder 502 and a loss function 504. It should be noted that the autoencoder 502 and the loss function 504 are merely particular examples of an autoencoder and a loss function, and are not intended to be limiting of the exemplary embodiment. In other words, the exemplary embodiment can utilize different autoencoders and loss functions.
[0067] To transform each row of data in a data asset, such as the input data asset 418 of FIG. 4, a data distribution preservator, such as the data distribution preservator 218 of FIG. 2, selects the columns of interest that include confidential data in the data asset that requires transformation as the input to an autoencoder 502 that includes a loss function 504. The number of layers and the depth of the bottleneck of the autoencoder 502 and the loss function 504 determine how each row of data is transformed. The loss function 504 further determines a measure of dispersion for each of the columns of interest. When executed for each corresponding row, the autoencoder 502 stores the original distribution of the values of the input data asset in the output data asset.
[0068] It should be noted that a set of autoencoders can cover some combinations of types of data - asset sequences. Further, the loss - function parameter coefficients provide variability and can be controlled by a policy. For example, when the value of a particular parameter coefficient that evaluates the variance within a sequence between the input data - asset and the output data - asset is not zero (0), the autoencoder 502 is considering the projection sequence. The data - distribution preservator needs to determine the canonical sequence and generate an autoencoder that can cover all combinations of autoencoders.
[0069] The loss function (LF) 504 has three parameter coefficients arranged as weighted functions. A particular weight on the data - asset variance provides variability in the depth of distribution preservation. The parameter coefficient rho (ρ) of the loss function evaluates the column - specific variance across the data - asset. In other words, rho minimizes the distance between the means of the columns of interest that contain the relevant data - classes across the entire data - asset. The parameter coefficient phi (φ) of the loss function evaluates the in - row / column - specific variance. In other words, phi minimizes the distance between the columns of interest within a particular row of data. The parameter coefficient tau (τ) of the loss function evaluates the data - asset - specific variance. In other words, tau minimizes the orthogonal distance of the output rows to the best - fit hyperplane of the data - asset. The data - distribution preservator generates a hyperplane based on the relevant data - classes of the data - asset. Thus, LF = ρ (variance within the column of interest)+φ (distance between columns within the row of interest)+τ (orthogonal distance from the best - fit hyperplane). It should be noted that the loss function 504 can be a calculated pre - mapping or post - mapping from the actual row values to the pseudo - row values.
[0070] Next, referring to FIG. 6, a flowchart showing a process of data classification and autoencoder-based training according to an exemplary embodiment is shown. The process shown in FIG. 6 can be implemented on a computer such as, for example, server 104 of FIG. 1 or data processing system 200 of FIG. 2. For example, the process shown in FIG. 6 can be implemented on data distribution storage 218 of FIG. 2.
[0071] This process begins when a computer receives a data asset as input (step 602). The computer uses a data profiler to profile the actual data of the data asset (step 604). Based on the profiling of the actual data of the data asset, the computer detects the data classes of interest corresponding to the data asset for each column (step 606).
[0072] The computer generates a best-fit hyperplane of the data asset that separates data values based on the data classes of interest corresponding to the data asset (step 608). The computer maintains the best-fit hyperplane corresponding to the data asset (step 610). The best-fit hyperplane represents the base distribution signature of the data asset and the 0% distribution preservation distance threshold of the data asset. The computer uses the best-fit hyperplane during data distribution to calculate the parameter coefficients of the loss function together with a specified distribution preservation directive. Note that a 100% distribution preservation directive represents the complete preservation of the input data asset to be reflected in the output data asset, and a 0% distribution preservation directive is the "don't care" point when generating data transformation by an autoencoder of the data classes of interest. The computer calculates a loss function threshold by scaling the worst or maximum orthogonal distance observed between all rows of data in the input data asset and the generated best-fit hyperplane. The computer stores this maximum Cartesian distance from the input data asset and the 0% distribution preservation directive when scaling the "don't care" point or policy as a percentage.
[0073] The computer searches a library of autoencoders for a set of autoencoders corresponding to the data classes of interest that are subject to anonymization, based on the historical enforcement and simulated enforcement of policies related to the data assets (step 612). The computer performs a determination in this search as to whether a set of autoencoders corresponding to all of the data classes of interest has been found (step 614). If the computer determines that a set of autoencoders corresponding to all of the data classes of interest has been found in this search, that is, if the output of step 614 is "yes", then the process ends. If the computer determines that a set of autoencoders corresponding to all of the data classes of interest has not been found in this search, that is, if the output of step 614 is "no", the computer generates new randomly initialized autoencoders for all of the data classes among the data classes of interest, based on the historical enforcement and simulated enforcement of policies related to the data assets (step 616). Further, the computer base-trains the newly randomly initialized autoencoders using an inverse mapping from the transformed row cell values obtained by classical deterministic transformation of the input rows to pseudo-row cell values suitable for autoencoder training (step 618). This classical deterministic transformation is a ground truth row cell value generation method that uses raw input row cell values. Thereafter, the computer adds the newly base-trained autoencoders to the library of autoencoders (step 620). Thereafter, the process ends.
[0074] Next, referring to FIGS. 7A-7C, a flowchart showing a process of policy enforcement for converting confidential data according to an exemplary embodiment is shown. The process shown in FIGS. 7A-7C can be implemented on a computer such as the server 104 of FIG. 1 or the data processing system 200 of FIG. 2. For example, the process shown in FIGS. 7A-7C can be implemented in the data distribution storage 218 of FIG. 2.
[0075] This process begins when a computer receives, via a network, from a client device of a particular user, a request asking the computer to access data of a particular input data asset (step 702). This particular input data asset is a rectangular data set. In response to the computer receiving the request, the computer's policy enforcement point requests a policy enforcement decision from the computer's policy decision point regarding the particular input data asset, the particular user, and the context associated with the request (step 704). The computer's policy decision point generates a policy enforcement decision regarding the particular input data asset, the particular user, and the context of the request based on a currently set set of policies (step 706). It should be noted that this policy enforcement decision may further include a percentage of desired data distribution preservation for a data class of interest that requires transformation by a selected autoencoder from a library of autoencoders. This provides the data classes to be transformed as directed by the policy and the data distribution preservation instructions or thresholds. Further, the computer calculates a loss function threshold of a loss function based on this policy enforcement decision (step 708).
[0076] Furthermore, the computer selects an autoencoder from a library of autoencoders (step 710) to perform the necessary data anonymization on the columns of interest that contain sensitive data in a particular input data asset, based on the policy enforcement decision. Thereafter, the computer selects one row from the particular input data asset (step 712). The computer posts the original input data values of the selected row to a temporary buffer (step 714). The computer performs the necessary data anonymization on the selected row using the selected autoencoder to transform the data values of the sensitive data in the associated set of row cells for the columns of interest and place them in a transformation buffer (step 716). The computer further generates a loss function value for the data anonymization of the selected row using a loss function that includes the parameter coefficients specified in the policy enforcement decision (step 718).
[0077] The computer compares the generated loss function value with the calculated loss function threshold (step 720). Based on this comparison, the computer makes a determination as to whether the generated loss function value is less than the calculated loss function threshold (step 722).
[0078] Based on this comparison, if the computer determines that the generated loss function value is greater than or equal to the calculated loss function threshold, i.e., if the output of step 722 is "no", the computer posts the transformed data values in the transformation buffer to an output and labeled output buffer using the forward mapping to the actual row cell values suitable for a particular user (step 724). Thereafter, the process proceeds to step 726. Based on this comparison, if the computer determines that the generated loss function value is less than the calculated loss function threshold, i.e., if the output of step 722 is "yes", the computer transfers the output and labeled output buffer to the next row for the output of that particular input data asset (step 726).
[0079] Furthermore, the computer makes a determination as to whether its temporary buffer is qualified for use in the autoencoder active retraining process based on the calculated loss function threshold and the stored moving average threshold (step 728). If the computer determines that the temporary buffer is not qualified for use in the autoencoder active retraining process, i.e., if the output of step 728 is "no", the computer makes a determination as to whether there is another row in the specific input data asset (step 730). If the computer determines that there is another row in the specific input data asset, i.e., if the output of step 730 is "yes", the process returns to step 712 and the computer selects another row from the specific input data asset. If the computer determines that there is no other row in the selected column, i.e., if the output of step 730 is "no", the computer sends the output for the specific input data asset to the client device of the specific user via the network (step 732). Thereafter, the process ends.
[0080] Returning to step 728, if the computer determines that the temporary buffer is qualified for use in the autoencoder active retraining process, i.e., if the output of step 728 is "yes", the computer makes a determination as to whether the generated loss function value is less than the calculated loss function threshold (step 734). If the computer determines that the generated loss function value is greater than or equal to the calculated loss function threshold, i.e., if the output of step 734 is "no", the computer labels the temporary buffer of the original input data values as a good candidate (step 736). Thereafter, the process proceeds to step 740. If the computer determines that the generated loss function value is less than the calculated loss function threshold, i.e., if the output of step 734 is "yes", the computer labels the temporary buffer of the original input data values as a reject candidate (step 738).
[0081] Thereafter, the computer saves the labeled output lines to a training buffer (step 740). Further, the computer passes the training buffer through a data sampler to form a sampled training buffer (step 742). In an exemplary embodiment, the data sampler is a random data sampler. The computer maintains the sampled training buffer (step 744).
[0082] Further, the computer retrains a particular autoencoder asynchronously using an inverse mapping from the transformed row cell values obtained by classical deterministic transformation of the input lines of the training buffer to pseudo row cell values suitable for autoencoder training (step 746). The computer additionally retrains a particular autoencoder in a library of autoencoders using the sampled training buffer based on a regularization or generalization method for enhancing good candidates and avoiding overfitting for rejection candidates (step 748). It should be noted that over time, the retraining of the autoencoder uses backpropagation and converges to a data distribution preservation threshold most frequently used for a particular combination of data class and autoencoder. Further, in an exemplary alternative embodiment, the threshold can index the library of autoencoders so that it can also play a role in the selection process when the autoencoders are explored in the library. The computer further saves the moving-averaged calculated loss function threshold along with the value used for a particular input data asset (step 750). Thereafter, the process returns to step 730 where the computer determines whether there is another row in that particular input data asset.
[0083] Accordingly, an exemplary embodiment of the present invention provides a computer-implemented method, a computer system, and a computer program product for dynamically and in real-time transforming confidential data of a data asset during distribution of the data to a requesting data consumer. A particular transformation on the requested data during the anonymization operation is achieved while preserving the transformed values of the output data distribution to a desired degree with respect to the input data. The exemplary embodiment declaratively controls this process based on a policy using an autoencoder that includes a parameterized loss function. The descriptions of the various embodiments of the present invention are presented for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, practical applications, or technical improvements not found in commercially available technologies, or to enable those of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
**Claim 1** A method for preserving the distribution of data values of a data asset in a data anonymization operation through computer information processing, the method comprising: executing data anonymization of selected rows in the data asset using an autoencoder to convert the data values of the sensitive data in a set of row cells associated with a column of interest and place them in a conversion buffer; generating a loss function value for the data anonymization of the selected rows using a loss function that includes a parameter coefficient specified in a policy enforcement decision; comparing the loss function value with a loss function threshold; in response to determining that the loss function value is greater than the loss function threshold based on the comparing, transferring the converted data values in the conversion buffer to an output buffer labeled as output using a forward mapping to actual row cell values suitable for a particular user; and transferring the output buffer labeled as output to the next row for output of the data asset A method comprising. **Claim 2** further comprising sending the output for the data asset to a client device of a requesting user via a network. The method according to claim 1, further comprising. **Claim 3** receiving, via a network, from a client device of the particular user a request for accessing data of the data asset, which is a rectangular data set consisting of columns and rows; requesting a policy enforcement decision regarding the data asset, the particular user, and the context of the request for accessing the data of the data asset; generating a policy enforcement decision regarding the data asset, the particular user, and the context of the request for accessing the data of the data asset; and calculating the loss function threshold of the loss function based on the policy enforcement decision The method according to claim 1 or 2, further comprising. **Claim 4** selecting the autoencoder from a library of autoencoders to perform the data anonymization on the column of interest that includes the sensitive data in the data asset based on the policy enforcement decision; and selecting one row out of a plurality of rows in the data asset to form the selected row The method according to any one of claims 1 to 3, further comprising.
5. determining whether a temporary buffer of the original input data values transcribed for the selected row is qualified for use in an autoencoder active retraining process; determining whether the loss function value was less than the loss function threshold in response to determining that the temporary buffer is qualified for use in the autoencoder active retraining process; labeling the temporary buffer with a good candidate label in response to determining that the loss function value was greater than or equal to the loss function threshold, or labeling the temporary buffer with a rejected candidate label in response to determining that the loss function value was less than the loss function threshold, and saving the labeled output row to a training buffer The method according to any one of claims 1 to 4, further comprising.
6. passing the training buffer through a data sampler to form a sampled training buffer; asynchronously retraining a particular autoencoder using an inverse mapping from the transformed row cell values obtained by classical deterministic transformation of the input rows of the training buffer to pseudo row cell values suitable for autoencoder training, and further retraining the particular autoencoder using the sampled training buffer based on regularization for strengthening good candidates and avoiding overfitting for rejected candidates The method according to claim 5, further comprising.
7. profiling the real data of the data asset; detecting, for each column, a data class of interest corresponding to the data asset based on the profiling of the real data of the data asset, and generating a best-fit hyperplane of the data asset that separates data values based on the data class of interest corresponding to the data asset The method according to any one of claims 1 to 6, further comprising.
8. In response to determining that no autoencoder corresponding to the data class of interest was found in the search of the autoencoder library, generating a new randomly initialized autoencoder for the data class of interest based on the historical implementation of the policy associated with the data asset and the simulated implementation, and base training the newly randomly initialized autoencoder using an inverse mapping from the transformed row cell values obtained by the classical deterministic transformation of the input row to pseudo row cell values suitable for autoencoder training The method according to claim 7, further comprising. **Claim 9** The method according to any one of claims 1 to 8, wherein a Laplace noise function is applied to the anonymized data values as differential privacy adjustment to prevent distribution inference. **Claim 10** The method according to any one of claims 1 to 9, wherein the policy implementation decision is based on the entropy of the data asset. **Claim 11** The method according to any one of claims 1 to 10, wherein the policy implementation decision is based on the sensitivity of the data asset with respect to the transformation change. **Claim 12** A computer system for preserving the distribution of data values of a data asset in a data anonymization operation, the computer system comprising a bus system, a storage device connected to the bus system and storing program instructions, the storage device, a processor connected to the bus system and the processor executes the program instructions to perform data anonymization of selected rows in the data asset using an autoencoder to transform the data values of the confidential data in a set of row cells associated with the column of interest and place them in a transformation buffer, generate a loss function value for the data anonymization of the selected rows using a loss function including parameter coefficients specified in the policy implementation decision, compare the loss function value with a loss function threshold, in response to determining that the loss function value is greater than the loss function threshold, transfer the transformed data values in the transformation buffer to an output buffer labeled as output using a forward mapping to actual row cell values suitable for a particular user For output of the data asset, transfer the output buffer labeled as output to the next line, A computer system. **Claim 13** The processor further executes the program instructions to Send the output for the data asset to the client device of the user making the request via a network. The computer system according to claim 12. **Claim 14** The processor further executes the program instructions to Receive, via a network, from the client device of the specific user, a request for accessing data of the data asset which is a rectangular data set consisting of columns and rows, Request a policy enforcement decision regarding the context of the data asset, the specific user, and the request for accessing the data of the data asset, Generate a policy enforcement decision regarding the context of the data asset, the specific user, and the request for accessing the data of the data asset, Calculate the loss function threshold of the loss function based on the policy enforcement decision. The computer system according to claim 12 or 13. **Claim 15** A computer program for preserving the distribution of data values of a data asset in a data anonymization operation, causing a computer system to Perform data anonymization of selected rows in the data asset using an autoencoder to convert data values of confidential data in a set of related row cells of columns of interest and place them in a conversion buffer, Generate a loss function value for the data anonymization of the selected rows using a loss function including parameter coefficients specified in a policy enforcement decision, Compare the loss function value with a loss function threshold, In response to determining that the loss function value is greater than the loss function threshold based on the comparison, transfer the converted data values in the conversion buffer to an output buffer labeled as output using a forward mapping to actual row cell values suitable for a specific user, For output of the data asset, transfer the output buffer labeled as output to the next line A computer program for causing the above to be done.
Citation Information
Patent Citations
System and method for detecting network intrusion
JP2007179542A
Privacy protection data providing system and privacy protection data providing method
JP2018097467A
Systems and methods for generating data interpretations in neural networks and related systems
JP2019520655A
Computer-implemented method, system, computer program, and storage medium for data anonymization
JP2021504798A
Model learning device, model learning method, and program
WO2019138655A1