Anonymize a network using network attributes and entity-based access rights
By utilizing the factors of network attributes, node attributes and edge attributes in computer networks, an anonymization system and method was developed, which solved the problem of difficulty in effectively anonymizing network information in the existing technology, and realized the privacy protection and utility optimization of network information.
Patent Information
- Application Number
- CN202080068518.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2020-09-17
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-09-17
AI Technical Summary
It is difficult for the prior art to effectively anonymize computer networks, especially when protecting the privacy of network information and meeting the utility constraints of consumer needs.
Develop a system and method to anonymize network information based on factors such as network attributes, node attributes and edge attributes. The system includes anonymization components, constraint components and optimal components, which can generate privacy and utility constraints to optimize the anonymization process of network information.
It realizes effective anonymization of network information, protects the privacy of network information, and meets consumers' needs for the utility of anonymized network, and improves the overall utility of anonymized network.
Smart Images

Figure CN114514522B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to computer networks and, more particularly, to anonymizing computer networks. Summary of the Invention
[0002] The following presents an overview to provide a basic understanding of one or more embodiments of the present invention. This overview is not intended to identify key or critical elements or to delineate any scope of particular embodiments or any scope of the claims. Its sole purpose is to present concepts in a simplified form as a prelude to the more detailed description that is presented later. In one or more embodiments described herein, devices, systems, methods, and computer-implemented methods are described that are capable of facilitating the anonymization of a network based on factors including network attributes, node attributes, and edge attributes of connections between nodes.
[0003] According to an embodiment, a system may include: a memory capable of storing computer-executable components; and a processor capable of executing the computer-executable components stored in the memory, wherein the computer-executable components include: an anonymization component capable of anonymizing network information of a network based on network attributes of the network and node attributes of a first node of the network, thereby generating an anonymized network. According to a variant of the above embodiment, connection information corresponding to a connection between the first node and a second node may include edge attributes, and anonymizing the network information may further be based on the edge attributes. According to another variant of the above embodiment, anonymizing the network information may further be based on a privacy constraint that enforces a privacy level of the anonymized network.
[0004] In an additional embodiment of the above system, the computer-executable components of the system may further include: a constraint component capable of generating the privacy constraint based on the network information. In a variant of the additional embodiment, the constraint component may further generate the privacy constraint based on privacy rules, and wherein the privacy rules apply to the first node. In other embodiments, the constraint component may further determine an access right of the first node, and the constraint component may further generate the privacy constraint based on the access right.
[0005] In yet another variant of the above system, anonymizing the network information can be further based on utility constraints for the anonymized network, and wherein the utility constraints include a first measure of the utility of the anonymized network based on consumer requirements. In another embodiment adding additional features to this variant, the computer-executable components of the system can further include: an optimality component that can select utility constraints that optimize or increase the optimality of anonymizing the network information, this optimization being based on anonymization characteristics of anonymizing the network information. In some embodiments, the anonymization characteristics can include a second measure of information loss during anonymizing the network information. In other embodiments, the anonymization characteristics can include the number of edits of the network performed during anonymizing the network information.
[0006] In an embodiment of the above system, the network information can include medical information of multiple patients. Alternatively, in one or more embodiments, the network can include a social media network of multiple users.
[0007] One or more embodiments of a computer-implemented method can include: generating network attributes of a network by a device operatively coupled to a processor. Additionally, the method can include: anonymizing network information of the network by the device based on the network attributes and node attributes of a first node of the network, thereby producing an anonymized network.
[0008] In a variant of the above method, connection information corresponding to a connection between the first node and a second node can include edge attributes, and anonymizing the network information can be further based on the edge attributes. In an additional variant of the method, anonymizing the network information can be further based on privacy constraints that enforce a privacy level of the anonymized network.
[0009] In a variant of the above computer-implemented method, the method can further include: generating the privacy constraints based on privacy rules, and the privacy rules can be applied to the first node. In additional embodiments, anonymizing the network information can be further based on utility constraints for the anonymized network, and the utility constraints can include a measure of the utility of the anonymized network based on consumer requirements. In an embodiment of the method with additional operations, the method can further include: selecting utility constraints that optimize or increase the optimality of anonymizing the network information based on anonymization characteristics of anonymizing the network information.
[0010] In another set of embodiments, a computer program product enables anonymization of network information, wherein the computer program product includes a computer-readable storage medium storing program instructions executable by a processor to cause the processor to: generate, by the processor, network attributes of a network of connected nodes, wherein a plurality of the connected nodes among the plurality of connected nodes include node attributes, thereby producing a plurality of node attributes. Further, generating the network attributes can be based on the plurality of node attributes. The processor can further be caused to: generate, by the processor, privacy constraints based on the network information of the connected nodes, the privacy constraints capable of enforcing a privacy level of the anonymized network. In this embodiment, the processor can further be caused to: anonymize, by the processor, the network information of the connected nodes based on the network attributes, the plurality of node attributes, and the privacy constraints, thereby producing an anonymized network.
[0011] In a variation of the above embodiment, anonymizing the network information can further be based on utility constraints for the anonymized network, the utility constraints capable of including a first measure of the utility of the anonymized network based on consumer requirements. In another embodiment based on the above computer program product, the instructions further cause the processor to: select, by the processor, utility constraints that optimize anonymizing the network information or increase the optimality of anonymizing the network information based on a second measure of information loss during anonymizing the network information.
[0012] According to one aspect, a system is provided that includes: a memory storing computer-executable components; and a processor executing the computer-executable components stored in the memory, wherein the computer-executable components include: an anonymization component that anonymizes network information of a network based on network attributes of the network and node attributes of a first node of the network, thereby producing an anonymized network.
[0013] According to another aspect, a computer-implemented method is provided that includes: generating, by a device operatively coupled to a processor, network attributes of a network; and anonymizing, by the device, network information of the network based on the network attributes and node attributes of a first node of the network, thereby producing an anonymized network.
[0014] According to another aspect, there is provided a computer program product for facilitating the anonymization of network information. The computer program product includes a computer-readable storage medium storing program instructions that are executable by a processor to cause the processor to: generate, by the processor, network attributes of a network of connected nodes, wherein a plurality of the connected nodes among the plurality of connected nodes include node attributes, thereby generating a plurality of node attributes, and wherein generating the network attributes is based on the plurality of node attributes; generate, by the processor, privacy constraints based on network information of the connected nodes, wherein the privacy constraints enforce a privacy level of the anonymized network; and anonymize the network information of the connected nodes by the processor based on the network attributes, the plurality of node attributes, and the privacy constraints, thereby generating an anonymized network. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Embodiments of the present invention will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0016] Figure 1 A block diagram of an example of a non-limiting network anonymization system according to one or more embodiments described herein is shown, which is capable of facilitating the anonymization of a network based on factors including network attributes, node attributes, and edge attributes describing connections between nodes;
[0017] Figure 2 A block diagram of an example of a non-limiting system according to one or more embodiments described herein is shown, which is capable of facilitating the anonymization of a network based on factors including network attributes, node data, and edge attributes describing connections between nodes;
[0018] Figure 3 A block diagram of an example of a non-limiting network anonymization system according to one or more embodiments described herein is shown, which is capable of facilitating the anonymization of a network based on factors including network attributes, node attributes, and edge attributes describing connections between nodes;
[0019] Figure 4 A block diagram of an example of a non-limiting system according to one or more embodiments that is capable of facilitating the publication of an anonymized network by an anonymization component for use by a consumer is shown;
[0020] Figure 5 A block diagram of an example of a non-limiting network anonymization system according to one or more embodiments described herein is shown, which is capable of facilitating the anonymization of a network based on factors including network attributes and using privacy data from network nodes;
[0021] Figure 6A block diagram illustrating an example non-limiting system 600 capable of facilitating anonymizing a network based on factors including network attributes, node attributes, and edge attributes describing connections between nodes in accordance with one or more embodiments described herein;
[0022] Figure 7 depicts a table of sample data from a network to be anonymized using different methods according to one or more embodiments;
[0023] Figures 8A - 8B Describes the use of Figure 7 Two alternative approaches to data anonymization depicted in;
[0024] Figure 9 A flow chart illustrating an example non-limiting computer-implemented method capable of using network attributes and entity-based access permissions to facilitate anonymizing a network in accordance with one or more embodiments described herein;
[0025] Figure 10 A block diagram is shown of an example non-limiting operating environment in which one or more embodiments described herein can be facilitated. DETAILED DESCRIPTION
[0026] The following detailed description is illustrative only and is not intended to limit the embodiments and / or the application or uses of the embodiments. In addition, it is not intended to be bound by any express or implied information presented in the previous background or summary or detailed description.
[0027] One or more embodiments are now described with reference to the accompanying drawings, wherein the same reference numerals are used throughout to represent the same elements. In the following description, for the purpose of explanation, many specific details are set forth in order to provide a more thorough understanding of one or more embodiments. However, in various cases, it is apparent that one or more embodiments may be practiced without these specific details. Note that the drawings of the present application are for illustrative purposes only, and therefore the drawings are not drawn to scale.
[0028] It should be understood that the embodiments of the present disclosure depicted in the various figures disclosed herein are for illustration only, and therefore, the architecture of these embodiments is not limited to the systems, devices and / or components depicted herein.
[0029] One or more embodiments are described herein in the context of sharing a certain amount of data from an otherwise restricted access network. Some embodiments described herein apply to a restricted network with some data sought for public and private reasons. For illustration purposes, an example of this type of network frequently used herein is a network that stores and provides medical data (e.g., information about hospitals, patients, treatments, and other relevant information).
[0030] Because healthcare networks can provide data that is both legally and ethically restricted and sought for public and private purposes, these networks can provide an example context for the application of the principles described herein. For example, a network where anonymization can be used to share patient data with researchers while complying with all legal and ethical requirements. As discussed further below, in many cases, one or more embodiments can evaluate privacy and utility constraints when selecting an anonymization method and the anonymization parameters to be used (such as the k value in the k-anonymity method described below).
[0031] It should be noted that although healthcare networks are sometimes used herein to illustrate the concepts of embodiments, these are non-limiting examples, and other types of networks, both now in use and those developed in the future, can be used with one or more of the embodiments described herein. For example, association member networks, human resources networks, transaction data networks, and social networks can also contain sought-after data that is in many cases protected from disclosure.
[0032] Generally speaking, as described herein, one or more embodiments are capable of facilitating network anonymization, which can incorporate different types of attributes for anonymization, such as node attributes, edge attributes, and network attributes. In addition, one or more embodiments are capable of facilitating the anonymization of network data such that the utility of the anonymized network is improved while meeting the privacy requirements of the original individual nodes, for example, as a trade-off between privacy constraints and utility constraints.
[0033] Figure 1 A block diagram of an example 100 of a non-limiting network anonymization system 102 in accordance with one or more embodiments described herein is shown. The system can facilitate anonymizing a network based on factors including network attributes, node attributes, and edge attributes that describe the connections between nodes. For the sake of brevity, repeated descriptions of the same elements and / or processes employed in the various embodiments are omitted. Some embodiments can include a network anonymization system 102. In some embodiments, the network anonymization system 102 can include an attribute component 108, an anonymization component 107, and any other components associated with the network anonymization system 102 disclosed herein.
[0034] It should be understood that the embodiments of the present disclosure depicted in the various figures herein are for illustrative purposes only, and thus, the architectures of these embodiments are not limited to the systems, devices, and / or components depicted herein. For example, in some embodiments, the network anonymization system 102 can also include those described herein with reference to the operating environment 1000 and Figure 10The various computers and / or computing-based elements described, e.g., in some embodiments, the network anonymization system 102 may further include a memory 104, a processor 106, and a bus 112. In multiple embodiments, such computers and / or computing-based elements may be combined to implement one or more of the systems, devices, components, and / or computer-implemented operations shown and described in connection with Figure 1 or other figures disclosed herein.
[0035] According to multiple embodiments, the memory 104 may store one or more computer and / or machine-readable, writable, and / or executable components and / or instructions that, when executed by the processor 106, may facilitate the execution of operations defined by the (one or more) executable components and (one or more) instructions. For example, the memory 104 may store computer and other machine-readable, writable, and / or executable components and / or instructions that, when executed by the processor 106, may facilitate the execution of various functions related to the network anonymization system 102, the attribute component 108, the anonymization component 107, and any other components associated with the network anonymization system 102 as described herein (with or without reference to the various figures of the present disclosure).
[0036] In some embodiments, the memory 104 may include volatile memory (e.g., random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), etc.) and / or non-volatile memory (e.g., read-only memory (ROM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), etc.) that may employ one or more memory architectures. Further examples of the memory 104 are described below with reference to the system memory 1016 and Figure 10 described further examples of the memory 104 that may be used to implement any embodiment of the present disclosure.
[0037] According to multiple embodiments, the processor 106 may include one or more types of processors and / or electronic circuits, which may implement one or more computer and / or machine-readable, writable, and / or executable components and / or instructions storable on the memory 104. For example, the processor 106 may perform various operations that may be specified by such computer and / or machine-readable, writable, and / or executable components and / or instructions, including but not limited to logical, control, input / output (I / O), arithmetic, and / or similar operations. In some embodiments, the processor 106 may include one or a combination of different central processing units, multi-core processors, microprocessors, dual microprocessors, microcontrollers, system-on-a-chip (SOC), array processors, vector processors, and any other type of processor. Further examples of the processor 106 are described below with reference to the processing unit 1014 and Figure 10 are described. These examples of the processor 106 may be used to implement any embodiment of the present disclosure.
[0038] In some embodiments, the elements of the network anonymization system 102 (including but not limited to the memory 104, the processor 106, the attribute component 108, the anonymization component 107, and any other components of the network anonymization system 102 as described herein) may be communicatively, electrically, and / or operatively coupled to each other via the bus 112 to perform the functions of the network anonymization system 102 and any other components coupled thereto. In multiple embodiments, the bus 112 may include one or more of a memory bus, a memory controller, a peripheral bus, an external bus, a local bus, or another type of bus that may employ various bus architectures. Further examples of the bus 112 are described below with reference to the system bus 1018 and Figure 10 are described, and these examples of the bus 112 may be used to implement any embodiment of the present disclosure.
[0039] In some embodiments, the network anonymization system 102 may include any type of component, machine, device, facility, apparatus, and / or instrument that includes a processor and / or is capable of communicating effectively and / or operably with a wired and / or wireless network. All such embodiments are foreseeable. For example, the network anonymization system 102 may include server devices, computing devices, general-purpose computers, special-purpose computers, quantum computing devices (e.g., quantum computers, quantum processors, etc.), tablet computing devices, handheld devices, server-class computing machines, and / or databases, laptop computers, notebook computers, desktop computers, cellular phones, smart phones, consumer appliances and / or instruments, industrial and / or commercial devices, digital assistants, multimedia Internet-enabled telephones, multimedia players, and / or another type of device.
[0040] In some embodiments, the network anonymization system 102 may be coupled (e.g., communicatively, electrically, operatively, etc.) to one or more external systems, sources, and / or devices (e.g., computing devices, communication devices, etc.) via a data cable (e.g., coaxial cable, High-Definition Multimedia Interface (HDMI), Recommended Standard (RS) 232, Ethernet cable, etc.). In some embodiments, the network anonymization system 102 may be coupled (e.g., communicatively, electrically, operatively, etc.) to one or more external systems, sources, and / or devices (e.g., computing devices, communication devices, etc.) via a network.
[0041] According to multiple embodiments, the network anonymization system 102 may include one or more computer and / or machine-readable, writable, and / or executable components and / or instructions that, when executed by the processor 106, may facilitate the execution of operations defined by such components and / or instructions. Additionally, in many embodiments, as described herein with or without reference to the various figures of the present disclosure, any component associated with the network anonymization system 102 may include one or more computer and / or machine-readable, writable, and / or executable components and / or instructions that, when executed by the processor 106, may facilitate the execution of operations defined by such components and / or instructions. For example, the attribute component 108 and the anonymization component 107, as well as any other component associated with the network anonymization system 102 disclosed herein (e.g., communicatively, electronically, and / or operatively coupled to and / or employed by the network anonymization system 102), may include such computer and / or machine-readable, writable, and / or executable components and / or instructions. Thus, according to many embodiments, the network anonymization system 102 and / or any component associated therewith as disclosed herein may employ the processor 106 to execute such computer and / or machine-readable, writable, and / or executable components and / or instructions to facilitate the execution of one or more operations described herein with reference to the network anonymization system 102 and / or any such component associated therewith.
[0042] Returning to Figure 1 the operation of the components depicted in Figure 2 As further discussed, different attributes of the nodes of the network may be used to facilitate the anonymization process and improve the utility of the anonymized network 280. For a medical network, different attributes of the network nodes may include, but are not limited to, patient data such as identifiers, names, diagnoses, and other similar data. As referenced Figures 3 - 4For further discussion, other node attributes that can be used include the characteristics of the entities stored by a particular node. For example, a node can have security requirements based on the entity type. For example, in a medical system, patient data can have more restricted access requirements than human resources data and general data about different medical facilities.
[0043] In addition to using existing node attributes for anonymization, one or more embodiments can analyze network data and identify attributes based on the network data, such as network attributes. For example, in one or more embodiments, the attribute component 108 can generate network attributes of a network based on an analysis of the connected nodes. In one or more embodiments, network attributes generally refer to attributes that describe the entire network or the relationships between the nodes of the network and other nodes. Example network attributes that can be used by one or more embodiments described herein include, but are not limited to, the total number of nodes, the total number of connections, the average degree of the nodes within the network, and different measures of the network centrality of the nodes. Example network centrality attributes can include measures of the in-degree and out-degree of a node, such as the count of data flows entering and leaving the node, respectively. Other examples of network centrality include betweenness centrality, for example, a measure of the centrality of a node in a graph based on the shortest paths between nodes.
[0044] Another example network attribute that can describe the network centrality of a node includes the closeness centrality of the network node, for example, how many nodes are between the node and other nodes of the network. In one or more embodiments, this value can be generated, for example, by the attribute component 108 by taking the average of the shortest path lengths from the node to each other node in the network. It should be understood that in alternative embodiments, instead of generating network attributes, the attribute component 108 can obtain previously generated network attributes for use in anonymization.
[0045] As further discussed below in connection with Figure 2 the example medical network depicted therein, by incorporating network attributes in the anonymization process, in one or more embodiments, these network attributes can be made available to consumers querying the anonymized network. As further described below, although node and edge attributes can be used for network analysis, by incorporating overall network attributes in the anonymization process, one or more embodiments are able to facilitate the generation of additional types of network analysis data from the anonymized network.
[0046] For example, in one or more embodiments, a measure of a node's network centrality can be incorporated into the anonymization process, and this can result in the centrality value being made available to consumers of the anonymized network 290 along with anonymized node attributes and edge attributes. Given the description herein, those skilled in the relevant art will understand that, in addition to being used by consumers independently of other anonymized network data, for some applications, different network attributes can be advantageously combined with anonymized node data. For example, after identifying anonymized nodes using medical data indicating exposure to an infectious disease condition, an epidemiological researcher can further query the anonymized network for different measures of the network centrality of the identified nodes, e.g., to analyze the role of the node in spreading the condition to other nodes within the anonymized network.
[0047] Regarding other aspects of one or more embodiments discussed below, Figure 2 Depicts a process of transforming a network into an anonymized network using anonymization component 107 and network attributes that can be generated by attribute component 108 as described above, by employing processes similar to those discussed above. Figures 3 - 4 Describes a constraint component 310 that can be used by one or more embodiments to generate constraints for use by anonymization component 107. Figures 5 - 6 Describes an optimality component 510 that can be used by one or more embodiments to anonymize a network based on the utility provided to consumer 515 and the constraints generated by constraint component 310. Further, for the operation of an example implementation of optimality component 510, Figures 7 - 8B Depicts an example manner in which the utility of an anonymized network can be adjusted based on the requirements of a consumer, given the required constraints on the anonymized data.
[0048] Figure 2 Illustrates a block diagram of an example non - limiting system 200 according to one or more embodiments described herein, which can facilitate anonymizing network 280 based on factors including data of network attributes 277, nodes 220A - 220D, and connections 225A - B between the nodes. For brevity, repeated descriptions of the same elements and / or processes employed in the various embodiments are omitted.
[0049] In some embodiments, system 200 can include a network 280 that is transformed into an anonymized network 290 through an anonymization 275 process. In one or more embodiments, network 280 can include a plurality of nodes, such as nodes 220A - D, that are differently connected by connections including connections 225A - B. In one or more embodiments, anonymized network 290 can include nodes 222A - D anonymized from nodes 220A - D, respectively.
[0050] In one example, network 280 is a medical network that includes hospitals, where each node 220A-D in system 200 represents a hospital, and the connection 225A-B between two nodes represents a collaborative link between hospital nodes. In this example, the node attributes of hospital nodes 220A-D can include the location of the hospital, the specialties of the hospital, the number of doctors allowed into the hospital per specialty, the number of patients in the hospital, and metrics corresponding to the performance of the hospital, such as the number of readmissions or the average cost per treatment. In one or more embodiments, the level of detail at the node attribute level can be referred to as detail at the micro level. After a more detailed discussion of the anonymization 275 process, the attribute levels of network 280 are discussed for the medium level (e.g., edge attributes that can provide data on hospital clusters within the network) and the macro level (e.g., network attributes 277 that can provide summary data of the network, and data that describes the placement of individual nodes within the network and their relationships with other nodes).
[0051] Continuing the general discussion of one or more embodiments, as described above, once generated by the attribute component 108 or otherwise identified by one or more embodiments, different attributes can be combined in the anonymization process. For example, in one or more embodiments, the anonymization component 107 can facilitate anonymizing the network information of the network based on the above attributes (e.g., attribute types including but not limited to node attributes, edge attributes, and network attributes 277).
[0052] In another example, network 280 can be a social network that interconnects users, where each node 220A-D in system 200 represents a user, and the connection 225A-B between two nodes represents a connection between user nodes. In this example, the node attributes of user nodes 220A-D can include the location of the user, connection information, posts, age, and other user information. Those skilled in the relevant art will understand, given the description herein, how the network attributes 277 described for the medical network example can be applied to other networks, such as this social network example.
[0053] Examples of why and when anonymization may be employed for one or more embodiments will be understood by those skilled in the art. Given the disclosure herein, data that is restricted from being provided can generally be transformed in some way so as not to be subject to the same restrictions. One type of transformation that can be employed by one or more embodiments is the process of anonymizing data to the extent that the data meets an anonymization level, e.g., the released data cannot be used to individually identify a particular node that is stored. For example, if a set of patient data is to be released to a researcher, the released data should not be able to individually identify a particular patient. For example, attributes such as patient identifiers, social security numbers, phone numbers, email addresses, and other similar data can be used in many cases to identify an individual. In one or more embodiments, this type of data can be removed from the anonymization network or obfuscated.
[0054] Another level of anonymization that can be used to protect a particular type of data prohibits the release of information that, when combined with other reasonably available data, can identify an entity within a particular probability. For example, if the anonymized patient data includes two or more different types of data (e.g., birthday, zip code, gender), then in some cases, these values can be combined with publicly available birthday, zip code, and gender data to identify an individual patient. Figures 7 - 8B Examples of how combinations of birthday, zip code, and gender can be used to identify an individual are provided. One or more embodiments can prevent the direct and associative identification of node data in anonymized data when used with different anonymization methods.
[0055] The underlying anonymization technology may not be important for one or more embodiments, and thus one or more embodiments can potentially be applied to different existing methods, such as k-anonymity, l-diversity, etc. Those skilled in the art will understand how these methods are used given the description herein, and methods developed in the future will also be able to be used with one or more embodiments.
[0056] An example anonymization method that can be used by one or more embodiments is the k-anonymity method. In some embodiments, generally speaking, the k-anonymity method can use different transformations to anonymize data such that for the anonymized data, at least k entities are similar enough that they cannot be distinguished from one another. For example, for a value of k = 3, there are at least three entities in the anonymized network where an attribute is removed or transformed such that at least three records cannot be distinguished from one another. Thus, for a value of k = 3, the data can be anonymized such that, for example, for a given zip code, if a complete birthday is included in the anonymized data, then no fewer than three records have the same combination. For example, in large zip codes (e.g., with a population greater than 50,000 people), there may always be multiple individuals with the same birthday, but this is less likely for typical zip code sizes, and not every individual in a large zip code can be included in the data to be anonymized.
[0057] Multiple examples of the k-anonymization process are described in the following reference Figures 7 - 8B where Figure 7 the dataset shown is transformed by an example k-anonymity method. For example, Figure 8A depicts an example use of the k-anonymity method where k = 3, Figure 8B and depicts an example for k = 4. Other example values of k that can be used by one or more embodiments include k values from 3 to 15, or any other value determined to be beneficial for the process.
[0058] As described above, one or more embodiments can incorporate the network attribute 277 of node 220A (e.g., a measure of node centrality) in the anonymized data of a node (e.g., node 222A of the anonymized network). It should be noted that based on the above anonymization method, one or more embodiments can anonymize the network such that incorporating the network attribute 277 in the anonymized network data maintains the level of anonymity settings required by the anonymization process, e.g., as described for K-anonymity above.
[0059] It should be noted that, as further described below with reference to Figure 5 during the anonymization process, one or more embodiments can automatically adjust the method used by the anonymization method in order to optimize or increase the optimality of the anonymization process, e.g., the utility of the anonymized network, given the type of queries that the anonymized network is designed to support.
[0060] Return Figure 2Regarding the discussion of the medical network 280, in the example, an anonymized version of the network 280 that can provide network analysis in a privacy-preserving manner is sought. According to the example requirements, the anonymized data should support medium-level queries. For example, queries that can identify hospital communities that collaborate more frequently and the execution of such collaborations. In one or more embodiments, the example edge attributes of the connection 225A-B can support these requirements by generating and including edge attributes for hospital pairs linked by the corresponding connection, such as the number of shared patients, the average performance when sharing patients, and performance metrics for treating shared patients.
[0061] In this example, additional requirements for the anonymized network 290 can include data that supports macro-level queries, such as the network attributes discussed above with reference to Figure 1 In one or more embodiments, by incorporating the overall network attributes 277 into the anonymization 275 process of the network 280, additional data can be made available in the anonymized network 290, which can provide a more comprehensive view of the network and, in some cases, improve the overall utility of the anonymized network. For example, insights that can be revealed using the network attributes from the anonymized network include, but are not limited to, the characteristics of one or more nodes in the network that affect the characteristics of other nodes.
[0062] Additional queries can relate to the degree sequence of hospitals, such as the distribution of the collaboration load among hospitals. For all these additional requirements at the macro level, network attributes 277 can be generated and included in the anonymization 275 process. For example, leveraging network attributes 277 such as betweenness centrality and clustering coefficient during anonymization can provide data that can highlight hospitals with high influence in their network. For example, a high betweenness centrality of a hospital can occur because links between other hospitals pass through the high-influence hospital.
[0063] Another example network attribute 277 that can support macro-level data analysis measures the in-degree and out-degree centrality of hospitals. For example, in some cases, a high out-degree centrality of a hospital can indicate a hospital that tends to send out many of its patients, while a high in-degree centrality can indicate a hospital that receives transferred patients.
[0064] It should be noted that the nodes 220A-D are depicted as black or white, and these colors indicate that the network 280 includes nodes 220A-D that can have different privacy levels. As described below in conjunction with Figures 3 - 4 One or more embodiments can provide entity-based access rights to the anonymized network 290
[0065] In some embodiments, the network anonymization system 102 can be associated with various technologies. For example, the network anonymization system 102 can be associated with medical technology, database technology, computing technology, artificial intelligence (AI) model technology, machine learning (ML) model technology, cloud computing technology, Internet of Things (IoT) technology, and / or other technologies.
[0066] According to multiple embodiments, networks such as network 280 and anonymization network 290 can include wired and wireless networks, including but not limited to local area networks (LANs), cellular networks, wide area networks (WANs) (e.g., the Internet), or storage area networks (SANs). For example, the network anonymization system 102 can communicate (and vice versa) with one or more external systems, sources, and / or devices (e.g., computing devices) using virtually any desired wired or wireless technology, including but not limited to: Wi-Fi (Wireless Fidelity), Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), Worldwide Interoperability for Microwave Access (WiMAX), Enhanced General Packet Radio Service (Enhanced GPRS), 3rd Generation Partnership Project (3GPP) Long Term Evolution (LTE), 3rd Generation Partnership Project 2 (3GPP2), Ultra Mobile Broadband (UMB), High Speed Packet Access (HSPA), Zigbee, and other 802.XX wireless technologies and / or traditional telecommunications technologies, Session Initiation Protocol (SIP), RF4CE protocol, WirelessHART protocol, 6LoWPAN (IPv6 over Low-Power Wireless Personal Area Network), Z-Wave, ANT, Ultra Wideband (UWB) standard protocol, and / or other proprietary and non-proprietary communication protocols. In such an example, the network anonymization system 102 can thus include hardware (e.g., a central processing unit (CPU), transceiver, decoder), software (e.g., a set of threads, a set of processes, software in execution), or a combination of hardware and software that facilitates the transfer of information between the network anonymization system 102 and external systems, sources, and / or devices (e.g., computing devices, communication devices, etc.).
[0067] Discussed together below Figure 3 and 4 to highlight the use of different components during the anonymization process. Figure 3 A block diagram of an example 300 of a non-limiting network anonymization system according to one or more embodiments described herein is shown, which can facilitate anonymizing a network based on factors including network attributes, node attributes, and edge attributes describing connections between nodes. In some embodiments, the network anonymization system 102 can include an attribute component 108, an anonymization component 107, a constraint component 310, and any other components associated with the network anonymization system 102 disclosed herein.
[0068] Figure 4 A block diagram showing an example 400 non - limiting system that can facilitate the release 406 of an anonymization network 280 by an anonymization component 107 for use by a consumer 450. The utility of the anonymization network 280 for the consumer 450 can be quantified by a utility metric 455. A constraint component 310 can generate privacy constraints for the anonymization component 107 based on privacy rules 428, which can be for the entire network 280 or specific to individual nodes of the network 280. For the sake of brevity, repeated descriptions of similar elements and / or processes employed in the various embodiments are omitted.
[0069] For the above Figure 2 , the anonymization of additional combinations of data is described, such as network attributes 277. For Figure 3 , the generated privacy constraints are described, and the generated privacy constraints can limit the anonymization component 107 in the creation of the anonymization network 290. In one or more embodiments, these privacy constraints can limit how the anonymization 275 process is performed, for example, specifying the value of K when the anonymization component 107 uses the K - anonymity method described above.
[0070] In additional embodiments, the constraint component 310 can facilitate providing entity - based access rights for accessing the anonymization network 290. In one or more embodiments, the method can allow the anonymization component 107 to enforce privacy based on entity - based access rights, thereby limiting unauthorized data access or inference. Thus, in one or more embodiments, the privacy constraints can be based on the data access rights of each node or entity stored with the node, where these privacy constraints are protected in the anonymization network.
[0071] In one or more embodiments, using privacy constraints that can be based on the data access rights of each node or entity stored with the node can provide flexibility in handling access to different types of data. For example, compared to many standard hospitals, in one or more embodiments, some mental health hospitals protect the privacy of nodes based on different node types rather than on the access rights granted to users. Additionally, it should be noted that since access rights are received from the node attributes of the individual nodes, in one or more embodiments, the method does not require using the role of the data requester to determine the level of anonymization to be applied by the anonymization component 107. It should be noted that in alternative embodiments, user - and role - based access rights can also be used exclusively or in combination with the above - described entity - based method.
[0072] Figure 5FIG. 500 is a block diagram of an example 500 of a non-limiting network anonymization system 102 in accordance with one or more embodiments described herein that is capable of facilitating anonymization of a network based on factors including network attributes and using privacy data from network nodes. For brevity, repeated descriptions of the same elements and / or processes employed in the various embodiments are omitted. In some embodiments, network anonymization system 102 may include an attributes component 108, an anonymization component 107, a constraints component 310, an optimality component 510, and any other components associated with network anonymization system 102 disclosed herein.
[0073] In one or more embodiments, anonymization of network 280 by anonymization component 107 may be further based on a combination of data for which the anonymized network is required to support acquisition. As described above, one or more embodiments support acquisition of data from anonymized network 290 by including the original or modified version of the data in the anonymization process. Additionally, these combinations of data to be acquired from anonymized network 290 may be embodied in the requirements of the anonymized network.
[0074] In an example, the requirements of a researcher may specify aspects of the granularity of the data to be acquired, thereby specifying the granularity of the data stored during anonymization. In this example, the researcher is conducting research based on the season of the year in which a person was born. Since, as described above, birthdays can be identifying characteristics that are rarely included in anonymized data in their original form, one or more anonymization methods may use bins of different sizes to introduce ambiguity into the birthday attribute. In this example, since the research is based on birth season, the granularity of the anonymized data may be the three-month length of a season. Thus, similar to the example Figures 7 - 8B shown below, for anonymization, the exact birthday may be binned according to the researcher's requirements into bins corresponding to three-month season intervals. Additionally, due to this aggregation of birthdays, in this example, the anonymity requirements of constraints component 310 may also be satisfied.
[0075] One or more embodiments may evaluate the overall effectiveness of the anonymized network and subsequently generate utility constraints that may limit the capabilities of anonymized network 290. For example, based on alternative anonymization strategies for meeting the privacy constraints generated by constraints component 310, an example of a reduction in the utility of the anonymized network is described below with reference to Figure 8B description of the reduction in the utility of the anonymized network.
[0076] In additional embodiments, to evaluate and adjust both utility and privacy constraints, the computer-executable component may further include an optimality component 510. In one approach, because privacy constraints for some types of data can be strict, the optimality component 510 may select utility constraints that can optimize the anonymization process or increase the optimality of the anonymization process based on some selected anonymization characteristics of the anonymization process of the anonymization component 107. In one or more embodiments, anonymization characteristics may include, but are not limited to, the degree of information loss during anonymizing network information. For example, in an example anonymization method, the optimality component 510 may select the minimum k value determined to satisfy the privacy constraints of the network to be anonymized. For instance, because increasing the number of entities (k) required to be indistinguishable necessarily results in increased data loss in the anonymized network 290.
[0077] Other anonymization characteristics used for optimization may include, but are not limited to, the number of network edits performed during the anonymization process, e.g., changes to the structure of the original network data to support anonymization. Example graph edits to the anonymized network may include, but are not limited to, edge addition / removal, edge swapping between any two nodes, or node addition / removal, edge attribute aggregation, or anonymization. For example, increasing the iterations towards the desired level of anonymity. In addition to the above, as discussed above and as depicted in Figure 8A -B, network edits may include node attribute aggregation, such as by aggregating attribute values in different bins to increase the granularity of the stored anonymized data.
[0078] Figure 6 FIG. shows a block diagram of an example non-limiting system 600 according to one or more embodiments described herein that is capable of facilitating anonymizing a network 280 based on factors including network attributes, node attributes, and edge attributes describing connections between nodes. For brevity, repeated descriptions of the same elements and / or processes employed in the various embodiments are omitted.
[0079] In some embodiments, the system 600 may include a network 280 transformed by the anonymization component 107 and published 406 as the anonymized network 290. In one or more embodiments, the network 280 and the anonymized network 290 may include a plurality of connected nodes, e.g., as described above with reference to Figure 2 described. In one or more embodiments, the constraint component 310 may be communicatively connected to the anonymization component 107. Additionally, the anonymized network consumer 450 may be communicatively coupled to the anonymized network 290, where the utility of the access is described by the utility metric 455.
[0080] For illustrative purposes, the following of Figure 6The description illustrates another method of operation of the optimality component 510 according to one or more embodiments. In this example, a utility constraint (U) based on a combination of one or more of node attributes, edge attributes, and network attributes can be used, where the node attributes, edge attributes, and network attributes are preferred attributes to be provided by the anonymization network 290 in response to a request from the consumer 515. Further for this example, a privacy constraint (P) can be defined based on the current privacy levels of the nodes in the network 280, e.g., based on access rights defined for users or entities of different nodes.
[0081] Based on the above constraints, the optimality component 510 can analyze the network 280 to identify different anonymization methods, where P is a hard constraint and U is a soft constraint, e.g., because privacy constraints are typically strict legal or ethical rules. Different anonymization methods can be evaluated according to criteria such as an objective function (e.g., information loss during anonymization, graph editing, or other optimization criteria).
[0082] During the evaluation of the method of the optimality component 510, penalties can be used to oppose the selection of anonymization methods that cannot maintain the preferred attributes in U. In one or more embodiments, for example, for node, edge, and network attributes, employing a penalty function can systematically improve the overall utility of the anonymized network 290 without compromising privacy. In some embodiments, this interaction between P and U can be described as a trade-off between P and U.
[0083] In another more detailed example of the operation of the optimality component 510, in one or more embodiments, for an example network G = (V, E), where V and E represent the set of nodes and the set of edges respectively, let X V 、X E and X N represent the node attributes, edge attributes, and network attributes of G respectively. In addition, let A be the set of access rights of all nodes, let A v be the access right defined by the node v, and where v ∈ V, let P and U. Let X G be the set of network attributes to be retained, e.g., the preferred attributes for meeting consumer requirements, as described above.
[0084] In operation, according to one or more embodiments, U can be applied based on an implementation selection to an ordered list of X V 、X E and X G , and for each v ∈ V, based on A v , update X Vv 、X Ev and X NvPermissible attributes. Continuing with this example, the privacy constraint P can be set by the constraint component 310 based on the set of access rights A for anonymization. The data is anonymized such that it satisfies P as a hard constraint while U (such as a specific network attribute, e.g., the betweenness centrality of a node set) remains as similar as possible in G and in the anonymized G'.
[0085] In conjunction with the anonymization process, in one or more embodiments, a penalty function F(G) can be implemented, which can preserve the nature of node, edge, or network attributes without compromising privacy constraints. In additional methods, some of the optimality methods described herein can assign weight attributes based on different factors (including the priority of maintaining their utility). In some implementations, to enforce the above constraints for each iteration, a change in the structure of the anonymized network 290 can be determined, which in some cases moves the result of anonymization towards satisfaction of P while penalizing changes in U (such as the network centrality attribute of a node set).
[0086] Figure 7 Table 700 depicts sample data from a network to be anonymized using different methods according to one or more embodiments. Table 700 includes 11 patients 710A - 710K, each patient having attributes corresponding to an ID, name 720, birthday 730, postal code 740, and gender 750.
[0087] Analyzing this data, it can be noted that patients 710A - 710K are from three different postal codes 740 and have unique IDs and names 720. For reference in the discussion Figures 8A - 8B of the anonymization methods herein, the highlighted attribute combination 790 emphasizes this identification combination of birthday 730, postal code 740, and gender 750.
[0088] Figures 8A - 8B depicts two alternative methods of anonymizing the data depicted in Figure 7 For brevity, repeated descriptions of the same elements and / or processes employed in the various embodiments are omitted.
[0089] Figure 8A depicts the anonymization of Table 700 using the k - anonymity method, where k = 3. It should be noted that for illustrative purposes, Figures 7 - 8B the examples herein include two types of attributes, namely, attributes that can directly identify an individual (e.g., ID and name 720), and publicly available attributes, and combinations of these attributes can sometimes be used to identify a specific person.
[0090] As described above, different embodiments can adjust the anonymization method based on the requirements of the data requester. In this example, different instances (not shown) seeking to study the condition of patients 712A - 712K are considered, and will include information about the condition and demographic data that can be used for the study, such as age determined from birthday 730, zip code 740 for residence information, and gender 750. In this instance, the age of the individual in months is requested.
[0091] In this example, the anonymization method begins by removing the ID and name 720, as non - essential data that can identify an individual. It should be noted that, similar to the name, the birthday is also unique for patients in the system, but since this data is requested to be available, different anonymization methods can be used by one or more embodiments that do not require removing the data. For example, as depicted in birthday 730, in this example, the months of the year are placed into groups of four months for anonymization. As Figure 8A emphasized, for this data set, the method meets the requirement of K = 3, e.g., at least K patients 712A - 712K are similar enough such that they cannot be distinguished from one another, e.g., patients 712I - 712K.
[0092] Returning to the Figures 5 - 6 discussion of the requirements and utility discussed, in this example, the age in months is a utility constraint, and k = 3 is a privacy constraint. Additionally, it should be noted that by choosing the three - month grouping method used for birthday 730, the utility value for the researcher may have been reduced, e.g., the age in months is only known within the three - month range used. The significance of this increased imprecision can depend on the specific circumstances of the study being conducted. In cases where there are more attributes in the real - world data set, anonymization options can be obtained where the birthday is represented by month and year.
[0093] Figure 8B Another example anonymization of table 700 according to one or more embodiments is depicted. In this example, the privacy requirement is more stringent than the previous example, e.g., k = 4. When choosing an alternative method, one or more embodiments can determine that only two different zip code prefix values (e.g., 550 and 650) are included in the data. Thus, instead of further grouping the birthday 730 and correspondingly reducing the utility of the anonymized age data, two groups of zip code prefixes can be used.
[0094] Thus, as depicted, patients 715A - 715D and patients 715I - 715K were previously distinguishable from each other by different ZIP code 740 values. However, since these differences are not in the ZIP code prefix, using this method, the above groups form block 890D of 7 indistinguishable patients 715A - K. Since records 715E - 715H form block 890E of 4 indistinguishable patient records, the new k value for the table is k = 4.
[0095] It should be noted that the detailed discussion of the anonymization methods included above is particularly aimed at illustrating how requirements and privacy restrictions can be related to a specific dataset. For example, if the data is different, such as the number of years of birth being different between patients, different methods will have to be used.
[0096] Another concept that should be understood based on the above discussion is the various different anonymization methods that can be used even for small, simple datasets. In one or more embodiments, using the various methods available to meet utility requirements and privacy requirements, the optimality component 510 can in some cases make minor adjustments, which can improve the performance of the anonymization network.
[0097] It should be understood that one or more embodiments described herein can utilize various combinations of electrical components, mechanical components, mass storage, and circuits that cannot be replicated in a human mind or performed by a human. For example, the analysis, processing, and anonymization network of relevant information is beyond the capabilities of the human mind. For example, the operation of the anonymization component 107 can be performed faster than what a human mind can perform within the same time period.
[0098] According to multiple embodiments, the network anonymization system 102 can also be fully operable to perform one or more other functions (e.g., fully powered on, fully executed, etc.) while also performing the various operations described herein. It should be understood that this simultaneous multi - operation execution is beyond the capabilities of the human mind. It should also be understood that the network anonymization system 102 can acquire, analyze, and process information that cannot be manually acquired, analyzed, and processed by an entity such as a human user. For example, the type, quantity, and / or variety of information included in the network anonymization system 102, the attribute component 108, the anonymization component 107, and any other components associated with the network anonymization system 102 as disclosed herein can be more complex than the information that can be manually obtained by a human user.
[0099] Figure 9A flow chart of an example non-limiting computer-implemented method 900 is shown that can facilitate anonymizing a network 280 based on factors including network attributes 227, node attributes, and connections 225A-B between nodes 220A-B, according to one or more embodiments described herein. For the sake of brevity, repeated descriptions of the same elements and / or processes employed in various embodiments are omitted.
[0100] In some embodiments, at 902, computer-implemented method 900 may include generating, by a device operatively coupled to a processor, network attributes of a network. For example, one or more embodiments may generate, by a device (e.g., network anonymization system 102) operatively coupled to a processor (e.g., processor 106), network attributes (e.g., network attributes 277) of a network (e.g., network 280).
[0101] In some embodiments, at 904, the computer-implemented method 900 may include anonymizing, by a device, network information of a network based on network attributes and node attributes of a first node of the network, thereby generating an anonymized network, wherein connection information corresponding to a connection between the first node and the second node includes edge attributes, and anonymizing the network information is further based on the edge attributes. For example, one or more embodiments may anonymize, by a device (e.g., anonymization component 107 in network anonymization system 102) network information of a network (e.g., network 280) based on network attributes (e.g., network attributes 277) and node attributes of a first node (e.g., node 220A) of the network (e.g., network 280), thereby generating an anonymized network (e.g., anonymized network 290), wherein connection information corresponding to a connection (e.g., connection 225B) between the first node (e.g., node 220A) and the second node (e.g., node 220B) includes edge attributes, and anonymizing the network information is further based on the edge attributes.
[0102] For the sake of simplicity of explanation, the computer-implemented method is depicted and described as a series of actions. It is understood and appreciated that the present invention is not limited by the actions and / or order of actions shown, for example, the actions can occur in various orders and / or concurrently, and can occur together with other actions not presented and described herein. In addition, not all the actions shown are necessary for realizing the computer-implemented method according to the disclosed subject matter. In addition, it will be understood and appreciated by those skilled in the art that the computer-implemented method can be represented as a series of interrelated states via state diagrams or events instead. In addition, it should also be understood that the computer-implemented method disclosed hereinafter and throughout this specification can be stored on an article of manufacture, so that these computer-implemented methods are transmitted and transferred to a computer. The term "article of manufacture" as used herein is intended to encompass a computer program accessible from any computer-readable device or storage medium.
[0103] Figure 10 Illustrates an example context of aspects of the disclosed subject matter in accordance with one or more embodiments. For example, this figure, and the following discussion, are intended to provide a general description of a suitable environment in which aspects of the disclosed subject matter may be implemented. For brevity, repeated descriptions of the same elements and processes employed in the various embodiments are omitted.
[0104] Figure 10 Shows a block diagram of an example non - restrictive operating environment in which one or more embodiments described herein may be facilitated. For brevity, repeated descriptions of similar elements employed in other embodiments described herein are omitted.
[0105] Referring Figure 10 , a suitable operating environment 1000 for implementing aspects of the present disclosure may also include a computer 1012. The computer 1012 may also include a processing unit 1014, a system memory 1016, and a system bus 1018. The system bus 1018 couples system components including, but not limited to, the system memory 1016 to the processing unit 1014. The processing unit 1014 may be any of a variety of available processors. Dual microprocessors and other multiprocessor architectures may also be used as the processing unit 1014. The system bus 1018 may be any of a variety of types of bus structures, including a memory bus or memory controller, a peripheral bus or external bus, and / or a local bus using any of a variety of available bus architectures, including, but not limited to, Industry Standard Architecture (ISA), Micro Channel Architecture (MSA), Extended ISA (EISA), Intelligent Drive Electronics (IDE), VESA Local Bus (VLB), Peripheral Component Interconnect (PCI), Card Bus, Universal Serial Bus (USB), Advanced Graphics Port (AGP), FireWire (IEEE 1394), and Small Computer System Interface (SCSI).
[0106] The system memory 1016 may also include volatile memory 1020 and non - volatile memory 1022. The Basic Input / Output System (BIOS) (containing basic routines such as transferring information between elements within the computer 1012 during startup) is stored in the non - volatile memory 1022. The computer 1012 may also include removable / non - removable, volatile / non - volatile computer storage media. For example, Figure 10A disk storage device 1024 is shown. The disk storage device 1024 may also include, but is not limited to, devices such as disk drives, floppy disk drives, tape drives, Jaz drives, Zip drives, LS-100 drives, flash cards, or memory sticks. The disk storage device 1024 may also include a separate storage medium or a storage medium in combination with other storage media. To facilitate connection of the disk storage device 1024 to the system bus 1018, a removable or non-removable interface such as interface 1026 is typically used. Figure 10 Software that acts as an intermediary between a user and the basic computer resources described in a suitable operating environment 1000 is also depicted. Such software may also include, for example, an operating system 1028. The operating system 1028, which may be stored on the disk storage device 1024, is used to control and allocate the resources of the computer 1012.
[0107] System applications 1030 utilize the operating system 1028 for resource management through, for example, program modules 1032 and program data 1034 stored in the system memory 1016 or the disk storage device 1024. It should be understood that the present disclosure may be implemented with various operating systems or combinations of operating systems. A user inputs commands or information into the computer 1012 through an input device 1036. The input device 1036 includes, but is not limited to, pointing devices such as a mouse, trackball, stylus, touch pad, keyboard, microphone, joystick, game pad, satellite dish, scanner, TV tuner card, digital camera, digital video camera, web camera, etc. These and other input devices are connected to the processing unit 1014 via the system bus 1018 through an interface port 1038. The interface port 1038 includes, for example, serial ports, parallel ports, game ports, and Universal Serial Bus (USB). Some of the (one or more) output devices 1040 use the same type of ports as the (one or more) input devices 1036. Thus, for example, a USB port may be used to provide input to the computer 1012 and output information from the computer 1012 to the output device 1040. An output adapter 1042 is provided to account for the presence of certain output devices 1040, such as monitors, speakers, and printers, and other output devices 1040 that require a dedicated adapter. By way of example and not limitation, the output adapter 1042 includes video cards and sound cards that provide a means of connection between the output device 1040 and the system bus 1018. It should be noted that other devices and / or systems of devices provide input and output capabilities, such as a remote computer 1044.
[0108] Computer 1012 can operate in a networked environment using a logical connection to one or more remote computers, such as remote computer 1044. Remote computer 1044 can be a computer, server, router, network PC, workstation, microprocessor-based appliance, peer device, or other common network node, etc., and typically can also include many or all of the elements described relative to computer 1012. For simplicity, only memory storage device 1046 is shown with remote computer 1044. Remote computer 1044 is logically connected to computer 1012 through network interface 1048 and then physically connected via communication connection 1050. Network interface 1048 includes wired and / or wireless communication networks, such as local area network (LAN), wide area network (WAN), cellular network, etc. LAN technologies include Fiber Distributed Data Interface (FDDI), Copper Distributed Data Interface (CDDI), Ethernet, Token Ring, etc. WAN technologies include, but are not limited to, point-to-point links, circuit-switched networks like Integrated Services Digital Network (ISDN) and its variants, packet-switched networks, and Digital Subscriber Line (DSL). Communication connection 1050 refers to the hardware / software for connecting network interface 1048 to system bus 1018. Although shown inside computer 1012 for clarity, it can also be outside computer 1012. For illustrative purposes only, the hardware / software for connection to network interface 1048 can also include internal and external technologies, such as modems including conventional telephone-grade modems, cable modems, and DSL modems, ISDN adapters, and Ethernet cards.
[0109] The present invention can be a system, method, apparatus, and / or computer program product at any possible level of integration of technical details. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention. The computer-readable storage medium can be a tangible device that is capable of retaining and storing instructions for use by an instruction execution device. The computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium can further include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or raised structures in a groove having instructions recorded thereon, and any appropriate combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0110] The computer-readable program instructions described herein can be downloaded to a corresponding computing / processing device from a computer-readable storage medium or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device. The computer-readable program instructions for carrying out operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, in order to carry out aspects of the present invention, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit.
[0111] Aspects of the present invention are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions. These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture including instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram. The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0112] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a special-purpose hardware-based system that performs the specified functions or acts or a combination of special-purpose hardware and computer instructions.
[0113] Although the subject matter has been described above in the general context of computer-executable instructions of a computer program product running on a computer and / or computers, those skilled in the art will recognize that the present disclosure may also be implemented in conjunction with, or be capable of being implemented in conjunction with, other program modules. Generally, program modules include routines, programs, components, data structures, etc. that perform particular tasks and / or implement particular abstract data types. In addition, those skilled in the art will appreciate that the computer-implemented methods of the present invention may be practiced with other computer system configurations, including single-processor or multi-processor computer systems, minicomputers, mainframe computers, and computers, hand-held computing devices (e.g., PDAs, telephones), microprocessor-based or programmable consumer or industrial electronic products, etc. The aspects shown may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network. However, some aspects of the present disclosure, if not all aspects, may be practiced on a stand-alone computer. In a distributed computing environment, program modules may be located in local and remote memory storage devices.
[0114] As used in this application, the terms "component", "system", "platform", "interface", etc. may refer to and / or may include computer-related entities or entities related to an operating machine with one or more specific functions. The entities disclosed herein may be hardware, a combination of hardware and software, software, or software in execution. For example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, an application running on a server and the server can both be components. One or more components may reside within a process and / or an execution thread, and a component may be located on one computer and / or distributed between two or more computers. In another example, the corresponding components may execute from various computer-readable media on which various data structures are stored. These components may communicate via local and / or remote processes, such as in accordance with a signal having one or more data packets (e.g., data from one component that interacts with another component in a local system, a distributed system, and / or interacts with other systems via a network such as the Internet). As another example, a component may be a device having a particular function provided by a mechanical component operated by an electrical or electronic circuit, where the electrical or electronic circuit is operated by a software or firmware application executed by a processor. In such a case, the processor may be inside or outside the device and may execute at least a portion of the software or firmware application. As yet another example, a component may be a device that provides a particular function through an electronic component rather than a mechanical component, where the electronic component may include a processor or other device to execute software or firmware that at least partially imparts the function of the electronic component. In one aspect, a component may emulate an electronic component via a virtual machine, such as within a cloud computing system.
[0115] In addition, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless specified otherwise or clear from the context, "X employs A or B" is intended to mean any natural inclusive permutation. That is, if X employs A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied in any of the foregoing instances. Further, unless specified otherwise or clear from the context to refer to the singular form, the articles "a" and "an" as used in this specification and the drawings shall generally be construed to mean "one or more". As used herein, the terms "example" and / or "exemplary" are used to mean serving as an example, instance, or illustration. To avoid doubt, the subject matter disclosed herein is not limited by these examples. Further, any aspect or design described herein as "example" and / or "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects or designs, nor does it imply exclusion of equivalent exemplary structures and techniques known to those of ordinary skill in the art.
[0116] As used in this specification, the term "processor" can refer to substantially any computing processing unit or device, including but not limited to a single-core processor; a single processor with software multithreading execution capabilities; a multi-core processor; a multi-core processor with software multithreading execution capabilities; a multi-core processor with hardware multithreading technology; a parallel platform; and a parallel platform with distributed shared memory. Additionally, a processor can refer to an integrated circuit, an application specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic controller (PLC), a complex programmable logic device (CPLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. Further, a processor can employ nanoscale architectures, such as but not limited to molecular and quantum dot based transistors, switches, and gates, in order to optimize space usage or enhance the performance of user equipment. A processor can also be implemented as a combination of computing processing units.
[0117] In the present disclosure, terms such as "storage", "database", and substantially any other information storage component related to the operation and functionality of components are used to refer to "memory components", entities embodied in "memory", or components that include memory. It should be understood that the memory and / or memory components described herein can be volatile memory or non-volatile memory or can include both volatile and non-volatile memory. By way of illustration and not limitation, non-volatile memory can include read only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), flash memory, or non-volatile random access memory (RAM) (e.g., ferroelectric RAM (FeRAM)). Volatile memory can include RAM, which can be used as, for example, an external cache memory. By way of illustration and not limitation, RAM can be obtained in many forms, such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), direct Rambus RAM (DRRAM), direct Rambus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM). In addition, the memory components of the systems or computer-implemented methods disclosed herein are intended to include, but are not limited to, including these and any other suitable types of memory.
[0118] The foregoing description includes only examples of systems and computer-implemented methods. Of course, it is not possible to describe every conceivable combination of components or computer-implemented methods for the purpose of describing the present disclosure, but one of ordinary skill in the art will recognize that many other combinations and permutations of the present disclosure are possible. In addition, with respect to the use of the terms "including", "having", "owning", etc. in the detailed description, the claims, the appendices, and the drawings, these terms are intended to be inclusive in a manner similar to the way the term "comprising" is interpreted when used as a transitional word in the claims.
[0119] The description of the various embodiments has been presented for purposes of illustration, but the description is not exhaustive or intended to be limited to the disclosed embodiments. Many modifications and variations will be apparent to one of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein were chosen to best explain the principles of the embodiments, the practical application, or a technical improvement over technologies found in the marketplace, or to enable one of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A system for anonymizing a computer network, comprising: a memory that stores computer-executable components; a processor that executes the computer-executable components stored in the memory, wherein the computer-executable components include: an anonymization component that is configured to: select an anonymization method from a set of anonymization methods using an objective function with optimization criteria, based on privacy constraints, utility constraints, and a penalty function based on network attributes, node attributes, and edge attributes of a node network; and use the anonymization method to anonymize network information of the node network based on network attributes of the node network and node attributes of a first node in the node network, thereby generating an anonymized network.
2. The system according to claim 1, wherein connection information corresponding to a connection between the first node and a second node includes edge attributes, and wherein anonymizing the network information is further based on the edge attributes.
3. The system according to claim 1, wherein anonymizing the network information is further based on a privacy constraint that enforces a privacy level of the anonymized network.
4. The system according to claim 3, wherein the computer-executable components further include: a constraint component that generates the privacy constraint based on the network information.
5. The system according to claim 4, wherein the constraint component further determines an access right of the first node, and wherein the constraint component further generates the privacy constraint based on the access right.
6. The system according to claim 3, wherein anonymizing the network information is further based on a utility constraint for the anonymized network, and wherein the utility constraint includes a first measure of the utility of the anonymized network based on consumer-based requirements.
7. The system according to claim 6, wherein the computer-executable components further include: an optimality component that selects a utility constraint that optimizes or increases the optimality of anonymizing the network information based on anonymization characteristics of anonymizing the network information.
8. The system according to claim 7, wherein the anonymization characteristics include a second measure of information loss during anonymizing the network information.
9. The system according to claim 7, wherein the anonymization characteristics include the number of edits of the network performed during anonymizing the network information.
10. The system according to claim 1, wherein the network information includes medical information of a plurality of patients.
11. The system according to claim 1, wherein the network includes a social media network of a plurality of users.
12. A method for anonymizing a computer network, comprising: generating, by a device operatively coupled to a processor, network attributes of a node network; selecting, by the device, an anonymization method from a set of anonymization methods using an objective function with optimization criteria, based on privacy constraints, utility constraints, and a penalty function based on the network attributes, node attributes, and edge attributes of the node network; and The device uses the anonymization method to anonymize the network information of the node network based on the network attributes and the node attributes of the first node in the node network, thereby generating an anonymized network.
13. The method according to claim 12, wherein, the connection information corresponding to the connection between the first node and the second node includes edge attributes, and wherein anonymizing the network information is further based on the edge attributes.
14. The method according to claim 12, wherein, anonymizing the network information is further based on a privacy constraint that enforces the privacy level of the anonymized network.
15. The method according to claim 14, further comprising: generating the privacy constraint based on a privacy rule, and wherein the privacy rule is applied to the first node.
16. The method according to claim 14, further comprising: generating the privacy constraint based on the network information.
17. The method according to claim 16, comprising: determining the access rights of the first node, and wherein the constraint component further generates the privacy constraint based on the access rights.
18. The method according to claim 14, wherein, anonymizing the network information is further based on a utility constraint for the anonymized network, and wherein the utility constraint includes a measure of the utility of the anonymized network based on consumer requirements.
19. The method according to claim 18, further comprising: selecting a utility constraint that optimizes or increases the optimality of anonymizing the network information based on the anonymization characteristics of anonymizing the network information.
20. The method according to claim 19, wherein, the anonymization characteristic includes a second measure of information loss during anonymizing the network information.
21. The method according to claim 19, wherein, the anonymization characteristic includes the number of edits of the network performed during anonymizing the network information.
22. The method according to claim 12, wherein, the network information includes medical information of multiple patients.
23. The method according to claim 12, wherein, the network includes a social media network of multiple users.
24. The method according to claim 12, wherein, the network attribute is for a network of connected nodes, wherein multiple of the connected nodes among the multiple connected nodes include node attributes, thereby generating multiple node attributes, and wherein generating the network attribute is based on the multiple node attributes, the method further includes: generating, by the processor, a privacy constraint based on the network information of the connected nodes, wherein the privacy constraint enforces the privacy level of the anonymized network; and anonymizing the network information of the connected nodes based on the network attribute, the multiple node attributes, and the privacy constraint, thereby generating an anonymized network.
25. A computer program product, the computer program product includes program instructions that can be executed by a processor to cause the processor to execute the method according to any one of claims 12 to 24.
26. A computer-readable storage medium storing program code which, when run on a computer, causes the computer to perform the method according to any one of claims 12 to 24.
Citation Information
Patent Citations
Network information collection and access control system
US20130160138A1
Guaranteeing anonymity of linked data graphs
US20140325666A1