Method, computing system, computer program and system (privacy-preserving graph analytics on hybrid cloud environments)

The method addresses privacy concerns in triangle counting by dividing graph data, processing it in parallel across a public cloud, and securely combining results within a hybrid cloud environment, ensuring efficient and private triangle counting.

JP2025092435APending Publication Date: 2025-06-19INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024202070
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-08
Filing Date
2024-11-20
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Triangle counting in graphs raises privacy concerns, particularly in hybrid cloud environments, as it can violate edge privacy by revealing sensitive relationships between individuals.

Method used

A method is provided that divides graph data into subgraphs, disperses them across separate servers in a public cloud, expands each subgraph with new vertices and random connections, calculates triangle counts in parallel, and combines results securely on-premises to protect privacy.

Benefits of technology

This method effectively protects privacy by ensuring that exact triangle counts remain within the private cloud, while parallel processing in the public cloud reduces computational resource usage and minimizes data exposure to potential eavesdroppers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025092435000001_ABST
    Figure 2025092435000001_ABST
Patent Text Reader

Abstract

To provide a method for preserving privacy by counting triangles on a graph for hybrid cloud environments.SOLUTION: The method includes partitioning data elements of the graph for hybrid cloud environments into a plurality of non-overlapping subgraphs, modifying each of the non-overlapping subgraphs to generate a plurality of server induced subgraphs, distributing each of the server induced subgraphs to a separate server located on a public cloud environment, computing a resultant number of triangles associated with each of the server induced subgraphs for each of the separate servers located on the public cloud environment, and computing a final number of triangles associated with the resultant number of triangles via an on-premise server.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to a method for counting triangles in a graph, and more specifically, to a method for protecting privacy by counting triangles in a graph of a hybrid cloud environment.

Summary of the Invention

Problems to be Solved by the Invention

[0002] Triangle counting in graph mining is a data analysis technique widely used in social network analysis, recommendation systems, and other complex fields. Accordingly, triangle counting is used to address important problems in the field of graph mining. For example, triangle counting is a critical parameter when mining relationships among people in social network analysis. Unfortunately, counting the number of triangles in a graph is a fundamental problem for many applications. This is because counting triangles is (a) counting cycles of a given length, (b) counting complete subgraphs of a given size, or (c) a special case of certain other small subgraphs called motifs. One such problem with triangle counting is that it can raise privacy concerns, such as violating edge privacy, such as sensitive relationships between individuals.

Means for Solving the Problems

[0003] A method for protecting privacy by counting triangles in a graph of a hybrid cloud environment is provided. The method includes: dividing data elements of a graph G into a plurality of non-overlapping subgraphs; dispersing each of the plurality of non-overlapping subgraphs to a separate server located in a public cloud environment; expanding each of the plurality of non-overlapping subgraphs using new vertices and random edge connections to generate p induced subgraphs; calculating, for each of the servers located in the public cloud environment, the number of triangles (p being an integer) obtained as a result associated with each of the plurality of non-overlapping subgraphs; transmitting the p integer to an on-premises server; and calculating the final number of triangles associated with the subgraph formed by combining the p induced subgraphs by adding the p integer to the number of triangles associated with the subgraph formed by combining the p induced subgraphs and subtracting the number of triangles having at least one additional edge connection to the previously added random edge connections, via the on-premises server.

[0004] Embodiments of the present invention also relate to a computer-implemented method and a computer program product having substantially the same characteristics and functions as the above computer system.

[0005] Further technical features and advantages are realized by the technology of the present invention. Embodiments and aspects of the present invention are described in detail herein and are considered part of the claimed subject matter. For a better understanding, refer to the detailed description and the drawings.

Brief Description of the Drawings

[0006] The details of the exclusive rights described herein are particularly pointed out and are clearly claimed in the claims at the conclusion of the specification. The foregoing and other features and advantages of embodiments of the present invention will be apparent from the following detailed description when read in conjunction with the accompanying drawings.

[0007]

Figure 1

[0008]

Figure 2

[0009]

Figure 3

[0010]

Figure 4

DETAILED DESCRIPTION OF THE INVENTION

[0011] As briefly described above, the present invention relates to a method for counting triangles in a graph of a hybrid cloud environment while protecting the privacy of graph analysis.

[0012] It should be understood that counting the number of triangles in a graph is a fundamental problem for many applications. This is because counting triangles is (a) counting cycles of a given length, (b) counting complete subgraphs of a given size, and / or (c) a special case of certain other small subgraphs called motifs. More importantly, counting triangles is a fundamental problem in network analysis. One important application involves the fact that a high clustering coefficient, which indicates "tight communities" of people, means that most pairs of people with mutual friends are friends themselves. Such communities are expected to have interesting properties such as an unusually high degree of "trust", where many of your other friends are likely to hear about it, strongly discouraging you from doing something wrong to someone you are connected to. Another important reason involves the fact that a low clustering coefficient can indicate "structural holes" containing vertices that are well-connected by different communities but not otherwise connected to each other. Such vertices are in potentially advantageous situations, such as generating innovation by combining two different skill sets or at least transferring ideas from one community to another.

[0013] Protecting edge privacy (and other privacy concerns) in triangle counting is a challenge, particularly when triangle counting involves hybrid and cloud environments, due to strong correlations between data from different clients. It is even more difficult to provide a clear solution to the current problem, as there are hundreds, if not thousands, of algorithms available to choose from to solve these problems, and these algorithms are broadly classified into exact methods (i.e., calculating the exact number of triangles) or approximate methods (i.e., calculating a faster estimate that is 90% to 95% accurate). In one embodiment, the method of the present invention focuses on the first category, and the method is modified to provide an approximate answer, thereby also achieving this complexity. The method of the present invention uses a new procedure to avoid the use of costly privacy mechanisms by addressing eavesdropping at the numerical level. In other words, the approach of the present method does not provide a faster scheme than current algorithms, but rather provides a method that is expected to require far fewer resources to accurately complete its task, as a hybrid cloud environment is assumed. Accordingly, even if some data is recovered from one of the servers (assuming there are multiple such servers) due to eavesdropping caused by additional incorrect vertices, the eavesdropper cannot actually know which data is correct and which data is superficial.

[0014] Accordingly, embodiments of the method utilize public cloud resources to execute parallel tasks that represent the major bulk of the overall computational complexity, where a private cloud environment is used to execute the remaining tasks. This approach introduces at least two advantages, namely: a) the exact number of triangles is not communicated outside the private cloud environment, and b) the tasks executed on the private cloud environment are also non-parallelizable tasks, so that instead of consuming time and money communicating data to the server, on-premises processing capabilities are utilized. The method of the present invention assumes a hybrid cloud environment, where all data is stored on-premises (i.e., not decentralized), and it should be understood that only a portion of the data is transferred to an external server (i.e., a public server) for parallel computing. Additionally, the method assumes that all data is erased after the parallel computations on the data transferred to the external server are completed. Further, to protect the data being transferred, the method further introduces a portion of dummy vertices and connections on top of standard encryption. In one embodiment, the method essentially divides the data and generates parallel subtasks on public and private servers.

[0015] Any two of the vertices are connected by an edge of the graph G (3-clique), and the matrix A ∈ {0,1} n×n is the corresponding adjacency matrix, where considering the situation where a triangle is a set of three vertices, nnz(A) indicates the number of non-zero entities (i.e., edges). Therefore, the number of triangles in the graph G is

Number

Number

Number

[0016] It should be understood that one problem that can be addressed includes generating marketing opportunities. Since a high transitivity ratio suggests similarity between nodes, marketing opportunities in an e - commerce platform (e.g., if they form triangles, proposing to user i what you proposed to users j and k) can be generated by calculating the transitivity ratio. Another problem includes community detection (i.e., identifying clusters of vertices) by finding adjacent vertices with high triangle involvement.

[0017] Various aspects of the present disclosure are illustrated by the description, flowchart, block diagram of a computer system, and / or block diagram of machine logic included in embodiments of a computer program product (CPP). For any flowchart, depending on the technology involved, operations may be executed in an order different from that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be executed in reverse order, as a single integrated step, simultaneously, or at least partially in a time - overlapping manner.

[0018] An embodiment of a computer program product (referred to herein as a "CPP embodiment" or "CPP") is a term used in this disclosure to describe any set of one or more storage media (also referred to as "media") collectively included in a set of one or more storage devices that collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. By way of non-limiting example, a computer-readable storage medium may be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media are floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on the major surfaces of discs), or any suitable combination of the foregoing. A computer-readable storage medium shall not be construed as storage in the form of a transitory signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses passing through an optical fiber cable, electrical signals communicated through a wire, and / or other transmission media, as the term is used in this disclosure. As will be understood by those skilled in the art, data is typically moved during normal operation of a storage device at some irregular points in time, such as during access, defragmentation, or garbage collection, but the data is not transitory while it is stored, and thus the storage device is not considered to be transitory for the purposes of the foregoing.

[0019] Computing environment 100 includes an example of an environment for at least some execution of computer code involved in performing the inventive method, such as 150 for solving a contextual bandit problem having a trend reward function. In addition to block 150, computing environment 100 includes, for example, computer 101, wide area network (WAN) 102, end user device (EUD) 103, remote server 104, public cloud 105, and private cloud 106. In this embodiment, computer 101 includes a processor set 110 (including processing circuit 120 and cache 121), communication fabric 111, volatile memory 112, persistent storage 113 (including operating system 122 and block 150 as shown above), a set of peripheral devices 114 (including user interface (UI), device set 123, storage 124, and Internet of Things (IoT) sensor set 125), and network module 115. Remote server 104 includes remote database 130. Public cloud 105 includes gateway 140, cloud orchestration module 141, host physical machine set 142, virtual machine set 143, and container set 144.

[0020] Computer 101 can take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch, or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or hereafter developed that is capable of executing programs, accessing a network, or querying a database such as remote database 130. As is well understood in the field of computer technology and depending on the technology, the execution of computer implementation methods can be distributed among multiple computers and / or between multiple locations. On the other hand, in this description of computing environment 100, for the sake of brevity as much as possible, the detailed discussion focuses on a single computer, specifically computer 101. Although computer 101 is not shown within the cloud in FIG. 1, it may be located within the cloud. On the other hand, computer 101 does not need to exist within the cloud except for any range that can be affirmatively shown.

[0021] Processor set 110 includes one or more computer processors of any type now known or hereafter developed. Processing circuitry 120 can be distributed among multiple packages, for example, multiple tuned integrated circuit chips. Processing circuitry 120 can implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within a processor chip package and is typically used for data or code that should be available for fast access by threads or cores executing on processor set 110. Cache memory is typically organized into multiple levels depending on its relative proximity to the processing circuitry. Alternatively, some or all of the cache for the processor set can be located "off-chip". In some computing environments, processor set 110 can be designed to operate using qubits and execute quantum computing.

[0022] Computer-readable program instructions cause a set of operation steps to be executed by the processor set 110 of computer 101, thereby realizing a computer-implemented method that is normally loaded onto computer 101. As a result, the instructions thus executed instantiate the method specified in the flowchart and / or description of the computer-implemented method (collectively referred to as "the method of the present invention") included in this document. These computer-readable program instructions are stored in various types of computer-readable storage media such as cache 121 and other storage media discussed below. The program instructions and related data are accessed by the processor set 110 to control and direct the execution of the method of the present invention. In computing environment 100, at least some of the instructions for executing the method of the present invention may be stored in block 150 within persistent storage 113.

[0023] Communication fabric 111 is a signal conduction path that enables various components of computer 101 to communicate with each other. Typically, this fabric is made up of switches and conductive paths such as buses, bridges, switches that make up physical input / output ports, and conductive paths. Other types of signal communication paths such as optical fiber communication paths and / or wireless communication paths may be used.

[0024] Volatile memory 112 is any type of volatile memory known currently or developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory is characterized by random access, but this is not essential unless affirmatively indicated. In computer 101, volatile memory 112 is located within a single package and exists inside computer 101. Alternatively or additionally, volatile memory may be distributed across multiple packages and / or located external to computer 101.

[0025] The persistent storage 113 is any form of non-volatile storage for a computer, known currently or developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether power is directly supplied to the computer 101 and / or to the persistent storage 113. The persistent storage 113 can be read-only memory (ROM), but usually at least a portion of the persistent storage enables writing of data, deletion of data, and re-writing of data. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 122 may take several forms, such as various known proprietary operating systems or an open-source portable operating system interface type of operating system that utilizes a kernel. The code included in block 150 typically includes at least some of the computer code involved in the execution of the method of the present invention.

[0026] The peripheral device set 114 includes a set of peripheral devices of the computer 101. The data communication connections between the peripheral devices of the computer 101 and other components may be implemented in various ways, such as a Bluetooth (registered trademark) connection, a Near Field Communication (NFC) connection, a connection formed by a cable (such as a Universal Serial Bus (USB) type cable), an insertion type connection (e.g., a Secure Digital (SD) card), a connection formed through a local area communication network, and even a connection formed through a wide area network such as the Internet. In various embodiments, the UI device set 123 may include components such as a display screen, a speaker, a microphone, wearable devices (such as glasses and smartwatches), a keyboard, a mouse, a printer, a touchpad, a game controller, and a haptic device. The storage 124 is an external storage such as an external hard drive or an insertable storage such as an SD card. The storage 124 may be persistent and / or volatile. In some embodiments, the storage 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 101 is required to have a large amount of storage (e.g., when the computer 101 locally stores and manages a large-scale database), this storage may be provided by a peripheral storage device designed to store a very large amount of data, such as a Storage Area Network (SAN) shared by a plurality of geographically distributed computers. The IoT sensor set 125 is composed of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer, and another sensor may be a motion detector.

[0027] The network module 115 is an aggregation of computer software, hardware, and firmware that enables the computer 101 to communicate with other computers via the WAN 102. The network module 115 may include hardware such as a modem or a Wi-Fi (registered trademark) signal transceiver, software for packetizing and / or depacketizing data for communication over a communication network, and / or web browser software for communicating data over the Internet. In some embodiments, the network control function and the network transfer function of the network module 115 are executed on the same physical hardware device. In other embodiments (e.g., embodiments that utilize Software-Defined Networking (SDN)), the control function and the transfer function of the network module 115 are executed on physically separate devices such that the control function manages several different network hardware devices. The computer-readable program instructions for executing the method of the present invention can typically be downloaded to the computer 101 from an external computer or an external storage device through a network adapter card or a network interface included in the network module 115.

[0028] The WAN 102 is any wide area network (e.g., the Internet) that can communicate computer data over a non-local distance by any technique for communicating computer data that is currently known or developed in the future. In some embodiments, the WAN can be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi network. The WAN and / or the LAN typically includes computer hardware such as copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.

[0029] An end-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of an enterprise operating computer 101) and can take any of the forms described above in relation to computer 101. The EUD 103 typically receives beneficial and useful data from the operation of computer 101. For example, in a virtual case where computer 101 is designed to provide recommendations to an end user, this recommendation will typically be communicated from the network module 115 of computer 101, via the WAN 102, to the EUD 103. In this way, the EUD 103 can display or otherwise present the recommendation to the end user. In some embodiments, the EUD 103 can be a client device such as a thin client, a thick client, a mainframe computer, and a desktop computer, etc.

[0030] A remote server 104 is any computer system that provides at least some data and / or functions to computer 101. The remote server 104 can be controlled and used by the same entity that operates computer 101. The remote server 104 represents a machine that collects and stores useful data for use by other computers such as computer 101. For example, in a virtual case where computer 101 is designed and programmed to provide recommendations based on historical data, this historical data can be provided from the remote database 130 of the remote server 104 to computer 101.

[0031] The public cloud 105 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computer functions, particularly data storage (cloud storage) and computing capabilities, without direct active management by the user. Cloud computing typically exploits resource sharing to achieve coherence and economies of scale. The direct active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments that run on various computers that make up the host physical machine set 142, which is the universe of physical computers within and / or available in the public cloud 105. The virtual computing environment (VCE) typically takes the form of virtual machines from the virtual machine set 143 and / or containers from the container set 144. It is understood that these VCEs can be stored as images and transferred either as images or after instantiation of the VCE among and within various physical machine hosts. The cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of the VCE, and manages the active instantiation of VCE deployments. The gateway 140 is a collection of computer software, hardware, and firmware that enables the public cloud 105 to communicate via the WAN 102.

[0032] Here, some further explanations of virtual computing environments (VCEs) are provided. A VCE can be stored as an "image". A new active instance of a VCE can be instantiated from the image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of the operating system where the kernel enables the existence of multiple isolated instances of user space, called containers. These isolated instances of user space typically behave as actual computers from the perspective of the programs running within them. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices allocated to the container, and this feature is known as containerization.

[0033] The private cloud 106 is similar to the public cloud 105, except that computing resources are available only for use by a single enterprise. The private cloud 106 is shown as being in communication with the WAN 102, but in other embodiments, the private cloud may be completely disconnected from the Internet and accessible only via a local / private network. A hybrid cloud is a composite of multiple different types of clouds (e.g., types of private clouds, community clouds, or public clouds), and is often implemented by different vendors. Each of the multiple clouds remains a separate discrete entity, but the larger hybrid cloud architecture is coupled by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.

[0034] One or more embodiments described in this specification may utilize machine learning techniques to perform tasks. More specifically, one or more embodiments described in this specification may incorporate and utilize rule-based decision-making and artificial intelligence (AI) inference to perform various operations described herein, namely, to implement containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of the operating system that enables the kernel to allow the existence of multiple isolated instances of user space, called containers. These isolated instances of user space typically behave as actual computers from the perspective of the programs running within them. A computer program running on a normal operating system may utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices assigned to the container, and this feature is known as containerization.

[0035] An ANN can be embodied as a so-called "neuromorphic" system of interconnected processor elements that act as simulated "neurons" and exchange messages with each other in the form of electrical signals. Similar to the so-called "plasticity" of synaptic neurotransmitter connections that carry messages between biological neurons, the connections in an ANN that carry electronic messages between simulated neurons are provided with numerical weights corresponding to the strength or weakness of a given connection. These weights can be adjusted and regulated based on experience to adapt the ANN to the input and enable learning. For example, an ANN for handwritten character recognition is defined by a set of input neurons that can be activated by the pixels of an input image. After being weighted and transformed by a function determined by the designer of the network, the activation of these input neurons is passed on to other downstream neurons, which are often referred to as "hidden" neurons. This process is repeated until the output neurons are activated. The activated output neurons determine which character was input. It should be understood that these same techniques can be applied in the case of containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of the operating system that enables the kernel to allow for the existence of multiple isolated instances of user space called containers. These isolated instances of user space typically behave as actual computers from the perspective of the programs running within them. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices allocated to the container, and this feature is known as containerization.

[0036] Due to the large scale of modern networks, it should be understood that counting the number of triangles in the associated graph requires distributed memory computing. In one embodiment, the method of the present invention contemplates an asynchronous system for counting triangles in a hybrid cloud environment. In particular, the method employs a hybrid cloud approach that divides the computation into two different parts, namely: 1) one part that is computed in a public cloud environment, and 2) one part that is computed in a private cloud environment. In addition to the execution speedup, the method maintains the state where the exact number of triangles is hidden from adversarial attacks even when the data encryption protocol is lost. Since counting the triangles of a large graph can be very time-consuming, the total number of triangles can be computed faster by decomposing the original graph into p ∈ Z partitions, thereby applying the divide-and-conquer method.

[0037] Consider, for example, a situation where the permuted adjacency matrix of the partitioned graph can be described as follows.

Number

Number

Number

Number

Number

[0038] According to an embodiment, as shown in FIG. 2, a method 300 for protecting privacy by counting triangles in a graph of a hybrid cloud environment is provided. Method 300 includes, as shown in operation block 302, dividing the data elements of graph G into p non-overlapping subgraphs. Referring to FIG. 3, this can be achieved by dividing the data elements of graph G into subgraphs based on a predefined parameter that separates the data elements within the graph from each other. Referring to FIG. 4, method 300 includes, as shown in operation block 304, dispersing each of the subgraphs to a separate server (from p servers) located in a public cloud environment, where each of the subgraphs is extended with a new set of vertices and additional random edge connections. Method 300 includes, as shown in operation block 306, calculating, at each of the p servers, the number of triangles associated with the dispersed subgraph. In this regard, the resulting number of triangles (i.e., p integers) is returned to the on-premises (i.e., private) server.

[0039] Method 300 also includes, as shown in operation block 308, calculating, at the on-premises server, the number of triangles associated with the subgraph formed by the combination of the p induced subgraphs. Method 300 includes, as shown in operation block 310, adding the number of triangles associated with the p servers to the number of triangles associated with the subgraph formed by the combination of the p induced subgraphs and subtracting the number of triangles having at least one additional edge connection to the previously added random edge connections.

[0040] It should be understood that method 300 protects the privacy of data by applying a first security layer and a second security layer. The first security layer includes a hybrid cloud environment where the total number of triangles is Δ = Δ public +Δ private equal to. Thus, the private cloud Δ privateAs long as it is secure, it is generally impossible to intercept Δ. However, eavesdroppers may potentially intercept the public cloud Δ public and matrix B j Accordingly, the second security layer involves applying redundant and random data. To address the eavesdropping problem of eavesdroppers, method 300 expands matrix Bj by introducing a random number of dummy vertices and introducing the actual vertices of the j-th partition into the random connections between these newly introduced vertices. These quantities are then transferred to the public cloud, and the number of resulting triangles

Number

Number

Number

Number

Number

[0041] Various embodiments of the present invention are described herein with reference to the accompanying drawings. Alternative embodiments of the present invention can be devised without departing from the scope of the present invention. In the following description and the drawings, various connection relationships and positional relationships (e.g., above, below, adjacent, etc.) between elements are described. These connection and / or positional relationships can be direct or indirect unless otherwise specified, and the present invention is not intended to be limited in this regard. Thus, the coupling of entities can refer to a direct or indirect coupling, and the positional relationship between entities can be a direct or indirect positional relationship. Also, the various tasks and process steps described herein can be incorporated into additional steps or more comprehensive procedures or processes having functions not detailed herein.

[0042] For the sake of brevity, the prior art related to the manufacture and use of aspects of the present invention may or may not be described in detail herein. In particular, aspects of computing systems and specific computer programs for implementing the various technical features described herein are well known. Thus, for the purpose of brevity, many details of conventional implementations are only briefly mentioned herein or are completely omitted without providing details of well-known systems and / or processes.

[0043] In some embodiments, various functions or operations may or may not occur at a given location and / or in relation to the operation of one or more devices or systems. In some embodiments, a portion of a given function or operation may be performed at a first device or location, and the remaining portion of the function or operation may be performed at one or more additional devices or locations.

[0044] The terms used in this specification are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The terms "comprises" and / or "comprising", when used in this specification, specify the presence of the stated feature, integer, step, operation, and / or element component, but do not preclude the presence or addition of one or more other features, integers, steps, operations, element components, and / or groups thereof.

[0045] In the following claims, any means-plus-function element or step-plus-function element's corresponding structure, material, act, and equivalents are intended to include any structure, material, or act for performing the function in combination with other claimed elements that are specifically claimed. Although this disclosure has been presented for purposes of illustration and description, it is not intended to be exhaustive or limited to the disclosed form. It will be apparent to those skilled in the art that many modifications and variations can be made without departing from the scope and spirit of this disclosure. Embodiments were selected and described in order to best explain the principles of this disclosure and its practical application and to enable others skilled in the art to understand this disclosure for various embodiments with various modifications as are suited to the particular use contemplated.

[0046] The drawings illustrated in this specification are exemplary. Without departing from the spirit of this disclosure, many changes can be made to the figures or the steps (or operations) described therein. For example, the actions can be performed in a different order, and the actions can be added, deleted, or modified. Also, the term "coupled" indicates that there is a signal path between two elements and does not imply a direct connection between elements without intervening elements / connections therebetween. All of these variations are considered to be a part of this disclosure.

[0047] The following definitions and abbreviations are used in the interpretation of the claims and the specification. As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," "contains," or "containing," or any other variation thereof, are intended to cover non-exclusive inclusion. For example, a composition, mixture, process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements, but may include other elements not expressly listed or other elements inherent to such composition, mixture, process, method, article, or apparatus.

[0048] In addition, the term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs. The terms "at least one" and "one or more" are understood to include any integer greater than or equal to one, i.e., one, two, three, four, etc. The term "a plurality" is understood to include any integer greater than or equal to two, i.e., two, three, four, five, etc. The term "connected" may include both indirect and direct "connections."

[0049] The terms "about," "substantially," "approximately," and variations thereof are intended to include the degree of error associated with the measurement of a particular quantity based on the equipment available at the time of filing. For example, "about" can include a range of ±8% or 5% or 2% of a given value.

[0050] The present invention may be an integrated system, method, and / or computer program product at any possible technical detail level. The computer program product may include one or more computer-readable storage media having computer-readable program instructions for causing a processor to execute aspects of the present invention.

[0051] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes, hereinafter, namely a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a punch card, or a mechanically encoded device such as a raised structure in a groove in which instructions are recorded, and any suitable combination of the foregoing. A computer-readable storage medium, as used herein, should not be construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through an electrical wire.

[0052] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium in each respective computing / processing device.

[0053] Computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk® or C++, and procedural programming languages such as the “C” programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit in order to carry out aspects of the present invention.

[0054] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0055] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium containing the instructions comprises a manufacture including instructions for implementing the aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0056] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, whereby the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0057] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions that include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may be performed in an order different from that noted in the drawings. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or, depending on the related functions, these blocks may sometimes be executed in the reverse order. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0058] The description of the various embodiments of the present invention is presented for purposes of illustration and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application, or the technical improvement in the marketplace, or to enable other ordinary skill in the art to understand the embodiments described herein. Also, embodiments or portions of embodiments may be combined with each other, either in whole or in part, without departing from the scope of the present invention.

Claims

1. 1. A method for protecting privacy by counting triangles in a graph G in a hybrid cloud environment, the method comprising: partitioning the data elements of the graph G into a number of non-overlapping subgraphs; modifying each of the plurality of non-overlapping subgraphs to generate a plurality of p induced subgraphs; distributing each of the plurality of p induced subgraphs onto a separate server located in a public cloud environment; calculating the number of resulting triangles (p integers) associated with each of the plurality of p induced subgraphs of the distinct servers located in the public cloud environment; and calculating, via an on-premise server, a final number of triangles associated with the resulting number of triangles (p integer); A method comprising:

2. The method of claim 1 , wherein splitting comprises splitting the data elements responsive to a set of graph parameters.

3. The method of claim 2 , wherein the graph parameters include at least one of a clustering coefficient and a transitivity ratio.

4. The method of claim 1 , wherein distributing comprises expanding each of the plurality of non-overlapping subgraphs with at least one of a set of new vertices and random edge connections.

5. The method of claim 1 , further comprising transmitting the resulting number of triangles (p integer) to the on-premise server.

6. The final number of triangles is calculated as follows: adding the p integer to the number of triangles associated with the plurality of p induced subgraphs to generate a semi-final number of triangles; and subtracting from the number of semi-final triangles the number of triangles that have at least one added edge connection to a previously added random edge connection. The method of any one of claims 1 to 5, comprising:

7. 1. A method for protecting privacy by counting triangles in a graph of a hybrid cloud environment, the method comprising: partitioning the data elements of the graph G into a number of non-overlapping subgraphs; distributing each of the plurality of non-overlapping subgraphs onto a separate server located in a public cloud environment; augmenting each of the plurality of non-overlapping subgraphs with new vertices and random edge connections to generate p induced subgraphs; calculating a number of resulting triangles (p integers) associated with each of the plurality of non-overlapping subgraphs for each of the servers located in the public cloud environment; transmitting the p integer to an on-premise server; and calculating a final number of triangles associated with the subgraph formed by combining the p induced subgraphs via the on-premise server by adding p integers to the number of triangles associated with the subgraph formed by combining the p induced subgraphs and subtracting the number of triangles having at least one added edge connection to a previously added random edge connection; A method comprising:

8. The method of claim 7 , wherein splitting comprises splitting the data elements responsive to a set of graph parameters.

9. The method of claim 8 , wherein the graph parameters include at least one of a clustering coefficient and a transitivity ratio.

10. 10. The method of claim 7, wherein distributing comprises expanding each of the plurality of non-overlapping subgraphs with at least one of a set of new vertices and random edge connections.

11. 1. A machine learning system for implementing a method for preserving privacy by counting triangles in a graph G of a hybrid cloud environment, the machine learning system comprising: partitioning the data elements of the graph G into a number of non-overlapping subgraphs; modifying each of the plurality of non-overlapping subgraphs to generate a plurality of p induced subgraphs; distributing each of the plurality of p induced subgraphs onto a separate server located in a public cloud environment; Calculating the number of resulting triangles (p integers) associated with each of the plurality of p induced subgraphs of the distinct servers located in the public cloud environment; and Calculating, via an on-premise server, a final number of triangles associated with the resulting number of triangles (p integer). a machine learning system configured to: A computing system comprising:

12. The computing system of claim 11 , wherein partitioning comprises partitioning the data elements responsive to a set of graph parameters.

13. The computing system of claim 12 , wherein the graph parameters include at least one of a clustering coefficient and a transitivity ratio.

14. 12. The computing system of claim 11, wherein distributing comprises expanding each of the plurality of non-overlapping subgraphs with at least one of a set of new vertices and random edge connections.

15. The computing system of claim 11 , further comprising transmitting the resulting number of triangles (p integer) to the on-premise server.

16. Calculating the final number of triangles is adding the p integer to the number of triangles associated with the plurality of p induced subgraphs to generate a semi-final number of triangles; and subtracting from the number of semi-final triangles the number of triangles that have at least one added edge connection to a previously added random edge connection.

16. A computing system according to any one of claims 11 to 15, comprising:

17. 1. A computer program comprising program instructions executable by a processor to cause the processor to perform operations for implementing a method for preserving privacy by counting triangles in a graph G of a hybrid cloud environment, the method comprising: partitioning the data elements of the graph G into a number of non-overlapping subgraphs; modifying each of the plurality of non-overlapping subgraphs to generate a plurality of p induced subgraphs; distributing each of the plurality of p induced subgraphs onto a separate server located in a public cloud environment; calculating the number of resulting triangles (p integers) associated with each of the plurality of p induced subgraphs for each of the distinct servers located in the public cloud environment; and calculating, via an on-premise server, a final number of triangles associated with the resulting number of triangles (p integer); Equipped with The final number of triangles is calculated as follows: adding the p integer to a number of triangles associated with the plurality of p induced subgraphs to generate a semi-final number of triangles; and subtracting from the number of semi-final triangles the number of triangles that have at least one added edge connection to a previously added random edge connection. Including, a computer program.

18. 20. The computer program product of claim 17, wherein partitioning comprises partitioning the data elements responsive to a set of graph parameters.

19. 20. The computer program product of claim 18, wherein the graph parameters include at least one of a clustering coefficient and a transitivity ratio.

20. 20. The computer program product of claim 17, wherein distributing comprises expanding each of the plurality of non-overlapping subgraphs with at least one of a set of new vertices and random edge connections.

21. 21. The computer program product of claim 17, further comprising transmitting the resulting number of triangles (p integer) to the on-premise server.

22. A memory having computer-readable instructions for implementing a method for protecting privacy by counting triangles in a graph G of a hybrid cloud environment; and One or more processors for executing the computer readable instructions, the computer readable instructions comprising: partitioning the data elements of said graph G into a number of non-overlapping subgraphs; modifying each of the plurality of non-overlapping subgraphs to generate a plurality of p induced subgraphs; distributing each of the plurality of p induced subgraphs onto a separate server located in a public cloud environment; Calculating the number of resulting triangles (p integers) associated with each of the plurality of p induced subgraphs of the distinct servers located in the public cloud environment; and calculating, via an on-premise server, a final number of triangles associated with said resulting number of triangles (p integers); one or more processors to control the one or more processors to perform operations comprising: A system comprising:

23. 23. The system of claim 22, wherein partitioning comprises partitioning the data elements responsive to a set of graph parameters, where the graph parameters include at least one of a clustering coefficient and a transitivity ratio.

24. 23. The system of claim 22, wherein distributing comprises expanding each of the plurality of non-overlapping subgraphs with at least one of a set of new vertices and random edge connections.

25. The steps to calculate the final number of triangles are: adding the p integers to the triangle counts associated with the plurality of p induced subgraphs to generate a semi-final triangle count; and subtracting from the number of semi-final triangles the number of triangles that have at least one added edge connection to a previously added random edge connection.

25. The system of any one of claims 22 to 24, comprising: