Method and System for Privacy-Preserving Information Exchange

By using two-dimensional vectors and masking technology in electronic devices, the problem of protecting the privacy of local participants in the information aggregation of multiple participants is solved, and privacy protection information exchange is realized, allowing local participants to safely utilize other participants' data and computing resources.

CN115087994BActive Publication Date: 2025-06-13TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080096457.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-12
Filing Date
2020-04-16
Publication Date
2025-06-13
Estimated Expiration
2040-04-16

AI Technical Summary

Technical Problem

In machine learning and other applications, information aggregation of multiple participants requires protecting the privacy of local participants and preventing the central entity from identifying the original source of information.

Method used

By implementing a method for privacy-protecting information exchange in an electronic device, multiple values ​​are stored using two-dimensional vectors and transmitted by masking the aggregator to prevent the aggregator from decoding. Masked vectors are aggregated with other local parties, allowing the aggregation of vectors to be decoded.

Benefits of technology

The privacy-protected information exchange between local participants and aggregators is realized, allowing local participants to make full use of other participants’ data and aggregators’ computing resources without damaging their privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115087994B_ABST
    Figure CN115087994B_ABST
Patent Text Reader

Abstract

Methods and systems for privacy - protected information exchange in a network of electronic devices are disclosed. In one embodiment, a method for information exchange between a local participant and another electronic device acting as an aggregator is implemented in an electronic device serving as the local participant. The method includes storing a plurality of values in a 2D vector, where the first dimension of the 2D vector is based on the number of values and where each position in the first dimension has a unique value. The method further includes transmitting the 2D vector to the aggregator with a mask for the aggregator to prevent the aggregator from decoding the 2D vector, where aggregating the masked 2D vector with masked 2D vectors from other local participants allows decoding of the aggregated 2D vector.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to International Application No. PCT / IB2020 / 051159, filed on February 12, 2020, which is incorporated herein by reference. Technical field

[0003] Embodiments of the invention relate to the field of information sharing; more specifically, to privacy - protected information exchange in a network of electronic devices. Background art

[0004] In machine learning and other applications, the aggregation of information from multiple parties will make learning more efficient and / or accurate. For example, some parties may have collected information from their own subject - related sources (e.g., from their clients, their own research, operational data, etc.) and such parties can be referred to as "local parties". Aggregating information from all these parties in the aggregation will better characterize the subject. Preferably, there is a central entity to aggregate the information. However, the central entity (referred to as the "aggregator") will learn which originating local party provides which information when the local party transmits its information to the aggregator. The local party may prefer to protect its privacy while sharing its local information about an object with other local parties and the aggregator. Summary of the invention

[0005] Embodiments include a method for privacy - protected information exchange implemented in an electronic device. In one embodiment, a method for privacy - protected information exchange between an electronic device acting as a local party and another electronic device acting as an aggregator is implemented, where the aggregator exchanges information with a plurality of local parties including the local party. The method includes storing a plurality of values in a two - dimensional (2D) vector, where the first dimension of the 2D vector is based on the number of values, and where each position in the first dimension has a unique value within the plurality of values. The method further includes transmitting the 2D vector to the aggregator with a mask for the aggregator to prevent the aggregator from decoding the 2D vector, where aggregating the masked 2D vector with masked 2D vectors from other local parties allows decoding of the aggregated 2D vector.

[0006] Embodiments include an electronic device for privacy - protected information exchange. In one embodiment, the electronic device is to be used as a local participant for privacy - protected information exchange between the local participant and another electronic device acting as an aggregator, where the aggregator exchanges information with a plurality of local participants including the local participant. The electronic device includes a processor and a non - transitory machine - readable storage medium having stored instructions that, when executed by the processor, can cause the electronic device to perform storing a plurality of values in a two - dimensional (2D) vector, where the first dimension of the 2D vector is based on the number of values, and where each position in the first dimension has a unique value within the plurality of values. The instructions can further cause the electronic device to perform transmitting the 2D vector to the aggregator with masking of the aggregator to prevent the aggregator from decoding the 2D vector, where aggregating the masked 2D vector with masked 2D vectors from other local participants allows decoding of the aggregated 2D vector.

[0007] Embodiments include a non - transitory machine - readable storage medium for privacy - protected information exchange. In one embodiment, the non - transitory machine - readable storage medium has stored instructions that, when executed by a processor of an electronic device, can cause the electronic device to perform storing a plurality of values in a two - dimensional (2D) vector, where the first dimension of the 2D vector is based on the number of values, and where each position in the first dimension has a unique value within the plurality of values. The instructions can further cause the electronic device to perform transmitting the 2D vector to the aggregator with masking of the aggregator to prevent the aggregator from decoding the 2D vector, where aggregating the masked 2D vector with masked 2D vectors from other local participants allows decoding of the aggregated 2D vector.

[0008] These embodiments provide a set of data structures for privacy - protected information exchange between a local participant and an aggregator. The set of data structures with masking allows the local participant to transmit information to the aggregator without disclosing which local participant contributed what data, yet the aggregator can decode the aggregated data from the local participants and make determinations based on the aggregated data. Such privacy - protected information exchange allows local participants to fully utilize data from other local participants and the computational resources of the aggregator without sacrificing its privacy and has a wide range of applications such as machine learning and artificial intelligence. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The invention can be best understood by reference to the following description and the drawings used to illustrate embodiments of the invention. In the drawings:

[0010] Figure 1 A privacy - protected information exchange system according to some embodiments is illustrated.

[0011] Figure 2Describes unmasking through 2D vector aggregation according to some embodiments.

[0012] Figure 3 Describes the aggregation of 2D vectors without value conflicts according to some embodiments.

[0013] Figure 4A Describes the aggregation of 2D vectors with value conflicts according to some embodiments.

[0014] Figure 4B Describes that the retransmission of rows with value conflicts within a 2D vector according to some embodiments results in the resolution of value conflicts.

[0015] Figure 4C Describes value identification after more than one iteration in the first embodiment.

[0016] Figure 4D Describes value identification after more than one iteration in the second embodiment.

[0017] Figure 4E Describes value identification after more than one iteration in the third embodiment.

[0018] Figure 4F Describes value identification after more than one iteration in the fourth embodiment.

[0019] Figure 5 Describes tree generation according to some embodiments.

[0020] Figure 6 Describes the operation of determining split point values based on split point value candidates from multiple local participants according to some embodiments.

[0021] Figure 7 Is a flowchart showing the operations for privacy - protected information exchange between a local participant and another electronic device acting as an aggregator for an electronic device acting as a local participant according to some embodiments.

[0022] Figure 8A -B is a flowchart showing the operations for privacy - protected information exchange between multiple electronic devices each acting as a local participant for an electronic device acting as an aggregator according to some embodiments.

[0023] Figure 9 Describes the connectivity between network devices (NDs) within an exemplary network and three exemplary implementations of NDs according to some embodiments. Detailed Description

[0024] The following description describes methods and apparatus for privacy-preserving information exchange in a network of electronic devices. In the following description, numerous specific details are set forth such as logic implementations, resource partitioning / sharing / copying implementations, types and interrelationships of system components, and logic partitioning / integration choices in order to provide a more thorough understanding of the present invention. However, one of ordinary skill in the art will realize that the invention may be practiced without such specific details. In other instances, control structures, gate-level circuits, and full software instruction sequences have not been shown in detail in order not to obscure the invention. One of ordinary skill in the art will be able to implement appropriate functionality using the included description without undue experimentation.

[0025] Privacy-protected information exchange network

[0026] Figure 1 A privacy-preserving information exchange system in accordance with some embodiments is illustrated. System (also referred to as a network) 100 includes a plurality of local participants 102-106 and an aggregator 112, with each local participant and the aggregator implemented in an electronic device (defined below herein). The local participants and the aggregator communicate via a communication network 190 (the communication network is discussed in more detail below herein).

[0027] At tag 152, each local participant conveys a vector to aggregator 112 using a mask to prevent the aggregator from identifying the original local participant of the value. Each vector includes a portion of the total value to be aggregated at aggregator 112. In one embodiment, as shown at tag 154, the vector is two-dimensional (2D). In one embodiment, the 2D vector includes a first dimension 132 based on the number of values (also referred to as "local values") at the local participant and a second dimension 134 based on the number of local participants.

[0028] For simplicity of explanation, the first dimension is referred to as the rows of the 2D vector and the second dimension is referred to as the columns of the 2D vector. Clearly, embodiments may include reversed row and column designations. Additionally, so long as the values are included within the 2D vector, the values may be conveyed in higher-dimensional vectors (e.g., 3D or higher).

[0029] In one embodiment, the number of rows is equal to the number of local values to be aggregated and the number of columns is not less than the number of local parties. In this way, each local value can take a row and select one of the columns of the row. In one embodiment, a row of the 2D vector can be randomly selected for a local value and a column within the row can also be randomly selected for the local value. In an alternative embodiment, another selection strategy is used to select either a row or a column. For example, rows can be selected based on the values - the lowest value takes row 1, the second lowest value takes row 2, and so on (or vice versa), and columns can be selected such that the local value from the first local party (assuming each local party is indexed to a number and sorted based on the index) takes the first column, the local value from the second local party takes the second column, and so on.

[0030] In some embodiments, the 2D vectors of each local party have the same size, and each local party has the same number of local values to be aggregated. Alternatively, depending on the number of local values and the way the column size is determined, the 2D vectors from local parties can have different sizes.

[0031] Note that when the size of the rows is large and rows are randomly selected for values, the chance that two local parties select the same row and the same column (referred to as a value conflict) is reduced, so a larger number of rows of the 2D vector reduces the chance of value conflicts. When aggregating the values from local parties at the aggregator, value conflicts make the value aggregation ambiguous, so the aggregator can request the local parties to retransmit the conflicting values. Because of this, a large number of rows can be selected for the 2D vector to reduce retransmissions. For example, in some embodiments, the size of the rows is not less than a multiple of the number of local values.

[0032] In some embodiments, when the number of local parties is large, the local parties can be divided into subgroups, each subgroup including a subset of the local parties. In that case, the size of the 2D vector is based on the size of the subset of local parties and the values to be aggregated within the subset of local parties. For example, the number of rows can be equal to the number of local values of the subgroup multiplied by the number of subgroups to be aggregated and the number of columns is not less than the number of local parties in the subgroup.

[0033] In some embodiments (e.g., when the number of local parties is large), the number of columns can be fixed. For example, we have m parties that want to share n values, rather than sharing a 2D matrix with m rows and ≥ n (e.g., order n 2 ) columns, having m × n / s rows and ≥ s columns (e.g., order s 2A 2D array of ( ) might be better, where s is the number of local parties sharing any given row. s Must be an integer > 1. The higher it is, the more information the local parties must exchange outside the protocol with the aggregator to break confidentiality, and the more expensive the transmission becomes. Each local party has an impact on those m × n / s rows in m and must know who else has an impact on those rows to use only the masks they share with them rather than all the masks. Thus, each party contributes m a row array, and then the aggregator maps them to m × n / s a row array and adds them together. As n grows, the number of columns can then remain stable, which allows the method to still work with a large number of parties.

[0034] In some embodiments, the number of rows can be based on both the number of values and the number of local parties, while the number of columns can be independent of both and be a fixed number. Additionally, it is possible for the local parties to send only some of the rows (although they need to know who else will be sending these rows to apply the appropriate masks), which reduces the amount of communication required.

[0035] In some embodiments, the local parties can be broken into subgroups and different parts of the 2D array assigned to different subgroups. In this way, the size of the second dimension can be controlled at the cost of expanding the first dimension.

[0036] For example, if an embodiment is to have 20 local parties and 50 split point candidate values, and if the number of columns is based on the number of local parties, the size of the 2D vector can be 50× n , where n ≥ 50. If, for example n = 100, the size of the 2D vector can be 50×100 = 5,000. In another embodiment, the number of columns can be fixed. For example, if an embodiment is to have 20 local parties and 50 desired split point candidate values, the group of 20 local parties can, for example, be split into 4 subgroups of 5 local parties each. Each row can have a fixed number of columns. For example, each row can have 10 (> 5, the number of local parties in a subgroup) columns. Each of the exemplary 4 subgroups can have 50 different rows, each of which can be used for split point candidate values, such that the size of the 2D vector will then be reduced to 4×50×10 = 2,000. Such a reduction results in less bandwidth consumption between the local parties and the aggregator and less computation at the local parties / aggregator.

[0037] Each 2D vector from a local party is masked to prevent the aggregator 112 from identifying the original local party of the value. The masked 2D vectors are transmitted to the aggregator 112 via the communication network 190. Since the masking prevents the aggregator 112 from identifying the original local party of the value, the aggregation as shown at label 156 is a secure aggregation from local parties. At label 158, the aggregator aggregates the masked 2D vectors from other local parties, and the aggregation allows the decoding of the aggregated 2D vectors. That is, although each mask prevents the aggregator 112 from identifying the original local party of the value, the aggregation of the masked vectors allows the aggregator to obtain the aggregated value without such identification. In this way, the aggregator obtains the values from the local parties (only in the aggregation without knowing the contribution of each party) and the local parties protect their privacy, thus achieving a privacy-preserving information exchange between the local parties and the aggregator.

[0038] Masking vector and unmasked aggregated vector

[0039] In a privacy-preserving information exchange, the values from local parties are masked to the aggregator such that the aggregator cannot decode the values on its own, yet the masking allows the aggregation of values from multiple local parties to be decoded. Many ways can achieve such masking and unmasking. For example, a privacy-preserving machine learning mechanism is disclosed in "Practical Secure Aggregation for Privacy-Preserving Machine Learning" by Bonawitz et al. (hereinafter "Bonawitz") published in 2017, which is incorporated herein by reference.

[0040] To briefly explain, the privacy-preserving information exchange uses masking at the local parties and unmasking at the aggregator by aggregating the masked local data. Each local party knows its own set of cryptographic keys (referred to as masks) used for masking and neither other local parties nor the aggregator knows the set of cryptographic keys, such that once the values are masked using the set of cryptographic keys (e.g., by encryption using the set of cryptographic keys), other local parties and the aggregator cannot decode the values. However, the masks are designed such that the aggregation of the masks cancels the masks, such that the aggregation of the masked values returns the aggregation of the values before masking.

[0041] Such masking and unmasking can be applied to the aggregation of 2D vectors from local parties. Figure 2 Illustrates the unmasking by 2D vector aggregation according to some embodiments. Each local party sends a masked 2D vector, which is shown as two components, such as the vector itself shown at label 252 ( x a tox d ), and its corresponding mask as shown at label 254. By applying the mask, no party other than the sending local party (neither the aggregator nor other local parties) knows the values within the vector it sends. However, the aggregation of the masked vectors cancels out the effects of the individual masks, and as shown at label 256, the aggregation of the masked vectors at the aggregator results in an unmasked vector aggregation. Note that while encryption is shown as a simple addition of masks, more complex encryption can be used, including different sets of symmetric and asymmetric keys.

[0042] Value conflict and resolution

[0043] Figure 3 Illustrates the aggregation of 2D vectors without value conflicts according to some embodiments. As explained above herein, local parties can randomly store values in 2D vectors. As shown at labels 302 to 306, the 2D vectors S 1 , S 2 and S 3 at parties A to C have the same dimension 5×4, where the number of columns (5) is greater than the number of local parties (3) and the number of rows (4) is equal to the number of local values. Each value takes a row (i.e., each row has a unique value) and the values randomly take columns within the row. Randomization of the column positions of the local values can be generated by a random number generator, a quasi-random number generator, or a pseudo-random number generator.

[0044] In this example, the 2D vectors from parties A to C are individually masked and then aggregated at the aggregator. As discussed above herein, the aggregation of the masked vectors cancels out the effects of the individual masks, and the aggregation results in an aggregated 2D vector that, as shown at label 310, does not have value conflicts.

[0045] Figure 4A Illustrates the aggregation of 2D vectors with value conflicts according to some embodiments. As shown at labels 402 to 406, the 2D vectors S 1 , S 2 and S 3 at parties A to C have the same dimension 5×4 and the same values, but the difference is that the randomization of column selection results in some values being stored in columns different from those at labels 302 to 306. In this example, as shown at labels 404 and 406, the value d 2 in S 2 and the value b 3 in S 2 are moved to new column positions.

[0046] Because the aggregator obtains the aggregation of the values of the aggregated 2D vectors without knowing which value comes from which local party, the aggregator cannot determine value conflicts based on the row positions of each incoming 2D vector. Instead, the aggregator can detect value conflicts in the aggregated 2D vectors by counting the number of non-zero values in each row. Since the system has three local parties, each row will have three non-zero values when each local party has a unique value in the row of the aggregated 2D vector. In this case, each of rows 2 and 3 of the aggregated 2D vector has three non-zero values, so the aggregator determines that rows 2 and 3 of the aggregated 2D vector do not have valid conflicts. Instead, each of rows 1 and 4 of the aggregated 2D vector has only two non-zero values, so the aggregator determines that these rows have value conflicts. The aggregator then requests retransmission of rows 1 and 4 of the 2D vectors from all local parties at labels 452 and 454 respectively (e.g., including the row ID of the row to be retransmitted in the request to the local parties). Moreover, because the aggregator does not know which value comes from which local party, it cannot determine the local party that caused the value conflict, so it requests retransmission of the conflicting rows from all local parties.

[0047] Figure 4B Illustrated is that retransmission of rows with value conflicts within a 2D vector according to some embodiments results in resolution of the value conflicts. Once the local parties are notified to retransmit the rows with value conflicts, it randomizes the column positions of the values to be retransmitted. In this example, as shown at labels 462 to 476, the randomization at each party results in a change in the column position of each value in rows 1 and 4. The updated rows are then transmitted from all local parties to the aggregator and result in an updated 2D vector without value conflicts as shown at label 482.

[0048] Although Figure 4B it is shown that one set of retransmissions resolves the value conflicts for all values, sometimes multiple iterations may be necessary to resolve the value conflicts for all values. In that case, the aggregator will repeat the detection of value conflicts in the aggregated 2D vector after retransmission by counting the number of non-zero values in each row and request the local parties to retransmit the conflicting row(s). The local parties will perform the retransmission based on the new request and mask the retransmitted values. The process continues until all value conflicts within the aggregated 2D vector at the aggregator are resolved.

[0049] Note that in Figure 4B the embodiments, each retransmitted masked vector is shown as a one-dimensional (1D) vector. In some embodiments, multiple retransmitted values from local parties may form a 2D vector. For example, instead of sending two 5×1 1D vectors at labels 462 and 472, Party A may send one 5×2 2D vector storing the same values.

[0050] In some embodiments, instead of the aggregator requesting all local participants to retransmit rows with conflicting values until no value conflicts are detected, the aggregator can reduce the number of retransmissions. When a conflicting row is detected, the aggregator knows how many values have conflicted. For example, if there are m values detected in a row and assuming that n is n-m , then there are conflicting values (which are incorrect) between 1 and 2m-1 and valid values between m-1 and

[0051] After more than one iteration (e.g., 2 iterations), if we assume that it is impossible for any value to be equal to the sum of a set of other values, the aggregator can identify some or all of those values (which should be the case if the values are given with high precision).

[0052] Figure 4C illustrates value identification after more than one iteration in the first embodiment. As shown at reference numeral 412, five participants each randomly transmit the values in their rows in iteration i (e.g., the first iteration) and that results in the aggregator receiving aggregated values with one conflicting value and three valid values. The aggregator expects 5 values but receives 4 values as shown at reference numeral 414. The aggregator requests the local participants to retransmit. The retransmission in the next iteration (iteration i + 1) again results in 4 values as shown at reference numeral 414. The aggregator determines that one value ( b + d ) is the sum of the first and third values in the previous iteration, so it can determine that values b and d are valid values and, because of this, the remaining values a , e and c are also valid since 4 values were received.

[0053] Figure 4D illustrates value identification after more than one iteration in the second embodiment. The result of iteration i is the same as Figure 4C and it is shown at reference numeral 412. In the retransmission of the next iteration, 4 values are received instead of the expected 5 as shown at reference numeral 416. In this iteration, the aggregator determines that values b and e are from the previous iteration and, since 4 values were received and there is only one conflict in the iteration, the valuesb and e are valid, and the aggregator can request that the corresponding sending local participants B and E do not retransmit, and only the remaining local participants need to retransmit (again using the randomized positions).

[0054] Figure 4E Illustrates value identification after more than one iteration in the third embodiment. Iteration i The result of Figure 4C is the same and it is shown at marker 412. In the retransmission of the next iteration, 3 values are received instead of the expected 5. The aggregator determines the values e and a + c are from the previous iteration and the value ( b + d ) is the sum of two values from the previous iteration. Thus, the aggregator determines that the values b and d are valid values and retransmission for them is unnecessary, and the aggregator can inform the corresponding local participants to so indicate.

[0055] In Figure 4D and 4E example, even if the aggregator does not completely eliminate more iterations of local participant retransmissions, by determining which participant (or which participants) no longer need to retransmit, fewer participants will perform retransmissions in future iterations, thereby reducing the bandwidth consumption between the local participants and the aggregator and the computation at the local participant / aggregator.

[0056] Figure 4F Illustrates value identification after more than one iteration in the fourth embodiment. Iteration i The result of Figure 4C is the same and it is shown at marker 412. In the retransmission of the next iteration, 4 values are received instead of the expected 5. In this iteration, all the values at marker 419 were previously seen, and the aggregator thus determines that the same conflict has occurred and the values cannot be determined by the iteration, and the aggregator requests all local participants to retransmit.

[0057] Using the values aggregated from all local participants, an aggregator can obtain values from all local participants without knowing which local participants the values originated from. Compared with the computing resources of individual local participants, the aggregator can have superior computing resources and can use all the data from all local participants to make better / faster decisions. Thus, each local participant can make full use of the data from other local participants and the computing resources of the aggregator without compromising its privacy, and such advantages are useful in many applications. For example, privacy-preserving information exchange can be used in machine learning and artificial intelligence in networking systems that combine lists of objects (such as names / identifiers, values / variables) without revealing who contributed what or in messaging systems that aggregate anonymous messages. For example, a computer of an organization can send such an array at a predetermined frequency. It is typically a mask on an empty array, but when someone has an anonymous message to send, it will be at a random position in the array.

[0058] Exemplary Application: Machine Learning

[0059] One application of using privacy-preserving information exchange is in machine learning, particularly in the training of machine learning models such as gradient boosting. XGBoost is an example of a gradient boosting technique that has gained traction. For example, XGBoost was disclosed in "XGBoost: A Scalable Tree Boosting System" by Chen et al. (hereinafter "Chen") published in 2016, which is incorporated herein by reference.

[0060] The basic idea of gradient boosting trees is to generate an ensemble of decision trees that overall comprise a model for a regression or classification problem. Then the predictions of each tree are added together, and the sum is the prediction of the model. The performance of the model is measured by a given loss function. The loss function is a measure of the actual value of the data and the predicted value of the data. Additionally, a regularization function that is a function of the number of leaf nodes in the ensemble and the weights of the leaf nodes in the ensemble can be used. In this case, the model is trained using a regularization objective that is the sum of the loss function and the regularization function.

[0061] In XGBoost, the model is trained one tree at a time in an additive manner. After t-1 trees have been trained, the algorithm trains the tree according to the following objective t :

[0062] (1)

[0063] In formula (1), g i is at the data point iThe first derivative of the loss function with respect to the prediction, evaluated at the predicted value f t (x i ) is the prediction of the tree i on the data point t and h i is the second derivative of the loss function with respect to the prediction, evaluated at the predicted value of the data point i and Ω(f t ) is the value of the regularization function applied to the tree t . The tree is generated in a greedy manner t where, for each feature, the training model evaluates multiple split point value candidates and the loss reduction at each split point value candidate is given by

[0064] (2)

[0065] In formula (2), I L is the set of all data points to the left of the split point value candidate, I R is the set of all data points to the right of the split point value candidate, and λ and γ are the parameters of the regularization function

[0066] An example of the regularization function applied to the tree t is as follows

[0067] (3)

[0068] In formula (3), T is the number of leaf nodes in the tree and ||w|| 2 is the sum of the squares of the weights of the leaf nodes

[0069] For datasets of sufficiently large size, testing every single split point value candidate for each feature of the tree can become computationally infeasible. Therefore, the XGBoost algorithm considers a search over a subset of the split point value candidates for each feature. This subset of split point value candidates is described by a data structure called a weighted quantile sketch, which consists of a controllable number of points k and approximately describes the k - quantile split distribution of the data, where each point i has a weight i that can be determined by the second derivative of the loss function at the current prediction of the point i at the point W i .

[0070] The weighted quantile sketch Q includes the following components: (1) S = the set of values in the sketch; (2) x = the weight of each w value; (3) x = the rank-decreasing function, substantially the sum of the weights of the values < ; and (4) y = the rank-increasing function, substantially the sum of the weights of the values ≤ y .

[0071] The rank function can be estimated for values not in the sketch by interpolating with the ranks and weight values of the points closest to the desired value, i.e., if , then:

[0072] (4).

[0073] Thus, to test a split point value candidate, the required data includes (i) the split point value candidate x ; (ii) the weight w for the split point value candidate; (iii) the ranks determined by the rank-decreasing and rank-increasing functions; (iv) the value of the first derivative of the loss function (see, e.g., formula (1)); and (v) the value of the second derivative of the loss function (see, e.g., formula (1)). As shown in formula (1), the values based on the first and second derivatives can be a set of sums of the derivatives of the loss function for a decision tree, where in some embodiments, each sum aggregates the values between adjacent split point value candidates. In alternative embodiments, different formulas can be used, applying the first derivative and / or the second derivative of the loss function to derive the values in (iv) and (v).

[0074] Figure 5Describes tree generation according to some embodiments. Decision tree 510 includes a plurality of nodes, each node splitting according to a feature. For example, node 512 maps to the feature of age and node 514 maps to the feature of gender. Each node maps to a feature, but a feature can be mapped to multiple nodes - for example, the feature of age can be mapped to a node for ages greater than a certain age and another node for ages younger than another age. For node 512, multiple split point value candidates 502 can be tested, including split point value candidates 15, 18, 21, and 22. After machine learning training, a single split point value 18 is selected for the feature at node 512. Machine learning training can use privacy-preserving information exchange. For example, masked 2D vectors can be used to transmit the split value candidates from multiple local parties to an aggregator, and the aggregator aggregates the masked 2D vectors to decode the aggregated 2D vectors using the aggregated values and extract the aggregated values in each position of the aggregated 2D vectors without identifying the local party from which the values originated. From the aggregated values, the aggregator can determine that the split point value for the feature is 18.

[0075] Figure 6 Describes the operation of determining a split point value based on split point value candidates from multiple local parties according to some embodiments. As shown, the system includes local parties 602 and 604 and an aggregator 612. Obviously, the system can include a large number of local parties, and the description of two parties is for ease of explanation.

[0076] At markings 622 and 632, local parties 602 and 604 store the split point value candidates for the feature into their respective 2D vectors, each local party having one 2D vector and the determination of the dimensions of the 2D vector being explained above in this document in relation to Figure 1 .

[0077] At markings 662 and 672, local parties 602 and 604 transmit their respective masked 2D vectors to aggregator 612. At marking 652, aggregator 612 aggregates the masked 2D vectors to unmask the aggregated values without identifying the local party from which the values originated. Masking, unmasking, and aggregation of values are explained above in this document in relation to Figure 2 and Figure 3 .

[0078] Optionally, a value conflict is detected in the aggregated 2D vector, and the aggregator 612 identifies the value conflict at tag 653. At tags 682 and 683 the aggregator 612 then requests the local participants 602 and 604 to retransmit the conflicting values (e.g., by identifying the (one or more) conflicting rows). At tags 664 and 674 the local participants 602 and 604 then each retransmit a masked vector of the values with the (one or more) earlier conflicts. The positions of the retransmitted values in the vector can be randomized. Value conflicts and retransmissions were explained above in this document in relation to Figure 4A -B.

[0079] Once the aggregator 612 has received all the split point value candidates of the features from all the local participants, at tags 684 and 685 the aggregator 612 transmits all the aggregated split point value candidates to all the local participants. As shown at tags 666 and 676, using masking, each local participant then transmits the quantile sketch information of all the split point value candidates it has to the aggregator 612. Once the aggregator 612 has received the masked quantile sketch information, at tag 656 it aggregates them to unmask the quantile sketch information from the local participants. Then at tag 658 the aggregator 612 determines the split point value based on the quantile sketch information from the local participants. In one embodiment, the determined split point value is a single value from all the split point value candidates.

[0080] Regarding the quantile sketch information, as explained above in this document in relation to formulas (1) to (4), in addition to the split point value candidates themselves (item (i) for testing split point value candidates explained above), the quantile sketch information can additionally include at least one of the following: (1) weights and / or ranks (items (ii) and (iii) for testing split point value candidates explained above) and (2) values based on the first derivative and / or second derivative of the loss function (items (iv) and (v) for testing split point value candidates explained above). Although in some embodiments the quantile sketch information such as (1) and (2) is transmitted together from the local participants, in other embodiments, only (1) or (2) is needed for the determination of the split point value, in which case only (1) or (2) is transmitted to the aggregator 612.

[0081] Additionally, after receiving some of the quantile sketch information, the aggregator can decide to prune the entire list of split point value candidates based on the quantile sketch information. In that case, the aggregator can send the reduced list of split point value candidates to the local participants after pruning, and the local participants will send additional quantile sketch information only for the remaining split point value candidates.

[0082] Using Figure 3As an example, aggregator 612 will send non-zero values in the aggregated 2D vectors to all local parties (e.g., operations at tags 684 and 685), and in S return ={ a 1 ,b 1 ,c 1 ,d 1 ,a 2 ,b 2 ,c 2 ,d 2 ,a 3 ,b 3 ,c 3 , d 3} includes non-zero values. The local parties use masking to provide the initial quantile sketch information about these split point candidate values back to the aggregator, e.g., (1) for the weights and / or ranks as discussed above herein for S return . The aggregator obtains the initial quantile sketch information and determines that a subset of S return are feasible split point candidate values, so it trims the list into a smaller list, e.g., S’ return ={ a 1 ,b 1 ,c 1 ,b 2 ,c 2 ,d 2 ,c 3 ,d 3}. The aggregator will send the smaller list to all local parties, and the local parties can use masking to provide additional quantile sketch information about these remaining split point candidate values back to the aggregator, e.g., (2) for the S’ returnThe values of the first derivative and / or the second derivative of the loss function. By partitioning the quantile sketch information into a first batch (i.e., initial quantile sketch information) and a second batch (i.e., additional quantile sketch information) such that the latter is only used for the list of split point value candidates pruning, the system reduces (1) the bandwidth consumption between the local parties and the aggregator and / or (2) the computing resources of the local parties and the aggregator (e.g., the local parties do not need to compute the additional quantile sketch information for the split point value candidates pruned by the aggregator based on the first batch and the aggregator does not perform additional computations for these removed split point candidates). Although in this embodiment, the initial quantile sketch information is (1) the weights and / or ranks discussed above herein and the additional quantile sketch information is (2) the values of the first derivative and / or the second derivative of the loss function discussed above herein, other embodiments may reverse the order such that the information in (2) is the initial quantile sketch information and the information in (1) is the additional quantile sketch information.

[0083] By operations related to Figure 6 the system protects the privacy of the local parties while using the aggregated information from the local parties to generate a decision tree model, enabling machine learning to be successfully performed without compromising the privacy of the local parties.

[0084] Some embodiments

[0085] Figure 7 is a flowchart showing operations for privacy-preserving information exchange between an electronic device serving as a local party and another electronic device serving as an aggregator according to some embodiments. The electronic device may be one of the network devices 902 to 906 discussed below herein.

[0086] At block 702, a plurality of values are stored in a two-dimensional (2D) vector, where the first dimension of the 2D vector is based on the number of values, and where each position in the first dimension has a unique value within the plurality of values. In some embodiments, and the second dimension of the 2D vector is based on the number of local parties. The determination of the dimensions of the 2D vector is explained above herein in relation to Figure 1 For example, in some embodiments, the first dimension of the 2D vector is equal to the number of the first plurality of split point value candidates and the second dimension of the 2D vector is not less than the number of local parties.

[0087] At block 704, the 2D vector is transmitted to the aggregator with masking for the aggregator to prevent the aggregator from decoding the 2D vector, where aggregating the masked 2D vector with the masked 2D vectors from other local parties allows decoding of the aggregated 2D vector. Above herein in relation to Figure 2 and Figure 3Explained herein is masking and unmasking and aggregation of values.

[0088] In some embodiments, the exchanged information is decision tree information for decision tree learning. The aggregator is to generate a decision tree, where a plurality of values are first plurality of split point value candidates for at least one feature of the decision tree, and where the aggregator is to determine a single split point value for a node of the decision tree based on the aggregated 2D vector. In some embodiments, each of the plurality of split point value candidates is mapped to a sketch of data for the feature at a local party.

[0089] Value conflicts can be detected in the aggregated 2D vector, in which case, at label 706, the local party retransmits one or more values in response to a request from the aggregator, each of the values being stored in a randomized location within another vector, where each retransmission uses a mask to the aggregator to prevent the aggregator from decoding the other vector, and where aggregating the masked vectors with masked vectors from other local parties allows decoding of the aggregated vector. Effective conflicts and retransmissions were discussed hereinabove in connection with Figure 4A -B. Note that the other vector can be a 1D or 2D vector as discussed hereinabove in connection with Figure 4B as discussed.

[0090] Additionally, optionally at label 708, a second plurality of split point value candidates are received from the aggregator, and at label 710, the local party transmits, using a mask, quantile sketch information for the second plurality of split point value candidates mapped to the feature to the aggregator to prevent the aggregator from decoding the quantile sketch information, where aggregating the masked quantile sketch information with quantile sketch information from other local parties allows decoding of the aggregated quantile sketch information. The second plurality of split point value candidates can be all split point value candidates for the feature.

[0091] In some embodiments, the transmitted quantile sketch information can include all quantile sketch information for the second plurality of split point value candidates. In alternative embodiments, the transmitted quantile sketch information can include only the initial quantile sketch information as discussed hereinabove in connection with Figure 6 as discussed, in which case the aggregator can perform pruning, and at label 712, the local party receives a third plurality of split point value candidates from the aggregator. In some embodiments, the third plurality of split point value candidates is a subset of the second plurality of split point value candidates.

[0092] Then at marker 714, the local party transmits additional quantile sketch information for the third plurality of split point value candidates that map to features to the aggregator using masking to prevent the aggregator from decoding the additional quantile sketch information, and wherein aggregating the masked additional quantile sketch information with additional quantile sketch information from other local parties allows decoding of the aggregated additional quantile sketch information, and wherein the additional quantile sketch information is based on the derivative of a loss function for a decision tree. The operation of the transmission of quantile sketch information (such as initial quantile sketch information and additional quantile sketch information) was discussed above in this document in relation to Figure 6 the operation of the transmission of quantile sketch information (such as initial quantile sketch information and additional quantile sketch information) was discussed above in this document in relation to

[0093] Figure 8A -B is a flowchart showing the operation of an electronic device serving as an aggregator for privacy-preserving information exchange among a plurality of electronic devices each serving as a local party. The electronic device can be one of network devices 902 to 906 discussed below in this document.

[0094] At Figure 8A marker 802, the aggregator receives a plurality of two-dimensional (2D) vectors, each two-dimensional vector from one of a plurality of local parties, having the same first and second dimensions, and each two-dimensional vector containing a plurality of values, wherein each value from a local party takes a unique position in the first dimension of the 2D vector from the local party, and wherein masking is applied to each 2D vector to prevent the aggregator from decoding the 2D vector. The dimensions of the 2D vector were explained above in this document in relation to Figure 1 the dimensions of the 2D vector were explained above in this document in relation to. For example, in some embodiments, the first dimension of the 2D vector is equal to the number of the first plurality of split point value candidates, and the second dimension of the 2D vector is not less than the number of local parties.

[0095] At marker 804, the 2D vectors are aggregated, wherein aggregation of the 2D vectors allows decoding of the aggregated 2D vector and extraction of the aggregated values at each position of the aggregated 2D vector without identifying the local party from which the values originated.

[0096] In some embodiments, the information exchanged is decision tree information for decision tree learning. The aggregator is to generate a decision tree, wherein the plurality of values are the first plurality of split point value candidates for at least one feature of the decision tree, and wherein the aggregator is to determine a single split point value for a node of the decision tree based on the aggregated 2D vector. In some embodiments, each of the plurality of split point value candidates maps to a sketch of data for a feature at the local party.

[0097] Value conflicts can be detected in the aggregated 2D vectors and the process goes to label 806, where the aggregator identifies one or more positions in the aggregated 2D vectors through which at least two local parties have transmitted their values. Then, at label 808, the aggregator requests the local parties to retransmit the identified values (e.g., using the 1D or 2D vectors discussed above herein in connection with Figure 4B ).

[0098] Optionally, the process goes to label 810, where the aggregator sends a second plurality of split point value candidates to each local party, and where the number of the second plurality of split point value candidates is the sum of all split point value candidates of the features from the plurality of local parties.

[0099] At label 812, the aggregator receives from the local parties the quantile sketch information mapped to the second plurality of split point value candidates of the feature, where masking is applied to each quantile sketch information to prevent the aggregator from decoding the quantile sketch information. At label 814, the quantile sketch information is aggregated, where the aggregation of the quantile sketch information allows decoding of the aggregated quantile sketch information and extraction of the aggregated quantile sketch information without identifying the local party from which the aggregated quantile sketch information originated.

[0100] In some embodiments, the transmitted quantile sketch information may include all quantile sketch information regarding the second plurality of split point value candidates. In alternative embodiments, the transmitted quantile sketch information may only include the initial quantile sketch information discussed above herein in connection with Figure 6 . In that case, the aggregator may perform pruning at label 816 of Figure 8B , where the aggregator selects a subset of split point value candidates from the second plurality of split point value candidates to become a third plurality of split point value candidates, where the selection is based on the aggregated quantile sketch information. Then at label 818, the aggregator sends the third plurality of split point value candidates to each local party.

[0101] Then at label 820, the aggregator receives, with masking, additional quantile sketch information mapped to the third plurality of split point value candidates of the feature to the aggregator, where the additional quantile sketch information is based on the derivative of the loss function for the decision tree. At label 822, the aggregator determines a single split point value for a node based on the additional quantile sketch information.

[0102] Network environment in which embodiments can operate

[0103] Figure 9 Illustrates the connectivity between network devices (NDs) within an exemplary network and three exemplary implementations of an ND according to some embodiments. Figure 9Shows ND 900A-H and their connectivity in a manner that includes the lines between 900A-900B, 900B-900C, 900C-900D, 900D-900E, 900E-900F, 900F-900G, and 900A-900G, as well as the lines between 900H and each of 900A, 900C, 900D, and 900G. These NDs are physical devices, and the connectivity between these NDs can be wireless or wired (often referred to as a link). Additional lines extending from NDs 900A, 900E, and 900F illustrate that these NDs act as entry and exit points of the network (and thus, these NDs are sometimes referred to as edge NDs; while other NDs can be referred to as core NDs).

[0104] Figure 9 Two of the exemplary ND implementations are: 1) a dedicated network device 902 that uses a custom application-specific integrated circuit (ASIC) and a dedicated operating system (OS); and 2) a general-purpose network device 904 that uses an off-the-shelf (COTS) processor and a standard OS.

[0105] The dedicated network device 902 includes networking hardware 910, which includes one or more processors 912, (one or more) forwarding resources 914 (which typically include one or more ASICs and / or network processors), and physical network interfaces (NIs) 916 (through which network connections are made, such as those shown by the connectivity between ND 900A-H), and a collection of non-transitory machine-readable storage media 918 in which networking software 920 is stored. During operation, the networking software 920 can be executed by the networking hardware 910 to instantiate a collection of one or more networking software instances 922. Each of the (one or more) networking software instances 922 and the portion of the networking hardware 910 that executes that networking software instance (which is a time slice of the hardware dedicated to that networking software instance and / or shared with other networking software instances in the (one or more) networking software instances 922 over time) form separate virtual network elements 930A-R. Each of the (one or more) virtual network elements (VNEs) 930A-R includes control communication and configuration modules 932A-R (sometimes referred to as local control modules or control communication modules) and (one or more) forwarding tables 934A-R, such that a given virtual network element (e.g., 930A) includes a control communication and configuration module (e.g., 932A), a collection of one or more forwarding tables (e.g., 934A), and the portion of the networking hardware 910 that executes the virtual network element (e.g., 930A). In one embodiment, the networking software 920 includes a federated learning coordinator 928. The federated learning coordinator 928 can perform the operations described with reference to earlier figures. The federated learning coordinator 928 can generate one or more federated learning coordinator instances 953, each for a virtual network element (e.g., a virtual switch). The federated learning coordinator 928 can be implemented in the local participants or aggregators discussed above in this document. When implemented in a local participant, it performs local participant operations (e.g., operations related to Figure 7 ); and when implemented in an aggregator, it performs aggregator operations (e.g., operations related to Figure 8A -B).

[0106] The dedicated network device 902 is often physically and / or logically considered to include: 1) an ND control plane 924 (sometimes referred to as the control plane) including one or more processors 912 that execute one or more control communication and configuration modules 932A-R; and 2) an ND forwarding plane 926 (sometimes referred to as the forwarding plane, data plane, or media plane) including one or more forwarding resources 914 and a physical NI 916 that utilize one or more forwarding tables 934A-R. As an example, where ND is a router (or implementing routing functionality), the ND control plane 924 (one or more processors 912 that execute one or more control communication and configuration modules 932A-R) is typically responsible for participating in controlling how data (e.g., packets) are to be routed (e.g., the next hop for the data and the outbound physical NI for that data) and storing that routing information in one or more forwarding tables 934A-R, and the ND forwarding plane 926 is responsible for receiving that data on the physical NI 916 and forwarding that data out to the appropriate physical NI in the physical NI 916 based on one or more forwarding tables 934A-R.

[0107] The general network device 904 includes hardware 940, which includes a collection of one or more processors 942 (which are often COTS processors), a physical NI 946, and a non-transitory machine-readable storage medium 948 in which software 950 is stored. During operation, the (one or more) processors 942 execute the software 950 to instantiate one or more collections of one or more applications 964A-R. Although one embodiment does not implement virtualization, alternative embodiments may use different forms of virtualization. For example, in one such alternative embodiment, the virtualization layer 954 represents the kernel of the operating system (or a shim executing on top of the base operating system) that takes into account the creation of multiple instances 962A-R of what are referred to as software containers, each of which can be used to execute one (or more) of the collections of applications 964A-R; where the multiple software containers (also referred to as virtualization engines, virtual private servers, or jails) are user spaces (usually virtual memory spaces) that are separate from each other and from the kernel space in which the operating system runs; and where the collections of applications running in a given user space cannot access the memory of other processes unless explicitly permitted. In another such alternative embodiment, the virtualization layer 954 represents a hypervisor (sometimes referred to as a virtual machine monitor (VMM)) or a hypervisor executing on top of the host operating system, and each collection of applications 964A-R runs on top of a guest operating system within instances 962A-R (which in some cases can be considered a tightly isolated form of software container) that run on top of the hypervisor - the guest operating system and applications may not know that they are running on a virtual machine rather than on a "bare metal" host electronic device, or through paravirtualization, the operating system and / or applications may know of the existence of virtualization for optimization purposes. In still other alternative embodiments, one, some, or all of the applications are implemented as (one or more) monoliths, which can be generated by leveraging the application to directly compile only a limited set of libraries (such as from a library operating system (LibOS) that includes OS services) that provide the specific OS services required by the application. Since a monolith can be implemented to run directly on the hardware 940, directly on the hypervisor (in which case the monolith is sometimes described as running within a LibOS virtual machine), or within a software container, embodiments can be fully realized by leveraging monoliths that run directly on the hypervisor represented by the virtualization layer 954, monoliths that run within the software containers represented by the instances 962A-R, or as a combination of monoliths and the above techniques (e.g., monoliths and virtual machines both running directly on the hypervisor, collections of applications running in different software containers and monoliths). Note that the networking software 950 includes a federated learning coordinator 928, the operation of which is discussed herein.In some embodiments, the federated learning coordinator 928 may be instantiated in the virtualization layer 954.

[0108] The instantiation and virtualization (if implemented) of one or more sets of one or more applications 964A-R are collectively referred to as the (one or more) software instances 952. Each set of applications 964A-R, the corresponding virtualized constructs (e.g., instances 962A-R) (if implemented), and the portion of the hardware 940 that executes them (which is hardware dedicated to that execution and / or a time slice of the hardware that is shared in time) form the (one or more) individual virtual network elements 960A-R.

[0109] The (one or more) virtual network elements 960A-R perform functionality similar to that of the (one or more) virtual network elements 930A-R - e.g., similar to the (one or more) control communication and configuration modules 932A and the (one or more) forwarding tables 934A (this virtualization of the hardware 940 is sometimes referred to as network function virtualization (NFV)). Thus, NFV can be used to consolidate many network device types onto industry-standard high-volume server hardware, physical switches, and physical storage devices that can be located in data centers, NDs, and customer premise equipment (CPE). While the embodiments are illustrated using each instance 962A-R corresponding to one VNE 960A-R, alternative embodiments may implement this correspondence at a finer-grained level (e.g., line card virtual machines virtualize line cards, control card virtual machines virtualize control cards, etc.); it should be understood that the techniques described herein with reference to the correspondence of instances 962A-R to VNEs also apply to embodiments that use such finer-grained levels and / or single cores.

[0110] In certain embodiments, the virtualization layer 954 includes a virtual switch that provides forwarding services similar to a physical Ethernet switch. Specifically, this virtual switch forwards traffic between the instances 962A-R and the (one or more) physical NIs 946 and optionally between the instances 962A-R; additionally, this virtual switch can enforce network isolation between VNEs 960A-R that are not permitted to communicate with each other according to a policy (e.g., by honoring virtual local area networks (VLANs)).

[0111] Figure 9 A third exemplary ND implementation is the hybrid network device 906, which includes both a custom ASIC / dedicated OS and a COTS processor / standard OS in a single ND or a single card within the ND. In certain embodiments of such a hybrid network device, the platform VM (i.e., the VM that implements the functionality of the dedicated network device 902) can supply para-virtualization to the networking hardware present in the hybrid network device 906.

[0112] Regardless of the above exemplary implementations of the ND, when considering a single VNE among multiple VNEs implemented by the ND (e.g., only one VNE among the VNEs is part of a given virtual network) or where only a single VNE is currently being implemented by the ND, the shortened term network element (NE) is sometimes used to refer to that VNE. Also, in all of the exemplary implementations among the above exemplary implementations, each VNE among the VNEs (e.g., the VNEs 930A-R, VNE 960A-R, and those in the hybrid network device 906) receives data on a physical NI (e.g., 916, 946) and forwards that data out to an appropriate physical NI among the physical NIs (e.g., 916, 946). For example, a VNE implementing IP router functionality forwards IP packets based on some of the IP header information in the IP packet; where the IP header information includes a source IP address, a destination IP address, a source port, a destination port (where "source port" and "destination port" refer to protocol ports herein, as opposed to the physical ports of the ND), a transport protocol (e.g., User Datagram Protocol (UDP), Transmission Control Protocol (TCP), and Differentiated Services Code Point (DSCP) value).

[0113] Figure 9 The ND can form part of the Internet or a private network; and other electronic devices (not shown; such as end-user devices including workstations, laptops, netbooks, tablets, palmtop computers, mobile phones, smartphones, phablets, multimedia phones, Voice over Internet Protocol (VOIP) phones, terminals, portable media players, GPS units, wearable devices, gaming systems, set-top boxes, Internet-enabled household appliances) can be coupled (either directly or through other networks such as access networks) to the network to communicate with each other (either directly or through servers) and / or access content and / or services over the network (e.g., the Internet or a virtual private network (VPN) overlaying (e.g., tunneling) the Internet). Such content and / or services are typically provided by one or more servers (not shown) belonging to a service / content provider or one or more end-user devices (not shown) participating in a peer-to-peer (P2P) service, and such content and / or services can include, for example, public web pages (e.g., free content, storefronts, search services), private web pages (e.g., web pages providing username / password access to an email service), and / or corporate networks on a VPN. For example, an end-user device can be coupled (e.g., through customer premise equipment (CPE) (wired or wirelessly) coupled to an access network) to an edge ND, which is coupled (e.g., through one or more core NDs) to other edge NDs, which are coupled to electronic devices acting as servers. However, through computing and storage virtualization, in Figure 9One or more of the electronic devices acting as NDs can also host one or more such servers (e.g., in the case of the general network device 904, one or more of the software instances 962A - R can operate as servers; the same applies to the hybrid network device 906; in the case of the dedicated network device 902, one or more such servers can also run on the virtualization layer executed by the (one or more) processors 912); in this case, the server is said to be co - located with the VNE of that ND.

[0114] A virtual network is a logical abstraction of a physical network (such as the physical network in Figure 9 that provides network services (such as L2 and / or L3 services). A virtual network can be implemented as an overlay network (sometimes referred to as a network virtualization overlay) that provides network services (such as layer 2 (L2, data link layer) and / or layer 3 (L3, network layer) services) on an underlying network (e.g., an L3 network, such as an Internet Protocol (IP) network that uses tunnels (such as Generic Routing Encapsulation (GRE), Layer 2 Tunneling Protocol (L2TP), IPSec) to create the overlay network).

[0115] The Network Virtualization Edge (NVE) is located at the edge of the underlying network and participates in implementing network virtualization; the network - facing side of the NVE uses the underlying network to tunnel frames to and from other NVEs; the outside - facing side of the NVE sends data to and receives data from systems outside the network. A Virtual Network Instance (VNI) is a specific instance of a virtual network on an NVE (e.g., an NE / VNE on an ND, a part of an NE / VNE on an ND, where that NE / VNE is divided into multiple VNEs by emulation); one or more VNIs can be instantiated on an NVE (e.g., as different VNEs on an ND). A Virtual Access Point (VAP) is a logical connection point on the NVE for connecting external systems to the virtual network; a VAP can be a physical or virtual port identified by a logical interface identifier (such as a VLAN ID).

[0116] Examples of network services include: 1) Ethernet LAN emulation services (Ethernet-based multipoint services similar to Internet Engineering Task Force (IETF) Multiprotocol Label Switching (MPLS) or Ethernet VPN (EVPN) services), where external systems are interconnected across a network through a LAN environment on an underlying network (e.g., NVE provides separate L2 VNIs (Virtual Switching Instances) for different such virtual networks, and L3 (e.g., IP / MPLS) tunnels the encapsulation across the underlying network); and 2) virtualized IP forwarding services (similar to IETF IP VPN (e.g., Border Gateway Protocol (BGP) / MPLS IP VPN) from a service definition perspective), where external systems are interconnected across a network through an L3 environment on an underlying network (e.g., NVE provides separate L3 VNIs (Forwarding and Routing Instances) for different such virtual networks and L3 (e.g., IP / MPLS) tunnels the encapsulation across the underlying network)). Network services can also include quality of service capabilities (e.g., traffic classification marking, traffic conditioning, and scheduling), security capabilities (e.g., filters used to protect customer premises from network-initiated attacks and to avoid malformed routing announcements), and management capabilities (e.g., full detection and handling).

[0117] Some NDs include functionality for authentication, authorization, and accounting (AAA) protocols (e.g., RADIUS (Remote Authentication Dial-In User Service), Diameter, and / or TACACS+ (Terminal Access Controller Access Control System Plus)). AAA can be provided through a client / server model, where the AAA client is implemented on the ND and the AAA server can be implemented locally on the ND or on a remote electronic device coupled to the ND. Authentication is the process of identifying and verifying a subscriber. For example, a subscriber can be identified through a combination of a username and password or through a unique key. Authorization determines what a subscriber can do after being authenticated, such as obtaining access to certain electronic device information resources (e.g., by using an access control policy). Accounting is the recording of user activities. By way of a general example, an end-user device can be coupled (e.g., through an access network) through an edge ND (supporting AAA processing), which is coupled to a core ND, which is coupled to an electronic device implementing the servers of a service provider / content provider. AAA processing is performed to identify a subscriber record stored in the AAA server for that subscriber. The subscriber record includes a set of attributes used during the processing of that subscriber's service (e.g., subscriber name, password, authentication information, access control information, rate limit information, policy information).

[0118] Inside certain NDs (e.g., certain edge NDs) are represented end-user devices that use subscriber circuits (or sometimes customer premise equipment (CPE) such as residential gateways (e.g., routers, modems)). The subscriber circuits uniquely identify subscriber sessions within the ND and typically exist for the lifetime of the session. Thus, when a subscriber connects to an ND, that ND typically allocates a subscriber circuit, and when that subscriber disconnects, the ND deallocates that subscriber circuit accordingly. Each subscriber session represents a distinguishable packet flow that is transferred between the ND and the end-user device (or sometimes the CPE such as a residential gateway or modem) using a protocol such as the Point-to-Point Protocol over another protocol (PPPoX) (e.g., where X is Ethernet or Asynchronous Transfer Mode (ATM)), Ethernet, 802.1Q Virtual Local Area Network (VLAN), Internet Protocol, or ATM. Various mechanisms (e.g., manual provisioning, Dynamic Host Configuration Protocol (DHCP), DHCP / Clientless Internet Protocol Service (CLIPS), or Media Access Control (MAC) address tracking) can be used to initiate subscriber sessions. For example, the Point-to-Point Protocol (PPP) is commonly used for Digital Subscriber Line (DSL) services and requires the installation of a PPP client that enables the subscriber to enter a username and password, which in turn can be used to select a subscriber record. When DHCP is used (e.g., for cable modem services), a username is typically not provided; however, in such cases, other information is provided (e.g., information including the MAC address of the hardware in the end-user device (or CPE)). DHCP and CLIPS are used on the ND to capture MAC addresses and use these addresses to distinguish subscribers and access their subscriber records.

[0119] A virtual circuit (VC), synonymous with virtual connection and virtual channel, is a connection-oriented communication service that is delivered via packet-mode communication. Virtual circuit communication is similar to circuit switching in that both are connection-oriented, meaning that data is delivered in the correct order in both cases and there is a signaling overhead during the connection establishment phase. Virtual circuits can exist at different layers. For example, at layer 4, a connection-oriented transport layer data link protocol such as the Transmission Control Protocol (TCP) can rely on a connectionless packet-switching network layer protocol such as IP, where different packets can be routed via different paths and thus be delivered out of order. Where a reliable virtual circuit is established using TCP over the underlying unreliable and connectionless IP protocol, the virtual circuit is identified by a source and destination network socket address pair (i.e., the sender and receiver IP addresses and port numbers). However, virtual circuits are possible because TCP includes segment numbering and reordering on the receiver side to prevent out-of-order delivery. Virtual circuits are also possible at layer 3 (network layer) and layer 2 (data link layer); such virtual circuit protocols are based on connection-oriented packet switching, meaning that data is always delivered along the same network path (i.e., via the same NE / VNE). In such protocols, packets are not routed individually and complete addressing information is not provided in the header of each data packet; only a small virtual channel identifier (VCI) is required in each packet; and during the connection establishment phase, the routing information is transmitted to the NE / VNE; the switching only involves looking up the virtual channel identifier in a table rather than analyzing the complete address. Examples of network layer and data link layer virtual circuit protocols where data is always delivered via the same path: X.25, where the VC is identified by a virtual channel identifier (VCI); Frame Relay, where the VC is identified by a VCI; Asynchronous Transfer Mode (ATM), where the circuit is identified by a virtual path identifier (VPI) and virtual channel identifier (VCI) pair; General Packet Radio Service (GPRS); and Multiprotocol Label Switching (MPLS), which can be used for IP over virtual circuits (each circuit is identified by a label).

[0120] Some NDs (such as some edge NDs) use a hierarchical structure of circuits. The leaf nodes of the hierarchical structure of circuits are subscriber circuits. The subscriber circuits have parent circuits in the hierarchical structure, and the parent circuits typically represent the aggregation of multiple subscriber circuits and thus represent the aggregation of network segments and components used to provide access network connectivity from those end-user devices to the ND. These parent circuits can represent the physical or logical aggregation of subscriber circuits (such as virtual local area networks (VLANs), permanent virtual circuits (PVCs) (such as for asynchronous transfer mode (ATM)), circuit groups, channels, pseudowires, the physical NIs of the ND, and link aggregation groups). A circuit group is a virtual construct that allows various collections of circuits to be grouped together for configuration purposes; for example, aggregated rate control. A pseudowire is an emulation of a layer 2 point-to-point connection-oriented service. A link aggregation group is a virtual construct that combines multiple physical NIs for the purpose of bandwidth aggregation and redundancy. Thus, the parent circuits physically or logically encapsulate the subscriber circuits.

[0121] Each VNE (such as a virtual router, a virtual bridge (which can act as a virtual switch instance in a virtual private LAN service (VPLS))) is typically independently manageable. For example, in the case of multiple virtual routers, each virtual router in the virtual routers can share system resources, but is separate from other virtual routers in terms of its management domain, AAA (authentication, authorization, and accounting) namespace, IP address, and (one or more) routing databases. Multiple VNEs can be used in an edge ND to provide direct network access and / or different classes of services for subscribers of services and / or content providers.

[0122] Within some NDs, an "interface" independent of the physical NI can be configured as part of a VNE to provide higher-layer protocol and service information (such as layer 3 addressing). In addition to other subscriber configuration requirements, the subscriber record in the AAA server also identifies which context (such as which in the VNE / NE) the corresponding subscriber should be bound to within the ND. As used herein, binding forms an association between a physical entity (such as a physical NI, a channel) or a logical entity (such as a circuit such as a subscriber circuit or a logical circuit (a collection of one or more subscriber circuits)) and the interface of the context that is configured for that context via its network protocol (such as a routing protocol, a bridging protocol). When a certain higher-layer protocol interface is configured and associated with a physical entity, subscriber data flows on that physical entity.

[0123] Note that an electronic device uses machine-readable media (also referred to as computer-readable media), such as machine-readable storage media (e.g., magnetic disks, optical disks, solid state drives, read-only memory (ROM), flash memory devices, phase change memory) and machine-readable transmission media (also referred to as carriers) (e.g., electrical, optical, radio, acoustic, or other forms of propagated signals - such as carrier waves, infrared signals), to store and transmit (internally and / or together with other electronic devices over a network) code (which consists of software instructions and which is sometimes referred to as computer program code or a computer program) and / or data. Thus, an electronic device (e.g., a computer) includes hardware and software, such as a collection of one or more processors (e.g., where the processors are microprocessors, controllers, microcontrollers, central processing units, digital signal processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), other electronic circuitry, or a combination of one or more of the foregoing) coupled to one or more machine-readable storage media to store code for execution on the collection of processors and / or to store data. For example, an electronic device may include non-volatile memory containing code, since non-volatile memory can continue to hold code / data even when the electronic device is turned off (when power is removed). When the electronic device is turned on, that portion of the code to be executed by the (one or more) processors of the electronic device is typically copied from the slower non-volatile memory into the volatile memory (e.g., dynamic random access memory (DRAM), static random access memory (SRAM)) of the electronic device. A typical electronic device also includes a collection of one or more physical network interfaces ((one or more) NIs) to establish a network connection with other electronic devices (thereby using propagated signals to transmit and / or receive code and / or data). For example, the collection of physical NIs (or the collection of physical NIs in combination with the collection of processors executing code) may perform any formatting, decoding, or conversion to allow the electronic device to send and receive data over either wired and / or wireless connections. In some embodiments, the physical NI may include radio circuitry capable of (1) receiving data from other electronic devices over a wireless connection and / or (2) transmitting data over a wireless connection to other devices. This radio circuitry may include (one or more) transmitters, (one or more) receivers, and / or (one or more) transceivers suitable for radio frequency communication. The radio circuitry may convert digital data into radio signals with appropriate parameters (e.g., frequency, timing, channel, bandwidth, etc.). The radio signals may then be transmitted via an antenna to the (one or more) appropriate recipients. In some embodiments, the collection of (one or more) physical NIs may include (one or more) network interface controllers (NICs), also referred to as network interface cards, network adapters, or local area network (LAN) adapters.One or more network interface cards (NICs) may facilitate connecting an electronic device to other electronic devices, allowing them to communicate over a wire by plugging a cable into a physical port connected to the NIC. Different combinations of software, firmware, and / or hardware may be used to implement one or more portions of the embodiments.

[0124] A network node / device is an electronic device. Some network devices are "multi-service network devices" that provide support for multiple networking functions (such as routing, bridging, switching, layer 2 aggregation, session border control, quality of service, and / or subscriber management) and / or provide support for multiple application services (such as data, voice, and video). Examples of network nodes also include NodeB, base station (BS), multi-standard radio (MSR) radio nodes (such as MSR BS, eNodeB, gNodeB, MeNB, SeNB), integrated access backhaul (IAB) nodes, network controllers, radio network controllers (RNCs), base station controllers (BSCs), relays, donor nodes controlling relays, base transceiver stations (BTSs), central units (such as in a gNB), distributed units (such as in a gNB), baseband units, centralized baseband, C-RAN, access points (APs), transmission points, transmission nodes, remote radio units (RRUs), remote radio heads (RRHs), nodes in a distributed antenna system (DAS), core network nodes (such as MSC, MME, etc.), O&M, OSS, SON, positioning nodes (such as E-SMLC), etc.

[0125] A communication network (such as communication network 190) may include any type of communication, telecommunication, data, cellular, and / or radio network or other similar type of system and / or may be interfaced with any type of communication, telecommunication, data, cellular, and / or radio network or other similar type of system. In some embodiments, the communication network may be configured to operate according to a specific standard or other type of predefined rules or procedures. Thus, a particular embodiment of a communication network may implement communication standards such as Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), Long Term Evolution (LTE), and / or other suitable 2G, 3G, 4G, or 5G standards; wireless local area network (WLAN) standards such as Institute of Electrical and Electronics Engineers (IEEE) 802.11 standards; and / or any other suitable wireless communication standards such as Worldwide Interoperability for Microwave Access (WiMax), Bluetooth, Z-Wave, and / or ZigBee standards.

[0126] A communication network may include one or more backhaul networks, core networks, IP networks, public switched telephone networks (PSTN), packet data networks, optical networks, wide area networks (WANs), local area networks (LANs), wireless local area networks (WLANs), wired networks, wireless networks, metropolitan area networks, and other networks that enable communication between devices.

[0127] The specification mentions "an embodiment", "embodiment", "exemplary embodiment", etc., indicating that the described embodiments may include specific features, structures, or characteristics, but each embodiment may not necessarily include the specific features, structures, or characteristics. In addition, such phrases do not necessarily refer to the same embodiment. Further, when a specific feature, structure, or characteristic is described in relation to an embodiment, it is considered within the knowledge of those skilled in the art to vary such feature, structure, or characteristic in relation to other embodiments, whether or not explicitly described.

[0128] Bracketed text and boxes with dashed boundaries (such as long dashes, short dashes, dotted lines, and dots) may be used herein to illustrate optional operations for adding additional features to an embodiment. However, such symbols should not be considered to mean that these are the only options or optional operations and / or that boxes with solid boundaries are not optional in some embodiments.

[0129] In the description, embodiments, and claims, the terms "coupled" and "connected" may be used along with their derivatives. It should be understood that these terms are not intended as synonyms for each other. "Coupled" is used to indicate that two or more elements, which may or may not be in direct physical or electrical contact with each other, cooperate or interact with each other. "Connected" is used to indicate the establishment of communication between two or more elements that are coupled to each other. As used herein, "set" refers to any positive integer number of items including one item.

[0130] Alternative embodiments

[0131] Although the invention has been described in terms of several embodiments, those skilled in the art will recognize that the invention is not limited to the described embodiments and that the invention can be practiced with modifications and variations within the spirit and scope of the appended claims. The description is thus to be regarded as illustrative rather than limiting.

Claims

1. A method for privacy-preserving information exchange between a local party, implemented in an electronic device acting as the local party, and another electronic device acting as an aggregator, wherein the aggregator exchanges information with a plurality of local parties including the local party, the method comprises: storing (702) a plurality of values in a two-dimensional (2D) vector, wherein a first dimension of the 2D vector is based on how many values there are in the plurality of values, and wherein each position in the first dimension has a unique value within the plurality of values, and wherein each unique value within the plurality of values is located in a randomly selected position in a second dimension, the second dimension being based on the number of the plurality of local parties including the local party; and transmitting (704) the 2D vector to the aggregator with masking of the aggregator to prevent the aggregator from decoding the 2D vector to determine the plurality of values transmitted by the local party, wherein aggregating each position of the masked 2D vector with a corresponding position of masked 2D vectors from other local parties allows unmasking of the plurality of values in the 2D vector without identifying the local party from which the values originated.

2. The method according to claim 1, wherein the exchanged information is decision tree information for decision tree learning, wherein the aggregator is to generate a decision tree, wherein the plurality of values are a first plurality of split point value candidates for at least one feature of the decision tree, and wherein the aggregator is to determine a single split point value for a node of the decision tree based on the aggregated 2D vector.

3. The method according to claim 2, wherein the first dimension of the 2D vector is equal to the number of the first plurality of split point value candidates, and wherein the second dimension of the 2D vector is not less than the number of local parties.

4. The method according to claim 2, wherein each of the plurality of split point value candidates is mapped to a sketch of data for the feature at the local party.

5. The method according to claim 2, further comprises: receiving (708) a second plurality of split point value candidates from the aggregator; and transmitting (710) quantile sketch information of the second plurality of split point value candidates mapped to the feature to the aggregator with masking to prevent the aggregator from decoding the quantile sketch information, wherein aggregating the masked quantile sketch information with quantile sketch information from other local parties allows decoding of the aggregated quantile sketch information.

6. The method according to claim 5, further comprises: receiving (712) a third plurality of split point value candidates from the aggregator; and Transmit (714) additional quantile sketch information of the third plurality of split point value candidates mapped to the feature to the aggregator using masking to prevent the aggregator from decoding the additional quantile sketch information, and wherein aggregating the masked additional quantile sketch information with additional quantile sketch information from other local participants allows decoding of the aggregated additional quantile sketch information, wherein the additional quantile sketch information is based on the derivative of a loss function for the decision tree.

7. The method of claim 6, wherein the third plurality of split point value candidates is a subset of the second plurality of split point value candidates.

8. The method of claim 1 or 2, further comprising: retransmit (706) one or more values in response to a request from the aggregator, each of the values being stored in a randomized location within another vector, wherein each retransmission uses masking of the aggregator to prevent the aggregator from decoding the other vector, and wherein aggregating the masked vectors with masked vectors from other local participants allows decoding of the aggregated vectors.

9. An electronic device (902, 904) acting as a local participant for privacy-preserving information exchange between the local participant and another electronic device acting as an aggregator, wherein the aggregator exchanges information with a plurality of local participants including the local participant, the electronic device (902, 904) comprising: a processor (912, 942) and a non-transitory machine-readable storage medium (918, 948) storing instructions that, when executed by the processor, cause the electronic device to perform: store (702) a plurality of values in a two-dimensional (2D) vector, wherein a first dimension of the 2D vector is based on how many values are in the plurality of values, and wherein each position in the first dimension has a unique value within the plurality of values, and wherein each unique value within the plurality of values is located in a randomly selected position in a second dimension, the second dimension being based on the number of the plurality of local participants including the local participant; and transmit (704) the 2D vector to the aggregator using masking of the aggregator to prevent the aggregator from decoding the 2D vector to determine the plurality of values transmitted by the local participant, wherein aggregating each position of the masked 2D vector with a corresponding position of masked 2D vectors from other local participants allows unmasking of the plurality of values in the 2D vector without identifying the local participant from which the values originated.

10. The electronic device of claim 9, wherein the exchanged information is decision tree information for decision tree learning, wherein the aggregator is to generate a decision tree, wherein the plurality of values are a first plurality of split point value candidates for at least one feature of the decision tree, and wherein the aggregator is to determine a single split point value for a node of the decision tree based on the aggregated 2D vector.

11. The electronic device according to claim 10, wherein the first dimension of the 2D vector is equal to the number of the first plurality of split point value candidates, and wherein the second dimension of the 2D vector is not less than the number of local participants.

12. The electronic device according to claim 10, wherein each of the plurality of split point value candidates is mapped to a sketch of data for the feature at the local participant.

13. The electronic device according to claim 10, wherein the instructions can further cause the electronic device to perform: Receiving (708) a second plurality of split point value candidates from the aggregator; and Transmitting (710) quantile sketch information of the second plurality of split point value candidates mapped to the feature to the aggregator using a mask to prevent the aggregator from decoding the quantile sketch information, wherein aggregating the masked quantile sketch information with quantile sketch information from other local participants allows decoding of the aggregated quantile sketch information.

14. The electronic device according to claim 9 or 10, wherein the instructions can further cause the electronic device to perform: Retransmitting (706) one or more values according to a request from the aggregator, each of the values being stored in a randomized position within another vector, wherein each retransmission uses a mask for the aggregator to prevent the aggregator from decoding the other vector, and wherein aggregating the masked vectors from other local participants allows decoding of the aggregated vector.

15. A non-transitory machine-readable storage medium (918, 948) storing instructions that, when executed by a processor of an electronic device acting as a local participant, can cause the electronic device to perform operations for privacy-preserving information exchange between the local participant and another electronic device acting as an aggregator, wherein the aggregator exchanges information with a plurality of local participants including the local participant, the operations comprising: Storing (702) a plurality of values in a two-dimensional (2D) vector, wherein the first dimension of the 2D vector is based on how many values there are in the plurality of values, and wherein each position in the first dimension has a unique value within the plurality of values, and wherein each unique value within the plurality of values is located in a randomly selected position in the second dimension, the second dimension being based on the number of the plurality of local participants including the local participant; and Transmitting (704) the 2D vector to the aggregator using a mask for the aggregator to prevent the aggregator from decoding the 2D vector to determine the plurality of values transmitted by the local participant, wherein aggregating each position of the masked 2D vector with the corresponding position of the masked 2D vectors from other local participants allows unmasking of the plurality of values in the 2D vector without identifying the local participant from which the values originated.

16. The non-transitory machine-readable storage medium according to claim 15, wherein the information exchanged is decision tree information for decision tree learning, wherein the aggregator is to generate a decision tree, wherein the plurality of values are first plurality of split point value candidates for at least one feature of the decision tree, and wherein the aggregator is to determine a single split point value for a node of the decision tree based on the aggregated 2D vectors.

17. The non-transitory machine-readable storage medium according to claim 16, wherein the first dimension of the 2D vector is equal to the number of the first plurality of split point value candidates, and wherein the second dimension of the 2D vector is not less than the number of local participants.

18. The non-transitory machine-readable storage medium according to claim 16, wherein each of the plurality of split point value candidates is mapped to a sketch of data for the feature at the local participant.

19. The non-transitory machine-readable storage medium according to claim 16, wherein the instructions can further cause the electronic device to perform: receiving (708) from the aggregator a second plurality of split point value candidates; and transmitting (710) quantile sketch information of the second plurality of split point value candidates mapped to the feature to the aggregator using a mask to prevent the aggregator from decoding the quantile sketch information, wherein aggregating the masked quantile sketch information with quantile sketch information from other local participants allows decoding of the aggregated quantile sketch information.

20. The non-transitory machine-readable storage medium according to claim 15 or 16, wherein the instructions can further cause the electronic device to perform: retransmitting (706) one or more values according to a request from the aggregator, each of the values being stored in a randomized position within another vector, wherein each retransmission uses a mask for the aggregator to prevent the aggregator from decoding the other vector, and wherein aggregating the masked vectors with masked vectors from other local participants allows decoding of the aggregated vectors.

21. A computer program product comprising instructions that, when executed by a processor of an electronic device acting as a local participant, can cause the electronic device to perform the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Communication efficient federated learning

    CN107871160A