Method and System for Accelerating Distributed Principal Component Analysis Using a Noisy Channel
By using SGD-based aerial aggregation technology and area adaptive power control in mobile networks, the problems of data privacy and communication delay in PCA calculation are solved, and efficient PCA calculation and accelerated convergence are achieved.
Patent Information
- Application Number
- CN202210575990.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-25
- Filing Date
- 2022-05-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-05-25
AI Technical Summary
The prior art faces data privacy issues and high communication latency when implementing principal component analysis (PCA) in mobile networks, especially when uploading large amounts of data.
Aerial aggregation technology based on stochastic gradient descent (SGD) is adopted to accelerate the gradient descent of PCA by adding moderate noise to channel noise, and optimize learning performance in combination with regional adaptive power control.
It effectively solves the problems of data privacy and communication delay, realizes efficient PCA calculation in mobile networks, and uses noise to accelerate convergence in the saddle area.
Smart Images

Figure CN115474266B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 193,042, filed on May 25, 2021, entitled "Methods and Systems for Accelerating Distributed Principal Components with Noisy Signals", the entire content of which is incorporated herein by reference. Technical Field
[0003] This application generally relates to principal component analysis (PCA) techniques applicable in the context of networked computing devices, and more particularly, to PCA techniques for calculating low - dimensional subspaces for data distributed on edge devices connected to a network. Background Art
[0004] In recent years, the large amounts of data distributed on edge devices such as smart phones, Internet of Things (IoT) sensors, and various other devices, along with ubiquitous connectivity, have triggered a paradigm shift towards distributed machine learning and large - scale data analysis.
[0005] Principal component analysis (PCA), as a standard technique in data analysis, provides a simple way to discover low - dimensional subspaces called principal components, which minimizes the information loss of high - dimensional data sets. This is useful for data compression, simplification, and feature extraction. For these reasons, PCA is applied in almost all scientific fields from wireless communication to machine learning.
[0006] Common methods of PCA are based on the singular value decomposition (SVD) of a data table that includes all data samples by rows. However, the required data set makes this method infeasible for implementing PCA in mobile networks because uploading mobile data violates the privacy of data samples.
[0007] The above background is only intended to provide an overview of some current problems and is not intended to be exhaustive. Other contextual information may become more apparent after reading the following detailed description. Brief Description of the Drawings
[0008] The non - limiting and non - exhaustive embodiments of the present disclosure are described with reference to the following drawings, in which like reference numerals refer to like components throughout the various views unless otherwise specified.
[0009] Figure 1 An example architecture of the AirPCA system according to one or more embodiments described herein is shown.
[0010] Figure 2 An example transmitter design for an edge device according to one or more embodiments described herein is shown.
[0011] Figure 3 Shows an example receiver design for an edge server according to one or more embodiments described herein.
[0012] Figure 4A 、 Figure 4B and Figure 4C Show three example types of regions according to one or more embodiments described herein.
[0013] Figure 5A Is a graph providing an example comparison of noiseless AirPCA, AirPCA with power control, and centralized PCA according to one or more embodiments described herein.
[0014] Figure 5B Is a graph providing an example comparison of AirPCA with power control, AirPCA with fixed power, and centralized PCA according to one or more embodiments described herein.
[0015] Figure 6A Is a first graph providing an example first learning performance comparison using a first data set.
[0016] Figure 6B Is a graph providing an example second learning performance comparison using a second data set.
[0017] Figure 7A Is a first graph showing an example impact of a power consumption coefficient on the learning performance of AirPCA with region - adaptive power control.
[0018] Figure 7B Is a second graph showing an example impact of a power consumption coefficient on the learning performance of AirPCA with region - adaptive power control.
[0019] Figure 8A Is a graph showing an example impact of the number of devices on the learning performance of AirPCA with region - adaptive power control.
[0020] Figure 8B Is a graph showing an example impact of a truncation threshold on the learning performance of AirPCA with region - adaptive power control.
[0021] Figure 9 Is a flowchart representing an example operation of a device participating in joint principal component analysis according to various aspects and embodiments of the present disclosure.
[0022] Figure 10 Is a flowchart representing an example operation of a server participating in joint principal component analysis according to various aspects and embodiments of the present disclosure.
[0023] Figure 11 is an example computing device that can implement any of the various devices cited herein, according to one or more embodiments described herein. Detailed Description
[0024] Aspects or features of the present disclosure are described with reference to the accompanying drawings, in which like reference numerals are always used to indicate like elements. In this specification, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it should be understood that certain aspects of the present disclosure may be practiced without these specific details or with other methods, components, materials, etc. In other instances, well-known structures and devices are shown in block diagram form to facilitate describing the present disclosure.
[0025] One or more aspects of the techniques described herein generally relate to an efficient design of joint PCA in a wireless system based on federated learning (FL) over the air, which utilizes the waveform superposition property of the multiple access channel to achieve low latency for over-the-air aggregation. This design is referred to herein as over-the-air PCA, or "AirPCA". Additionally, a power control scheme is disclosed that adapts the transmission power of a device to stochastic gradient descent (SGD) such that channel noise is transformed into an accelerator for the descent.
[0026] FL techniques can be used to protect data privacy for distributed PCA, also known as joint PCA. Joint PCA can help compress and simplify data distributed at the network edge (e.g., data generated by vehicle sensing or augmented reality / virtual reality (AR / VR) applications and collected by different devices) for convenient storage and their further use in edge learning.
[0027] The FL framework involves devices that update a prediction model. Each device can use local data to generate local updates, and the local updates (instead of the data) can be sent to a server. The server can aggregate the updates and update the global model. In this way, data privacy is protected.
[0028] Federated learning can also protect the data ownership of devices by avoiding uploading raw data, while providing a mechanism to fully utilize distributed mobile data. The server can request each device to upload updates for the global model, which are calculated using local training data. The updates do not directly expose the content of the local data and contain much less information than the local data, thus protecting the data ownership of the users.
[0029] In the "one-shot method", joint PCA can involve the device calculating an estimate of the principal components of the device via SVD of the device's local data and uploading the local estimate of the device to the server for aggregation to obtain a global estimate. The drawback of the one-shot method is that sharing local principal components raises data privacy issues. Additionally, when the number of devices increases, uploading all SVD results leads to high communication latency. By moderately reducing the dimension of the local subspace estimate, the communication latency problem can be alleviated. However, reducing the dimension of the local subspace estimate results in a biased error, which distorts the global estimate when the local data sets are highly non-independent and identically distributed.
[0030] Another solution for joint PCA is to apply the power method, which can be integrated with air aggregation to provide fast convergence and negligible communication latency. However, the power method is sensitive to noise perturbations, making it infeasible in wireless networks, especially when the signal-to-noise ratio (SNR) is low.
[0031] Embodiments of the present disclosure can apply SGD-based processing to solve joint PCA as an optimization problem of finding a subspace (principal components) to minimize the error function of data compression by projecting onto the subspace. In the context of joint PCA, the difficulty in applying SGD comes from the unitary / orthogonal constraint on the optimization variable that is the subspace, which makes the optimization problem non-decomposable. FL cannot be directly applied to non-decomposable optimization problems. This difficulty can be overcome by the discovery that the solution to the unconstrained problem without the unitary / orthogonal constraint also solves the original constrained problem.
[0032] Embodiments of the present disclosure can utilize the observation that the SGD method is robust to channel noise. Additionally, in the presence of channel noise, the SGD method can guarantee convergence to the global optimum, thus the SGD method improves upon the power method. Furthermore, by adopting air aggregation in the gradient upload phase, the communication latency problem can also be solved, and thus the disclosed SGD method can outperform the one-shot method, especially when the number of devices is large.
[0033] In scenarios with many devices and high-dimensional data, uploading local model updates from the devices can lead to a communication bottleneck in FL (including joint PCA). Various techniques have emerged to overcome the bottleneck, including source coding and resource management. Air FL is a class of techniques that achieves air aggregation by superimposing analog-modulated model updates sent simultaneously by the devices. Compared with digital orthogonal access, air aggregation that supports simultaneous access has the advantage of reducing multi-access latency when the number of devices is large.
[0034] However, uncoded analog transmission used in the air FL may expose the received signal to perturbations of channel noise, which may potentially degrade the learning performance. Embodiments of the present disclosure can turn this disadvantage into an advantage of AirPCA by leveraging the characteristics of the error function of AirPCA.
[0035] To train a model (e.g., a deep neural network) using FL, the (prediction) loss function is dataset-dependent and does not have a known expression. In contrast, the PCA error function is well-defined and its theoretical properties are well understood. The PCA error function has a finite number of stationary points, which include a global optimum and multiple discrete saddle points.
[0036] Therefore, the regions along the descent path belong to one of three types: (1) saddle regions centered on relevant saddle points, (2) non-stationary regions with relatively large slopes, and (3) optimal regions centered on the global optimum. These characteristics indicate that if the descent path encounters a saddle region, gradient descent can be trapped at a saddle point with zero gradient. One solution is to add artificial noise to the gradient to escape from the saddle point. On the other hand, the noise slows down the descent outside the saddle region and reduces the convergence accuracy. Instead of adding artificial noise, embodiments of the present disclosure can utilize the channel noise present in the received signal in AirPCA to help escape from the saddle point by amplifying the influence of the channel noise in the received signal in AirPCA in the saddle region but reducing the influence of the channel noise in the received signal in AirPCA in other types of regions on the descent path.
[0037] Embodiments can utilize the noise by designing region-adaptive power control for AirPCA. To implement region-adaptive power control, descent speed analysis can be performed. Embodiments can apply the framework of descent speed analysis to AirPCA, which is based on a martingale-based analysis method for centralized PCA training. This framework can consider wireless propagation and technologies, including orthogonal frequency division multiplexing (OFDM), air aggregation, channel fading, and noise.
[0038] The descent speed of AirPCA can be measured by reducing the expected error function over a given number of communication rounds. Using this framework and leveraging the above characteristics of the error function, the descent speed in different regions on the descent path can be mathematically characterized.
[0039] Consider gradient descent in the non-stationary region. In the presence of fading, a lower bound on the descent speed can be derived, which is shown to be a monotonically increasing function of the expected received signal-to-noise ratio (SNR) and the expected number of active devices. Due to signal amplitude alignment in air aggregation, the expected received SNR can be uniform for all devices.
[0040] Conversely, the descent speed in the saddle region is a monotonically decreasing function of these two variables because their decrease amplifies the noise effect and accelerates the escape from the saddle point. Finally, under the influence of channel noise, as long as the step size is small enough, the descent path can eventually enter the optimal region with probability.
[0041] Based on the analysis results of the descent speed analysis, by coordinating the transmission power of the device, a simple scheme for online power control can be designed to adapt the uniform received SNR to the type of the current descent region. Thus, the gradient descent of AirPCA can be accelerated. When the saddle region is detected, the received SNR can be fixed at the minimum value to amplify the noise effect so that the descent path can escape from the saddle point. This promotes power saving under the average power constraint.
[0042] On the other hand, when the non-stationary region or the optimal region is detected, the received SNR is enhanced by exhausting all the power savings of the previous round in the current round (referred to as single-save consumption) or by distributing the power savings over multiple rounds using a decreasing geometric sequence with a common ratio that controls the save-dissipation speed (referred to as gradual save consumption).
[0043] Experiments can be conducted using several well-known real-world datasets (e.g., the modified National Institute of Standards and Technology (MNIST) dataset, the 10-class Canadian Institute for Advanced Research (CIFAR-10) dataset, and the face image dataset known as the AR dataset) to evaluate the learning performance of AirPCA. The disclosed region-adaptive power control is shown to be effective in escaping from the saddle point and accelerating the convergence of AirPCA. The region-adaptive power control can also achieve the convergence accuracy of centralized PCA. In addition, it is found that if the common ratio is optimized, gradual save consumption can outperform single-save consumption. The effects of other parameters such as the number of devices and the channel truncation threshold can also be studied.
[0044] The following sections of this disclosure are organized as follows. First, an example AirPCA system is described. Next, the descent speed of AirPCA is analyzed. Subsequently, an example design of region-adaptive power control is described. Subsequently, example experimental results are given. Finally, a general overview and an example computing device are described.
[0045] Example AirPCA System
[0046] Figure 1An example architecture of the AirPCA system according to one or more embodiments described herein is shown. The example architecture 100 includes devices 101, 102, 103, and an edge server 110. Devices 101, 102, 103 can modulate signals through linear analog modulation and can transmit simultaneously. The server 110 can receive the aggregated signals (superimposed waveforms) on each sub-channel. Since the global gradient is the average of the local gradients, the server 110 can directly estimate the global gradient based on the aggregated signals.
[0047] A. Air Aggregation System
[0048] The example architecture 100 may include a broadband air aggregation system that supports AirPCA. The example broadband air aggregation system 100 may include K devices (101, 102, 103) communicating with a single server 110. The communication may include multiple rounds, and each round may be divided into an uplink transmission phase and a downlink transmission phase.
[0049] Figure 2 An example transmitter design for an edge device according to one or more embodiments described herein is shown. The example transmitter design may be incorporated into, for example, Figure 1 devices 101, 102, 103 as shown. The example transmitter design includes: local subspace calculation 201, analog amplitude modulation 202, serial-to-parallel converter 204, power control 205, truncated channel inversion 206, inverse fast Fourier transform (IFFT) 208, adding CP and parallel-to-serial converter 210, a mixer 212 that can combine the output from 210 with a carrier 211, and an antenna 214.
[0050] Figure 3 An example receiver design for an edge server according to one or more embodiments described herein is shown. The example receiver design may be incorporated into, for example, Figure 1 the edge server 110 as shown. The example receiver design includes an antenna 301, a mixer / demixer 302, a carrier 303, a superimposed waveform 304, removing CP and parallel-to-serial converter 306, fast Fourier transform 308, serial-to-parallel converter 310, a product 1 / K for the i-th parameter n (i) 312, gradient descent 314, and global subspace update 315.
[0051] Reference Figure 1 , Figure 2 and Figure 3, consider an uplink phase of any round. Each of the devices 101, 102, 103 sends a fixed number (denoted as c) of symbols to the server via M (frequency) subchannels generated by OFDM. To this end, the c symbols are divided into c / M blocks. Each block is transmitted within one OFDM symbol duration, and each subchannel is modulated with one symbol using linear analog modulation. The transmissions of all the devices are simultaneous to achieve in-air aggregation. Then, the i-th aggregated symbol (denoted as y n (i) ) received by the server in the n-th round of communication is given by:
[0052]
[0053] where, denotes the symbol sent by device k, and Gaussian random variable and denote the gain and noise of the corresponding subchannel respectively, is the precoding coefficient. Let P k,n denote the power consumption of the wideband transmission of device k in round n:
[0054] The transmission of each of the devices 101, 102, 103 may be subject to an average power constraint for a given constant P:
[0055]
[0056] In-air aggregation requires channel inversion such that each received symbol is the desired sum of the transmitted symbols. Embodiments may employ existing schemes designed to satisfy an average power constraint known as truncated channel inversion. Figure 2 The design of the truncated channel inversion 206 in Figure 3 can support signal alignment at the receiver of
[0057]
[0058] where, as described below, the controllable received power P n rx and the constant G are referred to as the signal amplitude calibration factor and the truncation threshold respectively. The factor P n rx scales the amplitude of the aggregated signal at the receiver and forms a power control sequence {P n rx} throughout the process of controlling the received power under the constraint of equation (2). Given the same distribution of subchannel gains, it can be obtained that:
[0059]
[0060] Among them, is the exponential integral function. On the other hand, the truncation threshold G avoids excessive power consumption caused by the inversion of subchannels with deep fading. To implement a fixed transmission delay, the symbols assigned to the truncated subchannels are discarded. The probability that a subchannel avoids truncation (or equivalently, its symbol is transmitted) is called the activation probability and is denoted by ζ act It can be obtained as:
[0061]
[0062] ζ act The value of reflects the reliability of the wireless channel.
[0063] After receiving the aggregated message, the server 110 can update the global model and further broadcast the updated global model in the downlink, which can be the same for all devices 101, 102, 103. Since the transmit power and bandwidth are usually large for broadcasting, we consider it as a high SNR condition and ignore the distortion during the downlink broadcast.
[0064] In practice, some devices 101, 102, 103 may occasionally disconnect from the server 110, which is called the outage effect. We consider the disconnection as a special case of channel truncation where all subchannels are truncated. In addition, when a device in the outage reconnects to the server 110, it first receives the latest subspace broadcast from the server 110, then continues to calculate the local gradient and rejoins AirPCA again.
[0065] B. Distributed PCA Problem and Algorithm
[0066] (1) Distributed PCA problem: In the example distributed PCA problem, assume that the global dataset including L samples is evenly distributed on K devices. Let D k denote the local dataset of device k generated by uniformly sampling the global dataset. The local datasets have a uniform size: where L = Kl0. Assume that the local datasets are pre-acquired and remain unchanged during the processing duration, which is a common setting. The distributed PCA problem is to find the low-dimensional subspace of the data space (called the principal components) to compress the distributed dataset under the minimum distortion criterion. Let d and D (D >> d) denote the dimension of the principal components and the dimension of the data space respectively. Let the i-th sample be denoted as In addition, the d-dimensional principal components are represented by the unitary / orthogonal real matrix The sample x i can use its projection in the subspace W Tx i is approximated by the projection onto WW T x i . To minimize the approximation error, the distributed PCA problem can be formulated as:
[0067]
[0068] s.t. W T W = I.
[0069] where, is the data set, and is used as aggregation. If all devices can upload their local data to the server, the problem (P1) can be solved by applying SVD to the centralized data set X = [x1; x2; …; x L . However, for the distributed PCA scenario, direct data upload is not feasible under the constraint of data privacy. Different SGD-based solutions can be used.
[0070] (2) Distributed PCA processing: In the example distributed PCA processing technique, to simplify the notation, let the objective function of problem (P1) be represented as:
[0071]
[0072] F(W) has a stationary point in the form of W = U d Q, where the column vectors of are the d different eigenvectors of the covariance matrix R = XX T , and is an arbitrary unitary matrix. If the Hessian matrix has positive and negative eigenvalues, then W is called a saddle point. Except for the point where U d contains the d principal eigenvectors of R, all stationary points of F(W) are saddle points. The above properties indicate that F(W) includes three types of regions as shown in Figure 4A , Figure 4B and Figure 4C . If the descent process can avoid being trapped at the saddle point, the gradient descent algorithm can effectively solve the following optimization problem, which is a simplified version of (P1) without its unitary / orthogonal constraint:
[0073]
[0074] A method to escape from the saddle point is to add artificial noise to the gradient. Then the column space of the optimal point W* solves the problem (P1).
[0075] As a special case of FL, the iterative algorithm for distributed PCA can be based on SGD. To describe the algorithm, consider any communication round of the algorithm. At the beginning, the server broadcasts the current principal component W to all devices to calculate the gradient based on all local data samples. For this purpose, the local objective function of device k is given as:
[0076]
[0077] In addition, the data covariance matrix at device k is defined as R k =X k X k T , where the matrix X of D×l0 k includes the samples in the local dataset D k . Then, the local gradient F k (W) at device k is calculated as:
[0078]
[0079] The devices upload their local gradients to the server for aggregation, and then update the principal component W. Note that the gradient of the global objective function F(W) can be written according to the local gradients as:
[0080]
[0081] However, the received gradients are deliberately perturbed by noise to escape from the saddle point:
[0082]
[0083] where z is a random vector representing the noise. Then, the principal component W in the current round (e.g., round n) n is updated by the server:
[0084]
[0085] where μ is a fixed step size. The above process for each round is repeated until W converges.
[0086] (3) AirPCA implementation: AirPCA can implement distributed PCA in an air aggregation system. The implementation in the nth round is described as follows. To facilitate the transmission on the in-phase and quadrature channels, the local and global gradients (matrices) and are complex vectorized using the mapping functions g k (.) and g(.), where the results are denoted as and each including elements. Given an independent and identically distributed (i.i.d.) data distribution across devices, the following assumptions for unbiased estimation are common in the literature on distributed learning and estimation.
[0087] Assumption 1 (Unbiased Estimation): The local gradients computed at each device can be assumed to be unbiased estimates of the global gradient:
[0088] g k (W) = g(W) + Δk, 1 ≤ k ≤ K, (11)
[0089] where the estimation error vector Δ k is called data noise and for a given constant k 2 satisfies:
[0090]
[0091] Note that the data noise {Δ k} across different devices is correlated.
[0092] To enable air aggregation, each device uses linear analog modulation to transmit its local gradient. Following a model with independent and identically distributed (i.i.d.) data distribution, the symbols at device k, i.e., the elements of the local gradient g k (W), can be modeled as random variables with the same distribution having mean η and variance v 2 ; the statistics for all devices are the same and they are known. For power control in equation (2), the untruncated symbols are normalized in the n-th round to have zero mean and unit variance, i.e., and then transmitted over the subchannel; otherwise, the symbol 0 is transmitted. Synchronized in time (i.e., using timing advance in 3GPP) and using truncated channel inversion in equation (3), all devices simultaneously transmit their OFDM symbols with aligned boundaries to perform air aggregation. This results in a symbol vector received by the server:
[0093]
[0094] Then, the received symbols are de-normalized to give the elements of the noisy global gradient, denoted as
[0095]
[0096] where K n (i) is defined as the number of devices transmitting the i-th gradient element in the n-th round, K n (i) denotes the set of devices, i.e., |K n(i) | = K n (i) This quantity follows a binomial distribution, K n (i) ~ B(K; ζ act ), ζ act is the activation probability in equation (5). Equation (14) implies that K n (i) is non-zero. This is reasonable because when ζ act is close to one and / or K is large, Pr(K n (i) = 0) = (1 - ζ act ) K is close to zero. Substituting the normalization equation and equation (13) into equation (14) gives the noisy global gradient received by the server:
[0097]
[0098] where the noise vector ξ n combines channel and data noise and is defined element-wise as:
[0099]
[0100] By unvectorizing the solution in equation (15) into a matrix as in equation (10), the principal components are updated, completing the n-th round of AirPCA. As in equation (10), the principal components are updated, completing the n-th round of AirPCA.
[0101] AirPCA Convergence Analysis
[0102] In this section, the convergence of AirPCA is quantified based on the descent speed and convergence accuracy in different types of regions. The results are useful for designing power control. Figure 4A 、 Figure 4B and Figure 4C illustrate three example types of regions according to one or more embodiments described herein. Figure 4A illustrates an example of a non-stationary region type. Figure 4B illustrates an example of a saddle region type. Figure 4C illustrates an example of an optimal region type.
[0103] A. Definitions and Assumptions
[0104] For tractable analysis, several definitions and assumptions are given as follows. First, as discussed, the objective function F(W) of the PCA problem in (P1) contains discrete saddle points, a global optimum without local optima. Such functions belong to the family of strict saddle functions defined as follows.
[0105] Definition 1 (Strict Saddle Function): A twice-differentiable function \(F(W)\) is called - strictly saddle if at least one of the following conditions is true for any point \(W\): 2) Consider the Hessian matrix whose minimum eigenvalue for some positive constant \(\gamma\) 3) Let \(W^*\) be the global minimum of \(F(W)\), and \(\delta\) and \(\alpha\) be positive constants. In the \(\delta\)-neighborhood of \(W^*\), the function \(F(W)\) is \(\alpha\)-strongly convex, i.e.,
[0106] The above definition allows three types of regions of \(F(W)\) as shown in Figure 4A , Figure 4B and Figure 4C to be mathematically defined as follows.
[0107] Definition 2 (Region Type): A region of \(F(W)\) belongs to one of the following three types.
[0108] Denoted as \(R\) ns the non-stationary region (see Figure 4A ) is the region where condition 1) holds and can thus be defined as
[0109] Denoted as \(R\) sa the saddle region (see Figure 4B ) is the region where conditions 1) and 2) hold and can thus be defined as
[0110] Denoted as \(R\) op the globally optimal region (see Figure 4C ) is the region where condition 3) holds and can thus be defined as
[0111] For ease of handling, several assumptions can be made about \(F(W)\) that introduce additional properties that typically hold in practice.
[0112] Assumption 2: The function \(F(W)\) has several additional properties:
[0113] (1) (Boundedness) The function \(F(W)\) and its gradient norm are both bounded: for all \(W\) and some constants \(B\) and \(C\), \(\|F(W)\|\leq B\) and \(\|g(W)\|\leq C\).
[0114] (2) (Smoothness) The function \(F(W)\) is \(\beta\)-Lipschitz smooth: for some positive constant \(\beta\):
[0115] \|g(W_1)-g(W_2)\|\leq\beta\|W_1 - W_2\| \ (17)
[0116] (3) (Hessian Smoothness) The Hessian matrix of F(W) is X-Lipschitz smooth: for some positive constant X:
[0117]
[0118] B. Characterizing Gradient Descent in Different Regions
[0119] (1) Descent in the non-stationary region: The descent speed is measured by the expected reduction in the error function over a given number of rounds, which is called the expected error reduction. The descent speed in the non-stationary region is related to the received signal power and other parameters as follows.
[0120] Theorem 1 (Descent speed in the non-stationary region): Consider n rounds of gradient descent in the non-stationary region R ns where the corresponding principal component state and the received power is controlled as If the step size where β specifies the smoothness of the error function, the expected error reduction over n rounds can be lower bounded by:
[0121]
[0122] where,
[0123] First, it can be observed from Equation (19) that the expected error reduction is proportional to the order nμ of the descent distance. Next, the three terms enclosed in parentheses on the right-hand side of Equation (19) quantify the effects of the slopes of the error function, data noise, and channel noise, respectively, which are explained as follows. The first term is proportional to the square of the minimum slope ns of the error function in R . Since it is negative, the second term reduces the descent speed by an amount proportional to the data noise variance k 2 and inversely proportional to the expected number of devices performing air aggregation (i.e., Kζ act ). This latter scaling law is due to more precise distributed estimation resulting from a larger global dataset with more devices.
[0124] The last term regarding the channel noise effect is new in the literature on distributed PCA. It can be observed that the reduction in the descent speed due to channel noise is inversely proportional to , which can be interpreted as the expected received SNR per device. This is evident in the case of fixed received power , for all m, is reduced to On the other hand, over-the-air aggregation results in the expected amplitude of the aggregated signal at the server being 1 / 2 relative to the expected number of devices Kζ act Therefore, the expected SNR after aggregation increases proportionally (Kζ act ) 2 times, so that the channel noise term in equation (19) is reduced as an inverse function of the factor. In addition, as a sanity check, by setting the channel noise variance σ 2 = 0 and activation probability ζ act = 1, the result in Theorem 1 converges to the existing assumed reliable channel. This also applies to Theorem 2 and Theorem 3.
[0125] Based on the result in Theorem 1, we can conclude that we hope to increase the effective received signal power (i.e. ) to suppress the influence of channel noise. In particular, given the power sequence If another sequence Element-wise greater than but This results in a larger expected reduction in the error function over n epochs.
[0126] (2) Decline in the saddle region: The rate of decline in the saddle region is related to the received signal power and other parameters as follows.
[0127] Theorem 2 (Decline rate in the saddle region): Consider the saddle region R sa The n rounds of gradient descent in the corresponding principal component state are The limited received power is Define two constants and If the step size and number of rounds satisfy equation (20), then the expected error reduction over n rounds can be lower bounded as equation (21).
[0128]
[0129]
[0130] In the saddle area (see Figure 4B ), gradient descent may not be feasible in some dimensions (e.g., the dimension in which the error function is convex and the current point is the minimum), and descent is only possible in the dimension corresponding to the concave minimum eigenvalue. The result in Theorem 2 shows that if the step size is small enough and the number of rounds is large enough, the gradient perturbations caused by data and channel noise have the beneficial effect of guaranteeing the expected descent (or equivalently a strictly positive expected error reduction). This leads to a randomization of the descent direction due to the noise in the direction corresponding to a high probability of decline in the dimension. In the parentheses on the right side of equation (21), the first term and the last two terms respectively represent the positive effects of data and channel noise on the decline rate, contrary to their negative effects in the non-stationary region (see Theorem 1).
[0131] An important observation for power control that can be obtained from equation (21) is that by reducing the received signal power to enhance the channel noise, the desired error reduction is enhanced. Therefore, it is desirable to set the power to its minimum value, As a result, the bound of the desired error reduction can be simplified to:
[0132]
[0133] where
[0134] On the other hand, it should be emphasized that the received signal power should not be too low, because too strong noise may make the aggregated gradient (or equivalently, the descent direction) completely random, making it impossible to truly escape from the saddle point in the long term, that is, repeatedly returning to that point.
[0135] (3) Convergence possibility and accuracy: The results in Theorems 1 and 2 show that the gradient descent of AirPCA is not trapped in any non-stationary region or saddle region. Therefore, the descent path will eventually almost surely enter the optimal region, promoting learning convergence. The convergence possibility can be characterized mathematically by the following theorem, where the constants V max and N max follow the constants defined in Theorem 2.
[0136] Theorem 3: Consider n rounds of gradient descent of AirPCA from an arbitrary initial point, with the step size μ satisfying and Let ε N denote the event that the descent path enters the optimal region within N rounds: ε N ={there exists some n such that 0 ≤ n ≤ N - 1 and If N = mN max , and then the probability of ε N can be lower bounded by:
[0137]
[0138] where the constant and B is the upper bound of the error function norm.
[0139] Theorem 3 shows that if the step size μ is small enough and the number of rounds is large enough, then by ensuring Pr(ε N)Approaching one guarantees convergence in probability. Although the descent path may deviate from the optimal region due to unexpected strong noise, according to Theorem 3, the descent path will almost surely return to R op .
[0140] The standard analysis method of SGD can be applied to characterize the convergence accuracy. For example, if the number of rounds is large enough, in equation (15), the learned principal component W n and the optimal point W * The distance between (i.e., ||W n -W * || 2 ) is linearly proportional to , where ξ is the data plus channel noise sample in (15).
[0141] Region Adaptive Power Control
[0142] Based on the convergence analysis in the previous part, a region adaptive power control scheme for accelerating AirPCA is designed in this part. The scheme consists of two component schemes: online detection of the descent region and online power control. They are described sequentially in the following subsections.
[0143] A. Online Detection of the Descent Region
[0144] Online detection of the type of the current descent region is the key to implementing the proposed region adaptive power control scheme. The main challenge lies in detecting the saddle region caused by conflicts. Consider any round, say the nth round. On the one hand, from the definition of the region, the type of the region can be detected by estimating the minimum eigenvalue of the Hessian matrix (i.e., ) and evaluating the value of this minimum eigenvalue relative to a given negative constant -γ. If a saddle region is detected, the channel noise should be enhanced so that the descent path can escape from the saddle point and not be trapped at the saddle point. On the other hand, the estimation of the Hessian matrix H(W n ) is difficult. Specifically, the server at most knows a descent path, which only provides partial knowledge of H(W n ), but calculating its eigenvalues requires all knowledge. Due to the difficulty of detecting the saddle region based on its definition, we propose a simple and effective online detection scheme as described below. Again consider the nth round, in which it is found that the norm of the aggregated gradient is lower than a given threshold while the norm in the previous round is higher than This indicates that the descent path has entered a region that is either a saddle region or an optimal region. By default, this region is detected as a saddle region, and then the received signal power is reduced to amplify the noise effect, so as to make the path deviate from the saddle point. Given the reduced SNR, gradient descent continues for N0 rounds, where N0 is a design parameter. Then, with respect to the positive threshold f0, the expected error reduction obtained in N0 rounds is evaluated, that is If the detection of the saddle region is correct, then according to Theorem 2, the deviation from the saddle point will result in a significant error reduction, so Otherwise, the detection is incorrect, and this region should be an optimal region. Assuming that the reduced SNR is not too low such that the descent path remains within this region after N0 rounds, then power control is suitable for the optimal region to reduce the noise, thus ensuring a small error after convergence. Finally, the detection of the non-stationary region is straightforward, and the criterion is
[0145] The scheme for online descent region detection is summarized in Algorithm 1 below.
[0146] Algorithm 1: Online Descent Region Detection
[0147]
[0148] B. Online Power Control
[0149] Based on the above online region detection scheme, the principle of region adaptive power control is to reduce the received signal power when the descent path enters the saddle region, and increase the power when the path enters the non-stationary region or the optimal region. The former uses the channel noise to help the path deviate from the saddle point (see Theorem 2), while the latter overcomes the noise to approach the steepest descent (see Theorem 1).
[0150] Consider the case where the saddle region R sa is detected. Then, each device controls the truncated channel inversion in Equation (3) such that during the sojourn in R sa , the received signal power is fixed at the selected parameter Mathematically, for all For example, the parameter should be carefully selected using subsequent experiments As discussed, although should be low enough to utilize the noise effect, but being too low may endanger finding the correct descent path. Under the average power constraint in Equation (4), it is necessary to select such that it is less than the maximum average received power This saves the power used in other types of regions. Let N sa denote the number of rounds of descent within R sa Then the power saving is given as
[0151] Next, consider the case where a non - stationary region or an optimal region (denoted as R ns / op ) is detected. The power control strategy is the same for both types of regions. The key feature of the power control strategy is to consume the accumulated power savings in the current region on accelerating the descent. Let n0, n1, …, n N-1 denote the rounds within R ns / op , where N represents the total number of rounds. The accumulated power savings can be written as We propose to control the received signal power in the current region as where n0 ≤ n ≤ n N-1 . The coefficient is called the power consumption coefficient and can be set using one of the following two designs.
[0152] 1) Single - shot power savings consumption: When the descent path enters R ns / op , all the accumulated power savings are used in the first round, i.e., for n = n1, …, n N-1 , a n0 = 1 and a n = 0. In other words, for n = n1, …, n N-1 , and
[0153] 2) Gradual power savings consumption: The accumulated power savings are consumed according to in all rounds, where 0 ≤ j ≤ N - 1 and q ∈ (0, 1). Since if N is large or q is close to zero, all the accumulated power savings are consumed in R ns / op . Otherwise, only a part of the power savings is used while the remaining part is reserved for subsequent regions along the descent path.
[0154] Finally, it should be emphasized that the above - mentioned scheme for online power control guarantees meeting the average power constraint. In addition, the computational complexity of the power control scheme is
[0155] Example Results
[0156] Figures 5A to 8B shows the exemplary results that can optionally be achieved using some embodiments of the present disclosure. Figure 5A and 5B show the usefulness of channel noise for AirPCA to escape from the saddle point. Figure 6A and 6B show the learning performance comparison using different datasets. Figure 7A and7B Shows the impact of the power consumption coefficient on the learning performance of AirPCA with regional adaptive power control. Figure 8A And 8B Shows the impact of the number of devices and the truncation threshold on the learning performance of AirPCA with regional adaptive power control. Figures 5A to 8B Shows example results based on various training datasets, including the modified National Institute of Standards and Technology (MNIST) dataset, the Canadian Institute for Advanced Research 10-class (CIFAR-10) dataset, and the AR dataset (facial image database).
[0157] Reference Figure 5A And Figure 5B , by comparing AirPCA with regional adaptive power control, noiseless AirPCA, and AirPCA with fixed power, the effectiveness of channel noise for AirPCA to escape from the saddle point can be observed. The MNIST dataset was used. To demonstrate the benefit of channel noise, in Figure 5A , the curves of PCA error versus the number of rounds were plotted for AirPCA with channel noise and regional adaptive power control (labeled "AirPCA with power control") and AirPCA without channel noise (labeled "noiseless AirPCA"). The curve for centralized PCA was also plotted for comparison. After about 2000 rounds, it was observed that the principal components of the learned noisy AirPCA converged to those of the centralized PCA, while this was not possible in the noiseless case. The reason is that the (gradient) descent path of the former escaped from the saddle point with the help of channel noise, while the descent path of the latter was trapped at that point.
[0158] Next, in Figure 5B , the learning performances of AirPCA with regional adaptive power control, AirPCA with fixed power, and centralized PCA were compared, where the curves of PCA error versus the number of rounds were plotted. It can be observed that the proposed power control scheme effectively accelerates convergence compared to the case with fixed power. For example, to achieve a PCA error 7% higher than the centralized PCA level (i.e., an error of 5.6) (i.e., an error of 5.2), the learning delay is approximately 1170 rounds compared to 1740 rounds for AirPCA with fixed power, i.e., the learning delay is reduced by 33%.
[0159] In addition, in Figure 6A And Figure 6B , the two other datasets CIFAR-10 and AR were also used to compare the learning performances. As in the last comparison, the same observation can be obtained, i.e., regional adaptive power control accelerates convergence. Finally, it is worth mentioning that for the initial part of the descent process for MNIST (see FigureFigure 5A , 5B ) is relatively steeper compared to the initial part of the descent process for other datasets (see Figure 6A , 6B ). The reason is that the data samples in MNIST are black and white images of handwritten letters, and their data information is more concentrated in the subspace of the principal components than the data information of CIFAR-10 and AR, which are composed of color images and grayscale images respectively. Generally, the descent speed depends on the power distribution of the components, which varies for different datasets.
[0160] Refer to Figure 7A and Figure 7B , which shows two power consumption coefficient designs in terms of their impact on learning performance in the proposed scheme of region adaptive power control, namely single-shot power saving consumption ( Figure 7A ) and gradual power saving consumption ( Figure 7B ). Both the MNIST and CIFAR-10 datasets are used, and the descent step sizes are set to μ = 0.005 and μ = 0.02 respectively. It can be seen that the gradual consumption of power saving in the non-stationary region and the optimal region with optimized parameters (i.e., q = 0.8) achieves faster convergence than the single-shot scheme or the gradual scheme with replacement values of q (e.g., 0.5 or 0.995). It can be observed that their different impacts on convergence are located in the non-stationary region and the optimal domain, rather than in the saddle region where the signal power is not affected by the power consumption coefficient. In addition, the convergence accuracy is not affected.
[0161] Figure 8A and Figure 8B show the impact of other system parameters. Figure 8A shows the impact of the number of devices on the learning performance of AirPCA with region adaptive power control. Figure 8B shows the impact of the truncation threshold on the learning performance of AirPCA with region adaptive power control.
[0162] Considering AirPCA with region adaptive power control, in Figure 8AThe curves of PCA error against the number of rounds are plotted for different numbers (K = {10; 20; 50}) of devices. Each device is provided with 10 data samples randomly drawn from the dataset. Thus, the total data used in AirPCA / Centralized PCA is proportional to the number of devices. We use the CIFAR-10 dataset for the experiments with a step size of 0.02. For the cases of K = {20; 50}, with a larger number of devices, the learning performance is better. On the other hand, when the number is small (e.g., K = 10), the SGD-based AirPCA fails to converge due to the combined effect of limited data for suppressing channel noise and insufficient aggregation gain (see Equation (16)). In contrast, the centralized PCA using SVD does not encounter such a problem. A possible solution to prevent divergence is to reduce the step size in AirPCA at the cost of slower convergence.
[0163] Next, we study the effect of the channel truncation threshold G in Equation (3) on the learning performance of AirPCA with region-adaptive power control. To this end, in Figure 8B the curves of PCA error against the number of rounds are plotted for various values of the truncation threshold G (G = {0.001; 0.2; 0.5}) for the CIFAR-10 dataset. Note that G controls the expected ratio of the truncated subchannels. It can be seen that setting G too small or too large results in divergence. The former is due to too small received signal power under the constraint of amplitude alignment on the active subchannels for aerial aggregation (see Equation (3)), and the latter is due to too many truncated subchannels, which severely distort the uploaded local gradients. This suggests the need to optimize G in some embodiments.
[0164] Figure 9 is a flowchart showing an example operation of a device participating in joint principal component analysis according to various aspects and embodiments of the present disclosure. The shown boxes may represent actions performed in a method, functional components of a computing device, or instructions implemented in a machine-readable storage medium executable by a processor. Although the operations are shown in an example order, in some embodiments, the operations may be eliminated, combined, or reordered.
[0165] Figure 9 The shown operations may be performed, for example, by a computing device such as Figure 1The apparatus 101 of the apparatus shown performs, and the apparatus 101 may include a first apparatus among a group of other apparatuses (such as apparatuses 102, 103). Example operation 902 includes receiving, by the first apparatus 101, an updated matrix from the server 110, where the first apparatus 101 is a participant in the joint principal component analysis. Example operation 904 includes determining, by the first apparatus 101, a local gradient with respect to the updated matrix, where the local gradient is based on local data stored at the first apparatus 101. Example operation 906 includes modulating, by the first apparatus 101, the local gradient using linear analog modulation to produce a modulated local gradient. Example operation 908 includes adjusting, by the first apparatus 101, the transmission power to produce an adjusted transmission power. Example operation 910 includes receiving, by the first apparatus 101, synchronization information to synchronize the first wireless signal with the second wireless signal. Example operation 912 includes transmitting, by the first apparatus 101, the modulated local gradient to the server 110 via the first wireless signal, where the first wireless signal includes the adjusted transmission power, and where the first wireless signal is synchronized with a second wireless signal transmitted by a second apparatus (e.g., apparatus 102).
[0166] In some embodiments according to Figure 9 adjusting the transmission power at operation 908 includes, for example, reducing the transmission power in response to detecting a potential saddle region at the server 110. The reduction of the transmission power may be performed to increase the noise in the first wireless signal, so as to enable escape from the potential saddle region detected at the server 110. As disclosed herein, adjusting the transmission power at operation 908 may further include increasing the transmission power to use the power saved previously by reducing the transmission power.
[0167] In some embodiments according to Figure 9 the illustrated method may be performed in multiple repetition cycles. The method may be repeated to enable determination at the server 110 of a low-dimensional subspace containing information from high-dimensional data (including local data stored at the first apparatus 101 and other local data stored at other apparatuses 102, apparatus 103).
[0168] In some embodiments according to Figure 9In some embodiments, the operation can be allowed to participate in joint principal component analysis by device 101 among multiple devices 101, 102, and 103, where the joint principal component analysis includes multiple communication rounds, and each communication round of the multiple communication rounds includes simultaneous wireless transmissions from the multiple devices 101, 102, and 103 to the server 110. The operation of each communication round can further include receiving an updated matrix from the server 110, determining a local gradient with respect to the updated matrix (where the local gradient is based on local data stored at device 101), modulating the local gradient using linear analog modulation to generate a modulated local gradient, adjusting the transmit power to generate an adjusted transmit power, and simultaneously wirelessly transmitting the modulated local gradient to the server 110 via a wireless signal, where the wireless signal includes the adjusted transmit power, and where the wireless signal is simultaneous with multiple wireless signals transmitted by the multiple devices 101, 102, and 103.
[0169] Figure 10 is a flowchart showing an example operation of a server participating in joint principal component analysis according to various aspects and embodiments of the present disclosure. The illustrated blocks can represent actions performed in a method, functional components of a computing device, or instructions implemented in a machine-readable storage medium executable by a processor. Although the operations are shown in an example order, in some embodiments the operations can be eliminated, combined, or reordered.
[0170] Figure 10 The operations shown in can be performed, for example, by a server such as Figure 1 the edge server 110 shown in. Example operation 1002 includes receiving an aggregated signal including a global gradient. The global gradient includes a combination of respective local gradients calculated at respective devices 101, 102, and 103, where the respective local gradients are simultaneously wirelessly transmitted by the respective devices 101, 102, and 103 for in-air combination of the respective local gradients to form the aggregated signal.
[0171] Example operation 1004 includes updating a matrix based on the global gradient to generate an updated matrix. Example operation 1006 includes determining a region type associated with the global gradient, e.g., a non-stationary region, a saddle region, or an optimal region.
[0172] Example operation 1008 includes determining power adjustments for application by respective devices 101, 102, 103 based on zone type. The power adjustments can include, for example, reducing the transmission power applied by respective devices 101, 102, 103. The reduction in transmission power is responsive to the zone type including a potential saddle zone. The reduction in transmission power can effect an increase in the noise included in the wireless transmissions of respective devices 101, 102, 103, where the increase in noise enables disengagement from the potential saddle zone. Alternatively, the power adjustments can include, for example, increasing the transmission power in order to enable respective devices 101, 102, 103 to use power previously saved by reducing the transmit power.
[0173] Example operation 1010 includes sending updated matrix, power adjustment, and synchronization information to respective devices 101, 102, 103. The synchronization information can synchronize the simultaneous wireless transmission of respective local gradients performed by respective devices 101, 102, 103.
[0174] With Figure 9 the same, Figure 10 the operations can be performed in multiple repetition cycles according to a defined frequency. The method can be repeated to effect PCA calculations, for example, calculating a low-dimensional subspace containing information included in high-dimensional data (respective local data stored at respective devices 101, 102, 103).
[0175] To provide context for various aspects of the disclosed subject matter, Figure 11 and the following discussion is intended to provide a brief general description of a suitable environment in which the various aspects of the disclosed subject matter can be implemented. Although the subject matter has been described herein in the general context of computer-executable instructions of a computer program running on one and / or more computers, those skilled in the art will recognize that the disclosed subject matter can also be implemented in combination with other program modules. Generally, program modules include routines, programs, components, data structures, etc. that perform particular tasks and / or implement particular abstract data types.
[0176] In this specification, terms such as "store", "storage", "database", and substantially any other information storage component related to the operation and functionality of a component refer to a "memory component" or an entity embodied in a "memory" or a component that includes the memory. It will be understood that the memory components described herein can be volatile memory or non-volatile memory, or can include both volatile and non-volatile memory, by way of illustration and not limitation, volatile memory 1120, non-volatile memory 1122, disk storage 1124, solid state memory devices, and memory storage 1146. Additionally, non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM), which acts as an external cache. By way of illustration and not limitation, RAM can take many forms, such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct Rambus RAM (DRRAM). Additionally, the disclosed memory components of the systems or methods herein are intended to include, but are not limited to, these and any other suitable types of memory.
[0177] Furthermore, it will be noted that the disclosed subject matter can be implemented with other computer system configurations, including single-processor or multi-processor computer systems, minicomputers, mainframe computers, and personal computers, handheld computing devices (e.g., PDAs, telephones, watches, tablet computers, netbook computers), microprocessor-based or programmable consumer or industrial electronic products, and the like. The aspects shown can also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network. However, some aspects (if not all aspects) of the disclosed subject matter can be practiced on an independent computer. In a distributed computing environment, program modules can be located in local and remote memory storage devices.
[0178] Figure 11A block diagram of a computing system 1100 is shown. The computing system is configured, for example, to operate as a controller 818 and is operable to execute the disclosed systems and methods according to an embodiment. A computer 1112 can be, for example, part of the hardware of system 1100, which includes a processing unit 1114, a system memory 1116, and a system bus 1118. The system bus 1118 couples system components, including but not limited to the system memory 1116, to the processing unit 1114. The processing unit 1114 can be any of a variety of available processors. Dual microprocessors and other multiprocessor architectures can also be used as the processing unit 1114.
[0179] The system bus 1118 can be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus or external bus, and / or a local bus using any of a variety of available bus architectures, including but not limited to Industry Standard Architecture (ISA), Micro Channel Architecture (MSA), Extended ISA (EISA), Intelligent Drive Electronics, VESA Local Bus (VLB), Peripheral Component Interconnect (PCI), Card Bus, Universal Serial Bus (USB), Advanced Graphics Port (AGP), Personal Computer Memory Card International Association Bus (PCMCIA), FireWire (IEEE1494), and Small Computer System Interface (SCSI).
[0180] The system memory 1116 can include volatile memory 1120 and non-volatile memory 1122. A Basic Input / Output System (BIOS) can be stored in the non-volatile memory 1122, which contains routines for transferring information between elements within the computer 1112 during startup. By way of illustration and not limitation, the non-volatile memory 1122 can include ROM, PROM, EPROM, EEPROM, or flash memory. The volatile memory 1120 includes RAM, which serves as an external cache. By way of illustration and not limitation, RAM has many forms, such as SRAM, Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus Direct RAM (RDRAM), Direct Rambus Dynamic RAM (DRDRAM), and Rambus Dynamic RAM (RDRAM).
[0181] The computer 1112 can also include removable / non-removable, volatile / non-volatile computer storage media. For example, Figure 11Disk storage 1124 is shown. Disk storage 1124 includes, but is not limited to, devices such as disk drives, floppy disk drives, tape drives, flash memory cards, or memory sticks. Additionally, disk storage 1124 may include storage media, either alone or in combination with other storage media, which include, but are not limited to, optical disk drives such as compact disk ROM devices (CD-ROM), CD recordable drives (CD-R drives), CD rewritable drives (CD-RW drives), or digital versatile disk ROM drives (DVD-ROM). To facilitate connection of disk storage 1124 to system bus 1118, a removable or non-removable interface such as interface 1126 is typically used.
[0182] Computing devices generally include a variety of media, which may include computer-readable storage media or communication media, where these two terms are used differently from one another as follows herein.
[0183] Computer-readable storage media can be any available storage media accessible by a computer and includes volatile and non-volatile media, removable and non-removable media. By way of example and not limitation, computer-readable storage media can be implemented in conjunction with any method or technology for storing information such as computer-readable instructions, program modules, structured data, or unstructured data. Computer-readable storage media can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD ROM, digital versatile disks (DVD) or other optical disk storage, magnetic tape cartridges, magnetic tape, disk storage or other magnetic storage devices, or other tangible media that can be used to store the desired information. In this regard, the term "tangible" as applied to storage, memory, or computer-readable media herein should be understood to exclude only propagating intangible signals per se as a modifier and not to forego coverage of all standard storage, memory, or computer-readable media that do not consist solely of propagating intangible signals. In one aspect, tangible media can include non-transitory media, where the term "non-transitory" as applied to storage, memory, or computer-readable media herein should be understood to exclude only propagating transitory signals per se as a modifier and not to forego coverage of all standard storage, memory, or computer-readable media that do not consist solely of propagating transitory signals. To avoid doubt, the term "computer-readable storage device" is used and defined herein to exclude transitory media. Computer-readable storage media can be accessed by one or more local or remote computing devices, for example, via an access request, query, or other data retrieval protocol, for performing various operations on the information stored by the media.
[0184] A communication medium typically embodies computer-readable instructions, data structures, program modules, or other structured or unstructured data in a data signal such as a modulated data signal (e.g., a carrier wave or other transmission mechanism), and includes any information delivery or transmission medium. The term "modulated data signal" or signal refers to a signal in which one or more characteristics are set or changed in order to encode information in one or more signals. By way of example, and not limitation, communication media include wired media (such as, for example, a wired network or direct-wired connection) and wireless media (such as, acoustic, RF, infrared, and other wireless media).
[0185] It can be noted that Figure 11 Software is described that acts as an intermediary between a user and computer resources described in a suitable operating environment 1100. Such software includes an operating system 1128. The operating system 1128 may be stored on disk storage 1124 for controlling and allocating resources of the computer system 1112. It should be noted that the disclosed subject matter may be implemented with a variety of operating systems or combinations of operating systems.
[0186] System applications 1130 utilize the operating system 1128 for resource management through program modules 1132 and program data 1134 stored in system memory 1116 or disk storage 1124. In some embodiments, a gas sensor control application 1131 may control operations described in connection with Figure 11 the operations described to perform gas sensor measurements and identify the measured gas or gas concentration. The gas sensor control application 1131 may use components of a gas-sensitive FET array as described herein to control the measurements, and may record the measurement data as data 1134.
[0187] A user may input commands or information to the computer 1112 through an input device 1136 (including, for example, a fingertip pointer as described herein). By way of example, a mobile device and / or a portable device may include a user interface implemented in a touch-sensitive display panel that allows the user to interact with the computer 1112. The input device 1136 includes, but is not limited to, pointer devices such as a mouse, trackball, stylus, touchpad, etc., a keyboard, a microphone, a joystick, a gamepad, a satellite dish, a scanner, a TV tuner card, a digital camera, a digital video camera, a web camera, a cellular phone, a smart phone, a tablet computer, etc. These and other input devices are connected to the processing unit 1114 through a system bus 1118 through an interface port 1138. The interface port 1138 includes, for example, a serial port, a parallel port, a game port, a universal serial bus (USB), an infrared port, a Bluetooth port, an IP port, or a logical port associated with a wireless service, etc. Some of the same types of ports used by the input device 1136 are used by the output device 1140.
[0188] Thus, for example, a USB port can be used to provide input to the computer 1112 and output information from the computer 1112 to an output device 1140. An output adapter 1142 is provided to illustrate that there are some output devices 1140, such as monitors, speakers, and printers, as well as other output devices 1140 that use special adapters. By way of illustration and not limitation, the output adapter 1142 includes a graphics card and a sound card that provide connection means between the output device 1140 and the system bus 1118. It should be noted that other devices and / or systems of devices provide input and output capabilities, such as a remote computer 1144.
[0189] The computer 1112 can operate in a networked environment using a logical connection to one or more remote computers, such as the remote computer 1144. The remote computer 1144 can be a personal computer, server, router, network PC, cloud storage, cloud service, workstation, microprocessor-based appliance, peer device, or other common network node, etc., and generally includes many or all of the elements described with respect to the computer 1112.
[0190] For simplicity, only the memory storage 1146 is shown together with the remote computer 1144. The remote computer 1144 is logically connected to the computer 1112 through a network interface 1148 and then physically connected to the computer 1112 through a communication connection 1150. The network interface 1148 includes wired and / or wireless communication networks, such as local area networks (LANs) and wide area networks (WANs). LAN technologies include Fiber Distributed Data Interface (FDDI), Copper Distributed Data Interface (CDDI), Ethernet, Token Ring, etc. WAN technologies include, but are not limited to, point-to-point links, circuit-switched networks like Integrated Services Digital Network (ISDN) and its variants, packet-switched networks, and Digital Subscriber Line (DSL). As described below, wireless technologies can be used to supplement or replace the above technologies.
[0191] The communication connection 1150 refers to the hardware / software used to connect the network interface 1148 to the bus 1118. Although, for clarity of illustration, the communication connection 1150 is shown inside the computer 1112, it can also be outside the computer 1112. The hardware / software used to connect to the network interface 1148 can include, for example, internal and external technologies, such as modems, including conventional telephone-grade modems, cable modems, and DSL modems, ISDN adapters, and Ethernet cards.
[0192] The foregoing description of the disclosed embodiments (including the description in the abstract) is not intended to be exhaustive or to limit the disclosed embodiments to the precise forms disclosed. While specific embodiments and examples are described herein for illustrative purposes, various modifications within the scope of such embodiments and examples are possible as would be recognized by those of ordinary skill in the relevant art.
[0193] In this regard, while the disclosed subject matter has been described in conjunction with various embodiments and the corresponding drawings, it should be understood that, where applicable, other similar embodiments may be used or modifications and additions may be made to the described embodiments to perform the same, similar, alternative, or substitute functions of the disclosed subject matter without departing from the disclosed subject matter. Accordingly, the disclosed subject matter should not be limited to any single embodiment described herein, but rather should be construed in accordance with the breadth and scope of the appended claims.
[0194] As used in this specification, the term "processor" can refer to substantially any computing processing unit or device, including but not limited to a single-core processor, a single processor with software multithreading execution capabilities, a multi-core processor, a multi-core processor with software multithreading execution capabilities, a multi-core processor with hardware multithreading technology, a parallel platform, and a parallel platform with distributed shared memory. Additionally, a processor can refer to an integrated circuit, an application specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic controller (PLC), a complex programmable logic device (CPLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor can employ nanoscale architectures such as, but not limited to, molecular and quantum dot based transistors, switches, and gates, in order to optimize space usage or enhance performance of a user device. A processor can also be implemented as a combination of computing processing units.
[0195] In this specification, terms such as "storage," "memory," "data storage," "data memory," "database," and substantially any other information storage component associated with the operation and functionality of a component refer to a "memory component" or an entity embodied in a "memory" or a component that includes the memory. It will be understood that the memory components described herein can be volatile memory or non-volatile memory, or can include both volatile memory and non-volatile memory.
[0196] As used in this application, the terms "component", "system", "platform", "layer", "selector", "interface", etc. are intended to refer to a computer-related entity or an entity related to an operating device having one or more specific functions, where the entity can be hardware, a combination of hardware and software, software, or software in execution. By way of example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable, an execution thread, a program, and / or a computer. By way of illustration and not limitation, an application running on a server and the server can both be components. One or more components can reside within a process and / or an execution thread, and a component can be located on one computer and / or distributed between two or more computers. Additionally, these components can execute from various computer-readable media, device-readable storage devices, or machine-readable media storing various data structures. These components can communicate via local and / or remote processes, such as in accordance with a signal having one or more data packets (e.g., data from one component that interacts with another component in a local system, a distributed system, and / or across a network (e.g., the Internet) via the signal with other systems). As another example, a component can be a device having a specific function provided by a mechanical component operated by an electrical or electronic circuit, where the electrical or electronic circuit is operated by a software or firmware application executed by a processor, and the processor can be internal or external to the device and executes at least a portion of the software or firmware application. As yet another example, a component can be a device that provides a specific function through an electronic component rather than a mechanical component, and the electronic component can include a processor therein to execute software or firmware that at least partially imparts the function of the electronic component.
[0197] Furthermore, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified, or clear from the context, "X employs A or B" is intended to mean any natural inclusive permutation. That is, if X uses A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied in any of the foregoing instances. Additionally, unless otherwise specified or clear from the context to refer to the singular form, the articles "a" and "an" as used in this specification and the drawings shall generally be construed to mean "one or more".
[0198] While the invention is susceptible to various modifications and alternative constructions, certain illustrative implementations of the invention are shown in the drawings and have been described in detail above. However, it should be understood that the invention is not intended to be limited to the specific forms disclosed, but on the contrary, the invention will cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the invention.
[0199] In addition to the various implementations described herein, it should be understood that other similar implementations may be used, or modifications and additions may be made to the described implementations to perform the same or equivalent functions of the corresponding implementations without departing from the corresponding implementations. Therefore, the present invention is not limited to any single implementation, but should be construed in accordance with the breadth, spirit, and scope of the appended claims.
Claims
1. A method for accelerating distributed principal component analysis using a noisy channel, comprising: Receiving, by a first device, an updated matrix from a server, wherein the first device participates in joint principal component analysis; Determining, by the first device, a local gradient with respect to the updated matrix, wherein the local gradient is based on local data stored at the first device; Modulating, by the first device, the local gradient using linear analog modulation to generate a modulated local gradient; Adjusting, by the first device, the transmission power to generate an adjusted transmission power; and Transmitting, by the first device, the modulated local gradient to the server via a first wireless signal, wherein the first wireless signal includes the adjusted transmission power, and wherein the first wireless signal is synchronized with a second wireless signal transmitted by a second device, wherein adjusting the transmission power includes: reducing the transmission power in response to detecting a potential saddle region at the server.
2. The method according to claim 1, wherein Performing the step of reducing the transmission power so as to increase the noise in the first wireless signal, thereby enabling to escape from the potential saddle region detected at the server.
3. The method according to claim 1, wherein, The step of adjusting the transmission power further includes: increasing the transmission power in response to detecting a non-stationary region or an optimal region at the server, so as to use the power saved previously by reducing the transmission power.
4. The method according to claim 1, wherein Performing the method in a plurality of repeating cycles.
5. The method according to claim 1, wherein Repeating the method to be able to determine, at the server, a low-dimensional subspace containing information from high-dimensional data, the high-dimensional data including the local data stored at the first device and other local data stored at other devices.
6. The method according to claim 1 further comprises: Receiving, by the first device, synchronization information for synchronizing the first wireless signal with the second wireless signal.
7. A server device configured to participate in joint principal component analysis, the server device comprising: A processor; And A memory storing executable instructions that, when executed by the processor, cause operations to be performed, the operations including: Receiving an aggregated signal including a global gradient, wherein the global gradient includes a combination of corresponding local gradients calculated at respective devices, wherein the corresponding local gradients are wirelessly transmitted simultaneously by the respective devices to perform an in-air combination of the corresponding local gradients to form the aggregated signal; Updating a matrix based on the global gradient to generate an updated matrix; Determining a region type associated with the global gradient; Determining a power adjustment applied by the respective devices based on the region type; and Sending the updated matrix and the power adjustment to the respective devices, wherein the power adjustment includes: reducing the transmission power applied by the respective devices in response to the region type including a potential saddle region.
8. The server device according to claim 7, wherein, Reducing the transmission power causes an increase in the noise included in the wireless transmissions of the respective devices, and wherein the increase in the noise enables to escape from the potential saddle region.
9. The server device according to claim 7, wherein, The power adjustment further includes: in response to the region type including a non-stationary region or an optimal region, increasing the transmission power so that each of the devices can use the power saved by previously reducing the transmission power.
10. The server device according to claim 7, wherein, The operation is performed in a plurality of repeated cycles according to a predetermined frequency.
11. The server device according to claim 7, wherein, The operation is repeated to enable calculation of a low-dimensional subspace containing information from high-dimensional data, the high-dimensional data including corresponding local data stored at each of the devices.
12. The server device according to claim 7, wherein, The operation further includes: sending synchronization information for synchronizing simultaneous wireless transmissions of the corresponding local gradients of the respective devices.
13. A non-transitory machine-readable medium includes executable instructions that, when executed by a processor, cause an operation to be performed, the operation including: Devices among a plurality of devices participate in joint principal component analysis, wherein the joint principal component analysis includes a plurality of communication rounds, wherein each of the plurality of communication rounds includes a simultaneous wireless transmission from the plurality of devices to a server, and wherein the operation of each communication round further includes: Receiving an updated matrix from the server; Determining a local gradient with respect to the updated matrix, where the local gradient is based on local data stored at the device; Modulating the local gradient using linear analog modulation to generate a modulated local gradient; Adjusting the transmission power to generate an adjusted transmission power; Simultaneously wirelessly transmitting the modulated local gradient to the server via a wireless signal, where the wireless signal includes the adjusted transmission power, and where the wireless signal occurs simultaneously with a plurality of wireless signals transmitted by the plurality of devices, and where adjusting the transmission power includes: in response to detecting a potential saddle region at the server, reducing the transmission power.
14. The non-transitory machine-readable medium according to claim 13, wherein, Performing an operation of reducing the transmission power so as to increase the noise in the wireless signal, thereby enabling avoidance of the potential saddle region detected at the server.
15. The non-transitory machine-readable medium according to claim 13, wherein, Adjusting the transmission power further includes: in response to detecting a non-stationary region or an optimal region at the server, increasing the transmission power to use the power saved by previously reducing the transmission power.
16. The non-transitory machine-readable medium according to claim 13, wherein, The joint principal component analysis enables determination of a low-dimensional subspace containing information from high-dimensional data at the server, the high-dimensional data including local data stored at the plurality of devices.