Quantum-secure multiparty deep learning
The quantum-secure multiparty deep learning protocol uses a coherent linear algebra engine to ensure secure data transmission and computation, maintaining high accuracy and confidentiality by leveraging quantum properties, addressing security risks in cloud-based deep learning.
Patent Information
- Application Number
- PCT/US2025/040025
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2025-07-31
- Publication Date
- 2026-02-05
AI Technical Summary
Current deep learning implementations are computationally intensive, requiring cloud-based servers that pose significant security risks, especially in sensitive domains like healthcare, due to vulnerabilities in data privacy and security.
A quantum-secure multiparty deep learning protocol using a coherent linear algebra engine that leverages quantum properties of light to ensure secure data transmission and computation, ensuring data and model weights remain confidential by exploiting the no-cloning theorem, allowing clients to perform inference without revealing sensitive information.
Maintains high classification accuracy (95% or higher) while ensuring robust security, guaranteeing that client data and server weights remain secure, addressing fundamental security challenges in multiparty computation.
Smart Images

Figure US2025040025_05022026_PF_FP_ABST
Abstract
Description
Attorney Docket No. MIT-26054WO01 Quantum-Secure Multiparty Deep Learning CROSS-REFERENCE TO RELATED APPLICATION(S)
[0001] This application claims the priority benefit, under 35 U.S.C.119(e), of U.S. Application No.63 / 677,972, which was filed on July 31, 2024, and is incorporated herein by reference in its entirety for all purposes. BACKGROUND
[0002] Although deep learning has revolutionized multiple domains, its applications are constrained by the increasing computational demands on the hardware. Given the high energy consumption required for state-of-the-art deep neural network (DNN) inference, it is common to delegate inference workloads from the edge to centralized server clusters. Unfortunately, this paradigm introduces vulnerabilities that compromise data security, which is detrimental in applications such as business, finance, and healthcare. This situation is reminiscent of a central challenge in secure computation, where multiple parties perform a joint evaluation of multivariate functions across distributed resources while preserving the privacy of their inputs.
[0003] Modern secure computation schemes are built on homomorphic encryption, which allows universal computation on encrypted data and preserves the security of the input and output of the computation. Recent developments have adapted state-of-the-art homomorphic encryption schemes for secure machine learning, but their real-world applications are limited due to their massive computational overhead and security vulnerabilities. Moreover, these encryption schemes typically depend on computational assumptions and are not information- theoretically secure.
[0004] Information-theoretically secure computation has been studied extensively through the use of quantum primitives, including superposition and entanglement, and has encompassed various protocols such as bit commitment, coin flipping, two-party function evaluation and oblivious transfer, multiparty quantum computation, and blind quantum computing. Despite these significant advancements, their practical applicability for real-world computational tasks, particularly in the domain of deep learning, remains elusive due to strict requirements on the quantum hardware.Attorney Docket No. MIT-26054WO01 SUMMARY
[0005] Here, we introduce a linear algebra engine for secure computation supported by quantum information theorems. Our engine can be used for multiparty deep learning. Applied to the MNIST classification task, hardware simulations for our engine obtain test accuracies of more than 95% while leaking fewer than 0.1 bits per weight symbol and 0.1 bits per data symbol. This weight leakage is an order of magnitude below the state-of-the-art minimum bit precision desired for deep learning.
[0006] Our engine uses a photonic computation architecture to realize efficient optical matrix- vector multiplication. It can be used to execute a quantum-secure multiparty deep learning process in an edge computing setting where DNN weights are streamed from a server to clients that perform inference. This process shields both client data from exposure to the server and shields the weights from exposure to the clients. In this process, each client performs inference using the weights and then returns residual light to the server as a certificate of the security of the weights. Leveraging the quantum nature of light, the server can compute an upper-bound on their weight leakage using the Holevo bound and the client can upper-bound their data leakage using the Cramér–Rao bound. This approach enables double-blind operations, which form the cornerstone of our quantum-secure multiparty deep learning process.
[0007] In some aspects, the techniques described herein relate to a method of quantum-secure, multiparty deep learning inference, the method including: receiving, from a server at a client, a first sequence of coherent pulses of light encoded with a weight vector for a deep-learning model, each coherent pulse in the first sequence of coherent pulses of light having a photon number of no more than 10; modulating, at the client, the first sequence of coherent pulses of light with an input vector to the deep-learning model to yield a second sequence of coherent pulses of light; performing, at the client, a homodyne measurement of a mode of the second sequence of coherent pulses of light, the homodyne measurement representing an inner product of the weight vector and the input vector; modulating, at the client, the second sequence of coherent pulses of light to produce a third sequence of coherent pulses of light encoded with the weight vector and having a variance greater than a variance of the first sequence of coherent pulses of light and below a threshold based on the photon number; and transmitting, by the client, the third sequence of coherent pulses of light to the server for verification that the client did not copy the weight vector.Attorney Docket No. MIT-26054WO01
[0008] In some aspects, the techniques described herein relate to a method, wherein receiving the first sequence of coherent pulses of light includes receiving between 1 photon and 6 photons per coherent pulse from the server.
[0009] In some aspects, the techniques described herein relate to a method, wherein: modulating the first sequence of coherent pulses of light includes performing a first unitary transformation on the first sequence of coherent pulses of light, and modulating the second sequence of coherent pulses of light includes performing a second unitary transformation on the second sequence of coherent pulses of light, the second unitary transformation being a conjugate transpose of the first unitary transformation.
[0010] In some aspects, the techniques described herein relate to a method, wherein performing the homodyne measurement includes interfering the mode of the second sequence of coherent pulses of light and a local oscillator signal.
[0011] In some aspects, the techniques described herein relate to a method, wherein performing the homodyne measurement includes filtering the second sequence of coherent pulses of light with a narrowband filter.
[0012] In some aspects, the techniques described herein relate to a method, further including: amplifying, with an optical amplifier, the second sequence of coherent pulses at a gain selected to achieve a desired classification accuracy of a deep learning inference computation.
[0013] In some aspects, the techniques described herein relate to a method, further including: amplifying, with an optical amplifier, the second sequence of coherent pulses at a gain selected to limit leakage of the input vector and / or of the weight vector.
[0014] In some aspects, the techniques described herein relate to a method, further including: measuring, at the server, the variance of the third sequence of coherent pulses of light.
[0015] In some aspects, the techniques described herein relate to a client for quantum-secure, multiparty deep learning inference, the client including: a first modulator to modulate a first sequence of coherent pulses of light encoded with a weight vector for a deep-learning model with first modulation based on an input vector to the deep-learning model to yield a second sequence of coherent pulses of light, each coherent pulse in the first sequence of coherent pulses of light having a photon number of no more than 10; a homodyne detector, in optical communication with the first modulator, to perform a homodyne measurement of a mode of the second sequence of coherent pulses of light, the homodyne measurement representing anAttorney Docket No. MIT-26054WO01 inner product of the weight vector and the input vector; and a second modulator, in optical communication with at least one of the first modulator or the homodyne detector, to modulate the second sequence of coherent pulses of light with second modulation based on the input vector to produce a third sequence of coherent pulses of light encoded with the weight vector and having a variance greater than a variance of the first sequence of coherent pulses of light and below a threshold based on the photon number.
[0016] In some aspects, the techniques described herein relate to a client, wherein the first sequence of coherent pulses of light includes between 1 photon and 6 photons per coherent pulse.
[0017] In some aspects, the techniques described herein relate to a client, wherein the first modulation represents a first unitary transformation and the second modulation represents a second unitary transformation, the second unitary transformation being a conjugate transpose of the first unitary transformation.
[0018] In some aspects, the techniques described herein relate to a client, further including: a local oscillator, in optical communication with the homodyne detector, to generate a local oscillator signal for performing the homodyne measurement.
[0019] In some aspects, the techniques described herein relate to a client, further including: a narrowband filter, in optical communication with the first modulator, the homodyne detector, and the second modulator, to direct the mode of the second sequence of coherent pulses of light from the first modulator to the homodyne detector and to direct at least one other mode of the second sequence of coherent pulses of light to the second modulator.
[0020] In some aspects, the techniques described herein relate to a client, further including: an optical amplifier, in optical communication with the first modulator, to amplify the second sequence of coherent pulses of light.
[0021] In some aspects, the techniques described herein relate to a client, wherein the optical amplifier has a gain selected to achieve a desired classification accuracy of a deep learning inference computation.
[0022] In some aspects, the techniques described herein relate to a client, wherein the optical amplifier has a gain selected to limit leakage of the input vector and / or of the weight vector.Attorney Docket No. MIT-26054WO01
[0023] In some aspects, the techniques described herein relate to a system including: the client; and a server, in optical communication with the client, to transmit the first sequence of coherent pulses of light to the client.
[0024] In some aspects, the techniques described herein relate to a system, wherein the server is further configured to receive the third sequence of coherent pulses of light from the client and to verify that the client did not copy the weight vector based on the variance of the third sequence of coherent pulses of light.
[0025] In some aspects, the techniques described herein relate to a system, wherein the client is a first client, the inner product is a first inner product, and further including: a second client, in optical communication with the first client and the server, to receive the third sequence of coherent pulses of light from the first client, to compute a second inner product with the weight vector, and to return a fourth sequence of coherent pulses of light to the server for verification that neither the first client nor the second client copied the weight vector.
[0026] In some aspects, the techniques described herein relate to a system, wherein the client is a first client, the inner product is a first inner product, and further including: a second client, in optical communication with the first client and the server, to receive a fourth sequence of coherent pulses of light encoding the weight vector from the server, to perform a second inner product with the weight vector, and to return a fifth sequence of coherent pulses of light to the server for verification that the second client did not copy the weight vector.
[0027] All combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are part of the inventive subject matter disclosed herein. All combinations of claimed subject matter appearing at the end of this disclosure are part of the inventive subject matter disclosed herein. The terminology used herein that also may appear in any disclosure incorporated by reference should be accorded a meaning most consistent with the concepts disclosed herein. BRIEF DESCRIPTIONS OF THE DRAWINGS
[0028] The skilled artisan will understand that the drawings primarily are for illustrative purposes and are not intended to limit the scope of the inventive subject matter described herein. The drawings are not necessarily to scale; in some instances, various aspects of the inventive subject matter disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate an understanding of different features. In the drawings, like referenceAttorney Docket No. MIT-26054WO01 characters generally refer to like features (e.g., elements that are functionally and / or structurally similar).
[0029] FIG.1A shows quantum-secure, deep-learning inference by a client using weights from a server.
[0030] FIG. 1B is another view of quantum-secure, deep-learning inference by a client that uses weights from a server and provides a verification state to the server for verifying that the weights remain secure.
[0031] FIG. 2A is a block diagram of a server and client suitable for quantum-secure, multiparty deep-learning.
[0032] FIG. 2B is a plot of the real parts of the time-domain input waveform ^⃗⃗^ , output waveform⃗⃗⃗^⃗^′ (return or verification state), input waveform after filtering (magnified 50 times for clarity).
[0033] FIG. 2C is a plot of information leakage versus filter bandwidth for encoding bandwidths of 25 GHz (upper trace) and 50 GHz (lower trace).
[0034] FIG.3 is a block diagram of a server and clients with time-domain and spatial-domain encoding and decoding for quantum-secure, multiparty deep-learning.
[0035] FIG.4A is a plot of classification accuracy versus weight and data leakage for a secure optical neural network that uses one of the coherent linear algebra engines (clients) shown in FIGS.1B or 3.
[0036] FIG. 4B is a plot of classification accuracy versus physical scaling parameter in quantum-secure, multiparty deep-learning.
[0037] FIG.5A is a plot of information leakage versus round-trip channel loss for both weights (upper trace) and data (lower trace) in quantum-secure, multiparty deep-learning.
[0038] FIG. 5B is a plot of information leakage versus neurons per layer for both weights (upper trace) and data (lower trace) in quantum-secure, multiparty deep-learning.
[0039] FIG.5C is a plot of information leakage versus technical noise and round-trip channel loss in quantum-secure, multiparty deep-learning.Attorney Docket No. MIT-26054WO01
[0040] FIG. 6 illustrates quantum-secure deep learning with multiple clients. The dark and light shading represent the communication patterns of symmetric and asymmetric protocols, respectively.
[0041] FIG.7 is a plot of classification accuracy of the client versus physical scaling parameter. DETAILED DESCRIPTION
[0042] Deep-learning models are being used in many fields, from health care diagnostics to financial forecasting. However, current implementations of these models are so computationally intensive that they require the use of powerful cloud-based servers. This reliance on cloud computing poses significant security risks, particularly in areas like health care, where hospitals may be hesitant to use AI tools to analyze confidential patient data due to privacy concerns.
[0043] The present technology includes a security protocol that leverages the quantum properties of light to guarantee that data sent to and from a cloud server remain secure during deep-learning computations. By encoding data in light (e.g., laser light used in fiber optic communications systems), this security protocol exploits the fundamental principles of quantum mechanics, making it impossible for attackers to copy or intercept the information without detection. Moreover, inventive techniques guarantee security without compromising the accuracy of the deep-learning models—they can maintain accuracies of 95% or higher while ensuring robust security.
[0044] Our protocol can be used for secure inference or for secure training. For example, our coherent linear algebra engine (client) can guarantee a secure inner product for federated learning, where multiple clients collaborate to train a model while ensuring that their data remain secure. Our approach addresses a fundamental security challenge in multiparty computation, paving the way for information-theoretical security at various stages of the machine-learning pipeline. The introduction of our protocol enables the secure application of deep learning on private datasets and models, establishing a rigorous standard of secure computation in sectors such as finance, healthcare and business.
[0045] FIGS. 1A and 1B illustrate how an inventive security protocol can be employed for quantum-secure, cloud-based deep learning inference involving two parties—in this case, a client 110 with confidential data, like medical images, and a central server 120 that controls a deep learning model (e.g., a DNN). The client 110 wants to use the deep-learning model toAttorney Docket No. MIT-26054WO01 make a prediction, such as whether a patient has cancer based on medical images, without revealing information about the patient. At the same time, the server 120 does not want to reveal the weights of the deep-learning model to the client 110.
[0046] In digital computation, a bad actor could easily copy the data sent from the server or the client. Quantum information, on the other hand, cannot be copied perfectly per the no- cloning theorem, which holds that it is impossible to create an independent and identical copy of an arbitrary unknown quantum state. The inventive security protocol exploits the no-cloning theorem to ensure that the client 110 cannot copy the weights that it receives from the server 120 by encoding the weights in the quantum state of a coherent pulse of light. (As used herein, the term “state” (e.g., as in “quantum state”) is the complete description of how the light is arranged and the term “mode” (e.g., as in “optical mode”) is an optical channel (e.g., a spatial, temporal, or polarization channel).) Instead of measuring all the incoming light from the server, the client 110 measures only the light that is necessary to run the deep neural network. Then the client 110 sends the residual light back to the server 120 for security checks. Due to the no- cloning theorem, the client 110 unavoidably introduces tiny errors into the quantum-state representations of the weights encoded in the residual light while measuring its result. When the server 120 receives the residual light from the client 110, the server 120 can measure these errors to determine if any information was leaked. This residual light does not reveal the client data, ensuring that both the weights and the client data remain secure.
[0047] In more concrete terms, illustrated in FIG.1B, the client 110 (e.g., at a hospital) holds sensitive data ^^ , and the server 120 holds a confidential DNN with weights ^^. To perform secure inference, the server 120 encodes the weights ^^ of a DNN into the complex amplitude of a coherent pulse of light (top), represented in FIG.1B as coherent states ^^^^, with a variance of 1 shot noise unit (SNU). (A shot-noise unit is a standard energy unit in quantum cryptography. One SNU is the built-in background noise of empty light (vacuum): the half- photon of zero-point energy in a single optical mode.) The client 110 uses these weights for inference with local data ^^ and transmits the residual light, called the return state or verificationstate, back to the server 120. The client 110 calculates the inner product ^⃗⃗^ ⋅ ^^ using: (i) a firstunitary transformation ^^^̂^, (ii) measurement and feedforward of the complex amplitude ^⃗⃗^ ⋅ ^̂^,which adds excess noise to the last mode, and (iii) a second unitary transformation ^^† ^̂^, which is the conjugate transpose of the first unitary transformation After applying ^^† ^̂^, the excess noise from the measure and feedforward step is spread over the output modes (right). The return stateAttorney Docket No. MIT-26054WO01has the same mean as the original DNN weight ^^^^ but a higher variance of (1 + ^^^^) SNU,where ^^^^is the excess noise. Only the client 110 learns the prediction; the server 120 observes no output. The return state allows the server 120 to certify that the weights have not been copied without affecting the inference outcome.
[0048] The secure inference problem addressed by our quantum-secure, multiparty inference technology can be formalized as follows. Consider a ^^-class classification task with trainingand tests sets: (^^^^ , ^^^^) ^^^^=1 and labelsTheserver 120 holds a DNN specified by weight matrices= 1,… , ^^.
[0049] For an input ^^^^, the layer‐wise activations arewhere ^^:ℝ → ℝ is a public nonlinearity (e.g., a ReLU applied coordinate‐wise). The predictedclass for ^^^^isWe measure performance by the classification accuracy on the test set: ^^+^^ Acc =1 ^^∑ ^^ {^̂^(^^^^) = ^^^^},^^=^^+1 where ^^{⋅} is the indicator function. The system leaks strictly less than ℐ^^bits of information about the server’s weights and less than ℐ^^bits about the client’s data while achieving aclassification accuracy of Acc(ℐ^^, ℐ^^).Coherent Linear Algebra Engine
[0050] Our quantum-secure, multiparty deep learning protocol implements a DNN in a layer- by-layer fashion using a coherent optical linear algebra engine utilized by the two parties, the client 110 and the server 120, in a network. To implement the ^^-th layer, the server 120 holds matrix of dimensions ^^^^×^^^^−1while the client 110 holds the activation vector ^^(^^(^^−1)) of length. For ease of exposition, the notation here is independent of layer since the procedure is the same for all ^^. We denote ^^ by ^^, ^^^^−1by ^^,by ^^, ^^(^^(^^−1)) by ^^ .Attorney Docket No. MIT-26054WO01
[0051] The matrix-vector product ^^^^ is computed through ^^ inner products ^^^^ ⋅ ^^ for ^^ ∈{1, … ,^^}, each of which involves three steps:1. The server 120 transmits a weight vector ^^^^of length ^^ by encoding the components into the complex amplitudes of ^^ coherent states denoted jointly by ^⃗⃗^ . The amplitudes of these physical states, ^^^^, encode the numerical weights as ^^^^= √^^^^^^^^ / ∥ ^^ ∥, where ∥ ^^ ∥ denotes the root mean square of the elements in ^^. Thisencoding ensures that the average photon number is ^^, a value chosen by the server. These coherent states have a variance of 1 SNU in each quadrature, as illustrated in FIG.1B. 2. The client 110 operates on these amplitudes as follows as shown in FIG.1B: (i) The client 110 calculates the inner product of the incoming weight vector ^⃗⃗^ with its local input data vector ^^ using a unitary transformation ^^^̂^based on ^^ that yields ^⃗⃗^ ⋅ ^̂^ into one mode of ^^^̂^^⃗⃗^ , where ^̂^ is obtained by the ℓ2normalization of ^^ . This mode is called the result mode; (ii) The client 110 diverts the result mode, and measures both quadratures of the result mode to obtain ^⃗⃗^ ⋅ ^̂^. Then, the client 110 performs a feed-forwardprotocol by reinjecting light with the same quadratures back into the result mode. The feed-forwarded state preserves the measured state’s expectation value ^⃗⃗^ ⋅ ^̂^ but, due to quantum noise, includes additional Gaussian noise ofvariance 1 SNU. The output of this measure and feed-forward step is labeled(iii)The client 110 applies the unitary transformation ^^† ^̂^based on ^^ to ℳ(^^^̂^^⃗⃗^ ) and returns ^^^^ = ^^†^̂^ℳ(^^^̂^^⃗⃗^ ) as a return state, also called a verification state, to the server 120 for security checks. By the definition of ℳ(^^^̂^^⃗⃗^ ) from the previous point, the first moment of ^^†^̂^ℳ(^^^̂^^⃗⃗^ ) is ^⃗⃗^ , while the variance of each mode ^^ has increased by ^^^^, see Fig.1(c). 3. Finally, the server 120 measures the variance of each mode ^^ of the return state in both quadratures, (1 + ^^^^) SNU, and calculates an upper bound for the leakage ofAttorney Docket No. MIT-26054WO01 each weight ^^^^to the client 110. Before this security check, the return state can be re- used by other clients to perform their own inference as discussed below.
[0052] The measure and feed-forward operation ℳ outlined above can be generalized to an operation ^^ that includes a combination of phase-insensitive amplification ^^ by an opticalamplifier and a beam splitter with a split ratio 1 −1 1the former approach is obtained fromthe latter in the limit ^^ ≫ 1. In this general case, the client 110 applies a gain ^^ to thecomponent ^⃗⃗^ ⋅ ^̂^, passes it through a beam splitter with the above splitting ratio, and uses theoutput of the 1 −1 ^^ port for its inner product computation readout while reinjecting the output 1 of the ^^port back into the result mode that originally carried ^⃗⃗^ ⋅ ^̂^. To obtain the exact innerproduct, the client 110 may scale the result digitallyThis scaling by the client, however, is superfluous for classification tasks.
[0053] The gain ^^ contributes amplification noise to the result mode. This variance is larger than standard shot noise by a quantity called the excess noise. This excess noise is spread via the operation of ^^† ^̂^in the return states modes, such that the excess noise in the ^^-th mode of the return state, ^^^^, is weighted by|^̂^^^|2:where ^̂^^^is the ^^th element of the normalized data vector ^̂^.
[0054] The signal-to-noise ratio (SNR) in the client’s measurement after applying theamplification and beam splitting is SNR =|^⃗⃗ |2 | |2^ ⋅ ^̂^ . Since ^⃗⃗^ is proportional to theaverage photon occupation number ^^, the SNR is proportional to ^^.
[0055] The gain ^^ is controlled by the client 110 and the average weight occupation ^^ is controlled by the server 120. These parameters dictate the return state excess noise, ^^^^, and theSNR of the inner product measurement. In the large gain limit ^^ ≫ 1, the excess noise, ^^^^, andthe SNR reduce to 2^̂^^21 ^ and 2|^⃗⃗^ ⋅ ^̂^|2, respectively. This result aligns with the measure andfeedforward special case. In the small gain limit ^^ ≈ 1, the excess noise ^^^^ and the SNR reduceto 2(^^ − 1)^̂^2^^ and (^^ − 1)|^⃗⃗^ ⋅ ^̂^|2, respectively. Reducing the gain disturbs the return stateless and reduces the SNR of the inner product obtained by the client 110. Optical Client and Server ImplementationsAttorney Docket No. MIT-26054WO01
[0056] There are several ways to implement the client 110 and the server 120. The client 110 can have a gain knob (amplifier) for controlling data leakage, as shown in FIG.3, or it can omit the gain knob, as shown in FIG. 2A, and use spectral filtering to introduce noise into the verification / return state.
[0057] FIGS.2A–2C illustrate a first implementation of the client 110 and the server 120. In this server 120 includes a laser 122, an optional attenuator 124, an in-phase / quadrature (I / Q) modulator 126 (either before or after the attenuator 124), one or more beam splitters or couplers 127, and a differential photodetector 128. The client 110 includes an encoding module with a push-pull Mach-Zehnder modulator 212 and a circulator 214; a detection module with a reflective narrowband or notch filter 216 (e.g., implemented by introducing a phase shift within the stop band of a fiber Bragg grating to create a sharp resonance), local oscillator laser 217 (e.g., a laser), and a differential photodetector 218; and a decoding module with another push- pull Mach-Zehnder modulator 219.
[0058] In operation, the server encodes and transmits its deep learning model weights (e.g., DNN weights) as a sequence of coherent pulses of light, where the weights are represented by an amplitude-modulated coherent state of the photons in the coherent pulses with just a few photons each (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 photons per pulse). To generate these coherent pulses, the laser 122 emits coherent light, either as a continuous-wave (cw) beam or as coherent pulses. If this coherent light is too intense, it is attenuated by the attenuator 124 to reduce the number of photons that are (or will be) in each coherent pulse. This attenuator 124 can be fixed or variable, with an attenuation depending on the laser’s output intensity and the desired average weight occupation ^^. The I / Q modulator 126 modulates the coherent light to generate a train or sequence of coherent pulses with complex amplitudes encoding the DNN weights (e.g., a train of 20 picosecond-long Gaussian pulses spaced by 200 ps, with complex amplitudes encoding 100 real-valued DNN weights). The intensity of the laser output, the attenuation, and the modulation parameters (e.g., pulse duration and pulse amplitude) can be selected to generate pulses with the average weight occupation ^^. The couplers 127 tap a portion of the coherent pulse train for checking the return or verification state as described below and transmit the rest to the client 110.
[0059] The client 110 receives the coherent pulse train and feeds it into the encoding Mach- Zehnder interferometer 212, which includes a beam splitter, two phase shifters (PS), and a second beam splitter. Driving the upper and lower arms of the encoding Mach-ZehnderAttorney Docket No. MIT-26054WO01interferometer 212 with ^^1,2 = arg(^^ ) ± sin−1(|^^|)is equivalent to performing the desired unitary transformation and modulates the coherent pulses to produce an output at the upper port of this MZI 212 that is proportional to the element-wise product of the weights and theinput vector, ^⃗⃗^ ⋅ ^^ . The circulator 214 directs this output, which is itself a coherent pulse train,to the narrowband filter 216, which transmits one spectral mode of the coherent pulse train— the portion of the coherent pulse train within a very narrow passband (e.g., a passband of 100 MHz at a center wavelength of 1550 nm)—to a balanced photodetector 218, which detects the homodyne interference of the transmitted coherent pulses and a local oscillator beam from a local oscillator laser 217. The spectral mode of the coherent pulse train transmitted by the narrowband filter 216 interferes with a LO signal from the LO signal 217 at a beam splitter whose outputs are coupled to the client’s differential photodetector 218, which measures one quadrature of the resulting interference. Integrating the output of the differential photodetector218 over time, e.g., with an integrating capacitor, yields the inner product ^⃗⃗^ ⋅ ^^ .
[0060] The narrowband filter 216 reflects the out-of-band portion of the coherent pulse train back to the circulator 214, which directs the other spectral modes of the coherent pulse train— the out-of-band portion of the coherent pulse train—to the second Mach-Zehnder interferometer 219. Returning fewer than all of the spectral modes of the coherent pulse train effectively adds noise in the time domain, which asymptotically vanishes as the number of coherent pulses increases. The second Mach-Zehnder interferometer 219 inverts the action of the first Mach-Zehnder interferometer 212, producing a verification state or return state⃗⃗⃗^⃗^′ that is noisier than the weight vector ^⃗⃗^ because it is missing a portion of the spectral information encoded in the weight vector ^⃗⃗^ . The client 110 sends the verification state back to the server 120, which makes a balanced homodyne measurement of the verification state interfered with the reserved portion of the outgoing coherent pulses encoded with the weights ^⃗⃗^ . If the weights have not been copied and the outgoing coherent pulses ^⃗⃗^ have a variance of 1 SNU per pulse, then the verification state⃗⃗⃗^⃗^′ should be similar to ^⃗⃗^ for a large number of pulses. This allows the server 120 to verify that the client 110 has not gained any additional information beyond what it uses for inference.
[0061] FIG. 2B is a plot that illustrates the effect of filtering the output of the first Mach- Zehnder interferometer 212 with the narrowband filter 216. It shows the real part of the input waveform ^⃗⃗^ , as well as the output vector⃗⃗⃗^⃗^′ and the transmission from the filter, magnified by a factor of 50 for clarity. The output waveform closely follows the original signal, andAttorney Docket No. MIT-26054WO01 integrating the real value of the filtered waveform yields the desired inner product. The client’s classification accuracy is independent of the filter’s passband because any wavelength other than the optical carrier wavelength effectively vanishes during real-valued integration, making no contribution to the result.
[0062] FIG. 2C shows the information leakage for different values of the filter bandwidth σ and for encoding bandwidths of BW = 25 GHz and BW = 50 GHz, corresponding to pulse durations of 0.44 / BW under the assumption of transform-limited Gaussian pulses. The information leakage increases with the filter bandwidth and decreases with the modulation bandwidth. Furthermore, to prevent data leakage over multiple uses of the model, the server 120 uses invariants of the DNN, ensuring that different sets of individual weights still represent the same overall function.
[0063] FIG.3 illustrates two alternative optical client implementations 310a and 310b suitable for performing quantum-secure multiparty deep learning with the server 120, where the client can control the data leakage by changing the gain. Again, the I / Q modulator 126 in the server 120 modulates the neural network weights onto a train of weak coherent states (pulses) produced by attenuating the output of a continuous-wave laser 122 with an attenuator 124 to the few-photon limit. The server’s differential photodetector 128 measures the modulated quadratures of the incoming return or verification state using homodyne detection with a reference local oscillator, which may be tapped off the output coherent pulses or generated with a separate LO laser 322. The optical power difference between the two output arms is proportional to either the in-phase or quadrature component of the return / verification state, depending on the phase of the local oscillator signal.
[0064] The client 310a implements the unitaries ^^ † ^̂^and ^^^̂^for encoding and decoding the weight signals from the server 120 in the time domain encoding using optical loops 312a connected with optical switches 306 and Mach-Zehnder interferometers 308. The other client 310b implements the unitaries †and ^^^̂^in the spatial domain with a mesh of interferometers 312b or free-space multi-plane light converters. Time-domain implementations are promising for large-scale universal optical information processing because the number of elements does not scale with the number of optical modes. Optical loss may limit arbitrary unitaries in both the time and spatial domains to a few dozen modes.
[0065] Fortunately, our quantum-secure, multiparty deep learning does not require the fullprogrammability of an arbitrary ^^ × ^^ unitary. Conventional photonic meshes realize productsAttorney Docket No. MIT-26054WO01^^^^ by encoding ^^ ∈ ^^(^^) in ^^(^^2) tunable phase shifters while injecting a length-^^ vector^^ optically. In contrast, our process involves encoding the client’s input ^^ with ^^(^^)tunable elements and then injecting a length-^^ weight vector ^⃗⃗^ (one row of ^^) so that the mesh outputsthe inner product ^⃗⃗^ ⋅ ^^ . If hardware constraints limit the number of optical modes, the innerproduct can be partitioned into ^^ sub-products of length ^^ / ^^, each processed on existing meshes and summed digitally; although this increases the excess noise—and hence the leakage bound—linearly with ^^, it markedly reduces the physical resource requirements. Finally, the engine works natively with complex numbers by packing each real length-^^ vector into a length-^^ / 2 complex vector.
[0066] In the time-domain client 320a, optical modes are defined by temporally separate coherent pulses of light. The first unitary operation ^^^̂^is based on ^̂^ and implemented using a Mach-Zehnder interferometer 308a (e.g., with electro-optic phase shifters) and a fiber loop 312a. This structure interferes successive pulses from the incoming pulse train ^⃗⃗^ with a pulse in the fiber loop 312a that contains the running sum of the inner product, resulting in the finalinner product ^⃗⃗^ ⋅ ^̂^ being written into the complex amplitude of the last output pulse (temporalmode). The light in this mode is then amplified by a factor of ^^ using a phase-insensitive amplifier 309a, such as an erbium-doped fiber amplifier. The gain can be controlled dynamically by sending only some of the pulses (e.g., the last pulse, or temporal mode) through the amplifier 309a. A beam splitter divides the amplified light from the amplifier 309a with asplitting ratio of 1 −1 1 ^^ : ^^ for the transmitted and reflected ports. The transmitted port is then measured by a differential photodetector 320a using homodyne detection in the real quadratures to estimate the desired inner product. The reflected port is fed into a nested Mach- Zehnder fiber loop 314a that implements the unitary ^^† ^̂^, also based on ^̂^, and routes the resultant return state back to the server 120.
[0067] In the spatial-domain client 320b, optical modes are defined by spatially separated waveguides. Here, the client 320b implements the same unitaries ^^ † ^̂^and ^^^̂^based on the input vector ^̂^ via two series of Mach-Zehnder interferometers 312a and 314a, respectively, weighted according to the elements of the input vector ^̂^. These Mach-Zehnder interferometers (e.g., with thermo-optic phase shifters) mix spatially adjacent modes to shift the desired inner product^⃗⃗^ ⋅ ^^ into the top mode, which is then amplified with an optical amplifier 309b. One spatialmode of the amplified pulse train is measured via coherent homodyne detection by a differential photodetector 320b. The client 310b transmits the other mode(s) emitted by theAttorney Docket No. MIT-26054WO01 uppermost Mach-Zehnder interferometer as this return state to the server 120, which measures the excess noise to compute the leakage of its weights. Quantum Description Phase-Insensitive Amplifier
[0068] We use the quantum electromagnetic field ladder operators (or annihilation and creation operators), denoted by ^̂^ and ^̂^†, respectively, to analyze the behavior of the system at the quantum level. The action of these operators on the photon number states |^^^ is given by: ^̂^|^^^ = √^^|^^ − 1^ and ^̂^†|^^^ = √^^ + 1|^^ + 1^
[0069] The input-output relationship of the ladder operators in the Heisenberg picture for a phase-insensitive amplifier with gain ^^ is :where ^̂^inis the annihilation operator for the input mode, ^̂^ampis the annihilation operator for the amplified output mode, and ^̂^vis the annihilation operator for the auxiliary vacuum noise mode introduced during amplification.
[0070] The quadrature operators are defined as: ^̂^ = ^̂^ + ^̂^†, ^̂^ =1 (†)^^^̂^ − ^̂^The variance for an operator ^̂^ is given by:Then, the following is true for the input mode: 2 ^(^^^̂^in)^ = 1, ^(^^^̂^in)2 ^= 1For the output mode of the phase-insensitive amplifier (e.g., amplifiers 309a and 309b in FIG. 3):Substituting the expression for ^̂^amp:^̂^amp = √^^^̂^in − √^^ − 1^̂^vAttorney Docket No. MIT-26054WO01 The variance of the quadratures after amplification is:This shows that the noise is amplified along with the signal, and for large gain ^^, the noise variance scales linearly with ^^. Weighted Beam Splitter
[0071] In both clients 310, the Mach-Zehnder interferometer 308 between the amplifier 309 (gain fiber) and the detector 320 acts as a variable beam splitter with a splitting ratio of(1 −1 1 ^^ : ^^). In the spatial domain client 310b, the input modes for the beam splitter are the amplified output ^̂^ampand a vacuum mode, labeled ^̂^vac. The weighted beam splitter transform yields:Substituting the above expressions for ^̂^amp, ^̂^amp:and similarly for ^̂^out2.
[0072] Since ^̂^vacis the quadrature of the vacuum mode, it has unit variance:= 1.
[0073] Two vacuum modes have been introduced: one during amplification to preserve the commutation relation of the amplified output mode, and one at the dark input port of the beam splitter. The two vacuum modes are independent and uncorrelated, so they can be summed inquadrature. One can see this directly by considering the reparameterization: √^^ ↦ cosh(^^)and √^^ − 1 ↦ sinh(^^) (note that cosh2(^^) = sinh2(^^) + 1 for any ^^). Their addition inquadrature forms a single vacuum noise mode with unit variance with a coefficient determined by the Pythagorean theorem sum of their composite variances.
[0074] Alternatively, we can derive the noise variance directly as follows:Attorney Docket No. MIT-26054WO01Th excess noise on the amplified signal after beam splitting is thus (2 −
[0075] For ^̂^out1:and similarly for ^̂^out1. Thus,
[0076] The optical power at the client’s different photodetector 320, after gain and splitting, is^^ = ^^ (1 − |^⃗⃗^ ⋅ ^̂^|2. The SNR of the client’s measurement is therefore:SNR = ^^(^^ − 1)2^^2 − |^⃗⃗^ ⋅ ^̂^|23^^ + 2Return State Excess Noise
[0077] First, we consider the effect of the overall unitary transform pair {^^^̂^, ^^^†̂^} on the variance of the return state. The initial input ^⃗⃗^ is a set of coherent states with a unit shot-noise variance in each mode, that is, the noise itself is uncorrelated and isotropic. We describe the noise in these ^^ modes and the correlations between them, through the covariance matrix oftheir quadratures ^^^^. The (^^, ^^)-th element of this matrix is given by= ^^[^^^^^^^^] −
[0078] Since a coherent state is equivalent to a displacement operation on the vacuum state, the covariance matrix of ^⃗⃗^ is ^^^^^^ = ^^. Since we want the first output mode (the result mode)of ^^ ^† ^̂^⃗⃗^ to carry the inner product ^⃗⃗^ ⋅ ^̂^, ^^^̂^ and ^^^̂^take the form:where ^̂^ is represented as a row vector and ^^ is a (non-unique) rectangular matrix which is chosen to ensure unitarity of ^^^̂^.
[0079] The covariance matrix of ^^^̂^^⃗⃗^ is:Attorney Docket No. MIT-26054WO01Next, we apply the amplify-and-split operation ^^(⋅) on the vector of modes ^^^̂^^⃗⃗^ . This operation involves applying a phase-insensitive amplification with gain ^^ on the result modeand then splitting it with the ratio 1 −1 1 1: ^^. The output of the ^^ port is fed back to the mesh that implements ^^† ^̂^. The operation ^^(⋅)leaves all the other modes of ^^^̂^^⃗⃗^ untouched. The resultant covariance matrix after amplification and splitting is the diagonal matrix∼diag(^^2, 1,1, … ,1), where ^^2 = 1 + (2 −is the variance of the result mode.
[0080] Finally, we operate with the transform ^^† ^̂^:
[0081] Since ^̂^^̂^†is a dense matrix, the various components of the return state ^^†^̂^^^(^^^̂^^⃗⃗^)are all correlated with one another. However, the variance of the ^^-th output mode only contains information from ^̂^^^; recalling that the excess variance of the ^^-th output mode is the excess noise ^^^^yields:Classification Accuracy
[0082] In this section, we calculate the classification accuracy of a secure neural network that uses our coherent linear algebra engine on the standard MNIST classification task. To this end, we first trained a digital noiseless neural network on the MNIST dataset and obtained a test classification accuracy of 98%. Then we fed the trained weights into a PyTorch model of our optical architecture to evaluate the test accuracy of our secure optical neural network.
[0083] We calculate the classification accuracy as a function of the average photon number occupation ^^ and the amplification gain ^^. We obtain 95% test accuracy with an averagephoton number of less than ^^ = 6 per weight and an amplification gain of ^^ = 1.3 (see“Classification Accurecy Calculations” below for more details).
[0084] In contrast to previous studies in optical machine learning, we do not measure all of thelight sent to the client but only the portion corresponding to |(^⃗⃗^ ⋅ ^̂^)|2 and yet obtain similarhigh accuracies. This follows from the fact that the SNR in the inner product is similar whetherAttorney Docket No. MIT-26054WO01 one optically routes the inner product amplitude to one mode and measures it versus measuring all the modes and calculating the inner product digitally. Security Analysis
[0085] Below, we analyze the security of the server’s information, assuming an honest server and a malicious client. We also analyze the security of the client’s data, assuming an honest client and a malicious server. These two results allow us to analyze the classification accuracyas a function of the information leakage of each party Acc(ℐ^^, ℐ^^).Weights Leakage
[0086] For assessing leakage of the weights, we use the return state ^^^^returned to the server by the client to calculate an upper bound on the amount of information about ^⃗⃗^ , in number of bits, that could be learned by a malicious client. A malicious client may try to learn the weights using arbitrary operations other than the operations defined by our protocol.
[0087] Here, we assume that dishonest operations are independent and identically distributed over all the incoming modes of ^⃗⃗^ . Our security analysis for the weight leakage relies on results for quantum secure communication and proceeds through the Holevo theorem.
[0088] We reduce the weight leakage analysis in the secure computation problem to key information leakage analysis in secure communication by mapping our two computational parties, namely the server and the client, to the three communication parties, namely Alice (transmitter), Bob (reviewer), and Eve (eavesdropper). Alice and Bob are both on the server side, while Eve represents the client. Alice is the part of the server that sends the weights. The client is either honestly using the transmitted information (the weights) to perform some computations or maliciously trying to extract extra information from them, so from the perspective of the server, the client acts like an eavesdropper, Eve. Bob is the part of the server that receives the return state to check if Eve obtained more information than they should have in an honest inner-product computation.
[0089] In secure communication, Bob checks for any leakage of the message sent by Alice, while in our secure computation protocol, Bob checks if the leakage exceeds what is to be expected from an honest inner-product computation. Any malicious client deviating from the protocol will be immediately detected by the Bob part of the server, which can abort the computation immediately. For example, if the client measures the weights, they will be unableAttorney Docket No. MIT-26054WO01 to provide a low-noise return state and, consequently, the server can detect the client’s dishonesty immediately.
[0090] We present here the mathematical result for the weight leakage. Representing the weight occupation by ^^ and defining the quantities:the weight leakage ℐwis bounded by the Holevo theorem (see “Weights Leakage Calculations” below for more details): (bits),where:This inequality is shown to be tight by analyzing the entangling cloner attack.
[0091] The leakage is calculated for one query, that is, a single exposure of the client to the weights. To prevent information leakage after multiple uses of the model, the server sends the client shuffled and scaled of the DNN model, which include different individual weights but implement the same model function (therefore, conserving the inference output). Intuitively, this method associates with each query a one-time pad that was randomly selected from all of the transformations under which the function is invariant. Because of this invariance, the client can run the computation without knowing the one-time pad and will never obtain the same model twice. Data Leakage
[0092] To calculate the data leakage, we consider an honest client and a dishonest server. Inthis setting, the server sends a weight matrix ^^ of shape ^^ × ^^ to the client and aims toevaluate the client’s data vector ^^ of length ^^ using ^^ measurements of the ^^ return states.
[0093] Mode ^^ of the return state ^^†^^^^(^^^^ ^⃗⃗^ ) has a mean ^^^^ and variance (1 + ^^^^ ). In otherwords, information about the client’s data ^^ is leaked to the server not through the mean of the return state but through its variance. The client may control this leakage by reducing the gainAttorney Docket No. MIT-26054WO01 ^^ in its amplification-and-splitting step. An honest client announces the true gain ^^ to the server.
[0094] We phrase dishonest attempts of the server to learn ^^ through ^^ measurements as the following concrete statistical estimation problem: Let ^^^^, the return state sent back to the server by the client, be a Gaussian random vector with mean ^⃗⃗^ and a covariance matrix whose ^^-th diagonal element is ^^2^^ = 1 + ^^^^ = 1 + (2 −⋅ |^̂^^^|2. What is the amount of information, in terms of bits, that the server can obtain about ^^ from ^^ independent measurements of ^^^^, given that it knows ^^?
[0095] We answer this question by using the Cramér–Rao inequality to compute a lower bound on the variance ^^^̂^^^of the server’s estimator ^̂^^^of the true ^̂^^^: ^^^̂^^^where ^^(^̂^^^)is the Fisher information of the true normalized data element ^̂^^^.
[0096] Next, we assume that the server’s estimator ^̂^^^is a Gaussian random variable with mean ^̂^^^and variance given the above lower bound and use the formula for the communication capacity of a Gaussian channel to calculate the amount of mutual information ℐ^^, in number of bits, between the true unnormalized ^^^^and the estimator ^̂^^^. By the data processing inequality, the information ℐ^^^^leaked about ^^^^is upper-bounded by the information ℐ^̂^leaked about ^̂^^^. In the case where the server has access to quantum operations, we use the quantum Cramér–Rao inequality to obtain the leakage bound.
[0097] Next, we turn to bound the leakage for collective attacks, where the server is allowed to perform fully correlated measurements across all returned modes. The ^^ identical outputsconstitute an ^^-copy Gaussian state with covariance ^^ + (2 −a rank-one spike addedto the identity. The inverse covariance follows from the Sherman–Morrison formula, and substituting it into the Gaussian-state quantum Fisher information expression yields a closed- form Quantum Fisher Information (QFI) that is isotropic on the N-1 dimensional tangent space of the unit sphere. In the asymptotic limit, where the dimensions are high, the entropies of the uniform prior on the sphere and the posterior are approximated by those of Gaussians with the same covariance, so the data leakage upper bound reduces to:Attorney Docket No. MIT-26054WO01 so an eavesdropper, even with optimal collective measurements, gains only a bounded amount of information about each element of the secret direction. (Please see “Data Leakage Calculations” below for more details.) Secure Classification
[0098] Now that we have a handle on both the weight leakage and the data leakage, we can see how they trade off against each other for a constant test accuracy. For this purpose, we revisit the classification accuracy that we computed numerically as a function of the two hardware configuration parameters of the system, the server average photon occupation per weight and the client gain ^^. The weight leakage and the data leakage are both functions of the hardware configuration parameters. The data leakage is calculated for the quantum adversary capable of correlated measurement.
[0099] FIG. 4A is a plot of the classification accuracy of the secure optical neural network, which uses our coherent linear algebra engine, versus weight and data leakage. This classification accuracy is numerically calculated as a function of the average photon occupation per weight and the amplification gain. These parameters serve as upper bounds on the weight leakage ℐ^^^^and the data leakage ℐ^^^^. The classification accuracy increases with both leakages and achieves the digital noiseless accuracy. For small data leakages, the weight leakage is inversely proportional to the data leakage for any given fixed classification accuracy. Ourprotocol achieves a classification accuracy of > 95% while leaking less than ℐ^^ = 0.1 bits perweight symbol, less than ℐ^^ = 0.1 bits per data symbol.
[0100] FIG. 4A depicts a clear tradeoff of the data leakage and weight leakage to maintain classification accuracy. This is because, when the server sends less energy per weight in order to limit weight leakage, the client has to increase the gain to preserve accuracy of the computation, thereby exposing more of its data. The data leakage increases monotonically with the gain and saturates, achieving the measure and feed-forward leakage. Therefore, as the server laser power is reduced and the client gain is increased, one moves from right to left on the accuracy isocontours in FIG.4A. For example, the 95% classification accuracy isocontourshows a weight leakage of less than ℐ^^ = 0.1 bits per symbol and data leakage of less thanℐ^^ = 0.1 bits per symbol.
[0101] FIG.4B is a plot of the classification accuracy versus amplification gain, which is a hardware parameter of the client, and the average photon occupation (photon number) for each coherent pulse, which is a hardware parameter of the server. The amplification gain is setAttorney Docket No. MIT-26054WO01 by the amplifier (and optical loss) in the client and the average photon occupation is set by the laser output power, attenuation, and modulation (pulse amplitude and duration) at the server. The classification accuracy increases with gain and average photon occupation (photon number) and asymptotically achieves digital noiseless accuracy. At low gain values, the gain is inversely proportional to the photon occupation number for any given fixed classification accuracy.
[0102] Reliable low-precision inference typically takes at least 8 bits per weight or as little as 1 bit using state-of-the-art quantization techniques. This is an order of magnitude higher than the upper bound on weight leakage for our protocol. Furthermore, if a dishonest client does not have access to the training dataset, they should not be able to use the leaked information to infer the rest of the model through training.
[0103] The optical loss present between the different parties in real-life communication networks affects these leakage calculations. Weight leakage increases with the channel loss, while data leakage is independent of the channel loss. To conserve the classification accuracy one can adjust the photon occupation number.
[0104] FIGS.5A and 5B illustrates the weight leakage (upper traces) and data leakage (lower traces) versus round-trip channel loss and neurons per layer of the DNN, respectively,at an amplification gain of ^^ = 1.3 and an average photon number of ^^ = 6 for an MNISTclassification accuracy of 95%. FIG. 5A shows that, in the presence of loss, the server can increase the average photon occupation per weight to conserve classification accuracy. This increases the weights leakage while the data leakage remains unaffected. The client cannot compensate for high losses, as the classification accuracy saturates as a function of the gain. A weight leakage of up to 4 bits per weight is obtained for standard losses in local-area networks (LAN) and metropolitan-area networks (MAN). For networks with fiber losses of up to 3 dB (e.g., LANs), the model leakage is below 2 bits per weight symbol.
[0105] FIG.5B shows that both weight and data leakages diminish with an increasing number of neurons per layer. Interestingly, the leakage vanishes with the number of neurons per layer. This is because the amplification excess noise is spread over more modes in the return state, leading to less weight and data leakages. This suggests that DNNs with more neurons that those used here may suffer even smaller leakages.
[0106] FIG. 5C is a plot of weight leakage as a function of round-trip channel losses and (excess) technical noise from 0.0 to 0.2 SNUs. We assume the same technical noise forAttorney Docket No. MIT-26054WO01 both the server and the client and adjust the server power to compensate for the client’s noise in order to conserve the classification accuracy. The dashed horizontal line in FIG.5C indicates a realistic excess technical noise value of around 2% SNU. At this realistic noise level, the technical noise does not contribute significantly to weight leakage. This analysis considers strict assumptions, while a lower leakage could be obtained in loose assumptions where the malicious party has limited access to the system components. Quantum-Secure Deep Learning Inference with Multiple Clients
[0107] FIG. 6 illustrates quantum-secure deep learning inference in a multi-user network with a server 120 and two or more clients 110, where each party assumes that all other parties are dishonest and collaborate against them, in both symmetric and asymmetric network architectures. In a symmetric network, indicated by dark arrows, the server 120 splits the light encoding of the weights W in half and sends each portion to a different client 110. The clients 110 use the same gain for inference and transmit return states R to the server 120. The server 120 coherently combines the return states R and measures the noise in the output. The added excess noise is doubled relative to the return state R from a single client 110, resulting in an increase in weight leakage. However, the data leakage for each client 110 does not change compared to the single-client scenario. Since the light received by each client is halved in power, the SNR drops by a factor of two for each client 110, leading to a reduction in classification accuracy. In an asymmetric network (lighter arrows), the first client 110 uses all the light for inference and transmits its entire return state R to the second client 110. The second client 110, after performing its own inference, sends the doubly disturbed return state R back to the server 120, which measures the excess noise. One advantage of the asymmetric architecture is that it can be performed without coherently combining the return states from both clients 110 at the server 120. Classification Accuracy Calculation
[0108] To analyze our protocol’s performance for deep learning, we compute the classification accuracy for the standard MNIST classification task using PyTorch. To this end, we write custom neural network layers that implement our coherent linear algebra engine (client) presented above; several of these custom layers were strung together to form a secure neural network. The homodyne current detected in each custom layer is sampled from a unit Gaussian distribution ^^(0,1) SNU to account for the quantum shot noise in each quadrature.Attorney Docket No. MIT-26054WO01 The resultant noisy sample is then fed to the non-linear ReLU activation function before being passed on to the next custom layer.
[0109] Separately, we trained a digital model composed of standard PyTorch layers on preprocessed MNIST images. The preprocessing of the data set included the flattening of the28 × 28 images to vectors of size 784, followed by a centering and scaling of each vector tothe range [-1, 1]. Then we trained a standard 2-layer network with 784 inputs, 784 hidden neurons, and 10 outputs, obtaining a classification accuracy of 98%. We transferred these trained digital weights into the custom secure neural network and computed the accuracy achieved by the secure protocol as a function of the weight pulse energy. We found that the accuracy of the digital and analog models agreed up to variations caused by quantum shot noise.
[0110] The real operations for standard deep learning tasks can be implemented on our complex-valued hardware by encoding the real input vector ^^(ℝ)of length ^^ and the realweight matrix ^^(ℝ) of size ^^ × ^^ into a complex input vector ^^ (ℂ) of length ^^ / 2 and acomplex weight matrix ^^(ℂ) of shape ^^ × ^^ / 2 using the following procedure:The hardware computes the real value of the productwhich is precisely the matrix- vector product ^^(ℝ)^^(ℝ)desired for the DNN computations.
[0111] We define the signal-to-noise ratio (SNR) at the homodyne detector at the neural network outputs as ^^(^^ − 1)SNR^^= 2^^2 − 3^^ + 2 |^⃗⃗^ ⋅ ^̂^|2,where the amplitudes of these physical states, ^^^^, encode the logical weights as ^^^^ = √^^^^^^^^ / ∥^^ ∥, where ∥ ^^ ∥ denotes the root mean square of the elements in ^^. Then we have:The prefactor is the result of a beam splitter with a ratio of 1 / ^^: 1 − 1 / ^^ right after theamplifier ^^ that dictates how much of the power is used for the computation and how much for the return state. Therefore, SNR does not monotonically increase with ^^; in fact, it increasesfrom 1 to 2 + √2 and then decreases.Attorney Docket No. MIT-26054WO01
[0112] To obtain the desired inner product, the client scales the result digitally by ∥^^ ∥ ∥^^∥√^^−1√^^. For deep learning, the scaling step for the client is superfluous because the client normalizes between layers regardless. A physical scaling parameter ^^ that depends solely on the gain and average photon occupation captures the hardware-dependent prefactor in the SNR equation,
[0113] Varying the physical scaling factor ^^ changes the SNR of the inner product outputs, thus affecting the accuracy of the secure neural network. Defining this parameter enables the calculation of the classification accuracy as a function of the gain ^^ and average photon occupation ^^ without directly executing the intensive numerical calculations in the two- dimensional space spanned by these two parameters. Instead, we calculate the classification accuracy within the space spanned by ^^.
[0114] FIG. 7 is a plot of the classification accuracy achieved by the secure neural network versus the physical scaling parameter ^^. The classification accuracy is calculated considering an additive noise to each inner product, using weights trained for a digital noiseless model. The classification accuracy increases monotonically with ^^, asymptotically obtaining the digital noiseless accuracy. This physical scaling parameter captures the hardware- dependent prefactor in the SNR, allowing to calculate the mutual effect of the average photon number ^^ and the amplification gain ^^. The following logistic function faithfully fits the variation of the classification accuracy as a function of ^^:with Root-Mean-Square-Error (RMSE) of 1.1%. This fit allows us to use a convenient analytical formula to predict the classification accuracy as a function of the hardware parameters ^^ and ^^, instead of using look-up tables. Security Analysis Calculation
[0115] As explained above, we analyzed the security of the server following the standard security analysis for continuous-variable quantum key distribution (CVQKD), and we analyzed the security of the client by phrasing our problem as a statistical estimation problem.Attorney Docket No. MIT-26054WO01
[0116] First, we express the setting of our problem in the continuous-variable quantum key distribution (CVQKD) framework. In a standard CVQKD setting, two legitimate parties Alice and Bob wish to communicate a message under possible attack by an external third-party eavesdropper Eve. In the CVQKD protocol, normal operation results in Bob receiving the message from Alice, and the mutual information between Alice and Eve (along with Eve’s Holevo information) should be bounded below an arbitrary threshold.
[0117] In the context of our multiparty quantum-secured DNN protocol, normal operation results in the transfer of a specific amount of information from the server to the client. More specifically, the client extracts the inner product of the server’s weights with the client’s data to satisfy the accuracy requirements. Here, both parties (server and client) have the opportunity to act maliciously.
[0118] When the server is honest and the client is malicious, the information that is at risk of leakage is the weight vector sent from the server to the client. When the server is malicious and the client is honest, the information that is at risk of leakage is the client data vector used for the client’s inner product. Our protocol protects the honest party’s information. Weights Leakage Calculations
[0119] From the perspective of a malicious client and honest server, the server prepares displaced coherent states with quadrature components ^^ and ^^ that are realizations of two independent and identically distributed (i.i.d.) random variables ^^ and ^^. These randomvariables are normally distributed as ^^, ^^~^^(0, ^̃^mod), where ^̃^mod denotes the modulationvariance. Here, the information leakage upper bound, quantified by the Holevo information between Eve and Alice, is maximized when the state shared by Alice and Bob is Gaussian.
[0120] The prepare-and-measure covariance matrix, after transmission through a channel with loss ^^ and noise ^^ into Bob’s lab who performs heterodyne measurement, is:where ^^mod = 4^̃^mod.Attorney Docket No. MIT-26054WO01
[0121] Using the source replacement method commonly applied in QKD security analyses, the state sent by the server can be described as an entangled state. In this scenario, the server prepares a two-mode squeezed vacuum state (TMSVS), measures both quadratures of one mode, and sends the other mode to the client. In shot-noise units, the TMSVS is characterized by the covariance matrix:Here, ^^^^is the Pauli matrix and ^^ represents the variance of the quadrature operators. Theupper and lower diagonal 2 × 2 blocks of ^^TMSVS are the covariance matrices of the quadratureoperators of Alice and Bob, respectively, while the off-diagonal block matrices are the covariances between the quadratures of Alice and those of Bob. The variance ^^ is related tothe mean photon number per pulse by the relation ^^ =1 2(^^ − 1). This variance is the sum ofAlice’s actual modulation variance of her two quadrature components and the vacuum shot- noise variance.
[0122] After Bob’s mode is transmitted through the channel with transmission coefficient ^^ and excess noise ^^, the covariance matrix transforms to:
[0123] Without loss of generality, we consider |^^^as a bipartite state with Alice’s and Bob’s joint subsystem comprising one part and Eve’s subsystem the other. Any pure bipartitestate can be expressed using the Schmidt decomposition as |^^^ = ∑^^ ^^^^ |^^^^^^^|^^^^^. Tracing outeither subsystem yields a mixed state with a von Neumann entropy that depends only on the Schmidt coefficients ^^^^: log(^^^^)
[0124] We use the Holevo theorem to bound the amount of information that Eve can obtain. The Holevo theorem provides an upper bound on the accessible information that can be extracted from a quantum system. Tracing out Eve’s subspace from the joint stateyieldsAlice’s and Bob’s joint substate, that is, Tr^^(|^^^^^^|) = ^^^^^^. Mathematically, the Holevobound ^^ for the information held by Eve’s substate ^^^^ = Tr^^^^(|^^^^^^|) about the quantumstate ^^^^^^shared between Alice and Bob is given by:Attorney Docket No. MIT-26054WO01where ^^^^|^^is the conditional von Neumann entropy of the density matrix ^^^^given the density matrix ^^^^.
[0125] The von Neumann entropy of a quantum state can be expressed in terms of the symplectic eigenvalues of that state’s covariance matrix. The symplectic eigenvalues of a2^^ × 2^^ matrix ^^ are defined as the absolute values of the ordinary eigenvalues of the matrix^‾^ = ^^^^^^ where ^^ is the 2^^ × 2^^ matrix given by the direct sum:10) It follows that a Gaussian state with covariance matrix ^^ has a von Neumann entropy given by^^ = ∑^^ ^^ (^^^^), where and ^^^^is the ^^-th symplectic eigenvalue of ^^.
[0126] To calculate ^^^^^^we calculate the symplectic eigenvalues of ^^^^^^: 1 ^^1,2=( [ ])2^^ ± ^^ − ^^whereThen, we have ^^^^^^ = ^^(^^1) + ^^(^^2).
[0127] We calculate ^^^^|^^assuming that Bob implements homodyne detection. The covariance matrix ^^^^|^^is given by:where ^^^^ = ^^^^, ^^^^ = ^^^^^^, ^^^^ = ^^, and ^^, ^^, and ^^ are defined as above. Using thesedefinitions, we obtainIts symplectic eigenvalue is:Attorney Docket No. MIT-26054WO01This gives = ^^(^^3), which yields− ^^(^^3).
[0128] To calculate the Holevo information one should determine the intensity transmittance ^^ from Alice to Bob (server to client and back to the server) and the excess noise in the intensity ^^. To this aim we follow and define the variance of the weights sent by Alice, ^̃^mod, the return state light amplitude variance ^^^^measured using homodyne detection, and the covariance between the encoded weights and the return state light amplitudes: ^^^^^^. Then, ing ^^mod : = 4^̃^mod, the transmittance is obtained from: ^^ = 2 ( ^^^2defin^^^^^mod) and the excess noiseis obtained from: ^^ = 2(^^^^ − 1) − ^^^^mod. In summary, ^^mod, ^^^^ , and ^^^^^^ are measured while^^ and ^^ are inferred from the measurements using the formulas above. Weight Leakage Accumulation Due to Multiple Queries
[0129] A client gains limited information on the DNN weights through a single broadcast of the weights. The server can control the accumulation of this information during multiple queries. Specifically, the server can manipulate the weights of the DNN in every broadcast in a way that preserves the neural network function while drastically reducing the accumulation of weight information at the client. Although neural network functions are not invariant to general linear transformations (e.g., translation), there exist isomorphisms between weight matrices that do not alter the neural network function.
[0130] Let ^^1, ^^2 ∈ ℝ^^×^^(this procedure works equally well for rectangular matrices) be two weight matrices with ^^(⋅)implementing an elementwise nonlinearity, so ^^2^^(^^1)implements two successive layers of a neural network. For any permutation matrix^^ and its inverse ^^−1:Thus, there should be ^^! different permutations possible that return the same computation. The server can encode any of these ^^^^ matrices onto the coherent pulses that the server sends to the client(s). This process can be repeated with ^^2and the next matrix of the network ^^3and so on. For the common case of^^(^^) = ReLU(^^) = max(0, ^^), ^^ may also include multiplication by a random scalar, or evenmultiplication by a random diagonal matrix, limiting the information in one measurement by the entropy of these random manipulations. By choosing a maximally entropic distribution over the transformation space, Alice can limit Eve’s ability to infer weight information between broadcasts.Attorney Docket No. MIT-26054WO01
[0131] The idea of using permutations to hide information has a rich history in the recent privacy and statistics literature. Permutations can hide (obfuscate) neural network functions in conjunction with homomorphic encryption to achieve computational security. Permutation recovery is statistically hard in situations where the SNR is limited. In parallel, in the field of Differential Privacy (DP), random shuffling (permutations) amplifies differential privacy guarantees over non-shuffling baselines.
[0132] In DP, privacy leakage is measured via two leakage parameters ^^, ^^; the smallerthey are, the more private the protocol is. If the per-row leakage ^^rowstays below a threshold log2^^ (where ^^ is the number of rows in the matrix), shuffle-DP reduces the effective leakage ^^shfor an entire query by roughly 1 / √^^. Once ^^rowexceeds the logarithmic threshold, the permutation can be reconstructed and the shuffle no longer helps precisely the phase transition observed in DP studies. Our protocol enables below-threshold operation via tuning of the server’s laser power and the client’s amplification gain.
[0133] To raise privacy beyond what permutations and positive scalings can deliver, one should select an activation that commutes with a larger group of linear masks, such as block-radial gates, where each ^^-channel group is rescaled commute with every block-orthogonal matrix ^^(^^), adding ^^(^^ − 1) / 2 masking degrees of freedom per block. Anotherexample is hinge activations, which reflect a vector across a learned normal c whenever ^^^|^^^ <0. The price of these non-element-wise nonlinearities is extra compute but in exchange they enable a mask space almost as large as the linear layers themselves, tightening information- theoretic leakage bounds. Data Leakage Calculations
[0134] In this scenario, the client operates legitimately and the server acts maliciously. Under our eavesdropper model, we assume that the server can perform any arbitrary individual attack upon receiving the return state ^^^^. In particular, the server could be equipped to make quantum measurements, observe quantum correlations, and transmit non-Gaussian states.
[0135] The eavesdropper may operate under different models based on different capabilities; here, we present analysis for individual attacks and collective attacks. Under individual attacks, the attacker performs independent and identically distributed (i.i.d.) attacks on all incoming modes; that is, the attacker prepares separable ancilla states that interact individually with one of the signal modes in the quantum channel. The ancilla states are thenAttorney Docket No. MIT-26054WO01 stored in a quantum memory until measurement and are measured independently of each other. Under collective attacks, the attacker similarly stores the states in a quantum memory, but now is able to interact with multiple states at a time, including by performing entanglement-based attacks to correlate multiple states prior to collective measurement.
[0136] We also consider the case where the server transmits non-Gaussian states. Non- Gaussian states are not fully characterized by their covariance matrices, i.e., the inter-mode correlations of non-Gaussian states are not restricted to their first- and second-order moments. In particular, the client’s Gaussian operations on the server’s non-Gaussian states also influence the higher-order moments, which could further expose the client’s data. However, the effect of Gaussian operations on higher-order moments of non-Gaussian states is mediated through the effect on the first- and second-order moments. By the data processing inequality, this implies that the amount of information contained in measurements of higher-order moments is less than or equal to the amount of information in the first two moments. Thus, the QFI extracted from the higher-order moments of a given non-Gaussian state is bounded by the QFI extracted from the Gaussian state with the same first- and second-order moments. Individual Attacks
[0137] Recall from above that the excess noise in the ^^-th mode in ^^^^is:In the following classical leakage analysis, we characterize the leakage of each individual data element ^^ (subscript ^^ suppressed for convenience) under the possibility of an individual classical attack from a malicious server.
[0138] Let ^^ be a Gaussian random variable corresponding to the ^^-th mode of the return state with mean ^^ and variance ^^2, wherefor a known parameter ^^ and an unknown parameter ^^.
[0139] We aim to calculate the precision in estimating ^^ using the Cramér-Rao bound for ^^ measurements and use this precision to determine the number of bits of information obtained by the server about ^^.
[0140] The probability density function (PDF) of ^^ is given by:Attorney Docket No. MIT-26054WO01where ^^2is given above.
[0141] The Fisher information ℐ^^(^^) for a single measurement with respect to a given distribution on ^^ is given by:
[0142] Noting thatthe derivative of the log-likelihood function is:Because ^^ and ^^ are constants with respect to the distribution over ^^, the Fisher information isUsing the properties of the moments of Gaussian distributions: ^^[(^^ − ^^)2] = ^^2 and^^[(^^ − ^^)4] = 3^^4 yields:For ^^ independent measurements, the Fisher information is additive:The Cramér-Rao bound for the variance of any unbiased estimator ^̂^ of ^^ is: VarThe precision (standard deviation) for estimating ^^ is:Attorney Docket No. MIT-26054WO01The number of bits of information ℐ^^that the estimator ^̂^ contains about ^^ can be computed by treating ^̂^ as the received signal when ^^ is sent through an additive Gaussian channel with noise variance Var(^̂^):This provides the number of bits of information obtained by the server from ^^ measurements of the client’s return state. Collective Attacks
[0143] Consider the information an adversary can obtain when she is allowed collective measurements on all ^^ identical copies of the returned quantum state. Let the hidden parameterbe a unit vector ^^ ∈ ℂ^^ with the Haar distribution on the unit sphere. Conditioned on ^^ = ^^the server obtains the ^^-mode Gaussian density operatorwhose displacement ^^ is independent of ^^, so that all ^^-dependence resides in the rank-one spike of the covariance.
[0144] For a zero-mean Gaussian state, the QFI matrix is ^^^^^^ = 1⁄ 2 Tr[^^−1(∂^^^^) ^^−1(∂^^^^)]. Because ^^(^^) differs from the identity by a rank-one term, the Sherman–Morrison formula gives ^^ ^^†.Substituting yieldsan isotropic block on the (^^ − 1)-dimensional tangent space of the unit sphere. Since ^^ itselfis a rank-one perturbation of the identity, the quantum Cramér–Rao inequality becomesAttorney Docket No. MIT-26054WO01 CovWith ‖^^‖ = 1 the per-coordinate variance simplifies toThis multivariate variance bound is higher than or equal to the univariate variance bound by the AM-GM inequality.
[0145] In the asymptotic limit of a large number of modes and return states, one canget a simple leakage formula. First, the uniform-sphere prior has covarianceand its entropy differs from that of the centered Gaussian with the same covariance only by an ^^(ln^^) term. Second, for large ^^, the Bayesian posterior entropy converges to the Gaussian entropy with the same conditional covariance. 2and any finite-block-length correction can be evaluated exactly as in the finite-size analyses of QKD protocols. Substituting these two Gaussian entropies gives the leading mutual- information termBecause each mode contributes two real parameters, the leakage per real component obeysso an eavesdropper, even with optimal collective measurements, gains only a bounded amount of information about each element of the secret direction.
[0146] To consider leakage from hidden layers, assume the server adds a normalizedmap= 1, representing the function that maps the activations to thenext layer. The adversary then receivesand the corresponding QFI matrix with respect to the real coordinates of ^^ is:Attorney Docket No. MIT-26054WO01 |^^.
[0147] A malicious server could learn additional information by enlarging certain directional derivatives of ^^. The client counters this by estimating the entries of ^^^^(^^) while the network weights are being transmitted. To this aim, the client evaluates ^^ at several points that are close to ^^ and numerically evaluates the derivatives. Such small perturbations leave the classification outcome unchanged yet allow the client to evaluate the derivatives at ^^. If any estimated derivative exceeds a preset threshold, the client aborts before sending a return state; otherwise, the measured values bound the relevant Jacobian components, ensuring that the information that the server can extract from any additional batch of copies remains provably limited. Conclusion
[0148] While various inventive embodiments have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the inventive teachings is / are used. Those skilled in the art will recognize or be able to ascertain, using no more than routine experimentation, many equivalents to the specific inventive embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specifically described and claimed. Inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the inventive scope of the present disclosure.
[0149] Also, various inventive concepts may be embodied as one or more methods, of which an example has been provided. The acts performed as part of the method may be orderedAttorney Docket No. MIT-26054WO01 in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.
[0150] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0151] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”
[0152] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0153] As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of” or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e. “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the claims, shall have its ordinary meaning as used in the field of patent law.
[0154] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one elementAttorney Docket No. MIT-26054WO01 selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0155] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of” and “consisting essentially of” shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.
Claims
Attorney Docket No. MIT-26054WO01 CLAIMS 1. A method of quantum-secure, multiparty deep learning inference, the method comprising: receiving, from a server at a client, a first sequence of coherent pulses of light encoded with a weight vector for a deep-learning model, each coherent pulse in the first sequence of coherent pulses of light having a photon number of no more than 10; modulating, at the client, the first sequence of coherent pulses of light with an input vector to the deep-learning model to yield a second sequence of coherent pulses of light; performing, at the client, a homodyne measurement of a mode of the second sequence of coherent pulses of light, the homodyne measurement representing an inner product of the weight vector and the input vector; modulating, at the client, the second sequence of coherent pulses of light to produce a third sequence of coherent pulses of light encoded with the weight vector and having a variance greater than a variance of the first sequence of coherent pulses of light and below a threshold based on the photon number; and transmitting, by the client, the third sequence of coherent pulses of light to the server for verification that the client did not copy the weight vector.
2. The method of claim 1, wherein receiving the first sequence of coherent pulses of light comprises receiving between 1 photon and 6 photons per coherent pulse from the server.
3. The method of claim 1, wherein: modulating the first sequence of coherent pulses of light comprises performing a first unitary transformation on the first sequence of coherent pulses of light, and modulating the second sequence of coherent pulses of light comprises performing a second unitary transformation on the second sequence of coherent pulses of light, the second unitary transformation being a conjugate transpose of the first unitary transformation.
4. The method of claim 1, wherein performing the homodyne measurement comprises interfering the mode of the second sequence of coherent pulses of light and a local oscillator signal.
5. The method of claim 1, wherein performing the homodyne measurement comprises filtering the second sequence of coherent pulses of light with a narrowband filter.Attorney Docket No. MIT-26054WO01 6. The method of claim 1, further comprising: amplifying, with an optical amplifier, the second sequence of coherent pulses at a gain selected to achieve a desired classification accuracy of a deep learning inference computation.
7. The method of claim 1, further comprising: amplifying, with an optical amplifier, the second sequence of coherent pulses at a gain selected to limit leakage of the input vector and / or of the weight vector.
8. The method of claim 1, further comprising: measuring, at the server, the variance of the third sequence of coherent pulses of light.
9. A client for quantum-secure, multiparty deep learning inference, the client comprising: a first modulator to modulate a first sequence of coherent pulses of light encoded with a weight vector for a deep-learning model with first modulation based on an input vector to the deep-learning model to yield a second sequence of coherent pulses of light, each coherent pulse in the first sequence of coherent pulses of light having a photon number of no more than 10; a homodyne detector, in optical communication with the first modulator, to perform a homodyne measurement of a mode of the second sequence of coherent pulses of light, the homodyne measurement representing an inner product of the weight vector and the input vector; and a second modulator, in optical communication with at least one of the first modulator or the homodyne detector, to modulate the second sequence of coherent pulses of light with second modulation based on the input vector to produce a third sequence of coherent pulses of light encoded with the weight vector and having a variance greater than a variance of the first sequence of coherent pulses of light and below a threshold based on the photon number.
10. The client of claim 9, wherein the first sequence of coherent pulses of light includes between 1 photon and 6 photons per coherent pulse.
11. The client of claim 9, wherein the first modulation represents a first unitary transformation and the second modulation represents a second unitary transformation, the second unitary transformation being a conjugate transpose of the first unitary transformation.Attorney Docket No. MIT-26054WO01 12. The client of claim 9, further comprising: a local oscillator, in optical communication with the homodyne detector, to generate a local oscillator signal for performing the homodyne measurement.
13. The client of claim 9, further comprising: a narrowband filter, in optical communication with the first modulator, the homodyne detector, and the second modulator, to direct the mode of the second sequence of coherent pulses of light from the first modulator to the homodyne detector and to direct at least one other mode of the second sequence of coherent pulses of light to the second modulator.
14. The client of claim 9, further comprising: an optical amplifier, in optical communication with the first modulator, to amplify the second sequence of coherent pulses of light.
15. The client of claim 14, wherein the optical amplifier has a gain selected to achieve a desired classification accuracy of a deep learning inference computation.
16. The client of claim 14, wherein the optical amplifier has a gain selected to limit leakage of the input vector and / or of the weight vector.
17. A system comprising: the client of claim 9; and a server, in optical communication with the client, to transmit the first sequence of coherent pulses of light to the client.
18. The system of claim 17, wherein the server is further configured to receive the third sequence of coherent pulses of light from the client and to verify that the client did not copy the weight vector based on the variance of the third sequence of coherent pulses of light.
19. The system of claim 17, wherein the client is a first client, the inner product is a first inner product, and further comprising: a second client according to claim 9, in optical communication with the first client and the server, to receive the third sequence of coherent pulses of light from the first client, to compute a second inner product with the weight vector, and to return a fourth sequence of coherent pulses of light to the server for verification that neither the first client nor the second client copied the weight vector.Attorney Docket No. MIT-26054WO01 20. The system of claim 17, wherein the client is a first client, the inner product is a first inner product, and further comprising: a second client according to claim 9, in optical communication with the first client and the server, to receive a fourth sequence of coherent pulses of light encoding the weight vector from the server, to perform a second inner product with the weight vector, and to return a fifth sequence of coherent pulses of light to the server for verification that the second client did not copy the weight vector.