Method for determining a quantum communication setup, quantum communication setup, computer program, and data processing system

The method uses reinforcement learning to optimize quantum communication setups in QKD, enhancing key generation rates by automating component selection and post-selection, addressing the challenge of exponential complexity in QKD setups.

JP7705659B2Active Publication Date: 2025-07-10TERRA QUANTUM AG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022142512
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-14
Filing Date
2022-09-07
Publication Date
2025-07-10
Estimated Expiration
2042-09-07

AI Technical Summary

Technical Problem

In quantum key distribution (QKD), determining an optimized quantum communication setup that maximizes key rate while minimizing resource usage is challenging due to the exponential increase in possible experimental implementations and optical modes.

Method used

A method utilizing reinforcement learning to automate the design of a quantum communication setup by optimizing quantum communication components and post-selection procedures, including beam splitters and phase shifters, to enhance key generation rate.

Benefits of technology

The method efficiently handles complexity in QKD setups, optimizing component selection and post-selection to achieve higher key rates while limiting information leakage to eavesdroppers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705659000087
    Figure 0007705659000087
  • Figure 0007705659000088
    Figure 0007705659000088
  • Figure 0007705659000089
    Figure 0007705659000089
Patent Text Reader

Abstract

To provide a method for determining a quantum communication setup that allows increased key rates in an efficient and resource-conserving manner.SOLUTION: The method comprises: providing a component set indicative of a quantum communication setup comprising quantum communication components of at least one of a first communication device 40, a second communication device 43 and an eavesdropping device 44; selecting actions each indicative of a further quantum communication component and each selectable with a selection probability depending on the component set and a reward value; including the selected further quantum communication component in the component set; determining a quantum model of at least one of the first communication device 40, the second communication device 43 and the eavesdropping device 44; determining a maximum key rate by optimizing an optimization parameter set comprising quantum communication component parameters; adjusting the reward value depending on the size of the maximum key rate in relation to a previous maximum key rate; and repeating the steps until a termination criterion is satisfied.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method for determining a quantum communication setup, which is executed in a data processing system. Further, a quantum communication setup, a computer program, and a data processing system are disclosed.

Background Art

[0002] In quantum key distribution (QKD), legitimate parties, conventionally called Alice and Bob, desire to establish a secret shared key that is completely unknown to potential eavesdroppers (conventionally called Eve). Finding an optimized component setup and post-processing design that provides a high key rate is generally a difficult task in QKD. This is because the number of possible experimental implementations increases exponentially with the number of components and, in the case of an optical system, the number of optical modes. The higher the key rate, the more information can be securely exchanged between legitimate parties.

[0003] A method for the automatic design of optical experiments has been proposed by Melnikov et al., Phys. Rev. Lett., 125(16), 2020, which has as its basic goal to find a high Bell score in a scenario with a fixed number of measurements and results.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to provide a method for determining a quantum communication setup that enables an increase in the key rate in an efficient and resource-saving manner.

Means for Solving the Problems

[0005] To solve this problem, a method for determining a quantum communication setup according to independent claim 1 is provided. Further, a quantum communication setup, a computer program, and a data processing system are provided according to independent claims 13, 14, and 15, respectively. Further embodiments are disclosed in the dependent claims.

[0006] As a result, it is possible to design a QKD protocol using a novel procedure including optimization regarding the secret key generation rate. This novel method further enables, simultaneously, devising the QKD protocol itself (such as encoding, post-selection rules, etc.), optimizing its implementation, and limiting the leakage of information to eavesdroppers.

[0007] According to one aspect, a method for determining a quantum communication setup, which is executed in a data processing system, is provided. The method includes providing a set of components representing a quantum communication setup including at least one quantum communication component of a first communication device (Alice), a second communication device (Bob), and an eavesdropping device (Eve); selecting an operation of a set of operations, each representing a further quantum communication component, wherein each operation is selectable with a selection probability depending on the set of components and a reward value; including the selected further quantum communication component in the set of components; determining a quantum model including at least one quantum state of at least one of the first communication device, the second communication device, and the eavesdropping device from the set of components; determining a maximum key rate by optimizing with respect to a set of optimization parameters including quantum communication component parameters from the set of components and the quantum model; adjusting the reward value according to the size of the maximum key rate with respect to a previous maximum key rate; and repeatedly repeating the above steps until an end criterion is met, resulting in an optimal set of components and an optimal set of setup parameters.

[0008] According to a second aspect, a quantum communication setup is provided that includes quantum communication components, and the quantum communication setup is created according to an optimal set of components and an optimal set of setup parameters generated by a method for determining the quantum communication setup.

[0009] According to another aspect, a computer program and / or a computer program product are provided, which, when executed in a data processing system, include instructions that cause the data processing system to perform the steps of a method for determining a quantum communication setup.

[0010] According to a further aspect, a data processing system is provided that is configured to determine a quantum communication setup by performing the following steps: providing a set of components representing a quantum communication setup that includes at least one quantum communication component of a first communication device, a second communication device, and an eavesdropping device; selecting an operation from a set of operations, each of which represents a further quantum communication component, where each operation is selectable with a selection probability that depends on the set of components and a reward value; including the selected further quantum communication component in the set of components; determining a quantum model from the set of components that includes at least one quantum state of at least one of the first communication device, the second communication device, and the eavesdropping device; determining a maximum key rate by optimizing an optimization parameter set that includes quantum communication component parameters with respect to the set of components and the quantum model; adjusting the reward value according to the size of the maximum key rate with respect to a previous maximum key rate; and repeatedly performing the above steps until an end criterion is met and an optimal set of components and an optimal set of setup parameters are obtained.

[0011] In the proposed method that may include concepts from reinforcement learning, the complexity that arises when increasing the number of components and the number of optical modes can be efficiently handled. This makes it possible to automate the design of a quantum communication setup suitable for QKD and optimize the corresponding setup parameters. The method can further make it possible to determine an optimal post-selection procedure for increasing the key rate and select optimal parameters for a known level of intensity captured by a potential eavesdropper. The resulting optimal set of components can be physically implemented in a QKD setup.

[0012] The quantum communication components may include at least one of optical components, particularly a beam splitter and a phase shifter.

[0013] The quantum communication components may be components of a first communication device and / or a second communication device, and / or an eavesdropping device.

[0014] The selection of actions may be performed by a reinforcement learning agent, preferably a projective simulation agent.

[0015] The reinforcement learning agent may comprise a projective simulation network. The reinforcement learning agent may comprise a two-layer network. Alternatively, the reinforcement learning agent may include a deep feedforward neural network.

[0016] The first layer of the network may include first nodes, each of which can represent a state indicating one of the possible sets of components. The second layer of the network may include second nodes, each of which can represent one of the actions of the set of actions (each action indicating an additional quantum communication component, particularly an optical component). The reinforcement learning agent may be part of a learning agent module.

[0017] Within the context of the present disclosure, the agent-environment interaction step between a reinforcement learning agent and an environment for exchanging quantities such as operational / quantum communication components or sets of components may be referred to as a round. A series of rounds that result in a reward value as a feedback signal may be referred to as a trial.

[0018] The termination criterion may be satisfied when the determined maximum key rate reaches the target key rate. The termination criterion may also be satisfied when the maximum number of trials is reached. The maximum number of trials may be an integer value between 50 and 1,000, preferably between 75 and 200, and particularly 100. The maximum number of rounds per trial may be an integer value between 10 and 100, preferably between 15 and 25, and particularly 20.

[0019] (Projection simulation) The network can include a plurality of edges with weights, and the weights are preferably trial-dependent and / or round-dependent.

[0020] The selection probability (policy) for selecting one of the set of operations can correspond to a reinforcement learning policy. Preferably, the selection probability for selecting one of the set of operations can depend on an exponential function of the weights that depend on the set of components and the operations.

[0021] Each selection probability can be updated after each round. Each weight can depend on a weight decay term and / or a grow term. A larger weight decay term can result in a smaller weight in the next round. A larger grow term can result in a larger weight in the next round.

[0022] This method may further include the step of removing the same quantum communication component from the set of components.

[0023] Preferably, the same quantum communication component can be removed from the set of components after including the selected additional quantum communication component in the set of components and before determining the quantum model.

[0024] Removing the same quantum communication components from the set of components may be performed in the compression module.

[0025] The quantum state may include a plurality of optical modes. Preferably, a random number including randomly generated information bits may be encoded using the optical modes.

[0026] The random number may be generated in the first communication device. The random number may be generated by a hardware random number generator, in particular, a quantum random number generator or a classical random number generator. The random number may be selected from a set of integers having a discrete uniform distribution.

[0027] The random number may be encoded in the phase of the phase shifter of the quantum communication setup. Additionally or alternatively, the random number may be encoded in the beam splitter parameters of the beam splitter or in the component parameters of other optical elements of the quantum communication setup.

[0028] The quantum state may depend on the set of components and / or the set of setup parameters. The quantum state is the first quantum state |Ψ a > related to the first communication device, and the signal state related to transmitting the first quantum state |Ψ a > via the quantum channel

Number

Number

Number

[0029] The optical mode can include a time bin mode and / or a frequency mode. The determination of the quantum model can be performed in a physical simulation module.

[0030] The quantum model can further include a plurality of quantum measurements performed in the first communication device and / or the second communication device.

[0031] The quantum model can preferably include a quantum operator acting on a quantum state. The quantum operator can be a POVM (positive operator-valued measure) element. Each POVM element can be composed of a partial operator acting on one optical mode state.

[0032] The method can further include determining, from the quantum model, a post-selection set including measurement results received by the first communication device and the second communication device. The post-selection set can preferably be determined by determining, for each measurement result, first information regarding the corresponding information bit provided to the second communication device and leakage information regarding the corresponding information bit provided to the eavesdropping device.

[0033] The first information and the leakage information can be determined by establishing a first conditional entropy and a second conditional entropy, respectively. The maximum key rate can depend on the post-selection set.

[0034] The post-selection set can be determined by including measurement results where the leakage information is smaller than the first information and discarded measurement results where the leakage information is larger than the first information.

[0035] The post-selection set can be determined via an optimization routine, particularly via an annealing method or a tree search.

[0036] The included measurement results can be indicated by the unit post-selection mask bits of the post-selection mask sequence. The discarded measurement results can be indicated by the zero post-selection mask bits of the post-selection mask sequence.

[0037] The maximum key rate can be determined by a global optimization method, in particular an annealing method.

[0038] During optimization, the optimization parameters of the set of optimization parameters can be sampled from a finite interval having a (continuous) uniform distribution. The length of the interval can depend on the type of optimization parameter. The optimization can be performed in an optimization module.

[0039] The set of optimization parameters may include at least one of an intensity value, a phase shift value, and a post-selection mask bit. Additionally or alternatively, the set of optimal setup parameters may include at least one of an optimal intensity value and an optimal phase shift value.

[0040] The set of optimization parameters and / or the set of optimal setup parameters may further include additional (quantum communication) setup parameters / quantum communication component parameters, in particular optical component parameters, i.e., beam splitter parameters and / or phase shifter parameters and / or displacement amplitudes.

[0041] Before adjusting the reward value, if the maximum key rate is less than or equal to the previous key rate, the method can further include repeatedly including another element in the set of components until the maximum number of iterations is reached or the maximum key rate becomes greater than the previous key rate, and updating the quantum key distribution model, the post-selection set, and the maximum key rate.

[0042] The maximum number of iterations can be the maximum number of iterations per round / agent-environment interaction step.

[0043] When the maximum number of iterations is reached, the reward value can be set to 0 and / or the set of components can be reset to an empty set.

[0044] If the maximum key rate is greater than the previous key rate, the reward value can be set to a non-zero value, preferably 1, and the set of components can be retained.

[0045] The method can further include, external to the data processing system, creating a quantum communication setup according to an optimal set of components and an optimal set of setup parameters.

[0046] The step of creating a quantum communication setup can include arranging and adjusting phase shifters and / or beam splitters according to an optimal set of components and an optimal set of setup parameters.

[0047] The foregoing embodiments regarding a method for determining a quantum communication setup can be provided corresponding to a data processing system configured to determine a quantum communication setup.

Brief Description of the Drawings

[0048] Hereinafter, by way of example, embodiments will be described with reference to the drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Best Mode for Carrying Out the Invention

[0049] FIG. 1 shows a graphic representation of an exemplary embodiment of the proposed method, where establishing an optical setup and a post-selection scheme that result in a high raw key generation rate for quantum key distribution (QKD) in an automated way is treated as a reinforcement learning task.

[0050] The input quantities 10, 11, 12 of the present method include the available operations and their physical models, the hyperparameters 10 of the (reinforcement) learning agent module 13, the characteristics of the quantum communication channel (as a physical model) 11, and the hyperparameters 12 of the optimization module 14. As the output 19 of the present method, it is possible to provide an optimal set of components indicating the optimal quantum communication setup and an optimal set of setup parameters.

[0051] The exemplary method, referring to FIG. 4, aims to automate the design and implementation of a QKD setup for achieving a high raw key rate between a first communication device 40 (conventionally called Alice) and a second communication device 43 (conventionally called Bob). The method utilizes a model of the quantum channel between the first communication device 40 and the second communication device 43, as well as a set of possible operations that can be executed on the physical systems associated with the first and second communication devices.

[0052] In particular, in the QKD measurement and generation protocol, a quantum system encoded with a randomly generated value a is created at a first communication device 40 and transmitted to a second communication device 43 via a quantum channel. Subsequently, at the second communication device 43, a quantum measurement is performed according to a measurement setting y, and a measurement result b is obtained. In a post-processing step, a particular result b is accepted (post-selected) by the second communication device 43, and a particular event (a subset of the possible combinations (a, y)) is accepted (shifted) by both the first communication device 40 and the second communication device 43 after partial information regarding a and y has been publicly revealed.

[0053] The automatic design task can be cast as an interaction between a reinforcement learning agent module 13 and an environment 13a including environment modules 14 to 18. As input, the learning agent module 13 takes the maximum number of operations and the maximum length of the experiment. At each interaction step, the learning agent module 13 determines which action to select according to the current state of the environment 13a. Subsequently, the state of the environment 13a is updated, and the learning agent module 13 receives a feedback reward value 19a and the updated state of the environment 13a. After the learning agent module 13 determines an action, the set of components is updated and subsequently (in the compression module 15) compressed, i.e., cleared of redundant components. In the physical simulation module 16, the encoding procedure at the first communication device and the measurement procedure at the second communication device are simulated using the established (compressed) set of components. In the post-processing module 17, the post-selection procedure is executed. Using the respective results from the physical simulation module 16 and the post-processing module 17, the corresponding (raw) key rate is calculated in the security proof module 18. The final secret key with the corresponding final key rate can be obtained from the shared raw key using standard error correction and privacy amplification techniques. The key rate is used to calculate the reward value fed back as a feedback signal 19a to the reinforcement learning agent module 13, thus enabling reinforcement learning.

[0054] [General Description of Method Steps] In this section, after each step of the method is described in general terms, it will be described in detail as an example in subsequent sections. For clarity, the steps are disclosed together with modules 13 to 19. As will be understood by those skilled in the art, the described method steps can also be executed without such modularization.

[0055] The learning agent module 13 takes as input N different optical elements, the maximum component set length L max , the current component set s k (t) (after compression), the reward value r(t), and the hyperparameters, and outputs the index of the action / further quantum component a k (t). The reinforcement learning agent module 13 interacts with the environment 13a in round (agent - environment interaction step) k. A series of rounds k that result in the reward value r(t) as the feedback signal 19a are called trials t. In FIG. 1, 1 round and 1 optimization step are visualized. At the start of each trial t, the learning agent module 13 receives as an input state the component set s0(t) which is a representation of the initial (empty) optical component setup. Subsequently, the learning agent module 13 selects an action a1(t) corresponding to adding a further quantum communication component to the component set s k (t) from among the set of available actions. The set of available actions may include applying optical elements and may vary according to a particular experimental embodiment. The selection process is adjusted by the current policy π k of the learning agent module 13, but the policy π k is modified at the end of the round according to the feedback reward value received by the learning agent module 13.

[0056] The compression module 15 takes the further quantum communication component and the component set s k-1 (t) of the previous round corresponding to the action a kReceive (t) (the index of) as input, and the current set of components s after compression k Output (t). The quantum communication component a selected by the reinforcement learning agent module 13 k (t) is the set of components s k-1 contained in (t), and the new set of components

Number

Number

[0057] The physical simulation module 16 takes as input the current set of components s after compression k (t), a list of parameter values of each quantum communication component, each state creation a, and the measurement setting y, and constructs the quantum model {|Ψ k (t) from the set of components s a 〉, E b∨y} is output. The quantum model {|Ψ a 〉, E b∨y} includes the first quantum state |Ψ a 〉 of the first communication device and the quantum measurement value / POVM {E b∨y} executed in the second communication device 43. In the first communication device 40, a random number a is generated, encoded into the first quantum state |Ψ a〉, which is then transmitted via a quantum channel using a quantum communication setup corresponding to the set of components provided by the learning agent module 13 to a second communication device. The received quantum state is measured at the second communication device via a measurement setup that yields a measurement result b. The measurement setting y may also depend on, for example, a locally generated random number y that can parameterize the measurement basis being used. In the physical simulation module, the optical setup is simulated, whereby a first quantum state |Ψ a 〉 corresponding to the random number a encoded by the first communication device 40 is obtained, and further, measurement values {E b∨y} obtained at the second communication device 43 are obtained.

[0058] The post - processing module 17 takes as input the set of components s k (t), the possible number of state creations a, the measurement setting y, and the measurement result b, and outputs a post - processing strategy. In QKD, the post - processing includes a post - selection procedure that provides the legitimate parties (the first communication device 40 and the second communication device 43) with more information to share than a potential eavesdropping device 44 (traditionally called Eve) can intercept. The first communication device 40 and the second communication device 43 can publicly reveal information about their random inputs a and y and decide to abort the communication round if the revealed information does not match a pattern (basis reconciliation / sifting). Further, the second communication device 43 may skip and not publish a round if the measurement result b does not meet a predetermined requirement (post - selection). The post - processing module 17 provides an optimal post - processing strategy for the input data, i.e., a binary string corresponding to the post - selection mask, where, for example, the value "1" (unity value) corresponds to a retained measurement result and the value "0" corresponds to a discarded measurement result.

[0059] The security proof module 18 uses as input the quantum model {|Ψ a 〉, E b∨y} and the post - processing strategy, and outputs a key rate R, specifically the current key rate corresponding to iteration i, round k, and trial t

Number

[0060] Prove the security of the QKD protocol resulting from the quantum communication setup utilized and the current key rate

Number

Number

[0061]

Number

[0062] wherein the symbol

Number

Number

[0063] There are several possibilities for implementing the security proof module 18. The security of a given protocol can be determined using an accurate model of the quantum channel that may be available via an external communication channel control technique. Alternatively, the entropy of the eavesdropping device 44 [Number] can be numerically bounded, for example, via explicit constrained minimization over multiple quantum channels or via a genetic algorithm for explicitly designing an attack from the eavesdropping device 44.

[0064] The optimization module 14 takes as input the component set s k (t), the post-processing strategy, hyperparameters, the maximum number of optimization iterations, the previous maximum key rate R(t - 1), and the current key rate [Number] and outputs several possible quantum state creations and measurements, the parameters of the quantum communication components within the component set s k (t) for all encodings and measurements, the post-processing method, and the reward value r(t). The optimization module 14 enables optimizing the parameters of the provided quantum communication setup and post-processing strategy such that the key rate is maximized. The optimization algorithm utilized enables the reinforcement learning agent module 13 to determine a quantum communication setup and post-processing strategy for a QKD protocol that provides a higher key rate. Through optimization, the optimal parameters of all quantum communication components within the quantum communication setup determined by the learning agent module 13 are established such that the measurement results yield the highest possible key rate. Furthermore, the optimal parameters for the post-processing strategy can be established. The optimization continues until the maximum number of optimization iterations is reached or the round key rate R k (t) (all current key rates for iteration i of round k [Number] covering), but the previous maximum key rate R(t - 1) = max from the previous trial k · ,t ·R k ·advances until it becomes larger than (t‘ ∨ t’ ≤ t - 1). R k If R(t) > R(t - 1), the reward value is set to r(t) = 1, the trial t ends, and the learning agent module 13 starts the next trial t + 1 from the initial configuration s(t + 1) with an empty set of components and the updated key rate R(t).

[0065] [Detailed description of the method steps as an example] In the following, a method for determining a quantum communication setup will be further described by applying it to a specific quantum communication setup.

[0066] FIG. 2 shows a graphic representation of a projection simulation agent utilized as a reinforcement learning agent in the learning agent module 13. The projection simulation agent comprises a two - layer projection simulation network 20. The simulation network 20 comprises a first layer of first nodes (perception) 21 representing a state / component set s(k,t) and a second layer of second nodes 22 representing an action a(k,t). The first nodes 21 and the second nodes 22 are connected by directed edges 23 having trial - dependent and round - dependent weights h k (s,a,t). The weights h k (s,a,t) together completely define a policy function π k . The weights h k (s,a,t) to the policy π k of the policy function π k (s,a,t) can be embodied in various ways within the projection simulation model, for example, as follows.

[0067] [Number]

[0068] Policy k (s, a, t) is updated by performing a change to the weight h k after each round / agent - environment interaction step k. That is, it is as follows.

[0069] h k+1 (s, a, t) = h k (s, a, t) - γ PS (h k (s, a, t) - 1) + g k+1 (s, a, t)r(t),(3)

[0070] Here, the initial weight h1(s, a, 1) is set to h1(s, a, 1) = 1 for all (s, a) pairs. The reward value r(t) is binary, that is, if the reinforcement learning agent discovers a set of components that result in a higher current key rate

Number

[0071] The second term on the right - hand side of equation (3), γ PS (h k(s,a,t)-1) is a weight decay term corresponding to the "forgetting" of the reinforcement learning agent. The second meta-parameter γ at a small level PS allows the reinforcement learning agent to switch to a completely new strategy when designing a quantum communication setup within one learning execution consisting of a series of trials. The second meta-parameter γ PS is also considered to be fixed. The first meta-parameter and the second meta-parameter are, for example, η PS = 0.3 and γ PS = 10 -3 respectively. Further meta-parameters such as the maximum number of trials N tr and the maximum number of rounds per trial N r can be set to 100 and 20 respectively, for example. Furthermore, the number of operations n act can be set to 6, and the maximum length L max of the component set can be set to 8. The reinforcement learning agent determines which action a k is selected according to the policy function π k (t). The policy π k (s,a,t) represents the probability that action a is selected at round k and trial t considering the state s k (t).

[0072] After one action a in the set of actions is selected with the probability corresponding to the relevant policy π k (s,a,t), the corresponding quantum communication component is included in the component set.

[0073] Figure 3 shows a table including an exemplary set of actions a (labeled with integer values in column 1) and the corresponding quantum communication components (column 2). Here, the quantum communication components are optical components, namely beam splitters (BS) and phase shifters (PS). Each of the beam splitters acts on 2 out of 3 optical modes, and each phase shifter acts on 1 out of 3 optical modes (see the graphical representation in column 4. Creation and annihilation operators

Number

[0074] Subsequently, the compression module 15 simplifies / compresses the set of components by removing the same components from the set of components s k The compressed set of components is then passed to the physical simulation module 16 and the post-processing module 17, both of which connect the set of components s k to the quantum communication setup to be optimized.

[0075] FIG. 4 shows a graphical representation of a general quantum communication setup for quantum key distribution. In a first communication device 40 (“Alice”), a random number a is generated by a quantum random number generator QRNG 41 from a set {1,..., K} of integers K. The random number a is encoded in a first quantum state |Ψ〉 a of d optical modes created in the creation unit 42.

[0076] The encoding is performed, for example, as follows: First, an initial state |Ψ0〉 that is a fixed coherent state of the optical mode is created.

[0077]

Number

[0078] Second, a unitary transformation

Number

[0079]

Number

[0080] Subsequently, the first quantum state |Ψ a 〉 is transmitted as a signal state via a potentially public quantum channel to the second communication device 43 (“Bob”). (The information from the random number a) is encoded by a unitary transformation

Number

Number

Number

[0081] The signal state in the joint AS quantum system

Number

[0082]

Number

[0083] Here, system A represents the classical number a transmitted from the first communication device 40 to the second communication device 43, and system S represents the corresponding signal. The random number a is randomly selected with probability p(a).

[0084] The first communication device 40 can be identified by a reinforcement learning agent to modify the first communication device 40 such that the key rate R is maximized according to the fraction r E of the strength of the transmitted signal that can be intercepted by the eavesdropping device 44 (“Eve”). The most appropriate set of unitary transformations and initial states

Number

Number

[0085] There are several physical degrees of freedom that can realize d different optical modes. One option is time binning. The advantage of such encoding is that different optical modes do not suffer from relative phase dissipation during transmission through the optical fiber of the quantum channel. The disadvantage is that, depending on the situation, the unitary transformation

Number

[0086] Another option for implementing optical modes is to utilize different frequencies. The optical modes can be slightly different from each other with respect to their respective frequencies. As a result, the optical modes can be coupled to the signal in a single optical fiber by using a multiplexer. Therefore, a non-trivial unitary transformation

Number

[0087] Assuming that the eavesdropping device 44 is limited to a beam splitter attack, after the attack of the eavesdropping device 44, the second quantum state of the joint ABE quantum system containing the random number (A)

Number

[0088]

Number

[0089] Here, transmission T = 10 -λD / 10 where λ represents the loss in dB / km on a quantum channel with length D.

[0090] A displacement operation is applied to the remaining signal received by the second communication device 43

Number

[0091]

Number

[0092] This is the quantum state after displacement

Number

[0093]

Number

[0094] Here, α is the amplitude of the displacement. By the shift operator, the local state in the phase space can be displaced by an amount α.

[0095] The third quantum state

Number

[0096]

Number

[0097] wherein

Number

[0098] In general, several types of measurements are possible that depend on the parameter y selected by the second communication device (43). Click / No click measurements can be performed and y can be made fixed. In this case, E b∨y corresponds to E b In a more general case, different y's can be selected by utilizing an ensemble of different measurement values related to several probability distributions.

[0099] The first communication device 40 and the second communication device 43 corresponding to legitimate parties then perform a post-selection via a classical authentication channel by discarding all non-deterministic measurement results in the second communication device 43, which provides an additional information advantage for legitimate parties over the eavesdropping device 44. Thus, the legitimate parties can decide whether to accept a particular result b

Number

[0100] Figure 5 shows a graphical representation of an embodiment of a quantum communication setup having three optical modes |α (1) 〉, |α (2) 〉, |α (3) 〉. In the first communication device 40, a random number a is encoded into a unitary transformation

Number

Number

Number

[0101] Each displacement operation

Number

[0102]

Number

[0103] The post - processing is post - selection, i.e., a subset of all measurement results accepted by the second communication device 43

Number

[0104] H(A∨b)(11)

[0105] The leakage information can be represented using the second conditional entropy.

[0106]

Number

[0107] Condition

Number

Number

[0108] To determine an appropriate post-selection procedure, the second conditional entropy

Number

Number

Number

[0109]

Number

[0110] In the formula,

Number

Number

Number

Number

[0111] The second conditional entropy

Number

Number

Number

[0112] For a given input a, the probability of post-selection can be calculated as follows.

[0113]

Number

[0114] By inserting into Equation (14), it becomes as follows.

[0115]

Number

[0116] Equation (15) is

Number

Number

Number

[0117]

Number

Number

[0118] The second conditional entropy H(A|b) is completely determined by the measurement result b and does not depend on the post-selection.

[0119] At this time, the key rate R is as follows.

[0120]

Number

[0121] The key rate R depends on the post-selection set. The most appropriate and thus the post-selection set used is [Number] can be. The post-selection module 17 results in a binary string (post-selection mask) of length [Number] such that the unit values correspond to the accepted measurement results that will be retained.

[0122] This post-selection mask, and thus the post-selection set, is determined by optimization. Subsequently, classical error correction and privacy amplification protocols can be executed on the accepted positions of the post-selection set, which can, in principle, be done at the Shannon limit.

[0123] Component set s k (t), corresponding physical model {|Ψ a 〉, E b∨y} and the post-selection set [Number] Once established, the maximum key rate R k (t) is determined at a given round k of trial t by optimizing over an optimization parameter set Φ(k,t) that includes quantum communication component parameters. The optimization can be done by simulated annealing. The optimization hyperparameters include the maximum number of iterations N it (t) per round k at trial t and the initial annealing temperature T. Alternatively, the optimization may be performed via, for example, brute-force search or grid search.

[0124] The optimization parameter set Φ(k,t) includes the component set parameters of the component set s(k,t) and the post-selection mask parameters, i.e., the bit values, of whether the measurement results are accepted or rejected. Since the component set s(k,t) is compressed, its length l k(t) (corresponding to the number of phase shifters and beam splitters each described by one parameter) is l k (t) satisfies (t) ≤ k. Therefore, the set of optimization parameters Φ(k,t) = {Φ1(k,t), Φ2(k,t),..., Φ z (k,t)} is in a z-dimensional parameter space, where z = 1 + x + l k (t), and x indicates the length of the post-selection mask. The additional dimension +1 corresponds to the intensity of the quantum state created in the first communication device 40, particularly the initial quantum state |Ψ0〉.

[0125] The optimization routine can start, for example, at an initial point having a photon number equal to 1 for a small value of r E ≪ 1, or, in another case,

Number

[0126] Subsequently, one of the optimization parameters Φ m (k,t) is modified. For a more detailed description, refer to Algorithm 1 below. Then, for the resulting modified set of optimization parameters Φ’(k,t), the key rate

Number

[0127] The key rate value

Number

[0128] The iteration continues until the maximum number of iterations N it (t) (e.g., N_it(t) = 19990 + 10t), or until the current key rate [Number] becomes greater than the previous maximum key rate R(t - 1), i.e., [Number] up to.

[0129] [Number] In that case, trial t ends, the component set s k (t) is set to s0(t + 1), and the maximum key rate is [Number] updated to, and the reward value r(t) = 1 is supplied to the reinforcement learning agent.

[0130] If the optimization routine does not achieve a value greater than R(t - 1) after N it (t) iterations, the round / agent-environment interaction step k ends, but trial t continues by adding another quantum communication component with the action a k+1 (t) to the component set. However, at the last round of the trial (i.e., k = k max, for example, k max = 20), or the current component set length l k (t) is already maximum (i.e., l k (t)=L max , the maximum component set length L max has, for example, a value of L max = 8), the trial t ends with a reward value r(t) = 0, and the component set is reset to the empty set.

[0131] Algorithm 1: Simulated Annealing Input: Φ(k,t), R k (t), R(t - 1), T i (t) Output: Φ(k,t), R k (t), R(t - 1), f Set the flag f indicating whether to interrupt the annealer to false. Uniformly select one of the optimization parameters Φ(k,t) from the set of optimization parameters Φ(k,t). If Φ(k,t) is intensity, Change Φ(k,t) to Φ’(k,t)=Φ(k,t)+δ ∫ , where δ ∫∈R is sampled from a uniform distribution in the interval (-100, 100); Otherwise, if Φ(k,t) is phase, Change Φ(k,t) to Φ’(k,t)=Φ(k,t)+δ ∫ , where δ ∫∈R is sampled from a uniform distribution in the interval (-0.05, 0.05); Otherwise, Φ(k,t) is a candidate selection mask bit, and invert Φ(k,t). For the modified set of optimization parameters Φ’(k,t), calculate the key rate

Number

Number

[0132] When the termination criterion is met, the output maximum key rate R, as well as the optimal set of components s and the optimal set of setup parameters Φ, are returned.

[0133] Figure 6 shows a plot of the output key rate R as a function of the ratio r E of the intercepted intensity (d = 3 optical modes, K = 3 states, T = 0.9). For a given r E each, the proposed method determines a different quantum communication setup with the maximum key rate. When the value of r E is small, the highest output key rate R can be achieved, but when the value of r E is large, a lower key rate is determined.

[0134] Figure 7 shows d = 3 optical modes and rE For \( = 01\), a graphic representation of the quantum communication setup obtained by the present method is shown. At the start of the reinforcement learning process, the optical setup (a) provides a key rate \(R = 10\) -5 and the optical setup (b) yields a key rate \(R = 1.68\), while the optical setup (c), i.e., the optimal quantum communication setup, gives the highest key rate, i.e., an output key rate \(R = 2.45\).

[0135] In FIG. 7(a), two optical components 70, 71 corresponding to the beam splitter BS 23 and the phase shifter PS2 are shown respectively. The optical components and setup parameters corresponding to FIG. 7 are as follows (see the table in FIG. 3).

[0136]

Table 1

[0137] After establishing the set of optimal components and the set of optimal setup parameters, the corresponding quantum communication setup can be physically implemented for quantum key distribution.

[0138] The features disclosed in this specification, the figures and / or the claims may be materials for realizing various embodiments, either alone or in various combinations thereof.

Explanation of Signs

[0139] 10 Meta - parameter 11 Characteristics of the quantum communication channel 12 Hyper - parameter 13 Learning agent module 13a Environment 14 Optimization module 15 Compression module 16 Physical simulation module 17 Post - processing module 18 Security proof module 19 Output 19a Feedback signal 20 Simulation network 21 First node 22 Second node 23 Edge 40 First communication device 41 Quantum random number generator 42 Creation unit 43 Second communication device 44 Eavesdropping device 45 Measurement unit 50 Eavesdropping beam splitter 51 Further beam splitter 52 Displacement operation 53 Detector 70,71 Optical components

Claims

1. A method for determining a quantum communication setup, the method being executed on a computer comprised in a data processing system, providing a set of components representing a quantum communication setup including at least one quantum communication component of a first communication device (40), a second communication device (43), and an eavesdropping device (44); selecting an operation from a set of operations, each representing a further quantum communication component and each being selectable with a selection probability depending on the set of components and a reward value; including the further quantum communication component in the set of components; determining a quantum model including at least one quantum state of at least one of the first communication device (40), the second communication device (43), and the eavesdropping device (44) from the set of components including the further quantum communication component; determining a maximum key rate by optimizing with respect to a set of optimization parameters including quantum communication component parameters from the set of components including the further quantum communication component and the quantum model; adjusting the reward value according to the size of the maximum key rate with respect to a previous maximum key rate; iteratively repeating the above steps until an end criterion is met and an optimal set of components and an optimal set of setup parameters for determining the quantum communication setup are provided; A method, comprising.

2. The method according to claim 1, wherein the quantum communication component includes at least one of an optical component, in particular a beam splitter and a phase shifter.

3. The method according to claim 1, wherein the selection of the operation is performed by a reinforcement learning agent.

4. The method according to claim 1, wherein the selection probability for selecting one of the set of operations corresponds to a reinforcement learning policy.

5. The method according to claim 1, further comprising removing identical quantum communication components from the set of components.

6. The method according to claim 1, wherein the quantum state includes a plurality of optical modes.

7. The method according to claim 1, wherein the quantum model further includes a quantum measurement executed in the first communication device (40) and / or the second communication device (43).

8. The method according to claim 1, further comprising determining, from the quantum model, a post-selection set including measurement results received by the first communication device (40) and the second communication device (43).

9. The method according to claim 1, wherein the maximum key rate is determined by a global optimization method, in particular an annealing method.

10. The set of optimization parameters includes at least one of an intensity value, a phase shift value, and a post-selection mask bit, and / or the set of optimal setup parameters includes The method according to claim 1, including at least one of an optimal intensity value and an optimal phase shift value.

11. Before adjusting the reward value, if the maximum key rate is less than or equal to the previous maximum key rate, another additional quantum communication component is repeatedly included in the set of components until the maximum number of repetitions is reached or the maximum key rate becomes greater than the previous maximum key rate, further including updating the quantum model, the post-selection set, and the maximum key rate. The method according to claim 8.

12. A quantum communication setup comprising a quantum communication component that is an optical component, created according to an optimal set of components and an optimal set of setup parameters generated by the method according to any one of claims 1 to 11.

13. A computer program, wherein when the computer program is executed on a computer comprised in a data processing system, the computer-readable instructions cause the computer to execute the method according to any one of claims 1 to 11. A computer program.

14. A data processing system, The data processing system comprises a computer, The computer is providing a set of components representing a quantum communication setup including at least one quantum communication component of a first communication device (40), a second communication device (43), and an eavesdropping device (44); selecting an operation from a set of operations, each representing an additional quantum communication component and each selectable with a selection probability depending on the set of components and a reward value; including the additional quantum communication component in the set of components; Determining a quantum model including at least one quantum state among the first communication device (40), the second communication device (43), and the eavesdropping device (44) from the set of components including the additional quantum communication components; Determining a maximum key rate by optimizing with respect to a set of optimization parameters including quantum communication component parameters from the set of components including the additional quantum communication components and the quantum model; Adjusting the reward value according to the size of the maximum key rate with respect to the previous maximum key rate; Repeatedly repeating the above steps until an end criterion is satisfied and an optimal set of components and an optimal set of setup parameters for determining the quantum communication setup are obtained; A data processing system configured to determine the quantum communication setup by performing the above.

Citation Information

Patent Citations

  • CVQKD (Continuous Variable Quantum Key Distribution) real-time performance optimization method and system based on machine learning

    CN107612688A

  • Quantum key distribution parameter optimization method based on random forest algorithm

    CN110798314A

  • Quantum cryptography system stability control method based on GRU model

    CN111738426A

  • Learning device, information processing system, learning method, and learning program

    WO2019225011A1