Quantum transmitter, optimization method, quantum transmitter, training method and quantum communication system

A reinforcement learning agent optimizes quantum transmitters by navigating control parameter spaces, addressing the complexity and cost of manual tuning in QKD systems, enabling efficient and scalable optimization across similar systems.

JP7797580B2Active Publication Date: 2026-01-13KK TOSHIBA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024122425
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-02-01
Filing Date
2024-07-29
Publication Date
2026-01-13
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

Optimizing quantum key distribution (QKD) systems is complex due to many coupled variables and nonlinear behavior, requiring manual tuning of control parameters which is time-consuming and costly, and new designs complicate manufacturing by introducing more manual optimization tasks.

Method used

Implementing a reinforcement learning (RL) agent to navigate a control parameter space and optimize quantum transmitters by learning to set control parameters, allowing self-tuning and generalization across similar systems without repeating the optimization procedure.

Benefits of technology

Enables fast and efficient optimization of QKD systems, reducing manual effort and costs, and allowing adaptation to new systems with similar dynamics, thus improving manufacturing scalability and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007797580000006
    Figure 0007797580000006
  • Figure 0007797580000007
    Figure 0007797580000007
  • Figure 0007797580000008
    Figure 0007797580000008
Patent Text Reader

Abstract

To provide a quantum transmitter enabling self-tuning of the hardware of a QKD transmitter, a method for optimizing, a quantum transmitter, a method for training, and a quantum communication system.SOLUTION: A method according to the present invention is a method for optimizing a transmitter using a reinforcement learning (RL) agent including a control unit to which a plurality of control signals defined by a set of a plurality of control parameters are applied and an optimization unit. The optimization unit sets the plurality of control parameters by performing reinforcement learning using a policy and the agent. The agent receives a plurality of observations of an environment, by moving through a control parameter space and obtaining scores indicative of the quality of the transmitter for a plurality of neighboring locations in the control parameter space, navigates a path through the control parameter space from first estimates of the plurality of control parameters to target settings for the plurality of control parameters, and is trained using a plurality of different quantum transmitters.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] SUMMARY OF THE INVENTION The embodiments described herein relate to quantum transmitters, methods of optimizing and quantum transmitters, methods of training and quantum communication systems. [Background technology]

[0002] In quantum communication systems, information is sent between a transmitter and a receiver by a single encoded quantum, such as a single photon. Each photon carries one bit of information, which can be encoded on a property of the photon, such as its polarization.

[0003] Quantum key distribution (QKD) is a technique that results in the sharing of a cryptographic key between two parties: a transmitter, often referred to as "Alice," and a receiver, often referred to as "Bob." The appeal of this technique is that it provides a test for whether any part of the key can be known to an unauthorized eavesdropper, often referred to as "Eve." In many forms of quantum key distribution, Alice and Bob use two or more non-orthogonal bases that encode bit values. The laws of quantum mechanics dictate that Eve's measurement of photons without prior knowledge of the respective encoding bases will inevitably result in changes to the state of some of the photons. These changes to the photon's state will introduce errors into the bit values ​​sent between Alice and Bob. By comparing portions of their common bit string, Alice and Bob can thereby determine whether Eve has acquired information. [Brief explanation of the drawings]

[0004] [Figure 1A] FIG. 1A is a schematic diagram of a Quantum Transmitter, according to one embodiment. [Figure 1B] FIG. 1B is a plot of the input and output signals of the components of FIG. 1A. [Figure 2]Figure 2A is a map of the experimentally measured QBER of the BB84 QKD protocol as a function of laser detuning frequency and injection ratio, where the box indicates a promising region for optimal operation. Figure 2B is a zoomed-in plot of the boxed region. [Figure 3] Figure 3 shows a schematic diagram of an agent in a reinforcement learning setup. [Figure 4] Figure 4 is a schematic diagram of how an RL agent can be used to navigate a parameter space. [Figure 5] Figure 5 is a map of the parameter space with the agent showing its nearest neighbors. [Figure 6] FIG. 6 is a schematic diagram of a system for training an RL agent. [Figure 7] FIG. 7 is a flowchart illustrating a method for training an RL agent. [Figure 8] Figure 8 is a plot showing the training of an RL agent with success rate plotted against episodes. [Figure 9] Figure 9 is a plot of how the RL agent is used to optimize the transmitter. [Figure 10] Figure 10 is a plot of QBER against the number of steps taken by the agent. [Figure 11] FIG. 11 is a plot of success rate versus the threshold QBER used as the criterion for success for RL agents trained using the method of FIG. 7 and for random agents. [Figure 12] FIG. 12 is a map of the parameter space showing the path taken by the agent from the starting point to the target point. [Figure 13] FIG. 13 is a schematic diagram of a point-to-point QKD system. [Figure 14] FIG. 14 is a schematic diagram of a point-to-point QKD system with control electronics. [Figure 15]FIG. 15 is a schematic diagram of a QKD emitter with a CW laser. [Figure 16A] FIG. 16A is a schematic diagram of a QKD emitter with a primary and secondary laser seed configuration. [Figure 16B] FIG. 16B is a schematic diagram of the corresponding drive signal and light output. [Figure 17A] FIG. 17A is a schematic diagram of the primary and secondary laser configuration. [Figure 17] Figure 17B is a plot of the optical frequency of the first-order laser under a small perturbation of the control gain of duration tm. Figure 17C is a plot of the optical phase trajectory with and without perturbation of the first-order laser. Figure 17D is a plot of the output pulse of the second-order laser. [Figure 17E] FIG. 17E is a schematic diagram of a gain modulation circuit for driving the primary and secondary lasers. [Figure 17F] FIG. 17F is a series of five time-dependent plots, beginning with the top plot: modulation of the primary laser, carrier density of the primary laser, output of the primary laser, modulation of the secondary laser, and output of the secondary laser. DETAILED DESCRIPTION OF THE INVENTION

[0005] In one embodiment, there is provided a quantum transmitter comprising: a plurality of transmitter elements, each of the transmitter elements comprising a radiation source; a control unit configured to apply a plurality of control signals defined by a plurality of sets of control parameters to at least one of the plurality of transmitter components; Optimization Unit and Equipped with wherein the optimization unit is configured to set a plurality of control parameters by performing reinforcement learning using a policy and an agent, wherein the agent moves through a control parameter space, and the agent: receiving a plurality of observations of its environment by obtaining scores for a plurality of neighboring locations of its in the control parameter space, the scores indicating a quality of the quantum transmitter; Navigating a path through a control parameter space from first estimates of a plurality of control parameters to target settings for the plurality of control parameters It is configured as follows: Here, the agent provides a quantum transmitter that has been trained using a number of different quantum transmitters.

[0006] The transmitter described above enables self-tuning of complex QKD transmitter hardware. Previously, determining the optimal control parameters and drive signals for all QKD hardware was performed manually. While much of this was done during the system design phase based on known hardware specifications, other components must be adjusted during system operation in response to environmental changes. The situation is further complicated by inherent variations among components; even if one QKD system is fully optimized manually, optimal parameters for systems with similar components may differ slightly. This means that the current manufacturing process for QKD systems involves various manual tasks to fine-tune the settings for each device. Such optimization is a time-consuming task requiring trained users with expertise in optics, electronics, and QKD system design. This requirement imposes high costs and impacts manufacturing scalability.

[0007] Furthermore, ongoing research leads to the development of new and improved QKD system designs with better performance / fewer components. However, these benefits are often achieved by precisely exploiting laser dynamics and electrically controlled light-matter interactions, and therefore such "improved" designs (such as phase seeds) may actually complicate the manufacturing phase of QKD technology by introducing more manual optimization tasks.

[0008] The task of optimization is extremely complex because the underlying problem involves many coupled variables and nonlinear behavior (e.g., laser dynamics is fundamentally nonlinear). Therefore, optimal performance cannot be easily found by scanning all of the parameter space or by applying simple feedback loops to each variable in turn (i.e., the problem is not convex). Optimization of the complete system requires a holistic approach that simultaneously considers many variables while also measuring multiple output properties that affect overall QKD performance.

[0009] The above provides a method for allowing any quantum transmitter to tune its own parameters. By training the RL model across multiple different environments, the new RL model learns how to learn.

[0010] The above allows generalizable optimization to be performed on quantum transmitters, allowing the optimization, once trained, to be applicable to new systems without having to repeat the entire optimization procedure. Deep reinforcement learning (RL) is implemented when a neural network can be trained and the knowledge acquired during optimization can be retained. This training allows the neural network to adapt and apply the learned optimization strategy to new, but similar, quantum transmitter systems. Framing the optimization problem as a pathfinding problem in which the RL agent must identify an efficient route to locate the goal, which is the optimal operating region of the quantum transmitter in parameter space, allows for effective generalization across different environments, e.g., different lasers, but shares the same underlying dynamics.

[0011] In one embodiment, the quantum bit error rate (QBER) will be used as a "reward" in guiding the RL agent to locate optimal operating conditions, while the RL agent will adjust quantum transmitter parameters such as laser bias current, optical injection power, etc. During the training phase, the RL agent will be trained on a variety of similar QKD transmitters to derive a general optimization strategy. After training, during the inference phase, the trained RL agent may then be applied to new quantum transmitters of the same type and perform optimization on the direct parameters by applying the derived strategy.

[0012] In an implementation, this optimization can be implemented in the software layer of the product in the form of a trained neural network, which does not require any additional hardware.

[0013] The above is an optimization technique that allows fast optimization to be performed directly after training, which can be applied to quantum communication systems, including chip-based systems, where optimal operating parameters need to be determined.

[0014] The target setting for a control parameter is the setting at which the score satisfies a condition, such as being below a threshold if the score is related to the error rate. This is sometimes called the optimal condition.

[0015] In one embodiment, the radiation source is a source of pulsed radiation. In a further embodiment, the pulsed radiation source comprises a primary laser and a secondary laser, the primary laser configured to apply a seed pulse to the secondary laser.

[0016] In one embodiment, the agent: receiving a plurality of observations of a current state thereof; and performing a next action for the agent based on the RL policy and a plurality of observations of the agent in a current state; receiving a reward for the agent; The method is configured to find a path through the control parameter space by:

[0017] A state refers to a complete description of the status of the environment at a given time. It is the precise information that describes the current situation in which the agent finds itself. Thus, in this case, the state is a complete description of the laser's dynamics—for example, the photon density and carrier density fluctuations in the laser. This is something that cannot be directly observed. However, observations can be made. An observation is what the agent perceives from the state of the environment. In a fully observable environment, the agent can observe everything about the state of the environment (e.g., playing a chess game), so observations are effectively the same as state. However, in the case of the transmitter above, we only have a partially observable environment, where observations are partial information about the state. When the agent updates its control parameters, its position in the parameter space will change accordingly, and as the laser's dynamics change, so will the state. In the case of the transmitter above, the observations made by the transmitter contribute to the optimization process. If the observations are the QBER (or other parameters related to the score) in its neighboring grid cells (and its own grid cell), the agent can generalize and perform well in similar environments.

[0018] A reward may be calculated from the score. The reward function may be shaped to heavily reward good updates to the state. For example, the reward function may be given as follows: Reward r=-QBER+R*s, QBER≦QBER optimum If so, then s=1, otherwise s=0. R is for QBER optimum is a constant for a large reward when R is reached. For example, R can be a value between 5 and 100, e.g., 10. In one embodiment, QBER optimummay be set at 3.5%, but other values ​​that allow good performance of the quantum transmitter may be used.

[0019] The above used QBER as the score. However, the above QBER can be replaced by other ways of monitoring the score. In one embodiment, the score is selected from at least one of a quantum bit error rate (QBER), a secure bit rate (SBR), phase information of the coded pulses, count rates for all pulses received at the receiver's detectors, count rates for pulses with predetermined coding, shapes of the received pulses, or arrival times of the pulses at the detectors.

[0020] In further embodiments, the reward may take the form of: r=-log(f(x,y)+1)+K*s, f(x,y)= <f opt If so, then s=1, otherwise s=0. where f(x,y) is the score at location x,y. It will be appreciated that this is an example for a 2D environment. In the case of an n-dimensional environment, then f describes the n-dimensional space. opt is the optimal value of the score, and K is a constant that heavily weights low-scoring QBER. For example, K may be 100. Other forms of reward functions may be used.

[0021] In one embodiment, the observations of neighboring locations are scores at neighboring locations in the control parameter space. Observations at the current agent location can be used in addition to observations in neighboring grid elements. These observations can also include the following additional parameters: Using the agent's surrounding terrain as observations allows the agent to be trained to recognize the terrain. Actions are therefore based on its local view rather than its current location. When the map environment changes slightly, the agent will still be able to detect its common features. Every additional "view" is one additional measurement.

[0022] In some embodiments, multiple additional observations are given to the agent, for example: P fr -Free running laser power P inj -Injected power Δf - frequency detuning between lasers

[0023] The RL agent's observations are designed to incorporate performance metrics that may be quantum bit error rate (QBER) and other variables, e.g., laser output power, that correlate with the underlying dynamics of the system to allow generalization in optimization. In one embodiment, the RL framework is as follows: · Environment={env(1),env(2),env(3)...} Observation = (QBER,Pfr,Pinj,Δf)(a,b), (QBER,Pfr,Pinj,Δf)(c,d)...where (a,b), (c,d) are the neighboring coordinates of the agent. Action = {+1,+0,-1} for each parameter Reward r=-QBER+R*s, and QBER≦QBER optimum If s = 1, otherwise s = 0, and R is a constant for large reward when the optimum is reached. where env(.) refers to the different environments that will be trained, and P fr is the free-running laser power, and P inj is the injection power and Δf is the frequency detuning between the lasers. optimum is the optimal QBER. During training, an environment is chosen randomly for each episode, allowing the agent to derive a generalized policy capable of solving all given environments.

[0024] In one embodiment, the plurality of control signals are a plurality of electronic control signals for a pulsed radiation source, the plurality of control signals comprising at least one of an intensity of the electronic control signal, a shape of the electronic control signal, and a DC offset of the electronic control signal. The pulsed radiation source may be provided on a temperature control element. In this situation, at least one of the control signals may be a control signal to the temperature control element. The set of control parameters may comprise a plurality of control parameters for both the primary laser and the secondary laser. The control signal may also be selected from at least one of a DC bias for the primary laser, a DC bias for the secondary laser, an injection power, a wavelength of the secondary laser, a peak-to-peak voltage of the laser, and a delay between injection-locked lasers.

[0025] In one embodiment, the plurality of transmitter components further comprises a modulation unit, the modulation unit configured to randomly encode the plurality of pulses of radiation.

[0026] As mentioned above, the quantum transmitter may form part of a quantum communication system, which may be selected from a point-to-point quantum communication system, a measurement device independent quantum communication system, and a twin-field quantum communication system.

[0027] In a further embodiment, there is provided a method of optimizing a quantum transmitter, the quantum transmitter comprising: a plurality of transmitter elements, each of the transmitter elements comprising a radiation source; a control unit configured to apply a plurality of control signals defined by a plurality of sets of control parameters to at least one of the plurality of transmitter components; Equipped with The method comprises: setting a plurality of control parameters by performing reinforcement learning using a policy and an agent, wherein the agent moves through a control parameter space, and the agent: receiving a plurality of observations of its environment by obtaining scores for a plurality of neighboring locations of the quantum transmitter, the scores indicating a quality of the quantum transmitter; Navigating a path through a control parameter space from first estimates of a plurality of control parameters to target settings for the plurality of control parameters It is configured as follows: Here, a method is provided in which an agent is trained using a plurality of different quantum transmitters.

[0028] In a further embodiment, there is provided a method of training an RL agent for use in optimizing a quantum transmitter, the quantum transmitter comprising: a plurality of transmitter elements, each of the transmitter elements comprising a radiation source; a control unit configured to apply a plurality of control signals defined by a plurality of sets of control parameters to at least one of the plurality of transmitter components; Equipped with The method comprises: training an RL agent to optimize a plurality of control parameters by performing reinforcement learning using the policy, wherein the RL agent is trained to move through a control parameter space and navigate a path from first estimates of the plurality of control parameters to target settings of the plurality of control parameters; Here, a method is provided in which an RL agent is trained using multiple training transmitters.

[0029] Training an RL agent to optimize multiple control parameters is selecting a training transmitter from a plurality of training transmitters; and navigating a path in a control parameter space to a target setting of a plurality of control parameters; and Here, navigating a route involves: receiving a plurality of observations of the RL agent at its current state in the selected transmitter by obtaining scores for a plurality of its neighboring locations in the control parameter space, the scores indicating the quality of the quantum transmitter; performing a next action for the RL agent based on the policy and a plurality of observations of the RL agent in a current state; receiving a reward for the RL agent; and Updating the RL agent's policy and Equipped with Here, navigating the route is performed multiple times for each of a plurality of training transmitters.

[0030] In one embodiment, training is provided using actual measurements from multiple transmitters. However, responses from one or more of the transmitters may be simulated by software to replicate the response of a real system. In one embodiment, responses from all training transmitters are simulated. In one embodiment, three or more training transmitters are used. In one embodiment, for a batch of quantum transmitter chips, an agent may be trained on a portion of the chips such that it is then applied to the remaining batch.

[0031] The above has focused on quantum transmitters. However, the principles may be applied to other components such as optical modulators or receivers. Thus, according to a further embodiment, there is provided a quantum communication system comprising: a transmitter comprising a plurality of transmitter elements, the plurality of transmitter elements comprising a pulsed radiation source and a modulation unit, the modulation unit configured to randomly encode the plurality of pulses of radiation; a receiver comprising a plurality of receiver elements, the plurality of receiver elements comprising a demodulator and a detector configured to decode and detect the plurality of randomly coded pulses; Equipped with The quantum communication system further comprises a control unit and an optimization unit, wherein the control unit is configured to apply a plurality of control signals defined by a plurality of sets of control parameters to at least one of the plurality of transmitter components and the plurality of receiver components; wherein the optimization unit is configured to set a plurality of control parameters by performing reinforcement learning using a policy and an agent, wherein the agent moves through a control parameter space, and the agent: receiving a plurality of observations of its environment by obtaining scores for a plurality of neighboring locations of its in the control parameter space, the scores indicative of a quality of the quantum communication system; Navigating a path through a control parameter space from first estimates of a plurality of control parameters to target settings for the plurality of control parameters It is configured as follows: Here, a quantum communication system is provided in which agents are trained using a plurality of different quantum communication systems.

[0032] The control signal for the modulator is an electronic control signal for the modulator comprising at least one of a magnitude of the electronic control signal, a shape of the electronic control signal, and a DC offset of the electronic control signal.

[0033] The control signal may be an electronic control signal for the demodulator comprising at least one of a magnitude of the electronic control signal, a shape of the electronic control signal, and a DC offset of the electronic control signal.

[0034] The control signal may be an electronic control signal for the detector comprising at least one of a magnitude of the electronic control signal, a shape of the electronic control signal, and a DC offset of the electronic control signal.

[0035] The optimization unit may be provided in either or both the transmitter and the receiver. It is also possible for the transmitter and / or the receiver to have their own optimization unit. Thus The above goes beyond learning from past experiences and applying them to similar scenarios in the future.

[0036] In one embodiment, a performance metric of our system, e.g., QBER, and optionally an experimental observable correlated with the map's underlying structure, e.g., laser output power, are used to guide an agent through the map. The agent's observations are a combination of experimental observables related to QBER performance within its local neighborhood in parameter space. This essentially allows us to train the agent to recognize local terrain. Thus, when training across different maps with similar terrain, the agent will be forced to learn to capture the map's underlying features that remain invariant—a generalization process. After training, the agent will be able to navigate through previously unseen maps that have similar underlying features. In effect, this is equivalent to optimizing a new quantum transmitter system that shares the underlying dynamics as the previous one.

[0037] This allows for fast and efficient optimization. Once trained, the model can be applied to a similar system and directly tune the system to its optimal settings without having to repeat the entire parameter search procedure. All the information needed for this optimization method is readily accessible in practice; for example, QBER is constantly measured in quantum communication systems.

[0038] In one embodiment, the control signal is selected from at least one of a DC bias for the primary laser, a DC bias for the secondary laser, an injection power, a wavelength of the secondary laser, a peak-to-peak voltage of the laser, and a delay between injection-locked lasers. FIG. 1A illustrates a self-tuning QKD transmitter according to one embodiment. A specific example is described to visualize the process. However, other types of transmitters or other components may be supported. The example of FIG. 1A enables autonomous optimization of an optical injection laser (OIL) system. In one embodiment, the method is applied to optimize QBER for the BB84 protocol. However, optimization could also be applied to pulse-to-pulse phase coherence in addition to, or instead of, QBER optimization.

[0039] In this example, the transmitter comprises two distributed feedback (DFB) lasers in an OIL configuration, with light from the primary laser injected into the cavity of the secondary laser via an optical circulator. In some texts, the primary laser is referred to as the "master" laser and the secondary laser as the "slave" laser. A variable optical attenuator (VOA) is used to control the injection power. Each laser wavelength is thermally stabilized using an integrated thermoelectric cooler, which can be adjusted via a controller. An RF signal and a DC bias from a current source are combined using a bias tee to drive the two lasers. The primary laser is gain-switched to generate a 1 GHz pulse train. Between pulses, the primary laser is driven below the lasing threshold to ensure each generated pulse has a random phase. The primary laser pulse is injected into the secondary laser, which is gain-switched at 2 GHz, to generate short pulses with a duration of ~70 ps. The RF signals of the two lasers are shown in Figure 1B. The two lasers are time-aligned such that each primary pulse seeds two secondary laser pulses, forming early and late bins of a single clock cycle (i.e., a single qubit) that share the same globally random phase.

[0040] To encode the relative phase between two secondary laser pulses, the RF signal of the primary laser is modulated by adding a small amplitude perturbation during the time interval between the secondary laser pulses. The perturbation alters the carrier density of the primary laser cavity, which in turn alters its emission frequency and its phase evolution. When seeded by injected primary photons, the secondary laser pulse inherits the phase of the primary pulse. The induced phase difference in the primary laser pulse is then transferred to the phase between successive secondary laser pulses, thereby achieving direct phase encoding. The applied phase shift can be precisely controlled by changing the amplitude of the electrical perturbation signal. A VOA is used to attenuate the pulses before sending them to the quantum channel.

[0041] In the receiver, an asymmetric Mach-Zehnder interferometer (AMZI) is used to decode the relative phase between the secondary laser pulses. The long arm of the AMZI has a 500 ps delay that matches the temporal separation of the successive secondary laser pulses. Depending on their relative phase, the successive secondary laser pulses can interfere constructively or destructively, thus allowing bits "0" and "1" to be assigned to the two output ports. The AMZI output is measured using a photodiode or single-photon detector. Control electronics control all of the laser electronics and can remotely set the values ​​for each parameter. The output of each parameter set is then measured and serves as feedback to the evaluation algorithm.

[0042] An example of laser parameters that can be optimized is listed in Table I. Generally, to achieve a stable OIL, the injection power from the primary laser and the frequency detuning between the primary and free-running secondary lasers, which depend on temperature and bias current, must be carefully selected. The OIL dynamics become more complex under gain-switching operation. The injection power from the primary laser must be strong enough to overcome the effects of spontaneous emission noise on the phase in the secondary laser, but excessive injection power can create undesirable parasitic effects and degrade performance. Furthermore, the bias current of the primary laser must be set to a level that allows the laser to be driven below threshold between each pulse for phase randomization, which also affects other important output properties, such as pulse phase and duration. To transfer phase, the two lasers must be time-aligned, and the duration of the primary laser pulse must be long enough to seed the generation of two consecutive secondary laser pulses. When phase modulation is considered, the implemented phase depends on the amplitude modulation applied to the drive signal of the primary laser. Therefore, it is necessary to adjust all of these parameters to take advantage of the benefits of OIL.

[0043] To investigate the complexity of the laser dynamics, we measured the QBER for the BB84 QKD protocol as a function of the frequency detuning between the two lasers and the injection ratio (defined as the ratio between the injected primary power and the free-running secondary power) with all other parameters fixed at predetermined optimal values, as shown in Figure 2A. The promising operating region is indicated by the box in Figure 2A and further expanded in Figure 2B. While QBER is affected by many factors, the fringes observed in Figure 2A are likely due to changes in the phase relationship between the primary and secondary lasers, which results in encoding errors in the relative phase between the secondary laser pulses. From Figure 2B, the sparsity of the optimal regime can be clearly observed. Mapping QBER can take longer than 8 hours to complete, even when limited to two parameters. This therefore highlights the need for efficient methods to determine the optimal operating regime, especially in large parameter spaces.

[0044] Phase encoding is widely used in QKD protocols, where a secret bit is encoded in the relative phase between successive pulses. Therefore, high phase coherence between pulses contributes to the quality of the system. Furthermore, a key requirement for secure quantum communication is that the phase of each qubit, consisting of early and late bins, is uniformly random. This allows the coherent states of the decaying pulses to be treated as photon number states, and security proofs against the most common attacks can be obtained.

[0045] OIL combined with gain switching represents a highly efficient way to generate pulses that meet these requirements. As explained above, gain switching allows each primary pulse to carry a random phase, while optical injection seeding allows the phase manipulation on the primary pulse to be coherently transferred to the relative phase between successive secondary laser pulses.

[0046] To investigate phase coherence, the primary laser is pulsed without additional modulation. Therefore, two secondary laser pulses seeded by the same primary pulse are in phase, and constructive and destructive interference can be obtained. In contrast, secondary laser pulses seeded by different primary laser pulses do not have a clear phase relationship, and therefore, interference should result in minimal visibility. To satisfy these conditions, the following fitness function is used, and the algorithm aims to maximize it by optimizing the parameters shown in Table I.

[0047] [Table 1] Table I: Input parameters for phase coherence and QBER optimization

[0048]

number

[0049] where V coherent (V random ) is the coherent visibility of a phase-coherent (phase-randomized) secondary laser pulse.

[0050] As mentioned above, there are many parameters that can affect the quality of a quantum transmitter. In this embodiment, QBER will be used. QBER is the quantum bit error rate, and is measured using the apparatus of claim 1 by determining whether the measured bits correspond to the sent bits.

[0051] In this example, again, assume the control parameters are: DC bias for the master laser DC bias for the slave laser injection power Slave laser wavelength

[0052] However, other parameters may be used, such as the peak-to-peak voltage of the primary and secondary lasers, and the delay between the primary and secondary lasers.

[0053] In one embodiment, the control parameters are set using a reinforcement learning setting. Figure 3 shows a reinforcement learning (RL) setting with an RL agent 201 interacting with an environment 203 during discrete time steps. At each time step t, the agent learns the current state s of the environment. t Observe and act a t Following this action, the environment provides an immediate reward, t Given a new state s t+1 The agent's objective is to maximize the cumulative reward it receives during its interactions with the environment. To achieve this, the agent learns through trial and error, gathering information from the environment and determining the best action to take in response to each observation. The objective is to train the agent to discover an optimal policy that maximizes the reward. To visualize the above, one can think of the agent as a robot that observes its surroundings and attempts to discover how it can avoid an environment 203, such as a maze, with the outcome indicating whether the robot has encountered an obstacle.

[0054] In one embodiment, to apply RL to optimize transmitter parameters, the task is first framed as a pathfinding problem in which an agent needs to learn to find the best path to travel from a source to a destination in a grid world. The objective is for the agent to learn a policy that maximizes the total reward.

[0055] In this framework, it is possible to consider the laser optimization landscape as a discrete grid world in which an agent must identify an optimal operating position, navigated by system performance as a form of reward. In one embodiment, the metric that quantifies the system performance of a quantum transmitter is QBER, which depends on the transmitter's driving parameters. However, other metrics, such as phase visibility or the contrast between on and off signals, could be used.

[0056] Similar to coordinates in a physical location space, the control parameters of a quantum transmitter can be imagined in a multidimensional control parameter space, with each parameter representing a dimension. Control parameter selection is performed by allowing agents to discover paths between control parameter settings.

[0057] Figure 4 shows a simplified diagram of a problem in which an agent 211 can move horizontally or vertically on a grid 213. The grid 213 is shown as simply 2D due to the variation of only two parameters. However, the dimensionality of the grid will expand with the number of parameters, and the agent is only allowed to move one step in one direction per action. However, in other embodiments, the agent may also be able to move more than one step diagonally during an action. In the current implementation, an action is a number between -1 and 1. This can be set between -1 and -0.3 to move the agent back one step, between -0.3 and 0.3 to stop the agent, or between 0.3 and 1 to move the agent forward one step. Further options can be added for the agent by adding more slices between -1 and 1.

[0058] An RL agent must learn to navigate an environment by making decisions at each step while moving from a start point to a goal. To aid in learning, the agent will receive a positive reward for reaching the goal and a negative reward for hitting an obstacle or taking too many steps. The primary objective in a pathfinding problem is for the agent to learn a policy that maximizes the total reward, which usually means finding the shortest or most efficient path to the goal without hitting any obstacles. This can be summarized as follows: 1. Environment = Grid World 2. Observation = Agent coordinates (x,y) 3.Action={up, down, right, left} 4. Reward = given by the value of each grid cell

[0059] In Figure 4, the reward for moving to each square of the grid is shown as a number on each grid. The goal is to reach the square with a reward of +10.

[0060] The environment for each transmitter will be slightly different, and the optimal parameters for each transmitter will be different; otherwise, it will be possible to apply the optimal parameters for each transmitter. Therefore, an RL policy is trained that will enable the agent to discover the best path through an unknown environment. To achieve this, the policy of the RL model is trained for multiple different transmitters.

[0061] The task is therefore not just to locate the optimal operating region of the laser, but also to generalize the learning so that the RL agent can still perform the optimization effectively when encountering new systems with similar underlying dynamics.

[0062] An RL agent can be trained using a variety of similar environments that share the same underlying dynamics. Thus, the agent will be exposed to different environments during its training, making it possible to derive a generalized policy for navigating this type of environment, thereby autonomously minimizing QBER.

[0063] To model this approach, a 2D environment is simulated using a test function known as the Beale function, a function in the study of optimization problems, to mimic complex laser dynamics. The function is given below and has a minimum at (3,0.5), where f(3,0.5)=0. Applying different transformations to the function, e.g., scaling variables and changing the values ​​of constants, can alter the contours of the function and thus shift the location of the minimum. Note that the Beale function is given to simulate QBER (or other score); however, QBER is measured when deployed for a real transmitter.

[0064] The agent's goal is to locate a minimum. When it reaches this target, it will receive a large reward. The value of the function f(x,y) is also used in the reward function as feedback to guide the agent through the environment. In a traditional pathfinding setting, the dimensionality of the grid world translates naturally into the optimization parameters of interest, and the current parameter values ​​are represented by the agent's coordinates in the (possibly multidimensional) grid world.

[0065] The RL settings are currently as follows: 1. Environment = Beale function f(x,y)=(1.5-x+xy) 2 +(2.25-x+xy 2 ) 2 +(2.625-x+xy 3 ) 3 2. Observation = Agent coordinates (x,y) 3.Action={+1,+0,-1} 4. The reward is r = -log(f(x,y)+1)+100*s, where s = 1 if f(x,y) = 0, and s = 0 otherwise.

[0066] However, in these settings, RL agents fail because the agent's observations are tied to coordinates in the environment. During training, the agent attempts to find the best route leading to a destination given a set of coordinates in the environment. When the environment changes, the agent simply has no way to distinguish where it is currently located in the environment. To overcome this problem, the representation of the agent's observations is reconstructed. Instead of using its own coordinates in the grid, the surrounding terrain in parameter space is used as the observed terrain in parameter space when observing it, as shown in Figure 5. In Figure 5, square 191 represents the neighborhood of agent 193.

[0067] The settings are summarized below: 1. Environment = Beale function f(x,y)=(1.5-x+xy) 2 +(2.25-x+xy 2 ) 2 +(2.625-x+xy 3 ) 3 2. Observation = ΣiΣjf(x+i,y+j), where (x,y) are the coordinates of the agent and i,j iterate over the coordinates of the local terrain surrounding the agent in the environment. 3.Action={+1,+0,-1} 4. The reward is r = -log(f(x,y) + 1) + 100s, where s = 1 if f(x,y) = 0, and s = 0 otherwise.

[0068] In addition to the above observations, the following additional observations are given: P fr -Free running laser power P inj -Injected power Δf - frequency detuning between lasers Additional observations can be given to the agent to help it learn. For example, when an agent sets a laser to, say, 50mA, informing the agent about how much power is actually coming out of the laser when it is powered at 50mA will help the agent understand the behavior of this particular laser, as each laser's output power is slightly different. Therefore, giving the agent these additional measurements can help the agent "understand" more about its environment.

[0069] Figure 6 is a schematic diagram of a system according to one embodiment that can be used to train an agent. The agent is trained to discover an optimal policy. During training, the agent establishes for itself a "strategy," which is a mapping between the actions it takes and the states it observes. In the system of Figure 6, there are three transmitters 201, 203, and 205 used for training, and a fourth quantum transmitter 207 used for validation.

[0070] Each of the four transmitters 201, 203, 205, and 207 is connected to control electronics 209, which can be used to vary the parameters of the transmitters (as described above with respect to FIG. 1). The control electronics 209 is connected to an RF switch 211, which allows the control electronics 209 to selectively communicate with each of the four transmitters.

[0071] The outputs of the four transmitters are connected via an optical switch to receiver 215. Receiver 215 is a basic receiver of the type described with reference to Figure 1 and is used in this example to measure the QBER.

[0072] During training, for each episode, one of the three transmitters 201, 203, and 205 used for training is randomly selected. Thus, each of the three training transmitters is optimized while the next action is randomly selected from each transmitter. Training is performed in this manner because it is necessary for the policy to be trained using all inputs from three different transmitters. If training were to be performed first on one transmitter and then sequentially, by the time transmitter 205 is used, the policy would begin to forget the training of the first transmitter, thus losing the benefit of training the policy to adapt to different environments because the policy would be dominated by the last environment the system was trained on.

[0073] 7 is a flowchart illustrating a training method according to one embodiment. In step S301, a training transmitter is selected. For example, referring again to FIG. 6, training transmitter 201, 203, or 205 may be selected. The training transmitter may be selected randomly, for example, via a switch.

[0074] In step S303, the agent receives observations given by the environment, which are values ​​of a region surrounding the agent in parameter space. When the control parameters of the quantum transmitter are updated, the current state of the agent is updated. When training begins, the control parameters may be set to initial values, which may be randomly selected or fixed values ​​such as the midpoint of the control parameter range.

[0075] The agent then receives an observation of its state for the selected transmitter. In this embodiment, the observation is the neighboring state, i.e., the QBER for its surrounding and central regions in the parameter space, e.g., the four neighboring cells in FIG. 5 plus the cell where the agent is shown (for a total of five cells). Thus, the parameter is set to a value in these five cells, and the corresponding QBER values ​​in these locations are measured and fed back to the agent. In other embodiments, a subset of neighbors may be used.

[0076] As described above, control parameters are envisioned in a multidimensional control parameter space, with each dimension defining a control parameter. Each control parameter is a variable over a range, and the range is divided into steps to define elements in the grid. The term "element" is used because the grid is n-dimensional. In some cases, to aid understanding, the term "square" may be used when a 2D view of the grid is shown. In one embodiment, neighbors are determined as control parameters one step away from the selected control parameter in both the positive and negative directions per axis. However, in other embodiments, different criteria may be used to determine neighbors, e.g., two or more steps per parameter, or grid elements diagonally adjacent to the selected control parameter grid element (center element).

[0077] To obtain an observation, the QBER is measured for each of the adjacent grid elements. Returning to the configuration of Figure 6, this is accomplished by connecting a selected transmitter to a receiver 215 via an optical switch 213. For each observation, a number of measurements are performed to establish the QBER. The observations are performed as follows: Observation = {f(x i ,y i )}, where (x i ,y i ) are the adjacent coordinates.

[0078] In step S305, the agent takes an action to update its state, which is selected from {+1, +0, -1} for each parameter.

[0079] Actions are selected according to the policy of the RL system. In one embodiment, the Soft Actor-Critic (SAC) algorithm, implemented using the Stable Baseline3 library, is used as the RL architecture. Generally, SAC implementations use separate neural networks for the actor (close to the policy) and the evaluator (close to the value function). The actor network consists of a multi-layer perceptron (MLP) with two hidden layers with 256 nodes per layer. Two evaluator networks are also used, both based on an MLP architecture with two hidden layers and 256 nodes per layer.

[0080] In step S307, the agent receives the reward. In one embodiment, the reward takes the form of: r=-log(f(x,y)+1)+100s, if f(x,y)=0 then s=1, otherwise s=0 where f(x,y) is the QBER for the state simulated by the Beale function.

[0081] In one embodiment, the policy is modeled as a state-action value function Q(s, a) defined as follows:

[0082]

number

[0083] When the agent takes action a to move from state s to s', the returned Q value, Q(s, a) = immediate reward r(s, a) + discounted maximum Q value (estimated future reward) in the next state s'. The policy is then updated in step S309.

[0084] Thus, through interactions with the environment, the agent gains experience from rewards received as its policy is updated.

[0085]

number

[0086] In step S311, it is determined whether an episode termination condition has been reached. An episode is a time during which a policy is updated using a single selected transmitter. The condition could be, for example, when a target is located (i.e., performance exceeds a threshold) or when the maximum number of time steps for the episode is reached.

[0087] If the episode termination condition is not met, the method loops back to step S303 and the agent makes an observation in its current state.

[0088] If an episode end condition is reached, the method then determines whether an end-of-training condition has been reached in S313. The end-of-training condition may be set to the number of episodes in which the target was successfully located. Once the end-of-training condition is met, the agent may be trained and deployed on a new transmitter.

[0089] The method then returns to step S301 where a transmitter is randomly selected. A single policy is updated for all three transmitters in step S311. In one embodiment, each transmitter is selected an equal number of times during training. Transmitters can be selected randomly, but they can also be selected sequentially. However, it becomes necessary to switch between transmitters to avoid training one transmitter dominating the policy.

[0090] When the Q-value function converges to its maximum value, the optimal policy can be obtained directly from it as follows:

[0091]

number

[0092] The end of training can be determined in several ways. For example, changes to the policy can be monitored until convergence. The QBER of the transmitter during training can be monitored until convergence. In another example, training runs for a set number of iterations.

[0093] Figure 8 shows a plot of the "success rate" during training for the system of Figure 6 trained according to the method of Figure 7. A "successful event" is defined as an event with a QBER < 3.5%. There is one trace per transmitter, and the QBER is measured for each transmitter. As the number of episodes increases beyond 5000, the QBER success rate is seen to approach 1 for all three training transmitters.

[0094] 9 is a schematic diagram of a flowchart for the inference stage, in which a trained policy is used to enable a new transmitter to learn its optimal parameters. In step S351, the transmitter is started. In step S353, the transmitter selects initial control parameters. These may be selected randomly, selected at the midpoint of a range, selected from previously known values, or selected from factory settings.

[0095] In step S355, the agent receives observations. As described above, for training, observations are received for the transmitter about its states. In this embodiment, the observations are QBERs for neighboring states (grid elements). As described above, the control parameters are envisioned in a multidimensional control parameter space, with each dimension defining a control parameter. Each control parameter is a variable that spans a range, and the range is divided into steps to define an element in the grid. The term "element" is used because the grid is n-dimensional. In some cases, to aid understanding, the term "square" may be used when a 2D view of the grid is shown. In one embodiment, neighbors are determined as control parameters one step away from the selected control parameter in both the positive and negative directions per axis. However, in other embodiments, different criteria may be used to determine neighbors, e.g., two or more steps per parameter, or grid elements diagonally adjacent to the selected control parameter grid element (center element).

[0096] The trained RL policy is then used in step S357 to take actions, obtain rewards, and select the next values ​​(next states) of the control parameters.

[0097] In step S359, the agent receives a reward. The reward is defined as described above during training. Next, in step S361, it is determined whether the agent successfully reached its target. It may be ascertained whether a quality parameter (QBER) has reached a threshold. For example, if QBER is less than a threshold (e.g., 3.5%), it is determined that the transmitter has found its optimal conditions. In other embodiments, step S361 would need to be satisfied a fixed number of times in succession, e.g., five times, before it is determined that the optimal operating conditions for the quantum transmitter have been obtained.

[0098] If in step S361 it is determined that the target has not been reached, the method loops back to step S355 and an observation is received for the state obtained as a result of the previous action.

[0099] The agent again takes action based on its RL policy and observations. A reward is then obtained in S359, and it is determined whether the success condition is met in S361. If it is determined that the transmission has reached its target (an acceptable QBER), the process is stopped.

[0100] The above method allows any quantum transmitter to learn its optimal parameters using reinforcement learning.

[0101] Figure 10 shows plots of QBER for eight trials showing the optimization performed by the agent. QBER is plotted against steps, with the time taken for 10 steps being approximately 5 minutes. The agent starts from a random state in each trial. In some trials, QBER is temporarily increased before decreasing to the optimum. This suggests that when QBER drops, the agent is actively predicting the optimum rather than simply following a path toward it.

[0102] Figure 11 is a plot of success rate versus the threshold for success, QBER. The top line shows the model, and the bottom line shows the results for a random RL agent. The success rate is approximately 80%. The higher the threshold, the easier the task and therefore the higher the success rate, and vice versa.

[0103] Figure 12 is a plot of the control space showing the agent's path from a random starting point 401 to an optimal point 403. We can see that the agent moves stepwise towards the optimal point.

[0104] Thus, the agent can be applied to new, unseen maps where the location of the minimum is different from those in the training set. Starting from a random location in the grid world, the agent can identify the location of the minimum with a high success rate, suggesting that the agent has successfully generalized its experience and transferred it successfully to newly encountered environments.

[0105] By closely examining the movement of an agent through the grid world, it can be observed that the agent recognizes certain patterns and learns how to navigate them, for example, the presence of ridges, troughs, and local minima. Regardless of its starting position, the agent is always observed moving toward the center.

[0106] In one embodiment, the above applies to quantum key distribution. Figure 13 shows a quantum key distribution (QKD) system. The simplified configuration of Figure 13 comprises a transmitter 1 and a receiver 3 connected by an optical communication channel 5 (e.g., optical fiber or free space).

[0107] Transmitter 1 (often referred to as "Alice") comprises a pulsed radiation source 7 and a state encoder 9. Transmitter 1 generates quantum states, which in one embodiment are coherent states formed by the pulsed laser emission of pulsed radiation source 7 and state encoder 9 (which modulates light on a random basis (e.g., X-based or Y-based) with both bit values ​​(0 or 1)). While a variety of different encoding schemes exist (e.g., polarization encoding, time bin encoding, etc.), all QKD protocols rely on fast, high-quality state generation (i.e., negligible error between expected and actual encoded values).

[0108] The encoded light is transmitted along a communication channel 5 where it may experience optical phenomena that change some of the pulse properties, for example, dispersion may broaden the pulse, polarization fluctuations may change the polarization state, and channel timing delays may change the expected pulse arrival time.

[0109] At receiver 3 (“Bob”), the quantum state is measured, which involves randomly choosing a basis for the measurement using a demodulator (not shown) (e.g., by applying some unitary operation to the quantum state), and then detecting the signal using a single-photon detector (not shown) to form quantum state measurement 11. Various detection techniques exist, but they typically rely on the ability to accurately determine the coded bit value while rejecting noise sources such as detector noise or other noise introduced by the channel.

[0110] Following the generation, transmission, and measurement of the quantum states, a post-processing step is performed involving authenticated classical communication between Alice and Bob, where they reveal a subset of their random selections to perform sifting, information alignment, error correction, and privacy amplification. This yields information distributed on both remote nodes theoretically secure quantum keys, which can then be distributed to other devices, for example, for data encryption.

[0111] Next, a basic quantum communication protocol using polarized light will be described. However, it should be noted that this is not meant to be limiting and other protocols may also be used. For simplicity, this protocol refers to polarized light, but phase-based protocols could also be used.

[0112] The protocol uses two bases, where each base is described by two orthogonal states: in this example, an H / V base and a D / A base, although an L / R base could also be chosen.

[0113] The transmitter in the protocol prepares a state with one of the H, V, D, or A polarizations. In other words, the prepared state is selected from two orthogonal states (H and V or D and A) in one of the two bases H / V and D / A. This can be thought of as sending a signal of 0 and 1 in one of the two bases, for example, H=0, V=1 in the H / V base and D=0, A=1 in the D / A base. The pulses are attenuated so that they comprise less than one photon on average. Therefore, if a measurement is made on the pulse, it will be destroyed. Also, it is not possible to split the pulse.

[0114] The receiver uses a measurement base for the polarization of the pulse selected from H / V base or D / A base. The selection of the measurement base can be active or passive. In passive selection, the base is selected using a fixed component such as a beam splitter. In "active" base selection, the receiver determines which base to measure using, for example, a modulator with an electrical control signal. If the base used to measure the pulse in the receiver is the same as the base used to encode the pulse, the receiver's measurement of the pulse will be accurate. However, if the receiver selects the other base to measure the pulse, there will be a 50% error in the result measured by the receiver.

[0115] To establish the key, the sender and receiver compare the bases used for encoding and measuring (decoding). If they match, the result is kept; if they do not match, the result is discarded. The above method is very secure. If an eavesdropper intercepts and then measures a pulse, the eavesdropper must prepare another pulse to send to the receiver. However, the eavesdropper will not know the correct measurement base and therefore will only have a 50% chance of measuring the pulse correctly. Any pulse reproduced by an eavesdropper will cause a larger error rate at the receiver, which can be used to prove the presence of an eavesdropper. The sender and receiver compare a small portion of the key to determine the error rate (QBER) and therefore the presence of an eavesdropper.

[0116] However, the above description of a QKD system only describes the basic operation. In reality, in addition to the core quantum state generation / encoding / measurement devices, there are also components included to compensate for channel disturbances or non-idealities in the real-world components (e.g., polarization adjustment, power adjustment, dispersion compensation, timing compensation, etc.). Precise electronic control of all the optoelectronic hardware is also required to ensure good synchronization between Alice and Bob (and all their internal hardware).

[0117] FIG. 14 illustrates a QKD system that more explicitly shows these additional components / stages that are important for ensuring high performance operation in the presence of real-world disturbances. Note that the exact implementation may vary, with compensation performed at Alice and / or Bob, and compensation performed in any order. To avoid unnecessary repetition, the same reference numbers as in FIG. 13 have been used where appropriate. Additionally, in FIG. 14, control electronics 17 and signal compensation unit 13 are provided in the transmitter. Control electronics 19 and signal compensation unit 15 are provided in the receiver.

[0118] Various optical / electronic system designs are possible for implementing QKD. Figure 15 shows one example of such an approach. Here, the transmitter comprises a continuous-wave (CW) laser 21, followed by modulators 23, 25, and 27 for chopping pulses in time and encoding the phase between the pulses. A first intensity modulator 23 is used to chop the CW output into pulses. A phase modulator 25 is then used to encode the pulses. In one embodiment, the phase modulator comprises an asymmetric Mach-Zehnder interferometer (AMZI), with a phase modulation component provided in an arm of the AMZI. A second intensity modulator 27 is then provided to attenuate the pulses output by the phase modulator 25 to pulses with one or fewer photons on average. In a further embodiment, the second intensity modulator is also used to implement decoy states, which are required for many QKD protocols. Decoy states are similar to the other encoded states but have a reduced amplitude (i.e., a smaller average number of photons used for the decoy pulse), which allows them to be used to evade attacks such as photon number splitting attacks.

[0119] To achieve short (e.g., picosecond-scale) well-defined pulses, intensity modulators 23 and 27 with high dynamic range are used. Regarding phase, in one embodiment, QKD requires that each quantum state be phase-randomized. Phase encoding is used, whereby if a quantum state is encoded by the phase between two time bins, the phase between the two bins is precisely controlled.

[0120] This places stringent requirements on coherent and highly stable light emission from all modulators 23, 25, and 27 and laser 21. The quality of the output can be improved by precise control of the electronic drive signals applied to modulators 23, 25, and 27 and laser 21. For example, the quality of the output can be improved by adjusting the bias / control of laser 21 and the control of the modulator. The modulator has the added complexity of being driven by an RF signal with the appropriate amplitude and timing to generate a properly encoded train of quantum states. The laser is also driven by an RF signal in some embodiments of the QKD transmitter. In the case of a "phase-seeded" configuration (i.e., a gain-switched primary-secondary laser configuration), both the primary and secondary lasers have an applied RF signal that generates optical pulses. Additional electrical control signals with many customizable parameters may also need to be adjusted, such as control signals for other signal compensation devices (not shown), such as polarization controllers, attenuators, etc.

[0121] Figure 16A shows an alternative transmitter design using a "phase seed." In a phase-seed QKD transmitter, multiple modulators are replaced by a primary-secondary configuration (whereby the primary laser injection locks the secondary, and the current in the primary is modulated to impose phase modulation on the secondary). This has the benefit of eliminating the modulators, which are generally bulky, costly, and prevent chip-scale integration of the design on a photonic chip (as lithium niobate-based modulators are not compatible with photonic integration). Figure 16B shows a drive scheme for such a configuration using primary and secondary lasers.

[0122] In the above embodiment of Figures 16A and 16B, a configuration is shown in which one primary pulse seeds two secondary pulses. Figures 17A to 17F explain some of the physics behind laser seeding. Here, for simplicity, we will describe the situation in which one primary pulse seeds one secondary pulse.

[0123] Figures 17A through 17F are used to describe a particular type of source that allows for the creation of an ultra-compact, high-performance QKD transmitter. This source, shown in Figure 17A, uses gain switching and optical injection locking, and has directly phase-modulated pulses without the need for an external phase modulator.

[0124] The source comprises a pulsed secondary laser 103 into which pulses from the primary laser 103 are injected to define the phase between the output pulses of the secondary laser based on optical injection locking. The primary laser 101 performs the task of phase preparation, while the secondary laser 103 performs the task of pulse generation. The description of Figures 17A to 17D focuses more on the control of the primary laser. The description of Figures 17E and 17F focuses more on the secondary laser and the combination of the two lasers.

[0125] As shown schematically in FIG. 17A, the primary laser diode 101 is connected to the secondary laser diode 103 via an optical circulator 105. Note that a circulator is not required. Further embodiments use different methods to inject light from one laser into another depending on the laser type / packaging. For example, FIG. 16A shows direct injection into one side of the laser with the emission power taken from the other. Such a configuration is possible, for example, when the laser has partially reflective facets on both sides of the cavity. Note that the primary laser diode 101 and the secondary laser diode 103 may be the same, and the terms "primary" and "secondary" are used merely for clarity and do not imply any physical difference between the primary 101 and secondary 103 laser diodes.

[0126] The primary laser 101 is used for phase preparation, which is directly modulated to generate long pulses from quasi-steady-state emission. Each of these pulses coherently seeds a block of two or more secondary short optical pulses emitted by gain-switching the secondary or pulse-generating laser 103. The phase-ready laser 101 is biased to generate quasi-steady-state optical pulses with shallow intensity modulation on the nanosecond scale, which also modulates the optical phase. For clock rates greater than 1 GHz, the pulse width is shorter than 1 ns. The gain-switched pulse-generating laser 103 emits short optical pulses that inherit the optical phase prepared by the phase-ready laser. The duration of each phase-ready laser pulse can be varied to seed pulse trains of different lengths.

[0127] The relative phase between the secondary pulses depends on the phase evolution of the primary pulse and can be set to any value by directly modulating the drive current applied to the primary or phase-ready laser 101 .

[0128] For example, the relative phase Φ between two secondary pulses can be obtained by introducing a small perturbation in the drive signal of the phase-ready laser in Figure 17 A. Similarly, the relative phases between three secondary pulses can be set to Φ and Φ by adding two small perturbations to the drive signal of the primary laser 101.

[0129] In principle, such perturbations in the drive signal could cause harmful variations in the intensity and frequency of the primary pulses. However, these can be avoided by switching off the gain of the secondary laser 103 in response to the perturbation signal. Effectively, the secondary laser 103 also acts as a filter to remove residual modulation.

[0130] To understand how the optical phase is set by perturbing the drive signal applied to a phase-ready laser, it is useful to consider an above-threshold continuous-wave laser emitting at a central frequency υ.

[0131] FIG. 17B shows the duration t m 17A and 17B are plots of the optical frequency of the phase-ready laser under small perturbations of .mu.m and .mu.m, respectively. Figure 17C is a plot of the optical phase trajectory with and without perturbation of the phase-ready laser.

[0132] When a small perturbation is applied to the drive signal, the optical frequency shifts by an amount Δν, changing the course of the phase evolution. When the perturbation is switched off, the frequency is restored to its initial value ν. This perturbation creates a phase difference.

[0133] Δφ=2πΔυt m where t m is the duration of the perturbation. Through optical injection, this phase difference is transferred onto the pair of secondary pulses emitted by the pulse-generating laser as shown in Figure 17D.

[0134] The perturbation signal here is an electrical voltage modulation applied to the phase-prepared laser. The change in optical frequency arises from the effect of carrier density on the refractive index in the laser-active medium within the primary laser diode 101. Laser cavity confinement allows the optical field to oscillate back and forth within the cavity, experiencing a change in refractive index for the entire duration of the perturbation. The laser cavity enhancement allows the phase modulation half-wave voltage to remain below 1 V. This cavity feature is absent in conventional phase modulators, where light makes only a single pass across the electro-optic medium, thus limiting the interaction distance to the length of the device.

[0135] Small changes in the primary light source's electrical controller signal (less than 1 volt, much less than required by conventional lithium niobate phase modulators) result in transient changes in the output frequency of the primary light source's output, which can then change the output phase of the secondary laser's optical output.

[0136] In this embodiment, the primary laser 101 is configured to output a series of optical pulses comprising a series of pairs. The phases of the pulses output by the primary laser are controlled such that the phase between pulses in the same pair is randomly selected from one of a set of phase differences, and there is a random phase difference between pulses from different pairs. In one embodiment, the set of phase differences may be selected from one of 0, π / 2, -π / 2, and π.

[0137] The secondary laser 103 , seeded by the primary laser, will output a series of pulse pairs having the same phase difference as the series of pulses output by the primary laser 101 .

[0138] Pulse injection seeding occurs whenever the secondary laser 103 is switched above the lasing threshold. In this case, the generated secondary light pulse has a fixed phase relationship with the injected primary light pulse. Because only one secondary light pulse is generated for each injected primary light pulse, the phase relationship between the pulses output by the secondary laser will be the same as the relationship between the pulses injected into the secondary laser.

[0139] Under the operating conditions described below with respect to Figures 17E and 17F, the secondary laser 103 generates a new series of pulses comprising a series of pairs. The phases between pulses in the same pair are randomly selected from one of a set of phase differences, and there are random phase differences between pulses from different pairs. These pulses will also have a smaller time jitter τ'<τ relative to the pulses output by the primary laser 101. The reduced jitter time improves interference visibility due to the low time jitter of the secondary light pulses.

[0140] For pulsed injection seeding to occur, the frequency of the light pulses from the primary laser 101 must match to a certain extent the frequency of the secondary laser 103. In one embodiment, the difference between the frequency of the light provided by the primary laser 101 and the frequency of the secondary laser 103 is less than 30 GHz. In some embodiments, when the secondary laser 103 is a distributed feedback (DFB) laser diode, the frequency difference is less than 100 GHz.

[0141] For successful pulse injection seeding, the relative power of the output optical pulse of primary laser 101 entering the optical cavity of secondary laser 103 must be within several limits that depend on the type of light source used. In one embodiment, the optical power of the injected optical pulse will be at least 1000 times lower than the optical output of secondary laser 103. In one embodiment, the optical power of the injected optical pulse will be at least 100 times lower than the optical output of secondary laser 103.

[0142] In one embodiment, the secondary laser 103 and the primary laser 101 are electrically driven gain-switched semiconductor laser diodes. In one embodiment, the secondary light source and the primary light source have the same bandwidth. In one embodiment, both light sources have a bandwidth of 10 GHz. In one embodiment, both light sources have a bandwidth of 2.5 GHz. Here, bandwidth refers to the highest bit rate achievable using gain-switched laser diodes under direct modulation. Lasers with a certain bandwidth can be operated at a lower clock rate.

[0143] FIG. 17E is a schematic diagram of a drive scheme for a phase-randomized light source 500 in which both the primary laser 503 and the secondary laser 502 are driven using a single gain modulation unit 509. The gain modulation unit 509 and delay line 510 are an example of a controller configured to apply a time-varying drive signal to the secondary laser 502 so that only one optical pulse is generated during each time period in which an optical pulse is received. The primary laser 503 is connected to the secondary laser 502 via an optical connection 505. The optical connection 505 can be a waveguide, e.g., an optical fiber. Alternatively, the optical pulses can travel between the primary laser 503 and the secondary laser 502 through free space. The optical connection can include additional components, such as an optical circulator or a beam splitter, as shown in the configuration of FIG. 17A.

[0144] The gain modulation unit 509 drives both the primary laser 503 and the secondary laser 502 to generate pulses of light. A delay line 510 is used to synchronize the devices. The delay line can be, for example, a fixed-length cable. The gain modulation unit is directly connected to the primary laser 503. For example, if the primary laser 503 is a semiconductor laser, the gain modulation circuit is electrically connected to the primary laser 503. The gain modulation unit 509 is connected to the secondary laser 502 through the delay line 510.

[0145] FIG. 17F shows the time sequence for the single gain modulation scheme shown in FIG. 17E. The upper graph shows the gain modulation applied to the primary light source 503. The current applied to the laser is shown on the vertical axis, and time on the horizontal axis. The gain modulation is a time-varying drive signal with the form of a square wave that, when applied to the primary light source, varies the carrier density above and below the lasing threshold. In other words, the gain modulation is a series of pulses. In the middle of the pulse, the gain has a minimum value, which is the gain bias, shown by the dotted line. The wave in this case is a square-type waveform. A different gain modulation signal can be used, for example, a sinusoidal or non-periodic time-varying signal. In this case, the current is not reduced to zero in the middle of the current modulation pulse, but only to the bias value (shown by the dotted line).

[0146] A current modulation signal is applied to the laser, periodically switching the laser's gain above and below the lasing threshold. The second graph shows the carrier density of the laser on the vertical axis against time on the horizontal axis. The lasing threshold is indicated by the dashed horizontal line. When a current modulation pulse is applied to the laser, the injected carriers increase the carrier density, resulting in an increase in photon density.

[0147] The laser output generated by the modulation signal is shown in the graph below. The vertical axis shows laser intensity and the horizontal axis shows time. When the carrier density rises above the lasing threshold, the laser outputs light. Photons generated by spontaneous emission in the laser cavity are amplified sufficiently by stimulated emission to produce the output signal. The length of the delay between the application of the current modulation pulse and the generation of the output light depends on several parameters, such as the laser type, cavity length, and pump power.

[0148] The rapid increase in photon density causes a decrease in carrier density. This, in turn, decreases photon density, which increases carrier density. At this point, the current modulation pulse is timed to switch back down to the DC bias level, causing the laser emission to quickly extinguish. The laser output therefore consists of a train of short laser pulses, as shown in the bottom graph.

[0149] To generate longer pulses, the gain bias is chosen to be closer to the lasing threshold. This means that the carrier density crosses the lasing threshold sooner, giving the light pulse more time to evolve. Initially, the light intensity overshoots, rapidly reducing the carrier density. This, in turn, reduces the photon density, increases the carrier density, and then increases the light intensity. This competitive process results in an oscillation of the light intensity at the beginning of the pulse, which quickly reaches a strongly damped steady state where the intensity is constant. The oscillation is called relaxation oscillation. When the current pulse ends, the laser pulse ends and the current is switched back to the bias value.

[0150] The following graph shows the output of the primary laser 503. Each time the carrier density increases above the lasing threshold, one optical pulse is output. As explained above, there may be a delay between when the gain increases and when the optical pulse is output. The optical pulse output from the primary laser has a large time jitter τ.

[0151] The following graph shows the gain modulation applied to the secondary laser 502. The gain modulation is the same as that applied to the primary laser 503, but with a time delay indicated by the arrow. The gain modulation is a time-varying drive signal applied to the secondary laser. In other words, the gain modulation applied to the secondary laser 502 is shifted in time with respect to the gain modulation applied to the primary laser 503. Each periodic increase in gain is applied to the secondary laser 502 later than it is applied to the primary laser 503. The delay, in this case, is approximately half a period of the gain modulation signal. The delay means that the periodic increase in gain is applied to the secondary laser 502 after the optical pulse is injected. Thus, an optical pulse from the primary laser 503 is present in the laser cavity of the secondary laser when the gain increase is applied, and the resulting secondary laser 502 generates an optical pulse by stimulated emission from the primary optical pulse. This means that the light pulses generated from the secondary laser have a fixed phase relationship to the light pulses injected into the secondary laser from the primary laser.

[0152] The secondary laser 502 is switched above the lasing threshold after the optical pulse from the primary laser is injected so that the pulse from the secondary laser is initiated by stimulated emission caused by the injected optical pulse. The timing of the onset of the gain bias of the secondary laser 502 is controlled via a delay line 510. The last graph shows the output of the secondary laser 502. Each time the carrier density increases above the lasing threshold, a single optical pulse is output. Again, there may be a delay between the increase in gain modulation and the output optical pulse. The time jitter of the output optical pulse from the secondary laser is lower than that of the optical pulse from the primary laser.

[0153] 17E, the gain modulation unit 509 applies a time-varying gain modulation to the secondary light source 502 such that it switches above the lasing threshold only once during the time that each light pulse from the primary laser is incident. The switching of the secondary laser 502 is synchronized with the arrival of the light pulse from the primary laser.

[0154] In the system shown in Figure 17F, the time-varying gain modulated signal has a square-type waveform, however, the time-varying gain modulated signal can comprise a signal with any pulse shape.

[0155] When the light source is a gain-switched semiconductor laser, the gain modulation signal is an applied current or voltage. In one embodiment, the gain modulation signal is an applied current or voltage with a square-type waveform. In an alternative embodiment, the time-varying current or voltage is an electrical sine wave generated by a frequency synthesizer. In one embodiment, the frequency of the gain modulation signal is less than or equal to 4 GHz. In one embodiment, the frequency is 2.5 GHz. In one embodiment, the frequency is 2 GHz.

[0156] Gain-switched semiconductor lasers have a good extinction ratio between the "off" state and the state when the pulse is emitted. They can be used to generate ultrashort pulses. In one embodiment, the duration of each pulse output from the secondary laser is less than 200 ps. In one embodiment, the duration of each pulse output from the secondary laser is less than 50 ps. In one embodiment, the duration of each pulse output from the secondary laser is on the order of a few picoseconds. In one embodiment, if the time-varying current or voltage is a square wave current or voltage with a frequency of 2 GHz, the short optical pulses will be spaced 500 ps apart.

[0157] In the light sources shown in these figures, the primary and secondary lasers share the same electrical driver for gain modulation. However, the primary and secondary lasers can also be driven by separate gain modulation units 509. By driving the gain modulation by separate units, it is possible to generate longer optical pulses output from the primary laser than those shown in FIG. 17F because the gain bias value is closer to the lasing threshold. This means that the carrier density crosses the lasing threshold sooner, giving the optical pulses more time to evolve. This can also be used to reduce jitter.

[0158] For the self-tuning of control parameters described herein, in one embodiment, there is autonomous setting of the control signals driving both lasers, i.e., each laser is set with a different optimal DC bias, AC waveform amplitude, waveform shape and signal timing, and in addition each laser requires precise wavelength setting (e.g., by controlling a temperature controller connected to each laser) to achieve good injection locking.

[0159] Other QKD transmitter designs are possible, and the methods described below can be applied to many other QKD system embodiments.

[0160] The above mentioned only the transmitter. However, complex multi-component optoelectronic designs are required for QKD receivers that also perform better if the control parameters for these receivers are adjusted.

[0161] The above can be used in point-to-point QKD or other systems such as MDI QKD and twin-field QKD. The above transmitter can be applied to several different quantum systems.

[0162] While several embodiments have been described, these embodiments are presented by way of example only and are not intended to limit the scope of the present invention. Indeed, the novel apparatus and methods described herein may be embodied in a variety of other forms, and various omissions, substitutions, and modifications of the forms of the devices, methods, and products described herein may be made without departing from the spirit of the present invention. The appended claims and their equivalents are intended to cover such forms or modifications as would fall within the scope and spirit of the present invention.

Claims

1. a plurality of transmitter elements, each of said plurality of transmitter elements comprising a radiation source; a control unit configured to apply a plurality of control signals defined by a plurality of sets of control parameters to at least one of the plurality of transmitter components; Optimization Unit and 1. A quantum transmitter comprising: wherein the optimization unit is configured to set the plurality of control parameters by performing reinforcement learning (RL) using a policy and an agent, wherein the agent moves in a control parameter space corresponding to an environment, the control parameter space being based on a plurality of vectors constituted by the plurality of control parameters, the policy is a guideline indicating an action of the agent to move in the control parameter space depending on a state, and the state s is a value of the plurality of control parameters at a certain time point; The agent: receiving a plurality of observations regarding states in the environment by obtaining scores assigned to each of a plurality of neighboring locations of the agent in the control parameter space, the scores indicative of a quality of the quantum transmitter; a training program to find a policy that maximizes a reward determined by a plurality of scores included in the plurality of observations; navigating a path in a control parameter space from a first estimate of the plurality of control parameters to a target where the plurality of control parameters are set to a target, wherein the target is a location in the control parameter space having an acceptable quality score; It is configured as follows: A quantum transmitter, wherein the agent is trained using a plurality of different quantum transmitters.

2. The agent: receiving a plurality of observations of a current state of the agent; performing a next action for the agent based on the policy and the plurality of observations of the agent in the current state; receiving compensation for said agent; 10. A quantum transmitter as claimed in claim 1, configured to find a path through the control parameter space by:

3. 2. The quantum transmitter of claim 1, wherein the score is selected from at least one of a quantum bit error rate (QBER), a secure bit rate (SBR), phase information of coded pulses, count rates for all pulses received at detectors of a receiver, count rates for pulses with a predetermined coding, shape of the received pulses, or arrival times of the pulses at the detectors.

4. 10. The quantum transmitter of claim 1, wherein the plurality of observations at the plurality of neighboring locations are the scores at the plurality of neighboring locations.

5. 10. The Quantum Transmitter of claim 1, wherein a plurality of additional observations of the operation of the Quantum Transmitter are provided to the agent.

6. 10. The quantum transmitter of claim 1, wherein the radiation source is a pulsed radiation source.

7. 7. The quantum transmitter of claim 6, wherein the pulsed radiation source is provided on a temperature control element, and wherein the control unit further provides a control signal to the temperature control element.

8. 7. The quantum transmitter of claim 6, wherein the plurality of control signals are a plurality of electronic control signals for the pulsed radiation source comprising at least one of an electronic control signal magnitude, a shape of the electronic control signal, and a DC offset of the electronic control signal.

9. 7. The quantum transmitter of claim 6, wherein the pulsed radiation source comprises a primary laser and a secondary laser, the primary laser configured to apply a seed pulse to the secondary laser.

10. 10. The quantum transmitter of claim 9, wherein the set of control parameters comprises control parameters for both the primary laser and the secondary laser.

11. 10. The quantum transmitter of claim 1, wherein the plurality of transmitter components further comprise a modulation unit, the modulation unit configured to randomly encode the plurality of pulses of radiation.

12. A quantum communication system comprising the quantum transmitter of claim 1.

13. 13. The quantum communication system of claim 12, wherein the quantum communication system is selected from a point-to-point quantum communication system, a measurement device independent quantum communication system, and a twin-field quantum communication system.

14. 1. A method of optimizing a quantum transmitter, said quantum transmitter comprising: a plurality of transmitter elements, each of said plurality of transmitter elements comprising a radiation source; a control unit configured to apply a plurality of control signals defined by a plurality of sets of control parameters to at least one of the plurality of transmitter components; Equipped with The method comprises: setting the plurality of control parameters by performing reinforcement learning using a policy and an agent, wherein the agent moves in a control parameter space corresponding to an environment, the control parameter space being based on a plurality of vectors constituted by the plurality of control parameters, the policy being a guideline indicating an action of the agent to move in the control parameter space depending on a state, and the state being values ​​of the plurality of control parameters at a certain point in time; The agent: receiving a plurality of observations regarding states in the environment by obtaining scores assigned to each of a plurality of neighboring locations of the agent in the control parameter space, the scores indicative of a quality of the quantum transmitter; a training program to find a policy that maximizes a reward determined by a plurality of scores included in the plurality of observations; navigating a path in a control parameter space from a first estimate of the plurality of control parameters to a target where the plurality of control parameters are set to a target, wherein the target is a location in the control parameter space having an acceptable quality score; It is configured as follows:

10. The method of claim 1, wherein the agent is trained using a plurality of different quantum transmitters.

15. 1. A method of training a reinforcement learning (RL) agent for use in optimizing a quantum transmitter, the quantum transmitter comprising: a plurality of transmitter elements, each of said plurality of transmitter elements comprising a radiation source; a control unit configured to apply a plurality of control signals defined by a plurality of sets of control parameters to at least one of the plurality of transmitter components; Equipped with The method comprises: training the RL agent to optimize the plurality of control parameters by performing reinforcement learning using a policy, wherein the RL agent moves through a control parameter space corresponding to an environment, the control parameter space being based on a plurality of vectors composed of the plurality of control parameters, the policy being a guideline for the agent's movement through the control parameter space depending on a state, the state being a value of the plurality of control parameters at a certain point in time, receiving a plurality of observations regarding states in the environment by obtaining scores assigned to each of a plurality of neighboring locations of the agent in the control parameter space, the scores indicating a quality of the quantum transmitter, the RL agent being trained to discover a policy that maximizes a reward determined by the plurality of scores included in the plurality of observations, and navigating a path through the control parameter space from a first estimate of the plurality of control parameters until the plurality of control parameters are set to a target; A method of training an RL agent, wherein the target is a location in a control parameter space having a score for which the quality is acceptable, and the RL agent is trained using multiple training transmitters.

16. Training the RL agent to optimize the plurality of control parameters includes: selecting a training transmitter from the plurality of training transmitters; navigating a path through a control parameter space until the plurality of control parameters are set to a target; Equipped with Here, navigating a route involves: receiving a plurality of observations of the RL agent at a current state of the RL agent in the selected transmitter by obtaining the score; performing a next action for the RL agent based on the policy and the plurality of observations of the RL agent in the current state; receiving a reward for the RL agent; updating the policy of the RL agent; Equipped with 16. The method of training an RL agent of claim 15, wherein navigating a route is performed multiple times for each of the plurality of training transmitters.

17. The method of claim 16 , wherein the plurality of observations at the plurality of neighboring locations are the scores at the plurality of neighboring locations.

18. a transmitter comprising a plurality of transmitter elements, the plurality of transmitter elements comprising a pulsed radiation source and a modulation unit, the modulation unit configured to randomly encode a plurality of pulses of radiation; a receiver comprising a plurality of receiver elements, the plurality of receiver elements comprising a demodulator and a detector configured to decode and detect the plurality of randomly encoded pulses; A quantum communication system comprising: the quantum communication system further comprising a control unit and an optimization unit, the control unit configured to apply a plurality of control signals defined by a set of a plurality of control parameters to at least one of the plurality of transmitter components and the plurality of receiver components. wherein the optimization unit is configured to set the plurality of control parameters by performing reinforcement learning using a policy and an agent, wherein the agent moves in a control parameter space corresponding to an environment, the control parameter space being based on a plurality of vectors constituted by the plurality of control parameters, the policy is a guideline indicating an action of the agent to move in the control parameter space depending on a state, and the state is a value of the plurality of control parameters at a certain time point; The agent: receiving a plurality of observations regarding states in the environment by obtaining scores assigned to each of a plurality of neighboring locations of the agent in the control parameter space, the scores indicative of a quality of the quantum communication system; a training program to find a policy that maximizes a reward determined by a plurality of scores included in the plurality of observations; navigating a path in a control parameter space from a first estimate of the plurality of control parameters to a target where the plurality of control parameters are set to a target, wherein the target is a location in the control parameter space having an acceptable quality score; It is configured as follows:

10. A quantum communication system, wherein the agent is trained using a plurality of different quantum communication systems.

Citation Information

Patent Citations

  • Optical transmitter, optical modulation control circuit, and optical modulation control method

    JP2013255263A

  • Light source, method for generating optical pulse, quantum communication system, and quantum communication method

    JP2022019522A

  • Quantum communication system, transmitter for quantum communication system, receiver for quantum communication system, and method for controlling quantum communication system

    JP2023130309A