Atomic oscillator, control method, control device, and program

The atomic oscillator uses reinforcement learning to stabilize oscillation frequency by dynamically controlling environmental factors, addressing instability issues caused by temperature shifts and fluctuations.

JP2025173990APending Publication Date: 2025-11-28NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024079927
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-16
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing atomic oscillators face instability in oscillation frequency due to temperature shifts and environmental fluctuations, limiting the stability of the resonance frequency.

Method used

An atomic oscillator employing reinforcement learning to control environmental states using a reward-based system, adjusting components like lasers, gas cells, and oscillators to stabilize the oscillation frequency.

Benefits of technology

Enhances the stability of the oscillation frequency by optimizing environmental control actions through reinforcement learning, ensuring consistent performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025173990000001_ABST
    Figure 2025173990000001_ABST
Patent Text Reader

Abstract

To solve the problem in impossibility of further improving the stability of an oscillation frequency.SOLUTION: An atomic oscillator includes: a gas cell in which an alkali metal atom is sealed; a light generator that irradiates the gas cell with irradiation light having at least two different frequency components; a light detector that detects transmitted light transmitted through the gas cell; and a controller that determines a resonance frequency based on a light amount of the transmitted light of the gas cell, and controls an oscillation frequency by the oscillator based on the determined resonance frequency. The atomic oscillator includes an agent that performs reinforcement learning to output an action for controlling a state of an environment according to the acquired state of the environment of the atomic oscillator using a reward corresponding to a difference between a preset reference frequency and an oscillation frequency.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an atomic oscillator, a control method, a control device, and a program. [Background technology]

[0002] An atomic oscillator is a device that measures time accurately based on the natural frequency of an atom. Compact atomic clocks typically measure the natural frequency of an atom by using coherent population trapping (CPT), a quantum interference effect that occurs when an alkali metal atomic gas is irradiated with excitation light of two frequencies. In CPT, when the difference in frequency between the two excitation lights matches the transition frequency between the ground levels of the alkali metal, the amount of transmitted light increases without absorption of the excitation light. Atomic oscillators based on CPT sweep the difference in frequency between the two excitation lights, and determine the resonant frequency, the difference between the frequencies at which the amount of transmitted light is maximized, as the natural frequency of the atom. Whether or not this resonant frequency, the natural frequency of the atom, can be stably obtained is one of the performance indicators of an atomic oscillator.

[0003] One of the causes of performance degradation in the above-mentioned atomic oscillator is temperature shift, which causes the resonance frequency to fluctuate due to temperature changes in the alkali metal atomic gas. In other words, when a temperature change occurs inside the oscillator, the optical transition characteristics of the atoms fluctuate, reducing the stability of the oscillation frequency. To address this problem, Patent Document 1 describes providing a temperature measuring element and a heater outside the alkali metal cell. As a result, Patent Document 1 describes how the temperature of the alkali metal cell can be kept constant by using temperature information from the alkali metal cell to heat the cell with the heater. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-011680 Summary of the Invention [Problem to be solved by the invention]

[0005] However, if the temperature of the alkali metal cell cannot be kept constant, the resonance frequency changes due to temperature shift, resulting in a problem of reduced stability of the oscillation frequency. Furthermore, if the environmental conditions of the atomic oscillator, not limited to the temperature of the alkali metal cell, cannot be kept constant, the stability of the oscillation frequency also decreases. As a result, there is a problem that the stability of the oscillation frequency of the atomic oscillator cannot be further improved.

[0006] Therefore, an object of the present disclosure is to provide an atomic oscillator that can solve the above-mentioned problem of not being able to further improve the stability of the oscillation frequency. [Means for solving the problem]

[0007] An atomic oscillator according to one embodiment of the present disclosure includes: a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a controller that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; An atomic oscillator comprising: an agent that performs reinforcement learning to output an action that controls a state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; The structure is as follows. Moreover, an atomic oscillator according to an embodiment of the present disclosure includes: a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a controller that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; An atomic oscillator comprising: an agent that has undergone reinforcement learning to output an action that controls a state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; the agent outputs the behavior in accordance with the acquired state of the environment of the atomic oscillator; the controller controls a state of the environment of the atomic oscillator based on the action output by the agent. The structure is as follows. Furthermore, a control method according to an embodiment of the present disclosure includes: a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; A control method by the control device in an atomic oscillator comprising: an agent included in the control device performs reinforcement learning so as to output an action that controls the state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; The structure is as follows. Furthermore, a control method according to an embodiment of the present disclosure includes: a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; A control method by the control device in an atomic oscillator comprising: an agent provided in the control device that has performed reinforcement learning to output an action to control the state of the environment of the atomic oscillator according to the acquired state of the environment using a reward according to the difference between a preset reference frequency and the oscillation frequency outputs the action according to the acquired state of the environment of the atomic oscillator, and the control device controls the state of the environment of the atomic oscillator based on the action output from the agent; The structure is as follows. Furthermore, a control device according to an embodiment of the present disclosure includes: a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; The control device in an atomic oscillator comprising: an agent that performs reinforcement learning to output an action that controls a state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; The structure is as follows. Furthermore, a control device according to an embodiment of the present disclosure includes: a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; The control device in an atomic oscillator comprising: an agent that has undergone reinforcement learning to output an action that controls the state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency, and the agent outputs the action according to the acquired state of the environment of the atomic oscillator, and controls the state of the environment of the atomic oscillator based on the output action; The structure is as follows. Furthermore, a program according to an embodiment of the present disclosure includes: a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; The control device in the atomic oscillator includes: an agent included in the control device performs reinforcement learning so as to output an action that controls the state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; Execute the process, The structure is as follows. Furthermore, a program according to an embodiment of the present disclosure includes: a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; The control device in the atomic oscillator includes: an agent provided in the control device that has performed reinforcement learning to output an action to control the state of the environment of the atomic oscillator according to the acquired state of the environment using a reward according to the difference between a preset reference frequency and the oscillation frequency, outputs the action according to the acquired state of the environment of the atomic oscillator, and controls the state of the environment of the atomic oscillator based on the output action; Execute the process, The structure is as follows. [Effects of the Invention]

[0008] With the above-described configuration, the present disclosure can provide an atomic oscillator that can further improve the stability of the oscillation frequency. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram illustrating an example of the configuration of an atomic oscillator according to the present disclosure. [Figure 2] FIG. 1 is a diagram illustrating an overview of the processing of an atomic oscillator according to the present disclosure. [Figure 3] 10 is a flowchart illustrating an example of a processing operation of an atomic oscillator according to the present disclosure. [Figure 4] FIG. 1 is a block diagram illustrating an example of the configuration of an atomic oscillator according to the present disclosure. [Figure 5] 10 is a flowchart illustrating an example of a processing operation of an atomic oscillator according to the present disclosure. [Figure 6] FIG. 1 is a block diagram illustrating an example of the configuration of an atomic oscillator according to the present disclosure. [Figure 7] 10 is a flowchart illustrating an example of a processing operation of an atomic oscillator according to the present disclosure. [Figure 8] 10 is a flowchart illustrating an example of a processing operation of an atomic oscillator according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0010] <Embodiment 1> A first embodiment of the present disclosure will be described with reference to the drawings, which may be relevant to any embodiment.

[0011] [composition] 1, the atomic oscillator in the present disclosure includes a light generator 1 equipped with a laser 11, a gas cell 21, a photodetector 31, a processor 4, and an oscillator 51. The processor 4 is configured as an information processing device (control device (controller)) equipped with an arithmetic unit and a storage device, and each functional unit of the processor 4, which will be described later, is realized by executing a program in the arithmetic unit.

[0012] The light generator 1 generates excitation light (irradiation light), which is light having at least two different frequencies, and irradiates the gas cell 21 with the excitation light. For example, the light generator 1 generates excitation light with a single wavelength, for example, 894.5812 nm, based on a setting value specified by the processor 4 using the laser 11, and generates excitation light having two different frequency components by frequency-modulating the excitation light with the single wavelength. The light generator 1 then irradiates the generated excitation light (irradiation light) onto the gas cell 21, and the transmitted light, which is light that has passed through the gas cell 21, reaches and is detected by the photodetector 31, converted into an electrical signal, etc., and sent to the processor 4.

[0013] In this case, the irradiated light irradiated from the light generator 1 to the gas cell 21 has at least two different frequency components. Note that the irradiated light irradiated from the light generator 1 may have three or more different frequency components, but the difference frequency between two of these frequency components is approximately equal to the transition frequency between specific quantum states that form the CPT resonance of the alkali metal atom.

[0014] Here, the light generator 1 further includes a laser environment control unit 10. The laser environment control unit 10 includes a laser environment sensor 12 that measures laser environment values ​​(measured values) that represent the state of the laser 11, and also has the function of controlling the state of the laser 11. For example, the state of the laser 11 may include the driving current of the laser 11 and the temperature of the laser 11, but any type of state related to the laser 11 may be used. As described below, the laser environment control unit 10 notifies the processor 4 of the laser environment value measured by the laser environment sensor 12 and controls the state of the laser 11 in accordance with a laser environment control value, which is a control command from the processor 4. As an example, the laser environment control unit 10 includes a temperature control device for the laser 11, such as a resistance heater, and uses the temperature control device to control the state of the laser 11, i.e., the temperature, in accordance with the laser environment control value. The state of the laser 11 is one of the environmental states of the atomic oscillator.

[0015] The gas cell 21 is configured by sealing alkali metal atoms in a container. The alkali metal atoms sealed in the gas cell 21 may be, for example, cesium atoms, rubidium atoms, sodium atoms, or potassium atoms. The material constituting the container of the gas cell 21 is preferably a transparent material such as glass that has a high transmittance for the irradiated light generated from the light generator 1. In addition to the alkali metal atoms, the gas cell 21 may also be filled with a buffer gas that does not contribute to absorbing the irradiated light, in order to reduce the effect of collisions between the container wall surface and the gaseous alkali metal atoms.

[0016] The gas cell 21 is also equipped with a gas cell temperature adjustment device that adjusts its own temperature. The gas cell temperature adjustment device may be configured, for example, as a resistance heater, and is configured as a device that has a heating or heating / cooling function and can adjust the temperature of the gas cell 21. The gas cell 21 is also equipped with a magnetic field application device (not shown). The magnetic field application device generates a magnetic field at a predetermined position inside the gas cell 21 in a direction parallel or anti-parallel to the irradiated light. The magnetic field application device is configured, for example, as a coil arranged to cover the gas cell 21, and by adjusting the direction and magnitude of the current applied to the coil, the direction and strength of the static magnetic field applied to the predetermined position inside the gas cell 21 can be controlled.

[0017] Here, the atomic oscillator further includes a gas cell environment control unit 20. The gas cell environment control unit 20 includes a gas cell environment sensor 22 that measures a gas cell environment value (measured value) that indicates the state of the gas cell 21, and also has a function of controlling the state of the gas cell 21. For example, the state of the gas cell 21 may be a magnetic field applied to the gas cell 21 or the temperature of the gas cell 21, but any type of state related to the gas cell 21 may be used. As will be described later, the gas cell environment control unit 20 notifies the processor 4 of the gas cell environment value measured by the gas cell environment sensor 22, and also controls the state of the gas cell 21 in accordance with a gas cell environment control value, which is a control command from the processor 4. The state of the gas cell 21 is one of the environmental states of the atomic oscillator.

[0018] The photodetector 31 has a device for detecting transmitted light, which is light that has passed through the gas cell 21. The photodetector 31 is realized, for example, by using a photodiode, but may be realized by any light detection means. Information on the light detected by the photodetector 31 is converted into an electric signal or the like and input to the processor 4.

[0019] Here, the atomic oscillator further includes a photodetector environment control unit 30. The photodetector environment control unit 30 includes a photodetector environment sensor 32 that measures a photodetector environment value (measured value) that indicates the state of the photodetector 31, and also has a function of controlling the state of the photodetector 31. For example, the state of the photodetector 31 may be the temperature of the photodetector 31, but any type of state related to the photodetector 31 may be used. As will be described later, the photodetector environment control unit 30 notifies the processor 4 of the photodetector environment value measured by the photodetector environment sensor 32, and also controls the state of the photodetector 31 in accordance with a photodetector environment control value, which is a control command from the processor 4. The state of the photodetector 31 is one of the environmental states of the atomic oscillator.

[0020] The processor 4 determines the resonance frequency from the amount of transmitted light input from the photodetector 31 and controls the oscillation frequency of the oscillator 51 based on the determined resonance frequency. Specifically, the processor 4 sweeps the difference frequency of the irradiated light and determines the resonance frequency from the transmitted light spectrum. Once the resonance frequency is determined, the processor 4 adjusts the control voltage of the oscillator 51 so that the error signal of the locked-in detected transmitted light spectrum is at a predetermined signal level. Here, the oscillator 51 is composed of a VCXO (voltage-controlled crystal oscillator) that oscillates at approximately 10 MHz. The oscillator 51 generates an oscillation signal according to the control voltage output from the processor 4 and applied to it, and outputs this to the external device 8 as the oscillation frequency, which is the external output of the atomic oscillator. As a result, the oscillation frequency is stabilized at 10 MHz unless the resonance frequency changes. The difference frequency of the irradiated light is generated by converting the VCXO oscillation signal into a signal of several GHz using a multiplier and input to the light generator 1.

[0021] Here, the atomic oscillator further includes an oscillator environment control unit 50. The oscillator environment control unit 50 includes an oscillator environment sensor 52 that measures an oscillator environment value (measured value) that indicates the state of the oscillator 51, and also has a function of controlling the state of the oscillator 51. For example, the state of the oscillator 51 may include the control voltage of the oscillator 51 and the temperature of the oscillator 51, but any type of state related to the oscillator 51 may be used. As will be described later, the oscillator environment control unit 50 notifies the processor 4 of the oscillator environment value measured by the oscillator environment sensor 52, and also controls the state of the oscillator 51 in accordance with an oscillator environment control value that is a control command from the processor 4. The state of the oscillator 51 is one of the environmental states of the atomic oscillator.

[0022] The atomic oscillator further includes an external environment sensor 61. The external environment sensor 61 measures external environment values ​​(measured values) that represent the state of the atomic oscillator or the state around the atomic oscillator. The state of the atomic oscillator may include the acceleration of the atomic oscillator itself, the magnetic field and temperature around (external to) the atomic oscillator, but any type of state related to the atomic oscillator may be used. As will be described later, the external environment sensor 61 notifies the processor 4 of the measured external environment value. The state of the atomic oscillator is one of the environmental states of the atomic oscillator.

[0023] The atomic oscillator also includes an agent 41 that performs reinforcement learning to output control commands that control the state of the atomic oscillator itself and each of the components, i.e., actions that are each environmental control value, in accordance with each measured environmental value that represents the state of the atomic oscillator itself and each of the components. The agent 41 is constructed by a computing device executing a program, and is provided, for example, in a control device equipped with the above-mentioned processor 4. The agent 41 then performs machine learning in cooperation with the processor 4, specifically, performs reinforcement learning as described below. Note that FIG. 2 shows an overview of the process when the agent 41 performs reinforcement learning.

[0024] First, the processor 4 acquires various environmental values ​​that represent the state of the atomic oscillator measured by the various sensors described above. As an example, the processor 4 acquires each measured environmental value (measurement value), such as the drive current of the laser 11, which is a laser environmental value, from the laser environment sensor 12, the temperature of the gas cell 21, which is a gas cell environmental value, from the gas cell environment sensor 22, the temperature of the photodetector 32, which is a photodetector environmental value, from the photodetector environment sensor 32, the control voltage of the oscillator 51, which is an oscillator environmental value, from the oscillator environment sensor 52, and the external temperature and external magnetic field, which are external environmental values, from the external environment sensor 61. Then, the processor 4 passes each acquired environmental value to the agent 41 as the state S of the atomic oscillator.

[0025] The agent 41, which has acquired the state S of the atomic oscillator, outputs each environmental control value, which is an action A corresponding to each environmental value, which is the state S, according to a policy π, which is a function that can be optimized by reinforcement learning. At this time, the policy π of the agent 41 outputs, for example, an action A, which is each environmental control value that changes each measured environmental value. As an example, the agent 41 may output environmental control values ​​consisting of change rates, such as +1% for the drive voltage of the laser 11, +2% for the temperature control voltage of the gas cell 21, and -1% for the control voltage of the oscillator 51, or may output specific voltage values ​​corresponding to the measured environmental values.

[0026] The processor 4, which receives each environmental control value output from the agent 41, controls the configuration of the atomic oscillator so that it is in the state of each environmental control value. That is, the processor 4 controls the state of the voltage values ​​applied to the laser 11, gas cell 21, photodetector 31, oscillator 51, etc., for each environmental control unit 10, etc., according to each environmental control value. The processor 4 then acquires the oscillation frequency, which is the output signal from the atomic oscillator in a state controlled according to each environmental control value. At this time, the processor 4 acquires the reference frequency output from the frequency standard 71 via the frequency reference receiver 72. The reference frequency is a target value for the oscillation frequency of the atomic oscillator, such as 10 MHz. The frequency reference receiver 72 may receive the reference frequency via wireless communication or using a global positioning system (GPS), and may receive the reference frequency by any method and pass it to the processor 4.

[0027] Then, the processor 4 calculates the difference between the oscillation frequency and the reference frequency, calculates a reward R based on this difference, and passes it to the agent 41 to be used in reinforcement learning. At this time, the processor 4 sets the reward R so that the smaller the absolute value (|Δf|) of the difference between the oscillation frequency and the reference frequency, the larger the value becomes. Furthermore, the processor 4 also sets the reward R according to the passage of time t of the reinforcement learning being performed by the agent 41, as will be described later. For example, the processor 4 sets the reward R so that the value becomes smaller as the time t of the reinforcement learning passes. As an example, the processor 4 calculates the reward R using the following formula 1, with γ (<1) being the discount rate.

number

[0028] The agent 41 uses the reward R received from the processor 4 to perform reinforcement learning on a policy π for outputting an action A from a state S, which is the acquired environmental value. For example, the agent 41 performs Q-learning to update an action value function Q shown in the following formula 2, and updates the policy π. At this time, the agent 41 performs reinforcement learning by giving multiple actions A for one state S. Here, α is a learning rate, which is greater than 0 and less than 1.

number

[0029] In addition, the agent 41 may update the policy π by performing DQN (Deep Q-Learning) in the above-mentioned reinforcement learning. Also, the agent 41 may update the policy π by using other machine learning techniques such as a neural network, using the above-mentioned rewards.

[0030] By performing reinforcement learning as described above, the agent 41 is configured to output an environmental control value that is an optimal action A according to the state S of the atomic oscillator. For example, the manufacturer of the atomic oscillator may perform the above-described reinforcement learning on the agent 41 before shipping. Alternatively, even after shipping by the manufacturer of the atomic oscillator, the agent 41 may perform reinforcement learning as described above at a preset timing or at an arbitrary timing, and be updated so as to always output the optimal action A. In this case, the atomic oscillator is configured to acquire a reference frequency, for example, by wireless communication or GPS.

[0031] [Operation] Next, the processing operation of the agent 41 during reinforcement learning using the atomic oscillator described above will be described. First, the processor 4 sets various parameters at the start of reinforcement learning for the agent 41. For example, the processor 4 sets the time T for one episode, which is the period from the start to the end of behavior in an environment given to reinforcement learning, as T=200, the number of episodes N=200, the current time t=0, and the number of current episodes n=0 (step S1 in FIG. 3).

[0032] Next, the processor 4 acquires various measured environmental values, which are the state S of the atomic oscillator, and passes them to the agent 41 (step S2 in FIG. 3). For example, the processor 4 acquires each measured environmental value (measured value), such as the drive current of the laser 11, which is a laser environmental value, the temperature of the gas cell 21, which is a gas cell environmental value, the temperature of the photodetector 32, which is a photodetector environmental value, the control voltage of the oscillator 51, which is an oscillator environmental value, and the external temperature and external magnetic field, which are external environmental values, and passes them to the agent 41.

[0033] Next, the agent 41 outputs each environmental control value, which is an action A corresponding to each environmental value, which is the received state S, according to the policy π, which is a function that can be optimized by reinforcement learning, and passes it to the processor 4 (step S3 in FIG. 3). For example, as an example of action A, the agent 41 outputs environmental control values ​​consisting of change rates such as +1% for the drive voltage of the laser 11, +2% for the temperature control voltage of the gas cell 21, and -1% for the control voltage of the oscillator 51.

[0034] Next, the processor 4 controls the configuration of the atomic oscillator so that it is in the state of each environmental control value corresponding to the outputted action A (step S4 in FIG. 3). That is, the processor 4 controls the state of the voltage values ​​applied to the laser 11, gas cell 21, photodetector 31, oscillator 51, etc. for each environmental control unit 10, etc., in accordance with each environmental control value corresponding to the outputted action A.

[0035] Next, the processor 4 acquires the oscillation frequency, which is the output signal from the atomic oscillator in a state controlled according to each environmental control value, calculates the reward S according to the difference Δf between the oscillation frequency and the reference frequency and the elapsed time t of learning, and passes it to the agent 41. At this time, the processor 4 calculates the reward S, for example, using the above-mentioned formula 1, so that the reward R becomes larger as the absolute value (|Δf|) of the difference between the oscillation frequency and the reference frequency becomes smaller, and becomes smaller as the time t of reinforcement learning passes.

[0036] Next, the agent 41 uses the reward R received from the processor 4 to perform reinforcement learning and update the policy π for outputting action A from the state S, which is the acquired environmental value (step S6 in FIG. 3). Then, the processor 4 and the agent 41 perform reinforcement learning by repeating the above-mentioned process until the various parameters satisfy the set conditions as described above (steps S7 to S10). As a result, the policy π of the agent 41 is configured to output an environmental control value, which is the optimal action A, depending on the state S of the atomic oscillator.

[0037] Next, the processing operation when using an atomic oscillator in which the agent 41 of the atomic oscillator has undergone reinforcement learning as described above will be described. In this case, since the agent 41 of the atomic oscillator has undergone reinforcement learning, the atomic oscillator does not need to be equipped with the configuration required for the reinforcement learning described above, as shown in FIG. 4. For example, an atomic oscillator shipped from a manufacturer may have the configuration shown in FIG. 4. However, even if the agent 41 has undergone reinforcement learning, the atomic oscillator may have the configuration shown in FIG. 1, and reinforcement learning of the agent 41 may be performed after the atomic oscillator is shipped from the manufacturer, and the agent 41 may be updated.

[0038] First, when the use of the atomic oscillator is started (step S21 in FIG. 5), the processor 4 acquires various measured environmental values, which are the state S of the atomic oscillator, and passes them to the agent 41 (step S22 in FIG. 5). For example, the processor 4 acquires each measured environmental value (measured value), such as the drive current of the laser 11, which is a laser environmental value, the temperature of the gas cell 21, which is a gas cell environmental value, the temperature of the photodetector 32, which is a photodetector environmental value, the control voltage of the oscillator 51, which is an oscillator environmental value, and the external temperature and external magnetic field, which are external environmental values, and passes them to the agent 41.

[0039] Next, the agent 41 outputs each environmental control value, which is an action A corresponding to each environmental value, which is the received state S, according to the policy π, which is a function optimized by reinforcement learning, and passes it to the processor 4 (step S23 in FIG. 5). For example, as an example of action A, the agent 41 outputs environmental control values ​​consisting of change rates such as +1% for the drive voltage of the laser 11, +2% for the temperature control voltage of the gas cell 21, and -1% for the control voltage of the oscillator 51.

[0040] Next, the processor 4 controls the configuration of the atomic oscillator so that it is in the state of each environmental control value corresponding to the output action A (step S24 in FIG. 5). That is, the processor 4 controls the state of the voltage values ​​applied to the laser 11, gas cell 21, photodetector 31, oscillator 51, etc. for each environmental control unit 10, etc., in accordance with each environmental control value corresponding to the output action A. Then, the above-mentioned control continues until an interrupt instruction is input (step S25 in FIG. 5).

[0041] As described above, the atomic oscillator of the present disclosure performs reinforcement learning using a reward according to the difference between the oscillation frequency and the reference frequency so that the agent 41 outputs the optimal control behavior according to the current state of the atomic oscillator. Therefore, by controlling the atomic oscillator based on the behavior output from the agent 41 that has completed reinforcement learning, it is possible to improve the frequency stability of the atomic oscillator.

[0042] <Embodiment 2> Next, a second embodiment of the present disclosure will be described with reference to the drawings. In this embodiment, an outline of the configuration of the atomic oscillator etc. described in the above embodiment is shown. Note that the drawings may be relevant to any embodiment.

[0043] As shown in FIG. 6, the atomic oscillator 100 in the present disclosure includes a gas cell 101 containing alkali metal atoms, a light generator 102 that irradiates the gas cell with irradiation light having at least two different frequency components, a photodetector 103 that detects the transmitted light that has passed through the gas cell, a controller 104 that determines a resonant frequency based on the amount of light transmitted through the gas cell and controls the oscillation frequency of the oscillator based on the determined resonant frequency, and an agent 105 that performs reinforcement learning using a reward according to the difference between a preset reference frequency and the oscillation frequency to output behavior that controls the environmental state of the atomic oscillator according to the acquired environmental state.

[0044] Then, in the atomic oscillator having the above configuration, the agent 105 performs reinforcement learning so as to output an action that controls the state of the environment of the acquired atomic oscillator according to the state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency (step S101 in FIG. 7).

[0045] Furthermore, in the atomic oscillator having the above configuration, the agent 105 that has performed reinforcement learning outputs an action according to the acquired state of the environment of the atomic oscillator, and the controller 104 controls the state of the environment of the atomic oscillator based on the action output from the agent (step S201 in FIG. 8).

[0046] As described above, in the present disclosure, reinforcement learning is performed using a reward according to the difference between the oscillation frequency and the reference frequency so that the agent outputs the optimal control behavior according to the current state of the atomic oscillator. Therefore, by controlling the atomic oscillator based on the behavior output from the agent that has completed reinforcement learning, it is possible to improve the frequency stability of the atomic oscillator.

[0047] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each of the above-described embodiments can be combined with other embodiments as appropriate.

[0048] <Additional Notes> Some or all of the above-described embodiments can be described as follows: The following provides an overview of the configuration of the atomic oscillator and the like in the present disclosure. However, the present disclosure is not limited to the configuration described in the following supplementary notes. Note that the configurations described in Supplements 2 to 8.2 that are dependent on Supplementary Note 1 above and some or all of the functions of the configurations may also be dependent on other Supplements 9 to 15 in the same dependent relationship as Supplements 2 to 8.2. Furthermore, not limited to Supplements 1, 9 to 15, but also within the scope of each of the above-mentioned embodiments, the configurations described as Supplements and some or all of the functions of the configurations may be made dependent on similar hardware, software, various recording means for recording software, or systems. (Appendix 1) a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a controller that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; An atomic oscillator comprising: an agent that performs reinforcement learning to output an action that controls a state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; Atomic oscillator. (Appendix 2) 2. The atomic oscillator according to claim 1, the controller sets the reward to a larger value as the absolute value of the difference between the oscillation frequency and the reference frequency decreases, The agent performs reinforcement learning using the reward. Atomic oscillator. (Appendix 3) 2. The atomic oscillator according to claim 1, the controller sets the reward according to the passage of time in the reinforcement learning; The agent performs reinforcement learning using the reward. Atomic oscillator. (Appendix 4) 4. The atomic oscillator according to claim 3, the controller sets the reward to a value that decreases as time passes during the reinforcement learning. Atomic oscillator. (Appendix 5) 2. The atomic oscillator according to claim 1, the controller sets a reward according to a difference between the oscillation frequency and the reference frequency when the state of the environment of the atomic oscillator is controlled based on the action output by the agent; The agent performs reinforcement learning using the reward. Atomic oscillator. (Appendix 6) 2. The atomic oscillator according to claim 1, the agent outputs the behavior in accordance with the acquired state of the environment of the atomic oscillator; the controller controls a state of the environment of the atomic oscillator based on the action output by the agent. Atomic oscillator. (Appendix 7) 2. The atomic oscillator according to claim 1, the environmental state acquired by the agent is at least one of measurements measured from the gas cell, the light generator, the light detector, and the oscillator; Atomic oscillator. (Appendix 8) 2. The atomic oscillator according to claim 1, the action output by the agent is a control value that controls a state of at least one of the gas cell, the light generator, the light detector, and the oscillator. Atomic oscillator. (Appendix 8.1) 2. The atomic oscillator according to claim 1, The agent performs Q-learning as the reinforcement learning. Atomic oscillator. (Appendix 8.2) 2. The atomic oscillator according to claim 1, a receiver for receiving the reference frequency from an external device; Atomic oscillator. (Appendix 9) a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a controller that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; An atomic oscillator comprising: an agent that has undergone reinforcement learning to output an action that controls a state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; the agent outputs the behavior in accordance with the acquired state of the environment of the atomic oscillator; the controller controls a state of the environment of the atomic oscillator based on the action output by the agent. Atomic oscillator. (Appendix 10) a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; A control method by the control device in an atomic oscillator comprising: an agent included in the control device performs reinforcement learning so as to output an action that controls the state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; Control method. (Appendix 11) a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; A control method by the control device in an atomic oscillator comprising: an agent provided in the control device that has performed reinforcement learning to output an action to control the state of the environment of the atomic oscillator according to the acquired state of the environment using a reward according to the difference between a preset reference frequency and the oscillation frequency outputs the action according to the acquired state of the environment of the atomic oscillator, and the control device controls the state of the environment of the atomic oscillator based on the action output from the agent; Control method. (Appendix 12) a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; The control device in an atomic oscillator comprising: an agent that performs reinforcement learning to output an action that controls a state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; Control device. (Appendix 13) a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; The control device in an atomic oscillator comprising: an agent that has undergone reinforcement learning to output an action that controls the state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency, and the agent outputs the action according to the acquired state of the environment of the atomic oscillator, and controls the state of the environment of the atomic oscillator based on the output action; Control device. (Appendix 14) a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; The control device in the atomic oscillator includes: an agent included in the control device performs reinforcement learning so as to output an action that controls the state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; A program that executes a process. (Appendix 15) a gas cell in which alkali metal atoms are sealed; a light generator that irradiates the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; In the control device in the atomic oscillator having the same, An agent included in the control device that performs reinforcement learning to output an action for controlling the state of the environment according to the state of the environment of the acquired atomic oscillator using a reward corresponding to the difference between a preset reference frequency and the oscillation frequency outputs the action according to the state of the environment of the acquired atomic oscillator and controls the state of the environment of the atomic oscillator based on the output action. A program for executing the process.

Explanation of Signs

[0049] 1 Light generator 10 Laser environment control unit 11 Laser 12 Laser environment sensor 20 Gas cell environment control unit 21 Gas cell 22 Gas cell environment sensor 30 Photodetector environment control unit 31 Photodetector 32 Photodetector environment sensor 4 Processor 41 Agent 50 Oscillator environment control unit 51 Oscillator 52 Oscillator environment sensor 61 External environment sensor 71 Frequency reference device 72 Frequency reference receiver 8 External device 100 Atomic oscillator 101 Gas cell 102 Light generator 103 Photodetector 104 Controller 105 Agent

Claims

1. a gas cell in which alkali metal atoms are sealed; a light generator for irradiating the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a controller that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; An atomic oscillator comprising: an agent that performs reinforcement learning to output an action that controls a state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; Atomic oscillator.

2. 2. The atomic oscillator according to claim 1, the controller sets the reward to a larger value as the absolute value of the difference between the oscillation frequency and the reference frequency decreases, The agent performs reinforcement learning using the reward. Atomic oscillator.

3. 2. The atomic oscillator according to claim 1, the controller sets the reward according to the passage of time in the reinforcement learning; The agent performs reinforcement learning using the reward. Atomic oscillator.

4. 4. The atomic oscillator according to claim 3, the controller sets the reward to a value that decreases as time passes during the reinforcement learning. Atomic oscillator.

5. 2. The atomic oscillator according to claim 1, the controller sets a reward according to a difference between the oscillation frequency and the reference frequency when the state of the environment of the atomic oscillator is controlled based on the action output by the agent; The agent performs reinforcement learning using the reward. Atomic oscillator.

6. 2. The atomic oscillator according to claim 1, the agent outputs the behavior in accordance with the acquired state of the environment of the atomic oscillator; the controller controls a state of the environment of the atomic oscillator based on the action output by the agent. Atomic oscillator.

7. 2. The atomic oscillator according to claim 1, the environmental state acquired by the agent is at least one of measurements measured from the gas cell, the light generator, the light detector, and the oscillator; Atomic oscillator.

8. 2. The atomic oscillator according to claim 1, the action output by the agent is a control value that controls a state of at least one of the gas cell, the light generator, the light detector, and the oscillator. Atomic oscillator.

9. a gas cell in which alkali metal atoms are sealed; a light generator for irradiating the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a controller that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; An atomic oscillator comprising: an agent that has undergone reinforcement learning to output an action that controls a state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; the agent outputs the behavior in accordance with the acquired state of the environment of the atomic oscillator; the controller controls a state of the environment of the atomic oscillator based on the action output by the agent. Atomic oscillator.

10. a gas cell in which alkali metal atoms are sealed; a light generator for irradiating the gas cell with light having at least two different frequency components; a photodetector that detects transmitted light that has passed through the gas cell; a control device that determines a resonance frequency based on the amount of light transmitted through the gas cell and controls an oscillation frequency of an oscillator based on the determined resonance frequency; A control method by the control device in an atomic oscillator comprising: an agent included in the control device performs reinforcement learning so as to output an action that controls the state of the environment of the atomic oscillator according to the acquired state of the environment, using a reward according to the difference between a preset reference frequency and the oscillation frequency; Control method.

Citation Information

Patent Citations

  • Atomic oscillator

    JP2017011680A