A method and apparatus for constructing an intelligent quantum sensor

By constructing a fundamental physical system model for intelligent quantum sensors and utilizing digital twins of characterization networks and reinforcement learning networks, the problem of insufficient measurement accuracy caused by noise interference was solved, and high-precision measurement by intelligent quantum sensors was achieved.

CN120832818BActive Publication Date: 2026-04-03SHANGHAI JIAOTONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing intelligent quantum sensors are open systems, which are susceptible to environmental noise interference, resulting in insufficient measurement accuracy. Furthermore, the unknown type and intensity of the noise make it impossible to directly design sensing solutions.

Method used

The basic physical system model of the intelligent quantum sensor is determined based on the parameters to be measured. A digital twin is constructed using a trained representation network. The target control strategy is determined through a reinforcement learning network, and the basic physical system model is optimized to compensate for errors caused by noise.

Benefits of technology

By introducing representation networks and reinforcement learning networks, the measurement accuracy of intelligent quantum sensors is improved through real-time adaptive compensation of system errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832818B_ABST
    Figure CN120832818B_ABST
Patent Text Reader

Abstract

This application provides a method and apparatus for constructing an intelligent quantum sensor, applicable to the field of quantum precision measurement technology. The method includes: determining a fundamental physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured; constructing a digital twin corresponding to the intelligent quantum sensor using a trained representation network based on the fundamental physical system model; determining a target control strategy using a trained reinforcement learning network based on the digital twin, and optimizing the fundamental physical system model based on the target control strategy to obtain a target physical system model corresponding to the intelligent quantum sensor. Thus, by introducing a representation network and a reinforcement learning network, and using the digital twin constructed by the representation network to infer the system's random errors, the method ensures that the reinforcement learning network can adaptively compensate for these errors in real time, thereby improving the measurement accuracy of the intelligent quantum sensor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of quantum precision measurement technology, and in particular to a method and apparatus for constructing an intelligent quantum sensor. Background Technology

[0002] Quantum sensors are physical devices designed based on the laws of quantum mechanics. They can achieve high-precision measurements through quantum effects, and their core technologies include phenomena such as quantum entanglement and atomic spin interference.

[0003] Without considering noise, the Heisenberg limit for quantum sensors can be achieved using coherence and entanglement. However, existing quantum sensing systems are inherently open systems, meaning that environmental or noise interference can cause decoherence. Furthermore, the unknown type and intensity of noise lead to an unknown evolutionary process for the quantum sensor system, making it impossible to directly design sensing schemes for existing intelligent quantum sensors, thus resulting in insufficient measurement accuracy.

[0004] Therefore, improving the measurement accuracy of intelligent quantum sensors is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] To address the aforementioned problems, this application provides a method and apparatus for constructing an intelligent quantum sensor, comprising:

[0006] Determine the underlying physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured;

[0007] Based on the aforementioned fundamental physical system model, a digital twin corresponding to the intelligent quantum sensor is constructed using a trained representation network;

[0008] Based on the digital twin, a target control strategy is determined using a trained reinforcement learning network, and the basic physical system model is optimized based on the target control strategy to obtain the target physical system model corresponding to the intelligent quantum sensor.

[0009] Optionally, determining the fundamental physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured includes:

[0010] Determine the physical platform corresponding to the intelligent quantum sensor based on the parameters to be measured;

[0011] A quantum mechanical model corresponding to the physical platform is established, and the encoding function of the parameter to be measured in the quantum mechanical model is calibrated.

[0012] Optionally, the step of constructing a digital twin corresponding to the intelligent quantum sensor based on the fundamental physical system model and using a trained representation network includes:

[0013] Obtain the expected values ​​of the initial observables corresponding to the basic physical system model;

[0014] The trained representation network is used to predict the expected value of the initial observable, and the predicted value is used as the digital twin of the intelligent quantum sensor.

[0015] Optionally, the representation network is trained using the following method:

[0016] Obtain the training dataset;

[0017] Using the training dataset, the initial representation network is used to make predictions, and the prediction results are obtained.

[0018] Construct a first loss function corresponding to the prediction result;

[0019] The first gradient information of the first loss function with respect to the initial parameters of the initial representation network is determined by the backpropagation algorithm;

[0020] Based on the first gradient information and the first preset termination condition, the initial parameters of the initial representation network are updated using the gradient descent method, and the trained representation network is obtained.

[0021] Optionally, obtaining the training dataset includes:

[0022] Initialize the basic physical system model;

[0023] A random control pulse is applied to the basic physical system model, causing the basic physical system model to evolve to a random initial state;

[0024] As the underlying physical system model continues to evolve from a random initial state to a final state, the expected values ​​of observable quantities are recorded based on continuous weak measurements, and these expected values ​​are used as the training dataset.

[0025] Optionally, the step of combining the training dataset, using the initial representation network for prediction, and obtaining the prediction result includes:

[0026] Based on the expected values ​​of the initial observables and the continuous weak measurement signals in the training dataset, the prediction results corresponding to the expected values ​​of the initial observables are generated through the forward propagation process of the initial representation network.

[0027] Optionally, the reinforcement learning network is trained using the following method:

[0028] Using the digital twin, the target control action is determined by an initial reinforcement learning network;

[0029] Construct a second loss function based on the target control action;

[0030] The second gradient information of the second loss function with respect to the initial parameters of the initial reinforcement learning network is determined by the backpropagation algorithm;

[0031] Based on the second gradient information and the second preset termination condition, the initial parameters of the initial reinforcement learning network are updated using the gradient descent method, and a trained reinforcement learning network is obtained.

[0032] Optionally, the step of combining the digital twin and using an initial reinforcement learning network to determine the target control action includes:

[0033] Extract real-time feature data from the digital twin;

[0034] The real-time feature data is input into the initial reinforcement learning network, and the first Q value of each control action is calculated through multi-layer nonlinear transformation.

[0035] Based on the first Q value, the first control action is determined using a greedy algorithm;

[0036] A first reward function is constructed by combining the predicted value in the digital twin and the ideal target state corresponding to the first control action, and the target control action is determined based on the first reward function.

[0037] Optionally, constructing the second loss function based on the target control action includes:

[0038] Sample the target control action and the corresponding batch of transfer data from the experience pool;

[0039] Based on the target control action and the same batch of transfer data, the corresponding second Q value and second reward function are calculated according to the forward propagation of the initial reinforcement learning network.

[0040] The mean squared error is calculated based on the second Q value and the second reward function, and the second loss function is obtained.

[0041] Secondly, embodiments of this application provide an apparatus for constructing an intelligent quantum sensor, comprising:

[0042] The determination module is used to determine the underlying physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured.

[0043] A construction module is used to construct a digital twin corresponding to the intelligent quantum sensor based on the basic physical system model and using a trained representation network;

[0044] An optimization module is used to determine a target control strategy based on the digital twin using a trained reinforcement learning network, and to optimize the basic physical system model based on the target control strategy to obtain the target physical system model corresponding to the intelligent quantum sensor.

[0045] As can be seen from the above technical solutions, compared with the prior art, this application has the following advantages:

[0046] The method for constructing an intelligent quantum sensor provided in this application includes first determining the underlying physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured. Then, based on the underlying physical system model, a digital twin corresponding to the intelligent quantum sensor is constructed using a trained representation network. Finally, based on the digital twin, a target control strategy is determined using a trained reinforcement learning network, and the underlying physical system model is optimized based on the target control strategy to obtain the target physical system model corresponding to the intelligent quantum sensor. Thus, by introducing a representation network and a reinforcement learning network, and using the digital twin constructed by the representation network to infer the system's random errors, the method ensures that the reinforcement learning network can adaptively compensate for these errors in real time, thereby improving the measurement accuracy of the intelligent quantum sensor. Attached Figure Description

[0047] Figure 1 A flowchart illustrating a method for constructing an intelligent quantum sensor, as provided in this application embodiment;

[0048] Figure 2 A flowchart of training data collection is provided for an embodiment of this application;

[0049] Figure 3 A flowchart of a characterization network training process is provided as an embodiment of this application;

[0050] Figure 4 A flowchart of a characterization network inference provided in an embodiment of this application;

[0051] Figure 5 A flowchart illustrating a digital twin technology provided in this application embodiment;

[0052] Figure 6 A flowchart illustrating the training process of a reinforcement learning network provided in an embodiment of this application;

[0053] Figure 7 An inference flowchart of a reinforcement learning network provided in an embodiment of this application;

[0054] Figure 8 This is a schematic diagram of a device for constructing an intelligent quantum sensor, provided in an embodiment of this application. Detailed Implementation

[0055] As mentioned earlier, existing intelligent quantum sensors suffer from insufficient measurement accuracy. Specifically, since existing quantum sensor systems are inherently open systems, they are susceptible to environmental or noise interference. Noise-induced decoherence impairs sensing accuracy, and the unknown type and intensity of noise lead to an unknown system evolution process, making it impossible to directly design sensing schemes. Consequently, existing intelligent quantum sensors suffer from insufficient measurement accuracy.

[0056] To address the aforementioned issues, embodiments of this application provide a method for constructing an intelligent quantum sensor, comprising: first, determining the underlying physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured; then, constructing a digital twin corresponding to the intelligent quantum sensor using a trained representation network based on the underlying physical system model; and finally, determining a target control strategy using a trained reinforcement learning network based on the digital twin, and optimizing the underlying physical system model based on the target control strategy to obtain the target physical system model corresponding to the intelligent quantum sensor.

[0057] Thus, by introducing representation networks and reinforcement learning networks, the random errors of the system can be inferred through the digital twins constructed by the representation networks, thereby ensuring that the reinforcement learning networks can adaptively compensate for these errors in real time, thus improving the measurement accuracy of the intelligent quantum sensor.

[0058] It should be noted that the method and apparatus for constructing an intelligent quantum sensor provided in this application embodiment can be applied to the field of quantum precision measurement technology. The above are merely examples and do not limit the application field of the method and apparatus for constructing an intelligent quantum sensor provided in this application embodiment.

[0059] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0060] Figure 1 This is a flowchart illustrating a method for constructing an intelligent quantum sensor, as provided in an embodiment of this application. (In conjunction with...) Figure 1 As shown, the method for constructing a smart quantum sensor provided in this application embodiment may include:

[0061] S101: Determine the basic physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured.

[0062] In practical applications, the quantum sensor system first needs to be initialized, which involves selecting the (basic) physical system model corresponding to the intelligent quantum sensor based on the properties of the parameters to be measured. This model can be mainly divided into two categories: one is a discrete variable system based on an atomic platform (such as diamond NV centers, cold atom ensembles, or ion traps), suitable for measuring discrete quantities such as magnetic fields and electric fields; the other is a continuous variable system based on an optical platform (such as photomechanical systems or squeezed-state light fields), suitable for measuring continuously changing physical quantities such as displacement and temperature. Further, a quantum mechanical model corresponding to the physical system model is established, including the system's Hamiltonian, decoherence mechanism, and mathematical description of the measurement scheme. Understandably, the model selection also needs to consider the constraints of the measurement environment, such as operating temperature and vacuum level, to ensure that the sensor model matches the actual application scenario.

[0063] Furthermore, since the methods for determining the underlying physical system model corresponding to the intelligent quantum sensor are not entirely the same, this application embodiment can describe one possible determination method.

[0064] In one case, S101: Determine the fundamental physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured, which may specifically include:

[0065] Determine the physical platform corresponding to the intelligent quantum sensor based on the parameters to be measured;

[0066] A quantum mechanical model corresponding to the physical platform is established, and the encoding function of the parameter to be measured in the quantum mechanical model is calibrated.

[0067] In practical applications, the choice of a fundamental physical system model is related to its response to unknown parameters. Understandably, the diversity of the parameters to be measured requires us to select appropriate quantum sensors based on their characteristics. For example, diamond NV centers are suitable for DC-microwave magnetic fields, Rydberg atoms for GHz-THz electric fields, and superconducting qubits for weak magnetic field detection; mechanical parameter measurements can utilize optomechanical systems (nanoscale displacement), spin-mechanical coupling systems (micrometer vibrations), or surface acoustic wave resonators (mass changes); temperature sensing can employ quantum dots (nanoscale), rare-earth ion crystals (wide temperature range), or nuclear magnetic resonance systems (bioimaging). Thus, after selecting a physical platform, its corresponding quantum mechanical model is established. Then, the parameters to be measured are determined through a combination of experimental calibration (changing parameters under controllable conditions and measuring the system response) and theoretical modeling (establishing a coupling model based on quantum mechanics). Encoding function in system Hamiltonian ,in, For storing coupling quantities, and coupling parameters Accurate calibration requires curve fitting and error analysis (assessing uncertainty, noise effects, and establishing correction models). Furthermore, in practical applications, it may be necessary to optimize sensitivity, adjust dynamic range, and improve environmental adaptability. These systematic parameter-platform matching and accurate calibration methods are key to realizing high-precision quantum sensors.

[0068] S102: Based on the aforementioned fundamental physical system model, construct a digital twin corresponding to the intelligent quantum sensor using the trained representation network.

[0069] In practical applications, a well-trained representation network can accurately represent the derivation process of a fundamental physical system model. The actual sampled values ​​of the fundamental physical system model are input into the trained representation network, which then outputs predicted values. It will be established as the digital twin of this intelligent quantum sensor.

[0070] Furthermore, since there are different ways to obtain digital twins, this application embodiment can describe one possible method of obtaining them.

[0071] In one scenario, S102: Based on the aforementioned fundamental physical system model, a digital twin corresponding to the intelligent quantum sensor is constructed using a trained representation network, which may specifically include:

[0072] Obtain the expected values ​​of the initial observables corresponding to the basic physical system model;

[0073] The trained representation network is used to predict the expected value of the initial observable, and the predicted value is used as the digital twin of the intelligent quantum sensor.

[0074] In practical applications, a control impulse is given to a fundamental physical system model, causing it to evolve into a random state, which is then used as the initial state to be sampled. This initial state is then sampled to obtain the expected values ​​of the initial observables corresponding to the fundamental physical system model. The trained representation network is used to predict the final state of the fundamental physical system model, that is, the predicted value corresponding to the expected value of the initial observables. and the predicted value This serves as the digital twin of the intelligent quantum sensor.

[0075] S103: Based on the digital twin, a target control strategy is determined using a trained reinforcement learning network, and the basic physical system model is optimized based on the target control strategy to obtain the target physical system model corresponding to the intelligent quantum sensor.

[0076] In practical applications, a well-trained reinforcement learning network can accurately determine its optimal control strategy based on the derivation process of the fundamental physical system model, thereby mitigating the impact of noise. Combined with the aforementioned well-trained representation network, which can accurately represent the derivation process of the fundamental physical system model, the digital twin, as the output of the representation network, has its corresponding actual feature data (including observable expected values). This leads to the (trained reinforcement learning network) determining the target control strategy (optimal control strategy) for the underlying physical system model based on the digital twin. The target control strategy is then used to optimize the underlying physical system model, resulting in a target physical system model that mitigates the effects of noise and improves the measurement accuracy of the intelligent quantum sensor.

[0077] Furthermore, since the methods for training representation networks are not entirely the same, the embodiments of this application can describe one possible training method.

[0078] In one case, the representation network is trained using the following method:

[0079] S201: Obtain the training dataset.

[0080] In practical applications, training datasets can be obtained from existing sample data or based on fundamental physical system models.

[0081] Furthermore, since there are different ways to obtain training datasets, this application embodiment can describe one possible acquisition method.

[0082] In one scenario, S201: Obtain the training dataset, specifically including:

[0083] Initialize the basic physical system model;

[0084] A random control pulse is applied to the basic physical system model, causing the basic physical system model to evolve to a random initial state;

[0085] As the underlying physical system model continues to evolve from a random initial state to a final state, the expected values ​​of observable quantities are recorded based on continuous weak measurements, and these expected values ​​are used as the training dataset.

[0086] Figure 2 This is a flowchart illustrating a training data collection process as provided in an embodiment of this application. (In conjunction with...) Figure 2 As shown, the quantum sensor (a fundamental physical system model) is first initialized, and its initial state is then randomly sampled. Since the initial state of a quantum sensor is usually deterministic (e.g., the ground state),... Therefore, it is necessary to apply a random quantum gate or a random (control) pulse to evolve it into a random state, which can then be used as the random initial state to be sampled. This ensures that the data covers the possible evolutionary behavior of the sensor. Then, the sensor is allowed to undergo a period of free / controlled stochastic master equation evolution, eventually reaching the final state. Furthermore, to reflect the influence of continuous weak measurements, the data recorded during the evolution from the initial random state to the final state will be preserved (this data contains key features of the random evolution during decoherence) as part of the training dataset, i.e., outputting the final state and collecting measurement data. This is done to obtain the expected value of the desired observable. and the expected value of the final state observable This requires repeatedly preparing the initial and final states and performing projection measurements on them until the maximum dataset is reached, at which point the process ends and the training dataset is acquired. Understandably, if the maximum dataset is not reached, the above process is repeated. Furthermore, the training dataset can be constructed in the following way: As the input to the neural network, the corresponding As a label. This design effectively utilizes the implicit random temporal information in the measurement data while maintaining computational efficiency. It is a continuous weak measurement signal.

[0087] S202: Using the training dataset, make predictions using the initial representation network and obtain the prediction results.

[0088] In practical applications, representation networks can predict the next state of a fundamental physical system model based on its current state. Therefore, by using the training dataset as input to the initial representation network, the corresponding prediction results can be obtained.

[0089] Furthermore, since the methods for obtaining prediction results are not entirely the same, this application embodiment can describe one possible method of obtaining the results.

[0090] In one scenario, S202: Combining the training dataset, prediction is performed using the initial representation network to obtain the prediction result, which may specifically include:

[0091] Based on the expected values ​​of the initial observables and the continuous weak measurement signals in the training dataset, the prediction results corresponding to the expected values ​​of the initial observables are generated through the forward propagation process of the initial representation network.

[0092] In practical applications, based on the aforementioned training dataset, the initial representation network receives input data. Then, prediction results are generated through the forward propagation process. This refers to the predicted value during the training process. This prediction process can be represented as... , where the function Represents the nonlinear mapping relationship of a neural network. This represents all trainable parameters of the network (including the weight matrix and bias vector of each layer).

[0093] S203: Construct a first loss function corresponding to the prediction result.

[0094] In practical applications, the above prediction results, i.e., the predicted values, are... Compared with the actual measured value A comparative evaluation is performed to construct the first loss function. The commonly used one is the root mean square error (RMSE), whose mathematical expression is as follows:

[0095] ;

[0096] In the formula Indicates the number of training samples. The specific training samples are indexed. Thus, the first loss function effectively reflects the overall deviation between the predicted and true values, facilitating subsequent backpropagation algorithm optimization of network parameters. It is understandable that the aforementioned first loss function possesses good differentiability mathematically. However, in certain specific application scenarios, other forms of loss functions (such as mean absolute error (MAE) or customized loss functions tailored to the characteristics of quantum sensor systems) can be considered to better adapt to the specific needs of the problem.

[0097] S204: Determine the first gradient information of the first loss function with respect to the initial parameters of the initial representation network through the backpropagation algorithm.

[0098] In practical applications, the first loss function is first calculated using the backpropagation algorithm. For the initial parameters (characterizing all trainable parameters of the network) gradient information .

[0099] S205: Based on the first gradient information and the first preset termination condition, the initial parameters of the initial representation network are updated using the gradient descent method, and the trained representation network is obtained.

[0100] In practical applications, gradient descent is used to update the initial parameters. The expression is: Among them, the learning rate This is used to control the update step size and convergence speed, while combining adaptive optimization strategies (such as Adam), gradient clipping, momentum terms, and other techniques to improve the training effect of the representation network. Furthermore, the training process terminates when the network iteration meets the first preset termination condition (when the loss function converges (i.e., the change in loss value over several consecutive iterations is less than a preset threshold) or when the preset maximum number of iterations is reached), resulting in a trained representation network.

[0101] In summary, Figure 3 This is a flowchart illustrating a characterization network training process provided in an embodiment of this application. (Combined with...) Figure 3 As shown, after the training of the representation network begins, the neural network (representation network) is first initialized, then batches of data are randomly sampled, the loss function is calculated, and then the network parameters are updated using the gradient descent method. Once the loss function converges or the number of iterations is reached, the training ends, and the trained representation network is obtained. If the loss function does not converge or the number of iterations is not reached, data is sampled and the network parameters are updated until convergence or the number of iterations is reached.

[0102] Furthermore, once the training process of the network to be represented terminates, the predicted values ​​output by its model... It will be established as a digital twin of the intelligent quantum sensor for subsequent quantum state prediction and parameter estimation tasks. Figure 4 This is a flowchart illustrating a network characterization inference process provided in an embodiment of this application. (Combined with...) Figure 4 As shown, the quantum sensor is first initialized, then allowed to undergo a period of free / controlled stochastic master equation evolution, eventually reaching a final state. During this process, continuous weak measurements output the final state and receive measurement data, which is then input into a time network (a trained representation network). The representation network outputs predicted sensor features. This concludes the representation network inference, and a digital twin corresponding to the intelligent quantum sensor is constructed based on this predicted value.

[0103] Furthermore, since there are different ways to train reinforcement learning networks, this application embodiment can describe one possible training method.

[0104] In one case, the reinforcement learning network is trained using the following method:

[0105] S301: Combine the digital twin with the initial reinforcement learning network to determine the target control action.

[0106] In practical applications, a well-trained representation network can accurately represent the derivation process of a fundamental physical system model. A well-trained reinforcement learning network is then used to determine the optimal control action based on observable expected values. Therefore, during the training of the reinforcement learning network, a digital twin constructed from the representation network's predictions can be used to first determine the target control action using the initial reinforcement learning network, and then the parameters of the reinforcement learning network can be updated.

[0107] Furthermore, since the methods for determining the target control action are not entirely the same, this application embodiment can describe one possible determination method.

[0108] In one scenario, S301: Combining the digital twin, the target control action is determined using an initial reinforcement learning network, which may specifically include:

[0109] Extract real-time feature data from the digital twin;

[0110] The real-time feature data is input into the initial reinforcement learning network, and the first Q value of each control action is calculated through multi-layer nonlinear transformation.

[0111] Based on the first Q value, the first control action is determined using a greedy algorithm;

[0112] A first reward function is constructed by combining the predicted value in the digital twin and the ideal target state corresponding to the first control action, and the target control action is determined based on the first reward function.

[0113] In practical applications, the first step is to randomly initialize all parameters of the reinforcement learning network (including weight matrices and bias vectors) using a Gaussian or uniform distribution to ensure the network has sufficient exploratory capabilities in the early stages of training. Then, real-time feature data (including the expected values ​​of observables) is extracted from the digital twin of the intelligent quantum sensor. These data are used as input to the initial reinforcement learning network, which calculates the (first) Q-value for each possible control action through multiple layers of nonlinear transformations, evaluating its potential for long-term cumulative reward in the current state. Then, based on a greedy algorithm... The "greedy strategy" selects the final control action (i.e., the first control action) from among various control actions. Specifically, this process can involve randomly selecting a control action with probability ε to facilitate exploration, or using probability... The control action with the highest current Q-value is selected to achieve optimal control. This first control action is then converted into a specific control impulse or quantum gate operation and applied to a real quantum sensor (which could be a model of the underlying physical system), driving it to evolve towards the target state. Simultaneously, a (first) reward function is introduced. Based on this reward function, the effectiveness of the first control action is evaluated using the predicted value in the digital twin and the ideal target state, thus determining the target control action. Here, the predicted value represents the network output, while the ideal target state is the expected value of the final observable of the real quantum sensor. Furthermore, the design of the first reward function typically employs key indicators from quantum information theory, such as quantum Fisher information or classical Fisher information, converting these indicators into scalar reward values. The magnitude of these function values ​​directly reflects the contribution of the control action to improving sensor performance. Further, the reward function may also include a regularization term to constrain the energy or smoothness of the control impulse, ensuring its feasibility in physical experiments. Thus, through this design, the reward mechanism can effectively guide the reinforcement learning strategy towards optimizing sensor sensitivity and accuracy.

[0114] In summary, Figure 5 A flowchart illustrating a digital twin technology provided in an embodiment of this application. (In conjunction with...) Figure 5 As shown, a real quantum sensor system is constructed, and the system is monitored and the measurement data is stored through continuous weak measurements. Then, a representation network is trained using the measurement data as training data to obtain a trained representation network. Based on the measurement data of the real quantum sensor system, the trained representation network can generate prediction results through a forward propagation process and construct a digital twin. The digital twin maps to the actual values ​​of the real quantum sensor system. Further, the digital twin serves as input to a reinforcement learning network, which generates an output (the first control action), which is then converted into a control pulse input to the real quantum sensor system. At this point, the quantum sensor state is updated, and under the actual action of the first control action, the quantum sensor undergoes controlled dynamic evolution. The corresponding measurement data (the continuously weak measurement signal acquired in real time) again serves as input to the representation network, which makes predictions, outputs predicted values, and updates the digital twin. The updated digital twin then serves as input to the reinforcement learning network again, thereby generating a second control action. This forms a closed loop. By introducing a first reward function, the above process is iterated and eventually converges to obtain the final target control action.

[0115] S302: Construct a second loss function based on the target control action.

[0116] In practical applications, after selecting the target control action, a second loss function can be constructed based on it.

[0117] Furthermore, since the methods for constructing the second loss function are not entirely the same, this application embodiment can describe one possible construction method.

[0118] In one scenario, S302: Constructing a second loss function based on the target control action, which may specifically include:

[0119] Sample the target control action and the corresponding batch of transfer data from the experience pool;

[0120] Based on the target control action and the same batch of transfer data, the corresponding second Q value and second reward function are calculated according to the forward propagation of the initial reinforcement learning network.

[0121] The mean squared error is calculated based on the second Q value and the second reward function, and the second loss function is obtained.

[0122] In practical applications, the initial reinforcement learning network parameters can be updated using a temporal difference (TD) learning algorithm based on experience replay. Specifically, a batch of transition data (including state, action, reward function, and next state) is first randomly sampled from the experience pool. If the selected action is the target control action, its corresponding state, reward, and next state are included as the transition data in the same batch. Then, the mean squared error is calculated together with the (second) Q-value and (second) reward function calculated by the forward propagation of the initial reinforcement learning network as the TD loss function (i.e., the second loss function). In other words, the Q-values ​​of the current state and the next state can be calculated using the current Q-network and the target Q-network, respectively. Then, a TD target value is constructed based on the Bellman equation. By comparing the mean squared error between the predicted Q-value and the TD target value, a second loss function is constructed to guide the initial reinforcement learning network in parameter optimization.

[0123] S303: Determine the second gradient information of the second loss function with respect to the initial parameters of the initial reinforcement learning network through the backpropagation algorithm.

[0124] In practical applications, the second gradient information of the second loss function with respect to the initial parameters of the initial reinforcement learning can be calculated using the backpropagation algorithm. Then, the weights of the Q-network can be updated using gradient descent (such as the Adam optimizer). This offline learning strategy not only improves data efficiency but also breaks down the correlation between data points through random sampling, significantly enhancing training stability.

[0125] S304: Based on the second gradient information and the second preset termination condition, the initial parameters of the initial reinforcement learning network are updated using the gradient descent method, and the trained reinforcement learning network is obtained.

[0126] In practical applications, the initial parameters of the initial reinforcement learning network are updated using gradient descent. The training process terminates when the reinforcement learning network iterations meet the second preset termination condition (the rate of change of the loss function is lower than a preset threshold (indicating policy convergence) or the preset maximum number of training rounds is reached), resulting in a trained reinforcement learning network. In each iteration, the system fully records key metrics such as average reward, TD error magnitude, and policy entropy to evaluate training progress. Understandably, if the termination condition is not met, a new training cycle is initiated. First, control impulses are generated based on the updated control actions. Then, the evolution and state monitoring of the real quantum sensor system are performed. Next, the digital twin and reward evaluation are updated through the representation network. Finally, the reinforcement learning network is optimized through experience replay. This iterative process continues until the reinforcement learning network reaches the required control accuracy or exhausts the computational budget. The resulting agent model (the target physical system model that determines the target control policy) can be directly deployed in the adaptive control task of a real quantum sensor.

[0127] In summary, Figure 6 This is a flowchart illustrating the training process of a reinforcement learning network, as provided in an embodiment of this application. (Combined with...) Figure 6 As shown, the reinforcement learning network, experience pool, and intelligent quantum sensor states are first initialized. Then, the current state features (i.e., real-time feature data) are obtained from the digital twin. Next, a greedy strategy is used to select the first control action, calculate its corresponding reward function, and thus determine the target control action. Then, data from the same batch as the selected action are sampled from the experience pool, the TD loss is calculated, and the network parameters are iteratively updated. Simultaneously, the sensor is updated based on the updated network parameters and the target control action, and its corresponding real data is stored. Finally, it is determined whether the current iteration has reached the maximum time step, and whether the network has reached the target accuracy or training epochs. If so, training ends, and the trained reinforcement learning network is obtained; otherwise, iterative training continues according to the above steps.

[0128] Figure 7 This is a flowchart illustrating the inference process of a reinforcement learning network, as provided in an embodiment of this application. (Combined with...) Figure 7 As shown, the inference process of a reinforcement learning network first initializes the state of the quantum sensor, then obtains the current state features from the digital twin. Next, it selects a control action based on the saved model, updates the sensor according to this action, and stores the corresponding real data. Finally, it determines whether the current iteration has reached the maximum time step; if so, training ends, resulting in a trained reinforcement learning network; otherwise, iterative training continues according to the above steps.

[0129] In summary, the method for constructing an intelligent quantum sensor provided in this application includes first determining the fundamental physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured. Then, based on the fundamental physical system model, a digital twin corresponding to the intelligent quantum sensor is constructed using a trained representation network. Finally, based on the digital twin, a target control strategy is determined using a trained reinforcement learning network, and the fundamental physical system model is optimized based on the target control strategy to obtain the target physical system model corresponding to the intelligent quantum sensor. Thus, by introducing a representation network and a reinforcement learning network, and using the digital twin constructed by the representation network to infer the system's random errors, the method ensures that the reinforcement learning network can adaptively compensate for these errors in real time, thereby improving the measurement accuracy of the intelligent quantum sensor.

[0130] Figure 8 This is a schematic diagram of a construction device for an intelligent quantum sensor provided in an embodiment of this application. (Combined with...) Figure 8 As shown, the construction apparatus 800 for the intelligent quantum sensor may include:

[0131] The determination module 801 is used to determine the basic physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured.

[0132] The construction module 802 is used to construct a digital twin corresponding to the intelligent quantum sensor based on the basic physical system model and using a trained representation network;

[0133] The optimization module 803 is used to determine the target control strategy based on the digital twin using a trained reinforcement learning network, and optimize the basic physical system model based on the target control strategy to obtain the target physical system model corresponding to the intelligent quantum sensor.

[0134] As one implementation method, regarding how to determine the fundamental physical system model corresponding to the intelligent quantum sensor, the aforementioned determining module 801 is specifically used for:

[0135] Determine the physical platform corresponding to the intelligent quantum sensor based on the parameters to be measured;

[0136] A quantum mechanical model corresponding to the physical platform is established, and the encoding function of the parameter to be measured in the quantum mechanical model is calibrated.

[0137] As one implementation method, regarding how to construct a digital twin corresponding to the intelligent quantum sensor, the aforementioned construction module 802 is specifically used for:

[0138] Obtain the expected values ​​of the initial observables corresponding to the basic physical system model;

[0139] The trained representation network is used to predict the expected value of the initial observable, and the predicted value is used as the digital twin of the intelligent quantum sensor.

[0140] As one implementation method, the above-mentioned intelligent quantum sensor construction device 800 further includes: a first training module, which is used to train the representation network;

[0141] The first training module is used to obtain the training dataset;

[0142] Using the training dataset, the initial representation network is used to make predictions, and the prediction results are obtained.

[0143] Construct a first loss function corresponding to the prediction result;

[0144] The first gradient information of the first loss function with respect to the initial parameters of the initial representation network is determined by the backpropagation algorithm;

[0145] Based on the first gradient information and the first preset termination condition, the initial parameters of the initial representation network are updated using the gradient descent method, and the trained representation network is obtained.

[0146] The acquisition of the training dataset includes:

[0147] Initialize the basic physical system model;

[0148] A random control pulse is applied to the basic physical system model, causing the basic physical system model to evolve to a random initial state;

[0149] As the underlying physical system model continues to evolve from a random initial state to a final state, the expected values ​​of observable quantities are recorded based on continuous weak measurements, and these expected values ​​are used as the training dataset.

[0150] The step of combining the training dataset, using the initial representation network to make predictions, and obtaining prediction results includes:

[0151] Based on the expected values ​​of the initial observables and the continuous weak measurement signals in the training dataset, the prediction results corresponding to the expected values ​​of the initial observables are generated through the forward propagation process of the initial representation network.

[0152] As one implementation method, the above-mentioned intelligent quantum sensor construction device 800 further includes a second training module for training the representation network;

[0153] The second training module is used to combine the digital twin and use the initial reinforcement learning network to determine the target control action;

[0154] Construct a second loss function based on the target control action;

[0155] The second gradient information of the second loss function with respect to the initial parameters of the initial reinforcement learning network is determined by the backpropagation algorithm;

[0156] Based on the second gradient information and the second preset termination condition, the initial parameters of the initial reinforcement learning network are updated using the gradient descent method, and a trained reinforcement learning network is obtained.

[0157] The step of combining the digital twin and using an initial reinforcement learning network to determine the target control action includes:

[0158] Extract real-time feature data from the digital twin;

[0159] The real-time feature data is input into the initial reinforcement learning network, and the first Q value of each control action is calculated through multi-layer nonlinear transformation.

[0160] Based on the first Q value, the first control action is determined using a greedy algorithm;

[0161] A first reward function is constructed by combining the predicted value in the digital twin and the ideal target state corresponding to the first control action, and the target control action is determined based on the first reward function.

[0162] The construction of the second loss function based on the target control action includes:

[0163] Sample the target control action and the corresponding batch of transfer data from the experience pool;

[0164] Based on the target control action and the same batch of transfer data, the corresponding second Q value and second reward function are calculated according to the forward propagation of the initial reinforcement learning network.

[0165] The mean squared error is calculated based on the second Q value and the second reward function, and the second loss function is obtained.

[0166] In summary, the method for constructing an intelligent quantum sensor provided in this application includes first determining the fundamental physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured. Then, based on the fundamental physical system model, a digital twin corresponding to the intelligent quantum sensor is constructed using a trained representation network. Finally, based on the digital twin, a target control strategy is determined using a trained reinforcement learning network, and the fundamental physical system model is optimized based on the target control strategy to obtain the target physical system model corresponding to the intelligent quantum sensor. Thus, by introducing a representation network and a reinforcement learning network, and using the digital twin constructed by the representation network to infer the system's random errors, the method ensures that the reinforcement learning network can adaptively compensate for these errors in real time, thereby improving the measurement accuracy of the intelligent quantum sensor.

[0167] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for constructing an intelligent quantum sensor, characterized in that, The method includes: Determine the underlying physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured; Based on the aforementioned fundamental physical system model, a digital twin corresponding to the intelligent quantum sensor is constructed using a trained representation network; Based on the digital twin, a target control strategy is determined using a trained reinforcement learning network, and the basic physical system model is optimized based on the target control strategy to obtain the target physical system model corresponding to the intelligent quantum sensor. The reinforcement learning network was trained using the following method: Using the digital twin, the target control action is determined by an initial reinforcement learning network; Construct a second loss function based on the target control action; The second gradient information of the second loss function with respect to the initial parameters of the initial reinforcement learning network is determined by the backpropagation algorithm; Based on the second gradient information and the second preset termination condition, the initial parameters of the initial reinforcement learning network are updated using the gradient descent method, and a trained reinforcement learning network is obtained. The process of combining the digital twin and using an initial reinforcement learning network to determine the target control action includes: Extract real-time feature data from the digital twin; The real-time feature data is input into the initial reinforcement learning network, and the first Q value of each control action is calculated through multi-layer nonlinear transformation. Based on the first Q value, the first control action is determined using a greedy algorithm; A first reward function is constructed by combining the predicted value in the digital twin and the ideal target state corresponding to the first control action, and the target control action is determined based on the first reward function; The construction of the second loss function based on the target control action includes: Sample the target control action and the corresponding batch of transfer data from the experience pool; Based on the target control action and the same batch of transfer data, the corresponding second Q value and second reward function are calculated according to the forward propagation of the initial reinforcement learning network. The mean squared error is calculated based on the second Q value and the second reward function, and the second loss function is obtained.

2. The method according to claim 1, characterized in that, The method for determining the fundamental physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured includes: Determine the physical platform corresponding to the intelligent quantum sensor based on the parameters to be measured; A quantum mechanical model corresponding to the physical platform is established, and the encoding function of the parameter to be measured in the quantum mechanical model is calibrated.

3. The method according to claim 1, characterized in that, The construction of a digital twin corresponding to the intelligent quantum sensor based on the fundamental physical system model and using a trained representation network includes: Obtain the expected values ​​of the initial observables corresponding to the basic physical system model; The trained representation network is used to predict the expected value of the initial observable, and the predicted value is used as the digital twin of the intelligent quantum sensor.

4. The method according to claim 1, characterized in that, The representation network was trained using the following method: Obtain the training dataset; Using the training dataset, the initial representation network is used to make predictions, and the prediction results are obtained. Construct a first loss function corresponding to the prediction result; The first gradient information of the first loss function with respect to the initial parameters of the initial representation network is determined by the backpropagation algorithm; Based on the first gradient information and the first preset termination condition, the initial parameters of the initial representation network are updated using the gradient descent method, and the trained representation network is obtained.

5. The method according to claim 4, characterized in that, The acquisition of the training dataset includes: Initialize the basic physical system model; A random control pulse is applied to the basic physical system model, causing the basic physical system model to evolve to a random initial state; As the underlying physical system model continues to evolve from a random initial state to a final state, the expected values ​​of observable quantities are recorded based on continuous weak measurements, and these expected values ​​are used as the training dataset.

6. The method according to claim 5, characterized in that, The step of combining the training dataset, using the initial representation network to make predictions, and obtaining prediction results includes: Based on the expected values ​​of the initial observables and the continuous weak measurement signals in the training dataset, the prediction results corresponding to the expected values ​​of the initial observables are generated through the forward propagation process of the initial representation network.

7. A device for constructing an intelligent quantum sensor, characterized in that, include: The determination module is used to determine the underlying physical system model corresponding to the intelligent quantum sensor based on the parameters to be measured. A construction module is used to construct a digital twin corresponding to the intelligent quantum sensor based on the basic physical system model and using a trained representation network; An optimization module is used to determine a target control strategy based on the digital twin using a trained reinforcement learning network, and to optimize the basic physical system model based on the target control strategy to obtain the target physical system model corresponding to the intelligent quantum sensor. The second training module is used to combine the digital twin and use the initial reinforcement learning network to determine the target control action; Construct a second loss function based on the target control action; The second gradient information of the second loss function with respect to the initial parameters of the initial reinforcement learning network is determined by the backpropagation algorithm; Based on the second gradient information and the second preset termination condition, the initial parameters of the initial reinforcement learning network are updated using the gradient descent method, and a trained reinforcement learning network is obtained. The process of combining the digital twin and using an initial reinforcement learning network to determine the target control action includes: Extract real-time feature data from the digital twin; The real-time feature data is input into the initial reinforcement learning network, and the first Q value of each control action is calculated through multi-layer nonlinear transformation. Based on the first Q value, the first control action is determined using a greedy algorithm; A first reward function is constructed by combining the predicted value in the digital twin and the ideal target state corresponding to the first control action, and the target control action is determined based on the first reward function; The construction of the second loss function based on the target control action includes: Sample the target control action and the corresponding batch of transfer data from the experience pool; Based on the target control action and the same batch of transfer data, the corresponding second Q value and second reward function are calculated according to the forward propagation of the initial reinforcement learning network. The mean squared error is calculated based on the second Q value and the second reward function, and the second loss function is obtained.

Citation Information

Patent Citations

  • Multi-agent motion control method based on interpretable reinforcement learning

    CN118689094A

  • Nuclear power station digital twinborn model implementation method based on physical-data hybrid driving

    CN118709527A