Construction method and device of intelligent quantum sensor

By constructing the basic physical system model of the intelligent quantum sensor and utilizing the digital twins of the characterization network and reinforcement learning network to optimize the control strategy, the problem of insufficient measurement accuracy caused by noise interference was solved and higher measurement accuracy was achieved.

CN120832818AActive Publication Date: 2025-10-24SHANGHAI JIAOTONG UNIV +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510993138.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-24
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

Existing intelligent quantum sensors are open systems, which are susceptible to decoherence due to environmental noise interference, making it impossible to directly design sensing schemes and resulting in insufficient measurement accuracy.

Method used

The basic physical system model is determined based on the parameters to be measured. A digital twin is constructed using a trained representation network. The target control strategy is determined through a reinforcement learning network to optimize the basic physical system model. Representation networks and reinforcement learning networks are introduced to compensate for noise errors.

Benefits of technology

By adaptively compensating for noise errors in real time, the measurement accuracy of the intelligent quantum sensor is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832818A_ABST
    Figure CN120832818A_ABST
Patent Text Reader

Abstract

The invention provides a construction method and device of an intelligent quantum sensor, and can be applied to the technical field of quantum precision measurement, and the method comprises the steps: determining a basic physical system model corresponding to the intelligent quantum sensor based on a to-be-measured parameter; based on the basic physical system model, utilizing a trained representation network to construct a digital twinborn body corresponding to the intelligent quantum sensor; and determining a target control strategy by using a trained reinforcement learning network based on the digital twin, and optimizing the basic physical system model based on the target control strategy to obtain a target physical system model corresponding to the intelligent quantum sensor. Therefore, the representation network and the reinforcement learning network are introduced, and the random error of the system is deduced through digital twinning constructed by the representation network, so that the reinforcement learning network can adaptively compensate the error in real time, and the measurement precision of the intelligent quantum sensor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of quantum precision measurement, in particular to a construction method and device of an intelligent quantum sensor. BACKGROUND

[0002] A quantum sensor is a physical device designed according to the laws of quantum mechanics, which can achieve high-precision measurement through quantum effects. Its core technologies include quantum entanglement, atomic spin interference and other phenomena.

[0003] Without considering noise, using coherence and entanglement can achieve the Heisenberg limit of quantum sensors. However, existing quantum sensing systems are open systems, which means that the decoherence of the system is caused by the interference of the environment or noise. In addition, due to the unknown types and intensities of noise, the evolution process of the quantum sensor system is unknown, and the existing intelligent quantum sensor cannot directly design a sensing scheme, thereby causing the problem of insufficient measurement accuracy of the intelligent quantum sensor.

[0004] Therefore, how to improve the measurement accuracy of the intelligent quantum sensor is a problem that those skilled in the art urgently need to solve. SUMMARY

[0005] Based on the above problems, the present application provides a construction method and device of an intelligent quantum sensor, comprising:

[0006] Based on the to-be-measured parameters, a basic physical system model corresponding to the intelligent quantum sensor is determined;

[0007] Based on the basic physical system model, a digital twin corresponding to the intelligent quantum sensor is constructed using a trained representation network;

[0008] Based on the digital twin, a target control strategy is determined using a trained reinforcement learning network, and the basic physical system model is optimized based on the target control strategy to obtain a target physical system model corresponding to the intelligent quantum sensor.

[0009] Optionally, the basic physical system model corresponding to the intelligent quantum sensor is determined based on the to-be-measured parameters, comprising:

[0010] Based on the to-be-measured parameters, a physical platform corresponding to the intelligent quantum sensor is determined;

[0011] A quantum mechanics model corresponding to the physical platform is established, and an encoding function of the to-be-measured parameters in the quantum mechanics model is calibrated.

[0012] Optionally, the digital twin corresponding to the intelligent quantum sensor is constructed based on the basic physical system model using a trained representation network, comprising:

[0013] obtaining an initial expected value of an observable corresponding to the underlying physical system model;

[0014] predicting, by using the trained representation network, a predicted value corresponding to the initial expected value of the observable, and taking the predicted value as the digital twin corresponding to the smart quantum sensor.

[0015] Optionally, the representation network is trained by the following method:

[0016] obtaining a training data set;

[0017] combining the training data set, predicting by using an initial representation network, and obtaining a prediction result;

[0018] constructing a first loss function corresponding to the prediction result;

[0019] determining, by using a back propagation algorithm, first gradient information of the first loss function with respect to initial parameters of the initial representation network;

[0020] based on the first gradient information and a first preset termination condition, updating the initial parameters of the initial representation network by using a gradient descent method, and obtaining a trained representation network.

[0021] Optionally, the obtaining of the training data set comprises:

[0022] initializing the underlying physical system model;

[0023] applying random control pulses to the underlying physical system model, so that the underlying physical system model evolves to a random initial state;

[0024] during a process in which the underlying physical system model continues to evolve from the random initial state to a final state, recording expected values of observables based on a continuous weak measurement mode, and taking the expected values of the observables as the training data set.

[0025] Optionally, the combining of the training data set, the predicting by using the initial representation network, and the obtaining of the prediction result comprise:

[0026] based on the initial expected values of the observables in the training data set and continuous weak measurement signals, generating, by using a forward propagation process of the initial representation network, a prediction result corresponding to the initial expected values of the observables.

[0027] Optionally, the reinforcement learning network is trained by the following method:

[0028] combining the digital twin, determining a target control action by using an initial reinforcement learning network;

[0029] constructing a second loss function based on the target control action;

[0030] determining second gradient information of the second loss function on initial parameters of the initial reinforcement learning network through a back propagation algorithm;

[0031] updating the initial parameters of the initial reinforcement learning network by using a gradient descent method based on the second gradient information and a second preset termination condition, and obtaining a trained reinforcement learning network.

[0032] Optionally, the determining the target control action by using the initial reinforcement learning network in combination with the digital twin includes:

[0033] extracting real-time feature data from the digital twin;

[0034] inputting the real-time feature data into the initial reinforcement learning network, and calculating a first Q value of each control action through a plurality of layers of nonlinear changes;

[0035] determining a first control action based on a greedy algorithm in combination with the first Q value;

[0036] constructing a first reward function in combination with a predicted value in the digital twin and an ideal target state corresponding to the first control action, and determining a target control action based on the first reward function.

[0037] Optionally, the constructing the second loss function based on the target control action includes:

[0038] sampling the target control action and same-batch transfer data corresponding to the target control action from an experience pool;

[0039] calculating a corresponding second Q value and a second reward function according to the initial reinforcement learning network forward propagation based on the target control action and the same-batch transfer data;

[0040] performing mean square error calculation based on the second Q value and the second reward function, and obtaining a second loss function.

[0041] In a second aspect, an embodiment of the present application provides a construction device of an intelligent quantum sensor, including:

[0042] A determination module is configured to determine a basic physical system model corresponding to an intelligent quantum sensor based on a to-be-measured parameter.

[0043] A construction module is configured to construct a digital twin corresponding to the intelligent quantum sensor by using a trained representation network based on the basic physical system model.

[0044] An optimization module is configured to determine a target control strategy based on the digital twin and the trained reinforcement learning network, and optimize the base physical system model based on the target control strategy to obtain a target physical system model corresponding to the intelligent quantum sensor.

[0045] From the above technical solutions, compared with the prior art, the present application has the following advantages:

[0046] The construction method of the intelligent quantum sensor provided by the present application comprises the following steps: first, determining a base physical system model corresponding to the intelligent quantum sensor based on the to-be-measured parameter; then, constructing a digital twin corresponding to the intelligent quantum sensor based on the base physical system model and using a trained representation network; finally, determining a target control strategy based on the digital twin and using a trained reinforcement learning network, and optimizing the base physical system model based on the target control strategy to obtain a target physical system model corresponding to the intelligent quantum sensor. In this way, the representation network and the reinforcement learning network are introduced, the random errors of the system are inferred through the digital twin constructed by the representation network, and then the reinforcement learning network can compensate for these errors in real time and adaptively, thereby improving the measurement accuracy of the intelligent quantum sensor. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 A flowchart of a construction method of an intelligent quantum sensor provided by an embodiment of the present application;

[0048] Figure 2 A flowchart of a training data collection provided by an embodiment of the present application;

[0049] Figure 3 A flowchart of a representation network training provided by an embodiment of the present application;

[0050] Figure 4 A flowchart of a representation network inference provided by an embodiment of the present application;

[0051] Figure 5 A flowchart of a digital twin technology provided by an embodiment of the present application;

[0052] Figure 6 A flowchart of a reinforcement learning network training provided by an embodiment of the present application;

[0053] Figure 7 A flowchart of a reinforcement learning network inference provided by an embodiment of the present application;

[0054] Figure 8 A structural schematic diagram of a construction device of an intelligent quantum sensor provided by an embodiment of the present application. DETAILED DESCRIPTION

[0055] As described above, the existing intelligent quantum sensor has the problem of insufficient measurement accuracy. Specifically, since the existing quantum sensor system is necessarily an open system, there is interference from the environment or noise. Noise-induced decoherence can damage the sensing accuracy, and the unknown type and intensity of the noise can make the evolution process of the system unknown, making it impossible to directly design a sensing scheme, and thus leading to the problem of insufficient measurement accuracy of the existing intelligent quantum sensor.

[0056] To solve the above problems, the embodiments of the present application provide a method for constructing an intelligent quantum sensor, which comprises first determining a basic physical system model corresponding to the intelligent quantum sensor based on the to-be-measured parameter. Then, based on the basic physical system model, a digital twin corresponding to the intelligent quantum sensor is constructed using a trained representation network. Finally, based on the digital twin, a target control strategy is determined using a trained reinforcement learning network, and the basic physical system model is optimized based on the target control strategy to obtain a target physical system model corresponding to the intelligent quantum sensor.

[0057] In this way, the representation network and the reinforcement learning network are introduced, and the digital twin constructed by the representation network is used to infer the random errors of the system, so as to ensure that the reinforcement learning network can compensate for these errors in real time and adaptively, thereby improving the measurement accuracy of the intelligent quantum sensor.

[0058] It should be noted that the method and device for constructing an intelligent quantum sensor provided by the embodiments of the present application can be applied to the field of quantum precision measurement. The above is only an example and does not limit the application field of the method and device for constructing an intelligent quantum sensor provided by the embodiments of the present application.

[0059] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme of the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0060] Figure 1 The flowchart of the method for constructing an intelligent quantum sensor provided by the embodiments of the present application is shown in FIG. 1. In combination with FIG. 1, the method for constructing an intelligent quantum sensor provided by the embodiments of the present application can comprise: Figure 1

[0061] S101: determining a basic physical system model corresponding to the intelligent quantum sensor based on the to-be-measured parameter.

[0062] ​In practical applications, the quantum sensor system needs to be initialized first, that is, the corresponding (basic) physical system model of the intelligent quantum sensor is selected according to the nature of the to-be-measured parameter. The model can be mainly divided into two categories. One is a discrete variable system based on an atomic platform (such as a diamond NV color center, a cold atomic ensemble, or an ion trap, etc.), which is suitable for discrete quantity measurement such as magnetic field and electric field. The other is a continuous variable system based on an optical platform (such as an optomechanical system or a squeezed state light field, etc.), which is suitable for continuous change of physical quantity measurement such as displacement and temperature. Further, a quantum mechanical model corresponding to the physical system model is established, including the mathematical description of the system Hamiltonian, the dephasing mechanism, and the measurement scheme. It can be understood that the selection of the model also needs to consider the constraint conditions of the measurement environment, such as working temperature, vacuum degree, etc., to ensure that the sensor model matches the actual application scenario.

[0063] In addition, since the way to determine the corresponding basic physical system model of the intelligent quantum sensor is not the same, the embodiments of the present application can explain one possible determination method.

[0064] In one case, S101: determining the corresponding basic physical system model of the intelligent quantum sensor based on the to-be-measured parameter, which can specifically include:

[0065] Determining the physical platform corresponding to the intelligent quantum sensor based on the to-be-measured parameter;

[0066] Establishing a quantum mechanical model corresponding to the physical platform, and calibrating the encoding function of the to-be-measured parameter in the quantum mechanical model.

[0067] In practical applications, the selection of the basic physical system model is related to its response to the unknown parameter. It can be understood that the diversity of the to-be-measured parameter requires us to select a suitable quantum sensor according to the parameter characteristics. For example, a diamond NV color center is suitable for DC-microwave magnetic field, a Rydberg atom is suitable for GHz-THz electric field, and a superconducting quantum bit is suitable for weak magnetic field detection; mechanical parameter measurement can select an optomechanical system (nanometer displacement), a spin-mechanical coupling system (micron vibration), or a surface acoustic wave resonator (mass change); temperature sensing can use quantum dots (nanometer scale), rare earth ion crystals (wide temperature range), or nuclear magnetic resonance systems (biological imaging). In this way, after selecting the physical platform, the corresponding quantum mechanical model is established. Further, the to-be-measured parameter is determined by combining experimental calibration (changing the parameter under controllable conditions and measuring the system response) and theoretical modeling (establishing a coupling model based on quantum mechanics). Encoding function in the system Hamiltonian , wherein, is the storage coupling quantity, and the coupling parameter The calibration needs to be refined by curve fitting and error analysis (evaluation of uncertainty, noise influence and establishment of correction model). Further, in practical applications, the sensitivity, dynamic range and environmental adaptability may also need to be optimized, and these systematic parameter-platform matching and accurate calibration methods are the key to realizing high-precision quantum sensors.

[0068] S102: Based on the basic physical system model, a digital twin corresponding to the intelligent quantum sensor is constructed using the trained representation network.

[0069] In practical applications, the trained representation network can accurately represent the deduction process of the basic physical system model. The actual sampling value of the basic physical system model is input into the trained representation network, and the predicted value output by the representation network is The digital twin corresponding to the intelligent quantum sensor is established.

[0070] In addition, since the way of obtaining the digital twin is not the same, the embodiments of the present application can be described with respect to one possible way of obtaining.

[0071] In one case, S102: Based on the basic physical system model, a digital twin corresponding to the intelligent quantum sensor is constructed using the trained representation network, which can specifically include:

[0072] Obtain the expected value of the initial observable corresponding to the basic physical system model;

[0073] Use the trained representation network to predict the predicted value corresponding to the expected value of the initial observable, and use the predicted value as the digital twin corresponding to the intelligent quantum sensor.

[0074] In practical applications, a control pulse is given to the basic physical system model to evolve it to a random state, which is used as the initial state to be sampled. Further, the initial state is sampled to obtain the expected value of the initial observable corresponding to the basic physical system model The trained representation network is used to predict the final state of the basic physical system model, i.e. the predicted value corresponding to the expected value of the initial observable , and the predicted value is used as the digital twin corresponding to the intelligent quantum sensor.

[0075] S103: Based on the digital twin, a target control strategy is determined using the trained reinforcement learning network, and the basic physical system model is optimized based on the target control strategy to obtain a target physical system model corresponding to the intelligent quantum sensor.

[0076] In practical applications, the trained reinforcement learning network can accurately determine the optimal control strategy of the underlying physical system model according to the deduction process of the underlying physical system model, thereby being used to alleviate the impact of noise. In combination with the above trained representation network, the deduction process of the underlying physical system model can be accurately represented, and the digital twin as the output of the representation network has its corresponding actual feature data (including observable expected values ), and then determines the target control strategy (optimal control strategy) of the underlying physical system model based on the digital twin (trained reinforcement learning network). Then, the target control strategy is used to optimize the underlying physical system model to obtain a target physical system model that can alleviate the impact of noise and improve the measurement accuracy of the intelligent quantum sensor.

[0077] In addition, since the training methods of the representation network are not the same, the embodiments of the present application can be described in terms of one possible training method.

[0078] In one case, the representation network is trained by the following method:

[0079] S201: Obtain a training data set.

[0080] In practical applications, the training data set can be obtained from existing sample data or based on the underlying physical system model.

[0081] In addition, since the training data set is obtained in different ways, the embodiments of the present application can be described in terms of one possible training method.

[0082] In one case, S201: obtaining a training data set, specifically comprising:

[0083] Initialize the underlying physical system model;

[0084] Apply random control pulses to the underlying physical system model to evolve the underlying physical system model to a random initial state;

[0085] During the process of the underlying physical system model evolving from the random initial state to the final state, record the expected value of the observable based on the continuous weak measurement method, and take the observable expected value as the training data set.

[0086] Figure 2 A training data collection flowchart is provided for the embodiments of the present application. As shown in Figure 2 , first, the quantum sensor (underlying physical system model) is initialized, and the initial state is continued to be randomly sampled. Since the initial state of the quantum sensor is usually determined (such as the ground state ), so it is necessary to evolve it to a random state by applying random quantum gates or applying random (control) pulses as a random initial state to be sampled , ensuring that the data covers the possible evolution behavior of the sensor. Then, let the sensor experience a period of free / controlled random master equation evolution, eventually reaching the final state . Further, in order to reflect the influence of continuous measurement, the data recorded by continuous weak measurement during the evolution from the random initial state to the final state will be saved (this data contains the key characteristics of random evolution in the decoherence process) as part of the training data set, that is, the output final state, collect measurement data. In order to obtain the required expected value of the observable and the expected value of the observable of the final state , it is necessary to repeat the preparation process of the initial state and the final state multiple times and project the measurement, until the maximum data set is reached, the process is completed, and the training data set is obtained. It can be understood that if the maximum data set is not reached, the above process is repeated. In addition, the training data set can be constructed in the following way: taking as the input of the neural network, and the corresponding as the label. This design effectively utilizes the implicit random time information in the measurement data while maintaining computational efficiency. Among them, is the continuous weak measurement signal.

[0087] S202: Use the initial representation network to make predictions in combination with the training data set, and obtain a prediction result.

[0088] In practical applications, the representation network can predict the next state of the underlying physical system based on the current state of the underlying physical system. Therefore, taking the training data set as the input of the initial representation network, the corresponding prediction result can be obtained.

[0089] In addition, since the way to obtain the prediction result is not the same, the embodiments of the present application can describe one possible way to obtain the prediction result.

[0090] In one case, S202: Use the initial representation network to make predictions in combination with the training data set, and obtain a prediction result, which can specifically include:

[0091] Based on the expected value of the initial observable in the training data set and the continuous weak measurement signal, the expected value of the initial observable is generated by the forward propagation process of the initial representation network. The prediction result corresponding to the initial observable.

[0092] In practical applications, in combination with the above training data set, the initial representation network receives input data , and then generates a prediction result through the forward propagation process, that is, the prediction value in the training process. The prediction process can be represented as where the function represents the nonlinear mapping relationship of the neural network, denotes all the trainable parameters of the network (including the weight matrices and bias vectors of each layer).

[0093] S203: Construct a first loss function corresponding to the prediction result.

[0094] In practical applications, the prediction result, i.e., the predicted value is compared and evaluated with the true value measured by actual measurement, and a first loss function is constructed. The commonly used is the root mean square error (RMSE), and its mathematical expression is as follows:

[0095] ;

[0096] In the formula, denotes the number of training samples, and the index represents a specific training sample. In this way, the overall deviation between the predicted value and the true value is effectively reflected by the first loss function, and the subsequent back propagation algorithm is optimized to optimize the network parameters. It can be understood that the first loss function has good differentiability in mathematical properties, and in some specific application scenarios, other forms of loss functions (such as mean absolute error (MAE) or customized loss functions combined with quantum sensor system characteristics) can be considered to better adapt to the needs of specific problems.

[0097] S204: Determine the first gradient information of the first loss function on the initial parameters of the initial representation network by the back propagation algorithm.

[0098] In practical applications, the first loss function is first calculated by the back propagation algorithm to obtain the gradient information of the initial parameters (all the trainable parameters of the network) .

[0099] S205: Update the initial parameters of the initial representation network based on the first gradient information and a first preset termination condition using the gradient descent method, and obtain a trained representation network.

[0100] In practical applications, the initial parameters are updated using the gradient descent method, and the expression is: . Wherein, the learning rate This is used to control the update step size and convergence rate, while also combining adaptive optimization strategies (such as Adam), gradient clipping, momentum terms, and other techniques to improve the training effect of the representation network. Furthermore, when the network iteration meets the first preset termination condition (when the loss function converges (i.e., the change in loss value for several consecutive iterations is less than a preset threshold) or reaches the preset maximum number of iterations), the training process terminates, resulting in a trained representation network.

[0101] Combined with the above, Figure 3 A flow chart of network training is provided for the embodiment of the present application. Figure 3 As shown in the figure, after the representation network training starts, the neural network (representation network) is initialized first, and then batch data is randomly sampled, the loss function is calculated, and then the network parameters are updated by the gradient descent method. After the loss function converges or reaches the number limit, the training is ended to obtain the trained representation network; if the loss function does not converge or does not reach the number limit, continue to sample data and update the network parameters until convergence or the number limit is reached.

[0102] Furthermore, the training process of the network to be represented is terminated, and the predicted value of its model output is It will be established as the digital twin of the intelligent quantum sensor for subsequent quantum state prediction and parameter estimation tasks. Figure 4 A flow chart of network inference provided in the embodiment of the present application. Figure 4 As shown, the quantum sensor is first initialized and then allowed to undergo a period of free / controlled stochastic master equation evolution, ultimately reaching a final state. During this process, continuous weak measurements output the final state and are influenced by the measured data. This measurement data is then fed into the temporal network (a trained representation network), which then outputs predicted sensor feature values. This concludes the representation network inference, and a digital twin corresponding to the smart quantum sensor is constructed based on these predicted values.

[0103] In addition, since there are different ways to train reinforcement learning networks, the embodiments of the present application can illustrate a possible training method.

[0104] In one embodiment, the reinforcement learning network is trained by:

[0105] S301: Determine a target control action using an initial reinforcement learning network in combination with the digital twin.

[0106] In practical applications, the trained representation network can accurately represent the deduction process of the underlying physical system model. The trained reinforcement learning network is used to determine the optimal control action according to the observable expected value. Therefore, during the training process of the reinforcement learning network, the initial reinforcement learning network can be used to determine the target control action based on the digital twin constructed by the representation network according to the predicted value, and then the parameter update of the reinforcement learning network is performed.

[0107] In addition, since the ways of determining the target control action are different, the embodiments of the present application can be described in terms of one possible determination method.

[0108] In one case, S301: in combination with the digital twin, the initial reinforcement learning network is used to determine the target control action, which can specifically include:

[0109] Extracting real-time feature data from the digital twin;

[0110] Inputting the real-time feature data into the initial reinforcement learning network, and calculating the first Q value of each control action through multi-layer nonlinear transformation;

[0111] Determining the first control action based on the greedy algorithm in combination with the first Q value;

[0112] Combining the predicted value in the digital twin and the ideal target state corresponding to the first control action to construct a first reward function, and determining the target control action based on the first reward function.

[0113] In practical applications, first, all parameters (including weight matrix and bias vector, etc.) of the reinforcement learning network need to be randomly initialized through Gaussian distribution or uniform distribution to ensure that the network has sufficient exploration ability at the initial stage of training. Then, real-time feature data (including expected value of observable quantity ) is extracted from the digital twin of the intelligent quantum sensor, and these data are used as the input of the initial reinforcement learning network. The network calculates the (first) Q value of each possible control action through multi-layer nonlinear transformation, and evaluates its potential for long-term cumulative reward in the current state. Then, based on the greedy algorithm (greedy strategy), the final control action (i.e. the first control action) is selected from each control action. The specific process can be to randomly select a control action with a probability of ε to promote exploration, or to select a control action with a probability of The control action with the highest current Q value is selected to achieve optimal control. At this time, the selected first control action will be converted into a specific control pulse or quantum gate operation and applied to the real quantum sensor (which can be a basic physical system model) to drive it to evolve towards the target state. At the same time, a (first) reward function is introduced, based on which the effectiveness of the first control action is evaluated by the predicted value in the digital twin and the ideal target state, and the target control action is determined. The predicted value is represented by the network output, and the ideal target state is the expected value of the final state observable of the real quantum sensor. In addition, the design of the first reward function usually adopts key indicators in quantum information theory, such as quantum Fisher information or classical Fisher information, which are converted into scalar reward values, and the magnitude of the function value directly reflects the contribution of the control action to improving sensor performance. Further, the reward function can also include a regularization term to constrain the energy or smoothness of the control pulse, ensuring its realizability in physical experiments. In this way, the reward mechanism can effectively guide the reinforcement learning strategy to optimize in the direction of improving sensor sensitivity and accuracy.

[0114] In combination with the above, Figure 5 A flowchart of a digital twin technology is provided for embodiments of the present application. As shown in Figure 5 The real quantum sensor system is constructed, and the system is monitored by continuous weak measurement and the measurement data is stored. Then, the measurement data is used as training data to train the representation network to obtain the trained representation network. Based on the measurement data of the real quantum sensor system, the trained representation network can generate a predicted result through a forward propagation process, and a digital twin is constructed. The digital twin and the real value of the real quantum sensor system are mutually mapped. Further, the digital twin is input into the reinforcement learning network to generate an output (first control action), and then the first control action is converted into a control pulse and input into the real quantum sensor system. At this time, the quantum sensor state is updated, and the quantum sensor undergoes controlled dynamic evolution under the actual action of the first control action. The corresponding measurement data (real-time acquisition of continuous weak measurement signals) is again input into the representation network, and the representation network outputs a predicted value and updates the digital twin. The updated digital twin is again input into the reinforcement learning network to generate a second control action. In this way, a closed loop is formed, and the above process is iterated by introducing a first reward function to finally converge and obtain the final target control action.

[0115] S302: Construct a second loss function based on the target control action.

[0116] In practical applications, after the target control action is selected, a second loss function can be constructed based on it.

[0117] In addition, since the second loss function is constructed in different ways, the embodiments of the present application can be described in terms of one possible construction method.

[0118] In one case, S302: constructing a second loss function based on the target control action, which can specifically include:

[0119] Sampling the target control action and the same batch of transfer data corresponding to the target control action from the experience pool;

[0120] Based on the target control action and the same batch of transfer data, the corresponding second Q value and second reward function are calculated according to the forward propagation of the initial reinforcement learning network.

[0121] Based on the second Q value and the second reward function, the mean square error is calculated, and the second loss function is obtained.

[0122] In practical applications, the update of the initial reinforcement learning network parameters can use an experience replay-based temporal difference (TD) learning algorithm. Specifically, first, a batch of transition data (including state, action, reward function, and next state) is randomly sampled from the experience pool. If the selected action is the target control action, the state, reward, and next state corresponding to the target control action are used as the same batch of transfer data. Then, the (second) Q value calculated by the forward propagation of the initial reinforcement learning network and the (second) reward function are used to calculate the mean square error as the TD loss function (i.e., the second loss function). In other words, the current Q network and the target Q network can be used to calculate the Q values of the current state and the next state, respectively. Then, based on the Bellman equation, the TD target value is constructed. By comparing the mean square error of the predicted Q value and the TD target value, the second loss function is constructed to guide the parameter optimization of the initial reinforcement learning network.

[0123] S303: determining the second gradient information of the second loss function on the initial parameters of the initial reinforcement learning network through a backpropagation algorithm.

[0124] In practical applications, the second gradient information of the second loss function on the initial parameters of the initial reinforcement learning network can be calculated by a backpropagation algorithm, and then the gradient descent method (such as the Adam optimizer) is used to update the weights of the Q network. In this way, through this offline learning method, the data efficiency is improved, and the correlation between data is broken by random sampling, significantly improving the stability of the training.

[0125] S304: updating the initial parameters of the initial reinforcement learning network using the gradient descent method based on the second gradient information and the second preset termination condition, and obtaining a trained reinforcement learning network.

[0126] In practical applications, the initial parameters of the initial reinforcement learning network are updated using the gradient descent method. When the reinforcement learning network iteration meets the second preset termination condition (the change rate of the loss function is lower than the preset threshold (indicating that the policy has converged) or reaches the preset maximum number of training rounds), the training process is terminated, and a trained reinforcement learning network is obtained. In each iteration, the system will record key indicators such as average reward, TD error amplitude, and policy entropy, etc. for evaluating the training progress. It can be understood that if the termination condition is not met, a new training cycle is started, first based on the updated control action to generate a control pulse, then the evolution and state monitoring of the real quantum sensor system are performed, then the digital twin is updated through the representation network and the reward evaluation, and finally the reinforcement learning network is optimized through experience replay. This cyclic iteration process continues until the reinforcement learning network reaches the required control accuracy or exhausts the computing budget, and the final intelligent agent model (the target physical system model that determines the target control policy) can be directly deployed in the adaptive control task of the actual quantum sensor.

[0127] In combination with the above, Figure 6 A training flowchart of a reinforcement learning network is provided for the embodiments of the present application. As shown in Figure 6 , the reinforcement learning network, experience pool and intelligent quantum sensor state are first initialized. Then the current state features (i.e. real-time feature data) are obtained from the digital twin. Then the first control action is selected through the greedy strategy, the corresponding reward function is calculated, and the target control action is determined. Then the same batch of data corresponding to the selected action is sampled from the experience pool, the TD loss is calculated, the network parameters are iteratively updated, and the sensor is updated according to the updated network parameters and the target control action, and the corresponding real data is stored. Finally, it is determined whether the current iteration reaches the maximum time step, whether the network reaches the target accuracy or the training round, if so, the training is ended, and a trained reinforcement learning network is obtained, if not, the iteration training is continued according to the above steps.

[0128] Figure 7 A reinforcement learning network inference flowchart is provided for the embodiments of the present application. As shown in Figure 7 , the inference process of the reinforcement learning network first initializes the quantum sensor state, and then obtains the current state features from the digital twin. Then the control action is selected according to the saved model, and the sensor is updated according to the control action, and the corresponding real data is stored. Finally, it is determined whether the current iteration reaches the maximum time step, if so, the training is ended, and a trained reinforcement learning network is obtained, if not, the iteration training is continued according to the above steps.

[0129] In summary, the construction method of the intelligent quantum sensor provided in the application comprises the following steps: first, determining a basic physical system model corresponding to the intelligent quantum sensor based on a to-be-measured parameter; then, constructing a digital twin corresponding to the intelligent quantum sensor by using a trained representation network based on the basic physical system model; and finally, determining a target control strategy by using a trained reinforcement learning network based on the digital twin, and optimizing the basic physical system model based on the target control strategy to obtain a target physical system model corresponding to the intelligent quantum sensor. In this way, the representation network and the reinforcement learning network are introduced, the digital twin constructed by the representation network is used to infer the random errors of the system, and then the reinforcement learning network can compensate for these errors in real time and adaptively, thereby improving the measurement accuracy of the intelligent quantum sensor.

[0130] Figure 8 FIG. 1 shows a structural schematic diagram of a construction device of an intelligent quantum sensor provided in an embodiment of the application. As shown in the figure, the construction device 800 of the intelligent quantum sensor can comprise: Figure 8

[0131] a determination module 801 configured to determine a basic physical system model corresponding to the intelligent quantum sensor based on a to-be-measured parameter;

[0132] a construction module 802 configured to construct a digital twin corresponding to the intelligent quantum sensor by using a trained representation network based on the basic physical system model;

[0133] an optimization module 803 configured to determine a target control strategy by using a trained reinforcement learning network based on the digital twin, and optimize the basic physical system model based on the target control strategy to obtain a target physical system model corresponding to the intelligent quantum sensor.

[0134] As an implementation form, for how to determine the basic physical system model corresponding to the intelligent quantum sensor, the determination module 801 is specifically configured to:

[0135] determine a physical platform corresponding to the intelligent quantum sensor based on the to-be-measured parameter;

[0136] establish a quantum mechanics model corresponding to the physical platform, and calibrate an encoding function of the to-be-measured parameter in the quantum mechanics model.

[0137] As an implementation form, for how to construct the digital twin corresponding to the intelligent quantum sensor, the construction module 802 is specifically configured to:

[0138] obtain an expected value of an initial observable corresponding to the basic physical system model;

[0139] ​predict a prediction value corresponding to the expected value of the initial observable by using the trained representation network, and take the prediction value as a digital twin corresponding to the intelligent quantum sensor.

[0140] As an implementation, for how to train the representation network, the construction device 800 of the above intelligent quantum sensor further includes a first training module.

[0141] The first training module is configured to obtain a training data set.

[0142] In combination with the training data set, an initial representation network is used for prediction, and a prediction result is obtained.

[0143] A first loss function corresponding to the prediction result is constructed.

[0144] The first gradient information of the first loss function on the initial parameters of the initial representation network is determined by a back propagation algorithm.

[0145] Based on the first gradient information and a first preset termination condition, the initial parameters of the initial representation network are updated by using a gradient descent method, and a trained representation network is obtained.

[0146] The obtaining of the training data set includes:

[0147] The basic physical system model is initialized.

[0148] Random control pulses are applied to the basic physical system model, so that the basic physical system model evolves to a random initial state.

[0149] In the process of continuing evolution of the basic physical system model from the random initial state to the final state, the expected value of the observable is recorded based on the continuous weak measurement, and the expected value of the observable is taken as the training data set.

[0150] The prediction by using the initial representation network in combination with the training data set and the obtaining of the prediction result include:

[0151] Based on the expected value of the initial observable in the training data set and the continuous weak measurement signal, a prediction result corresponding to the expected value of the initial observable is generated through a forward propagation process of the initial representation network.

[0152] As an implementation, for how to train the representation network, the construction device 800 of the above intelligent quantum sensor further includes a second training module.

[0153] The second training module is configured to determine a target control action by using an initial reinforcement learning network in combination with the digital twin.

[0154] constructing a second loss function based on the target control action;

[0155] determining second gradient information of the second loss function on initial parameters of the initial reinforcement learning network through a back propagation algorithm;

[0156] updating the initial parameters of the initial reinforcement learning network based on the second gradient information and a second preset termination condition by using a gradient descent method, and obtaining a trained reinforcement learning network.

[0157] wherein, the initial reinforcement learning network is used to determine the target control action in combination with the digital twin, including:

[0158] extracting real-time feature data from the digital twin;

[0159] inputting the real-time feature data into the initial reinforcement learning network, and calculating a first Q value of each control action through a plurality of layers of nonlinear changes;

[0160] determining a first control action based on a greedy algorithm in combination with the first Q value;

[0161] constructing a first reward function in combination with a predicted value in the digital twin and an ideal target state corresponding to the first control action, and determining the target control action based on the first reward function.

[0162] the second loss function is constructed based on the target control action, including:

[0163] sampling the target control action and same batch transfer data corresponding to the target control action from an experience pool;

[0164] calculating a corresponding second Q value and a second reward function according to the initial reinforcement learning network forward propagation based on the target control action and the same batch transfer data;

[0165] performing mean square error calculation based on the second Q value and the second reward function, and obtaining a second loss function.

[0166] To sum up, the construction method of the intelligent quantum sensor provided in the application comprises the following steps: firstly, determining a basic physical system model corresponding to the intelligent quantum sensor based on a to-be-measured parameter; then, constructing a digital twin corresponding to the intelligent quantum sensor by using a trained representation network based on the basic physical system model; and finally, determining a target control strategy by using a trained reinforcement learning network based on the digital twin, and optimizing the basic physical system model based on the target control strategy to obtain a target physical system model corresponding to the intelligent quantum sensor. In this way, the representation network and the reinforcement learning network are introduced, the random errors of the system are inferred through the digital twin constructed by the representation network, and then it is ensured that the reinforcement learning network can compensate for these errors in real time and adaptively, thereby improving the measurement accuracy of the intelligent quantum sensor.

[0167] The above description of disclosed embodiments enables one of ordinary skill in the art to make or use the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method of constructing an intelligent quantum sensor, characterized by, The method comprises: determining a basic physical system model corresponding to the intelligent quantum sensor based on a to-be-measured parameter; constructing a digital twin corresponding to the intelligent quantum sensor based on the basic physical system model and using a trained representation network; determining a target control strategy based on the digital twin and optimizing the basic physical system model based on the target control strategy to obtain a target physical system model corresponding to the intelligent quantum sensor.

2. The method of claim 1, wherein, The determination of the basic physical system model corresponding to the intelligent quantum sensor based on the to-be-measured parameter comprises: determining a physical platform corresponding to the intelligent quantum sensor based on the to-be-measured parameter; establishing a quantum mechanics model corresponding to the physical platform and calibrating an encoding function of the to-be-measured parameter in the quantum mechanics model.

3. The method of claim 1, wherein, The construction of the digital twin corresponding to the intelligent quantum sensor based on the basic physical system model and using the trained representation network comprises: obtaining an expected value of an initial observable corresponding to the basic physical system model; predicting a predicted value corresponding to the expected value of the initial observable by using the trained representation network, and taking the predicted value as the digital twin corresponding to the intelligent quantum sensor.

4. The method of claim 1, wherein, The representation network is trained by the following method: obtaining a training data set; predicting by using an initial representation network in combination with the training data set and obtaining a prediction result; constructing a first loss function corresponding to the prediction result; determining first gradient information of the first loss function on initial parameters of the initial representation network by a back propagation algorithm; updating the initial parameters of the initial representation network by a gradient descent method based on the first gradient information and a first preset termination condition, and obtaining a trained representation network.

5. The method of claim 4, wherein, The obtaining of the training data set comprises: initializing the basic physical system model; applying random control pulses to the basic physical system model to evolve the basic physical system model to a random initial state; recording expected values of observables based on a continuous weak measurement mode in a process in which the basic physical system model continues to evolve from the random initial state to a final state, and taking the expected values of the observables as the training data set.

6. The method of claim 5, wherein, The prediction by using the initial representation network in combination with the training data set and the obtaining of the prediction result comprise: generating a prediction result corresponding to the expected value of the initial observable by a forward propagation process of the initial representation network based on the expected value of the initial observable and a continuous weak measurement signal in the training data set.

7. The method of claim 1, wherein, The reinforcement learning network is trained by the following method: determining a target control action by using an initial reinforcement learning network in combination with the digital twin; constructing a second loss function based on the target control action; determining second gradient information of the second loss function on initial parameters of the initial reinforcement learning network by a back propagation algorithm; updating the initial parameters of the initial reinforcement learning network by a gradient descent method based on the second gradient information and a second preset termination condition, and obtaining a trained reinforcement learning network.

8. The method of claim 7, wherein, The determination of the target control action by using the initial reinforcement learning network in combination with the digital twin comprises: extracting real-time feature data from the digital twin; inputting the real-time feature data into an initial reinforcement learning network, and calculating a first Q value of each control action through multi-layer nonlinear transformation; determining a first control action based on a greedy algorithm in combination with the first Q value; constructing a first reward function based on a predicted value in the digital twin and an ideal target state corresponding to the first control action, and determining a target control action based on the first reward function.

9. The method of claim 7, wherein, the second loss function is constructed based on the target control action, including: sampling the target control action and same-batch transfer data corresponding to the target control action from an experience pool; calculating a corresponding second Q value and second reward function according to the initial reinforcement learning network forward propagation based on the target control action and the same-batch transfer data; performing mean square error calculation based on the second Q value and the second reward function, and obtaining a second loss function.

10. A construction device of an intelligent quantum sensor, characterized in that, including: a determination module configured to determine a basic physical system model corresponding to an intelligent quantum sensor based on a to-be-measured parameter; a construction module configured to construct a digital twin corresponding to the intelligent quantum sensor based on the basic physical system model and using a trained representation network; an optimization module configured to determine a target control strategy based on the digital twin and using a trained reinforcement learning network, and to optimize the basic physical system model based on the target control strategy to obtain a target physical system model corresponding to the intelligent quantum sensor.

Citation Information

Patent Citations

  • Sparse reinforcement learning-based sensor network optimization method

    CN103702349A

  • Multi-agent motion control method based on interpretable reinforcement learning

    CN118689094A

  • Nuclear power station digital twinborn model implementation method based on physical-data hybrid driving

    CN118709527A

  • Power transmission line forest fire risk assessment and early warning system and method

    CN119811051A

  • Intelligent inspection path planning method and system based on reinforcement learning

    CN119990496A