Process tensor tomography model training method and apparatus, and process tensor tomography method and apparatus
By introducing a deep learning model into quantum tomography, considering non-Markovian properties and time-dependent characteristics, the problem of insufficient accuracy in reconstructing open quantum systems in existing technologies is solved, achieving higher reconstruction accuracy and reduced dependence on measurement data.
Patent Information
- Application Number
- PCT/CN2025/078865
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-18
- Filing Date
- 2025-02-24
- Publication Date
- 2025-12-26
AI Technical Summary
Existing quantum tomography techniques have limitations in reconstructing open quantum systems, especially considering non-Markovian properties and time-dependent characteristics, resulting in insufficient reconstruction accuracy.
A deep learning model is introduced and trained using the measurement probabilities at multiple time points and their corresponding true values of the process tensor. A loss function is constructed to update the parameters of the process tensor tomography model, taking into account non-Markovian characteristics and time correlation.
It improves the accuracy of process tensor tomography, enabling better reconstruction of the dynamic processes of open quantum systems and reducing dependence on the original measurement data, especially showing better results in high-dimensional systems.
Smart Images

Figure CN2025078865_26122025_PF_FP_ABST
Abstract
Description
Process tensor tomography model training and process tensor tomography methods and apparatus Technical Field
[0001] This disclosure relates to the field of quantum information technology, specifically to the fields of quantum computing, deep learning, etc., and particularly to a process tensor tomography model training and process tensor tomography method and apparatus. Background Technology
[0002] Quantum tomography is a technique used to measure and reconstruct quantum systems, including quantum state tomography (QST) and quantum process tomography (QPT). QST reconstructs the state of a quantum system based on measurement data, typically represented as obtaining the density matrix ρ of the quantum states. QPT reconstructs the process of a quantum system based on measurement data, typically represented as obtaining process information ε.
[0003] Tensor networks are mathematical tools for representing high-dimensional data and are commonly used in quantum information and machine learning. When tensor networks are introduced into quantum tensor tomography (QPT), the resulting QPT can be called Process Tensor Tomography (PTT). PTT obtains process tensors based on measurement data.
[0004] For PTT, the problem of how to obtain the process tensor needs to be solved. Summary of the Invention
[0005] This disclosure provides a process tensor tomography model training method, apparatus, equipment, and medium for process tensor tomography.
[0006] According to one aspect of this disclosure, a method for training a process tensor tomography model is provided, comprising: obtaining measurement probabilities and the true values of process tensors corresponding to the measurement probabilities from a pre-constructed dataset; processing the measurement probabilities using a process tensor tomography model to obtain predicted values of the process tensors; constructing a loss function based on the true values of the process tensors and the predicted values of the process tensors; and updating the parameters of the process tensor tomography model based on the loss function; wherein the dataset includes measurement probabilities at multiple time points and the corresponding true values of the process tensors; for the current time point among the multiple time points, the measurement probability at the current time point is obtained by: performing a current operation on the input state of the quantum system at the current time point to obtain the output state of the quantum system at the current time point, and measuring the output state to obtain the measurement probability at the current time point; wherein, if the current time point is not the initial time point, the input state of the quantum system at the current time point is obtained by mapping the output state of the quantum system at the previous time point and the environmental state at the previous time point; and the true value of the process tensor at the current time point is obtained based on the measurement probability at the current time point.
[0007] According to another aspect of this disclosure, a process tensor tomography method is provided, comprising: obtaining measurement probabilities of a quantum system to be reconstructed; processing the measurement probabilities using a pre-trained process tensor tomography model to obtain a process tensor of the quantum system to be reconstructed; wherein the process tensor tomography model is trained using any of the training methods described in any of the foregoing aspects.
[0008] According to another aspect of this disclosure, a process tensor tomography model training apparatus is provided, comprising: an acquisition module for acquiring measurement probabilities and the true values of process tensors corresponding to the measurement probabilities from a pre-constructed dataset; a prediction module for processing the measurement probabilities using the process tensor tomography model to obtain predicted values of the process tensors; a construction module for constructing a loss function based on the true values of the process tensors and the predicted values of the process tensors; and an update module for updating the parameters of the process tensor tomography model based on the loss function; wherein the dataset includes measurement probabilities at multiple time points and the corresponding true values of the process tensors; and for... The measurement probability of the current time point among multiple time points is obtained as follows: The current operation at the current time point is used to operate on the input state of the quantum system at the current time point to obtain the output state of the quantum system at the current time point, and the output state is measured to obtain the measurement probability of the current time point; wherein, if the current time point is not the initial time point, the input state of the quantum system at the current time point is obtained by mapping the output state of the quantum system at the previous time point and the environmental state at the previous time point; the true value of the process tensor at the current time point is obtained based on the measurement probability of the current time point.
[0009] According to another aspect of this disclosure, a process tensor tomography apparatus is provided, comprising: an acquisition module for acquiring measurement probabilities of a quantum system to be reconstructed; and a processing module for processing the measurement probabilities using a pre-trained process tensor tomography model to obtain a process tensor of the quantum system to be reconstructed; wherein the process tensor tomography model is trained using any of the training methods described in any of the preceding aspects.
[0010] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to said at least one processor; wherein the memory stores instructions executable by said at least one processor, said instructions being executed by said at least one processor to enable said at least one processor to perform the method as described in any of the foregoing aspects.
[0011] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method according to any of the preceding aspects.
[0012] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to any of the preceding aspects.
[0013] This disclosure can improve the accuracy of process tensor chromatography.
[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0015] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0016] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0017] Figure 2 is a schematic diagram of the quantum circuit for process tensor tomography provided according to an embodiment of the present disclosure;
[0018] Figure 3 is a schematic diagram of the preparation process of the Choi matrix according to an embodiment of the present disclosure;
[0019] Figure 4 is a schematic diagram of the structure of the process tensor tomography model provided according to an embodiment of the present disclosure;
[0020] Figure 5 is a schematic diagram according to a second embodiment of the present disclosure;
[0021] Figure 6 is a schematic diagram according to a third embodiment of the present disclosure;
[0022] Figure 7 is a schematic diagram according to the fourth embodiment of the present disclosure;
[0023] Figure 8 is a schematic diagram according to the fifth embodiment of the present disclosure;
[0024] Figure 9 is a schematic diagram of an electronic device used to implement the process tensor tomography model training method or process tensor tomography method of the present disclosure embodiments. Detailed Implementation
[0025] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0026] To better understand the embodiments of this disclosure, the terms used in the embodiments of this disclosure are explained as follows:
[0027] CPTP: CP stands for Completely Positive Map, and TP stands for Trace-Preserving Map. CPTP is an important matrix mapping relationship in the field of quantum information and is often used to characterize the features of quantum channels.
[0028] Quantum state tomography: obtaining the density matrix ρ in quantum computing through multiple experimental measurements.
[0029] Quantum process tomography: reconstructing the quantum operation ε in quantum computing through multiple experimental measurements.
[0030] Markov processes: They lack memory features, and their conditional probabilities are only related to the current state of the system, and are independent of the historical and future states.
[0031] Choi isomorphism: a homomorphic mapping that describes the relationship between quantum channels. The Choi matrix involved in the mapping relationship is also called a Choi state or Choi representation.
[0032] Tensor networks: countable sets of tensors that are connected by contraction, often used in modern mathematics, computer science, quantum information science and other fields.
[0033] Process tensors are mathematical tools used to describe the dynamics of open quantum systems. In quantum mechanics, quantum systems often interact with their external environment, making their evolution more complex. Process tensors provide a framework for describing the dynamics of such open systems.
[0034] Neural networks: An adaptive learning system inspired by the human brain, composed of a large number of interconnected neurons. It consists of numerous interconnected processing units that learn the mapping relationship from input data to output results by adjusting weights and biases.
[0035] Quantum tomography is a method for measuring and reconstructing the state of a quantum system. This technique can be applied to a variety of quantum systems, and its purpose is to obtain enough information to completely determine the state of the quantum system through a series of precise measurement experiments.
[0036] Among related technologies, QSP and QPT perform well in reconstructing simplified or ideal quantum systems, but they have certain limitations when dealing with real quantum systems, especially open quantum systems that have complex interactions with their environment.
[0037] In realizing this disclosure, the inventors discovered that the current QSP or QPT has limitations in reconstructing quantum systems because it only obtains a state or process representation at a single point in time, while ignoring the non-Markovian properties that evolve over time and the correlations between different points in time.
[0038] Therefore, in this embodiment of the disclosure, for QPT using tensor networks, i.e. for PTT, when obtaining the process tensor, the aforementioned non-Markovian properties and time dependence will be considered.
[0039] Specifically, by introducing deep learning models into the process tensor reconstruction process, the deep learning models will contain non-Markovian properties and temporal correlation information. Thus, when using deep learning models to obtain process tensors, non-Markovian properties and temporal correlation will be taken into account.
[0040] To ensure that the deep learning model incorporates non-Markovian properties and temporal correlation information, the measurement probabilities at multiple time points and the corresponding true values of the process tensor are used as samples for model training. Since samples from different time points exhibit non-Markovian properties and temporal correlation, the model trained based on these samples will contain this information. Furthermore, the process tensor reconstructed based on this model takes into account these non-Markovian properties and temporal correlation, thus improving the accuracy of PTT (Programme for Time-To-Test).
[0041] In this embodiment of the disclosure, the input of the deep learning model is the measurement probability, and the output is the process tensor. Therefore, the deep learning model can be called the process tensor tomography model.
[0042] The following examples will describe the model training process and the model application process, respectively.
[0043] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure. This embodiment provides a method for training a process tensor tomography model, as shown in Figure 1, the method includes:
[0044] 101. Obtain the measurement probability and the true value of the process tensor corresponding to the measurement probability from the pre-constructed dataset.
[0045] 102. The measurement probability is processed using a process tensor tomography model to obtain the predicted value of the process tensor.
[0046] 103. Construct a loss function based on the true value of the process tensor and the predicted value of the process tensor.
[0047] 104. Update the parameters of the process tensor tomography model based on the loss function.
[0048] The dataset includes the measurement probabilities at multiple time points and the true values of their corresponding process tensors;
[0049] For the current time point among multiple time points, the measurement probability of the current time point is obtained in the following way: using the current operation at the current time point, the input state of the quantum system at the current time point is operated on to obtain the output state of the quantum system at the current time point, and the output state is measured to obtain the measurement probability of the current time point; wherein, if the current time point is not the initial time point, the input state of the quantum system at the current time point is obtained by mapping the output state of the quantum system at the previous time point and the environmental state at the previous time point;
[0050] The true value of the process tensor at the current time point is obtained based on the measurement probability at the current time point.
[0051] Suppose a pre-constructed dataset is denoted by D, which includes multiple training samples, represented as D = {{x1,z1},{x2,z2},...{x...z1}. i ,z i},...,{x N ,z N}}, where {x i ,z i Let} represent the i-th training sample group, where i = 1, 2, ..., N, and N is the number of samples.
[0052] When training a process tensor tomography model (or simply model), a certain number of training samples are selected from the dataset based on a preset proportion (e.g., 75%).
[0053] Any set of training samples used for model training is denoted by {x, z}. During training, x is input into the model, and the output is the predicted value, denoted by z'.
[0054] Then, a loss function is constructed based on the true value z and the predicted value z'.
[0055] After constructing the loss function, the model parameters are updated based on the loss function. The updated parameters include the weight coefficient w and the bias coefficient b.
[0056] The model training process has been briefly described above. In this embodiment, x is the measurement probability, z is the true value of the process tensor, and z' is the predicted value of the process tensor.
[0057] In this embodiment of the disclosure, the measured probability and the true value of its corresponding process tensor are derived from multiple time points.
[0058] Specifically, any continuous-time quantum random process can be discretized into several time points: T k ={t0,t1,…,t k-1}, in order to capture each t i ∈T k Statistical data can be obtained by altering the control operations that determine the information completeness of a quantum process. Each control operation changes the state of the quantum process, influencing its evolutionary trajectory over time and mapping it to the output. Through these control operations, the initial state of the quantum system is mapped to a new output state; this mapping relationship is a key issue in quantum process tomography. By analyzing the data at different time points t... i By applying different control operations to the system and observing the results, it is possible to infer the relationship with T. k The process tensor tomography method identifies all possible quantum state trajectories in a consistent manner. Once this process is completed, all possible joint statistics can be constructed, and the output of the quantum process under any general sequence can be predicted.
[0059] The quantum evolution process at different points in time can be seen in Figure 2.
[0060] Figure 2 is a schematic diagram of the quantum circuit for process tensor tomography provided according to an embodiment of the present disclosure, where S represents the quantum system (which may be simply referred to as the system), i.e., the part of interest whose dynamics are to be understood and controlled. E represents the environment, which interacts with the system but is usually not directly controlled and has an impact on the system state. This represents the sequence of operations from time 0 to k-1, which is a series of operations performed by the experimenter on the system. This represents the initial state of the system. This is a unitary mapping on SE space. The dashed lines distinguish between operations that the experimenter can control and external factors of the system that they cannot control, such as environmental influences. The experimenter associates these operations with the system's state at a later time by performing a sequence of operations on the system. This framework describes the system's dynamics by distinguishing between controllable and uncontrollable objects (represented by dashed lines). The process tensor is used to comprehensively describe quantum random processes, mapping the set of control operations to the system's output states. The process tensor contains all the information about the system-environment initial state and interactions, which can be inferred solely from the system's dynamics, providing a new tool for understanding and representing non-Markovian dynamics.
[0061] Based on the quantum circuit described above, the measurement probabilities at multiple time points and the true values of their corresponding process tensors can be collected. Then, the collected data is stored in the dataset as samples for subsequent model training.
[0062] Multiple time points using T k ={t0,t1,…,t k-1The time series is defined as follows: any one of the multiple time points is called the current time point, denoted as ti (i = 0, 1, ..., k). The length k of the time series is an artificially set value, usually determined according to experimental requirements and the dynamic characteristics of the system, with the aim of fully capturing the characteristics of the quantum system's evolution over time.
[0063] The measurement probability at the current time point and the true value of its corresponding process tensor are obtained in the following way:
[0064] Measurement probability at the current time point: Using the current operation at the current time point, the input state of the quantum system at the current time point is operated on to obtain the output state of the quantum system at the current time point, and the output state is measured to obtain the measurement probability at the current time point; wherein, if the current time point is the initial time point, the input state of the quantum system at the current time point is a preset initial state, or, if the current time point is not the initial time point, the input state of the quantum system at the current time point is obtained by mapping the output state of the quantum system at the previous time point and the environmental state at the previous time point.
[0065] Specifically, let's assume the current time point is t. i This indicates that the current operation uses... This is represented by t0, where t0 represents the initial time point, and the others are non-initial time points. Regarding the initial time point t0: the input state of the quantum system at that initial time point is a preset initial state, i.e. For the current time point which is not the initial time point t i For example, i≠0, taking the current time point t1, the input state of the quantum system at t1 is obtained by mapping the output state of the quantum system at t0 and the environmental state at t0. This mapping is done using a unitary matrix. express.
[0066] Taking the current time point t1 as an example, the measurement probability p1 of t1 is obtained after measuring the output state of the quantum system at t1. The output state of the quantum system at t1 is obtained by using the current operation. This is obtained by manipulating the input state of the quantum system t1.
[0067] The true value of the process tensor at the current time point is obtained based on the measurement probability at the current time point. The specific calculation formula can be:
[0068] Where, p j It is the measurement probability at the current time point j.
[0069] Δ jis the j-th dual basis vector of the Orthogonal Operator-Valued Measure (POVM) operator; L is the number of dual basis vectors of the POVM operator;
[0070] ω i It is the i-th dual basis vector of the quantum density of states matrix; the superscript T denotes the transpose operation; n is the number of dual basis vectors of the quantum density of states matrix.
[0071] The two dual basis vectors mentioned above are constructed from the relevant measurement operators or density matrices using linear algebra methods. They are mathematically complete, ensuring that they are complete and orthogonal in the defined space.
[0072] The quantum state density matrix described above is the density matrix of the quantum system's input states (or simply input states) at the current time point. At the initial time point t0, the input states are the preset initial states, represented as follows: For non-initial time point t i The input state is obtained by mapping the output state (or simply output state) and the environment state (or simply environment state) of the quantum system at the previous time point. For example, at time point t1, the input state is the result of unitary mapping between the output state and the environment state at t0.
[0073] In process tensor tomography, temporal information is reflected through the measurement probabilities and operation sequences at different time points. The process tensor changes over time, reflecting the state evolution of the quantum system at different time points. By changing the operation sequences at different time points and recording the measurement results, the temporal evolution trajectory of the entire quantum process, i.e., the process tensor, can be reconstructed.
[0074] The current operation described above is obtained by arranging the pre-acquired operation base.
[0075] For example, if the pre-acquired operation base is N, then after permutation, there are N! (N factorial) possible arrangements. Each arrangement corresponds to one current operation, thus yielding N! current operations. At each time point t... i The above N! types of current operations can be executed. After each current operation, a set of measurement probabilities can be obtained, and the true value of the process tensor can be obtained based on the measurement probabilities. Then, the measurement probabilities and their corresponding true values of the process tensors are stored in the dataset for subsequent model training.
[0076] In this way, by arranging the operation basis to obtain the current operation, a large number of current operations can be obtained. Based on multiple current operations, multiple operations can be performed, thereby obtaining more measurement probabilities and the true values of their corresponding process tensors, thus enriching the number of training samples and improving the model performance.
[0077] It is understandable that other methods can also be used to obtain the current operation based on the operation bases. For example, one or more of the N operation bases can be selected as the current operation, or the selected partial operation bases can be arranged differently to obtain the current operation, etc.
[0078] The operation bases can be represented by the measurement bases |i><j|. The measurement bases are one of the forms of the above-mentioned POVM operators, which are a set of POVMs that can completely describe the quantum state and are known.
[0079] During the evolution process at different time points, the output state is obtained by mapping the input state and the environmental state. The mapping relationship corresponding to the mapping can be represented by the Choi matrix.
[0080] Before quantum tomography, it is first necessary to prepare the Choi matrix, that is, to model the quantum channel or quantum operation as a CP (Completely Positive Map, CP) map. This mapping relationship ensures that when mapping a quantum state to a quantum state, it satisfies the properties of preserving positivity and trace non-increase, ensuring that no non-physical results are generated during the quantum operation process, and also maintaining these properties when coupled with any auxiliary system. Using the Choi isomorphism (Choi-Jamiolkowski isomorphism, CJI), a CP map can be given by the positive matrix representation of a quantum state.
[0081] Figure 3 is a schematic diagram of the preparation process of the Choi matrix provided according to an embodiment of the present disclosure.
[0082] As shown in Figure 3, the Choi matrix can be prepared in the following manner:
[0083] 301. Create a quantum channel.
[0084] 302. Construct a maximally entangled state.
[0085] 303. Based on the quantum channel and the maximally entangled state, construct a composite system.
[0086] 304. Calculate the Choi matrix for characterizing the composite system.
[0087] Among them, creating a quantum channel: The quantum channel is represented by ε, which describes the mapping from one quantum state to another quantum state. Specifically, it can be realized by constructing a specific quantum operation, such as using quantum gates to define this channel.
[0088] Constructing a maximally entangled state: Introduce a maximally entangled state |Φ > in the Hilbert space + >, which is usually a Bell state and can be expressed as: where d is the dimension of the Hilbert space, and |ii> represents in The ground state in; H A and H B denote two Hilbert spaces, and their tensor product represents the Hilbert space of a composite system composed of two subsystems.
[0089] Construct the composite system: Apply the quantum channel ε to one of the subsystems and the identity operation I to the other subsystem, obtaining the following expression for the composite system:
[0090] where ε is the created quantum channel, and |Φ + > is the constructed maximally entangled state.
[0091] Calculate the Choi matrix, and the calculation formula is: For each set of measurement bases |i><j|, calculate the result of the quantum channel ε acting on the projection operator |i><j|, and then multiply it by |i><j|. Sum over all terms to obtain the Choi matrix. Here, the measurement bases are one of the forms of the above POVM operators, which is a set of POVMs that can completely describe the quantum state and is known, that is, the above operation basis, d is the number of measurement bases, and ε is the quantum channel.
[0092] In this way, the Choi matrix can be prepared. This Choi matrix is used to characterize the composite system and can reflect the information of an open quantum system that has an interaction relationship with the environment. Furthermore, based on this Choi matrix, the time-evolution process of the composite system can be reflected to obtain samples at different time points of the time series, including the measurement probabilities and their corresponding true values of the process tensors.
[0093] The above describes the sample acquisition process. After obtaining the samples, the samples will be used to train the model.
[0094] FIG. 4 is a schematic structural diagram of a process tensor tomography model according to an embodiment of the present disclosure.
[0095] As shown in FIG. 4, the process tensor tomography model is a deep learning network model, including an input layer, a hidden layer, and an output layer.
[0096] The model input is the measurement probability, and the model output is the process tensor (during the training process, this output is called the predicted value of the process tensor; when the model is applied, this output is used as the process tensor of the quantum system to be reconstructed that needs to be obtained finally).
[0097] The measurement probability is the probability corresponding to the measurement results under a series of quantum states, and its quantity is denoted by n; the number of model inputs and outputs is the same, so the number of process tensors is also denoted by n.
[0098] During the training process, the process of obtaining the predicted value of the process tensor is as follows:
[0099] The input layer is used to process the measured probability to obtain the output of the input layer.
[0100] The hidden layer is used to process the output of the input layer to obtain the output of the hidden layer. The hidden layer processing includes activation processing of the hidden layer, and the activation function used in the activation processing of the hidden layer is the ReLU function.
[0101] The output layer is used to process the output of the hidden layer to obtain the predicted value of the process tensor. The output layer processing includes the activation processing of the output layer, and the activation function used in the activation processing of the output layer is the tangent function.
[0102] Specifically, the process tensor tomography model is a deep learning network model that consists of multiple layers. For the k-th layer of the model, the input-output relationship is represented as: z (k) =f(y (k) ); y (k) =W (k) z (k-1) +b (k) ;
[0103] Among them, z (k-1) It is the input of the k-th layer of the model, that is, the output of the (k-1)-th layer; the input of the first layer of the model, that is, the initial input, is the measurement probability.
[0104] W (k) b is the weight coefficient of the k-th layer of the model. (k) These are the bias coefficients of the k-th layer of the model;
[0105] y (k) It is the output of the linear transformation of the k-th layer of the model;
[0106] f() is the activation function;
[0107] z (k) It is the output of the k-th layer of the model.
[0108] For the hidden layer, the activation function is the ReLU function, expressed as f(x) = max(0,x);
[0109] For the output layer, the activation function is the tangent function, expressed as f(x) = tanh(x).
[0110] The input layer here does not contain an activation function; it only indicates that the neural network receives and transmits data.
[0111] Using the above formula, the model output z' can be obtained based on the model input x, where x is the measurement probability and z' is the predicted value of the process tensor.
[0112] Next, a loss function, Loss, is constructed based on the true and predicted values of the process tensor. In this embodiment, the mean squared error function is used, and the calculation formula is as follows:
[0113] <·> represents the average value between the predicted and actual values of the neural network, and the weight and bias coefficients are continuously optimized based on the results of each training round; m is the number of samples in the current training batch; r k It is the true value of the k-th sample, that is, the true value of the process tensor of the k-th sample; It is the predicted value of the k-th sample, that is, the predicted value of the process tensor of the k-th sample; These are the model parameters, including the weight coefficients w and the bias coefficients b; f is the activation function.
[0114] In this embodiment, the Adam algorithm (Adaptive Moment Estimation, Adam) is used to update the model parameters based on the loss function. This method has advantages such as simple implementation, high computational efficiency, and low memory consumption. The parameter update process is as follows: m t =β1m t-1 +(1-β1)g t
[0115] in, These are the updated model parameters. These are the model parameters before the update;
[0116] It is the change in parameters;
[0117] m t and v t Let these represent the first and second integrals of the gradient, respectively; and They represent m respectively t and v t The unbiased estimate; α represents the learning rate; β1 and β2 represent the decay rates of the exponentially weighted average, respectively; ε is a small non-negative quantity used to ensure that the denominator is non-zero; g t The gradient represents the partial derivative of the loss with respect to the parameters. When training a neural network using the Adam algorithm, a parameter set {α, β1, β2, ε} can be set, for example, choosing a set of parameter values {0.001, 0.9, 0.999, 10...}. -7}
[0118] Before training begins, parameter initialization is necessary. Incorrectly initialized weights can lead to vanishing or exploding gradients, negatively impacting the training process. Since ReLU is chosen as the activation function, the He initialization method, suitable for ReLU networks, is selected. The specific method for He initialization involves using weights with a mean of 0 and a standard deviation of... Weights are randomly selected from a normal distribution, where n is the number of neurons in the previous layer. This initialization of weights helps to better propagate gradients during training. Its characteristics are: during forward propagation, the variance of the state values remains constant; during backward propagation, the variance of the gradient with respect to the activation values remains constant.
[0119] The above-mentioned parameter tuning using the Adam algorithm and parameter initialization using the He initialization algorithm can be understood to mean that other parameter tuning algorithms can also be used, such as stochastic gradient descent (SGD), root mean square propagation (RMSProp), adaptive gradient (AdaGrad), accelerated gradient descent (Nesterov), etc.; parameter initialization can also use Xavier initialization, uniform random initialization, etc.
[0120] The model training process is a continuous process of updating the model parameters. The model parameters are continuously updated based on the above formula until the preset termination condition is reached. The model that reaches the termination condition is taken as the final trained model.
[0121] Termination conditions can be set according to actual needs, such as reaching a preset number of iterations or model convergence.
[0122] In some embodiments, the termination conditions include: reaching a preset number of iterations and the model performance meeting preset conditions.
[0123] Based on the above termination conditions, loss function, activation function, and parameter initialization method, this disclosure also provides a model training method, as shown in Figure 5.
[0124] Figure 5 is a schematic diagram according to a second embodiment of the present disclosure, which provides a process tensor model training method, the method comprising:
[0125] 501. Initialization.
[0126] Initialization can include: initializing model parameters, setting the maximum number of iterations, and setting the conditions that the model performance needs to meet.
[0127] For model parameter initialization, the He initialization method can be used.
[0128] The maximum number of iterations can be set as needed; for example, maximum number of iterations = 100.
[0129] Model performance and the conditions it needs to meet can also be set according to actual needs. For example, model performance can be selected as the mean squared error of the probability of the predicted value. That is, for each sample, not only is the predicted value of the process tensor obtained, but also the predicted probability corresponding to that predicted value is obtained. The difference between the predicted probability and the true probability of each sample is calculated, and the mean of the sum of squares of the differences is used as the performance index. When the performance index is less than or equal to a preset threshold, it indicates that the model performance meets the preset conditions.
[0130] 502. Obtain the measurement probability of the current iteration process and the true value of the process tensor corresponding to the measurement probability from the pre-constructed dataset.
[0131] 503. Input the measurement probability into the initialized process tensor tomography model, and the output is the predicted value of the process tensor.
[0132] In the activation process, the activation function of the hidden layer can be the ReLU function, and the activation function of the output layer can be the tangent function.
[0133] 504. Based on the true value and the predicted value of the process tensor, construct the mean squared error function as the loss function.
[0134] 505. Based on the loss function, the Adam algorithm is used to update the parameters of the process tensor tomography model.
[0135] 506. Determine if the preset maximum number of iterations has been reached. If yes, proceed to 507; otherwise, repeat 502 and its subsequent steps.
[0136] 507. Determine whether the model performance meets the preset conditions. If yes, proceed to 508; otherwise, repeat 502 and its subsequent steps.
[0137] 508. Use the current process tensor tomography model as the final process tensor tomography model.
[0138] The above describes the model training process. After training, the final process tensor tomography model is obtained, which can be used to perform process tensor tomography on the quantum system to be reconstructed.
[0139] Figure 6 is a schematic diagram according to a third embodiment of the present disclosure, which provides a process tensor tomography method, the method comprising:
[0140] 601. Obtain the measurement probability of the quantum system to be reconstructed.
[0141] 602. The measurement probability is processed using a pre-trained process tensor tomography model to obtain the process tensor of the quantum system to be reconstructed.
[0142] The process tensor tomography model is trained using the model training method of any of the above embodiments.
[0143] Specifically, after preparing the known input state, the input state is input into the quantum system to be reconstructed, the output state is measured, the measurement probability of each measurement result is obtained, and the measurement probability p is then input into the pre-trained process tensor tomography model. The model output is the process tensor ε of the quantum system to be reconstructed.
[0144] In this embodiment, the process tensor of the quantum system is reconstructed through a process tensor tomography model. Compared to methods such as linear inversion, this approach considers the non-Markovian properties that evolve over time and the correlations between different time points, improving the accuracy of process tensor tomography. Tensor process operators constructed based on the Choi matrix greatly aid in the tomographic construction of process information, allowing for a more complete acquisition of the non-Markovian properties during the quantum system's evolution. Deep reinforcement learning methods based on neural networks can learn complex mappings and achieve good results even in multi-time-node and noisy open quantum environments. Neural networks possess strong nonlinear fitting capabilities, enabling them to learn and represent complex quantum processes, handling systems deeply interrelated with their environment in quantum process tensor tomography tasks. Traditional quantum tomography methods are highly dependent on measurement sample data and require numerous repetitive measurement processes. Especially as the system dimension increases, the measurement complexity grows exponentially, making quantum tomography of large systems particularly difficult. Process tensor tomography significantly reduces the dependence on original measurement data, achieving better process tomography results even with low sample measurement requirements. Quantum processes in open systems may be affected by randomness, and neural networks can better capture this randomness by learning the statistical properties of data. This is very helpful for robustness against noise and handling the uncertainty of experimental data. By establishing a neural network model of the tensor matrix from input measurement results to output process, an end-to-end learning process can be achieved, simplifying the modeling process of complex systems. It is more concise and efficient for handling high-dimensional systems and can automatically extract system features and learn complex mapping relationships.
[0145] Figure 7 is a schematic diagram according to a third embodiment of the present disclosure, which provides a process tensor tomography model training device 700. The device includes: an acquisition module 701, a prediction module 702, a construction module 703, and an update module 704.
[0146] The acquisition module 701 is used to acquire the measurement probability and the true value of the process tensor corresponding to the measurement probability from a pre-constructed dataset; the prediction module 702 is used to process the measurement probability using a process tensor tomography model to obtain the predicted value of the process tensor; the construction module 703 is used to construct a loss function based on the true value of the process tensor and the predicted value of the process tensor; and the update module 704 is used to update the parameters of the process tensor tomography model based on the loss function.
[0147] The dataset includes measurement probabilities at multiple time points and the true values of their corresponding process tensors. For the current time point, the measurement probability is obtained as follows: The current operation at the current time point is used to operate on the input state of the quantum system at the current time point to obtain the output state of the quantum system at the current time point, and the output state is measured to obtain the measurement probability at the current time point. If the current time point is not the initial time point, the input state of the quantum system at the current time point is obtained by mapping the output state of the quantum system at the previous time point and the environmental state at the previous time point. The true value of the process tensor at the current time point is obtained based on the measurement probability at the current time point.
[0148] In some embodiments, the current operation is obtained by arranging a pre-acquired operation base.
[0149] In some embodiments, the mapping is performed based on a Choi matrix, and the apparatus further includes a preparation module for:
[0150] Create a quantum channel;
[0151] Construct the maximum entangled state;
[0152] Based on the quantum channel and the maximally entangled state, a composite system is constructed;
[0153] Calculate the Choi matrix used to characterize the composite system.
[0154] In some embodiments, the process tensor tomography model includes an input layer, a hidden layer, and an output layer; the prediction module is further used for:
[0155] The input layer is used to process the measured probability to obtain the output of the input layer.
[0156] The hidden layer is used to process the output of the input layer to obtain the output of the hidden layer. The hidden layer processing includes activation processing of the hidden layer, and the activation function used in the activation processing of the hidden layer is the ReLU function.
[0157] The output layer is used to process the output of the hidden layer to obtain the predicted value of the process tensor. The output layer processing includes the activation processing of the output layer, and the activation function used in the activation processing of the output layer is the tangent function.
[0158] In some embodiments, the construction module is further configured to: construct a mean squared error function based on the true value of the process tensor and the predicted value of the process tensor, and use the mean squared error function as the loss function; and / or,
[0159] The update module is further configured to: update the parameters of the process tensor tomography model based on the loss function using the Adam algorithm; and / or,
[0160] The initial values of the parameters of the process tensor tomography model are obtained using the He initialization method.
[0161] Figure 8 is a schematic diagram according to a fourth embodiment of the present disclosure, which provides a process tensor tomography apparatus 800. The apparatus includes an acquisition module 801 and a processing module 802.
[0162] The acquisition module 801 is used to acquire the measurement probability of the quantum system to be reconstructed; the processing module 802 is used to process the measurement probability using a pre-trained process tensor tomography model to obtain the process tensor of the quantum system to be reconstructed; wherein the process tensor tomography model is trained using the model training method of any of the above embodiments.
[0163] It is understood that the same or similar content in different embodiments of this disclosure can be referred to each other.
[0164] It is understood that the terms "first" and "second" in the embodiments of this disclosure are only used for distinction and do not indicate the degree of importance or the order of events.
[0165] It is understandable that, unless otherwise specified, the order of steps in the process indicates that the temporal relationship between these steps is not limited.
[0166] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0167] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0168] Figure 9 illustrates a schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure. The electronic device 900 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0169] As shown in Figure 9, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 902 or a computer program loaded into a random access memory (RAM) 903 from a storage unit 908. The RAM 903 can also store various programs and data required for the operation of the electronic device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0170] Multiple components in electronic device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of displays, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows electronic device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0171] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the process tensor tomography model training method or the process tensor tomography method. For example, in some embodiments, the process tensor tomography model training method or the process tensor tomography method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the process tensor tomography model training method or the process tensor tomography method described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured by any other suitable means (e.g., by means of firmware) to perform a process tensor tomography model training method or a process tensor tomography method.
[0172] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0173] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0174] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0175] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0176] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0177] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0178] It should be understood that the various forms of processes shown above can be used to reorder, add, or cancel steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0179] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for training a process tensor tomography model, characterized in that, include: Obtain the measurement probabilities and the true values of the process tensors corresponding to the measurement probabilities from a pre-constructed dataset; The measurement probability is processed using a process tensor tomography model to obtain the predicted value of the process tensor; A loss function is constructed based on the true value and the predicted value of the process tensor. The parameters of the process tensor tomography model are updated based on the loss function; The dataset includes the measurement probabilities at multiple time points and the true values of their corresponding process tensors; For the current time point among multiple time points, the measurement probability of the current time point is obtained in the following way: using the current operation at the current time point, the input state of the quantum system at the current time point is operated on to obtain the output state of the quantum system at the current time point, and the output state is measured to obtain the measurement probability of the current time point; wherein, if the current time point is not the initial time point, the input state of the quantum system at the current time point is obtained by mapping the output state of the quantum system at the previous time point and the environmental state at the previous time point; The true value of the process tensor at the current time point is obtained based on the measurement probability at the current time point.
2. The method according to claim 1, characterized in that, The current operation is obtained by arranging the pre-acquired operation base.
3. The method according to claim 1, characterized in that, The mapping is performed based on the Choi matrix, which is prepared as follows: Create a quantum channel; Construct the maximum entangled state; Based on the quantum channel and the maximally entangled state, a composite system is constructed; Calculate the Choi matrix used to characterize the composite system.
4. The method according to claim 1, characterized in that, The process tensor tomography model includes: an input layer, a hidden layer, and an output layer; The process tensor tomography model is used to process the measurement probability to obtain the predicted value of the process tensor, including: The input layer is used to process the measured probability to obtain the output of the input layer. The hidden layer is used to process the output of the input layer to obtain the output of the hidden layer. The hidden layer processing includes activation processing of the hidden layer, and the activation function used in the activation processing of the hidden layer is the ReLU function. The output layer is used to process the output of the hidden layer to obtain the predicted value of the process tensor. The output layer processing includes the activation processing of the output layer, and the activation function used in the activation processing of the output layer is the tangent function.
5. The method according to claim 1, characterized in that, The step of constructing a loss function based on the true value and the predicted value of the process tensor includes: constructing a mean squared error function based on the true value and the predicted value of the process tensor, and using the mean squared error function as the loss function; And / or, The step of updating the parameters of the process tensor tomography model based on the loss function includes: using the Adam algorithm to update the parameters of the process tensor tomography model based on the loss function; And / or, The initial values of the parameters of the process tensor tomography model are obtained using the He initialization method.
6. A process tensor tomography method, characterized in that, include: Obtain the measurement probability of the quantum system to be reconstructed; A pre-trained process tensor tomography model is used to process the measurement probabilities to obtain the process tensor of the quantum system to be reconstructed. The process tensor tomography model is trained using the method described in any one of claims 1-5.
7. A training device for a process tensor tomography model, characterized in that, include: The acquisition module is used to acquire the measurement probability and the true value of the process tensor corresponding to the measurement probability from a pre-constructed dataset; The prediction module is used to process the measurement probability using a process tensor tomography model to obtain the predicted value of the process tensor. The construction module is used to construct a loss function based on the true value and the predicted value of the process tensor; An update module is used to update the parameters of the process tensor tomography model based on the loss function; The dataset includes the measurement probabilities at multiple time points and the true values of their corresponding process tensors; For the current time point among multiple time points, the measurement probability of the current time point is obtained in the following way: using the current operation at the current time point, the input state of the quantum system at the current time point is operated on to obtain the output state of the quantum system at the current time point, and the output state is measured to obtain the measurement probability of the current time point; wherein, if the current time point is not the initial time point, the input state of the quantum system at the current time point is obtained by mapping the output state of the quantum system at the previous time point and the environmental state at the previous time point; The true value of the process tensor at the current time point is obtained based on the measurement probability at the current time point.
8. The apparatus according to claim 7, characterized in that, The current operation is obtained by arranging the pre-acquired operation base.
9. The apparatus according to claim 7, characterized in that, The mapping is performed based on the Choi matrix, and the apparatus further includes a preparation module, which is used for: Create a quantum channel; Construct the maximum entangled state; Based on the quantum channel and the maximally entangled state, a composite system is constructed; Calculate the Choi matrix used to characterize the composite system.
10. The apparatus according to claim 7, characterized in that, The process tensor tomography model includes: an input layer, a hidden layer, and an output layer; The prediction module is further used for: The input layer is used to process the measured probability to obtain the output of the input layer. The hidden layer is used to process the output of the input layer to obtain the output of the hidden layer. The hidden layer processing includes activation processing of the hidden layer, and the activation function used in the activation processing of the hidden layer is the ReLU function. The output layer is used to process the output of the hidden layer to obtain the predicted value of the process tensor. The output layer processing includes the activation processing of the output layer, and the activation function used in the activation processing of the output layer is the tangent function.
11. The apparatus according to claim 7, characterized in that, The construction module is further configured to: construct a mean squared error function based on the true value and the predicted value of the process tensor, and use the mean squared error function as the loss function; and / or, The update module is further configured to: update the parameters of the process tensor tomography model based on the loss function using the Adam algorithm; and / or, The initial values of the parameters of the process tensor tomography model are obtained using the He initialization method.
12. A process tensor chromatography apparatus, characterized in that, include: The acquisition module is used to acquire the measurement probability of the quantum system to be reconstructed. The processing module is used to process the measurement probability using a pre-trained process tensor tomography model to obtain the process tensor of the quantum system to be reconstructed. The process tensor tomography model is trained using the method described in any one of claims 1-5.
13. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Quantum noise elimination method and device in quantum operation, electronic equipment and medium
CN115310618A
Process tensor chromatography model training method and device and process tensor chromatography method and device
CN118350473A
Cost function deformation in quantum approximate optimization
US20190164079A1
System and method of quantum enhanced accelerated neural network training
US20210342730A1
Data processing method, machine learning framework system and related device
WO2023125858A1