Multi-dimensional data adaptive decision optimization method and system based on quantum reinforcement learning
Through quantum wavelet multi-scale decomposition and dynamic qubit allocation, combined with synchronization gate and annealing regularization terms to optimize quantum line parameters, the problems of low resource utilization and insufficient adaptability of traditional quantum enhancement learning in multi-dimensional data processing are solved, and efficient and real-time multi-dimensional data decision optimization is achieved.
Patent Information
- Application Number
- CN202510539598.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-27
AI Technical Summary
When traditional quantum enhancement learning methods process multi-dimensional data, static encoding and static resource allocation strategies cannot adapt to the heterogeneity and dynamics of multi-dimensional data, resulting in low quantum resource utilization, difficulty in achieving effective extraction of multi-scale features, and lack dynamic adaptation to hardware resource constraints, making it difficult to meet the high-precision and real-time decision-making requirements in complex scenarios.
By performing quantum wavelet multi-scale decomposition of multi-dimensional input data, dynamically allocate the number of qubits, encode the feature components into quantum state eigenvectors, and insert synchronization gates into the block variable component quantum circuit, and modify the quantum natural gradient in combination with annealing regularization terms, generate the optimized quantum state eigenvectors, and output the optimization decision strategy.
It significantly improves the accuracy and efficiency of multi-dimensional data decision optimization, enhances the stability and dynamic adaptability of strategies, and can achieve high-precision and real-time decision-making in complex environments. It is suitable for finance, medical care and intelligent transportation and other fields.
Smart Images

Figure CN120449095A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and in particular relates to a multidimensional data adaptive decision optimization method and system based on quantum enhanced learning. Background Art
[0002] With the convergence of quantum computing and machine learning technologies, quantum-enhanced learning (QEL) has emerged as an emerging approach for solving complex decision-making optimization problems. By combining the high parallelism of quantum computing with the policy optimization of classical reinforcement learning, this technique significantly improves the efficiency of high-dimensional data processing and real-time decision-making. Traditional QEL methods typically employ fixed quantum circuit architectures and static resource allocation strategies, such as predefined parameterized models based on quantum neural networks or classification decision mechanisms based on quantum support vector machines, to handle classification and regression tasks for structured data or time-series signals. However, when processing multidimensional data, these techniques' static encoding and resource allocation strategies are unable to adapt to the heterogeneity and dynamic nature of multidimensional data, resulting in low quantum resource utilization. Furthermore, this technique relies on fixed quantum circuit topologies, such as fully connected entangled structures, and lacks dynamic adaptation to hardware resource constraints, making it difficult to effectively extract multi-scale features with a limited number of qubits. Furthermore, in the face of dynamic environmental changes, this technique struggles to achieve coordinated optimization of features and policies through real-time feedback, resulting in insufficient decision accuracy and optimization efficiency, making it unable to meet the demands of high-precision, real-time decision-making in complex scenarios. Summary of the Invention
[0003] Based on this, it is necessary to provide a multi-dimensional data adaptive decision optimization method and system based on quantum reinforcement learning to address the above technical problems, so as to improve the quantum resource utilization efficiency and decision-making accuracy of multi-dimensional data processing, and enhance the strategy stability and dynamic adaptability in complex scenarios.
[0004] In a first aspect, the present application provides a multi-dimensional data adaptive decision optimization method based on quantum enhancement learning, comprising:
[0005] Performing quantum wavelet multiscale decomposition on multidimensional input data, dynamically allocating the number of quantum bits based on the variance of each dimension, and encoding the decomposed characteristic components into quantum state characteristic vectors, wherein the multidimensional input data includes at least one of structured table data, time series sensor data, and multimodal image data;
[0006] The quantum state eigenvector is input into a quantum strategy network constructed by block variational quantum circuits. A synchronization gate is inserted into the block variational quantum circuit to output the probability distribution of strategy actions. The operating parameters of the synchronization gate are dynamically adjusted according to the clock signal deviation between the quantum processor and the classical controller.
[0007] Calculate the error between the probability distribution of the policy action and the target value function, use the annealing regularization term to correct the quantum natural gradient, optimize the parameters of the block variational quantum circuit, and generate updated parameters;
[0008] Execute the adaptive decision action corresponding to the probability distribution of the strategy action, obtain the reward signal from the environment by executing the adaptive decision action, and dynamically trigger the quantum feature recoding according to the cumulative error between the updated parameters and the reward signal, generate the optimized quantum state feature vector, and output the optimized decision strategy.
[0009] In one embodiment, multidimensional input data is subjected to quantum wavelet multi-scale decomposition, the number of quantum bits is dynamically allocated based on the variance of each dimension, and the decomposed characteristic components are encoded into quantum state characteristic vectors, including:
[0010] Perform quantum Haar wavelet transform on multi-dimensional input data to generate multi-resolution feature components;
[0011] Calculate the variance of each dimension of the multidimensional input data, and dynamically allocate the number of qubits based on the variance using the following formula to obtain the number of qubits after dynamic allocation:
[0012]
[0013] Among them, B j is the number of quantum bits after dynamic allocation, is the variance of the jth dimension of the multidimensional input data, Q is the total number of available quantum bits in the system, is the variance of the i-th dimension of the multidimensional input data;
[0014] If the total number of quantum bits after dynamic allocation exceeds the total number of available quantum bits in the system, the number of quantum bits after dynamic allocation is proportionally compressed to obtain the compressed number of quantum bits;
[0015] Based on the compressed number of quantum bits, the multi-resolution feature components are encoded into quantum state feature vectors.
[0016] In one embodiment, the calculation formula of the quantum state eigenvector is:
[0017]
[0018] Among them, |ψ(w i )> represents the encoded quantum state feature vector, which is used to describe the quantum state of the i-th sample after encoding, B j ′ is the number of quantum bits after compression, R y is a rotation gate around the y-axis, used to apply a rotation operation around the y-axis to the quantum bit, w ij To represent the multi-resolution feature component w kThe j-th dimension data of the i-th sample, w i is the multi-resolution feature component w k The i-th sample in .
[0019] In one embodiment, a quantum state feature vector is input into a quantum strategy network constructed by a block variational quantum circuit, a synchronization gate is inserted into the block variational quantum circuit, and a strategy action probability distribution is output, including:
[0020] Input the quantum state feature vector into the initial quantum register of the block variational quantum circuit as the input state of the quantum strategy network;
[0021] According to the dimensional distribution of the quantum state eigenvector, the block variational quantum circuit is split into multiple sub-circuits using the following formula:
[0022] M=D / Q block
[0023] Where M is the total number of subcircuits, D is the dimensional distribution of the quantum state eigenvector, Q block is the maximum quantum bit capacity of each subcircuit, and each subcircuit contains a parameterized quantum gate, which is composed of a combination of rotation gates around the Z axis and the X axis;
[0024] Insert CZ entanglement gates between adjacent sub-circuits and construct block variational quantum circuits using the following formula:
[0025]
[0026] Among them, U(θ) represents the overall operation of the block variational quantum circuit, which consists of multiple sub-operations, U m (θ m ) is the parameterized quantum gate operation of the mth subcircuit, θ m is the parameter set of the subcircuit, and Entangle(m,m+1) is the CZ entanglement gate inserted between adjacent subcircuits m and m+1.
[0027] According to the clock signal deviation between the clock signal of the quantum processor and the clock signal of the classical controller, the time parameter of the synchronization gate is calculated, and the synchronization gate is constructed by the following formula:
[0028]
[0029] in is the Hamiltonian of the synchronous gate, is the Z Pauli operator of the i-th quantum bit;
[0030] A synchronization gate is inserted into the block variational quantum circuit, and the clock signal deviation is calibrated to be less than the preset deviation threshold through quantum state tomography technology, and the probability distribution of the strategy action is output.
[0031] In one embodiment, the error between the probability distribution of the policy action and the target value function is calculated, the quantum natural gradient is corrected using an annealing regularization term, and the parameters of the block variational quantum circuit are optimized to generate updated parameters, including:
[0032] The mean square error between the policy action probability distribution and the target value function is calculated using the following formula:
[0033] L(θ)=E s,a [(Q(s,a)-V(s)) 2 ]
[0034] Among them, L(θ) is the mean square error loss function, E s,a is the mathematical expectation of state s and action a, Q(s,a) is the action value function obtained by quantum phase estimation, which represents the expected return of performing action a in state s, and V(s) is the target value function, which represents the long-term value of state s;
[0035] According to the gradient of the probability distribution of the strategy action, the quantum Fisher information matrix is calculated, and the pseudo-inverse of the quantum Fisher information matrix is calculated to obtain the pseudo-inverse matrix;
[0036] generating a dynamic annealing coefficient according to the chaos sensitivity of the parameters of the block variational quantum circuit, wherein the chaos sensitivity is quantified by a Lyapunov exponent, and the annealing coefficient is generated based on the Lyapunov exponent;
[0037] Based on the pseudo-inverse matrix and the annealing coefficient, the annealing regularization term is used to correct the quantum natural gradient using the following formula to obtain the corrected quantum natural gradient:
[0038]
[0039] Among them, F -1 (θ) is the pseudo-inverse matrix, is the gradient of the original loss function L(θ) with respect to the parameter θ, indicating the direction of loss decrease, and λ(t) is the annealing coefficient;
[0040] Based on the modified quantum natural gradient, multi-task gradient conflicts are suppressed through quantum interference path selection, the parameters of the block variational quantum circuit are updated, and the updated parameters are generated.
[0041] In one embodiment, an adaptive decision action corresponding to the probability distribution of the strategic action is executed, a reward signal is obtained from the environment by executing the adaptive decision action, and quantum feature recoding is dynamically triggered based on the cumulative error between the updated parameters and the reward signal to generate an optimized quantum state feature vector and output an optimized decision strategy, including:
[0042] According to the probability distribution of the strategic action, the adaptive decision action is selected through the Softmax decision function. The decision probability in the Softmax decision function is determined by the probability distribution of the strategic action after being smoothed by the temperature coefficient.
[0043] Perform adaptive decision actions, obtain reward signals from the environment, and calculate the cumulative error of the reward signal using the following formula:
[0044]
[0045] in, is the predicted reward value based on the parameters of the block variational quantum circuit at time t, r t is the reward signal at time t;
[0046] If the cumulative error is greater than a preset threshold, the updated parameters are loaded into the block variational quantum circuit, triggering quantum feature recoding;
[0047] The encoding weights are adjusted according to the gradient direction of the reward prediction loss function. The reward prediction loss is the mean square error between the predicted reward value and the corresponding reward signal, and the optimized quantum state feature vector is generated by the following formula:
[0048]
[0049] Among them, X new is the optimized quantum state eigenvector, X is the original quantum state eigenvector before optimization, ⊙ represents the Hadamard product, Encoding weight gradient, used to reflect the reward prediction loss L r The direction and rate of change relative to the coding weight X;
[0050] Based on the optimized quantum state eigenvector and updated parameters, the optimization decision strategy is output.
[0051] In one embodiment, a synchronization gate is inserted into the block variational quantum circuit, and the clock signal deviation is calibrated to be less than a preset deviation threshold through quantum state tomography technology, and the strategy action probability distribution is output, where the preset deviation threshold is 0.1ns, including:
[0052] The quantum state after the synchronization gate is measured using quantum state tomography technology, and the fidelity between the quantum state and the ideal quantum state is calculated using the following formula:
[0053]
[0054] Where Tr is the trace operation, which is used to sum the diagonal elements of the matrix, ρ ideal is the density matrix of the ideal quantum state, ρ out is the density matrix of the quantum state after the synchronization gate;
[0055] If the fidelity is less than 0.99, adjust the timing parameters of the synchronization gate using the following formula:
[0056] t←t·(1+0.1·(1-F))
[0057] Repeatedly adjust the timing parameters of the synchronization gate until the fidelity is greater than or equal to 0.99 and the clock signal deviation is less than 0.1ns. The corresponding calibrated synchronization gate parameters are obtained and the probability distribution of the strategy action is output.
[0058] In a second aspect, the present application also provides a multi-dimensional data adaptive decision optimization system based on quantum enhanced learning, comprising:
[0059] A quantum feature encoding module is used to perform quantum wavelet multiscale decomposition on multidimensional input data, dynamically allocate the number of quantum bits based on the variance of each dimension, and encode the decomposed feature components into quantum state feature vectors. The multidimensional input data includes at least one of structured table data, time series sensor data, and multimodal image data.
[0060] The quantum strategy network module is used to input the quantum state feature vector into the quantum strategy network constructed by the block variational quantum circuit, insert the synchronization gate into the block variational quantum circuit, and output the probability distribution of the strategy action. The operating parameters of the synchronization gate are dynamically adjusted according to the clock signal deviation between the quantum processor and the classical controller;
[0061] The quantum gradient optimization module is used to calculate the error between the probability distribution of policy actions and the target value function, use the annealing regularization term to correct the quantum natural gradient, optimize the parameters of the block variational quantum circuit, and generate updated parameters;
[0062] The adaptive decision feedback module is used to execute adaptive decision actions corresponding to the probability distribution of strategy actions. By executing adaptive decision actions, it obtains reward signals from the environment and dynamically triggers quantum feature recoding based on the cumulative error between the updated parameters and the reward signal. It generates an optimized quantum state feature vector and outputs the optimized decision strategy.
[0063] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the first aspect when executing the computer program.
[0064] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the first aspect when executed by a processor.
[0065] The aforementioned quantum-enhanced learning-based multidimensional data adaptive decision optimization method and system effectively overcomes the limitations of traditional quantum-enhanced learning methods, such as low feature encoding efficiency and difficulty capturing complex data correlations, when processing high-dimensional, multimodal data by performing quantum wavelet multiscale decomposition on the multidimensional input data and dynamically allocating the number of qubits. The multidimensional input data, which includes various types such as structured tabular data, time-series sensor data, and multimodal image data, can reflect real-world scene information from multiple dimensions. Quantum encoding fully utilizes the superposition and entanglement properties of quantum states to map data features into quantum space, providing a rich representation of information for subsequent decision-making. Secondly, the quantum state feature vectors are input into a quantum strategy network constructed using block-variational quantum circuits. Dynamic parameter adjustment via synchronization gates leverages the advantages of quantum computing in parallel processing and rapid search, resulting in an optimized probability distribution for strategic actions. The synchronization gates adjust parameters based on clock signal deviations, effectively resolving synchronization issues when quantum processors and classical controllers work together, further ensuring the stability of the decision-making process.
[0066] Furthermore, this method utilizes an annealing regularization term to modify the quantum natural gradient to optimize the block variational quantum circuit parameters, overcoming the shortcomings of traditional gradient optimization methods, which are prone to falling into local optimality and low optimization efficiency in high-dimensional parameter spaces. By combining the quantum natural gradient with the annealing mechanism, it is possible to accelerate convergence while improving the adaptability of the policy network to complex environments. Finally, quantum feature recoding is dynamically triggered based on the reward signal fed back from the environment, enabling real-time perception of environmental changes and timely adjustment of the quantum state eigenvectors, thereby outputting an optimized decision-making strategy that is more suitable for the actual scenario, ensuring the timeliness and accuracy of the decision-making strategy.
[0067] Compared with traditional quantum enhanced learning methods, this method significantly improves the accuracy, efficiency and environmental adaptability of multi-dimensional data decision optimization through quantum feature encoding, quantum strategy network optimization and dynamic feedback adjustment, providing strong technical support for complex decision-making scenarios in finance, medical care, intelligent transportation and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0069] Figure 1 A flowchart of a multi-dimensional data adaptive decision optimization method based on quantum enhancement learning provided by an exemplary embodiment of the present invention;
[0070] Figure 2 A schematic diagram of the structure of a multi-dimensional data adaptive decision optimization system based on quantum reinforcement learning provided as an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0071] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0072] In one embodiment, Figure 1 As shown, a multi-dimensional data adaptive decision optimization method based on quantum enhanced learning is provided. This embodiment uses the method applied to a terminal as an example. It is understandable that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0073] S101: Perform quantum wavelet multi-scale decomposition on multi-dimensional input data, dynamically allocate the number of quantum bits based on the variance of each dimension, and encode the decomposed characteristic components into quantum state characteristic vectors, where the multi-dimensional input data includes at least one of structured table data, time series sensor data, and multimodal image data.
[0074] Specifically, structured table data can be transaction record tables related to the financial sector, containing multiple columns of structured data such as customer information, transaction amount, and transaction time. Time series sensor data can be temperature, humidity, and air pressure data recorded chronologically by various sensors in a weather station. Multimodal image data can be CT images, MRI images, and other medical imaging data, but this is not limited here. Quantum wavelet multiscale decomposition technology can then be used to decompose the different dimensions of the multidimensional input data according to different scales and frequencies, thereby extracting features of the data at different scales, facilitating more accurate subsequent analysis and processing.
[0075] Secondly, the number of qubits can be dynamically allocated based on the variance of each dimension. Variance reflects the degree of data dispersion in a particular dimension. A large variance indicates greater data fluctuations in that dimension and contains richer information. Therefore, more qubits can be allocated to dimensions with large variance, while fewer qubits can be allocated to dimensions with small variance. This allows for more detailed encoding of dimensions with large amounts of information within limited quantum resources, improving the overall data representation. In quantum computing, quantum states possess properties such as superposition and entanglement, enabling efficient storage and processing of information. Finally, the decomposed characteristic components are encoded as quantum state eigenvectors, which are converted into quantum state forms that can be processed by quantum computing, providing a solid foundation for subsequent quantum computing and decision-making.
[0076] S102: Input the quantum state feature vector into the quantum strategy network constructed by the block variational quantum circuit, insert a synchronization gate into the block variational quantum circuit, and output the strategy action probability distribution. The operating parameters of the synchronization gate are dynamically adjusted according to the clock signal deviation between the quantum processor and the classical controller.
[0077] Specifically, a block variational quantum circuit is a quantum circuit structure based on quantum gate operations. It can divide the entire quantum circuit into multiple subcircuits, each of which can be independently parameterized to better adapt to different problems and data. A synchronization gate is a quantum gate used to coordinate the operation between the quantum processor and the classical controller. In actual quantum computing systems, the clock signals of the quantum processor and the classical controller may deviate, affecting the accuracy and stability of quantum computing. Therefore, a synchronization gate is inserted into the block variational quantum circuit. The operating parameters of the synchronization gate can be dynamically adjusted based on the clock signal deviation between the quantum processor and the classical controller. Through real-time monitoring and feedback mechanisms, the synchronization of the quantum processor and the classical controller is ensured, thereby improving the reliability of quantum computing. After processing by the block variational quantum circuit and the synchronization gate, the quantum policy network can output a policy action probability distribution. This policy action probability distribution represents the probability of various possible decision actions occurring under a given quantum state eigenvector, providing a basis for subsequent decision action execution.
[0078] S103: Calculate the error between the probability distribution of the strategic action and the target value function, use the annealing regularization term to correct the quantum natural gradient, optimize the parameters of the block variational quantum circuit, and generate updated parameters.
[0079] Specifically, the error between the strategy action probability distribution obtained in step S102 and the target value function is calculated. The target value function is a predefined function that represents the optimal decision-making effect expected to be achieved in a specific decision-making scenario. The error reflects the gap between the decision result of the current quantum strategy network and the optimal decision result. Therefore, in order to narrow the error, the annealing regularization term can be used to correct the quantum natural gradient. The quantum natural gradient is a method used to optimize parameters in quantum computing. It takes into account the geometric structure of the quantum state and can guide the update of parameters more efficiently. The annealing regularization term is a regularization method that can avoid overfitting during the parameter update process by introducing a parameter-related penalty term, thereby further improving the generalization ability of the model. Based on the corrected quantum natural gradient, the parameters of the block variational quantum circuit are optimized to generate updated parameters, which can enable the quantum strategy network to output a better strategy action probability distribution in the next iteration.
[0080] S104: Execute the adaptive decision action corresponding to the probability distribution of the strategy action, obtain the reward signal from the environment by executing the adaptive decision action, and dynamically trigger the quantum feature recoding according to the cumulative error between the updated parameters and the reward signal, generate the optimized quantum state feature vector, and output the optimized decision strategy.
[0081] Specifically, the adaptive decision action corresponding to the strategy action probability distribution in step S102 is executed. This adaptive decision action is the action most likely to achieve the optimal decision effect, selected based on the strategy action probability distribution. In practical applications, depending on different decision-making scenarios, the adaptive decision action can be a specific operating instruction, control strategy, etc. Furthermore, during execution, a reward signal is obtained from the environment in real time. This reward signal is the environment's feedback on the decision action, indicating the execution effect of the decision action. The reward signal can be positive, indicating that the decision action has achieved good results, or negative, indicating that the decision action has had adverse effects. The cumulative error in the reward signal reflects the overall gap between the decision results of the quantum strategy network and the optimal decision result over multiple decision-making processes. When the cumulative error reaches a certain threshold, quantum feature recoding is triggered. This quantum feature recoding is the process of recoding the quantum state eigenvector obtained in step S101. By adjusting the parameters of the quantum state to make it more consistent with the current environment and decision-making requirements, an optimized quantum state eigenvector is generated, and an optimized decision strategy is output. This optimized decision strategy, based on the optimized quantum state eigenvector, is a decision solution that can more accurately adapt to environmental changes and achieve the optimal decision effect.
[0082] In this method, by performing quantum wavelet multiscale decomposition on multidimensional input data and dynamically allocating the number of quantum bits based on the variance of each dimension to encode the quantum state feature vector, it can effectively extract key features and rationally allocate quantum resources for different types of data, such as structured tabular data, time-series sensor data, and multimodal image data. This provides high-quality input for subsequent data processing, making the data feature representation more precise and conducive to improving decision accuracy. Secondly, the quantum state feature vector is input into a quantum strategy network constructed by block variational quantum circuits and inserted with synchronization gates, and the strategy action probability distribution is output. On the one hand, the structure of the block variational quantum circuit helps to enhance the expressive power and flexibility of the quantum network, enabling better fitting of complex decision-making strategies. On the other hand, the setting of the synchronization gate and the dynamic adjustment of the operating parameters effectively solve the problem of clock signal deviation between the quantum processor and the classical controller, ensuring the stable operation and output reliability of the quantum strategy network, and ensuring the accuracy and effectiveness of the strategy action probability distribution.
[0083] Furthermore, after calculating the error between the policy action probability distribution and the target value function, this method uses an annealing regularization term to modify the quantum natural gradient to optimize the parameters of the block variational quantum circuit. This method can balance the model's fitting and generalization capabilities during the optimization process, avoiding problems such as overfitting, making the parameter optimization of the block variational quantum circuit more reasonable and effective, thereby improving the performance of the quantum policy network and laying the foundation for generating better decision-making strategies. Finally, by executing adaptive decision actions corresponding to the policy action probability distribution to obtain reward signals, and dynamically triggering quantum feature recoding based on the cumulative error between the updated parameters and the reward signal, it can timely adjust and optimize the quantum state eigenvector to better reflect environmental changes and the effectiveness of decisions, thereby enabling the output optimized decision strategy to continuously adapt to new situations and needs, achieving continuous optimization of decisions, and improving the intelligence and adaptability of the decision-making process under complex and changing environments and tasks.
[0084] In one embodiment, multidimensional input data is subjected to quantum wavelet multi-scale decomposition, the number of quantum bits is dynamically allocated based on the variance of each dimension, and the decomposed characteristic components are encoded into quantum state characteristic vectors, including:
[0085] Perform quantum Haar wavelet transform on multi-dimensional input data to generate multi-resolution feature components;
[0086] Calculate the variance of each dimension of the multidimensional input data, and dynamically allocate the number of qubits based on the variance using the following formula to obtain the number of qubits after dynamic allocation:
[0087]
[0088] Among them, B j is the number of quantum bits after dynamic allocation, is the variance of the jth dimension of the multidimensional input data, Q is the total number of available quantum bits in the system, is the variance of the i-th dimension of the multidimensional input data;
[0089] If the total number of quantum bits after dynamic allocation exceeds the total number of available quantum bits in the system, the number of quantum bits after dynamic allocation is proportionally compressed to obtain the compressed number of quantum bits;
[0090] Based on the compressed number of quantum bits, the multi-resolution feature components are encoded into quantum state feature vectors.
[0091] Specifically, the quantum Haar wavelet transform is a specific implementation of quantum wavelet multiscale decomposition. In traditional signal processing, the Haar wavelet transform is a wavelet transform method that decomposes signals into approximate and detailed components at different scales. In a quantum computing environment, the quantum Haar wavelet transform can leverage the superposition and entanglement properties of quantum states to decompose multidimensional input data at different scales, generating multi-resolution feature components. These feature components contain information about the data at different scales, from macroscopic approximate features to microscopic detailed features, providing rich information for subsequent data processing and analysis. Subsequently, the variance of each dimension of the multidimensional input data can be calculated and the number of qubits can be dynamically allocated based on the variance using the formula above. Specifically, the total number of available qubits in the system is allocated based on the proportion of the variance of each dimension to the total variance. Dimensions with large variances are assigned more qubits to more accurately represent their characteristics; dimensions with small variances are assigned fewer qubits.
[0092] If the total number of dynamically allocated qubits exceeds the total number of available qubits in the system during the actual allocation process, a compression ratio can be calculated and multiplied by the number of dynamically allocated qubits in each dimension to obtain a compressed qubit number, ensuring that the total number of qubits allocated across all dimensions does not exceed the total number of available qubits in the system. Based on the compressed qubit number, the values of the multi-resolution characteristic components can then be mapped to the amplitude or phase of the quantum state. Using a preset encoding method, such as amplitude encoding or phase encoding, the characteristic component values are converted into the corresponding parameters of the quantum state, resulting in a quantum state eigenvector. This quantum state eigenvector can fully leverage the advantages of quantum computing for efficient information processing and analysis.
[0093] In one embodiment, the calculation formula of the quantum state eigenvector is:
[0094]
[0095] Among them, |ψ(w i )> represents the encoded quantum state feature vector, which is used to describe the quantum state of the i-th sample after encoding, B j ′ is the number of quantum bits after compression, R y is a rotation gate around the y-axis, used to apply a rotation operation around the y-axis to the quantum bit, w ij To represent the multi-resolution feature component w k The j-th dimension data of the i-th sample, w i is the multi-resolution feature component w k The i-th sample in .
[0096] In one embodiment, a quantum state feature vector is input into a quantum strategy network constructed by a block variational quantum circuit, a synchronization gate is inserted into the block variational quantum circuit, and a strategy action probability distribution is output, including:
[0097] Input the quantum state feature vector into the initial quantum register of the block variational quantum circuit as the input state of the quantum strategy network;
[0098] According to the dimensional distribution of the quantum state eigenvector, the block variational quantum circuit is split into multiple sub-circuits using the following formula:
[0099] M=D / Q block
[0100] Where M is the total number of subcircuits, D is the dimensional distribution of the quantum state eigenvector, Q block is the maximum quantum bit capacity of each subcircuit, and each subcircuit contains a parameterized quantum gate, which is composed of a combination of rotation gates around the Z axis and the X axis;
[0101] Insert CZ entanglement gates between adjacent sub-circuits and construct block variational quantum circuits using the following formula:
[0102]
[0103] Among them, U(θ) represents the overall operation of the block variational quantum circuit, which consists of multiple sub-operations, U m (θ m ) is the parameterized quantum gate operation of the mth subcircuit, θ m is the parameter set of the subcircuit, and Entangle(m,m+1) is the CZ entanglement gate inserted between adjacent subcircuits m and m+1.
[0104] According to the clock signal deviation between the clock signal of the quantum processor and the clock signal of the classical controller, the time parameter of the synchronization gate is calculated, and the synchronization gate is constructed by the following formula:
[0105]
[0106] in is the Hamiltonian of the synchronous gate, is the Z Pauli operator of the i-th quantum bit;
[0107] A synchronization gate is inserted into the block variational quantum circuit, and the clock signal deviation is calibrated to be less than the preset deviation threshold through quantum state tomography technology, and the probability distribution of the strategy action is output.
[0108] Specifically, the block variational quantum circuit is the core execution unit of the quantum strategy network, and the initial quantum register serves as the input interface of the quantum circuit, responsible for receiving the quantum state characteristic vector. In addition, in the quantum computing system, the quantum register is composed of multiple quantum bits. When the quantum state characteristic vector is input into the initial quantum register, the mapping operation of the quantum state is used to enable the initial state of the register to accurately reflect the characteristic information of the original data, thereby providing the input state for subsequent quantum computing. The dimensional distribution of the quantum state characteristic vector reflects the complexity and amount of information of the data characteristics. In order to improve the computational efficiency and resource utilization of the quantum circuit, the block variational quantum circuit can be split according to the dimension of the quantum state characteristic vector using the above formula. This formula ensures that all characteristic dimensions can be effectively processed by rounding up the calculations, where the maximum quantum bit capacity Q of each sub-circuit is block It can be pre-set based on the hardware performance and computing resource limitations of the quantum processor. Each of the split sub-circuits contains a parameterized quantum gate. This parameterized quantum gate is composed of a combination of rotation gates around the Z and X axes. Adjusting the rotation angle can change the phase and probability amplitude distribution of the quantum state, thereby enabling targeted processing of the local features of the quantum state eigenvector, enabling the extraction and transformation of feature information.
[0109] Specifically, the CZ entanglement gate is a two-qubit gate that can generate an entangled state between two qubits. By inserting CZ entanglement gates between adjacent subcircuits, a block variational quantum circuit can be constructed based on the above formula. In this construction process, the parameterized quantum gate operation U of each subcircuit is m (θ m) prioritizes processing local features, then uses CZ entanglement gates to facilitate information transfer and fusion between subcircuits. This allows the quantum circuit to process quantum state feature vectors at a holistic level, forming a complete feature processing and decision mapping. Furthermore, to eliminate the impact of clock signal deviation, the timing parameters of the synchronization gate can be calculated using the above formula to construct a synchronization gate. This synchronization gate is then inserted into the block variational quantum circuit. Its position and insertion method can be optimized based on the overall structure and computational requirements of the quantum circuit to ensure that the synchronization gate effectively coordinates the operation of the quantum processor and the classical controller. After the synchronization gate is inserted, the quantum circuit output can be calibrated using quantum state tomography. Quantum state tomography is a method for reconstructing quantum states through multiple measurements. By measuring and analyzing the output state of the quantum circuit, complete information about the quantum state can be obtained. During the calibration process, the goal is to keep clock signal deviation below a preset deviation threshold. The timing parameters of the synchronization gate and the parameters of the quantum circuit's parameterized quantum gates can be adjusted to optimize the quantum circuit output. Once calibration is complete, the quantum strategy network outputs the probability distribution of the strategy action through measurement and probability calculation of the final quantum state. This distribution characterizes the probability of each decision action occurring under the current quantum state feature vector input, providing a quantitative basis for subsequent decision optimization.
[0110] In one embodiment, a synchronization gate is inserted into the block variational quantum circuit, and the clock signal deviation is calibrated to be less than a preset deviation threshold through quantum state tomography technology. The strategy action probability distribution is output, where the preset deviation threshold is 0.1ns, including:
[0111] The quantum state after the synchronization gate is measured using quantum state tomography technology, and the fidelity between the quantum state and the ideal quantum state is calculated using the following formula:
[0112]
[0113] Where Tr is the trace operation, which is used to sum the diagonal elements of the matrix, ρ ideal is the density matrix of the ideal quantum state, ρ out is the density matrix of the quantum state after the synchronization gate;
[0114] If the fidelity is less than 0.99, adjust the timing parameters of the synchronization gate using the following formula:
[0115] t←t·(1+0.1·(1-F))
[0116] Repeatedly adjust the timing parameters of the synchronization gate until the fidelity is greater than or equal to 0.99 and the clock signal deviation is less than 0.1ns. The corresponding calibrated synchronization gate parameters are obtained and the probability distribution of the strategy action is output.
[0117] Specifically, in this embodiment, a preset deviation threshold of 0.1ns is set, and fidelity is introduced to further measure the effectiveness of the synchronization gate. The fidelity can be calculated using the above formula, with a value ranging from 0 to 1. The closer it is to 1, the closer the actual quantum state is to the ideal quantum state. When the calculated fidelity is less than 0.99, it means that there is a significant deviation between the quantum state after the synchronization gate and the ideal quantum state. The timing parameters of the synchronization gate can then be repeatedly adjusted using the above formula until the preset fidelity and clock signal deviation requirements are met. This means that the synchronization gate can be considered calibrated to the ideal state, and the block variational quantum circuit can operate stably and accurately.
[0118] In one embodiment, the error between the probability distribution of the policy action and the target value function is calculated, the quantum natural gradient is corrected using the annealing regularization term, the parameters of the block variational quantum circuit are optimized, and the updated parameters are generated, including:
[0119] The mean square error between the policy action probability distribution and the target value function is calculated using the following formula:
[0120] L(θ)=E s,a [(Q(s,a)-V(s)) 2 ]
[0121] Among them, L(θ) is the mean square error loss function, E s,a is the mathematical expectation of state s and action a, Q(s,a) is the action value function obtained by quantum phase estimation, which represents the expected return of performing action a in state s, and V(s) is the target value function, which represents the long-term value of state s;
[0122] According to the gradient of the probability distribution of the strategy action, the quantum Fisher information matrix is calculated, and the pseudo-inverse of the quantum Fisher information matrix is calculated to obtain the pseudo-inverse matrix;
[0123] generating a dynamic annealing coefficient according to the chaos sensitivity of the parameters of the block variational quantum circuit, wherein the chaos sensitivity is quantified by a Lyapunov exponent, and the annealing coefficient is generated based on the Lyapunov exponent;
[0124] Based on the pseudo-inverse matrix and the annealing coefficient, the annealing regularization term is used to correct the quantum natural gradient using the following formula to obtain the corrected quantum natural gradient:
[0125]
[0126] Among them, F -1 (θ) is the pseudo-inverse matrix, is the gradient of the original loss function L(θ) with respect to the parameter θ, indicating the direction of loss decrease, and λ(t) is the annealing coefficient;
[0127] Based on the modified quantum natural gradient, multi-task gradient conflicts are suppressed through quantum interference path selection, the parameters of the block variational quantum circuit are updated, and the updated parameters are generated.
[0128] Specifically, in the above mean squared error calculation formula, state s includes system environment information described by multidimensional input data, such as market conditions in financial transactions or equipment parameters in industrial control. Action a corresponds to an executable decision option, such as an investment strategy adjustment or a device operation instruction. Q(s,a) is obtained through quantum phase estimation technology. This technique leverages the superposition and interference properties of quantum states to quantitatively assess the expected return after executing action a in state s, enabling efficient calculation of complex function values. This formula is then used to calculate the mean squared error between the probability distribution of the strategy actions and the target value function.
[0129] Specifically, the quantum Fisher information matrix is an important tool for measuring the distinguishability of quantum state parameters. Its calculation relies on the gradient of the probability distribution of policy actions. In quantum computing, small changes in parameters can lead to complex evolution of the quantum state. The quantum Fisher information matrix can capture the sensitivity of these changes, providing key information for further optimization. The gradient of the policy action probability distribution can be derived with respect to the parameters of the block variational quantum circuit. This gradient reflects the direction and extent of the impact of parameter changes on the policy output. Based on this gradient, the quantum Fisher information matrix can be constructed. Since the quantum Fisher information matrix can be singular, that is, irreversible, its pseudo-inverse can be calculated and used to correct the quantum natural gradient for more efficient parameter updates. Furthermore, during the optimization process of the block variational quantum circuit, its parameter space may exhibit chaotic characteristics, meaning that small changes in parameters can lead to drastic fluctuations in the output. In this embodiment, the chaotic sensitivity of the circuit parameters is quantified using the Lyapunov exponent. The Lyapunov exponent reflects the speed at which adjacent orbits in the system separate in phase space; larger values indicate a higher degree of chaos in the system. Based on the calculated Lyapunov exponent, a specific algorithm can be used to generate an annealing coefficient. This annealing coefficient plays a regulatory role during the optimization process, similar to the temperature parameter in simulated annealing. It gradually decreases as the optimization progresses to balance global search and local convergence. This annealing coefficient ensures adaptive adjustment based on the chaotic characteristics of the line parameters.
[0130] Specifically, based on the pseudo-inverse matrix and the annealing coefficient, the annealing regularization term can be used to correct the quantum natural gradient through the above formula. In this formula, the pseudo-inverse matrix can map the original gradient to the natural gradient direction on the quantum state manifold. Compared with the traditional gradient, this natural gradient takes into account the geometric structure of the quantum state space, can more efficiently guide parameter updates, and reduce redundant calculations. The corrected quantum natural gradient integrates the information of the quantum state geometry and the dynamic regularization constraint, providing an optimization direction for the parameter update of the block variational quantum circuit. However, in practical applications, the quantum strategy network may handle multiple optimization tasks simultaneously, such as maximizing benefits and minimizing risks, and the gradient directions of different tasks may conflict with each other. Therefore, during the update process, a quantum interference path selection mechanism can be further introduced to utilize the interference characteristics of quantum states to suppress gradient conflicts in multi-task scenarios. That is, quantum interference path selection simulates the interference phenomenon of quantum states on different paths and dynamically adjusts the weights of the gradients of each task, so that parameter updates can comprehensively consider multiple goals, avoid optimization stagnation or performance degradation caused by gradient conflicts, and ultimately generate optimized parameters, which can enable the quantum policy network to better adapt to environmental changes, output a better probability distribution of policy actions, and achieve the goal of multi-dimensional data adaptive decision optimization.
[0131] In one embodiment, an adaptive decision action corresponding to the probability distribution of the strategic action is executed, a reward signal is obtained from the environment by executing the adaptive decision action, and quantum feature recoding is dynamically triggered based on the cumulative error between the updated parameters and the reward signal to generate an optimized quantum state feature vector and output an optimized decision strategy, including:
[0132] According to the probability distribution of the strategic action, the adaptive decision action is selected through the Softmax decision function. The decision probability in the Softmax decision function is determined by the probability distribution of the strategic action after being smoothed by the temperature coefficient.
[0133] Perform adaptive decision actions, obtain reward signals from the environment, and calculate the cumulative error of the reward signal using the following formula:
[0134]
[0135] in, is the predicted reward value based on the parameters of the block variational quantum circuit at time t, r t is the reward signal at time t;
[0136] If the cumulative error is greater than the preset threshold, the updated parameters are loaded into the block variational quantum circuit, triggering quantum feature recoding. The encoding weights are adjusted according to the gradient direction of the reward prediction loss function. The reward prediction loss is the mean square error between the predicted reward value and the corresponding reward signal. The optimized quantum state feature vector is generated using the following formula:
[0137]
[0138] Among them, X new is the optimized quantum state eigenvector, X is the original quantum state eigenvector before optimization, ⊙ represents the Hadamard product, Encoding weight gradient, used to reflect the reward prediction loss L r The direction and rate of change relative to the coding weight X;
[0139] Based on the optimized quantum state eigenvector and updated parameters, the optimization decision strategy is output.
[0140] Specifically, the temperature coefficient is an adjustable hyperparameter whose value directly affects the randomness and determinism of the decision. The decision probability in the Softmax decision function is determined by smoothing the policy action probability distribution with the temperature coefficient. This balances exploring new decisions and leveraging existing experience, adaptively selecting the optimal decision action based on actual needs. In practical applications, the reward signal can be set based on the specific task. For example, in financial investment decisions, the reward signal can be set as the investment return. To evaluate the effectiveness of the decision strategy, the squared error between the predicted reward value and the actual reward signal at each moment can be accumulated using the above formula to calculate the cumulative error. This cumulative error provides a comprehensive and dynamic measure of the accuracy and stability of the decision strategy. Furthermore, when the cumulative error exceeds a preset threshold, indicating a significant deviation between the current decision strategy's predictions and the actual environmental feedback, the updated parameters can be loaded into the block variational quantum circuit, enabling the circuit to perform calculations based on the newly optimized parameters. The original quantum state eigenvector is then re-encoded. Specifically, the encoding weights are adjusted according to the gradient direction of the reward prediction loss function using the above formula to generate an optimized quantum state eigenvector, further adapting to environmental changes and optimizing decision needs. Based on the optimized quantum state eigenvectors and updated block variational quantum circuit parameters, the corresponding output optimization decision strategy can more accurately adapt to the dynamically changing environment, achieving efficient and accurate decision-making for multidimensional data.
[0141] Based on the same inventive concept, Figure 2 As shown, the embodiment of the present application also provides a multi-dimensional data adaptive decision optimization system 200 based on quantum enhanced learning. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme described in the above method, so the specific limitations of one or more embodiments of the multi-dimensional data adaptive decision optimization system based on quantum enhanced learning provided below can be found in the above limitations of the multi-dimensional data adaptive decision optimization method based on quantum enhanced learning, and will not be repeated here. The system includes:
[0142] a quantum feature encoding module 201 for performing quantum wavelet multiscale decomposition on multidimensional input data, dynamically allocating the number of quantum bits based on the variance of each dimension, and encoding the decomposed feature components into quantum state feature vectors, wherein the multidimensional input data includes at least one of structured table data, time series sensor data, and multimodal image data;
[0143] A quantum strategy network module 202 is configured to input the quantum state feature vector into a quantum strategy network constructed by block variational quantum circuits, insert synchronization gates into the block variational quantum circuits, and output a probability distribution of strategy actions. The operating parameters of the synchronization gates are dynamically adjusted according to the clock signal deviation between the quantum processor and the classical controller.
[0144] The quantum gradient optimization module 203 is used to calculate the error between the probability distribution of the policy action and the target value function, use the annealing regularization term to correct the quantum natural gradient, optimize the parameters of the block variational quantum circuit, and generate updated parameters;
[0145] The adaptive decision feedback module 204 is used to execute the adaptive decision action corresponding to the probability distribution of the strategy action, obtain the reward signal from the environment by executing the adaptive decision action, and dynamically trigger the quantum feature recoding based on the cumulative error between the updated parameters and the reward signal, generate the optimized quantum state feature vector, and output the optimized decision strategy.
[0146] In the aforementioned quantum-enhanced learning-based multidimensional data adaptive decision optimization system 200, the quantum feature encoding module 201 can perform quantum wavelet multiscale decomposition on multidimensional input data, effectively extracting key features of different data types, dynamically allocating the number of quantum bits based on the variance of each dimension, rationally allocating quantum resources, and encoding the decomposed feature components into quantum state feature vectors, providing a richer and more efficient information representation for the subsequent decision-making process. The quantum strategy network module 202 can input the quantum state feature vectors into a quantum strategy network constructed from block variational quantum circuits and insert synchronization gates therein. The synchronization gate operating parameters are dynamically adjusted based on the clock signal deviation between the quantum processor and the classical controller to output the probability distribution of the strategy action. This fully leverages the advantages of quantum computing in parallel processing and rapid search, enabling the rapid and efficient generation of a more optimal strategy action probability distribution. Furthermore, by dynamically adjusting the synchronization gate parameters, the synchronization problem when the quantum processor and the classical controller work together is effectively resolved, ensuring the stability and reliability of the decision-making process.
[0147] The quantum gradient optimization module 203 can calculate the error between the probability distribution of the policy action and the target value function and use the annealing regularization term to correct the quantum natural gradient. This not only speeds up the convergence speed, but also improves the adaptability of the quantum policy network to complex environments, enabling the system to find the optimal solution more quickly and accurately. The adaptive decision feedback module 204 executes the adaptive decision action corresponding to the probability distribution of the policy action, obtains the reward signal from the environment, and dynamically triggers the quantum feature recoding based on the cumulative error between the updated parameters and the reward signal. This enables the system to perceive environmental changes in real time and adjust the quantum state feature vector in a timely manner to ensure that the output decision strategy can closely fit the actual scenario, further improving the timeliness and accuracy of the decision strategy.
[0148] Through the collaborative work of the above modules, the system realizes full-process optimization from multi-dimensional data encoding, policy network optimization to decision feedback adjustment, significantly improving the accuracy, efficiency and adaptability of multi-dimensional data decision optimization to complex environments, and providing strong technical support for complex decision-making scenarios in finance, medical care, intelligent transportation and other fields.
[0149] In an exemplary embodiment, the present invention further provides a computer device comprising a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the multi-dimensional data adaptive decision optimization method based on quantum reinforcement learning described herein. A multi-core processor is preferred to improve the system's parallel processing capabilities. The memory provides sufficient temporary storage space to support program execution and data processing. The memory capacity should be large enough to accommodate large amounts of supply information and computing tasks.
[0150] In an exemplary embodiment, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multi-dimensional data adaptive decision optimization method based on quantum enhanced learning of the present application.
[0151] The above-described embodiments merely represent several implementation methods of the embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the concept of the embodiments of the present application, and these modifications and improvements fall within the scope of protection of the embodiments of the present application.
Claims
1. A multi-dimensional data adaptive decision optimization method based on quantum enhanced learning, characterized in that: The method comprises: performing quantum wavelet multiscale decomposition on multidimensional input data, dynamically allocating the number of quantum bits based on the variance of each dimension, and encoding the decomposed characteristic components into quantum state characteristic vectors, wherein the multidimensional input data includes at least one of structured table data, time series sensor data, and multimodal image data; Inputting the quantum state eigenvector into a quantum strategy network constructed by a block variational quantum circuit, inserting a synchronization gate into the block variational quantum circuit, and outputting a strategy action probability distribution, wherein the operating parameters of the synchronization gate are dynamically adjusted according to the clock signal deviation between the quantum processor and the classical controller; Calculating the error between the probability distribution of the strategy action and the target value function, using the annealing regularization term to correct the quantum natural gradient, optimizing the parameters of the block variational quantum circuit, and generating updated parameters; Execute an adaptive decision action corresponding to the probability distribution of the strategy action, obtain a reward signal from the environment by executing the adaptive decision action, and dynamically trigger quantum feature recoding according to the cumulative error between the updated parameters and the reward signal to generate an optimized quantum state feature vector, and output an optimized decision strategy.
2. The method according to claim 1, characterized in that The method performs quantum wavelet multi-scale decomposition on the multi-dimensional input data, dynamically allocates the number of quantum bits based on the variance of each dimension, and encodes the decomposed characteristic components into quantum state characteristic vectors, including: Performing a quantum Haar wavelet transform on the multidimensional input data to generate multi-resolution feature components; Calculate the variance of each dimension of the multidimensional input data, and dynamically allocate the number of qubits according to the variance using the following formula to obtain the number of qubits after dynamic allocation: Among them, B j is the number of quantum bits after the dynamic allocation, is the variance of the j-th dimension of the multidimensional input data, Q is the total number of available quantum bits in the system, is the variance of the i-th dimension of the multidimensional input data; If the total number of qubits after the dynamic allocation exceeds the total number of qubits available in the system, compressing the number of qubits after the dynamic allocation proportionally to obtain a compressed number of qubits; Based on the compressed number of quantum bits, the multi-resolution feature component is encoded into the quantum state feature vector.
3. The method according to claim 2, characterized in that The calculation formula of the quantum state eigenvector is: Among them, |ψ(w i )> represents the encoded quantum state feature vector, which is used to describe the quantum state of the i-th sample after encoding, B j ′ is the number of quantum bits after compression, R y is a rotation gate around the y-axis, used to apply a rotation operation around the y-axis to the quantum bit, w ij To represent the multi-resolution feature component w k The j-th dimension data of the i-th sample, w i is the multi-resolution feature component w k The i-th sample in .
4. The method according to claim 1, wherein The step of inputting the quantum state feature vector into a quantum strategy network constructed by a block variational quantum circuit, inserting a synchronization gate into the block variational quantum circuit, and outputting a strategy action probability distribution comprises: Inputting the quantum state feature vector into the initial quantum register of the block variational quantum circuit as the input state of the quantum strategy network; According to the dimensional distribution of the quantum state eigenvector, the block variational quantum circuit is split into multiple sub-circuits using the following formula: M=D / Q block Where M is the total number of sub-circuits, D is the dimensional distribution of the quantum state eigenvector, Q block is the maximum quantum bit capacity of each of the sub-circuits, and each of the sub-circuits includes a parameterized quantum gate, and the parameterized quantum gate is composed of a combination of rotation gates around the Z axis and the X axis; Insert CZ entanglement gates between adjacent sub-circuits, and construct the block variational quantum circuit using the following formula: Among them, U(θ) represents the overall operation of the block variational quantum circuit, which consists of multiple sub-operations, U m (θ m ) is the parameterized quantum gate operation of the mth subcircuit, θ m is the parameter set of the subcircuit, and Entangle(m,m+1) is the CZ entanglement gate inserted between adjacent subcircuits m and m+1. According to the clock signal deviation between the clock signal of the quantum processor and the clock signal of the classical controller, the time parameter of the synchronization gate is calculated, and the synchronization gate is constructed by the following formula: in is the Hamiltonian of the synchronous gate, is the Z Pauli operator of the i-th quantum bit; The synchronization gate is inserted into the block variational quantum circuit, and the clock signal deviation is calibrated to be less than a preset deviation threshold through quantum state tomography technology, and the strategy action probability distribution is output.
5. The method according to claim 1, characterized in that The calculation of the error between the probability distribution of the strategy action and the target value function, the use of the annealing regularization term to correct the quantum natural gradient, the optimization of the parameters of the block variational quantum circuit, and the generation of updated parameters include: The mean square error between the strategy action probability distribution and the target value function is calculated using the following formula: L(θ)=E s,a [(Q(s,a)-V(s)) 2 ] Among them, L(θ) is the mean square error loss function, E s,a is the mathematical expectation of state s and action a, Q(s,a) is the action value function obtained by quantum phase estimation, which represents the expected return of performing action a in state s, and V(s) is the target value function, which represents the long-term value of state s. Calculating a quantum Fisher information matrix according to the gradient of the probability distribution of the strategy action, and calculating a pseudo-inverse of the quantum Fisher information matrix to obtain a pseudo-inverse matrix; generating a dynamic annealing coefficient according to the chaotic sensitivity of the parameters of the block variational quantum circuit, wherein the chaotic sensitivity is quantified by a Lyapunov exponent, and generating the annealing coefficient based on the Lyapunov exponent; Based on the pseudo-inverse matrix and the annealing coefficient, the annealing regularization term is used to correct the quantum natural gradient using the following formula to obtain a corrected quantum natural gradient: Among them, F -1 (θ) is the pseudo-inverse matrix, is the gradient of the original loss function L(θ) with respect to the parameter θ, indicating the direction of loss decrease, and λ(t) is the annealing coefficient; Based on the modified quantum natural gradient, multi-task gradient conflicts are suppressed through quantum interference path selection, and the parameters of the block variational quantum circuit are updated to generate updated parameters.
6. The method according to claim 1, characterized in that The adaptive decision-making action corresponding to the probability distribution of the strategic action is executed, a reward signal is obtained from the environment by executing the adaptive decision-making action, and quantum feature recoding is dynamically triggered according to the cumulative error between the updated parameter and the reward signal to generate an optimized quantum state feature vector, and an optimized decision-making strategy is output, including: According to the strategy action probability distribution, the adaptive decision action is selected through a Softmax decision function, where the decision probability in the Softmax decision function is determined by smoothing the strategy action probability distribution with a temperature coefficient; Execute the adaptive decision action, obtain the reward signal from the environment, and calculate the cumulative error of the reward signal using the following formula: in, is the predicted reward value based on the parameters of the block variational quantum circuit at time t, r t is the reward signal at time t; If the cumulative error is greater than a preset threshold, the updated parameters are loaded into the block variational quantum circuit, and the quantum feature recoding is triggered; The encoding weight is adjusted according to the gradient direction of the reward prediction loss function, where the reward prediction loss is the mean square error between the predicted reward value and the corresponding reward signal, and the optimized quantum state feature vector is generated by the following formula: Among them, X new is the optimized quantum state eigenvector, X is the original quantum state eigenvector before optimization, ⊙ represents the Hadamard product, Encoding weight gradient, used to reflect the reward prediction loss L r The direction and rate of change relative to the coding weight X; Based on the optimized quantum state eigenvector and the updated parameters, the optimization decision strategy is output.
7. The method according to claim 4, characterized in that The step of inserting the synchronization gate into the block variational quantum circuit and calibrating the clock signal deviation to be less than a preset deviation threshold by using quantum state tomography technology, and outputting the strategy action probability distribution, wherein the preset deviation threshold is 0.1 ns, includes: The quantum state after the synchronization gate is measured by the quantum state tomography technique, and the fidelity between the quantum state and the ideal quantum state is calculated by the following formula: Where Tr is the trace operation, which is used to sum the diagonal elements of the matrix, ρ ideal is the density matrix of the ideal quantum state, ρ out is the density matrix of the quantum state after the synchronization gate acts; If the fidelity is less than 0.99, the time parameter of the synchronous gate is adjusted by the following formula: t←t·(1+0.1·(1-F)) Repeatedly adjust the time parameters of the synchronization gate until the fidelity is greater than or equal to 0.99 and the clock signal deviation is less than 0.1ns, obtain the corresponding calibrated synchronization gate parameters, and output the strategy action probability distribution.
8. A multi-dimensional data adaptive decision optimization system based on quantum enhanced learning, characterized in that: The system comprises: a quantum feature encoding module for performing quantum wavelet multiscale decomposition on multidimensional input data, dynamically allocating the number of qubits based on the variance of each dimension, and encoding the decomposed feature components into quantum state feature vectors; a quantum strategy network module, configured to input the quantum state eigenvector into a quantum strategy network constructed by a block variational quantum circuit, insert a synchronization gate into the block variational quantum circuit, and output a probability distribution of strategy actions, wherein the operating parameters of the synchronization gate are dynamically adjusted according to the clock signal deviation between the quantum processor and the classical controller; A quantum gradient optimization module is used to calculate the error between the probability distribution of the policy action and the target value function, use the annealing regularization term to correct the quantum natural gradient, optimize the parameters of the block variational quantum circuit, and generate updated parameters; An adaptive decision feedback module is used to execute the adaptive decision action corresponding to the probability distribution of the strategy action, obtain a reward signal from the environment by executing the adaptive decision action, and dynamically trigger quantum feature recoding based on the cumulative error between the updated parameters and the reward signal to generate an optimized quantum state feature vector and output the optimized decision strategy.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Fuzzy semantic coding method, device and equipment based on electroencephalogram signals and medium
CN118964964A
Internet of Things intelligent detection method for electric power system
CN119827899A
Cited By
Control method and optimization system for energy loss of intelligent power plant
CN121028718A