A Multidimensional Data Adaptive Decision Optimization Method and System Based on Quantum Reinforcement Learning

The quantum policy network optimization method based on quantum wavelet decomposition and dynamic coding solves the problems of low resource utilization and insufficient decision accuracy in multidimensional data processing of traditional quantum reinforcement learning, and achieves efficient and stable multidimensional data decision optimization, which is applicable to fields such as finance, healthcare and intelligent transportation.

CN120449095BActive Publication Date: 2025-10-28SHANGHAI CIVIL AVIATION VOCATIONAL & TECH COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510539598.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-10-28
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

Traditional quantum reinforcement learning methods cannot adapt to the heterogeneity and dynamism of multidimensional data when processing it due to static encoding and resource allocation strategies. This results in low quantum resource utilization, making it difficult to achieve high-precision, real-time decision-making. Furthermore, they lack dynamic adaptation to hardware resources and collaborative optimization for dynamic environmental changes.

Method used

By performing quantum wavelet multi-scale decomposition on multidimensional input data, dynamically allocating the number of qubits, encoding the feature components into quantum state feature vectors, inserting synchronization gates in the block variable quantum circuits, using annealing regularization terms to correct the quantum natural gradient, optimizing the parameters of the quantum policy network, and dynamically triggering quantum feature recoding in conjunction with environmental feedback to generate an optimized decision strategy.

Benefits of technology

It improves the efficiency of quantum resource utilization, enhances the stability and dynamic adaptability of strategies, and improves the accuracy and efficiency of multidimensional data decision-making, making it suitable for complex decision-making scenarios such as finance, healthcare, and intelligent transportation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449095B_ABST
    Figure CN120449095B_ABST
Patent Text Reader

Abstract

This application relates to a multidimensional data adaptive decision optimization method and system based on quantum reinforcement learning. The method includes: performing quantum wavelet multi-scale decomposition on multidimensional input data and encoding the decomposed feature components into quantum state feature vectors; inputting this vector into a quantum policy network constructed from block variable quantum circuits containing synchronization gates, outputting a policy action probability distribution; calculating the error between this distribution and the target value function; correcting the quantum natural gradient using an annealing regularization term to generate updated parameters; executing the adaptive decision action corresponding to this distribution, obtaining a reward signal, and dynamically triggering quantum feature recoding in conjunction with the updated parameters to generate an optimized quantum state feature vector, and outputting an optimized decision policy. This method significantly improves the accuracy and efficiency of multidimensional data decision-making and enhances adaptability to complex dynamic environments through quantum feature encoding, quantum policy networks, parameter optimization, and adaptive feedback mechanisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer technology, and in particular relates to a multidimensional data adaptive decision optimization method and system based on quantum reinforcement learning. Background Technology

[0002] With the convergence of quantum computing and machine learning technologies, quantum-enhanced learning is gradually emerging as a new direction for solving complex decision optimization problems. This technology significantly improves the efficiency of high-dimensional data processing and real-time decision-making by combining the high parallelism of quantum computing with the policy optimization of classical reinforcement learning. Traditional quantum reinforcement learning methods typically employ fixed quantum circuit architectures and static resource allocation strategies, such as predefined parameterized models based on quantum neural networks or classification decision mechanisms based on quantum support vector machines, to handle classification and regression tasks for structured data or time-series signals. However, when processing multidimensional data, the static encoding and static resource allocation strategies cannot adapt to the heterogeneity and dynamism of multidimensional data, resulting in low quantum resource utilization. Furthermore, this technology relies on fixed quantum circuit topologies such as fully connected entangled structures, lacking dynamic adaptation to hardware resource constraints, making it difficult to effectively extract multi-scale features with a limited number of qubits. In addition, facing dynamic environmental changes, this technology struggles to achieve synergistic optimization of features and policies through real-time feedback, resulting in insufficient decision accuracy and optimization efficiency, failing to meet the high-precision, real-time decision-making requirements in complex scenarios. Summary of the Invention

[0003] Therefore, it is necessary to provide a multidimensional data adaptive decision optimization method and system based on quantum reinforcement learning to address the above-mentioned technical problems, so as to improve the efficiency of quantum resource utilization and decision accuracy in multidimensional data processing, and enhance the stability and dynamic adaptability of strategies in complex scenarios.

[0004] Firstly, this application provides a multidimensional data adaptive decision optimization method based on quantum reinforcement learning, including:

[0005] Quantum wavelet multiscale decomposition is performed on multidimensional input data, and the number of qubits is dynamically allocated based on the variance of each dimension. The decomposed feature components are encoded into quantum state feature vectors. The multidimensional input data includes at least one of structured tabular data, time-series sensor data, and multimodal image data.

[0006] The quantum state feature vector is input into the quantum policy network constructed by block variable quantum circuits. A synchronization gate is inserted into the block variable quantum circuits to output the policy action probability distribution. The operating parameters of the synchronization gate are dynamically adjusted according to the clock signal deviation between the quantum processor and the classical controller.

[0007] The error between the probability distribution of the strategy action and the target value function is calculated. The quantum natural gradient is corrected using the annealing regularization term. The parameters of the block variable quantum circuit are optimized to generate updated parameters.

[0008] The adaptive decision action corresponding to the probability distribution of the execution strategy action is executed. The reward signal is obtained from the environment by executing the adaptive decision action, and the quantum feature recoding is dynamically triggered according to the cumulative error between the updated parameters and the reward signal to generate the optimized quantum state feature vector and output the optimized decision strategy.

[0009] In one embodiment, the multidimensional input data is decomposed using quantum wavelet multi-scale decomposition. The number of qubits is dynamically allocated based on the variance of each dimension. The decomposed feature components are then encoded into quantum state feature vectors, including:

[0010] Quantum Haar wavelet transform is performed on multidimensional input data to generate multi-resolution feature components;

[0011] Calculate the variance of each dimension of the multidimensional input data, and dynamically allocate the number of qubits based on the variance using the following formula to obtain the dynamically allocated number of qubits:

[0012]

[0013] Among them, B j The number of qubits after dynamic allocation. Let be the variance of the j-th dimension of the multidimensional input data, and Q be the total number of usable qubits in the system. Let be the variance of the i-th dimension of the multidimensional input data;

[0014] If the total number of qubits after dynamic allocation exceeds the total number of available qubits in the system, the number of qubits after dynamic allocation is compressed proportionally to obtain the compressed number of qubits.

[0015] Based on the compressed number of qubits, the multi-resolution feature components are encoded into quantum state feature vectors.

[0016] In one embodiment, the formula for calculating the quantum state eigenvector is:

[0017]

[0018] Where, |ψ(w i )> represents the encoded quantum state feature vector, used to describe the encoded quantum state of the i-th sample, B j ′ represents the number of compressed qubits, R y A rotation gate about the y-axis, used to apply rotation operations about the y-axis to qubits, w ij To represent multi-resolution feature components w kThe j-th dimension of the i-th sample, w i For multi-resolution feature components w k The i-th sample.

[0019] In one embodiment, the quantum state feature vector is input into a quantum policy network constructed from block variable quantum circuits, a synchronization gate is inserted into the block variable quantum circuits, and the policy action probability distribution is output, including:

[0020] The quantum state eigenvectors are input into the initial quantum register of the block variable quantum circuit as the input state of the quantum policy network;

[0021] Based on the dimensionality distribution of the quantum state eigenvectors, the block-variable quantum circuit is divided into multiple sub-circuits using the following formula:

[0022] M = D / Q block

[0023] Where M is the total number of sub-circuits, D is the dimensional distribution of the quantum state eigenvectors, and Q... block The maximum qubit capacity of each sub-circuit is given, and each sub-circuit contains parameterized quantum gates, which are composed of a combination of rotating gates around the Z-axis and X-axis.

[0024] By inserting CZ entanglement gates between adjacent sub-lines, block-variable component sub-lines are constructed using the following formula:

[0025]

[0026] Where U(θ) represents the overall operation of the block variable component sub-line, which consists of multiple sub-operations, U m (θ m ) represents the parameterized quantum gate operation of the m-th sub-circuit, θ m The parameter set of the sub-line is Entangle(m,m+1), which is the CZ entanglement gate inserted between adjacent sub-lines m and m+1.

[0027] Based on the clock signal deviation between the quantum processor's clock signal and the classical controller's clock signal, the timing parameters of the synchronization gate are calculated, and the synchronization gate is constructed using the following formula:

[0028]

[0029] in This is the Hamiltonian for the synchronization gate. Let Z be the Pauli operator for the i-th qubit;

[0030] A synchronization gate is inserted into the block variable quantum circuit, and the clock signal deviation is calibrated to be less than a preset deviation threshold using quantum state tomography, and the probability distribution of the strategy action is output.

[0031] In one embodiment, the error between the probability distribution of the policy action and the target value function is calculated, the quantum natural gradient is corrected using an annealing regularization term, the parameters of the block-variable quantum circuit are optimized, and updated parameters are generated, including:

[0032] The mean squared error between the probability distribution of strategy actions and the target value function is calculated using the following formula:

[0033] L(θ)=E s,a [(Q(s,a)-V(s)) 2 ]

[0034] Where L(θ) is the mean square error loss function, E s,a Let Q(s,a) be the mathematical expectation of state s and action a, and let V(s) be the action value function obtained through quantum phase estimation, which represents the expected reward of performing action a in state s. Let V(s) be the target value function, which represents the long-term value of state s.

[0035] Based on the gradient of the probability distribution of the policy action, calculate the quantum Fisher information matrix and then calculate the pseudo-inverse of the quantum Fisher information matrix to obtain the pseudo-inverse matrix.

[0036] Dynamic annealing coefficients are generated based on the chaotic sensitivity of the parameters of the block variable quantum circuit, wherein the chaotic sensitivity is quantified by the Lyapunov exponent and the annealing coefficients are generated based on the Lyapunov exponent.

[0037] Based on the pseudo-inverse matrix and annealing coefficients, and using the annealing regularization term, the quantum natural gradient is corrected using the following formula, yielding the corrected quantum natural gradient:

[0038]

[0039] Among them, F -1 (θ) is the pseudo-inverse matrix. Let L(θ) be the gradient of the original loss function L(θ) with respect to the parameter θ, representing the direction of loss descent, and λ(t) be the annealing coefficient.

[0040] Based on the corrected quantum natural gradient, the parameters of the block variable quantum circuit are updated by suppressing multi-task gradient conflicts through quantum interference path selection, and the updated parameters are generated.

[0041] In one embodiment, an adaptive decision action corresponding to the probability distribution of the execution strategy action is performed. This adaptive decision action acquires a reward signal from the environment, and based on the cumulative error between the updated parameters and the reward signal, quantum feature recoding is dynamically triggered to generate an optimized quantum state feature vector. The optimized decision strategy is then output, including:

[0042] Based on the probability distribution of the strategy action, an adaptive decision action is selected through the Softmax decision function. The decision probability in the Softmax decision function is determined by smoothing the probability distribution of the strategy action with a temperature coefficient.

[0043] Execute adaptive decision-making actions, obtain reward signals from the environment, and calculate the cumulative error of the reward signals using the following formula:

[0044]

[0045] in, Let r be the predicted reward value based on the parameters of the block variable component quantum line at time t. t The reward signal is given at time t;

[0046] If the accumulated error exceeds the preset threshold, the updated parameters will be loaded into the block variable quantum circuit and quantum feature recoding will be triggered.

[0047] The encoding weights are adjusted according to the gradient direction of the reward prediction loss function, where the reward prediction loss is the mean square error between the predicted reward value and the corresponding reward signal. The optimized quantum state feature vector is generated using the following formula:

[0048]

[0049] Among them, X new Let X be the optimized quantum state eigenvector, and let ⊙ represent the Hadamard product. To encode the weight gradient, which reflects the reward prediction loss L r The direction and rate of change relative to the encoding weight X;

[0050] Based on the optimized quantum state eigenvectors and updated parameters, an optimized decision-making strategy is output.

[0051] In one embodiment, a synchronization gate is inserted into the block variable quantum circuit, and the clock signal deviation is calibrated to be less than a preset deviation threshold using quantum state tomography. The output strategy action probability distribution is then generated, where the preset deviation threshold is 0.1 ns, and includes:

[0052] The quantum state after synchronization gate action is measured using quantum state tomography, and the fidelity between the quantum state and the ideal quantum state is calculated using the following formula:

[0053]

[0054] Where Tr is the trace operation, used to sum the diagonal elements of a matrix, and ρ ideal Let ρ be the density matrix of an ideal quantum state. out This is the density matrix of the quantum states after the synchronization gate is applied;

[0055] If the fidelity is less than 0.99, adjust the timing parameters of the synchronization gate using the following formula:

[0056] t←t·(1+0.1·(1-F))

[0057] Repeatedly adjust the timing parameters of the synchronization gate until the fidelity is greater than or equal to 0.99 and the clock signal deviation is less than 0.1ns, to obtain the corresponding calibrated synchronization gate parameters and output the probability distribution of the strategy action.

[0058] Secondly, this application also provides a multidimensional data adaptive decision optimization system based on quantum reinforcement learning, comprising:

[0059] The quantum feature encoding module is used to perform quantum wavelet multi-scale decomposition on multidimensional input data, dynamically allocate the number of qubits based on the variance of each dimension, and encode the decomposed feature components into quantum state feature vectors. The multidimensional input data includes at least one of structured tabular data, time-series sensor data, and multimodal image data.

[0060] The quantum policy network module is used to input the quantum state feature vector into the quantum policy network constructed by block variable quantum circuits, insert synchronization gates in the block variable quantum circuits, and output the policy action probability distribution. The operating parameters of the synchronization gates are dynamically adjusted according to the clock signal deviation between the quantum processor and the classical controller.

[0061] The quantum gradient optimization module is used to calculate the error between the probability distribution of policy actions and the target value function. It uses the annealing regularization term to correct the quantum natural gradient, optimizes the parameters of the block variable quantum circuit, and generates updated parameters.

[0062] The adaptive decision feedback module is used to execute adaptive decision actions corresponding to the probability distribution of policy actions. By executing adaptive decision actions, it obtains reward signals from the environment, and dynamically triggers quantum feature recoding based on the cumulative error between the updated parameters and the reward signals, generating optimized quantum state feature vectors and outputting optimized decision strategies.

[0063] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the first aspect.

[0064] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the first aspect.

[0065] The aforementioned multidimensional data adaptive decision optimization method and system based on quantum reinforcement learning effectively overcomes the limitations of traditional quantum reinforcement learning methods in handling high-dimensional, multimodal data, such as low feature encoding efficiency and difficulty in capturing the complex internal correlations of the data. This is achieved by performing quantum wavelet multi-scale decomposition on multidimensional input data and dynamically allocating the number of qubits. The multidimensional input data includes various types such as structured tabular data, time-series sensor data, and multimodal image data, reflecting real-world scene information from multiple dimensions. Quantum encoding fully utilizes the superposition and entanglement properties of quantum states, mapping data features to quantum space and providing rich information representation for subsequent decision-making. Furthermore, by inputting the quantum state feature vectors into a quantum policy network constructed from block-variable quantum circuits and dynamically adjusting parameters through synchronization gates, the advantages of quantum computing in parallel processing and fast search are fully leveraged, resulting in a better policy action probability distribution. The synchronization gates adjust parameters based on clock signal deviations, effectively solving the synchronization problem when the quantum processor and classical controller work together, further ensuring the stability of the decision-making process.

[0066] Furthermore, this method utilizes annealing regularization to modify the quantum natural gradient to optimize block-variable quantum circuit parameters, overcoming the shortcomings of traditional gradient optimization methods that are prone to getting trapped in local optima and have low optimization efficiency in high-dimensional parameter spaces. By combining quantum natural gradient and annealing mechanisms, the convergence speed can be accelerated while improving the adaptability of the policy network to complex environments. Finally, the quantum feature recoding is dynamically triggered based on the reward signal from environmental feedback, enabling real-time perception of environmental changes and timely adjustment of quantum state feature vectors. This results in outputting optimization decision strategies that are more closely aligned with the actual scenario, ensuring the timeliness and accuracy of the decision strategy.

[0067] Compared with traditional quantum reinforcement learning methods, this method significantly improves the accuracy, efficiency, and environmental adaptability of multidimensional data decision optimization through quantum feature encoding, quantum policy network optimization, and dynamic feedback adjustment, providing strong technical support for complex decision-making scenarios in fields such as finance, healthcare, and intelligent transportation. Attached Figure Description

[0068] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0069] Figure 1 A flowchart of a multidimensional data adaptive decision optimization method based on quantum reinforcement learning is provided as an exemplary embodiment of the present invention.

[0070] Figure 2 This is a schematic diagram of a multidimensional data adaptive decision optimization system based on quantum reinforcement learning, provided as an exemplary embodiment of the present invention. Detailed Implementation

[0071] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0072] In one embodiment, such as Figure 1 As shown, a multidimensional data adaptive decision optimization method based on quantum reinforcement learning is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0073] S101: Perform quantum wavelet multi-scale decomposition on multidimensional input data, dynamically allocate the number of qubits based on the variance of each dimension, and encode the decomposed feature components into quantum state feature vectors. The multidimensional input data includes at least one of structured tabular data, time-series sensor data, and multimodal image data.

[0074] Specifically, structured tabular data can be transaction record tables related to the financial sector, containing multiple columns of structured data such as customer information, transaction amount, and transaction time. Time-series sensor data can be data such as temperature, humidity, and air pressure recorded chronologically by various sensors in a weather station. Multimodal image data can be data such as CT images and MRI images from medical imaging, without limitation. Subsequently, quantum wavelet multi-scale decomposition technology can be used to decompose the data of different dimensions in the multidimensional input data according to different scales and frequencies, thereby extracting the features of the data at different scales, which helps to facilitate more accurate analysis and processing in the future.

[0075] Secondly, the number of qubits can be dynamically allocated based on the variance of each dimension. Variance reflects the dispersion of data in a certain dimension; a large variance indicates greater data fluctuation and richer information content in that dimension. Therefore, more qubits can be allocated to dimensions with large variance and fewer qubits to dimensions with small variance. This allows for more detailed encoding of information-rich dimensions with limited quantum resources, improving the overall data representation. In quantum computing, quantum states possess properties such as superposition and entanglement, enabling efficient storage and processing of information. Finally, the decomposed feature components are encoded into quantum state feature vectors, transforming them into a quantum state form that can be processed by quantum computing, providing a solid foundation for subsequent quantum computing and decision-making.

[0076] S102: Input the quantum state feature vector into the quantum policy network constructed by block variable quantum circuits, insert a synchronization gate in the block variable quantum circuits, and output the policy action probability distribution. The operating parameters of the synchronization gate are dynamically adjusted according to the clock signal deviation between the quantum processor and the classical controller.

[0077] Specifically, block variable quantum circuits are quantum circuit structures based on quantum gate operations. They divide the entire quantum circuit into multiple sub-circuits, each of which can be independently parameterized, better adapting to different problems and data. Synchronization gates, on the other hand, are quantum gates used to coordinate the operation between the quantum processor and the classical controller. In practical quantum computing systems, the clock signals of the quantum processor and the classical controller may deviate, affecting the accuracy and stability of quantum computing. Therefore, by inserting synchronization gates into the block variable quantum circuits, and dynamically adjusting the operating parameters of these gates based on the clock signal deviation between the quantum processor and the classical controller, a real-time monitoring and feedback mechanism ensures that the quantum processor and the classical controller can work synchronously, thereby improving the reliability of quantum computing. After processing by the block variable quantum circuits and synchronization gates, the quantum policy network can output a policy action probability distribution. This policy action probability distribution represents the probability of various possible decision actions occurring under a given quantum state eigenvector, providing a basis for subsequent decision actions.

[0078] S103: Calculate the error between the probability distribution of the strategy action and the target value function, use the annealing regularization term to correct the quantum natural gradient, optimize the parameters of the block variable quantum circuit, and generate updated parameters.

[0079] Specifically, the error between the policy action probability distribution obtained in step S102 and the target value function is calculated. The target value function is a predefined function representing the expected optimal decision outcome in a specific decision scenario. This error reflects the gap between the current decision result of the quantum policy network and the optimal decision result. Therefore, to reduce this error, the quantum natural gradient can be corrected using an annealing regularization term. The quantum natural gradient is a method used in quantum computing to optimize parameters, considering the geometric structure of quantum states and guiding parameter updates more efficiently. The annealing regularization term is a regularization method that introduces a penalty term related to the parameters to avoid overfitting during parameter updates, further improving the model's generalization ability. Based on the corrected quantum natural gradient, the parameters of the block variable quantum circuit are optimized to generate updated parameters, enabling the quantum policy network to output a better policy action probability distribution in the next iteration.

[0080] S104: Execute the adaptive decision action corresponding to the probability distribution of the execution strategy action. Obtain the reward signal from the environment by executing the adaptive decision action, and dynamically trigger quantum feature recoding based on the cumulative error between the updated parameters and the reward signal to generate the optimized quantum state feature vector and output the optimized decision strategy.

[0081] Specifically, the adaptive decision action corresponding to the probability distribution of the strategy action in step S102 is executed. This adaptive decision action is the action most likely to achieve the optimal decision effect, selected based on the probability distribution of the strategy action. In practical applications, depending on different decision scenarios, the adaptive decision action can be a specific operation command, control strategy, etc. Furthermore, during execution, reward signals are acquired from the environment in real time. This reward signal is the feedback from the environment to the decision action, representing the execution effect of the decision action. The reward signal can be positive, indicating that the decision action has achieved a good effect, or negative, indicating that the decision action has produced a negative impact. The cumulative error of the reward signal reflects the overall gap between the decision result of the quantum policy network and the optimal decision result in multiple decision-making processes. When the cumulative error reaches a certain threshold, quantum feature recoding is triggered. This quantum feature recoding is the process of recoding the quantum state feature vector obtained in step S101. By adjusting the parameters of the quantum state to better suit the current environment and decision requirements, an optimized quantum state feature vector is generated, and an optimized decision strategy is output. This optimized decision strategy is a decision scheme obtained based on the optimized quantum state feature vector, which can more accurately adapt to environmental changes and achieve the optimal decision effect.

[0082] In the aforementioned method, by performing quantum wavelet multi-scale decomposition on multidimensional input data and dynamically allocating the number of qubits based on the variance of each dimension to encode the quantum state feature vector, it is possible to effectively extract key features and rationally allocate quantum resources for different types of data, such as structured tabular data, time-series sensor data, and multimodal image data. This provides high-quality input for subsequent data processing, making the feature representation of the data more accurate and conducive to improving the accuracy of decision-making. Secondly, the quantum state feature vector is input into a quantum policy network constructed from block variable quantum circuits and with synchronization gates inserted, outputting the policy action probability distribution. On the one hand, the block variable quantum circuit structure helps to enhance the expressive power and flexibility of the quantum network, enabling it to better fit complex decision-making policies; on the other hand, the setting of the synchronization gate and the dynamic adjustment of the operating parameters effectively solve the problem of clock signal deviation between the quantum processor and the classical controller, ensuring the stable operation of the quantum policy network and the reliability of the output, thus ensuring the accuracy and effectiveness of the policy action probability distribution.

[0083] Furthermore, after calculating the error between the policy action probability distribution and the target value function, this method uses an annealing regularization term to correct the quantum natural gradient to optimize the parameters of the block-variable quantum circuit. This balances the model's fitting and generalization abilities during optimization, avoiding overfitting and making the parameter optimization of the block-variable quantum circuit more reasonable and effective. This improves the performance of the quantum policy network and lays the foundation for generating better decision-making strategies. Finally, by executing adaptive decision actions corresponding to the policy action probability distribution to obtain reward signals, and dynamically triggering quantum feature recoding based on the cumulative error between the updated parameters and the reward signals, the quantum state feature vector can be adjusted and optimized in a timely manner. This allows it to better reflect environmental changes and the effectiveness of decision-making, enabling the output optimized decision-making strategy to continuously adapt to new situations and needs. This achieves continuous optimization of decision-making and improves the intelligence and adaptability of the decision-making process under complex and ever-changing environments and tasks.

[0084] In one embodiment, quantum wavelet multi-scale decomposition is performed on multidimensional input data, the number of qubits is dynamically allocated based on the variance of each dimension, and the decomposed feature components are encoded into quantum state feature vectors, including:

[0085] Quantum Haar wavelet transform is performed on multidimensional input data to generate multi-resolution feature components;

[0086] Calculate the variance of each dimension of the multidimensional input data, and dynamically allocate the number of qubits based on the variance using the following formula to obtain the dynamically allocated number of qubits:

[0087]

[0088] Among them, B j The number of qubits after dynamic allocation. Let be the variance of the j-th dimension of the multidimensional input data, and Q be the total number of usable qubits in the system. Let be the variance of the i-th dimension of the multidimensional input data;

[0089] If the total number of qubits after dynamic allocation exceeds the total number of available qubits in the system, the number of qubits after dynamic allocation is compressed proportionally to obtain the compressed number of qubits.

[0090] Based on the compressed number of qubits, the multi-resolution feature components are encoded into quantum state feature vectors.

[0091] Specifically, the quantum Haar wavelet transform is a concrete implementation of quantum wavelet multi-scale decomposition. In traditional signal processing, the Haar wavelet transform is a wavelet transform method that decomposes a signal into approximate and detail components at different scales. In a quantum computing environment, the quantum Haar wavelet transform can utilize the superposition and entanglement properties of quantum states to decompose multidimensional input data at different scales, generating multi-resolution feature components. These feature components contain information about the data at different scales, from macroscopic approximate features to microscopic detail features, providing rich information for subsequent data processing and analysis. Subsequently, the variance of each dimension of the multidimensional input data can be calculated, and the number of qubits can be dynamically allocated according to the variance using the formula mentioned above. That is, the total number of usable qubits in the system is allocated according to the proportion of each dimension's variance to the total variance. Dimensions with large variances are allocated more qubits to more accurately represent the features of that dimension; dimensions with small variances are allocated fewer qubits.

[0092] In actual allocation, if the total number of dynamically allocated qubits exceeds the total number of available qubits in the system, a compression ratio can be calculated. This ratio is then multiplied by the dynamically allocated qubits in each dimension to obtain the compressed qubit count, ensuring that the total number of allocated qubits in all dimensions does not exceed the total number of available qubits. Subsequently, based on the compressed qubit count, the values ​​of the multi-resolution feature components can be mapped to the amplitude or phase of the quantum state. Using a preset encoding method, such as amplitude encoding or phase encoding, the values ​​of the feature components are converted into the corresponding parameters of the quantum state, resulting in a quantum state feature vector. This quantum state feature vector fully leverages the advantages of quantum computing for efficient information processing and analysis.

[0093] In one embodiment, the formula for calculating the quantum state eigenvector is:

[0094]

[0095] Where, |ψ(w i )> represents the encoded quantum state feature vector, used to describe the encoded quantum state of the i-th sample, B j ′ represents the number of compressed qubits, R y A rotation gate about the y-axis, used to apply rotation operations about the y-axis to qubits, w ij To represent multi-resolution feature components w k The j-th dimension of the i-th sample, w i For multi-resolution feature components w k The i-th sample.

[0096] In one embodiment, the quantum state feature vector is input into a quantum policy network constructed from block variable quantum circuits, a synchronization gate is inserted into the block variable quantum circuits, and the policy action probability distribution is output, including:

[0097] The quantum state eigenvectors are input into the initial quantum register of the block variable quantum circuit as the input state of the quantum policy network;

[0098] Based on the dimensionality distribution of the quantum state eigenvectors, the block-variable quantum circuit is divided into multiple sub-circuits using the following formula:

[0099] M = D / Q block

[0100] Where M is the total number of sub-circuits, D is the dimensional distribution of the quantum state eigenvectors, and Q... block The maximum qubit capacity of each sub-circuit is given, and each sub-circuit contains parameterized quantum gates, which are composed of a combination of rotating gates around the Z-axis and X-axis.

[0101] By inserting CZ entanglement gates between adjacent sub-lines, block-variable component sub-lines are constructed using the following formula:

[0102]

[0103] Where U(θ) represents the overall operation of the block variable component sub-line, which consists of multiple sub-operations, U m (θ m ) represents the parameterized quantum gate operation of the m-th sub-circuit, θ m The parameter set of the sub-line is Entangle(m,m+1), which is the CZ entanglement gate inserted between adjacent sub-lines m and m+1.

[0104] Based on the clock signal deviation between the quantum processor's clock signal and the classical controller's clock signal, the timing parameters of the synchronization gate are calculated, and the synchronization gate is constructed using the following formula:

[0105]

[0106] in This is the Hamiltonian for the synchronization gate. Let Z be the Pauli operator for the i-th qubit;

[0107] A synchronization gate is inserted into the block variable quantum circuit, and the clock signal deviation is calibrated to be less than a preset deviation threshold using quantum state tomography, and the probability distribution of the strategy action is output.

[0108] Specifically, the block-variable quantum circuit is the core execution unit of the quantum policy network. The initial quantum register, as the input interface of the quantum circuit, is responsible for receiving quantum state feature vectors. Furthermore, in a quantum computing system, the quantum register consists of multiple qubits. When the quantum state feature vector is input into the initial quantum register, through quantum state mapping operations, the initial state of the register accurately reflects the feature information of the original data, thus providing the input state for subsequent quantum computation. The dimensional distribution of the quantum state feature vector reflects the complexity and information content of the data features. To improve the computational efficiency and resource utilization of the quantum circuit, the block-variable quantum circuit can be decomposed according to the dimension of the quantum state feature vector using the formula mentioned above. This formula, by rounding up, ensures that all feature dimensions can be effectively processed, where the maximum qubit capacity Q of each sub-circuit is... block The parameters can be pre-defined based on the hardware performance and computational resource limitations of the quantum processor. Furthermore, each sub-circuit after splitting contains a parameterized quantum gate. This parameterized quantum gate is composed of a combination of rotation gates around the Z-axis and X-axis, i.e., rotation gates around the Z-axis and rotation gates around the X-axis. By adjusting the rotation angle, the phase and probability amplitude distribution of the quantum state can be changed, thereby enabling targeted processing of local features of the quantum state eigenvector, achieving the extraction and transformation of feature information.

[0109] Specifically, the CZ entanglement gate is a two-qubit gate that can create an entangled state between two qubits. By inserting CZ entanglement gates between adjacent sub-circuits, block-variable quantum circuits can be constructed based on the above formula. During this construction process, the parameterized quantum gate operation U of each sub-circuit... m (θ mThe process prioritizes processing local features, then uses CZ entanglement gates to achieve information transfer and fusion between sub-circuits, enabling the quantum circuit to process quantum state feature vectors at the overall level, forming a complete feature processing and decision mapping. To eliminate the influence of clock signal deviation, the timing parameters of the synchronization gate can be calculated using the aforementioned formula to construct the synchronization gate. This synchronization gate is then inserted into the block-variable quantum circuit; its position and insertion method can be optimized according to the overall structure and computational requirements of the quantum circuit to ensure that the synchronization gate can effectively coordinate the operation of the quantum processor and the classical controller. After inserting the synchronization gate, the output of the quantum circuit can be calibrated using quantum state tomography (QT). QT is a method of reconstructing quantum states through multiple measurements, obtaining complete information about the quantum state by measuring and analyzing the output state of the quantum circuit. During calibration, the clock signal deviation is aimed at being less than a preset deviation threshold. The output of the quantum circuit can be optimized by adjusting the timing parameters of the synchronization gate and the parameterized quantum gate parameters of the quantum circuit. After calibration, the quantum policy network outputs the probability distribution of policy actions by measuring and calculating the probability of the final quantum state. This distribution characterizes the probability of each decision action occurring under the current quantum state eigenvector input, providing a quantitative basis for subsequent decision optimization.

[0110] In one embodiment, a synchronization gate is inserted into the block variable quantum circuit, and the clock signal deviation is calibrated to be less than a preset deviation threshold using quantum state tomography. The output strategy action probability distribution is then generated, where the preset deviation threshold is 0.1 ns, and includes:

[0111] The quantum state after synchronization gate action is measured using quantum state tomography, and the fidelity between the quantum state and the ideal quantum state is calculated using the following formula:

[0112]

[0113] Where Tr is the trace operation, used to sum the diagonal elements of a matrix, and ρ ideal Let ρ be the density matrix of an ideal quantum state. out This is the density matrix of the quantum states after the synchronization gate is applied;

[0114] If the fidelity is less than 0.99, adjust the timing parameters of the synchronization gate using the following formula:

[0115] t←t·(1+0.1·(1-F))

[0116] Repeatedly adjust the timing parameters of the synchronization gate until the fidelity is greater than or equal to 0.99 and the clock signal deviation is less than 0.1ns, to obtain the corresponding calibrated synchronization gate parameters and output the probability distribution of the strategy action.

[0117] Specifically, in this embodiment, a preset deviation threshold of 0.1 ns is set, and fidelity is introduced to further measure the effect of the synchronization gate. The fidelity can be calculated using the above formula, and its value ranges from 0 to 1. The closer it is to 1, the closer the actual quantum state is to the ideal quantum state. When the calculated fidelity is less than 0.99, it means that there is a large deviation between the quantum state after the synchronization gate is applied and the ideal quantum state. In this case, the timing parameters of the synchronization gate can be repeatedly adjusted using the above formula until the preset fidelity and clock signal deviation requirements are met. This indicates that the synchronization gate has been calibrated to an ideal state, and the block variable quantum circuit can operate stably and accurately.

[0118] In one embodiment, the error between the probability distribution of the policy action and the target value function is calculated, the quantum natural gradient is corrected using an annealing regularization term, the parameters of the block-variable quantum circuit are optimized, and updated parameters are generated, including:

[0119] The mean squared error between the probability distribution of strategy actions and the target value function is calculated using the following formula:

[0120] L(θ)=E s,a [(Q(s,a)-V(s)) 2 ]

[0121] Where L(θ) is the mean square error loss function, E s,a Let Q(s,a) be the mathematical expectation of state s and action a, and let V(s) be the action value function obtained through quantum phase estimation, which represents the expected reward of performing action a in state s. Let V(s) be the target value function, which represents the long-term value of state s.

[0122] Based on the gradient of the probability distribution of the policy action, calculate the quantum Fisher information matrix and then calculate the pseudo-inverse of the quantum Fisher information matrix to obtain the pseudo-inverse matrix.

[0123] Dynamic annealing coefficients are generated based on the chaotic sensitivity of the parameters of the block variable quantum circuit, wherein the chaotic sensitivity is quantified by the Lyapunov exponent and the annealing coefficients are generated based on the Lyapunov exponent.

[0124] Based on the pseudo-inverse matrix and annealing coefficients, and using the annealing regularization term, the quantum natural gradient is corrected using the following formula, yielding the corrected quantum natural gradient:

[0125]

[0126] Among them, F -1 (θ) is the pseudo-inverse matrix. Let L(θ) be the gradient of the original loss function L(θ) with respect to the parameter θ, representing the direction of loss descent, and λ(t) be the annealing coefficient.

[0127] Based on the corrected quantum natural gradient, the parameters of the block variable quantum circuit are updated by suppressing multi-task gradient conflicts through quantum interference path selection, and the updated parameters are generated.

[0128] Specifically, in the above formula for calculating the mean squared error, state s includes system environment information described by multi-dimensional input data, such as market conditions in financial transactions or equipment parameters in industrial control; action a corresponds to executable decision options, such as investment strategy adjustments or equipment operation instructions. Q(s,a) is obtained through quantum phase estimation technology, which utilizes the superposition and interference properties of quantum states to quantitatively evaluate the expected return after performing action a in state s, enabling efficient calculation of complex function values. This formula then allows for the calculation of the mean squared error between the probability distribution of strategy actions and the target value function.

[0129] Specifically, the quantum Fisher information matrix is ​​an important tool for measuring the distinguishability of quantum state parameters, and its calculation depends on the gradient of the policy action probability distribution. In quantum computing, small changes in parameters can lead to complex evolutions in the quantum state. The quantum Fisher information matrix can capture the sensitivity to these changes, providing crucial information for further optimization. The gradient can be obtained by differentiating the policy action probability distribution with respect to the parameters of the block-variable quantum circuit. This gradient reflects the direction and extent of the influence of parameter changes on the policy output, and the quantum Fisher information matrix can be constructed based on this gradient. Since the quantum Fisher information matrix may be a singular matrix, i.e., a non-invertible matrix, its pseudo-inverse matrix can be calculated to correct the quantum natural gradient, thereby achieving more efficient parameter updates. Furthermore, during the optimization process of the block-variable quantum circuit, its parameter space may exhibit chaotic characteristics, meaning that small changes in parameters can lead to drastic fluctuations in the output results. In this embodiment, the chaotic sensitivity of the circuit parameters is quantified by the Lyapunov exponent. The Lyapunov exponent reflects the separation velocity of neighboring orbits in the phase space; a larger value indicates a higher degree of chaos in the system. Based on the calculated Lyapunov exponent, an annealing coefficient can be generated using a specific algorithm. This annealing coefficient plays a regulatory role during the optimization process, similar to the temperature parameter in simulated annealing algorithms. It gradually decreases as the optimization progresses to balance global search and local convergence capabilities. This annealing coefficient further ensures that the optimization can adaptively adjust according to the chaotic characteristics of the circuit parameters.

[0130] Specifically, based on the pseudo-inverse matrix and annealing coefficients, the quantum natural gradient can be corrected using the annealing regularization term and the formula above. In this formula, the pseudo-inverse matrix maps the original gradient to the natural gradient direction on the quantum state manifold. Compared to the traditional gradient, this natural gradient considers the geometric structure of the quantum state space, guiding parameter updates more efficiently and reducing redundant computation. The corrected quantum natural gradient integrates quantum state geometric information with dynamic regularization constraints, providing an optimization direction for parameter updates in block-variable quantum circuits. However, in practical applications, quantum policy networks may simultaneously handle multiple optimization tasks such as maximizing rewards and minimizing risks, and the gradient directions for different tasks may conflict. Therefore, during the update process, a quantum interference path selection mechanism can be further introduced. By utilizing the interference properties of quantum states, gradient conflicts in multi-task scenarios can be suppressed. That is, quantum interference path selection dynamically adjusts the weights of gradients for each task by simulating the interference phenomenon of quantum states on different paths. This allows parameter updates to comprehensively consider multiple objectives, avoid optimization stagnation or performance degradation caused by gradient conflicts, and ultimately generate optimized parameters. This enables the quantum policy network to better adapt to environmental changes, output a better policy action probability distribution, and achieve the goal of multi-dimensional data adaptive decision optimization.

[0131] In one embodiment, an adaptive decision action corresponding to the probability distribution of the execution policy action is performed. This adaptive decision action acquires a reward signal from the environment, and based on the cumulative error between the updated parameters and the reward signal, quantum feature recoding is dynamically triggered to generate an optimized quantum state feature vector. The optimized decision policy is then output, including:

[0132] Based on the probability distribution of the strategy action, an adaptive decision action is selected through the Softmax decision function. The decision probability in the Softmax decision function is determined by smoothing the probability distribution of the strategy action with a temperature coefficient.

[0133] Execute adaptive decision-making actions, obtain reward signals from the environment, and calculate the cumulative error of the reward signals using the following formula:

[0134]

[0135] in, Let r be the predicted reward value based on the parameters of the block variable component quantum line at time t. t The reward signal is given at time t;

[0136] If the accumulated error exceeds a preset threshold, the updated parameters are loaded into the block-variable quantum circuit, and quantum feature recoding is triggered. The encoding weights are adjusted according to the gradient direction of the reward prediction loss function, where the reward prediction loss is the mean square error of the predicted reward value and the corresponding reward signal. The optimized quantum state feature vector is generated using the following formula:

[0137]

[0138] Among them, X new Let X be the optimized quantum state eigenvector, and let ⊙ represent the Hadamard product. To encode the weight gradient, which reflects the reward prediction loss L r The direction and rate of change relative to the encoding weight X;

[0139] Based on the optimized quantum state eigenvectors and updated parameters, an optimized decision-making strategy is output.

[0140] Specifically, the temperature coefficient is an adjustable hyperparameter whose value directly affects the randomness and determinism of decision-making. The decision probability in the Softmax decision function is determined by smoothing the policy action probability distribution with the temperature coefficient, achieving a balance between exploring new decisions and utilizing existing experience, and adaptively selecting the optimal decision action according to actual needs. In practical applications, the reward signal can be set according to specific tasks; for example, in financial investment decisions, the reward signal can be set as investment returns. Furthermore, to evaluate the effectiveness of the decision-making strategy, the cumulative error can be calculated by continuously accumulating the squared error between the predicted reward value and the actual reward signal at each time step using the above formula. This cumulative error can comprehensively and dynamically measure the accuracy and stability of the decision-making strategy. In addition, when the cumulative error exceeds a preset threshold, it indicates that there is a significant deviation between the prediction result of the current decision-making strategy and the actual environmental feedback. In this case, the updated parameters can be loaded into the block variable quantum circuit, enabling the circuit to perform calculations based on the latest optimized parameters and re-encode the original quantum state feature vector. That is, the encoding weights are adjusted according to the gradient direction of the reward prediction loss function using the above formula to generate an optimized quantum state feature vector, further adapting to environmental changes and optimizing decision-making needs. Based on the optimized quantum state eigenvectors and the updated block variable quantum circuit parameters, the corresponding output optimization decision strategy can more accurately adapt to the dynamically changing environment, achieving efficient and accurate decision-making on multidimensional data.

[0141] Based on the same inventive concept, such as Figure 2 As shown, this application also provides a multidimensional data adaptive decision optimization system 200 based on quantum reinforcement learning. The solution provided by this system is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of a multidimensional data adaptive decision optimization system based on quantum reinforcement learning provided below can be found in the limitations of the multidimensional data adaptive decision optimization method based on quantum reinforcement learning described above, and will not be repeated here. The system includes:

[0142] The quantum feature encoding module 201 is used to perform quantum wavelet multi-scale decomposition on multidimensional input data, dynamically allocate the number of qubits based on the variance of each dimension, and encode the decomposed feature components into quantum state feature vectors. The multidimensional input data includes at least one of structured tabular data, time-series sensor data, and multimodal image data.

[0143] The quantum policy network module 202 is used to input the quantum state feature vector into the quantum policy network constructed by the block variable quantum circuit, insert a synchronization gate in the block variable quantum circuit, and output the policy action probability distribution. The operating parameters of the synchronization gate are dynamically adjusted according to the clock signal deviation between the quantum processor and the classical controller.

[0144] The quantum gradient optimization module 203 is used to calculate the error between the probability distribution of policy actions and the target value function, correct the quantum natural gradient using the annealing regularization term, optimize the parameters of the block variable quantum circuit, and generate updated parameters.

[0145] The adaptive decision feedback module 204 is used to execute adaptive decision actions corresponding to the probability distribution of policy actions. By executing adaptive decision actions, it obtains reward signals from the environment, and dynamically triggers quantum feature recoding based on the cumulative error between the updated parameters and the reward signals to generate optimized quantum state feature vectors and output optimized decision strategies.

[0146] In the aforementioned multidimensional data adaptive decision optimization system 200 based on quantum reinforcement learning, the quantum feature encoding module 201 can perform quantum wavelet multi-scale decomposition on multidimensional input data, effectively extract key features of different types of data, dynamically allocate the number of qubits based on the variance of each dimension, rationally allocate quantum resources, and encode the decomposed feature components into quantum state feature vectors, providing richer and more efficient information representation for the subsequent decision-making process. The quantum policy network module 202 can input the quantum state feature vectors into a quantum policy network constructed from block variable quantum circuits, insert synchronization gates, and dynamically adjust the synchronization gate operation parameters according to the clock signal deviation between the quantum processor and the classical controller to output the policy action probability distribution. This fully leverages the advantages of quantum computing in parallel processing and fast search, enabling the rapid and efficient generation of better policy action probability distributions. Furthermore, by dynamically adjusting the synchronization gate parameters, it effectively solves the synchronization problem when the quantum processor and the classical controller work together, ensuring the stability and reliability of the decision-making process.

[0147] The quantum gradient optimization module 203 can correct the quantum natural gradient by calculating the error between the policy action probability distribution and the target value function and using the annealing regularization term. This not only accelerates the convergence speed but also improves the adaptability of the quantum policy network to complex environments, enabling the system to find the optimal solution more quickly and accurately. The adaptive decision feedback module 204 executes adaptive decision actions corresponding to the policy action probability distribution, obtains reward signals from the environment, and dynamically triggers quantum feature recoding based on the cumulative error between the updated parameters and the reward signals. This allows the system to perceive environmental changes in real time and adjust the quantum state feature vector accordingly, ensuring that the output decision policy closely matches the actual scenario and further improving the timeliness and accuracy of the decision policy.

[0148] Through the collaborative work of the above modules, the system achieves end-to-end optimization from multidimensional data encoding and policy network optimization to decision feedback adjustment, significantly improving the accuracy, efficiency, and adaptability to complex environments of multidimensional data decision optimization, and providing strong technical support for complex decision-making scenarios in fields such as finance, healthcare, and intelligent transportation.

[0149] In one exemplary embodiment, the present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the multidimensional data adaptive decision optimization method based on quantum reinforcement learning proposed in this application. A multi-core processor is preferred to improve the parallel processing capability of the system. The memory provides sufficient temporary storage space to support program execution and data processing. The memory capacity should be large enough to accommodate a large amount of information and computational tasks.

[0150] In one exemplary embodiment, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the multidimensional data adaptive decision optimization method based on quantum reinforcement learning of this application.

[0151] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A multidimensional data adaptive decision optimization method based on quantum reinforcement learning, characterized in that, The method includes: The multidimensional input data is decomposed into quantum wavelet multiscale decomposition, and the number of qubits is dynamically allocated based on the variance of each dimension. The decomposed feature components are encoded into quantum state feature vectors. The multidimensional input data includes structured tabular data, time-series sensor data and multimodal image data. The quantum state feature vector is input into a quantum policy network constructed from block variable quantum circuits. A synchronization gate is inserted into the block variable quantum circuits to output the policy action probability distribution. The operating parameters of the synchronization gate are dynamically adjusted according to the clock signal deviation between the quantum processor and the classical controller. The error between the probability distribution of the policy action and the target value function is calculated, the quantum natural gradient is corrected using the annealing regularization term, the parameters of the block variable quantum circuit are optimized, and the updated parameters are generated. The adaptive decision action corresponding to the probability distribution of the strategy action is executed. The reward signal is obtained from the environment by executing the adaptive decision action. The quantum feature recoding is dynamically triggered according to the cumulative error between the updated parameters and the reward signal to generate an optimized quantum state feature vector and output the optimized decision strategy.

2. The method according to claim 1, characterized in that, The process of performing quantum wavelet multi-scale decomposition on multidimensional input data, dynamically allocating the number of qubits based on the variance of each dimension, and encoding the decomposed feature components into quantum state feature vectors includes: The multidimensional input data is subjected to quantum Haar wavelet transform to generate multi-resolution feature components; Calculate the variance of each dimension of the multidimensional input data, and dynamically allocate the number of qubits according to the variance using the following formula to obtain the dynamically allocated number of qubits: Among them, B j The number of qubits after dynamic allocation. Let Q be the variance of the j-th dimension of the multidimensional input data, and let Q be the total number of usable qubits in the system. Let V be the variance of the i-th dimension of the multidimensional input data, and D be the total number of dimensions of the multidimensional input data; If the total number of dynamically allocated qubits exceeds the total number of available qubits in the system, the number of dynamically allocated qubits is compressed proportionally to obtain the compressed number of qubits. Based on the compressed number of qubits, the multi-resolution feature components are encoded into the quantum state feature vector.

3. The method according to claim 2, characterized in that, The formula for calculating the quantum state eigenvector is as follows: Where, |ψ(w i )> represents the encoded quantum state feature vector, used to describe the encoded quantum state of the i-th sample, B j ′ represents the number of compressed qubits, R y A rotation gate about the y-axis, used to apply rotation operations about the y-axis to qubits, w ij To represent the multi-resolution feature components w k The j-th dimension of the i-th sample, w i For the multi-resolution feature components w k The i-th sample.

4. The method according to claim 1, characterized in that, The process of inputting the quantum state feature vector into a quantum policy network constructed from block variable quantum circuits, inserting synchronization gates into the block variable quantum circuits, and outputting the policy action probability distribution includes: The quantum state feature vector is input into the initial quantum register of the block variable quantum circuit as the input state of the quantum policy network; Based on the dimensionality distribution of the quantum state eigenvectors, the block-variable quantum circuit is divided into multiple sub-circuits using the following formula: M=D / Q block Where M is the total number of the sub-circuits, D is the dimensional distribution of the quantum state eigenvectors, and Q... block The maximum qubit capacity of each of the sub-circuits is defined, and each of the sub-circuits contains parameterized quantum gates, which are composed of a combination of rotating gates about the Z-axis and the X-axis. A CZ entanglement gate is inserted between adjacent sub-lines, and the block variable component sub-lines are constructed using the following formula: Wherein, U(θ) represents the overall operation of the block variable component sub-line, which consists of multiple sub-operations, U m (θ m ) represents the parameterized quantum gate operation of the m-th sub-circuit, θ m The parameter set of the sub-line is Entangle(m,m+1), which is the CZ entanglement gate inserted between adjacent sub-lines m and m+1. Based on the clock signal deviation between the quantum processor's clock signal and the classical controller's clock signal, the timing parameters of the synchronization gate are calculated, and the synchronization gate is constructed using the following formula: Among them, H sync Let be the Hamiltonian of the synchronization gate. Let Z be the Pauli operator for the i-th qubit; The synchronization gate is inserted into the segmented variable quantum circuit, and the clock signal deviation is calibrated to be less than a preset deviation threshold using quantum state tomography, and the probability distribution of the strategy action is output.

5. The method according to claim 1, characterized in that, The calculation of the error between the probability distribution of the strategy action and the target value function, the correction of the quantum natural gradient using the annealing regularization term, the optimization of the parameters of the block-variable quantum circuit, and the generation of updated parameters include: The mean square error between the probability distribution of the strategy actions and the target value function is calculated using the following formula: L(θ)=E s,a [(Q(s,a)-V(s)) 2 ] Where L(θ) is the mean square error loss function, E s,a Let Q(s,a) be the mathematical expectation of state s and action a, and let V(s) be the action value function obtained through quantum phase estimation, which represents the expected reward of performing action a in state s. V(s) is the target value function, which represents the long-term value of state s. Based on the gradient of the probability distribution of the policy action, calculate the quantum Fisher information matrix, and calculate the pseudo-inverse of the quantum Fisher information matrix to obtain the pseudo-inverse matrix; Dynamic annealing coefficients are generated based on the chaotic sensitivity of the parameters of the block variable quantum circuit, wherein the chaotic sensitivity is quantized by the Lyapunov exponent and the annealing coefficients are generated based on the Lyapunov exponent. Based on the pseudo-inverse matrix and the annealing coefficients, and using the annealing regularization term, the quantum natural gradient is corrected using the following formula to obtain the corrected quantum natural gradient: Among them, F -1 (θ) is the pseudo-inverse matrix. Let λ(t) be the gradient of the original loss function L(θ) with respect to the parameter θ, representing the direction of loss descent, and let λ(t) be the annealing coefficient. Based on the corrected quantum natural gradient, the parameters of the block variable quantum circuit are updated by suppressing multi-task gradient conflicts through quantum interference path selection, thus generating the updated parameters.

6. The method according to claim 1, characterized in that, The adaptive decision action corresponding to the probability distribution of the strategy action is executed. This involves obtaining a reward signal from the environment by executing the adaptive decision action, dynamically triggering quantum feature recoding based on the cumulative error between the updated parameters and the reward signal, generating an optimized quantum state feature vector, and outputting an optimized decision strategy, including: Based on the probability distribution of the strategy action, the adaptive decision action is selected through the Softmax decision function, wherein the decision probability in the Softmax decision function is determined by smoothing the probability distribution of the strategy action with a temperature coefficient. The adaptive decision-making action is executed, the reward signal is obtained from the environment, and the cumulative error of the reward signal is calculated using the following formula: in, Let r be the predicted reward value at time t based on the parameters of the segmented variable quantum line. t The reward signal at time t; If the accumulated error is greater than a preset threshold, the updated parameters are loaded into the block variable quantum circuit, and the quantum feature recoding is triggered. The encoding weights are adjusted according to the gradient direction of the reward prediction loss function, where the reward prediction loss is the mean square error of the predicted reward value and the corresponding reward signal, and the optimized quantum state feature vector is generated using the following formula: Among them, X new Let X be the optimized quantum state eigenvector, and let ⊙ represent the Hadamard product. To encode the weight gradient, which reflects the reward prediction loss L r The direction and rate of change relative to the encoding weight X; Based on the optimized quantum state feature vector and the updated parameters, the optimized decision strategy is output.

7. The method according to claim 4, characterized in that, The process involves inserting the synchronization gate into the segmented variable quantum circuit and calibrating the clock signal deviation to less than a preset deviation threshold using quantum state tomography, then outputting the strategy action probability distribution. The preset deviation threshold is 0.1 ns, and includes: The quantum state after the synchronization gate is activated is measured using the quantum state tomography technique, and the fidelity between the quantum state and the ideal quantum state is calculated using the following formula: Where Tr is the trace operation, used to sum the diagonal elements of a matrix, and ρ ideal Let ρ be the density matrix of the ideal quantum state. out Let be the density matrix of the quantum states after the synchronization gate is applied; If the fidelity is less than 0.99, the timing parameter of the synchronization gate is adjusted using the following formula: t←t·(1+0.1·(1-F)) Repeatedly adjust the timing parameters of the synchronization gate until the fidelity is greater than or equal to 0.99 and the clock signal deviation is less than 0.1ns, to obtain the corresponding calibrated synchronization gate parameters, and output the probability distribution of the strategy action.

8. A multidimensional data adaptive decision optimization system based on quantum reinforcement learning, characterized in that, The system includes: The quantum feature encoding module is used to perform quantum wavelet multi-scale decomposition on multidimensional input data, dynamically allocate the number of qubits based on the variance of each dimension, and encode the decomposed feature components into quantum state feature vectors. The multidimensional input data includes structured table data, time-series sensor data, and multimodal image data. A quantum policy network module is used to input the quantum state feature vector into a quantum policy network constructed from block variable quantum circuits, insert a synchronization gate into the block variable quantum circuits, and output the policy action probability distribution. The operating parameters of the synchronization gate are dynamically adjusted according to the clock signal deviation between the quantum processor and the classical controller. The quantum gradient optimization module is used to calculate the error between the probability distribution of the policy action and the target value function, correct the quantum natural gradient using the annealing regularization term, optimize the parameters of the block variable quantum circuit, and generate updated parameters. The adaptive decision feedback module is used to execute adaptive decision actions corresponding to the probability distribution of the policy action. By executing the adaptive decision actions, it obtains reward signals from the environment, dynamically triggers quantum feature recoding based on the cumulative error between the updated parameters and the reward signals, generates an optimized quantum state feature vector, and outputs an optimized decision strategy.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Fuzzy semantic coding method, device and equipment based on electroencephalogram signals and medium

    CN118964964A

  • Internet of Things intelligent detection method for electric power system

    CN119827899A