Control method and device of quantum voltage metering device, terminal equipment and storage medium

By using an adaptive control model and reinforcement learning algorithm to dynamically adjust compressor parameters, the problem of low compressor control efficiency in quantum voltage metering devices is solved, achieving efficient and stable cooling effects and fault prediction, thus improving the accuracy and reliability of the device.

CN121028539APending Publication Date: 2025-11-28MEASUREMENT CENT OF GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511176937.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing electronic voltage metering devices have low compressor control efficiency, making it difficult to respond to temperature requirements and environmental changes in real time, resulting in unstable cooling and a tendency to malfunction.

Method used

An adaptive control model combined with reinforcement learning algorithm is adopted. By acquiring compressor operating data in real time, target control action matching and optimal control action selection are performed, and compressor parameters, including flow rate and frequency, are dynamically adjusted to achieve efficient cooling.

Benefits of technology

This improves the operating efficiency and stability of the compressor, reduces the failure rate, and ensures the high accuracy and reliability of the quantum voltage metering device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121028539A_ABST
    Figure CN121028539A_ABST
Patent Text Reader

Abstract

The invention discloses a control method and device of a quantum voltage metering device, terminal equipment and a storage medium, and relates to the technical field of compressor control, and the method comprises the steps: obtaining compressor real-time operation data of the quantum voltage metering device; inputting the real-time operation data of the compressor into a self-adaptive control model to enable the self-adaptive control model to perform target control action matching on the real-time operation data of the compressor, and taking the successfully matched target control action as an optimal control action; the target control action comprises a control action corresponding to the operation data of each compressor at the maximum return evaluation value; and performing operation control on the quantum voltage metering device based on the optimal control action. Therefore, by implementing the invention, the problem of low control precision of the compressor of the quantum voltage metering device in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of compressor control, and in particular to a control method and device of a quantum voltage measurement device, a terminal device and a storage medium. BACKGROUND

[0002] The quantum voltage measurement device is a key technology for realizing voltage reference, and is widely used in national metrology standard institutions, high-precision measurement laboratories and metrology scientific research fields. As the core component of a low-temperature refrigeration system, the compressor provides necessary ultra-low temperature by compressing and circulating the refrigerant, so as to ensure the stability of the quantum effect of the superconducting device, thereby ensuring the precision of the quantum voltage. The fixed mode cooling system currently used is a traditional refrigeration technology widely used in industrial and laboratory environments, and its operation is based on preset parameters such as the frequency of the compressor, the refrigerant flow and the pressure. This system maintains the required cooling effect through fixed control logic, has the characteristics of simple structure and easy implementation, and is widely used in scenes with relatively stable demand.

[0003] The current quantum voltage measurement device usually adopts a fixed mode cooling system, and the operating parameters of the compressor such as frequency and refrigerant flow are mostly preset values. The operating parameters of the compressor are adjusted by manual adjustment, but this method has a lag and is difficult to respond to changes in temperature demand, system load and environmental temperature in real time. Therefore, the quantum voltage measurement device of the prior art has the problem of low control efficiency of the compressor, and the compressor is prone to cooling, which may cause the quantum voltage measurement device to malfunction. SUMMARY

[0004] The present application provides a control method and device of a quantum voltage measurement device, a terminal device and a storage medium, which can solve the problem of low control efficiency of the compressor of the quantum voltage measurement device in the prior art.

[0005] The present application provides a control method of a quantum voltage measurement device, comprising:

[0006] obtaining real-time operating data of a compressor of the quantum voltage measurement device;

[0007] inputting the real-time operating data of the compressor into an adaptive control model, so that the adaptive control model matches the target control action of the real-time operating data of the compressor, and the target control action that is successfully matched is taken as an optimal control action; the target control action includes a control action corresponding to each compressor operating data at a maximum return evaluation value;

[0008] performing operating control on the quantum voltage measurement device based on the optimal control action;

[0009] The training of the adaptive control model comprises: initializing compressor operation data and reward evaluation values of the quantum voltage metering device to obtain initial compressor operation data and initial reward evaluation values; taking the initial compressor operation data and the initial reward evaluation values as inputs of a first reinforcement learning algorithm training operation, repeatedly performing the reinforcement learning algorithm training operation, then stopping the reinforcement learning algorithm training operation when the updated reward evaluation values are unchanged, and selecting a target control action occurring in the repeated reinforcement learning algorithm training operation as an optimal control action corresponding to each compressor operation data, thereby completing the training of the adaptive control model; wherein, in each execution of the reinforcement learning algorithm training operation, a control action is selected according to the current input, the compressor operation data and the reward evaluation values are updated by executing the control action.

[0010] As an improvement of the above scheme, the reinforcement learning algorithm training operation specifically comprises:

[0011] selecting a control action based on the current input through a greedy strategy;

[0012] after executing the control action on the quantum voltage metering device corresponding to the current input, obtaining updated compressor operation data of the quantum voltage metering device after executing the action, and calculating the reward of the quantum voltage metering device after executing the action;

[0013] determining updated reward evaluation values based on the reward and the reward evaluation values corresponding to the current input;

[0014] judging whether the updated reward evaluation values are the same as the reward evaluation values corresponding to the current input;

[0015] if the same, the current updated reward evaluation values are unchanged;

[0016] if not the same, taking the updated compressor operation data and the updated reward evaluation values as inputs of the next reinforcement learning algorithm training operation.

[0017] As an improvement of the above scheme, the control action comprises one of a random action or an optimal action; and the selecting a control action based on the current input through a greedy strategy comprises:

[0018] obtaining a plurality of executable actions;

[0019] determining an optimal action according to the reward evaluation values corresponding to the current input;

[0020] randomly selecting an action from all executable actions as a random action;

[0021] In the optimal action and the random action, a preset probability is selected, and a control action is obtained based on a selected result; wherein a selection probability of the optimal action is a first probability threshold, and a selection probability of the random action is a second probability threshold.

[0022] As an improvement of the above scheme, the reward and the return evaluation value corresponding to the current input are substituted into an update formula to obtain an updated return evaluation value; wherein the update formula is specifically:

[0023] As an improvement of the above scheme, the reward and the return evaluation value corresponding to the current input are substituted into an update formula to obtain an updated return evaluation value; wherein the update formula is specifically:

[0024]

[0025] In the formula, Q'(s t ,a t ) is the updated return evaluation value, Q(s t ,a t ) is the return evaluation value corresponding to the current input, r t is the reward, is the predicted future return of the compressor update operation data corresponding to the optimal action, a is the learning rate, s t is the current input of the compressor operation data, a t is the current control action, s t+1 is the compressor update operation data, and a' is the next control action.

[0026] As an improvement of the above scheme, the reward of the quantum voltage metering device after the execution of the action is calculated, including:

[0027] The temperature deviation, flow deviation, load and vibration data of the quantum voltage metering device after the execution of the action are obtained, and the temperature deviation, flow deviation, load and vibration data are substituted into a reward calculation formula to obtain the reward of the quantum voltage metering device after the execution of the action; wherein the reward formula is specifically:

[0028]

[0029] In the formula, r t is the reward, ΔT is the temperature deviation, ΔF is the flow deviation, P t is the load, V t is the vibration data, w1 is the first weight coefficient, w2 is the second weight coefficient, w3 is the third weight coefficient, w4 is the fourth weight coefficient, T target is the optimal temperature threshold, and F optimal is the optimal flow threshold.

[0030] As an improvement of the above scheme, the embodiment further includes:

[0031] The compressor operation data is input into a fault prediction model to obtain a fault prediction result;

[0032] Based on the prediction result, a fault maintenance strategy is determined, and the quantum voltage metering device is maintained based on the fault maintenance strategy.

[0033] As an improvement of the above scheme, the fault prediction model satisfies the following conditions:

[0034]

[0035] In the formula, P fault (s t ) is the probability of failure of the current s t , f i (s t ) is a feature function related to failure, and alpha i is a weight coefficient, and sigma is an activation function.

[0036] Another embodiment of the application also provides a control device for a quantum voltage metering device, comprising a data acquisition module, a matching module, and a control module.

[0037] The data acquisition module is configured to acquire real-time operation data of a compressor of the quantum voltage metering device.

[0038] The matching module is configured to input the real-time operation data of the compressor into an adaptive control model, so that the adaptive control model matches a target control action of the real-time operation data of the compressor, and the target control action that is successfully matched is taken as an optimal control action; the target control action includes a control action corresponding to each compressor operation data at a maximum return evaluation value.

[0039] The control module is configured to perform operation control on the quantum voltage metering device based on the optimal control action.

[0040] The training of the adaptive control model includes: initializing the compressor operation data and the return evaluation value of the quantum voltage metering device to obtain initial compressor operation data and initial return evaluation value; taking the initial compressor operation data and the initial return evaluation value as inputs of a first reinforcement learning algorithm training operation, repeatedly performing the reinforcement learning algorithm training operation, and then stopping the reinforcement learning algorithm training operation when the updated return evaluation value is unchanged, and selecting a target control action that appears in the repeated reinforcement learning algorithm training operation as an optimal control action corresponding to each compressor operation data, to complete the training of the adaptive control model; wherein, in each execution of the reinforcement learning algorithm training operation, a control action is selected according to the current input, and the compressor operation data and the return evaluation value are updated by executing the control action.

[0041] The application further provides a terminal device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the computer program is executed by the processor, the steps of the control method of the quantum voltage metering device are implemented.

[0042] The application further provides a computer readable storage medium, comprising a stored computer program, and when the computer program is executed, the device where the computer readable storage medium is located executes the steps of the control method of the quantum voltage metering device.

[0043] The application has the following beneficial effects:

[0044] The application obtains the current compressor operation data of the quantum voltage metering device, inputs the compressor operation data into the adaptive control model, and determines the optimal control action. The adaptive control model is trained through repeated execution of the reinforcement learning algorithm training operation, and the initial operation data and the initial reward evaluation value are determined by initializing the compressor operation data and the reward evaluation value, and are used as the input of the first reinforcement learning algorithm training operation to start the reinforcement learning algorithm training operation. In each execution of the reinforcement learning algorithm training operation, the compressor operation data and the reward evaluation value are updated by executing the selected control action, the reward evaluation value is used as the judgment standard of the control action, and the control action corresponding to each compressor operation data at the maximum reward evaluation value is used as the optimal control action, and the training of the adaptive control model is completed. The adaptive control model responds to the compressor operation data to automatically select the control action, can respond to the operation change of the compressor at any time, and improves the operation efficiency of the compressor. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0046] Figure 1 is a flowchart of the control method of the quantum voltage metering device provided by an embodiment of the application;

[0047] Figure 2 is a structural schematic diagram of the control device of the quantum voltage metering device provided by an embodiment of the application;

[0048] Figure 3 is an algorithm flowchart provided by the application.

[0049] Figure 4 is a comparison of the algorithm provided by an embodiment of the present application and a conventional algorithm in terms of refrigeration efficiency;

[0050] Figure 5 is a comparison of the algorithm provided by an embodiment of the present application and a conventional algorithm in terms of failure rate. DETAILED DESCRIPTION

[0051] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this application; the description and the claims of this application and the above description of drawings, the terms "comprising" and "having", and any variations thereof, are intended to cover not exclusively containing.

[0053] In the description of the embodiments of the present application, the technical terms "first", "second", etc. are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "a plurality of" is two or more, unless otherwise explicitly and specifically limited.

[0054] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearance of the phrase in various places in the specification does not necessarily all refer to the same embodiment, or necessarily alternatives to other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0055] In the description of the embodiments of the present application, the term "and / or" is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " in this paper generally represents that the front and rear associated objects are a "or" relationship.

[0056] In the description of the embodiments of the present application, the term "a plurality of" refers to two or more (including two), and similarly, "a plurality of groups" refers to two or more groups (including two groups), and "a plurality of pieces" refers to two or more pieces (including two pieces).

[0057] In the description of the embodiments of the present application, unless otherwise explicitly specified and limited, the technical terms "mounting", "connecting", "connecting", "fixing" and the like should be understood in a broad sense, for example, it can be fixedly connected, or it can be detachably connected, or it can be integrated; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the internal communication of two elements or the interaction relationship between two elements. For those skilled in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to the specific circumstances.

[0058] Reference Figure 1 To solve the problem of low compressor control precision of quantum voltage measurement device in the prior art, an embodiment of the present application provides a control method of quantum voltage measurement device, comprising:

[0059] 101, obtaining real-time operation data of the compressor of the quantum voltage measurement device;

[0060] In a specific embodiment, the real-time operation data of the compressor includes but is not limited to temperature, coolant flow, pressure and vibration.

[0061] 102, inputting the real-time operation data of the compressor into an adaptive control model, so that the adaptive control model matches the target control action of the real-time operation data of the compressor, and the target control action that matches successfully is taken as the optimal control action; the target control action includes the control action corresponding to each compressor operation data at the maximum return evaluation value;

[0062] The training of the adaptive control model includes: initializing the compressor operation data and the return evaluation value of the quantum voltage measurement device to obtain the initial operation data and the initial return evaluation value of the compressor; taking the initial operation data and the initial return evaluation value of the compressor as the input of the first reinforcement learning algorithm training operation, repeatedly executing the reinforcement learning algorithm training operation, then stopping executing the reinforcement learning algorithm training operation when the updated return evaluation value is unchanged, and selecting the target control action appearing in the repeated execution of the reinforcement learning algorithm training operation as the optimal control action corresponding to each compressor operation data, to complete the training of the adaptive control model; wherein, at each execution of the reinforcement learning algorithm training operation, a control action is selected according to the current input, and the compressor operation data and the return evaluation value are updated by executing the control action.

[0063] In a specific embodiment, the embodiment stores the basic data and actions by building state space and action space, specifically:

[0064] State space construction: the compressor operating data is stored in the state space, and various operating parameters of the compressor are collected in real time through sensors and monitoring systems. These parameters include but are not limited to temperature, flow, pressure and vibration; among them, temperature: monitor the internal temperature of the compressor to ensure its stable cooling state; flow: detect the flow of internal coolant to ensure that the flow remains within a reasonable range; pressure: monitor the internal pressure of the equipment to avoid excessive or low pressure affecting system operation; vibration: detect the vibration of the equipment to discover equipment abnormalities, equipment damage or unstable operation in time. Based on the above parameters, the state space of the compressor is constructed, that is, different operating state combinations are defined by using these key parameters, so that the state of the compressor at different time points can be accurately described.

[0065] Action space definition: the control actions that can be used to adjust the compressor are stored in the action space, which defines a series of control actions that can be used to adjust the compressor. These actions will be the decision output of the reinforcement learning algorithm, mainly including: flow adjustment and compression frequency adjustment; among them, flow adjustment: adjust the flow of cooling capacity according to the current cooling demand to maintain the target temperature; compression frequency adjustment: adjust the working frequency of the server according to the load condition to optimize energy efficiency. These actions specify the adjustment measures that the server control system can take in different states, so that the system has flexible response ability.

[0066] It can be understood that the construction of state space and action space provides basic data for the subsequent training of reinforcement learning algorithm. Through real-time collection of state data and corresponding action selection, the algorithm can continuously learn and adjust in actual operation to meet different working condition requirements and improve the cooling effect and efficiency of the radiator.

[0067] As an improvement of the above scheme, the reinforcement learning algorithm training operation is specifically:

[0068] Based on the current input, a control action is selected by a greedy strategy;

[0069] After performing the control action on the quantum voltage measurement device corresponding to the current input, the updated operating data of the quantum voltage measurement device after performing the action is obtained, and the reward of the quantum voltage measurement device after performing the action is calculated;

[0070] Based on the reward and the return evaluation value corresponding to the current input, an updated return evaluation value is determined;

[0071] Determine whether the updated return evaluation value is the same as the return evaluation value corresponding to the current input;

[0072] If the same, the current update return evaluation value is unchanged;

[0073] If not the same, the compressor update running data and the update return evaluation value are taken as inputs of the next reinforcement learning algorithm training operation.

[0074] As an improvement of the above scheme, the control action comprises one of a random action or an optimal action; the control action is selected by a greedy strategy based on the current input, comprising:

[0075] Obtaining a plurality of executable actions;

[0076] Determining the optimal action according to the return evaluation value corresponding to the current input;

[0077] Randomly selecting an action as the random action from all executable actions;

[0078] Selecting the optimal action and the random action according to a preset probability, and obtaining the control action based on the selected result; wherein the selection probability of the optimal action is a first probability threshold, and the selection probability of the random action is a second probability threshold.

[0079] As an improvement of the above scheme, the update return evaluation value is determined based on the reward and the return evaluation value corresponding to the current input, comprising:

[0080] Substituting the reward and the return evaluation value corresponding to the current input into an update formula to obtain the update return evaluation value; wherein the update formula is specifically:

[0081]

[0082] In the formula, Q'(s t ,a t ) is the update return evaluation value, Q(s t ,a t ) is the return evaluation value corresponding to the current input, r t is the reward, is the predicted future return of the optimal action corresponding to the compressor update running data, α is a learning rate, s t is the current input compressor running data, a t is the current control action, s t+1 is the compressor update running data, and a' is the next control action.

[0083] In a specific embodiment, when a completely new state s t+1At this time, it is necessary to solve the problem of unknown Q value by initialization. Specifically, before the learning process begins, the algorithm will create a data table for storing the Q value of all "state-action" pairs, or approximate it by a function. It should be noted that this table is not empty at the beginning, but is initialized to a specific value, that is, all Q values are initialized to zero or some small random numbers.

[0084] Therefore, the corresponding initial value can be obtained even if an unknown state is encountered. Even if this calculation is not accurate, a learning cycle is completed, and Q'(s t ,a t ) is more accurate.

[0085] As an improvement of the above scheme, the reward of the quantum voltage metering device after performing the action includes:

[0086] Obtain the temperature deviation, flow deviation, load and vibration data of the quantum voltage metering device after performing the action, and substitute the temperature deviation, flow deviation, load and vibration data into the reward calculation formula to obtain the reward of the quantum voltage metering device after performing the action; wherein the reward formula is specifically:

[0087]

[0088] In the formula, r t is the reward, ΔT is the temperature deviation, ΔF is the flow deviation, P t is the load, V t is the vibration data, w1 is the first weight coefficient, w2 is the second weight coefficient, w3 is the third weight coefficient, w4 is the fourth weight coefficient, T target is the optimal temperature threshold, F optimal is the optimal flow threshold.

[0089] In a specific embodiment, the core of the adaptive control model is to use the Q-learning algorithm to perform adaptive control on the compressor, so that it can dynamically adjust the operating parameters under different working conditions to achieve the best cooling effect and energy efficiency.

[0090] Step 2.1, initialization: define the state space S of the compressor, the action space A, the reward function R (used to calculate the reward), the discount factor γ, the learning rate α and other parameters. Initialize the Q value function Q(s, a) (used to calculate the return evaluation value). Usually, the Q value is random or set to zero at the beginning.

[0091] Step 2.2, state collection: at each time step t, the system collects the current state s t, including temperature, coolant flow, pressure, vibration, etc. These data come from sensors and monitoring systems to determine the current compressor state.

[0092] Step 2.3, select action: based on the current state s t and existing Q value (i.e. the reward evaluation value corresponding to the current input), an action a t is selected using the ε-greedy strategy

[0093] Step 2.3.1. Select a random action with probability ε (i.e. the second probability threshold) to maintain the possibility of exploring new strategies.

[0094] Step 2.3.2. Select the action with the highest current Q value (i.e. the optimal action) with probability 1-ε (i.e. the first probability threshold) to achieve the use of learned strategies.

[0095] Step 2.3.3 The action a t may be adjusting the cooling flow, compression frequency, load, etc. control measures.

[0096] Step 2.4, execute action and observe reward: execute action s t in state s t , the compressor enters the next state s t+1 (i.e. the compressor updates the running data), and obtains the immediate reward r t . The reward function is calculated according to the temperature stability, energy efficiency, vibration level, etc. index, to encourage the compressor to tend to the optimized running state.

[0097] Step 2.5, update Q value (i.e. update reward evaluation value): update Q1(s t ,a t ) by the following formula;

[0098] Q2(s t ,a t )←Q1(s t ,a t )+α[Q1(s t ,a t )]β

[0099] where Q2(s t ,a t ) is the updated reward evaluation value without reward, Q1(s t ,a t ) is the reward evaluation value without reward corresponding to the current input, r t is the reward, α is the learning rate, s t is the compressor running data of the current input, a t is the current control action, st+1 Update the compressor operation data.

[0100] Step 2.6, repeat steps 2.2-2.5, the system obtains a new state at each time step, selects and executes an action, and updates the Q value. In continuous interaction, the control strategy of the compressor is gradually optimized to achieve the goal of maximizing cooling effect, improving energy efficiency, and reducing fault risk.

[0101] Step 2.7, policy optimization is complete: after a large number of iterations, the Q value converges, and the system forms an optimal policy that can automatically select the optimal action in different states to ensure efficient operation of the compressor.

[0102] Based on the adaptive control model, the application of reinforcement learning algorithm is further refined by designing a comprehensive reward function to guide the Q-learning algorithm to optimize the control strategy of the compressor. The design goal of the reward function is to balance multiple factors, including: cooling effect: encourage the compressor to control the temperature within the target range to ensure cooling performance. Energy efficiency: encourage the compressor to reduce energy consumption and improve overall operating efficiency. Load and vibration control: encourage the compressor to reduce load and vibration to extend the service life of the equipment. Fault prevention: encourage the compressor to avoid entering high-risk states to ensure the safety and stability of the system.

[0103] Step 3.1, determine the design goal of the reward function. The design of the reward function mainly revolves around the following four core goals:

[0104] Step 3.1.1 Cooling effect: keep the temperature of the compressor within the target temperature T target nearby, ensure cooling performance.

[0105] Step 3.1.2 Energy efficiency: minimize the energy consumption of the compressor to improve overall operating efficiency.

[0106] Step 3.1.3 Load and vibration control: reduce excessive load and vibration of the compressor to extend the service life of the equipment.

[0107] Step 3.1.4 Fault prevention: avoid the compressor entering high-risk states to ensure the safety and stability of the system.

[0108] Step 3.2, define the performance indicators related to the goals. In order to quantify each goal, the corresponding performance indicators need to be defined: temperature deviation ΔT: represents the difference between the current temperature and the target temperature, usually wants to be as close to 0 as possible. Flow deviation ΔF: represents the difference between the current coolant flow and the optimal flow, used to measure the accuracy of flow control. Load P t : the current working load of the compressor, excessive load will increase energy consumption and accelerate equipment wear. Vibration level V tVibration can pose a threat to the stability and safety of the equipment, so it needs to be monitored.

[0109] Step 3.3, design the formula of the reward function, convert the above objectives into specific mathematical expressions in the reward function. The reinforcement learning model is driven by the reward function to achieve optimal control, usually using a weighted combination formula to integrate various objectives to ensure that the compressor achieves stable and efficient operation:

[0110]

[0111] In the formula, the temperature control part w4, when the temperature deviation ΔT decreases, the reward value increases. This encourages the model to keep the temperature close to the target value. The flow control part When the flow deviation ΔF decreases, the reward value increases, ensuring flow optimization. The vibration control part -w3·|V t |, the higher the vibration level, the less the reward, encouraging the model to reduce mechanical vibration. The load control part -w4·P t , the higher the load, the less the reward, encouraging the model to reduce the compressor load while meeting work requirements.

[0112] Step 3.4, determine the weight coefficients w1, w2, w3, w4, select appropriate weight coefficients to balance the importance of each objective in the reward function. These coefficients can be tuned through experiments or historical data to ensure that the control strategy meets the expected effect in actual operation. For example: if the cooling effect has high priority, you can increase w1, if the long-term health and risk of failure of the equipment are critical, you can appropriately increase w3 and w4.

[0113] Step 3.5, real-time adjustment and optimization of the reward function, in actual operation, the reward function can be dynamically adjusted according to changes in the environment and working conditions to ensure that the reinforcement learning model can effectively optimize the control strategy under different load conditions.

[0114] Apply the results of the comprehensive reward function to the Q-learning update algorithm, continuously update the Q value to gradually optimize the compressor control strategy, and ultimately achieve efficient and stable operation. Through Q value updating and strategy optimization, the present application can effectively improve the operating efficiency and control accuracy of the compressor in the quantum voltage measurement device, providing reliable protection for high-precision voltage measurement, and has important application value and innovation. Specifically:

[0115] Step 4.1, define the Q value function, Q value represents the long-term expected return obtained by taking action in a certain state. As training progresses, the Q value function will gradually approach the optimal Q value, i.e. the optimal return in each state-action pair.

[0116] Step 4.2, Action Selection, at each time step t, the reinforcement learning agent (i.e., the subject performing the operations of the reinforcement learning algorithm) selects an action a t based on the current Q-values Q(s t ). Typically, a greedy policy is used, with Exploration: selecting a random action with probability ∈, and Exploitation: selecting the action with the highest current Q-value, i.e., the optimal action, with probability 1-∈, maximizing the current reward. Through the combination of exploration and exploitation, the agent can avoid falling into local optimal solutions while gradually optimizing the policy.

[0117] Step 4.3, Action Execution and Observation of New State and Reward, the selected action a t is executed in state s t , the compressor system adjusts its state according to the control instruction and enters the next state s t+1 , and returns an immediate reward. This reward is calculated by the reward function designed earlier, reflecting the impact of the action on the system's operation.

[0118] Step 4.4, Q-value Update Formula, after executing the action, the Q-learning update formula is used to iteratively update the Q-value, calculating the new Q-value:

[0119]

[0120] where Q'(s t ,a t ) is the updated reward evaluation value, Q(s t ,a t ) is the current input corresponding reward evaluation value, r t is the reward, is the predicted future reward of the optimal action corresponding to the compressor updated running data, α is the learning rate, s t is the current input compressor running data, a t is the current control action, s t+1 is the compressor updated running data, and a' is the next control action.

[0121] Step 4.5, Iterative Update, gradually optimizing the policy The Q-value is closer to the optimal Q-value after each update. After multiple iterations, the Q-value function gradually converges, eventually forming a set of stable Q-values, which encode the optimal decisions in each state.

[0122] Step 4.6, Policy Extraction, after the Q-value function converges, the optimal policy π * can be directly extracted from the Q-value using the greedy policy:

[0123] π *(s) = argmax a Q(s, a)

[0124] That is, in each state S, the action that maximizes the Q value is selected, which is the optimal strategy for that state.

[0125] Step 4.7, the cycle with policy optimization is completed, and the process will continue to iterate until the Q value function converges. In operation, the reinforcement learning model will continuously update the Q value and optimize the policy according to the changes in the environment to adapt to the control requirements of the compressor under different working conditions.

[0126] Finally, through Q value update and policy optimization, the agent can learn an adaptive control strategy to achieve efficient operation of the compressor system and reduce the risk of failure.

[0127] As an improvement of the above scheme, the embodiment further comprises:

[0128] inputting the compressor operation data into the fault prediction model to obtain a fault prediction result;

[0129] determining a fault maintenance strategy based on the prediction result, and performing fault maintenance on the quantum voltage metering device based on the fault maintenance strategy.

[0130] As an improvement of the above scheme, the fault prediction model satisfies the following conditions:

[0131]

[0132] In the formula, P fault (s t ) is the probability of failure occurring at the current s t , f i (s t ) is a feature function related to failure, a i is a weight coefficient, and σ is an activation function. n is the total number of features.

[0133] In a specific embodiment, the feature function f i (s t ) is a special function covering different running times, temperatures, flow rates, loads, and vibrations, which needs to be designed based on engineering experience and operation requirements.

[0134] For example: f1(s t ) = f(current temperature - standard working temperature), f2(s t ) = f(equipment maintenance-free working time);

[0135] Through the current state s t and f i (s t) are calculated and weighted summed by a i and finally mapped to the range 0-1 by a sigmoid function, resulting in the current s t probability of failure.

[0136] Step 103: Perform operational control on the quantum voltage metering device based on the optimal control action.

[0137] In a specific embodiment, based on a large amount of operation data accumulated in the above steps, a fault prediction model is established to realize active monitoring and prediction of the health status of the compressor. The model identifies potential failure risks by comparing the current state with the normal operation mode, and performs predictive maintenance such as adjusting control parameters, replacing wear parts, and scheduling maintenance plans according to the prediction results, thereby improving the reliability and safety of the equipment and prolonging the service life of the equipment. Among them, the establishment of the fault prediction model realizes the active monitoring and prediction of the health status of the compressor, which is specifically:

[0138] Step 5.1, data accumulation and learning: the system accumulates long-term operation data of the compressor, including state parameters, control parameters, fault records, etc. The reinforcement learning model uses these data to learn the normal operation mode and potential failure characteristics of the compressor.

[0139] Step 5.2, state monitoring: the system monitors the state parameters of the compressor in real time, such as temperature, coolant flow, pressure, vibration, etc., and analyzes them. By comparing the current state with the normal operation mode, potential failure risks are identified.

[0140] Step 5.3, fault prediction model: based on the reinforcement learning model and the state monitoring results, a fault prediction model is established. The model can predict the probability of failure of the compressor and issue an early warning.

[0141]

[0142] where P fault (s t ) is the probability of failure of the current state s t , f i (s t ) is a feature function related to failure, a i is a weight coefficient, and σ is an activation function that outputs a failure prediction probability between 0 and 1.

[0143] Step 5.4, predictive maintenance: based on the fault prediction results, the system can take maintenance measures in advance, including adjusting control parameters, replacing wear parts, and scheduling maintenance plans.

[0144]

[0145] where T is a preset period, M(s t ) represents a maintenance decision made based on the current state and historical experience, λ is a decay factor, and Q(s t ,a t ) is the Q value of each state-action pair.

[0146] As shown in Figure 2 , on the basis of the above method embodiment, corresponding device embodiments are provided.

[0147] An embodiment of the present application provides a control device of a quantum voltage metering device, comprising a data acquisition module 201, a matching module 202 and a control module 203.

[0148] The data acquisition module is used for acquiring real-time operation data of a compressor of the quantum voltage metering device.

[0149] The matching module is used for inputting the real-time operation data of the compressor into an adaptive control model, so that the adaptive control model matches a target control action of the real-time operation data of the compressor, and the target control action that is matched successfully is used as an optimal control action.

[0150] The control module is used for performing operation control on the quantum voltage metering device based on the optimal control action.

[0151] The training of the adaptive control model comprises the following steps: initializing compressor operation data and a reward evaluation value of the quantum voltage metering device to obtain initial compressor operation data and an initial reward evaluation value; taking the initial compressor operation data and the initial reward evaluation value as input of a first reinforcement learning algorithm training operation, repeatedly performing the reinforcement learning algorithm training operation, then stopping the reinforcement learning algorithm training operation when the updated reward evaluation value is unchanged, and selecting a target control action that appears in the repeated reinforcement learning algorithm training operation as an optimal control action corresponding to each compressor operation data, thereby completing the training of the adaptive control model.

[0152] In a specific embodiment, referring to Figure 3, by constructing a space containing parameters such as temperature, coolant state, pressure and vibration, and defining the corresponding control action, through the reinforcement learning algorithm (Q-learning), the system can automatically select the optimal control strategy under different states, dynamically adjust the router operating parameters to optimize the cooling effect and energy efficiency, while reducing the failure rate and prolonging the service life of the equipment, thereby improving the accuracy of the quantum voltage measurement device. The specific implementation is as follows.

[0153] The ultimate goal of the present application is to improve the efficiency and accuracy, which is also a summary and sublimation of the results of the previous steps. Through the application of reinforcement learning algorithm, the compressor realizes adaptive control and can dynamically adjust the operating parameters under different working conditions to achieve the best cooling effect and energy efficiency. Through fault prediction, predictive maintenance is further realized, and the reliability and safety of the equipment are improved. These achievements ensure that the compressor is always in the best operating state, providing an efficient and stable low-temperature environment for the quantum voltage measurement device, thereby improving the accuracy and reliability of the quantum voltage reference, and providing reliable protection for high-precision voltage measurement. By introducing reinforcement learning algorithm, the operation control and state monitoring of the compressor in the quantum voltage measurement device are optimized. Specifically, first, define the operating state space of the compressor, including temperature, coolant flow, pressure, vibration and other key parameters, and build a dynamic state monitoring model based on these parameters. Then, through reinforcement learning algorithm, especially Q-learning, based on real-time collected operating data, dynamically adjust the control strategy of the compressor, such as cooling flow, compression frequency, etc., to achieve the best cooling effect and energy efficiency. Ensure efficient operation of the compressor under different working conditions. By continuously updating the Q value function, the reinforcement learning algorithm can adaptively adjust the compressor parameters and optimize the control strategy according to the real-time feedback of the system. In addition, the system also has a fault prediction function, which can identify potential fault risks through long-term data accumulation and realize predictive maintenance, effectively prolonging the service life of the equipment. This technical solution optimizes the operation control of the compressor through intelligentization, improves the stability and accuracy of the quantum voltage measurement device, reduces energy consumption and reduces the risk of failure downtime, and has important application value and innovation.

[0154] For better illustration, see Figure 4 , 5 , respectively, the comparison of refrigeration efficiency between the algorithm of the present application and the traditional algorithm, and the comparison of failure rate between the algorithm of the present application and the traditional algorithm.

[0155] It can be understood that the above device item embodiments correspond to the method item embodiments of the present application, which can realize the control method of the quantum voltage measurement device provided by any one of the above method item embodiments.

[0156] It should be noted that the apparatus embodiments described above are only illustrative, and part or all of the modules can be selected to achieve the purpose of the embodiment. In addition, the connection relationship between the modules in the apparatus embodiments provided by the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement it without creative labor.

[0157] On the basis of the above-mentioned embodiment of the control method of the quantum voltage measurement device, another embodiment of the present application provides a terminal device, which comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, the control method of the quantum voltage measurement device of any one of the embodiments of the present application is realized.

[0158] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the terminal device.

[0159] The terminal device can be a desktop computer, a notebook computer, a palm computer, a cloud server and other computing devices. The terminal device can include, but is not limited to, a processor and a memory.

[0160] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, and connects all parts of the terminal device through various interfaces and lines.

[0161] On the basis of the above method embodiment, another embodiment of the present application provides a computer readable storage medium, comprising a stored computer program, wherein the computer program controls the device where the computer readable storage medium is located to execute the control method of the quantum voltage metering device according to any one of the above method embodiments of the present application when running.

[0162] The modules / units integrated in the device / terminal equipment, if realized in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of the above-mentioned various method embodiments can be realized. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0163] The above is the preferred embodiment of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements are also considered to be within the scope of protection of the present application.

Claims

1. A control method for a quantum voltage metering device, characterized in that, include: Acquire real-time operating data of the compressor in the quantum voltage metering device; The real-time operating data of the compressor is input into the adaptive control model so that the adaptive control model matches the real-time operating data of the compressor with the target control action and takes the successfully matched target control action as the optimal control action. The target control actions include: the control actions corresponding to each compressor operating data at the maximum return evaluation value; The quantum voltage metering device is operated and controlled based on the optimal control action described above. The training of the adaptive control model includes: initializing the compressor operating data and reward evaluation value of the quantum voltage metering device to obtain the initial compressor operating data and initial reward evaluation value; using the initial compressor operating data and initial reward evaluation value as input for the first reinforcement learning algorithm training operation, repeatedly executing the reinforcement learning algorithm training operation, and then stopping the execution of the reinforcement learning algorithm training operation when the updated reward evaluation value remains unchanged, and selecting the target control action that appears in the repeated execution of the reinforcement learning algorithm training operation as the optimal control action corresponding to each compressor operating data, thus completing the training of the adaptive control model; wherein, in each execution of the reinforcement learning algorithm training operation, a control action is selected according to the current input, and the compressor operating data and reward evaluation value are updated respectively by executing the control action.

2. The control method for the quantum voltage metering device as described in claim 1, characterized in that, The reinforcement learning algorithm training operation is specifically as follows: Based on the current input, select a control action using a greedy strategy; After the quantum voltage metering device corresponding to the current input executes the control action, obtain the compressor update operation data of the quantum voltage metering device after the action is executed, and calculate the reward of the quantum voltage metering device after the action is executed; Based on the reward and the reward evaluation value corresponding to the current input, determine the updated reward evaluation value after the quantum voltage metering device performs a control action under the compressor operating data corresponding to the current input; Determine whether the updated return assessment value is the same as the return assessment value corresponding to the current input; If they are the same, the current updated return assessment value remains unchanged; If they are different, the compressor update operation data and update reward evaluation value will be used as input for the next reinforcement learning algorithm training operation.

3. The control method for the quantum voltage metering device as described in claim 2, characterized in that, The control action includes either a random action or an optimal action; the selection of a control action based on the current input using a greedy strategy includes: Obtain several executable actions; Determine the optimal action based on the reward assessment value corresponding to the current input; From all executable actions, one action is randomly selected as the random action. Among the optimal action and the random action, a selection is made according to a preset probability, and a control action is obtained based on the selection result; wherein, the selection probability of the optimal action is a first probability threshold, and the selection probability of the random action is a second probability threshold.

4. The control method for the quantum voltage metering device as described in claim 3, characterized in that, The process of determining the updated reward assessment value for the compressor operating data corresponding to the current input after executing the control action, based on the reward and the reward assessment value corresponding to the current input, includes: Substitute the reward and the corresponding return assessment value of the current input into the update formula to obtain the updated return assessment value; wherein, the update formula is specifically: In the formula, Q′(s t ,a t To update the return assessment value, Q(s) t ,a t ) represents the return assessment value corresponding to the current input, r t As a reward, Let Δ be the expected future reward for the optimal action corresponding to the compressor's updated operating data, and let s be the learning rate. t For the currently input compressor operating data, a t For the current control action, s t+1 Update the compressor's operating data; a′ is the next control action.

5. The control method for the quantum voltage metering device as described in claim 4, characterized in that, The reward for the quantum voltage metering device after the calculation execution action includes: After executing the action, acquire the temperature deviation, flow rate deviation, load, and vibration data of the quantum voltage metering device. Substitute these data into the reward calculation formula to obtain the reward for the quantum voltage metering device after executing the action. The specific reward formula is as follows: In the formula, r t For rewards, ΔT represents temperature deviation, ΔF represents flow rate deviation, and P represents... t For load, V t For vibration data, w1 is the first weighting coefficient, w2 is the second weighting coefficient, w3 is the third weighting coefficient, w4 is the fourth weighting coefficient, and T... target For the optimal temperature threshold, F optimal This is the optimal flow threshold.

6. The control method for the quantum voltage metering device as described in claim 5, characterized in that, Also includes: The compressor operating data is input into the fault prediction model to obtain the fault prediction results; Based on the prediction results, a fault maintenance strategy is determined, and fault maintenance is performed on the quantum voltage metering device based on the fault maintenance strategy.

7. The control method for the quantum voltage metering device as described in claim 6, characterized in that, The fault prediction model satisfies the following conditions: In the formula, P fault (s t ) is the current s t The probability of failure occurring, f i (s t α is a characteristic function related to the fault. i σ is the weight coefficient, and σ is the activation function.

8. A control device for a quantum voltage metering device, characterized in that, include: Data acquisition module, matching module, and control module; The data acquisition module is used to acquire real-time operating data of the compressor of the quantum voltage metering device; The matching module is used to input the real-time operating data of the compressor into the adaptive control model, so that the adaptive control model matches the real-time operating data of the compressor with the target control action, and takes the successfully matched target control action as the optimal control action. The target control actions include: the control actions corresponding to each compressor operating data at the maximum return evaluation value; The control module is used to control the operation of the quantum voltage metering device based on the optimal control action; The training of the adaptive control model includes: initializing the compressor operating data and reward evaluation value of the quantum voltage metering device to obtain the initial compressor operating data and initial reward evaluation value; using the initial compressor operating data and initial reward evaluation value as input for the first reinforcement learning algorithm training operation, repeatedly executing the reinforcement learning algorithm training operation, and then stopping the execution of the reinforcement learning algorithm training operation when the updated reward evaluation value remains unchanged, and selecting the target control action that appears in the repeated execution of the reinforcement learning algorithm training operation as the optimal control action corresponding to each compressor operating data, thus completing the training of the adaptive control model; wherein, in each execution of the reinforcement learning algorithm training operation, a control action is selected according to the current input, and the compressor operating data and reward evaluation value are updated respectively by executing the control action.

9. A terminal device, characterized in that, The device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements a control method for the quantum voltage metering device as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, include: A stored computer program, wherein, when the computer program is executed, the device containing the computer-readable storage medium is controlled to perform a control method for the quantum voltage metering device as described in any one of claims 1-7.