Adaptive Vibration Suppression Method and Device for Compressors Based on Reinforcement Learning
By applying reinforcement learning methods in the compressor, using the rotational fluctuation amplitude and vibration trend parameters to build a reinforcement learning model to generate a compensation current signal, solving the problem that traditional feedback control is difficult to suppress the vibration caused by gas pulsation, and achieving more efficient vibration suppression and equipment stability.
Patent Information
- Application Number
- CN202510479601.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-17
AI Technical Summary
In traditional compressors, it is difficult to effectively suppress vibration caused by gas pulsation through feedback control, especially when load fluctuations are not allowed to cope with complex working conditions.
Using a reinforcement learning method, by obtaining the rotation fluctuation amplitude and vibration trend quantization parameters of the compressor motor rotation cycle, a reinforcement learning model of the PPO algorithm is constructed to generate a compensation current signal to suppress vibration.
Without sensors, the compressor vibration can be accurately suppressed, the equipment operation stability and service life can be improved, and the equipment can be dynamically responded to different working conditions.
Smart Images

Figure CN119982484B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of machine learning, and particularly to a method for adaptively suppressing compressor vibration based on reinforcement learning, and a device for adaptively suppressing compressor vibration based on reinforcement learning. Background Art
[0002] The working process of a compressor is usually controlled by a frequency converter. The frequency converter controls the operation of the compressor by adjusting the speed of the motor. The compressor sucks in low-pressure gas and compresses the gas through mechanical components inside the compressor. During the compression process, the volume of the gas decreases and the pressure increases. The frequency converter dynamically adjusts the speed of the motor according to the demand to achieve efficient operation of the compressor under different loads. For example, when the load is light, the frequency converter reduces the speed to reduce energy consumption; while when the load increases, the frequency converter increases the speed to maintain the stability of the system pressure.
[0003] A compressor works by compressing gas. During the working process of the compressor, the compressed gas experiences fluctuations in pressure and flow rate in each working cycle. This pulsation is transmitted to the mechanical components of the compressor, causing the mechanical components to bear periodic mechanical stress in each pulsation impact. This repeated impact force accelerates the wear of the components, further causing the mechanical components to generate irregular mechanical vibrations. The periodic load changes caused by pulsation also lead to relative movement and impact of the compressor components. Especially when the pulsation frequency is close to the natural frequency of the compressor, more serious resonance phenomena may be triggered, resulting in increased vibration. The long-term accumulation of gas pulsation and vibration will lead to a decrease in compressor efficiency, shortening of component life, increased maintenance frequency, and even possible equipment failures.
[0004] In traditional compressors, pulsation is compensated by manually adjusting the frequency converter or PID control, that is, feedback control. By monitoring the vibration or other relevant variables of the compressor and adjusting control variables such as speed to reduce the deviation, the pulsation and vibration can be reduced to a certain extent. However, the effect of feedback control is very limited, especially when facing the pulsation fluctuations of the load, and it cannot cope with the complex working conditions in practical applications. Summary of the Invention
[0005] In view of the above problems, this application designs and provides a method for adaptively suppressing compressor vibration based on reinforcement learning.
[0006] A method for adaptively suppressing compressor vibration based on reinforcement learning includes the following steps: obtaining vibration quantity quantization parameters, where the vibration quantity quantization parameters include: the rotational fluctuation amplitude of the compressor motor rotation period , the rotational fluctuation amplitude of the compressor motor rotation period The difference between the maximum rotational speed and the minimum rotational speed in the rotational speed sequence corresponding to each rotation cycle of the compressor motor; obtaining a vibration trend quantization parameter based on the vibration quantity quantization parameter, where the vibration trend quantization parameter includes one or more of a sliding mean, a change rate, an accumulated change amount, and a standard deviation; constructing a reinforcement learning model based on the vibration quantity quantization parameter and the vibration trend quantization parameter; training the reinforcement learning model; deploying the reinforcement learning model, and according to the compensation torque current parameter and the compensation angle parameter , generating a compensation current signal , where the compensation current signal is , is a time variable, is an angular frequency, , is the rotational frequency of the compressor motor rotor; the frequency converter receives the compensation current signal , and superimposes the compensation current signal with the compressor motor current control signal, and uses the superimposed current control signal to drive the compressor motor to operate, so as to suppress the vibration caused by the fluctuations of the pressure and flow experienced by the compressed gas in each working cycle.
[0007] Further, the reinforcement learning model is constructed based on the PPO algorithm, and constructing the reinforcement learning model includes the following steps: constructing a state space, where the state space includes: the vibration quantity quantization parameter, the sliding mean, the change rate, the accumulated change amount, the standard deviation, the average rotational speed of the compressor motor, and the load current; constructing an action space, where the action space includes the compensation torque current parameter and the compensation angle parameter ; constructing a reward function, where the reward function is: ; where, is the immediate reward obtained at the time step , , and are weight coefficients; constructing a policy network; constructing a value network.
[0008] Further, training the reinforcement learning model includes the following steps: using the policy network of the current state to generate actions corresponding to the action space; evaluating the value of each action using generalized advantage estimation , and constructing a policy loss function based on the value; the policy loss function is:
[0009] ;
[0010] where, is the time step The expected value; is the probability that the current policy selects an action in the state ; is the probability that the old policy selects an action in the same state ; is a clipping operation that restricts the ratio within ;
[0011] Construct a PPO loss function, and the PPO loss function includes an overall loss function , and the overall loss function is weighted based on the policy loss function , the value loss function and the entropy regularization function ; the overall loss function satisfies:
[0012] ;
[0013] and are weight coefficients;
[0014] Among them, the value loss function is:
[0015] ;
[0016] Among them, is the expected value at time step ; is the predicted value of the value network in the state , is the target value output by the value network;
[0017] The entropy regularization function is:
[0018] ;
[0019] Among them, is the expected value at time step ; is the probability that the current policy selects an action in the state ; is the probability that the current policy selects an action in the state ;
[0020] Use the gradient descent algorithm to minimize the overall loss and update the parameters of the policy network and the value network;
[0021] After completing a policy update, save the parameters of the current policy as the parameters of the old policy until the preset number of policy updates is reached.
[0022] Further, the step of using the policy network of the current state to generate an action corresponding to the action space includes: keeping the compressor in an operating state, and collecting the current state once per rotation cycle of each compressor motor. ; According to the current state Output an action through the policy network ; Execute the action , obtain the time step of the immediate reward and the next state ; Store the current state , action , immediate reward and the next state in the experience replay buffer; randomly select samples from the experience replay buffer; use generalized advantage estimation to evaluate the value of each action of the randomly selected samples , and construct a policy loss function based on the value.
[0023] Further, the policy network includes four fully connected layers, each layer including 64 neurons; the value network includes four fully connected layers, each layer including 64 neurons.
[0024] The second aspect of the present application provides a compressor vibration adaptive suppression device based on reinforcement learning, including: a first acquisition unit, which is used to acquire vibration quantity quantization parameters, and the vibration quantity quantization parameters include: the rotational fluctuation amplitude of the compressor motor rotation cycle , the rotational fluctuation amplitude of the compressor motor rotation cycle is the difference between the maximum speed and the minimum speed in the speed sequence corresponding to each compressor motor rotation cycle;
[0025] A second acquisition unit, which is used to acquire vibration trend quantization parameters based on the vibration quantity quantization parameters, and the vibration trend quantization parameters include one or more of a sliding mean, a change rate, an accumulated change amount, and a standard deviation;
[0026] A construction unit, which is used to construct a reinforcement learning model based on the vibration quantity quantization parameters and the vibration trend quantization parameters;
[0027] A training unit, which is used to train the reinforcement learning model;
[0028] An application unit, which is used to deploy the reinforcement learning model, and according to the compensation torque current parameters output by the reinforcement learning model And compensation angle parameter to generate a compensation current signal The compensation current signal is , is a time variable, is the angular frequency, , is the rotational frequency of the compressor motor rotor;
[0029] A compensation unit, which is used to make the frequency converter receive the compensation current signal , and superimpose the compensation current signal with the compressor motor current control signal, and use the superimposed current control signal to drive the compressor motor to run, so as to suppress the vibration caused by the fluctuations of the pressure and flow experienced by the compressed gas in each working cycle.
[0030] Furthermore, the construction unit is used to construct a reinforcement learning model based on the PPO algorithm, which includes:
[0031] A state space construction part, which is used to construct a state space, and the state space includes: the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change amount, the standard deviation, the average rotational speed of the compressor motor, and the load current; an action space construction part, which is used to construct an action space, and the action space includes the compensation torque current parameter and compensation angle parameter ; a reward function construction part, which is used to construct a reward function, and the reward function is: , where is the time step The immediate reward obtained, , and are weight coefficients; a policy network construction part, which is used to construct a policy network;
[0032] A value network construction part, which is used to construct a value network.
[0033] Furthermore, the training unit includes: an action generation part, which is used to generate an action corresponding to the action space using the policy network of the current state; an evaluation part, which is used to evaluate the value of each action using the generalized advantage estimation , and construct a policy loss function based on the value; the policy loss function is:
[0034] ;
[0035] where is the expected value at time step ; is the current policy, and at state select action The probability; is the probability that the old policy selects an action in the same state ; is a clipping operation that restricts the ratio within the range;
[0036] A loss function construction unit that constructs a PPO loss function, where the PPO loss function includes an overall loss function , and the overall loss function is weighted based on the policy loss function , the value loss function and the entropy regularization function ; the overall loss function satisfies:
[0037] ;
[0038] and are weight coefficients;
[0039] Among them, the value loss function is:
[0040] ;
[0041] Among them, is the expected value of the time step ; is the predicted value of the value network in the state , is the target value output by the value network;
[0042] The entropy regularization function is:
[0043] ;
[0044] Among them, is the expected value of the time step ; is the probability that the current policy selects an action in the state ; is the logarithm of the probability that the current policy selects an action in the state ;
[0045] An update unit that uses the gradient descent algorithm to minimize the overall loss and update the parameters of the policy network and the value network;
[0046] A storage unit, which is used to save the parameters of the current policy as the parameters of the old policy after a policy update is completed; until a preset number of policy update times is reached.
[0047] Further, the steps for the action generation unit to generate an action corresponding to the action space using the policy network of the current state include: keeping the compressor in an operating state, and collecting the current state once per rotation cycle of each compressor motor ; according to the current state output an action through the policy network ; execute the action , and obtain the immediate reward at the time step and the next state ; store the current state , the action , the immediate reward and the next state in the experience replay buffer; randomly select samples from the experience replay buffer; the evaluation unit uses generalized advantage estimation to evaluate the value of each action of the randomly selected samples , and construct a policy loss function based on the value.
[0048] Further, the policy network includes four fully connected layers, and each layer includes 64 neurons; the value network includes four fully connected layers, and each layer includes 64 neurons.
[0049] Through real-time learning and optimization, the present application can accurately suppress compressor vibration without sensors. The present application uses the rotation fluctuation amplitude and vibration trend quantization parameters of the compressor motor to construct a reinforcement learning model and generate a compensation current signal, thereby effectively reducing the vibration caused by the pressure and flow fluctuations of the compressor during the working cycle, and improving the stability and service life of the equipment operation.
[0050] After reading the specific embodiments of the present application in conjunction with the accompanying drawings, other features and advantages of the present application will become clearer. Description of the Drawings
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 is a flowchart of the compressor vibration adaptive suppression method based on reinforcement learning provided by the present application;
[0053] Figure 2 The network architecture diagram of a specific example of the policy network;
[0054] Figure 3 The network architecture diagram of a specific example of the value network;
[0055] Figure 4 The structural schematic block diagram of the compressor vibration adaptive suppression method based on reinforcement learning provided by the present application. Specific implementation manners
[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0057] In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore, should not be construed as a limitation to the present application.
[0058] In the description of the present application, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "electrical connection" should be understood in a broad sense. For example, it can be a fixed electrical connection, a detachable electrical connection, or an integral electrical connection. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations. In the description of the implementation manners, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0059] The terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0060] In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.
[0061] Compressors are widely used in multiple industries and are mainly used for gas compression, playing a crucial role in refrigeration systems, achieving refrigeration (heating) effects by compressing gases. In the field of household appliances, the compressor-based core refrigeration system is not only applied in traditional air conditioners or refrigerators (including freezers), but also in new products such as dryers and heat pump systems. For example, in a dryer, the compressor is used to circulate the refrigerant, thereby achieving heat recovery and temperature control during the drying process, providing the required temperature and humidity for drying.
[0062] In a compressor, the motor is the core component that drives the compressor to operate. The motor provides power through rotation, driving the mechanical parts of the compressor (such as the rotor, etc.) to compress the gas. The rotational speed of the compressor motor is adjusted by an inverter (also known as the compressor drive circuit). The inverter can adjust the rotational speed of the motor in real time according to the load demand, enabling the motor to operate at the best efficiency.
[0063] During the operation of the compressor, the compressed gas will experience fluctuations in pressure and flow rate in each working cycle; this pulsation will be transmitted to the mechanical components of the compressor, causing the mechanical components to bear periodic mechanical stresses during each pulsation impact, resulting in periodic micro-vibrations. The repeated impact forces will accelerate the wear of the mechanical components in the compressor, further causing the mechanical components to generate irregular mechanical vibrations. In addition, the periodic load changes caused by the pulsation will also lead to relative movements and impacts of the compressor components. Especially when the pulsation frequency is close to the natural frequency of the compressor, more serious resonance phenomena may be triggered, resulting in increased vibrations. The long-term accumulation of gas pulsation and vibration will lead to a decrease in compressor efficiency, a shortening of the lifespan of mechanical components, an increase in maintenance frequency, and may even cause equipment failures. Especially for devices equipped with both a compressor and a high-speed rotating motor, such as in a dryer, the vibration of the compressor may be further transmitted through the support structure, superimposed with the vibrations generated by other high-speed rotating devices in the dryer, resulting in a significant increase in noise, having a negative impact on the user experience.
[0064] To solve this problem, the first aspect of this application provides a method for adaptively suppressing compressor vibration based on reinforcement learning, especially for compensating for the vibrations caused by the fluctuations in pressure and flow rate experienced by the compressed gas in each working cycle, that is, suppressing the vibrations caused by compressor pulsation.
[0065] The method for adaptively suppressing compressor vibration based on reinforcement learning provided by this application is implemented by a processing unit running a program. In some embodiments of this application, the processing unit is implemented by an embedded computing platform; in other embodiments of this application, the processing unit is implemented by a combination of an embedded computing platform and a GPU, or a combination of an embedded computing platform and an FPGA.
[0066] In some embodiments of the present application, the compressor vibration adaptive suppression method based on reinforcement learning includes the following steps as shown in Figure 1 the following:
[0067] Step S101: Obtain the vibration quantity quantization parameter.
[0068] In some embodiments of the present application, the vibration quantity quantization parameter includes: the rotational speed fluctuation amplitude of the compressor motor rotation period.
[0069] The rotational speed fluctuation amplitude of the compressor motor rotation period is calculated based on the rotational speed sequence of each compressor motor rotation period.
[0070] The rotational speed sequence is the rotational speed time sequence of each compressor motor rotation period. The compressor motor rotation period is the unit time for the compressor motor to rotate one week.
[0071] The rotation of the compressor motor for one week corresponds to the actions of the internal mechanical components of the compressor. For example, in a rotary compressor, the rotary compressor uses a rotating rotor to compress gas. The rotation of the compressor motor directly drives the rotor to rotate, thereby realizing the suction, compression, and discharge of gas; while in a scroll compressor, two scroll disks, one fixed and the other rotating above it, the gas is compressed between the scroll disks. The rotation of the compressor motor drives the movement of the rotating scroll disk, and the rotation of the compressor motor also corresponds to the movement of the scroll disk. Therefore, the change in the pulsating load caused by gas pulsation will also be periodically reflected in the rotational speed of the compressor motor. The rotational speed fluctuation amplitude of the compressor motor rotation period can be regarded as the vibration quantity quantization parameter.
[0072] Set the sampling frequency, which defines the number of rotational speed data collected per second, so as to obtain sufficient rotational speed data points within each compressor motor rotation period, and convert the time stamp and the corresponding rotational speed value into a time series format.
[0073] The rotational speed sequence of each compressor motor rotation period can be expressed as: .
[0074] Calculate the rotational speed fluctuation amplitude of the compressor motor rotation period. The rotational speed fluctuation amplitude is the difference between the maximum rotational speed and the minimum rotational speed in the rotational speed sequence corresponding to each compressor motor rotation period.
[0075] The rotational speed fluctuation amplitude of one compressor motor rotation period is denoted as:
[0076] ;
[0077] where ;
[0078] Record multiple vibration quantity quantization parameters for each rotation cycle of the compressor motor to form a time series of vibration quantity quantization parameters.
[0079] ;
[0080] where is the current moment, is the length of the historical window.
[0081] Step S102: Obtain vibration trend quantization parameters based on the vibration quantity quantization parameters.
[0082] In some embodiments of the present application, the vibration trend quantization parameters include: moving average.
[0083] The moving average is calculated based on the time series of vibration quantity quantization parameters.
[0084] Exemplarily, for the time series of vibration quantity quantization parameters ;
[0085] The moving average can be calculated by the following formula:
[0086] ;
[0087] The moving average reflects the average level of vibration caused by gas pulsation within the historical window and eliminates the interference of instantaneous fluctuations.
[0088] In some embodiments of the present application, the vibration trend quantization parameters further include: rate of change.
[0089] The rate of change is calculated based on the time series of vibration quantity quantization parameters.
[0090] Exemplarily, the rate of change can be calculated by the following formula:
[0091] ;
[0092] The rate of change represents the change of the vibration quantity quantization parameters within the historical window. A positive value indicates that the vibration is deteriorating, and a negative value indicates that the vibration is improving.
[0093] In some embodiments of the present application, the vibration trend quantization parameters further include: cumulative change amount.
[0094] The cumulative change amount is calculated based on the time series of vibration quantity quantization parameters.
[0095] Exemplarily, the cumulative change amount can be calculated by the following formula:
[0096] ;
[0097] The cumulative change amount is the total change amount of the vibration quantity quantization parameter within the historical window. If the cumulative change amount fluctuates up and down within a certain range and the positive and negative values change frequently, it indicates that the vibration quantity quantization parameter is continuously oscillating; while if it gradually approaches a specific value, it shows a state of tending to be stable.
[0098] In some embodiments of the present application, the vibration trend quantization parameter includes: standard deviation.
[0099] The standard deviation is calculated based on the vibration quantity quantization parameter time series and the moving average.
[0100] Exemplarily, the standard deviation can be calculated by the following formula:
[0101] ;
[0102] The standard deviation is used to measure the stability and volatility of the vibration quantity quantization parameter.
[0103] Step S103: Construct a reinforcement learning model based on the vibration quantity quantization parameter and the vibration trend quantization parameter.
[0104] In the present application, a reinforcement learning model is constructed based on the PPO algorithm.
[0105] Constructing a reinforcement learning model based on the PPO algorithm includes the following steps:
[0106] Step S1031: Construct the state space
[0107] Taking the vibration quantity quantization parameter, the moving average, the change rate, the cumulative change amount, the standard deviation, the average rotational speed of the compressor motor, and the load current as input features, construct the state space. That is, the state space includes: the vibration quantity quantization parameter, the moving average, the change rate, the cumulative change amount, the standard deviation, the average rotational speed of the compressor motor, and the load current;
[0108] The state space can be expressed as:
[0109] ;
[0110] Wherein, is the average rotational speed of the compressor motor, is the load current of the compressor motor.
[0111] Step S1032: Construct the action space
[0112] Define a continuous action space, and the action space includes two dimensions of the compensated torque current parameter and the compensated angle parameter.
[0113] The action space can be expressed as:
[0114] ;
[0115] Among them, is the compensation torque current parameter, defined within the range and can be obtained under experimental conditions; can be obtained under experimental conditions; is the compensation angle parameter, defined within the range i.e., varying between 0 and 360°.
[0116] Step S1033: Construct the reward function
[0117] ;
[0118] is the time step and the immediate reward obtained;
[0119] In the reward function, the rotational speed fluctuation amplitude is taken as part of the negative reward, and excessive compensation torque current and compensation angle also mean unstable system load, so they are also taken as part of the negative reward, thus encouraging stable and efficient operation. , and are weight coefficients used to balance the importance of various factors. The weight coefficients are set according to the requirements of actual applications to optimize the influence of different indicators on the reward.
[0120] Step S1034: Construct the policy network.
[0121] The policy network is used to output the action value of an action according to the current state. By continuously training and updating, the policy is optimized so that the intelligent agent can make better decisions in the environment.
[0122] In this application, by way of example, as Figure 2 shown, the policy network includes four fully connected layers and ReLU activation functions to capture the complex relationship between the output state and the output action. Among them, the input layer is used to receive state information, including the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change amount, the standard deviation, the average rotational speed of the compressor motor, and the load current. Each hidden layer has 64 neurons and uses the ReLU activation function; the output layer outputs the specific compensation torque current parameter , the compensation angle parameter , and the variance or standard deviation related to this action (variance parameterization, as shown by log_std in Figure 2 ). In the continuous action space, the action is sampled from a probability distribution, and the variance represents the degree of dispersion of this action. The design of variance parameterization allows the intelligent agent to not only select a definite action but also generate a random sample of an action according to the variance to increase exploration.
[0123] Step S1035: Construct a value network.
[0124] The value network is used to estimate the value of a certain state or state-action pair, thereby helping the agent make better decisions.
[0125] In this application, by way of example, as Figure 3 shown, the value network includes four fully connected layers and ReLU activation functions to output a corresponding value estimate. Among them, the input layer is used to receive state information, including vibration quantity quantization parameters, sliding mean, change rate, cumulative change amount, standard deviation, average rotational speed of the compressor motor, and load current. The number of neurons in each hidden layer is 64, and the ReLU activation function is used; the output layer outputs the value of the state, representing the expected return in the state.
[0126] Step S104: Train the reinforcement learning model.
[0127] Step S1041: Use the policy network of the current state to generate actions corresponding to the action space.
[0128] The steps of using the policy network of the current state to generate actions corresponding to the action space specifically include: keeping the compressor in the running state, collecting the state once every compressor motor rotation cycle (for example, 200 ms) , that is, the environmental state in which the agent is located at each time step; according to the current state output actions through the policy network , that is, the actions selected by the agent in each state. After executing the actions obtain the next state and the time step and the immediate reward , where the next state is associated with the actions taken , and the reward is the immediate feedback obtained by the agent at each time step. Store the current state , actions , immediate reward and the next state in the experience replay buffer, storing data of multiple trajectories . A trajectory is a sequence of states, actions, and rewards experienced by the agent when interacting with the environment. Trajectories are used to help the agent learn how to select optimal actions based on the current state to maximize the cumulative reward.
[0129] Continuously execute the above steps until a set number of trajectory data is obtained. Randomly select samples from the experience replay buffer.
[0130] Step S1042: Evaluate the value of each action using Generalized Advantage Estimation (GAE) , and construct a policy loss function based on the value. For randomly selected samples from the experience replay buffer, evaluate the value of each action of the randomly selected samples using Generalized Advantage Estimation (GAE) , and construct a policy loss function based on the value.
[0131] The steps of evaluating the value of each action using Generalized Advantage Estimation (GAE) specifically include:
[0132] Calculate the temporal difference error, which is used to measure the difference between the value of the current state and the value predicted by the immediate reward and the value of the next state.
[0133] The calculation formula of the temporal difference error can be expressed as:
[0134] ;
[0135] where is the time step obtained immediate reward; is the discount factor, between [0, 1], used to represent the importance of future rewards; is the value of the next state; is the value of the current state.
[0136] Among them, and are both obtained from the trained Critic network. Through training, the Critic network will learn how to output a value according to the input state.
[0137] The Critic network can be trained using algorithms disclosed in the prior art, and the weights of the Critic network are updated using the rewards obtained from the environment and the value of the next state. For example, it can be trained through temporal difference learning, and the Critic network will learn a suitable value function during the training process.
[0138] Based on the temporal difference error calculate the advantage function of Generalized Advantage Estimation.
[0139] The calculation formula of the advantage function of Generalized Advantage Estimation can be expressed as:
[0140] ;
[0141] The advantage function of Generalized Advantage Estimation measures the quality of the action in the state , and reflects the reward obtained after executing the action relative to the average reward in this state (i.e., the state value The gain of ( ); where, is the total length of the trajectory; is the accumulated number of steps; is the GAE parameter, which is used to control the smoothness of the temporal difference estimation. When , GAE degenerates into completely relying on future rewards, with high variance; when , GAE is a standard temporal difference method, completely relying on the current temporal difference error, with low variance, but may produce bias; by adjusting , a balance can be achieved between bias and variance, providing a more stable advantage estimation. Reasonable variance can avoid the oscillation in the learning process of the agent, prevent the agent from making extreme decisions in some steps, and also enable each sample to have a greater impact on the policy update, improve the utilization efficiency of samples, learn more information from the environment within a limited number of interactions, and reduce the demand for a large number of samples.
[0142] Step S1043: Construct the PPO loss function. PPO includes the overall loss function , and the overall loss function is weighted based on the policy loss function , the value loss function and the entropy regularization function .
[0143] Construct the policy loss function based on value. The policy loss function is expressed as:
[0144] ;
[0145] where, represents the expected value at time step , which is the average of samples in the entire trajectory; is the probability of the current policy selecting action in state ; is the probability of the old policy (the policy in the previous round of training) selecting action in the same state ; is the advantage function of GAE; is the clipping operation, which limits the ratio within the range of to avoid large updates causing drastic changes in the policy, is the clipping threshold, preferably set to 0.2.
[0146] The value loss function is expressed as:
[0147] ;
[0148] Among them, represents the expected value at time step , which is the average of samples in the entire trajectory; is the predicted value of the current value network at state , representing the agent's value estimation of this state. is the target value, calculated based on the discounted sum of future rewards starting from the current state. The loss function is expressed in the form of mean squared error, penalizing large prediction errors and helping the network converge faster. is output by the value network.
[0149] The value function is used to evaluate the value of each action, enabling the value network to more accurately reflect the true value of the state and helping the agent select better actions.
[0150] Define the entropy regularization function, which is defined as:
[0151] ;
[0152] represents the expected value at time step , which is the average of samples in the entire trajectory; is the probability of the current policy selecting action at state ; is the logarithm of the probability of the current policy selecting action at state . The entropy regularization function can calculate the entropy contribution of a specific action. The higher the entropy, the more diverse the actions selected in that state and the stronger the exploration ability of the agent. The entropy regularization function can increase network diversity and avoid premature convergence.
[0153] Overall loss function:
[0154] ;
[0155] Among them, and are weight coefficients, , .
[0156] Step S1044: Use the gradient descent algorithm (such as Adam) to minimize the overall loss , and update the parameters of the policy network and the value network.
[0157] Before obtaining training samples under environmental interaction, initialize the parameters of the policy network and the value network, usually in a random initialization manner, and further set hyperparameters, such as learning rate (e.g., "3e-4 / 1e-3"), discount factor (e.g., 0.99), weight coefficient, clipping threshold, batch size (e.g., 128), number of iterations, GAE parameters, etc.
[0158] Step S1045: After completing one policy update, save the parameters of the current policy as the old policy parameters to prepare for the next round of training. Repeatedly use the policy network of the current state to generate actions, evaluate the value of each action using generalized advantage estimation , and construct a policy loss function based on the value, use the gradient descent algorithm to minimize the overall loss, and update the parameters of the policy network and the value network until reaching a predetermined number of policy updates (e.g., 4 times) or performance criteria. After training is completed, use the trained policy to evaluate in the environment. Evaluate the performance of the agent by observing metrics such as average reward and success rate.
[0159] Step S105: Deploy the reinforcement learning model, and generate a compensation current signal according to the compensation torque current parameter and the compensation angle parameter output by the reinforcement learning model.
[0160] The compensation current signal is expressed as:
[0161] ;
[0162] is the time variable, is the angular frequency, , is the frequency, for example, the rotational frequency of the compressor motor rotor can be selected.
[0163] Step S106: The frequency converter receives the compensation current signal , and superimposes the compensation current signal with the compressor motor current control signal, and uses the superimposed current control signal to drive the compressor motor to operate. The superimposed current control signal can suppress periodic vibration.
[0164] Continuously collect data during the actual operation of the compressor, and regularly update the network parameters of the reinforcement learning model.
[0165] In some embodiments of the present application, it is preferably to increase the saturation limit so that the compensation torque current parameter does not exceed 20% of the maximum current of the compressor motor.
[0166] After testing, when using the compressor vibration adaptive suppression method based on reinforcement learning provided by this application, the vibration amplitude is stabilized within ±0.02 mm, and the parameter self-tuning time is shortened to within 30 seconds.
[0167] The compressor vibration adaptive suppression method based on reinforcement learning provided by this application indirectly quantifies vibration by using the rotational speed fluctuation amplitude of the compressor motor rotation period, without the need for sensor vibration detection; at the same time, it generates compensation torque current parameters and compensation angle parameters, covering a wider vibration spectrum; and realizes parameter self-tuning under dynamic working conditions through reinforcement learning, avoiding manual intervention.
[0168] In some embodiments of this application, the processing unit is deployed in a cloud server, and the cloud server is communicatively connected to a frequency converter to output a compensation current signal to the frequency converter to drive the compressor motor to operate.
[0169] In some embodiments of this application, the compressor vibration adaptive suppression method based on reinforcement learning can be applied to the compressor motor in a dryer. The operating conditions of a dryer are different from those of traditional refrigeration equipment such as refrigerators. There is also another motor in the dryer that drives the drying drum to rotate. The motor that drives the drying drum to rotate may operate in a high-speed mode, resulting in a problem of vibration superposition; on the other hand, the reinforcement learning model requires relatively high hardware resources, and an embedded system cannot provide them.
[0170] To address this problem, in some embodiments of this application, the compressor vibration adaptive suppression method based on reinforcement learning further includes the following steps:
[0171] Construct a PPO model with a smaller network structure as a backup model. Use fewer model layers, fewer neurons, and / or a simpler activation function. The other network architectures of the backup model are the same as those of the trained reinforcement learning model. Different from the reinforcement learning model, the state space of the backup model includes: vibration quantity quantization parameters, sliding mean, change rate, cumulative change amount, standard deviation, average motor speed, load current, and dryer operating mode. The dryer operating mode at least includes the mode in which the drying drum motor is rotating at high speed and the compressor motor is operating. Preferably, it also includes the mode in which the drying drum motor is rotating at medium speed and the compressor motor is operating, and the mode in which the drying drum motor is rotating at low speed and the compressor motor is operating. In different dryers, the defined names of the modes may be different, and will not be further defined here. Through state space expansion, the backup model can better capture the operating states when the two motors are running simultaneously.
[0172] Under the dryer operating mode, use the reinforcement learning model in step S105 for actual operation, collect data of state-action pairs, record the outputs of the reinforcement learning model under different operating modes, and form soft labels, which will provide valuable information for the backup model.
[0173] Construct a distillation loss function. The distillation loss function is used to compare the policy outputs of the reinforcement learning model and the backup model in the same state. The mean squared error can be selected as the distillation loss function.
[0174] Further combine the distillation loss function with the overall loss function in step S1043 to form a comprehensive loss function. The comprehensive loss function can combine the distillation loss function and the overall loss function in a weighted manner.
[0175] Obtain the training samples of the backup model under environment interaction: Use the backup model to generate actions, and interact with the environment to obtain rewards and the next state.
[0176] Calculate the comprehensive loss and update the parameters of the backup model.
[0177] Repeat the steps from obtaining the training samples of the backup model under environment interaction to updating the parameters of the backup model until convergence or reaching a predetermined number of training rounds to obtain a trained backup model.
[0178] Evaluate the performance of the backup model under different operating modes of the dryer and adjust the hyperparameters of the backup model to make the performance of the backup model meet the expectations to obtain a trained backup model.
[0179] Deploy the backup model on the embedded system of the dryer to ensure that it can run in actual work.
[0180] Through the backup model, on the one hand, the reinforcement learning model can be optimized and extended according to the operating characteristics of the dryer, increasing the support for different operating modes and performing better on the dryer; on the other hand, the reinforcement learning model can be compressed into a smaller alternative model without losing its functions, enabling it to be deployed on a local embedded system and run on a resource-constrained embedded system, improving the model's transferability and meeting the requirements of the embedded system for real-time performance and efficiency.
[0181] The second aspect of this application provides a compressor vibration adaptive suppression device based on reinforcement learning, as Figure 4 shown, including:
[0182] A first acquisition unit 101, which is used to acquire vibration quantity quantization parameters, and the vibration quantity quantization parameters include: the rotational fluctuation amplitude of the compressor motor rotation period , and the rotational fluctuation amplitude of the compressor motor rotation period is the difference between the maximum rotational speed and the minimum rotational speed in the rotational speed sequence corresponding to each compressor motor rotation period;
[0183] The second acquisition unit 102 is configured to obtain a vibration trend quantization parameter based on the vibration quantity quantization parameter, where the vibration trend quantization parameter includes one or more of a sliding mean, a change rate, an accumulated change amount, and a standard deviation;
[0184] The construction unit 103 is configured to construct a reinforcement learning model based on the vibration quantity quantization parameter and the vibration trend quantization parameter;
[0185] The training unit 104 is configured to train the reinforcement learning model;
[0186] The application unit 105 is configured to deploy the reinforcement learning model, and generate a compensation current signal according to the compensation torque current parameter and the compensation angle parameter output by the reinforcement learning model, where the compensation current signal is , where is a time variable, is an angular frequency, is the rotational frequency of the compressor motor rotor; , ;
[0187] The compensation unit 106 is configured to enable the frequency converter to receive the compensation current signal , superimpose the compensation current signal with the compressor motor current control signal, and drive the compressor motor to operate using the superimposed current control signal to suppress the vibration caused by the fluctuations in the pressure and flow experienced by the compressed gas in each working cycle.
[0188] Further, the construction unit is configured to construct a reinforcement learning model based on the PPO algorithm, and it includes:
[0189] A state space construction part, which is configured to construct a state space, where the state space includes: the vibration quantity quantization parameter, the sliding mean, the change rate, the accumulated change amount, the standard deviation, the average rotational speed of the compressor motor, and the load current;
[0190] An action space construction part, which is configured to construct an action space, where the action space includes the compensation torque current parameter and the compensation angle parameter ;
[0191] A reward function construction part, which is configured to construct a reward function, and the reward function is: , where is the immediate reward obtained at the time step , , and are weight coefficients;
[0192] The policy network construction department, which is used to construct the policy network;
[0193] The value network construction department, which is used to construct the value network.
[0194] Furthermore, the training unit includes:
[0195] The action generation department, which is used to generate actions corresponding to the action space using the policy network of the current state;
[0196] The evaluation department, which is used to evaluate the value of each action using generalized advantage estimation , and construct a policy loss function based on the value;
[0197] The policy loss function is:
[0198] ;
[0199] Where, is the expected value at time step ; is the probability of selecting action in state by the current policy; is the probability of selecting action in the same state by the old policy; is the clipping operation, which limits the ratio within the range of ;
[0200] The loss function construction department, which is used to construct the PPO loss function, and the PPO loss function includes the overall loss function , and the overall loss function is weighted and obtained based on the policy loss function , the value loss function and the entropy regularization function ; The overall loss function satisfies:
[0201] ;
[0202] and are weight coefficients;
[0203] Where, the value loss function is:
[0204] ;
[0205] Where, is the expected value at time step ; is the predicted value of the value network in state , which is the target value output by the value network;
[0206] The entropy regularization function is:
[0207] ;
[0208] where is the expected value at time step ; is the probability of selecting action in state by the current policy; is the logarithm of the probability of selecting action in state by the current policy;
[0209] An update unit, which is used to minimize the overall loss using the gradient descent algorithm and update the parameters of the policy network and the value network;
[0210] A storage unit, which is used to save the parameters of the current policy as the parameters of the old policy after one policy update; until the preset number of policy updates is reached.
[0211] Furthermore, the steps for the action generation unit to generate an action using the policy network of the current state include:
[0212] Keep the compressor in the running state, and collect the current state once per rotation cycle of each compressor motor ;
[0213] According to the current state output the action through the policy network;
[0214] Execute the action , and obtain the immediate reward at time step and the next state ;
[0215] Store the current state , the action , the immediate reward and the next state in the experience replay buffer;
[0216] Randomly select samples from the experience replay buffer;
[0217] The evaluation unit uses generalized advantage estimation to evaluate the value of each action of the randomly selected samples , and construct a policy loss function based on the value.
[0218] Further, the policy network includes four fully connected layers, and each layer includes 64 neurons;
[0219] The value network includes four fully connected layers, and each layer includes 64 neurons.
[0220] Through real-time learning and optimization, this application can accurately suppress the vibration of the compressor without sensors. This application uses the rotation fluctuation amplitude and vibration trend quantization parameters of the compressor motor to construct a reinforcement learning model and generate a compensation current signal, thereby effectively reducing the vibration caused by the pressure and flow fluctuations in the working cycle of the compressor, and improving the stability and service life of the equipment operation.
[0221] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, for those of ordinary skill in the art, it is still possible to modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions required to be protected by this application.
Claims
1. A method for adaptively suppressing compressor vibration based on reinforcement learning, characterized in that: The following steps are involved: Obtain vibration quantity quantification parameters, the vibration quantity quantification parameters include: the rotation fluctuation amplitude of the compressor motor rotation period , the rotation fluctuation amplitude of the compressor motor rotation period The difference between the maximum speed and the minimum speed in the speed sequence corresponding to each compressor motor rotation cycle; Acquire vibration trend quantification parameters based on the vibration quantity quantification parameters, wherein the vibration trend quantification parameters include sliding mean, change rate, cumulative change and standard deviation; Building a reinforcement learning model based on the vibration quantity quantified parameter and the vibration trend quantified parameter; training the reinforcement learning model; Deploy the reinforcement learning model, and calculate the compensation torque current parameter according to the output of the reinforcement learning model. and compensation angle parameters , generating a compensation current signal , the compensation current signal is , is the time variable, is the angular frequency, , is the rotation frequency of the compressor motor rotor; The frequency converter receives the compensation current signal , the compensation current signal Superimposed with the compressor motor current control signal, the compressor motor is driven to operate using the superimposed current control signal to suppress vibration caused by fluctuations in pressure and flow experienced by the compressed gas in each working cycle; The reinforcement learning model is constructed based on the PPO algorithm, and constructing the reinforcement learning model includes the following steps: Constructing a state space, the state space including: the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change, the standard deviation, the compressor motor average speed and the load current; Construct an action space, the action space including the compensation torque current parameter and the compensation angle parameter ; Construct a reward function, which is: ;in, is the time step Instant rewards received, , and is the weight coefficient; A strategy network is constructed, wherein the strategy network is used to output an action value of an action according to the current state; the strategy network includes four fully connected layers and a ReLU activation function, wherein the input layer is used to receive state information, including the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change, the standard deviation, the average speed of the compressor motor and the load current; the number of neurons in each hidden layer is 64, and the ReLU activation function is used; the output layer outputs the compensation torque current parameter , the compensation angle parameter , and the variance or standard deviation associated with that action; A value network is constructed, wherein the value network is used to estimate the value of a state or a state-action pair; the value network includes four fully connected layers and a ReLU activation function to output a corresponding value estimate; wherein the input layer is used to receive state information, including the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change, the standard deviation, the average speed of the compressor motor and the load current, the number of neurons in each hidden layer is 64, and a ReLU activation function is used; the output layer outputs the value of the state, which represents the expected return under the state.
2. The method for adaptively suppressing compressor vibration based on reinforcement learning according to claim 1, characterized in that: Training the reinforcement learning model includes the following steps: Using the policy network in the current state to generate an action corresponding to the action space; Assess the value of each action using generalized advantage estimation , and construct a strategy loss function based on the value; The strategy loss function is: ; in, is the time step Expected value; For the current strategy in state Next select action probability; Is the old policy in the same state Next select action probability; is the cropping operation, which changes the ratio Restricted to within the scope; Construct a PPO loss function, which includes an overall loss function , the overall loss function Based on the strategy loss function , value loss function and entropy regularization function Weighted to obtain; the overall loss function satisfy: ; and is the weight coefficient; Wherein, the value loss function is: ; in, is the time step Expected value; is the predicted value of the value network in state, is the target value output by the value network; The entropy regularization function is: ; in, is the time step Expected value; For the current strategy in state Next select action probability; For the current strategy in state Next select action The logarithm of the probability of Minimize the overall loss using a gradient descent algorithm and update the parameters of the policy network and the value network; After completing a strategy update, the parameters of the current strategy are saved as the parameters of the old strategy; until the preset number of strategy updates is reached.
3. The method for adaptively suppressing compressor vibration based on reinforcement learning according to claim 2, characterized in that: The step of using the policy network of the current state to generate an action corresponding to the action space comprises: Keep the compressor in operation and collect the current status once every compressor motor rotation cycle ; According to the current status Output actions through the policy network ; Execute an action , get the time step Instant Rewards and the next state ; The current state ,action , Instant Rewards and the next state Stored in the experience replay buffer; randomly selecting samples from the experience replay buffer; Use generalized advantage estimation to evaluate the value of each action in a randomly selected sample , and construct a policy loss function based on the value.
4. A compressor vibration adaptive suppression device based on reinforcement learning, characterized in that: include: The first acquisition unit is used to acquire the vibration quantity quantitative parameters, wherein the vibration quantity quantitative parameters include: the rotation fluctuation amplitude of the compressor motor rotation period , the rotation fluctuation amplitude of the compressor motor rotation period The difference between the maximum speed and the minimum speed in the speed sequence corresponding to each compressor motor rotation cycle; A second acquisition unit, which is used to acquire a vibration trend quantification parameter based on the vibration quantity quantification parameter, wherein the vibration trend quantification parameter includes a sliding mean, a change rate, a cumulative change, and a standard deviation; A construction unit, which is used to construct a reinforcement learning model based on the vibration amount quantification parameter and the vibration trend quantification parameter; A training unit, which is used to train the reinforcement learning model; An application unit is used to deploy the reinforcement learning model and calculate the compensation torque current parameter output by the reinforcement learning model. and compensation angle parameters , generating a compensation current signal , the compensation current signal is , is the time variable, is the angular frequency, , is the rotation frequency of the compressor motor rotor; A compensation unit, which is used to enable the frequency converter to receive the compensation current signal , the compensation current signal Superimposed with the compressor motor current control signal, the compressor motor is driven to operate using the superimposed current control signal to suppress vibration caused by fluctuations in pressure and flow experienced by the compressed gas in each working cycle; The construction unit is used to construct a reinforcement learning model based on the PPO algorithm, which includes: A state space construction unit, which is used to construct a state space, wherein the state space includes: the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change, the standard deviation, the average speed of the compressor motor and the load current; An action space construction unit, which is used to construct an action space, wherein the action space includes the compensation torque current parameter and the compensation angle parameter ; A reward function construction unit, which is used to construct a reward function, wherein the reward function is: ,in, is the time step Instant rewards received, , and is the weight coefficient; A strategy network construction unit is used to construct a strategy network; wherein the strategy network is used to output an action value of an action according to the current state; the strategy network includes four fully connected layers and a ReLU activation function, wherein the input layer is used to receive state information, including the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change, the standard deviation, the average speed of the compressor motor and the load current; the number of neurons in each hidden layer is 64, and the ReLU activation function is used; the output layer outputs the compensation torque current parameter , the compensation angle parameter , and the variance or standard deviation associated with the action; A value network construction unit is used to construct a value network; wherein the value network is used to estimate the value of a state or a state-action pair; the value network includes four fully connected layers and a ReLU activation function to output a corresponding value estimate; wherein the input layer is used to receive state information, including the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change, the standard deviation, the average speed of the compressor motor and the load current, the number of neurons in each hidden layer is 64, and the ReLU activation function is used; the output layer outputs the value of the state, which represents the expected return under the state.
5. The compressor vibration adaptive suppression device based on reinforcement learning according to claim 4, characterized in that: The training unit comprises: An action generation unit, configured to generate an action corresponding to the action space using the policy network in a current state; The evaluation part is used to evaluate the value of each action using generalized advantage estimation , and construct a strategy loss function based on the value; The strategy loss function is: ; in, is the time step Expected value; For the current strategy in state Next select action probability; Is the old policy in the same state Next select action probability; is the cropping operation, which changes the ratio Restricted to within the scope; A loss function construction unit, which is used to construct a PPO loss function, wherein the PPO loss function includes an overall loss function , the overall loss function Based on the strategy loss function , value loss function and entropy regularization function Weighted to obtain; the overall loss function satisfy: ; and is the weight coefficient; Wherein, the value loss function is: ; in, is the time step Expected value; For the value network in state The predicted value under is the target value output by the value network; The entropy regularization function is: ; in, is the time step Expected value; For the current strategy in state Next select action probability; For the current strategy in state Next select action The logarithm of the probability of An updating unit, which is used to minimize the overall loss using a gradient descent algorithm and update the parameters of the policy network and the value network; The storage unit is used to save the parameters of the current strategy as the parameters of the old strategy after completing a strategy update; until a preset number of strategy updates is reached.
6. The compressor vibration adaptive suppression device based on reinforcement learning according to claim 5, characterized in that: The step of the action generation unit using the current state of the policy network to generate an action corresponding to the action space includes: Keep the compressor in operation and collect the current status once every compressor motor rotation cycle ; According to the current status Output actions through the policy network ; Execute an action , get the time step Instant Rewards and the next state ; The current state ,action , Instant Rewards and the next state Stored in the experience replay buffer; randomly selecting samples from the experience replay buffer; The evaluation unit uses generalized advantage estimation to evaluate the value of each action of the randomly selected sample , and construct a policy loss function based on the value.
Citation Information
Patent Citations
Control method and system of compressor
CN103470483A
Motor drive device and compressor using the same
CN103684165A