Compressor vibration self-adaptive suppression method and device based on reinforcement learning
Through a method based on reinforcement learning, a reinforcement learning model is constructed using vibration quantization parameters to generate a compensation current signal, which solves the problem that traditional feedback control is difficult to suppress compressor vibration, and achieves a more efficient vibration suppression effect.
Patent Information
- Application Number
- CN202510479601.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
In traditional compressors, it is difficult to effectively suppress vibration caused by gas pulsation through feedback control, especially when load fluctuations are not allowed to cope with complex working conditions.
Using a method based on reinforcement learning, a reinforcement learning model is constructed by obtaining vibration quantization parameters and vibration trend quantization parameters, and a compensation current signal is generated to adjust the speed of the compressor motor to suppress vibration.
It realizes precisely suppressing compressor vibration without sensors, improving the stability and service life of equipment operation.
Smart Images

Figure CN119982484A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of machine learning, and in particular to a method for adaptively suppressing compressor vibration based on reinforcement learning, and a device for adaptively suppressing compressor vibration based on reinforcement learning. Background Art
[0002] The working process of the compressor is usually controlled by a frequency converter. The frequency converter controls the operation of the compressor by adjusting the speed of the motor. The compressor sucks in low-pressure gas and compresses the gas through the mechanical components inside the compressor. During the compression process, the volume of the gas decreases and the pressure increases. The frequency converter dynamically adjusts the speed of the motor according to demand to achieve efficient operation of the compressor under different loads. For example, when the load is light, the frequency converter will reduce the speed to reduce energy consumption; when the load increases, the frequency converter will increase the speed to keep the system pressure stable.
[0003] Compressors work by compressing gas. During the operation of the compressor, the compressed gas will experience fluctuations in pressure and flow in each working cycle. This pulsation will be transmitted to the mechanical parts of the compressor, causing the mechanical parts to be subjected to periodic mechanical stress in each pulsation impact. This repeated impact force will accelerate the wear of the components and further cause irregular mechanical vibrations in the mechanical parts. The periodic load changes caused by pulsation will also cause relative movement and impact of the compressor components, especially when the pulsation frequency is close to the natural frequency of the compressor, which may cause more serious resonance and lead to intensified vibration. Long-term accumulation of gas pulsation and vibration will lead to reduced compressor efficiency, shortened component life, increased maintenance frequency, and may even cause equipment failure.
[0004] In traditional compressors, pulsation is compensated by manually adjusting the frequency converter or PID control, which is also called feedback control. By monitoring the vibration or other related variables of the compressor, the deviation is reduced by adjusting control variables such as the speed, thereby reducing pulsation and vibration to a certain extent. However, the effect of feedback control is very limited, especially when facing load pulsation fluctuations, and it cannot cope with complex working conditions in actual applications. Summary of the invention
[0005] In view of the above problems, the present application designs and provides a compressor vibration adaptive suppression method based on reinforcement learning.
[0006] A method for adaptively suppressing compressor vibration based on reinforcement learning, comprising the following steps: obtaining vibration quantity quantification parameters, wherein the vibration quantity quantification parameters include: the rotation fluctuation amplitude of the compressor motor rotation period; , the rotation fluctuation amplitude of the compressor motor rotation period The invention relates to a method for generating a vibration trend quantification parameter based on the vibration quantity quantification parameter, wherein the vibration trend quantification parameter includes one or more of a sliding mean, a rate of change, a cumulative change, and a standard deviation; a reinforcement learning model is constructed based on the vibration quantity quantification parameter and the vibration trend quantification parameter; the reinforcement learning model is trained; the reinforcement learning model is deployed, and the compensation torque current parameter output by the reinforcement learning model is calculated. and compensation angle parameters , generating a compensation current signal , the compensation current signal is , is the time variable, is the angular frequency, , is the rotation frequency of the compressor motor rotor; the frequency converter receives the compensation current signal , the compensation current signal The superimposed current control signal is used to drive the compressor motor to operate, thereby suppressing the vibration caused by the fluctuation of pressure and flow rate experienced by the compressed gas in each working cycle.
[0007] Furthermore, the reinforcement learning model is constructed based on the PPO algorithm, and the construction of the reinforcement learning model includes the following steps: constructing a state space, the state space including: the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change, the standard deviation, the compressor motor average speed and load current; constructing an action space, the action space including the compensation torque current parameter and compensation angle parameters ; Construct a reward function, the reward function is: ;in, is the time step Instant rewards, , and is the weight coefficient; construct a strategy network; construct a value network.
[0008] Furthermore, training the reinforcement learning model includes the following steps: using the current state of the policy network to generate actions corresponding to the action space; using generalized advantage estimation to evaluate the value of each action , and construct a strategy loss function based on the value; the strategy loss function is: ; in, is the time step Expected value; For the current strategy in state Next select action probability; Is the old policy in the same state Next select action probability; is the cropping operation, which changes the ratio Restricted to within the scope; Construct a PPO loss function, which includes an overall loss function , the overall loss function Based on the strategy loss function , value loss function and entropy regularization function Weighted to obtain; the overall loss function satisfy: ; and is the weight coefficient; Wherein, the value loss function is: ; in, is the time step Expected value; For the value network in state The predicted value under is the target value output by the value network; The entropy regularization function is: ; in, is the time step Expected value; For the current strategy in state Next select action probability; For the current strategy in state Next select action The logarithm of the probability of Minimize the overall loss using a gradient descent algorithm and update the parameters of the policy network and the value network; After completing a strategy update, the parameters of the current strategy are saved as the parameters of the old strategy; until the preset number of strategy updates is reached.
[0009] Furthermore, the step of using the strategy network of the current state to generate an action corresponding to the action space includes: keeping the compressor in a running state, collecting the current state once in each compressor motor rotation cycle ; According to the current state Output actions through the policy network ; Execute action , get the time step Instant Rewards and the next state ; The current state ,action , Instant Rewards and the next state Store in an experience replay buffer; randomly select samples from the experience replay buffer; evaluate the value of each action of the randomly selected samples using generalized advantage estimation , and construct a policy loss function based on the value.
[0010] Furthermore, the strategy network includes four fully connected layers, each layer includes 64 neurons; the value network includes four fully connected layers, each layer includes 64 neurons.
[0011] The second aspect of the present application provides a compressor vibration adaptive suppression device based on reinforcement learning, comprising: a first acquisition unit, which is used to acquire a vibration quantity quantification parameter, wherein the vibration quantity quantification parameter includes: a rotation fluctuation amplitude of a compressor motor rotation period; , the rotation fluctuation amplitude of the compressor motor rotation period The difference between the maximum speed and the minimum speed in the speed sequence corresponding to each compressor motor rotation cycle; a second acquisition unit, configured to acquire a vibration trend quantification parameter based on the vibration quantity quantification parameter, wherein the vibration trend quantification parameter includes one or more of a sliding mean, a change rate, a cumulative change, and a standard deviation; A construction unit, which is used to construct a reinforcement learning model based on the vibration amount quantification parameter and the vibration trend quantification parameter; A training unit, which is used to train the reinforcement learning model; An application unit is used to deploy the reinforcement learning model and calculate the compensation torque current parameter output by the reinforcement learning model. and compensation angle parameters , generating a compensation current signal , the compensation current signal is , is the time variable, is the angular frequency, , is the compressor motor rotor rotation frequency; A compensation unit, which is used to enable the frequency converter to receive the compensation current signal , the compensation current signal The superimposed current control signal is used to drive the compressor motor to operate, thereby suppressing the vibration caused by the fluctuation of pressure and flow rate experienced by the compressed gas in each working cycle.
[0012] Furthermore, the construction unit is used to construct a reinforcement learning model based on the PPO algorithm, which includes: A state space construction unit is used to construct a state space, wherein the state space includes: the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change, the standard deviation, the compressor motor average speed and the load current; an action space construction unit is used to construct an action space, wherein the action space includes the compensation torque current parameter and compensation angle parameters ; A reward function construction unit, which is used to construct a reward function, wherein the reward function is: ,in, is the time step Instant rewards, ,and is a weight coefficient; a strategy network construction unit, which is used to construct a strategy network; The value network construction department is used to construct the value network.
[0013] Furthermore, the training unit includes: an action generation unit, which is used to generate an action corresponding to the action space using the current state of the policy network; an evaluation unit, which is used to evaluate the value of each action using generalized advantage estimation , and construct a strategy loss function based on the value; the strategy loss function is: ; in, is the time step Expected value; For the current strategy in state Next select action The probability of Is the old policy in the same state Next select action The probability of is the cropping operation, which changes the ratio Restricted to within the scope; A loss function construction unit, which is used to construct a PPO loss function, wherein the PPO loss function includes an overall loss function , the overall loss function Based on the strategy loss function , value loss function and entropy regularization function Weighted to obtain; the overall loss function satisfy: ; and is the weight coefficient; Wherein, the value loss function is: ; in, is the time step Expected value; For the value network in state The predicted value under is the target value output by the value network; The entropy regularization function is: ; in, is the time step Expected value; For the current strategy in state Next select action probability; For the current strategy in state Next select action The logarithm of the probability of An updating unit, which is used to minimize the overall loss using a gradient descent algorithm and update the parameters of the policy network and the value network; The storage unit is used to save the parameters of the current strategy as the parameters of the old strategy after completing a strategy update; until a preset number of strategy updates is reached.
[0014] Further, the action generation unit generates an action corresponding to the action space using the strategy network of the current state, including: keeping the compressor in a running state, collecting the current state once in each compressor motor rotation cycle ; According to the current state Output actions through the policy network ; Execute action , get the time step Instant Rewards and the next state ; The current state ,action , Instant Rewards and the next state Store in the experience replay buffer; randomly select samples from the experience replay buffer; the evaluation unit uses generalized advantage estimation to evaluate the value of each action of the randomly selected sample , and construct a policy loss function based on the value.
[0015] Furthermore, the strategy network includes four fully connected layers, each layer includes 64 neurons; the value network includes four fully connected layers, each layer includes 64 neurons.
[0016] This application can accurately suppress compressor vibration without sensors through real-time learning and optimization. This application uses the rotation fluctuation amplitude and vibration trend quantification parameters of the compressor motor to build a reinforcement learning model and generate a compensation current signal, thereby effectively reducing the vibration caused by pressure and flow fluctuations in the compressor during the working cycle and improving the stability and service life of the equipment.
[0017] After reading the specific implementation manner of the present application in conjunction with the accompanying drawings, other features and advantages of the present application will become more clear. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 A flowchart of a method for adaptively suppressing compressor vibration based on reinforcement learning provided in this application; Figure 2 A network architecture diagram for a specific example of a policy network; Figure 3 A network architecture diagram for a specific example of a value network; Figure 4 This is a schematic block diagram of the structure of the compressor vibration adaptive suppression method based on reinforcement learning provided in this application. DETAILED DESCRIPTION
[0020] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0021] In the description of the present application, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore, should not be understood as a limitation on the present application.
[0022] In the description of the present application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "electrically connected" should be understood in a broad sense. For example, it can be a fixed electrical connection, a detachable electrical connection, or an integral electrical connection. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances. In the description of the implementation method, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0023] The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.
[0024] In the description of the present application, unless otherwise specified, “plurality” means two or more.
[0025] Compressors are widely used in many industries, mainly for gas compression, and play a vital role in refrigeration systems, achieving cooling (heating) effects by compressing gas. In the field of household appliances, refrigeration systems with compressors as the core are not only used in traditional air conditioners or refrigerators (including freezers), but also in new products such as dryers and heat pump systems. For example, in dryers, compressors are used to circulate refrigerants, thereby achieving heat recovery and temperature control during the drying process, and providing the temperature and humidity required for drying.
[0026] In a compressor, the motor is the core component that drives the compressor. The motor provides power through rotation, driving the mechanical part of the compressor (such as the rotor, etc.) to compress the gas. The speed of the compressor motor is adjusted by the inverter (also known as the compressor drive circuit). The inverter can adjust the motor speed in real time according to the load demand, so that the motor can operate at the best efficiency.
[0027] During the operation of the compressor, the compressed gas will experience fluctuations in pressure and flow in each working cycle; this pulsation will be transmitted to the mechanical parts of the compressor, causing the mechanical parts to be subjected to periodic mechanical stress in each pulsation impact, resulting in periodic micro-vibrations. Repeated impact forces will accelerate the wear of mechanical parts in the compressor, further causing irregular mechanical vibrations in the mechanical parts. In addition, the periodic load changes caused by pulsation will also cause relative movement and impact of compressor parts, especially when the pulsation frequency is close to the natural frequency of the compressor, which may cause more serious resonance phenomena and lead to increased vibration. Long-term accumulation of gas pulsation and vibration will lead to reduced compressor efficiency, shortened mechanical parts life, increased maintenance frequency, and may even cause equipment failure. Especially for devices equipped with both compressors and high-speed rotating motors, such as dryers, the vibration of the compressor may be further transmitted through the supporting structure, superimposed with the vibrations generated by other high-speed rotating devices of the dryer, resulting in a significant increase in noise, which has a negative impact on the user experience.
[0028] To solve this problem, the first aspect of the present application provides a method for adaptively suppressing compressor vibration based on reinforcement learning, which is particularly used to compensate for vibrations caused by fluctuations in pressure and flow experienced by the compressed gas in each working cycle, that is, to suppress vibrations caused by compressor pulsation.
[0029] The compressor vibration adaptive suppression method based on reinforcement learning provided in the present application is implemented by a processing unit running a program. In some embodiments of the present application, the processing unit is implemented by an embedded computing platform; in other embodiments of the present application, the processing unit is implemented by a combination of an embedded computing platform and a GPU, or a combination of an embedded computing platform and an FPGA.
[0030] In some embodiments of the present application, the compressor vibration adaptive suppression method based on reinforcement learning includes the following steps: Figure 1 The following steps are shown: Step S101: Obtain vibration quantity quantization parameter.
[0031] In some embodiments of the present application, the vibration quantity quantification parameter includes: the speed fluctuation amplitude of the compressor motor rotation period.
[0032] The speed fluctuation amplitude of the compressor motor rotation cycle is obtained based on the speed sequence calculation of each compressor motor rotation cycle.
[0033] The speed sequence is a speed time sequence of each compressor motor rotation cycle. The compressor motor rotation cycle is the unit time for the compressor motor to rotate one circle.
[0034] One revolution of the compressor motor corresponds to the movement of the mechanical parts inside the compressor. For example, in a rotary compressor, the rotary compressor uses a rotating rotor to compress the gas, and the rotation of the compressor motor directly drives the rotor to rotate, thereby achieving the suction, compression and discharge of the gas; while in a scroll compressor, there are two scrolls, one fixed and the other rotating above it, and the gas is compressed between the scrolls. The rotation of the compressor motor drives the movement of the rotating scroll, and the rotation of the compressor motor corresponds to the movement of the scroll. Therefore, the pulsating load changes caused by gas pulsation will also be periodically reflected in the compressor motor speed, and the speed fluctuation amplitude of the compressor motor rotation cycle can be regarded as a vibration quantity quantification parameter.
[0035] Set the sampling frequency and define the number of speed data collected per second so that enough speed data points can be obtained in each compressor motor rotation cycle, and convert the timestamp and the corresponding speed value into a time series format.
[0036] The speed sequence for each compressor motor rotation cycle can be expressed as: .
[0037] The speed fluctuation amplitude of the compressor motor rotation cycle is calculated, and the speed fluctuation amplitude is the difference between the maximum speed and the minimum speed in the speed sequence corresponding to each compressor motor rotation cycle.
[0038] The speed fluctuation amplitude of a compressor motor rotation cycle is recorded as: ;
[0039] in, ;
[0040] A plurality of vibration quantity quantification parameters are recorded in each compressor motor rotation cycle to form a vibration quantity quantification parameter time series.
[0041] ;
[0042] in For the current moment, is the history window length.
[0043] Step S102: Obtaining a vibration tendency quantification parameter based on the vibration amount quantification parameter.
[0044] In some implementations of the present application, the vibration trend quantification parameter includes: a sliding mean.
[0045] The sliding mean is calculated based on the time series of the vibration quantity quantification parameters.
[0046] For example, for the vibration quantity quantization parameter time series ;
[0047] The sliding mean can be calculated as follows: ;
[0048] The sliding mean reflects the average level of vibration caused by gas pulsation within the historical window, eliminating the interference of instantaneous fluctuations.
[0049] In some implementations of the present application, the vibration trend quantification parameter further includes: a rate of change.
[0050] The change rate is calculated based on the time series of the vibration quantity quantified parameter.
[0051] Exemplarily, the rate of change can be calculated by the following formula: ;
[0052] The rate of change is the change in the vibration quantity parameter within the historical window. A positive value means that the vibration is getting worse, and a negative value means that the vibration is improving.
[0053] In some implementations of the present application, the vibration trend quantification parameter further includes: cumulative change.
[0054] The cumulative change is calculated based on the vibration quantity quantification parameter time series.
[0055] Exemplarily, the cumulative change can be calculated by the following formula: ;
[0056] The cumulative change is the total change of the vibration quantity parameter in the historical window. If the cumulative change fluctuates within a certain range and changes positively and negatively frequently, it means that the vibration quantity parameter continues to oscillate; if it gradually approaches a certain value, it shows a stable state.
[0057] In some embodiments of the present application, the vibration trend quantification parameter includes: standard deviation.
[0058] The standard deviation is obtained based on the vibration quantity quantification parameter time series and the sliding mean calculation.
[0059] For example, the standard deviation can be calculated by the following formula: ;
[0060] The standard deviation is used to measure the stability and volatility of the vibration quantity quantitative parameters.
[0061] Step S103: constructing a reinforcement learning model based on the vibration quantity quantification parameter and the vibration trend quantification parameter.
[0062] In this application, a reinforcement learning model is constructed based on the PPO algorithm.
[0063] Building a reinforcement learning model based on the PPO algorithm includes the following steps: Step S1031: Constructing state space The vibration quantity quantification parameter, sliding mean, change rate, cumulative change, standard deviation, average speed of the compressor motor and load current are used as input features to construct the state space. That is, the state space includes: vibration quantity quantification parameter, sliding mean, change rate, cumulative change, standard deviation, average speed of the compressor motor and load current; The state space can be expressed as: ;
[0064] in, is the average speed of the compressor motor, is the compressor motor load current.
[0065] Step S1032: Constructing an action space A continuous action space is defined, and the action space includes two dimensions: compensation torque current parameter and compensation angle parameter.
[0066] The action space can be expressed as: ;
[0067] in, The torque compensation current parameter is defined in the range Within, can be obtained under experimental conditions; To compensate the angle parameter, define the range within, that is, varies between 0 and 360°.
[0068] Step S1033: Construct reward function ;
[0069] is the time step Instant rewards received; In the reward function, the speed fluctuation amplitude is taken as part of the negative reward, while excessive compensation torque current and compensation angle also mean that the system load is unstable, so they are also taken as part of the negative reward, thereby encouraging stable and efficient operation. , and It is a weight coefficient used to balance the importance of various factors. The weight coefficient is set according to the needs of the actual application to optimize the impact of different indicators on rewards.
[0070] Step S1034: construct a policy network.
[0071] The policy network is used to output the action value of an action based on the current state. Through continuous training and updating, the strategy is optimized so that the agent can make better decisions in the environment.
[0072] In this application, exemplary examples include Figure 2 As shown in the figure, the policy network includes four fully connected layers and ReLU activation function to capture the complex relationship between output state and output action. The input layer is used to receive state information, including vibration quantity quantization parameters, sliding mean, change rate, cumulative change, standard deviation, compressor motor average speed and load current. The number of neurons in each hidden layer is 64, and the ReLU activation function is used; the output layer outputs specific compensation torque current parameters , compensation angle parameters , and the variance or standard deviation associated with that action (variance parameterized as Figure 2 In the continuous action space, actions are sampled from probability distributions, and the variance represents the dispersion of the action. The design of variance parameterization allows the agent to not only select a certain action, but also generate a random sample of actions based on the variance to increase exploration.
[0073] Step S1035: construct a value network.
[0074] The value network is used to estimate the value of a state or state-action pair, thereby helping the agent make better decisions.
[0075] In this application, for example, Figure 3 As shown in the figure, the value network includes four fully connected layers and a ReLU activation function to output a corresponding value estimate. The input layer is used to receive state information, including vibration quantity quantization parameters, sliding mean, change rate, cumulative change, standard deviation, compressor motor average speed and load current. The number of neurons in each hidden layer is 64, and the ReLU activation function is used; the output layer outputs the value of the state, which represents the expected return under the state.
[0076] Step S104: training a reinforcement learning model.
[0077] Step S1041: Use the current state of the policy network to generate actions corresponding to the action space.
[0078] The steps of using the current state policy network to generate actions corresponding to the action space specifically include: keeping the compressor in operation, collecting the state once every compressor motor rotation cycle (e.g. 200ms) , that is, the state of the environment in which the agent is located at each time step; according to the current state Output actions through the policy network , which is the action selected by the agent in each state. Execute action Then get the next state and time step Instant Rewards , where the next state is related to the action taken Related, Reward is the instant feedback obtained by the agent at each time step. ,action , Instant Rewards and the next state Stored in the experience replay buffer, storing data for multiple trajectories A trajectory is a sequence of states, actions, and rewards that an agent experiences when interacting with the environment. Trajectories are used to help the agent learn how to choose the best action based on the current state to maximize the cumulative reward.
[0079] Continue to perform the above steps until the set amount of trajectory data is obtained. Randomly select samples from the experience replay buffer.
[0080] Step S1042: Evaluate the value of each action using generalized advantage estimation , and construct a policy loss function based on the value. For randomly selected samples from the experience replay buffer, the value of each action of the randomly selected sample is evaluated using the generalized advantage estimate , and construct a policy loss function based on the value.
[0081] Evaluate the value of each action using Generalized Advantage Estimation (GAE) The steps specifically include: Calculate the temporal difference error, which measures the difference between the value of the current state and the value predicted by the immediate reward and the value of the next state.
[0082] The calculation formula of time difference error can be expressed as: ;
[0083] Where, is the time step Instant rewards received; is a discount factor between [0,1], used to indicate the importance of future rewards; is the value of the next state; is the value of the current state.
[0084] in, and All of them are obtained by training the Critic network. Through training, the Critic network will learn how to output a value based on the input state.
[0085] The critic network can be trained using algorithms disclosed in the prior art, using rewards obtained from the environment and the value of the next state to update the weights of the critic network. For example, training can be performed through time difference learning. The critic network will learn the appropriate value function by itself during the training process.
[0086] Based on time difference error Compute the advantage function for the generalized advantage estimate.
[0087] The calculation formula of the advantage function of generalized advantage estimation can be expressed as: ;
[0088] The advantage function of generalized advantage estimation Measuring actions in state The quality of the state value reflects the reward obtained after executing the action relative to the average reward in the state (that is, the state value ) of the gain; where is the total length of the trajectory; is the number of accumulated steps; is a GAE parameter used to control the smoothness of the temporal difference estimate. When , GAE degenerates into being completely dependent on future rewards and has a higher variance; when When , GAE is a standard time difference method, which is completely dependent on the current time difference error and has a lower variance, but may produce deviations; by adjusting , can strike a balance between bias and variance, providing a more stable advantage estimate. Reasonable variance can avoid oscillations in the agent's learning process and prevent the agent from making extreme decisions in certain steps. It can also make each sample have a greater impact in the strategy update, improve the efficiency of sample utilization, learn more information from the environment within a limited number of interactions, and reduce the need for a large number of samples.
[0089] Step S1043: Construct the PPO loss function, PPO includes the overall loss function , the overall loss function is based on the policy loss function , value loss function and entropy regularization function Weighted.
[0090] The strategy loss function is constructed based on the value. The strategy loss function is expressed as: ;
[0091] in, Represents the time step The expected value of is the average of the samples in the entire trajectory; For the current strategy in state Next select action The probability of The old strategy (the strategy in the previous round of training) is in the same state Next select action The probability of is the advantage function of GAE; is the cropping operation, which changes the ratio Restricted to range to avoid drastic changes in strategy caused by excessive updates. is the clipping threshold, preferably set to 0.2.
[0092] The value loss function is expressed as: ;
[0093] in, Represents the time step The expected value of is the average of the samples in the entire trajectory; is the current value network in state The predicted value under , represents the agent’s estimate of the value of this state. is the target value, calculated based on the discounted sum of future rewards from the current state. The loss function is expressed as a squared error, penalizing large prediction errors and helping the network converge faster. Output from the value network.
[0094] The value function is used to evaluate the value of each action, so that the value network can more accurately reflect the true value of the state and help the agent choose better actions.
[0095] Define the entropy regularization function. The entropy regularization function is defined as: ;
[0096] Represents the time step The expected value of is the average of the samples in the entire trajectory; Is the current strategy in state Next select action The probability of For the current strategy in state Next select action The entropy regularization function can calculate the entropy contribution of a specific action. The higher the entropy, the more diverse the actions selected in this state, and the stronger the exploratory nature of the agent. The entropy regularization function can increase network diversity and avoid premature convergence.
[0097] Overall loss function: ;
[0098] in, and is the weight coefficient, , .
[0099] Step S1044: Use a gradient descent algorithm (such as Adam) to minimize the overall loss , update the parameters of the policy network and the value network.
[0100] Before obtaining training samples under environmental interaction, initialize the parameters of the policy network and the value network, usually by random initialization, and further set hyperparameters such as learning rate (e.g. 3e-4 / 1e-3”), discount factor (e.g. 0.99), weight coefficient, clipping threshold, batch size (e.g. 128), number of cycles, GAE parameters, etc.
[0101] Step S1045: After completing a strategy update, save the parameters of the current strategy as the old strategy parameters to prepare for the next round of training. Repeatedly use the current state of the strategy network to generate actions and use generalized advantage estimation to evaluate the value of each action. , and construct a policy loss function based on the value, use the gradient descent algorithm to minimize the overall loss, and update the parameters of the policy network and the value network until a predetermined number of policy updates (e.g., 4 times) or performance criteria are reached. After the training is completed, the trained policy is used for evaluation in the environment. The performance of the agent is evaluated by observing indicators such as average reward and success rate.
[0102] Step S105: deploying a reinforcement learning model, and generating a compensation current signal according to the compensation torque current parameter and the compensation angle parameter output by the reinforcement learning model.
[0103] The compensation current signal is expressed as: ;
[0104] is the time variable, is the angular frequency, , is the frequency, for example the compressor motor rotor rotation frequency can be selected.
[0105] Step S106: The inverter receives the compensation current signal , the compensation current signal The superimposed current control signal is used to drive the compressor motor to operate. The superimposed current control signal can suppress periodic vibration.
[0106] Continuously collect data during the actual operation of the compressor and regularly update the network parameters of the reinforcement learning model.
[0107] In some embodiments of the present application, it is preferred to increase the saturation limit so that the compensation torque current parameter does not exceed 20% of the maximum current of the compressor motor.
[0108] After testing, it was found that by using the compressor vibration adaptive suppression method based on reinforcement learning provided in this application, the vibration amplitude was stabilized within ±0.02mm, and the parameter self-tuning time was shortened to within 30 seconds.
[0109] The reinforcement learning-based adaptive vibration suppression method for compressors provided in this application utilizes the speed fluctuation amplitude of the compressor motor rotation cycle to indirectly quantify the vibration without the need for sensor vibration detection; it simultaneously generates compensation torque current parameters and compensation angle parameters to cover a wider vibration spectrum; and it achieves parameter self-tuning under dynamic conditions through reinforcement learning to avoid manual intervention.
[0110] In some embodiments of the present application, the processing unit is deployed in a cloud server, and the cloud server is communicatively connected to the inverter to output the compensation current signal to the inverter to drive the compressor motor to operate.
[0111] In some embodiments of the present application, the adaptive vibration suppression method for compressors based on reinforcement learning can be applied to compressor motors in clothes dryers. The operating conditions of clothes dryers are different from those of traditional refrigeration equipment such as refrigerators. There is another motor in the clothes dryer that drives the dryer drum to rotate. The motor that drives the dryer drum to rotate may work in high-speed mode, which may cause vibration superposition problems. On the other hand, reinforcement machine learning models require higher hardware resources, which embedded systems cannot provide.
[0112] To address this problem, in some embodiments of the present application, the compressor vibration adaptive suppression method based on reinforcement learning further includes the following steps: Construct a PPO model with a smaller network structure as a backup model, using fewer model layers, fewer neurons and / or simpler activation functions, and the other network architectures of the backup model are the same as the trained reinforcement learning model. Unlike the reinforcement learning model, the state space of the backup model includes: vibration quantity quantization parameters, sliding mean, rate of change, cumulative change, standard deviation, average motor speed, load current and dryer operation mode. The dryer operation mode at least includes a mode in which the dryer motor is in high-speed rotation and the compressor motor is running, and preferably also includes a mode in which the dryer motor is in medium-speed rotation and the compressor motor is running, and a mode in which the dryer motor is in low-speed rotation and the compressor motor is running. In different dryers, the definition names of the modes may be different and are not further defined here. By expanding the state space, the backup model can better capture the operating status of the two motors when they are running at the same time.
[0113] In the dryer operation mode, the reinforcement learning model in step S105 is used for actual operation, the data of state-action pairs are collected, and the output of the reinforcement learning model in different operation modes is recorded to form soft labels, which will provide valuable information for the backup model.
[0114] Construct a distillation loss function. The distillation loss function is used to compare the policy outputs of the reinforcement learning model and the backup model under the same state. The distillation loss function can be selected as the mean square error.
[0115] The distillation loss function is further combined with the overall loss function in step S1043 to form a comprehensive loss function. The comprehensive loss function can combine the distillation loss function and the overall loss function in a weighted manner.
[0116] Get training samples of the backup model under environment interaction: Use the backup model to generate actions and interact with the environment to obtain rewards and next states.
[0117] Calculate the comprehensive loss and update the parameters of the backup model.
[0118] The steps of obtaining training samples of the backup model under the environment interaction and updating the parameters of the backup model are repeated until convergence or a predetermined number of training rounds is reached to obtain a trained backup model.
[0119] The performance of the backup model is evaluated under different operating modes of the dryer, and the hyperparameters of the backup model are adjusted so that the performance of the backup model reaches expectations, thereby obtaining a trained backup model.
[0120] The backup model was deployed on the dryer’s embedded system to ensure that it could run in real-world conditions.
[0121] Through the backup model, on the one hand, the reinforcement learning model can be optimized and expanded according to the operating characteristics of the dryer, and support for different operating modes can be increased to perform better on the dryer; on the other hand, the reinforcement learning model can be compressed into a smaller backup model without losing functionality, so that it can be deployed in a local embedded system and run on a resource-constrained embedded system, improving the portability of the model and meeting the real-time and efficiency requirements of the embedded system.
[0122] The second aspect of the present application provides a compressor vibration adaptive suppression device based on reinforcement learning, such as Figure 4 As shown, including: The first acquisition unit 101 is used to acquire vibration quantity quantification parameters, wherein the vibration quantity quantification parameters include: the rotation fluctuation amplitude of the compressor motor rotation period; , the rotation fluctuation amplitude of the compressor motor rotation period The difference between the maximum speed and the minimum speed in the speed sequence corresponding to each compressor motor rotation cycle; A second acquisition unit 102, which is used to acquire a vibration trend quantification parameter based on the vibration quantity quantification parameter, wherein the vibration trend quantification parameter includes one or more of a sliding mean, a change rate, a cumulative change, and a standard deviation; A construction unit 103, which is used to construct a reinforcement learning model based on the vibration quantity quantization parameter and the vibration trend quantization parameter; A training unit 104, which is used to train the reinforcement learning model; The application unit 105 is used to deploy the reinforcement learning model and calculate the compensation torque current parameter according to the output of the reinforcement learning model. and compensation angle parameters , generating a compensation current signal , the compensation current signal is , is the time variable, is the angular frequency, , is the rotation frequency of the compressor motor rotor; The compensation unit 106 is used to enable the inverter to receive the compensation current signal , the compensation current signal The superimposed current control signal is used to drive the compressor motor to operate, thereby suppressing the vibration caused by the fluctuation of pressure and flow rate experienced by the compressed gas in each working cycle.
[0123] Furthermore, the construction unit is used to construct a reinforcement learning model based on the PPO algorithm, which includes: A state space construction unit, which is used to construct a state space, wherein the state space includes: the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change, the standard deviation, the average speed of the compressor motor and the load current; An action space construction unit, which is used to construct an action space, wherein the action space includes the compensation torque current parameter and compensation angle parameters ; A reward function construction unit, which is used to construct a reward function, wherein the reward function is: ,in, is the time step Instant rewards, , and is the weight coefficient; A strategy network construction unit, which is used to construct a strategy network; The value network construction department is used to construct the value network.
[0124] Furthermore, the training unit includes: An action generation unit, which is used to generate an action corresponding to the action space using the policy network in the current state; The evaluation part is used to evaluate the value of each action using generalized advantage estimation , and construct a strategy loss function based on the value; The strategy loss function is: ;
[0125] in, is the time step Expected value; For the current strategy in state Next select action The probability of Is the old policy in the same state Next select action The probability of is the cropping operation, which changes the ratio Restricted to within the scope; A loss function construction unit, which is used to construct a PPO loss function, wherein the PPO loss function includes an overall loss function , the overall loss function Based on the strategy loss function , value loss function and entropy regularization function Weighted to obtain; the overall loss function satisfy: ;
[0126] and is the weight coefficient; Wherein, the value loss function is: ;
[0127] in, is the time step Expected value; For the value network in state The predicted value under is the target value output by the value network; The entropy regularization function is: ;
[0128] in, is the time step Expected value; For the current strategy in state Next select action probability; For the current strategy in state Next select action The logarithm of the probability of An updating unit, which is used to minimize the overall loss using a gradient descent algorithm and update the parameters of the policy network and the value network; The storage unit is used to save the parameters of the current strategy as the parameters of the old strategy after completing a strategy update; until a preset number of strategy updates is reached.
[0129] Furthermore, the step of the action generation unit using the current state of the policy network to generate an action includes: Keep the compressor in operation and collect the current status once every compressor motor rotation cycle ; According to the current status Output actions through the policy network ; Execute an action , get the time step Instant Rewards and the next state ; The current state ,action , Instant Rewards and the next state Stored in the experience replay buffer; randomly selecting samples from the experience replay buffer; The evaluation unit uses generalized advantage estimation to evaluate the value of each action of the randomly selected sample , and construct a policy loss function based on the value.
[0130] Furthermore, the policy network includes four fully connected layers, each layer including 64 neurons; The value network includes four fully connected layers, each layer includes 64 neurons.
[0131] This application can accurately suppress compressor vibration without sensors through real-time learning and optimization. This application uses the rotation fluctuation amplitude and vibration trend quantification parameters of the compressor motor to build a reinforcement learning model and generate a compensation current signal, thereby effectively reducing the vibration caused by pressure and flow fluctuations in the compressor during the working cycle and improving the stability and service life of the equipment.
[0132] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit the same. Although the present application has been described in detail with reference to the aforementioned embodiments, it is still possible for a person skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to replace some of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions claimed to be protected by the present application.
Claims
1. A method for adaptively suppressing compressor vibration based on reinforcement learning, characterized in that: The following steps are involved: Obtain vibration quantity quantification parameters, the vibration quantity quantification parameters include: the rotation fluctuation amplitude of the compressor motor rotation period , the rotation fluctuation amplitude of the compressor motor rotation period The difference between the maximum speed and the minimum speed in the speed sequence corresponding to each compressor motor rotation cycle; Acquire a vibration trend quantification parameter based on the vibration quantity quantification parameter, wherein the vibration trend quantification parameter includes one or more of a sliding mean, a change rate, a cumulative change, and a standard deviation; Building a reinforcement learning model based on the vibration quantity quantified parameter and the vibration trend quantified parameter; training the reinforcement learning model; Deploy the reinforcement learning model, and calculate the compensation torque current parameter according to the output of the reinforcement learning model. and compensation angle parameters , generating a compensation current signal , the compensation current signal is , is the time variable, is the angular frequency, , is the rotation frequency of the compressor motor rotor; The frequency converter receives the compensation current signal , the compensation current signal The superimposed current control signal is used to drive the compressor motor to operate, thereby suppressing the vibration caused by the fluctuation of pressure and flow rate experienced by the compressed gas in each working cycle.
2. The method for adaptively suppressing compressor vibration based on reinforcement learning according to claim 1, characterized in that: The reinforcement learning model is constructed based on the PPO algorithm, and constructing the reinforcement learning model includes the following steps: Constructing a state space, the state space including: the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change, the standard deviation, the compressor motor average speed and the load current; Construct an action space, the action space including the compensation torque current parameter and compensation angle parameters ; Construct a reward function, which is: ;in, is the time step Instant rewards received, , and is the weight coefficient; Build a strategic network; Build a value network.
3. The method for adaptively suppressing compressor vibration based on reinforcement learning according to claim 2, characterized in that: Training the reinforcement learning model includes the following steps: Using the policy network in the current state to generate an action corresponding to the action space; Assess the value of each action using generalized advantage estimation , and construct a strategy loss function based on the value; The strategy loss function is: ; in, is the time step Expected value; For the current strategy in state Next select action probability; Is the old policy in the same state Next select action probability; is the cropping operation, which changes the ratio Restricted to within the scope; Construct a PPO loss function, which includes an overall loss function , the overall loss function Based on the strategy loss function , value loss function and entropy regularization function Weighted to obtain; the overall loss function satisfy: ; and is the weight coefficient; Wherein, the value loss function is: ; in, is the time step Expected value; is the predicted value of the value network in state, is the target value output by the value network; The entropy regularization function is: ; in, is the time step Expected value; For the current strategy in state Next select action probability; For the current strategy in state Next select action The logarithm of the probability of Minimize the overall loss using a gradient descent algorithm and update the parameters of the policy network and the value network; After completing a strategy update, the parameters of the current strategy are saved as the parameters of the old strategy; until the preset number of strategy updates is reached.
4. The method for adaptively suppressing compressor vibration based on reinforcement learning according to claim 3, characterized in that: The step of using the policy network of the current state to generate an action corresponding to the action space comprises: Keep the compressor in operation and collect the current status once every compressor motor rotation cycle ; According to the current status Output actions through the policy network ; Execute an action , get the time step Instant Rewards and the next state ; The current state ,action , Instant Rewards and the next state Stored in the experience replay buffer; randomly selecting samples from the experience replay buffer; Use generalized advantage estimation to evaluate the value of each action in a randomly selected sample , and construct a policy loss function based on the value.
5. The method for adaptively suppressing compressor vibration based on reinforcement learning according to any one of claims 2 to 4, characterized in that: The policy network includes four fully connected layers, each layer includes 64 neurons; The value network includes four fully connected layers, each layer includes 64 neurons.
6. A compressor vibration adaptive suppression device based on reinforcement learning, characterized in that: include: The first acquisition unit is used to acquire the vibration quantity quantitative parameters, wherein the vibration quantity quantitative parameters include: the rotation fluctuation amplitude of the compressor motor rotation period , the rotation fluctuation amplitude of the compressor motor rotation period The difference between the maximum speed and the minimum speed in the speed sequence corresponding to each compressor motor rotation cycle; a second acquisition unit, configured to acquire a vibration trend quantification parameter based on the vibration quantity quantification parameter, wherein the vibration trend quantification parameter includes one or more of a sliding mean, a change rate, a cumulative change, and a standard deviation; A construction unit, which is used to construct a reinforcement learning model based on the vibration amount quantification parameter and the vibration trend quantification parameter; A training unit, which is used to train the reinforcement learning model; An application unit is used to deploy the reinforcement learning model and calculate the compensation torque current parameter output by the reinforcement learning model. and compensation angle parameters , generating a compensation current signal , the compensation current signal is , is the time variable, is the angular frequency, , is the rotation frequency of the compressor motor rotor; A compensation unit, which is used to enable the frequency converter to receive the compensation current signal , the compensation current signal The superimposed current control signal is used to drive the compressor motor to operate, thereby suppressing the vibration caused by the fluctuation of pressure and flow rate experienced by the compressed gas in each working cycle.
7. The compressor vibration adaptive suppression device based on reinforcement learning according to claim 6, characterized in that: The construction unit is used to construct a reinforcement learning model based on the PPO algorithm, which includes: A state space construction unit, which is used to construct a state space, wherein the state space includes: the vibration quantity quantization parameter, the sliding mean, the change rate, the cumulative change, the standard deviation, the average speed of the compressor motor and the load current; An action space construction unit, which is used to construct an action space, wherein the action space includes the compensation torque current parameter and compensation angle parameters ; A reward function construction unit, which is used to construct a reward function, wherein the reward function is: ,in, is the time step Instant rewards received, , and is the weight coefficient; A strategy network construction unit, which is used to construct a strategy network; The value network construction department is used to construct the value network.
8. The compressor vibration adaptive suppression device based on reinforcement learning according to claim 7, characterized in that: The training unit comprises: An action generation unit, configured to generate an action corresponding to the action space using the policy network in a current state; The evaluation part is used to evaluate the value of each action using generalized advantage estimation , and construct a strategy loss function based on the value; The strategy loss function is: ; in, is the time step Expected value; For the current strategy in state Next select action probability; Is the old policy in the same state Next select action probability; is the cropping operation, which changes the ratio Restricted to within the scope; A loss function construction unit, which is used to construct a PPO loss function, wherein the PPO loss function includes an overall loss function , the overall loss function Based on the strategy loss function , value loss function and entropy regularization function Weighted to obtain; the overall loss function satisfy: ; and is the weight coefficient; Wherein, the value loss function is: ; in, is the time step Expected value; For the value network in state The predicted value under is the target value output by the value network; The entropy regularization function is: ; in, is the time step Expected value; For the current strategy in state Next select action probability; For the current strategy in state Next select action The logarithm of the probability of An updating unit, which is used to minimize the overall loss using a gradient descent algorithm and update the parameters of the policy network and the value network; The storage unit is used to save the parameters of the current strategy as the parameters of the old strategy after completing a strategy update; until a preset number of strategy updates is reached.
9. The compressor vibration adaptive suppression device based on reinforcement learning according to claim 8, characterized in that: The step of the action generation unit using the current state of the policy network to generate an action corresponding to the action space includes: Keep the compressor in operation and collect the current status once every compressor motor rotation cycle ; According to the current status Output actions through the policy network ; Execute an action , get the time step Instant Rewards and the next state ; The current state ,action , Instant Rewards and the next state Stored in the experience replay buffer; randomly selecting samples from the experience replay buffer; The evaluation unit uses generalized advantage estimation to evaluate the value of each action of the randomly selected sample , and construct a policy loss function based on the value.
10. The compressor vibration adaptive suppression device based on reinforcement learning according to any one of claims 7 to 9, characterized in that: The policy network includes four fully connected layers, each layer includes 64 neurons; The value network includes four fully connected layers, each layer includes 64 neurons.
Citation Information
Patent Citations
Control method and system of compressor
CN103470483A
Motor drive device and compressor using the same
CN103684165A
Air conditioner inverter compressor frequency-domain constant torque control system and method
CN104728090A
Compressor low-frequency vibration suppression method and system
CN111277189A
Air conditioner and method for restraining low-frequency vibration of compressor
CN114517937A
Cited By
Compressor energy-saving operation control method and system based on reinforcement learning
CN121165467A
A compressor energy-saving operation control method and system based on reinforcement learning
CN121165467B