Compressor control method, controller, air conditioner, heat management system and vehicle

By applying reinforcement learning models in the compressor control system and learning the best control strategy, the problems of sensor dependence and delay in the prior art are solved, and faster and more efficient compressor control is achieved.

CN120056681APending Publication Date: 2025-05-30BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311630813.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing compressor control methods rely on the accuracy of temperature and pressure sensors, and there are control errors and delays, making it difficult to quickly meet user needs.

Method used

The reinforcement learning model is used to learn the compressor speed control, and the model is inputted through environmental parameters to obtain the target speed control instructions, reducing the dependence on sensor performance.

Benefits of technology

It achieves faster control speed, reduces dependence on sensor performance, avoids delay, and improves user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120056681A_ABST
    Figure CN120056681A_ABST
Patent Text Reader

Abstract

The invention discloses a compressor control method, a controller, an air conditioner, a heat management system and a vehicle, and the compressor control method comprises the steps that environmental parameters are obtained, and the environmental parameters are used for being input into a reinforcement learning model; obtaining an output value of the reinforcement learning model, wherein the output value represents a target rotating speed control instruction of the compressor; and controlling the rotating speed of the compressor according to the target rotating speed control instruction of the compressor. By adopting the method, the control of the rotating speed of the compressor can be learned through the reinforcement learning model, and a control strategy with the lowest energy consumption and the fastest time consumption is realized, so that the control method has the advantages of good self-learning property, relatively high robustness and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of vehicles, and in particular, to a compressor control method, a controller, an air conditioning system, a thermal management system, and a vehicle. Background Art

[0002] In related technologies, the current compressor control method collects the in-vehicle temperature and the low pressure of the air conditioning system to judge the temperature change trend; within a control period, according to the deviation value △T between the in-vehicle temperature and the user-set temperature and the temperature change trend, the compressor capacity demand coefficient K is calculated, and then the compressor output frequency is calculated; within a compensation period, according to the deviation value △P between the low pressure of the air conditioning system and the target low pressure, the low pressure compensation control method is used to calculate the low pressure compensation value, and the compressor output frequency is changed according to the low pressure compensation value. This method depends on the temperature and pressure collection measuring points in the air conditioning system and their accuracy. Once the sensor fails, it will cause control errors; temperature feedback is required, and there is a certain time delay, making it difficult to quickly meet user needs. Summary of the Invention

[0003] The present invention aims to solve at least one of the technical problems existing in the prior art. For this reason, the first object of the present invention is to propose a compressor control method. With this method, the compressor speed control can be learned through a reinforcement learning model, and the best control strategy can be learned relatively quickly. During the application process, the control speed is faster and the dependence on the sensor performance is less.

[0004] The second object of the present invention is to propose a controller.

[0005] The third object of the present invention is to propose an air conditioning system.

[0006] The fourth object of the present invention is to propose a thermal management system.

[0007] The fifth object of the present invention is to propose a vehicle.

[0008] To solve the above problems, an embodiment of the first aspect of the present invention provides a compressor control method, which is characterized by including: obtaining environmental parameters, where the environmental parameters are used to input a reinforcement learning model; obtaining an output value of the reinforcement learning model, where the output value represents a target speed control instruction for the compressor; and controlling the compressor speed according to the target speed control instruction of the compressor.

[0009] According to the compressor control method of the embodiment of the present invention, by using a reinforcement learning model, the best control strategy can be learned relatively quickly. During the application process of controlling the compressor, the control speed is faster and the dependence on the sensor performance is less. Temperature feedback is not required, avoiding delay and improving user satisfaction.

[0010] In some embodiments, the reinforcement learning model is constructed by learning the variation relationship between the space temperature of the air-conditioning temperature control space corresponding to the compressor speed and / or the compressor energy consumption.

[0011] In some embodiments, when the reinforcement learning model is constructed, the variation relationship between the compressor speed control and the space temperature and / or the compressor energy consumption is learned through simulation, experiment or vehicle test.

[0012] In some embodiments, the reinforcement learning model adopts a value function approximation Q-network; in the value function approximation Q-network, the speed control command of the compressor is an action, and the time for the space temperature of the air-conditioning temperature control space to reach the target temperature and the compressor energy consumption during the temperature control process are used as the reward values.

[0013] In some embodiments, the reinforcement learning model is used to determine the target value function value according to the environmental parameters;

[0014] The target control command of the compressor is the action in the system state where the compressor is located corresponding to the target value function value.

[0015] In some embodiments, the target value function value is the value function value corresponding to the maximum reward value of the system state where the compressor is located under the environmental parameters.

[0016] In some embodiments, in the value function approximation Q-network, the immediate reward corresponding to each system state where the compressor is located is a fixed value.

[0017] In some embodiments, in the value function approximation Q-network, the immediate reward of the system state where the compressor is located at the (t + 1) -th moment is the sum of a first reciprocal and a second reciprocal; the first reciprocal is the reciprocal of the temperature difference between the space temperature at the (t + 1) -th moment and the target temperature; the second reciprocal is the reciprocal of the increase in the compressor energy consumption after the system state where the compressor is located is changed from the t-th moment to the (t + 1) -th moment.

[0018] In some embodiments, when the space temperature reaches the target temperature, the final reward of the value function approximation Q-network is the sum of a third reciprocal and a fourth reciprocal; the third reciprocal is the reciprocal of the time required for the space temperature to reach the target temperature; the fourth reciprocal is the reciprocal of the energy consumption of the compressor when the space temperature reaches the target temperature.

[0019] In some embodiments, when training the reinforcement learning model, at least one of the following constraint conditions is satisfied: the temperature adjustment time exceeds the benchmark temperature adjustment time or the compressor energy consumption is higher than the benchmark energy consumption, the reward value is negative and the model training is terminated; during the process of controlling the change of the compressor speed, the changes in the system state where the compressor is located all satisfy the system safety constraints; the compressor speed control command fluctuates within an allowable range, the duration of a single change in the compressor speed control command is fixed, and the change in the compressor speed control command increases or decreases monotonically with time.

[0020] In some embodiments, the hyperparameters of the reinforcement learning model include: the activation function is the Sigmoid function, the weight initialization method is the normal distribution, the loss function is the mean squared error, the optimization algorithm is Adam optimization, and the batch size for network training is 100.

[0021] In some embodiments, the hyperparameters of the reinforcement learning model include that the learning rate is between 0.0001 and 0.001, the number of neurons is between 50 and 100, and the number of network layers is between 1 and 3.

[0022] An embodiment of the second aspect of the present invention provides a controller, which is characterized by including: a processor configured with the reinforcement learning model; a memory communicatively connected to the processor; the memory stores a computer program executable by the processor, and when the processor executes the computer program, it implements the compressor control method described in the above embodiments.

[0023] According to the controller of the embodiment of the present invention, after obtaining the output value of the reinforcement learning model, the controller controls the compressor speed through the output value of the reinforcement learning model, realizes the control strategy with the shortest time, has less dependence on the sensor performance, has small latency, and improves the user experience.

[0024] An embodiment of the third aspect of the present invention provides an air conditioning system, which is characterized by including: a compressor; the controller described in the above embodiments, and the controller is connected to the compressor.

[0025] According to the air conditioning system of the embodiment of the present invention, after the controller obtains the output value of the reinforcement learning model, it controls the compressor speed through the output value of the reinforcement learning model, comprehensively considers the temperature adjustment speed and the process energy consumption, realizes the control strategy with the lowest energy consumption and the shortest time, and improves the user satisfaction.

[0026] An embodiment of the fourth aspect of the present invention provides a thermal management system, which is characterized by including the air conditioning system described in the above embodiments.

[0027] According to the thermal management system of the embodiment of the present invention, by adopting the air conditioning system of the above embodiment, the thermal management can be realized more quickly, and the user satisfaction is improved.

[0028] An embodiment of the fifth aspect of the present invention provides a vehicle, which is characterized by including the air conditioning system described in the above embodiment.

[0029] For the vehicle according to the embodiment of the present invention, by adopting the air conditioning system of the above embodiment, the control speed is faster during application, the dependence on the sensor performance is less, the energy consumption of the whole vehicle is reduced, and the user satisfaction is improved.

[0030] The additional aspects and advantages of the present invention will be partly given in the following description, partly become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0031] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:

[0032] Figure 1 is a flowchart of a compressor control method according to an embodiment of the present invention;

[0033] Figure 2 is a schematic diagram of the parameters of a vehicle thermal management system according to an embodiment of the present invention;

[0034] Figure 3 is a schematic diagram of the update of a compressor speed control instruction according to an embodiment of the present invention;

[0035] Figure 4 is a schematic diagram of the Q-network training according to an embodiment of the present invention;

[0036] Figure 5 is a flowchart of an intelligent adjustment scheme for the compressor speed of a reinforcement learning model according to an embodiment of the present invention;

[0037] Figure 6 is a structural block diagram of a controller according to an embodiment of the present invention;

[0038] Figure 7 is a structural block diagram of an air conditioning system according to an embodiment of the present invention;

[0039] Figure 8 is a structural block diagram of a thermal management system according to an embodiment of the present invention;

[0040] Figure 9 is a structural block diagram of a vehicle according to an embodiment of the present invention.

[0041] Reference Signs:

[0042] Controller 10; Air conditioning system 20; Thermal management system 30; Vehicle 40;

[0043] Processor 1; Memory 2; Compressor 3. Detailed implementation

[0044] The embodiments of the present invention will be described in detail below. The embodiments described with reference to the accompanying drawings are exemplary. The embodiments of the present invention will be described in detail below.

[0045] Figure 1 is a flowchart of a compressor control method according to an embodiment of the present invention. As Figure 1 shown, the compressor control method includes steps S1 - S3.

[0046] S1, Obtain environmental parameters, which are used to input into the reinforcement learning model.

[0047] Specifically, build a reinforcement learning model for the compressor control method. After obtaining the environmental parameters, input them into the reinforcement learning model, that is, the environmental parameters serve as the input of the reinforcement learning model. Among them, the reinforcement learning model can be pre - constructed, for example, through model training to meet the requirements, and the model is pre - stored in the processor.

[0048] In some embodiments, the reinforcement learning model may include an agent, an environment, a state, an action, and a reward, etc. After the agent executes an action, the environment will transition to a new state. For this new state, the environment will give a reward signal (positive reward or negative reward); subsequently, the agent, according to the new state and the reward feedback from the environment, executes a new action according to a certain strategy; through reinforcement learning, the agent can know in what state it should take what kind of action to maximize its own reward.

[0049] S2, Obtain the output value of the reinforcement learning model, and the output value represents the target speed control instruction for the compressor.

[0050] Specifically, after obtaining the environmental parameters, input them into the reinforcement learning model. The reinforcement learning model processes them according to the internal strategy to obtain the output value. In the embodiments of the present invention, the output value of the reinforcement learning model represents the target speed control instruction for the compressor; the target speed of the compressor can be understood as the best working speed of the compressor in the current system state, for example, the compressor speed corresponding to the maximum reward. Determine the current target speed of the compressor according to the output value of the reinforcement learning model.

[0051] S3, Control the compressor speed according to the target speed control instruction of the compressor.

[0052] Specifically, after inputting the current environmental parameters into the reinforcement learning model, the reinforcement learning model processes data based on the model parameters, structure, and policy, and then outputs an output value. At this time, the output value of the reinforcement learning model is the optimal operating speed of the compressor under the current state of the system. Using the output value of the reinforcement learning model as the target speed control command, the compressor is controlled to operate at the target speed.

[0053] According to the compressor control method of the embodiments of the present invention, a reinforcement learning model is adopted. The reinforcement learning model can obtain a relatively fast training speed relying on the number of computing cores, and the model occupies a small space, has high prediction accuracy and fast prediction speed, and can be conveniently and quickly applied to different environments. During the application process of controlling the compressor, the control speed is faster, the dependence on sensors is small, no temperature feedback is required, the latency is small, and the user satisfaction is improved.

[0054] In some embodiments, the reinforcement learning model is constructed by learning the variation relationship between the space temperature of the air-conditioning temperature adjustment space corresponding to the compressor speed and / or the compressor energy consumption.

[0055] Specifically, the temperature change of the air-conditioning temperature adjustment space and the change of the compressor energy consumption brought by the action of determining the compressor speed are determined through the reinforcement learning model. Different compressor speeds correspond to different space temperatures and / or compressor energy consumptions; the air-conditioning temperature adjustment space can be the passenger compartment of a vehicle, and the space temperature is the passenger compartment temperature; that is, the reinforcement learning model learns by controlling the temperature change of the passenger compartment of the vehicle and / or the change of the compressor energy consumption by adjusting the compressor speed, and while making the temperature of the passenger compartment of the vehicle reach the target temperature, the compressor speed is reduced as much as possible to reduce the compressor energy consumption.

[0056] In some embodiments, when constructing the reinforcement learning model, the variation relationship between the compressor speed control and the corresponding space temperature and / or the compressor energy consumption is learned through simulation, experiment, or vehicle test.

[0057] Specifically, the reinforcement learning model can be a vehicle thermal management system simulation model, such as Figure 2 shown, including main components such as a compressor, a condenser, a valve, an evaporator, etc. The input parameters mainly include the compressor speed, and the output parameters mainly include the temperature change of the passenger compartment and the compressor power. The valve opening is adjusted through the operation mode (such as adjusting the valve opening to control the superheat at the evaporator outlet during refrigeration and adjusting the valve opening to control the suction superheat of the compressor during heating), and the blower speed automatically adjusts the gear according to the operation mode. Therefore, by controlling the compressor speed, the compressor power, the suction and discharge pressures can be changed, thereby changing the internal cooling or internal evaporation temperature and then controlling the temperature change process in the passenger compartment; the simulation model can be replaced by experiments and vehicle tests to learn the variation relationship between the compressor speed control and the corresponding space temperature and / or the compressor energy consumption.

[0058] In some embodiments, the reinforcement learning model employs a value function approximation Q-network; in the value function approximation Q-network, the rotational speed control command of the compressor is the action, and the time for the space temperature in the air-conditioning temperature control space to reach the target temperature and the compressor energy consumption during the temperature control process are used as the reward values.

[0059] Specifically, the Q-network is a method that combines a neural network and Q-learning. The state is directly used as the input of the neural network, and the neural network calculates the action values of all actions and selects the maximum value as the output, or both the state and the action are used as the input of the neural network to directly output the corresponding Q-value; as long as there is a Q-network, reinforcement learning can be performed; with the value function, it is possible to decide which action to take and perform policy improvement according to the value function; therefore, in the value function approximation Q-network, the rotational speed control command of the compressor can be used as the action, and the time for the space temperature in the air-conditioning temperature control space to reach the target temperature and the compressor energy consumption during the temperature control process can be used as the reward values.

[0060] In some embodiments, the reinforcement learning model is used to determine the target value function value according to the environmental parameters; the target control command of the compressor is the action in the system state where the compressor is located corresponding to the target value function value.

[0061] Specifically, the Q-network is initialized with the goal of inputting the state St at time t (S includes the occupant compartment temperature and the compressor rotational speed at this moment, that is, the environmental parameters) and outputting the target value function values of each compressor rotational speed control command in the state St.

[0062] Since the state of the compressor rotational speed control during the temperature control process is a continuous state space, the target value function Q(S,a) cannot be represented in tabular form. Therefore, given the state St at time t and the compressor rotational speed control command at, the empirical knowledge of each compressor rotational speed control command in the state St, that is, the target value function value, will be output. Thus, a value function approximation Q-network (abbreviated as Q-network) will be used to approximate the target value function value Q(S,a), where the Q-network will be written as Q(S,a,w), and w is the parameter to be trained; before training, the Q-network is initialized to 0 or a non-zero minimum value, and at the same time, the parameter w to be trained is initialized.

[0063] In some embodiments, the target value function value is the value function value corresponding to the maximum reward value of the system state where the compressor is located under the environmental parameters.

[0064] Specifically, the reward is the feedback information obtained after performing an action. For example Figure 3As shown, using the ε-greedy strategy, select a compressor speed control instruction at (this instruction can be the value of speed change or the speed value), input at into the environment to obtain the new state St+1 and Rt+1; set an ε value for determination, and then randomize a value. If the value is less than ε, randomly select the compressor speed control instruction at, and if the value is greater than ε, select the compressor speed control instruction at with the largest known target value function value.

[0065] During the training process, this ε value can be gradually reduced to reduce the probability of the network randomly selecting values, and ensure the stability of the model in the later stage of training; usually, ε is set to a very small value, and 1 - ε may be 90%, that is, there is a 90% probability of determining the action according to the Q function, but there is a 10% probability of being random. Usually, ε decreases over time in implementation. At the very beginning, because we don't know which action is better, we will spend more effort on exploration. Next, as the number of training times increases, we are more certain about which Q is better, so we will reduce exploration and make the value of ε smaller. Therefore, as the number of training times increases, the target value function value can be determined as the value function value corresponding to the maximum reward value of the system state where the compressor is located under the environmental parameters.

[0066] In some embodiments, in the value function approximation Q network, the immediate reward corresponding to each system state where the compressor is located is a fixed value.

[0067] Specifically, input the compressor control instruction at obtained from the previous strategy into the model in the first step to obtain the state St+1 that can be obtained by using the compressor speed control instruction at in the state St and the obtained immediate reward Rt+1; further, the immediate reward Rt+1 can be given according to a fixed value.

[0068] In some embodiments, in the value function approximation Q network, the immediate reward of the system state where the compressor is located at the (t + 1)th moment is the sum of the first reciprocal and the second reciprocal; the first reciprocal is the reciprocal of the temperature difference between the space temperature and the target temperature at the (t + 1)th moment; the second reciprocal is the reciprocal of the increase in compressor energy consumption after the system state where the compressor is located changes from the tth moment to the (t + 1)th moment.

[0069] Specifically, the immediate reward Rt+1 can be given according to a fixed value, or can be calculated according to the fact that the immediate reward of the system state where the compressor is located at the (t + 1)th moment is the sum of the first reciprocal and the second reciprocal. The first reciprocal is the reciprocal of the temperature difference between the space temperature and the target temperature at the (t + 1)th moment; the second reciprocal is the reciprocal of the increase in compressor energy consumption after the system state where the compressor is located changes from the tth moment to the (t + 1)th moment. Calculate the immediate reward of the system state where the compressor is located at the (t + 1)th moment according to the first reciprocal and the second reciprocal.

[0070] In some embodiments, when the space temperature reaches the target temperature, the final reward of the value function approximation Q-network is the sum of a third reciprocal and a fourth reciprocal; the third reciprocal is the reciprocal of the time required for the space temperature to reach the target temperature; the fourth reciprocal is the reciprocal of the energy consumption of the compressor when the space temperature reaches the target temperature.

[0071] Specifically, when the space temperature reaches the target temperature, that is, when the temperature of the passenger compartment reaches the preset temperature, the final reward of the value function approximation Q-network is the sum of a third reciprocal and a fourth reciprocal. The third reciprocal is the reciprocal of the time required for the space temperature to reach the target temperature; the fourth reciprocal is the reciprocal of the energy consumption of the compressor when the space temperature reaches the target temperature. The final reward is determined by the sum of the three reciprocals and the fourth reciprocal.

[0072] In some embodiments, when training the reinforcement learning model, at least one of the following constraint conditions is satisfied: when the temperature adjustment time exceeds the reference temperature adjustment time or the compressor energy consumption is higher than the reference energy consumption, the reward value is negative and the model training is terminated; during the period of controlling the change of the compressor speed, the changes in the system state where the compressor is located all satisfy the system safety constraints; the compressor speed control command is maintained within the allowable range of fluctuations, the duration of a single change in the compressor speed control command is fixed, and the change in the compressor speed control command is monotonically increasing or decreasing with time.

[0073] Specifically, if the temperature adjustment time exceeds the reference temperature adjustment time, or the process energy consumption is higher than the reference energy consumption (the reference value can be given by the existing control strategy), a relatively large negative value is added to the policy reward value and the training is terminated, thereby reducing the cost required for training.

[0074] During the period of controlling the change of the compressor speed, the changes in the system state where the compressor is located all satisfy the system safety constraints, and the threshold value of the safety constraints can be determined by the user's safety level. For example, the speed change during the process does not exceed 1000 rpm / s, and the pressure change at each measuring point does not exceed 5 bar / s, etc.

[0075] The compressor speed control command needs to be maintained within the allowable range of fluctuations, for example, between 0 - 7000 rpm; the duration of a single change is fixed, for example, 10 s; the change in the compressor speed control command is monotonically increasing or decreasing with time, that is, the command will not reciprocate during the control process and is allowed to remain unchanged.

[0076] In some embodiments, the hyperparameters of the reinforcement learning model include: the activation function is the Sigmoid function, the weight initialization method is the normal distribution, the loss function is the squared loss, the optimization algorithm is the Adam optimization, and the batch size for network training is 100.

[0077] Specifically, train the Q-network and update the parameters to be trained so that the output value of the value function is close to the target value function.

[0078] As Figure 4 shown, calculate the target value yt of the Q network as yt = Rt+1 + γ·maxaQ(St+1, a; w), and calculate the loss function as the difference between yt and Q(S, a; w); then use the gradient descent strategy to update the parameter w in the Q network so that yt and Q(St, at) gradually approach and converge to the expected accuracy; further, determine the hyperparameters to be used in the deep learning model: the hyperparameters mainly include fixed parameters, such as the activation function being the Sigmoid function, the weight initialization method being the normal distribution, the loss function being the squared loss such as L = 1 / 2[yt - Q(S, a; w)]2, the optimization algorithm being Adam optimization, and the batch size of network training being 100.

[0079] In some embodiments, the hyperparameters of the reinforcement learning model include that the learning rate is between 0.0001 and 0.001, the number of neurons is between 50 and 100, and the number of network layers is between 1 and 3.

[0080] Specifically, the hyperparameters of the learning model include that the learning rate is between 0.0001 and 0.001, the number of neurons is between 50 and 100, and the number of network layers is between 1 and 3, and the neural network model is adjusted by regulating the variable parameters.

[0081] Next, with reference to Figure 5 shown, an example of the compressor control method according to the embodiments of the present invention will be described, and the specific content is as follows.

[0082] Step S4, start.

[0083] Step S5, build a simulation, experiment or vehicle model, and determine the change in the occupant compartment temperature and energy consumption brought about by the action of changing the compressor speed.

[0084] Step S6, initialize the Q network, with the goal of realizing inputting the state St at time t (S includes the occupant compartment temperature and energy consumption at this moment) and outputting the Q value of each compressor speed control instruction in the St state.

[0085] Step S7, use the ε-greedy strategy to select a compressor speed control instruction at, input at into the environment, and obtain the new state St+1 and Rt+1.

[0086] Step S8, put (St, at, rt, St+1) into the experience replay buffer, then randomly extract a batch of data for subsequent training, and update the experience replay buffer.

[0087] Further, the batch can be set to 128, and the size of the experience replay buffer can be set to 1024.

[0088] Step S9: Train the Q-network and update w to make the Q-value close to the target Q-value.

[0089] Step S10: When the temperature in the occupant compartment reaches the preset temperature, the final reward of this strategy is the sum of the reciprocals of the required time and the consumed energy, so as to determine the rotation speed control strategy with the fastest temperature adjustment time and the lowest energy consumption during the temperature adjustment process.

[0090] Step S11: End.

[0091] An embodiment of the second aspect of the present invention provides a controller, as Figure 6 shown, the controller 10 includes: a processor 1 and a memory 2.

[0092] Wherein, the memory 2 stores a computer program executable by the processor 1, and when the processor 1 executes the computer program, it implements the compressor control method of the above embodiment.

[0093] According to the controller of the embodiment of the present invention, after obtaining the output value of the reinforcement learning model, the controller controls the rotation speed of the compressor through the output value of the reinforcement learning model, realizes the control strategy with the fastest time, has less dependence on the sensor performance, has small latency, and improves the user experience.

[0094] An embodiment of the third aspect of the present invention provides an air conditioning system, as Figure 7 shown, the air conditioning system 20 includes: a compressor 3 and a controller 10.

[0095] Wherein, the controller 10 is connected to the compressor 3.

[0096] Specifically, the controller 10 in the air conditioning system 20 obtains the output value of the reinforcement learning model, represents the output value as the target rotation speed control instruction of the compressor 3, and when the controller 10 obtains the output value of the reinforcement learning model, controls the compressor 3 to operate with the output value as the target rotation speed of the compressor.

[0097] According to the air conditioning system of the embodiment of the present invention, after the controller obtains the output value of the reinforcement learning model, it controls the rotation speed of the compressor through the output value of the reinforcement learning model, comprehensively considers the temperature adjustment speed and the process energy consumption, realizes the control strategy with the lowest energy consumption and the fastest time, and improves the user satisfaction.

[0098] An embodiment of the fourth aspect of the present invention provides a thermal management system, as Figure 8 shown, the thermal management system 30 includes: an air conditioning system 20.

[0099] Specifically, the thermal management system 30 adjusts the space temperature through the air conditioning system 20. During the adjustment process, the controller in the air conditioning system 20 obtains the output value of the reinforcement learning model, represents the output value as the target rotational speed control command for the compressor, and when the controller obtains the output value of the reinforcement learning model, it controls the compressor to operate with the output value as the target rotational speed of the compressor.

[0100] According to the thermal management system of the embodiment of the present invention, by adopting the air conditioning system of the above embodiment, thermal management can be achieved more quickly, improving user satisfaction.

[0101] An embodiment of the fifth aspect of the present invention provides a vehicle, as Figure 9 shown, the vehicle 40 includes: an air conditioning system 20.

[0102] Specifically, the vehicle 40 adjusts the temperature of the vehicle occupant compartment through the air conditioning system 20. During the adjustment process, the controller in the air conditioning system 20 obtains the output value of the reinforcement learning model, represents the output value as the target rotational speed control command for the compressor, and when the controller obtains the output value of the reinforcement learning model, it controls the compressor to operate with the output value as the target rotational speed of the compressor.

[0103] According to the vehicle of the embodiment of the present invention, by adopting the air conditioning system of the above embodiment, the control speed is faster during application, the dependence on sensor performance is less, the vehicle energy consumption is reduced, and user satisfaction is improved.

[0104] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example.

[0105] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A compressor control method, characterized in that, it includes: Obtain environmental parameters, where the environmental parameters are used to input into a reinforcement learning model; Obtain the output value of the reinforcement learning model, where the output value represents the target speed control instruction for the compressor; Control the speed of the compressor according to the target speed control instruction of the compressor.

2. The compressor control method according to claim 1, characterized in that, The reinforcement learning model is constructed by learning the variation relationship between the space temperature and / or the compressor energy consumption in the air-conditioning temperature adjustment space corresponding to the compressor speed.

3. The compressor control method according to claim 2, characterized in that, When constructing the reinforcement learning model, the variation relationship between the space temperature and / or the compressor energy consumption corresponding to the compressor speed is learned through simulation or experiment or vehicle test.

4. The compressor control method according to any one of claims 1-3, characterized in that, The reinforcement learning model adopts a value function approximation Q-network; In the value function approximation Q-network, the speed control instruction of the compressor is the action, and the time for the space temperature in the air-conditioning temperature adjustment space to reach the target temperature and the compressor energy consumption during the temperature adjustment process are used as the reward values.

5. The compressor control method according to claim 4, characterized in that, The reinforcement learning model is used to determine the target value function value according to the environmental parameters; The target control instruction of the compressor is the action in the system state where the compressor is located corresponding to the target value function value.

6. The compressor control method according to claim 4, characterized in that, The target value function value is the value function value corresponding to the maximum reward value of the system state where the compressor is located under the environmental parameters.

7. The compressor control method according to claim 4, characterized in that, In the value function approximation Q-network, the immediate reward corresponding to each system state where the compressor is located is a fixed value.

8. The compressor control method according to claim 4, characterized in that, In the value function approximation Q-network, the immediate reward of the system state where the compressor is located at the (t + 1)th moment is the sum of a first reciprocal and a second reciprocal; The first reciprocal is the reciprocal of the temperature difference between the space temperature at the (t + 1)th moment and the target temperature; The second reciprocal is the reciprocal of the increase in the compressor energy consumption after the system state where the compressor is located is changed from the tth moment to the (t + 1)th moment.

9. The compressor control method according to claim 4, characterized in that, When the space temperature reaches the target temperature, the final reward of the value function approximation Q-network is the sum of a third reciprocal and a fourth reciprocal; The third reciprocal is the reciprocal of the time required for the space temperature to reach the target temperature; The fourth reciprocal is the reciprocal of the energy consumption of the compressor when the space temperature reaches the target temperature.

10. The compressor control method according to claim 4, characterized in that, When training the reinforcement learning model, it satisfies at least one of the following constraint conditions: When the temperature adjustment time exceeds the reference temperature adjustment time or the compressor energy consumption is higher than the reference energy consumption, the reward value is negative and the model training is terminated; During the period of controlling the change of the compressor speed, the changes in the system state where the compressor is located all satisfy the system safety constraints; The compressor speed control instruction fluctuates within an allowable range, the duration of a single change in the compressor speed control instruction is fixed, and the change in the compressor speed control instruction increases or decreases monotonically with time.

11. The compressor control method according to claim 1, characterized in that, The hyperparameters of the reinforcement learning model include: the activation function is the Sigmoid function, the weight initialization method is the normal distribution, the loss function is the mean squared error, the optimization algorithm is Adam optimization, and the batch size of network training is 100.

12. The compressor control method according to claim 1, characterized in that, The hyperparameters of the reinforcement learning model include that the learning rate is between 0.0001 and 0.001, the number of neurons is between 50 and 100, and the number of network layers is between 1 and 3.

13. A controller, characterized in that, comprises: a processor, the processor is configured with the reinforcement learning model; a memory, communicatively connected to the processor; The memory stores a computer program executable by the processor, and when the processor executes the computer program, it implements the compressor control method according to any one of claims 1-12.

14. An air conditioning system, characterized in that, comprises: a compressor; The controller according to claim 13, the controller is connected to the compressor.

15. A thermal management system, characterized in that, comprises the air conditioning system according to claim 14.

16. A vehicle, characterized in that, comprises the air conditioning system according to claim 14.