A DDPG-based air conditioning and refrigeration control method for passenger compartment of pure electric vehicles

Through the DDPG-based reinforcement learning method, the compressor and blower of the pure electric vehicle air-conditioning system are automatically adjusted, solving the problems of insufficient human thermal comfort and high energy consumption in the existing system, and achieving thermal comfort optimization and energy consumption reduction in different environments.

CN116729060BActive Publication Date: 2025-10-03ZHONGQIYAN AUTOMOBILE INSPECTION CENT (CHANGZHOU) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202310591773.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2025-10-03
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

The existing air-conditioning and refrigeration systems of pure electric vehicles lack adaptability to human thermal comfort, resulting in suboptimal thermal comfort and high energy consumption.

Method used

A reinforcement learning method based on DDPG is used to train the passenger compartment air conditioning and refrigeration control system on a simulation platform. The compressor speed, blower speed, and damper opening are automatically adjusted. Combined with the passenger compartment three-dimensional model and the human thermal comfort evaluation model, the energy consumption and thermal comfort of the air conditioning system are optimized.

Benefits of technology

Achieve passenger thermal comfort targets in a variety of environments, reduce HVAC system energy consumption, and reduce engineer workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116729060B_ABST
    Figure CN116729060B_ABST
Patent Text Reader

Abstract

The present invention provides a DDPG-based passenger compartment air conditioning and cooling control method for pure electric vehicles. The method includes a passenger compartment air conditioning and cooling control module based on the DDPG algorithm, a reinforcement learning training environment, and a passenger compartment thermal flow and thermal comfort module. The reinforcement learning training environment includes a one-dimensional model of the vehicle air conditioning system and a passenger compartment thermal comfort prediction model; the passenger compartment thermal flow and thermal comfort module includes a three-dimensional passenger compartment model and a human thermal comfort model; and the passenger compartment air conditioning and cooling control module based on the DDPG algorithm includes an action network, an evaluation network, and an experience recovery pool. The passenger compartment thermal flow and thermal comfort module is converted into a passenger compartment thermal comfort prediction model in the reinforcement learning training environment using a deep learning approach. The passenger compartment air conditioning and cooling control module based on the DDPG algorithm continuously interacts with the reinforcement learning training environment to achieve effective training results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automobile dynamic control and artificial intelligence technology, and in particular to a DDPG-based air conditioning and refrigeration control method for a passenger compartment of a pure electric vehicle. Background Art

[0002] With the development of science and technology and the continuous improvement of people's living standards, cars, as an indispensable means of transportation, are increasingly entering every aspect of people's daily lives. As one of the main components affecting the comfort and safety performance of automobiles, automobile air conditioning can adjust the air temperature in the car cabin to improve the thermal comfort of passengers. Thermal comfort reflects the human body's subjective perception of the ambient thermal state in a sealed space. In the car cockpit, whether the air conditioning is operating in mode affects the passenger's perception of the air conditioning system and the interior design. Especially when the passenger is in the car for a long time, thermal comfort affects both the physiological and psychological feelings of the person, thus affecting driving safety. Therefore, passenger thermal comfort is a very important research direction when developing automobile air conditioning systems.

[0003] Traditional automotive air conditioning control is essentially temperature control, that is, the in-vehicle ambient temperature or the evaporator and surface temperatures reach target values. This control method ignores human thermal comfort to a certain extent. Traditional automotive air conditioning control methods such as PID control, fuzzy PID control, and PSO-based fuzzy PID control are conservative and cannot automatically adapt to complex and changing environmental conditions. To achieve accurate temperature control and avoid excessive energy consumption, automotive air conditioning calibration engineers are required to calibrate the controller. Therefore, the calibration workload of air conditioning controllers is huge and requires extremely high engineer experience.

[0004] Patent No. CN201310246901.4 invents a pure electric vehicle air-conditioning control method and its control system. The shortcomings of this invention patent are as follows: the control method adopted for the car is temperature-based control, which only determines whether the temperature inside the car can be stabilized at the set temperature when the air-conditioning is running, and does not consider the impact of wind speed, light, and humidity on comfort; Patent No. CN201820616898.9 invents a pure electric vehicle air-conditioning control system. The common shortcomings of these two inventions are as follows: the invention only considers the use requirements when the air-conditioning is running, and does not consider energy consumption factors. It cannot reduce power consumption while meeting the use requirements to achieve energy-saving effects; Patent No. CN202211160279.0 invents a pure electric vehicle air-conditioning control method. The shortcomings of this invention are as follows: the control method adopted for the car is target evaporator temperature control. This control method requires a large number of actual vehicle tests to calibrate the control strategy, which is a huge workload and very costly.

[0005] The existing air conditioning and refrigeration system of pure electric vehicles is a thermostat-type control, which lacks adaptability to human thermal comfort, so that the thermal comfort of the human body cannot be optimized. Therefore, there is an urgent need to provide an air conditioning and refrigeration control method that can well adapt to human thermal comfort. Summary of the Invention

[0006] The technical problem to be solved by the present invention is: in order to overcome the shortcomings of the existing technology, the present invention provides a pure electric vehicle passenger compartment air conditioning and refrigeration control method based on DDPG (Deep Deterministic Policy Gradient, deep deterministic policy gradient algorithm), which adopts reinforcement learning to realize the formulation of the control system, belonging to the field of automotive dynamic control and artificial intelligence.

[0007] The technical solution employed by the present invention to solve its technical problems is a DDPG-based method for controlling air conditioning and cooling in the passenger compartment of a pure electric vehicle. The technical concept involves training a DDPG-based reinforcement learning model on a simulation platform, interacting with the model through a virtual environment, and achieving the desired control effect by setting appropriate action spaces, state spaces, and reward functions. The trained air conditioning control system automatically adjusts the compressor speed, blower speed, and damper opening based on solar radiation intensity, interior and exterior vehicle temperatures, and vehicle speed, achieving a two-way optimization of improving passenger compartment thermal comfort and reducing vehicle air conditioning system energy consumption. After training, the code is compiled and flashed into the pure electric vehicle air conditioning controller to optimize the control of the actual vehicle air conditioning system.

[0008] The control method specifically includes the following steps:

[0009] S1: Build a passenger cabin human thermal comfort prediction model

[0010] S1.1: Construct a 3D passenger cabin model and a human thermal comfort evaluation model in 3D design software. The 3D passenger cabin model and the human thermal comfort evaluation model constitute the passenger cabin thermal flow field and thermal comfort module.

[0011] The 3D passenger compartment model, or 3D passenger compartment simulation model, is obtained by extracting the passenger compartment with the air conditioning system from the vehicle digital model. This model is simplified and meshed before being imported into the 3D simulation software. The 3D digital model is inspected for integrity within the 3D design software, and the relevant passenger compartment components are extracted. The 3D digital model is simplified and meshed, and the air conditioning dummy model is added to generate a volume mesh. Regions are set and named within the volume mesh. A physical model is created within the volume mesh, along with the physical model and boundary conditions. Multiple thermal comfort monitoring points are also set on the dummy model.

[0012] The human thermal comfort evaluation model, also known as the passenger cabin thermal comfort evaluation model, simulates the body's thermophysiological regulation mechanisms in different temperature environments. It calculates passenger thermal comfort evaluation values ​​by inputting key passenger physical characteristics and the air temperature, air velocity, mean radiant temperature, and relative humidity near thermal comfort monitoring points obtained through passenger cabin CFD simulation. There are 14 thermal comfort monitoring points: the passenger's head, torso, left forearm, left upper arm, left hand, right forearm, right upper arm, right hand, left thigh, left calf, left foot, right thigh, right calf, and right foot.

[0013] For details about the passenger cabin 3D model and the human thermal comfort evaluation model, please refer to the passenger cabin 3D simulation model and the passenger cabin thermal comfort evaluation model in the invention application with publication number CN 114757116 A, which will not be repeated here.

[0014] S1.2: Set the characteristic parameters of the passenger cabin thermal flow field and thermal comfort module based on the requirements of deep learning neural network training;

[0015] S1.3: Through a joint simulation of the passenger cabin three-dimensional model and the human thermal comfort evaluation model, extract the numerical values ​​corresponding to the characteristic parameters described in the simulation results as the data set for deep learning neural network training. Preprocess the data set and divide it into a training set and a validation set; the training set is used to train the model, and the validation set is used to verify the prediction effect of the trained model.

[0016] S1.4: Neural network training: Build a deep learning network based on the defined model structure, including input layer, hidden layer and output layer, initialize the corresponding weights and thresholds, and set the hyperparameters of neural network training, including optimizer, learning rate, number of iterations, time step and batch size; use the data of the training set to train the deep learning model, and use the backpropagation algorithm to update the weights and thresholds; obtain the passenger compartment human thermal comfort prediction model, and use the validation set to evaluate the prediction effect of the passenger compartment human thermal comfort prediction model.

[0017] S2: Build a passenger compartment air conditioning and cooling control strategy

[0018] S2.1: Build the passenger compartment air conditioning and refrigeration control module

[0019] A reinforcement learning model is defined based on the automobile air conditioning and refrigeration system. The state s, action a and reward r of the MDP process in the reinforcement learning model are determined. The passenger compartment air conditioning and refrigeration control module is determined based on the reinforcement learning model.

[0020] S2.2: Building the environment required for reinforcement learning training

[0021] The environment required for reinforcement learning training refers to a virtual environment established for training the reinforcement learning model, including a one-dimensional model of the automobile air-conditioning and refrigeration system and a passenger compartment human thermal comfort prediction model (a three-dimensional passenger compartment model and a thermal comfort evaluation model) obtained in step S1, wherein the one-dimensional model of the automobile air-conditioning and refrigeration system is used to simulate the operation of components in the automobile air-conditioning and refrigeration system according to the interior temperature, exterior temperature, solar radiation intensity, air humidity, vehicle speed, and control instructions of the automobile air-conditioning and refrigeration system, and output the air velocity and temperature data after the evaporator, the interior temperature data, and the energy consumption data of the air-conditioning system; the passenger compartment human thermal comfort prediction model is used to predict the human thermal comfort evaluation results based on the interior temperature, exterior temperature, solar radiation intensity, air humidity, the air velocity and temperature data after the evaporator, and feed back the human thermal comfort evaluation results to the passenger compartment thermal comfort control module.

[0022] S2.3: Reinforcement learning training for the cabin air conditioning and cooling control module

[0023] In the reinforcement learning training environment of step S3, the passenger compartment air conditioning and refrigeration control module constructed in step S2 is trained using a reinforcement learning control structure network based on the DDPG algorithm. During the training process, sample data is collected, and the passenger compartment air conditioning and refrigeration control module is updated and optimized based on the sample data. Once the passenger compartment air conditioning and refrigeration control module reaches a convergence state, the training is completed. At this point, the control strategy of the passenger compartment air conditioning and refrigeration control module is the target strategy - the passenger compartment air conditioning and refrigeration control strategy of the automobile air conditioning and refrigeration system.

[0024] The reinforcement learning model is trained on a simulation platform. Action a is input into the training environment. The training environment then provides the reinforcement learning model with state feedback based on action a. The reinforcement learning model then uses a reward strategy to determine the merits of the state change, thereby assessing the effectiveness of action a. To collect more rewards, the reinforcement learning model continuously explores, records, and summarizes the optimal behavioral decisions for each step. After sufficient training, the reinforcement learning model replaces the controller, accurately outputting the optimal action in various situations. Each interaction between the reinforcement learning model and the virtual environment requires a simulation cycle. The thermal flow field simulation of the three-dimensional passenger compartment model is extremely time-consuming, so a deep learning model prediction approach is used to replace the three-dimensional passenger compartment model and the human thermal comfort model. Given the complexity, nonlinearity, and coupling of automotive air conditioning and refrigeration systems, which generate a large amount of high-dimensional nonlinear data during operation, the present invention utilizes the DDPG algorithm to construct the reinforcement learning system.

[0025] S3: Application of control strategies

[0026] The trained passenger compartment air conditioning and refrigeration control module's control strategy is converted into code and burned into the vehicle's air conditioning controller, which serves as the actual vehicle's air conditioning and refrigeration control system to control and adjust the thermal comfort of the passenger compartment.

[0027] Furthermore, the characteristic parameters described in step S1.2 include the outdoor temperature of the vehicle, the solar radiation intensity, the air temperature behind the evaporator, the air flow velocity behind the evaporator, the air temperature of each air-conditioning outlet, the air flow velocity of each air-conditioning outlet, the air temperature of each part of the human body surface, the air flow velocity of each part of the human body surface, the average radiation temperature of each part of the human body surface, the relative humidity of each part of the human body surface, and the thermal comfort evaluation results of the human body.

[0028] Furthermore, step S1.3 specifically further includes the following steps:

[0029] Data collection: Using a three-dimensional passenger cabin model and a human thermal comfort evaluation model, we simulate and collect data on interior and exterior temperatures, solar radiation intensity, air humidity, air velocity and temperature after the evaporator, and passenger thermal comfort evaluation results.

[0030] The dataset is preprocessed by denoising the initial sample dataset, eliminating outliers, and interpolating missing values. The min-max standardization method is then used for normalization. The specific formula is as follows:

[0031]

[0032] In the formula, y is the normalized data; x is the original data; x min is the minimum value in the original data set; x max is the maximum value in the original data set;

[0033] Dataset division: The dataset is divided into a training set and a validation set in a ratio of 8:2. The training set is used to train the model, and the validation set is used to adjust the model parameters.

[0034] Furthermore, the neural network training in step S1.4 specifically includes the following steps:

[0035] S1.4.1: Build a deep learning network: Based on the defined model structure, build a deep learning network consisting of an input layer, a hidden layer, and an output layer. The input layer consists of six neurons, corresponding to the interior temperature, exterior temperature, solar radiation intensity, air humidity, and air velocity and temperature after the evaporator. The hidden layer consists of four neurons, using the ReLU activation function to extract features from the input data. The output layer consists of one neuron, which outputs the passenger thermal comfort evaluation results.

[0036] S1.4.2: Model preprocessing: Initialize the weights between the input layer and the hidden layer, the weights between the hidden layer and the output layer, and the thresholds of the hidden layer and the output layer. Use the Bayesian regularization algorithm for neural network training, the Adam optimizer, a learning rate of 0.001, 200 epochs, a batch size of 32, and a timestep of 2.

[0037] S1.4.3: Training model: Use the data from the training set to train the deep learning model, and use the backpropagation algorithm to update the weights and thresholds. The output error, that is, the difference between the expected output and the actual output, is calculated through the original path, reversed through the hidden layer to the input layer. During the backpropagation process, the error is distributed to each unit in each layer, and the error signal of each unit in each layer is obtained, which is used as the basis for correcting the weights of each unit. This calculation process is completed using the gradient descent method. After continuously adjusting the weights and thresholds of neurons in each layer, the error signal is reduced to a minimum.

[0038] S1.4.4: Model evaluation: The deep learning model is evaluated using yearly-based and station-based validation methods. The evaluation metrics include MSE, RMSE (Root Mean Squared Error), MAE (Meanabsolute Error), and R-Squared coefficient of determination. The formulas are as follows:

[0039]

[0040]

[0041]

[0042]

[0043] in, is the predicted value, y i is the true value, is the mean value, m is the number of samples;

[0044] When R 2 The larger the value and the smaller the values ​​of other indicators, the better the prediction effect of the model.

[0045] Furthermore, in step S1.4, when training the neural network, the cost function of the neural network is trained using Bayesian regularization to minimize the training error, wherein the cost function is:

[0046]

[0047] Where α1 and α2 are Bayesian hyperparameters that specify the direction the learning process seeks, i.e., minimizing the error or weight; n is the number of training samples; Y i is the actual value of the i-th; Y i ′ is the i-th predicted value of the neural network; m is the number of weights in the neural network, w j is the jth weight.

[0048] Furthermore, the definition of the reinforcement learning model in step S2.1 specifically includes the following process:

[0049] (1) Define the state s of the MDP process

[0050] The state information of the automobile air conditioning and refrigeration system is obtained, and the state s of the MDP (Markov Decision Processes) process is defined as: s = [s1, s2, s3, s4, s5, s6, s7], where s1 is the ambient temperature outside the vehicle, s2 is the solar radiation intensity, s3 is the temperature inside the vehicle, s4 is the vehicle speed, s5 is the air humidity, s6 is the passenger thermal comfort evaluation result, and s7 is the energy consumption per minute of the vehicle air conditioning system. Among them, s1, s2, s3, s4, and s5 are random inputs with a limited range, providing a learning environment under different working conditions.

[0051] (2) Define the action a of the MDP process

[0052] According to the output control instructions of the automobile air conditioning refrigeration system, the action a of the MDP process is defined as: a = [a1, a2, a3], where a1 is the blower speed, a2 is the compressor speed, and a3 is the damper opening.

[0053] (3) Define the reward r of the MDP process

[0054] Based on the main performance evaluation indicators of the automotive air conditioning and refrigeration system, the reward r of the MDP process is defined as: r = -E - λΔT, where E is the total energy consumption per minute of the automotive air conditioning and refrigeration system components when the passenger compartment heat load is balanced, and E takes a negative value; λ is the thermal comfort penalty function coefficient; and ΔT is the difference between the current thermal comfort evaluation result and the target thermal comfort evaluation result.

[0055] Considering that the main performance evaluation indicators of automobile air conditioning and refrigeration are divided into two parts: ① Human thermal comfort evaluation λΔT: the difference between the thermal comfort evaluation results of passengers in the vehicle and the optimal thermal comfort value (0); ② Energy consumption E: the total energy consumption per minute of the components of the automobile air conditioning and refrigeration system when the passenger compartment heat load is balanced. Therefore, the reward part is set to r = -E-λΔT. Since energy consumption should be minimized as much as possible, E takes a negative value. At the same time, in order to ensure the stability and effectiveness of the vehicle's thermal comfort control, a penalty function related to the difference between the thermal comfort evaluation results of passengers in the vehicle and the optimal thermal comfort value (0) is added. λ is the thermal comfort penalty function coefficient.

[0056] Specifically, the reinforcement learning control structure network of the DDPG algorithm described in step S2.3 includes an action network and an evaluation network, the action network includes a current policy network and a target policy network, and the evaluation network includes a current Q-value network and a target Q-value network, wherein the input information of the previous policy network is state s, and the output information is action a; the input and output of the target policy network are the same as the current policy network, and the current policy network parameters are copied periodically; the input information of the current Q-value network is state s and action a, and the output information is value Q; the input and output of the target Q-value network are the same as the current Q-value network, and the current Q-value network parameters are copied periodically.

[0057] Specifically, the training process in step S2.3 is:

[0058] (1) Update of action network

[0059] The current policy network interacts with the reinforcement learning training environment. The state s is input to the current policy network to obtain action a. Action a is applied to the reinforcement learning training environment, and the reinforcement learning training environment returns the state s' and reward r at the next moment. The sample data (s, a, r, s') at this time is collected and placed in the experience recycling pool. The target policy network is responsible for selecting the optimal next action a|s' based on the next state s' sampled in the experience recycling pool. The network structure of the target policy network is the same as that of the current policy network. The parameters of the target policy network are regularly copied from the parameters of the current policy network. When the current policy network applies action a to the reinforcement learning training environment, random action noise needs to be added to avoid excessive errors in training.

[0060] (2) Update of evaluation network

[0061] The current Q-value network is responsible for iteratively updating the value network parameter ω. Input S and a in (s, a, r, s') into the current Q-value network, calculate the value Q(s, a, ω) of the current Q-value network, input s' in (s, a, r, s') into the target policy network, obtain action a', and input s' and a' together into the target Q-value network to calculate the target Q value y i=r+YQ'(s',a',ω'), where Y is the discount factor; the target Q-value network is responsible for calculating the Q'(s',a',ω') part of the Q-value. The network structure is the same as the current Q-value network, and the network parameters are regularly copied from the current Q-value network. The following formula is used to calculate the loss function Loss of the current Q-value network:

[0062]

[0063] Meaning of the parameters: y i is the target Q value; i is the number of cycles; a' is the action output by the target policy network; ω' is the parameter of the value network; Q' is the Q value calculated by the target Q value network.

[0064] Among them, the role of the loss function is to describe the gap between the model's predicted value and the true value, and guide the model to move towards convergence during the training process.

[0065] Furthermore, when the current policy network applies action a to the reinforcement learning training environment, random action noise needs to be added to avoid excessive errors in training.

[0066] Existing pure electric vehicle air conditioning and refrigeration systems use thermostat-type controls, lacking adaptability to human thermal comfort, resulting in suboptimal thermal comfort. The present invention provides a DDPG-based pure electric vehicle passenger compartment air conditioning and refrigeration control method. The vehicle air conditioning and refrigeration system primarily comprises an expansion valve, a compressor, a blower, an evaporator, a front-end cooling module fan, an air conditioning damper, a solar radiation sensor, an outside temperature sensor, an inside temperature sensor, a vehicle speed sensor, and a vehicle air conditioning controller. The present invention involves programming trained DDPG-based pure electric vehicle passenger compartment thermal comfort control code into the vehicle air conditioning controller. The controller inputs data collected by the solar radiation sensor, outside temperature sensor, inside temperature sensor, and vehicle speed sensor. By controlling the compressor speed, blower speed, and HVAC damper opening, the controller adjusts the operating state of the vehicle air conditioning and refrigeration system, ensuring optimal thermal comfort for passengers in various environmental conditions.

[0067] The beneficial effects of the present invention are:

[0068] 1. Instead of using temperature as the control target in traditional automotive air conditioning control systems, this innovative system uses passenger thermal comfort as the control target, ensuring passenger thermal comfort in all environmental conditions.

[0069] 2. Taking the energy consumption of each component of the vehicle's air conditioning and refrigeration system into consideration can effectively reduce energy consumption while meeting passenger thermal comfort requirements;

[0070] 3. Innovatively apply reinforcement learning methods to automotive air conditioning and refrigeration control systems. The generalization performance of reinforcement learning enables the air conditioning and refrigeration system to dynamically and adaptively adjust to cope with various complex environments, and can also effectively reduce the workload of engineers. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] The present invention will be further described below with reference to the accompanying drawings and examples.

[0072] Figure 1 Schematic diagram of the air conditioning and refrigeration control system for the passenger compartment of a pure electric vehicle based on DDPG.

[0073] Figure 2 It is a training framework for the DDPG-based air conditioning and refrigeration control method of the passenger compartment of pure electric vehicles.

[0074] Figure 3 The figure is a flow chart of the air conditioning and refrigeration control method for the passenger compartment of a pure electric vehicle based on DDPG.

[0075] In the figure: 1 - expansion valve; 2 - front cooling module fan; 3 - condenser and receiver-drier assembly; 4 - compressor; 5 - blower; 6 - evaporator; 7 - air conditioning damper; 8 - HVAC; 9 - solar radiation intensity sensor; 10 - outside temperature sensor; 11 - inside temperature sensor; 12 - vehicle speed sensor; 13 - passenger compartment air conditioning and cooling control strategy; 14 - vehicle air conditioning controller;

[0076] 15-One-dimensional model of automobile air-conditioning and refrigeration system; 16-Passenger compartment thermal comfort prediction model; 17-Three-dimensional model of passenger compartment; 18-Human body thermal comfort model; 19-Current policy network; 20-Target policy network; 21-Current Q-value network; 22-Target Q-value network; 23-Experience recovery pool; 24-Action noise; 25-Action network; 26-Evaluation network; 27-Reinforcement learning training environment; 28-Passenger compartment air-conditioning and refrigeration control module; 29-Passenger compartment thermal flow field & thermal comfort module. DETAILED DESCRIPTION

[0077] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams that illustrate the basic structure of the present invention only in a schematic manner. Therefore, they only show components relevant to the present invention, and directions and references (e.g., up, down, left, right, etc.) may be used solely to facilitate the description of features in the drawings. The following detailed description is therefore not to be taken in a limiting sense, and the scope of the claimed subject matter is defined solely by the appended claims and their equivalents.

[0078] like Figure 1Figure 2 shows the hardware structure of a DDPG-based pure electric vehicle passenger compartment air conditioning and cooling control system. The system includes an expansion valve, a front-end cooling module fan, a condenser and receiver-drier assembly, a compressor, a blower, an evaporator, an air conditioning damper, a solar radiation intensity sensor, an outside temperature sensor, an inside temperature sensor, a vehicle speed sensor, and an automotive air conditioning controller. The blower, evaporator, and air conditioning damper constitute the HVAC (Heating, Ventilation, and Air Conditioning) system. The passenger compartment air conditioning and cooling control strategy is converted into code and burned into the automotive air conditioning controller. Data such as solar intensity, outside temperature, inside temperature, and vehicle speed collected by the solar radiation intensity sensor, outside temperature sensor, inside temperature sensor, and vehicle speed sensor are input into the automotive air conditioning controller. The automotive air conditioning controller executes the control strategy based on this data and outputs control commands to the compressor, blower, and air conditioning damper. This controls the compressor and blower speeds, as well as the damper opening in the HVAC system, thereby adjusting the operating state of the automotive air conditioning and cooling system to ensure good thermal comfort for passengers in various environmental conditions.

[0079] The expansion valve, front-end cooling module fan, condenser and receiver-drier assembly, compressor, evaporator, and other components comprise the front-end cooling system of an automotive air conditioner, responsible for cooling the passenger compartment. The compressor compresses low-pressure refrigerant into high-pressure gas, raising its temperature. The condenser cools this high-temperature, high-pressure gas through the radiator, converting it into a high-pressure liquid. The expansion valve controls the refrigerant flow and pressure, expanding the high-pressure liquid into a low-pressure liquid, reducing its temperature. The evaporator blows the low-pressure, low-temperature refrigerant through the fan, absorbing heat from the vehicle interior and converting it into a low-temperature, low-pressure gas. The front-end cooling module fan blows interior air through the condenser, exchanging heat with the refrigerant. The receiver-drier primarily filters and dries the refrigerant, preventing air, moisture, and impurities from entering the air conditioning system.

[0080] like Figure 2 and Figure 3 As shown in FIG, a DDPG-based pure electric vehicle passenger compartment air conditioning and refrigeration control method of the present invention includes the following steps:

[0081] S1: Build a passenger cabin human thermal comfort prediction model

[0082] S1.1: Construct a three-dimensional passenger cabin model and a human thermal comfort evaluation model in the three-dimensional design software. The three-dimensional passenger cabin model and the human thermal comfort evaluation model constitute the passenger cabin thermal flow field & thermal comfort module.

[0083] The 3D passenger compartment model, or 3D passenger compartment simulation model, is obtained by extracting the passenger compartment with the air conditioning system from the vehicle digital model. This model is simplified and meshed before being imported into the 3D simulation software. The 3D digital model is inspected for integrity within the 3D design software, and the relevant passenger compartment components are extracted. The 3D digital model is simplified and meshed, and the air conditioning dummy model is added to generate a volume mesh. Regions are set and named within the volume mesh. A physical model is created within the volume mesh, along with the physical model and boundary conditions. Multiple thermal comfort monitoring points are also set on the dummy model.

[0084] The human thermal comfort evaluation model, also known as the passenger cabin thermal comfort evaluation model, simulates the human body's thermophysiological regulation mechanisms in different temperature environments. It calculates the passenger thermal comfort evaluation value by inputting key passenger physical characteristics and the air temperature, air velocity, mean radiant temperature, and relative humidity near thermal comfort monitoring points obtained through passenger cabin CFD simulation. There are 14 thermal comfort monitoring points: the passenger's head, torso, left forearm, left upper arm, left hand, right forearm, right upper arm, right hand, left thigh, left calf, left foot, right thigh, right calf, and right foot. By inputting the air velocity, temperature, mean radiant temperature, and relative humidity of each surface area, the human thermal comfort evaluation result is obtained. The human thermal comfort evaluation result ranges from [-3, 3], with negative numbers indicating excessive cooling and positive numbers indicating excessive heating. Values ​​closer to 0 indicate better comfort.

[0085] For details about the passenger cabin 3D model and the human thermal comfort evaluation model, please refer to the passenger cabin 3D simulation model and the passenger cabin thermal comfort evaluation model in the invention application with publication number CN 114757116 A, which will not be repeated here.

[0086] S1.2: Feature Parameter Setting: Manual labeling of some important feature parameters is required during deep learning neural network training. This can effectively increase model prediction accuracy with limited datasets. Based on the requirements of deep learning neural network training, the feature parameters of the passenger compartment thermal flow field and thermal comfort module are set. These feature parameters include exterior vehicle temperature, solar radiation intensity, air temperature after the evaporator, air velocity after the evaporator, air temperature at each air conditioning outlet, air velocity at each air conditioning outlet, air temperature at each human body surface, air velocity at each human body surface, average radiation temperature at each human body surface, relative humidity at each human body surface, and human thermal comfort evaluation results. The "human body parts" refer to the 14 thermal comfort monitoring points defined on the human body.

[0087] S1.3: Dataset extraction: Through the joint simulation of the passenger cabin three-dimensional model and the human thermal comfort evaluation model, the numerical values ​​corresponding to the characteristic parameters described in the simulation results are extracted as the data set for deep learning neural network training. The data set is preprocessed and divided into a training set and a validation set; wherein,

[0088] Data collection: Using a three-dimensional passenger cabin model and a human thermal comfort evaluation model, we simulate and collect data on interior and exterior temperatures, solar radiation intensity, air humidity, air velocity and temperature after the evaporator, and passenger thermal comfort evaluation results.

[0089] The dataset is preprocessed by denoising the initial sample dataset, eliminating outliers, and interpolating missing values. The min-max standardization method is then used for normalization. The specific formula is as follows:

[0090]

[0091] In the formula, y is the normalized data; x is the original data; x min is the minimum value in the original data set; x max is the maximum value in the original data set;

[0092] Dataset division: The dataset is divided into a training set and a validation set in a ratio of 8:2. The training set is used to train the model, and the validation set is used to adjust the model parameters and verify the prediction effect of the trained model. The validation set does not participate in model training.

[0093] By changing the input conditions, different human thermal comfort evaluation results will be output. During the simulation process, the temperature inside the vehicle will change as the simulation runs. The outside temperature, solar radiation intensity, and the air temperature and speed after the evaporator can all be set as changing curve inputs. The output human thermal comfort and the above-set characteristic parameters are also changing curves. That is, a single simulation can extract multiple sets of data sets for training neural networks in deep learning.

[0094] S1.4: Neural network training, specifically including the following steps,

[0095] S1.4.1: Build a deep learning network based on the defined model structure, including an input layer, hidden layer, and output layer, and initialize the corresponding weights and thresholds. The input layer includes six neurons, corresponding to the interior temperature, exterior temperature, solar radiation intensity, air humidity, and air velocity and temperature after the evaporator. The hidden layer includes four neurons, using the ReLU activation function to extract features from the input data. The output layer includes one neuron, which outputs the passenger thermal comfort evaluation results.

[0096] S1.4.2: Model preprocessing: Initialize the weights between the input layer and the hidden layer, the weights between the hidden layer and the output layer, and the thresholds of the hidden layer and the output layer. Use the Bayesian Regularization algorithm for neural network training, the Adam optimizer, the learning rate set to 0.001, the number of epochs to 200, the batch size to 32, and the input time step to 2. Reasonably set the hyperparameters of neural network training: optimizer, learning rate, number of iterations, time step, batch size, and number of neurons to improve the accuracy and effectiveness of model predictions.

[0097] S1.4.3: Training model: Use the data from the training set to train the deep learning model, and use the backpropagation algorithm to update the weights and thresholds. The output error, that is, the difference between the expected output and the actual output, is calculated through the original path, reversed through the hidden layer to the input layer. During the backpropagation process, the error is distributed to each unit in each layer, and the error signal of each unit in each layer is obtained, which is used as the basis for correcting the weights of each unit. This calculation process is completed using the gradient descent method. After continuously adjusting the weights and thresholds of neurons in each layer, the error signal is reduced to a minimum.

[0098] S1.4.4: Model evaluation: The deep learning model is evaluated using yearly-based and station-based validation methods. The evaluation metrics include MSE, RMSE (Root Mean Squared Error), MAE (Meanabsolute Error), and R-Squared coefficient of determination. The formulas are as follows:

[0099]

[0100]

[0101]

[0102]

[0103] in, is the predicted value, y i is the true value, is the mean value, m is the number of samples;

[0104] When R 2 The larger the value and the smaller the values ​​of other indicators, the better the prediction effect of the model.

[0105] The passenger compartment human thermal comfort prediction model can directly predict the human thermal comfort evaluation results by inputting the interior temperature, exterior temperature, solar radiation intensity, air humidity, air velocity and temperature after the evaporator. To avoid data overfitting during training, the cost function of the neural network is trained using Bayesian regularization to minimize the training error. The cost function is:

[0106]

[0107] Where α1 and α2 are Bayesian hyperparameters that specify the direction the learning process seeks, i.e., minimizing the error or weight; n is the number of training samples; Y i is the actual value of the i-th; Y i ′ is the i-th predicted value of the neural network; m is the number of weights in the neural network, w j is the jth weight.

[0108] S2: Build a passenger compartment air conditioning and cooling control strategy

[0109] S2.1: Build the passenger compartment air conditioning and refrigeration control module

[0110] A reinforcement learning model is defined based on the automobile air-conditioning and refrigeration system. The state s, action a, and reward r of the MDP process in the reinforcement learning model are determined. The passenger compartment air-conditioning and refrigeration control module is determined based on the reinforcement learning model. The passenger compartment air-conditioning and refrigeration control module includes an action network, an evaluation network, and an experience recovery pool.

[0111] Defining a reinforcement learning model involves the following steps:

[0112] (1) Define the state s of the MDP process

[0113] Obtain the state information of the vehicle air conditioning and refrigeration system and define the state s of the MDP (Markov Decision Processes) process as: s = [s1, s2, s3, s4, s5, s6, s7], where s1 is the ambient temperature outside the vehicle, s2 is the solar radiation intensity, s3 is the temperature inside the vehicle, s4 is the vehicle speed, s5 is the air humidity, s6 is the passenger thermal comfort evaluation result, and s7 is the energy consumption per minute of the vehicle air conditioning system. S1, s2, s3, s4, and s5 are random inputs with a limited range, providing a learning environment under different working conditions.

[0114] (2) Define the action a of the MDP process

[0115] According to the output control instructions of the automobile air conditioning refrigeration system, the action a of the MDP process is defined as: a = [a1, a2, a3], where a1 is the blower speed, a2 is the compressor speed, and a3 is the damper opening;

[0116] (3) Define the reward r of the MDP process

[0117] Based on the main performance evaluation indicators of the automotive air conditioning and refrigeration system, the reward r of the MDP process is defined as: r = -E - λΔT, where E is the total energy consumption per minute of the automotive air conditioning and refrigeration system components when the passenger compartment heat load is balanced, and E takes a negative value; λ is the thermal comfort penalty function coefficient; and ΔT is the difference between the current thermal comfort evaluation result and the target thermal comfort evaluation result.

[0118] Considering that the main performance evaluation indicators of automobile air conditioning and refrigeration are divided into two parts: ① Human thermal comfort evaluation λΔT: the difference between the thermal comfort evaluation result of the passengers in the car and the optimal thermal comfort value (0); ② Energy consumption E: the total energy consumption per minute of the components of the automobile air conditioning and refrigeration system when the passenger compartment heat load is balanced, mainly including the energy consumption of the compressor, blower and front-end cooling module fan. Therefore, the reward part is set to r = -E-λΔT. Since energy consumption should be minimized as much as possible, E takes a negative value. At the same time, in order to ensure the stability and effectiveness of the car's control of thermal comfort in the car, a penalty function related to the difference between the thermal comfort evaluation result of the passengers in the car and the optimal thermal comfort value (0) is added. λ is the coefficient of the thermal comfort penalty function.

[0119] In this embodiment, the total input of the reinforcement learning training environment is the solar radiation intensity, the temperature inside and outside the vehicle, the vehicle speed, and the reinforcement learning execution action a = [a1, a2, a3]. The output is the reward r = -E-λΔT and the state s = [s1, s2, s3, s4, s5, s6, s7], forming a complete reinforcement learning training cycle.

[0120] S2.2: Building the environment required for reinforcement learning training

[0121] The environment required for reinforcement learning training includes a one-dimensional model of the automobile air-conditioning and refrigeration system and a passenger compartment human thermal comfort prediction model obtained in step S1, wherein the one-dimensional model of the automobile air-conditioning and refrigeration system is used to simulate the operation of components in the automobile air-conditioning and refrigeration system according to the interior temperature, exterior temperature, solar radiation intensity, air humidity, and control instructions of the automobile air-conditioning and refrigeration system, and output air velocity and temperature data after the evaporator, interior temperature data, and energy consumption data of the air-conditioning system; the passenger compartment human thermal comfort prediction model is used to predict the human thermal comfort evaluation result based on the interior temperature, exterior temperature, solar radiation intensity, air humidity, air velocity and temperature data after the evaporator, and feed the human thermal comfort evaluation result back to the passenger compartment thermal comfort control module; wherein, the construction of the one-dimensional model of the automobile air-conditioning and refrigeration system is referred to CN114757116 A.

[0122] S2.3: Reinforcement learning training for the cabin air conditioning and cooling control module

[0123] In the reinforcement learning training environment of step S3, the passenger compartment air conditioning and refrigeration control module constructed in step S2 is trained using a reinforcement learning control structure network based on the DDPG algorithm. During the training process, sample data is collected, and the passenger compartment air conditioning and refrigeration control module is updated and optimized based on the sample data. Once the passenger compartment air conditioning and refrigeration control module reaches a convergence state, the training is completed. At this point, the control strategy of the passenger compartment air conditioning and refrigeration control module is the target strategy - the passenger compartment air conditioning and refrigeration control strategy of the automobile air conditioning and refrigeration system.

[0124] Among them, the reinforcement learning control structure network based on the DDPG algorithm includes an action network and an evaluation network, the action network includes a current policy network and a target policy network, and the evaluation network includes a current Q-value network and a target Q-value network, wherein the input information of the previous policy network is state s, and the output information is action a; the input and output of the target policy network are the same as the current policy network, and the current policy network parameters are copied periodically; the input information of the current Q-value network is state s and action a, and the output information is value Q; the input and output of the target Q-value network are the same as the current Q-value network, and the current Q-value network parameters are copied periodically.

[0125] The training process in step S2.3 is:

[0126] (1) Update of action network

[0127] The current policy network interacts with the reinforcement learning training environment. The state s is input to the current policy network to obtain action a. Action a is applied to the reinforcement learning training environment, and the reinforcement learning training environment returns the state s' and reward r at the next moment. The sample data (s, a, r, s') at this time is collected and placed in the experience recycling pool. The target policy network is responsible for selecting the optimal next action a|s' based on the next state s' sampled in the experience recycling pool. The network structure of the target policy network is the same as that of the current policy network. The parameters of the target policy network are regularly copied from the parameters of the current policy network. When the current policy network applies action a to the reinforcement learning training environment, random action noise needs to be added to avoid excessive errors in training.

[0128] (2) Update of evaluation network

[0129] The current Q-value network is responsible for iteratively updating the value network parameter ω. Input S and a in (s, a, r, s') into the current Q-value network, calculate the value Q(s, a, ω) of the current Q-value network, input s' in (s, a, r, s') into the target policy network, obtain action a', and input s' and a' together into the target Q-value network to calculate the target Q value y i=r+YQ'(s',a',ω'), where Y is the discount factor; the target Q-value network is responsible for calculating the Q'(s',a',ω') part of the Q-value. The network structure is the same as the current Q-value network, and the network parameters are regularly copied from the current Q-value network. The following formula is used to calculate the loss function Loss of the current Q-value network:

[0130]

[0131] Meaning of the parameters: y i is the target Q value; i is the number of cycles; a' is the action output by the target policy network; ω' is the parameter of the value network; Q' is the Q value calculated by the target Q value network.

[0132] Among them, the role of the loss function is to describe the gap between the model's predicted value and the true value, and guide the model to move towards convergence during the training process.

[0133] The reinforcement learning model is trained. The reinforcement learning interacts with the virtual environment to collect various sampling data, and the action network and evaluation network are continuously updated. The experience data is stored in the experience recovery pool. After the model converges, the comfort-based automobile air-conditioning and refrigeration control strategy is developed after the training is completed. According to different indoor and outdoor temperatures, solar radiation intensity and driving speed information, the appropriate compressor speed, blower speed and damper opening can be directly output to achieve two-way optimization of passenger thermal comfort and air-conditioning and refrigeration system energy consumption.

[0134] S3: Application of control strategies

[0135] The trained passenger compartment air conditioning and refrigeration control module's control strategy is converted into code and burned into the vehicle's air conditioning controller, which serves as the actual vehicle's air conditioning and refrigeration control system to control and adjust the thermal comfort of the passenger compartment.

[0136] The present invention's DDPG-based passenger compartment air conditioning and refrigeration control method for pure electric vehicles features a training framework comprised of three components: a passenger compartment air conditioning and refrigeration control module based on the DDPG algorithm, a reinforcement learning training environment, and a passenger compartment thermal flow and thermal comfort module. The reinforcement learning training environment includes a one-dimensional model of the vehicle's air conditioning system and a passenger compartment thermal comfort prediction model; the passenger compartment thermal flow and thermal comfort module includes a three-dimensional passenger compartment model and a human thermal comfort model; and the DDPG-based passenger compartment air conditioning and refrigeration control module includes an action network, an evaluation network, and an experience recovery pool. The passenger compartment thermal flow and thermal comfort module is transformed into a passenger compartment thermal comfort prediction model within the reinforcement learning training environment using deep learning. The DDPG-based passenger compartment air conditioning and refrigeration control module continuously interacts with the reinforcement learning training environment to achieve effective training.

[0137] With the above-described preferred embodiments of the present invention as inspiration, and with reference to the above description, relevant personnel may make various changes and modifications without departing from the scope of the present invention. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.

Claims

1. A DDPG-based air conditioning and refrigeration control method for a passenger compartment of a pure electric vehicle, characterized by: The following steps are involved: S1: Build a passenger cabin human thermal comfort prediction model S1.1: Construct a 3D passenger cabin model and a human thermal comfort evaluation model in 3D design software. The 3D passenger cabin model and the human thermal comfort evaluation model constitute the passenger cabin thermal flow field and thermal comfort module. S1.2: Set the characteristic parameters of the passenger cabin thermal flow field and thermal comfort module based on the requirements of deep learning neural network training; S1.3: Through a joint simulation of the passenger cabin 3D model and the human thermal comfort evaluation model, extract the numerical values ​​corresponding to the characteristic parameters described in the simulation results as a dataset for deep learning neural network training. Preprocess the dataset and divide it into a training set and a validation set. S1.4: Neural network training: Build a deep learning network based on the defined model structure, including input, hidden, and output layers; initialize the corresponding weights and thresholds; and set the neural network training hyperparameters, including the optimizer, learning rate, number of iterations, time step, and batch size. Train the deep learning model using the training set data, and update the weights and thresholds using the backpropagation algorithm. Develop a passenger compartment human thermal comfort prediction model and evaluate its prediction performance using a validation set. S2: Build a passenger compartment air conditioning and cooling control strategy S2.1: Build the passenger compartment air conditioning and refrigeration control module Define a reinforcement learning model based on the automotive air conditioning and refrigeration system, determine the state s, action a, and reward r of the MDP process in the reinforcement learning model, and determine the passenger compartment air conditioning and refrigeration control module based on the reinforcement learning model; S2.2: Building the environment required for reinforcement learning training The environment required for reinforcement learning training includes a one-dimensional model of the automobile air-conditioning and refrigeration system and a passenger compartment human thermal comfort prediction model obtained in step S1, wherein the one-dimensional model of the automobile air-conditioning and refrigeration system is used to simulate the operation of components in the automobile air-conditioning and refrigeration system based on the interior temperature, exterior temperature, solar radiation intensity, air humidity, vehicle speed, and control instructions of the automobile air-conditioning and refrigeration system, and output air velocity and temperature data after the evaporator, interior temperature data, and energy consumption data of the air-conditioning system; The passenger compartment human thermal comfort prediction model is used to predict the human thermal comfort evaluation results based on the vehicle interior temperature, vehicle exterior temperature, solar radiation intensity, air humidity, air velocity and temperature after the evaporator, and feed the human thermal comfort evaluation results back to the passenger compartment thermal comfort control module; S2.3: Reinforcement learning training for the cabin air conditioning and cooling control module In the reinforcement learning training environment of step S2.2, the passenger compartment air conditioning and cooling control module constructed in step S2.1 is trained using a reinforcement learning control structure network based on the DDPG algorithm. During the training process, sample data is collected and the passenger compartment air conditioning and cooling control module is updated and optimized based on the sample data. Once the passenger compartment air conditioning and cooling control module reaches a convergence state, the training is completed. At this point, the control strategy of the passenger compartment air conditioning and cooling control module becomes the target strategy—the passenger compartment air conditioning and cooling control strategy of the vehicle air conditioning and cooling system. S3: Application of control strategies The trained passenger compartment air conditioning and refrigeration control module's control strategy is converted into code and burned into the vehicle's air conditioning controller, which serves as the actual vehicle's air conditioning and refrigeration control system to control and adjust the thermal comfort of the passenger compartment.

2. The DDPG-based pure electric vehicle passenger compartment air conditioning and refrigeration control method according to claim 1, characterized in that: The characteristic parameters described in step S1.2 include the outdoor temperature of the vehicle, the solar radiation intensity, the air temperature behind the evaporator, the air flow velocity behind the evaporator, the air temperature of each air-conditioning outlet, the air flow velocity of each air-conditioning outlet, the air temperature of each part of the human body surface, the air flow velocity of each part of the human body surface, the average radiation temperature of each part of the human body surface, the relative humidity of each part of the human body surface, and the thermal comfort evaluation results of the human body.

3. The DDPG-based pure electric vehicle passenger compartment air conditioning and refrigeration control method according to claim 1, characterized in that: Step S1.3 specifically also includes the following steps: Data collection: Using a three-dimensional passenger cabin model and a human thermal comfort evaluation model, we simulated and collected data on interior and exterior temperatures, solar radiation intensity, air humidity, air velocity and temperature after the evaporator, and passenger thermal comfort evaluation results. The dataset is preprocessed by denoising the initial sample dataset, eliminating outliers, and interpolating missing values. The min-max standardization method is then used for normalization. The specific formula is as follows: In the formula, y is the normalized data; x is the original data; x min is the minimum value in the original data set; x max is the maximum value in the original data set; Dataset division: The dataset is divided into a training set and a validation set in a ratio of 8:

2. The training set is used to train the model, and the validation set is used to adjust the model parameters.

4. The DDPG-based pure electric vehicle passenger compartment air conditioning and refrigeration control method according to claim 1, characterized in that: The neural network training in step S1.4 specifically includes the following steps: S1.4.1: Build a deep learning network: Based on the defined model structure, build a deep learning network consisting of an input layer, a hidden layer, and an output layer. The input layer consists of six neurons, corresponding to the interior temperature, exterior temperature, solar radiation intensity, air humidity, and air velocity and temperature after the evaporator. The hidden layer consists of four neurons, using the ReLU activation function to extract features from the input data. The output layer consists of one neuron, which outputs the passenger thermal comfort evaluation results. S1.4.2: Model preprocessing: Initialize the weights between the input layer and the hidden layer, the weights between the hidden layer and the output layer, and the thresholds of the hidden layer and the output layer. The neural network training algorithm uses the Bayesian regularization algorithm, the optimizer uses Adam, the learning rate is set to 0.001, the number of epochs is 200, the batch size is 32, and the input time step is 2. S1.4.3: Training Model: Use the data from the training set to train the deep learning model. Use the backpropagation algorithm to update the weights and thresholds. The output error (i.e., the difference between the expected output and the actual output) is calculated by backpropagating through the original path, back through the hidden layer, and finally to the input layer. During the backpropagation process, the error is distributed to each unit in each layer, and the error signal of each unit in each layer is obtained. This error signal is used as the basis for correcting the weights of each unit. S1.4.4: Model evaluation: The deep learning model is evaluated using yearly-based and station-based validation methods. The evaluation metrics include MSE, RMSE root mean square error, MAE mean absolute error, and R-Squared coefficient of determination. The formulas are as follows: in, is the predicted value, y i is the true value, is the mean value, m is the number of samples; When R 2 The larger the value and the smaller the values ​​of other indicators, the better the prediction effect of the model.

5. The DDPG-based pure electric vehicle passenger compartment air conditioning and refrigeration control method according to claim 4, characterized in that: In step S1.4, when training the neural network, the cost function of the neural network is trained using Bayesian regularization to minimize the training error, where the cost function is: Where α1 and α2 are Bayesian hyperparameters that specify the direction the learning process seeks, i.e., minimizing the error or weight; n is the number of training samples; Y i is the actual value of the i-th; Y′ i is the i-th predicted value of the neural network; m is the number of weights in the neural network, w j is the jth weight.

6. The DDPG-based pure electric vehicle passenger compartment air conditioning and refrigeration control method according to claim 1, characterized in that: Defining the reinforcement learning model in step S2.1 specifically includes the following steps: (1) Define the state s of the MDP process Obtain the state information of the automobile air conditioning and refrigeration system and define the state s of the MDP process as: s = [s1, s2, s3, s4, s5, s6, s7], where s1 is the ambient temperature outside the vehicle, s2 is the solar radiation intensity, s3 is the temperature inside the vehicle, s4 is the vehicle speed, s5 is the air humidity, s6 is the passenger thermal comfort evaluation result, and s7 is the energy consumption per minute of the vehicle air conditioning system. s1, s2, s3, s4, and s5 are random inputs with a limited range, providing a learning environment under different working conditions. (2) Define the action a of the MDP process According to the output control instructions of the automobile air conditioning refrigeration system, the action a of the MDP process is defined as: a = [a1, a2, a3], where a1 is the blower speed, a2 is the compressor speed, and a3 is the damper opening; (3) Define the reward r of the MDP process Based on the main performance evaluation indicators of the automotive air conditioning and refrigeration system, the reward r of the MDP process is defined as: r = -E - λΔT, where E is the total energy consumption per minute of the automotive air conditioning and refrigeration system components when the passenger compartment heat load is balanced, and E takes a negative value; λ is the thermal comfort penalty function coefficient; and ΔT is the difference between the current thermal comfort evaluation result and the target thermal comfort evaluation result.

7. The DDPG-based pure electric vehicle passenger compartment air conditioning and refrigeration control method according to claim 1, characterized in that: The reinforcement learning control structure network of the DDPG algorithm described in step S2.3 includes an action network and an evaluation network, wherein the action network includes a current policy network and a target policy network, and the evaluation network includes a current Q-value network and a target Q-value network, wherein the input information of the current policy network is state s, and the output information is action a; the input and output of the target policy network are the same as the current policy network, and the parameters of the current policy network are copied periodically; the input information of the current Q-value network is state s and action a, and the output information is value Q; the input and output of the target Q-value network are the same as the current Q-value network, and the parameters of the current Q-value network are copied periodically.

8. The DDPG-based pure electric vehicle passenger compartment air conditioning and refrigeration control method according to claim 7, characterized in that: The training process in step S2.3 is: (1) Update of action network The current policy network interacts with the reinforcement learning training environment. The input state s is fed into the current policy network to obtain action a. Action a is applied to the reinforcement learning training environment, which returns the state s' and reward r at the next moment. The sample data (s, a, r, s') at this moment is collected and placed in the experience recycling pool. The target policy network is responsible for selecting the optimal next action a|s' based on the next state s' sampled from the experience recycling pool. The network structure of the target policy network is the same as that of the current policy network, and the parameters of the target policy network are regularly copied from the parameters of the current policy network. (2) Update of evaluation network The current Q-value network is responsible for iteratively updating the value network parameter ω. Input S and a in (s, a, r, s') into the current Q-value network, calculate the value Q(s, a, ω) of the current Q-value network, input s' in (s, a, r, s') into the target policy network, obtain action a', and input s' and a' together into the target Q-value network to calculate the target Q value y i =r+ΥQ'(s',a',ω'), where Υ is the discount factor; the target Q-value network is responsible for calculating the Q'(s',a',ω') part of the Q-value. The network structure is the same as the current Q-value network, and the network parameters are regularly copied from the current Q-value network. The following formula is used to calculate the loss function Loss of the current Q-value network: Meaning of the parameters: y i is the target Q value; i is the number of cycles; a' is the action output by the target policy network; ω' is the parameter of the value network; Q' is the Q value calculated by the target Q value network.

9. The DDPG-based pure electric vehicle passenger compartment air conditioning and refrigeration control method according to claim 7, characterized in that: When the current policy network applies action a to the reinforcement learning training environment, random action noise needs to be added to avoid excessive errors in training.

Citation Information

Patent Citations

  • Control method and system of air conditioner of electric vehicle

    CN103342091A

  • System and method for evaluating and optimizing thermal comfort of passenger compartment with heat source equivalence

    CN114757116A

  • Electric automobile air conditioner control method

    CN115489259A

  • Electric automobile air conditioner control system

    CN209441142U

  • Passenger compartment temperature control method based on artificial neural network algorithm

    CN109435630A