A control method and device of a heat pump system and a vehicle
By configuring independent intelligent agents in the heat pump system of electric vehicles for distributed collaborative control, the problem of the disconnect between temperature control accuracy and energy consumption optimization in traditional control methods is solved, and efficient, safe and comfortable temperature regulation of the heat pump system is achieved.
Patent Information
- Application Number
- CN202610853690.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-07-24
AI Technical Summary
Traditional heat pump systems in electric vehicles cannot simultaneously meet the requirements of passenger cabin temperature control and comfort, precise regulation of power battery temperature, and energy consumption optimization, leading to reduced vehicle range and safety hazards.
By employing a distributed cooperative control method, independent intelligent agents based on prediction bias are configured to acquire global operating data of the heat pump system, construct input feature vectors, and realize independent decision-making and cooperative optimization of each actuator group, thereby reducing energy consumption and improving temperature control accuracy.
It achieves coordinated optimization of temperature control accuracy and system energy consumption of multiple actuator groups without the need for real-time communication, reduces control complexity, meets the real-time and reliability requirements of the vehicle controller, and ensures passenger cabin comfort and power battery safety.
Smart Images

Figure CN122443152A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of heat pump system control technology, and specifically to a control method, device, and vehicle for a heat pump system. Background Technology
[0002] With increasing energy shortages and stringent environmental requirements, electric vehicles (EVs) are gaining widespread adoption due to their energy-saving and environmentally friendly advantages. Unlike traditional gasoline vehicles, EVs require a thermal management system that not only ensures passenger cabin comfort but also precisely controls the temperature of the battery. The optimal operating temperature for lithium batteries is 20°C to 40°C. Excessive temperature can lead to thermal runaway and other safety issues, while excessively low temperatures reduce battery charging and discharging efficiency, impacting overall vehicle performance. Currently, EVs primarily rely on heat pump systems for thermal management and temperature regulation. However, traditional heat pump control methods are relatively rigid, resulting in poor temperature control accuracy, high energy consumption, and even significant range reduction when the air conditioning is running. These systems are ill-suited to the complex thermal management requirements of battery safety, cabin comfort, and low-energy operation under demanding conditions. Summary of the Invention
[0003] The purpose of this invention is to provide a control method, device, and vehicle for a heat pump system. By configuring independent intelligent agents based on prediction bias for each temperature regulation dimension, the distributed collaborative control of the heat pump system can be achieved, thereby reducing control complexity, improving response accuracy, and significantly saving energy consumption.
[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows: Firstly, a control method for a heat pump system is provided. The heat pump system includes multiple actuator groups, each actuator group being used to regulate the temperature of a controlled area it affects. The method includes: acquiring current global operating data of the heat pump system; the current global operating data includes: the current control state of each actuator in the heat pump system, the current actual temperature and set temperature of the area controlled by the heat pump system, and the operating energy consumption of the heat pump system; extracting local operating data related to each actuator group from the current global operating data; the local operating data includes at least: the current temperature error of the controlled area affected by the corresponding actuator group, and the current control state of the actuator group; based on the actuator group's control state, the method further analyzes the temperature of the controlled area and regulates the temperature of the controlled area. Based on the local operating data related to the actuator group and the operating energy consumption of the heat pump system, the input feature vector of the agent associated with the actuator group is determined; wherein, each actuator group corresponds to one agent; the agent is used to: determine the control action to be performed by the corresponding actuator group with the goal of reducing the current temperature error of the controlled area and the operating energy consumption of the heat pump system; the input feature vector is used to characterize: the current operating state of the local area corresponding to the agent and the operating energy consumption of the heat pump system; the corresponding input feature vector is input to the agent corresponding to the actuator group, and the target control action of the actuator group is output; based on the target control action of each actuator group, the operation of the heat pump system is controlled.
[0005] Based on the aforementioned technical means, the heat pump system control method provided in this application acquires current global operating data, including the current control state of each actuator, the current actual temperature and set temperature of the controlled area, and the system operating energy consumption. It then extracts local operating data related to each actuator group from this data, constructs an input feature vector based on the local operating data and the system operating energy consumption, and allows each agent to simultaneously output target control actions based on the input feature vector, with the goal of reducing the current temperature error of the controlled area and the system operating energy consumption, thereby controlling the operation of the heat pump system. This method directly incorporates system operating energy consumption into the decision input of the agents, enabling each agent to simultaneously perceive temperature deviation and energy consumption level when outputting control actions. This achieves synergistic optimization between precise temperature control and system energy saving, avoiding the problem of temperature control and energy consumption optimization being isolated from each other in traditional methods. At the same time, each agent only needs to make decisions and output control actions and state predictions based on local operating data related to its corresponding actuator group and the overall system energy consumption. Distributed collaborative control of multiple actuator groups is achieved without the need for real-time communication between agents, effectively reducing the input dimension and decision complexity of each agent, meeting the real-time and reliability requirements of the vehicle controller, and thus optimizing the overall energy consumption of the heat pump system while ensuring passenger cabin comfort and power battery thermal safety.
[0006] Furthermore, the agent is also used to: estimate the predicted operating data after the actuator group performs the control action based on the determined control action to be performed by the actuator group; the estimated operating data is used to compare with the actual local operating data at the next time step to calculate the state prediction deviation in the input feature vector at the next time step; based on the local operating data related to the actuator group and the operating energy consumption of the heat pump system, determine the input feature vector of the agent associated with the actuator group, and also includes: based on the local operating data related to the actuator group, the operating energy consumption of the heat pump system, and the estimated operating data output by the agent associated with the actuator group at the previous time step, determine the input feature vector of the agent associated with the actuator group, so that the agent outputs the target control action of the actuator group and the estimated operating data; the input feature vector also includes: the state prediction deviation of the agent for the local area at the previous time step.
[0007] Based on the aforementioned technical means, by endowing the agent with dual "decision-prediction" output capabilities, it outputs predicted data on the future operating state after the execution of the target control action, and incorporates the predicted operating data from the previous moment into the input feature vector of the current moment, thus constructing a "decision-prediction-feedback" closed-loop mechanism. This mechanism enables the agent to perceive the accuracy of its own prediction of the dynamic characteristics of the heat pump system, effectively cope with the large inertia and hysteresis characteristics of the heat pump system, reduce over-adjustment and temperature oscillation caused by system response delay, and achieve synergistic optimization of temperature control accuracy and system energy consumption by multiple actuator groups without the need for real-time communication between agents, thereby improving the foresight and robustness of the control strategy.
[0008] Furthermore, the multiple actuator groups include: a compressor actuator group for driving refrigerant circulation and adjusting the heating or cooling capacity of the heat pump system; an electronic expansion valve actuator group for adjusting the refrigerant throttling opening and controlling the refrigerant flow on the evaporator or condenser side; a fan actuator group for adjusting the air-side heat exchange intensity and affecting the temperature response of the passenger compartment; and a water pump actuator group for driving coolant circulation and adjusting the temperature of the power battery or drive motor.
[0009] Based on the aforementioned technical means, by specifically dividing the actuator group into compressor actuator group, electronic expansion valve actuator group, fan actuator group, and water pump actuator group, temperature can be coordinated and regulated from four different dimensions: refrigerant circulation drive, refrigerant throttling flow distribution, air-side heat exchange intensity adjustment, and coolant circulation drive. This decouples the complex thermal management task of the heat pump system into multiple functionally defined and mutually cooperating sub-control tasks, making the control objectives of each intelligent agent clearer and improving the pertinence of the control strategy and the coordination efficiency between the actuator groups.
[0010] Furthermore, when the actuator group is a compressor actuator group, the current temperature error includes the passenger compartment temperature error and the battery temperature error, and the current control state quantity includes the compressor current speed; when the actuator group is an electronic expansion valve actuator group, the current temperature error includes the suction superheat error and the evaporator outlet temperature error; when the actuator group is a fan actuator group, the current temperature error includes the condenser heat exchange state and the passenger compartment temperature error; when the actuator group is a water pump actuator group, the current temperature error includes the battery temperature error and the coolant temperature error.
[0011] Based on the aforementioned technical means, by targeting the control functions and physical characteristics of different actuator groups, specific state variables directly related to the control objectives of the actuator group are extracted from the global operating data as local operating data. This enables the input characteristics of each agent to be highly matched with its control task, effectively filtering out redundant information irrelevant to the current control decision, further reducing the input dimension, improving the targeting and accuracy of agent decisions, thereby accelerating inference speed and enhancing the robustness of the control strategy.
[0012] Furthermore, multiple expert sample data are acquired; the expert sample data includes the system state vector and the corresponding expert control actions; the behavior cloning pre-training is performed on multiple agents corresponding to multiple actuator groups using the multiple expert sample data, so as to minimize the deviation between the actions output by each agent and the corresponding expert control actions.
[0013] Based on the above technical means, by acquiring expert sample data consisting of system state vectors and expert control actions, and using it to perform behavioral cloning pre-training on multiple agents, each agent can learn the safe and stable control laws of the preset controller under various working conditions before formally entering the joint fine-tuning of reinforcement learning, and obtain a set of feasible basic strategies. This effectively avoids the problems of low training efficiency and unsafe initial actions caused by random exploration from scratch, significantly shortens the training time and improves training stability.
[0014] Furthermore, multiple online interaction experience sample data are acquired; multiple agents are jointly trained using multiple online interaction experience sample data; during the joint training process, multiple agents are trained based on the same global reward signal, which is used to reflect temperature control error, heat pump system energy consumption and safety constraints.
[0015] Based on the aforementioned technical means, by acquiring online interactive experience sample data and using it to jointly train multiple agents, and with each agent optimizing based on the same global reward signal, each agent not only focuses on local control effects when optimizing its own strategy, but also needs to consider the impact of the actions of other actuator groups on the overall system performance. This promotes effective cooperation among multiple agents, avoids overall performance degradation caused by each agent acting independently, and ultimately learns a cooperative control strategy that balances temperature control accuracy, system energy consumption, and safety.
[0016] Furthermore, during joint training, the training samples include expert sample data and online interaction experience sample data. The expert sample data is used to minimize the deviation between the actions output by each agent and the corresponding expert control actions. As the number of training rounds increases, the proportion of expert sample data gradually decreases, while the proportion of online interaction experience sample data gradually increases.
[0017] Based on the aforementioned technical means, by dynamically adjusting the sampling ratio of expert sample data and online interactive experience sample data during the joint training process, the training initially focuses on expert sample data to ensure the safety and stability of policy updates, while the later training focuses on online interactive experience sample data to promote autonomous exploration and policy optimization. This achieves a smooth transition from "imitating experts" to "autonomous optimization" in stages. It inherits the engineering reliability and interpretability of the preset controller, while giving full play to the global optimization capability of reinforcement learning, effectively avoiding problems such as policy oscillation or convergence difficulties during the training process.
[0018] Furthermore, the global reward signal includes at least a temperature control reward item, an energy consumption penalty item, and a safety constraint penalty item; the temperature control reward item is used to reflect the deviation between the passenger compartment temperature and the set value, as well as the deviation between the power battery temperature and the target temperature; the energy consumption penalty item is used to reflect the total power consumption of the compressor, fan, and water pump; and the safety constraint penalty item is used to reflect situations where the battery temperature, refrigerant pressure, or actuator operating range exceeds the safety boundary.
[0019] Based on the aforementioned technical means, by designing the global reward signal to include at least temperature control reward items, energy consumption penalty items, and safety constraint penalty items, each agent is simultaneously guided by three dimensions during training: temperature tracking accuracy, system operating energy consumption, and safety boundary compliance. This enables the agent to minimize the total power consumption of the compressor, fan, and water pump while pursuing precise temperature control of the passenger compartment and power battery, and to actively avoid dangerous operating conditions such as battery temperature exceeding limits and refrigerant pressure exceeding standards. Ultimately, the agent learns a collaborative control strategy that achieves the optimal balance between comfort, energy efficiency, and safety.
[0020] Furthermore, multiple expert sample data are acquired, including: under multiple historical operating conditions, each preset controller controls the corresponding actuator group to operate, generating multiple preset trajectories; wherein, each intelligent agent corresponds to a preset controller, and the preset controller controls one actuator group in the heat pump system. The preset controller is used to generate control actions for the corresponding actuator group based on preset control logic and the deviation between the current state data and the target state data of the heat pump system under multiple historical operating conditions; based on the parameter set of each preset trajectory, multiple preset trajectories are scored for quality; the parameter set includes at least one of the following: temperature overshoot, steady-state error, total system power consumption, action smoothness, and number of safety constraint violations; preset trajectories with quality scores less than a preset score threshold or meeting preset rejection conditions are removed, and expert sample data is determined based on the multiple preset trajectories after removal.
[0021] Based on the aforementioned technical means, multiple preset trajectories are generated by each preset controller under various historical operating conditions. The trajectories are then scored and screened based on multi-dimensional indicators such as temperature overshoot, steady-state error, total system power consumption, action smoothness, and number of safety constraint violations. This process can automatically eliminate low-quality trajectories with poor control quality, high energy consumption, or potential safety hazards from a large number of preset trajectories, while retaining high-quality trajectories with excellent overall performance as expert sample data. This provides reliable, diverse, and secure expert prior data for subsequent behavior cloning pre-training, ensuring the quality of the model's initial strategy.
[0022] Furthermore, the preset rejection conditions include: the number of safety constraint violations exceeds a preset threshold; the actuator saturation ratio exceeds a preset saturation threshold; the absolute value of the difference between the actual temperature and the desired temperature of the passenger compartment exceeds a preset temperature deviation threshold for a number of cycles reaching a preset continuous cycle threshold; and the absolute value of the difference between the actual temperature and the desired battery temperature exceeds a preset temperature deviation threshold for a number of cycles reaching a preset continuous cycle threshold.
[0023] Based on the aforementioned technical means, by setting multiple hard rejection conditions such as the number of safety constraint violations, actuator saturation ratio, and continuous and significant deviations in the temperature of the occupant compartment and power battery, preset trajectories that meet the quality score but have serious single defects can be vetoed. This effectively prevents trajectories with local defects such as battery temperature runaway, refrigerant pressure exceeding limits, or actuator saturation for a long time from being mixed into the expert sample data, further ensuring the safety and reliability of the expert sample data and ensuring that the control laws learned by each agent in the pre-training stage strictly comply with engineering safety requirements.
[0024] Furthermore, the method also includes: performing preprocessing operations on the global operating data of the heat pump system to extract local operating data from the preprocessed global operating data; wherein the preprocessing operations include at least one of the following: outlier detection and correction operations, low-pass filtering operations, and normalization processing.
[0025] Based on the above technical means, by performing preprocessing operations such as outlier detection and correction, low-pass filtering and normalization on the global operating data of the heat pump system, sensor noise and abnormal jump interference can be effectively eliminated, high-frequency fluctuations of continuous state quantities can be smoothed, and the scale of physical quantities with different dimensions can be unified, so that the quality of the extracted local operating data is higher. This helps to improve the inference accuracy and stability of each agent policy network in actual deployment, while ensuring the consistency of input data distribution during the training and deployment phases.
[0026] Secondly, this application provides a control device for a heat pump system, comprising: a data acquisition module, a local extraction module, an action execution module, and a system operation module; the data acquisition module is used to acquire the current global operating data of the heat pump system; the global operating data includes: the current control state of each actuator in the heat pump system, the current actual temperature and set temperature of the area controlled by the heat pump system, and the operating energy consumption of the heat pump system; the local extraction module is used to extract local operating data related to each actuator group from the current global operating data; the local operating data includes at least: the current temperature error of the controlled area affected by the corresponding actuator group, and the current control state of the actuator group; the local extraction module is further used to, based on the local data related to the actuator group, extract local operating data... The system uses operational data to determine the input feature vectors of the agents associated with each actuator group. Each actuator group corresponds to one agent. The agent is used to: determine the control actions required by the corresponding actuator group, aiming to reduce the current temperature error of the controlled area and the operating energy consumption of the heat pump system; and to estimate the operating state of the local area or controlled object corresponding to the agent after executing the control actions. The input feature vectors characterize the current operating state of the local area corresponding to the agent and the operating energy consumption of the heat pump system. The action execution module inputs the corresponding input feature vectors into the agent corresponding to the actuator group and outputs the target control actions of the actuator group. The system operation module controls the operation of the heat pump system based on the target control actions of each actuator group.
[0027] Thirdly, this application provides an electronic device comprising: a processor and a memory; the memory storing processor-executable instructions. When the processor is configured to execute the instructions, the electronic device implements the method described in the first aspect.
[0028] Fourthly, this application provides a vehicle that includes the electronic equipment described in the third aspect.
[0029] Fifthly, this application provides a computer-readable storage medium that, when the instructions in the computer-readable storage medium are executed by a vehicle's processor, enables the vehicle to perform the methods described in the first aspect and any of their possible embodiments.
[0030] In a sixth aspect, this application provides a computer program product including computer instructions that, when executed on a vehicle, cause the vehicle to perform the method described in the first aspect and any possible implementation thereof.
[0031] It should be noted that the technical effects of any of the implementation methods in aspects two through six can be found in the technical effects of the corresponding implementation methods in aspect one, and will not be repeated here.
[0032] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0033] Figure 1 A schematic diagram of the composition of a heat pump system provided in this application; Figure 2 A schematic diagram of another heat pump system provided in this application; Figure 3 A flowchart illustrating a heat pump system control method provided in this application; Figure 4 This application provides a flowchart illustrating the process of acquiring multiple expert sample data. Figure 5 A schematic diagram illustrating the process of pre-training multiple expert sample data provided in this application; Figure 6 A flowchart illustrating a training process using online interactive experience sample data provided in this application; Figure 7 A complete flowchart of an agent training method provided in this application is shown. Figure 8 A schematic diagram of a reinforcement learning training logic provided in this application; Figure 9 A flow chart of a heat pump system control device provided in this application; Figure 10 This is a schematic diagram of the composition of an electronic device provided in this application. Detailed Implementation
[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0035] It should be noted that in the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0036] In the embodiments of this application, the terms "first," "second," "third," "fourth," "fifth," and "sixth" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," "third," "fourth," "fifth," and "sixth" may explicitly or implicitly include one or more of that feature.
[0037] In embodiments of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. For "A and / or B," this includes three combinations: A only, B only, and a combination of A and B.
[0038] Energy shortages and environmental policies have driven the widespread adoption of electric vehicles. Their heat pump systems must simultaneously manage passenger compartment temperature and protect the power battery. Lithium batteries are best suited for operating temperatures between 20°C and 40°C; abnormal temperature fluctuations can lead to safety hazards or reduced performance. Currently, most electric vehicles use heat pump systems for thermal management, as traditional control methods struggle to balance temperature control, energy consumption, and range requirements.
[0039] Relevant control strategies can be mainly divided into three categories: rule-based, model-based optimization, and deep reinforcement learning. Among them, rule-based methods such as ON / OFF, PID, and fuzzy control are simple to implement, but they rely too much on human experience, have poor adaptability to operating conditions, and have high energy consumption. Model-based optimization methods such as dynamic programming and genetic algorithms rely on pre-built system mechanism models for calculation and prediction, which not only has high modeling requirements, but also suffers from large computational load and insufficient real-time performance.
[0040] Therefore, there is an urgent need for a heat pump system control method that can adapt to multiple operating conditions, perform real-time decoupled control, and balance temperature control accuracy and energy efficiency.
[0041] Based on this, this application proposes a control method for a heat pump system. By acquiring current global operating data, including the current control state of each actuator, the current actual temperature and set temperature of the controlled area, and the system operating energy consumption, local operating data related to each actuator group is extracted from the data. An input feature vector is constructed based on the local operating data and the system operating energy consumption. Each agent aims to reduce the current temperature error of the controlled area and the system operating energy consumption, and outputs a target control action based on the input feature vector, thereby controlling the operation of the heat pump system. This method directly incorporates system operating energy consumption into the decision input of the agents, enabling each agent to simultaneously perceive temperature deviation and energy consumption level when outputting control actions. This achieves synergistic optimization between precise temperature control and system energy saving, avoiding the problem of temperature control and energy consumption optimization being isolated from each other in traditional methods. At the same time, each agent only needs to make decisions and output control actions and state predictions based on local operating data related to its corresponding actuator group and the overall system energy consumption. Distributed collaborative control of multiple actuator groups is achieved without the need for real-time communication between agents, effectively reducing the input dimension and decision complexity of each agent, meeting the real-time and reliability requirements of the vehicle controller, and thus optimizing the overall energy consumption of the heat pump system while ensuring passenger cabin comfort and power battery thermal safety.
[0042] The embodiments of this application are described below with reference to the accompanying drawings.
[0043] Please see Figure 1 , Figure 1 The present application provides a heat pump system comprising a plurality of actuator groups 101, a sensor group 102, and a controller 103. The controller 103 is connected to the plurality of actuator groups 101 and the sensor group 102.
[0044] The actuator groups 101 may include compressor actuator groups, electronic expansion valve actuator groups, fan actuator groups, and water pump actuator groups. Each actuator group is used to regulate the temperature of the affected controlled area.
[0045] In one implementation, the compressor actuator assembly can be arranged in the refrigerant circulation loop of the heat pump system. As one possible implementation, the compressor actuator assembly includes an electric compressor and its drive module for driving the refrigerant circulation and regulating the heating or cooling capacity of the heat pump system.
[0046] In one implementation, an electronic expansion valve actuator assembly can be arranged on the refrigerant line between the condenser and evaporator of a heat pump system. As one possible implementation, the electronic expansion valve actuator assembly includes an electronic expansion valve and its stepping drive unit for adjusting the refrigerant throttling opening and controlling the refrigerant flow rate on the evaporator or condenser side.
[0047] In one implementation, the fan actuator assembly can be arranged on the condenser and / or evaporator side of the heat pump system. As one possible implementation, the fan actuator assembly includes a condenser fan and an evaporator fan for regulating the air-side heat transfer intensity, thus affecting the cabin temperature response.
[0048] In one implementation, the water pump actuator assembly can be arranged in the battery cooling circuit or other heat exchange circuit of the heat pump system. As one possible implementation, the water pump actuator assembly includes a coolant circulating water pump for driving coolant circulation and regulating the temperature of the power battery or drive motor.
[0049] Sensor group 102 may include a temperature sensor, a pressure sensor, a flow sensor, a current sensor, and a voltage sensor. In one possible implementation, the temperature sensor is used to collect real-time data on the actual temperature of the passenger compartment, the power battery temperature, the drive motor temperature, the battery coolant temperature, the ambient temperature, and the refrigerant temperatures at the inlet and outlet of the condenser and evaporator; the pressure sensor is used to collect real-time suction and exhaust pressures; the flow sensor is used to collect real-time coolant flow rate; and the current and voltage sensors are used to collect real-time current and voltage data for each actuator group to calculate real-time power consumption. The data collected by these sensors collectively constitutes the current global operating data of the heat pump system.
[0050] The controller 103 can be a vehicle controller or a dedicated thermal management controller. In one possible implementation, the controller 103 is used to deploy multiple trained agents, each agent corresponding to an actuator group. The controller 103 receives current global operating data collected by the sensor group 102, extracts local operating data related to each actuator group from the current global operating data, inputs the local operating data into the corresponding agents, outputs control actions for each actuator group, and controls the operation of the corresponding actuator group in the heat pump system based on each control action.
[0051] Please see Figure 2 , Figure 2 A schematic diagram of another heat pump system provided in this application includes: a thermal management execution module 201, a state observation module 202, a multi-agent control module 203, an expert experience generation module 204, and a reinforcement learning training module 205.
[0052] and Figure 1Corresponding to the system architecture shown, the thermal management execution module 201 corresponds to multiple actuator groups 101, which is used to drive the heat pump system to operate according to control actions; the state observation module 202 corresponds to the sensor group 102, which is used to collect current global operating data; the multi-agent control module 203 is deployed in the controller 103, which is used to output control actions according to local operating data; the expert experience generation module 204 and the reinforcement learning training module 205 are modules for the offline training stage, which are used to generate expert sample data and perform multiple rounds of reinforcement learning training, respectively.
[0053] The thermal management execution module 201 includes a compressor actuator group, an electronic expansion valve actuator group, a fan actuator group, and a water pump actuator group. Each actuator group is used to regulate the temperature of the affected controlled area, receive control actions output by the multi-agent control module 203, drive the heat pump circulation loop to operate, and regulate the temperature of the passenger compartment and the power battery.
[0054] The status observation module 202 includes a temperature sensor, a pressure sensor, a flow sensor, a current sensor, and a voltage sensor. It is used to collect the current global operating data of the heat pump system in real time, and send the preprocessed global operating data to the multi-agent control module 203.
[0055] The multi-agent control module 203 is deployed in the controller 103 and includes multiple agents. Each agent corresponds to an actuator group and is used to extract local operating data related to its corresponding actuator group from the current global operating data, and output the control action of the corresponding actuator group based on the local operating data.
[0056] The expert experience generation module 204 is used to control the corresponding actuator group to run under multiple historical working conditions by each preset controller, generate multiple preset trajectories, and construct expert sample data after quality scoring and screening of the preset trajectories.
[0057] The reinforcement learning training module 205 is used to jointly train the agents in the multi-agent control module 203 using expert sample data and online interaction experience sample data. During the joint training process, the sampling ratio of expert sample data gradually decreases and the sampling ratio of online interaction experience sample data gradually increases. Finally, the trained agents are deployed to the multi-agent control module 203.
[0058] It should be noted that, in the embodiments of this application, Figure 1 or Figure 2 The illustrated structure does not constitute a limitation on the heat pump system of this application. The system may include more or fewer components than shown in the figures, or combine or separate certain components, or employ different component arrangements. The components shown in the figures can be implemented in hardware, software, or a combination of both.
[0059] Please see Figure 3 This application provides a flowchart of a heat pump system control method, applied to the above-mentioned... Figure 1 The controller and control methods of a heat pump system include: S301. Obtain the current global operating data of the heat pump system.
[0060] Among them, the current global operating data refers to the set of parameters that the heat pump system collects in real time through various sensors arranged in the heat pump system during the current control cycle, which are used to characterize the complete operating status of the heat pump system.
[0061] In this embodiment of the application, the current global operating data includes: the current control status of each actuator in the heat pump system, the current actual temperature and set temperature of the area controlled by the heat pump system, and the operating energy consumption of the heat pump system.
[0062] In the heat pump system, the current control status parameters of each actuator group refer to the actual operating parameters of each actuator group in the current control cycle, reflecting the current operating status of each actuator group. The current control status parameters of the actuator group include: the current speed of the compressor actuator group, which can be acquired through a speed sensor mounted on the compressor or obtained through feedback from the compressor drive module; the current opening degree of the electronic expansion valve actuator group, which can be obtained through feedback from the stepper drive unit of the electronic expansion valve; the current speed of the fan actuator group, which can be obtained through feedback from the fan drive module or acquired through a speed sensor; and the current duty cycle or flow rate of the water pump actuator group, which can be obtained through feedback from the water pump drive module or acquired through a flow sensor. In addition, the current control status parameters may also include the control actions of each actuator group in the previous control cycle, reflecting the historical control information of the actuator group.
[0063] The current actual temperature and set temperature of the area controlled by the heat pump system reflect the system's temperature control target and current control effect. The current actual temperature of the controlled area includes: the actual temperature of the passenger compartment, which can be obtained in real time through temperature sensors located within the passenger compartment; and the actual temperature of the power battery, which can be obtained in real time through temperature sensors located within the power battery pack. The set temperature includes: the passenger compartment set temperature, set by the user through the vehicle's air conditioning control panel or automatically provided by the vehicle's control strategy; and the power battery target temperature, determined by the battery management system based on the battery's optimal operating temperature range. By comparing the current actual temperature with the set temperature, the temperature control deviation can be calculated, providing a basis for subsequent control decisions by the intelligent agent.
[0064] The current operating power consumption of the heat pump system reflects the system's energy consumption level during the current control cycle. This current operating power consumption includes the real-time power consumption of the compressor actuator group, fan actuator group, and water pump actuator group. The real-time power consumption of each actuator group is obtained by collecting current and voltage values in real time using current and voltage sensors located in the power supply circuit of each actuator group. The real-time power of each actuator group is then calculated, and the sum of the real-time power of each actuator group yields the total current power consumption of the system. This current operating power consumption serves as a crucial input for the energy consumption penalty term, guiding each agent to minimize system energy consumption while meeting temperature control requirements.
[0065] As one possible implementation, the current global operating data of the heat pump system is obtained by: real-time acquisition of information such as ambient temperature, passenger compartment set temperature, passenger compartment actual temperature, power battery temperature, drive motor temperature, battery coolant temperature, compressor speed, suction pressure, exhaust pressure, condenser inlet and outlet refrigerant temperature, evaporator inlet and outlet refrigerant temperature, fan speed, water pump speed, and control actions of each actuator group at the previous moment through temperature sensors, pressure sensors, flow sensors, current sensors, voltage sensors, and actuator feedback sensors arranged in the heat pump system, to form the current global operating data.
[0066] It should be understood that after acquiring the raw sensor data, preprocessing operations can be performed on the current global operating data. These preprocessing operations include outlier detection and correction, low-pass filtering, rate of change calculation, and normalization to form an enhanced state representation, thereby improving the accuracy and stability of subsequent agent inference. The specific steps of the preprocessing operations are consistent with the preprocessing operations performed on the system state vector and the second sample operating data during the training phase, to ensure the consistency of the input data distribution between the training and deployment phases.
[0067] In this embodiment, the heat pump system includes multiple actuator groups, each used for temperature regulation from different dimensions. These different dimensions refer to the controlled area affected, i.e., different actuator groups.
[0068] As one possible implementation, multiple actuator groups regulate temperature from different dimensions. This means that each actuator group undertakes different thermal management functions in the heat pump system and jointly achieves coordinated control of the passenger compartment temperature and the power battery temperature through different physical means.
[0069] One implementation includes multiple actuator groups: a compressor actuator group for driving refrigerant circulation and adjusting the heating or cooling capacity of the heat pump system; an electronic expansion valve actuator group for adjusting the refrigerant throttling opening and controlling the refrigerant flow rate on the evaporator or condenser side; a fan actuator group for adjusting the air-side heat exchange intensity and affecting the temperature response of the passenger compartment; and a water pump actuator group for driving coolant circulation and adjusting the temperature of the power battery or drive motor.
[0070] The compressor actuator group provides the basic power for temperature regulation from the refrigerant side by driving the refrigerant circulation to adjust the heating or cooling capacity of the system; the electronic expansion valve actuator group controls the refrigerant flow on the evaporator or condenser side by adjusting the refrigerant throttling opening, affecting heat exchange efficiency from the perspective of refrigerant flow distribution; the fan actuator group affects the temperature response speed of the passenger compartment by adjusting the heat exchange intensity on the air side, achieving rapid temperature regulation of the passenger compartment from the perspective of air-side heat exchange; and the water pump actuator group regulates the temperature of the power battery or drive motor by driving the coolant circulation, ensuring battery thermal safety from the perspective of liquid-side circulation.
[0071] It should be understood that by assigning temperature regulation tasks of different dimensions to corresponding actuator groups, and having the agents corresponding to each actuator group make independent decisions based on local operating data, the complex multi-objective thermal management problem can be decoupled into multiple mutually cooperating sub-control problems. This reduces the decision-making difficulty of individual agents while achieving global optimization of the overall performance of the heat pump system.
[0072] S302. Extract the local operating data related to each actuator group from the current global operating data.
[0073] As one possible implementation, local runtime data related to the actuator group corresponding to each agent is extracted from the current global runtime data, including: for the first... Each agent uses the same state extraction function as during the training phase to filter out state variables related to the control decisions of the actuator group corresponding to that agent from the current global running data, thus forming the agent's local running data. The training process is described below and will not be repeated here.
[0074] For example, the compressor intelligent agent extracts state variables related to refrigerant circulation, such as passenger compartment temperature, power battery temperature, suction pressure, and discharge pressure; the electronic expansion valve intelligent agent extracts state variables related to throttling control, such as suction superheat and evaporator inlet and outlet refrigerant temperatures; the fan intelligent agent extracts state variables related to air-side heat exchange, such as condenser outlet temperature and passenger compartment temperature; and the water pump intelligent agent extracts state variables related to coolant circulation, such as power battery temperature and battery coolant temperature.
[0075] It should be understood that each agent only needs to obtain local operational data that is directly related to its control decisions, rather than all global operational data. This reduces the input dimension of each agent, which is conducive to improving the computational efficiency and real-time performance of online inference. At the same time, it also allows each agent to operate independently, meeting the real-time requirements of the vehicle controller.
[0076] In this embodiment of the application, the local operating data includes at least: the current temperature error of the controlled area affected by the corresponding actuator group, and the current control state of the actuator group.
[0077] The current temperature error of the controlled area affected by the corresponding actuator group refers to the deviation between the actual value and the set value of the temperature control target that the actuator group can directly or indirectly affect through its own control actions. Since different actuator groups perform different temperature regulation functions in the heat pump system, the controlled areas affected by each actuator group and the corresponding temperature errors also differ. The current temperature error can be calculated from the current actual temperature and the set temperature in the global operating data.
[0078] In one implementation, the compressor actuator assembly affects the heat exchange capacity of the passenger compartment and the power battery by driving refrigerant circulation. Therefore, the current temperature error of the controlled area affected by this assembly includes both passenger compartment temperature error and battery temperature error. The passenger compartment temperature error is the difference between the actual passenger compartment temperature and the set passenger compartment temperature, while the battery temperature error is the difference between the actual power battery temperature and the target power battery temperature. The electronic expansion valve actuator assembly affects the refrigerant state at the evaporator outlet by adjusting the refrigerant throttling opening. The current temperature error of the controlled area affected by this assembly includes the suction superheat error, i.e., the difference between the actual suction superheat and the target suction superheat. The fan actuator assembly affects the passenger compartment temperature response and condenser heat exchange effect by adjusting the air-side heat exchange intensity. The current temperature error of the controlled area affected by this assembly includes both the passenger compartment temperature error and the condenser outlet temperature error. The water pump actuator assembly affects the heat dissipation effect of the power battery or drive motor by driving coolant circulation. The current temperature error of the controlled area affected by this assembly includes either the battery temperature error or the drive motor temperature error.
[0079] The current control status parameters of the actuator group refer to the actual operating parameters of the actuator group in the current control cycle, reflecting the current working status of the actuator group. These current control status parameters can be collected by sensors arranged on each actuator group or obtained through feedback from the actuator drive module. The current control status parameters for the compressor actuator group include the current compressor speed; for the electronic expansion valve actuator group, they include the current opening degree of the electronic expansion valve; for the fan actuator group, they include the current fan speed; and for the water pump actuator group, they include the current water pump duty cycle or current flow rate.
[0080] It should be understood that each agent only needs to extract local operational data directly related to its corresponding actuator group, rather than all global operational data. By decoupling the global operational data according to the control function of the actuator group, each agent only focuses on the temperature error of the controlled area it can influence and the current state of its own actuator group. This effectively reduces the input dimension of each agent, reduces interference from irrelevant information, and helps improve the targeting and computational efficiency of agent reasoning.
[0081] In one possible implementation, when the actuator group is a compressor actuator group, the current temperature error includes the passenger compartment temperature error and the battery temperature error, and the current control state quantity includes the compressor current speed; when the actuator group is an electronic expansion valve actuator group, the current temperature error includes the suction superheat error and the evaporator outlet temperature error; when the actuator group is a fan actuator group, the current temperature error includes the condenser heat exchange state and the passenger compartment temperature error; when the actuator group is a water pump actuator group, the current temperature error includes the battery temperature error and the coolant temperature error.
[0082] As one possible implementation, local operating data related to each actuator group is extracted, including: when the actuator group is a compressor actuator group, extracting the passenger compartment temperature error, battery temperature error, and current compressor speed as local operating data; when the actuator group is an electronic expansion valve actuator group, extracting the suction superheat error and evaporator outlet temperature as local operating data; when the actuator group is a fan actuator group, extracting the condenser heat exchange status and passenger compartment temperature error as local operating data; and when the actuator group is a water pump actuator group, extracting the battery temperature error and coolant temperature as local operating data.
[0083] Based on this implementation, by selectively extracting local operational data directly related to the control target of the actuator group from the global operational data according to the control functions and physical characteristics of different actuator groups, the input characteristics of each agent are highly matched with its control task. This effectively filters out redundant information that is irrelevant to the current control decision, further reduces the input dimension, improves the pertinence and accuracy of agent decision-making, thereby accelerating the inference speed and enhancing the robustness of the control strategy.
[0084] Based on S302, by extracting the local operational data corresponding to each agent from the current global operational data, each agent can independently complete inference based on only a small amount of state information directly related to its control decisions. This effectively reduces the input dimension, improves the computational efficiency and real-time performance of online inference, and provides a foundation for the distributed and independent deployment of each agent, meeting the real-time and reliability requirements of the vehicle controller.
[0085] S303. Based on the local operating data related to the actuator group and the operating energy consumption of the heat pump system, determine the input feature vector of the intelligent agent associated with the actuator group.
[0086] As one possible implementation, the intelligent agent is used to: determine the control actions that the corresponding actuator group needs to perform with the goal of reducing the current temperature error of the controlled area and the operating energy consumption of the heat pump system.
[0087] As one possible implementation, based on the local operating data related to the actuator group and the operating energy consumption of the heat pump system, the input feature vector of the agent associated with the actuator group is determined, including: concatenating the local operating data extracted from the global operating data at the current moment with the operating energy consumption related data of the heat pump system to form the input feature vector.
[0088] For example, for the first An intelligent agent, at any time Input feature vector The following relationship must be satisfied:
[0089] in, For a moment Extracted from the current global runtime data, and related to the first The local operating data corresponding to each intelligent agent includes at least the current temperature error of the controlled area affected by the actuator group corresponding to the intelligent agent and the current control state of the actuator group. For a moment Energy consumption data related to the operation of the heat pump system, including the current total power consumption or normalized energy consumption value of the compressor, fan and water pump.
[0090] Taking the agent corresponding to the compressor actuator group as an example, its input feature vector includes: This includes passenger compartment temperature error, battery temperature error, and current compressor speed; The current total power consumption of the compressor, fan, and water pump is given. The agent aims to reduce cabin temperature error, battery temperature error, and total system power consumption, and determines the target compressor speed based on this input feature vector.
[0091] It should be understood that by directly concatenating local operational data with system energy consumption as input feature vectors, each agent can perceive not only the temperature control deviation of the controlled area and the current operating status of its own actuators, but also the overall energy consumption level of the heat pump system when making decisions. This input design allows the agent to proactively consider energy-saving effects while pursuing precise temperature control when outputting control actions. This avoids the problem of "excessive energy consumption in pursuit of temperature control accuracy" that may occur when relying solely on temperature errors for decision-making. It achieves built-in synergy between temperature control and energy consumption optimization, while maintaining a simple structure that facilitates engineering implementation and real-time inference.
[0092] As one possible embodiment, the intelligent agent is also used to: estimate the estimated operating data after the actuator group performs the control action based on the determined control action to be performed by the actuator group.
[0093] The estimated operational data refers to the predicted value of the operational state of the local region or controlled object corresponding to the agent at a future time after the agent executes the target control action. The estimated operational data is used to compare with the actual local operational data extracted from the current global operational data in the next control cycle to calculate the state prediction deviation. This state prediction deviation will be incorporated into the input feature vector at the next time step, forming a closed-loop correction mechanism of "decision-prediction-feedback".
[0094] In one implementation, the estimated operating data includes at least one of the following: the estimated temperature error of the controlled area affected by the actuator group at the next moment, the estimated control state quantity of the actuator group at the next moment, and the estimated operating energy consumption of the heat pump system at the next moment.
[0095] For example, the estimated operating data output by the agents corresponding to different actuator groups varies depending on their control functions. The estimated operating data output by the agent corresponding to the compressor actuator group includes: estimated passenger compartment temperature error and estimated battery temperature error, i.e., the difference between the passenger compartment temperature predicted at the current moment and the set temperature at the next moment after the compressor reaches its target speed, and the difference between the power battery temperature and the target temperature. The estimated operating data output by the agent corresponding to the electronic expansion valve actuator group includes: estimated suction superheat error, i.e., the difference between the suction superheat predicted at the current moment and the target suction superheat at the next moment after the electronic expansion valve reaches its target opening. The estimated operating data output by the agent corresponding to the fan actuator group includes: estimated passenger compartment temperature error and estimated condenser heat exchange status, i.e., the difference between the passenger compartment temperature predicted at the current moment and the set temperature at the next moment after the fan reaches its target speed, and the condenser heat exchange effect. The estimated operating data output by the intelligent agent corresponding to the water pump actuator group includes: estimated battery temperature error, which is the difference between the predicted power battery temperature and the target temperature at the next moment after the target duty cycle of the water pump is executed.
[0096] It should be understood that by enabling the agent to output predicted data about future operating states while simultaneously outputting control actions, the agent becomes not only a passive feedback controller but also a forward-looking decision-making unit with an internal predictive model. In the next control cycle, after the system collects real local operating data through sensors, it compares this data with the predicted operating data output at the current moment to calculate the state prediction deviation. The larger the deviation, the less accurate the agent's understanding of the dynamic characteristics of the heat pump system, requiring greater correction in the next decision-making step; the smaller the deviation, the more accurate the agent's prediction of system behavior, and the more reliable its output control actions. This self-evaluation and self-correction mechanism can effectively address the large inertia and hysteresis characteristics of the heat pump system, reduce over-adjustment and temperature oscillations caused by system response delays, and improve the adaptability and robustness of the control strategy.
[0097] As one possible implementation, the input feature vector of the agent associated with the actuator group is determined based on the local operating data related to the actuator group and the operating energy consumption of the heat pump system. This further includes: determining the input feature vector of the agent associated with the actuator group based on the local operating data related to the actuator group, the operating energy consumption of the heat pump system, and the estimated operating data output by the agent associated with the actuator group at the previous moment, so that the agent outputs the target control action of the actuator group and the estimated operating data; the input feature vector also includes: the state prediction deviation of the agent for the local area at the previous moment.
[0098] The state prediction deviation of the agent for the local region at the previous moment refers to the difference between the estimated operating data output by the agent at the previous moment and the actual local operating data extracted from the current global operating data of the heat pump system at the current moment. This difference is used to quantify the accuracy of the agent's prediction of the dynamic characteristics of the heat pump system.
[0099] In one implementation, the state prediction deviation includes at least one of the following: temperature prediction deviation of the controlled area, i.e., the difference between the predicted temperature error of the controlled area at the previous moment and the actual temperature error of the controlled area at the current moment; actuator state prediction deviation, i.e., the difference between the predicted actuator control state quantity at the previous moment and the actual actuator control state quantity at the current moment; and operating energy consumption prediction deviation, i.e., the difference between the predicted heat pump system operating energy consumption at the previous moment and the actual heat pump system operating energy consumption at the current moment.
[0100] For example, taking the agent corresponding to the compressor actuator group as an example, the state prediction deviation calculated at the current moment includes: occupant cabin temperature prediction deviation, which is the difference between the predicted occupant cabin temperature error output by the agent at the previous moment and the actual occupant cabin temperature error collected from the sensor at the current moment; battery temperature prediction deviation, which is the difference between the predicted battery temperature error at the previous moment and the actual battery temperature error at the current moment; and compressor speed prediction deviation, which is the difference between the predicted compressor speed at the previous moment and the actual compressor speed at the current moment. The closer the deviation value is to zero, the more accurate the agent's prediction of the dynamic behavior of the heat pump system; the larger the deviation value, the less accurate the agent's understanding of the system's dynamic characteristics, and the more significant the correction needs to be made in the current decision.
[0101] It should be understood that incorporating the state prediction deviation from the previous moment into the input feature vector at the current moment allows the agent to perceive the changing trend of its own prediction capability. When the state prediction deviation remains large, it indicates a significant discrepancy between the agent's internal prediction model and the actual dynamic characteristics of the heat pump system. The agent may then tend towards a more conservative control strategy in its current decision-making, reducing over-adjustment caused by inaccurate predictions. When the state prediction deviation gradually converges to a smaller value, it indicates that the agent has accurately grasped the dynamic response law of the heat pump system, allowing it to more confidently adopt optimization strategies in its current decision-making. This self-evaluation and adaptive mechanism enables each agent to dynamically adjust its decision-making strategy based on its own prediction accuracy without external intervention, effectively addressing changes in the dynamic characteristics of the heat pump system under different operating conditions and improving the robustness and environmental adaptability of the control strategy.
[0102] As one possible implementation, based on the local operating data related to the actuator group and the operating energy consumption of the heat pump system, the target control action of the actuator group is output through the intelligent agent corresponding to the actuator group. This further includes: determining the input feature vector of the intelligent agent associated with the actuator group based on the local operating data related to the actuator group, the operating energy consumption of the heat pump system, and the estimated operating data output by the intelligent agent associated with the actuator group at the previous moment; wherein the input feature vector is used to characterize: the current operating state of the local area corresponding to the intelligent agent, the current energy consumption level of the heat pump system, and the state prediction deviation of the intelligent agent for the local area at the previous moment; the estimated operating data is used to compare with the actual local operating data at the next moment to calculate the state prediction deviation in the input feature vector at the next moment; the corresponding input feature vector is input into the intelligent agent corresponding to the actuator group, and the target control action and estimated operating data of the actuator group are output.
[0103] As one possible implementation, the input feature vector of the agent associated with the actuator group is determined based on the local operating data related to the actuator group, the operating energy consumption of the heat pump system, and the estimated operating data output by the agent associated with the actuator group at the previous moment.
[0104] Each actuator group corresponds to an intelligent agent. The intelligent agent is used to: determine the control actions that the corresponding actuator group needs to perform, and predict the operating state of the local area or controlled object corresponding to the intelligent agent after the execution of the control actions, with the goal of reducing the current temperature error of the controlled area and the operating energy consumption of the heat pump system.
[0105] It should be understood that the intelligent agent in this application differs from traditional controllers that only output a single control action; it possesses a dual output capability of "decision-prediction." On one hand, based on the local operating data perceived at the current moment and historical feedback information, the intelligent agent outputs target control actions to adjust the corresponding actuator group, directly participating in the real-time control of the heat pump system. On the other hand, the intelligent agent simultaneously outputs predicted data for future operating states, reflecting its expectation of the effects of its control actions. By comparing the predicted data with the actual collected data, the state prediction deviation can be obtained. This deviation serves as a quantitative indicator of the accuracy of the intelligent agent's understanding of the dynamic characteristics of the heat pump system and can be fed back to the intelligent agent's input at the next moment, forming a self-correcting closed-loop mechanism. This dual-output design makes the intelligent agent not only a passive feedback controller but also an intelligent decision-making unit with an internal predictive model, effectively addressing the large inertia and hysteresis characteristics of the heat pump system and reducing over-adjustment and temperature oscillations caused by system response delays.
[0106] The input feature vector is used to characterize: the current operating state of the local region corresponding to the agent, the agent's prediction deviation of the local region's state in the previous moment, the current energy consumption level of the heat pump system, and the instantaneous reward feedback obtained by the agent after executing the control action in the previous moment.
[0107] It should be understood that the input feature vector is the sole information basis for the agent's control decisions and state predictions, and its design directly determines the quality of the agent's decisions. This application expands the input feature vector into a comprehensive representation containing four complementary types of information: first, the current local operating state, providing the basis for the agent's real-time reactive decision-making; second, state prediction bias, enabling the agent to perceive its own inadequacy in prediction capabilities and compensate for it in decision-making; third, the current energy consumption level, allowing the agent to directly perceive the system's energy consumption status during decision-making and internalize energy-saving goals into every action selection; and fourth, immediate reward feedback, serving as a quantitative evaluation of the merits of the previous action, guiding the agent to prefer action strategies that have historically resulted in better temperature control and lower energy consumption. These four types of information complement each other, collectively constituting the agent's complete perception of the current situation. This allows the agent to upgrade from reactive decision-making that "only looks at the present" to forward-looking decision-making that "perceives the past, evaluates the present, and predicts the future," thereby learning the dynamic characteristics of the heat pump system more quickly and improving the accuracy, stability, and energy-saving effect of the control strategy.
[0108] As one possible implementation, based on the local operating data related to the actuator group, the operating energy consumption of the heat pump system, and the estimated operating data output by the agent associated with the actuator group at the previous moment, the input feature vector of the agent associated with the actuator group is determined. This includes: concatenating the local operating data extracted from the global operating data at the current moment, the operating energy consumption related data of the heat pump system, the estimated operating data output by the agent at the previous moment, the state prediction deviation between the actual local operating data at the current moment and the estimated operating data at the previous moment, and the reward scalar obtained after executing the control action at the previous moment to form the input feature vector.
[0109] Among them, local operation data is used to reflect the temperature error of the controlled area and the working status of the actuator itself as perceived by the actuator group corresponding to the agent at the current moment, which is the basis for the agent to make real-time decisions.
[0110] The energy consumption data of the heat pump system is used to reflect the energy consumption level of the system at the current moment, so that the intelligent agent can directly perceive the current energy consumption status when making decisions, and thus actively consider the energy-saving effect while pursuing temperature control accuracy.
[0111] The predicted operating data from the previous moment is the agent's prediction of the future state at the current moment in the previous control cycle. Incorporating it into the current input allows the agent to compare the difference between its prediction and the actual result.
[0112] State prediction bias is the difference between the actual local operating data at the current moment and the estimated operating data at the previous moment. It is used to quantify the accuracy of the agent's prediction of the dynamic characteristics of the heat pump system. The larger the bias, the less accurate the agent's understanding of the system's dynamic behavior, and the greater the need for correction in the decision-making at the current moment.
[0113] The reward scalar is determined based on the operating energy consumption of the heat pump system and the temperature error of the controlled area. It is used to reflect the immediate evaluation signal of environmental feedback after the control action was executed at the previous moment. The principle for determining the reward scalar is: the smaller the temperature error of the controlled area, the larger the value of the reward scalar; the lower the operating energy consumption of the heat pump system, the larger the value of the reward scalar.
[0114] One implementation method is to reward scalars. It can be determined in the following ways: ,in, The temperature error is the temperature error of the controlled area affected by the actuator group corresponding to the agent at the previous moment. This temperature error is the absolute value of the difference between the actual temperature of the controlled area and the set temperature. The smaller the temperature error, the better the temperature control effect. This represents the operating energy consumption of the heat pump system at the previous moment, including the total power consumption of the compressor, fan, and water pump. This is a temperature error evaluation function. When the temperature error is less than the preset allowable range, it outputs a positive reward; when the temperature error exceeds the preset allowable range, it outputs a negative penalty. This is an energy consumption evaluation function; the lower the energy consumption, the larger its output value, in order to encourage energy-saving behavior. and These are weighting coefficients for temperature control evaluation and energy consumption evaluation, used to adjust their relative importance in the reward scalar. This application does not impose specific limitations on this.
[0115] By incorporating the reward scalar determined based on temperature error and operating energy consumption into the input feature vector at the current moment, the agent can directly perceive the actual effect of its previous decision on temperature control and energy saving. When the reward scalar is high, it tends to continue similar action strategies, and when the reward scalar is low, it adjusts and corrects the action strategies, thereby achieving rapid online adaptive optimization of the strategy without waiting for a complete training cycle.
[0116] For example, for the first An intelligent agent, at any time Input feature vector The following relationship must be satisfied:
[0117] in, For a moment Extracted from the current global runtime data, and related to the first Local operational data corresponding to each intelligent agent; For a moment Energy consumption data related to the operation of heat pump systems; For the first The estimated operating data output by each agent in the previous moment; The state prediction deviation at the previous moment; This is the global reward signal from the previous moment.
[0118] It should be understood that by simultaneously incorporating five types of information—local operating data, operating energy consumption, estimated operating data, state prediction deviation, and reward scalar—into the input feature vector at the current moment, the agent can not only make reactive decisions based on the current real-time state, but also make forward-looking decisions using its own prediction model of the environment, its ability to perceive prediction errors, and its real-time evaluation of the merits and demerits of historical actions. This closed-loop mechanism of "decision-prediction-feedback-evaluation" enables the agent to learn more quickly the large inertia and hysteresis characteristics of the heat pump system, effectively distinguish which action sequences can bring better temperature control and lower energy consumption, improve the prediction accuracy and response speed of the control strategy, thereby reducing system oscillations caused by over-adjustment while reducing temperature errors and system energy consumption in the controlled area, and enhancing the robustness and engineering applicability of the control strategy.
[0119] S304. Input the corresponding input feature vector into the agent corresponding to the actuator group, and output the target control action of the actuator group.
[0120] Each actuator group corresponds to one intelligent agent. The intelligent agent is used to determine the target control action to be performed by the corresponding actuator group in the heat pump system based on the input feature vector related to its own actuator group, and to estimate the operating state of the local area or controlled object corresponding to the intelligent agent after the target control action is performed.
[0121] As one possible implementation, the corresponding input feature vector is input to the agent corresponding to the actuator group, and the target control action and estimated operation data of the actuator group are output, including: for the first... An intelligent agent, at time... Input feature vector Input the agent's policy network The policy network outputs the target control action corresponding to the agent through forward inference. and estimated operating data ,Right now:
[0122] in, For the first An intelligent agent at time The target control action to be output; For the first An intelligent agent at time The output estimated operating data is used to compare with the actual local operating data at the next time step to calculate the state prediction deviation. The target control actions of each agent are combined to form the target control actions of each actuator group in the current control cycle.
[0123] For example, the compressor agent outputs the compressor target speed and the estimated crew cabin temperature error and battery temperature error; the electronic expansion valve agent outputs the electronic expansion valve target opening degree and the estimated suction superheat; the fan agent outputs the fan target speed and the estimated crew cabin temperature error and condenser outlet temperature; and the water pump agent outputs the water pump target duty cycle or flow rate and the estimated battery temperature error.
[0124] It should be understood that the agent simultaneously outputs the target control action and the estimated operating data, enabling it to not only make control decisions but also predict the effectiveness of its own control. The estimated operating data will be used in the next control cycle to calculate the state prediction deviation, forming a closed-loop mechanism of "decision-prediction-feedback." This helps the agent learn the large inertia and hysteresis characteristics of the heat pump system, improving the prediction accuracy and response speed of the control strategy.
[0125] As one possible implementation, after the agent outputs the target control action, it is also necessary to perform physical range mapping and rate limiting processing on the target control action, map the standardized action value output by the policy network to the actual physical range of the corresponding actuator group, and limit the change amplitude of the action between adjacent control cycles to avoid system oscillation or mechanical shock caused by sudden changes in actuator commands.
[0126] Based on S304, by inputting an input feature vector that integrates current local operating data, previous estimated operating data, and state prediction deviation into each agent, each agent independently infers and simultaneously outputs the target control action and estimated operating data. This achieves distributed parallel decision-making and self-prediction feedback for multiple actuator groups. It can complete the generation of control commands and state prediction without real-time communication between agents, meeting the real-time and reliability requirements of the vehicle controller. At the same time, the feedback mechanism of prediction deviation improves the adaptability of each agent to the dynamic characteristics of the heat pump system.
[0127] S305. Control the operation of the heat pump system based on the target control action of each actuator group in multiple actuator groups.
[0128] As one possible implementation, the operation of the heat pump system is controlled based on the target control action of each actuator group in multiple actuator groups, including: sending each target control action after physical range mapping and rate limiting processing to the corresponding actuator group through a control bus or drive signal, and having each actuator group execute the corresponding control command to drive the heat pump system to operate in order to regulate the temperature of the passenger compartment and the power battery.
[0129] For example, the compressor actuator group adjusts the operating speed of the electric compressor according to the compressor target speed command; the electronic expansion valve actuator group adjusts the step position of the electronic expansion valve according to the electronic expansion valve target opening command; the fan actuator group adjusts the speed of the condenser fan and the evaporator fan through PWM signal according to the fan target speed command; and the water pump actuator group adjusts the operating flow rate of the coolant circulating water pump according to the water pump target duty cycle command.
[0130] It should be understood that during the process of issuing control commands to each actuator group, when communication abnormalities, sensor failures, critical state out-of-bounds errors, or strategy outputs exceeding safety boundaries are detected, the system can automatically switch to the control mode corresponding to the preset controller to ensure the thermal safety of the power battery, the comfort of the passenger cabin, and the operational reliability of the heat pump system.
[0131] Based on S301-S305, this application acquires the current global operating data of the heat pump system and extracts the local operating data corresponding to each agent. By using the trained multiple agents to independently output the target control actions of each actuator group, the distributed collaborative control of multiple actuator groups is realized. Under the premise of ensuring the safe operation of the heat pump system, the collaborative optimization of passenger cabin comfort, power battery thermal safety and system energy consumption can be effectively taken into account.
[0132] The following describes the training process of the intelligent agent in the embodiments of this application. This training process is used to obtain the intelligent agents corresponding to each actuator group in the above control method. Through offline training, each intelligent agent learns a control strategy that can achieve collaborative optimization of multiple actuator groups.
[0133] The next step is to use expert sample data to clone and train the agent.
[0134] As one possible implementation, multiple expert sample data are acquired; wherein, the expert sample data includes system state vectors and corresponding expert control actions; the behavior cloning pre-training is performed on multiple agents corresponding to multiple actuator groups using the multiple expert sample data, so that the deviation between the actions output by each agent and the corresponding expert control actions is minimized.
[0135] Among them, expert sample data refers to high-quality state-action pair samples generated by the preset controller through closed-loop interaction with the heat pump system under multiple historical operating conditions, which are used to provide supervision signals for the pre-training of behavioral clones of each agent. The preset controller is a proportional-integral controller designed based on engineering calibration experience, which can output safe, stable control actions that conform to physical constraints.
[0136] The system state vector refers to the global operating data of the heat pump system at a certain moment, which is used to characterize the complete operating state of the heat pump system at that moment. It includes at least information such as passenger compartment temperature, power battery temperature, compressor speed, suction pressure, exhaust pressure, fan speed, water pump speed, electronic expansion valve opening, and ambient temperature.
[0137] The corresponding expert control action refers to the control action vector jointly output by each preset controller according to the deviation between the current system state vector and the preset target state data, and according to the preset control logic. This includes the compressor target speed, the electronic expansion valve target opening, the fan target speed, and the water pump target duty cycle or flow rate.
[0138] It should be understood that the system state vector in the expert sample data corresponds one-to-one with the corresponding expert control action, together forming a "state-action" supervision signal pair, which is used to train the policy network of each agent, so that it learns the mapping relationship from system state to control action, and thus has a basic control capability similar to the preset controller before formally entering the joint fine-tuning of reinforcement learning.
[0139] As one possible implementation method, please refer to Figure 4 To obtain multiple expert sample data, the specific steps include: S401. Under multiple historical operating conditions, each preset controller controls the operation of the corresponding actuator group to generate multiple preset trajectories.
[0140] Each intelligent agent corresponds to a preset controller; the preset controller controls an actuator group in the heat pump system. The preset controller is used to generate control actions for the corresponding actuator group based on preset control logic and the deviation between the current state data and the target state data under multiple historical operating conditions of the heat pump system.
[0141] As one possible implementation, the multiple historical operating conditions include at least: low-temperature heat pump heating condition, high-temperature cooling condition, parking preheating condition, fast-charging battery cooling condition, high-rate discharge battery temperature control condition, and conditions with different vehicle speeds and passenger compartment heat loads.
[0142] It should be understood that by running the preset controller under a wealth of historical operating conditions, it is possible to ensure that the generated expert sample data has sufficient environmental coverage and strategy diversity.
[0143] As one possible implementation, the preset controller controls the operation of an actuator group in the heat pump system. Specifically, this includes: constructing a corresponding preset controller for each actuator group in the heat pump system; and generating control actions for the corresponding actuator group based on preset control logic and the deviation between the current state data and the target state data of the heat pump system.
[0144] For example, the preset control logic corresponding to each preset controller can be proportional-integral control logic. Each preset controller is configured with output limiting parameters, rate limiting parameters, and anti-integral saturation parameters, which are used to constrain the generated control actions within the physical allowable range of the corresponding actuator group, so as to avoid system instability or actuator damage due to control quantity exceeding limits, sudden action, or integral saturation, thereby ensuring that the expert control actions in the expert sample data are all within the safe boundary of engineering implementation.
[0145] As one possible implementation, the proportional-integral control logic corresponding to each preset controller specifically includes: Step 1: Calculate the deviation and the original output.
[0146] In one implementation, a preset controller calculates the deviation between the target value (i.e., the target state data mentioned above, as described below) and the current state value (i.e., the current state data mentioned above, as described below). The controller then performs proportional and integral calculations on the deviation to obtain the original output value. The proportional control term is proportional to the current deviation and is used to quickly adjust the actuator output based on the magnitude of the deviation. The integral control term integrates the cumulative deviation over time to eliminate steady-state error.
[0147] For example, the preset controller is based on the target value Compared with the current state value Calculate the deviation The deviation is then proportionally and integrally calculated to obtain the original output value. The specific calculation formula is as follows:
[0148]
[0149] in, It is a proportionality coefficient, which is directly proportional to the deviation and is used to quickly adjust the actuator output according to the current error. The larger the value, the faster the system response rate, but excessively large values... This value may cause the system to have a large overshoot, thus affecting the system's dynamic performance. The coefficient of integration is the integral coefficient, which is related to the integral term. Together they work to eliminate the steady-state error of the system. The larger the value, the faster the steady-state error can be eliminated, but it is also more prone to overshoot and oscillation.
[0150] Step 2: Calculation of integral terms and correction against integral saturation.
[0151] In one implementation, during the integral operation, when the controller output has reached the limit boundary, the further accumulation of the integral term is suppressed by the anti-integral saturation parameter, so as to avoid overshoot and oscillation caused by excessive accumulation of integral terms leading to the system still outputting an excessively large control quantity after the error has decreased.
[0152] For example, the integral term The calculation formula is as follows: .in, The sampling period. To prevent integral saturation, the integral term is used to suppress excessive accumulation of the integral term when the controller output reaches the limiting boundary, thereby avoiding integral saturation. When the system error persists for a long time and the control output has reached the upper or lower limit, if the integral term continues to accumulate, it will cause the controller to output an excessively large control quantity even after the error decreases, resulting in a large overshoot and oscillation. Therefore, when the output reaches the saturation boundary, the integral term is adjusted using equation (3). The term provides a reverse correction to the integral term to suppress further deviation of the integral term.
[0153] For example, in the continuous time domain, the above-mentioned anti-integral saturation mechanism can be described by an equation representing the rate of change of the integral term. In this case, the integral term... Dynamic satisfaction:
[0154] in, The derivative of the integral term with respect to time represents the rate of accumulation of the integral; This represents the current deviation. To resist integral saturation gain; This is the controller's raw output; This is the output after limiting. When the controller has not reached saturation... The integral speed is entirely determined by the deviation, which is conventional integral control. When the controller output exceeds physical limits... When, the correction term in the formula This immediately suppresses the integral rate, preventing the integral term from continuously increasing in the saturation direction, thus achieving smooth anti-saturation control in the continuous domain. In discretization, the differential equation can be discretized using methods such as Euler discretization to obtain the aforementioned iterative anti-integral saturation update formula.
[0155] Step 3: Output limiting processing.
[0156] In one implementation, the original output value is limited to obtain a limited control quantity, which is restricted to the minimum and maximum output values allowed by the actuator.
[0157] For example, the original output value is obtained. Then, it is subjected to amplitude limiting to ensure that the actuator control commands are always within their physical limits, avoiding the actuator from exceeding its working limits due to excessive or insufficient control input. Let... and Let these be the minimum and maximum allowable output values of the actuator, respectively. Then, the controlled quantity after limiting... Represented as:
[0158] Step 4: Rate limiting processing.
[0159] After output limiting, the rate of the limited control quantity is further limited to constrain the rate of change of the control output between adjacent time points, thus preventing system oscillation, mechanical shock, or unstable response caused by abrupt changes in actuator commands. Let... This is the final control output from the previous moment. This represents the maximum allowable increase within a single sampling period. This represents the maximum allowable decrease within a single sampling period, and , The final control output after rate limiting Represented as:
[0160] Through steps one through four above, the preset controller first calculates the original output based on the deviation between the target value and the current state value. Then, the original output is subjected to amplitude limiting to obtain The integral term is corrected using an anti-integral saturation mechanism based on the output deviation before and after the limiting, and finally the final control output sent to the actuator is obtained through rate limiting. Therefore, each preset controller can generate smooth, stable motion commands that conform to physical constraints, resulting in preset trajectories with high quality and engineering usability.
[0161] Based on the general proportional-integral control logic described above, the following explanation will further elaborate on the specific control requirements and parameter configurations of each actuator group.
[0162] In one possible implementation, the preset controller for the compressor takes a weighted combination of the passenger compartment temperature error and the battery temperature error as input and outputs the target compressor speed; the preset controller for the electronic expansion valve takes the suction superheat error or the evaporator outlet temperature error as input and outputs the target opening degree of the electronic expansion valve; the preset controller for the fan takes the condenser heat exchange state and the passenger compartment temperature error as input and outputs the target fan speed; and the preset controller for the water pump takes the battery coolant temperature error or the battery temperature error as input and outputs the target water pump duty cycle or flow rate.
[0163] It should be understood that the input error definitions of the above-mentioned preset controllers are determined based on the core control functions of each actuator group in the heat pump system.
[0164] As the core power source of the heat pump system, the compressor's speed affects both the cooling / heating capacity of the passenger compartment and the refrigerant supply to the battery cooling circuit. Therefore, a weighted combination of passenger compartment temperature error and battery temperature error is used as the comprehensive error input to coordinate and respond to the thermal management needs of multiple temperature zones. The main function of the electronic expansion valve is to control the superheat of the evaporator outlet by adjusting its opening degree, ensuring sufficient refrigerant evaporation within the evaporator and preventing liquid slugging. Therefore, the suction superheat error is used as the primary input; in some system configurations, the evaporator outlet temperature error can also be used as a substitute or supplementary input. The fan changes the heat exchange airflow on the condenser or evaporator side by adjusting its speed, thus affecting the condensing pressure and passenger compartment heat exchange effect. Therefore, the condenser outlet temperature error and passenger compartment temperature error are used together as inputs to balance system energy efficiency and passenger compartment comfort. The water pump drives the coolant to circulate in the battery cooling circuit; directly using the battery temperature error as input can effectively achieve closed-loop control of the battery temperature.
[0165] The rationality of the above error definition has been fully verified in engineering practice, and it can ensure that each preset controller achieves stable and efficient basic control performance on its respective control target.
[0166] For example, under the above proportional-integral control logic framework, the specific input error of each preset controller is defined as follows:
[0167]
[0168]
[0169]
[0170] in, Set the temperature for the crew cabin (i.e., the desired crew cabin temperature). For a moment The actual cabin temperature in the crew cabin The target temperature (i.e., the desired battery temperature) for the power battery. For a moment The temperature of the power battery, and These are the weights for crew cabin temperature error and battery temperature error, respectively. Set the intake superheat value. For a moment The actual intake superheat; For a moment The condenser outlet temperature, This is the target temperature at the condenser outlet.
[0171] The specific parameter configurations of each preset controller can be found in Table 1 below. The data in Table 1 are only some embodiments and do not impose specific limitations on this application.
[0172] Table 1
[0173] It should be noted that the items listed in Table 1 The proportional coefficient, integral coefficient, and anti-integral saturation parameter in the above formulas correspond to the upper and lower boundaries in the output limiting, respectively. and The upper and lower boundaries in the rate limit correspond to the formulas above. and By configuring according to the parameters, each preset controller can operate stably on its corresponding actuator group and generate high-quality control actions that meet engineering requirements, thereby providing a reliable data foundation for subsequent trajectory screening and expert sample data generation.
[0174] As one possible implementation, under multiple historical operating conditions, each preset controller controls the operation of its corresponding actuator group to generate multiple preset trajectories, specifically including: Under each historical operating condition, the heat pump system is placed in the corresponding initial state and boundary conditions. In each control cycle, each preset controller calculates its corresponding raw output value according to the proportional-integral control logic based on the deviation between the currently collected global state data and the preset target state data. This raw output value is then processed sequentially through output limiting and rate limiting to generate the final control action of the corresponding actuator group. After each actuator group executes the final control action, the operating state of the heat pump system shifts, generating new global state data. This process repeats until the operating duration of the condition reaches a preset value or a preset termination condition is met. The global state data collected in each control cycle throughout the entire operation, the joint control actions output by each preset controller, the instantaneous reward value, and the global state data at the next moment are recorded in chronological order to form a preset trajectory.
[0175] For example, after expanding multiple preset trajectories obtained under various historical operating conditions by time step, the global operation data, joint control actions, instantaneous reward values, and global operation data at the next moment corresponding to each time step are extracted to form a source set of expert sample data. :
[0176] in, For a moment The system state vector, i.e., the state vector of the heat pump system at time t, Global runtime data; For each preset controller at time Jointly output expert control actions; The immediate reward value obtained after performing the expert-controlled action; The global runtime data for the next moment after the expert control action is executed; This refers to the total number of samples after expanding all preset trajectories under different historical conditions according to time steps.
[0177] The above samples are organized according to a one-to-one correspondence of "system state vector - expert control action - immediate reward - global operation data at the next moment" to characterize the one-step state transition process of the heat pump system under different historical operating conditions, and to provide standardized expert sample data for subsequent behavior cloning pre-training.
[0178] It should be understood that the preset controller continuously performs closed-loop adjustment based on the deviation between the real-time state and the target state of the heat pump system. Its output control actions are always constrained by output amplitude and rate limits. Therefore, the generated preset trajectory not only contains rich operating condition coverage information, but also naturally carries engineering-based safe, smooth and stable control prior knowledge, providing reliable material for subsequent quality scoring and expert sample data screening.
[0179] Based on S401, multiple preset trajectories are generated by running a preset controller configured with output limiting, rate limiting and anti-integral saturation mechanism under multiple historical operating conditions. This can efficiently and automatically generate high-quality preset trajectories that cover a variety of operating scenarios and ensure safe and smooth control actions. This provides a reliable data source for subsequent screening of high-quality expert sample data and reduces the cost and inconsistency of manual collection and annotation.
[0180] S402. Based on the parameter set of each preset trajectory, perform quality scoring on multiple preset trajectories.
[0181] The parameter set includes at least one of the following: temperature overshoot, steady-state error, total system power consumption, smoothness of action, and number of safety constraint violations.
[0182] Temperature overshoot is used to characterize the maximum peak error of the actual temperature of the passenger compartment or the actual temperature of the power battery deviating from their respective target temperatures within a preset trajectory. This indicator reflects the maximum deviation of the preset controller during dynamic adjustment; the larger the overshoot, the more severe the instantaneous impact of the control process on comfort or battery safety.
[0183] Steady-state error characterizes the sustained deviation between the actual temperature of the passenger compartment and the set temperature, or the sustained deviation between the actual temperature of the power battery and the target temperature, during the stable operation phase of a preset trajectory. This indicator reflects the static control accuracy of the preset controller after reaching dynamic equilibrium. The smaller the steady-state error, the more accurately the control strategy can maintain the temperature within the desired range.
[0184] Total system power consumption characterizes the cumulative or average energy consumption of various actuator groups such as compressors, fans, and water pumps in a heat pump system over the entire operating time of a preset trajectory. This indicator reflects the energy economy of the corresponding control strategy; the lower the total power consumption, the more energy-efficient the strategy is while meeting temperature control requirements.
[0185] Motion smoothness characterizes the variation in control actions of each actuator group within a preset trajectory over adjacent control cycles. This metric reflects the smoothness of the control output; more drastic changes in motion are more likely to cause mechanical shocks, system oscillations, or noise problems. Smooth motion sequences help extend actuator life and improve system stability.
[0186] The number of safety constraint violations characterizes the total number of times critical operating parameters exceed preset safety boundaries within a pre-defined trajectory. Pre-defined safety boundaries include, but are not limited to, upper limits for battery temperature, upper and lower limits for refrigerant pressure, and actuator saturation limits. This indicator directly reflects the safety and reliability of the control strategy; fewer violations indicate a stronger ability of the strategy to ensure the system operates within the safety domain.
[0187] As one possible implementation, based on the parameter set of each preset trajectory, a quality score is performed on multiple preset trajectories, including: for each preset trajectory, calculating its temperature overshoot, steady-state error, total system power consumption, smoothness index, and number of safety constraint violations; normalizing each index according to a preset normalization upper limit to obtain a normalized score for each index; and weighting and summing the normalized scores according to the preset weights corresponding to each index to obtain the comprehensive quality score of the preset trajectory.
[0188] For example, suppose the first The time length of the trajectory is , for the The process of scoring the quality of the preset trajectory is as follows: (1) Calculation of temperature overshoot.
[0189] One implementation method is the crew cabin temperature overshoot. For the first The maximum positive deviation of the actual temperature in the passenger compartment from the set temperature in the preset trajectory, and the overshoot of the power battery temperature. For the first The maximum positive deviation of the actual battery temperature from the target temperature in the preset trajectory:
[0190]
[0191] in, and They are time points The actual temperature of the passenger compartment and the actual temperature of the power battery are in degrees Celsius (°C). and These represent the desired crew cabin temperature and the desired battery temperature, respectively, in degrees Celsius (°C); function This indicates that only positive deviation values are taken. When the actual temperature is lower than or equal to the set temperature, the overshoot contribution at that moment is zero.
[0192] It should be understood that the temperature overshoot mentioned above is defined as the maximum positive deviation, applicable to situations where the temperature exceeds the target value under heating conditions. For cooling conditions, the negative deviation where the actual temperature is lower than the target temperature can be used as the overshoot index, and the definition of the overshoot can be flexibly adjusted according to the specific operating condition.
[0193] (2) Calculation of steady-state error.
[0194] One implementation method is to set the steady-state evaluation window length as... In the first The last of the preset trajectory Within each time step, the average absolute value of the deviation between the actual temperature and the set temperature is calculated as the steady-state error:
[0195]
[0196] in, The steady-state error of the crew cabin for the kth preset trajectory is expressed in degrees Celsius (°C). The smaller the value, the higher the control accuracy of the crew cabin temperature in the steady-state stage. The steady-state error of the battery for the kth preset trajectory is expressed in degrees Celsius (°C). The smaller the value, the higher the control accuracy of the battery temperature in the steady-state stage. The steady-state evaluation window length represents the number of sampling points selected at the end of the trajectory to evaluate steady-state performance.
[0197] It should be understood that the steady-state assessment window The value of should ensure that the selected time period is in the dynamic equilibrium stage of the system, avoiding the inclusion of dynamic fluctuations during the initial adjustment phase in the assessment of steady-state error. One implementation method is... Based on the total length of the trajectory The proportion is determined, for example, by taking 20% of the time steps at the end of the trajectory, or by taking a fixed value based on experience. .
[0198] (3) Calculation of total system power consumption.
[0199] One implementation method is to calculate the first... The total cumulative energy consumption of the system during the entire duration of a preset trajectory satisfies the following relationship:
[0200] in, The total cumulative energy consumption of the system for the k-th preset trajectory, expressed in kilojoules (kJ) or kilowatt-hours (kWh). For a moment The total power consumption of all actuator groups in a heat pump system, including compressors, fans, and water pumps, is expressed in kilowatts (kW). This represents the total runtime of the k-th preset trajectory. The integral represents the continuous sum of the system's instantaneous total power consumption over the entire runtime.
[0201] It should be understood that in actual engineering implementation, since sensor data is sampled discretely, the above integral can be approximated by summing the discrete power consumption sequence using numerical integration methods (such as the trapezoidal rule or the rectangular method) to ensure computability on discrete systems.
[0202] (4) Calculation of motion smoothness.
[0203] One implementation involves calculating the average amplitude of the motion changes of all actuator groups at adjacent time points; a smaller amplitude indicates smoother motion. The smoothness of motion is determined by the following relationship:
[0204] in, The motion smoothness index for the k-th preset trajectory; the smaller the value, the smoother the control motion sequence. For the first A preset controller at time Output control actions; For the first A preset controller at time Output control actions; This is the preset total number of controllers, i.e., the number of actuator groups; The total number of time steps for the k-th preset trajectory; It represents the absolute change in action between two adjacent moments.
[0205] It should be understood that the motion smoothness index is essentially an average of the motion variation amplitudes of all actuators at all adjacent moments, which can comprehensively reflect the smoothness of the motion along the entire trajectory. The lower the index, the smoother the motion sequence output by the preset controller, which is more conducive to reducing mechanical wear of the actuators and transient shocks to the system.
[0206] (5) Calculation of actuator saturation ratio.
[0207] One implementation involves calculating the total percentage of time that the outputs of all actuator groups reach the amplitude limit boundary, satisfying the following relationship:
[0208] in, The actuator saturation ratio for the k-th preset trajectory, with a value range of [value missing]. The closer the value is to 1, the longer the actuator is in a saturated state. This is an indicator function; it takes the value 1 if the condition inside the parentheses is true, and 0 otherwise. and The first The maximum and minimum amplitude limits of the actuator group's actions.
[0209] It should be understood that the actuator saturation ratio reflects the control margin of the preset controller under the corresponding operating conditions. A high saturation ratio usually means that the preset controller has been operating at its limit capacity for a long time under the current operating conditions, with insufficient adjustment margin, and may be unable to cope with sudden disturbances.
[0210] (6) Calculation of the number of times safety constraints are violated.
[0211] One implementation involves calculating the total number of safety constraint violations along the entire trajectory, satisfying the following relationship:
[0212] in, The number of times the safety constraints of the k-th preset trajectory are violated is a non-negative integer. The more violations, the worse the safety of the trajectory. For a moment The safety constraint violation indicator is triggered when there is a violation of battery temperature limits, refrigerant pressure limits, actuator movement limits, or other safety constraints at that moment. ,otherwise ; This is an indicator function.
[0213] It should be understood that the criteria for determining safety constraints can be preset based on the actual system's safe operating boundaries. For example, the upper limit for battery temperature safety can be set to 45℃ or 50℃, and the upper limit for refrigerant high pressure safety can be set to 3.0MPa, etc. When any critical state exceeds the corresponding safety threshold, it is considered a violation of a safety constraint.
[0214] (7) Overall quality score.
[0215] Divide each of the above indicators by its corresponding normalized upper limit, multiply by the preset weight, and sum the results. This sum is then used as a deduction from the full score of 100 points to obtain the result. Overall quality score of the trajectory :
[0216] in, The overall quality score for the k-th preset trajectory is given, with a maximum score of 100 points. The higher the score, the better the overall control quality of the trajectory. to The weighting coefficients for each indicator satisfy the following conditions: This is used to adjust the importance of different indicators in the overall score; , , , , , , These are the normalization upper limits for each indicator, used to map indicator values of different dimensions to a comparable scale.
[0217] It should be understood that the design philosophy of the aforementioned comprehensive scoring function is as follows: starting from a maximum score of 100, points are deducted according to the normalized deviation value of each indicator based on its weight; the larger the deviation, the more points are deducted, and the lower the final score, the worse the trajectory quality. This design makes the scoring results intuitive and easy to understand, and facilitates the setting of a uniform scoring threshold for trajectory selection.
[0218] For example, Table 2 provides an example of the normalization upper limit and recommended weight of each scoring indicator. In actual applications, these can be adjusted according to the specific vehicle model and heat pump system characteristics. This example does not impose specific restrictions on this application.
[0219] Table 2
[0220] Based on S402, each preset trajectory is quantitatively scored using multi-dimensional indicators, which can objectively and comprehensively evaluate the overall performance of the preset controller under multiple objectives such as dynamic response, static accuracy, energy economy, motion smoothness and operational safety. This allows for the automatic selection of high-quality trajectories with balanced performance and excellent overall quality, providing reliable and diverse expert prior data for subsequent behavior cloning pre-training.
[0221] S403. Remove preset trajectories whose quality scores are less than the preset score threshold or meet the preset removal conditions, and determine expert sample data based on the multiple preset trajectories after removal.
[0222] It should be understood that the comprehensive quality score calculated in S402 reflects the overall performance of the preset trajectory under multiple indicators. Setting a preset score threshold can initially screen out trajectories that meet the comprehensive performance standards. Based on the comprehensive score meeting the standards, further preset rejection conditions are set to conduct a single-item veto for key safety and comfort indicators such as the number of safety constraint violations, actuator saturation ratio, and continuous significant temperature deviations. This can effectively prevent trajectories with serious local defects from being mixed into the expert sample data, thereby ensuring the quality and reliability of the expert experience base.
[0223] As one possible implementation, a preset trajectory that satisfies the condition that the quality score is less than a preset score threshold is removed from the multiple preset trajectories.
[0224] For example, This score is used to represent the overall quality score of the k-th preset trajectory, with a maximum score of 100. If the score is below the preset score threshold of 80 points, the trajectory whose overall quality does not reach a good level will be directly removed.
[0225] As one possible implementation, the number of safety constraint violations exceeds a preset threshold; the actuator saturation ratio exceeds a preset saturation threshold; the absolute value of the difference between the actual temperature and the desired temperature of the passenger compartment exceeds a preset temperature deviation threshold, and the number of consecutive cycles that meet this threshold reaches a preset consecutive cycle threshold; the absolute value of the difference between the actual temperature and the desired battery temperature exceeds a preset temperature deviation threshold, and the number of consecutive cycles that meet this threshold reaches a preset consecutive cycle threshold.
[0226] For example, for the k-th preset trajectory, if it meets any of the following conditions, it is determined to be a low-quality trajectory and is removed, as follows:
[0227] in, This is used to represent the number of times the safety constraints of the k-th preset trajectory are violated. This indicates that there is a safety constraint violation event in the k-th preset trajectory. The preset threshold for the number of violations is 0, meaning that no safety violations are allowed. This represents the actuator saturation ratio for the k-th preset trajectory, with a value ranging from [0,1]. This indicates that the actuator reaches the limit boundary for more than 30% of the time, and the trajectory that has been running at its limit for a long time and lacks control margin will be eliminated.
[0228] This indicates that the absolute value of the difference between the actual temperature and the desired temperature of the passenger cabin exceeds the preset temperature deviation threshold of 10°C, and this state is established for 20 consecutive control cycles. That is, the number of consecutive deviation cycles reaches the preset consecutive cycle threshold of 20, indicating that the preset controller has lost its ability to effectively regulate the temperature of the passenger cabin.
[0229] This indicates that the absolute value of the difference between the actual temperature of the power battery and the desired battery temperature exceeds the preset temperature deviation threshold of 10°C, and this state is established for 20 consecutive control cycles. That is, the number of consecutive deviation cycles reaches the preset consecutive cycle threshold of 20, indicating that the preset controller cannot maintain the battery temperature in a safe and efficient range.
[0230] It is understood that the preset scoring threshold, preset saturation threshold, preset temperature deviation threshold and preset continuous cycle threshold mentioned above are all exemplary settings. Those skilled in the art can adjust them according to actual needs such as the thermal inertia of the actual system, the length of the control cycle, the comfort standards of the passenger cabin and the upper limit of the battery safety temperature, so as to achieve a flexible balance between the strictness of the preset trajectory screening and the number of samples.
[0231] As one possible implementation, the expert sample data is determined based on the multiple preset trajectories after removal, including: Each preset trajectory retained after rejection is expanded by time step, and the global operating state vector and joint control action vector corresponding to each time step are extracted to construct a state-action pair. The global operating state vector serves as the system state vector, and the joint control action vector serves as the expert control action. All state-action pairs extracted from each retained trajectory are summarized to constitute the expert sample data.
[0232] It should be understood that the above extraction method ensures that each sample in the expert sample data comes from a high-quality trajectory that has been screened by both quality scoring and elimination criteria. Its state covers a variety of typical historical working conditions, and its actions have engineering safety, stability and rationality, thus providing high-quality supervision signals for behavior cloning pre-training.
[0233] Based on S401-S403, this application uses a screening method that combines multi-dimensional index quantitative scoring with preset scoring thresholds and single-item elimination conditions to automatically select high-quality expert trajectories with excellent comprehensive performance, providing reliable and safe prior data for subsequent training and ensuring the quality of the model's initial strategy and the stability of training.
[0234] As one possible implementation, behavioral cloning pre-training is performed on multiple agents corresponding to multiple actuator groups using multiple expert sample data to minimize the deviation between the actions output by each agent and the corresponding expert control actions. Specifically, this includes the following steps: Extract system state vectors and expert control action pairs from expert sample data to construct a sample set for behavior cloning pre-training. For the first... The first agent extracts local operational data related to the actuator group corresponding to that agent from the system state vector, constituting the agent's local operational data. This local operational data is then input into the first... A policy network for each agent outputs predicted control actions. The parameters of the policy network are updated using gradient descent, with the objective of minimizing the deviation between the predicted control actions output by the policy network and the corresponding expert control actions.
[0235] In some embodiments, please refer to Figure 5 As shown, multiple agents are cloned and pre-trained using multiple expert sample data, including: S501. From the system state vector and corresponding expert control actions in the expert sample data, filter the local operation data and corresponding expert control actions related to the actuator group corresponding to each agent.
[0236] For example, system state vectors and corresponding expert control action pairs are extracted from expert sample data to construct a sample set for behavior cloning pre-training:
[0237] in, For a moment The system state vector; For each preset controller at time The expert control actions output jointly satisfy: ,in, Indicates the first A preset controller at time The output expert control action, where i takes values from 1 to n. This is the preset total number of controllers.
[0238] For the An intelligent agent, from the system state vector Extract the local operational data related to the actuator group corresponding to the agent to form the local observation state:
[0239] in, For the first An intelligent agent at time Local runtime data; For the first The state extraction function corresponding to each agent is used to filter state variables related to the corresponding actuator group from the global running data.
[0240] It should be understood that the aforementioned state extraction functions are designed according to the control requirements of different actuator groups. For example, the compressor agent extracts state quantities related to refrigerant circulation, such as passenger compartment temperature, battery temperature, suction pressure, and discharge pressure, while the fan agent extracts state quantities related to air-side heat exchange, such as condenser outlet temperature and passenger compartment temperature. This allows each agent to focus only on local information that directly affects its control decisions, reducing the input dimension and improving learning efficiency.
[0241] S502. Based on the local operation data, with the goal of minimizing the deviation between the predicted control action output by the target agent and the corresponding expert control action, perform multiple rounds of cloning pre-training on the target agent.
[0242] The target agent can be any one of the multiple agents.
[0243] For example, the first The policy network of individual agents uses local operational data. As input, output predictive control action:
[0244] in, For the first Policy network of individual agents These are the parameters of the policy network. For policy networks based on local operational data The predicted control actions.
[0245] The goal of behavior cloning pre-training is to make the predicted control action output by the policy network as close as possible to the corresponding expert control action. Therefore, the first... The behavior cloning loss function of an agent Defined as:
[0246] in, Indicates the first A preset controller at time The output of expert-controlled actions, For policy networks based on local operational data The predicted control action, For the sample size, This represents the sum of squares of the deviations between the predicted control action and the expert control action.
[0247] Based on the behavioral cloning loss of each agent, the total behavioral cloning loss function of the multi-agent system is defined as:
[0248] in, For the first The loss weights corresponding to each agent are used to adjust the proportion of the action error of different actuator groups in the total loss.
[0249] During training, the policy network parameters of each agent are updated using gradient descent.
[0250] in, For the first Policy network parameters for each agent; The learning rate for the behavior cloning pre-training phase; This represents the gradient of the behavior cloning loss function with respect to the policy network parameters.
[0251] It should be understood that by continuously reducing the deviation between the predicted control actions output by the policy network and the expert control actions, and by performing behavioral cloning pre-training on the policy networks of each agent, each agent can learn basic control laws similar to the preset controller before formally entering the joint fine-tuning of reinforcement learning. This significantly reduces the ineffective training time and unstable actions caused by subsequent random exploration, and provides a high-quality initial policy for the joint fine-tuning of reinforcement learning.
[0252] As one possible implementation, based on the pre-training of the policy network, expert sample data can be further used to initialize the shared feature extraction layer or value network to improve the sample utilization efficiency, convergence speed and training stability in the subsequent reinforcement learning joint training stage.
[0253] The shared network layer refers to the feature extraction part shared among the policy networks of multiple agents, which is used to extract the common state features of each actuator group from the system state vector; the evaluation network refers to the network used in the joint training phase of reinforcement learning to receive global running data and joint control actions and output the joint action value evaluation value, which is used to guide each agent to update in the direction of maximizing global reward.
[0254] As one possible implementation, the behavior cloning pre-training uses the Adam optimizer, with the total behavior cloning loss function of the above multi-agent as the optimization objective, and is trained according to the hyperparameters shown in Table 3.
[0255] Table 3
[0256] For example, the training and validation sets are divided in an 8:2 ratio. The training set is used to update the policy network parameters, and the validation set is used to monitor overfitting during pre-training. Setting the L2 regularization coefficient and gradient clipping threshold can prevent the policy network parameters from overfitting expert control actions, thereby improving the generalization ability of the policy network after pre-training.
[0257] Training is terminated when one of the following three conditions is met: the validation set loss decreases by less than 1 × 10 for 10 consecutive training epochs. -4 Total behavioral cloning loss Less than 1×10 - ³, The number of training rounds reaches 200. Setting the above termination conditions ensures that pre-training stops promptly after the agent converges, avoiding ineffective training and wasted computational resources.
[0258] Based on S501-S502, this application extracts the local operation data and expert control actions corresponding to each agent from expert sample data, and performs behavioral cloning pre-training on each agent with the goal of minimizing the deviation between the predicted control actions and the expert control actions. This enables each agent to have basic control capabilities consistent with engineering experience before fine-tuning by reinforcement learning, effectively shortening the training time and improving the safety and stability of the initial strategy.
[0259] It should be understood that through the above-mentioned behavior cloning pre-training, the policy network of each agent can learn the control rules of the preset controller under different operating conditions from expert sample data. This enables the policy network to have a set of safe, stable and engineering-feeling basic control capabilities before formally entering the joint fine-tuning of reinforcement learning. This significantly reduces the number of invalid interactions and unsafe actions caused by random exploration in subsequent reinforcement learning training, effectively shortens the training time and improves training stability.
[0260] In summary, by using expert sample data to perform behavioral cloning pre-training on each agent, each agent possesses basic control capabilities similar to the preset controller before formally entering the joint fine-tuning of reinforcement learning. This effectively shortens the subsequent training time and avoids the potential safety hazards that may arise from random exploration starting from scratch.
[0261] The next step is to jointly train multiple agents using online interactive experience sample data.
[0262] As one possible implementation, multiple online interaction experience sample data are acquired; multiple agents are jointly trained using the multiple online interaction experience sample data.
[0263] It should be understood that online interactive experience sample data refers to the state transition samples generated by the agent in the process of online interaction with the simulation environment or real heat pump system during reinforcement learning training. This includes the global operating data at the current moment, the joint control actions output by each agent, the global reward signal obtained after executing the actions, and the global operating data at the next moment.
[0264] In the joint training process, multiple agents are trained based on the same global reward signal, which is used to reflect temperature control error, heat pump system energy consumption and safety constraints.
[0265] It should be understood that using a unified global reward signal to jointly train multiple agents enables each agent to consider not only the local control effect when optimizing its own strategy, but also the impact of the actions of other actuator groups on the overall system performance. This promotes effective cooperation among multiple agents and avoids overall performance degradation caused by each agent acting independently.
[0266] As one possible implementation, during joint training, the training samples include expert sample data and online interaction experience sample data. The expert sample data is used to minimize the deviation between the actions output by each agent and the corresponding expert control actions; and as the training rounds increase, the sampling proportion of expert sample data gradually decreases, while the sampling proportion of online interaction experience sample data gradually increases.
[0267] One implementation method involves using expert sample data derived from high-quality trajectories generated by a pre-defined controller under multiple historical operating conditions. This data carries engineering-safe, stable, and reasonable prior knowledge of control. During training, behavioral cloning loss is used to constrain the predicted actions output by each agent to approximate the expert control actions. This allows each agent to quickly acquire a set of feasible basic policies in the early stages of training, avoiding the inefficient and unsafe actions resulting from random exploration from scratch.
[0268] One implementation method involves using online interactive experience sample data derived from the online interactions between each agent and the simulation environment or a real heat pump system. With the goal of maximizing global rewards, each agent is guided to conduct autonomous exploration and strategy improvement, enabling the agent to break through the performance boundaries of the preset controller and discover better collaborative control strategies.
[0269] One approach involves using expert sample data as the primary training data in the early stages. This allows the policy networks of each agent to be updated closely around the expert's prior knowledge, ensuring stability and security in the early stages of training. As training progresses, the proportion of online interactive experience sample data is gradually increased, enabling each agent to gradually transition from imitating expert behavior to autonomous policy optimization based on environmental feedback.
[0270] It should be understood that by dynamically adjusting the sampling ratio of expert sample data and online interactive experience sample data, a smooth transition from supervised imitation to reinforcement learning is achieved. This not only inherits the stability and interpretability of the preset controller, but also fully leverages the global optimization capability of reinforcement learning. It effectively avoids problems such as policy oscillation, performance degradation, or convergence difficulties during training, and ultimately obtains a collaborative control strategy that balances safety, comfort, and energy efficiency.
[0271] As one possible implementation, multiple agents are jointly trained using multiple online interactive experience sample data, including: extracting local operational data related to the actuator group corresponding to each agent from the online interactive experience sample data; and performing multiple rounds of reinforcement learning training on the target agent based on the target local operational data; wherein the target agent is any one of the multiple agents, and the target local operational data is used to guide the target agent to optimize the policy with the goal of maximizing the global reward.
[0272] As one possible implementation, the joint training employs a multi-agent, dual-delay, deep deterministic policy gradient algorithm.
[0273] The multi-agent dual-delay deep deterministic policy gradient algorithm employs a training architecture of centralized training and distributed execution. Each agent has an independent policy network, and a dual value evaluation network is set up during the training phase to alleviate the problem of overestimation of action value and improve training stability. During the training phase, the value evaluation network calculates the joint action value function based on global runtime data and joint control actions to alleviate the environmental non-stationarity problem in multi-agent training and guide each agent's policy to update towards the system-level global optimum. During the deployment phase, each agent independently outputs control actions based only on locally observable runtime data to meet the real-time requirements of the actual controller.
[0274] One implementation method employs a training logic based on a policy network-value evaluation network for multi-agent reinforcement learning joint training. Each policy network corresponds to an actuator group such as a compressor, electronic expansion valve, fan, and water pump, and is used to output control actions for the actuator group based on its corresponding local operational data. The value evaluation network receives global operational data and the joint control actions of each agent during the training phase, and outputs a joint action value function. The value evaluation network adopts a centralized value evaluation structure, with global operational data as its input. and joint control actions The output is the joint action value function. Among them, joint control actions This can be represented as the intelligent agent corresponding to each actuator group at time... Combination of output control actions:
[0275] in, For the compressor intelligent agent at all times Output control actions, For the electronic expansion valve intelligent agent at any time Output control actions, For wind turbine intelligent agents at all times Output control actions, For the water pump intelligent agent at all times Output control actions.
[0276] For example, the structural parameters of the above-mentioned strategy network and value assessment network are shown in Table 4 below.
[0277] Table 4
[0278] in, Indicates the first The dimension of the local operational data corresponding to each agent. This represents the dimension of the globally running data. This represents the dimension of the joint control action.
[0279] It should be understood that the policy network output layer uses the tanh function to restrict the output to a certain value. Within the specified range, the values are then scaled to the physical action range of the corresponding actuator group using a linear mapping, ensuring that the output control action always remains within the executable range of the actuator group. The value assessment network output layer uses linear output, directly outputting the value estimate of the current joint control action without performing any additional nonlinear transformations.
[0280] For example, the hyperparameters for the joint fine-tuning phase of reinforcement learning are shown in Table 5 below.
[0281] Table 5
[0282] It should be understood that the hyperparameter settings above follow these principles: the policy network learning rate is less than the value evaluation network learning rate to avoid training oscillations caused by excessively rapid policy network updates; the discount factor is set to 0.99 to balance the long-term energy consumption optimization objective of the heat pump system with temperature control accuracy; the soft update rate is set to 0.005, the policy delay update step size is set to 2, the standard deviation of the target policy smoothing noise is set to 0.2, and the target policy noise truncation threshold is set to 0.5 to suppress overestimation of action values and improve training stability; the mini-batch sample size is set to 128, and the experience replay pool capacity is set to 1×10⁻⁶. 7 These parameters are used to balance sample diversity and training efficiency. The parameters can be adjusted based on the number of actuator groups, the dimensionality of local runtime data, reward weights, and simulation complexity.
[0283] It can be understood that the network structure parameters and hyperparameter configurations given in the two tables above are recommended settings that have been experimentally verified. In practical applications, they can be appropriately adjusted according to the specific configuration of the heat pump system, the number of actuator groups, and the control accuracy requirements. For example, for complex heat pump systems with more actuator groups, the network width or depth can be appropriately increased; for application scenarios with higher real-time requirements, the network size can be appropriately reduced.
[0284] In some embodiments, please refer to Figure 6 The agent is trained through multiple rounds of reinforcement learning using online interaction experience sample data, including: S601: Input the global operation data from the online interactive experience sample data into the agent for prediction, and output the predicted control actions of each actuator group.
[0285] As one possible implementation, global operational data from online interactive experience sample data is input into an agent trained on a policy network-value evaluation network, and the predictive control actions of each actuator group are output.
[0286] For example, in the joint training phase of reinforcement learning, let the joint control action output by each agent at time t be:
[0287] in, This represents the control action output by the i-th agent at time t, where n is the number of agents.
[0288] Based on S601, each agent extracts its corresponding local operation data from the global operation data of online interaction experience sample data, and independently outputs predictive control actions through the policy network to form the joint control action at the current moment, realizing distributed decision-making of multiple actuator groups.
[0289] S602. Control the operation of the heat pump system based on the predictive control action corresponding to each actuator group, and obtain the actual temperature of the passenger compartment and the actual temperature of the power battery after the operation.
[0290] As one possible implementation, the operation of the heat pump system is controlled based on the predictive control actions corresponding to each actuator group, including: after the predictive control actions output by each agent are processed by physical range mapping and rate limiting, they are sent to the corresponding actuator group for execution through the control bus to drive the operation of the heat pump system.
[0291] As one possible implementation, obtaining the actual temperature of the passenger compartment and the actual temperature of the power battery after execution includes: real-time acquisition of the passenger compartment temperature and the power battery temperature by temperature sensors arranged in the heat pump system, and feeding them back to the intelligent agent after preprocessing.
[0292] Based on S602, by applying predictive control actions to the heat pump system and collecting feedback on the actual temperature, a closed-loop interaction between the agent and the environment is realized, providing real feedback signals for subsequent reward calculation and policy updates.
[0293] S603 determines the temperature deviation reward based on the set temperature of the passenger compartment and the actual temperature of the passenger compartment, and the target temperature of the power battery and the actual temperature of the power battery, and determines the total system power consumption penalty of the heat pump system based on the predictive control actions of each actuator group.
[0294] As one possible implementation, the global reward signal includes at least a temperature control reward, an energy consumption penalty, and a safety constraint penalty.
[0295] Among them, the temperature control reward item is used to reflect the deviation between the passenger compartment temperature and the set value, as well as the deviation between the power battery temperature and the target temperature; the energy consumption penalty item is used to reflect the total power consumption of the compressor, fan and water pump; and the safety constraint penalty item is used to reflect the situation where the battery temperature, refrigerant pressure or actuator action range exceeds the safety boundary.
[0296] One implementation method includes a temperature deviation reward consisting of a passenger compartment temperature deviation reward and a battery temperature deviation reward. The passenger compartment temperature deviation reward is determined based on the deviation between the actual temperature of the passenger compartment and the set temperature of the passenger compartment, while the battery temperature deviation reward is constructed based on the deviation between the actual temperature of the power battery and the target temperature of the power battery.
[0297] For example, temperature control reward items The following relationship must be satisfied:
[0298]
[0299] in, The controlled object is either the battery or the crew compartment; For a moment The absolute value of the deviation between the actual temperature and the target temperature; This is the temperature error penalty coefficient. For example, a reward coefficient for achieving the temperature target. You can take 0.5. A value of 10 can be chosen. The above piecewise reward function penalizes the agent when the temperature error is large, and rewards it positively when the temperature error falls within the allowable range, thereby improving temperature tracking accuracy and accelerating policy convergence.
[0300] One implementation method, wherein the energy consumption penalty term is determined based on the predictive control actions of each actuator group, includes: calculating the corresponding power consumption value according to the predictive control actions of each actuator group, summing the power consumption values of each actuator group to obtain the total power consumption of the system, and then normalizing the total power consumption of the system to obtain the energy consumption penalty term.
[0301] For example, the energy consumption penalty term can be determined based on normalized system energy consumption. Normalized system energy consumption The following relationship must be satisfied:
[0302] in, For actuator assemblies such as compressors, fans, and water pumps at all times The total energy consumption The maximum energy consumption is preset for the system. The total energy consumption can be calculated by integrating the power consumption of each actuator group:
[0303] in, This indicates the total cumulative energy consumption of the heat pump system during its operation. Indicates the total running time; Indicates time Real-time power consumption of the compressor actuator assembly; Indicates time Real-time power consumption of the fan actuator assembly; Indicates time Real-time power consumption of the pump actuator assembly; This represents a time infinitesimal element.
[0304] The safety constraint penalty term is used to reflect situations where the battery temperature, refrigerant pressure, or actuator operating range exceeds the preset safety boundary. When a safety constraint is violated, the safety constraint penalty term takes a positive value to reduce the value of the global reward signal.
[0305] The real-time power consumption of each actuator group can be calculated based on the control actions output by the corresponding intelligent agent (such as compressor speed, fan speed, pump duty cycle, etc.) combined with the power consumption characteristic curve or power consumption model of the actuator group. In actual engineering implementation, since the control actions are usually discretely sampled, the above continuous integral can be approximated by a numerical integration method that sums the discrete power consumption sequence to ensure computability in discrete control systems.
[0306] It should be understood that by normalizing the total system power consumption to The interval ensures that the dimensions of the energy consumption reward item are consistent with those of other reward items, making it easier to perform weighted combination on the same scale.
[0307] S604. Determine the global reward signal based on temperature control reward items, energy consumption penalty items, and safety constraint items.
[0308] As one possible implementation, a global reward signal is generated based on a temperature control reward term and an energy consumption penalty term. The following method can be used to determine the optimal system energy consumption: weighted combination of passenger cabin temperature deviation reward, battery temperature deviation reward, and normalized system energy consumption, so that each intelligent agent takes into account passenger cabin comfort, power battery thermal safety, and system energy consumption during the optimization process.
[0309] It should be understood that the design of the global reward signal directly determines the optimization direction of each agent. The above reward structure places temperature control accuracy in the positive part of the reward item and system energy consumption in the penalty item, so that each agent can reduce energy consumption as much as possible while pursuing precise temperature control, thereby achieving synergistic optimization of comfort and energy saving.
[0310] As another possible implementation, the global reward signal is determined based on temperature control reward and energy consumption penalty, and further includes: acquiring the change in action between the predicted control actions output by each actuator group at adjacent time points, and acquiring the safety constraint violation during the operation of the heat pump system; determining the action smoothness penalty based on the change in action; determining the safety constraint penalty based on the safety constraint violation; and determining the global reward signal based on the temperature control reward, energy consumption penalty, action smoothness penalty, and safety constraint penalty.
[0311] For example, the global reward signal It can be determined in the following ways:
[0312] in, This is an award for crew cabin temperature control. For battery temperature control, To normalize the system energy consumption, As a penalty for smoothness of movement, Violation of safety regulations will result in penalties. and The weights for the crew cabin temperature control bonus and the battery temperature control bonus are respectively. Used to adjust the intensity of energy consumption penalties; Constraint strength used to adjust the smoothness of motion; Used to increase the severity of penalties when safety constraints are violated. Used to characterize whether battery temperature, refrigerant pressure, actuator operating range, or other critical states exceed preset safety boundaries; when the system does not violate safety constraints... Set to 0 when a safety constraint is violated. Take the positive value.
[0313] One implementation method is a motion smoothness penalty term. The following relationship must be satisfied:
[0314] in, For the first An intelligent agent at time Output control actions, For the first The control action output by the intelligent agent in the previous moment. The number of agents.
[0315] It should be understood that by introducing a motion smoothness penalty term into the global reward signal, drastic changes in control actions at adjacent time points can be suppressed, reducing mechanical wear of the actuator assembly and system oscillations; by introducing a safety constraint penalty term and setting a high penalty coefficient... This enables each intelligent agent to proactively avoid dangerous conditions such as battery temperature exceeding limits and refrigerant pressure exceeding standards during the exploration process, ensuring the safety of the training process and the final control strategy.
[0316] S605. Train each agent in the direction of maximizing the global reward signal.
[0317] As one possible implementation, each agent is trained in the direction of maximizing the global reward signal, including: calculating the joint action value based on a dual value evaluation network, and alternately updating the parameters of the value evaluation network and each policy network by minimizing the loss of the value evaluation network and maximizing the joint action value output by each policy network, so that each agent gradually learns the optimal cooperative control strategy that can obtain the maximum cumulative global reward signal.
[0318] For example, the dual-value assessment network is based on global operational data. With joint control actions Calculate the joint action value function. The outputs of the two value assessment networks are respectively... and To prevent overestimation of action value, the joint action value is taken as the smaller of the outputs of the two value assessment networks.
[0319] The first step is to update the value assessment network.
[0320] (1) Let the target joint control action at the next moment be: .
[0321] in, This indicates the target joint control action used to calculate the target value at the next moment; The target joint control action output by the target policy network in the next time step is the target policy network, which is the target network of the current policy network and is used for the update process of stable value estimation. To smooth the noise for the target policy, it is usually random noise that follows a truncated normal distribution. By adding a small amount of noise to the target action, the value evaluation network is prevented from overfitting to local peaks of the policy, thus improving training robustness.
[0322] (2) The Bellman objective value is defined as follows: .
[0323] in, Indicates time The Bellman objective value is the regression objective updated by the value assessment network; For a moment The global reward obtained after performing a joint control action; This is the discount factor, and its value range is... This is used to adjust the degree to which future rewards affect the current value estimate. The closer the value is to 1, the more emphasis is placed on long-term returns; and These are two target value assessment networks, one of which is the target network of the current value assessment network. Its parameters are slowly updated from the current network parameters through a soft update method. This is the global operational data for the next moment after the joint control action is executed; This means taking the smaller value from the outputs of the two target value assessment networks to suppress overestimation of action value.
[0324] (3) Based on the above Bellman objective value, the loss functions of the two value assessment networks are defined as follows: , .
[0325] in, and These are the mean squared error loss functions for the two value assessment networks, respectively. and These are the two value assessment networks for the joint control action at the current moment. The estimated value; The target value is the Bellman target value; the loss function minimizes the deviation between the estimated value and the target value, allowing the value assessment network to gradually approach the true joint action value function.
[0326] (4) The parameters of the two value assessment networks are updated according to the gradient descent method as follows: , .in, and These are the parameters of the two value assessment networks; The learning rate is used to evaluate the value of the network and to control the step size of parameter updates; and These are the gradients of the two loss functions with respect to their respective network parameters.
[0327] The second step is to update the policy network.
[0328] (1) The policy network aims to maximize the value of joint actions, and its loss function is defined as: .
[0329] in, Let be the loss function of the policy network. Taking a negative value makes the optimization direction maximize the value of the joint action. The first value assessment network estimates the value of the joint control action output by the policy network. This represents the joint control action output by each policy network under the corresponding local operating data; For a moment Global runtime data.
[0330] (2) The formula for updating the policy network parameters is as follows: .
[0331] in, These are the parameters of the policy network; The learning rate of the policy network is used to control the step size for updating the policy network parameters. This represents the gradient of the policy network loss function with respect to the policy network parameters.
[0332] The third step is to update the target network.
[0333] As one possible implementation, a delayed policy update mechanism is adopted, that is, the policy network updates the policy every interval. Each value assessment network update step performs a parameter update once; and the target network parameters are updated using a soft update method. , , .
[0334] in, , These are the parameters of two target value assessment networks; These are the parameters of the target policy network; This is a soft update parameter, and its value range is... , The smaller the target network parameter, the slower the update, which is beneficial to the stability of the training process; , These are the parameters of the current value assessment network; These are the parameters of the current policy network.
[0335] It should be understood that constructing the Bellman target value by using a dual-value evaluation network and taking the smaller value can effectively suppress the problem of overestimating action value that is prone to occur in a single-value evaluation network; adopting a delayed policy update mechanism, which allows the value evaluation network to be updated multiple times before the policy network is updated, is beneficial to the accuracy of value estimation; and using a soft update method to update the target network parameters, which allows the target network parameters to slowly track the current network parameters, is beneficial to the stability of the training process.
[0336] As one possible implementation, a hybrid experience replay mechanism combining the expert sample data and the online interactive experience sample data is employed during the reinforcement learning fine-tuning phase, with the expert-guided loss weight gradually decaying with training iterations. By maintaining a high level of constraint from the expert sample data in the early stages of training, deviations of each policy network from the engineering-feasible control region can be effectively prevented; gradually reducing the strength of the expert sample data constraint in the later stages of training helps each agent break through the local performance boundaries of the preset controller, further obtaining a better system-level cooperative control strategy.
[0337] In one implementation, during the joint fine-tuning phase of reinforcement learning, the parameters of each policy network and the value evaluation network are updated based on global runtime data, joint control actions, and the global reward signal. The training termination condition can be set according to actual needs; for example, the steady-state evaluation window length is set to... The steady-state error of the passenger compartment and power battery is defined as follows:
[0338]
[0339] in, This indicates the average temperature error of the crew cabin within the steady-state assessment window; This indicates the average temperature error of the power battery within the steady-state evaluation window; The steady-state evaluation window length is used to extract a continuous period of time in the later stages of training for steady-state performance evaluation. The total runtime of the current training round; For a moment The actual temperature of the crew cabin; Desired cabin temperature; For a moment The actual temperature of the power battery; The desired battery temperature; It is a time infinitesimal element.
[0340] Training terminates when both errors are less than a preset threshold or when the maximum number of training epochs is reached. For example, let the steady-state evaluation window length be... The value is taken as 300s, when the steady-state error of the crew cabin and battery steady-state error Training should be terminated when the temperature is below 0.5℃ or the maximum number of training rounds reaches 2000.
[0341] Based on S601-S605, this application inputs the online interactive experience sample data into the policy network of each agent to output predictive control actions and apply them to the heat pump system. Based on the real-time feedback of temperature deviation and system power consumption, a global reward signal that takes into account comfort, energy saving, smooth action and safety is constructed. The parameters of each policy network and the value evaluation network are updated in the direction of maximizing the global reward signal by using a dual value evaluation network and a delayed policy update mechanism. Under the guidance of the online interactive experience sample data, each agent gradually learns a system-level collaborative optimization strategy that breaks through the preset controller performance boundary.
[0342] As one possible implementation, the method of this application further includes: performing preprocessing operations on the global operating data of the heat pump system to extract local operating data from the preprocessed global operating data; wherein, the preprocessing operations include at least one of the following: outlier detection and correction operations, low-pass filtering operations, and normalization processing.
[0343] As one possible implementation, after acquiring the raw data in the sample database, the global operating data in the expert sample data and the global operating data in the online interactive experience sample data are preprocessed in the order of outlier detection and correction, low-pass filtering, rate of change calculation and normalization to obtain the enhanced state representation for each agent's input.
[0344] One implementation involves dividing key variables in the collected global operational data into multiple sub-dimensional vectors based on their physical attributes and control objectives. These include: an environment and demand dimension vector, comprising ambient temperature, cabin set temperature, operating mode, and the operating status of the heat pump system, used to define the system's external boundary conditions and user comfort requirements; a thermal safety and heat load dimension vector, comprising actual cabin temperature, power battery temperature, drive motor temperature, battery coolant temperature, and battery state of charge information, used to characterize the current thermal state and heat load change trends of the controlled object; a thermodynamic cycle dimension vector, comprising compressor speed, suction pressure, discharge pressure, evaporator inlet and outlet refrigerant temperatures, condenser inlet and outlet refrigerant temperatures, coolant flow rate, and refrigerant superheat or subcooling, used to reflect the internal thermodynamic cycle process and energy efficiency characteristics of the heat pump system; and an actuator status and historical feedback dimension vector, comprising the current control actions of the fan, water pump, and electronic expansion valve, the control actions at the previous moment, and the rate of change of key states. All of these dimension vectors together constitute a time-series data. The original global runtime data.
[0345] It should be understood that dividing the raw global operational data into multiple sub-dimensional vectors according to physical attributes and control objectives is beneficial for adopting differentiated preprocessing strategies for different types of state variables. It also makes it easier for each agent to extract the corresponding local operational data from the global operational data according to its own control requirements.
[0346] As one possible implementation, the data in the sample database will be preprocessed from the perspectives of outlier detection and correction, low-pass filtering, rate of change calculation, and normalization.
[0347] Step 1: Outlier detection and correction.
[0348] Let any original sampled state quantity be The sampling period is For the original sampled state quantity, outlier determination is first performed based on the physical range and the rate of change of adjacent samples. The determination is based on the following criteria: when the state quantity exceeds the allowable physical range, or the rate of change between adjacent sample points exceeds the corresponding threshold, the sampled value is determined to be an outlier.
[0349] outlier corrected state variables The following relationship must be satisfied:
[0350] in, For a moment State variables after outlier correction; For a moment The original sampled state quantity; This is the state variable after outlier correction from the previous time step; and These are the minimum and maximum values allowed for the state quantity, i.e., the lower and upper limits of the physical range; The maximum allowable rate of change for the state variable; The sampling period; For a symbolic function, it is defined as follows: when hour, When it is 1, hour, When it is 0, hour, for .
[0351] When three consecutive sampling cycles are determined to be abnormal values, the sensor corresponding to the state quantity is in a fault state, and the agent is triggered to switch to the control mode corresponding to the preset controller.
[0352] It should be understood that the above-mentioned outlier-corrected state variables The first branch indicates that if the sampled value is within the allowable physical range and the rate of change of adjacent samples does not exceed the threshold, the original sampled value is retained. The second branch indicates that if the sampled value exceeds the physical range, the corrected value from the previous moment is used instead. The third branch indicates that if the sampled value is within the physical range but the rate of change exceeds the limit, the change amplitude is limited to the maximum allowable rate of change. (Sign function) Used to maintain the original direction of change.
[0353] Step 2: Low-pass filtering operation.
[0354] After the outlier correction is completed, the continuous state variables are filtered using a first-order discrete low-pass filter. The filtered state variables are... The following relationship must be satisfied:
[0355] in, For a moment State variables after low-pass filtering; For a moment State variables after outlier correction; This is the state quantity after low-pass filtering at the previous moment; The weight of the current sample value is calculated using the following formula:
[0356] in, The filtering time constant is The sampling period. The larger, The smaller the value, the slower the filtered state variable responds to changes in the original sampled value, and the stronger the smoothing effect.
[0357] It should be noted that continuous feedback quantities such as temperature, pressure, flow rate, compressor speed, fan speed, water pump speed, and electronic expansion valve opening need to be low-pass filtered; however, the desired cabin temperature, operating mode, previous control action, and discrete status flags do not need to be low-pass filtered.
[0358] For example, the configuration of the range, normalization interval and filtering time constant of some key state quantities can be seen in Table 6 below.
[0359] Table 6
[0360] It should be understood that "—" in the table indicates that the state variable is not subject to low-pass filtering or maximum rate of change limitation. The above parameter configurations are for illustrative purposes only and can be adjusted according to sensor accuracy, system dynamic characteristics, and control cycle in actual applications.
[0361] Step 3: Calculate the rate of change.
[0362] The rate of change characteristics of the filtered state variables are further calculated using a sliding window difference method, with the general formula being:
[0363] in, For a moment The characteristics of the rate of change of state; For a moment Filtered state variables; For a moment Filtered state variables; The length of the sliding window; The sampling period.
[0364] For temperature variables, sliding window length Set to 5; for the pressure variable, the sliding window length is... Take 2; for actuator feedback, the sliding window length Set to 1. When the sampling period... At that time, the rate of change of temperature corresponds to a 5-second time window, the rate of change of pressure corresponds to a 2-second time window, and the rate of change of actuator feedback corresponds to a 1-second time window.
[0365] Therefore, the temperature gradient of the occupant compartment, the rate of temperature rise of the power battery, the rate of change of exhaust pressure, and the rate of change of intake pressure satisfy the following relationships: ; ; ; .
[0366] in, For a moment The temperature gradient of the crew cabin; For a moment The actual temperature of the crew cabin after filtering; For a moment The actual temperature of the crew cabin after filtering; For a moment The rate of temperature rise of the power battery; For a moment Temperature of the power battery after filtering; For a moment Temperature of the power battery after filtering; For a moment The rate of change of exhaust pressure; For a moment Filtered exhaust pressure; For a moment Filtered exhaust pressure; For a moment The rate of change of inspiratory pressure; For a moment Filtered intake pressure; For a moment The intake pressure after filtering.
[0367] It should be understood that extracting rate of change features helps the agent perceive the dynamic changing trends of key states such as temperature and pressure in the heat pump system, enhances the ability to model the system's large inertia, strong coupling, and hysteresis characteristics, thereby improving the foresight and response speed of the control strategy.
[0368] Step 4: Normalization.
[0369] One implementation involves normalizing state variables with clearly defined upper and lower bounds and non-negative values to the following method. Interval:
[0370] in, For a moment Normalized state variables; For a moment Filtered state variables; and These are the minimum and maximum values of the state variables, respectively. For state variables that do not undergo low-pass filtering, their filtered state variables are set to satisfy... .
[0371] For state variables such as error quantities and rates of change that require retaining positive and negative direction information, the following method is used for normalization to... Interval:
[0372] in, For a moment Normalized error or rate of change characteristics; For a moment The original error or rate of change characteristics to be normalized; and These represent the lower and upper limits of the value range for the corresponding error amount or rate of change characteristic, respectively.
[0373] For example, the range of values and normalization intervals for the error quantity and rate of change characteristics can be found in Table 7 below.
[0374] Table 7
[0375] It should be understood that normalizing state variables and rate of change features can unify physical quantities of different dimensions and magnitudes to the same scale, which helps eliminate dimensional differences between features and improves the training stability and convergence speed of subsequent policy networks and value evaluation networks. For non-negative state variables, normalization is applied... Normalization is applied to the error quantity and rate of change that contain information about positive and negative directions. Normalization can preserve the directional information of state deviation, making it easier for the agent to make directionally correct control decisions.
[0376] Through the sequential preprocessing steps of outlier detection and correction, low-pass filtering, rate of change calculation, and normalization, the system state vector in the expert sample data and the online interactive experience sample data are ultimately composed of the corrected and filtered state variables and the normalized rate of change features, forming an enhanced state representation. This provides high-quality, standardized input data for the subsequent training of each agent.
[0377] In summary, please refer to Figure 7 The present application provides a complete flowchart of an agent training method, which includes: collecting preset trajectories by each preset controller under multiple historical working conditions; constructing expert sample data after quality scoring and screening of the collected preset trajectories; subsequently, using the expert sample data to perform behavioral cloning pre-training on the policy network of each agent; and then conducting joint fine-tuning of multi-agent reinforcement learning by combining the expert sample data with online interaction experience sample data, and finally completing the policy deployment of each agent.
[0378] Please see Figure 8 A diagram illustrating the reinforcement learning training logic, including: the environment based on current global runtime data. Joint control actions output by the policy network of each agent Perform a state transition and return an immediate reward. and global runtime data at the next moment To form an interactive experience .
[0379] The interactive experience is stored in the experience replay pool. During training, samples are randomly taken from the experience replay pool. Each experience point forms a sampling batch, which is then input into the policy network and value assessment network modules for parameter updates.
[0380] For the policy network, the policy network optimizer is based on the deterministic policy gradient formula. Calculate the gradient and update the policy network parameters along the direction that maximizes the value of the joint action. After the update, the policy network adjusts its parameters according to the state at the next time step. Output new control actions Subsequently, the parameters of the policy network are slowly synchronized to the target policy network using a soft update method. The target policy network then updates according to its state at the next time step. Output the target motion and add smooth noise. ,get This is used for the subsequent calculation of the Bellman objective value.
[0381] For value assessment networks, the value assessment network optimizer That is, based on Bellman objective value Compared with the current value assessment network output Mean square error between The gradient is calculated, and the parameters of the value assessment network are updated in the direction that minimizes the error. After the outputs of the two value assessment networks are calculated using the Bellman objective, the parameters of the two value assessment networks are slowly synchronized to the corresponding two target value assessment networks through a soft update method.
[0382] As one possible implementation method, the dynamic adjustment process of the sample ratio between expert sample data and online interactive experience sample data in multi-round reinforcement learning training specifically includes: In the initial training phase, multiple agents are trained using only expert sample data, meaning the proportion of expert sample data is 100%, and the proportion of online interaction experience sample data is 0%. As the training rounds increase, the proportion of expert sample data gradually decreases from 100%, while the proportion of online interaction experience sample data gradually increases from 0%, until the proportion of expert sample data decreases to a preset minimum or the proportion of online interaction experience sample data increases to a preset maximum.
[0383] It should be understood that the initial training phase relies entirely on expert sample data, allowing the policy networks of each agent to fully learn the safe and stable control patterns of the preset controller under various historical operating conditions through behavioral cloning, establishing a reliable initial policy foundation. If online interactive experience sample data is introduced for online exploration at this stage, the actions generated may deviate from the safety boundaries because the policy networks do not yet possess basic control capabilities, leading to system instability or even triggering safety protection. As training rounds increase, each policy network masters the basic control logic, gradually increasing the proportion of online interactive experience sample data. This allows each agent to explore better control strategies through online interaction with the environment, breaking through the performance boundaries of the preset controller and ultimately achieving global collaborative optimization. This phased, smooth transition mechanism from "complete imitation" to "autonomous optimization" effectively avoids the low training efficiency and safety risks caused by random exploration in the early stages of training.
[0384] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the heat pump system control device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0385] This application embodiment can, according to the above method, exemplarily divide a heat pump system control device or electronic device into functional modules. For example, the heat pump system control device or electronic device may include functional modules corresponding to each functional division, or two or more functions may be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; in actual implementation, there may be other division methods.
[0386] Please see Figure 9This application provides a heat pump system control device, which includes: a data acquisition module 901, a local extraction module 902, an action execution module 903, and a system operation module 904. The data acquisition module 901 is used to acquire the current global operating data of the heat pump system. The global operating data includes: the current control state of each actuator in the heat pump system, the current actual temperature and set temperature of the area controlled by the heat pump system, and the operating energy consumption of the heat pump system. The local extraction module 902 is used to extract local operating data related to each actuator group from the current global operating data. The local operating data includes at least: the current temperature error of the controlled area affected by the corresponding actuator group, and the current control state of the actuator group. The local extraction module 904... 02 is also used to determine the input feature vector of the intelligent agent associated with the actuator group based on the local operating data related to the actuator group and the operating energy consumption of the heat pump system; wherein, each actuator group corresponds to one intelligent agent; the intelligent agent is used to: determine the control action to be performed by the corresponding actuator group with the goal of reducing the current temperature error of the controlled area and the operating energy consumption of the heat pump system; the input feature vector is used to characterize: the current operating state of the local area corresponding to the intelligent agent and the operating energy consumption of the heat pump system; the execution action module 903 is used to input the corresponding input feature vector into the intelligent agent corresponding to the actuator group and output the target control action of the actuator group; the system operation module 904 is used to control the operation of the heat pump system based on the target control action of each actuator group in the multiple actuator groups.
[0387] like Figure 10 As shown, the electronic device 1000 provided in this application includes, but is not limited to, a processor 1001 and a memory 1002.
[0388] The memory 1002 described above is used to store the executable instructions of the processor 1001. It is understood that the processor 1001 is configured to execute instructions to implement the methods in the above embodiments.
[0389] It should be noted that those skilled in the art will understand that Figure 10 The electronic device structure shown does not constitute a limitation on electronic device 1000; electronic devices may include, but are not limited to, other types of electronic devices. Figure 10 This may indicate more or fewer components, or combinations of certain components, or different component arrangements.
[0390] The processor 1001 is the control center of the electronic device 1000. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1002, and by calling data stored in the memory 1002, it performs various functions and processes data of the electronic device 1000, thereby providing overall monitoring of the electronic device 1000. The processor 1001 may include one or more processing units. Optionally, the processor 1001 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1001.
[0391] The memory 1002 can be used to store software programs and various data. The memory 1002 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, application programs required by at least one functional module (such as a determination unit, processing unit, etc.), etc. Furthermore, the memory 1002 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0392] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 1002 including instructions, which can be executed by a processor 1001 of an electronic device 1000 to implement the methods in the above embodiments.
[0393] Optionally, the computer-readable storage medium may be a non-transitory computer-readable storage medium, such as a read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.
[0394] In an exemplary embodiment, this application also provides a computer program product including one or more instructions, which can be executed by the processor 1001 of the electronic device 1000 to perform the methods described above.
[0395] It should be noted that when one or more instructions in the computer-readable storage medium or computer program product are executed by the processor of the electronic device, they implement the various processes of the above method embodiments and achieve the same technical effect as the above method. To avoid repetition, they will not be described again here.
[0396] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0397] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0398] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0399] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0400] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to related technologies, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0401] The above embodiments are merely preferred embodiments provided to fully illustrate this application, and the scope of protection of this application is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on this application are all within the scope of protection of this application.
Claims
1. A control method for a heat pump system, characterized in that, The heat pump system includes multiple actuator groups, each actuator group being used to regulate the temperature of the affected controlled area, and the method includes: Obtain the current global operating data of the heat pump system; the current global operating data includes: the current control status of each actuator in the heat pump system, the current actual temperature and set temperature of the area controlled by the heat pump system, and the operating energy consumption of the heat pump system; From the current global operating data, extract the local operating data related to each actuator group; the local operating data includes at least: the current temperature error of the controlled area affected by the corresponding actuator group, and the current control state quantity of the actuator group; Based on the local operating data related to the actuator group and the operating energy consumption of the heat pump system, the input feature vector of the intelligent agent associated with the actuator group is determined; wherein, each actuator group corresponds to one intelligent agent; the intelligent agent is used to: determine the control action to be performed by the corresponding actuator group with the goal of reducing the current temperature error of the controlled area and the operating energy consumption of the heat pump system; the input feature vector is used to characterize: the current operating state of the local area corresponding to the intelligent agent and the operating energy consumption of the heat pump system; The corresponding input feature vector is input into the agent corresponding to the actuator group, and the target control action of the actuator group is output. The heat pump system is controlled to operate based on the target control action of each of the plurality of actuator groups.
2. The control method according to claim 1, characterized in that, The intelligent agent is also used to: estimate the estimated operating data after the actuator group executes the control action based on the determined control action to be performed by the actuator group; the estimated operating data is used to compare with the actual local operating data at the next time step to calculate the state prediction deviation in the input feature vector at the next time step. The step of determining the input feature vector of the agent associated with the actuator group based on the local operating data related to the actuator group and the operating energy consumption of the heat pump system further includes: Based on the local operating data related to the actuator group, the operating energy consumption of the heat pump system, and the estimated operating data output by the agent associated with the actuator group at the previous moment, the input feature vector of the agent associated with the actuator group is determined so that the agent outputs the target control action of the actuator group and the estimated operating data; the input feature vector also includes: the state prediction deviation of the agent for the local area at the previous moment.
3. The control method according to claim 1, characterized in that, The plurality of actuator groups include: The compressor actuator assembly is used to drive the refrigerant circulation and regulate the heating or cooling capacity of the heat pump system; Electronic expansion valve actuator assembly is used to adjust the refrigerant throttling opening and control the refrigerant flow on the evaporator or condenser side; The fan actuator assembly is used to adjust the heat exchange intensity on the air side, which affects the temperature response of the passenger compartment. The water pump actuator assembly is used to drive the circulation of coolant and regulate the temperature of the power battery or drive motor.
4. The method according to claim 1, characterized in that, When the actuator group is a compressor actuator group, the current temperature error includes the crew cabin temperature error and the battery temperature error, and the current control state quantity includes the current compressor speed; When the actuator group is an electronic expansion valve actuator group, the current temperature error includes the suction superheat error and the evaporator outlet temperature error; When the actuator group is a fan actuator group, the current temperature error includes the condenser heat exchange status and the crew cabin temperature error; When the actuator group is a water pump actuator group, the current temperature error includes battery temperature error and coolant temperature error.
5. The method according to claim 1, characterized in that, The method further includes: Acquire multiple expert sample data; wherein, the expert sample data includes system state vectors and corresponding expert control actions; By using multiple expert sample data, behavioral cloning pre-training is performed on multiple agents corresponding to the multiple actuator groups to minimize the deviation between the actions output by each agent and the corresponding expert control actions.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Obtain sample data of multiple online interactive experiences; The multiple agents are jointly trained using multiple online interactive experience sample data; wherein, during the joint training process, the multiple agents are trained based on the same global reward signal, which is used to reflect temperature control error, heat pump system energy consumption and safety constraints.
7. The method according to claim 6, characterized in that, During the joint training process, the training samples include expert sample data and online interaction experience sample data; the expert sample data is used to minimize the deviation between the actions output by each agent and the corresponding expert control actions. Furthermore, as the number of training rounds increases, the sampling proportion of expert sample data gradually decreases, while the sampling proportion of online interactive experience sample data gradually increases.
8. The method according to claim 6, characterized in that, The global reward signal includes at least a temperature control reward item, an energy consumption penalty item, and a safety constraint penalty item; the temperature control reward item is used to reflect the deviation between the passenger compartment temperature and the set value, as well as the deviation between the power battery temperature and the target temperature; the energy consumption penalty item is used to reflect the total power consumption of the compressor, fan, and water pump; and the safety constraint penalty item is used to reflect situations where the battery temperature, refrigerant pressure, or actuator operating range exceeds the safety boundary.
9. The method according to claim 5, characterized in that, The acquisition of multiple expert sample data includes: Under multiple historical operating conditions, each preset controller controls the operation of the corresponding actuator group to generate multiple preset trajectories; wherein, each intelligent agent corresponds to a preset controller, and the preset controller controls one actuator group in the heat pump system. The preset controller is used to generate control actions for the corresponding actuator group based on preset control logic and according to the deviation between the current state data and the target state data of the heat pump system under the multiple historical operating conditions. Based on the parameter set of each preset trajectory, a quality score is given for multiple preset trajectories; the parameter set includes at least one of the following: temperature overshoot, steady-state error, total system power consumption, smoothness of action, and number of safety constraint violations; The expert sample data is determined based on the multiple preset trajectories whose quality scores are less than a preset score threshold or meet preset rejection conditions.
10. The method according to claim 9, characterized in that, The preset rejection conditions include: The number of times the security constraint is violated exceeds a preset threshold. The actuator saturation ratio exceeds the preset saturation threshold; The absolute value of the difference between the actual temperature and the expected temperature of the passenger cabin exceeds the preset temperature deviation threshold for a number of cycles, reaching the preset continuous cycle threshold. The absolute value of the difference between the actual battery temperature and the expected battery temperature exceeds the preset temperature deviation threshold for a number of cycles, reaching the preset continuous cycle threshold.
11. A control device for a heat pump system, characterized in that, The control device of the heat pump system includes: a data acquisition module, a local extraction module, an action execution module, and a system operation module; The data acquisition module is used to acquire the current global operating data of the heat pump system; the current global operating data includes: the current control status of each actuator in the heat pump system, the current actual temperature and set temperature of the area controlled by the heat pump system, and the operating energy consumption of the heat pump system; The local extraction module is used to extract local operating data related to each actuator group from the current global operating data; the local operating data includes at least: the current temperature error of the controlled area affected by the corresponding actuator group, and the current control state quantity of the actuator group; The local extraction module is further configured to determine the input feature vector of the agent associated with the actuator group based on the local operating data related to the actuator group; wherein each actuator group corresponds to one agent; the agent is configured to: determine the control action to be performed by the corresponding actuator group with the goal of reducing the current temperature error of the controlled area and the operating energy consumption of the heat pump system; the input feature vector is used to characterize: the current operating state of the local area corresponding to the agent, and the operating energy consumption of the heat pump system; The execution action module is used to input the corresponding input feature vector into the agent corresponding to the executor group and output the target control action of the executor group. The system operation module is used to control the operation of the heat pump system based on the target control action of each of the plurality of actuator groups.
12. A vehicle, characterized in that, The vehicle includes a controller for implementing the control method of the heat pump system as described in any one of claims 1 to 10.