Vehicle thermal management system control method and device, vehicle, medium and program product
By using driving data to dynamically generate candidate control strategies in the vehicle thermal management system, the problem of poor adjustment effect under fixed threshold control is solved, and more accurate thermal management and shorter development cycles are achieved.
Patent Information
- Application Number
- CN202510655534.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-27
AI Technical Summary
The existing vehicle thermal management system is controlled by setting a fixed threshold, resulting in poor adjustment effect under various operating conditions.
By dynamically generating multiple candidate control strategies based on vehicle driving data, an optimal target control strategy is selected to control the thermal management system, including the target opening degree of the active air intake grille and the target speed of the cooling fan.
More precise thermal management is achieved, the adjustment effect is improved, and the development cycle is reduced without the need for calibration thresholds.
Smart Images

Figure CN120207052A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of vehicles, and particularly to a control method, device, vehicle, medium and program product for a thermal management system of a vehicle. Background Art
[0002] The thermal management system of a vehicle is a technical system that controls the temperature to ensure that the vehicle's electronic control system, battery, motor, cockpit, etc. operate under optimal conditions. In the related art, a fixed threshold is set for the thermal management system, and the start and stop of the thermal management system are controlled according to the fixed threshold to control the temperature. Using a fixed threshold for control may result in poor adjustment effects of the thermal management system under various vehicle conditions. Summary of the Invention
[0003] To overcome the problems existing in the related art, the present disclosure provides a control method, device, vehicle, medium and program product for a thermal management system of a vehicle.
[0004] According to the first aspect of the embodiments of the present disclosure, a control method for a thermal management system of a vehicle is provided. The method includes: obtaining a target control strategy for controlling the thermal management system according to the driving data of the vehicle, where the target control strategy is one of multiple candidate control strategies obtained according to the driving data.
[0005] Dynamically generating multiple candidate control strategies, selecting a target control strategy from the multiple candidate control strategies, and using the target control strategy to control the thermal management system, which performs thermal management more precisely and improves the adjustment effect. And compared with the related art, there is no need to calibrate the threshold, shortening the development cycle.
[0006] In some possible implementation manners, the target control strategy includes the target opening degree of the active intake grille of the thermal management system and / or the target rotation speed of the cooling fan of the thermal management system.
[0007] In some possible implementation manners, the obtaining a target control strategy for controlling the thermal management system according to the driving data of the vehicle includes: obtaining multiple candidate control strategies according to the driving data of the vehicle through the policy module of the reinforcement learning model and the energy consumption large model, where each candidate control strategy includes the candidate opening degree of the active intake grille of the thermal management system and / or the candidate rotation speed of the cooling fan of the thermal management system; determining the target control strategy from the multiple candidate control strategies through the value module of the reinforcement learning model; where the energy consumption large model is trained through driving data samples and parameter samples corresponding to the driving data samples, and the reinforcement learning model is trained through parameter samples and candidate control strategy samples corresponding to the parameter samples.
[0008] Compared with the related technology, obtaining the target parameters through driving data and then obtaining the target control strategy according to the target parameters eliminates the need for threshold calibration, reduces the calibration cost, and enables flexible control of the thermal management system.
[0009] In some possible implementation manners, the number of the multiple candidate control strategies is N, where N is a positive integer. Obtaining multiple candidate control strategies according to the driving data of the vehicle through the policy module and the energy consumption large model of the reinforcement learning model includes: iteratively executing the policy acquisition process N times to obtain N candidate control strategies; wherein, in the first policy acquisition process of the N policy acquisition processes, it includes: obtaining target parameters through the policy module of the energy consumption large model based on the driving data; obtaining a candidate control strategy through the reinforcement learning model based on the target parameters; wherein, in the non-first policy acquisition process of the N policy acquisition processes, it includes: updating the target parameters through the energy consumption large model based on the driving data and the target parameters obtained in the previous policy acquisition process; obtaining a candidate control strategy through the policy module of the reinforcement learning model based on the updated target parameters.
[0010] Combining the reinforcement learning model and the energy consumption large model to obtain multiple candidate control strategies, and determining a better target control strategy from the multiple candidate control strategies for controlling the thermal management system, dynamically selecting a control strategy that better matches the vehicle, and improving the comfort of riding in the vehicle.
[0011] In some possible implementation manners, determining the target control strategy from the multiple candidate control strategies through the value module of the reinforcement learning model includes: determining the reward values respectively corresponding to the N candidate control strategies; determining the candidate control strategy corresponding to the maximum reward value as the target control strategy, and the maximum reward value is the maximum value among the reward values respectively corresponding to the N candidate control strategies.
[0012] In this implementation manner, according to the reward value corresponding to each of the multiple candidate control strategies, selecting the candidate control strategy corresponding to the maximum reward value as the target control strategy can significantly improve the performance and reliability of the thermal management system.
[0013] In some possible embodiments, determining the reward values corresponding to the N candidate control strategies respectively includes: sequentially taking the N candidate control strategies as the first candidate control strategy and performing the following steps: taking the sum of the first value corresponding to the first candidate control strategy and the second value corresponding to the first candidate control strategy as the reward value corresponding to the first candidate control strategy; wherein, when the candidate control strategy obtained in the first policy acquisition process is taken as the first candidate control strategy, the first value corresponding to the first candidate control strategy is 0, and when the candidate control strategy obtained in the m-th policy acquisition process is taken as the first candidate control strategy, the first value corresponding to the first candidate control strategy is: the reward value corresponding to the candidate control strategy obtained in the (m - 1)-th policy acquisition process, where m is a positive integer greater than 1 and less than N; the second value corresponding to the first candidate control strategy is: the score value of the first candidate control strategy.
[0014] In some possible embodiments, the score value of the first candidate control strategy is obtained based on the energy consumption corresponding to the candidate opening degree and / or candidate rotational speed of the first candidate control strategy.
[0015] In some possible embodiments, the energy consumption corresponding to the candidate opening degree and / or candidate rotational speed of the first candidate control strategy includes at least one type of energy consumption, and the score value of the first candidate control strategy is obtained based on the following steps: obtaining the intermediate score values corresponding to at least one type of energy consumption according to the at least one type of energy consumption corresponding to the candidate opening degree and / or candidate rotational speed of the first candidate control strategy and the energy consumption coefficients corresponding to the at least one type of energy consumption; obtaining the score value of the first candidate control strategy according to the intermediate score values corresponding to the at least one type of energy consumption.
[0016] In some possible embodiments, the at least one type of energy consumption corresponding to the candidate opening degree and / or candidate rotational speed of the first candidate control strategy is obtained based on the candidate opening degree and / or candidate rotational speed corresponding to the first candidate control strategy.
[0017] In some possible embodiments, the at least one type of energy consumption includes at least one of the following: the energy consumption of the compressor of the vehicle, the air resistance energy consumption, the energy consumption of the cooling fan.
[0018] In some possible embodiments, the air resistance energy consumption corresponding to the first candidate control strategy is obtained by the following method: obtaining the air resistance coefficient corresponding to the first candidate control strategy; obtaining the air resistance power according to the air resistance coefficient; obtaining the air resistance energy consumption corresponding to the first candidate control strategy according to the air resistance power.
[0019] In some possible embodiments, the first candidate control strategy includes the opening value of the active intake grille in the thermal management system, and obtaining the drag coefficient corresponding to the first candidate control strategy includes: obtaining the drag coefficient corresponding to the opening value according to a first preset mapping relationship, where the first preset mapping relationship includes the corresponding relationship between the opening value and the drag coefficient.
[0020] In some possible embodiments, obtaining the drag power according to the drag coefficient includes: obtaining the drag power corresponding to the drag-related parameters according to a second preset mapping relationship, where the second preset mapping relationship includes the corresponding relationship between the drag-related parameters and the drag power, and the drag-related parameters include the drag coefficient and at least one of the following: vehicle speed, air density, frontal area, drag coefficient.
[0021] In some possible embodiments, the target control strategy includes the target opening of the active intake grille of the thermal management system and / or the target speed of the cooling fan of the thermal management system.
[0022] Jointly control the active intake grille and the cooling fan according to the target control strategy to improve the effects of thermal management and energy consumption control.
[0023] In some possible embodiments, the energy consumption of the compressor of the vehicle corresponding to the first candidate control strategy is obtained by the following method: obtaining the power of the compressor corresponding to the first candidate control strategy; obtaining the energy consumption of the compressor corresponding to the first candidate control strategy according to the power of the compressor corresponding to the first candidate control strategy.
[0024] In some possible embodiments, the energy consumption of the cooling fan corresponding to the first candidate control strategy is obtained by the following method: obtaining the power of the cooling fan corresponding to the first candidate control strategy; obtaining the energy consumption of the cooling fan corresponding to the first candidate control strategy according to the power of the cooling fan corresponding to the first candidate control strategy.
[0025] In some possible embodiments, obtaining the power of the cooling fan corresponding to the first candidate control strategy includes: obtaining the candidate speed of the cooling fan corresponding to the first candidate control strategy; obtaining the power of the cooling fan corresponding to the candidate speed according to a third mapping relationship, where the third mapping relationship includes the corresponding relationship between the candidate speed and the power of the cooling fan.
[0026] According to a second aspect of the embodiments of the present disclosure, there is provided a control device for a thermal management system of a vehicle, which is applied to the method described in the first aspect.
[0027] According to a third aspect of the embodiments of the present disclosure, a vehicle is provided, including: a processor; a memory for storing instructions executable by the processor; wherein, when the processor is configured to execute the above instructions, the steps of the method described in the first aspect are implemented.
[0028] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the method provided in the first aspect of the present disclosure are implemented.
[0029] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0030] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0032] Figure 1 is one of the flow diagrams of the method for controlling the thermal management system of the vehicle provided by the present disclosure.
[0033] Figure 2 is an application scenario diagram of the method for controlling the thermal management system of the vehicle provided by the present disclosure.
[0034] Figure 3 is another flow diagram of the method for controlling the thermal management system of the vehicle provided by the present disclosure.
[0035] Figure 4 is a schematic structural diagram of the energy consumption large model provided by the present disclosure.
[0036] Figure 5 is another flow diagram of the method for controlling the thermal management system of the vehicle provided by the present disclosure.
[0037] Figure 6 is a schematic diagram of the TD3 large model for obtaining the target control strategy provided by the present disclosure.
[0038] Figure 7 is a schematic diagram of the large model deployed on the vehicle side provided by the present disclosure.
[0039] Figure 8 is a schematic structural diagram of the device for controlling the thermal management system of the vehicle provided by the present disclosure.
[0040] Figure 9 is a schematic structural diagram of the vehicle provided by the present disclosure. Detailed Implementation Modes
[0041] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0042] The thermal management system of a vehicle is a technical system that controls the temperature to ensure that the vehicle's electronic control system, battery, motor, cockpit, etc. operate under optimal conditions. The active grille system (AGS) and the cooling fan are the core components of the thermal management system. In addition, the thermal management system also includes devices such as a high-voltage coolant heater (HVCH), a warm core water pump, a compressor, an expansion valve, etc. The AGS controls the air flow entering the battery compartment and the motor compartment by adjusting the grille opening, thereby affecting the cooling effect and the overall vehicle aerodynamic drag; the cooling fan enhances the heat dissipation effect through forced convection. However, in the related art, the control of the AGS and the cooling fan is usually independent, lacking global optimization, and the thermal management and energy consumption control effects are poor.
[0043] Optimizing the energy efficiency of the vehicle's thermal management system has generally become one of the key technologies. The cooling systems in the related art usually adopt passive control strategies, which consume a lot of time to calibrate appropriate thresholds, set fixed thresholds for the thermal management system, and control the start and stop of the thermal management system according to the fixed thresholds to control the temperature. For example, a fixed temperature threshold is set, and the start and stop of the cooling fan are controlled according to the temperature threshold. For another example, a fixed AGS opening threshold is set. Or a fixed AGS opening threshold is used to adjust the intake air volume. Using fixed thresholds for control may result in less than ideal adjustment effects of the thermal management system under various vehicle operating conditions. For example, although the vehicle energy consumption is reduced after adjustment, the temperature adjustment is not ideal, or although the temperature is appropriate after adjustment, the vehicle energy consumption is high. Moreover, this control method in the related art cannot be dynamically adjusted according to the real-time operating conditions, resulting in energy waste and shortened driving range.
[0044] To solve the above problems, an embodiment of the present disclosure provides a control method for a vehicle's thermal management system. Please refer to Figure 1 ., the control method for the vehicle's thermal management system can be applied to Figure 8 the vehicle's thermal management system control device 300 shown in Figure 9The vehicle 600, computer program product, and computer-readable storage medium shown. In this embodiment, taking the application to a vehicle as an example, the vehicle can be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle can be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle, and this embodiment does not limit this. The following will be directed to Figure 1 the process shown will be elaborated in detail. The method for controlling the thermal management system of the vehicle can specifically include the following steps: Step S10: Obtain a target control strategy for controlling the thermal management system according to the driving data of the vehicle, where the target control strategy is one of multiple candidate control strategies obtained according to the driving data.
[0045] The driving data of the vehicle refers to the data related to the vehicle state, driving behavior, driving environment, etc. generated during the vehicle's driving. The driving data can include at least one of the following: frame ID (Identity document), driving time, saturated high pressure of the refrigerant in the thermal management system, compressor exhaust temperature, requested air temperature, positive temperature coefficient thermistor (PTC) power, HVCH inlet water temperature, HVCH outlet water temperature, warm core water pump flow rate, internal cooling outlet temperature, compressor inlet temperature, saturated low pressure of the refrigerant, compressor speed, expansion valve opening, AGS opening, cooling fan speed, vehicle speed, ambient temperature, occupant compartment temperature, compressor power, etc. The vehicle can save the driving data locally, and can also, as Figure 2 shown, the vehicle 210 uploads the driving data to the data platform 220.
[0046] In some possible implementation manners, multiple candidate control strategies are obtained according to the driving data of the vehicle, and one of them is determined as the target control strategy. Among them, the target control strategy can be the better one among the multiple candidate control strategies. For example, for each candidate control strategy among the multiple candidate control strategies, there is an evaluation index, and the candidate control strategies are evaluated through the evaluation index. The evaluation index can be confidence, score value, reward value, etc. According to the evaluation index, the target control strategy is determined from the multiple candidate strategies.
[0047] Obtain the target control strategy according to the driving data of the vehicle. The target control strategy can be obtained by combining with a large model. Input the driving data into the large model to obtain the target control strategy output by the large model.
[0048] In some possible implementation manners, please continue to refer to Figure 2, the driving data in the data platform 220 is input into the large model 230 to obtain the target control strategy output by the large model 230, and the thermal management system on the vehicle is controlled by the target control strategy, thereby realizing the control of the vehicle temperature. Optionally, the large model 230 can be deployed on the server running the data platform 220. The target control strategy is obtained through the large model 230 on the server, and then the server sends the target control strategy to the vehicle to control the thermal management system of the vehicle. Deploying the large model 230 on the server can reduce the memory occupancy of the vehicle 210, and with the help of the strong computing power of the server, a more accurate target control strategy can be obtained.
[0049] In some other possible implementation manners, please continue to refer to Figure 2 , the driving data collected by the vehicle 210 is input into the large model 230 to obtain the target control strategy output by the large model 230, and the thermal management system on the vehicle is controlled by the target control strategy, thereby realizing the control of the vehicle temperature. The large model 230 can be deployed locally on the vehicle, and the vehicle can quickly call the large model 230 to calculate the target control strategy, and the target control strategy can be quickly obtained to quickly control the thermal management system.
[0050] The control method of the thermal management system of the vehicle provided in this embodiment dynamically generates multiple candidate control strategies according to the driving data of the vehicle, selects the target control strategy from the multiple candidate control strategies, and uses the target control strategy to control the thermal management system, which performs thermal management more accurately and improves the adjustment effect. And compared with the related technology, there is no need to calibrate the threshold, shortening the development cycle.
[0051] In one implementation manner, please refer to Figure 3 , Figure 3 is a further explanation of step S10. Step S10 (obtaining a target control strategy for controlling the thermal management system according to the driving data of the vehicle) includes the following steps: Step S111: According to the driving data of the vehicle, through the policy modules of the energy consumption large model and the reinforcement learning model, obtain multiple candidate control strategies, and each candidate control strategy includes the candidate opening degree of the active intake grille of the thermal management system and / or the candidate rotation speed of the cooling fan of the thermal management system.
[0052] As a way, please refer to Figure 2 , the large model 230 includes an energy consumption large model 231, and the driving data is processed through the energy consumption large model 231 to obtain target parameters for the thermal management system. Among them, the energy consumption large model is trained through sample driving data and sample parameters corresponding to the sample driving data.
[0053] The number of target parameters can be N, where N is a positive integer. N target parameters are sequentially obtained through the high energy consumption model 231. Exemplarily, driving data is input into the high energy consumption model 231 to obtain the first target parameter output by the high energy consumption model 231. The first target parameter and the driving data are input into the high energy consumption model 231 to obtain the second target parameter output by the high energy consumption model 231. Similarly, the second target parameter and the driving data are input into the high energy consumption model 231 to obtain the third target parameter output by the high energy consumption model 231. In the same way, the Nth target parameter is obtained. For ease of understanding, the second to the Nth target parameters are all referred to as non-first target parameters. It is not difficult to understand that among the N target parameters obtained this time, the first target parameter is obtained based on the driving data, and the non-first target parameters are obtained based on the driving data and the previous target parameter.
[0054] N is a preset value. For example, N can be 20, 30, 50, etc., or the value range of N can be 30±n, where n is a preset value or a fixed value, such as 1, 2, 3, 5, 10, or 15, etc. The present disclosure does not limit this, or the value range of N can be 20≤N≤50, or the value range of N can be 20≤N≤40, or the value range of N can be 25≤N≤35, or the value range of N can be 28≤N≤32. The present disclosure does not limit this.
[0055] Optionally, the driving data is denoted as X(t) as the input of the high energy consumption model. X(t) can include the saturated low pressure at the previous moment, the compressor discharge temperature at the previous moment, the load request value of the occupant compartment in the thermal management system at the previous moment, the PTC power at the previous moment, the inlet water temperature of HVCH at the previous moment, the outlet water temperature of HVCH at the previous moment, the outlet temperature of the internal cooler at the previous moment, the compressor inlet temperature at the previous moment, the saturated low pressure, the ambient temperature, the compressor speed request at the previous moment, the current AGS opening, the current cooling fan speed, the vehicle speed, and the ambient temperature. The output is denoted as Y(t), and Y(t) includes the target saturated high pressure, the target saturated low pressure, the compressor speed, the opening of the constant pressure expansion valve, the internal cooling temperature, the compressor discharge temperature, the saturated high pressure, the saturated low pressure, the compressor inlet temperature, and the compressor power. The corresponding relationship between the input X(t) and the output Y(t) can be expressed by the following formula:
[0056] where W and b are set constants. The output Y(t) is used as the target parameter.
[0057] Exemplarily, the model structure of the high energy consumption model 231 can be as Figure 4As shown, the high energy consumption large model 231 may include a compressor target high pressure setting strategy model, a target low pressure setting model, a compressor control strategy model, a compressor high pressure prediction model, a compressor low pressure prediction model, and a compressor power model. The high energy consumption large model 231 obtains target parameters in the following manner: taking the driving data as the input quantity of the high energy consumption large model and denoting it as X(t), inputting the saturated high pressure at the previous moment, the load request value of the occupant compartment in the thermal management system at the previous moment, the compressor exhaust temperature at the previous moment, the PTC power at the previous moment, the water temperature at the outlet of the HVCH at the previous moment, the water temperature at the inlet of the HVCH at the previous moment, and the flow rate of the warm core water pump at the previous moment in X(t) into the compressor target high pressure setting strategy model, and the compressor target high pressure setting strategy model outputs the target saturated high pressure, where the target saturated high pressure is the saturated high pressure estimated by the target high pressure setting strategy model, and taking this target saturated high pressure as the output Y(t). Inputting the saturated low pressure at the previous moment, the current ambient temperature, and the compressor speed request at the previous moment in X(t) into the target low pressure setting model, and the target low pressure setting model outputs the target saturated low pressure, where the target saturated low pressure is the saturated low pressure estimated by the target low pressure setting model, and taking this target saturated low pressure as the output Y(t) as well. Inputting the difficulty at the outlet of the internal cooler at the previous moment, the compressor inlet temperature at the previous moment, and the saturated low pressure at the previous moment in X(t), together with the target saturated high pressure output by the compressor target high pressure setting strategy model and the target saturated low pressure output by the target low pressure setting model, into the compressor control strategy model, and the compressor control strategy model outputs the compressor speed and the opening degree of the constant pressure expansion valve, and taking the compressor speed and the opening degree of the constant pressure expansion valve output by the compressor control strategy model as the output Y(t) as well. Inputting the internal cooler outlet temperature at the previous moment, the compressor inlet temperature at the previous moment, and the saturated low pressure at the previous moment in X(t), together with the compressor speed and the opening degree of the constant pressure expansion valve output by the compressor control strategy model, into the compressor high pressure prediction model, and the compressor high pressure prediction model outputs the internal cooler temperature, the compressor exhaust temperature, the saturated high pressure, and the opening degree of the constant pressure expansion valve, and taking the internal cooler temperature, the compressor exhaust temperature, and the saturated high pressure output by the compressor high pressure prediction model as the output Y(t) as well. Inputting the current AGS opening degree, the current cooling fan speed, the current vehicle speed, and the current ambient temperature in X(t), together with the internal cooler temperature, the compressor exhaust temperature, the saturated high pressure, and the opening degree of the constant pressure expansion valve output by the compressor high pressure prediction model, into the compressor low pressure prediction model, and the compressor low pressure prediction model outputs the compressor speed, the saturated low pressure, the compressor inlet temperature, the saturated high pressure, and the compressor exhaust temperature, and taking the saturated low pressure and the compressor inlet temperature output by the compressor low pressure prediction model as the output Y(t). Inputting the compressor speed, the saturated low pressure, the compressor inlet temperature, the saturated high pressure, and the compressor exhaust temperature output by the compressor low pressure prediction model into the compressor power model, and the compressor power model outputs the compressor power, and taking the compressor power output by the compressor power model as the output Y(t). Taking Y(t) as the target parameter.
[0058] Please continue to refer to Figure 2 Figure 2 , the large model 230 further includes a reinforcement learning model 232. By using the reinforcement learning model 232 in the large model 230, for each of the N target parameters, a candidate control strategy corresponding to each target parameter is obtained, and a plurality of candidate control strategies are obtained.
[0059] The number of target parameters is N. The N target parameters are sequentially input into the reinforcement learning model to obtain a candidate control strategy corresponding to each target parameter, and N candidate control strategies are obtained. Then, through the reinforcement learning model, one of the N candidate control strategies is determined as the target control strategy, and the target control strategy is output through the reinforcement learning model. Optionally, the reinforcement learning model includes a policy module, and N candidate control strategies are generated through the policy module.
[0060] The above generation of target parameters and target control strategies can be a dynamic process. For example, the policy acquisition process is iteratively executed N times to obtain N candidate control strategies.
[0061] Among them, in the first policy acquisition process of the N policy acquisition processes: Based on the driving data, the target parameters are obtained through the energy consumption large model; based on the target parameters, a candidate control strategy is obtained through the policy module of the reinforcement learning model.
[0062] Among them, in the non-first policy acquisition process of the N policy acquisition processes: Based on the driving data and the target parameters obtained in the previous policy acquisition process, the target parameters are updated through the energy consumption large model; based on the updated target parameters, a candidate control strategy is obtained through the policy module of the reinforcement learning model.
[0063] Step S112: According to the plurality of candidate control strategies, through the value module of the reinforcement learning model, the target control strategy is determined from the plurality of candidate control strategies.
[0064] Among them, the energy consumption large model is obtained by training with driving data samples and parameter samples corresponding to the driving data samples, and the reinforcement learning model is obtained by training with parameter samples and candidate control strategy samples corresponding to the parameter samples.
[0065] In this embodiment, the reinforcement learning model and the energy consumption large model are combined to obtain a plurality of candidate control strategies, and a better target control strategy is determined from the plurality of candidate control strategies to control the thermal management system, dynamically selecting a control strategy that better matches the vehicle, and improving the comfort of riding in the vehicle.
[0066] Since the number of candidate policies obtained is N, it is necessary to determine a relatively optimal target control policy from the N candidate control policies to improve the control effect of the thermal management system. Please refer to Figure 5 , Figure 5 This is a further explanation of step S112. Step S112 (determine the target control policy from multiple candidate control policies through the value module of the reinforcement learning model according to the multiple candidate control policies). In one implementation, using the value module, the target control policy can be obtained in the following manner: Step S1121: Determine the reward values corresponding to the N candidate control policies respectively.
[0067] As a way, take the N candidate control policies as the first candidate control policy in turn and execute the following steps: Take the sum of the first value corresponding to the first candidate control policy and the second value corresponding to the first candidate control policy as the reward value corresponding to the first candidate control policy. The reward value corresponding to each candidate control policy can be used to evaluate the energy consumption corresponding to the candidate control policy.
[0068] Among them, when the candidate control policy obtained in the first policy acquisition process is used as the first candidate control policy, the first value corresponding to the first candidate control policy is 0. When the candidate control policy obtained in the m-th policy acquisition process is used as the first candidate control policy, the first value corresponding to the first candidate control policy is: the reward value corresponding to the candidate control policy obtained in the (m - 1)-th policy acquisition process, where m is a positive integer greater than 1 and less than N. The second value corresponding to the first candidate control policy is: the score value of the first candidate control policy. It is not difficult to understand that for the first candidate control policy, its reward value is the score value corresponding to itself. For non-first candidate control policies, its reward value is the sum of the score value corresponding to itself and the reward value corresponding to the previous candidate control policy. It can be understood that the reward value of non-first candidate control policies is obtained by accumulating the score value of itself and the score values of previous candidate control policies.
[0069] As another way, through the reinforcement learning model, determine the reward values corresponding to the N candidate control policies respectively, and obtain N reward values. Optionally, the reinforcement learning model includes a value module. Through the value module, determine the reward values corresponding to the N candidate control policies respectively, and obtain N reward values.
[0070] Step S1122: Determine the candidate control policy corresponding to the maximum reward value as the target control policy, where the maximum reward value is the maximum value among the reward values corresponding to the N candidate control policies.
[0071] Through a reinforcement learning model, the candidate control strategy corresponding to the maximum reward value among N reward values is determined as the target control strategy. Optionally, through a value module, the candidate control strategy corresponding to the maximum reward value among N reward values is determined as the target control strategy.
[0072] In this embodiment, according to the reward value corresponding to each of multiple candidate control strategies, the candidate control strategy corresponding to the maximum reward value is selected as the target control strategy, which can significantly improve the performance and reliability of the thermal management system.
[0073] In another embodiment, M candidate reward values with relatively large reward values can be determined from N reward values. It is not difficult to understand that any one of the M candidate reward values is greater than the other reward values among the N reward values except the M candidate reward values. The candidate control strategy corresponding to any one of the M candidate reward values is determined as the target control strategy.
[0074] In another embodiment, the candidate control strategy is used to control the thermal management system. For example, the candidate control strategy is used to control the opening degree of the active intake grille of the thermal management system and / or the rotation speed of the cooling fan of the thermal management system. In addition to the aforementioned consideration of energy consumption, the cost of hardware devices in the thermal management system can also be considered. For example, when the rotation speed of the cooling fan is greater than the preset rotation speed, it may cause relatively serious loss to the cooling fan. Considering the hardware cost of the candidate control strategies corresponding to the M candidate reward values, the cost value corresponding to each candidate control strategy among the M candidate control strategies is obtained, and M cost values are obtained. The candidate control strategy corresponding to the minimum cost value among the M cost values is determined as the target control strategy.
[0075] The aforementioned N target parameters, the candidate control strategy corresponding to each target parameter, and the target control strategy can be used as training samples, and the training samples are used to iteratively train the reinforcement learning model to optimize the reinforcement learning model. Optionally, the update frequencies of the policy module and the value module in the reinforcement learning model can be the same. The update frequencies of the two can also be different. For example, the update frequency of the policy module is lower than the update frequency of the value module, so that the value module can more accurately evaluate the policy module, reduce the instability of the policy module update, reduce the cumulative error, and improve the overall performance of the reinforcement learning model.
[0076] Exemplarily, the policy module can be the Actor in the TD3 (Twin Delayed Deep Deterministic Policy Gradient) model, and the value model can be the Critic in the TD3 model. Please refer to Figure 6, the number of target parameters is N, which can be S0 to SN. S0 to SN are sequentially input into TD3 for processing. For example, based on the target parameter St, the candidate control strategy at is obtained through the Actor module, and the Critic module combines the total energy consumption and the superheat control target to obtain the reward value Qt corresponding to the candidate control strategy at. Among them, the input target parameter St is any one between S0 and SN. Through the energy consumption large model, according to the driving data and the target parameter St, the target parameter St+1 is obtained, and the target parameter St+1 is the next target parameter of the target parameter St. Based on the target parameter St+1, the candidate control strategy at+1 is obtained through the Actor module, and the Critic module combines the total energy consumption and the superheat control target to obtain the reward value Qt+1 corresponding to the candidate control strategy at+1. Exemplarily, the Critic module combines the total energy consumption and the superheat control target to obtain the score value Rt+1 corresponding to the candidate control strategy at+1, and based on the sum of the score value Rt+1 and the reward value Qt corresponding to the previous candidate control strategy at, the reward value Qt+1 corresponding to the candidate control strategy at+1 is obtained.
[0077] Optionally, in order to improve the overall performance of the TD3 model, the Actor module and the Critic module can be updated by means of soft update.
[0078] Optionally, the score value of the first candidate control strategy is obtained based on the energy consumption corresponding to the candidate opening and / or candidate rotation speed corresponding to the first candidate control strategy. At least one of the energy consumptions corresponding to the candidate opening and / or candidate rotation speed corresponding to the first candidate control strategy is obtained based on the candidate opening and / or candidate rotation speed corresponding to the first candidate control strategy.
[0079] The energy consumption includes at least one type of energy consumption, and the energy consumption can be power consumption, power, etc. The score value of the first candidate control strategy is obtained based on the following steps: According to at least one energy consumption corresponding to the candidate opening and / or candidate rotation speed of the first candidate control strategy, and the energy consumption coefficients corresponding to at least one energy consumption respectively, the intermediate score values corresponding to at least one energy consumption respectively are obtained. Optionally, the product of the energy consumption and the energy consumption coefficient μ corresponding to the energy consumption is used as the intermediate score value. Then, based on the intermediate score values corresponding to at least one energy consumption respectively, the score value of the first candidate control strategy is obtained. In the case where there are multiple types of energy consumption, the score value of the first candidate control strategy is obtained based on the sum of the intermediate score values corresponding to multiple energy consumptions respectively.
[0080] For example, the at least one energy consumption includes at least one of the following: the energy consumption of the compressor of the vehicle, the air resistance energy consumption, the energy consumption of the cooling fan in the thermal management system, and the energy consumption corresponding to the superheat degree of the thermal management system.
[0081] Among them, the wind resistance energy consumption corresponding to the first candidate control strategy is obtained in the following manner: Obtain the wind resistance coefficient Ca corresponding to the first candidate control strategy. The wind resistance coefficient Ca mainly consists of form drag, surface friction drag, internal drag, interference drag, etc. of the vehicle. Among them, the internal drag is affected by factors such as air resistance generated by air passing through the radiator, air inlet, etc. of the thermal management system, and the AGS opening. The AGS opening and the wind resistance coefficient Ca are positively correlated, that is, the larger the AGS opening, the larger the wind resistance coefficient Ca, and vice versa, the smaller the AGS opening, the smaller the wind resistance coefficient Ca. Obtain the wind resistance power Fa(t) according to the wind resistance coefficient Ca. Obtain the wind resistance energy consumption in the target parameters corresponding to the first candidate control strategy according to the wind resistance power.
[0082] The first candidate control strategy includes the opening value ags of the active intake grille in the thermal management system. The obtaining of the wind resistance coefficient corresponding to the first candidate control strategy includes: Obtain the wind resistance coefficient corresponding to the opening value according to the first preset mapping relationship. Among them, the first preset mapping relationship includes the corresponding relationship between the opening value and the wind resistance coefficient. Among them, the first preset mapping relationship can be expressed by the following formula:
[0083] Among them, Ca is the wind resistance coefficient, ags is the opening of the active intake grille, and k1, k2, and k3 are coefficients, all of which are constants. ags can be a percentage. For example, the opening is 40%, 50%, etc.
[0084] For example, if gas is x1, then, combining the above formula, Ca can be calculated as .
[0085] The obtaining of the wind resistance power according to the wind resistance coefficient includes: Obtain the wind resistance power corresponding to the wind resistance-related parameters according to the second preset mapping relationship. Among them, the second preset mapping relationship includes the corresponding relationship between the wind resistance-related parameters and the wind resistance power. The wind resistance-related parameters include the wind resistance coefficient and at least one of the following: vehicle speed, air density, frontal area, wind resistance coefficient.
[0086] The second preset mapping relationship is expressed by the following formula:
[0087] Among them, Fa is the wind resistance power, and the unit of Fa can be watt (Watt, W), is the air density, the unit of can be kilogram per cubic meter (kg / m³), is the vehicle speed, the unit of v can be kilometer per second (km / s), Ca is the wind resistance coefficient, and A is the frontal area. The unit of A can be square meter (m³).
[0088] For example, when x2 is x2, v is x3, Ca is x4, and A is x5, combining the above formula, the aerodynamic drag power Fa is .
[0089] The energy consumption of the compressor of the vehicle corresponding to the first candidate control strategy is obtained in the following manner: Obtain the power of the compressor corresponding to the first candidate control strategy, and the power of the compressor can be directly obtained from the driving data. According to the power of the compressor corresponding to the first candidate control strategy, obtain the energy consumption of the compressor corresponding to the first candidate control strategy. For example, integrate the power of the compressor to obtain the energy consumption of the compressor.
[0090] The energy consumption of the cooling fan corresponding to the first candidate control strategy is obtained in the following manner: Obtain the power of the cooling fan corresponding to the first candidate control strategy. According to the power of the cooling fan corresponding to the first candidate control strategy, obtain the energy consumption of the cooling fan corresponding to the first candidate control strategy. For example, integrate the power of the cooling fan to obtain the energy consumption of the cooling fan.
[0091] The power of the cooling fan can be directly obtained from the driving data, and can also be obtained in the following manner: Obtain the candidate speed of the cooling fan corresponding to the first candidate control strategy; according to the third mapping relationship, obtain the power of the cooling fan corresponding to the candidate speed, where the third mapping relationship includes the corresponding relationship between the candidate speed and the power of the cooling fan. The third mapping relationship can be represented by the following formula:
[0092] where is the power of the cooling fan, fan is the candidate speed of the cooling fan, and k1, k2, and k3 are coefficients, and k1, k2, and k3 can all be constants.
[0093] The energy consumption corresponding to the superheat degree of the first candidate control strategy is obtained in the following manner: According to the difference between the superheat degree in the thermal management system and the target superheat degree, obtain the energy consumption corresponding to the superheat degree. The energy consumption corresponding to the superheat degree can be represented by the following formula:
[0094] where is the energy consumption corresponding to the superheat degree, is the actual superheat degree, is the target superheat degree. The target superheat degree is preset and can be 3.5.
[0095] Taking the energy consumption including the energy consumption of the compressor of the vehicle, the aerodynamic drag energy consumption, the energy consumption of the cooling fan in the thermal management system, and the energy consumption corresponding to the superheat degree of the thermal management system as an example. The corresponding energy consumption can be obtained according to the power. Exemplarily, the energy consumption can be obtained according to the integral of the power. For example, the energy consumption of the compressor can be obtained according to the integral of the power of the compressor, the aerodynamic drag energy consumption can be obtained according to the integral of the aerodynamic drag power, and the energy consumption of the cooling fan can be obtained according to the integral of the power of the cooling fan. According to the energy consumption J1 of the compressor and its corresponding coefficient μ1, the intermediate score value J1·μ1 corresponding to the energy consumption J1 of the compressor is obtained. For the aerodynamic drag energy consumption J2 and its corresponding coefficient μ2, the intermediate score value J2·μ2 corresponding to the aerodynamic drag energy consumption J2 is obtained. For the energy consumption J3 of the cooling fan in the thermal management system and its corresponding coefficient μ3, the intermediate score value J3·μ3 corresponding to the energy consumption J3 of the cooling fan is obtained. And the energy consumption corresponding to the superheat degree and its corresponding coefficient μ4, the energy consumption corresponding to the superheat degree The corresponding intermediate score value ·μ4. Calculate the sum of the above intermediate score values to obtain the score of the first candidate control strategy as J1·μ1 + J2·μ2 + J3·μ3 + ·μ4.
[0096] It should be noted that for the energy consumption with the same coefficient, when obtaining this type of energy consumption, the powers corresponding to this type of energy consumption can be added to obtain the sum of powers, and then the integral of the sum of powers is performed to obtain the total energy consumption. For example, μ1, μ2, and μ3 are the same, the powers corresponding to the energy consumption of μ1, μ2, and μ3 can be added, that is, the energy consumption of the compressor Fc(t), the aerodynamic drag power Fa(t), and the power of the cooling fan Ff(t) are added, and then the integral of the sum of powers Fc(t) + Fa(t) + Ff(t) is performed to obtain the total energy consumption . The intermediate score value corresponding to the total energy consumption is ·μ1. The score value of the first candidate control strategy is ·μ1 + ·μ4.
[0097] Optionally, the thermal management system includes an active intake grille and a cooling fan. The target control strategy includes the target opening degree of the active intake grille and / or the target rotational speed of the cooling fan. Joint control of the active intake grille and the cooling fan according to the target control strategy can improve the effects of thermal management and energy consumption control compared with the single adjustment method in the related art.
[0098] After obtaining the driving data of the vehicle, the driving data can also be preprocessed to obtain the preprocessed driving data. Then, according to the preprocessed driving data, the target control strategy for controlling the thermal management system is obtained.
[0099] The preprocessing of driving data can be carried out in the following way: Arrange the driving data in ascending order according to the vehicle frame ID and time, ensuring that the operation data of the same device on the vehicle is arranged together in chronological order. Process the sampling signals of different cycles into the same interval length. For example, the interval is 1 second. Then normalize the processed driving data, and adopt different normalization methods for different field types. For example, for values such as saturated high voltage, compressor exhaust temperature, requested air temperature, PTC power, HVCH inlet water temperature, HVCH outlet water temperature, warm core water pump flow, inner cold outlet temperature, etc., use min-max normalization. The processed driving data is X, the value range of the processed driving data is [Xmin, Xmax], and the normalized value is X’. X’ is calculated by the following formula:
[0100] Preprocessing the driving data of the vehicle can improve the quality of the driving data and provide more accurate and reliable data for subsequent obtaining of the target control strategy.
[0101] The energy consumption large model and the reinforcement learning model can be deployed on the vehicle or on the server connected to the vehicle.
[0102] In one implementation, please refer to Figure 7 , due to the high requirement for timeliness in controlling AGS and the cooling fan in the thermal management system and the need for real-time response, the reinforcement learning model can be deployed in the in-vehicle NPU (Neural Processing Unit) chip to facilitate the real-time estimation and output of the target control strategy. The estimation link is that the vehicle's CPU (Central Processing Unit) accesses the input signal (the input data can be driving data) from Carservice, processes the signal and sends it to the NPU for inference to obtain the target control strategy. The NPU returns the target control strategy to the CPU, and finally the CPU outputs the target control strategy to control the thermal management system. Among them, Carservice is the in-vehicle Android operating system.
[0103] Based on the same inventive concept, the present disclosure also provides a control device for the thermal management system of a vehicle, which is applied to the method for the control device of the thermal management system of the aforementioned vehicle. Figure 8 It is a block diagram of a control device for the thermal management system of a vehicle shown according to an exemplary embodiment. Please refer to Figure 8 , the control device 300 for the thermal management system of the vehicle includes: An acquisition module 310 is configured to obtain a target control strategy for controlling the thermal management system according to the driving data of the vehicle, where the target control strategy is one of multiple candidate control strategies obtained based on the driving data.
[0104] In a possible implementation, the target control strategy includes the target opening degree of the active intake grille of the thermal management system, and / or the target rotation speed of the cooling fan of the thermal management system.
[0105] In a possible implementation, the acquisition module 310 includes: A target parameter acquisition module is configured to obtain multiple candidate control strategies through an energy consumption large model and a policy module of a reinforcement learning model according to the driving data of the vehicle. Each candidate control strategy includes the candidate opening degree of the active intake grille of the thermal management system, and / or the candidate rotation speed of the cooling fan of the thermal management system; A target control strategy acquisition module is configured to determine the target control strategy from the multiple candidate control strategies through the value module of the reinforcement learning model; Wherein, the energy consumption large model is obtained by training with driving data samples and parameter samples corresponding to the driving data samples, and the reinforcement learning model is obtained by training with parameter samples and candidate control strategy samples corresponding to the parameter samples.
[0106] In a possible implementation, the number of the multiple candidate control strategies is N, where N is a positive integer. The target control strategy acquisition module includes: An iteration module is configured to iteratively execute the policy acquisition process N times to obtain N candidate control strategies; Wherein, in the first policy acquisition process of the N policy acquisition processes: Based on the driving data, obtain target parameters through the energy consumption large model; based on the target parameters, obtain a candidate control strategy through the policy module of the reinforcement learning model; Wherein, in the non-first policy acquisition process of the N policy acquisition processes: Based on the driving data and the target parameters obtained in the previous policy acquisition process, update the target parameters through the energy consumption large model; based on the updated target parameters, obtain a candidate control strategy through the policy module of the reinforcement learning model.
[0107] In a possible implementation, the acquisition module 310 includes: A reward value acquisition module is configured to determine the reward values corresponding to the N candidate control strategies respectively; A determination module, configured to determine the candidate control strategy corresponding to the maximum reward value as the target control strategy, where the maximum reward value is the maximum value among the reward values corresponding to the N candidate control strategies respectively.
[0108] In a possible implementation manner, the reward value acquisition module includes: A summation module, configured to use the sum of the first value corresponding to the first candidate control strategy and the second value corresponding to the first candidate control strategy as the reward value corresponding to the first candidate control strategy; Wherein, when the candidate control strategy obtained in the first policy acquisition process is used as the first candidate control strategy, the first value corresponding to the first candidate control strategy is 0; when the candidate control strategy obtained in the m-th policy acquisition process is used as the first candidate control strategy, the first value corresponding to the first candidate control strategy is: the reward value corresponding to the candidate control strategy obtained in the (m - 1)-th policy acquisition process, where m is a positive integer greater than 1 and less than N; The second value corresponding to the first candidate control strategy is: the score value of the first candidate control strategy.
[0109] In a possible implementation manner, the score value of the first candidate control strategy is obtained based on the energy consumption corresponding to the candidate opening degree and / or candidate rotation speed corresponding to the first candidate control strategy.
[0110] In a possible implementation manner, the energy consumption corresponding to the candidate opening degree and / or candidate rotation speed corresponding to the first candidate control strategy includes at least one type of energy consumption, and further includes: An intermediate score value acquisition module, configured to obtain intermediate score values corresponding to at least one type of energy consumption respectively according to at least one type of energy consumption corresponding to the candidate opening degree and / or candidate rotation speed of the first candidate control strategy and the energy consumption coefficients corresponding to the at least one type of energy consumption; A score value acquisition module, configured to obtain the score value of the first candidate control strategy according to the intermediate score values corresponding to the at least one type of energy consumption respectively.
[0111] In a possible implementation manner, the at least one type of energy consumption corresponding to the candidate opening degree and / or candidate rotation speed corresponding to the first candidate control strategy is obtained based on the candidate opening degree and / or candidate rotation speed corresponding to the first candidate control strategy.
[0112] In a possible implementation manner, the at least one type of energy consumption includes at least one of the following: the energy consumption of the compressor of the vehicle, the air resistance energy consumption, the energy consumption of the cooling fan.
[0113] In a possible implementation manner, the vehicle thermal management system control device 300 further includes: A wind resistance coefficient acquisition module, configured to acquire the wind resistance coefficient corresponding to the first candidate control strategy; A first power acquisition module, configured to acquire the wind resistance power according to the wind resistance coefficient; A first energy consumption acquisition module, configured to obtain the wind resistance energy consumption in the target parameters corresponding to the first candidate control strategy according to the wind resistance power.
[0114] In a possible implementation manner, the first candidate control strategy includes the opening value of the active intake grille in the thermal management system, and the wind resistance coefficient acquisition module includes: A first mapping module, configured to obtain the wind resistance coefficient corresponding to the opening value according to a first preset mapping relationship, where the first preset mapping relationship includes the corresponding relationship between the opening value and the wind resistance coefficient.
[0115] In a possible implementation manner, the first power acquisition module includes: A second mapping module, configured to obtain the wind resistance power corresponding to the wind resistance-related parameters according to a second preset mapping relationship, where the second preset mapping relationship includes the corresponding relationship between the wind resistance-related parameters and the wind resistance power, and the wind resistance-related parameters include the wind resistance coefficient and at least one of the following: vehicle speed, air density, frontal area, wind resistance coefficient.
[0116] In a possible implementation manner, the thermal management system control device 300 of the vehicle further includes: a second power acquisition module, configured to acquire the power of the compressor corresponding to the first candidate control strategy; A second energy consumption acquisition module, configured to obtain the energy consumption of the compressor corresponding to the first candidate control strategy according to the power of the compressor corresponding to the first candidate control strategy.
[0117] In a possible implementation manner, the thermal management system control device 300 of the vehicle further includes: A third power acquisition module, configured to acquire the power of the cooling fan corresponding to the first candidate control strategy; A third energy consumption acquisition module, configured to obtain the energy consumption of the cooling fan corresponding to the first candidate control strategy according to the power of the cooling fan corresponding to the first candidate control strategy.
[0118] In a possible implementation manner, the third power acquisition module includes: A rotational speed acquisition module, configured to acquire the candidate rotational speed of the cooling fan corresponding to the first candidate control strategy; A third mapping module, configured to acquire the power of the cooling fan corresponding to the candidate rotational speed according to a third mapping relationship, where the third mapping relationship includes the corresponding relationship between the candidate rotational speed and the power of the cooling fan.
[0119] Regarding the vehicle thermal management system control device 300 in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0120] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the vehicle thermal management system control method provided by the present disclosure are implemented.
[0121] Figure 9 It is a block diagram of a vehicle shown according to an exemplary embodiment. For example, the vehicle 600 can be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 600 can be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.
[0122] Please refer to Figure 9 , the vehicle 600 may include various subsystems. For example, the infotainment system 610, the perception system 620, the decision control system 630, the drive system 640, and the computing platform 650. Among them, the vehicle 600 may further include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 600 can be interconnected by wired or wireless means.
[0123] In some embodiments, the infotainment system 610 may include a communication system, an entertainment system, a navigation system, and the like.
[0124] The perception system 620 may include several types of sensors for sensing information about the environment around the vehicle 600. For example, the perception system 620 may include a global positioning system (the global positioning system can be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter wave radar, ultrasonic radar, and a camera device.
[0125] The decision control system 630 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.
[0126] The drive system 640 may include components that provide power movement for the vehicle 600. In one embodiment, the drive system 640 may include an engine, an energy source, a transmission system, and wheels. The engine can be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting the energy provided by the energy source into mechanical energy.
[0127] Some or all functions of vehicle 600 are controlled by computing platform 650. Computing platform 650 may include at least one processor 651 and a memory 652, and processor 651 may execute instructions 653 stored in memory 652.
[0128] Processor 651 may be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.
[0129] Memory 652 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0130] In addition to instructions 653, memory 652 may also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in memory 652 can be used by computing platform 650.
[0131] In an embodiment of the present disclosure, processor 651 may execute instructions 653 to complete all or part of the steps of the above-described method for controlling the thermal management system of the vehicle.
[0132] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program capable of being executed by a programmable device, and the computer program has a code portion for executing the above-described method for controlling the thermal management system of the vehicle when executed by the programmable device.
[0133] Optionally, the chip system may further include a memory for storing necessary computer programs and data.
[0134] Those skilled in the art can also understand that the various illustrative logical blocks and steps listed in the embodiments of the present application can be implemented by electronic hardware, computer software, or a combination of both. Whether such a function is implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art can use various methods to implement the described function for each specific application, but such implementation should not be construed as exceeding the scope protected by the embodiments of the present application.
[0135] In addition, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be understood as being advantageous compared to other aspects or designs. Instead, the use of the word exemplary is intended to present concepts in a concrete manner. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X applies A or B" is intended to mean any arrangement in a natural inclusive arrangement. That is, if X applies A; X applies B; or X applies both A and B, then "X applies A or B" is satisfied in any of the foregoing instances. Additionally, unless otherwise specified or clear from the context indicating a singular form, the articles "a" and "an" as used in this application and the appended claims are generally understood to mean "one or more".
[0136] Similarly, although the present disclosure has been shown and described with respect to one or more implementations, those skilled in the art will envision equivalent variations and modifications after reading and understanding this specification and the drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. Specifically with respect to the various functions performed by the components (e.g., elements, resources, etc.) described above, unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. Additionally, although a particular feature of the present disclosure may have been disclosed with respect to only one of several implementations, such a feature may be combined with one or more other features of other implementations as may be desired and advantageous for any given or particular application. Furthermore, with respect to the use of "comprises", "has", "includes", "contains", or variants thereof in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "includes".
[0137] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the art not disclosed in the present disclosure. The specification and examples are only illustrative, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
[0138] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
[0139] It should be understood that, unless otherwise specifically stated, the features of some embodiments of the various present disclosures described herein can be combined with each other. As used herein, the term "and / or" includes any one of the related listed items and any combination of any two or more of them; similarly, "at least one of......" includes any one of the related listed items and any combination of any two or more of them.
Claims
1. A method for controlling a thermal management system of a vehicle, characterized in that: The method comprises: A target control strategy for controlling the thermal management system is obtained based on the driving data of the vehicle, wherein the target control strategy is one of a plurality of candidate control strategies obtained based on the driving data.
2. The vehicle thermal management system control method according to claim 1, characterized in that: The target control strategy includes a target opening of an active air intake grille of the thermal management system and / or a target rotation speed of a cooling fan of the thermal management system.
3. The vehicle thermal management system control method according to claim 1 or 2, characterized in that: The step of obtaining a target control strategy for controlling the thermal management system according to the vehicle driving data includes: According to the driving data of the vehicle, a plurality of candidate control strategies are obtained through a strategy module of a reinforcement learning model and a large energy consumption model, each candidate control strategy including a candidate opening of an active air intake grille of the thermal management system and / or a candidate speed of a cooling fan of the thermal management system; According to the multiple candidate control strategies, determining the target control strategy from the multiple candidate control strategies through the value module of the reinforcement learning model; The large energy consumption model is obtained by training with driving data samples and parameter samples corresponding to the driving data samples, and the reinforcement learning model is obtained by training with parameter samples and candidate control strategy samples corresponding to the parameter samples.
4. The method according to claim 3, characterized in that The number of the plurality of candidate control strategies is N, and the plurality of candidate control strategies are obtained according to the driving data of the vehicle through the strategy module of the reinforcement learning model and the energy consumption model, including: Iterate the strategy acquisition process N times to obtain N candidate control strategies; Among them, in the N times of strategy acquisition process, the first strategy acquisition process includes: Based on the driving data, a target parameter is obtained through an energy consumption model; based on the target parameter, a candidate control strategy is obtained through a strategy module of the reinforcement learning model; Among them, in the N times of strategy acquisition process, the non-first strategy acquisition process includes: Based on the driving data and the target parameters obtained in the previous strategy acquisition process, the target parameters are updated through the energy consumption model; based on the updated target parameters, a candidate control strategy is obtained through the strategy module of the reinforcement learning model.
5. The method according to claim 3 or 4, characterized in that: The step of determining the target control strategy from the multiple candidate control strategies by using a value module of a reinforcement learning model according to the multiple candidate control strategies includes: Determine the reward values corresponding to the N candidate control strategies respectively; The candidate control strategy corresponding to the maximum reward value is determined as the target control strategy, and the maximum reward value is the maximum value of the reward values respectively corresponding to the N candidate control strategies.
6. The method according to claim 5, characterized in that The determining of the reward values corresponding to the N candidate control strategies respectively includes: The N candidate control strategies are sequentially used as the first candidate control strategy to perform the following steps: taking the sum of a first value corresponding to the first candidate control strategy and a second value corresponding to the first candidate control strategy as the reward value corresponding to the first candidate control strategy; Wherein, when the candidate control strategy obtained in the first strategy acquisition process is used as the first candidate control strategy, the first value corresponding to the first candidate control strategy is 0, and when the candidate control strategy obtained in the mth strategy acquisition process is used as the first candidate control strategy, the first value corresponding to the first candidate control strategy is: the reward value corresponding to the candidate control strategy obtained in the m-1th strategy acquisition process, where m is a positive integer greater than 1 and less than N; The second value corresponding to the first candidate control strategy is: the score value of the first candidate control strategy.
7. The method according to claim 6, characterized in that The score value of the first candidate control strategy is obtained based on the energy consumption corresponding to the candidate opening degree and / or the candidate speed of the first candidate control strategy.
8. The method according to claim 7, characterized in that The energy consumption corresponding to the candidate opening degree and / or candidate speed of the first candidate control strategy includes at least one energy consumption, and the score value of the first candidate control strategy is obtained based on the following steps: Obtaining an intermediate score value corresponding to at least one energy consumption according to at least one energy consumption corresponding to the candidate opening degree and / or the candidate speed of the first candidate control strategy and an energy consumption coefficient corresponding to at least one energy consumption; The score value of the first candidate control strategy is obtained according to the intermediate score values respectively corresponding to the at least one energy consumption.
9. The method according to claim 8, characterized in that At least one energy consumption corresponding to the candidate opening degree and / or the candidate speed corresponding to the first candidate control strategy is obtained based on the candidate opening degree and / or the candidate speed corresponding to the first candidate control strategy.
10. The method according to claim 8, characterized in that The at least one energy consumption includes at least one of the following: energy consumption of a compressor of the vehicle, wind resistance energy consumption, and energy consumption of the cooling fan.
11. The method according to claim 10, characterized in that The windage energy consumption corresponding to the first candidate control strategy is obtained in the following manner: Obtaining a drag coefficient corresponding to a first candidate control strategy; Obtaining wind resistance power according to the wind resistance coefficient; According to the wind resistance power, wind resistance energy consumption corresponding to the first candidate control strategy is obtained.
12. The method according to claim 11, characterized in that The first candidate control strategy includes an opening value of an active air intake grille in the thermal management system, and obtaining a drag coefficient corresponding to the first candidate control strategy includes: According to a first preset mapping relationship, a drag coefficient corresponding to the opening value is obtained, wherein the first preset mapping relationship includes a corresponding relationship between the opening value and the drag coefficient.
13. The method according to claim 11, characterized in that The obtaining of wind resistance power according to the wind resistance coefficient includes: According to the second preset mapping relationship, the wind resistance power corresponding to the wind resistance related parameters is obtained, wherein the second preset mapping relationship includes the correspondence between the wind resistance related parameters and the wind resistance power, and the wind resistance related parameters include the drag coefficient and at least one of the following: vehicle speed, air density, frontal area, and drag coefficient.
14. The method according to claim 10, characterized in that The energy consumption of the compressor of the vehicle corresponding to the first candidate control strategy is obtained in the following manner: Obtaining the power of the compressor corresponding to the first candidate control strategy; According to the power of the compressor corresponding to the first candidate control strategy, the energy consumption of the compressor corresponding to the first candidate control strategy is obtained.
15. The method according to claim 10, characterized in that The energy consumption of the cooling fan corresponding to the first candidate control strategy is obtained in the following manner: Obtaining the power of the cooling fan corresponding to the first candidate control strategy; According to the power of the cooling fan corresponding to the first candidate control strategy, the energy consumption of the cooling fan corresponding to the first candidate control strategy is obtained.
16. The method according to claim 15, characterized in that The obtaining the power of the cooling fan corresponding to the first candidate control strategy includes: Obtaining a candidate speed of the cooling fan corresponding to the first candidate control strategy; The power of the cooling fan corresponding to the candidate rotational speed is acquired according to a third mapping relationship, wherein the third mapping relationship includes a correspondence between the candidate rotational speed and the power of the cooling fan.
17. A thermal management system control device for a vehicle, characterized in that: The method applied to any one of claims 1 to 16.
18. A vehicle, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the steps of the method described in any one of claims 1 to 16 when executing the above instructions.
19. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 16 are implemented.
20. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 16.