Flying car oil cooler multi-physical field cooperative control method based on reinforcement learning

Through a multi-physics field collaborative control method based on reinforcement learning, dynamic optimization of the flying car oil cooler under different flight altitudes and environmental conditions is achieved, solving the problems of limited heat dissipation efficiency and energy waste in traditional control schemes, and improving the performance and reliability of the thermal management system.

CN120756657AActive Publication Date: 2025-10-10TIANJIN ZHONGRONG TIANYU AUTO PARTS CO LTD

Patent Information

Application Number
CN202510915396.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-10
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

The existing flying car oil cooler control scheme relies on the independent regulation of a single physical field, resulting in limited heat dissipation efficiency and an inability to meet the comprehensive heat dissipation needs across airspaces and multiple working conditions. In addition, the control logic is mismatched with the dynamic environment, resulting in an increased risk of excessive oil temperature and energy waste.

Method used

A multi-physics field collaborative control method based on reinforcement learning is adopted. By integrating power, environment and thermodynamic parameters, a multi-dimensional state vector with time and space correlation is constructed. Combined with the reinforcement learning algorithm, a three-dimensional continuous action space is designed. The fin opening, magnetic fluid flow rate and phase change material activation ratio are dynamically adjusted to achieve the coordinated optimization of magnetic fluid heat conduction, phase change material heat storage and air heat dissipation, and real-time control and online learning are carried out through edge computing nodes.

Benefits of technology

It significantly improves the thermal management efficiency of the flying car oil cooler in cross-airspace and multi-operating scenarios, takes into account both energy savings and equipment life extension, improves the system's environmental adaptability and control accuracy, and ensures oil temperature stability and safety.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to the technical field of hovercar oil coolers, in particular to a hovercar oil cooler multi-physics field cooperative control method based on reinforcement learning. A three-dimensional continuous action space (fin opening, magnetofluid flow velocity and phase change material activation proportion) is designed in combination with a reinforcement learning algorithm, and collaborative optimization of magnetofluid heat conduction, phase change material heat storage and air heat dissipation is realized in control of the oil cooler of the hovercar for the first time. According to the design, the limitation of traditional single physical field control is broken through, the thermal management efficiency of the system in cross-airspace and multi-working-condition scenes is remarkably improved, and meanwhile energy consumption is saved and the service life of equipment is prolonged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of flying car oil coolers, and in particular to a multi-physical field collaborative control method for a flying car oil cooler based on reinforcement learning. Background Art

[0002] With the continuous advancement of technology, traditional modes of transportation are no longer able to meet people's demands for travel efficiency and convenience. As an emerging mode of transportation, flying cars can effectively alleviate ground traffic congestion and improve travel efficiency. The oil cooler in a flying car is a core component of its thermal management system, responsible for efficiently dissipating heat generated by the powertrain (e.g., batteries, motors, and engines), while also adapting to the complex high-altitude, low-pressure, and inter-airspace operating conditions unique to flying cars. A magnetofluid-phase change coupled oil cooler uses a magnetic field to control the flow rate of a magnetic nanofluid, combined with phase change materials for heat storage. At high altitude and low pressure, the magnetofluid flow rate can be increased to twice that of conventional designs, compensating for heat loss through air. The phase change material absorbs transient heat during thermal shock, preventing oil temperature overshoot.

[0003] However, the existing flying car oil coolers have the following technical problems: 1. Traditional flying car oil cooler control schemes rely on the independent regulation of a single physical field (e.g., regulating only air cooling or magnetic fluid heat conduction), resulting in a lack of dynamic coordination between different cooling mechanisms. For example, in high-altitude, low-pressure scenarios, over-reliance on inefficient air cooling, coupled with a failure to dynamically activate the magnetic fluid and phase-change material for complementary optimization, limits cooling efficiency and accelerates equipment aging when a single component is fully loaded. This issue is essentially a conflict between local optimization and global thermal management requirements, and is unable to meet the comprehensive cooling requirements across multiple airspaces and operating conditions.

[0004] 2. Existing flying car oil cooler control systems use fixed rules or offline calibration parameters, without dynamically adjusting action weights based on flight altitude. For example, at high altitudes, the magnetic fluid flow rate is not prioritized to compensate for the decrease in air density, increasing the risk of oil temperature exceeding the specified limit. Furthermore, redundant fin openings are maintained even when humidity is excessively high at low altitudes, resulting in wasted energy. The root cause of this problem lies in the mismatch between the control logic and the dynamic environment. The system is unable to perceive and respond in real time to sudden changes in nonlinear environmental variables, such as low air pressure at high altitude and high humidity at low altitude. This results in low control accuracy and reduced energy efficiency.

[0005] Therefore, there is a need for an emergency blocking device that can solve the above problems. Summary of the Invention

[0006] The present invention aims to provide a multi-physics-field coordinated control method for a flying car oil cooler based on reinforcement learning. By integrating dynamic, environmental, and thermodynamic parameters to construct a spatiotemporally correlated multidimensional state vector, and combining it with a reinforcement learning algorithm to design a three-dimensional continuous action space, this method achieves the first coordinated optimization of magnetic fluid heat conduction, phase change material heat storage, and air heat dissipation in a flying car oil cooler. This design overcomes the limitations of traditional single-physics-field control, significantly improving the system's thermal management efficiency in cross-space and multi-operating scenarios while simultaneously balancing energy savings and equipment life extension.

[0007] The technical solution adopted by the present invention to solve the above technical problems is: a multi-physics field coordinated control method for a flying car oil cooler based on reinforcement learning, comprising the following steps: Step S1: The flying car sensor network collects power parameters, environmental parameters, and thermodynamic parameters in real time, normalizes the data, and constructs a multi-dimensional state vector with temporal and spatial correlation; Step S2: Constructing a reinforcement learning model based on the PPO algorithm, defining a state space including the multidimensional state vector, a three-dimensional continuous action space including fin opening, magnetic fluid flow rate, and phase change material activation ratio, and a reward function with oil temperature stability, energy efficiency, and equipment life as optimization objectives, and performing model training in a simulation environment; Step S3: Deploy the trained PPO model to the edge computing node, receive the real-time state vector output control parameters, dynamically adjust the action weights according to the flight altitude, and update the policy network through the online learning mechanism; Step S4: Execute the cooling strategy and handle exceptions: activate the phase change material to store or release heat according to the control parameters, trigger model replanning in response to sudden changes in thermal load, and switch to safety mode when the oil temperature deviation exceeds the preset threshold.

[0008] Furthermore, the step S1 specifically includes: Step S1-1: real-time acquisition of power parameters, including propeller speed and motor power; Step S1-2: synchronously collecting environmental parameters, including flight altitude, ambient temperature and humidity; Step S1-3: collecting thermodynamic parameters in parallel, including oil temperature at the inlet and outlet of the oil cooler, magnetic fluid flow rate, and phase change state of the phase change material; Step S1-4: Dynamically normalize the three types of parameters respectively; Step S1-5: The fused parameters are used to construct a multi-dimensional state vector containing a timestamp and a spatial position marker.

[0009] Furthermore, the model training in the simulation environment in step S2 includes: Step S2-1: Generate a low-altitude high-humidity working condition data set to simulate the impact of condensation effect on heat dissipation; Step S2-2: Construct a high-altitude low-pressure working condition data set to simulate the attenuation of heat dissipation efficiency caused by changes in air density; Step S2-3: Create transient thermal load impact scenarios when modules are separated and combined; Step S2-4: verifying the coupling relationship between the phase change rate of the phase change material and the thermal conductivity efficiency of the magnetic fluid.

[0010] Furthermore, the dynamic adjustment of action weights in step S3 includes: Step S3-1: When the flying car's altitude is greater than 1000 meters, the action weight of the magnetic fluid flow rate is increased by 30%-50%; Step S3-2: When the flying car's flight altitude is ≤300 meters, increase the decision priority of the fin opening action and add an energy consumption penalty factor.

[0011] Furthermore, triggering model replanning in step S4 includes: Step S4-1: monitor the oil temperature change rate in real time, and determine that the thermal load has suddenly changed when the change rate exceeds a threshold value α; Step S4-2: Freeze the current policy network and call the emergency control sequence of the historical scenario; Step S4-3: Using the current state as the initial point, perform short-term reinforcement learning rapid replanning; Step S4-4: Output the new control parameter combination and verify the stability.

[0012] Furthermore, the reward function construction of step S2 includes: Step S2-5: define the oil temperature stability index as the inverse of the variance between the inlet and outlet temperature difference and the set value; Step S2-6: defining the energy efficiency index as the weighted inverse of the sum of the magnetic fluid pump power and the heat sink drive power; Step S2-7: defining the device life index as a negative correlation function between the number of phase change cycles of the phase change material and the temperature fluctuation amplitude; Step S2-8: Use the entropy weight method to dynamically allocate weight coefficients of the three indicators.

[0013] Furthermore, the online learning mechanism in step S3 includes: Step S3-4: Storing the real-time flight data of the flying car into a circulating experience pool and marking abnormal operating condition data; Step S3-5: Trigger incremental training every time N pieces of data are accumulated, and prioritize replaying abnormal operating condition data; Step S3-6: Asynchronously updating the strategic network parameters using a dual network structure; Step S3-7: Prevent the policy update from deviating from the baseline model through the KL divergence constraint.

[0014] Furthermore, in step S4, activating the phase change material to store or release heat according to the control parameters, triggering model replanning in response to a sudden change in thermal load, and switching to a safe mode when the oil temperature deviation exceeds a preset threshold includes: S4-5: Activate safety mode when the oil temperature deviation is greater than ±3°C for 5 consecutive cycles; S4-6: Switch to the backup sensor group to verify data consistency; S4-7: Enable the preset PID controller to maintain basic cooling; S4-8: Generate a model diagnosis report and trigger a manual intervention flag.

[0015] Furthermore, the phase change state monitoring of the phase change material in step S1-3 includes: Step S1-3a: detecting the axial temperature gradient of the phase change material module by an embedded temperature sensor array; Step S1-3b: Determine the solid-liquid phase change ratio using the rate of change of the acoustic wave propagation velocity; Step S1-3c: quantize the phase change ratio into a continuous value of 0-1 and incorporate it into the state vector.

[0016] Furthermore, the energy consumption optimization in step S3-2 includes: Step S3-2a: Establishing a relationship model between the heat dissipation fin opening and the air resistance coefficient; Step S3-2b: Add a flight resistance penalty term to the reward function; Step S3-2c: Limit the selection probability of high-energy-consuming actions through the action mask mechanism.

[0017] The advantages of the present invention are: 1. This invention integrates dynamic, environmental, and thermodynamic parameters to construct a spatiotemporally correlated multidimensional state vector. Incorporating a reinforcement learning algorithm, this design employs a three-dimensional continuous action space (fin opening, magnetic fluid flow rate, and phase change material activation ratio). This achieves the first coordinated optimization of magnetic fluid heat conduction, phase change material heat storage, and air heat dissipation in a flying car oil cooler. This design transcends the limitations of traditional single-physics field control, significantly improving the system's thermal management efficiency across multiple airspaces and operating conditions, while simultaneously balancing energy savings and extending equipment life.

[0018] 2. To address the environmental differences between high altitude, low air pressure, and low altitude, high humidity, this paper proposes a dynamic adjustment mechanism for action weights based on flight altitude. This mechanism optimizes control parameters in real time through online learning at edge computing nodes. For example, at high altitude, the control priority of the magnetic fluid flow rate is increased to compensate for the impact of decreased air density. At low altitude, the decision weight for fin opening is increased and an energy penalty factor is introduced. This adaptive strategy effectively balances heat dissipation requirements and energy consumption during different flight phases, improving the system's environmental adaptability and control accuracy.

[0019] 3. This invention utilizes a short-term reinforcement learning replanning algorithm and a multi-level safety mode switching design. The system can rapidly generate new control strategies in response to sudden thermal load changes and seamlessly switch to safety mode when oil temperature exceeds tolerance. This mechanism not only ensures oil temperature stability under transient thermal shocks but also enables fault traceability through abnormal data labeling and manual intervention flags, comprehensively enhancing the safety and robustness of the flying car's thermal management system. DETAILED DESCRIPTION

[0020] The following is a clear and complete description of the technical solution of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.

[0021] Example 1: The present invention proposes a multi-physics field collaborative control method for a flying car oil cooler based on reinforcement learning, comprising the following steps: collecting power parameters, environmental parameters, and thermodynamic parameters in real time through a flying car sensor network, normalizing the data, and constructing a multidimensional state vector with spatiotemporal correlation; constructing a reinforcement learning model based on the PPO algorithm, defining a state space including the multidimensional state vector, a three-dimensional continuous action space including fin opening, magnetorheological fluid flow rate (referring to the flow rate of magnetorheological fluid in the oil cooler flow channel), and phase change material activation ratio, and a reward function with oil temperature stability, energy efficiency, and equipment life as optimization objectives, and training the model in a simulation environment; deploying the trained PPO model to an edge computing node, receiving the real-time state vector output control parameters, dynamically adjusting the action weights according to the flight altitude, and updating the policy network through an online learning mechanism; executing the cooling strategy and handling anomalies: activating the phase change material to store or release heat according to the control parameters, triggering model replanning in response to sudden changes in thermal load, and switching to a safe mode when the oil temperature deviation exceeds a preset threshold.

[0022] The flying car sensor network refers to a monitoring system composed of sensors distributed across key locations of the flying car. Specifically, these sensors can be implemented using temperature, humidity, pressure, and speed sensors, and are used to collect real-time operating parameters of the power system and external environmental parameters. Normalization involves converting monitoring data of varying dimensions into standardized values ​​of uniform dimensions. This can be achieved using a min-max scaling algorithm to eliminate the impact of magnitude differences between parameters on model training. The multidimensional state vector refers to a composite data structure that integrates time series and spatial position features. Specifically, it can be implemented using a tensor to store timestamped sensor data, constructing a state representation that reflects the dynamic changes in the system. The PPO algorithm is a policy gradient-based reinforcement learning method. Specifically, it uses a clipping objective function to implement policy updates, addressing the instability of traditional methods in training continuous action spaces. The three-dimensional continuous action space refers to the operational dimension that simultaneously controls the mechanical adjustment of the fins, the flow rate of the magnetic fluid, and the activation level of the phase change material. Specifically, a Gaussian distribution can be used to parameterize the action probability distribution, enabling coordinated control of multiple physical field parameters. The reward function is an evaluation metric that quantifies the effectiveness of oil cooler control. Specifically, an entropy weighting method can be used to dynamically assign weights for oil temperature stability, energy efficiency, and equipment lifespan, guiding the model to learn a multi-objective balancing strategy. The edge computing node refers to an embedded computing device deployed locally on the flying car. Specifically, a low-power GPU can be used for model inference to meet the low-latency requirements of real-time control. Dynamic action weight adjustment adaptively changes the priority of control parameters based on flight altitude. This can be achieved through a pre-trained altitude-heat dissipation efficiency mapping table to address the degradation of heat dissipation efficiency caused by low air pressure at high altitudes. The online learning mechanism is a training method that updates the policy network based on real-time data. Specifically, a recurring experience pool can be used to store abnormal operating condition data, improving the model's environmental adaptability through incremental training. The safety mode switch is a backup control logic activated in the event of an oil temperature anomaly. Specifically, a threshold trigger mechanism can be used to activate a pre-set PID controller to ensure the system's basic heat dissipation function in the event of model failure.

[0023] The core innovation of this invention lies in integrating the multi-physics field collaborative control of magnetofluid heat conduction, phase change heat storage and air heat dissipation through a reinforcement learning framework, combining dynamic weight adjustment of flight altitude with an online learning mechanism to achieve multi-objective optimization control of the oil cooler system under cross-airspace conditions. At the same time, an abnormal response and safety mode guarantee mechanism is constructed to solve the problems of insufficient environmental adaptability and safety of traditional single physical field control strategies.

[0024] The working process and principle of this invention are as follows: first, the flying car's sensor network collects power parameters, environmental parameters, and thermodynamic parameters in real time, normalizes these data, and constructs a multi-dimensional state vector with temporal and spatial correlation. This step achieves comprehensive perception of the flying car's operating status.

[0025] Next, a reinforcement learning model based on the PPO algorithm was constructed. A state space consisting of a multidimensional state vector was defined to reflect the current state of the system. A three-dimensional continuous action space consisting of fin opening, magnetic fluid flow rate, and phase change material activation ratio was defined as the control output. A reward function was defined with oil temperature stability, energy efficiency, and equipment life as optimization objectives to guide the model in learning the optimal control strategy. Model training was performed in a simulation environment to adapt the model to various operating conditions.

[0026] The trained PPO model is deployed to edge computing nodes for real-time control. The model receives a real-time state vector and outputs control parameters. Action weights are dynamically adjusted based on flight altitude to accommodate cooling requirements at varying altitudes. The policy network is continuously updated through online learning to maintain the state of the control strategy.

[0027] Finally, the cooling strategy is executed and abnormal conditions are handled. Active thermal management is achieved by activating the phase change material to store or release heat based on control parameters. When a sudden change in thermal load is detected, model replanning is triggered, rapidly adjusting the control strategy. If the oil temperature deviation exceeds a preset threshold, the system switches to safe mode to ensure flight safety.

[0028] The multi-physics field collaborative control method based on reinforcement learning proposed in the present invention realizes the dynamic collaborative optimization of magnetic fluid heat conduction, phase change material heat storage and air heat dissipation. It can adapt to the complex operating environment of flying cars across airspace and multiple working conditions, and improves thermal management efficiency and system reliability.

[0029] As a preferred embodiment, the solution of the present invention is specifically implemented as follows: The flying car's sensor network includes power parameter sensors, environmental parameter sensors, and thermodynamic parameter sensors. The power parameter sensors collect propeller speed and motor power data. The environmental parameter sensors collect flight altitude, ambient temperature, and humidity data. The thermodynamic parameter sensors collect data on the oil cooler inlet and outlet oil temperatures, magnetic fluid flow rate, and the phase change state of the phase change material.

[0030] Data normalization uses the min-max normalization method to map each parameter to the interval [0, 1]. The multidimensional state vector of spatiotemporal association contains a timestamp and a spatial position marker, forming a 16-dimensional vector.

[0031] The PPO algorithm uses an actor-critic architecture. The actor network outputs action probability distributions, while the critic network estimates state values. The state space is 16-dimensional, and the action space is a 3-dimensional continuous space. The reward function is R=w1S+w2E+w3*L, where S is the oil temperature stability indicator, E is the energy efficiency indicator, L is the equipment lifespan indicator, and w1, w2, and w3 are dynamic weights.

[0032] The simulation environment construction includes low-altitude high-humidity working conditions, high-altitude low-pressure working conditions, and transient thermal load impact scenarios when the modules are separated and combined. The training adopts an asynchronous parallel mode, and a single training iteration is 1 million steps.

[0033] The edge computing node adopts an embedded GPU platform, receives a state vector every 100 ms, and outputs control parameters. When the flight height is greater than 1000 meters, the action weight of the magnetic fluid flow rate is increased by 40%. When the flight height is less than or equal to 300 meters, the decision priority of the fin opening action is increased.

[0034] The online learning mechanism adopts the experience replay technology and stores the last 10,000 flight data. Incremental training is triggered every 1000 new data accumulations, and the policy network parameters are updated.

[0035] The thermal load mutation determination standard is that the oil temperature change rate exceeds 5℃ / min. The model re-planning adopts a short-time domain PPO algorithm, and the prediction time domain is 10s. The safety mode switching condition is that the oil temperature deviation is greater than ±3℃ in the last 5 cycles.

[0036] Through the above scheme, the present application realizes the cooperative control of the oil cooler of the flying car in multiple physical fields. The reinforcement learning model can adaptively adjust the fin opening, the magnetic fluid flow rate and the phase change material activation ratio, and keep the oil temperature stable under different flight altitudes and environmental conditions. The dynamically adjusted control strategy effectively deals with complex working conditions such as high-altitude low-pressure and low-altitude high-humidity, and improves the heat dissipation efficiency. The online learning mechanism enables the system to continuously optimize the control strategy and adapt to the performance changes during the long-term use of the flying car. The multi-level abnormal processing mechanism enhances the robustness of the system and effectively deals with abnormal conditions such as thermal load mutation, ensuring flight safety. The method breaks through the limitations of traditional single physical field control, and significantly improves the performance and reliability of the thermal management system of the flying car.

[0037] The present application further proposes a real-time data acquisition and processing method comprising the following steps: real-time acquisition of power parameters, including propeller speed and motor power; synchronous acquisition of environmental parameters, including flight altitude, environmental temperature and humidity; parallel acquisition of thermodynamic parameters, including oil cooler inlet and outlet oil temperature, magnetic fluid flow rate and phase change state of phase change material; dynamic normalization processing of the three types of parameters respectively; and fusion processing of the parameters to construct a multi-dimensional state vector containing time stamp and spatial position marker.

[0038] Dynamic parameters are sampled at the millisecond level using Hall sensors and power meters. Environmental parameters are captured synchronously using air pressure sensors and temperature and humidity sensors. Thermodynamic parameters are acquired in parallel using distributed temperature probes and ultrasonic flow meters. Dynamic normalization uses extreme value mapping within a sliding time window to eliminate sensor range variations. Timestamps are accurate to the millisecond level, and spatial location tags include the topological coordinates of the oil cooler module.

[0039] Specifically, propeller speed data is collected at a frequency of 500 Hz using a magnetic encoder, and motor power data is calculated using current transformers and voltage sensors. Flight altitude data is updated at a frequency of 50 Hz using a barometric altimeter, and ambient temperature and humidity data are sampled simultaneously using a digital sensor. Oil cooler inlet and outlet oil temperatures are measured using platinum resistance temperature sensors, magnetic fluid flow velocity is detected using a Doppler ultrasonic flowmeter, and the phase change material state is determined by combining temperature gradient and acoustic wave propagation velocity. During dynamic normalization, propeller speed is mapped to a range of 0-1, motor power is converted as a percentage of maximum rated power, and flight altitude is linearly scaled within a preset airspace range. Timestamps and spatial location markers are automatically generated using a hardware clock synchronization module and the device topology map. The resulting 128-dimensional state vector consists of eight dynamic parameter channels, five environmental parameter channels, six thermodynamic parameter channels, and a spatiotemporal marker. This multidimensional state vector is transmitted to edge computing nodes via Gigabit Ethernet and serves as the input data source for the reinforcement learning model, ensuring spatiotemporal consistency of the different physical quantities.

[0040] As a preferred embodiment, the solution of the present invention is specifically implemented as follows: Step S1-1: Real-time acquisition of power parameters, including propeller speed and motor power. Specifically, the propeller speed is measured using a Hall effect sensor mounted on the propeller shaft, with a sampling frequency of 100 Hz. The motor power is calculated using the current and voltage sensors built into the motor controller, with a sampling frequency of 200 Hz.

[0041] Step S1-2: Synchronously collect environmental parameters, including flight altitude, ambient temperature, and humidity. Furthermore, a pressure sensor is used to measure flight altitude with a sampling frequency of 10 Hz; a temperature and humidity composite sensor is used to measure ambient temperature and humidity with a sampling frequency of 1 Hz.

[0042] Step S1-3: Parallel collection of thermodynamic parameters, including the oil cooler inlet and outlet oil temperatures, the magnetic fluid flow rate, and the phase change state of the phase change material. The oil cooler inlet and outlet oil temperatures were measured using thermocouples at a sampling frequency of 5 Hz; the magnetic fluid flow rate was measured using an electromagnetic flowmeter at a sampling frequency of 50 Hz; and the phase change state of the phase change material was detected using differential scanning calorimetry at a sampling frequency of 0.1 Hz.

[0043] Step S1-4: Dynamically normalize each of the three parameters. Specifically, a sliding window method is used to calculate the local maximum and minimum values ​​of each parameter, mapping the raw data to the [0, 1] interval. The window size is dynamically adjusted based on the parameter's changing characteristics. For example, a 60-second window is used for propeller speed and a 600-second window is used for ambient temperature.

[0044] Step S1-5: The fused parameters are used to construct a multidimensional state vector containing timestamps and spatial location markers. The normalized parameters are then combined into a vector in a predefined order, with the Unix timestamp and GPS coordinates added as spatiotemporal markers. The resulting state vector has 15 dimensions, including three dynamic parameters, three environmental parameters, four thermodynamic parameters, one timestamp, and three spatial coordinates.

[0045] Through the above-mentioned technical solution, the present invention achieves real-time acquisition, normalization, and spatiotemporal correlation fusion of multi-source heterogeneous data for flying vehicles. This method improves the timeliness and consistency of data, providing high-quality input features for subsequent reinforcement learning models. Furthermore, dynamic normalization enhances the system's adaptability to parameter variations under different operating conditions, while spatiotemporal labeling provides the model with important contextual information, helping to capture long-term dependencies and spatial correlations.

[0046] The present invention further proposes model training in a simulation environment, including generating a low-altitude, high-humidity working condition data set to simulate the impact of the condensation effect on heat dissipation, constructing a high-altitude, low-pressure working condition data set to simulate the attenuation of heat dissipation efficiency caused by changes in air density, creating transient thermal load impact scenarios when modules are separated and combined, and verifying the coupling relationship between the phase change rate of phase change materials and the thermal conductivity efficiency of magnetic fluids.

[0047] When generating a dataset for low-altitude, high-humidity operating conditions, the condensate film formation process is simulated by adjusting the combined parameters of ambient temperature and humidity to quantify the degree of attenuation of the heat transfer coefficient on the heat sink surface. When constructing a dataset for high-altitude, low-pressure operating conditions, the air density variation curve with altitude is calculated based on the gas state equation, and the Reynolds number correction coefficient in the heat sink aerodynamic model is dynamically adjusted. When creating a transient thermal load shock scenario, a step function is used to simulate the sudden change in power of the power system, and the transient response curve of the oil cooler inlet and outlet oil temperatures is recorded. When verifying the coupling relationship, a nonlinear mapping model of the heat conduction rate of the two is established by changing the ratio of the magnetic fluid flow rate and the activation ratio of the phase change material.

[0048] Specifically, in low-altitude, high-humidity simulations, the humidity parameter was set to a range of 80%-95%. The condensate film thickness was updated in real time via the surface thermal resistance calculation module, reducing the fin heat dissipation efficiency by 15%-30%. In high-altitude, low-pressure simulations, the flight altitude was set to a range of 1,000-5,000 meters, and the air density was dynamically adjusted according to the International Standard Atmospheric Model. The fin heat dissipation efficiency decayed exponentially with altitude. In a transient thermal load shock scenario, the power system power suddenly increased by 200% within 0.5 seconds. The oil cooler heat load response curve was generated through finite element thermal simulation and used to train the reinforcement learning model's ability to suppress rapid thermal disturbances. In coupling relationship verification, the magnetic fluid flow rate was controlled within the range of 0.5-2.5 m / s, and the phase change material activation ratio was increased in steps of 10%. A heat flux density sensor was used to collect the comprehensive thermal conductivity of the two components under synergistic action, and a multidimensional parameter lookup table was established for the strategy network to learn.

[0049] As a preferred embodiment, the solution of the present invention is specifically implemented as follows: When training the model in a simulation environment, a dataset of low-altitude, high-humidity conditions was first generated to simulate the impact of condensation on heat dissipation. Specifically, computer simulations generated environmental parameters at altitudes of 200 to 500 meters and relative humidity of 80% to 95%. Combined with the geometric model of the flying car's oil cooler, the formation and evaporation of condensed water on the fin surface were calculated, and its impact on heat dissipation efficiency was analyzed.

[0050] Furthermore, a dataset for high-altitude, low-pressure operating conditions was constructed to simulate the degradation of heat dissipation efficiency caused by changes in air density. For example, the system simulated the change in environmental parameters from a flight altitude of 1,000 to 5,000 meters, calculated the air density and thermal conductivity at different altitudes, and evaluated their impact on the oil cooler's heat dissipation performance.

[0051] This allows for the creation of transient thermal load scenarios when modules are separated and combined. Specifically, the system simulates the sudden change in thermal load on the powertrain when switching between ground-based driving and aerial flight modes, analyzing the transient response characteristics of the oil temperature.

[0052] Finally, the coupling relationship between the phase change rate of the phase change material and the thermal conductivity of the magnetic fluid was verified. For example, using thermodynamic simulation software, a phase change kinetic model of the phase change material and a flow heat conduction model of the magnetic fluid were established. This allowed the influence of the magnetic fluid flow rate on the phase change process of the phase change material under different magnetic field intensities to be analyzed, as well as the contribution of the phase change latent heat of the phase change material to the overall heat dissipation efficiency of the system.

[0053] Through the above technical solution, the present invention realizes the simulation and evaluation of the performance of the flying car oil cooler under a variety of complex working conditions. By generating low-altitude high humidity and high-altitude low-pressure working condition data sets, the heat dissipation effect of the flying car under different environmental conditions is simulated, providing data support for the optimization of subsequent control strategies. Creating transient thermal load impact scenarios when modules are separated and combined helps improve the system's response to sudden thermal loads. Verifying the coupling relationship between the phase change rate of phase change materials and the thermal conductivity efficiency of magnetic fluids provides a theoretical basis for multi-physical field collaborative control. The simulation training process in the present invention significantly improves the adaptability and control accuracy of the reinforcement learning model to complex environments, laying the foundation for the practical application of flying car oil coolers.

[0054] The present invention further proposes that when the flying car's flight altitude is greater than 1,000 meters, the action weight of the magnetic fluid flow rate is increased by 30%-50%; when the flight altitude is less than or equal to 300 meters, the decision priority of the fin opening action is increased and an energy consumption penalty factor is added.

[0055] Among them, the adjustment of the magnetic fluid flow rate weight is achieved through the dynamic scaling of the policy network parameters in the edge computing node, and the weight coefficient is updated in real time based on the flight altitude sensor data. When the flight altitude exceeds 1000 meters, the air density decreases, resulting in a decrease in air heat dissipation efficiency. At this time, the action weight increase range is determined through simulation experiments, and the power share of the magnetic fluid pump is increased first to maintain thermal efficiency. When the flight altitude is lower than 300 meters, the priority of the heat dissipation fin opening is increased using an action mask mechanism. The energy consumption penalty factor is dynamically calculated based on the positive correlation between the fin opening value and the air drag coefficient. For example, for every 10% increase in the fin opening, the drag coefficient increases by 0.15, and the energy consumption penalty item in the corresponding reward function increases by 8%.

[0056] Specifically, when the flight altitude reaches above 1,000 meters, the edge computing node receives the altitude sensor signal, triggering the weight adjustment module to apply a 1.3-1.5x gain to the magnetic fluid flow rate action item output by the policy network. Through online reinforcement learning, the magnetic fluid flow rate control instruction is optimized to increase it to 1.8-2.2 times that of normal operating conditions in a low-pressure environment, compensating for air heat dissipation losses. When the flight altitude drops below 300 meters, the policy network prioritizes the fin opening adjustment action. At the same time, a secondary energy consumption evaluation function is superimposed on the fin opening instruction based on real-time humidity data. For example, when the humidity exceeds 85%, the energy consumption penalty factor increases by 0.2 for every 5% increase in fin opening, suppressing the occurrence of redundant actions. This mechanism is verified through joint training of a high-altitude, low-pressure dataset and a low-altitude, high-humidity dataset in a simulation environment to ensure that the weight adjustment parameters match the physical constraints of the flight conditions.

[0057] As a preferred embodiment, the solution of the present invention is specifically implemented as follows: In the method for the multi-physical field collaborative control of the air car oil cooler, the dynamic adjustment of the action weight comprises the following steps: When the flight height of the air car exceeds 1000 meters, the action weight of the magnetic fluid flow rate is increased by 40%. Specifically, the edge computing node receives the height sensor data in real time, and once it is detected that the flight height exceeds the threshold of 1000 meters, the weight adjustment algorithm is triggered immediately. The algorithm multiplies the original weight of the magnetic fluid flow rate by a factor of 1.4, thereby giving the magnetic fluid flow rate a higher priority in the decision-making process.

[0058] Further, when the flight height of the air car does not exceed 300 meters, the decision priority of the wing opening action is increased and an energy consumption penalty factor is added. In specific implementation, first, the original weight of the wing opening action is increased by 20%. Second, an energy consumption penalty term is introduced in the reward function, which is positively correlated with the wing opening. For example, the penalty factor can be set as 0.05 x wing opening percentage. In this way, the system will tend to adjust the wing opening moderately when flying at low altitude, while avoiding excessive opening that leads to energy waste.

[0059] Through the above technical solutions, the present application realizes the adaptive adjustment of the air car oil cooler control system to different flight heights. In high-altitude environment, by increasing the regulation weight of the magnetic fluid flow rate, the influence of the decrease of air density on the heat dissipation efficiency is effectively compensated. When flying at low altitude, the decision weight of the wing opening is increased and an energy consumption penalty mechanism is introduced, which not only ensures the heat dissipation effect, but also avoids unnecessary energy consumption. The dynamic weight adjustment strategy based on flight height proposed by the present application significantly improves the environmental adaptability and control accuracy of the oil cooler under cross-air-domain working conditions, and realizes the balance optimization of heat dissipation effect and energy utilization.

[0060] The present application further proposes a trigger model re-planning, which comprises: monitoring the oil temperature change rate in real time, determining thermal load mutation when the change rate exceeds the threshold a; freezing the current strategy network, calling the emergency control sequence of the historical scene; taking the current state as the initial point, performing short-time domain reinforcement learning for fast re-planning; outputting a new control parameter combination and verifying the stability.

[0061] Among them, the oil temperature change rate threshold a is dynamically adjusted according to the flight stage, and is set to 3℃ / s in high-altitude working conditions and 5℃ / s in low-altitude working conditions; the emergency control sequence is stored in the non-volatile memory of the edge computing node, and contains typical thermal shock scene data in different height intervals; the short-time domain reinforcement learning adopts time step compression technology to shorten the regular training period to 20% of the original period, and generates a candidate action sequence through the Monte Carlo tree search algorithm.

[0062] Specifically, when the flying car encounters a sudden increase in propeller power during its climb phase, causing the oil temperature to change at a rate of 4°C / s, the monitoring module triggers a thermal load mutation detection. At this point, the policy network freeze module immediately halts the current policy's gradient update to prevent policy drift. The edge computing node calls the stored low-altitude emergency control sequence and loads basic control parameters, including full fin opening and maximum ferrofluid flow rate. The reinforcement learning replanning module uses the current oil temperature, ferrofluid flow rate, and phase change material state as initial inputs, completing 100 policy iterations within 50ms and generating a new control parameter combination with an 82% fin opening, a 2.5 m / s ferrofluid flow rate, and a 65% phase change material activation ratio. During the verification phase, the digital twin system simulates the oil temperature curve over the next three seconds. After confirming that the new parameters can control the oil temperature deviation within ±1.5°C, they are then distributed to the actuators. This process enables the system to complete the entire process, from anomaly detection to policy update, within 200ms, reducing the response time by 76% compared to traditional PID control.

[0063] As a preferred embodiment, the solution of the present invention is specifically implemented as follows: The oil temperature change rate is monitored in real time. A thermal load mutation is detected when the rate exceeds a preset threshold α. For example, set α to 2°C / minute, sample the oil temperature every 10 seconds, and calculate the temperature difference between two consecutive samples divided by the time interval to determine the rate of change. If the rate of change exceeds α for three consecutive samples, a thermal load mutation is detected.

[0064] Freeze the current policy network and invoke the emergency control sequence for the historical scenario. Specifically, the parameters of the current policy network are saved as a checkpoint. Simultaneously, the historical scenario most similar to the current state is retrieved from the pre-stored emergency control database and the corresponding control sequence is extracted. For example, if the current state is a high-altitude rapid climb, the corresponding emergency control sequence is invoked, such as increasing the magnetic fluid flow rate to its maximum value and activating all phase change material modules.

[0065] Using the current state as the initial point, a short-term reinforcement learning fast replanning is performed. Furthermore, a truncated temporal difference learning method is used, using the current state as the initial state, to perform a fast policy search over the next 5-10 time steps. This generates a series of candidate action sequences, and the long-term reward of each sequence is evaluated through Monte Carlo tree search.

[0066] The new control parameter combination is output and its stability is verified. Specifically, the action sequence with the highest evaluation score is selected as the new control strategy. The stability of this strategy is then verified in a simulation environment, including checking whether the oil temperature converges to the target range and whether energy consumption meets the constraints. If verification passes, the new strategy is deployed to the actual system; otherwise, the plan is repeated again.

[0067] Through the above-mentioned technical solution, the present invention achieves rapid response and precise control of sudden thermal load changes. By employing short-term reinforcement learning for rapid replanning, the system can generate new control strategies within milliseconds, effectively preventing oil temperature overshoot and equipment damage. Furthermore, by freezing the current strategy network and invoking historical emergency sequences, system stability is ensured during the replanning process. Furthermore, simulation verification of the new strategy further enhances control reliability and safety.

[0068] The present invention further proposes a reward function construction scheme including the following steps: defining the oil temperature stability index as the inverse of the variance of the inlet and outlet temperature difference and the set value; defining the energy efficiency index as the inverse of the weighted sum of the magnetic fluid pump power and the heat sink fin drive power; defining the equipment life index as a negative correlation function of the number of phase change cycles of the phase change material and the temperature fluctuation amplitude; and using the entropy weight method to dynamically allocate the weight coefficients of the three indicators.

[0069] Among them, the oil temperature stability index quantifies the temperature control accuracy by calculating the inverse of the square difference between the actual temperature difference and the target temperature difference; the energy efficiency index uses the weighted sum and inverse of the driving power of the magnetohydrodynamic pump and the heat sink to reflect the comprehensive impact of energy consumption of different actuators; the equipment life index establishes a linear negative correlation function between the number of phase change cycles and temperature fluctuations, incorporating material fatigue loss into the optimization target; the entropy weight method calculates the information entropy value of each indicator based on real-time operating data, and dynamically adjusts the weight coefficient to adapt to the needs of the current flight phase.

[0070] Specifically, during the operation of the magnetic fluid pump, when low air pressure at high altitudes causes the efficiency of the cooling fins to decrease, the system automatically increases the weight coefficient of the oil temperature stability indicator through the entropy weighting method, prioritizing temperature differential control. In low-altitude, high-humidity environments, the algorithm increases the entropy value calculation based on the fluctuation range of the magnetic fluid pump power data, reducing the weight of the energy efficiency indicator to avoid overload operation. After each complete phase change cycle of the phase change material module, the output value of the equipment life indicator function decreases by 0.05. This value is determined by the experimentally measured fatigue coefficient of the phase change material. The entropy weighting method recalculates the indicator weights every 30 seconds, and triggers an immediate weight update when the ambient temperature sensor detects a sudden temperature change exceeding 5°C, ensuring that the weight distribution matches the real-time heat load.

[0071] As a preferred embodiment, the solution of the present invention is specifically implemented as follows: Step S2-5: Define the oil temperature stability index as the inverse of the variance between the inlet and outlet temperature difference and the set value. Specifically, collect the oil cooler inlet temperature T in and outlet temperature T out , calculate the temperature difference ΔT=T in -T out . Set the target temperature difference ΔT target , calculate the variance within N sampling periods σ^2=Σ(ΔT-ΔTtarget = 1 / N. The oil temperature stability index S temp = 1 / σ^2.

[0072] Step S2-6: define the energy consumption efficiency index as the weighted sum reciprocal of the magnetic fluid pump power and the heat dissipation fin driving power. Specifically, measure the magnetic fluid pump power P pump and the heat dissipation fin driving power P fin , calculate the weighted sum P total = w1P pump + w2P fin , where w1 and w2 are weight coefficients. The energy consumption efficiency index S energy = 1 / P total .

[0073] Step S2-7: define the device life index as a negative correlation function of the number of phase change cycles of the phase change material and the temperature fluctuation amplitude. Specifically, record the number of phase change cycles N cycle of the phase change material and the temperature fluctuation amplitude ΔT pcm . The device life index S life = exp(-kN cycle ΔT pcm ), where k is an attenuation coefficient.

[0074] Step S2-8: dynamically allocate the weight coefficients of the three indexes by using the entropy weight method. Calculate the information entropy of each index, determine the weight according to the information entropy, and obtain the comprehensive reward function R = αS temp + βS energy + γ * S life , where α, β, γ are dynamically adjusted weight coefficients.

[0075] Through the above technical solutions, the present application realizes comprehensive optimization of oil temperature stability, energy consumption efficiency and device life. The oil temperature stability is measured by variance reciprocal, which can effectively reflect the temperature fluctuation. The energy consumption efficiency is evaluated by weighted sum reciprocal, which takes into account the power consumption of the magnetic fluid pump and the heat dissipation fin. The device life index considers the number of cycles and temperature fluctuations of the phase change material, reflecting the influence of thermal fatigue on life. The entropy weight method dynamically allocates the weight, so that the reward function can adapt to the optimization needs under different working conditions, improving the flexibility and adaptability of the control strategy.

[0076] The present application further proposes a technical solution of storing real-time flight data of the flying car to a cycle experience pool and marking abnormal working condition data; triggering incremental training every N data accumulation and preferentially replaying abnormal working condition data; adopting a double network structure asynchronous update strategy network parameter; and preventing policy update from deviating from the benchmark model through KL divergence constraint.

[0077] The recurrent experience pool utilizes a circular buffer structure, with data storage capacity set to the last 24 hours of flight data. Abnormal operating condition data is automatically flagged using a preset oil temperature change rate threshold. Incremental training uses N=1000 as a trigger, increasing the probability of abnormal data replay to three times that of normal data. The dual-network architecture comprises an online network and a target network. The online network parameters are updated via gradient descent every 5 seconds, while the target network parameters are synchronized with the online network parameters every 30 minutes. A KL divergence constraint of 0.01 is used to maintain model stability during policy updates by constraining policy distribution differences.

[0078] Specifically, the real-time data generated during the flight is stored in a cyclic experience pool, and the abnormal condition data is stored in independent partitions. When the cumulative data volume reaches 1,000, the training module extracts batch data from the experience pool, of which abnormal data is sampled preferentially at a ratio of 75%. The online network receives the sampled data to calculate the gradient and update the parameters. The target network maintains the stability of the current control strategy and avoids policy oscillations through periodic synchronization. During the parameter update process, the KL divergence of the new and old strategies is calculated. When the divergence value exceeds 0.01, the learning rate is automatically reduced to ensure that the policy update amplitude is controlled. This mechanism enables the model to continuously absorb new data features, while improving its response capabilities to sudden conditions through abnormal data reinforcement learning. The dual network structure ensures that the online update process does not affect the execution of real-time control tasks.

[0079] As a preferred embodiment, the solution of the present invention is specifically implemented as follows: The online learning mechanism includes the following steps: Step S3-4: The flying car's real-time flight data is stored in a circulating experience pool, and abnormal operating condition data is marked. The circulating experience pool uses a first-in-first-out data structure and has a capacity of 10,000 data entries. Abnormal operating condition data includes conditions such as oil temperature exceeding 85°C and magnetic fluid flow rate below 0.5m / s.

[0080] Step S3-5: Trigger incremental training every time 1,000 new data points are accumulated, prioritizing the replay of abnormal operating condition data. Incremental training uses a mini-batch approach, with each batch containing 64 data points, of which abnormal operating condition data accounts for no less than 30%.

[0081] Step S3-6: A dual-network architecture is used to asynchronously update the policy network parameters. The primary network is used for online decision-making, and the target network is used to calculate the target Q value. The target network parameters are updated every 100 training steps using a soft update method with an update rate of 0.01.

[0082] Step S3-7: Prevent the policy update from deviating from the baseline model by using the KL divergence constraint. The KL divergence threshold is set to 0.01. When the KL divergence exceeds the threshold, the policy gradient is clipped.

[0083] Through the above-mentioned technical solution, the present invention achieves continuous optimization and adaptive capabilities for the flying car oil cooler control system. By utilizing a circulating experience pool and a prioritized replay mechanism, the system efficiently learns control strategies under abnormal operating conditions. The dual-network architecture and KL divergence constraints ensure the stability of policy updates, preventing excessive deviation from the validated baseline model. The online learning mechanism enables the control system to adapt to complex and changing flight environments, continuously improving control accuracy and robustness, thereby enhancing the safety and reliability of the flying car.

[0084] The present invention further proposes to activate the safety mode when the oil temperature deviation exceeds ±3°C in five consecutive cycles, switch to the backup sensor group to verify data consistency, enable the preset PID controller to maintain basic cooling, generate a model diagnosis report and trigger the manual intervention flag.

[0085] Among them, the trigger condition of the safety mode is set to five consecutive cycles of oil temperature deviation monitoring to avoid false triggering caused by transient interference; the backup sensor group adopts a redundant design to eliminate the impact of a single sensor failure through cross-comparison; the preset PID controller parameters are calibrated according to historical operating data, which can provide basic temperature regulation capabilities when the reinforcement learning model fails; the model diagnosis report contains the status parameters and control instruction sequence of the abnormal time period, providing data support for manual intervention.

[0086] Specifically, when the oil cooler control system's real-time monitoring module detects oil temperature deviations exceeding ±3°C for five consecutive control cycles, it determines the current control strategy has failed and automatically triggers a safe mode switch. During this process, the system first switches to an independently powered backup sensor group. By comparing the oil temperature data from the primary and backup sensors, it can eliminate misjudgments caused by sensor failure. Once the data verification passes, the system disables the control output of the reinforcement learning model and activates the PID control algorithm pre-stored in the edge computing node. This algorithm dynamically adjusts the output power of the cooling actuator based on the deviation between the current oil temperature and the setpoint to maintain basic heat dissipation. Simultaneously, the system compiles the operating data for ten cycles before and after the anomaly trigger to generate a diagnostic report, marking the anomaly with a timestamp and triggering a manual intervention prompt on the human-machine interface. The threshold for the number of consecutive cycles is set by combining the control cycle length and the thermal inertia of the oil cooler to ensure timely and accurate anomaly detection. The backup sensor group is spatially distributed across key temperature measurement points on the oil cooler to prevent misjudgments caused by local temperature failures. The PID controller parameters are pre-stored in two sets, corresponding to typical high-altitude and low-altitude operating conditions, and are automatically adjusted based on the actual flight altitude during switchover.

[0087] As a preferred embodiment, the present invention is implemented as follows: When the real-time monitoring module of the flying car's oil cooler detects an absolute oil temperature deviation exceeding 3°C for five consecutive control cycles, the system automatically triggers a safe mode switch. First, the main control unit shuts off the control signal output of the reinforcement learning model and activates the redundant backup sensor set for data acquisition. The oil temperature and flow rate parameters acquired by the primary and backup sensors are compared. If the data difference is within a preset tolerance, the abnormal condition is confirmed to be valid. Subsequently, the basic cooling control module switches to a pre-set PID controller, dynamically adjusting the magnetic fluid pump's baseline speed and heat sink fin opening based on the current oil temperature deviation to maintain the oil cooler's basic heat dissipation function. Simultaneously, the fault diagnosis module generates a historical data report containing the abnormality timestamp, environmental parameters, and control parameters. A red warning flag is triggered through the human-computer interface, prompting the operator to intervene and check the equipment's operating status.

[0088] Through the above technical solution, the present invention can quickly switch to a reliable basic control mode when the oil temperature continues to exceed the tolerance, avoiding the risk of equipment overheating or overcooling due to abnormal model output. The cross-check mechanism of the backup sensor effectively eliminates the interference of single-point data failure and ensures the accuracy of the safety mode triggering. The preset PID controller maintains the basic heat dissipation capacity through proportional-integral-differential adjustment, buying processing time for manual intervention. The generation of abnormal data reports and the triggering of warning signs realize the traceability of the fault process, facilitating subsequent fault location and system optimization.

[0089] The present invention further proposes specific steps for monitoring the phase change state of phase change materials, including detecting the axial temperature gradient of the phase change material module through an embedded temperature sensor array; determining the solid-liquid phase change ratio by using the rate of change of the sound wave propagation velocity; and quantifying the phase change ratio into a continuous value of 0-1 and incorporating it into a state vector.

[0090] An array of embedded temperature sensors is distributed at equal intervals along the axis of the phase change material module. Multi-point temperature acquisition is used to construct an axial temperature distribution curve, where the temperature gradient reflects the heat transfer direction during the phase change process. The rate of change in acoustic wave propagation velocity is determined by transmitting ultrasonic waves through a piezoelectric transducer and receiving echo signals. The difference in acoustic velocity between the solid and liquid phases is converted into a propagation time difference, which is then combined with the temperature gradient data to calculate the phase change interface position. The phase change ratio is quantified by weighted fusion of the axial temperature gradient and acoustic wave propagation velocity data to generate a continuous value from 0 to 1 to represent the overall phase change degree, where a completely solid state corresponds to 0 and a completely liquid state corresponds to 1.

[0091] Specifically, after the temperature sensor array detects and obtains the axial temperature distribution data, the cubic spline interpolation algorithm is used to fit the temperature gradient curve to determine the starting and ending positions of the phase change region. The time difference method is used to measure the acoustic wave propagation velocity. By comparing the acoustic wave propagation time difference of the solid and liquid phase change materials under the same path, the solid-liquid phase change ratio of the path is calculated. The phase change ratio and temperature gradient data of each path are input into the fuzzy logic system to generate a comprehensive phase change ratio quantization value. This quantization value is embedded in the multidimensional state vector for the reinforcement learning model to obtain the dynamic heat storage capacity of the phase change material in real time. The continuous nature of the quantization value avoids the control lag problem caused by traditional discrete processing. For example, when the quantization value reaches 0.8, the reinforcement learning model will prioritize activating the heat storage function of the phase change material to cope with the sudden change in heat load, and at the same time adjust the flow rate of the magnetic fluid to compensate for the heat dissipation demand.

[0092] A preferred embodiment is implemented as follows: three temperature sensor groups are arranged at equal intervals along the axial direction in the phase change material module, and each group contains four platinum resistance temperature sensors distributed in a ring. When the oil cooler is working, the temperature data of the phase change material at different axial positions are synchronously collected through the sensor group, and the temperature gradient difference between adjacent sensor groups is calculated. At the same time, ultrasonic transmitters and receivers are installed at both ends of the module to transmit pulse signals at a frequency of 1 MHz to measure the propagation time of the sound wave in the phase change material. According to the experimentally calibrated solid-liquid two-phase sound velocity curve, the measured sound velocity change rate is converted into a phase change ratio value. After the temperature gradient data and the sound velocity data are fused by Kalman filtering, a continuous phase variable in the range of 0-1 is generated, where 0 represents a completely solid state and 1 represents a completely liquid state. The phase variable is input into the state vector of the reinforcement learning model at an update frequency of 10 times per second.

[0093] Through the above-mentioned technical solution, the present invention achieves precise quantitative monitoring of the phase change state of phase change materials, resolving the control lag problem caused by traditional binary state judgment. By fusing multimodal data from temperature gradients and sound velocity changes, the material state in the phase transition zone is accurately identified, avoiding the risk of misjudgment by a single sensor at the phase transition critical point. This monitoring method provides continuous state input to the reinforcement learning model, improving the control accuracy of the activation ratio of the phase change material. This allows for timely response to transient thermal shocks while avoiding material fatigue damage caused by frequent phase changes.

[0094] The present invention further proposes that energy consumption optimization includes: establishing a relationship model between the opening of the heat sink fins and the air drag coefficient; adding a flight resistance penalty term to the reward function; and limiting the selection probability of high-energy-consuming actions through an action mask mechanism.

[0095] Among them, the relationship model between the opening of the cooling fin and the air drag coefficient is obtained by fitting the wind tunnel test data. Its function form is a quadratic polynomial relationship. Every 10% increase in the opening leads to an increase in the drag coefficient by 0.12-0.15; the flight drag penalty term is constructed as the product of the air drag coefficient and the square of the flight speed to dynamically reflect the actual flight energy consumption; the action mask mechanism is based on the real-time calculated drag-opening curve, and imposes a probability threshold limit on the action output when the opening value is in the drag-sensitive range.

[0096] Specifically, at flight altitudes ≤300 meters, increasing the cooling fin opening to above 40% causes a sharp increase in the drag coefficient. By establishing an opening-drag model and inputting the drag coefficient change into the policy network during the reinforcement learning model training phase, the action decision-making process automatically avoids high-drag intervals. The flight drag penalty term introduced in the reward function has a weight coefficient set between 0.3 and 0.5, forming a composite energy consumption indicator with the magnetohydrodynamic pump power, forcing the model to simultaneously optimize flight energy consumption when adjusting the fin opening. An action masking mechanism dynamically calculates the upper opening threshold based on the current flight speed. When the flight speed reaches 15 m / s, the range of opening actions selected is limited to 0-35%. High-drag actions are suppressed by modifying the probability distribution of the policy network's output layer. This combined approach reduces overall energy consumption by 18%-22% under low-altitude operating conditions while maintaining oil temperature stability within ±2°C.

[0097] As a preferred embodiment, the solution of the present invention is specifically implemented as follows: in the low-altitude flight stage, when the flight altitude is below 300 meters, a correlation model between the heat dissipation fin opening control parameters and the air drag coefficient is established. The model obtains air resistance data at different openings through wind tunnel experiments, and uses a polynomial regression method to fit the functional relationship between the opening and the drag coefficient. In the reward function calculation process of the reinforcement learning model, an additional negative reward term related to the real-time flight resistance is introduced, and its weight coefficient is dynamically adjusted through Monte Carlo tree search. In the action selection stage, a probability mask is applied to the output layer of the neural network, which will force the probability of the fin opening action that causes the air drag coefficient to be higher than the preset threshold to zero, while retaining the ability to adjust the magnetic fluid flow rate and the phase change material activation ratio.

[0098] Through the above-mentioned technical solution, the present invention effectively solves the energy waste problem caused by redundant fin openings during low-altitude flight. By coupling an aerodynamic model with a reinforcement learning reward mechanism, the system can autonomously avoid fin control strategies under high-drag conditions, reducing flight energy consumption while ensuring basic heat dissipation requirements. The implementation of the action mask mechanism further constrains the feasible domain of control parameters, preventing the policy network from falling into local optimal solutions and achieving a dynamic balance between heat dissipation efficiency and energy consumption.

[0099] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A multi-physics field collaborative control method for a flying car oil cooler based on reinforcement learning, characterized in that: The following steps are involved: Step S1: The flying car sensor network collects power parameters, environmental parameters, and thermodynamic parameters in real time, normalizes the data, and constructs a multi-dimensional state vector with temporal and spatial correlation; Step S2: Constructing a reinforcement learning model based on the PPO algorithm, defining a state space including the multidimensional state vector, a three-dimensional continuous action space including fin opening, magnetic fluid flow rate, and phase change material activation ratio, and a reward function with oil temperature stability, energy efficiency, and equipment life as optimization objectives, and performing model training in a simulation environment; Step S3: Deploy the trained PPO model to the edge computing node, receive the real-time state vector output control parameters, dynamically adjust the action weights according to the flight altitude, and update the policy network through the online learning mechanism; Step S4: Execute the cooling strategy and handle exceptions: activate the phase change material to store or release heat according to the control parameters, trigger model replanning in response to sudden changes in thermal load, and switch to safety mode when the oil temperature deviation exceeds the preset threshold.

2. The multi-physics field coordinated control method for a flying car oil cooler based on reinforcement learning according to claim 1 is characterized in that: The step S1 specifically includes: Step S1-1: real-time acquisition of power parameters, including propeller speed and motor power; Step S1-2: synchronously collecting environmental parameters, including flight altitude, ambient temperature and humidity; Step S1-3: collecting thermodynamic parameters in parallel, including oil temperature at the inlet and outlet of the oil cooler, magnetic fluid flow rate, and phase change state of the phase change material; Step S1-4: Dynamically normalize the three types of parameters respectively; Step S1-5: The fused parameters are used to construct a multi-dimensional state vector containing a timestamp and a spatial position marker.

3. The multi-physics field coordinated control method for a flying car oil cooler based on reinforcement learning according to claim 1 is characterized in that: The model training in the simulation environment in step S2 includes: Step S2-1: Generate a low-altitude high-humidity working condition data set to simulate the impact of condensation effect on heat dissipation; Step S2-2: Construct a high-altitude low-pressure working condition data set to simulate the attenuation of heat dissipation efficiency caused by changes in air density; Step S2-3: Create transient thermal load impact scenarios when modules are separated and combined; Step S2-4: verifying the coupling relationship between the phase change rate of the phase change material and the thermal conductivity efficiency of the magnetic fluid.

4. The multi-physics field coordinated control method for a flying car oil cooler based on reinforcement learning according to claim 1 is characterized in that: Dynamically adjusting the action weight in step S3 includes: Step S3-1: When the flying car's altitude is greater than 1000 meters, the action weight of the magnetic fluid flow rate is increased by 30%-50%; Step S3-2: When the flying car's flight altitude is ≤300 meters, increase the decision priority of the fin opening action and add an energy consumption penalty factor.

5. The multi-physics field coordinated control method for a flying car oil cooler based on reinforcement learning according to claim 1 is characterized in that: Triggering model replanning in step S4 includes: Step S4-1: monitor the oil temperature change rate in real time, and determine that the thermal load has suddenly changed when the change rate exceeds a threshold value α; Step S4-2: Freeze the current policy network and call the emergency control sequence of the historical scenario; Step S4-3: Using the current state as the starting point, perform short-term reinforcement learning for rapid replanning; Step S4-4: Output the new control parameter combination and verify the stability.

6. The multi-physics field coordinated control method for a flying car oil cooler based on reinforcement learning according to claim 1 is characterized in that: The reward function construction in step S2 includes: Step S2-5: define the oil temperature stability index as the inverse of the variance between the inlet and outlet temperature difference and the set value; Step S2-6: defining the energy efficiency index as the weighted inverse of the sum of the magnetic fluid pump power and the heat sink drive power; Step S2-7: defining the device life index as a negative correlation function between the number of phase change cycles of the phase change material and the temperature fluctuation amplitude; Step S2-8: Use the entropy weight method to dynamically allocate weight coefficients of the three indicators.

7. The multi-physics field coordinated control method for a flying car oil cooler based on reinforcement learning according to claim 1 is characterized in that: The online learning mechanism in step S3 includes: Step S3-4: Storing the real-time flight data of the flying car into a circulating experience pool and marking abnormal operating condition data; Step S3-5: Trigger incremental training every time N pieces of data are accumulated, and prioritize replaying abnormal operating condition data; Step S3-6: Asynchronously updating the strategic network parameters using a dual network structure; Step S3-7: Prevent the policy update from deviating from the baseline model through the KL divergence constraint.

8. The multi-physics field coordinated control method for a flying car oil cooler based on reinforcement learning according to claim 1 is characterized in that: In step S4, the phase change material is activated to store or release heat according to the control parameters, the model is triggered to re-plan in response to the sudden change of the thermal load, and the safety mode is switched when the oil temperature deviation exceeds the preset threshold. S4-5: Activate safety mode when the oil temperature deviation is greater than ±3°C for 5 consecutive cycles; S4-6: Switch to the backup sensor group to verify data consistency; S4-7: Enable the preset PID controller to maintain basic cooling; S4-8: Generate a model diagnosis report and trigger a manual intervention flag.

9. The multi-physics field coordinated control method for a flying car oil cooler based on reinforcement learning according to claim 2 is characterized in that: The phase change state monitoring of the phase change material in step S1-3 includes: Step S1-3a: detecting the axial temperature gradient of the phase change material module by an embedded temperature sensor array; Step S1-3b: Determine the solid-liquid phase change ratio using the rate of change of the acoustic wave propagation velocity; Step S1-3c: quantize the phase change ratio into a continuous value of 0-1 and incorporate it into the state vector.

10. The multi-physics field coordinated control method for a flying car oil cooler based on reinforcement learning according to claim 4 is characterized in that: The energy consumption optimization of step S3-2 includes: Step S3-2a: Establishing a relationship model between the heat dissipation fin opening and the air resistance coefficient; Step S3-2b: Add a flight resistance penalty term to the reward function; Step S3-2c: Limit the selection probability of high-energy-consuming actions through the action mask mechanism.

Citation Information

Patent Citations

  • Bearing system

    JP2022129900A

  • Intelligent mission thermal management system

    US20180354641A1

  • Fuel cell oxygen delivery system, method and apparatus for clean fuel electric aircraft

    WO2022035968A1

Cited By

  • Multi-cabin system transformer substation thermal environment simulation and phase change material configuration optimization method

    CN121960295A

  • Simulation of Thermal Environment and Optimization Method for Phase Change Material Configuration in Multi-Compartment Substations

    CN121960295B