A system and method for optimizing cooling of an automotive mold based on reinforcement learning

By optimizing the mold cooling system through reinforcement learning, and monitoring and optimizing cooling parameters in real time, the problem of uneven cooling of complex mold structures has been solved, achieving efficient temperature control and production stability.

CN121433379BActive Publication Date: 2026-05-19SHAANXI ZUNRONG INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHAANXI ZUNRONG INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2025-11-05
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Traditional mold cooling systems struggle to accurately capture local thermal characteristics when dealing with complex geometries, resulting in uneven cooling and unstable temperature distribution, which affects molding quality and production efficiency.

Method used

An automotive mold cooling optimization system based on reinforcement learning is adopted. Through a multi-dimensional basic cooling state database, a local thermal characteristic model, and a global control strategy, the system monitors and optimizes the cooling water flow and temperature regulation in real time. By combining simulation and reinforcement learning, the system dynamically adjusts the cooling parameters.

Benefits of technology

It improves the uniformity and stability of mold temperature control, significantly enhances the accuracy and efficiency of cooling response, reduces heat accumulation, and improves the stability of the production process and the consistency of finished product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121433379B_ABST
    Figure CN121433379B_ABST
Patent Text Reader

Abstract

The application relates to the field of mold cooling optimization, and discloses a vehicle mold cooling optimization system and method based on reinforcement learning, which comprises the following steps: collecting real-time temperature and flow data of each cooling loop, pre-processing and abnormity filtering the data, combining historical operation records and environmental parameters to construct a multi-dimensional basic cooling state database, performing spatial mapping and statistical analysis on the temperature and flow data in the basic database, identifying high-temperature areas with heat concentration and parts with significant temperature gradients to form a local heat feature model, establishing an independent control submodel, adopting a method combining simulation and reinforcement learning to optimize cooling water flow and temperature regulation parameters of each loop, fusing the local adjustment strategy and the overall cooling target, and optimizing main loop parameters through a global control algorithm, and continuously monitoring temperature and flow fluctuations during mold cooling and inputting real-time data into the optimization system. The application has the advantages of improving cooling precision and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mold cooling optimization, specifically to a reinforcement learning-based automotive mold cooling optimization system and method. Background Technology

[0002] In the production of automotive parts, the design and control of the mold cooling system have a significant impact on the molding quality and production efficiency of the finished products. As a technician with extensive experience in mold design and debugging, I am keenly aware of the importance of a rational cooling system layout. Traditional optimization methods mainly rely on engineers' experience or fixed simulation models, manually setting cooling water flow and temperature control parameters to improve mold temperature distribution. However, when the mold structure is complex, such as with large surface variations, deep cavities, or fine ribs, the internal heat conduction path becomes complex, and heat accumulation in local areas can easily occur, leading to uneven cooling and unstable temperature distribution, thus affecting molding quality and production efficiency. In recent years, intelligent algorithms such as reinforcement learning have been gradually introduced into the field of mold cooling optimization, enabling automatic adjustment of cooling parameters by learning mold temperature feedback in real time. Although single-layer reinforcement learning models have a certain effect on overall cooling optimization, they often struggle to accurately capture local thermal characteristics when dealing with complex geometries, resulting in unsatisfactory cooling effects in locally high-temperature areas. In my personal debugging experience, this problem is particularly pronounced in irregularly shaped molds or deep-cavity molds, often requiring multiple manual fine-tuning adjustments to meet design requirements. Therefore, it is essential to design a reinforcement learning-based automotive mold cooling optimization system and method to improve cooling accuracy and efficiency. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a reinforcement learning-based automotive mold cooling optimization system and method, which has the advantages of improving cooling accuracy and efficiency, and solves the problems mentioned in the background technology.

[0004] To achieve the aforementioned goals of improving cooling accuracy and efficiency, this invention provides the following technical solution: a reinforcement learning-based method for optimizing automotive mold cooling, comprising the following steps:

[0005] Real-time temperature and flow data of each cooling loop are collected, and the data is preprocessed and anomaly filtered. Combined with historical operation records and environmental parameters, a multi-dimensional basic cooling status database is constructed.

[0006] Based on the multidimensional basic cooling state database, spatial mapping and statistical analysis are performed on the temperature and flow data in the basic database to identify high-temperature areas with concentrated heat and areas with significant temperature gradients. Based on historical thermal response trends, potential heat concentration risks are predicted to form a local thermal characteristic model.

[0007] For the high-temperature regions identified in the local thermal feature model, an independent control sub-model is established. A combination of simulation and reinforcement learning is used to optimize the cooling water flow and temperature regulation parameters of each loop, forming a local adjustment strategy.

[0008] The local adjustment strategy is integrated with the overall cooling target. The main loop parameters are optimized through a global control algorithm, and a priority adjustment mechanism is introduced to generate a global control strategy.

[0009] During the mold cooling process, temperature and flow fluctuations are continuously monitored, and real-time data is input into the optimization system to dynamically correct local and global control strategies and generate executable adjustment instructions.

[0010] Preferably, the process of constructing a multidimensional basic cooling state database is as follows:

[0011] Temperature sensors and flow meters are installed in each cooling circuit to collect data synchronously, and temperature, flow rate, pressure and coolant characteristic parameters are obtained in real time.

[0012] By combining information on mold structure topology, material thermal conductivity, and environmental conditions, the collected data is synchronized in time and aligned in space.

[0013] By employing noise removal, anomaly detection, missing data compensation, and data normalization, a multidimensional basic cooling state database containing spatial coordinates, time series, and state variables is established.

[0014] Preferably, the process for identifying high-temperature regions with concentrated heat and areas with significant temperature gradients is as follows:

[0015] Based on the temperature, flow rate and time series data recorded in the multidimensional basic cooling state database, and combined with the three-dimensional geometric model of the mold, the cooling data is spatially mapped and associated with nodes to construct a multidimensional spatial temperature distribution matrix.

[0016] Statistical and time-series analysis is performed on the mapped data to calculate the average temperature, temperature gradient vector, fluctuation coefficient, and flow distribution consistency index of each region, comprehensively reflecting the local cooling uniformity and heat conduction characteristics.

[0017] Cluster analysis and threshold segmentation algorithms were used to decompose the temperature distribution matrix into regions, identifying areas of temperature anomaly, zones of significant gradient changes, and areas of flow imbalance.

[0018] By combining time series trends, the areas of sustained high temperatures and instantaneous heat peaks are further marked.

[0019] Preferably, the process of forming a local thermal feature model is as follows:

[0020] Based on the identified high-temperature regions with concentrated heat and areas with significant temperature gradients, the temperature, flow rate, and time series data of the corresponding spatial nodes are extracted from the multidimensional basic cooling state database.

[0021] By combining historical operation records and heat conduction simulation results, time series modeling of temperature dynamics in each high-temperature region is performed to calculate heat accumulation rate, local peak response time and cooling decay constant.

[0022] By using trend prediction and anomaly detection algorithms, the thermal stability, periodic fluctuations and potential heat concentration risks of different regions are assessed, and thermal response sensitive areas and high-risk feature points are identified.

[0023] For each high-temperature region, a structured feature matrix containing temperature response coefficient, gradient change rate, thermal inertia coefficient, cooling coupling degree, and risk weight is constructed to form a local thermal feature model.

[0024] Preferably, the process of establishing an independent control sub-model is as follows:

[0025] Based on the high-temperature region and thermal response parameters output by the local thermal characteristic model, an independent control sub-model is established for each region.

[0026] During the model building process, the state space, action space, and reward function are defined;

[0027] By combining the results of cooling simulation analysis with historical operating data, each independent control sub-model is trained and validated, and the control strategy parameters are continuously iterated and optimized using reinforcement learning algorithms.

[0028] Preferably, the process of optimizing the cooling water flow and temperature regulation parameters for each loop is as follows:

[0029] In the independent control sub-model, the flow and temperature control parameters are updated using a reinforcement learning strategy gradient algorithm based on the current local temperature deviation, gradient distribution, and control objective.

[0030] The control results are verified by simulation, taking into account the physical constraints of the cooling system, such as the upper limit of water pressure in the pipeline, the range of pump speed, and the flow balance requirements.

[0031] A local adjustment strategy is generated by detecting policy convergence and correcting constraints.

[0032] Preferably, the process of optimizing the main loop parameters using a global control algorithm is as follows:

[0033] Based on the local adjustment strategies of each loop and the optimization results of its cooling parameters, the flow distribution ratio, temperature regulation curve and control frequency characteristic parameters of the local loop are extracted.

[0034] By integrating the local optimization results with the overall cooling requirements, a global control algorithm model is established. The model inputs include key variables such as main loop flow rate, coolant temperature, pump speed, pressure loss, and energy consumption coefficient.

[0035] A priority adjustment mechanism based on reinforcement learning is introduced to dynamically allocate the adjustment of the main loop according to the thermal risk level, response delay and energy consumption feedback of each local loop;

[0036] Through multidimensional simulation and global optimization calculation, the strategy balance, thermal response stability and system constraint compliance are verified, and finally the optimized main loop parameters are output.

[0037] Preferably, the process of generating a global control strategy is as follows:

[0038] Based on the optimization results of the global control algorithm, the cooling flow, temperature setting, adjustment frequency and response characteristic data output from each local adjustment strategy are integrated to establish a unified strategy fusion framework.

[0039] During the strategy integration phase, the main loop parameter settings, local loop priority allocation, dynamic adjustment frequency and control timing are structured and integrated, and the global control objectives are dynamically adjusted based on the system energy consumption weight, cooling balance and thermal response stability indicators.

[0040] Based on the temperature response curves, flow fluctuation characteristics, and equipment operating constraints of each cooling circuit, the coupling relationship between various control parameters is analyzed.

[0041] In the policy generation phase, a reinforcement learning policy network is used to train and optimize the control instruction generation logic, forming a structured global control policy.

[0042] Preferably, the process of generating executable adjustment instructions is as follows:

[0043] During the mold cooling process, real-time temperature, flow rate, and environmental parameters of each circuit are continuously collected, and the collected data is input into the optimization system.

[0044] The optimization system is based on a global control strategy, combined with the priority of each loop, temperature response characteristics, flow fluctuations and equipment operation constraints in each local adjustment strategy;

[0045] Calculate the cooling water flow rate, temperature setting, and adjustment parameters for each loop, and generate execution commands that can be directly sent to the control device.

[0046] A reinforcement learning-based automotive mold cooling optimization system, characterized in that it includes:

[0047] Data acquisition model: Collect real-time temperature and flow data of each cooling loop, perform preprocessing and anomaly filtering, and build a multi-dimensional basic cooling status database by combining historical and environmental parameters;

[0048] Thermal analysis model: Based on a multidimensional basic cooling state database, it performs spatial mapping and statistical analysis on temperature and flow data to identify high-temperature areas and areas with significant temperature gradients, and predict potential heat concentration risks.

[0049] Local control model: For the high-temperature region identified in the local thermal feature model, an independent control sub-model is established, and the cooling water flow rate and temperature regulation parameters of each loop are optimized by combining simulation and reinforcement learning;

[0050] Global Coordination Model: This model integrates local adjustment strategies with overall cooling objectives, optimizes main loop parameters through a global control algorithm, and introduces a priority adjustment mechanism to generate a global control strategy.

[0051] Strategy execution model: During the mold cooling process, temperature and flow rate are continuously monitored, real-time data is input into the optimization system, local and global strategies are dynamically corrected, and executable adjustment instructions are generated.

[0052] Compared with existing technologies, this invention provides a reinforcement learning-based automotive mold cooling optimization system and method, which has the following beneficial effects:

[0053] This invention utilizes a multi-dimensional basic cooling state database to monitor and analyze the temperature and flow rate of each circuit in the mold in real time, enabling precise identification and dynamic control of local high-temperature areas. Simultaneously, it combines a global control strategy to optimize and adjust the main circuit parameters, effectively matching the synergistic effect of local and overall cooling, thereby improving the uniformity and stability of mold temperature control. Through precise management and real-time feedback correction of cooling water flow rate, temperature, and adjustment cycle, it significantly improves the response accuracy and adjustment efficiency of mold cooling, reduces heat accumulation, and enhances the stability of the production process and the consistency of finished product quality. Attached Figure Description

[0054] Figure 1 This is a schematic diagram of the method of the present invention;

[0055] Figure 2 This is a schematic diagram of the structure of the present invention. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] Example 1: Please refer to Figure 1 As shown in the embodiment of the present invention, a method for optimizing the cooling of automotive molds based on reinforcement learning includes the following steps:

[0058] S1: Collect real-time temperature and flow data of each cooling circuit, preprocess the data and filter anomalies, and build a multi-dimensional basic cooling status database by combining historical operation records and environmental parameters.

[0059] The process of constructing a multidimensional basic cooling status database in S1 is as follows:

[0060] Temperature sensors and flow meters are installed in each cooling circuit to synchronously acquire real-time data on temperature, flow rate, pressure, and coolant characteristics. Temperature sensors and flow meters are installed on each cooling circuit, with sensor placement based on the mold cooling water circuit design. A temperature acquisition point is set every 20 to 50 millimeters. Flow meters are installed on the main inlet and outlet pipes and key branch locations. The system achieves real-time synchronous data acquisition via industrial Ethernet or fieldbus, with a sampling frequency of once per second. It simultaneously records sensor number, circuit number, temperature, flow rate, pressure, and coolant physical characteristics, including density and specific heat capacity. During the acquisition process, sensor status is monitored, and abnormal signals such as drift, failure, or transient anomalies are marked. Environmental temperature and humidity are also collected.

[0061] By combining information on mold structure topology, material thermal conductivity, and environmental conditions, the collected data is synchronized in time and aligned in space. The collected temperature, flow rate, pressure, and coolant characteristic parameters are associated with the geometric information of the mold's 3D model. Each sensor corresponds to the spatial coordinates of the mold. Based on the topology of the cooling circuit and the direction of water flow, the collected data from different circuits and sensors are synchronized in time, mapping the collected data to a unified time reference. Sensor response delay and signal transmission delay are recorded. By combining the thermal conductivity and heat capacity information of various parts of the mold materials, as well as environmental condition parameters such as coolant inlet temperature and ambient temperature, the data is aligned at spatial nodes to achieve precise spatial and temporal correspondence.

[0062] By employing noise removal, anomaly detection, missing data compensation, and data normalization, a multidimensional basic cooling state database containing spatial coordinates, time series, and state variables is established. Signal processing is performed on the synchronized and aligned data, using low-pass filtering and median filtering to remove high-frequency noise and spike interference. Anomaly detection is performed by statistically analyzing the temperature and flow data of each loop node; values ​​exceeding three standard deviations of the average are marked as anomalies. Simultaneously, missing data is filled using neighboring node interpolation or historical curve fitting. Data normalization ensures that temperature, flow rate, pressure, and coolant characteristics are stored under a unified dimension, and the data is organized into multidimensional structured tables according to loop number, spatial coordinates, and time series. This results in a multidimensional basic cooling state database containing spatial coordinates, time series, and state variables, which can be directly used for local thermal feature modeling and global control strategy calculation.

[0063] S2: Based on the multidimensional basic cooling state database, spatial mapping and statistical analysis are performed on the temperature and flow data in the basic database to identify high-temperature areas with concentrated heat and areas with significant temperature gradients. Based on historical thermal response trends, potential heat concentration risks are predicted to form a local thermal characteristic model.

[0064] The process of identifying high-temperature regions with concentrated heat and areas with significant temperature gradients in S2 is as follows:

[0065] Based on the temperature, flow rate, and time series data recorded in the multidimensional basic cooling state database, and combined with the mold's three-dimensional geometric model, the cooling data is spatially mapped and associated with nodes to construct a multidimensional spatial temperature distribution matrix. The temperature, flow rate, and time series data of each loop are extracted from the multidimensional basic cooling state database, and each acquisition point is mapped to a spatial node of the mold's three-dimensional geometric model. According to the mold's water channel topology and coolant flow direction, the position of each sensor or measurement point in the three-dimensional grid is determined, and a spatial index relationship is established to achieve a one-to-one correspondence between data and physical location. The time series data is synchronously mapped to nodes according to the sampling time to form a multidimensional spatial temperature distribution matrix.

[0066] Statistical and time-series analysis is performed on the mapped data to calculate the average temperature, temperature gradient vector, fluctuation coefficient, and flow distribution consistency index of each region, comprehensively reflecting the local cooling uniformity and heat conduction characteristics. Node-by-node statistical analysis is performed on the spatial temperature distribution matrix to calculate the average temperature, standard deviation, and temperature fluctuation coefficient of each node. Simultaneously, node temperature gradient vectors are generated using finite difference or gradient calculation methods to quantify the local thermal change rate. Combined with the flow distribution data of each loop, the flow consistency index is calculated to reflect the uniformity of coolant at each node. Time-series analysis methods are used to calculate the temperature and gradient change trends over time, generating local thermal response curves.

[0067] Cluster analysis and threshold segmentation algorithms are used to decompose the temperature distribution matrix into regions, identifying areas of abnormal temperature clusters, zones of significant gradient changes, and areas of flow imbalance. The temperature, gradient, and flow consistency data obtained from statistical analysis are input into the clustering algorithm, and density clustering or hierarchical clustering methods are used to classify spatial nodes, identifying areas of high temperature concentration, areas of significant gradient changes, and areas of flow imbalance. Based on the set threshold, nodes whose temperature exceeds the average value by a certain multiple are marked as abnormal clusters, and nodes whose gradient changes are greater than the threshold form gradient significant zones. At the same time, the corresponding loop number and spatial coordinates are recorded.

[0068] By combining time series trends, the sustained high-temperature areas and instantaneous heat peak areas are further labeled; the clustering and threshold segmentation results are analyzed in conjunction with time series data to calculate the duration for which the temperature in each region continuously exceeds the threshold, and the continuous high-temperature areas are labeled as sustained high-temperature areas; nodes where the temperature rises rapidly to the peak within a short period of time are labeled as instantaneous heat peak areas. The start and end times, peak temperature, average temperature and temperature change rate of each area are recorded, and the cooling circuit number is associated with them.

[0069] The process of forming a local thermal feature model in S2 is as follows:

[0070] Based on the identified high-temperature regions with concentrated heat and areas with significant temperature gradients, the temperature, flow rate, and time-series data of the corresponding spatial nodes in the multidimensional basic cooling state database are extracted. Based on the high-temperature regions and areas with significant temperature gradients identified in the previous stage, the temperature, flow rate, and time-series data of the corresponding spatial nodes in the multidimensional basic cooling state database are completely extracted. The data for each node includes continuous sampling timestamps, instantaneous temperature values, flow rate values, coolant pressure, and relevant environmental parameters.

[0071] By combining historical operation records and heat conduction simulation results, time series modeling of temperature dynamics in each high-temperature region is performed to calculate the heat accumulation rate, local peak response time, and cooling decay constant. The extracted node temperature data is compared with historical operation records and heat conduction simulation results to construct a temperature time series model for each high-temperature region. The heat accumulation rate of each region is calculated using linear regression, exponential smoothing, and numerical integration methods to evaluate the occurrence time and duration of temperature peaks. At the same time, the cooling decay constant is calculated using a first-order or second-order heat decay model. Parallel analysis of flow rate data is performed to establish a dynamic response relationship between temperature and flow rate, ensuring that the causal characteristics of temperature changes are reflected in the feature matrix.

[0072] By employing trend prediction and anomaly detection algorithms, the thermal stability, periodic fluctuations, and potential heat concentration risks of different regions are assessed, identifying thermally sensitive areas and high-risk feature points. The sliding window method and autocorrelation analysis are applied to predict the temperature sequence of high-temperature regions, extracting periodic fluctuation characteristics and local anomalous peaks. Anomaly detection algorithms are used to identify abrupt changes and nodes exceeding statistical thresholds in the temperature change curve, calculating local thermal stability indices and potential heat concentration risks. Sensitivity weights are assigned to each node, and thermally sensitive areas and high-risk feature points are marked based on historical peak frequencies and gradient change amplitudes.

[0073] For each high-temperature region, a structured feature matrix is ​​constructed, including temperature response coefficient, gradient change rate, thermal inertia coefficient, cooling coupling degree, and risk weight, forming a local thermal feature model. The calculation results of each high-temperature region are organized into a structured feature matrix, including temperature response coefficient, temperature gradient change rate, thermal inertia coefficient, local cooling coupling degree, and risk weight. Each row in the matrix corresponds to a spatial node, and each column corresponds to different thermal characteristic parameters. Standardization and normalization processes are used to ensure that the dimensions of different parameters are consistent. The feature matrix also records the loop number, location coordinates, and time label of the node, forming a complete local thermal feature model.

[0074] S3: For the high-temperature areas identified in the local thermal feature model, an independent control sub-model is established. A combination of simulation and reinforcement learning is used to optimize the cooling water flow and temperature regulation parameters of each loop, forming a local adjustment strategy.

[0075] The process of establishing an independent control sub-model in S3 is as follows:

[0076] Based on the high-temperature region and thermal response parameters output by the local thermal characteristic model, an independent control sub-model is established for each region. Based on the high-temperature region and its thermal response parameters output by the local thermal characteristic model, a separate control sub-model is established for each high-temperature region. During the model initialization phase, the node list, peak temperature response, gradient change rate, and thermal inertia coefficient of the high-temperature region are used as input features, and the cooling loop number, spatial coordinates, and flow constraints are used as model boundary conditions. An independent data channel and control object are assigned to each region to ensure that there is no cross-interference between the sub-models.

[0077] During model building, a state space, action space, and reward function are defined. For each independent control sub-model, the state space includes local temperature, local flow rate, temperature gradient, thermal response rate, historical temperature fluctuations, and coolant pressure parameters. Each state variable is synchronously updated according to the sampling frequency, forming a multi-dimensional state vector. The action space consists of adjustable control parameters of the local loop, including the cooling water flow rate adjustment amplitude, temperature setpoint, and adjustment cycle. Each action component has upper and lower limits set, and constraints are applied based on pump speed, valve characteristics, and loop flow resistance physical constraints to ensure the feasibility of action execution. Within the control cycle, the action space supports continuous and discrete adjustment operations and can be optimized for action sequences by combining historical adjustment responses. The reward function is constructed based on local temperature deviation, gradient change rate, response speed, and cooling stability, quantitatively evaluating the control effect of the model at each time step. By numerically calculating the sum of squared errors between the temperature and the target temperature, combined with gradient change smoothness and response delay, a comprehensive reward value is generated and fed back to the strategy iteration algorithm for parameter updates. The reward function also considers local loop physical constraints to ensure that action adjustments do not lead to flow imbalance or overpressure.

[0078] By combining cooling simulation analysis results with historical operating data, each independent control sub-model is trained and validated. Reinforcement learning algorithms are used to continuously iterate and optimize the control strategy parameters. The independent control sub-models are trained and validated using reinforcement learning algorithms. The impact of control actions on temperature distribution is predicted through forward simulation, and the strategy parameters are then updated and adjusted in reverse based on the reward function. During training, the learning rate and iteration step size are dynamically adjusted, and the strategy is continuously optimized based on the response in high-temperature areas until the local control sub-model can stably output action sequences within the control cycle and meet the regional cooling requirements.

[0079] The process of optimizing the cooling water flow and temperature regulation parameters for each loop in S3 is as follows:

[0080] In the independent control sub-model, the flow and temperature control parameters are updated using a reinforcement learning strategy gradient algorithm based on the current local temperature deviation, gradient distribution, and control objective. In the independent control sub-model, the real-time temperature data of each node in the high-temperature region is first compared with the target cooling temperature to calculate the local temperature deviation. Simultaneously, a multi-dimensional state vector is generated by combining the temperature gradient distribution and historical response characteristics. This state vector includes the node's average temperature, maximum temperature gradient, local cooling flow rate, and temperature fluctuation information from the most recent several cycles, which is used as input for the subsequent reinforcement learning strategy. The data acquisition frequency is consistent with the control cycle to ensure that the model can capture the dynamic trend of temperature changes.

[0081] Combining the physical constraints of the cooling system, such as the upper limit of pipeline water pressure, pump speed range, and flow balance requirements, the control results are simulated and verified. Based on the aforementioned state vector, the cooling water flow rate, temperature setting, and adjustment cycle parameters in the local control sub-model are iteratively updated using a policy gradient reinforcement learning algorithm. The algorithm calculates the instantaneous reward value by simulating the impact of each action on the local temperature and gradient, associates the reward with the policy gradient, and executes parameter updates. During the iteration process, batch sampling and time series playback techniques are used to smooth policy updates and reduce noise interference, ensuring that the policy gradually converges.

[0082] A local adjustment strategy is generated through strategy convergence testing and constraint correction. Before being issued, the updated control parameters need to be verified in conjunction with the physical constraints of the cooling system, including the upper limit of pipeline water pressure, pump speed range, and loop flow balance. The pressure and flow distribution of each loop are calculated through simulation model. If over-limit or unbalanced conditions are found, the flow and temperature setting parameters are proportionally corrected or the constraints are trimmed. After the strategy convergence test is completed, the control parameters that have been constrained and corrected are organized into a structured local adjustment strategy file. This strategy includes the cooling water flow setting, temperature target value, action execution sequence, and adjustment cycle information for each loop, and can be directly issued to the execution control system.

[0083] S4: Integrate local adjustment strategies with overall cooling targets, optimize main loop parameters through global control algorithms, and introduce a priority adjustment mechanism to generate a global control strategy.

[0084] The optimization process of the main loop parameters in S4 using a global control algorithm is as follows:

[0085] Based on the local adjustment strategies and cooling parameter optimization results of each loop, the flow distribution ratio, temperature regulation curve, and control frequency characteristic parameters of the local loop are extracted. From the local adjustment strategies generated by each independent control sub-model, the cooling parameter characteristics of each loop are extracted, including the water flow distribution ratio, temperature regulation curve, and control frequency sequence. The water flow distribution ratio is calculated by the ratio of the design flow rate to the actual optimized flow rate of each loop. The temperature regulation curve is composed of the control input change sequence of the local node temperature over time. The control frequency is obtained by calculating the number of adjustments in each cycle. These characteristics are uniformly transformed into matrix or multidimensional tensor form.

[0086] By integrating local optimization results with overall cooling requirements, a global control algorithm model is established. The model inputs include key variables such as main loop flow rate, coolant temperature, pump speed, pressure loss, and energy consumption coefficient. After integrating local features, a global control model is established, using key variables such as main loop flow rate, coolant inlet temperature, pump speed, loop pressure loss, and cooling energy consumption coefficient as inputs. Flow rate, temperature, and pump speed are collected in real time by sensors and stored as time series. Pressure loss can be calculated from the relationship between flow rate and pipeline resistance. The energy consumption coefficient is quantified based on the pump power curve and circulating water flow rate. The model combines and maps the input variables to form a main loop parameter control space, providing a complete set of variables for subsequent optimization calculations.

[0087] A reinforcement learning-based priority adjustment mechanism is introduced to dynamically allocate the main loop adjustment based on the thermal risk level, response delay, and energy consumption feedback of each local loop. To ensure that the main loop allocation matches the local thermal load, a reinforcement learning-based priority adjustment mechanism is introduced. The thermal risk level of each loop is output by the local thermal characteristic model, including the peak temperature, gradient change rate, and thermal inertia coefficient. The response delay is calculated from historical temperature feedback data, and the energy consumption feedback is calculated from pump power and flow rate. The priority adjustment mechanism dynamically calculates the allocation weight of each loop in the total flow rate of the main loop based on the above indicators, and prioritizes the allocation of flow rate to loops with high risk and short response delay to ensure global cooling balance.

[0088] Through multidimensional simulation and global optimization calculation, the strategy balance, thermal response stability, and system constraint compliance are verified, and the optimized main loop parameters are finally output. The local adjustment strategy and priority weights are input into the global optimization calculation module, and the optimal combination of main loop parameters is solved through multidimensional numerical simulation. The simulation process includes the prediction of temperature response curves, the calculation of the impact of flow distribution on local node temperatures, and the evaluation of pump speed and pressure loss matching. Through step-by-step iterative optimization, the input variables are adjusted until the temperature stability index and system constraint requirements are met, such as the water pressure of each loop not exceeding the design upper limit and the pump power consumption being lower than the set threshold. After global optimization calculation and simulation verification, the final main loop parameters, including flow setting, inlet temperature, pump speed, and adjustment cycle, are generated into structured data output, which can be directly called by the global control strategy generation module. The output format is a parameter table or multidimensional matrix with timestamps, which can be used to issue commands to the control device, while recording input characteristics and optimization results.

[0089] The process of generating a global control strategy in S4 is as follows:

[0090] Based on the optimization results of the global control algorithm, the cooling flow rate, temperature setpoint, adjustment frequency, and response characteristic data output from each local adjustment strategy are integrated to establish a unified strategy fusion framework. The cooling flow rate setpoint, temperature target, adjustment frequency, and temperature response characteristics of each loop are extracted from the local adjustment strategies output by each independent control sub-model. The flow rate and temperature data are presented in time series format, the adjustment frequency is obtained by statistically analyzing the number of adjustments within each control cycle, and the response characteristics include peak temperature, gradient change, and local thermal inertia coefficient. These data are archived according to loop number and time order and stored uniformly in a multi-dimensional matrix.

[0091] In the strategy fusion phase, the main loop parameter settings, local loop priority allocation, dynamic adjustment frequency, and control timing are structurally integrated, and the global control objective is dynamically adjusted based on the system energy consumption weight, cooling balance, and thermal response stability index. In the fusion phase, the main loop parameter settings, local loop priority allocation, adjustment frequency, and control timing are structurally integrated. The main loop parameters are determined by calculating the proportion of each local flow and priority weight. The adjustment frequency and control timing are dynamically allocated based on the local response delay and historical temperature change sequence. The system energy consumption weight is calculated from the pump power and circulation flow. The cooling balance is calculated by calculating the average temperature deviation and gradient difference index of each loop. The thermal response stability is evaluated by the local thermal inertia coefficient and response delay. All parameters are uniformly input into the strategy fusion framework through matrix processing.

[0092] By combining the temperature response curves, flow fluctuation characteristics, and equipment operating constraints of each cooling loop, the coupling relationship between various control parameters is analyzed. By combining a multi-dimensional basic cooling state database and a local thermal characteristic model, a coupling analysis is performed on the temperature response curves, flow fluctuation characteristics, and equipment operating constraints of each loop. The temperature response curves are modeled by the dynamic response of each node temperature to flow changes through time series analysis. The flow fluctuation characteristics are obtained by statistically analyzing the flow fluctuation amplitude and frequency of historical control cycles. Equipment constraints include upper limit of water pressure, pump speed range, and pipeline flow balance. The coupling analysis establishes the interaction relationship between loops by calculating the temperature gradient influence coefficient and flow correlation matrix between loops.

[0093] In the strategy generation phase, a reinforcement learning policy network is used to train and optimize the control command generation logic, forming a structured global control strategy. The global control command generation logic is trained using this network, with the input being the integrated main loop parameter matrix and local loop priority matrix, and the output being an executable sequence of global control actions. The policy network uses forward propagation to calculate the control probability of each parameter, calculates the reward function using historical data and simulation results, adjusts the network weights using backpropagation, and iterates until the control strategy converges. During training, the learning rate is dynamically adjusted to ensure that the global strategy can generate reasonable control commands under different temperature fluctuations and flow conditions. The trained and optimized global control strategy is saved in structured data format, including main loop flow settings, coolant inlet temperature, pump speed configuration, loop priorities, adjustment frequencies, and control timing. The data uses timestamp indexing and matrix storage, supporting real-time distribution to the control device, and simultaneously recording input features, optimization results, and historical reference data during the strategy generation process.

[0094] S5: During the mold cooling process, continuously monitor temperature and flow fluctuations, input real-time data into the optimization system, dynamically correct local and global control strategies, and generate executable adjustment instructions.

[0095] The process of generating executable adjustment instructions in S5 is as follows:

[0096] During the mold cooling operation, real-time temperature, flow rate, and environmental parameters of each loop are continuously collected, and the collected data is input into the optimization system. During the mold cooling operation, temperature sensors, flow meters, and pressure sensors arranged in each cooling loop continuously collect real-time temperature, flow rate, and coolant status parameters within the loop, while simultaneously recording ambient temperature, humidity, and mold surface temperature information. The collection frequency is consistent with the control cycle, typically one to ten times per second, to ensure that the dynamic trends of temperature fluctuations and flow rate changes are captured. The collected data is archived using timestamps and loop numbers, and undergoes signal quality detection, noise filtering, and outlier labeling to form a multi-dimensional data matrix that can be directly input into the optimization system.

[0097] The optimization system is based on a global control strategy, combined with the priority, temperature response characteristics, flow fluctuations, and equipment operation constraints of each loop in each local adjustment strategy. The collected real-time data is integrated with the loop priority, historical temperature response characteristics, and flow fluctuation information in the global control strategy and each local adjustment strategy. The loop priority is calculated by historical thermal risk assessment and the duration of local temperature peaks. The flow fluctuation characteristics are generated by statistically analyzing the standard deviation and frequency of flow over the most recent several periods. The temperature response characteristics are obtained by calculating the node average temperature, gradient change, and cooling attenuation coefficient. The integrated data matrix is ​​used as input to calculate the cooling regulation requirements of each loop and to provide multi-dimensional input support for strategy optimization.

[0098] The system calculates the cooling water flow rate, temperature setpoint, and adjustment parameters for each loop, generating execution commands that can be directly sent to the control device. Based on the integrated input data, the optimization system calculates the cooling water flow rate setpoint, temperature target value, and adjustment cycle for each loop. During the calculation, a time series prediction and loop coupling analysis method is used, combined with local thermal inertia coefficient, response delay, and flow ratio to dynamically adjust the cooling parameters. Simultaneously, all calculation results must be verified by physical constraints, including upper limits for pipeline water pressure, pump speed range, flow balance, and upper limits for temperature. The calculated cooling water flow rate, temperature setpoint, and adjustment cycle parameters for each loop are then converted... The system transforms control commands into a directly quantifiable data structure. Each command includes a loop number, control timestamp, target flow rate, target temperature, adjustment range, and execution cycle. These commands are stored in a matrix or serialized format, supporting real-time reading by the control device. After the control command is issued and executed, the system continuously monitors the actual temperature, flow rate, and equipment operating status of each loop. It compares the feedback data with the expected command, records deviations, response delays, and abnormal events, and returns the feedback information to the optimization system in real time via the data bus. This allows for dynamic correction of the local and global control strategies for the next cycle, enabling iterative updates and ensuring the continuity and operability of control command generation.

[0099] A reinforcement learning-based automotive mold cooling optimization system includes:

[0100] Data acquisition model: Collect real-time temperature and flow data of each cooling loop, perform preprocessing and anomaly filtering, and build a multi-dimensional basic cooling status database by combining historical and environmental parameters;

[0101] Thermal analysis model: Based on a multidimensional basic cooling state database, it performs spatial mapping and statistical analysis on temperature and flow data to identify high-temperature areas and areas with significant temperature gradients, and predict potential heat concentration risks.

[0102] Local control model: For the high-temperature region identified in the local thermal feature model, an independent control sub-model is established, and the cooling water flow rate and temperature regulation parameters of each loop are optimized by combining simulation and reinforcement learning;

[0103] Global Coordination Model: This model integrates local adjustment strategies with overall cooling objectives, optimizes main loop parameters through a global control algorithm, and introduces a priority adjustment mechanism to generate a global control strategy.

[0104] Strategy execution model: During the mold cooling process, temperature and flow rate are continuously monitored, real-time data is input into the optimization system, local and global strategies are dynamically corrected, and executable adjustment instructions are generated.

[0105] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0106] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for optimizing cooling of automotive molds based on reinforcement learning, characterized in that, Includes the following steps: Real-time temperature and flow data of each cooling loop are collected, and the data is preprocessed and anomaly filtered. Combined with historical operation records and environmental parameters, a multi-dimensional basic cooling status database is constructed. Based on the multidimensional basic cooling state database, spatial mapping and statistical analysis are performed on the temperature and flow data in the basic database to identify high-temperature areas with concentrated heat and areas with significant temperature gradients. Based on historical thermal response trends, potential heat concentration risks are predicted to form a local thermal characteristic model. For the high-temperature regions identified in the local thermal feature model, an independent control sub-model is established. A combination of simulation and reinforcement learning is used to optimize the cooling water flow and temperature regulation parameters of each loop, forming a local adjustment strategy. The process of establishing an independent control sub-model is as follows: Based on the high-temperature region and thermal response parameters output by the local thermal characteristic model, an independent control sub-model is established for each region. During the model building process, the state space, action space, and reward function are defined; By combining the results of cooling simulation analysis with historical operating data, each independent control sub-model is trained and verified, and the control strategy parameters are continuously iterated and optimized using reinforcement learning algorithms. The process of optimizing the cooling water flow and temperature regulation parameters for each loop is as follows: In the independent control sub-model, the flow and temperature control parameters are updated using a reinforcement learning strategy gradient algorithm based on the current local temperature deviation, gradient distribution, and control objective. The control results are verified by simulation based on the physical constraints of the cooling system, including the upper limit of water pressure in the pipeline, the pump speed range, and the flow balance requirements. A local adjustment strategy is generated by detecting policy convergence and correcting constraints. The local adjustment strategy is integrated with the overall cooling target. The main loop parameters are optimized through a global control algorithm, and a priority adjustment mechanism is introduced to generate a global control strategy. The process of optimizing the main circuit parameters using a global control algorithm is as follows: Based on the local adjustment strategies of each loop and the optimization results of its cooling parameters, the flow distribution ratio, temperature regulation curve and control frequency characteristic parameters of the local loop are extracted. By integrating the local optimization results with the overall cooling requirements, a global control algorithm model is established. The model inputs include key variables such as main loop flow rate, coolant temperature, pump speed, pressure loss, and energy consumption coefficient. A priority adjustment mechanism based on reinforcement learning is introduced to dynamically allocate the adjustment of the main loop according to the thermal risk level, response delay and energy consumption feedback of each local loop; Through multi-dimensional simulation and global optimization calculation, the strategy balance, thermal response stability and system constraint compliance are verified, and the optimized main loop parameters are finally output. The process of generating a global control strategy is as follows: Based on the optimization results of the global control algorithm, the cooling flow, temperature setting, adjustment frequency and response characteristic data output from each local adjustment strategy are integrated to establish a unified strategy fusion framework. During the strategy integration phase, the main loop parameter settings, local loop priority allocation, dynamic adjustment frequency and control timing are structured and integrated, and the global control objectives are dynamically adjusted based on the system energy consumption weight, cooling balance and thermal response stability indicators. Based on the temperature response curves, flow fluctuation characteristics, and equipment operating constraints of each cooling circuit, the coupling relationship between various control parameters is analyzed. In the policy generation stage, a reinforcement learning policy network is used to train and optimize the control instruction generation logic to form a structured global control policy. During the mold cooling process, temperature and flow fluctuations are continuously monitored, and real-time data is input into the optimization system to dynamically correct local and global control strategies and generate executable adjustment instructions.

2. The method for optimizing automotive mold cooling based on reinforcement learning according to claim 1, characterized in that, The process of constructing a multidimensional basic cooling status database is as follows: Temperature sensors and flow meters are installed in each cooling circuit to collect data synchronously, and temperature, flow rate, pressure and coolant characteristic parameters are obtained in real time. By combining information on mold structure topology, material thermal conductivity, and environmental conditions, the collected data is synchronized in time and aligned in space. By employing noise removal, anomaly detection, missing data compensation, and data normalization, a multidimensional basic cooling state database containing spatial coordinates, time series, and state variables is established.

3. The method for optimizing automotive mold cooling based on reinforcement learning according to claim 2, characterized in that, The process of identifying high-temperature regions with concentrated heat and areas with significant temperature gradients is as follows: Based on the temperature, flow rate and time series data recorded in the multidimensional basic cooling state database, and combined with the three-dimensional geometric model of the mold, the cooling data is spatially mapped and associated with nodes to construct a multidimensional spatial temperature distribution matrix. Statistical and time-series analysis is performed on the mapped data to calculate the average temperature, temperature gradient vector, fluctuation coefficient, and flow distribution consistency index of each region, comprehensively reflecting the local cooling uniformity and heat conduction characteristics. Cluster analysis and threshold segmentation algorithms were used to decompose the temperature distribution matrix into regions, identifying areas of temperature anomaly, zones of significant gradient changes, and areas of flow imbalance. By combining time series trends, the areas of sustained high temperatures and instantaneous heat peaks are further marked.

4. The method for optimizing automotive mold cooling based on reinforcement learning according to claim 3, characterized in that, The process of forming a local thermal characteristic model is as follows: Based on the identified high-temperature regions with concentrated heat and areas with significant temperature gradients, the temperature, flow rate, and time series data of the corresponding spatial nodes are extracted from the multidimensional basic cooling state database. By combining historical operation records and heat conduction simulation results, time series modeling of temperature dynamics in each high-temperature region is performed to calculate heat accumulation rate, local peak response time and cooling decay constant. By using trend prediction and anomaly detection algorithms, the thermal stability, periodic fluctuations and potential heat concentration risks of different regions are assessed, and thermal response sensitive areas and high-risk feature points are identified. For each high-temperature region, a structured feature matrix containing temperature response coefficient, gradient change rate, thermal inertia coefficient, cooling coupling degree, and risk weight is constructed to form a local thermal feature model.

5. The method for optimizing automotive mold cooling based on reinforcement learning according to claim 1, characterized in that, The process of generating executable adjustment instructions is as follows: During the mold cooling process, real-time temperature, flow rate, and environmental parameters of each circuit are continuously collected, and the collected data is input into the optimization system. The optimization system is based on a global control strategy, combined with the priority of each loop, temperature response characteristics, flow fluctuations and equipment operation constraints in each local adjustment strategy; Calculate the cooling water flow rate, temperature setting, and adjustment parameters for each loop, and generate execution commands that can be directly sent to the control device.

6. A reinforcement learning-based automotive mold cooling optimization system, applied to the method described in claims 1-5, characterized in that, include: Data acquisition model: Collect real-time temperature and flow data of each cooling loop, perform preprocessing and anomaly filtering, and build a multi-dimensional basic cooling status database by combining historical and environmental parameters; Thermal analysis model: Based on a multidimensional basic cooling state database, it performs spatial mapping and statistical analysis on temperature and flow data to identify high-temperature areas and areas with significant temperature gradients, and predict potential heat concentration risks. Local control model: For the high-temperature region identified in the local thermal feature model, an independent control sub-model is established, and the cooling water flow rate and temperature regulation parameters of each loop are optimized by combining simulation and reinforcement learning; Global Coordination Model: This model integrates local adjustment strategies with overall cooling objectives, optimizes main loop parameters through a global control algorithm, and introduces a priority adjustment mechanism to generate a global control strategy. Strategy execution model: During the mold cooling process, temperature and flow rate are continuously monitored, real-time data is input into the optimization system, local and global strategies are dynamically corrected, and executable adjustment instructions are generated.