Injection mold temperature control method and device and computer equipment
Through the combination of the deep deterministic strategy gradient control network and the gated cycle unit, the problems of uneven temperature field distribution and insufficient control accuracy in the traditional injection mold temperature control method are solved, and high-precision and efficient temperature control are achieved.
Patent Information
- Application Number
- CN202510148234.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional injection mold temperature control methods are difficult to achieve uniformity and high-precision control of temperature field distribution, especially in the processing of multi-region temperature fields and the identification of dynamic response characteristics.
A deep deterministic strategic gradient control network is adopted, combined with the gated cycle unit and the fully connected layer, optimized and trained through the temperature data matrix to generate a temperature control model to achieve coordinated control of the heating stage and the stability stage.
It significantly improves the decision quality of temperature control, shortens the injection molding cycle, reduces energy consumption, and avoids the "dimensional curse" problem, achieving high-precision temperature field control.
Smart Images

Figure CN120038919A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of temperature control, and particularly to a temperature control method, device and computer equipment for an injection mold. Background Art
[0002] Traditional temperature control of injection molds mainly relies on the PID control method. However, due to the complex mold structure, the lag of heat conduction, and the thermal coupling relationship between regions, the temperature field distribution is uneven, affecting product quality.
[0003] With the continuous improvement of injection molding process requirements, the requirements for temperature control accuracy and response speed have also increased. However, the existing temperature control methods have limitations in dealing with the uniformity of the multi-region temperature field and dynamic response characteristics, and it is difficult to effectively solve problems such as uneven temperature field distribution and insufficient control accuracy. Especially in dealing with the optimization of the mixed action space and the identification of the dynamic characteristics of the temperature field, traditional methods are difficult to meet the high-precision control requirements. Summary of the Invention
[0004] The main purpose of the present invention is to provide a temperature control method, device and computer equipment for an injection mold, which realizes the coordinated control of the heating stage and the stable stage, shortens the injection molding cycle and reduces energy consumption.
[0005] To achieve the above purpose, the present invention provides a temperature control method for an injection mold, including the following steps: Collect temperature data of multiple temperature measurement points of the injection mold to obtain a temperature data matrix; Establish a heating and cooling input matrix and a temperature response matrix according to the temperature data matrix, and solve the optimization equation to obtain a heat transfer path coefficient matrix; Use the temperature data matrix as the state input, construct a deep deterministic policy gradient control network including a gated recurrent unit and a fully connected layer, and perform a linear mapping on the cooling water flow rate and the inlet water temperature to obtain a mixed action space; Input the heat transfer path coefficient matrix and the mixed action space into the training environment, iteratively train the deep deterministic policy gradient control network based on the experience replay mechanism, update the network parameters by calculating the reward function, and obtain a temperature control model; Deploy the temperature control model to the target control system, collect the real-time temperature data of each temperature measurement point, and calculate the real-time control parameter set according to the temperature control model.
[0006] The present invention also provides a temperature control device for an injection mold, including: A collection module for collecting temperature data of multiple temperature measurement points of the injection mold to obtain a temperature data matrix; A solution module, configured to establish a heating and cooling input matrix and a temperature response matrix based on the temperature data matrix, and solve an optimization equation to obtain a heat transfer path coefficient matrix; A mapping module, configured to use the temperature data matrix as a state input, construct a deep deterministic policy gradient control network including a gated recurrent unit and a fully connected layer, and perform a linear mapping on the cooling water flow rate and the inlet water temperature to obtain a mixed action space; A training module, configured to input the heat transfer path coefficient matrix and the mixed action space into a training environment, iteratively train the deep deterministic policy gradient control network based on an experience replay mechanism, update network parameters by calculating a reward function, and obtain a temperature control model; A deployment module, configured to deploy the temperature control model to a target control system, collect real-time temperature data of each temperature measurement point, and calculate a real-time control parameter set according to the temperature control model.
[0007] The present invention further provides a computer device, including a memory and a processor, where a computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0008] The present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0009] In summary, the technical solution provided by the present invention improves the depth deterministic policy gradient algorithm by introducing a gated recurrent unit, provides historical information on the temperature change of the mold for the DDPG agent, and improves the decision-making quality of temperature control; uses a linear mapping technology to process the mixed action space, solves the problem of making decisions in a mixed action space such as the flow rate and temperature of the cooling water path in an injection mold, and effectively avoids the "curse of dimensionality" problem; an optimized temperature field transfer path analysis method based on the total least squares method takes into account the measurement errors in the heating / cooling input and temperature response matrix, and provides a better solution for overdetermined linear equations under actual working conditions; through a deep reinforcement learning method with an experience replay mechanism and a soft update strategy, the convergence time of the temperature control system is significantly reduced, and a high-precision temperature field data acquisition and preprocessing system is constructed in combination with a fast-response temperature sensor and a data correction method, ensuring the accuracy of the temperature field data; adopts a hierarchical control strategy to achieve coordinated control in the heating stage and the stable stage, shorten the injection molding cycle and reduce energy consumption. Description of the Drawings
[0010] Figure 1 is a schematic diagram of the steps of an injection mold temperature control method in an embodiment of the present invention; Figure 2It is a structural block diagram of a temperature control device for an injection mold in an embodiment of the present invention; Figure 3 It is a structural schematic block diagram of a computer device in an embodiment of the present invention.
[0011] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0012] In order to make the object, technical solution and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0013] Referring to Figure 1 , this embodiment provides a method for controlling the temperature of an injection mold, including the following steps: S1, collect temperature data at multiple temperature measurement points of the injection mold to obtain a temperature data matrix; Among them, fast-response temperature sensors are reasonably arranged in the heating and cooling source areas and the target temperature measurement areas of the injection mold. Temperature data is collected in real time through these sensors. These sensors have high response speeds and accuracies to capture the dynamic changes in the mold temperature and record them as raw data in the form of a time series. The raw time series temperature data is standardized to eliminate numerical deviations caused by different temperature measurement points or measurement conditions, generating standard time series temperature data. Outlier detection and elimination are performed on the standard time series temperature data. By analyzing the data distribution, outlier points, such as excessively high or low values, are identified using statistical methods or machine learning algorithms, generating denoised temperature data. Data reconstruction is performed on the denoised temperature data using interpolation methods or other algorithms to restore the data to a continuous form, generating a target temperature data sequence that reflects the dynamic temperature changes at each temperature measurement point of the mold. According to the target temperature data sequence, the data is rearranged according to the number of the temperature measurement point and the sampling time to generate an initial temperature matrix. The initial temperature matrix is a two-dimensional structure with time and space as dimensions, where the rows represent the sampling time, the columns represent the numbers of the temperature measurement points, and each element of the matrix is the temperature value of a certain temperature measurement point at a certain time. On this basis, by determining the actual physical positions of the temperature measurement points in the mold structure diagram, a three-dimensional space coordinate system is established based on the position coordinates of each temperature measurement point. By associating the three-dimensional positions of the temperature measurement points with the columns of the initial temperature matrix, a spatial distribution model of the temperature measurement points is obtained. The spatial distribution model of the temperature measurement points is used to perform spatial mapping on the initial temperature matrix, and the temperature gradient and heat flux distribution between adjacent temperature measurement points are calculated. The calculation of the temperature gradient is based on the relationship between the temperature difference between adjacent temperature measurement points and their physical distance, and the heat flux distribution is calculated by combining material property parameters such as thermal conductivity with the temperature gradient. Through this step, a temperature field distribution data describing the internal heat conduction characteristics of the mold is obtained, reflecting the temperature change law of the entire mold. The temperature field distribution data is fused with the initial temperature matrix, and the temporal characteristics of the initial data and the spatial characteristics of the temperature field are integrated through weighted superposition or other fusion algorithms to finally generate a temperature data matrix that reflects the temperature status of each area of the mold.
[0014] S2. Establish a heating and cooling input matrix and a temperature response matrix based on the temperature data matrix, and solve the optimization equation to obtain a heat transfer path coefficient matrix; Specifically, the input power data of the heating and cooling sources are extracted from the temperature data matrix, and the data are organized in a time series manner to form the heating and cooling input matrix. Each column of this matrix corresponds to a specific time point, and each row represents the input power of different heating or cooling sources, reflecting the heat source operation characteristics of the system at different times. At the same time, the temperature data of each temperature measurement point in the temperature data matrix are extracted in the corresponding time series and reorganized along the time axis to form the temperature response matrix. The structure of this matrix is similar to that of the input matrix. Each row represents the temperature response of different temperature measurement points, and each column corresponds to a certain time point, recording the dynamic changes of the temperature field inside the mold. The heating and cooling input matrix and the temperature response matrix are substituted into the heat transfer path equation to construct the optimization objective function. The heat transfer path equation is based on the physical laws of mold heat conduction and describes the relationship between the input power and the temperature response. Its general form is , where is the temperature response matrix,[[]] is the heating and cooling input matrix,[[]] is the heat transfer path coefficient matrix to be solved. To ensure the accuracy of the optimization process, the least squares method is used to construct the optimization objective function, and its expression is . Through this function, the heat transfer path coefficient matrix that minimizes the temperature response error is found. During the optimization process, the singular value decomposition is performed on the heating and cooling input matrix. By singular value decomposition, the input matrix is decomposed into three parts: the left singular matrix, the singular value diagonal matrix, and the right singular matrix. This decomposition form effectively reduces the complexity of matrix calculation and extracts the main features of the matrix, making the optimization equation more stable when solving the heat transfer path coefficient. Based on the results of the singular value decomposition, the solution process of the heat transfer path equation is simplified, and the initial heat transfer path coefficient is obtained quickly. Analyze the stability and error characteristics of the initial heat transfer path coefficient. Through the perturbation analysis of the heat transfer path coefficient, the error distribution characteristics of the coefficient under different conditions are revealed. This process uses the Monte Carlo method or random simulation technology to perturb the initial heat transfer path coefficient multiple times and calculate the corresponding error range. Through the error distribution characteristics, the uncertainty of the initial coefficient is quantified. Based on the error distribution characteristics, the optimal regularization parameter is determined. The selection of the regularization parameter directly affects the accuracy and generalization ability of the heat transfer path coefficient. Through cross-validation or Bayesian optimization methods, the regularization parameter that can balance the fitting accuracy and error suppression is found and substituted into the optimization objective function for correction to obtain the corrected heat transfer path coefficient. The corrected heat transfer path coefficient is used to complete the matrix reconstruction to generate the heat transfer path coefficient matrix, which describes the heat transfer path between each heating and cooling input and each temperature measurement point in the mold and reveals the internal heat flow distribution law of the mold.
[0015] S3. Take the temperature data matrix as the state input, construct a deep deterministic policy gradient control network including a gated recurrent unit and a fully connected layer, and perform a linear mapping on the cooling water flow rate and the inlet temperature to obtain a mixed action space; It should be noted that the temperature data matrix is processed to extract the historical temperature change information through time series analysis, generate a temperature state sequence, and reflect the dynamic evolution of the mold temperature field over time. The sliding window method is used to separate continuous time periods from the temperature data matrix, and the sequence is preprocessed by techniques such as normalization, differencing, and denoising to improve the stability of the input data and the efficiency of network learning. A basic network framework is designed based on the temperature state sequence, which consists of an execution network and an evaluation network. In the design of the execution network, a three-layer gated recurrent unit (GRU) is used as the core module. The GRU can effectively process the time dependence and non-linear characteristics of the temperature data. The number of neurons in each layer's hidden state is set to 128 to provide sufficient capacity to capture complex dynamic relationships. Two fully connected layers are added to the GRU layer, and the ReLU activation function is used to perform non-linear mapping on the features to enhance the network's expressive ability. At the same time, the parameters of the basic network framework are initialized using a uniform distribution or the Xavier initialization method to ensure the stability in the initial stage of network training. An evaluation network is constructed based on the basic network framework. The design of the evaluation network adopts a four-layer fully connected layer structure, and the number of neurons connected in each layer decreases sequentially to extract higher-level value evaluations from the features generated by the execution network. Through this evaluation network, the deep deterministic policy gradient control network can learn the correlation between each set of control actions and the system performance, providing guidance for policy optimization. After integrating the execution network and the evaluation network, the deep deterministic policy gradient control network is obtained. A cooling water flow rate mapping function is constructed based on the deep deterministic policy gradient control network to map the flow rate control range [0, 20] L / min to the [-1, 1] interval, obtaining a standardized flow rate action space. Similarly, for the inlet temperature, a corresponding mapping function is constructed to map the temperature control range [15, 95] °C to the [-1, 1] interval, obtaining a standardized temperature action space. An initial mixed action space is constructed based on the standardized flow rate action space and the standardized temperature action space. The mixed action space combines two control variables, flow rate and temperature, providing flexible control options to adapt to different temperature control requirements. To prevent actions from exceeding the physically feasible range, action output limit conditions are introduced on the mixed action space, and the action range is constrained by the tanh function. The output range of the tanh function is in [-1, 1], which matches the standardized action space. At the same time, its smooth gradient characteristics contribute to the stable training of the network. Generate a mixed action space that meets the physical limit conditions.
[0016] S4. Input the heat transfer path coefficient matrix and the mixed action space into the training environment, and iteratively train the deep deterministic policy gradient control network based on the experience replay mechanism. Update the network parameters by calculating the reward function to obtain the temperature control model; Specifically, create a training environment that simulates the temperature field of the injection mold according to the heat transfer path coefficient matrix and the mixed action space. The heat transfer path coefficient matrix is used to describe the heat transfer relationship inside the system, while the mixed action space defines the control range and behavior space of the cooling water flow rate and the inlet water temperature. By combining these two parts, a temperature control environment model that can dynamically respond to control actions is established. After the environment is constructed, to ensure the stable training of the deep deterministic policy gradient control network, set the parameter update rates of the execution network and the evaluation network to 0.001, and at the same time adopt a soft update strategy to ensure the smoothness and convergence of network updates. At the beginning of training, use the current temperature state data in the training environment as input, and perform forward propagation calculations through the execution network of the deep deterministic policy gradient control network to generate control actions. The control actions include specific cooling water flow rate and inlet water temperature values. Combining with the heat transfer path coefficient matrix, simulate the change of the mold temperature state under the current control strategy to obtain the next state. After the state transition is completed, design a reward function to quantify the quality of the current action. The reward function is constructed based on the target temperature distribution and the control energy consumption, and its form is , where represents the target temperature distribution, is the current temperature distribution, is the control energy consumption, is the balance coefficient. Use the calculated reward value to form a state transition quadruple of the current state, control action, reward value, and next state. Store the state transition quadruple in the experience replay pool to support the diversity and stability of data during the training process. In each iteration, randomly sample from the experience replay pool Samples are used to construct a mini-batch dataset as the training sample set. For each sample, the target Q-value is calculated through the evaluation network, and the parameters of the evaluation network are updated using the temporal difference algorithm to optimize the prediction accuracy of the Q-value. During this process, the gradients of the evaluation network are recorded to guide the optimization of the execution network. The parameters of the execution network are updated using the policy gradient algorithm with the gradients of the evaluation network. The goal of the execution network is to maximize the expected reward value. During the optimization process, the gradient ascent method is used to enable the execution network to continuously learn and improve the policy, generating better control actions. As the training progresses, the gradients of the evaluation network and the execution network are accumulated. Through this joint optimization, the network can gradually converge to the optimal solution in complex temperature control tasks. The convergence criterion during training is that when the number of training episodes reaches a preset first target value and the change rate of the average reward value for N consecutive episodes is less than a set second target value, it is considered that the training has reached the convergence state. At this time, the deep deterministic policy gradient control network can stably generate high-quality control policies, ultimately forming a complete temperature control model. This model combines the physical description of the heat transfer path coefficient and the policy optimization ability of deep learning, and can dynamically adjust the cooling water flow rate and inlet water temperature according to the real-time mold temperature state, thereby minimizing energy consumption while ensuring the uniformity of the mold temperature.
[0017] Obtain the current temperature status data from the training environment, and perform sequence partitioning on it to form a temperature time-series feature sequence. This sequence can capture the variation law of the mold temperature field in the time dimension. To improve the feature expressiveness of the input data, ReLU activation processing is applied to the temperature time-series feature sequence. Nonlinear processing can effectively enhance the expression of positive features, while filtering out invalid data to generate temperature feature vectors. Input the temperature feature vectors into the execution network of the deep deterministic policy gradient control network. The execution network performs recurrent propagation calculations through its three-layer gated recurrent unit (GRU) to extract deep feature representations from the time-series data. The design of the GRU enables the network to effectively capture the temporal correlation of the temperature data and gradually refine the high-dimensional feature representations representing the temperature variation law. The extracted deep features are passed to the two-layer fully connected layer of the execution network, and through nonlinear mapping and feature combination, the deep features are transformed into action prediction vectors. The action prediction vectors contain the predicted values of the cooling water flow rate and the inlet water temperature. To convert the action prediction vectors into actual operable physical control values, interval mapping is performed. The action prediction vectors are smoothed through the tanh activation function and mapped to the limited range of the cooling water flow rate [0, 20] L / min and the range of the inlet water temperature [15, 95] °C. According to the generated control actions and the heat transfer path coefficient matrix, the mold heat transfer process is simulated using matrix operations, and the temperature responses of each temperature measurement point under this control strategy are calculated to obtain the temperature distribution of the next state. Based on this temperature distribution, the temperature gradient is calculated through the temperature difference between adjacent temperature measurement points to reflect the actual intensity of heat conduction inside the mold. At the same time, the temperature distribution of the next state is compared with the target temperature distribution, and the temperature deviation value of each temperature measurement point is calculated to evaluate the effect of the control strategy. To quantify the quality of the control strategy, a reward function is constructed to calculate the reward value. The reward function comprehensively considers the balance between temperature control accuracy and energy consumption, and obtains the comprehensive evaluation of the current strategy by assigning different weights to the deviation value and the energy consumption contribution. The magnitude of the reward value directly reflects the effect of the current control action. Combine the current temperature status data, control actions, reward values, and the temperature distribution of the next state into state transition samples, and convert them into tensor form for the use of the deep learning framework.
[0018] S5. Deploy the temperature control model to the target control system, collect the real-time temperature data of each temperature measurement point, and calculate the real-time control parameter set according to the temperature control model.
[0019] Among them, the temperature control model is imported into the target control system, and the data acquisition module in the system is configured to generate an operating parameter set. The generation of the operating parameter set includes links such as setting the data acquisition frequency, calibrating the positions of temperature measurement points, and debugging the signal transmission protocol, ensuring that the temperature data of all temperature measurement points can be accurately collected and timely transmitted to the system processing module. According to the operating parameter set, the temperature data of each temperature measurement point of the injection mold is collected in real time. Through the data acquisition module, the temperature data of each temperature measurement point is read in real time, reflecting the dynamic temperature distribution inside the mold. The collected real-time temperature data is input into the temperature control model, and the model constructs a historical temperature sequence for the data through an internal algorithm, generating a temperature feature sequence reflecting the temperature change trend. The temperature feature sequence is input into the execution network of the temperature control model for forward calculation. The model outputs the control values of the cooling water flow rate and the inlet water temperature by comprehensively considering factors such as historical temperature changes and heat transfer path coefficients. These control values constitute real-time control instructions for driving the subsequent actuators. The real-time control instructions are transmitted to the execution module of the target control system. The signal of the cooling water flow rate control instruction is converted by a frequency converter into a corresponding water pump speed control signal to obtain a flow rate adjustment instruction. The change in the speed of the water pump directly determines the flow rate of the cooling water, providing the required cooling capacity for the mold. At the same time, the inlet water temperature control instruction is transmitted to the electric control valve, and through the dynamic adjustment function of the control valve, the inlet water temperature is accurately regulated to generate a temperature adjustment instruction. The flow rate adjustment and temperature adjustment act together on the heating and cooling systems of the mold to ensure that the temperature field is maintained within the set range and the temperature inside the mold is evenly distributed. During the execution of the control instructions, the target control system monitors the execution status of the flow rate adjustment instruction and the temperature adjustment instruction in real time to ensure that the instructions can be accurately executed and the expected effects can be obtained. At the same time, the parameters and temperature response data related to each control are recorded to generate operating status data. The operating status data includes information such as flow rate, temperature, and system power consumption. The operating status data is stored in the database of the system to ensure that all key data in the control process are recorded and stored. A parameter set containing real-time control information is generated.
[0020] In one example, temperature data of multiple temperature measurement points of the injection mold is collected to obtain a temperature data matrix, including: Quick response temperature sensors are installed in the heating and cooling source areas and the target temperature measurement areas of the injection mold for data collection to obtain the original time series temperature data, and the original time series temperature data is standardized to obtain the standard time series temperature data; Outlier detection and elimination are performed on the standard time series temperature data to obtain the denoised temperature data, and data reconstruction is performed on the denoised temperature data to obtain the target temperature data sequence; The target temperature data sequence is rearranged into a matrix according to the numbering of temperature measurement points and the sampling time to obtain an initial temperature matrix, and a three-dimensional space coordinate system is established based on each temperature measurement point of the initial temperature matrix to obtain a spatial distribution model of temperature measurement points; According to the spatial distribution model of temperature measurement points, spatial mapping is performed on the initial temperature matrix, the temperature gradient and heat flow distribution between adjacent temperature measurement points are calculated to obtain temperature field distribution data, and data fusion is performed on the temperature field distribution data and the initial temperature matrix to obtain a temperature data matrix.
[0021] In this example, fast-response temperature sensors are arranged at the heating source, cooling source, and important target temperature measurement areas of the injection mold to collect the original data of temperature change over time. It is assumed that temperature measurement points are arranged in the mold, and the temperature data of each point is obtained at a fixed sampling interval to form time series data , where represents the numbering of the temperature measurement point, is the sampling time, and is the index of the time point. The original time series temperature data is standardized to eliminate numerical deviations caused by different temperature measurement points or sampling conditions. The standardization formula is: ; where, is the standardized temperature value, is the average temperature of the th temperature measurement point, and the calculation formula is: ; is the temperature standard deviation of the th temperature measurement point, which is expressed by the formula: ; where is the total number of samplings. After standardization, outlier detection and elimination are performed on the standard time series temperature data. Outlier detection adopts the rule based on three times the standard deviation, that is, when a certain data point satisfies: ; it is regarded as an outlier and eliminated. The missing data after elimination is reconstructed by interpolation method, such as linear interpolation method: ; to obtain complete denoised temperature data . The denoised temperature data is sorted out, and an initial temperature matrix is constructed with the temperature measurement point number as the row and the sampling time as the column, and its form is: ; Based on each temperature measurement point of the initial temperature matrix and combining with the mold structure, a three-dimensional space coordinate system is established. Assuming the position of the th temperature measurement point is , a spatial distribution model of the temperature measurement points is constructed. Based on this model, a spatial mapping of the initial temperature matrix is performed to calculate the temperature gradient and heat flux distribution between adjacent temperature measurement points. The calculation formula for the temperature gradient is: ; where and are the numbers of adjacent temperature measurement points, and the heat flux distribution is calculated according to the thermal conductivity of the material: ; The temperature gradient and heat flux distribution data calculated above are fused with the initial temperature matrix , and temperature field distribution data are generated through weighted average or interpolation methods, and finally a temperature data matrix is formed.
[0022] In an example, a heating and cooling input matrix and a temperature response matrix are established based on the temperature data matrix, and an optimization equation is solved to obtain a heat transfer path coefficient matrix, including: Performing time series extraction on the heating and cooling source input power data in the temperature data matrix to obtain a heating and cooling input matrix, and performing time series extraction on the temperature data of the temperature measurement points in the temperature data matrix to obtain a temperature response matrix; Substituting the heating and cooling input matrix and the temperature response matrix into the heat transfer path equation to construct an optimization objective function to obtain an optimization equation, and performing singular value decomposition operation on the heating and cooling input matrix to obtain a matrix decomposition result; Solving the heat transfer path coefficient according to the matrix decomposition result to obtain an initial heat transfer path coefficient, and performing perturbation analysis on the initial heat transfer path coefficient to obtain an error distribution characteristic; Based on the error distribution characteristic, determining an optimal regularization parameter, substituting the optimal regularization parameter into the optimization equation to obtain a corrected heat transfer path coefficient, and performing matrix reconstruction according to the corrected heat transfer path coefficient to obtain a heat transfer path coefficient matrix.
[0023] In this example, the input power data of the heating and cooling sources are extracted from the temperature data matrix. Assuming there are heating and cooling sources in the system, and the power input of each source at different times is , where Indicates the number of the heating or cooling source, represents the sampling time. Through time series extraction, these data are organized into a heating and cooling input matrix , in the form of: ; where is the total number of samplings. At the same time, similar time series extraction is performed on the temperature data of the temperature measurement points in the temperature data matrix. Assuming there are temperature measurement points, and their temperatures at different times are , which are also organized into a temperature response matrix : ; Substitute the input matrix and the response matrix into the heat transfer path equation to construct the optimization objective function. The heat transfer path equation describes the relationship between the input power and the temperature response, and its mathematical form is: ; where is the heat transfer path coefficient matrix to be solved, which describes the heat transfer relationship between the heating or cooling source and each temperature measurement point. To solve , an optimization objective function is constructed to minimize the error between the temperature prediction value and the actual temperature response matrix . The optimization objective function is expressed as: ; where represents the Frobenius norm of the matrix, that is, the square root of the sum of the squares of the elements, which is used to measure the error size between matrices. To reduce the computational complexity, the heating and cooling input matrix is subjected to singular value decomposition and decomposed into the product of three matrices: ; where, is the left singular matrix, which contains the main direction characteristics of the input matrix; is the diagonal singular value matrix, which represents the singular values of the input matrix; is the transpose of the right singular matrix, which is used to describe the orthogonal basis of the input matrix. Using the results of singular value decomposition, the optimization problem is simplified to a lower-dimensional calculation, thereby accelerating the solution of the heat transfer path coefficient matrix . According to the decomposition results, an initial solution of is obtained to get the initial heat transfer path coefficient matrix . Perturbation analysis is performed on the initial heat transfer path coefficient matrix . During the perturbation analysis, by Introduce small - scale random variations and observe their impact on the optimization objective function to determine the error distribution characteristics. These characteristics are quantified through statistical methods, such as calculating the range of changes in the objective function values before and after perturbation. Based on the error distribution characteristics, determine the optimal regularization parameter , to suppress errors and overfitting phenomena during the optimization process. After substituting the optimal regularization parameter into the optimization objective function, the optimization equation is updated as follows: ; where the regularization term is used to limit the excessive fluctuations of the heat transfer path coefficient matrix, ensuring the stability and physical meaning of the solution. By iteratively optimizing and solving the above - updated optimization equation, a corrected heat transfer path coefficient matrix is obtained. Use this matrix for matrix reconstruction to generate the final heat transfer path coefficient matrix, which is used to describe the heat transfer characteristics in the entire heating and cooling system.
[0024] In one example, take the temperature data matrix as the state input, construct a deep deterministic policy gradient control network including a gated recurrent unit and a fully - connected layer, and perform a linear mapping on the cooling water flow rate and the inlet water temperature to obtain a mixed action space, including: Conduct a temporal analysis and data processing of the temperature history changes of the temperature data matrix to obtain a temperature state sequence; Configure the basic network framework according to the temperature state sequence. The execution network in the basic network framework includes three layers of gated recurrent units and two layers of fully - connected layers, and initialize the parameters of the basic network framework, set the number of neurons in the hidden layer to 128, and the activation function to the ReLU function; Construct an evaluation network based on the basic network framework. The evaluation network branch is four layers of fully - connected layers to obtain a deep deterministic policy gradient control network; Construct a cooling water flow rate mapping function based on the deep deterministic policy gradient control network, map the flow rate control range [0, 20] L / min to the interval [-1, 1] to obtain a standardized flow rate action space, and construct an inlet water temperature mapping function based on the deep deterministic policy gradient control network, map the temperature control range [15, 95] °C to the interval [-1, 1] to obtain a standardized temperature action space; Construct an initial mixed action space based on the standardized flow rate action space and the standardized temperature action space, and set action output limit conditions according to the initial mixed action space, and perform action range constraint through the tanh function to obtain a mixed action space.
[0025] In this example, conduct a temporal analysis and data processing of the temperature data matrix. Assume the temperature data matrix represents in the mold temperature measurement points at The temperature distribution at a time step, with the structure: ; where is the temperature value of the th temperature measurement point at time . Through time series analysis, the temperature sequence of each temperature measurement point in the matrix is processed with a sliding window to capture the historical change characteristics of the temperature. Assuming the window size is , a subsequence is generated within each window: ; After integrating the results of all windows, a temperature state sequence is constructed, expressed as: ; This sequence captures the dynamic changes of the mold temperature in time and space. Based on the temperature state sequence , a basic network framework is configured to construct an execution network, which includes three layers of gated recurrent units (GRUs) and two layers of fully connected layers. GRU is a special recurrent neural network structure that can effectively handle the dependencies in time series data, and its state update formula is: ; where, is the current hidden state; is the update gate, representing the weight of the current state to historical information; is the candidate hidden state; represents element-wise multiplication. The number of neurons in the hidden layer of GRU is set to 128, and the activation function uses the ReLU function, whose mathematical expression is: ; This configuration can ensure that the network has sufficient expressive power while avoiding the problem of gradient vanishing. The deep features extracted by GRU are mapped to the action space through two layers of fully connected layers to provide a preliminary prediction for the control of the cooling water flow rate and the inlet water temperature. Based on the execution network, an evaluation network is constructed to evaluate the effect of each set of control actions. The evaluation network consists of four layers of fully connected layers, with the input coming from the output features of the execution network and the current state information, generating the Q value for each action, representing the quality of the action. The update of the Q value follows the Bellman equation, which is used to guide the policy optimization of the network. Based on the deep deterministic policy gradient control network, a mapping function for the cooling water flow rate and the inlet water temperature is constructed. The control range of the cooling water flow rate is [0, 20] L / min, which is mapped to the standardized interval [-1, 1], and the mapping formula is: ; where is the actual cooling water flow rate, is the flow rate action after standardization. The control range of the inlet water temperature is [15, 95] °C, and the mapping formula is: ; where is the actual inlet water temperature, is the temperature action after standardization. Through the above mapping, a standardized flow rate action space and a temperature action space are generated. Based on these two spaces, an initial mixed action space is constructed, and its form is: ; To ensure the feasibility of the control action in actual operation, the tanh function is used to constrain the range of the action output to obtain the mixed action space. The formula is: ; where is the action predicted by the network, is the final action after constraint.
[0026] In an example, the heat transfer path coefficient matrix and the mixed action space are input into the training environment, and the deep deterministic policy gradient control network is iteratively trained based on the experience replay mechanism. The network parameters are updated by calculating the reward function to obtain the temperature control model, including: Construct an environment for the heat transfer path coefficient matrix and the mixed action space to obtain the training environment, and set the update rates of the execution network and the evaluation network in the deep deterministic policy gradient control network to 0.001 to obtain the soft update parameters; Predict the action for the current temperature state data in the training environment, input the execution network of the deep deterministic policy gradient control network for forward propagation calculation to obtain the control action, calculate the next state according to the control action and the heat transfer path coefficient matrix, construct the reward function to calculate the reward value, and obtain the state transition quadruple; Store the state transition quadruple in the experience replay pool, randomly sample M samples from the experience replay pool to construct a mini-batch of data to obtain the training sample set, calculate the target Q value for the training sample set, and update the evaluation network parameters using the temporal difference algorithm to obtain the evaluation network gradient; Update the execution network parameters according to the evaluation network gradient using the policy gradient algorithm, and optimize the expected reward value using the gradient ascent method to obtain the execution network gradient; Accumulate the evaluation network gradient and the execution network gradient. When the number of training episodes reaches the first target value and the change rate of the average reward value for N consecutive episodes is less than the second target value, the temperature control model is obtained.
[0027] In this example, an environment is constructed for the heat transfer path coefficient matrix and the mixed action space, and a training environment is created to simulate the heat transfer behavior of the mold. Assume the heat transfer path coefficient matrix represents the heat transfer relationship between the heating and cooling sources and each temperature measurement point, and its structure is: ; where represents the heat transfer coefficient of the th heat source to the th temperature measurement point. Combining the mixed action space , which consists of the normalized actions of the cooling water flow rate and the inlet water temperature, is defined as: ; Through these inputs, the training environment can dynamically simulate the influence of each set of control actions on the temperature field of the mold, providing an interactive scenario for reinforcement learning. During the training process, the current temperature state data is input into the execution network of the deep deterministic policy gradient control network for forward propagation calculation. The structure of the execution network includes three layers of gated recurrent units (GRUs) and two layers of fully connected layers. Assume the input state is , and the output is the action , then the mathematical expression of the forward propagation is: ; where is the function of the execution network, includes the control values of the cooling water flow rate and the inlet water temperature, and calculates the next state temperature distribution through the heat transfer path coefficient matrix and the action : ; After obtaining the next state, according to the target temperature distribution and the current temperature distribution, a reward function is constructed to quantify the effect of the control strategy. Assume the target temperature is , the current temperature is , and the calculation formula of the reward function is: ; where, represents the sum of the squares of the temperature deviations, represents the control energy consumption, is the balance coefficient, which is used to adjust the weight between the temperature control accuracy and the energy consumption optimization. The generated state transition data is organized into state transition quadruples and stored in the experience replay pool. The experience replay pool avoids the time dependence of the data through random sampling, improving the stability and efficiency of training. In each training iteration, randomly sample from the experience replay pool A sample constitutes a small batch of data, and the target Q value is calculated ; Among them is the discount factor, representing the decay rate of future rewards, is the Q value predicted by the evaluation network. According to the target Q value , the parameters of the evaluation network are updated using the temporal difference (TD) algorithm : ; Among them is the learning rate, controlling the step size of parameter update. At the same time, using the gradient of the evaluation network, the parameters of the execution network are updated through the policy gradient algorithm , and the goal is to maximize the expected reward value: ; In each training iteration, the parameters of the execution network and the evaluation network are synchronized in a soft update manner, and the update rate is set to 0.001. The soft update formula is: ; Among them is the soft update coefficient, ensuring the stable update of the network. The termination condition of the training process is: when the number of training episodes reaches the preset target value , and the change rate of the average reward value in consecutive episodes is less than the set threshold , that is, satisfying: ; At this time, it is considered that the model has converged, and the final temperature control model is obtained.
[0028] Specifically, this embodiment further includes: interpolating and reconstructing the temperature measurement point data in the temperature data matrix according to three-dimensional space coordinates, establishing a two-dimensional temperature field image with a resolution of 128×128, and performing temporal combination on the temperature field images of 10 consecutive time steps to obtain a sequence of temperature field feature maps; constructing a physical information convolution kernel with a size of 5×5 according to Fourier's heat conduction law in heat transfer, embedding physical parameters such as thermal conductivity, specific heat capacity, and density into the convolution kernel weights, and discretizing the partial differential equation by the finite difference method to obtain a physical constraint convolution operator; inputting the sequence of temperature field feature maps into a three-layer physical information convolution network, setting 64 physical constraint convolution operators in each layer, with a convolution step size of 1, and adding a ReLU activation function after each layer of convolution to obtain a physical feature mapping matrix; performing feature dimensionality reduction on the physical feature mapping matrix through a 2×2 max pooling layer with a pooling step size of 2, and performing non-linear mapping on the dimensionality-reduced features through two fully connected layers with 256 hidden units to obtain a 128-dimensional physically enhanced feature vector; concatenating the physically enhanced feature vector with the temperature state feature output by the deep deterministic policy gradient control network in the channel dimension to construct a 256-dimensional hybrid feature space, and performing feature weighting through an attention mechanism to obtain a fused feature vector; correcting the action output of the deep deterministic policy gradient control network based on the fused feature vector, and restricting the consistency between the control action and the heat transfer law through a physical constraint layer to obtain a control instruction that satisfies the physical constraint; substituting the control instruction into the heat transfer model, and numerically solving the temperature field evolution equation by the Runge-Kutta method to predict the temperature distribution in the next 5 time steps to obtain a temperature prediction result; calculating the root mean square error between the temperature prediction result and the measured temperature data, and updating the parameters of the physical information convolution kernel based on the backpropagation algorithm. When the prediction error is less than 0.1°C, an optimized temperature control model is obtained.
[0029] In one example, action prediction is performed on the current temperature state data in the training environment. The execution network of the deep deterministic policy gradient control network is input for forward propagation calculation to obtain a control action. The next state is calculated based on the control action and the heat transfer path coefficient matrix, and a reward function is constructed to calculate the reward value, obtaining a state transition quadruple, including: The current temperature state data in the training environment is divided into sequences to construct a temperature temporal feature sequence, and the temperature temporal feature sequence is subjected to ReLU activation processing to obtain a temperature feature vector; The temperature feature vector is input into the three-layer gated recurrent unit of the execution network in the deep deterministic policy gradient control network for recurrent propagation calculation to obtain a deep feature representation, and the deep feature representation is subjected to mapping transformation through the two fully connected layers of the execution network in the deep deterministic policy gradient control network to obtain an action prediction vector; The action prediction vector is mapped to the intervals through the tanh activation function, corresponding to the range of [0, 20] L / min for the cooling water flow rate and the range of [15, 95] °C for the inlet temperature, respectively, to obtain the control action. According to the control action and the heat transfer path coefficient matrix, matrix multiplication is performed to calculate the temperature response of each temperature measurement point, obtain the next state temperature distribution, calculate the temperature gradient of adjacent temperature measurement points in the next state temperature distribution, and calculate the deviation from the target temperature to obtain the temperature deviation value. Based on the temperature deviation value, a reward function is constructed, and the weight coefficient α = 0.5 is set for weighted calculation to obtain the reward value. The current temperature state data, control action, reward value, and next state temperature distribution are combined into a state transition sample and converted into a tensor form to obtain a state transition quadruple.
[0030] In this example, the temperature data is divided into sequences to extract the temporal features of historical temperature changes. Assume that the current temperature state data is represented in matrix form as , where represents the current temperature distribution of temperature measurement points in the mold. In the sequence division, the data of the past time steps are combined to form a temperature temporal feature sequence , and its form is: ; Among them, is a matrix, and each column represents the temperature distribution at a certain moment. The temperature temporal feature sequence is subjected to non-linear activation processing. The ReLU (Rectified Linear Unit) activation function is used to effectively remove invalid data and enhance the expressiveness of positive features, and its mathematical expression is: ; Through ReLU processing, the temperature temporal feature sequence is converted into a temperature feature vector , that is: ; Among them, is the activated feature vector, which retains the meaningful information in the temperature temporal changes. The generated temperature feature vector is input into the execution network of the deep deterministic policy gradient control network, and cyclic propagation calculation is performed through three layers of gated recurrent units (GRUs). The design of the GRU can capture the time-dependent relationships in the temperature sequence, and its core formula is: ; Among them, is the hidden state at the current moment; is the update gate, which is used to control the proportion of historical information and new information; is the candidate hidden state, representing the information of the current input; represents element-wise multiplication. The number of neurons in the hidden layer of the GRU is set to 128. After three layers of recursive calculation, a deep feature representation is obtained , and its form is: ; where are the parameters of the GRU, , is the output dimension of the GRU. The deep feature representation is input into two fully connected layers of the execution network for mapping transformation. Assume that the weights of the first and second fully connected layers are and respectively, and the biases are and respectively. Then the calculation formula for the action prediction vector is: ; where represents the predicted values of the cooling water flow rate and the inlet water temperature. To map the predicted action to the actual control range, the tanh activation function is used for interval mapping. The mathematical expression of the tanh function is: ; The output range of tanh is mapped from [-1, 1] to the actual range: ; ; where and are the actual control values of the cooling water flow rate and the inlet water temperature respectively. According to the control action and the heat transfer path coefficient matrix , the next state temperature distribution is calculated as: ; where is the control action vector, is the heat transfer path coefficient matrix, describing the heat transfer relationship between the heating and cooling sources and the temperature measurement points. After obtaining the next state temperature distribution , the temperature gradient between adjacent temperature measurement points is calculated by finite difference, and the formula is: ; where and are the temperatures of the temperature measurement points and respectively, and are their spatial positions. After calculating the temperature gradient, compare it with the target temperature distribution to obtain the temperature deviation value : ; Construct a reward function based on the temperature deviation value, and set the weight coefficient in combination with the control energy consumption , and the reward value calculation formula is: ; Among them, is the energy consumption index, indicating the execution cost of the control action. Combine the current temperature state data , control action , reward value and the next state temperature distribution to form a state transition sample, and convert it into a tensor form to generate a state transition quadruple .
[0031] In an example, deploy the temperature control model to the target control system, collect the real-time temperature data of each temperature measurement point, and calculate the real-time control parameter set according to the temperature control model, including: Deploy the temperature control model to the target control system, and configure the data acquisition module of the target control system to obtain the operating parameter set; According to the operating parameter set, perform real-time temperature data acquisition on each temperature measurement point of the injection mold, and read the data through the data acquisition module to obtain the real-time temperature data of each temperature measurement point; Input the real-time temperature data of each temperature measurement point into the temperature control model to construct a historical temperature sequence, obtain the temperature feature sequence, and input the temperature feature sequence into the execution network of the temperature control model for forward calculation, output the cooling water flow rate and the inlet water temperature control value, and obtain the real-time control instruction; Perform signal conversion on the real-time control instruction through the frequency converter, convert the flow control instruction into a water pump speed control signal to obtain the flow regulation instruction, and input the flow regulation instruction into the electric control valve to dynamically adjust the inlet water temperature to obtain the temperature regulation instruction; Monitor the execution status of the flow regulation instruction and the temperature regulation instruction, record the control parameters and temperature response data to obtain the operating status data, and store the operating status data in the system database to record and store the control process data to obtain the real-time control parameter set.
[0032] In this example, load the trained temperature control model into the target control system, which includes a data acquisition module, an execution module, and a monitoring module. Configure the temperature control model with the data acquisition module in the control system, and set the operating parameter set, including the sampling frequency , sampling points position , as well as communication protocols, etc. These parameters ensure the spatio-temporal consistency of the real-time collected data. According to the configured set of operating parameters, real-time temperature data is collected for each temperature measurement point of the injection mold. Assume that there are temperature measurement points in the mold, and the real-time temperature data of each point is represented by , where is the time of the th sampling, and the temperature data matrix is expressed as: ; where is the total number of samplings. The data acquisition module reads in real time and inputs it into the temperature control model. The real-time temperature data is input into the temperature control model to construct a historical temperature sequence to capture the dynamic temperature change characteristics. Set the window length , and extract a continuous time series from the real-time data to form a temperature feature sequence : ; Integrate the feature sequences of all temperature measurement points to obtain the overall temperature feature matrix : ; Input the temperature feature matrix into the execution network of the temperature control model. The execution network includes multiple layers of gated recurrent units (GRUs) and fully connected layers, which are used to extract features and generate control instructions. Through forward propagation calculation, the control values of the cooling water flow and the inlet water temperature are output: ; where is the model of the execution network. Convert the obtained control values and into real-time control instructions. The control value of the cooling water flow is converted into the rotational speed control signal of the water pump through signal conversion by the frequency converter. Assume the linear relationship between the water pump rotational speed and the cooling water flow is: ; where is the water pump rotational speed, and is the proportional coefficient of the flow rate to the rotational speed. At the same time, input the inlet water temperature control value into the electric control valve to dynamically adjust the inlet water temperature through the valve opening. Assume the non-linear relationship between the valve opening and the inlet water temperature is: ; Among them, is the valve opening, is the proportionality coefficient, is the ambient temperature. By executing the flow rate adjustment instruction and the temperature adjustment instruction , the dynamic control of the temperature of the injection mold is realized. At the same time, the monitoring module of the target control system monitors the execution status of these instructions in real time and records the control parameters and temperature response data. The control parameters include the water pump speed , the valve opening , and the temperature response data is the real-time temperature change of the temperature measurement points in the mold . The recorded control parameters and temperature response data are sorted into operation status data : ; The operation status data is stored in the system database for subsequent analysis and optimization. The operation status data combines real-time control instructions and temperature feedback information to form a real-time control parameter set.
[0033] In this embodiment, heuristic optimization of the temperature control parameters is also included, specifically including: performing sensitivity analysis on the flow rate and temperature control variables in the real-time control parameter set, constructing an influence factor calculation function F(x)=∑(∂T / ∂xi)², where T is the temperature response value, xi is the control variable, calculating the partial derivative through the numerical difference method to obtain the parameter influence degree matrix M; sorting all control variables {x1, x2,..., xn} in descending order according to the influence degree value based on the parameter influence degree matrix M, selecting the first K variables with influence degree values greater than the preset threshold λ to construct a core variable subset, obtaining a priority variable set P; constructing a significance index function Si=ωi·|ΔTi| for each variable xi in the priority variable set P, where the weight coefficient ωi is determined by the analytic hierarchy process, and the temperature change amount ΔTi is statistically calculated based on historical data, normalizing all index values to obtain a variable significance index set S; based on the variable significance index set S, using the adaptive threshold method θ(t)=θ0·exp(-βt) to define the search neighborhood range, where θ0 is the initial threshold, β is the attenuation coefficient, and t is the number of iterations, determining the variables with significance indexes greater than the current threshold θ(t) as search objects to obtain a dynamic neighborhood space D; performing local optimization using the Tabu search strategy within the dynamic neighborhood space D, setting the taboo length to L and the neighborhood movement step size to γ, iteratively updating the variable values and recording the search trajectory to obtain a local optimal solution ; The local optimal solution Substitute into the heat transfer path coefficient matrix for temperature field simulation calculation. Based on the temperature uniformity index U = max|Ti - Tavg| and the control cost index C = ∑ui², conduct a comprehensive evaluation to obtain the improved solution set R. Screen the high-quality samples in the improved solution set R, add the screened samples {(si, ai, ri, si')} to the experience replay pool of the deep deterministic policy gradient network, and use the mini-batch stochastic gradient descent method to update the network parameters to obtain the optimized control strategy. ; According to the optimized control strategy Generate a new control instruction sequence {u1, u2,..., um}, online correct the flow rate adjustment instruction and the temperature adjustment instruction, and output the corrected control quantity through the actuator to obtain the optimized control instruction.
[0034] Refer to Figure 2 , this embodiment provides an injection mold temperature control device, including: The acquisition module 1 is used to collect temperature data of multiple temperature measurement points of the injection mold to obtain a temperature data matrix; The solution module 2 is used to establish a heating and cooling input matrix and a temperature response matrix according to the temperature data matrix, and solve the optimization equation to obtain the heat transfer path coefficient matrix; The mapping module 3 is used to use the temperature data matrix as the state input, construct a deep deterministic policy gradient control network including a gated recurrent unit and a fully connected layer, and perform a linear mapping on the cooling water flow rate and the inlet water temperature to obtain a mixed action space; The training module 4 is used to input the heat transfer path coefficient matrix and the mixed action space into the training environment, iteratively train the deep deterministic policy gradient control network based on the experience replay mechanism, and update the network parameters by calculating the reward function to obtain the temperature control model; The deployment module 5 is used to deploy the temperature control model to the target control system, collect the real-time temperature data of each temperature measurement point, and calculate the real-time control parameter set according to the temperature control model.
[0035] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to that described in the above method embodiment, and details will not be repeated here.
[0036] Refer to Figure 3 , this embodiment of the present invention also provides a computer device, which can be a server, and its internal structure can be as Figure 3As shown in the figure. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0037] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0038] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0039] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0040] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article or method comprising a series of elements not only includes those elements but also other elements not expressly listed, or further includes elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, apparatus, article or method comprising such element.
[0041] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, are similarly included in the patent protection scope of the present invention.
Claims
1. A method for controlling the temperature of an injection mold, characterized in that: The following steps are involved: Collect temperature data from multiple temperature measurement points of the injection mold to obtain a temperature data matrix; Establishing a heating and cooling input matrix and a temperature response matrix according to the temperature data matrix, and solving an optimization equation to obtain a heat transfer path coefficient matrix; Taking the temperature data matrix as state input, constructing a deep deterministic policy gradient control network including a gated recurrent unit and a fully connected layer, and linearly mapping the cooling water flow rate and the inlet water temperature to obtain a mixed action space; Inputting the heat transfer path coefficient matrix and the mixed action space into a training environment, iteratively training the deep deterministic policy gradient control network based on an experience replay mechanism, updating the network parameters by calculating a reward function, and obtaining a temperature control model; The temperature control model is deployed to the target control system, real-time temperature data of each temperature measuring point is collected, and a real-time control parameter set is calculated according to the temperature control model.
2. The injection mold temperature control method according to claim 1, characterized in that: The temperature data are collected at multiple temperature measurement points of the injection mold to obtain a temperature data matrix, including: Installing fast response temperature sensors in the heating and cooling source areas and target temperature measurement areas of the injection mold to collect data, obtain time series temperature raw data, and perform standardization on the time series temperature raw data to obtain standard time series temperature data; Performing outlier detection and elimination on the standard time series temperature data to obtain denoised temperature data, and performing data reconstruction on the denoised temperature data to obtain a target temperature data sequence; Rearranging the target temperature data sequence in a matrix according to the number and sampling time of the temperature measurement points to obtain an initial temperature matrix, and establishing a three-dimensional spatial coordinate system based on each temperature measurement point of the initial temperature matrix to obtain a temperature measurement point spatial distribution model; The initial temperature matrix is spatially mapped according to the temperature measurement point spatial distribution model, the temperature gradient and heat flow distribution between adjacent temperature measurement points are calculated to obtain temperature field distribution data, and the temperature field distribution data is fused with the initial temperature matrix to obtain a temperature data matrix.
3. The injection mold temperature control method according to claim 2, characterized in that: The heating and cooling input matrices and the temperature response matrix are established according to the temperature data matrix, and the optimization equation is solved to obtain the heat transfer path coefficient matrix, including: Performing time series extraction on the heating and cooling source input power data in the temperature data matrix to obtain a heating and cooling input matrix, and performing time series extraction on the temperature data of the temperature measuring points in the temperature data matrix to obtain a temperature response matrix; Substituting the heating and cooling input matrix and the temperature response matrix into the heat transfer path equation to construct an optimization objective function to obtain an optimization equation, performing a singular value decomposition operation on the heating and cooling input matrix to obtain a matrix decomposition result; Solving the heat transfer path coefficient according to the matrix decomposition result to obtain the initial heat transfer path coefficient, and performing a perturbation analysis on the initial heat transfer path coefficient to obtain an error distribution characteristic; Based on the error distribution characteristics, an optimal regularization parameter is determined, the optimal regularization parameter is substituted into the optimization equation to obtain a corrected heat transfer path coefficient, and a matrix reconstruction is performed according to the corrected heat transfer path coefficient to obtain a heat transfer path coefficient matrix.
4. The injection mold temperature control method according to claim 3, characterized in that: The temperature data matrix is used as the state input to construct a deep deterministic policy gradient control network including a gated recurrent unit and a fully connected layer, and the cooling water flow rate and the inlet water temperature are linearly mapped to obtain a mixed action space, including: Performing time series analysis and data processing of temperature history changes on the temperature data matrix to obtain a temperature state sequence; A basic network framework is configured according to the temperature state sequence, wherein the execution network in the basic network framework includes three layers of gated recurrent units and two layers of fully connected layers, and parameters of the basic network framework are initialized, the number of neurons in the hidden layer is set to 128, and the activation function is a ReLU function; An evaluation network is constructed according to the basic network framework, wherein the evaluation network is branched into four fully connected layers to obtain a deep deterministic policy gradient control network; A cooling water flow mapping function is constructed according to the deep deterministic policy gradient control network, and the flow control range [0,20]L / min is mapped to the interval [-1,1] to obtain a standardized flow action space; and an inlet water temperature mapping function is constructed according to the deep deterministic policy gradient control network, and the temperature control range [15,95]℃ is mapped to the interval [-1,1] to obtain a standardized temperature action space; An initial mixed action space is constructed based on the standardized flow action space and the standardized temperature action space, and action output restriction conditions are set according to the initial mixed action space. The action range is constrained by the tanh function to obtain a mixed action space.
5. The injection mold temperature control method according to claim 4, characterized in that: The heat transfer path coefficient matrix and the mixed action space are input into a training environment, the deep deterministic policy gradient control network is iteratively trained based on an experience replay mechanism, and the network parameters are updated by calculating a reward function to obtain a temperature control model, including: An environment is constructed for the heat transfer path coefficient matrix and the mixed action space to obtain a training environment, and the update rates of the execution network and the evaluation network in the deep deterministic policy gradient control network are set to 0.001 to obtain a soft update parameter; Performing action prediction on the current temperature state data in the training environment, inputting it into the execution network of the deep deterministic policy gradient control network for forward propagation calculation to obtain the control action, and calculating the next state according to the control action and the heat transfer path coefficient matrix, constructing a reward function to calculate the reward value, and obtaining a state transfer quadruple; The state transition quadruple is stored in an experience replay pool, M samples are randomly sampled from the experience replay pool to construct a small batch of data, a training sample set is obtained, a target Q value is calculated for the training sample set, and a temporal difference algorithm is used to update the evaluation network parameters to obtain an evaluation network gradient; According to the evaluation network gradient, the execution network parameters are updated by the policy gradient algorithm, and the expected reward value is optimized by the gradient ascent method to obtain the execution network gradient; The evaluation network gradient and the execution network gradient are accumulated, and when the number of training rounds reaches a first target value and the average reward value change rate of N consecutive rounds is less than a second target value, a temperature control model is obtained.
6. The method for controlling the temperature of an injection mold according to claim 5, characterized in that: The current temperature state data in the training environment is predicted, input into the execution network of the deep deterministic policy gradient control network for forward propagation calculation, and the control action is obtained. The next state is calculated according to the control action and the heat transfer path coefficient matrix, and a reward function is constructed to calculate the reward value to obtain a state transfer quadruple, including: Sequencing the current temperature state data in the training environment, constructing a temperature time series feature sequence, and performing ReLU activation processing on the temperature time series feature sequence to obtain a temperature feature vector; Input the temperature feature vector into the three-layer gated loop unit of the execution network in the deep deterministic policy gradient control network for loop propagation calculation to obtain a deep feature representation, and map the deep feature representation to the two fully connected layers of the execution network in the deep deterministic policy gradient control network to obtain an action prediction vector; The action prediction vector is interval-mapped by a tanh activation function, corresponding to the cooling water flow rate range of [0,20] L / min and the inlet water temperature range of [15,95] °C, respectively, to obtain a control action; Performing a matrix multiplication operation according to the control action and the heat transfer path coefficient matrix, calculating the temperature response of each temperature measuring point, obtaining a next-state temperature distribution, calculating the temperature gradient of adjacent temperature measuring points in the next-state temperature distribution, and calculating the deviation from the target temperature to obtain a temperature deviation value; A reward function is constructed based on the temperature deviation value, and the weight coefficient α is set to 0.5 for weighted calculation to obtain a reward value. The current temperature state data, the control action, the reward value and the next state temperature distribution are combined into a state transition sample and converted into a tensor form to obtain a state transition quaternion.
7. The injection mold temperature control method according to claim 6, characterized in that: The temperature control model is deployed to the target control system, real-time temperature data of each temperature measurement point is collected, and a real-time control parameter set is calculated according to the temperature control model, including: Deploying the temperature control model to a target control system, and configuring a data acquisition module of the target control system to obtain an operating parameter set; According to the operating parameter set, temperature data of each temperature measuring point of the injection mold is collected in real time, and data is read by the data collection module to obtain real-time temperature data of each temperature measuring point; The real-time temperature data of each temperature measuring point is input into the temperature control model to construct a historical temperature sequence to obtain a temperature characteristic sequence, and the temperature characteristic sequence is input into the execution network of the temperature control model for forward calculation, and the cooling water flow rate and the inlet water temperature control value are output to obtain a real-time control instruction; The real-time control instruction is converted into a signal by a frequency converter, and the flow control instruction is converted into a water pump speed control signal to obtain a flow regulation instruction, and the flow regulation instruction is input into an electric regulating valve to dynamically adjust the inlet water temperature to obtain a temperature regulation instruction; The execution status of the flow regulation instruction and the temperature regulation instruction is monitored, the control parameters and temperature response data are recorded, the operation status data is obtained, and the operation status data is stored in the system database, the control process data is recorded and stored, and a real-time control parameter set is obtained.
8. An injection mold temperature control device, characterized in that: For implementing the steps of the method according to any one of claims 1 to 7, the device comprises: The acquisition module is used to collect temperature data from multiple temperature measurement points of the injection mold to obtain a temperature data matrix; A solution module, used for establishing a heating and cooling input matrix and a temperature response matrix according to the temperature data matrix, and solving the optimization equation to obtain a heat transfer path coefficient matrix; A mapping module is used to use the temperature data matrix as a state input, construct a deep deterministic policy gradient control network including a gated recurrent unit and a fully connected layer, and linearly map the cooling water flow rate and the inlet water temperature to obtain a mixed action space; A training module, used for inputting the heat transfer path coefficient matrix and the mixed action space into a training environment, iteratively training the deep deterministic policy gradient control network based on an experience replay mechanism, updating the network parameters by calculating a reward function, and obtaining a temperature control model; The deployment module is used to deploy the temperature control model to the target control system, collect real-time temperature data of each temperature measurement point, and calculate the real-time control parameter set according to the temperature control model.
9. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Cooling adjusting method and system for injection mold
CN120422432A
A cooling adjustment method and system for injection mold
CN120422432B
Temperature control method and system for injection mold
CN120439538A
A method and system for temperature control of an injection mold
CN120439538B
Temperature control optimization method and system for injection mold
CN120620598A