Linear motor temperature control method and system based on air cooling heat dissipation

By dividing the linear motor into multiple temperature control zones and constructing a distributed control network, and utilizing LSTM and deep Q-learning to optimize airflow regulation, the problems of uneven heat distribution and energy efficiency of the linear motor are solved, achieving precise temperature control and energy consumption optimization.

CN121077355APending Publication Date: 2025-12-05SHENZHEN FERGUS ELECTROMECHANICAL EQUIP CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511384319.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing linear motor air-cooling technology cannot effectively solve the problem of uneven heat distribution in different areas, leading to local overheating or energy waste, and lacks the ability to predict temperature change trends and optimize the overall energy efficiency of the system.

Method used

The linear motor is divided into multiple temperature control zones, each with an independent control execution unit. The temperature change trend is predicted by LSTM, and a distributed control network based on multi-agent reinforcement learning is constructed. The deep Q-learning algorithm is used to optimize the airflow regulation strategy and dynamically adjust the electric regulating valve and fan parameters to achieve temperature balance and minimize energy consumption.

Benefits of technology

It achieves precise dynamic control of linear motor temperature, reduces the risk of overheating, improves equipment lifespan and system efficiency, and reduces energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121077355A_ABST
    Figure CN121077355A_ABST
Patent Text Reader

Abstract

The invention provides a linear motor temperature control method and system based on air cooling heat dissipation, and relates to the technical field of motor temperature control, and the method comprises the steps: dividing a linear motor into a plurality of temperature control regions, and configuring an independent control execution unit for each region; predicting a temperature trend based on the LSTM; a multi-agent reinforcement learning control network is constructed, and each agent calculates an optimal control strategy based on deep Q learning; generating an overall airflow regulation strategy through a collaborative optimization algorithm; dynamically adjusting the opening of the electric adjusting valve and the rotating speed of the fan. Accurate temperature control is achieved, the heat dissipation efficiency is improved, and the service life of equipment is prolonged.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to motor temperature control technology, and in particular to a linear motor temperature control method and system based on air cooling heat dissipation. BACKGROUND

[0002] Linear motors have been widely used in high-precision equipment, rail transportation and advanced manufacturing due to their high precision, high response speed and no transmission chain loss. However, the existing linear motor air cooling heat dissipation technology has the following defects and deficiencies: The linear motor is considered as a whole, ignoring the problem of uneven heat distribution in different areas. In actual application, due to the difference of load distribution and motion state, the temperature of each part of the linear motor often appears significant difference, and the unified control strategy cannot realize accurate temperature management, which is easy to cause local overheating or energy waste.

[0003] Simple threshold control or PID control method is mostly used, which only makes reactive adjustment based on the current temperature state, and lacks the prediction ability of temperature change trend. This passive response control method is difficult to cope with the dynamic temperature change of linear motor under complex working conditions, often leading to control lag and large temperature fluctuation.

[0004] There is a lack of optimization consideration of system overall energy efficiency. Most systems are difficult to achieve a good balance between heat dissipation demand and energy consumption, either over-consumption of energy to ensure temperature safety, or sacrifice temperature control accuracy to save energy, lack of intelligent decision-making ability of dynamic adjustment according to actual working conditions, and difficult to meet the dual requirements of temperature uniformity and energy efficiency. SUMMARY

[0005] The linear motor temperature control method and system based on air cooling heat dissipation provided by the embodiments of the present application can solve the problems in the prior art.

[0006] In a first aspect of the embodiments of the present application, a linear motor temperature control method based on air cooling heat dissipation is provided, comprising: The linear motor is divided into multiple temperature control areas, and each temperature control area is configured with an independent control execution unit; based on the temperature data and historical operation data of each temperature control area, the temperature change trend in the future operation process of the linear motor is predicted through LSTM; An independent agent is assigned to each of the temperature control areas, a distributed control network based on multi-agent reinforcement learning is constructed, each agent receives the temperature change trend and the current temperature state as input, based on the deep Q learning algorithm, takes the temperature uniformity and the minimization of energy consumption as the reward function, calculates the optimal control strategy of the region; based on the optimal control strategy of each region and the state information shared between agents, balance the control requirements of each region through a collaborative optimization algorithm, and generate an overall air flow adjustment strategy considering the coupling effect between regions; According to the overall air flow regulation strategy, the opening degree of the electric regulating valve and the rotation speed of the supply and exhaust fan of each temperature control area are dynamically adjusted until the temperature change rate and energy consumption ratio reach the preset efficiency threshold and meet the safety constraint of the linear motor temperature state.

[0007] The step of predicting the temperature change trend of the linear motor in the future operation process based on the temperature data and historical operation data of each temperature control area through the LSTM includes: Combining the temperature data, historical operation data of the target temperature control sub-area, and the temperature data of the adjacent temperature control sub-area to construct an input feature vector; Based on the temperature distribution characteristics and temperature change characteristics of the target temperature control sub-area, first and second adaptive coefficients are obtained respectively; the bias parameters of the LSTM network forget gate and input gate corresponding to the target temperature control sub-area are adjusted according to the first and second adaptive coefficients respectively, and an adaptive LSTM prediction network is established; Based on the thermal conductivity coefficient and distance parameter between the target temperature control sub-area and the adjacent temperature control sub-area, a thermal coupling matrix is constructed; The input feature vector is input into the adaptive LSTM prediction network to obtain the initial prediction state of the target temperature control sub-area; the historical prediction state of the adjacent temperature control sub-area is weighted and combined according to the thermal coupling matrix and added to the initial prediction state to obtain the temperature prediction value of the target temperature control sub-area.

[0008] Each agent receives the temperature change trend and the current temperature state as input, and based on the deep Q learning algorithm, takes the temperature uniformity and energy consumption minimization as the reward function, and calculates the optimal control strategy of the region. The steps include: The current temperature state of the temperature control area, the temperature change trend and the temperature state of the adjacent area are constructed into an agent state space, and the parameter adjustment interval of the control execution unit of the temperature control area is constructed into an action space; a deep Q learning network is used to map the state space and the action space; Based on the current temperature state and the temperature change trend, temperature dynamic features are extracted using adaptive weights, and state importance is calculated based on the temperature dynamic features, and state sampling probability is dynamically adjusted; The temperature deviation of the current temperature state is constructed into a temperature uniformity index; based on the energy efficiency characteristic curve of the control execution unit of the temperature control area under different loads, the cumulative energy consumption of the control sequence is evaluated, and an energy utilization efficiency index is constructed; the temperature state of the adjacent temperature control area is calculated to obtain a heat flow interference coefficient; the temperature uniformity index, energy utilization efficiency index and heat flow interference coefficient are constructed into a reward function, and the weight coefficient of the reward function is adaptively adjusted; Train the deep Q-learning network based on the state sampling probability and the reward function, use an importance-based experience replay mechanism for sample learning, and obtain an optimal control strategy for the temperature control region.

[0009] The step of using an importance-based experience replay mechanism for sample learning comprises: Obtain the rate of change of the temperature change trend, the control response time of the temperature control region, and the environmental parameter fluctuation amplitude, and calculate a temperature failure risk assessment value; Based on the temperature failure risk assessment value, divide the danger level of the experience sample into a high-risk level, a warning level, and a regular level; set a sampling weight according to the danger level of the experience sample, and use a roulette method to sample according to the sampling weight; Extract the temperature change feature sequence and the control parameter sequence of the historical failure data of the temperature control region, use a time series interpolation method to generate an intermediate state sequence of the failure evolution process, and supplement the intermediate state sequence to the experience replay pool as a high-risk level sample; When the temperature failure risk assessment value of the current state is greater than a safety threshold, the deep Q-learning network limits the exploration action within the historical safe action range during exploration.

[0010] Based on the optimal control strategy of each region and the state information shared between agents, balance the control requirements of each region through a collaborative optimization algorithm, and generate an overall air flow regulation strategy considering the coupling effects between regions, comprising: The state information shared between agents includes temperature field gradient information and air flow field characteristic quantities; calculate the state information entropy based on the state information, and determine the dynamic update period of the state information according to the state information entropy; Based on the temperature field gradient information, calculate the heat conduction coefficient in combination with the contact area and characteristic distance of adjacent regions; based on the air flow field characteristic quantities, calculate the air flow interference strength of adjacent regions on the shared boundary area; and weight and sum the heat conduction coefficient and the air flow interference strength to obtain the comprehensive coupling strength; Calculate the strategy adjustment amount according to the comprehensive coupling strength; superimpose the strategy adjustment amount on the optimal control strategy of each region to obtain an optimized control strategy, and set temperature equalization constraints and air flow equalization constraints for the optimized control strategy; Determine the priority of each region according to the dynamic update period, the comprehensive coupling strength, and the control execution unit response time of the temperature control region; perform conflict resolution on the optimized control strategy according to the priority to obtain a resolved control strategy; calculate the contribution degree of each region in the overall control target, and use the contribution degree as a fusion weight; and combine the resolved control strategy by weighting according to the fusion weight to generate an overall air flow regulation strategy.

[0011] The step of conflict resolution on the optimized control strategy according to the priority comprises: A regional temperature control performance parameter is calculated based on the optimized control strategy, and a temperature control comprehensive score of the region is obtained by weighting the regional temperature control performance parameter; A remaining adjustment capacity of the control execution unit is obtained, an adjustable air supply amount range and an air supply time range of each region are determined based on the priority, and an air supply strategy initial adjustment scheme is generated within the air supply amount range and the air supply time range according to a temperature control comprehensive score difference of adjacent regions; For the region whose priority is higher than a reference value, when the temperature control comprehensive score of the region is lower than a target threshold value, the air supply amount range and the air supply time range of the region are gradually expanded by a preset step, and an air supply strategy supplementary scheme is generated within the expanded range; the temperature control comprehensive score is calculated after each generation of the supplementary scheme, and the adjustment is stopped when the score change value is less than a preset change threshold value for a plurality of consecutive times. The air supply strategy initial adjustment scheme and the air supply strategy supplementary scheme are combined to form a final air supply strategy scheme, a temperature control comprehensive score corresponding to the final air supply strategy scheme is calculated, and the scheme with the highest score is selected to perform conflict resolution.

[0012] The step of adaptively adjusting the weight coefficient of the reward function comprises: The historical temperature state and the current temperature state of the temperature control region are input into a learnable parameter matrix, the historical temperature state and the current temperature state are mapped and executed through a hyperbolic tangent transformation, a state importance evaluation value is obtained through exponential normalization processing, a weight sensitivity is calculated according to the state importance evaluation value and the partial derivative of each component of the reward function, and an attention weighted gradient is obtained by multiplying the weight sensitivity and the temperature control performance gradient; A temperature state distribution probability of the temperature control region is obtained, a state dispersion degree is calculated based on the temperature state distribution probability, and a temperature control effect evaluation value is weighted and adjusted based on the state dispersion degree to obtain an adaptive learning rate, the temperature control effect evaluation value being obtained based on an actual temperature and a target temperature of the temperature control region; The adaptive learning rate and the attention weighted gradient are multiplied to update the weight coefficient of the reward function, and the updated weight coefficient is dynamically threshold-constrained to obtain the final weight coefficient.

[0013] In a second aspect of the embodiment of the present application, a linear motor temperature control system based on air cooling heat dissipation is provided, comprising: A first unit is configured to divide the linear motor into a plurality of temperature control regions, and each temperature control region is configured with an independent control execution unit; temperature data and historical operation data of each temperature control region are used to predict a temperature change trend in a future operation process of the linear motor through an LSTM; a second unit for assigning an independent agent to each of the temperature control regions, constructing a distributed control network based on multi-agent reinforcement learning, each agent receiving the temperature change trend and the current temperature state as input, calculating the optimal control strategy of the region based on the deep Q learning algorithm, taking temperature uniformity and energy consumption minimization as the reward function, balancing the control requirements of each region through a collaborative optimization algorithm based on the optimal control strategy of each region and the shared state information between agents, and generating an overall air flow adjustment strategy considering the coupling effects between regions; a third unit for dynamically adjusting the opening degree of the electric regulating valve and the rotation speed of the inlet and exhaust fans of each temperature control region according to the overall air flow adjustment strategy until the temperature change rate and energy consumption ratio reach the preset efficiency threshold and meet the safety constraints of the linear motor temperature state.

[0014] In a third aspect, the embodiment of the present application provides an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to execute the method described above.

[0015] In a fourth aspect, the embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions are executed by a processor to implement the method described above.

[0016] The embodiment of the present application realizes accurate prediction of the temperature change trend of the linear motor through multi-temperature control region division and LSTM prediction technology, improves the prediction ability of the system for temperature abnormalities, reduces the risk of overheating, and prolongs the service life of the equipment.

[0017] The embodiment of the present application adopts a distributed control network based on multi-agent reinforcement learning, each region makes independent decisions and collaborates to optimize, greatly improves the precision and response speed of temperature control, reduces control delay, makes temperature control more accurate and dynamic, and reduces system energy consumption.

[0018] The embodiment of the present application realizes balanced control of the temperature of each region of the linear motor, avoids the generation of local hot spots, meets different working condition requirements through dynamic adjustment of air flow parameters, ensures safe operation of the equipment while maximizing energy saving, and improves the overall operation efficiency and reliability of the system. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 Fig. 1 is a flowchart of a linear motor temperature control method based on air cooling heat dissipation according to an embodiment of the present application; Figure 2 Fig. 2 is a flowchart of an overall air flow adjustment strategy generation based on multi-agent collaborative optimization according to an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work are within the protection scope of the present application.

[0021] The technical solutions of the present application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and some embodiments may not be described again for the same or similar concepts or processes.

[0022] Figure 1 The flowchart of the linear motor temperature control method based on air cooling of the embodiments of the present application is shown in FIG. 1, which comprises the following steps. Figure 1 The linear motor is divided into multiple temperature control areas, and each temperature control area is configured with an independent control execution unit. Based on the temperature data and historical operation data of each temperature control area, the temperature change trend in the future operation of the linear motor is predicted through LSTM. An independent agent is assigned to each temperature control area, and a distributed control network based on multi-agent reinforcement learning is constructed. Each agent receives the temperature change trend and the current temperature state as input, and calculates the optimal control strategy of the region based on the deep Q learning algorithm, with temperature uniformity and energy consumption minimization as the reward function. Based on the optimal control strategy of each region and the shared state information between agents, the collaborative optimization algorithm is used to balance the control requirements of each region, and the overall air flow adjustment strategy considering the coupling effect between regions is generated. According to the overall air flow adjustment strategy, the opening degree of the electric regulating valve and the rotating speed of the air inlet and exhaust fan of each temperature control area are dynamically adjusted until the temperature change rate and the energy consumption ratio reach the preset efficiency threshold and meet the safety constraints of the temperature state of the linear motor.

[0023] ​For example, the linear motor is divided into multiple temperature control regions, and specifically, the linear motor can be divided into three main temperature control regions, namely, a stator coil region, a permanent magnet region, and a guide rail region. For the coil region, four groups of adjustable air inlet vents and two groups of variable frequency exhaust fans are configured; for the permanent magnet region, two groups of air inlets and one group of exhaust fans are configured; and for the guide rail region, two groups of air inlets and shared exhaust are configured. The control execution unit of each region includes a temperature sensor, an electric regulating valve, and a fan speed controller. The temperature sensor is a PT100 type platinum resistance sensor with an accuracy of ±0.1°C and a sampling frequency of 10 Hz. The electric regulating valve is a proportional regulating type with an opening range of 0-100% and an adjustment accuracy of 1%. The fan is an EC direct current variable frequency fan with a speed range of 600-3000 rpm and an adjustment accuracy of 10 rpm.

[0024] The temperature prediction uses a long short-term memory network (LSTM) structure, and the input layer includes temperature time series, motor load, and environmental temperature features. The LSTM network is composed of an input layer, a hidden layer, and an output layer, and the hidden layer includes 64 LSTM units with a time step of 30, corresponding to 300 seconds of historical data. The network training uses historical operation data, including temperature change records under different load conditions, with a training sample size of 10,000 groups and a division ratio of 80% for the training set and 20% for the validation set. For example, when running at 75% rated load, the coil zone temperature rises from the initial 25°C to 46.5°C, and the LSTM predicts that the temperature will be 48.2°C after 30 minutes, with the actual measured value being 48.5°C, a prediction error of 0.3°C, and a relative error of 0.6%.

[0025] In the distributed control network, each temperature control region is configured with an independent agent. Taking the coil zone agent as an example, its state space includes current temperature, temperature change rate, predicted temperature, load state, and other features, and the action space is a combination of air inlet opening (0-100%) and exhaust fan speed (600-3000 rpm). The deep Q network (DQN) structure includes four layers: an input layer (8 neurons), two hidden layers (64 and 32 neurons respectively), and an output layer (corresponding to 25 discrete actions). The activation function uses ReLU, and the optimizer is Adam. The reward function is composed of a temperature balance index and an energy consumption index. The temperature balance index is calculated as 1 minus the ratio of the actual temperature difference to the target temperature and the allowed temperature deviation, with the target temperature set to 40°C and the allowed temperature deviation set to ±10°C. When the actual temperature is 46.5°C, the temperature balance index is 1-6.5 / 10=0.35. The energy consumption index is calculated based on the fan power and operating time, with the fan power being 85W at 2400rpm and the energy consumption being 0.085kWh per hour of operation. The two indices are combined into the total reward according to the weights, with the initial weights being [0.6, 0.4].

[0026] To improve the control accuracy, the experience replay technique is adopted, with a buffer size of 10000 and each training batch containing 64 samples. The training process uses an ε-greedy strategy for exploration, with an initial ε value of 1.0 gradually decaying to 0.1 at a decay rate of 0.995. The target network is updated every 100 steps. After 5000 rounds of training, the coil zone agent can control the temperature within ±2°C of the target temperature, with a control stability of over 95%.

[0027] The collaborative optimization algorithm generates a global control strategy for multiple zones based on shared state information. The shared state information includes the current temperature, predicted temperature, and control parameters of each zone, which are exchanged through the communication network between zones with a communication period of 1 second. Taking the collaborative optimization of three zones as an example, when the coil zone temperature is 46.5°C, the permanent magnet zone temperature is 38.2°C, and the guide rail zone temperature is 35.7°C, the optimal action calculated by the coil zone agent is an inlet opening degree of 75% and an exhaust fan speed of 2400 rpm. However, considering the coupling effect between zones, the collaborative optimization algorithm adjusts the fan speed of the coil zone to 2600 rpm, the fan speed of the permanent magnet zone from 1800 rpm to 2000 rpm, and the inlet opening degree of the guide rail zone from 50% to 60%, forming a global airflow regulation strategy. The coupling influence coefficient matrix is determined based on thermal flow field simulation and actual test data, for example, the coupling coefficient from the coil zone to the permanent magnet zone is 0.3, which means that for every 10°C increase in the coil zone temperature, the temperature of the permanent magnet zone is affected by 3°C.

[0028] The execution control process uses a segmented PID algorithm to achieve precise control of the electric regulating valve and fan speed. Taking the electric regulating valve as an example, when the opening degree command is 75%, the controller uses a proportional coefficient of 2.5, an integral coefficient of 0.5, and a derivative coefficient of 0.1 to make the valve smoothly transition from the current 50% opening to the target opening, with a transition time of about 2.5 seconds and an overshoot during the adjustment process controlled within 3%. The fan speed control uses a ramp function for acceleration and deceleration, with a maximum acceleration of 200 rpm / s and a deceleration of 150 rpm / s. The process of adjusting from 1800 rpm to 2600 rpm takes 4 seconds, during which the fan current increases within 30% of the rated current, ensuring safe operation of the fan.

[0029] The temperature change rate and energy consumption ratio are continuously monitored during the control process. The ratio is calculated by dividing the temperature drop rate (°C / min) by the energy consumption per unit time (kWh / min). For example, during the control execution process, the coil area temperature drops at a rate of 0.3°C / min, and the fan energy consumption is 0.0014 kWh / min. The temperature change rate and energy consumption ratio is 214.3°C / kWh. The preset efficiency threshold is 200°C / kWh. When the ratio is higher than the threshold and the temperatures of each area meet the safety constraints (coil area ≤ 65°C, permanent magnet area ≤ 60°C, and guide rail area ≤ 50°C), it is considered that the control strategy achieves the optimization goal.

[0030] In an alternative embodiment, based on the temperature data and historical operation data of each temperature control area, the step of predicting the temperature change trend of the linear motor in the future operation process by LSTM includes: Combining the temperature data, historical operation data of the target temperature control sub-area, and the temperature data of the adjacent temperature control sub-area to construct an input feature vector; Based on the temperature distribution characteristics and temperature change characteristics of the target temperature control sub-area, first and second adaptive coefficients are obtained respectively. The bias parameters of the forgetting gate and the input gate of the LSTM network corresponding to the target temperature control sub-area are adjusted according to the first and second adaptive coefficients, respectively, to establish an adaptive LSTM prediction network; Based on the thermal conductivity coefficient and distance parameter between the target temperature control sub-area and the adjacent temperature control sub-area, a thermal coupling matrix is constructed; The input feature vector is input into the adaptive LSTM prediction network to obtain the initial prediction state of the target temperature control sub-area. The historical prediction state of the adjacent temperature control sub-area is weighted and combined according to the thermal coupling matrix, and added to the initial prediction state to obtain the temperature prediction value of the target temperature control sub-area.

[0031] For example, the temperature data and historical operation data of each temperature control sub-area are obtained. The temperature control sub-area can be divided into multiple blocks, such as the coil area, the guide rail area, the magnetic guide area, etc. The temperature data of each area includes temperature values in time series, for example, 24 hours of temperature data recorded at 5-minute sampling intervals. The historical operation data includes motor current values, speed values, and running load parameters, which are usually stored in a database in time series.

[0032] For a specific target temperature control sub-region, such as a coil area, temperature data of the region, related historical operation data (such as current value) and temperature data of adjacent regions (such as guide rail areas) are selected to construct an input feature vector. Specifically, data at the current time t and the previous n times (such as the previous 5 times) can be combined to form a feature vector. For example, if the temperature of the target region coil area at time t is 45°C, the previous 5 times are 44°C, 42°C, 40°C, 39°C and 38°C respectively; at the same time, the current value is 10A, and the temperature of the adjacent guide rail area is 35°C, then the input feature vector constructed is [45, 44, 42, 40, 39, 38, 10, 35].

[0033] In constructing the adaptive LSTM prediction network, the temperature distribution characteristics of the target temperature control sub-region are analyzed. Taking the coil area as an example, the uniformity of the temperature distribution is evaluated by calculating the coefficient of variation (the ratio of the standard deviation to the mean) of its temperature data. If the coefficient of variation is large (such as greater than 0.15), it indicates that the temperature distribution is uneven, and the first adaptive coefficient can be set to 0.7; if the coefficient of variation is small (such as less than 0.05), it indicates that the temperature distribution is uniform, and the first adaptive coefficient can be set to 0.3. The temperature change characteristics of the target temperature control sub-region are analyzed. Taking the coil area as an example, the stability of temperature change is evaluated by calculating the average value and standard deviation of the temperature change rate (such as 0.5°C per minute) at consecutive times. If the standard deviation of the temperature change rate is large (such as greater than 0.1°C per minute), it indicates that the temperature change is unstable, and the second adaptive coefficient can be set to 0.6; if the standard deviation is small (such as less than 0.05°C per minute), it indicates that the temperature change is stable, and the second adaptive coefficient can be set to 0.2.

[0034] Based on the two coefficients obtained, the bias parameters of the forget gate and the input gate of the prediction adjustment LSTM network are adjusted. For example, for the LSTM forget gate, the original bias parameter is 1.0, and according to the first adaptive coefficient 0.7, the new bias parameter is adjusted to 1.0+0.7=1.7, so that the LSTM network is more inclined to retain historical information; for the input gate, the original bias parameter is 0.5, and according to the second adaptive coefficient 0.6, the new bias parameter is adjusted to 0.5-0.6=-0.1, so that the LSTM network is more cautious in receiving new input information.

[0035] To consider the thermal conduction effect between temperature-controlled sub-regions, a thermal coupling matrix is constructed. Taking three temperature-controlled sub-regions (coil region, rail region, and magnetic conductor region) as an example, the thermal conduction coefficients are calculated according to the material thermal conductivity, the distance between regions, and the contact area. Obtain the thermal conductivity parameters of the materials in each region (such as copper coil 385 W / (m·K)), measure the physical distance and contact area between regions. Apply Fourier's law of heat conduction to calculate the thermal conduction coefficient, that is, thermal conduction coefficient = material thermal conductivity × contact area / distance. Assume that the thermal conduction coefficient between the coil region and the rail region is 0.4, and between the coil region and the magnetic conductor region is 0.2; the thermal conduction coefficient between the rail region and the magnetic conductor region is 0.3. To eliminate the scale effect, the thermal conduction coefficient is normalized. The final thermal coupling matrix is represented as a 3 × 3 matrix, where the diagonal elements are 1 (representing self-influence), and the non-diagonal elements represent the normalized thermal conduction coefficient between regions. This matrix can be calibrated by temperature change experimental data to ensure that it accurately reflects the actual thermal conduction relationship between regions.

[0036] After obtaining the input feature vector, it is input into the adaptive LSTM prediction network for forward calculation. Assume that the initial predicted temperature of the coil region after passing through the adaptive LSTM network is 48°C. Then, the historical prediction states of adjacent regions are combined by weighting using the thermal coupling matrix. Assume that the historical prediction temperatures of the rail region and the magnetic conductor region are 37°C and 30°C respectively, according to the thermal coupling matrix, their influences on the coil region are 0.4 × 37 = 14.8 and 0.2 × 30 = 6 respectively. Add these influences to the initial prediction state 48°C and normalize to get the final temperature prediction value of the coil region as (48 + 14.8 + 6) / (1 + 0.4 + 0.2) = 43.1°C. Continuous prediction can also be achieved by sliding window method. For example, after predicting the temperature at t+1 time, the prediction value is added to the feature vector, and the earliest time data is removed to construct a new input feature vector to predict the temperature at t+2 time.

[0037] Collect a sequence of temperature prediction values at multiple future time points (such as every 5 minutes within the next 30 minutes), and construct a rate of change sequence by calculating the temperature change rate (such as °C / min) between adjacent time points. Then, apply polynomial fitting or wavelet decomposition methods to smooth the rate of change sequence, eliminating the influence of short-term fluctuations. Based on the processed rate of change sequence, identify key feature points (such as extreme points, inflection points) and extract trend features, including temperature rise rate, stable time, peak temperature, and temperature fluctuation range. According to these features, classify the temperature change trend into "rapid temperature rise", "slow temperature rise", "temperature stable", "slow temperature drop", or "rapid temperature drop" types, and calculate the trend duration and change amplitude to form a complete temperature change trend description, providing decision basis for subsequent control strategy formulation.

[0038] The adaptive LSTM network of the application dynamically adjusts the gating parameters according to temperature distribution characteristics and change characteristics, and enhances the adaptability to different temperature states. The thermal coupling matrix effectively models the heat conduction relationship between the temperature control sub-regions, solving the defect that the traditional prediction method ignores the inter-regional thermal interaction. The method organically combines deep learning with thermodynamic principles, not only improves the prediction accuracy, but also adapts to the temperature dynamic changes under complex working conditions of the motor.

[0039] In an optional embodiment, each agent receives the temperature change trend and the current temperature state as input, and calculates the optimal control strategy of the region based on a deep Q learning algorithm, with temperature uniformity and energy consumption minimization as the reward function, the steps comprising: The current temperature state of the temperature control region, the temperature change trend and the temperature state of the adjacent region are constructed into an agent state space, and the parameter adjustment interval of the control execution unit of the temperature control region is constructed into an action space; a deep Q learning network is used to map the state space and the action space; Based on the current temperature state and the temperature change trend, temperature dynamic features are extracted using adaptive weights, and state importance is calculated based on the temperature dynamic features, and the state sampling probability is dynamically adjusted; The temperature deviation of the current temperature state is constructed into a temperature uniformity index; based on the energy efficiency characteristic curve of the control execution unit of the temperature control region under different loads, the cumulative energy consumption of the control sequence is evaluated, and an energy utilization efficiency index is constructed; the temperature state of the adjacent temperature control region is calculated to obtain a heat flow interference coefficient; the temperature uniformity index, the energy utilization efficiency index and the heat flow interference coefficient are constructed into a reward function, and the weight coefficient of the reward function is adaptively adjusted; The deep Q learning network is trained based on the state sampling probability and the reward function, and sample learning is performed using an importance-based experience replay mechanism to obtain the optimal control strategy of the temperature control region.

[0040] For example, the current temperature state of a temperature control area can be represented as a set of temperature values of multiple temperature measurement points in the area, such as 45°C, 46°C, 44°C, 47°C, and 45°C of five temperature measurement points in a certain coil area. The temperature change trend includes predicted temperature change direction, change rate, and stabilization time, etc. For example, it is predicted that the area will show a "rapid heating" trend in the next 30 minutes, with a heating rate of 0.5°C / min, and it is expected to reach a peak of 55°C after 20 minutes. The temperature state of the adjacent area is also represented as a set of temperature values, such as 38°C, 39°C, and 37°C of the temperature of the adjacent rail area. These state information combinations constitute the state space of the agent. The parameter adjustment interval of the control execution unit constitutes the action space, such as the opening range of the electric regulating valve, which is 0-100%, which can be divided into 11 discrete values (0%, 10%, 20%,..., 100%); the speed range of the supply and exhaust fan, which is 600-1800 rpm, which can be divided into 7 discrete values (600 rpm, 800 rpm,..., 1800 rpm). The agent can select combinations from these discrete actions to form a complete action space.

[0041] The deep Q-learning network is used to realize the mapping from the state space to the action space. The network adopts a multi-layer neural network structure, the number of nodes in the input layer is consistent with the dimension of the state vector, such as a 15-dimensional vector composed of temperature state, temperature trend, and adjacent area state; the hidden layer adopts a two-layer structure, each layer has 128 neurons, and uses ReLU activation function; the number of nodes in the output layer is consistent with the size of the action space, such as 77 actions formed by the combination of valve opening and fan speed. The network parameter initialization adopts the Xavier method to ensure that the gradient will not disappear or explode in the propagation process.

[0042] The temperature dynamic feature extraction process needs to consider the correlation between temperature state and change trend. Real-time temperature data of the temperature control area is obtained, for example, the temperatures of 5 temperature measurement points in a certain coil area are 45℃, 46℃, 44℃, 47℃, and 45℃ respectively, and the average temperature is calculated to be 45.4℃. The target temperature is set to 40℃, so the temperature deviation is calculated as 5.4℃. According to the temperature change trend prediction, the future temperature of this area shows a "rapid warming" trend, with a warming rate of 0.5℃ / min, which means that the temperature deviation will further increase without taking control measures. Combined with the temperature deviation and the warming rate, the temperature deviation trend index is calculated, which is calculated by multiplying the temperature deviation and the warming rate by a time coefficient of 1.5, and the temperature deviation trend index value is 5.4×0.5×1.5=4.05. The uniformity feature of temperature distribution is also extracted, and the range of multi-point temperature is calculated to be 47-44=3℃, the standard deviation is 1.1℃, and the coefficient of variation is 1.1÷45.4=0.024, which form the uniformity feature vector. Based on the analysis of historical control data, the stability of the temperature change rate is analyzed, and the fluctuation range of the temperature change rate in the past 30 minutes is calculated to be 0.3℃ / min to 0.7℃ / min, and the fluctuation amplitude is 0.4℃ / min, which is taken as the stability feature value. For the extracted features, the initial weights are assigned: temperature deviation weight 0.4, temperature deviation trend index weight 0.3, uniformity feature weight 0.2, and stability feature weight 0.1. These weights are automatically adjusted based on the temperature control effect during operation. The adjustment method is: calculate the temperature control error after each control, such as the temperature after control is 41.2℃, the error is 1.2℃; calculate the improvement degree of temperature uniformity, such as the standard deviation after control is reduced to 0.6℃, which is improved by 0.5℃; if the temperature control error increases, increase the temperature deviation weight and the temperature deviation trend index weight, such as increase to 0.45 and 0.35; if the uniformity is worse, increase the uniformity feature weight, such as increase to 0.3. The weight adjustment adopts a gradual method, and the maximum adjustment amplitude is not more than 0.05 each time, ensuring stability. Based on the weighted temperature dynamic features, the state importance score is calculated by multiplying each feature value by the corresponding weight and summing, and then mapping it to the range of 0-1 through the sigmoid function. For example, the state importance score of high temperature deviation (5.4℃), rapid change (0.5℃ / min), and non-uniformity (standard deviation 1.1℃) is 0.85; while the state importance score of low temperature deviation (1.2℃), slow change (0.1℃ / min), and uniformity (standard deviation 0.3℃) is 0.25. Divide the importance score by the average of all state importance scores to get the relative importance, and then set the sampling probability to the relative importance value. For example, the relative importance of importance score 0.85 is 1.7, and the sampling probability is 0.25; the relative importance of importance score 0.25 is 0.5, and the sampling probability is 0.07.

[0043] The temperature deviation of each temperature measurement point in the temperature control area from the set temperature (e.g., 40°C) is calculated, and the deviation values are 5°C, 6°C, 4°C, 7°C, and 5°C, respectively. The average value of these deviations, 5.4°C, is taken as the area temperature deviation index; at the same time, the standard deviation, 1.1°C, is taken as the temperature uniformity index. The temperature uniformity index takes into account both the temperature deviation and the temperature uniformity, and the normalized scores of these two factors are calculated respectively, and then combined with a weight to form the final index value.

[0044] The energy utilization efficiency index evaluates the energy consumption of the control execution unit. The energy consumption data of the electric regulating valve at different opening degrees is obtained, such as 15W at 30% opening and 25W at 70% opening; the energy consumption data of the inlet and exhaust fans at different speeds is obtained, such as 50W at 800rpm and 120W at 1500rpm. For a control sequence, such as the valve opening from 30% to 70% and the fan speed from 800rpm to 1500rpm for 15 minutes, the cumulative energy consumption is calculated as (25W+120W)×15 minutes=2175W·minutes. According to the control effect (such as a temperature deviation reduction of 4°C), the unit effect energy consumption is calculated as 2175÷4=543.75W·minutes / °C. Based on historical operation data, the benchmark value of unit effect energy consumption is determined, and the optimal energy consumption value is set as 300W·minutes / °C and the worst energy consumption value is set as 1200W·minutes / °C. The energy utilization efficiency index converts the unit effect energy consumption into an index value in the range of 0-1 through a nonlinear mapping function, and the calculation method is 1-((unit effect energy consumption-optimal energy consumption value) / (worst energy consumption value-optimal energy consumption value)) 2 . For the unit effect energy consumption of 543.75W·minutes / °C, the energy utilization efficiency index is calculated as 1-((543.75-300) / (1200-300)) 2 =1-(243.75 / 900) 2 =1-0.0735=0.9265, rounded to 0.93. The optimal and worst energy consumption values are updated every 100 hours. The energy utilization efficiency index is inversely proportional to the unit effect energy consumption, and the lower the energy consumption, the higher the index value.

[0045] The heat flow interference coefficient reflects the influence degree of the temperature of the adjacent area (such as the guide rail area) on the target area (the coil area). The temperature difference between the adjacent area (such as the guide rail area) and the target area (the coil area) is 7°C. Based on the heat conduction coefficient between the two areas of 0.4W / (m 2 ·℃), the heat flow interference coefficient is calculated as 7×0.4=2.8W / m 2The coefficient represents the heat transferred per square meter of surface area per second from the adjacent zone to the target zone, affecting the temperature control effect of the target zone. To facilitate the use in the reward function, the heat flow disturbance coefficient is normalized to convert it to a value in the range of 0-1. The normalization process first determines the maximum heat flow disturbance coefficient observed in the historical operation as 8 W / m2 2 The current heat flow disturbance coefficient is divided by the maximum value to obtain the relative disturbance intensity 2.8 ÷ 8 = 0.35. The disturbance direction is also considered. When the temperature of the adjacent zone is lower than that of the target zone, the disturbance helps to lower the temperature, and the normalized coefficient is taken as a negative value; when the temperature of the adjacent zone is higher than that of the target zone, the disturbance intensifies the temperature rise, and the normalized coefficient is taken as a positive value. In this example, the temperature of the adjacent zone is 38℃ lower than that of the target zone, but the target is to lower the temperature of the target zone to 40℃, so the disturbance is conducive to the control target, and the normalized coefficient is taken as -0.35. The maximum heat flow disturbance coefficient value is updated every hour to ensure that the normalization process adapts to environmental changes.

[0046] The reward function integrates the temperature uniformity index, the energy utilization efficiency index, and the heat flow disturbance coefficient. Let the temperature uniformity index value be 0.75 (full score is 1, the closer to 1, the more uniform the temperature), the energy utilization efficiency index value be 0.65 (full score is 1, the closer to 1, the higher the energy efficiency), and the normalized heat flow disturbance coefficient be 0.35 (range 0-1, the closer to 0, the smaller the disturbance). The initial weights of the three indexes are set as w1 = 0.5, w2 = 0.3, and w3 = 0.2. The reward value is calculated as 0.5 × 0.75 + 0.3 × 0.65 - 0.2 × 0.35 = 0.545. As the control process proceeds, the control effect is monitored and the weights are adjusted. For example, when the temperature control effect is poor but the energy consumption is low, the weight of the temperature uniformity index is increased to 0.7 and the weight of the energy utilization efficiency index is reduced to 0.2; when there is a large external disturbance, the weight of the heat flow disturbance coefficient is increased to 0.3 and the weights of the other indexes are correspondingly reduced.

[0047] The deep Q-learning network is trained using an importance-based experience replay mechanism. An experience replay pool with a capacity of 10000 is maintained to store state transition samples (s, a, r, s'), including the current state s, the executed action a, the reward r, and the next state s'. Each training extracts 256 samples from the replay pool to form a batch. According to the previously calculated state importance, each sample is assigned a sampling weight, and samples with high importance are more likely to be selected. For example, a sample with a state importance of 0.85 is assigned a sampling weight of 1.7, while a sample with an importance of 0.25 has a sampling weight of 0.5. The network is trained iteratively: the target Q value is calculated, which is the current reward plus the maximum Q value of the next state multiplied by the discount factor 0.95; the mean square error between the predicted Q value and the target Q value is calculated as the loss function; the Adam optimizer is used for gradient descent, with a learning rate of 0.001, and the learning rate is multiplied by 0.99 every 100 training steps to decay; the training process is repeated until the loss function converges or the preset 10000 training steps are reached. After training, the agent selects the optimal action in each state according to the learned Q value, forming the optimal control strategy for the temperature control area.

[0048] The present application deeply integrates deep reinforcement learning with temperature control professional knowledge, achieving intelligent control oriented to temperature uniformity and energy consumption minimization. By adaptively extracting temperature dynamic features and dynamically adjusting state sampling probability, the learning ability of the model for key states is enhanced. The multi-dimensional reward function design takes into account both temperature control effect and energy efficiency and adjacent area influence, achieving balanced optimization of control objectives and effectively improving the temperature control performance and reliability of the linear motor.

[0049] In an alternative embodiment, the step of learning samples using an importance-based experience replay mechanism includes: Obtain the rate of change of the temperature trend, the control response time of the temperature control area, and the fluctuation amplitude of the environmental parameters, and calculate the temperature failure risk assessment value; Based on the temperature failure risk assessment value, the danger level of the experience sample is divided into high-risk level, warning level and normal level; according to the danger level of the experience sample, the sampling weight is set, and the sample is extracted according to the sampling weight using the roulette method; Extract the temperature change feature sequence and control parameter sequence of the historical failure data of the temperature control area, and generate the intermediate state sequence of the failure evolution process using the time series interpolation method, and supplement the intermediate state sequence to the experience replay pool as high-risk level samples; When the deep Q-learning network explores, if the temperature failure risk assessment value of the current state is greater than the safety threshold, the exploration action is limited within the range of historical safe actions.

[0050] For example, when training the deep Q-learning network based on the state sampling probability and the reward function, the importance-based experience replay mechanism is used for sample learning, which can significantly improve the training efficiency and model performance. The rate of change of the temperature trend data is obtained, for example, the temperature change rate of a certain coil area increases from 0.2°C / min to 0.5°C / min in the past 30 minutes, and the change value of the rate of change is 0.3°C / min / hour. At the same time, the control response time of the temperature control area is measured, for example, after adjusting the valve from 30% opening to 70% opening, the delay time for the temperature to start responding is 45 seconds, and the time for the temperature to reach a stable state is 180 seconds. The fluctuation amplitude of the environmental parameters is obtained by monitoring, for example, the environmental temperature fluctuates in the range of 22°C to 28°C within 24 hours, and the fluctuation amplitude is 6°C; the environmental humidity fluctuation range is 45% to 65%, and the fluctuation amplitude is 20%. The normalized value is obtained by dividing the change value of the temperature change rate by the standard change rate of 0.1°C / min / hour, which is 3.0; the normalized value is calculated by dividing the control response time by the standard response time of 120 seconds, which is 1.5; the normalized value is obtained by dividing the environmental temperature fluctuation amplitude by the standard fluctuation amplitude of 4°C, which is 1.5. The three normalized values are weighted and summed according to the weights of 0.5, 0.3 and 0.2, and the temperature failure risk assessment value is 3.0*0.5+1.5*0.3+1.5*0.2=2.25.

[0051] The risk assessment value greater than 2.0 is high risk level, the risk assessment value between 1.0 and 2.0 is alert level, and the risk assessment value less than 1.0 is normal level. In this example, the risk assessment value 2.25 is classified as high risk level. The base weight of the high risk level sample is 5.0, the base weight of the alert level sample is 2.0, and the base weight of the normal level sample is 1.0. Further, the weight is adjusted according to the timeliness of the sample. For newly generated samples, the weight is multiplied by the timeliness coefficient 1.2; for samples that have existed for more than 1000 training steps, the weight is multiplied by the timeliness coefficient 0.9. For example, a newly generated high risk level sample has a final sampling weight of 5.0*1.2=6.0; an existing alert level sample has a final sampling weight of 2.0*0.9=1.8. The roulette method is used to sample the samples according to the sampling weight, and the specific steps are as follows: calculate the sum of the weights of all samples, for example, the weight sum of 10000 samples is 18000; generate a random number between 0 and 18000, for example, 12500; add the sample weights in order, and when the cumulative value first exceeds the random number, the current sample is selected. This method ensures that samples with high sampling weights have a higher probability of being selected, but low weight samples still have a chance of being selected.

[0052] In a certain temperature failure event, the temperature sequence 10 minutes before failure is recorded as [42℃, 43℃, 45℃, 48℃, 52℃] and the corresponding control parameter sequence is [valve opening 30%, fan speed 800 rpm], [valve opening 40%, fan speed 1000 rpm], [valve opening 50%, fan speed 1200 rpm], [valve opening 60%, fan speed 1400 rpm], [valve opening 70%, fan speed 1600 rpm]. The time series interpolation method is used to generate the intermediate state sequence of the failure evolution process. The linear interpolation method inserts 4 equally spaced intermediate states between each two adjacent time points, for example, [42.2℃, 42.4℃, 42.6℃, 42.8℃] is inserted between 42℃ and 43℃, and [valve opening 32%, fan speed 850 rpm], [valve opening 34%, fan speed 900 rpm], [valve opening 36%, fan speed 950 rpm], [valve opening 38%, fan speed 975 rpm] are inserted between the control parameters [valve opening 30%, fan speed 800 rpm] and [valve opening 40%, fan speed 1000 rpm]. For parts with obvious nonlinear changes, the cubic spline interpolation method is used to generate intermediate states that are more consistent with the physical change law. The generated intermediate state sequence is supplemented to the experience replay pool as a high-risk level sample, which expands the number of high-risk samples and improves the model's learning ability for failure conditions.

[0053] When the deep Q-learning network explores, it needs to balance the relationship between exploring new actions and using known good actions. The ε-greedy strategy is used to explore with a probability of ε to select a random action and with a probability of 1-ε to select the action with the maximum current Q value. When the temperature failure risk assessment value of the current state is greater than the safety threshold (for example, the safety threshold is set to 1.8), the exploration action is limited within the range of historical safe actions to avoid selecting dangerous actions that lead to failure. The range of historical safe actions is determined by analyzing past successful control records. For example, analysis shows that under high temperature conditions (> 45℃), the valve opening degree should not be less than 40% and the fan speed should not be less than 1000 rpm; when the temperature rises rapidly (> 0.3℃ / min), the valve opening degree increment should not exceed 20% and the fan speed increment should not exceed 400 rpm. When in a high-risk state (risk assessment value 2.25> safety threshold 1.8), the exploration action will be limited within these safety ranges. The specific implementation method is: after generating a random action, check whether the action is within the safety range; if not, project the action onto the nearest safe action. For example, the randomly generated action is [valve opening degree 20%, fan speed 900 rpm], since the current temperature is 47℃> 45℃, the valve opening degree should not be less than 40%, so the action is corrected to [valve opening degree 40%, fan speed 900 rpm]. This safety exploration mechanism ensures safety in high-risk states while retaining some exploration ability.

[0054] The present application realizes the focus on high-risk states in the sample learning process, improving the prevention ability of potential failure conditions. By generating intermediate state sequences from historical failure data and supplementing them to the experience replay pool, the training data is enriched, especially the scarce failure scenario data.

[0055] In an optional implementation, based on the optimal control strategies of each region and the state information shared among agents, the step of balancing the control requirements of each region through a collaborative optimization algorithm to generate an overall air flow adjustment strategy considering the coupling effects between regions includes: The state information shared among agents includes temperature field gradient information and air flow field characteristic quantities; the state information entropy is calculated based on the state information, and the dynamic update period of the state information is determined according to the state information entropy; Based on the temperature field gradient information, the heat conduction coefficient is calculated in combination with the contact area and characteristic distance of adjacent regions; based on the air flow field characteristic quantities, the air flow interference strength of adjacent regions on the shared boundary area is calculated; the heat conduction coefficient and the air flow interference strength are weighted and summed to obtain the comprehensive coupling strength; The strategy adjustment amount is calculated according to the comprehensive coupling strength; the optimized control strategy is obtained by superimposing the strategy adjustment amount on the optimal control strategy of each region, and the temperature equalization constraint and the air flow equalization constraint are set for the optimized control strategy. According to the dynamic update period, the comprehensive coupling strength and the control execution unit response time of the temperature control area, the priority of each area is determined; according to the priority, the optimized control strategy is conflict resolved to obtain a resolved control strategy; the contribution of each area in the overall control target is calculated, and the contribution is taken as a fusion weight; according to the fusion weight, the resolved control strategy is weighted combined to generate an overall air flow regulation strategy.

[0056] In combination Figure 2 The overall air flow regulation strategy generation flowchart based on multi-agent collaborative optimization is described. For example, the temperature field gradient information shared between agents is collected by a temperature sensor array distributed in each temperature control area of the linear motor. For example, 12 temperature sensors are arranged in the coil area of the linear motor to form a 3x4 sensor array. Each sensor collects the surrounding temperature value, and the temperature field gradient vector is calculated according to the discrete temperature values. The temperature gradient vector is represented as the temperature change per meter, such as 3.2℃ / m in the X direction, 2.1℃ / m in the Y direction, and 4.5℃ / m in the Z direction. The air flow field characteristic quantity is collected by a wind speed sensor and a pressure sensor, including the air flow velocity vector, the pressure distribution and the turbulence intensity. For example, at the shared boundary between the coil area and the magnetic track area, the average air flow velocity is 1.2m / s, the direction is from the coil area to the magnetic track area, the pressure difference at the boundary is 22Pa, and the turbulence intensity is 0.28. The temperature field gradient value and the air flow field characteristic quantity are discretized into several intervals, the occurrence frequency of each interval is counted, and the information entropy value is calculated. For example, the state information entropy calculated at a certain moment is 3.15 bits. According to the state information entropy, the dynamic update period of the state information is determined. The higher the state information entropy value, the more intense the state change, and the state information needs to be updated more frequently. The update period calculation method is as follows: the basic update period is 60 seconds, when the information entropy is lower than 1.8 bits, the update period is extended to 120 seconds; when the information entropy is higher than 2.8 bits, the update period is shortened to 30 seconds. In this example, the state information entropy is 3.15 bits, and the corresponding update period is 30 seconds.

[0057] With the coil area and the magnetic track area of the linear motor as an example, the contact area of the two areas is 2.2 square meters, and the characteristic distance is 0.35 meters. The projection value of the temperature field gradient on the contact surface is 3.5℃ / m. The calculation of the heat transfer coefficient considers the thermal conductivity of the contact material, and the ratio of the contact area to the characteristic distance, multiplied by the projection value of the temperature gradient. For example, the thermal conductivity of the contact material is 0.22 W / (m·℃), and the heat transfer coefficient is calculated as 0.22×(2.2 / 0.35)×3.5=4.82 W / ℃. The air flow interference intensity considers three factors: air flow speed, pressure difference, and turbulence intensity. The air flow speed contribution is the speed value multiplied by the flow direction coefficient (positive for the control direction, negative for the reverse control direction), such as 1.2 m / s×(-1)=-1.2; the pressure difference contribution is the pressure difference value divided by the reference pressure value 100 Pa, such as 22 / 100=0.22; the turbulence intensity contribution is the turbulence intensity value multiplied by the coefficient 2, such as 0.28×2=0.56. The weighted sum of the three (weights are 0.5, 0.3, and 0.2, respectively) gives the air flow interference intensity as -1.2×0.5+0.22×0.3+0.56×0.2=-0.472. The heat transfer coefficient and the air flow interference intensity are weighted and summed to obtain the comprehensive coupling intensity, and the weights are dynamically adjusted according to the current control target. For example, when the control target is mainly temperature equalization, the weight of the heat transfer coefficient is 0.7, and the weight of the air flow interference intensity is 0.3; when the control target is mainly air flow regulation, the weights are 0.3 and 0.7, respectively. In the case of temperature equalization control of the linear motor, the comprehensive coupling intensity is 4.82×0.7+(-0.472)×0.3=3.23.

[0058] The policy adjustment amount is proportional to the comprehensive coupling intensity, and the adjustment gain coefficient is also considered. For example, the adjustment gain coefficient of the fan control is 180 rpm / (unit coupling intensity), and the adjustment gain coefficient of the valve control is 12% / (unit coupling intensity). For a comprehensive coupling intensity of 3.23, the adjustment amount of the fan speed is 180×3.23=581 rpm, and the adjustment amount of the valve opening is 12%×3.23=38.8%. The optimized control strategy is obtained by superimposing the policy adjustment amount on the optimal control strategy of each area. For example, the optimal control strategy of the coil area is fan speed 1500 rpm and valve opening 40%, and after superimposing the adjustment amount, the optimized control strategy is fan speed 2081 rpm and valve opening 78.8%. The optimized control strategy is also set with temperature equalization constraints and air flow balance constraints. The temperature equalization constraint requires that the temperature deviation of each area of the linear motor does not exceed the set threshold, such as ±5℃; the air flow balance constraint requires that the air intake and exhaust of the entire linear motor are balanced, with a deviation of no more than 7%. If the optimized control strategy violates the constraint condition, the strategy will be modified according to the minimum adjustment principle. For example, if the optimized valve opening causes the air intake to exceed the air exhaust by 9%, the valve opening will be adjusted from 78.8% to 74.5% to reduce the air flow deviation to within the constraint range.

[0059] The priority calculation considers three factors: the shorter the update period, the higher the priority, with a contribution coefficient of 0.4; the greater the comprehensive coupling strength, the higher the priority, with a contribution coefficient of 0.3; the shorter the response time, the higher the priority, with a contribution coefficient of 0.3. For the linear motor coil area, the update period is 30 seconds, and the normalized value is 0.9; the comprehensive coupling strength is 3.23, and the normalized value is 0.85; the response time is 20 seconds, and the normalized value is 0.85. The priority score of the coil area is 0.9x0.4+0.85x0.3+0.85x0.3=0.87. Similarly, calculate the priority scores of other areas, such as the magnetic track area, which is 0.72, and the bearing area, which is 0.65. When the control strategies of two areas conflict, the requirements of the area with higher priority are met first, while the impact on the area with lower priority is minimized. For example, the linear motor coil area and the magnetic track area share an air inlet, and the coil area requires an increase in air volume, while the magnetic track area requires a decrease in air volume. Since the coil area has a higher priority (0.87>0.72), the coil area's requirements will be met, but the magnetic track area's requirements will be compensated for on other independent control units, such as increasing the speed of the magnetic track area's dedicated exhaust fan. The resolved control strategy ensures that the core control requirements of each area are met, while minimizing the negative impact of conflicts. The contribution calculation is based on the improvement in temperature deviation, energy efficiency, and the impact on adjacent areas. For example, the coil area's control strategy is expected to reduce the temperature deviation from 7°C to 1.8°C, an improvement of 5.2°C, with a normalized value of 0.82; energy efficiency is improved by 22%, with a normalized value of 0.78; and the positive impact on adjacent areas is scored at 0.65. The overall contribution is 0.82x0.5+0.78x0.3+0.65x0.2=0.777. Normalize the contribution of each area, such as the coil area's contribution of 0.777, the magnetic track area's contribution of 0.683, and the bearing area's contribution of 0.592, which are normalized to 0.38, 0.33, and 0.29, respectively. For shared control execution units, such as the main air inlet unit, the adjustment parameters are weighted by the contribution of each area, with the formula: main air volume = coil area suggestion value x 0.38 + magnetic track area suggestion value x 0.33 + bearing area suggestion value x 0.29. For example, the coil area suggests a main air volume of 120 m 3 / h, the magnetic track area suggests 100 m 3 / h, and the bearing area suggests 90 m 3 / h, the weighted main air volume is 120x0.38+100x0.33+90x0.29=104.8 m 3 / h. For area-specific control execution units, such as the coil area's local exhaust fan, the area's control suggestion value is retained, but is fine-tuned according to the overall air flow balance requirements. The adjustment ratio is based on the impact of the area on the overall air flow balance, calculated as the ratio of the area's air intake to the total air intake.

[0060] Taking the actual operation scene of a linear motor as an example, it is assumed that the temperature rise rate of the coil area is detected to accelerate, and it is predicted that the temperature will exceed the threshold value after 30 minutes. The coil area agent calculates the optimal control strategy as increasing the air intake by 25% and increasing the exhaust fan speed by 300 rpm. However, the adjacent magnetic track area needs to reduce the temperature fluctuation at the same time, and requires to reduce the air volume fluctuation of the shared air duct. Through the collaborative optimization algorithm, the comprehensive coupling strength between the coil area and the magnetic track area is calculated as 3.23, the priority of the coil area is 0.87, and the priority of the magnetic track area is 0.72. Based on these parameters, the original strategy of the coil area is adjusted: the air intake increase ratio is reduced from 25% to 20%, but the exhaust fan speed is increased from 300 rpm to 450 rpm, and the airflow guide plate angle is adjusted from 35° to 28°, and the airflow path is redistributed. This adjusted strategy not only meets the temperature control requirements of the coil area, but also reduces the interference to the magnetic track area. Additional auxiliary fan resources are allocated to the magnetic track area to compensate for the impact of the shared air duct adjustment. The final overall air flow adjustment strategy realizes the collaborative control of the linear motor multi-area temperature by precisely controlling the parameters of each execution unit.

[0061] The collaborative optimization algorithm of the present application realizes the overall coordinated control of the linear motor multi-temperature control area through the sharing of state information between agents and the evaluation of coupling effects. The dynamic update cycle mechanism effectively reduces the communication burden, the comprehensive coupling strength evaluation accurately captures the mutual influence between regions, and the conflict resolution based on priority and the strategy fusion based on contribution degree ensure the efficient allocation of resources, significantly improving the temperature control accuracy and energy efficiency of the linear motor.

[0062] In an optional implementation, the step of resolving conflicts of the optimized control strategy according to the priority includes: calculating a region temperature control performance parameter based on the optimized control strategy, and obtaining a region temperature control comprehensive score by weighting the region temperature control performance parameter; obtaining the remaining adjustment capability of the control execution unit, determining the adjustable air supply amount range and air supply time range of each region based on the priority, and generating an air supply strategy initial adjustment scheme within the air supply amount range and the air supply time range according to the temperature control comprehensive score difference of adjacent regions; for the region whose priority is higher than the reference value, when its temperature control comprehensive score is lower than the target threshold value, gradually expand the air supply amount range and the air supply time range of the region by a preset step, generate an air supply strategy supplementary scheme within the expanded range, and calculate the temperature control comprehensive score after generating the supplementary scheme each time, and stop adjusting when the score change value is less than a preset change threshold value for multiple times in succession; The initial adjustment scheme and the supplementary scheme are combined to form a final air supply strategy scheme, and a temperature control comprehensive score corresponding to the final air supply strategy scheme is calculated, and the scheme with the highest score is selected to execute conflict resolution.

[0063] For example, the optimized control strategy is used to calculate the regional temperature control performance parameters, including temperature control accuracy, response speed and energy efficiency. The temperature control accuracy refers to the deviation degree of the current temperature of the temperature control region from the target temperature and the stability of the temperature. The calculation method is the ratio of the absolute value of the temperature deviation to the allowed deviation range, and the standard deviation of the temperature fluctuation is also considered. Taking the coil area of the linear motor as an example, the target temperature is 65℃, the current temperature is 72℃, and the allowed deviation range is ±8℃. The temperature deviation ratio is |72-65| / 8=0.875. The standard deviation of the temperature fluctuation in the coil area in the last 5 minutes is 1.2℃, and the maximum allowed fluctuation is 5℃. The temperature stability index is 1-1.2 / 5=0.76. The final value of the temperature control accuracy is 0.875×0.7+0.76×0.3=0.839. The response speed refers to the matching degree of the temperature change rate and the target change rate and the acceleration of the temperature change. The target cooling rate of the coil area is 2.5℃ / min, the actual cooling rate is 1.8℃ / min, the change rate matching degree is 1.8 / 2.5=0.72. The cooling acceleration of the coil area is 0.4℃ / min 2 , the target acceleration is 0.5℃ / min 2 , and the acceleration matching degree is 0.4 / 0.5=0.8. The final value of the response speed is 0.72×0.6+0.8×0.4=0.752. The energy efficiency refers to the temperature regulation effect under unit energy consumption, and the calculation method is the percentage of the ratio of the temperature change amount to the energy consumption relative to the benchmark value. The coil area can be cooled by 12℃ per kilowatt-hour, and the benchmark value is 10℃ / kWh, so the energy efficiency is 12 / 10=1.2, i.e. 120%. The weights of the three parameters of the coil area are 0.4, 0.35 and 0.25 respectively, and the comprehensive score is 0.839×0.4+0.752×0.35+1.2×0.25=0.904. Similarly, the comprehensive score of the magnetic track area is 0.81, and the comprehensive score of the bearing area is 0.76.

[0064] The remaining regulation capacity of the control execution unit includes the fan speed regulation margin, the valve opening regulation margin, and the regulation time margin. The running state of the control execution unit is monitored in real time, and the immediately available regulation capacity is calculated. Taking the coil area exhaust fan as an example, the rated speed is 3000 rpm, and the current running speed is 1800 rpm. The preliminary remaining regulation capacity is (3000-1800) / 3000=40%. Considering the equipment protection factor, the maximum available speed of the exhaust fan will be reduced by 5% after continuous high-speed running for more than 30 minutes, i.e. to 2850 rpm. At the same time, the vibration sensor monitors the fan vibration amplitude of 0.3 mm, which is lower than the threshold of 0.5 mm, and no additional restrictions are needed. Comprehensive consideration, the actual maximum available speed is 2850 rpm, and the actual remaining regulation capacity is (2850-1800) / 3000=35%. The valve opening is currently 60%, the maximum opening is 100%, and the remaining regulation margin is (100%-60%) / 100%=40%. The regulation time margin refers to the percentage of the air supply time that can be extended or shortened under the premise of meeting the temperature control requirements, which is calculated to be ±25% according to the current temperature change rate and the target temperature. The distribution of air supply range and air supply time range adopts the priority proportional principle, that is, the higher the priority, the more regulation resources can be allocated. Under the condition that the remaining regulation capacity of the exhaust fan is 35%, the coil area (priority 0.87), the magnetic track area (priority 0.72), and the bearing area (priority 0.65) respectively obtain the upper limit of the regulation amount of 35% x 0.87 / (0.87+0.72+0.65)=13.8%, 35% x 0.72 / (0.87+0.72+0.65)=11.4%, and 35% x 0.65 / (0.87+0.72+0.65)=10.3%. Similarly, the air supply time range of each area is determined, such as the coil area air supply duration can be increased or decreased by 25% x 0.87 / (0.87+0.72+0.65)=9.8% based on the current value.

[0065] The temperature control comprehensive score difference calculation method of adjacent areas is: the score difference of the coil area and the magnetic track area is 0.904-0.81=0.094, and the score difference of the coil area and the bearing area is 0.904-0.76=0.144. The airflow between the areas of the linear motor exists complex mutual influence, and an airflow influence matrix is established to record the influence coefficient of the change of the air supply amount of each area on other areas. For example, the coil area air supply amount increases by 10%, which will cause the actual air flow of the magnetic track area to decrease by 3%, and at the same time, the air flow of the bearing area decreases by 2%; the magnetic track area air supply amount increases by 10%, which will cause the air flow of the coil area to decrease by 2.5%, and the air flow of the bearing area to decrease by 1.8%. The adjustment amount is proportional to the score difference, and the airflow influence matrix is considered. The initial air volume adjustment of the coil area relative to the magnetic track area is 13.8% x 0.094 / 0.144=9.0%, and the initial air volume adjustment of the coil area relative to the bearing area is 13.8% x 0.144 / 0.144=13.8%. After applying the airflow influence matrix, if the coil area air volume increases by 9.0%, it will cause the magnetic track area air volume to decrease by 9.0% x 0.3=2.7%; if it increases by 13.8%, it will cause the bearing area air volume to decrease by 13.8% x 0.2=2.76%. Iterative adjustment is needed until the actual air volume change of all areas meets the temperature control requirement. After three rounds of iterative calculation, it is finally determined that the coil area fan speed increases by 11.5%, that is, increases by 1800 x 11.5%=207 rpm, and is adjusted to 2007 rpm; the air supply time is extended by 8.5%, and the original air supply period is 120 seconds, which is extended to 120 x (1+8.5%) =130.2 seconds.

[0066] Assume the baseline priority value is 0.75 and the target comprehensive score threshold is 0.85. The coil area priority 0.87 is greater than the baseline value, and the temperature control comprehensive score 0.904 is greater than the target threshold, so no supplementary scheme needs to be generated. Consider the extreme case of sudden changes in the working environment temperature, the coil area temperature control comprehensive score will drop to 0.82, at which time a supplementary scheme needs to be generated. The preset step size is 20% of the original adjustment range, i.e. expanding the coil area adjustable air supply range and air supply time range by 2.76% (13.8% x 20%). In the expanded range, the coil area fan speed can be increased by (13.8% + 2.76%) x 1800 rpm = 298 rpm at most, and the air supply time can be extended by (9.8% + 1.96%) x 120 seconds = 14.1 seconds at most. The temperature control of the linear motor needs to meet both the temperature control and energy optimization targets. Multiple candidate supplementary schemes are generated, covering different combinations in the adjustment range, such as fan speed from 2050 rpm to 2300 rpm with a step size of 50 rpm; air supply time from 130 seconds to 145 seconds with a step size of 5 seconds, a total of 16 combinations. Calculate the comprehensive score for each combination, including the temperature control score (weight 0.7) and the energy consumption score (weight 0.3). The temperature control score is based on the predicted temperature control effect after the scheme is implemented, and the energy consumption score is based on the ratio of the predicted energy consumption to the baseline energy consumption, such as a 15% increase in energy consumption per hour, the energy consumption score is 1-0.15=0.85. After calculation, the optimal supplementary scheme is fan speed 2150 rpm and air supply time 140 seconds, with a comprehensive score of 0.89.

[0067] After the first adjustment, the coil area temperature control comprehensive score is increased from 0.82 to 0.87, with an increase of 0.05; in the second attempt, the fan speed is further increased to 2200 rpm and the air supply time is extended to 142 seconds, the score is increased to 0.88, and the increase is reduced to 0.01; the third adjustment has a score increase of only 0.005, which is less than the preset change threshold 0.01, at which time the adjustment is stopped and the second adjustment scheme is adopted. The state monitoring trigger mechanism is used to determine when to enable the supplementary scheme. Define multiple state indicators, such as temperature change rate, environmental temperature change and load change, etc. When the combined score of these indicators exceeds the threshold, the strategy switching is triggered. To avoid frequent switching causing oscillation, set the minimum duration to 60 seconds and the hysteresis interval (the score is between 0.9-1.1 to maintain the current strategy). When the strategy is switched, a linear interpolation method is used to achieve smooth transition, such as gradually adjusting the fan speed from 2007 rpm to 2150 rpm in 10 seconds, adjusting about 14.3 rpm per second, ensuring stability.

[0068] For the coil area, the initial adjustment scheme is a fan speed of 2007 rpm and a blowing time of 130.2 seconds, and the supplementary scheme is a fan speed of 2150 rpm and a blowing time of 140 seconds. The combined strategy needs to consider the compatibility and implementation conditions of the two schemes. The initial scheme is used when the environmental temperature is normal, and the supplementary scheme is switched to when the environmental temperature fluctuates greatly. The final scheme is set as follows: under normal conditions, the fan speed is 2007 rpm and the blowing time is 130.2 seconds; under abnormal temperature conditions (such as a temperature change rate exceeding 2°C / min or an environmental temperature change exceeding 5°C / hour), the fan speed is 2150 rpm and the blowing time is 140 seconds. To achieve smooth transition between the two strategies, when an abnormal condition is detected, a transition curve is calculated in advance to gradually change the control parameters according to the preset slope, avoiding oscillation caused by parameter mutation. For example, when the environmental temperature change rate is detected to be 4.5°C / hour (close to the threshold of 5°C / hour), the fan speed starts to gradually increase at a rate of 3 rpm per second, and the transition from 2007 rpm to 2150 rpm is completed in about 48 seconds. The final blowing strategy scheme is calculated for the comprehensive temperature control score, and the scheme with the highest score is selected to execute conflict resolution. Simulation calculations show that under different working conditions, the comprehensive score of the combined strategy always remains above 0.87, which is better than a single strategy. At the same time, the influence of this scheme on the magnetic track area and the bearing area is controllable, and it will not cause a significant decrease in the temperature control performance of these areas.

[0069] During the actual operation of the linear motor, fine tuning will also be performed according to the real-time state. For example, when it is detected that the coil area temperature rise rate is accelerating and it is predicted that the safety threshold will be exceeded after 20 minutes, the enhanced cooling mode will be started in advance, the blowing strategy will be switched from the initial scheme to the supplementary scheme, and the blowing strategies of adjacent areas will be adjusted at the same time to maintain overall air flow balance. For shared resources such as the air volume distribution of the main air duct, the distribution is performed according to the priority ratio, and a coordination mechanism is established between local areas. When the coil area needs additional air volume, the air volume of the bearing area with the lowest priority will be reduced first, and the total air volume will be appropriately increased to maintain overall air flow balance. The priority and resource allocation ratio will also be dynamically adjusted according to the temperature state change trend of each area to ensure that the core area temperature control demand is met while preventing temperature abnormalities in other areas.

[0070] The conflict resolution method of the present application scientifically quantifies the control requirements of each area by calculating refined temperature control performance parameters and temperature control comprehensive scores; reasonably allocates limited adjustment resources by using a priority-based resource allocation mechanism and an air flow influence matrix; improves adaptability and stability in response to environmental changes through the combination and smooth transition strategy of the initial scheme and the supplementary scheme; avoids oscillation caused by excessive adjustment by using multi-level adjustment and adaptive stop conditions, and realizes optimal control of linear motor multi-area temperature control under resource constraints.

[0071] In an alternative embodiment, the step of adaptively adjusting the weight coefficients of the reward function comprises: The historical temperature state and the current temperature state of the temperature control area are input into a learnable parameter matrix, the historical temperature state and the current temperature state are mapped and a hyperbolic tangent transformation is performed, and a state importance evaluation value is obtained through exponential normalization processing; a weight sensitivity is calculated according to the state importance evaluation value and the partial derivative of each component of the reward function; the weight sensitivity is multiplied by the temperature control performance gradient to obtain an attention weighted gradient; A temperature state distribution probability of the temperature control area is obtained, and a state dispersion degree is calculated based on the temperature state distribution probability; a temperature control effect evaluation value is weighted and adjusted based on the state dispersion degree to obtain an adaptive learning rate, and the temperature control effect evaluation value is obtained based on the actual temperature and the target temperature of the temperature control area; The adaptive learning rate is multiplied by the attention weighted gradient to update the weight coefficients of the reward function; the updated weight coefficients are dynamically threshold-constrained to obtain the final weight coefficients.

[0072] For example, a temperature state sequence of a temperature control area in the past period of time is obtained from a database. Taking a linear motor coil area as an example, temperature data in the past 30 minutes at intervals of 2 minutes is collected to form a historical temperature sequence of 15 sampling points, for example, [42.5℃, 43.1℃, 43.8℃, 44.2℃, 44.5℃, 44.7℃, 44.8℃, 45.0℃, 45.3℃, 45.5℃, 45.7℃, 45.8℃, 45.9℃, 46.0℃, 46.2℃], and the current temperature state is 46.5℃. The deviation of these temperature data from the target temperature 40℃ is converted into a normalized temperature deviation sequence [0.25, 0.31, 0.38, 0.42, 0.45, 0.47, 0.48, 0.50, 0.53, 0.55, 0.57, 0.58, 0.59, 0.60, 0.62, 0.65]. The learnable parameter matrix is initialized as a 4×16 matrix, where 4 is the feature dimension and 16 is the input sequence length. The normalized temperature deviation sequence is multiplied by the learnable parameter matrix to obtain a 4-dimensional feature vector, for example, [1.2, 0.8, -0.5, 0.3]. The hyperbolic tangent transformation is performed on the feature vector to map each element to the range of -1 to 1, obtaining [0.83, 0.66, -0.46, 0.29]. Through exponential normalization processing, the exponential value of each element is calculated and divided by the sum of all element exponential values to obtain the state importance evaluation value [0.42, 0.35, 0.12, 0.11], which represents the importance of temperature deviation size, temperature change trend, temperature distribution uniformity and control stability respectively.

[0073] The reward function is composed of three components: temperature uniformity index, energy utilization efficiency index, and heat flow disturbance coefficient, with current weights of [0.5, 0.3, 0.2]. For each component, the partial derivative with respect to the current state is calculated, which indicates the influence degree of temperature control decision on each component. Through numerical differentiation method, the partial derivative of temperature uniformity index is 0.08, indicating the improvement degree of current control strategy on temperature uniformity; the partial derivative of energy utilization efficiency index is -0.05, indicating that the current strategy slightly reduces energy efficiency; the partial derivative of heat flow disturbance coefficient is 0.03, indicating that the current strategy slightly increases heat flow disturbance. Multiply the first three elements of state importance evaluation value [0.42, 0.35, 0.12] with the three partial derivatives [0.08, -0.05, 0.03] respectively, to obtain the weight sensitivity [0.0336, -0.0175, 0.0036]. The weight sensitivity indicates the influence degree of changing the weight of each component on the overall control effect under the current state. The temperature control performance gradient refers to the rate at which the current temperature approaches the target temperature, with the current value being -0.15°C / min, and the negative value indicating that the temperature is decreasing. Multiply the weight sensitivity with the performance gradient to obtain the attention weighted gradient [-0.00504, 0.00263, -0.00054]. The attention weighted gradient indicates how to adjust the weight of each component to optimize the control effect under the current temperature control trend.

[0074] The temperature state of the temperature control area is divided into multiple intervals, for example, 35°C to 50°C is divided into 15 intervals, each interval is 1°C wide. Based on historical temperature data, the probability of the temperature falling into each interval is calculated. For the coil area, the temperature distribution probability in the past 24 hours is [0.05, 0.08, 0.12, 0.15, 0.18, 0.15, 0.10, 0.07, 0.05, 0.03, 0.01, 0.01, 0, 0, 0]. The state dispersion is an indicator to measure the concentration of temperature distribution, and the calculation method of information entropy is adopted, that is, the probability of each interval is taken logarithm and multiplied by the probability itself, then summed and taken negative value. The dispersion calculation result of the coil area temperature state is 2.15. The higher the dispersion, the more dispersed the temperature distribution, the more difficult the control. The temperature control effect evaluation value is calculated based on the difference between the actual temperature and the target temperature, the target temperature is 40°C, the actual temperature is 46.5°C, the difference is 6.5°C, and relative to the allowed deviation range ±10°C, the initial control effect evaluation value is 1-6.5 / 10=0.35. The initial evaluation value is adjusted by weighting with the state dispersion, the adjustment formula is to multiply the initial evaluation value by (1+the logarithmic value of the state dispersion ×0.05), the logarithmic value of the state dispersion 2.15 is about 0.76, so the adjustment coefficient is 1+0.76×0.05=1.038, and the adjusted temperature control effect evaluation value is 0.35×1.038=0.3633. The calculation method of the adaptive learning rate is to multiply the basic learning rate 0.01 by the adjusted temperature control effect evaluation value, so that the adaptive learning rate is 0.01×0.3633=0.00363, then multiplied by the state adaptation factor (1+state dispersion ×0.05)=1+2.15×0.05=1.1075, and the final adaptive learning rate is 0.00363×1.1075=0.00402.

[0075] The current reward function weight is [0.5, 0.3, 0.2], the adaptive learning rate is 0.00402, and the attention weighted gradient is [-0.00504, 0.00263, -0.00054]. After multiplication, the weight adjustment amount is [-0.0000203, 0.0000106, -0.0000022]. Add the adjustment amount to the current weight to get the updated weight [0.49998, 0.30001, 0.19999]. Dynamic threshold constraints include two aspects: one is to ensure that the sum of all weights is 1, and the other is to limit the change range of a single weight. Calculate the sum of the updated weights as 0.99998, and the difference from 1 is 0.00002. Distribute this difference to each weight according to the weight ratio to get the corrected weight [0.49999, 0.30001, 0.20000]. A single weight change limit is also set, for example, no more than 5% of the original weight per adjustment. Check the adjustment range this time, the temperature balance index weight change is -0.00001, which is 0.002% of the original weight; the energy utilization efficiency index weight change is 0.00001, which is 0.003% of the original weight; the heat flow interference coefficient weight change is 0, which has not changed. All changes are within the allowed range, so the final weight coefficient is [0.49999, 0.30001, 0.20000].

[0076] In practical applications, the temperature control effect is continuously monitored and the above weight adjustment process is repeatedly executed. As the control process progresses, the temperature control state changes, and the weight coefficient is also adjusted accordingly. For example, when the coil area temperature gradually approaches the target temperature of 40°C, the weight of the temperature balance index will decrease, and the weight of the energy utilization efficiency index will correspondingly increase, optimizing energy consumption. If the temperature fluctuates or is disturbed by external factors, the weight of the heat flow interference coefficient will increase, strengthening the suppression of interference. By adaptively adjusting the weight coefficient of the reward function, the temperature control strategy can flexibly adjust the priority of the control target according to the actual situation, achieving dynamic balance between temperature control and energy efficiency. For different working states of the linear motor, such as the startup phase, stable running phase and cooling down phase, the weight coefficient of the reward function will be automatically adjusted based on the control requirements of the current phase to optimize the control effect. For example, in the startup phase, the temperature balance index weight will be higher to ensure uniform temperature rise of each part; in the stable running phase, the energy utilization efficiency index weight will be increased to reduce energy consumption; in the cooling down phase, the heat flow interference coefficient weight will be increased to accelerate the cooling process using environmental heat dissipation.

[0077] The adaptive reward function weight adjustment method provided by the application extracts temperature state characteristics by introducing a learnable parameter matrix, calculates weight sensitivity in combination with an attention mechanism, and dynamically adjusts a learning rate by using state dispersion, so that accurate control of the weights of components of the reward function is realized; the method can automatically optimize the priority of the control target according to the real-time state and control demand of the temperature control area, and significantly improves the linear motor temperature control control precision, stability and energy utilization efficiency.

[0078] In a second aspect of the embodiment of the application, a linear motor temperature control system based on air cooling heat dissipation is provided, comprising: A first unit is configured to divide the linear motor into multiple temperature control areas, each temperature control area being configured with an independent control execution unit; based on temperature data and historical operation data of each temperature control area, an LSTM is used to predict a temperature change trend in a future operation process of the linear motor; A second unit is configured to assign an independent agent to each temperature control area, and construct a distributed control network based on multi-agent reinforcement learning; each agent receives the temperature change trend and the current temperature state as input, and based on a deep Q learning algorithm, takes temperature uniformity and energy consumption minimization as a reward function, and calculates an optimal control strategy for the area; based on the optimal control strategy of each area and the state information shared between agents, a collaborative optimization algorithm is used to balance the control requirements of each area, and a whole air flow adjustment strategy considering the coupling influence between areas is generated; A third unit is configured to dynamically adjust the opening degree of the electric regulating valve and the rotating speed of the air inlet and exhaust fan of each temperature control area according to the whole air flow adjustment strategy, until the temperature change rate and the energy consumption ratio reach a preset efficiency threshold and meet the temperature state In a third aspect of the embodiment of the application, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described above.

[0079] In a fourth aspect of the embodiment of the application, a computer readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the method described above.

[0080] The application can be a method, device, system and / or computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions loaded thereon for performing various aspects of the application.

[0081] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A temperature control method for a linear motor based on air cooling heat dissipation, characterized in that, The application relates to a temperature control method for a linear motor. The linear motor is divided into multiple temperature control areas, and each temperature control area is configured with an independent control execution unit. Based on temperature data and historical operation data of each temperature control area, the temperature change trend of the linear motor in future operation is predicted through LSTM. An independent agent is assigned to each temperature control area, and a distributed control network based on multi-agent reinforcement learning is constructed. Each agent receives the temperature change trend and the current temperature state as input, and calculates the optimal control strategy of the region based on the deep Q learning algorithm, with temperature uniformity and energy consumption minimization as the reward function.

2. The method of claim 1, wherein, Based on the optimal control strategy of each region and the shared state information between agents, the control requirements of each region are balanced through a collaborative optimization algorithm to generate an overall air flow adjustment strategy considering the coupling effect between regions. According to the overall air flow adjustment strategy, the opening degree of the electric regulating valve and the rotation speed of the inlet and exhaust fans of each temperature control area are dynamically adjusted until the temperature change rate and energy consumption ratio reach the preset efficiency threshold and meet the safety constraints of the temperature state of the linear motor. The step of predicting the temperature change trend of the linear motor in future operation based on temperature data and historical operation data of each temperature control area includes: Combining the temperature data, historical operation data of the target temperature control sub-area, and the temperature data of adjacent temperature control sub-areas to construct an input feature vector; Based on the temperature distribution characteristics and temperature change characteristics of the target temperature control sub-area, first and second adaptive coefficients are obtained respectively.

3. The method of claim 1, wherein, According to the first and second adaptive coefficients, the bias parameters of the forget gate and the input gate of the LSTM network corresponding to the target temperature control sub-area are adjusted respectively to establish an adaptive LSTM prediction network. Based on the thermal conductivity coefficient and distance parameter between the target temperature control sub-area and the adjacent temperature control sub-area, a thermal coupling matrix is constructed. The input feature vector is input into the adaptive LSTM prediction network to obtain the initial prediction state of the target temperature control sub-area. The historical prediction states of adjacent temperature control sub-areas are weighted and combined according to the thermal coupling matrix and added to the initial prediction state to obtain the temperature prediction value of the target temperature control sub-area. The step of each agent receiving the temperature change trend and the current temperature state as input, and calculating the optimal control strategy of the region based on the deep Q learning algorithm, with temperature uniformity and energy consumption minimization as the reward function, includes: The current temperature state of the temperature control area, the temperature change trend, and the temperature state of the adjacent area are constructed into an agent state space, and the parameter adjustment interval of the control execution unit of the temperature control area is constructed into an action space. A deep Q learning network is used to map the state space and the action space. Based on the current temperature state and the temperature change trend, adaptive weights are used to extract temperature dynamic features, and the state importance is calculated based on the temperature dynamic features to dynamically adjust the state sampling probability. construct a temperature deviation of the current temperature state, construct a temperature balance index based on the temperature deviation, evaluate cumulative energy consumption of the control sequence based on energy efficiency characteristic curves of the control execution unit of the temperature control region under different loads, and construct an energy utilization efficiency index, calculate a heat flow interference coefficient based on temperature states of adjacent temperature control regions, construct a reward function based on the temperature balance index, the energy utilization efficiency index, and the heat flow interference coefficient, and adaptively adjust weight coefficients of the reward function; train the deep Q learning network based on the state sampling probability and the reward function, perform sample learning based on an importance-based experience replay mechanism, and obtain an optimal control strategy of the temperature control region.

4. The method of claim 3, wherein, The step of performing sample learning based on the importance-based experience replay mechanism includes: obtain a change rate of the temperature change trend, a control response time of the temperature control region, and an environmental parameter fluctuation amplitude, and calculate a temperature failure risk evaluation value; divide a danger level of the experience sample into a high danger level, a warning level, and a normal level based on the temperature failure risk evaluation value, set a sampling weight according to the danger level of the experience sample, and perform sample extraction according to the sampling weight by using a roulette method; extract a temperature change feature sequence and a control parameter sequence of historical failure data of the temperature control region, generate an intermediate state sequence of a failure evolution process by using a time series interpolation method, and supplement the intermediate state sequence to an experience replay pool as a high danger level sample; when the temperature failure risk evaluation value of the current state is greater than a safety threshold, the deep Q learning network limits an exploration action to a historical safe action range during exploration.

5. The method of claim 1, wherein, The step of balancing control requirements of each region by a collaborative optimization algorithm based on the optimal control strategy of each region and state information shared between agents, and generating an overall air flow adjustment strategy considering coupling effects between regions includes: The state information shared between agents includes temperature field gradient information and air flow field characteristic quantities; calculate state information entropy based on the state information, and determine a dynamic update period of the state information according to the state information entropy; based on the temperature field gradient information, calculate a heat conduction coefficient in combination with a contact area and a characteristic distance of an adjacent region; based on the air flow field characteristic quantities, calculate air flow interference intensity of the adjacent region on a shared boundary area; and perform weighted summation on the heat conduction coefficient and the air flow interference intensity to obtain a comprehensive coupling strength; calculate a strategy adjustment amount according to the comprehensive coupling strength; superimpose the strategy adjustment amount on the optimal control strategy of each region to obtain an optimized control strategy, and set temperature balance constraints and air flow balance constraints for the optimized control strategy; determine a priority of each region according to the dynamic update period, the comprehensive coupling strength, and a control execution unit response time of the temperature control region; perform conflict resolution on the optimized control strategy according to the priority to obtain a resolved control strategy; calculate a contribution degree of each region in the overall control target, and use the contribution degree as a fusion weight; and perform weighted combination on the resolved control strategy according to the fusion weight to generate an overall air flow adjustment strategy.

6. The method of claim 5, wherein, The step of performing conflict resolution on the optimized control strategy according to the priority includes: Calculate a regional temperature control performance parameter based on the optimized control strategy, and obtain a temperature control comprehensive score of the region by weighting the regional temperature control performance parameter; Obtain a remaining adjustment capacity of the control execution unit, determine an adjustable air supply amount range and an air supply time range of each region based on the priority, and generate an initial adjustment scheme of the air supply strategy in the air supply amount range and the air supply time range according to a temperature control comprehensive score difference between adjacent regions; When the temperature control comprehensive score of the region with the priority higher than the reference value is lower than a target threshold value, gradually expand the air supply amount range and the air supply time range of the region by a preset step, generate a supplementary scheme of the air supply strategy in the expanded range, calculate the temperature control comprehensive score after each generation of the supplementary scheme, and stop adjusting when a score change value is smaller than a preset change threshold value for multiple times in succession; Combine the initial adjustment scheme of the air supply strategy and the supplementary scheme of the air supply strategy to form a final air supply strategy scheme, calculate a temperature control comprehensive score corresponding to the final air supply strategy scheme, and select a scheme with the highest score to execute conflict resolution.

7. The method of claim 3, wherein, The step of adaptively adjusting the weight coefficient of the reward function comprises: Input historical temperature states and the current temperature state of a temperature control region into a learnable parameter matrix, map the historical temperature states and the current temperature state, perform a hyperbolic tangent transformation, obtain a state importance evaluation value through exponential normalization processing, calculate a weight sensitivity according to the state importance evaluation value and a partial derivative of each component of the reward function, and multiply the weight sensitivity and a temperature control performance gradient to obtain an attention weighted gradient; Obtain a temperature state distribution probability of a temperature control region, calculate a state dispersion degree based on the temperature state distribution probability, and obtain an adaptive learning rate by weighting and adjusting a temperature control effect evaluation value based on the state dispersion degree, wherein the temperature control effect evaluation value is obtained based on an actual temperature and a target temperature of the temperature control region; Multiply the adaptive learning rate and the attention weighted gradient to update the weight coefficient of the reward function, and obtain a final weight coefficient by dynamically thresholding the updated weight coefficient.

8. A temperature control system for a linear motor based on air cooling for implementing the method according to any one of the preceding claims 1-7, characterized in that, The method comprises: A first unit is configured to divide a linear motor into a plurality of temperature control regions, and each temperature control region is configured with an independent control execution unit; Based on temperature data and historical operation data of each temperature control region, an LSTM is used to predict a temperature change trend in a future operation process of the linear motor; A second unit is configured to assign an independent agent to each temperature control region, construct a distributed control network based on multi-agent reinforcement learning, and each agent receives the temperature change trend and the current temperature state as input, calculates an optimal control strategy of the region based on a deep Q learning algorithm, takes temperature uniformity and energy consumption minimization as a reward function, balances control requirements of each region through a collaborative optimization algorithm based on the optimal control strategy of each region and shared state information between agents, and generates an overall air flow adjustment strategy considering the coupling influence between regions. The third unit is configured to dynamically adjust the opening degree of the electric regulating valve and the rotating speed of the air supply and exhaust fan of each temperature control area according to the overall air flow regulation strategy until the temperature change rate and the energy consumption ratio reach a preset efficiency threshold and meet the safety constraint of the linear motor temperature state.

9. An electronic device, comprising: The computer program instructions are executed by the processor to implement the method of any one of claims 1-7. The computer program instructions are executed by the processor to implement the method of any one of claims 1-7. The computer program instructions are executed by the processor to implement the method of any one of claims 1-7. ​ 10. A computer-readable storage medium having stored thereon computer program instructions, wherein, ​

Citation Information

Cited By

  • Cloud collaborative power consumption dynamic sensing and energy-saving optimization control method and system

    CN121578658A

  • Asymmetric air flow channel heat dissipation system based on AI computing power server

    CN121957307A