A roadway local cooling control method and system based on deep reinforcement learning
Through deep reinforcement learning methods, combined with temperature field, wind speed field and energy consumption data, a local cooling control system for the tunnel was constructed, which solved the shortcomings of existing technologies in local cooling control of deep tunnels and achieved accurate and efficient temperature field control and energy efficiency optimization.
Patent Information
- Application Number
- CN202510100353.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Existing technologies for local cooling control in deep tunnels suffer from several problems, including difficulty in accurately characterizing the dynamic characteristics of the temperature field, slow convergence speed of reinforcement learning strategies, lack of physical model constraints, and a disconnect between prediction and control leading to system response lag and low energy efficiency. Furthermore, the combination of deep learning and traditional control theory makes it difficult to simultaneously achieve control accuracy, system energy efficiency, and dynamic adaptability goals.
By constructing a roadway local cooling control method based on deep reinforcement learning, historical data of temperature field, wind speed field, humidity and energy consumption are obtained. After data preprocessing, a deep learning temperature field optimization and prediction model is constructed. A reinforcement learning control environment is designed, and a reinforcement learning policy network is constructed to achieve real-time optimization control of the local temperature of the roadway.
It achieves precise and efficient control of local temperature in the tunnel, improves the model's prediction accuracy and the system's response speed, ensures that the control strategy conforms to physical laws, and balances temperature control effect and energy efficiency.
Smart Images

Figure CN119882879B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent control, and particularly relates to a roadway local cooling control method and system based on deep reinforcement learning. BACKGROUND
[0002] With the development of mine exploitation and tunnel engineering, the high-temperature environment caused by the geothermal gradient has become a key factor restricting safe production. Deep roadway is narrow in space, complex in heat source, uneven in temperature distribution and difficult in ventilation, which makes local cooling control face severe challenges.
[0003] In the prior art, the roadway cooling is mainly carried out by fixed-point air supply, segmented refrigeration and full-face ventilation technology. The fixed-point air supply technology is simple to operate, but has air supply dead angles and is difficult to realize uniform temperature field regulation and control. The segmented refrigeration technology expands the refrigeration range by multi-point arrangement, but has high system energy consumption and interference between refrigeration points. The full-face ventilation technology has a wide coverage, but needs high-power equipment, seriously wastes energy and is difficult to meet the precise temperature control requirements of local areas. With the development of intelligent control technology, deep learning methods have been gradually applied in temperature field prediction and control. However, the method still has some problems: first, the deep neural network is difficult to accurately describe the dynamic characteristics of the complex temperature field; second, the reinforcement learning strategy has slow convergence speed and is easy to fall into local optimum; third, the lack of physical model constraints makes the prediction results violate the actual physical laws; fourth, the separation of prediction and control leads to system response lag and low energy efficiency. In addition, the combination scheme of deep learning and traditional control theory mainly stays at the simple combination level, and when dealing with the highly nonlinear and strongly coupled roadway temperature field problem, it is difficult to balance the control accuracy target, system energy efficiency target and dynamic adaptability target.
[0004] At present, there is little research on the local cooling control of deep roadway, and there is no specific temperature field control method combining deep learning, reinforcement learning and model predictive control. SUMMARY
[0005] In view of the defects in the prior art, the present application provides a roadway local cooling control method and system based on deep reinforcement learning.
[0006] In a first aspect, the present application provides a kind of local cooling control method of roadway based on deep reinforcement learning, comprising the following steps: obtaining the historical data of temperature field, wind speed field, humidity and energy consumption in roadway environment;The historical data is preprocessed to obtain standardized data;Based on the standardized data, an optimized prediction model of deep learning temperature field is constructed;Based on the optimized prediction model, a reinforcement learning control environment is designed;Based on the reinforcement learning control environment, a reinforcement learning strategy network is constructed;Based on the reinforcement learning strategy network, real-time optimization control of local temperature of roadway is realized.The present application collects the historical data of temperature field, wind speed field, humidity and energy consumption in roadway, provides a solid foundation for subsequent accurate prediction and control;By pre-processing the historical data, the standardization and consistency of data are ensured, and the prediction accuracy of the model is improved;By deep learning method, the optimized prediction model of temperature field is constructed, and the complex law of environmental change is efficiently captured;By optimizing the prediction model, the reinforcement learning control environment is designed, which provides a simulation and verification platform for the development of intelligent control strategy;By constructing reinforcement learning strategy network, real-time response to the change of roadway environment is realized, and the cooling strategy is automatically adjusted, to realize accurate and efficient control of local temperature of roadway.
[0007] Optionally, the pre-processing of the historical data to obtain standardized data comprises: data cleaning and outlier processing of the historical data to obtain optimized data;The optimized data is standardized to obtain standardized data.The present application effectively eliminates the error, repetition or unreasonable data points by data cleaning and outlier processing, ensures the quality and reliability of data, and provides a solid foundation for subsequent modeling and control;By standardizing the optimized data, the dimensional differences between different variables are eliminated, which not only helps to improve the training efficiency and prediction accuracy of the model, but also makes the model more robust and easy to interpret.
[0008] Optionally, the constructing an optimized prediction model of a deep learning temperature field based on the standardized data comprises: constructing a deep neural network structure based on the standardized data; designing an input layer of temperature field and wind speed field data based on the deep neural network structure; designing a feature extraction layer and a time series prediction layer based on the input layer; designing a loss function with physical constraints based on the feature extraction layer and the time series prediction layer; determining model training parameters and optimization methods based on the loss function with physical constraints; and constructing an optimized prediction model of a deep learning temperature field through the model training parameters and the optimization methods. The application fully captures the spatio-temporal characteristics of temperature field and wind speed field by designing a deep neural network structure, thereby improving the prediction ability of the model. The application ensures that temperature field and wind speed field data are accurately and efficiently input into the model by designing an input layer. The application deeply mines the potential rules and trends of data by designing a feature extraction layer and a time series prediction layer. The application ensures that the prediction results of the model conform to the physical laws by introducing a loss function with physical constraints. The application constructs an optimized prediction model of a deep learning temperature field by determining model training parameters and optimization methods, thereby providing strong support for real-time optimization control of local temperature in a roadway.
[0009] Optionally, the designing a reinforcement learning control environment based on the optimized prediction model comprises: designing an environment state space, an action space and a reward function according to the optimized prediction model, wherein the environment state space comprises temperature field distribution, wind speed field distribution, fan speed and guide vane angle, the action space comprises adjustment amplitude of fan speed and adjustment amplitude of guide vane angle, and the reward function comprises target temperature and current energy consumption. The application comprehensively perceives the environment for a reinforcement learning algorithm by designing an environment state space that covers temperature field distribution, wind speed field distribution, fan speed and guide vane angle information. The application provides an effective regulation and control means for a temperature control system by designing an action space that clearly defines adjustment amplitudes of fan speed and guide vane angle. The application ensures that a temperature control system pursues temperature control effect while also taking into account energy consumption efficiency by designing a reward function that comprehensively considers target temperature and current energy consumption.
[0010] Optionally, the constructing a reinforcement learning strategy network based on the reinforcement learning control environment comprises: designing a network structure and a loss function of deep reinforcement learning based on the reinforcement learning control environment; and constructing a reinforcement learning strategy network based on the network structure and the loss function of deep reinforcement learning. The application designs a network structure of deep reinforcement learning according to the characteristics of a reinforcement learning control environment, thereby efficiently processing complex environment states and action spaces and providing strong decision-making ability for a temperature control system. The application ensures that a strategy network constantly approaches an optimal strategy in a training process by designing a loss function that matches the strategy network. The application intelligently adjusts fan speed and guide vane angle according to real-time environment states by constructing a reinforcement learning strategy network, thereby achieving precise regulation and control of a roadway temperature field.
[0011] Optionally, the network structure and loss function based on the deep reinforcement learning are used to construct a reinforcement learning strategy network, including: based on the network structure and loss function of the deep reinforcement learning, using the experience replay pool to perform strategy training iterations to obtain iterative results; based on the iterative results, evaluating strategy performance indicators and constructing a reinforcement learning strategy network, wherein the strategy performance indicators include target area temperature control accuracy, temperature field distribution uniformity, wind speed field uniformity, system comprehensive energy consumption and control response time. The present invention effectively utilizes historical data, improves training efficiency and strategy stability, and obtains accurate iterative results by utilizing the experience replay pool to perform strategy training iterations; by comprehensively evaluating strategy performance indicators, including target area temperature control accuracy, temperature field distribution uniformity, wind speed field uniformity, system comprehensive energy consumption and control response time, the excellent performance of the strategy in practical applications is ensured; by constructing a reinforcement learning strategy network, the fan speed and wind guide plate angle are intelligently adjusted to achieve accurate and efficient control of the tunnel temperature field.
[0012] Optionally, the real-time optimization control of the local temperature of the tunnel based on the reinforcement learning strategy network includes: constructing a prediction model that takes into account the dynamic characteristics of the temperature field based on the reinforcement learning strategy network, and the prediction model sets a prediction time domain and a control time domain; based on the prediction model, designing an optimization objective function that includes temperature control accuracy and energy efficiency; generating a dynamic optimal control strategy through the optimization objective function and constraints; and realizing real-time optimization control of the local temperature of the tunnel according to the optimal control strategy. The present invention provides a reference in the time dimension for real-time optimization by constructing a prediction model that takes into account the dynamic characteristics of the temperature field and sets a prediction time domain and a control time domain; by designing an optimization objective function that includes temperature control accuracy and energy efficiency, it ensures that the optimization process pursues both high-precision temperature control and energy efficiency improvement; by optimizing the objective function and constraints, a dynamic optimal control strategy is generated, and accordingly, real-time optimization control of the local temperature of the tunnel is realized.
[0013] In a second aspect, the present application provides a roadway local cooling control system based on deep reinforcement learning, comprising an input device, a processor, an output device and a memory; the input device comprises a data acquisition module; the processor comprises a model prediction module, a decision control module and an execution control module; the input device, the processor, the output device and the memory are connected with each other; the memory is used for storing a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions. Through the modular design, each module can be independently developed, upgraded and maintained, thereby improving the overall flexibility and scalability of the system; by integrating the model prediction module, the decision control module and the execution control module, the data collected by the input device can be efficiently processed, and accurate decisions can be made according to the preset algorithm and logic, so as to ensure that the system can quickly respond when facing complex environments and realize accurate control; through the programmed control mode, the behavior of the system can be predicted and repeated, thereby enhancing the stability and reliability of the system.
[0014] Optionally, the sampling frequency of the data acquisition module is adjustable, and has data anomaly detection and automatic calibration functions, comprising a temperature sensor array and a wind speed sensor array, the temperature sensor array is arranged in a grid form at key positions of the roadway, and the wind speed sensor array is arranged at key sections of the airflow channel; the model prediction module has online learning and model updating functions, comprising a deep learning prediction unit and a reinforcement learning strategy unit, for predicting temperature field distribution and generating control strategies; the decision control module has multi-objective optimization and real-time response capabilities, comprising a model prediction control unit, for optimizing control strategies and generating control instructions; the execution control module has fault diagnosis and automatic protection functions, comprising a double-axial flow fan unit, an adjustable air deflector unit and a cold wall system, the double-axial flow fan unit and the air deflector are arranged in cooperation, the adjustable air deflector unit comprises independently adjustable air deflector units, and the cold wall system provides stable cold source output. The acquisition module of the present application ensures the accuracy and reliability of sensor data through flexible sampling frequency and anomaly detection and automatic calibration functions, providing a solid foundation for subsequent temperature field prediction and control; the model prediction module combines the deep learning prediction unit and the reinforcement learning strategy unit to realize accurate prediction of the temperature field and generation of intelligent control strategies, improving the prediction ability and control precision of the system; the decision control module and the execution control module ensure the accuracy of the control instructions and the safety of the execution process through multi-objective optimization and real-time response, as well as fault diagnosis and automatic protection functions, realizing accurate, efficient and safe control of the local temperature of the roadway.
[0015] Optionally, the double-shaft axial flow fan set comprises a main fan, an auxiliary fan, a variable frequency control unit, an angle adjusting mechanism and a position adjusting mechanism; the main fan is used to provide main airflow; the auxiliary fan is used to adjust the airflow direction; the variable frequency control unit is used to adjust the fan rotating speed; the angle adjusting mechanism is used to adjust the angle within the range of 0° to 90°; and the position adjusting mechanism is used to adjust the position of the air deflector. The main fan of the present application provides main airflow, ensuring the basic demand of air circulation in the roadway; the auxiliary fan further optimizes the air flow path by adjusting the airflow direction, improving the ventilation efficiency; the variable frequency control unit adjusts the fan rotating speed according to the actual demand, realizing the fine management of energy consumption, meeting the ventilation demand and reducing the energy consumption; and the angle adjusting mechanism and the position adjusting mechanism work cooperatively, precisely controlling the airflow distribution, ensuring the uniform temperature in each area of the roadway. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 a flow chart of the deep reinforcement learning-based local cooling control method for a roadway of an embodiment of the present application;
[0017] Figure 2 a schematic diagram of the overall structure of the deep reinforcement learning-based local cooling control system for a roadway of an embodiment of the present application;
[0018] Figure 3 a schematic diagram of the main structure of the deep reinforcement learning-based local cooling control system for a roadway of an embodiment of the present application;
[0019] Figure 4 a schematic diagram of the air flow direction of an embodiment of the present application;
[0020] Figure 5 a flow chart of the operation of the deep reinforcement learning-based local cooling control system for a roadway of an embodiment of the present application;
[0021] Figure 6 a schematic diagram of the system time sequence temperature field structure of an embodiment of the present application.
[0022] BRIEF DESCRIPTION OF DRAWINGS
[0023] 1 is an axial flow fan set; 2 is an air deflector set; 3 is a cold arm; 4 is a monitoring point; and 5 is a working face. DETAILED DESCRIPTION
[0024] The specific embodiments of the present application will be described in detail below, and it should be noted that the embodiments described herein are only used for illustration and do not limit the present application. In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present application. However, it is obvious to those skilled in the art that the present application does not have to be implemented with these specific details. In other instances, well-known circuits, software or methods have not been specifically described in order not to obscure the present application.
[0025] Reference throughout this specification to "one embodiment", "an embodiment", "one example", or "an example" means that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the application. The appearances of the phrases "in one embodiment", "in an embodiment", "one example" or "an example" in various places in the specification are not necessarily all referring to the same embodiment or example. Furthermore, the particular features, structures, or characteristics can be combined in any suitable
[0026] Reference throughout this specification to "one embodiment", "an embodiment", "one example", or "an example" means that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the application. The appearances of the phrases "in one embodiment", "in an embodiment", "one example" or "an example" in various places in the specification are not necessarily all referring to the same embodiment or example. Furthermore, the particular features, structures, or characteristics can be combined in any suitable Figure 1 The embodiments of the present application provide a local cooling control method for a roadway based on deep reinforcement learning, which comprises the following steps:
[0027] S1. Obtain historical data of temperature field, wind speed field, humidity and energy consumption in the roadway environment.
[0028] S1 further comprises the following steps:
[0029] S11. Define system parameters.
[0030] In one embodiment, the operating range of the fan is defined, including the upper and lower limits of the adjustable rotating speed, the rated air volume and the corresponding air pressure range, the layout position, quantity and spacing of the fan are reasonably planned to ensure the airflow coverage of the target area. For the air deflector, the adjustment range of the deflection angle is defined from zero to ninety degrees, and the kinetic energy recovery coefficient and friction coefficient required for air flow distribution are set in combination with the material. In terms of roadway parameters, the cross-sectional size, surface roughness and ventilation distance are defined to provide support for the calculation of along-path resistance and the prediction of air flow attenuation. These parameters lay the foundation for model construction and system optimization.
[0031] S12. Based on the system parameters, data acquisition configuration is performed.
[0032] Specifically, the data acquisition configuration includes temperature sensor arrangement, humidity sensor arrangement, wind speed sensor arrangement and sampling frequency setting.
[0033] The temperature sensor arrangement method is as follows: according to the temperature control requirements in the roadway, the appropriate type and quantity of sensors are selected to ensure coverage of key areas, including fan outlets, before and after air deflectors and target cooling areas. At the same time, the accuracy and range of the sensors are set to meet the monitoring needs of actual temperature changes. When arranging, the principle of reasonable distribution is followed: more sensors are arranged in areas with large temperature differences, and the number of sensors is appropriately reduced in areas with gentle temperature changes, so as to balance monitoring accuracy and cost efficiency.
[0034] The wind speed sensor arrangement method is as follows: focus on covering the fan outlet, the key position of the air deflector and the target cooling area to capture the dynamic changes of the wind speed distribution. According to the characteristics of the air flow in the roadway, the measurement range and accuracy of the sensor are determined to ensure that the fluctuations and distribution of the wind speed can be accurately reflected. The sensor positions are reasonably distributed to ensure the comprehensiveness of the data and take into account the complexity of the roadway structure and actual needs.
[0035] The humidity sensor arrangement method is as follows: mainly arranged in a small area near the fan to monitor the humidity changes around the fan in real time, providing a basis for subsequent energy consumption data acquisition.
[0036] The sampling frequency setting method is as follows: the temperature sensor, humidity sensor and wind speed sensor are sampled synchronously at a frequency of one (1 Hz) or higher to ensure the dynamic consistency of the data and provide real-time environmental feedback for the control system. At the same time, the storage path of the data should be configured to support the transmission of sensor data to the processing center through wireless or wired means to realize real-time monitoring and data analysis functions.
[0037] S13. Based on the data acquisition configuration, historical data of temperature field, wind speed field, humidity field and energy consumption in the roadway environment are acquired.
[0038] In one embodiment, by arranging temperature sensors, wind speed sensors and humidity sensors, historical data of temperature field, wind speed field and humidity in the roadway environment are directly acquired.
[0039] Further, the historical data of the temperature field, the wind speed field and the humidity are substituted into the energy consumption calculation expression of the fan system to obtain the historical data of the energy consumption.
[0040] S2. The historical data is preprocessed to obtain standardized data.
[0041] Wherein, S2 further includes the following steps:
[0042] S21. The historical data is subjected to data cleaning and outlier processing to obtain a processing result.
[0043] In one embodiment, the historical data is subjected to data cleaning.
[0044] Specifically, the collected data is imported into a data cleaning tool or a programming environment, such as the Pandas library of Python.
[0045] Further, the missing values in the data are checked. For records with few missing values, linear interpolation or mean filling is considered. If there are many missing values, these records need to be deleted or the data needs to be re-collected.
[0046] Further, duplicate values in the data are checked. For exactly identical records, only one is kept. If the records are similar but not exactly identical, further analysis is needed to determine whether to keep.
[0047] Further, all data is converted to a uniform format and unit. For example, temperature is converted to Celsius, wind speed is converted to meters / second, humidity is converted to percentage, and energy consumption is converted to kilowatt-hours.
[0048] In another embodiment, the historical data after data cleaning is subjected to outlier processing.
[0049] Specifically, visualization tools such as box plots and scatter plots are used to visually identify outliers. Box plots can show the distribution of data and the median, as well as the range between the first quartile and the third quartile. Scatter plots can show the relationship between variables and help identify possible outliers.
[0050] Further, outliers are judged in combination with the physical meaning and common sense of the roadway environment. For example, if a certain temperature value is much lower or much higher than other records and has no reasonable explanation, this value is judged as an outlier.
[0051] Further, for the identified outliers, replacement, deletion or marking processing is performed.
[0052] It should be noted that the processing of historical data of the temperature field includes: dividing the temperature field distribution data in the roadway into regions, extracting key region temperature features, ensuring balanced data distribution, and avoiding excessive or insufficient data in a region. The processing of historical data of the wind speed field includes: extracting features of the wind speed distribution in the roadway, capturing cross-sectional wind speed gradient data and smoothing wind speed data and performing denoising processing.
[0053] S22. Standardizing the processing result to obtain standardized data.
[0054] In one embodiment, the pre-processed historical data in step S21 is subjected to standardization processing, eliminating the dimensional difference of the data and improving the stability and convergence speed of the model training.
[0055] Further, the data distribution is adjusted to standardized data with a mean of 0 and a standard deviation of 1.
[0056] S3. Based on the standardized data, an optimized prediction model of deep learning is constructed.
[0057] Wherein, S3 further comprises the following steps:
[0058] S31. Based on the standardized data, a deep neural network structure is constructed.
[0059] wherein S31 further comprises the following steps:
[0060] S311. Based on the standardized data, a physical model of ventilation and cooling is established.
[0061] In one embodiment, based on the standardized data, a physical model of ventilation and cooling is established in combination with parameters of the fan system. The physical model includes a calculation model of ventilation resistance, a calculation model of wind speed distribution, and a temperature field prediction model.
[0062] Specifically, the expression of ventilation resistance is as follows:
[0063]
[0064] wherein, is the pressure drop caused by the resistance along the path, is the resistance coefficient along the path, is the air density, is the wind speed.
[0065] Further, the expression of the fan outlet flow is as follows:
[0066]
[0067] wherein, is the volume flow rate at the fan outlet, is the wind speed at the fan outlet, is the cross-sectional area at the fan outlet.
[0068] Further, the expression of the wind pressure is as follows:
[0069]
[0070] wherein, is the dynamic pressure at the fan outlet, is the air density, is the wind speed at the fan outlet.
[0071] Further, the expression of the deflector reflection and incident wind speed decomposition is as follows:
[0072]
[0073]
[0074] wherein, is the component of the incident wind speed in the normal (perpendicular to the deflector) direction, is the component of the incident wind speed in the tangential (parallel to the deflector) direction, is the angle between the deflector and the air flow of the fan.
[0075] Further, the reflected wind velocity component satisfies the following expression:
[0076]
[0077] wherein, is the normal component of the reflected wind velocity, is the recovery coefficient.
[0078] Further, the tangential component satisfies the following expression:
[0079]
[0080] wherein, is the tangential component of the reflected wind velocity, is the friction coefficient.
[0081] Further, the combined wind velocity and direction are obtained, and the horizontal component and the vertical component satisfy the following expressions:
[0082]
[0083]
[0084] wherein, , is the horizontal component of the combined wind velocity, , is the horizontal velocity component of the two wind flows, , is the vertical velocity component of the two wind flows, , is the mass flow of the two wind flows.
[0085] Further, the combined wind velocity satisfies the following expression:
[0086]
[0087] wherein, is the total combined wind velocity.
[0088] Further, the combined direction angle satisfies the following expression:
[0089]
[0090] wherein, is the combined wind flow direction angle.
[0091] Further, the temperature field distribution formula is as follows:
[0092]
[0093] wherein, is a temperature potential field, is a time variable, is a velocity field vector, is a diffusion term, is a thermal diffusivity, is a heat source sink function.
[0094] Further, the thermal diffusivity satisfies the following relationship:
[0095]
[0096] wherein, is a thermal conductivity, is a fluid density, is a specific heat at constant volume.
[0097] S312. According to the physical model of the ventilation and cooling, a performance index is determined.
[0098] In one embodiment, according to the method design target and actual demand, the core performance index of the air flow optimization and cooling system is defined to evaluate the cooling effect and optimize the space. These indexes mainly include cooling effect, air speed uniformity and system energy consumption.
[0099] Specifically, the temperature control deviation is the core index to evaluate the cooling effect, and the temperature control deviation satisfies the following expression:
[0100]
[0101] wherein, is a temperature control deviation index, is a total number of monitoring points, is an actual temperature of the th monitoring point, is a predicted temperature.
[0102] Further, the air speed distribution standard deviation is the core index to evaluate the air speed uniformity, and the air speed distribution standard deviation satisfies the following expression:
[0103]
[0104] wherein, is an air speed distribution standard deviation, is an air speed of the th monitoring point, is an average air speed in the cross section.
[0105] Further, the air speed direction is consistent, and the corresponding index expression is as follows:
[0106]
[0107] wherein, is a wind speed direction consistency index, is the target wind flow direction angle. is the stock wind flow direction angle, is the target wind flow direction angle.
[0108] Further, the energy consumption of the fan satisfies the following expression:
[0109]
[0110] wherein, is the total energy consumption of the fan, is the air volume, is the static pressure power consumption of the fan, is the fan efficiency, is the initial time, is the running end time.
[0111] Further, the energy consumption of the cooling system satisfies the following expression:
[0112]
[0113] wherein, is the total energy consumption of the cooling system, is the cooling power per unit time.
[0114] Further, the system energy efficiency ratio is as follows:
[0115]
[0116] wherein, is the refrigeration capacity of the system; is the total energy consumption of the system, satisfying the following expression:
[0117]
[0118] S32. Design an input layer of temperature field and wind speed field data based on the deep neural network structure.
[0119] In one embodiment, an input layer of temperature field and wind speed field data is designed based on the deep neural network structure. Among them, the design of the input layer takes the key features of the roadway environment as the core, adopts the temperature field (two-dimensional or three-dimensional matrix distribution), the wind speed field (horizontal and vertical components), and the fan parameters and the guide vane state control variables. The standardized data is organized and structured into a form that can be directly processed by the network, providing support for subsequent feature extraction and dynamic prediction.
[0120] S33. Design a feature extraction layer and a time series prediction layer based on the input layer to obtain a prediction model of the deep learning temperature field and wind speed field.
[0121] In one embodiment, based on the input layer, the design of the hidden layer is performed through convolution operation and time series processing. Among them, the feature extraction layer of the hidden layer is obtained through the convolution operation; the time series prediction layer of the hidden layer is obtained through the time series processing.
[0122] Specifically, the convolution operation is implemented through a convolution model which satisfies the following expression:
[0123]
[0124] Among them, is the output feature after the convolution operation, is the input feature, is the convolution kernel weight, is the bias term, , , , is the feature sequence before the convolution operation, , are two different maximum feature sequences, , is the sequence corresponding to the input feature,
[0125] The time series processing is implemented through a time series model which satisfies the following expression:
[0126]
[0127] Among them, is the hidden state of the current time step, is the output gate, is the hyperbolic tangent activation function, is the memory cell state of the current time step, and its update mechanism satisfies the following expression:
[0128]
[0129] Among them, is the memory cell state of the current time step, is the forget gate, is the memory cell state of the previous time step, is the input gate, is the hyperbolic tangent activation function, is the weight matrix, is the hidden state of the previous time step, is the input of the current time step, is the bias term. Through the time series processing, the time series prediction layer is obtained.
[0130] Further, according to the input layer and hidden layer of the design, combined with other layer structures of the neural network structure, a prediction model of the temperature field and the wind speed field of deep learning is constructed.
[0131] S34. The prediction model is trained to obtain an optimized prediction model.
[0132] In one embodiment, the prediction model of the temperature field and the wind speed field of deep learning is trained. Wherein, the implementation steps of training include: training data set division, learning parameter setting, model training process and training effect evaluation.
[0133] Specifically, the training data set division includes: dividing the processed temperature field and wind speed field data set into training set, validation set and test set. Wherein, the training set is used for model parameter optimization, accounting for 70%~80% of the total data; the validation set is used for evaluating the performance of the model during the training process, accounting for 10%~15% of the total data; the test set is used for evaluating the performance of the final model, accounting for 10%~15% of the total data. The data division adopts the method of random sampling, while ensuring that there is no cross sample between different data sets, so as to ensure the accuracy of the evaluation of the generalization ability of the model. In order to avoid the model bias caused by uneven distribution of data, the data is balanced sampled according to the regional characteristics of the temperature field and the wind speed field, so that the data proportion of each sub region is consistent, thereby improving the adaptability of the model to different regional environment.
[0134] The learning parameter setting includes: in the training process, a self-defined loss function is used, which combines the prediction error and the physical consistency constraint, and the loss function is as follows:
[0135]
[0136] Wherein, is the loss function, is the mean square error between the model prediction value and the true value, is the adjustment weight, is the momentum conservation equation of the temperature field heat diffusion law, which is defined as:
[0137]
[0138] Wherein, is the time rate of change of temperature, is the fluid velocity vector, is the gradient of the temperature field, is the thermal diffusion coefficient, is the Laplace operator of the temperature field, is the heat source term, is the square of two norm.
[0139] Further, the Adam optimizer is used to optimize the algorithm, and the parameter update is expressed as:
[0140]
[0141] where, is the time temperature value at time is the time temperature value at time is the learning rate, is the first-order momentum, is the second-order momentum, is a small value to prevent the denominator from being zero.
[0142] The model training process includes the following three key steps: forward propagation, error calculation, and back propagation. In the forward propagation stage, the training data passes through the input layer, hidden layer, and output layer in turn to generate the predicted value. Then, based on the designed loss function, the error between the predicted value and the true value is calculated. Next, the weights are updated by calculating the gradient of the loss with respect to the model parameters, and multiple iterations are performed using the optimization algorithm to gradually reduce the training error, thereby realizing back propagation. To accelerate convergence and avoid overfitting, the learning rate is dynamically adjusted, and the model performance is monitored on the validation set to trigger the early stopping mechanism and save the best model parameters. After training is complete, the model prediction accuracy and physical consistency are evaluated on the test set to ensure the model's generalization ability and reliability in real-world scenarios.
[0143] The training effect evaluation is completed through the validation dataset and the test dataset. After each complete training of a dataset, the validation data is input into the model, the validation error is calculated to evaluate the model's generalization ability, and the model's training state is determined according to the trend of the validation error. To avoid overfitting of the model, an early stopping mechanism is used, and when the validation error no longer decreases after multiple consecutive complete training, the training is stopped and the current optimal model parameters are saved. After training is complete, the test dataset is used to evaluate the model's final performance, the prediction error and physical consistency indicators are calculated to verify the model's prediction accuracy and compliance with physical laws in actual application scenarios, ensuring its reliability and applicability in practical applications. When the training effect is the best, the optimized prediction model is obtained.
[0144] It should be noted that in the optimization process of the prediction model, the performance of the model is determined by the size of the relative error, which satisfies the following expression:
[0145]
[0146] where, is the relative error, is the actual value, is the predicted value.
[0147] Furthermore, in the output layer, the predicted values of the temperature field and wind speed field are used to correct the difference with the measured values of the sensor to achieve real-time optimization of the prediction results. The expression of the difference correction is as follows:
[0148]
[0149] in, is the corrected temperature, is the initial temperature predicted by the model, is the actual temperature measured by the sensor, is the current temperature predicted by the model, is the corrected wind speed vector, is the initial wind speed vector predicted by the model, is the current wind speed vector predicted by the model, is the true wind speed vector measured by the sensor.
[0150] Furthermore, the fan speed and the wind deflector angle are adjusted. The adjustment expression is as follows:
[0151]
[0152] in, is the predicted value of wind speed field, is the target value of the speed field, To adjust the fan speed, is the current fan speed.
[0153] Furthermore, the angle of the air deflector is adjusted, and the expression for the adjustment is as follows:
[0154]
[0155] in, The angle deviation of the wind deflector needs to be adjusted. , is the wind speed vector in the horizontal direction , The weight, is the predicted angle, is the target angle.
[0156] S4. Based on the optimization prediction model, design a reinforcement learning control environment.
[0157] In one embodiment, the design of the reinforcement learning control environment is achieved by designing the state space, action space, and reward function.
[0158] Specifically, in the process of designing the state space, the state vector includes the current temperature field distribution, the wind speed field distribution, the fan speed and the angle of the deflector, and the specific form is as follows:
[0159]
[0160] wherein, is the state vector, is the temperature field distribution, is the wind speed field distribution, is the current fan speed, is the current angle of the deflector.
[0161] Further, in the process of designing the action space, the action vector includes the adjustment amplitude of the fan speed and the adjustment amplitude of the angle of the deflector, and the specific form is as follows:
[0162]
[0163] wherein, is the action vector, is the change of the fan speed, is the change of the angle of the deflector.
[0164] Further, in the process of designing the reward function, the reward function takes the temperature control effect and the energy consumption optimization as the target, and the specific form is as follows:
[0165]
[0166] wherein, is the reward function, , is the weight coefficient, is the current energy consumption, is the target temperature, is the current temperature.
[0167] S5. Based on the reinforcement learning control environment, a reinforcement learning strategy network is constructed.
[0168] wherein, S5 further includes the following steps:
[0169] S51. Based on the reinforcement learning control environment, a network structure and a loss function of deep reinforcement learning are designed.
[0170] In one embodiment, after the reinforcement learning control environment has been completed, the reinforcement learning strategy network is constructed, including constructing the network structure and the loss function.
[0171] Specifically, the network structure uses a deep reinforcement learning network, including an input layer, a hidden layer and an output layer. The loss function of the strategy network has the following specific form:
[0172]
[0173] wherein, is a loss function of the policy network, is a state distribution expectation under a sampled case, is a current policy, is a state of a time step, is an action of a time step, is a state distribution, is an action value function.
[0174] S52. Constructing a reinforcement learning policy network based on the deep reinforcement learning network structure and loss function.
[0175] wherein, S52 further comprises the following steps:
[0176] S521. Based on the deep reinforcement learning network structure and loss function, using an experience replay pool to perform a policy training iteration to obtain an iteration result;
[0177] In one embodiment, an experience replay and policy update method is used to train and optimize the reinforcement learning policy network.
[0178] Specifically, the experience replay method comprises storing historical quadruples in an experience replay pool and randomly sampling for training. The policy update method uses a policy update model, which satisfies the following expression:
[0179]
[0180] wherein, is a policy gradient, is an expected value of a subsequent random variable, is a policy objective function, is a parameterized policy, is a gradient of a policy probability, is an action value function.
[0181] S522. Based on the iteration result, evaluating a policy performance indicator to construct a reinforcement learning policy network, wherein the policy performance indicator comprises a target area temperature control accuracy, a temperature field distribution uniformity, a wind speed field uniformity, a system comprehensive energy consumption, and a control response time.
[0182] In one embodiment, the effect of the reinforcement learning strategy is evaluated, mainly from the following aspects: first, the temperature control effect is evaluated by the deviation between the target temperature and the actual temperature, and the temperature control accuracy of the system is measured; second, the performance of the strategy in energy consumption optimization is evaluated by calculating the total energy consumption of the fan and the cold source system; finally, the training effect and usability of the strategy network are judged by observing the convergence and stability of the strategy network in the training process. In the evaluation method, the reinforcement learning strategy is verified through simulation experiment before actual application, and the adaptability and stability of the strategy under different working conditions are tested by combining with the actual operation data after online deployment, so as to ensure the reliability and practicality of the strategy in real scene. The performance indicators of the strategy involved include target area temperature control accuracy, temperature field distribution uniformity, wind speed field uniformity, system comprehensive energy consumption and control response time. When the evaluation result of the strategy network is good, the strategy network is determined as the reinforcement learning strategy network.
[0183] S6. Based on the reinforcement learning strategy network, real-time optimization control of the local temperature of the roadway is realized.
[0184] S6. Based on the reinforcement learning strategy network, real-time optimization control of the local temperature of the roadway is realized.
[0185] S61. Based on the evaluation result, a prediction model considering the dynamic characteristics of the temperature field is constructed.
[0186] In one embodiment, on the basis of obtaining better results by evaluating the reinforcement learning strategy, a system prediction model considering the dynamic characteristics of the temperature field is constructed. The construction process of the system prediction model includes definition of system dynamic model, design of disturbance model, setting of prediction time domain and control time domain.
[0187] Specifically, for the definition of the system dynamic model, the main form is the state space equation, which is used to describe the variation law of the system state, and its specific form can be expressed as:
[0188]
[0189]
[0190] wherein, is an input variable, is a state variable, is an output variable, is a control variable, is a disturbance variable, , , , is a system parameter matrix, is time.
[0191] Further, the disturbance model is designed to describe the influence of the external environment on the system. The disturbance model satisfies the following expression:
[0192]
[0193] wherein, is a disturbance variable, is a disturbance function, is an external environment variable.
[0194] Further, the prediction time domain and the control time domain are set to make the system balance system performance, real-time performance and computational efficiency, so as to make the control decision more accurate, and to ensure the robustness and adaptability of the system in a complex dynamic environment, and to provide a reliable foundation prediction for the temperature control and energy consumption optimization of the roadway.
[0195] Specifically, the specific form of the prediction time domain is as follows:
[0196]
[0197] wherein, is a predicted state, is a state transition function, is a time change amount.
[0198] The specific form of the control time domain is as follows:
[0199]
[0200] wherein, is an optimal control input at the current time, is an objective function. is the number of input features, denotes the parameter taking the minimum value.
[0201] S62. Based on the prediction model, an optimization problem is constructed.
[0202] In one embodiment, the optimization problem is constructed by designing an objective function, defining a constraint condition, and optimizing variables and solving. The construction of the optimization problem is the core step of model predictive control, which aims to formalize the temperature control requirements and energy consumption constraints of the system into a mathematical model, and further process to generate an optimal control strategy.
[0203] Specifically, the purpose of designing the objective function is to determine the optimization direction and balance the temperature control accuracy and energy consumption, so as to achieve precise cooling while saving energy to the maximum extent. The objective function is in the following form:
[0204]
[0205] wherein, is the objective function value, is the predicted temperature, is the target temperature, is the control variable, is the temperature control deviation weight coefficient, is the energy consumption weight coefficient, is the time variation amount, is the total time variation amount.
[0206] Further, the purpose of defining the constraint condition is to ensure that the optimization result is within the device operation and physical model range, avoid unfeasible control scheme, and limit the variation range of temperature and air speed, to ensure the safety and reliability of the control strategy. The constraint condition is obtained according to the device operation limit, and the specific form is as follows:
[0207]
[0208]
[0209] wherein, , are the minimum and maximum values of the fan speed, respectively, , are the minimum and maximum values of the air deflector angle, respectively.
[0210] Further, the optimization variables include the fan speed and the air deflector angle, which are used to adjust the speed and direction of the system airflow, respectively. By setting the upper and lower limit ranges of the fan speed and the air deflector angle, the physical feasibility of the control strategy is ensured within the safe range of device operation. The optimization problem is designed based on the objective function and the constraint condition, and a quadratic programming or gradient optimization algorithm is used to solve it in real time to generate a dynamic optimal control strategy, to simultaneously meet the temperature control accuracy and energy minimization requirements of the system, and to ensure the rapid response capability of the control.
[0211] S63. Implementing a control strategy based on the optimization problem.
[0212] In one embodiment, based on the objective function and constraint conditions constructed based on the optimization problem, a specific application of rolling optimization, real-time solution and feedback adjustment mechanism is used to dynamically generate and execute the control strategy of fan speed and guide vane angle. At each time step, the control system takes the objective function (balance of temperature control accuracy and energy consumption) in S62 as the optimization core, combines the constraint conditions (device operating range and physical model limit), and adjusts the predicted temperature field and wind speed field distribution in real time. Only the current control quantity within the control time domain is implemented, and the remaining part will be recalculated in the next step of rolling optimization. The optimization problem defined in S62 is quickly solved by a quadratic programming or gradient optimization algorithm to generate a control strategy that meets the real-time requirements. After the implementation of the control quantity, the system bias is corrected using real-time sensor data based on the state transition formula and feedback error calculation mechanism, and the fan speed and guide vane angle are dynamically adjusted to form a closed-loop control. This process ensures the accuracy, stability and adaptability to complex environments of the control strategy.
[0213] S64. Based on the control strategy, the control effect verification is performed.
[0214] In one embodiment, based on the objective function and constraint conditions defined in S62, the actual effect of the optimization control strategy is comprehensively verified through control accuracy evaluation and performance index summary.
[0215] Specifically, the control accuracy is evaluated by calculating the temperature control error. The calculation model of the temperature control error satisfies the following expression:
[0216]
[0217] Wherein, is the mean square error, is the total number of data sampling points, is the target temperature, is the measured temperature, is the sampling point sequence.
[0218] Further, the overall performance of the system is evaluated by quantifying the control effect of energy efficiency. The calculation model of the energy efficiency satisfies the following expression:
[0219]
[0220] Wherein, is the energy efficiency, is the fan power, is the power of the cooling system, is the total time.
[0221] Further, according to the control effect, the control strategy is continuously optimized to achieve real-time optimization control of the local temperature in the roadway.
[0222] Please refer toFigure 2 , Figure 2 Figure 1 is a schematic diagram of the overall structure of a roadway local cooling control system based on deep reinforcement learning according to an embodiment of the present application. The system uses the roadway local cooling control method based on deep reinforcement learning. The system comprises an input device, a processor, an output device, and a memory. The memory is used to store a computer program, the computer program comprises program instructions, the processor is configured to invoke the program instructions, and the input device, the processor, the output device, and the memory are connected to each other. The main arrangement of the system is shown in Figure 2. Figure 3 Figure 3 The arrangement structure comprises a double-axial flow fan set 1, a guide vane set 2, a cold wall system 3, a monitoring point 4, and a working face 5. The cross-sectional size and point position of the working face are marked in the dashed box. Figure 3
[0223] In this embodiment, the input device comprises a data acquisition module.
[0224] Specifically, the sampling frequency of the data acquisition module is adjustable, and the data acquisition module has data anomaly detection and automatic calibration functions. The data acquisition module comprises a temperature sensor array and a wind speed sensor array. The temperature sensor array is arranged in a grid form at key positions in the roadway, and the wind speed sensor array is arranged at key cross sections of the airflow channel.
[0225] The processor comprises a model prediction module, a decision control module, and an execution control module.
[0226] Specifically, the model prediction module has online learning and model updating functions, and comprises a deep learning prediction unit and a reinforcement learning strategy unit, which are used to realize temperature field distribution prediction and control strategy generation. The decision control module has multi-objective optimization and real-time response capabilities, and comprises a model prediction control unit, which is used to optimize the control strategy and generate control instructions. The execution control module has fault diagnosis and automatic protection functions, and comprises a double-axial flow fan set 1, a guide vane set 2, and a cold wall system 3. The double-axial flow fan set 1 and the guide vane are arranged cooperatively to realize accurate control of the airflow. The cold wall system 3 is used to provide stable cold source output. The guide vane set 2 comprises independently adjustable guide vane units. The double-axial flow fan set 1 comprises a main fan for providing main airflow, an auxiliary fan for adjusting the airflow direction, a variable frequency control unit for realizing accurate adjustment of the fan speed, an angle adjustment mechanism for realizing angle adjustment within a range of 0° to 90°, and a position adjustment mechanism for adjusting the guide vane position according to the working face progress. During system operation, the airflow direction is as shown in Figure 3. Figure 4
[0227] The output device adopts a high-definition display screen for displaying the control results of the roadway local cooling.
[0228] The memory adopts a high-speed solid-state hard disk for storing data acquired by the input device, results displayed by the output device and results processed by the processor, and has the characteristics of fast read-write speed, large capacity and high reliability, and can meet the demand of large data storage.
[0229] Please refer to Figure 5 , Figure 5 The running flow chart of a roadway local cooling control system based on deep reinforcement learning according to an embodiment of the present application. The contents involved in the flow chart include system initialization and modeling, deep learning model development, model predictive control development, system running and feedback, system evaluation and maintenance, and reinforcement learning strategy development.
[0230] Please refer to Figure 6 , Figure 6 The system timing temperature field structure diagram according to an embodiment of the present application is shown, which respectively gives the timing temperature field distribution of 10 seconds, 20 seconds, 30 seconds, 1 minute, 2 minutes, 3 minutes, 5 minutes, 10 minutes, 30 minutes and 60 minutes. Figure 6 S represents seconds and Min represents minutes. The diagram reflects the change of the temperature field with time.
[0231] In the present embodiment, system running and feedback is the final execution and closed-loop optimization stage of the model predictive control system, aiming to ensure the efficient and stable operation of the system in a dynamic and complex environment through four core links of real-time data acquisition and processing, state prediction analysis, control decision execution and running effect analysis. It mainly includes the following steps:
[0232] Firstly, real-time data acquisition.
[0233] Specifically, real-time data acquisition is based on the sensor arrangement and sampling frequency configured in step S1, and through the temperature and humidity sensor and the wind speed sensor, the environmental data and the equipment running state (such as fan speed, deflector angle) are collected in real time, to ensure that the dynamic feedback of the data meets the real-time requirements of the system. At the same time, combined with the standardization method of data preprocessing in S2, abnormal values are detected and corrected, and interpolation or smoothing algorithm is used to remove noise, to ensure the effectiveness and integrity of the data.
[0234] Secondly, state prediction analysis.
[0235] Specifically, the state prediction analysis is based on the optimized prediction model of the temperature field and wind speed field constructed in step S3, and the future state of the system is predicted in multiple steps. The dynamic characteristics of the key area are extracted by combining trend analysis to identify potential risks in advance. At the same time, through the temperature control accuracy and energy consumption constraint conditions defined in step S62, the real-time state is evaluated to determine whether the target requirements are met. When the prediction or real-time data exceeds the safety threshold set by the constraint conditions in step S62, a warning message is generated to provide a basis for strategy adjustment in the next control decision execution.
[0236] Third, control decision execution.
[0237] Specifically, the control decision execution is based on the objective function and constraint conditions defined in step S62, combined with the rolling optimization and real-time solving mechanism in step S63 control strategy implementation, to generate dynamic adjustment suggestions for fan speed and guide vane angle through reinforcement learning strategy, and to calculate the optimal control instructions using model predictive control. These instructions are transmitted to the actuators to complete real-time operations, while the feedback adjustment mechanism in step S63 monitors the execution effect, corrects the deviation and dynamically updates the system state to ensure the efficiency and accuracy of the control strategy in complex environments.
[0238] Fourth, running effect analysis. Running effect analysis calculates the temperature control error (such as mean square error) and energy efficiency index (such as system energy consumption efficiency) through the objective function and constraint conditions defined in step S62 to verify the temperature control accuracy and energy optimization effect of the system. At the same time, combined with the rolling optimization mechanism and feedback adjustment process in step S63 control strategy implementation, the response time and recovery ability of the system when disturbed are analyzed to evaluate the stability and accuracy of the control strategy, providing data support for further optimizing control logic and improving system performance.
[0239] In this embodiment, system evaluation and maintenance is a key step to ensure the long-term stable operation of the model predictive control system, covering performance evaluation, model updating and optimization, and system maintenance. Through comprehensive analysis and optimization of system performance, combined with the state inspection of equipment and sensors, the long-term reliability and economy of the system in complex environments are ensured. Specifically, the following steps are included:
[0240] First, performance evaluation.
[0241] Specifically, performance evaluation combines the parameter settings of the system physical model in step S1 to define evaluation criteria for wind speed and temperature distribution, ensuring that the temperature control system design is consistent with actual requirements. Simultaneously, combined with the data processing and feature extraction results in step S3, a comprehensive analysis of temperature control effectiveness and energy efficiency is conducted. Temperature control effectiveness is verified by calculating the deviation between the target temperature and the actual temperature (e.g., mean square error). Energy efficiency is evaluated by combining the objective function and the system energy efficiency ratio calculation formula in step S62 to assess fan and cooling system energy consumption. Finally, based on the comprehensive analysis of system response stability and energy consumption costs during the control effectiveness verification in step S64, performance indicators are optimized to provide data support for the reliability and economic efficiency of the overall system.
[0242] The second step is model updating and optimization.
[0243] Specifically, the model update and optimization combines the deep learning model development in step S3 and the reinforcement learning strategy development in step S5. The deep learning model is retrained through historical data and real-time data to optimize the prediction accuracy of the temperature field and wind speed field. The experience replay pool and policy gradient method in reinforcement learning are used to dynamically adjust the reward function to enhance the adaptability and convergence speed of the policy network. At the same time, based on the rolling optimization mechanism and objective function setting in the model predictive control development in step S6, the prediction time domain, control time domain and parameters of the MPC are dynamically adjusted, and the weight coefficients in the objective function are optimized to make the control strategy better adapt to complex dynamic environments. Through the above comprehensive optimization, the prediction accuracy and control effect of the system are continuously improved, and its operational efficiency and adaptability are enhanced.
[0244] The third step is system maintenance.
[0245] Specifically, we regularly check the operating status of equipment such as fans, air deflectors, and cooling systems to ensure stable operation within safe ranges. We also calibrate temperature and wind speed sensors to ensure accurate and consistent data collection. We also incorporate fault diagnosis technology to analyze operating data and historical records to identify potential issues (such as equipment anomalies or sensor failures). We then develop a maintenance plan, including equipment cleaning, component replacement, and sensor calibration, and promptly implement preventive maintenance and repair measures to ensure long-term, reliable system operation.
[0246] In summary, the underground roadway local cooling control method based on deep learning and reinforcement learning provided by the present application provides a solid foundation for accurate prediction and control of the roadway temperature by collecting historical data of the temperature field, wind speed field, humidity and energy consumption in the roadway; the historical data is preprocessed to ensure the standardization and consistency of the data, thereby improving the prediction accuracy of the model; the deep learning method is used to construct an optimized prediction model of the temperature field, which efficiently captures the complex rules of environmental changes; the optimized prediction model is used to design a reinforcement learning control environment, providing a simulation and verification platform for the development of intelligent control strategies; the reinforcement learning strategy network is constructed to respond to changes in the roadway environment in real time, automatically adjust the cooling strategy, and achieve accurate and efficient control of the local temperature of the roadway.
[0247] The underground roadway local cooling control system based on deep learning and reinforcement learning provided by the present application is suitable for precise cooling requirements in high-temperature environments such as deep mine working faces and tunnel construction. The system realizes temperature regulation in complex environments by cooperatively arranging double-axial flow fans, air deflectors and cold wall systems. The system relies on temperature sensor arrays for real-time environmental data acquisition, uses deep learning technology to extract environmental features and predict temperature field and wind speed field changes, introduces physical constraints to ensure that the prediction conforms to the laws of thermodynamics, combines reinforcement learning algorithms to optimize control strategies, and realizes dynamic adjustment of fan speed, air deflector angle and cold source output through model predictive control. The combination design of double-axial flow fans and adjustable air deflectors ensures accurate air guidance, and the cold wall system provides stable cold source support. The system layout is flexible and can be adjusted according to the mining working face or construction progress, and has fault warning and self-adaptive adjustment capabilities. Actual application verification shows that the present application effectively improves the temperature control accuracy, reduces the system energy consumption, shortens the control response time, and provides a reliable solution for underground high-temperature treatment.
[0248] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for part or all of the technical features; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and they should be covered in the scope of the claims and description of the present application.
Claims
1. A tunnel local cooling control method based on deep reinforcement learning, characterized in that: The method comprises the following steps: Obtain historical data on temperature field, wind speed field, humidity and energy consumption in the tunnel environment; Preprocessing the historical data to obtain standardized data; Based on the standardized data, a deep learning temperature field optimization prediction model is constructed, including: Based on the standardized data, a deep neural network structure is constructed; Based on the deep neural network structure, design the input layer of temperature field and wind speed field data; Based on the input layer, design a feature extraction layer and a time series prediction layer; Based on the feature extraction layer and the time series prediction layer, designing a loss function with physical constraints; Determining model training parameters and optimization methods based on the loss function with physical constraints; By using the model training parameters and the optimization method, an optimization prediction model of deep learning temperature field is constructed; Based on the optimization prediction model, a reinforcement learning control environment is designed, including: Based on the optimization prediction model, an environmental state space, an action space, and a reward function are designed. The environmental state space includes the temperature field distribution, the wind speed field distribution, the fan speed, and the wind deflector angle. The action space includes the adjustment range of the fan speed and the adjustment range of the wind deflector angle. The reward function includes the target temperature and current energy consumption. Based on the reinforcement learning control environment, a reinforcement learning strategy network is constructed, including: Based on the reinforcement learning control environment, design the network structure and loss function of deep reinforcement learning; Based on the network structure and loss function of deep reinforcement learning, a reinforcement learning strategy network is constructed, including: Based on the network structure and loss function of the deep reinforcement learning, the experience replay pool is used to perform policy training iterations to obtain iterative results; Based on the iterative results, the strategy performance indicators are evaluated and a reinforcement learning strategy network is constructed. The strategy performance indicators include the temperature control accuracy of the target area, the uniformity of the temperature field distribution, the uniformity of the wind speed field, the comprehensive energy consumption of the system and the control response time; Based on the reinforcement learning strategy network, real-time optimization control of the local temperature of the roadway is achieved, including: Based on the reinforcement learning strategy network, a prediction model considering the dynamic characteristics of the temperature field is constructed, wherein the prediction model sets a prediction time domain and a control time domain; Based on the prediction model, an optimization objective function including temperature control accuracy and energy efficiency is designed; Generate a dynamic optimal control strategy through the optimization objective function and constraint conditions; According to the optimal control strategy, real-time optimization control of the local temperature of the tunnel is achieved.
2. A tunnel local cooling control method based on deep reinforcement learning according to claim 1, characterized in that: The preprocessing of the historical data to obtain standardized data includes: Performing data cleaning and outlier processing on the historical data to obtain optimized data; The optimized data is standardized to obtain standardized data.
3. A tunnel local cooling control system based on deep reinforcement learning, wherein the system uses a tunnel local cooling control method based on deep reinforcement learning according to any one of claims 1 to 2, characterized in that: The system includes an input device, a processor, an output device, and a memory; The input device includes a data acquisition module; The processor includes a model prediction module, a decision control module and an execution control module; The input device, the processor, the output device and the memory are connected to each other; The memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions.
4. A tunnel local cooling control system based on deep reinforcement learning according to claim 3, characterized in that: The data acquisition module has an adjustable sampling frequency and is equipped with data anomaly detection and automatic calibration functions. It includes a temperature sensor array and a wind speed sensor array. The temperature sensor array is arranged in a grid-like manner at key locations in the laneway, and the wind speed sensor array is arranged at key sections of the airflow channel. The model prediction module has online learning and model update functions, including a deep learning prediction unit and a reinforcement learning strategy unit, which are used to predict temperature field distribution and generate control strategies; The decision control module has multi-objective optimization and real-time response capabilities, including a model prediction control unit for optimizing control strategies and generating control instructions; The execution control module has fault diagnosis and automatic protection functions, including a dual-axial fan unit, an adjustable air guide plate group and a cold wall system. The dual-axial fan unit and the air guide plate are arranged in a coordinated manner. The air guide plate group includes independently adjustable air guide plate units. The cold wall system provides stable cold source output.
5. The tunnel local cooling control system based on deep reinforcement learning according to claim 4 is characterized in that: The dual axial flow fan unit includes a main blower, an auxiliary blower, a frequency conversion control unit, an angle adjustment mechanism and a position adjustment mechanism; The main blower is used to provide the main airflow; The auxiliary fan is used to adjust the direction of airflow; The frequency conversion control unit is used to adjust the fan speed; The angle adjustment mechanism is used to adjust the angle within the range of 0° to 90°; The position adjustment mechanism is used to adjust the position of the air guide plate.
Citation Information
Patent Citations
Air conditioner energy-saving control method and system based on reinforcement learning and digital twinborn model
CN118031385A
Closed environment energy-saving control method and device based on deep reinforcement learning, equipment and storage medium
CN118998926A