Concrete cooling water pipe dynamic flow adjusting method and device based on reinforcement learning

By using a pre-embedded sensor array and reinforcement learning model, combined with three-dimensional temperature field reconstruction and hybrid neural networks, the problems of temperature monitoring blind spots and pipeline blockage in concrete temperature control were solved, achieving high-precision and adaptive cooling water pipe flow regulation, and improving system stability and resource utilization efficiency.

CN121501047APending Publication Date: 2026-02-10HUBEI ENG CONSTR GRP THIRD CONSTR ENG CO LTD

Patent Information

Application Number
CN202511555836.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies cannot reconstruct the three-dimensional temperature field inside concrete in real time, lack intelligent learning capabilities, and cannot dynamically optimize the flow regulation of cooling water pipes. This results in blind spots in temperature monitoring, making it difficult to achieve high-precision, adaptive temperature control, and lacks proactive prevention of pipe blockage.

Method used

By employing a pre-embedded temperature sensor array and flow meter, combined with a reinforcement learning model, and through three-dimensional temperature field reconstruction, hybrid neural network feature extraction, and strategy optimization, dynamic adjustment of cooling water valves is achieved. Physical constraints and real-time monitoring are also introduced to automatically trigger cleaning measures.

Benefits of technology

It achieves high-resolution temperature distribution identification, dynamically optimizes cooling strategies, improves flow control accuracy and response speed, reduces manual inspection, and ensures system stability and cooling resource efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121501047A_ABST
    Figure CN121501047A_ABST
Patent Text Reader

Abstract

The invention provides a concrete cooling water pipe dynamic flow regulation method and device based on reinforcement learning, and relates to the technical field of concrete cooling water pipe dynamic flow regulation. The method comprises the following steps: reconstructing a three-dimensional temperature field based on Kriging, fusing the three-dimensional temperature field with material and environmental parameters into a state vector, inputting a near-end strategy optimized reinforcement learning model, outputting each valve opening vector by a strategy network, evaluating a state value by a value network, and performing physical constraint verification before issuing, an opening instruction realizes closed-loop flow distribution through coarse adjustment of the main valve and fine adjustment of the piezoelectric micro valve, meanwhile, flow attenuation and pressure difference are monitored in real time, and ultrasonic waves and mechanical scraper cleaning are automatically triggered to ensure long-term stable operation of the system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of dynamic flow regulation of concrete cooling water pipes, in particular to a dynamic flow regulation method and device for concrete cooling water pipes based on reinforcement learning. BACKGROUND

[0002] In the construction and maintenance process of large concrete structures, the poor thermal conductivity of concrete leads to significant temperature rise due to internal heat accumulation, while the external heat dissipation is fast, thereby forming a large temperature gradient and tensile stress inside the structure, which is the main source of inducing temperature cracks, affecting the durability and safety of the structure. Therefore, how to achieve fine, timely and energy-efficient control of temperature distribution in the concrete during the post-pouring and maintenance period has become a key problem in engineering practice. The traditional method mainly regulates the internal water flow to remove heat by arranging cooling water pipes, but the intelligent level of the regulation strategy directly determines the temperature control effect and resource efficiency.

[0003] The existing technology mainly tries to solve the concrete temperature control problem through two types of means: one type is based on discrete point temperature measurement and empirical control, using several embedded temperature points or surface measurement points to adjust the opening of a single path valve through empirical rules or traditional controllers, or to adjust the cooling water distribution regularly through artificial experience; the other type is based on physical modeling and offline simulation, using heat conduction equations and material parameters to design and parameterize the scheme before construction, in order to implement cooling according to the pre-set strategy during construction. In addition, the current practice in pipeline maintenance relies on regular inspection or simple trigger cleaning, lacking online intelligent identification and automatic cleaning means for pipeline flow degradation.

[0004] At the data perception level, the existing method cannot reconstruct the three-dimensional temperature field inside the concrete in real time, resulting in blind spots in temperature monitoring and difficulty in accurately identifying high temperature areas. Secondly, at the decision-making level, the existing technology lacks intelligent learning ability and cannot dynamically optimize the flow regulation strategy based on multi-source data, especially cannot embed engineering physical constraints such as temperature gradient limits into the decision-making process, which is prone to disconnect the control instructions from the actual needs. At the execution level, the existing valve control is extensive and cannot achieve regional and pipe group level flow coordination and fine adjustment, and lacks an active prevention mechanism for pipe blockage. These deficiencies make it difficult for the existing technology to meet the demand for high-precision and adaptive temperature control of mass concrete.

[0005] The above information disclosed in the background section is only used to enhance the understanding of the background of the present disclosure, therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0006] The purpose of this invention is to provide a method and apparatus for dynamic flow regulation of concrete cooling water pipes based on reinforcement learning, so as to solve the problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides the following technical solution: A method for dynamic flow regulation of concrete cooling water pipes based on reinforcement learning, comprising the following steps: Step 1: The temperature, cooling water flow and differential pressure data inside the concrete are collected by an array of temperature sensors embedded in the concrete and flow meters and differential pressure sensors installed on the cooling water pipe branches. Environmental parameters and material property parameters are also collected. The collected data are filtered and noise reduced. The three-dimensional temperature field inside the concrete is reconstructed based on the spatial interpolation algorithm. Step 2: After fusing the three-dimensional temperature field, cooling water flow rate, pressure difference data, material properties and environmental parameters into a state vector, it is input into a reinforcement learning model based on the near-end policy optimization algorithm. The model infers the optimal opening command of each cooling water valve through its policy network and evaluates the state value through its value network. Before outputting the optimal opening command, physical constraints are introduced to verify the feasibility. Step 3: The optimized opening command that has passed the feasibility verification is sent to the actuator. The actuator includes an electric proportional valve for coarse adjustment of flow rate at the regional level and a piezoelectric ceramic micro valve for fine adjustment of flow rate at the pipe group level. The valve opening is adjusted according to the optimized opening command to realize the dynamic distribution of cooling water flow. Step 4: Monitor the flow attenuation rate and pressure difference of each branch in real time. If the preset blockage conditions are met, the anti-blockage mechanism will be automatically triggered to clean the branch. The anti-blockage mechanism includes an ultrasonic cleaner and a mechanical rotating scraper.

[0008] Furthermore, multiple layers of temperature sensors are pre-embedded along the height direction in four key areas inside the concrete to form a three-dimensional monitoring network. The key areas include the hydration heat core area, the heat dissipation sensitive area, the stress danger area, and the boundary effect area. The arrangement density of temperature sensors in the hydration heat core area and the stress danger area is higher than that in the heat dissipation sensitive area and the boundary effect area. Temperature sensors are used to collect temperature data. Differential pressure sensors and flow meters are installed at the branch inlets and capillary junctions of each cooling water pipe to monitor pipeline differential pressure and cooling water flow. The collected raw temperature data is processed by a sliding window filtering algorithm for noise reduction, and a Kriging space interpolation algorithm is used to reconstruct the three-dimensional temperature field distribution.

[0009] Furthermore, the reinforcement learning model employs a hybrid neural network architecture for feature extraction. This architecture uses a three-dimensional convolutional neural network to extract spatial features from a three-dimensional temperature field. The convolutional kernel size of the three-dimensional convolutional neural network is 3×3×3, and the stride is 1. At the same time, a long short-term memory network is used to process the temporal features of flow rate, pressure difference, and temperature data, and a fully connected network is used to process environmental parameters and material property parameters. Finally, all features are fused and input into the policy network and the value network. Input the state vector of the reinforcement learning model The data is composed of the following multidimensional data fusion: the current three-dimensional temperature field obtained from reconstruction, the current flow rate value of each cooling water pipe branch and the current pressure difference value of the pipe inlet and outlet, the current material characteristic parameters and the current environmental parameters. The material characteristic parameters include the cement type coefficient and the decay factor that changes with age calculated from the time of concrete pouring. The environmental parameters include ambient temperature, relative humidity and wind speed, where t is the index of the current time.

[0010] Furthermore, the output of the policy network in the reinforcement learning decision model is an optimized opening instruction, which is generated by a truncated Gaussian probability distribution, and its formula is: ; in, The optimized opening command for the i-th valve, generated at the current time t, is normalized to... The range corresponds to the valve opening degree of 0%-100%. and It is the policy network in the input state vector The output is the Gaussian distribution mean and standard deviation corresponding to the i-th valve. The truncation function ensures that the output optimized opening command is within the valid range. Inside, For a standard normal distribution The random noise sampled in the middle, where i is the index of the cooling water valve.

[0011] Furthermore, the value network is implemented using a deep neural network, and its role is to evaluate the state vector. Given the expected long-term return achievable by the current strategy, the training objective of this network is to minimize the temporal difference error, and its loss function is... for: ; in, For the state vector The immediate reward obtained after executing the optimized opening command. This is a future reward discount factor, ranging from 0 to 1, used to measure the value of future rewards at the current moment. This is an estimate of the value function of the value network for the next time step, representing the total expected return that can be obtained starting from the next time step. Let be the state vector at the current time. State value function estimate; To suppress excessive internal temperature gradients in concrete, the reinforcement learning model uses an improved advantage function for policy optimization. The formula for this function in generalized advantage estimation is as follows: ; in, The advantage estimate at the current moment is used to measure the advantage in the state vector. The extent to which the optimized opening command is advantageous compared to the average level is considered. These are gradient penalty weights, used to control the proportion of the temperature gradient penalty term in the total reward. This is the temperature gradient modulus inside the concrete at the current moment, calculated based on the reconstructed three-dimensional temperature field. The maximum safe temperature gradient threshold; This is a temperature gradient penalty term, which only applies if the temperature gradient... Exceeding the maximum threshold The advantage estimate is negatively adjusted at that time; In the formula, To estimate the dominance value for the standard generalized dominance, , The truncation length represents the maximum number of steps that take into account future time-series difference errors when calculating the generalized dominance estimate. This is the step offset index, used in the summation formula to represent the number of future steps relative to the current time t. The generalized dominance estimation coefficient, ranging from 0 to 1, is used to balance the bias and variance of the dominance estimation. In the formula... , The time-series difference error at the current moment represents the difference between the current estimate and the future estimate; The temperature gradient magnitude Based on the reconstructed three-dimensional temperature field By calculating its spatial partial derivative, we obtain: .

[0012] Furthermore, the optimized opening command, after conversion, is output to an electric proportional valve for coarse flow adjustment at the regional level and a piezoelectric ceramic micro-valve for fine flow adjustment at the pipe group level. The electric proportional valve receives the voltage signal after range conversion, and its opening... With traffic The relationship, after nonlinear calibration, is described by the following equation: ; in, For flow coefficient, To characterize the nonlinear function of the valve characteristic curve, , This is the steepness coefficient. The pressure difference between the upstream and downstream sides of the valve. Specific gravity of the fluid; The piezoelectric ceramic microvalve receives an amplified drive voltage signal and outputs dynamic flow. With drive voltage signal The relationship is described by the following differential equation: ; in, The valve's electromechanical time constant. Piezoelectric gain represents the change in flow rate that can be generated per unit driving voltage. The pressure disturbance coefficient characterizes the upstream pressure difference. The degree of impact on output flow.

[0013] Furthermore, before outputting the optimized opening command, the logic for introducing physical constraints to verify feasibility is as follows: Physical constraint verification includes thermal stress constraints, hydraulic balance constraints, and mechanical protection constraints. Thermal stress constraints ensure that the maximum temperature difference inside the concrete does not exceed the material's allowable threshold, that is, ensure the maximum temperature difference between any two points inside the concrete. ,in This refers to the allowable temperature difference threshold for the material. Hydraulic balance constraints ensure the flow rate of the main pipeline The sum of the flow rates of each cooling water pipe branch The deviation is within 5%, that is... ; Mechanical protection constraints ensure that the rate of change of the optimized opening command does not exceed the maximum response capability of the actuator, and that the optimized opening command is within... Within the range; When the q-th cooling water pipe branch is detected to simultaneously meet the following preset blockage conditions, and the duration exceeds the preset threshold, When the ultrasonic cleaner and mechanical rotating scraper of the cooling water pipe branch are automatically triggered, the preset blockage conditions include: 1) the flow rate attenuation rate of the qth cooling water pipe branch exceeds the preset attenuation threshold; 2) the inlet and outlet pressure difference of the cooling water pipe branch exceeds the preset safety threshold. When the flow rate of the q-th cooling water pipe branch returns to its baseline flow rate If the above conditions are met, the cleaning is considered successful; otherwise, a maintenance warning requiring manual intervention will be issued. q is the index of the cooling water pipe branch.

[0014] The present invention also provides a reinforcement learning-based dynamic flow regulation device for concrete cooling water pipes, the device being used to execute the above-described reinforcement learning-based dynamic flow regulation method for concrete cooling water pipes, comprising: The data acquisition module is used to collect data on the internal temperature of the concrete, the flow rate of the cooling water and the differential pressure by an array of temperature sensors embedded in the concrete and flow meters and differential pressure sensors installed on the cooling water pipe branches. It also collects environmental parameters and material property parameters, filters and reduces noise on the collected data, and reconstructs the three-dimensional temperature field inside the concrete based on a spatial interpolation algorithm. The opening optimization module is used to fuse the three-dimensional temperature field, cooling water flow rate, pressure difference data, material properties and environmental parameters into a state vector, and then input it into a reinforcement learning model based on the near-end policy optimization algorithm. The model infers the optimal opening command of each cooling water valve through its policy network and evaluates the state value through its value network. Before outputting the optimal opening command, physical constraints are introduced to verify the feasibility. The valve opening adjustment module is used to send the optimized opening command that has passed the feasibility verification to the actuator. The actuator includes an electric proportional valve for coarse adjustment of flow rate at the regional level and a piezoelectric ceramic micro valve for fine adjustment of flow rate at the pipe group level. The valve opening is adjusted according to the optimized opening command to realize the dynamic distribution of cooling water flow. The monitoring and early warning module is used to monitor the flow attenuation rate and pressure difference changes of each branch in real time. If the preset blockage conditions are met, the anti-blockage mechanism is automatically triggered to clean the branch. The anti-blockage mechanism includes an ultrasonic cleaner and a mechanical rotating scraper.

[0015] Compared with the prior art, the beneficial effects of the present invention are: This invention uses a pre-embedded temperature sensor array and pipeline sensors, along with spatial interpolation to reconstruct a three-dimensional temperature field. This allows the control system to acquire high-resolution internal temperature distribution information, eliminating reliance on discrete points or extrapolation estimates. This enables accurate identification of high-cooling-requirement areas and high-temperature gradient zones within the concrete, providing reliable spatial evidence for subsequent decision-making. It eliminates misjudgments or omissions caused by sparse monitoring in traditional experience-based control. This step also suppresses measurement noise through data preprocessing and filtering, and combined with smooth reconstruction results, reduces the risk of false triggering caused by weak local sensors. This provides a stable and reliable state input for strategy training and online decision-making. This invention, based on a reinforcement learning model using a near-end policy optimization algorithm, achieves dynamic optimization of cooling strategies by integrating multi-dimensional data such as temperature field, flow rate, pressure difference, material properties, and environmental parameters. In particular, by embedding a temperature gradient penalty term in the design of the reward function and performing physical constraint verification before outputting commands, the system can both autonomously learn the optimal control strategy and ensure that the decision results meet engineering safety requirements. This fundamentally solves the problem of control commands being out of touch with actual needs caused by the lack of intelligent learning capabilities in traditional methods. This invention employs a hierarchical control architecture combining electric proportional valves and piezoelectric ceramic micro-valves, achieving coordinated coordination between regional-level coarse flow adjustment and pipe-group-level fine flow adjustment. This significantly improves the accuracy and response speed of flow control. Simultaneously, real-time flow attenuation and differential pressure monitoring, linked with ultrasonic and mechanical scraper cleaning measures, enable early intelligent identification and automatic handling of pipe blockages. This reduces the frequency of manual inspections and emergency maintenance while ensuring the long-term stability and resilience of the system. In summary, this invention achieves comprehensive beneficial effects in key aspects such as accurate risk identification, online adaptive control, physical safety assurance, and automated maintenance. It can effectively reduce structural risks caused by concrete temperature differences, improve cooling resource utilization efficiency, and significantly enhance system operational reliability. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the overall method flow of the present invention; Figure 2 This is a scatter plot of the estimated advantages of this invention. Figure 3 This is a line graph showing the estimated advantage value of the invention – the temperature gradient magnitude. Figure 4 This is a curve showing the fit between the standard generalized advantage estimate and the advantage estimate in this invention. Figure 5 This is a probability diagram of the advantage estimate of the present invention; Figure 6 This is a flowchart of the overall device structure of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0018] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0019] Example: Please see Figures 1-5 The present invention provides a technical solution: A method for dynamic flow regulation of concrete cooling water pipes based on reinforcement learning, comprising the following steps: Step 1: The temperature, cooling water flow and differential pressure data inside the concrete are collected by an array of temperature sensors embedded in the concrete and flow meters and differential pressure sensors installed on the cooling water pipe branches. Environmental parameters and material property parameters are also collected. The collected data are filtered and noise reduced. The three-dimensional temperature field inside the concrete is reconstructed based on the spatial interpolation algorithm. The temperature field of concrete is the core state of the controlled object. The flow rate and pressure difference of the water pipe network reflect the real-time working status and efficiency of the cooling system. Environmental factors such as air temperature, humidity, and wind speed directly affect the heat dissipation rate of concrete and are disturbances that must be considered in the control model. The cement type and age of the material determine the rate and amount of heat release from hydration and are prior knowledge for building an accurate model. Temperature sensors can only measure the temperature at discrete points. However, the temperature inside concrete is a continuous field. By inferring and generating the three-dimensional temperature distribution throughout the entire concrete volume through algorithms, it becomes the basis for accurately allocating cooling flow on demand, because decision-making requires knowing the temperature of every area (not just the sensor point). Kriging spatial interpolation is not just simple mathematical interpolation; it is a geostatistical method that can consider the spatial correlation of data (i.e., the temperatures of adjacent points are not independent), thus providing optimal, unbiased interpolation estimates. This reflects the temperature distribution inside concrete more accurately than linear or spline interpolation. Multiple layers of temperature sensors are pre-embedded along the height direction in four key areas inside the concrete to form a three-dimensional monitoring network. The key areas include the hydration heat core area, the heat dissipation sensitive area, the stress danger area, and the boundary effect area. The arrangement density of temperature sensors in the hydration heat core area and the stress danger area is higher than that in the heat dissipation sensitive area and the boundary effect area. The hydration heat core area is usually the thickest and worst-dissipated geometric center of the structure, and is the area where heat accumulates and the temperature rises the most. It must be closely monitored. The heat dissipation sensitive area is the area near the surface, where the temperature gradient is the largest. It is the main channel for heat loss and is also the area that is prone to tensile stress due to excessive heat dissipation. The stress danger area is the structural abrupt point where stress is prone to concentration and is extremely sensitive to temperature changes. Even a small temperature difference can cause cracks, requiring special monitoring. The boundary effect area is greatly affected by the formwork and curing conditions, and its temperature behavior is different from that of the interior. The three-dimensional monitoring network emphasizes that the layout is three-dimensional, not two-dimensional. The temperature field of concrete changes in three-dimensional space. Only a three-dimensional layout can truly capture its distribution and provide data support for three-dimensional reconstruction. Deploying more sensing resources (sensors) in areas with the most drastic temperature changes and the most critical to structural safety (the core area heats up the fastest and the stress danger zone is the most sensitive) can obtain higher resolution details. In areas with relatively gentle changes or less impact, they can be deployed sparsely to save costs. This reflects the economy and intelligence of the technical solution. Temperature sensors are used to collect temperature data. Differential pressure sensors and flow meters are installed at the branch inlets and capillary junctions of each cooling water pipe to monitor pipeline differential pressure and cooling water flow. The collected raw temperature data is processed for noise reduction using a sliding window filtering algorithm, and the three-dimensional temperature field distribution is reconstructed using a Kriging space interpolation algorithm. The temperature sensors are divided into four intervals in the vertical direction according to their depth from the top surface: 0.1m, 0.8m, 1.25m, and 2.0m, corresponding to the four key regions. Sliding window filtering is a real-time digital filtering technique that performs a weighted average (such as exponentially decaying weights) on continuously arriving data within a fixed-length window. It can effectively smooth out random impulse noise while ensuring good real-time performance, and will not produce excessive lag like simple global averaging. Sliding window filtering is a time-domain smoothing process that ensures the reliability of data in the time dimension, while Kriging spatial interpolation is a spatial estimation method that ensures the integrity of the state in the spatial dimension. The combination of the two transforms the scattered, noisy, and low-dimensional raw data collected from the field into a clean, continuous, high-dimensional "digital twin" that can truly reflect the internal working conditions of concrete—a three-dimensional temperature field. This "digital twin" is the objective world model upon which the subsequent reinforcement learning intelligent decision-making "eyes" and "brain" rely for thinking. The reconstructed three-dimensional temperature field is one of the core inputs of the reinforcement learning algorithm. The algorithm makes a globally optimal flow allocation decision based on this complete temperature map, rather than a few isolated points, thereby fundamentally solving the drawback of the traditional "one-size-fits-all" approach and achieving precise temperature control with small spatial resolution. Installing flow and differential pressure sensors at key nodes such as branch inlets and capillary junctions means that the system can achieve zoned metering and fault isolation. When the system detects an abnormal flow in a branch, it can immediately locate which control area is causing the problem and, combined with the temperature data of that area, quickly diagnose whether it is a problem with the control command or a pipe blockage. Step 2: After fusing the three-dimensional temperature field, cooling water flow rate, pressure difference data, material properties and environmental parameters into a state vector, it is input into a reinforcement learning model based on the near-end policy optimization algorithm. The model infers the optimal opening command of each cooling water valve through its policy network and evaluates the state value through its value network. Before outputting the optimal opening command, physical constraints are introduced to verify the feasibility. The three-dimensional temperature field is essentially a spatial volume data. The 3x3x3 convolution kernel of 3D CNN is like a "smart scanner" that can slide in three-dimensional space, automatically learn and extract local temperature distribution patterns, such as identifying where the core area of ​​the "hot spot" that is heating up rapidly and where the edge area with a gentle temperature change is. The heat of hydration of concrete is a dynamic process. Data such as temperature and flow rate are time series. LSTM networks remember historical information and can understand trends, such as identifying whether the current temperature is in an accelerated rising phase, a peak plateau phase, or a slow falling phase. Fully connected networks embed scattered, non-spatial scalar parameters such as cement type, age, and ambient temperature and humidity into a high-dimensional feature space, enabling them to be effectively integrated with spatial and temporal features. The kernel size of 3×3×3 is a key and specific structural parameter that defines the window size for feature scanning in three-dimensional space. The stride of 1 determines the distance the kernel moves in each step. A stride of 1 means performing dense, high-resolution feature extraction. The core function of 3D CNN is to automatically extract the spatial distribution features of the internal temperature field of concrete. For example, a 3×3×3 convolution kernel can learn to identify spatial patterns such as local high-temperature regions (hydration heat core region) and steep temperature gradient zones (heat dissipation sensitive areas). The hidden layer dimension of the LSTM network is 128-dimensional. This is a specific network structure parameter that defines the length of the internal state vector of the LSTM and determines its ability to memorize and process temporal information. The fully connected network takes environmental parameters (temperature, humidity, wind speed) and material property parameters (cement type coefficient, age decay factor) as input and outputs a feature dimension of 64-dimensional. This is also a specific network structure parameter that encodes these scalar parameters into a 64-dimensional feature vector. The fully connected network fuses and elevates various types of scalar parameters into a high-dimensional feature space, enabling it to be effectively concatenated and subsequently processed with spatial features from 3D CNNs and temporal features from LSTMs. The features extracted by these three sub-networks (spatial features, temporal features, and scalar features) will be concatenated together to form a comprehensive feature vector, which will be input into the subsequent policy network and value network for final decision-making. The reinforcement learning model uses a hybrid neural network architecture for feature extraction. This architecture uses a three-dimensional convolutional neural network to extract spatial features from a three-dimensional temperature field. The convolutional kernel size of the three-dimensional convolutional neural network is 3×3×3, and the stride is 1. At the same time, a long short-term memory network is used to process the temporal features of flow rate, pressure difference and temperature data, and a fully connected network is used to process environmental parameters and material property parameters. Finally, all features are fused and input into the policy network and the value network. Input the state vector of the reinforcement learning model It is composed of the following multi-dimensional data fusion: the current three-dimensional temperature field obtained by reconstruction, the current flow rate value of each cooling water pipe branch and the current pressure difference value of the pipe inlet and outlet, the current material characteristic parameters and the current environmental parameters. The material characteristic parameters include the cement type coefficient and the decay factor that changes with age calculated from the time of concrete pouring. The environmental parameters include ambient temperature, relative humidity and wind speed, where t is the index of the current time. Cement type coefficient It is a dimensionless correction factor used to quantify the differences in the heat release rate and total heat release of hydration reactions among different types and strength grades of cement. It uses the heat release characteristics of a certain benchmark cement (e.g., the most commonly used ordinary Portland cement, P.O42.5 grade) as a reference. Defined as 1.0), other cements are compared to obtain corresponding coefficients; This coefficient is obtained through standardized laboratory heat of hydration tests, specifically following these steps: Determine a reference cement (e.g., P.O42.5) with its coefficient... Prepare the cement samples to be tested, such as P.II52.5 and medium-heat cement. For both the reference cement and the cement to be tested, conduct tests using an isothermal calorimeter according to the national standard "Method for Determination of Heat of Hydration of Cement" or an equivalent international standard (such as ASTM C1702). The test conditions are at a standard temperature, such as 20°C ± 0.1°C. Measure the cumulative heat release of hydration of the cement on the 3rd and 7th days after pouring. These two age periods are chosen because they cover the critical period of heat release during hydration. The calculation formula is: ; in, and It represents the cumulative heat of hydration of the cement under test on days 3 and 7. and This refers to the cumulative heat of hydration of the benchmark cement on days 3 and 7; through the above calculations, a specific... Values, for example, ranging from 0.8 to 1.2, indicate that the heat release of some low-heat cements is only 80% of the benchmark, while the heat release of some early-strength cements reaches 120% of the benchmark. Age decay factor It is a dimensionless time function between 0 and 1, used to describe the self-pouring time of concrete. From age onwards, its hydration heat release capacity increases with age. The pattern of increasing and gradually decreasing (in days) occurs when... hour, Theoretically, it's 0 (no heat has been released yet); as time goes on, Gradually increases and approaches 1 (heat release is basically complete); This factor is obtained by fitting a cumulative hydration heat curve. The specific steps are as follows: Using the aforementioned isothermal calorimeter, a long-term tracking test is conducted on the selected cement (at a specific water-cement ratio). The cumulative hydration heat release curve is measured from several hours after pouring to 28 days or longer. This will yield a set of data points. ,in It is the age of maturity The total heat released at that time; normalize the cumulative heat of hydration data, let To calculate the final total heat of hydration (usually approximated by a 28-day value), the normalized heat release rate at each time point is calculated: ,at this time, This is the measured decay factor corresponding to that age. The cement hydration process can usually be well described by an exponential decay model, therefore the following mathematical model is used for fitting: ; in, and These are model parameters related to cement performance, obtained through measured data. Substituting the data into the model, the curve can be fitted using the nonlinear least squares method to determine the result; for example, by fitting the curve... , Such specific parameters; age It is a known continuous variable that starts timing from the moment the concrete is poured, and a pre-fitted value is obtained through experiments. and The parameters are stored in the database, and when it is necessary to calculate the attenuation factor at a certain moment, the formula can be directly called. Simply perform the calculation; In reinforcement learning decision models, the output of the policy network is an optimized opening instruction, which is generated by a truncated Gaussian probability distribution, and its formula is: ; in, For the optimized opening command generated at the current moment, corresponding to the i-th valve, normalize to The range corresponds to the valve opening degree of 0%-100%. and It is the policy network in the input state vector The output is the Gaussian distribution mean and standard deviation corresponding to the i-th valve. The truncation function ensures that the output optimized opening command is within the valid range. Inside, For a standard normal distribution The random noise sampled in the middle, where i is the index of the cooling water valve; and It is the output of the policy network (a deep neural network) which takes the fused feature vector as input. The output layer usually has two heads, which output the mean vector and the standard deviation vector respectively. Each dimension of the vector corresponds to a valve. The model determines the optimal opening degree for the i-th valve at the current time t. This is the final output of the control system and directly determines the flow rate of cooling water into the corresponding concrete area. The model judges that the greater the cooling demand of the area corresponding to the valve, the greater the cooling water flow rate is needed for cooling. The larger the value (closer to 1), the greater the valve opening is required by the system decision, thereby providing more cooling water flow rate to the corresponding concrete area and enhancing the cooling effect. The policy network is based on the current state The calculated baseline opening is what the network considers the optimal action. The more information it contains points to the need for cooling, such as high temperature and rapid temperature rise, Generally, the larger; This represents the network's uncertainty about its own decision-making or its willingness to explore, especially during the early stages of training or when its state is uncertain. The opening will be relatively large to encourage exploration of different opening sizes; as learning progresses, It will decrease, and the strategy will tend to stabilize; Random noise To facilitate exploration, randomness is introduced, allowing the system to attempt actions slightly deviating from the current optimal estimate, thereby discovering a better control strategy; truncation function. It is a safety mechanism that ensures the final output command is within the physical range that the valve can execute. (i.e., within 0% to 100%), it reflects the rigor of engineering practice; The value network is implemented using a deep neural network, and its function is to evaluate the state vector. Given the expected long-term return achievable by the current strategy, the training objective of this network is to minimize the temporal difference error, and its loss function is... for: ; in, For the state vector The immediate reward obtained after executing the optimized opening command. This is a future reward discount factor, ranging from 0 to 1, used to measure the value of future rewards at the current moment. This is an estimate of the value function of the value network for the next time step, representing the total expected return that can be obtained starting from the next time step. Let be the state vector at the current time. State value function estimate; Loss value of value network Reflects the value network The inaccuracy of its predictions is determined by the network's predicted values. Compared to the "true" target value The mean square error between them The larger the value network value, the greater its prediction error, meaning that its estimation of state value is very inaccurate and it urgently needs parameter updates to reduce the error. The larger the value network value, the worse the model's "judgment" is, and the less accurately it can assess the long-term performance of the current state and policy. State value estimation It is the dependent variable, the object that the network wants to optimize, and the supervision signal during training; it is equivalent to the immediate reward. (The direct benefits of the current action) plus the discounted value of the next moment. (Present value of potential future benefits); Target value The more specific and accurate, the better, through minimization. Value network The more accurate the prediction, the better. Able to accurately answer: "Current state" "So, what is the total return I can expect from now on?" This provides a benchmark for strategy evaluation; Discount factor The larger the value (closer to 1), the more "foresightful" the model is, and the higher the weight given to future returns. The smaller the size, the more "short-sighted" the model, focusing only on immediate rewards. This balances the immediate temperature control effect with long-term energy consumption and the risk of cracking; immediate rewards It is a direct feedback from the environment to the action; a reward function designed to reduce temperature, balance flow, and conserve water will... Increases when making good decisions; To suppress excessive internal temperature gradients in concrete, the reinforcement learning model uses an improved advantage function for policy optimization. The formula for this function in generalized advantage estimation is as follows: ; in, The advantage estimate at the current moment is used to measure the advantage in the state vector. The extent to which the optimized opening command is advantageous compared to the average level is considered. These are gradient penalty weights, used to control the proportion of the temperature gradient penalty term in the total reward. This is the temperature gradient modulus inside the concrete at the current moment, calculated based on the reconstructed three-dimensional temperature field. The maximum safe temperature gradient threshold; This is a temperature gradient penalty term, which only applies if the temperature gradient... Exceeding the maximum threshold The advantage estimate is negatively adjusted at that time; In the formula, To estimate the dominance value for the standard generalized dominance, , The truncation length represents the maximum number of steps that take into account future time-series difference errors when calculating the generalized dominance estimate. This is the step offset index, used in the summation formula to represent the number of future steps relative to the current time t. The generalized dominance estimation coefficient, ranging from 0 to 1, is used to balance the bias and variance of the dominance estimation. In the formula... , The time-series difference error at the current moment represents the difference between the current estimate and the future estimate; Known as the truncation length, it is a positive integer, such as L=10 or 20, representing the maximum number of steps to consider future time-series difference errors when calculating the generalized dominance estimate. In other words, it defines the upper limit of the number of steps to look backward from the current time t in the summation formula. The meaning of truncation is that in the theoretical formula, the summation of the generalized dominance estimate should be from the current time t to infinity, but in actual calculations, due to the finite sequence length or for computational efficiency, it is only calculated up to t+L-1 steps, and the part after that is ignored (i.e., truncation). This is an approximation that avoids infinite summation while still capturing most of the future information. For example, if L=5, the generalized dominance estimate only calculates the weighted sum of time-series difference errors from time t to t+4, and the influence of more distant futures is ignored. The generalized dominance estimation coefficient is a hyperparameter between 0 and 1, used to control the weight of future time-series difference errors in dominance estimation. Adjusting the balance between the bias and variance of the estimate, when When the generalized dominance estimation degenerates into single-step time-series difference error (high bias, low variance); when At that time, the generalized advantage estimation approximates the Monte Carlo estimation (low bias, high variance). The closer it is to 1, the greater the weight of future returns, resulting in a smoother estimate but higher variance; The closer to 0, the more dependent it is on the current immediate reward, resulting in a more stable estimate but also a larger deviation; in practical applications, The value is determined through experimental parameter tuning, for example, by setting it to 0.95; Step Offset Index It is a non-negative integer ( In the summation formula, it is used to represent the number of future steps relative to the current time t. Starting from 0, it represents the current time t itself; represents the next time step t+1, and so on; represents the index of a future time step, i.e., starting from the current time t, the th... The time after the step. For example, if t is the current time, then t+0 is t, t+1 is the next time, and t+2 is the time after that; l is the summation variable, and t+l constitutes a sequence of future times. In the generalized advantage estimation summation, l is traversed from 0 to L-1, and the temporal difference error of each future time t+l is calculated. and multiplied by the discount weight. This representation is a standard practice in reinforcement learning, used to construct weighted averages over time series. The temperature gradient magnitude Based on the reconstructed three-dimensional temperature field By calculating its spatial partial derivative, we obtain: ; Improved advantage function The value comprehensively measures the state The study assesses how well a particular optimized opening command is compared to the average level, while also accounting for the risk of the optimized opening command causing excessive temperature gradients. The larger the value, the higher the benefit of this optimized opening instruction ( It is a "good action" that should be encouraged and strengthened, as it is large in scale and brings small risks of temperature gradient (small penalty). A positive value indicates that the action is excellent, while a negative value indicates that the action is below average. Taking into account both the short-term and long-term returns of the action, The value reflects the long-term advantages of the action when evaluated solely from the perspective of control objectives (such as temperature and energy consumption). The larger the value, the better the action is from a traditional control perspective (cooling, energy saving); It reflects the maximum rate of temperature change per unit distance within concrete, and is the most direct driving force of thermal stress. The larger the value, the more severe the uneven heating and cooling inside the concrete, and the higher the risk of cracking. The larger the value, the more serious the structural safety is under threat. Temperature gradient penalty term This reflects the safety margin of the temperature gradient caused by the current action, when the gradient When the gradient is within the limit, this term is 0; when the gradient exceeds the limit, this term is negative. The penalty term itself is negative, so the larger its absolute value (i.e., the more negative it is), the more severe the gradient exceeds the limit. The larger the (current temperature gradient), especially when it exceeds... The greater the negative value of the penalty term, the more it affects the overall advantage. The more that is subtracted, the more the advantage score of the action is significantly reduced. This is the core mathematical mechanism for achieving "control instructions that suppress cracks". The greater the weight of safety violations in the overall decision-making process, the lower the tolerance for exceeding the temperature gradient limit, and the more severe the penalty for exceeding the limit. The specific data for some time numbers and advantage estimates are shown in Table 1.

[0020] Table 1. Statistics of Advantage Estimated Values

[0021] Through data analysis, it was observed that there are certain correlations between different characteristic parameters. For example, it can be seen from the data that there is a certain co-changing trend between the standard generalized dominance estimate and the final dominance estimate, but this relationship is significantly affected by the temperature gradient. When the temperature gradient does not exceed the safety threshold, the standard generalized dominance estimate and the final dominance estimate show a relatively consistent relationship. For example, the standard generalized dominance estimate for sample 1 is 1.23, and the corresponding final dominance estimate is 1.25, showing a close correlation. Similarly, the standard generalized dominance estimate for sample 6 is 1.56, and the final dominance estimate is 1.58, also maintaining a good correspondence. This indicates that when the temperature gradient is within the safe range, the system's dominance assessment mainly relies on the initial standard generalized dominance estimate.

[0022] However, when the temperature gradient exceeds the safety threshold, this relationship changes significantly. The standard generalized dominance estimate for sample number 2 is -0.45, but because the temperature gradient reaches 3.20, which significantly exceeds the safety threshold, the final dominance estimate is further reduced to -0.53. Similarly, the standard generalized dominance estimate for sample number 5 is 0.12, which should be positive, but because the temperature gradient is as high as 3.60, which exceeds the safety threshold by a large margin, the final dominance estimate is greatly reduced to 0.05, which is almost zero. Further analysis revealed that the influence of the temperature gradient on the final dominance estimate is also related to the setting of the gradient penalty weight. The standard generalized dominance estimate for sample number 4 is -1.67, and the temperature gradient is 2.80. Due to the higher gradient penalty weight, the final dominance estimate is further reduced to -1.71. Although the temperature gradient for sample number 12 reaches 3.40, the difference between the final dominance estimate and the standard generalized dominance estimate is not very significant due to the relatively small gradient penalty weight. It is worth noting that the case of sample number 9 is quite special. Although the temperature gradient is only 1.90, which does not exceed the safety threshold, the final advantage estimate is also maintained at a low level of -1.30 because the standard generalized advantage estimate itself is a large negative value. This indicates that when the temperature gradient is safe, the system advantage assessment depends more on the initial standard generalized advantage estimate.

[0023] Based on these observations, during the dynamic flow regulation of concrete cooling water pipes, it is necessary to consider both the fundamental role of the standard generalized dominance estimate and the regulating role of the temperature gradient. In particular, when the temperature gradient exceeds the safety threshold, the control strategy should be adjusted appropriately to ensure the stable operation of the system and the effective control of the temperature field.

[0024] The immediate reward consists of a weighted sum of temperature control reward, energy consumption penalty, water consumption penalty, and flow balance reward. The weighting coefficients are determined using a reinforcement learning parameter tuning method. The specific composition is as follows: ; in, This is a reward item for temperature control, where MSE represents the temperature at various points inside the concrete. With target temperature The mean square error, which is negative; the smaller the error (the less negative), the higher the reward; energy consumption penalty term. For the operating power of equipment such as water pumps, negative values ​​are directly used to encourage energy conservation; water consumption penalty item. The total cooling water flow rate is negative to encourage water conservation; flow balance incentive item. The standard deviation of the flow rate of each branch from the average flow rate is negative, which encourages even flow distribution; These are the weight coefficients for each item. These coefficients can be determined using conventional reinforcement learning parameter tuning methods, such as grid search and Bayesian optimization. Their specific values ​​do not affect the implementation of the method. Both the policy network and the value network are deep neural networks composed of fully connected layers. The architecture is as follows: input layer (receives the fused feature vector) → hidden layer 1 (256 neurons, using the ReLU activation function) → hidden layer 2 (128 neurons, using the ReLU activation function) → output layer. The strategy network output layer outputs a corresponding number of valves based on the number of valves. and (Implemented through two independent linear output layers), Value Network Output Layer: Outputs a scalar value ; It is based on the reconstructed three-dimensional temperature field matrix. The gradient penalty weights in the advantage function are obtained by calculating the spatial partial derivatives using the central difference method in numerical differentiation. Discount Factor Generalized dominance estimation coefficient These are hyperparameters, determined using grid search, random search, or Bayesian optimization methods, where: , , .

[0025] Step 3: The optimized opening command that has passed the feasibility verification is sent to the actuator. The actuator includes an electric proportional valve for coarse adjustment of flow rate at the regional level and a piezoelectric ceramic micro valve for fine adjustment of flow rate at the pipe group level. The valve opening is adjusted according to the optimized opening command to realize the dynamic distribution of cooling water flow. Electric proportional valves (regional coarse adjustment) are usually installed at the branch pipe level and are responsible for macroscopic and large-scale adjustment of the total cooling water flow in a large area. They have a relatively slow response speed but a large flow adjustment range. Piezoelectric ceramic micro valves (pipe group level fine adjustment) are usually installed at the capillary level and are responsible for microscopic and fine flow adjustment of a single or a small cluster of capillary loops. They have an extremely fast response speed. This architecture allows the system to increase the total flow rate through the main valve when the core area requires strong cooling, while reducing the flow rate in the edge area through micro valves to prevent overcooling; and vice versa, to achieve on-demand distribution, that is, the total cooling demand of a region is determined by an electric proportional valve, and then the piezoelectric ceramic micro valves accurately distribute the total flow rate to the core area that needs the most cooling and the edge area that needs insulation based on the finer temperature differences in that area. The optimized opening command, after conversion, is output to an electric proportional valve for coarse flow adjustment at the regional level and a piezoelectric ceramic microvalve for fine flow adjustment at the pipe group level. The electric proportional valve receives the voltage signal after range conversion, and its opening... With traffic The relationship, after nonlinear calibration, is described by the following equation: ; in, For flow coefficient, To characterize the nonlinear function of the valve characteristic curve, , This is the steepness coefficient. The pressure difference between the upstream and downstream sides of the valve. Specific gravity of the fluid; Volumetric flow rate It reflects the instantaneous volumetric flow rate of the cooling medium flowing through the valve, and it directly indicates the cooling capacity of the entire area controlled by that valve. The larger the value, the more heat is removed per unit time; the larger the value, the stronger the cooling force the system is applying to that area. Flow coefficient It is a constant provided by the valve manufacturer or calibrated experimentally, representing the inherent flow capacity of the valve itself. The larger the value, the greater the flow rate that the valve can pass under the same opening degree and pressure difference. The larger it is, the more linearly proportional it is to the number of elements in the universe. It is an empirical formula used to describe the nonlinearity caused by the mechanical structure of a valve, where It is the normalized valve opening degree and steepness coefficient. It is an index obtained by fitting experimental data, which determines the "speed" or "sensitivity" of flow rate as it changes with opening degree; when hour, For linear valves, when At this time, it is a fast-opening characteristic: opening degree A slight increase when the flow is small can cause a surge in traffic. The sharp increase in opening makes the valve highly responsive at low opening degrees; opening degree The larger, The larger the volume, the higher the flow rate. The larger the slope coefficient, the higher the steepness coefficient. This amplifies the speed of this change; As a power source to drive fluid flow, The larger the volume, the higher the flow rate. The larger it is, the more it is used in the formula. This reflects a common approximation in turbulent flow, indicating that flow rate is proportional to the square root of the pressure difference; fluid specific gravity... This is the ratio of the fluid's density to the density of water under standard conditions; it reflects the fluid's inertia. The larger the fluid (the "heavier" the fluid), the more difficult it is to be propelled under the same power, therefore the flow rate... The smaller it is, the more inversely proportional it is to each other; Its overall logic is traffic. With opening Valve capacity and driving pressure It is positively correlated with fluid inertia. The negative correlation allows this calibrated model to accurately translate the abstract opening commands output by the agent into the expected physical flow. The piezoelectric ceramic microvalve receives an amplified drive voltage signal and outputs dynamic flow. With drive voltage signal The relationship is described by the following differential equation: ; in, The valve's electromechanical time constant. Piezoelectric gain represents the change in flow rate that can be generated per unit driving voltage. The pressure disturbance coefficient characterizes the upstream pressure difference. The degree of impact on output flow rate; It reflects the instantaneous volumetric flow rate and its dynamic change process through the piezoelectric ceramic micro-valve. It is a quantitative indicator of the fine cooling capacity of the tube group and can reflect the transient process of flow establishment. This value changes with time and indicates the real-time cooling water distribution at the fine control level. The larger the value, the more cooling water the system is precisely injecting into the capillary group at this moment. Electromechanical time constant It is a parameter determined by the valve's physical structure, such as the response speed of the piezoelectric ceramic and the mass of the valve core. The unit is usually seconds, and it quantifies the valve's inertia. The smaller the value, the smaller the valve inertia, the faster the response speed, and the faster it can reach the target flow rate. The "seconds" indicates that this is a fast-response valve with a speed of milliseconds to seconds, making it ideal for frequent, fine adjustments. piezoelectric gain It is a key parameter calibrated experimentally, reflecting the valve's core driving efficiency, that is, the steady-state flow rate change capability generated per unit voltage, and the driving voltage. The larger the value, the more the right side of the equation... The larger the term, the more likely it is to produce a larger steady-state flow. This is a linear proportional relationship; Pressure interference coefficient It is also a parameter calibrated experimentally, reflecting upstream pressure fluctuations. The intensity of interference with the valve, and the increase in upstream pressure, will act as a "reverse drive," causing the right side of the equation to... The term decreases, thus tending to reduce the valve's output flow. This is a negative feedback mechanism. Introducing this allows the controller to anticipate and compensate for the impact of pressure fluctuations in the pipeline network, thereby achieving more precise control. Drive voltage It is a control command issued to the micro-valve. The larger the value, the stronger the driving term on the right side of the equation. The larger it is, the more likely it is to cause Increase; The larger the value, the more interference terms appear on the right side of the equation. The greater the negative impact, the more likely it is to cause... Decrease; The use of first-order differential equations instead of simple static formulas accurately describes the dynamic characteristics of the actuator, which is crucial for achieving high-frequency (e.g., once every 3 seconds) and high-precision control, because the algorithm needs to predict or understand the small delay in flow establishment after the command is issued to avoid overshoot oscillation. , , These parameters are inherent properties of piezoelectric ceramic valves, provided by the manufacturer or obtained through system identification via simple step response experiments. For example, by applying a step voltage to the valve and measuring the flow rate rise curve, the parameters can be fitted. The value of this is standard practice in the field of automatic control; The system parses and converts unified optimization instructions into signals suitable for different actuators. The range conversion voltage is given to the electric valve, and the amplified drive voltage is given to the piezoelectric valve. The electric proportional valve performs regional-level coarse flow adjustment based on its nonlinear static model, and the piezoelectric ceramic micro valve performs pipe-group-level fine flow adjustment based on its first-order dynamic model. The two valve models work together to accurately map the algorithm's decision into a high-resolution cooling flow distribution in the concrete interior space, thereby dynamically and accurately balancing the entire temperature field. Step 4: Monitor the flow attenuation rate and pressure difference of each branch in real time. If the preset blockage conditions are met, the anti-blockage mechanism will be automatically triggered to clean the branch. The anti-blockage mechanism includes an ultrasonic cleaner and a mechanical rotating scraper. Before outputting the optimized opening command, the logic for introducing physical constraints to verify feasibility is as follows: Physical constraint verification includes thermal stress constraints, hydraulic balance constraints, and mechanical protection constraints. Thermal stress constraints ensure that the maximum temperature difference inside the concrete does not exceed the material's allowable threshold, that is, ensure the maximum temperature difference between any two points inside the concrete. ,in This refers to the allowable temperature difference threshold for the material. Maximum temperature difference detected It reflects the temperature difference between the hottest and coldest points inside the entire concrete structure. It directly indicates the overall thermal stress level caused by uneven temperature inside the structure. Excessive temperature difference will cause uneven shrinkage of concrete and produce through cracks. The larger the temperature difference, the greater the thermal stress inside the concrete and the higher the risk of cracking. It is the difference between the maximum and minimum values ​​among all temperature points monitored by the system in real time. It is an observation result. When the control command of the agent causes the core area to be too cold or the edge area to be too hot, thus widening this difference, the constraint will be triggered. This is a constant determined by material properties such as concrete mix proportions, curing time, and admixtures, typically determined through material testing or by following industry standards, such as the "Standard for Construction of Mass Concrete." This constraint is an inequality check; any instruction output by the algorithm, if calculated through simulation or prediction, will lead to... If so, the instruction will be considered infeasible and blocked or modified; (Gradients) are used to prevent surface cracks caused by localized, rapid temperature changes; therefore, during training, they are used as "soft constraints" and guided by a reward function. (Temperature difference) is used to prevent overall and penetrating cracks, so it serves as an insurmountable "hard constraint" before implementation. Both work together on different scales to ensure that the concrete does not crack. Hydraulic balance constraints ensure the flow rate of the main pipeline The sum of the flow rates of each cooling water pipe branch The deviation is within 5%, that is... ; The flow deviation rate reflects the degree of mismatch between the pipeline supply flow and the sum of the demand flows of all branches. The larger the value, the more severely the system's mass conservation is violated, and all flow-based control and diagnostics will lose accuracy. It verifies whether the set of flow commands output by the algorithm physically satisfies the law of mass conservation, indicating a major anomaly in the system. (total flow) and (The sum of the flow rates of each branch) is a direct measurement value. This formula calculates their relative deviation. For example, if the sum of the commanded flow rates of each branch is much greater than the flow rate of the main pipeline, it will cause a sudden drop in system pressure and prevent the system from working properly. It is an engineering experience value used to accommodate flow meter measurement errors and dynamic fluctuations in the system, ensuring the engineering feasibility of the command; Mechanical protection constraints ensure that the rate of change of the optimized opening command does not exceed the maximum response capability of the actuator, and that the optimized opening command is within... Within the range; The optimized opening command change rate reflects the intensity of the control command. It indicates the strength of the impact of the system on the valve mechanism. An excessively high change rate is equivalent to making the valve "open and close rapidly". The larger the change rate, the greater the mechanical impact and wear on the actuator, which may damage the valve or cause water hammer. The maximum allowable change rate is a technical parameter provided by the valve manufacturer and is an inherent property of the valve. When the q-th cooling water pipe branch is detected to simultaneously meet the following preset blockage conditions, and the duration exceeds the preset threshold, When the ultrasonic cleaner and mechanical rotating scraper of the cooling water pipe branch are automatically triggered, the preset blockage conditions include: 1) the flow rate attenuation rate of the qth cooling water pipe branch exceeds the preset attenuation threshold; 2) the inlet and outlet pressure difference of the cooling water pipe branch exceeds the preset safety threshold. The flow decay rate reflects the dynamic process of blockage formation; a higher value indicates faster blockage development. The preset decay threshold is set to 0.1% / minute. Pressure differential reflects the steady-state consequences of blockage; blockages increase flow resistance, requiring a higher driving pressure differential to maintain flow. (This is related to logic and duration.) This is crucial for preventing false positives (e.g., 30 seconds). A single drop in flow rate can be caused by sensor malfunction or other reasons, but combining AND logic with the duration of the drop significantly improves the reliability of the diagnosis. It is an adjustable parameter used to avoid malfunctions in response to instantaneous fluctuations; it also monitors flow rate and differential pressure, which greatly improves the reliability of blockage diagnosis and avoids false alarms that occur with single-parameter diagnosis. When the flow rate of the q-th cooling water pipe branch recovers to its baseline flow rate If the above conditions are met, the cleaning is considered successful; otherwise, a maintenance warning requiring manual intervention will be issued. The baseline flow rate is a historical reference value of the flow rate of this branch under clean and normal conditions. It can be a fixed value or a value that is dynamically updated according to the opening of the main valve. It is a coefficient less than or equal to 1 (e.g., 0.95), which sets an acceptable recovery standard, does not require 100% restoration to the initial state, acknowledges slight residues or performance degradation, and reflects the practicality of the engineering; if the post-cleaning flow rate reaches or exceeds the baseline flow rate... If the value is doubled, the system is considered to have automatically repaired itself; otherwise, it is reported to a human to prevent the system from endlessly looping in ineffective cleaning. The cleaning sequence is performed according to the following steps: 1) Adjust the opening of relevant valves to ensure system safety during cleaning; 2) Start the ultrasonic generator to loosen the deposits using cavitation effect; 3) Start the built-in mechanical rotating scraper to remove stubborn deposits; 4) After a brief flush, restore the pipeline to normal monitoring status. Adjusting the relevant valves specifically refers to partially or completely closing the upstream electric proportional valve of the blocked branch to reduce pipeline pressure and avoid high-pressure water hammer impact during cleaning; maintaining or fine-tuning the opening of the downstream valve of the branch to ensure that impurities generated during the cleaning process can be carried out by the water flow instead of accumulating deep in the pipeline; ensuring system safety means ensuring that the pipeline pressure of the branch drops to a safe low pressure range. Starting the ultrasonic generator is the first round of non-contact cleaning. It uses high-frequency mechanical vibration to generate cavitation bubbles in the liquid. When the bubbles collapse, the resulting micro-jet and shock wave fatigue damages and peels off the scale and deposits on the pipe wall. The frequency is set to 20±2kHz, the duration is 30 seconds, and a frequency sweep mode is used to prevent the formation of standing waves that could lead to uneven cleaning. Activating the built-in mechanical rotating scraper is the second round of contact cleaning. It is used to physically scrape away deposits that have been loosened by ultrasonic waves but not washed away, or particularly stubborn deposits (such as hard limescale). The dual design of ultrasonic waves and mechanical scrapers forms a complete cleaning chain from "loosening" to "removal", which solves the problem that traditional single methods (such as ultrasonic waves alone) are not effective for stubborn blockages. The drive method is magnetic coupling transmission, which is an advanced way to achieve dynamic sealing, ensuring that the rotating parts can work reliably in water-filled pipes without leakage; After a brief flush, the final and verification steps are performed. The purpose of this step is to quickly remove the large amount of impurities removed in the first two steps from the pipeline and prevent them from entering other precision components (such as piezoelectric ceramic microvalves) with the water flow, causing secondary blockage or damage. A brief flush is, for example, slowly reopening the upstream valve that was closed in step 1 to restore the branch to the baseline flow rate and maintaining this state for 10-20 seconds. Returning to normal monitoring status means that after flushing, the system restores all valve openings adjusted for this cleaning to the state of the control commands to be executed before cleaning, reactivates the real-time monitoring logic for that branch, and initiates the cleaning success determination condition, namely, determining whether the flow rate of any cooling water pipe branch has returned to its baseline flow rate. above.

[0026] Please see Figure 6 The present invention also provides a reinforcement learning-based dynamic flow regulation device for concrete cooling water pipes, the device being used to execute the above-described reinforcement learning-based dynamic flow regulation method for concrete cooling water pipes, comprising: The data acquisition module is used to collect data on the internal temperature of the concrete, the flow rate of the cooling water and the differential pressure by an array of temperature sensors embedded in the concrete and flow meters and differential pressure sensors installed on the cooling water pipe branches. It also collects environmental parameters and material property parameters, filters and reduces noise on the collected data, and reconstructs the three-dimensional temperature field inside the concrete based on a spatial interpolation algorithm. The opening optimization module is used to fuse the three-dimensional temperature field, cooling water flow rate, pressure difference data, material properties and environmental parameters into a state vector, and then input it into a reinforcement learning model based on the near-end policy optimization algorithm. The model infers the optimal opening command of each cooling water valve through its policy network and evaluates the state value through its value network. Before outputting the optimal opening command, physical constraints are introduced to verify the feasibility. The valve opening adjustment module is used to send the optimized opening command that has passed the feasibility verification to the actuator. The actuator includes an electric proportional valve for coarse adjustment of flow rate at the regional level and a piezoelectric ceramic micro valve for fine adjustment of flow rate at the pipe group level. The valve opening is adjusted according to the optimized opening command to realize the dynamic distribution of cooling water flow. The monitoring and early warning module is used to monitor the flow attenuation rate and pressure difference changes of each branch in real time. If the preset blockage conditions are met, the anti-blockage mechanism is automatically triggered to clean the branch. The anti-blockage mechanism includes an ultrasonic cleaner and a mechanical rotating scraper.

[0027] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0028] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0029] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0030] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for dynamic flow regulation of concrete cooling water pipes based on reinforcement learning, characterized in that, The specific steps include: Step 1: The temperature, cooling water flow and differential pressure data inside the concrete are collected by an array of temperature sensors embedded in the concrete and flow meters and differential pressure sensors installed on the cooling water pipe branches. Environmental parameters and material property parameters are also collected. The collected data are filtered and noise reduced. The three-dimensional temperature field inside the concrete is reconstructed based on the spatial interpolation algorithm. Step 2: After fusing the three-dimensional temperature field, cooling water flow rate, pressure difference data, material properties and environmental parameters into a state vector, it is input into a reinforcement learning model based on the near-end policy optimization algorithm. The model infers the optimal opening command of each cooling water valve through its policy network and evaluates the state value through its value network. Before outputting the optimal opening command, physical constraints are introduced to verify the feasibility. Step 3: The optimized opening command that has passed the feasibility verification is sent to the actuator. The actuator includes an electric proportional valve for coarse adjustment of flow rate at the regional level and a piezoelectric ceramic micro valve for fine adjustment of flow rate at the pipe group level. The valve opening is adjusted according to the optimized opening command to realize the dynamic distribution of cooling water flow. Step 4: Monitor the flow attenuation rate and pressure difference of each branch in real time. If the preset blockage conditions are met, the anti-blockage mechanism will be automatically triggered to clean the branch. The anti-blockage mechanism includes an ultrasonic cleaner and a mechanical rotating scraper.

2. The method for dynamic flow regulation of concrete cooling water pipes based on reinforcement learning according to claim 1, characterized in that: Multiple layers of temperature sensors are pre-embedded along the height direction in four key areas inside the concrete to form a three-dimensional monitoring network. The key areas include the hydration heat core area, the heat dissipation sensitive area, the stress danger area, and the boundary effect area. The arrangement density of temperature sensors in the hydration heat core area and the stress danger area is higher than that in the heat dissipation sensitive area and the boundary effect area. Temperature sensors are used to collect temperature data. Differential pressure sensors and flow meters are installed at the branch inlets and capillary junctions of each cooling water pipe to monitor pipeline differential pressure and cooling water flow. The collected raw temperature data is processed by a sliding window filtering algorithm for noise reduction, and a Kriging space interpolation algorithm is used to reconstruct the three-dimensional temperature field distribution.

3. The method for dynamic flow regulation of concrete cooling water pipes based on reinforcement learning according to claim 1, characterized in that: The reinforcement learning model uses a hybrid neural network architecture for feature extraction. This architecture uses a three-dimensional convolutional neural network to extract spatial features from a three-dimensional temperature field. The convolutional kernel size of the three-dimensional convolutional neural network is 3×3×3, and the stride is 1. At the same time, a long short-term memory network is used to process the temporal features of flow rate, pressure difference and temperature data, and a fully connected network is used to process environmental parameters and material property parameters. Finally, all features are fused and input into the policy network and the value network. Input the state vector of the reinforcement learning model The data is composed of the following multidimensional data fusion: the current three-dimensional temperature field obtained from reconstruction, the current flow rate value of each cooling water pipe branch and the current pressure difference value of the pipe inlet and outlet, the current material characteristic parameters and the current environmental parameters. The material characteristic parameters include the cement type coefficient and the decay factor that changes with age calculated from the time of concrete pouring. The environmental parameters include ambient temperature, relative humidity and wind speed, where t is the index of the current time.

4. The method for dynamic flow regulation of concrete cooling water pipes based on reinforcement learning according to claim 3, characterized in that: In reinforcement learning decision models, the output of the policy network is an optimized opening instruction, which is generated by a truncated Gaussian probability distribution, and its formula is: ; in, The optimized opening command for the i-th valve, generated at the current time t, is normalized to... The range corresponds to the valve opening degree of 0%-100%. and It is the policy network in the input state vector The output is the Gaussian distribution mean and standard deviation corresponding to the i-th valve. The truncation function ensures that the output optimized opening command is within the valid range. Inside, For a standard normal distribution The random noise sampled in the middle, where i is the index of the cooling water valve.

5. The method for dynamic flow regulation of concrete cooling water pipes based on reinforcement learning according to claim 4, characterized in that: The value network is implemented using a deep neural network, and its function is to evaluate the state vector. Given the expected long-term return achievable by the current strategy, the training objective of this network is to minimize the temporal difference error, and its loss function is... for: ; in, For the state vector The immediate reward obtained after executing the optimized opening command. This is a future reward discount factor, ranging from 0 to 1, used to measure the value of future rewards at the current moment. This is an estimate of the value function of the value network for the next time step, representing the total expected return that can be obtained starting from the next time step. Let be the state vector at the current time. State value function estimate; To suppress excessive internal temperature gradients in concrete, the reinforcement learning model uses an improved advantage function for policy optimization. The formula for this function in generalized advantage estimation is as follows: ; in, The advantage estimate at the current moment is used to measure the advantage in the state vector. The extent to which the optimized opening command is advantageous compared to the average level is considered. These are gradient penalty weights, used to control the proportion of the temperature gradient penalty term in the total reward. This is the temperature gradient modulus inside the concrete at the current moment, calculated based on the reconstructed three-dimensional temperature field. The maximum safe temperature gradient threshold; This is a temperature gradient penalty term, which only applies if the temperature gradient... Exceeding the maximum threshold The advantage estimate is negatively adjusted at that time; In the formula, To estimate the dominance value for the standard generalized dominance, , The truncation length represents the maximum number of steps that take into account future time-series difference errors when calculating the generalized dominance estimate. This is the step offset index, used in the summation formula to represent the number of future steps relative to the current time t. The generalized dominance estimation coefficient, ranging from 0 to 1, is used to balance the bias and variance of the dominance estimation. In the formula... , The time-series difference error at the current moment represents the difference between the current estimate and the future estimate; The temperature gradient magnitude Based on the reconstructed three-dimensional temperature field By calculating its spatial partial derivative, we obtain: .

6. The method for dynamic flow regulation of concrete cooling water pipes based on reinforcement learning according to claim 5, characterized in that: The optimized opening command, after conversion, is output to an electric proportional valve for coarse flow adjustment at the regional level and a piezoelectric ceramic microvalve for fine flow adjustment at the pipe group level. The electric proportional valve receives the voltage signal after range conversion, and its opening... With traffic The relationship, after nonlinear calibration, is described by the following equation: ; in, For flow coefficient, To characterize the nonlinear function of the valve characteristic curve, , This is the steepness coefficient. The pressure difference between the upstream and downstream sides of the valve. Specific gravity of the fluid; The piezoelectric ceramic microvalve receives an amplified drive voltage signal and outputs dynamic flow. With drive voltage signal The relationship is described by the following differential equation: ; in, The valve's electromechanical time constant. Piezoelectric gain represents the change in flow rate that can be generated per unit driving voltage. The pressure disturbance coefficient characterizes the upstream pressure difference. The degree of impact on output flow.

7. The method for dynamic flow regulation of concrete cooling water pipes based on reinforcement learning according to claim 6, characterized in that: Before outputting the optimized opening command, the logic for introducing physical constraints to verify feasibility is as follows: Physical constraint verification includes thermal stress constraints, hydraulic balance constraints, and mechanical protection constraints. Thermal stress constraints ensure that the maximum temperature difference inside the concrete does not exceed the material's allowable threshold, that is, ensure the maximum temperature difference between any two points inside the concrete. ,in This refers to the allowable temperature difference threshold for the material. Hydraulic balance constraints ensure the flow rate of the main pipeline The sum of the flow rates of each cooling water pipe branch The deviation is within 5%, that is... ; Mechanical protection constraints ensure that the rate of change of the optimized opening command does not exceed the maximum response capability of the actuator, and that the optimized opening command is within... Within the range; When the q-th cooling water pipe branch is detected to simultaneously meet the following preset blockage conditions, and the duration exceeds the preset threshold, When the ultrasonic cleaner and mechanical rotating scraper of the cooling water pipe branch are automatically triggered, the preset blockage conditions include: 1) the flow rate attenuation rate of the qth cooling water pipe branch exceeds the preset attenuation threshold; 2) the inlet and outlet pressure difference of the cooling water pipe branch exceeds the preset safety threshold. When the flow rate of the q-th cooling water pipe branch returns to its baseline flow rate If the above conditions are met, the cleaning is considered successful; otherwise, a maintenance warning requiring manual intervention will be issued. q is the index of the cooling water pipe branch.

8. A dynamic flow regulation device for concrete cooling water pipes based on reinforcement learning, characterized in that: The device is used to execute a reinforcement learning-based dynamic flow regulation method for concrete cooling water pipes as described in any one of claims 1-7, comprising: The data acquisition module is used to collect data on the internal temperature of the concrete, the flow rate of the cooling water and the differential pressure by an array of temperature sensors embedded in the concrete and flow meters and differential pressure sensors installed on the cooling water pipe branches. It also collects environmental parameters and material property parameters, filters and reduces noise on the collected data, and reconstructs the three-dimensional temperature field inside the concrete based on a spatial interpolation algorithm. The opening optimization module is used to fuse the three-dimensional temperature field, cooling water flow rate, pressure difference data, material properties and environmental parameters into a state vector, and then input it into a reinforcement learning model based on the near-end policy optimization algorithm. The model infers the optimal opening command of each cooling water valve through its policy network and evaluates the state value through its value network. Before outputting the optimal opening command, physical constraints are introduced to verify the feasibility. The valve opening adjustment module is used to send the optimized opening command that has passed the feasibility verification to the actuator. The actuator includes an electric proportional valve for coarse adjustment of flow rate at the regional level and a piezoelectric ceramic micro valve for fine adjustment of flow rate at the pipe group level. The valve opening is adjusted according to the optimized opening command to realize the dynamic distribution of cooling water flow. The monitoring and early warning module is used to monitor the flow attenuation rate and pressure difference changes of each branch in real time. If the preset blockage conditions are met, the anti-blockage mechanism is automatically triggered to clean the branch. The anti-blockage mechanism includes an ultrasonic cleaner and a mechanical rotating scraper.

Citation Information

Patent Citations

  • Normal state-roller compacted concrete gravity dam combined damming safety assessment method

    CN111460548A

  • Liquid cooling server safety management system and method

    CN120152228A

  • Concrete multi-target proportioning optimization method and equipment based on reinforcement learning and medium

    CN120853709A

Cited By

  • Concrete mix proportion intelligent generation and online compliance checking method and system

    CN122245572A