Cross-scene intelligent refrigeration system energy efficiency optimization method based on multi-agent collaborative reinforcement learning
By combining multi-agent collaborative reinforcement learning and digital twin verification with a nanofluid four-parameter coupling model and Chinese population thermal comfort evaluation, a closed-loop optimization of a cross-scenario intelligent cooling system is formed, which solves the energy efficiency and stability problems of the cooling system under complex operating conditions and achieves sustainable optimization and user comfort across scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF JINAN
- Filing Date
- 2026-04-13
- Publication Date
- 2026-07-14
AI Technical Summary
Existing refrigeration systems struggle to achieve stable and sustainable energy efficiency optimization under complex operating conditions. Nanofluid solutions do not adequately consider the side effects caused by particle concentration. Single-agent control has poor stability. Digital twin technology is not deeply coupled, making cross-scenario deployment difficult. The issue of thermal comfort for people has not been adequately considered.
By employing multi-agent collaborative reinforcement learning control, a nanofluid four-parameter coupling model, digital twin pre-deployment verification, federated training and modular deployment, and combining Chinese population thermal comfort evaluation, a closed-loop optimization of a cross-scenario intelligent cooling system is formed.
It achieves system-level energy efficiency optimization and transport resistance trade-off, improves online control stability, reduces deployment risks, and enhances cross-scenario reusability and user comfort.
Smart Images

Figure CN122386670A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent optimization and control technology for refrigeration systems, and particularly to a cross-scenario intelligent refrigeration system energy efficiency optimization method that combines nanofluid heat transfer mechanism modeling, multi-agent collaborative reinforcement learning control, digital twin strategy verification, cross-scenario federated training, and modular deployment adaptation. This method is applicable to refrigeration scenarios requiring a balance between energy efficiency, stability, and feasibility, such as building central air conditioning, industrial process cooling, rail transit vehicle cooling, and data center cooling.
[0002] More specifically, this invention does not focus on the local control of a single device, but rather on building a closed-loop control chain around the entire refrigeration system, from physical parameter modeling, state characterization, intelligent decision-making, pre-deployment simulation verification, on-site execution to feedback updates, so that the refrigeration system can still maintain a good level of energy efficiency and operational stability under complex operating conditions. Background Technology
[0003] With the increasing demands for building energy conservation, industrial energy conservation, and high-reliability temperature control, refrigeration systems have gradually transformed from simple infrastructure providing cooling capacity into comprehensive control objects that simultaneously undertake multiple tasks such as energy efficiency optimization, comfort assurance, safe operation, and maintenance management. Existing refrigeration systems typically include compressors, evaporators, condensers, circulating pumps, fans, throttling devices, and sensors distributed across various key components. These devices exhibit significant thermodynamic coupling, fluid transport coupling, and control timing coupling relationships. Therefore, relying solely on local adjustments at a single measuring point or a single actuator often fails to achieve stable and sustainable energy-saving effects at the system level.
[0004] In existing technologies, one type of approach primarily focuses on improving the formulation of nanofluids, nanoparticles, or heat exchange media. For example, by changing the type, volume fraction, dispersion method, or surface modification of nanoparticles, the thermal conductivity of the fluid can be increased, enhancing the heat exchange effect. While this approach can improve local heat transfer capacity to some extent, it often prioritizes increasing thermal conductivity and fails to adequately consider the potential side effects of increased particle concentration, such as increased viscosity, transport resistance, and pump work. In other words, existing nanofluid-related solutions largely remain at the material or formulation level; the methods for simultaneously incorporating four variables—thermal conductivity, concentration, temperature, and viscosity—into a real-time control system and forming an executable system-level optimization strategy are still insufficiently disclosed.
[0005] Another type of approach focuses on using machine learning, reinforcement learning, or deep reinforcement learning to control central air conditioning, chiller systems, cooling towers, pumps, fans, or valves to achieve energy savings or load tracking. However, this type of approach generally suffers from three problems: First, most approaches use single-agent or loosely coupled multi-agent structures, making it difficult to maintain good training stability and online interpretability in refrigeration systems with high action dimensions and many controlled objects. Second, most approaches mainly construct reward functions around power consumption or temperature errors, rarely explicitly incorporating the physical constraints of the heat exchange medium itself into the state and constraint spaces. Third, candidate strategies are often directly output by the data-driven model and then executed on-site, lacking a reliable mechanism for verifying the strategies in a virtual environment before deployment, thus posing a certain risk of deployment when operating conditions change abruptly, abnormal loads occur, or boundary conditions deviate.
[0006] Another approach uses digital twin technology to model, monitor, or analyze the operation and maintenance of refrigeration systems. Digital twins can build models corresponding to physical devices in a virtual space and simulate operating states and fault evolution. However, many existing publicly available digital twins only focus on visualization, offline analysis, or fault diagnosis, without deep coupling with reinforcement learning control, and without forming a closed-loop process where "candidate policies are first verified in the twin environment, and then the verified policies are executed on-site."
[0007] Regarding cross-scenario deployment, existing intelligent control solutions are mostly customized for single buildings, single processes, or single equipment families, with strong scenario-dependent algorithm models, communication interfaces, and hardware deployment methods. This leads to difficulties in model migration between different scenarios, high development costs, and poor reusability. Meanwhile, pedestrian flow and energy consumption data in building scenarios, process data in industrial scenarios, and operational data in data center scenarios can all be highly sensitive. Uploading raw datasets centrally for unified training not only poses a risk of data leakage but may also encounter compliance obstacles.
[0008] In cooling scenarios involving human living environments, thermal comfort is equally important. While traditional PMV models are widely used, their parameter systems are largely based on general standards. When directly applied to Chinese populations with different regions, clothing habits, and activity levels, discrepancies may arise between the control objectives and users' subjective comfort. Simply using general comfort thresholds to drive control can easily lead to a deviation between energy-saving results and the user's subjective comfort experience.
[0009] Therefore, it is necessary to propose an integrated technical solution that couples nanofluidic physics models, hierarchical multi-agent control, digital twin pre-deployment verification, Chinese population thermal comfort constraints, federated learning security aggregation, and cross-scenario modular deployment along the same technical line. This would solve the problem of existing technologies where modules are isolated from each other and difficult to form a stable closed loop. Summary of the Invention
[0010] To address the aforementioned problems in existing technologies, this invention provides a cross-scenario intelligent cooling system energy efficiency optimization method based on multi-agent collaborative reinforcement learning. This method uses a nanofluidic four-parameter coupling model as its physical foundation, a hierarchical collaborative control framework consisting of a sensing agent, a decision-making agent, a coordinating agent, and an execution agent as its decision-making core, a digital twin model as its pre-deployment verification and online write-back tool, federated training as its cross-scenario update mechanism, and introduces Chinese-speaking population thermal comfort evaluation constraints into applicable scenarios, thus forming a technical solution with physical constraints, simulation pre-verification, closed-loop feedback, and transferable deployment capabilities.
[0011] To achieve the above objectives, the present invention adopts the following technical solution: First, acquire equipment operation data, environmental data, load data, and nanofluid-related parameters in the target refrigeration scenario, and establish a four-parameter coupled model of thermal conductivity, concentration, temperature, and viscosity; second, map the output of the four-parameter coupled model as part of the state vector, and construct a multi-agent collaborative control architecture to generate several candidate control strategies; third, establish a digital twin model corresponding to the target refrigeration scenario, perform operating condition simulation and boundary constraint verification on each candidate control strategy, and only determine the strategy that passes the verification as the target control strategy; then, send the target control strategy to the actual refrigeration equipment for execution; finally, collect the feedback data after execution, update the four-parameter coupled model, the digital twin model, and the multi-agent control strategy, forming a closed-loop optimization process.
[0012] In a preferred embodiment, the four-parameter coupling model uses thermal conductivity k, nanoparticle volume fraction φ, temperature T, and dynamic viscosity μ as core coupling variables, and can be corrected according to online concentration and viscosity deviations. Compared to schemes that only focus on improving thermal conductivity, this invention emphasizes the parallel consideration of improving thermal conductivity and viscosity constraints, thereby avoiding the loss of system energy-saving benefits due to increased transport resistance caused by high concentrations of nanoparticles.
[0013] In a preferred embodiment, the multi-agent collaborative control architecture adopts a hierarchical organizational form. The perception agent is responsible for collecting multi-source data from the equipment side and the environment side; the decision agent generates candidate actions for different sub-tasks such as compressors, pumps, fans, valves, and nanofluidic branches; the coordination agent resolves conflicts, filters for consistency, and fuses global goals for candidate actions from multiple decision agents; and the execution agent converts the filtered target strategy into control commands that can be recognized by the field controller.
[0014] In a preferred embodiment, the digital twin model includes at least a physical entity layer, a data mapping layer, a model layer, a simulation layer, and an execution feedback layer. In the offline phase, the digital twin model can be used to construct typical operating condition samples, extreme operating condition samples, and risky operating condition samples. In the online phase, it performs pre-deployment simulation verification of candidate control strategies, including but not limited to energy efficiency boundaries, temperature boundaries, pressure boundaries, viscosity boundaries, and abnormal risk boundaries. Only candidate strategies that pass the digital twin verification can be implemented in the field.
[0015] In scenarios applicable to human living environments, this invention further establishes a thermal comfort evaluation model suitable for the Chinese population. This model combines parameters such as regional adaptation coefficient, metabolic rate, clothing thermal resistance, ambient temperature, ambient humidity, and wind speed to correct general thermal comfort indices, and uses the corrected thermal comfort evaluation results as one of the state inputs and constraints in the optimization process.
[0016] Regarding cross-scenario updates, this invention preferably employs a federated training approach. Each scenario node completes model training locally based on its own scenario data, uploading only encrypted or perturbed model parameters or gradient information. The aggregation end then performs parameter fusion without exchanging the original runtime data. This approach enhances the generalization ability of cross-scenario strategies and also facilitates data security and engineering implementation.
[0017] At the deployment level, this invention preferably employs modular hardware and a protocol adaptation mechanism, enabling the same control framework to interface with different cooling devices through protocol adaptation. The modular hardware includes at least a control module, a sensing module, an execution module, and an edge computing module, and the protocol adaptation supports at least one of Modbus, Profinet, and EtherNet / IP.
[0018] Compared with existing technologies, this invention has at least the following beneficial effects: First, by introducing a four-parameter coupled nanofluid model into the control state space and constraint space, it is possible to comprehensively balance the side effects of enhanced heat transfer and transport resistance at the system level; Second, by adopting a multi-agent collaborative control architecture that separates perception, decision-making, coordination, and execution, it is beneficial to reduce the training difficulty in complex action spaces and improve the stability of online control; Third, by adopting a digital twin pre-deployment verification mechanism, it reduces the operational risks of directly deploying data-driven strategies; Fourth, by introducing Chinese population thermal comfort evaluation results into human settlement scenarios, it makes energy efficiency optimization more consistent with actual comfort experience; Fifth, by adopting federated training and encrypted aggregation to achieve cross-scenario knowledge sharing, it is beneficial to improve the model's generalization ability without sharing the original data; Sixth, by adopting a modular deployment and protocol adaptation mechanism, it improves the engineering reuse efficiency between different cooling scenarios. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the overall closed-loop architecture of the present invention.
[0020] Figure 2 This is a schematic diagram of four-parameter coupling and state mapping in this invention.
[0021] Figure 3 This is a schematic diagram of the multi-agent collaborative control and digital twin verification process in this invention.
[0022] Figure 4 This is a schematic diagram of cross-scenario federated training and modular deployment in this invention.
[0023] Explanation of reference numerals in the attached figures 100 Refrigeration System Body; 110 Data Acquisition Module; 120 Four-Parameter Coupled Model Module; 121 Thermal Conductivity Submodule; 122 Concentration Submodule; 123 Temperature Submodule; 124 Viscosity Submodule; 125 State Vector Generation Unit; 126 Constraint Boundary Generation Unit; 127 Candidate Action Set; 130 Multi-Agent Cooperative Control Module; 131 Perception Agent; 132 Decision Agent; 133 Coordination Agent; 134 Execution Agent; 140 Digital Twin Verification Module; 141 Digital Twin Model; 142 Safety Constraint Judgment Unit; 150 Execution Control Module; 160 Feedback Update Module; 171 Building Scene Node; 172 Industrial Scene Node; 173 Rail Transit Scene Node; 174 Encryption Aggregation Node; 175 Protocol Adaptation Module; 176 Control Module; 177 Sensing / Execution Module. Detailed Implementation
[0024] The present invention will now be described in further detail with reference to the accompanying drawings. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. For those skilled in the art, any equivalent substitutions made to specific parameters, algorithm types, device forms, and deployment methods without departing from the core concept of the present invention should be considered to fall within the scope of protection of the present invention.
[0025] like Figure 1 As shown, the overall closed-loop architecture of this invention includes at least a refrigeration system body 100, a data acquisition module 110, a four-parameter coupled model module 120, a multi-agent cooperative control module 130, a digital twin verification module 140, an execution control module 150, and a feedback update module 160. The data acquisition module 110 acquires real-time operating data, environmental data, and load data from the refrigeration system body 100; the four-parameter coupled model module 120 forms physical constraints based on the acquired nanofluid-related parameters; the multi-agent cooperative control module 130 generates candidate control strategies based on the current state vector; the digital twin verification module 140 performs operating condition simulation and boundary verification on the candidate control strategies; the execution control module 150 sends the verified target strategy to the refrigeration system body 100; and the feedback update module 160 updates the model and strategy based on the execution results.
[0026] The refrigeration system body 100 can adopt different equipment combinations according to different scenarios. In the building central air conditioning scenario, the refrigeration system body 100 may include a chiller, chilled water pump, cooling water pump, cooling tower, air handling unit, and fan coil unit; in the industrial process cooling scenario, the refrigeration system body 100 may include a process heat exchanger, circulating pump, compressor unit, and valve array; in the rail transit scenario, the refrigeration system body 100 may include an on-board compressor, fan, heat exchange module, and edge controller; in the data center scenario, the refrigeration system body 100 may include a cold plate, pump set, cabinet air duct device, and heat exchange loop.
[0027] The data acquisition module 110 may include temperature sensors, pressure sensors, flow sensors, power acquisition devices, concentration detection devices, viscosity estimation devices, ambient temperature and humidity sensors, and human flow or activity status sensing devices. The data acquired by these sensors may include supply liquid temperature, return liquid temperature, evaporation temperature, condensation temperature, evaporation pressure, condensation pressure, flow rate, compressor frequency, fan speed, valve opening, equipment power, indoor temperature and humidity, ambient temperature, load demand, nanoparticle concentration, and estimated dynamic viscosity. The data acquisition module 110 can also perform time alignment, outlier removal, filtering, and standardization on the raw data to improve the stability of subsequent modeling and control.
[0028] like Figure 2As shown, the four-parameter coupled model module 120 includes at least a thermal conductivity submodule 121, a concentration submodule 122, a temperature submodule 123, and a viscosity submodule 124, and is further connected to a state vector generation unit 125 and a constraint boundary generation unit 126. The thermal conductivity submodule 121 is used to characterize the thermal conductivity of the nanofluid under the current operating conditions; the concentration submodule 122 is used to characterize the volume fraction or concentration deviation of the nanoparticles; the temperature submodule 123 is used to characterize the fluid temperature, ambient temperature, and temperature conditions related to the heat transfer process; and the viscosity submodule 124 is used to characterize the fluid dynamic viscosity and the resulting transport resistance trend.
[0029] In one implementation, the four-parameter coupling relationship can be established by combining experimental sampling and simulation fitting. Specifically, the thermal conductivity and viscosity data of nanofluid samples can be measured first under different concentrations and temperatures, and then a fitting model can be established by combining the results of particle surface modification, particle size distribution, and dispersion stability. To improve the engineering usability of the model, the coupling relationship can be represented as a polynomial model, a piecewise function model, a lookup table interpolation model, or a neural network approximation model. As long as the model can simultaneously reflect the coupling effects between thermal conductivity, concentration, temperature, and viscosity, and can output state variables and constraint boundaries, it can be applied to this invention.
[0030] The state vector generation unit 125, based on the outputs of the thermal conductivity submodule 121, concentration submodule 122, temperature submodule 123, and viscosity submodule 124, concatenates physical parameters with equipment-side state parameters to form a control state vector. In addition to the four coupled variables, the control state vector may also include compressor frequency, pump frequency, fan speed, valve opening, supply and return liquid temperatures, ambient temperature and humidity, and load requirements. The constraint boundary generation unit 126 generates temperature boundaries, pressure boundaries, viscosity boundaries, energy consumption boundaries, and rate of change of motion boundaries based on the current operating conditions and preset safety boundaries. The candidate action set 127 represents the set of control actions that the multi-agent collaborative control module 130 can select under the current state and constraint conditions.
[0031] During actual operation, when the concentration submodule 122 detects a high volume fraction of nanoparticles, the constraint boundary generation unit 126 simultaneously increases its sensitivity to viscosity boundaries and restricts some actions that might significantly increase pump power. When the thermal conductivity submodule 121 indicates insufficient heat transfer enhancement while the viscosity submodule 124 remains within the allowable range, the candidate action set 127 is allowed to increase the flow rate of the nanofluid branch or adjust the ratio to improve the heat transfer effect. Thus, nanofluid parameters are no longer only used for offline material design but also participate in system optimization as part of real-time control.
[0032] like Figure 3As shown, the multi-agent cooperative control module 130 includes at least a perceptual agent 131, a decision-making agent 132, a coordinating agent 133, and an executive agent 134. The perceptual agent 131 receives data from the data acquisition module 110 and the state vector generation unit 125, and is responsible for forming a state representation suitable for reinforcement learning processing; the decision-making agent 132 generates candidate actions based on the state representation and local control tasks; the coordinating agent 133 combines the global objective and constraints to fuse and filter multiple candidate actions; and the executive agent 134 converts the final action into a field control command.
[0033] The perceptual agent 131 can group data according to equipment type, geographical location, or control timing. For example, data from compressors, pump sets, fans, and valves can be used to form local observations; data related to human comfort, process temperature, or vehicle disturbances can also be used as additional observation inputs. The decision-making agent 132 can be set up separately for different sub-tasks. For example, one set of decision-making agents can be set up for chillers and electronic expansion valves, another set for pump sets and fans, and yet another set for nanofluidic circulation branches and bypass valves. This allows the high-dimensional action space to be divided into several relatively trainable subspaces.
[0034] The coordinating agent 133 does not simply select one of several candidate actions, but rather uses preset priority rules, constraint satisfaction rules, consensus algorithms, or secondary optimization mechanisms to perform consensus screening of candidate actions. For example, one decision agent 132 may tend to increase the pump frequency to enhance heat exchange, while another decision agent 132 may tend to decrease the pump frequency to reduce energy consumption. In this case, the coordinating agent 133 needs to combine the viscosity boundary given by the four-parameter coupled model module 120, the operating condition prediction given by the digital twin model 141, and the overall energy efficiency target to select a better combination of actions.
[0035] In a preferred embodiment, the reward function of the decision agent 132 includes at least an energy efficiency improvement component, a thermal comfort satisfaction rate component, an energy consumption penalty component, and a viscosity deviation penalty component. For industrial scenarios, a temperature control accuracy penalty component and a safety boundary violation penalty component can also be added; for rail transit scenarios, a motion smoothness penalty component can also be added to reduce the impact of frequent and large-scale adjustments on on-board actuators.
[0036] Figure 3The digital twin model 141 and the safety constraint determination unit 142 constitute the core of the digital twin verification module 140. The digital twin model 141 includes at least a physical entity layer, a data mapping layer, a model layer, a simulation layer, and an execution feedback layer. The physical entity layer corresponds to the actual refrigeration system equipment, the data mapping layer is responsible for receiving standardized data from the data acquisition module 110, the model layer is responsible for establishing thermodynamic models, fluid transport models, and environmental response models, the simulation layer is used to perform rapid simulation of candidate actions within a future time window, and the execution feedback layer is used to receive on-site execution results and write back the model parameters.
[0037] In one implementation, after the decision-making agent 132 outputs several candidate actions, the digital twin model 141 predicts the temperature change, pressure change, energy consumption change, and viscosity trend of different candidate actions within the future prediction window. The safety constraint determination unit 142 performs boundary judgment on the prediction results; if there are situations such as temperature exceeding limits, pressure exceeding limits, viscosity exceeding allowable limits, action switching being too frequent, or comfort falling out of the acceptable range, the candidate action is marked as unacceptable and fed back to the coordinating agent 133 for reselection or reorganization of actions. Only the target action verified by the safety constraint determination unit 142 will be issued by the executing agent 134 to the execution control module 150.
[0038] The execution control module 150 is used to map the target action into a control quantity that can be recognized by the field controller. In different scenarios, the execution control module 150 may correspond to a compressor frequency converter, pump controller, fan drive, electronic expansion valve controller, valve actuator, or nanofluid proportioning execution unit, respectively. After receiving the target action, the execution control module 150 adjusts at least two of the following: compressor speed, pump frequency, fan speed, electronic expansion valve opening, nanofluid circulation flow rate, and bypass valve opening.
[0039] The feedback update module 160 is used to collect the actual running results after execution and update the four-parameter coupled model module 120, the digital twin model 141, and the multi-agent cooperative control module 130. The update content may include four-parameter model parameter correction, digital twin model parameter calibration, experience sample write-back, local policy parameter updates, and cross-scenario federated training cache updates. The feedback update module 160 can operate on a second, minute, or hourly cycle, with the specific update cycle set according to the scenario type and device inertia.
[0040] For scenarios related to human living environments, this invention preferably establishes a thermal comfort evaluation model applicable to the Chinese population. This thermal comfort evaluation model can comprehensively consider factors such as regional adaptability coefficient, metabolic rate, clothing thermal resistance, ambient temperature, ambient humidity, and wind speed to correct general thermal comfort indicators. The corrected comfort evaluation result can be used as input to the sensing agent 131 and as one of the constraint conditions of the safety constraint determination unit 142. When a candidate action with better energy efficiency may cause the comfort evaluation to fall out of the acceptable range, the safety constraint determination unit 142 can directly determine the candidate action as unacceptable.
[0041] like Figure 4 As shown, this invention, during cross-scene updates, includes at least a building scene node 171, an industrial scene node 172, a rail transit scene node 173, an encrypted aggregation node 174, a protocol adaptation module 175, a control module 176, and a sensing / execution module 177. Different scene nodes locally train multi-agent control strategies using their respective scene data, uploading only encrypted or perturbed model parameters or gradient information to the encrypted aggregation node 174, without uploading the original runtime data. After completing parameter fusion, the encrypted aggregation node 174 sends the new shared strategy or global parameters back to each scene node.
[0042] In a preferred embodiment, the upload process can employ at least one of homomorphic encryption and differential privacy, and can also combine at least one of gradient quantization compression, sparse sampling, and asynchronous reporting to reduce communication overhead during cross-scenario training and improve parameter aggregation efficiency. This approach enhances knowledge sharing and model generalization capabilities across different scenarios without sharing the original data.
[0043] Protocol adaptation module 175 enables the same control framework to interface with different devices and protocols. For example, in building scenarios, it can prioritize adaptation to commonly used building automation system protocols; in industrial scenarios, it can prioritize adaptation to industrial Ethernet and PLC interfaces; and in rail transit scenarios, it can prioritize adaptation to vibration-resistant input / output interfaces and vehicle communication interfaces. Control module 176 can be an embedded controller, industrial controller, or edge computing node, while sensing / actuation module 177 corresponds to the sensor and actuator components in the actual deployment. Through modular design, reused deployment across different scenarios can be achieved by simply replacing some units in protocol adaptation module 175 or sensing / actuation module 177.
[0044] The application of this invention will be further explained below with reference to specific scenarios.
[0045] Example 1: In an office building central air conditioning scenario, the refrigeration system 100 includes a chiller, chilled water pump, cooling tower, air handling unit, and fan coil units. The data acquisition module 110 collects information such as outdoor temperature, indoor temperature and humidity, supply and return water temperature, pump frequency, fan speed, valve opening, and nanoparticle concentration. The state vector generation unit 125 combines the above information with the outputs of the thermal conductivity submodule 121 and the viscosity submodule 124 to form the current state vector. The decision-making agent 132 generates candidate actions for the chiller, pump unit, and terminal fans, respectively. The digital twin model 141 verifies the changes in room temperature and energy consumption within a future prediction window. When a candidate action, although capable of reducing energy consumption, causes the comfort evaluation of the conference room area to fall outside the set range, the safety constraint judgment unit 142 determines the candidate action as unacceptable, and the coordinating agent 133 selects a target action that balances comfort and energy efficiency.
[0046] Example 2: In an industrial reactor cooling scenario, the refrigeration system 100 includes a process heat exchanger, a circulating pump, a compressor unit, and a high-temperature sensor. Since this scenario prioritizes process temperature accuracy and safety boundaries, thermal comfort evaluation results may not be the primary optimization objective, but the weights of viscosity and temperature boundaries are increased. The decision-making agent 132 focuses on adjusting the compressor speed, pump frequency, and nanofluid circulation flow rate. The digital twin model 141 predicts temperature overshoot, pressure anomalies, and pump load trends caused by different candidate actions; if a candidate action can improve heat exchange in a short time, but will cause viscosity to exceed the boundary or pump work to increase significantly at the end of the prediction window, the safety constraint determination unit 142 excludes it.
[0047] Example 3: In the rail transit vehicle-mounted cooling scenario, the cooling system 100 includes an onboard compressor, a fan, a heat exchange module, and an edge controller. The data acquisition module 110 collects information on the temperature and humidity of the carriage, equipment power, and fluid parameters, as well as the train's operating status and vibration information. Due to the rapid fluctuations in the scenario, the feedback update module 160 can use a shorter update cycle, while the execution control module 150 emphasizes smooth operation. Through federated training and sharing among the building scenario node 171, the industrial scenario node 172, and the rail transit scenario node 173, the rail transit scenario can inherit some experience regarding load fluctuations and energy efficiency optimization from other scenarios without exposing its own original operating data.
[0048] In summary, this invention integrates four-parameter coupled modeling, multi-agent collaborative control, digital twin pre-deployment verification, Chinese population thermal comfort constraints, federated training, and modular deployment into a unified closed-loop chain, enabling multiple technical means that were originally scattered across different technical fields to form a clear functional cooperation relationship on the same refrigeration system optimization problem.
Claims
1. A method for optimizing the energy efficiency of a cross-scenario intelligent refrigeration system based on multi-agent collaborative reinforcement learning, characterized in that, Includes the following steps: S1. Obtain equipment operation data, environmental data, load data, and thermal conductivity, concentration, temperature, and viscosity parameters related to nanofluids in the target cooling scenario, and establish a four-parameter coupled model to reflect the heat transfer characteristics of nanofluids. S2. Based on the four-parameter coupling model, construct the state vector and constraint boundary, and establish a multi-agent collaborative control architecture consisting of a perception agent, a decision agent, a coordination agent, and an execution agent. The decision agent generates at least two sets of candidate control strategies. S3. Establish a digital twin model corresponding to the target cooling scenario, use the digital twin model to perform operating condition simulation and boundary verification on each candidate control strategy, and determine the target control strategy from the candidate control strategies that pass the verification. S4. The target control strategy is sent to the execution end of the refrigeration equipment to adjust at least one of the following operating parameters: compressor speed, pump frequency, fan speed, valve opening, and nanofluid circulation branch. S5. Collect feedback data after execution, and update the four-parameter coupling model, digital twin model and multi-agent cooperative control strategy to form a closed-loop optimization process of collection, modeling, simulation, decision-making, execution and feedback. The target cooling scenario can be any one of the following: building scenario, industrial scenario, rail transit scenario, and data center scenario. The optimization objective of the multi-agent collaborative control strategy includes at least the energy efficiency index of the cooling system.
2. The method according to claim 1, characterized in that, The four-parameter coupling model uses thermal conductivity k, nanoparticle volume fraction φ, operating temperature T, and dynamic viscosity μ as coupling variables. During the model establishment process, the coupling relationship is determined based on experimental sampling results and simulation fitting results, and the coupling relationship is corrected based on the concentration deviation and viscosity deviation obtained from online detection.
3. The method according to claim 1, characterized in that, The equipment operation data collected by the sensing agent in step S2 includes at least several items from the following: supply and return liquid temperature, evaporation temperature, condensation temperature, pressure, flow rate, compressor frequency, fan speed, and valve opening; the state vector also includes thermal conductivity, concentration, temperature, and viscosity characteristics output by the four-parameter coupling model.
4. The method according to claim 1, characterized in that, The reward function of the multi-agent collaborative control strategy includes at least an energy efficiency improvement component, an energy consumption penalty component, and a viscosity deviation penalty component; when the target cooling scenario is a residential environment-related scenario, the reward function also includes a thermal comfort satisfaction rate component.
5. The method according to claim 1, characterized in that, The digital twin model mentioned in step S3 includes a physical entity layer, a data mapping layer, a model layer, a simulation layer, and an execution feedback layer; the boundary verification includes at least two of the following: temperature boundary verification, pressure boundary verification, viscosity boundary verification, and rate of change of motion boundary verification.
6. The method according to claim 1, characterized in that, It also includes establishing a thermal comfort evaluation model suitable for Chinese people; the thermal comfort evaluation model combines at least several of the following factors: regional adaptation coefficient, metabolic rate, clothing thermal resistance, ambient temperature, ambient humidity and wind speed, to correct the thermal comfort evaluation results, and uses the corrected thermal comfort evaluation results as the state input in step S2 and the constraint conditions in step S3.
7. The method according to claim 1, characterized in that, In step S5, when updating the multi-agent collaborative control strategy, a cross-scenario federated training method is adopted. After each scenario node completes model training locally, it only uploads model parameters or gradient information. The aggregation end aggregates the uploaded content to complete the cross-scenario strategy update without sharing the original running data.
8. The method according to claim 7, characterized in that, The uploaded content is processed using at least one of homomorphic encryption and differential privacy before uploading, and at least one of quantization compression, sparse sampling, or asynchronous reporting is performed before or during uploading.
9. The method according to claim 1, characterized in that, The method is deployed in different cooling scenarios through modular hardware and protocol adaptation. The modular hardware includes a control module, a sensing module, an execution module, and an edge computing module. The protocol adaptation supports at least one of Modbus, Profinet, and EtherNet / IP.
10. The method according to claim 1, characterized in that, The operating parameters adjusted in step S4 include at least two of the following: compressor speed, pump frequency, fan speed, electronic expansion valve opening, nanofluid circulation flow rate, and bypass valve opening; the feedback update cycle in step S5 is set to seconds, minutes, or hours depending on the scenario type.