Subsystem cooperative control method and system and storage medium
By optimizing subsystem control through deep reinforcement learning models and distributed decision-making frameworks, the environmental adaptability and real-time performance issues of intelligent control systems are solved, achieving multi-objective optimization and efficient energy consumption management.
Patent Information
- Application Number
- CN202511075383.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-12-05
AI Technical Summary
Existing intelligent control systems have poor environmental adaptability in complex spaces, frequent conflicts between system devices, high decision-making delays and single objectives, obvious data asynchrony problems, rigid strategies, and difficulty in meeting real-time and multi-objective optimization requirements.
A deep reinforcement learning model is used for multi-objective optimization. By acquiring spatial control parameters, spatiotemporal alignment and feature fusion, a set of collaborative strategies is generated. The strategy is then decomposed into a sequence of subsystem control commands using a distributed decision framework and adjusted in real time to optimize energy consumption, comfort and equipment lifespan.
It significantly improved the Pareto frontier coverage of comfort scores, energy efficiency, and equipment lifespan, reduced energy consumption and equipment aging risks, and enhanced environmental responsiveness and control strategy precision.
Smart Images

Figure CN121069832A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of intelligent control, and particularly relates to a subsystem cooperative control method and system and a storage medium. BACKGROUND
[0002] With the rapid development of the Internet of Things and intelligent control technology, multi-subsystem cooperative control has become a core technical requirement in the fields of intelligent manufacturing, smart cities, etc. In the prior art, the following technical bottlenecks exist in the complex spatial multi-subsystem cooperative control:
[0003] 1. Poor environmental adaptability and frequent conflicts between system devices
[0004] According to relevant experimental data, in an open space with dynamic personnel distribution, the energy consumption fluctuation range of a control system based on traditional intelligent control rules reaches ± 35%, and the user comfort qualification rate can only be maintained at 58-72%; and when the fresh air system and the air purifier are turned on at the same time, due to the lack of coupling relationship modeling, the probability of air pressure imbalance in the space is as high as 41%
[0005] 2. High decision delay and single target
[0006] When the number of systems or devices in the space to be controlled reaches more than 20, the calculation time increases exponentially, resulting in a decision delay of the intelligent control system of more than 30s, which cannot meet the real-time requirement; at the same time, the traditional cooperative control system usually simply aims to minimize energy consumption, which expands the life span variance of each subsystem and device to 0.6-0.8, accelerating the aging of key subsystems and devices.
[0007] Currently, some patents also study intelligent control technology, and the following deficiencies exist after verification:
[0008] 1. Obvious data asynchrony problem
[0009] The acquisition of parameter data depends on various sensors, and when the network delay of the sensors exceeds 200ms, the probability of control strategy failure will increase to more than 25%;
[0010] 2. Strategy is rigid
[0011] Most of the existing patents use Q-learning algorithm, which shows obvious local optimal tendency in tests, and the Pareto frontier coverage rate can only reach about 63%.
[0012] At the same time, it is also pointed out in the ASHRAE 2023 annual report that the current intelligent device cooperative control is facing similar challenges as the above problems.
[0013] Therefore, it is an urgent problem for those skilled in the art to research a cooperative control method or system which can overcome the above defects. SUMMARY
[0014] To solve the above problems, the present application provides a subsystem cooperative control method, system and storage medium to solve the problems in the prior art.
[0015] In order to achieve the above-mentioned purposes of the application, the present application provides a subsystem cooperative control method, characterized in that the method comprises the following steps:
[0016] S1, obtaining the regulation and control parameters of the target space in which the subsystem is located;
[0017] S2, performing spatio-temporal alignment and feature fusion on the regulation and control parameters to generate a dynamic feature vector containing interaction influence factors of each subsystem;
[0018] S3, performing multi-objective optimization on the dynamic feature vector based on a deep reinforcement learning model to generate a cooperative strategy set containing subsystem control instructions, execution priorities and weight distribution;
[0019] S4, decomposing the cooperative strategy set into a sequence of subsystem control instructions through a distributed decision-making framework and performing dynamic feedback adjustment.
[0020] Further, the regulation and control parameters in step S1 include spatial physical environment parameters and user behavior parameters, the user behavior parameters are based on millimeter wave radar point cloud data to identify user position distribution, and are obtained through a motion capture algorithm.
[0021] Further, step S2 specifically comprises:
[0022] S21, using an attention mechanism to perform spatio-temporal alignment on the regulation and control parameters in step S1 to eliminate phase difference caused by delay during device communication;
[0023] S22, extracting time series features through a bidirectional gate unit cycle to generate a 128-dimensional feature vector by fusing subsystem state parameters;
[0024] S23, constructing an energy consumption coupling relationship graph between subsystems, calculating a subsystem interaction influence factor matrix, and obtaining a dynamic feature vector.
[0025] Further, step S3 specifically comprises:
[0026] S31, establishing a multi-objective optimization function containing comfort score, energy efficiency and life consumption rate based on a deep reinforcement learning model;
[0027] S32, searching for a Pareto optimal solution set in the action space using an improved DDPG algorithm to obtain a candidate strategy;
[0028] S33, performing virtual deduction on the candidate strategies, and selecting the one with the highest comprehensive score as the cooperative strategy set.
[0029] Further, step S4 specifically includes:
[0030] S41, decomposing the cooperative strategy set into subsystem control instruction sequences through a distributed decision-making framework;
[0031] S42, dynamically adjusting the sending timing according to the communication delay of the subsystems, and preferentially issuing the subsystem control instruction sequence with the highest weight allocation;
[0032] S43, monitoring the execution state of the subsystems in real time, and when an execution deviation is detected, updating the cooperative strategy parameters online using a sliding window.
[0033] Further, the spatiotemporal alignment of the regulation parameters in step S21 is performed according to the following calculation:
[0034] For the i-th sensor data stream, the time delay compensation formula is as follows:
[0035]
[0036] where Δt k is the communication delay set, and α k is the compensation coefficient calculated through the attention mechanism:
[0037]
[0038] where q is the query vector and W is the trainable weight matrix.
[0039] Further, the subsystem interaction influence factor matrix in step S23 is as follows:
[0040] For the subsystem set D = {d1, d2,..., d n}, the calculation formula of the element C ij in the coupling coefficient matrix C is as follows:
[0041]
[0042] where σ is the Sigmoid activation function, w k is the weight parameter, P i is the current power of the subsystem i, R ij is the functional correlation degree between subsystems, and D ij is the physical distance between subsystems.
[0043] Further, the comfort score in step S31 is calculated as follows:
[0044] Scomfort = K1*T + K2*H + K3*L + K4*A + K5*N
[0045] Wherein K1 is a temperature weight coefficient, T is temperature, K2 is a humidity weight coefficient, H is humidity, K3 is an illumination weight coefficient, L is illumination, K4 is an air quality weight coefficient, A is air quality, K5 is a noise weight coefficient, and N is a noise control degree.
[0046] The application also provides a subsystem cooperative control system for implementing the above method, and the system comprises:
[0047] A parameter acquisition module is configured to collect in real time regulation and control parameters including physical environment parameters and user behavior parameters in a target space;
[0048] A data analysis module is configured to perform spatiotemporal alignment and feature fusion on the regulation and control parameters to generate a dynamic feature vector;
[0049] A strategy generation module is capable of performing multi-objective optimization on the dynamic feature vector based on a deep reinforcement learning model, and is configured to generate a cooperative strategy set including subsystem control instructions, execution priorities and weight allocation;
[0050] A distributed control module is configured to decompose the cooperative strategy set into a subsystem control instruction sequence and distribute the subsystem control instruction sequence to each subsystem for execution.
[0051] The application also provides a computer storage medium storing program instructions for executing the above method.
[0052] Compared with the prior art, the application has the following beneficial effects:
[0053] 1. The application can improve the Pareto frontier coverage rate in the three-dimensional space of comfort score, energy efficiency and life consumption rate to 89%, and can simultaneously achieve dynamic balance of multiple competitive indexes while reducing the energy consumption of each subsystem.
[0054] 2. By performing spatiotemporal alignment on the regulation and control parameters, the application effectively solves the data asynchronization problem caused by network delay, significantly improves the control strategy accuracy, and has excellent environmental mutation response capability.
[0055] 3. The application has good space environment regulation capability in an open space with dynamic personnel distribution, and helps to improve the intelligent control technology of office, commercial and industrial smart scenes. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 A step flowchart of the subsystem cooperative control method of the application;
[0057] Figure 2The system structure diagram of the subsystem cooperative control system in the application. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical scheme and advantages of the application more clear, the application is further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the application and do not limit the application.
[0059] Example 1
[0060] The embodiment is a subsystem cooperative control method, and the method is as follows:
[0061] S1, obtaining the regulation and control parameters of the target space where the subsystem is located;
[0062] The regulation and control parameters include space physical environment parameters and user behavior parameters, the user behavior parameters are based on millimeter wave radar point cloud data to identify user position distribution, and the user behavior parameters are obtained through a motion capture algorithm
[0063] S2, performing space-time alignment and feature fusion on the regulation and control parameters to generate a dynamic feature vector containing interaction influence factors of each subsystem;
[0064] S21, using an attention mechanism to perform space-time alignment on the regulation and control parameters in step S1 to eliminate the phase difference caused by the delay when the devices communicate;
[0065] Wherein the space-time alignment of the regulation and control parameters is performed according to the following calculation:
[0066] For the i th sensor data stream, the time delay compensation formula is as follows:
[0067]
[0068] Where, Δt k is a set of communication delays, and a k is a compensation coefficient calculated through the attention mechanism:
[0069]
[0070] Wherein, q is a query vector, and W is a trainable weight matrix
[0071] S22, extracting time series features through a bidirectional gate unit cycle, and generating a 128-dimensional feature vector by fusing the subsystem state parameters;
[0072] S23, constructing an energy coupling relationship graph between the subsystems, calculating a subsystem interaction influence factor matrix, and obtaining a dynamic feature vector;
[0073] Wherein the subsystem interaction influence factor matrix is as follows:
[0074] For a subsystem set D = {d1, d2,..., d n}, the calculation formula of the element C ij in the coupling coefficient matrix C is as follows:
[0075]
[0076] Wherein, σ is a Sigmoid activation function, w k is a weight parameter, P i is the current power of the subsystem i, R ij is the functional correlation between subsystems, D ij is the physical distance between subsystems;
[0077] S3, multi-objective optimization of the dynamic feature vector based on the deep reinforcement learning model to generate a collaborative strategy set containing subsystem control instructions, execution priority and weight allocation;
[0078] S31, based on the deep reinforcement learning model, a multi-objective optimization function containing comfort score, energy efficiency and life consumption rate is established;
[0079] The comfort score is calculated as follows:
[0080] S comfort = K1·T + K2·H + K3·L + K4·A + K5·N
[0081] Wherein K1 is the temperature weight coefficient, T is the temperature, K2 is the humidity weight coefficient, H is the humidity, K3 is the illumination weight coefficient, L is the illumination, K4 is the air quality weight coefficient, A is the air quality, K5 is the noise weight coefficient, and N is the noise control degree;
[0082] S32, an improved DDPG algorithm is used to search for a Pareto optimal solution set in the action space to obtain a candidate strategy;
[0083] S33, the candidate strategy is virtually deduced, and the one with the highest comprehensive score is selected as the collaborative strategy set
[0084] S4, the collaborative strategy set is decomposed into a subsystem control instruction sequence through a distributed decision-making framework, and dynamic feedback adjustment is performed;
[0085] S41, the collaborative strategy set is decomposed into a subsystem control instruction sequence through a distributed decision-making framework;
[0086] S42, the sending timing is dynamically adjusted according to the communication delay of the subsystem, and the subsystem control instruction sequence with the highest weight allocation is preferentially issued;
[0087] S43, the execution state of the real-time monitoring subsystem is monitored, and when an execution deviation is monitored, the cooperative strategy parameters are updated online using a sliding window.
[0088] at 200 m 2 Taking office space as an example, the existing intelligent control equipment / system is improved according to the method in the above embodiment, and is continuously operated for 1 month, and the following comparison is obtained:
[0089] Indicators Conventional control method Embodiment method Lift range Daily energy consumption / kWh 82.3 63.7 22.6% Comfort score 68.2 89.5 31.2% Equipment life variance 0.47 0.28 40.4%
[0090] From the above table, it can be seen that the control method provided in the embodiment has significant advantages in the three competitive indicators of comfort score, energy efficiency and life consumption rate.
[0091] According to the daily average energy consumption reduction rate of 22.6%, if it can be applied to 100,000 m 2 Large commercial complex, it is expected to save electricity fee 2.18 million yuan per year (calculated at 1.2 yuan / kWh), carbon emission reduction reaches 620 tons per year, and due to the reduction of 40.4% in the life variance of the equipment, the inspection frequency of the related equipment and subsystem can be reduced from 2 times per day to 1-2 times per week, and the related maintenance cost is expected to be reduced by about 35%.
[0092] In addition to the above three competitive indicators, the other performance quantification indicators of the embodiment are compared with the existing intelligent control equipment / system as follows:
[0093]
[0094] The above table shows that each performance quantification indicator of the embodiment has been significantly optimized compared with the corresponding parameters of the traditional control method, and meets the industry requirements.
[0095] Embodiment 2
[0096] The embodiment provides a subsystem cooperative control system for realizing the method described in embodiment 1, as shown in the figure, the system comprises: Figure 2 As shown in the figure, the system comprises:
[0097] The parameter acquisition module is used for real-time acquisition of the regulation and control parameters including the physical environment parameters and the user behavior parameters in the target space;
[0098] The data analysis module is used for spatio-temporal alignment and feature fusion of the regulation and control parameters to generate a dynamic feature vector;
[0099] The strategy generation module can perform multi-objective optimization on the dynamic feature vector based on a deep reinforcement learning model, and is used for generating a cooperative strategy set comprising subsystem control instructions, execution priority and weight allocation;
[0100] A distributed control module is configured to decompose the cooperative strategy set into subsystem control instruction sequences and distribute the subsystem control instruction sequences to each subsystem for execution.
[0101] Embodiment 3
[0102] The embodiment provides a computer storage medium, which stores program instructions, and the program instructions control a device where the computer storage medium is located to execute the method described above when the program instructions are executed.
[0103] It should be understood that, although each step in the flowchart of each embodiment of the present application is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in each embodiment can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.
[0104] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The above-mentioned program can be stored in a non-volatile computer readable storage medium, and the program can include the processes of the above-mentioned embodiment methods when executed. Any reference to memory, storage, database or other medium used in each embodiment provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0105] Any combination of the technical features in the above-described embodiments can be made, and for the sake of brevity, not all possible combinations are described, however, as long as there is no conflict, any combination of the technical features should be considered within the scope of the present disclosure.
[0106] The above-described embodiments only express several embodiments of the present application, which are described in a more specific and detailed manner, but should not be understood as limiting the scope of the patent of the present application. It should be noted that for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, which are all within the scope of protection of the present application. Therefore, the scope of protection of the patent of the present application should be subject to the appended claims.
Claims
1. A subsystem cooperative control method, characterized in that: The method is as follows: S1. Obtain the control parameters of the target space where the subsystem is located; S2. Perform spatiotemporal alignment and feature fusion on the control parameters to generate a dynamic feature vector containing the interaction factors of each subsystem; S3. Based on a deep reinforcement learning model, perform multi-objective optimization on dynamic feature vectors to generate a set of collaborative strategies that includes subsystem control instructions, execution priorities, and weight allocation. S4. The collaborative strategy set is decomposed into a sequence of subsystem control instructions through a distributed decision-making framework, and dynamic feedback adjustment is performed.
2. The subsystem cooperative control method according to claim 1, characterized in that: The control parameters in step S1 include spatial physical environment parameters and user behavior parameters. The user behavior parameters are based on millimeter-wave radar point cloud data to identify the user's location distribution and are obtained through motion capture algorithms.
3. The subsystem cooperative control method according to claim 1, characterized in that: Step S2 specifically includes: S21. Use an attention mechanism to perform spatiotemporal alignment on the control parameters in step S1 to eliminate the phase difference caused by delay during device communication. S22. Extract time series features cyclically through bidirectional gating units and fuse subsystem state parameters to generate a 128-dimensional feature vector; S23. Construct an energy consumption coupling relationship map between subsystems, calculate the interaction influence factor matrix of subsystems, and obtain dynamic feature vectors.
4. The subsystem cooperative control method according to claim 1, characterized in that: Step S3 specifically includes: S31. Establish a multi-objective optimization function based on a deep reinforcement learning model, which includes comfort score, energy efficiency and life loss rate. S32. Use the improved DDPG algorithm to search for Pareto optimal solution sets in the action space to obtain candidate strategies; S33. Perform virtual simulations on the candidate strategies and select the one with the highest comprehensive score as the collaborative strategy set.
5. The subsystem cooperative control method according to claim 1, characterized in that: Step S4 specifically includes: S41. The collaborative strategy set is decomposed into a sequence of subsystem control instructions through a distributed decision-making framework. S42. Dynamically adjust the transmission timing according to the subsystem communication delay, and prioritize the transmission of the subsystem control instruction sequence with the highest weight allocation; S43. Monitor the execution status of the subsystem in real time. When an execution deviation is detected, use a sliding window to update the collaborative strategy parameters online.
6. The subsystem cooperative control method according to claim 3, characterized in that: In step S21, the control parameters are spatiotemporally aligned according to the following calculation: For the data stream of the i-th sensor, the delay compensation formula is as follows: Where, Δt k Let α be the set of communication delays. k The compensation coefficients are calculated using the attention mechanism: Where q is the query vector and W is the trainable weight matrix.
7. The subsystem cooperative control method according to claim 3, characterized in that: The subsystem interaction influence factor matrix in step S23 is as follows: For the subsystem set D = {d1, d2, ..., d...} n }, the elements C in its coupling coefficient matrix C ij The calculation formula is as follows: Where σ is the Sigmoid activation function, w k P is the weight parameter. i R is the current power of subsystem i. ij D represents the functional correlation between subsystems. ij This refers to the physical distance between subsystems.
8. The subsystem cooperative control method according to claim 4, characterized in that: The comfort score in step S31 is calculated as follows: S comfort =K1·T+K2·H+K3·L+K4·A+K5·N Where K1 is the temperature weighting coefficient, T is the temperature, K2 is the humidity weighting coefficient, H is the humidity, K3 is the illuminance weighting coefficient, L is the illuminance, K4 is the air quality weighting coefficient, A is the air quality, K5 is the noise weighting coefficient, and N is the noise control degree.
9. A subsystem cooperative control system for implementing the method as described in any one of claims 1 to 8, characterized in that: The system includes: The parameter acquisition module is used to collect control parameters, including physical environment parameters and user behavior parameters, within the target space in real time. The data analysis module is used to perform spatiotemporal alignment and feature fusion of the control parameters to generate dynamic feature vectors; The policy generation module can perform multi-objective optimization on dynamic feature vectors based on a deep reinforcement learning model to generate a set of collaborative policies that includes subsystem control instructions, execution priorities, and weight allocation. The distributed control module is used to decompose the set of collaborative strategies into a sequence of subsystem control instructions and distribute them to each subsystem for execution.
10. A computer storage medium, characterized in that: The computer storage medium stores program instructions for performing the method according to any one of claims 1 to 8.