Building operation and maintenance agent strategy updating method and storage medium

CN121920409AActive Publication Date: 2026-04-24SHANGHAI CONSTRUCTION FOURTH CONSTRUCTION GROUP CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI CONSTRUCTION FOURTH CONSTRUCTION GROUP CO LTD
Filing Date
2025-12-23
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing building operation and maintenance management systems, the credit allocation among intelligent agents is unfair, making it difficult to assess the value of cross-system collaboration and lacking the ability to model dynamic collaborative relationships, resulting in unfair and inefficient optimization of operation and maintenance strategies.

Method used

By introducing Shapley value theory from cooperative game theory, and calculating the Shapley credit value of each agent by training the joint action value function, and combining the operation and maintenance collaboration affinity matrix and adaptive sampling strategy, the action selection of agents is optimized to achieve fair credit allocation and cross-system collaborative optimization.

Benefits of technology

It improves the fairness and efficiency of building operation and maintenance, can dynamically adapt to building usage patterns and seasonal changes, discover and strengthen effective cross-system collaboration patterns, and achieve comprehensive optimization of energy consumption, comfort and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121920409A_ABST
    Figure CN121920409A_ABST
Patent Text Reader

Abstract

The invention provides a building operation and maintenance intelligent agent strategy updating method and a storage medium, and the method fully considers cooperation values among subsystems of heating, ventilation and air conditioning, security and protection monitoring, equipment maintenance, environment control and the like, and accurately quantifies the contribution of each professional intelligent agent to the overall operation effect of a building. A Shapril value theory in a cooperative game theory is introduced into a building operation and maintenance CTDE framework, theoretically fair credit evaluation is provided for each operation and maintenance agent, and the fairness problem of a traditional method in multi-target tradeoff is solved; aiming at the situation that a large building possibly comprises dozens of professional agents, an efficient Shapril value approximate calculation method is provided, so that the method has feasibility in an actual building operation and maintenance system; through accurate credit distribution, strategy optimization of each operation and maintenance agent is guided, and comprehensive optimization of building energy consumption, comfort, safety and maintenance cost is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for updating building operation and maintenance intelligent agent strategies and a storage medium. Background Technology

[0002] In intelligent building operation and maintenance management systems, specialized intelligent agents such as security, energy, equipment, and environment need to collaborate to achieve multi-objective optimization of efficiency, safety, and energy conservation. However, the strong coupling of the system leads to severe challenges in credit allocation. There is cross-system collaboration complexity among subsystems; for example, the linkage between air conditioning, lighting, ventilation, and security can produce a synergistic effect of "1+1>2." Traditional counterfactual methods (such as COMA) only assess the marginal contribution of individuals, making it difficult to capture the collaborative value of the alliance. Simultaneously, preventative maintenance strategies have time delays and long-term benefits, making immediate reward mechanisms prone to underestimating their contribution. Furthermore, there are conflicting trade-offs among multiple objectives; improved comfort may increase energy consumption, requiring a fair assessment of the actual role of each intelligent agent in complex decision-making. Existing CTDE frameworks rely on centralized BMS coordination but lack the ability to model dynamic collaborative relationships, making it difficult to adapt to the evolving operating environment caused by building usage patterns, seasonal changes, and equipment aging. Summary of the Invention

[0003] The purpose of this invention is to provide a method for updating building operation and maintenance intelligent agent strategies and a storage medium.

[0004] To address the above problems, this invention provides a method for updating the strategy of a building operation and maintenance intelligent agent, comprising: Based on trainable parameters, train the function for the current task, and based on the completed training function, obtain the joint action value function of the intelligent agent of the building operation and maintenance system for the current task. Based on the joint action value function of the agents, calculate the Shapley credit value of each agent; Based on each agent's Shapley credit score, efficient and compliant actions are selected for each agent.

[0005] Furthermore, in the above method, based on trainable parameters, a function for the current task is trained; based on the trained function, the joint action value function of the building operation and maintenance system agent for the current task is obtained, including: The function of the current task for: ; in, The comprehensive reward for building operation and maintenance includes a multi-dimensional evaluation of energy efficiency, user comfort, safety indicators, and equipment health. It is a discount factor used to balance the weight of current rewards and future rewards; and These are the current global state and the combined action, respectively. This represents the value function of the joint actions at the current moment; the building operation and maintenance system contains N intelligent agents, each of which... Based on its own local observations Generate action During the training phase, the global state is received. Joint actions with all specialized intelligent agents It outputs a joint action value function at the current moment. ; and These are the global state and the combined action at the next moment, respectively. The value function representing the joint action at the next moment; Represents the value function of joint actions The model parameters can be used to train the parameters. It is by The determined value assessment function is optimized. ,let The valuation results are more in line with the actual benefits of building operation and maintenance; It represents the mathematical expectation.

[0006] Furthermore, in the above method, based on the joint action value function of the agents, the Shapley credit score of each agent is calculated, including: Based on the preset baseline strategy and joint action value function, the marginal contribution of the agent in the building operation and maintenance subsystem alliance is obtained; Based on the comprehensive reward for building operation and maintenance, the operation and maintenance collaboration affinity matrix among intelligent agents is obtained; Based on the marginal contribution of each agent in the building operation and maintenance subsystem alliance, the operation and maintenance collaboration affinity matrix between agents, and the preset adjustment terms, the adaptive sampling number of the agent is obtained. The Shapley credit value for each agent is obtained based on the marginal contribution of each agent in the building operations and maintenance subsystem alliance and the adaptive sampling number of each agent.

[0007] Furthermore, in the above method, based on the preset baseline strategy and joint action value function, the marginal contribution of the agent in the building operation and maintenance subsystem alliance is obtained, including: Based on preset baseline strategy and joint action value function To obtain the intelligent agent In the Building Operation and Maintenance Subsystem Alliance Marginal contribution The formula is as follows: ; Among them, baseline strategy This is the default policy; Each agent Each has a corresponding active current operation and maintenance policy; Indicates alliance The specialized operation and maintenance agent in the system executes the current operation and maintenance strategy; Indicates alliance All external agents adopt the building operations and maintenance baseline strategy. ; Indicates alliance Internal intelligent agents Other members execute the current operation and maintenance policy; Indicates that the intelligent agent The action replacement is its building operation and maintenance baseline strategy. ; Represents intelligent agents In the Building Operation and Maintenance Subsystem Alliance S, the marginal contribution of building operation and maintenance to the overall value of the system. The marginal contribution of building operation and maintenance is reflected by the value difference between the two scenarios, which is used by the intelligent agent. Function: First item Indicates the presence of intelligent agents The overall value of the system when the entire building operation and maintenance subsystem S executes the current strategy; Second item This indicates that only intelligent agents exist in the Building Operation and Maintenance Subsystem Alliance S. The overall value of the system when switching to the baseline strategy; and The difference is the agent's value. The additional value increment brought to the system by executing the current strategy.

[0008] Furthermore, in the above method, based on the comprehensive building operation and maintenance reward, an operation and maintenance collaboration affinity matrix is ​​obtained among professional operation and maintenance agents, including: Based on comprehensive building operation and maintenance rewards Obtain the operation and maintenance collaboration affinity matrix among professional operation and maintenance intelligent agents. The formula is as follows: ; in, The comprehensive reward for building operation and maintenance includes a multi-dimensional evaluation of energy efficiency, user comfort, safety indicators, and equipment health. For a moment At that time, intelligent agent With intelligent agents The degree of cooperation and affinity between them; For a moment At that time, intelligent agent With intelligent agents The degree of cooperation and affinity between them; It is to utilize the previous moment Affinity, combined arrive The actual interaction data of the stage is used to calculate the result through the update formula; This serves as a weighting factor, used to balance the impact of historical affinity and current collaborative contribution; This is a building operations and maintenance collaboration indicator function used to determine the first... In this interaction, the intelligent agent and Whether collaboration has occurred is indicated by a value of 1 if yes and 0 otherwise. Among them, building operation and maintenance collaboration indication function Based on the following building operation and maintenance scenario definition: When the energy-saving strategy of the energy control agent and the comfort control of the environmental condition agent cooperate, i.e., when collaboration occurs... =1; When the personnel detection of the security management intelligent agent and the area control linkage of the facility management intelligent agent occur, i.e., when collaboration takes place... =1; When the coordination between the equipment maintenance agent and the energy control and management agent in preventive maintenance, i.e., when collaboration occurs, =1.

[0009] Furthermore, in the above method, after obtaining the operational collaboration affinity matrix among professional operational intelligence agents, the following steps are also included: According to building usage patterns Dynamically adjust the operation and maintenance collaboration affinity matrix Update frequency: .

[0010] Furthermore, in the above method, based on the marginal contribution of each agent in the building operation and maintenance subsystem alliance, the operation and maintenance collaboration affinity matrix between agents, and a preset adjustment term, the adaptive sampling number of the agent is obtained, including: Based on the operational collaboration affinity matrix among intelligent agents and preset adjustment items To obtain the operation and maintenance intelligent agent Adaptive sampling number The formula is as follows: ; in, For intelligent agents The total number of samples, which is composed of , and This is obtained by adding the three sampling times together; The baseline number of samples is given for all agents, where the sampling probability P based on layer 1 is... layer1 =0.5, sampling probability P of layer 2 layer2 =0.3, sampling probability P of layer 3 layer3 =0.2 and the baseline sampling number are used to assign corresponding sampling numbers to agents at each layer; layer 1 includes: security management, energy control and environmental control agents; layer 2 includes: equipment maintenance and facility management agents; layer 3 includes: cleaning, parking and communication agents; For intelligent agents The variance of operational performance; The variance of the overall system operation and maintenance performance; and Based on the operation and maintenance collaboration affinity matrix Perform calculations; The additional sampling coefficient represents the marginal contribution of building operations and maintenance based on each agent. The marginal contribution variance of a single agent and the marginal contribution variance of the entire system are obtained; based on the marginal contribution variance of a single agent and the marginal contribution variance of the entire system, the coefficient of the additional sampling number is obtained. ; As a seasonal adjustment item, the number of samplings is increased during seasons with significant changes in building load.

[0011] Furthermore, in the above method, based on the marginal contribution of each agent in the building operation and maintenance subsystem alliance, the operation and maintenance collaboration affinity matrix between agents, and a preset adjustment term, the adaptive sampling number of the agent is obtained, including: The calculated seasonal parameters Used to adjust seasonal adjustment items ,Right now: , ; in, Baseline parameter values; It indicates the day of the year.

[0012] Furthermore, in the above method, based on each agent's marginal contribution in the building operations and maintenance subsystem alliance and the adaptive sampling number of each agent, the Shapley credit value of each agent is obtained, including: The Shapley credit value for each agent is obtained using the following formula. : ; in, For intelligent agents In the Building Operation and Maintenance Subsystem Alliance Marginal contribution in; Building Operation and Maintenance Subsystem Alliance There is One intelligent agent; For intelligent agents The adaptive number of samplings.

[0013] Furthermore, in the above method, based on the Shapley credit score of each agent, efficient and compliant actions are selected for each agent, including: The global value of each agent is calculated based on the Shapley credit value and joint action value function of each operation and maintenance agent. Based on the global value of each agent, efficient and compliant actions are selected for each agent.

[0014] Furthermore, in the above method, based on the Shapley credit score and joint action value function of each operational agent, the global value of each agent is calculated, including: Shapley credit score fitted based on value network intelligent agent The individual value network loss function is: ; in, For intelligent agents The local action value function, based on its own observation and actions Evaluate the value of actions; after value network optimization, Accurately reflects the intelligent agent The overall value; The Shapley's estimate of agent i is used to quantify the agent. Contributions in collaboration; These are the model parameters of the value network for agent i.

[0015] Furthermore, in the above method, based on the global value of each agent, efficient and compliant actions are selected for each agent, including: The loss function of the network based on the following policy Select efficient and compliant actions for each intelligent agent: ; Policy gradient term middle, Agent i is observing Select action The probability of the strategy; It is an intelligent agent. The local action value function; Operation and maintenance constraints middle, : These are building operation and maintenance constraints, including: safety constraints, energy consumption limits, and comfort requirements.

[0016] According to another aspect of the present invention, a computer-readable storage medium is also provided, having stored thereon computer-executable instructions, wherein when executed by a processor, the computer-executable instructions cause the processor to perform the method described in any of the preceding claims.

[0017] Compared with the prior art, the present invention has the following advantages: (1) Improved fairness of building operation and maintenance: Credit allocation based on Shapley value ensures that the contribution of each operation and maintenance subsystem is fairly evaluated, and avoids underestimation of some key but insignificant systems (such as preventive maintenance).

[0018] (2) Cross-system collaboration optimization: By evaluating the value of an alliance of operation and maintenance subsystems of any size, effective cross-system collaboration models can be discovered and strengthened, such as the joint energy-saving strategy of HVAC and smart lighting.

[0019] (3) Improved building operation and maintenance efficiency: The special design for building operation and maintenance scenarios enables each intelligent agent to learn strategies that are more in line with the actual building operation needs, significantly improving the overall operation and maintenance efficiency.

[0020] (4) Enhanced dynamic adaptability: Considering the seasonal and cyclical changes in building usage patterns, the credit allocation mechanism can adapt to changes in collaborative relationships under different operating scenarios. Attached Figure Description

[0021] Figure 1 This is a flowchart of a building operation and maintenance intelligent agent strategy update method according to an embodiment of the present invention. Detailed Implementation

[0022] The present invention will now be described in further detail with reference to the accompanying drawings.

[0023] In a typical configuration of this application, the terminal, the device of the service network, and the trusted party all include one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0024] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0025] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.

[0026] There is an urgent need for a credit allocation mechanism that can comprehensively evaluate the contributions of subsystem alliances of any size and integrate long-term effects and dynamic adaptability, so as to accurately identify the true value of agents in complex coupled environments and improve the fairness and optimization efficiency of overall operation and maintenance decisions.

[0027] This invention aims to address the technical problem of existing credit allocation methods in building operation and maintenance multi-agent systems failing to adequately assess cross-system collaborative relationships, leading to unfair optimization of operation and maintenance strategies. Specifically, it addresses the unique needs of the building operation and maintenance field, with the following objectives: (1) Fully consider the collaborative value among subsystems such as HVAC, security monitoring, equipment maintenance, and environmental control, and accurately quantify the contribution of each professional intelligent agent to the overall building operation effect.

[0028] (2) The Shapley value theory in cooperative game theory is introduced into the CTDE framework of building operation and maintenance, providing theoretically fair credit assessment for each operation and maintenance agent, and solving the fairness problem of traditional methods in multi-objective trade-offs.

[0029] (3) For large buildings that may contain dozens of specialized intelligent agents, an efficient method for approximating Shapley values ​​is provided to make it feasible in actual building operation and maintenance systems.

[0030] (4) By accurately allocating credits, guide the strategy optimization of each operation and maintenance agent to achieve comprehensive optimization of building energy consumption, comfort, safety and maintenance costs.

[0031] like Figure 1 As shown, the present invention provides a method for updating the strategy of a building operation and maintenance intelligent agent, the method comprising: Step 1: Construct a building operations and maintenance manager, based on trainable parameters. The function to train the current task Based on training completion function The joint action value function of the building operation and maintenance system agents for the current task is obtained. ; Here, the current task, such as detecting that the temperature inside the building is too high, requires adjusting the temperature; The building operation and maintenance system comprises N specialized operation and maintenance intelligent agents, covering security management intelligent agents, energy control intelligent agents, equipment maintenance intelligent agents, environmental control intelligent agents, facility management intelligent agents, etc., each intelligent agent... Based on its own local observations Generate action During the training phase, the building operations supervisor receives the global status. (or the collection of local observations from all agents) and the joint actions of all specialized agents. It outputs a joint action value function at the current moment. Joint action value function To evaluate the overall quality of the current joint actions, the building operations manager updates the trainable parameters by minimizing the temporal difference (TD) error. : ; in, The comprehensive reward for building operation and maintenance includes multi-dimensional evaluations such as energy efficiency, user comfort, safety indicators, and equipment health. It is a discount factor used to balance the weight of current rewards and future rewards; and These are the current global state and the combined action, respectively. This represents the value function of the joint actions at the current moment; and These are the global state and the combined action at the next moment, respectively. The value function representing the joint action at the next moment; Represents the value function of joint actions The model parameters can be used to train the parameters. It is by The determined value assessment function is optimized. ,let The valuation results are more in line with the actual benefits of building operation and maintenance; It is a function used to evaluate the global state. Execute joint actions At that time, the overall value obtained by the building operation and maintenance system integrates benefits from multiple dimensions such as energy consumption, comfort, and safety; while It's this function. The trainable parameters of the corresponding model (such as a neural network) are obtained by minimizing the loss function of the temporal difference error. To update these trainable parameters, thereby improving the joint action value function. It can more accurately assess the overall advantages and disadvantages of coordinated actions; It represents mathematical expectation.

[0032] The output of this step This will serve as the value benchmark function for the cooperative game in step 2.

[0033] Specifically, the building operations manager is responsible for overall coordination, resource allocation, strategy optimization, and performance monitoring.

[0034] Professional operation and maintenance intelligent agents may include: Security management intelligent agent: responsible for access control, monitoring management, intrusion detection, and emergency response; Energy control intelligent agent: responsible for air conditioning control, lighting management, power dispatching, and energy-saving optimization; Environmental condition agent: responsible for air quality control, temperature and humidity regulation, and ventilation management; Equipment maintenance intelligent agent: responsible for elevator dispatching, water supply control, equipment maintenance, and fault early warning; The global state This refers to the overall status parameters of the building's operation and maintenance system, such as: Total energy consumption of the building as a whole and grid load; Environmental parameters such as temperature, humidity, and air quality in all areas of the building; The operational status and health of all building equipment (such as air conditioners and elevators); The overall security level of the security system (such as whether there are any abnormal alarms). This is status information directly collected from a "system-wide perspective".

[0035] When the global state cannot be directly obtained, local observations from all specialized agents can also be used. A set to replace the global state .

[0036] Local observations of each agent It refers to local information within its scope of responsibility, such as energy control intelligent agents. It is regional energy consumption data, and an intelligent environmental control agent. It refers to regional temperature and humidity. By aggregating the local observations of all intelligent agents, we can stitch together the "panoramic information" of the system, which is equivalent to the global state. .

[0037] Whether it's directly collecting the global state or combining the local observations of all agents, the goal is to enable building operations managers to obtain complete operational information of the system, thereby accurately assessing the global value of the joint actions of multiple agents.

[0038] Preferred, global state The dimensions are 75, including: Environmental parameters (20 dimensions): temperature, humidity, CO2 concentration, etc. in each area.

[0039] Energy consumption data (15 dimensions): real-time power consumption, peak and off-peak electricity prices, etc.

[0040] Security status (15 dimensions): access control status, personnel flow, etc.

[0041] Equipment status (15 dimensions): equipment health, operating status, etc.

[0042] Service metrics (10 dimensions): user satisfaction, response time, etc.

[0043] The action space of each intelligent agent is set according to its responsibilities.

[0044] Central Value Network The structure is [75, 512, 256, 128, 1], with a learning rate of... .

[0045] During the training phase, the building operations supervisor receives the current global status. Joint actions of all intelligent agents Output value of joint actions .

[0046] The training process is performed by minimizing the TD error: The value assessment function is determined by optimization. ,let The valuation results are more in line with the actual benefits of building operation and maintenance. Among them, comprehensive rewards It is calculated using a multi-dimensional weighted approach.

[0047] After this step is completed, the trained... The function will serve as the value benchmark for the subsequent step 2, Shapley value calculation.

[0048] Step 2: Building Operation and Maintenance Shapley Credit Allocation Model: Incorporating the Joint Action Value Function As a characteristic function of cooperative game theory in building operation and maintenance, each professional operation and maintenance agent is considered as a player in the game, and the calculation of each agent's characteristics is performed. Shapley credit rating .

[0049] Here, step 2 utilizes the data generated in step 1. A function that calculates the contribution credit for each agent.

[0050] Preferably, each agent is computed. Shapley credit rating The process includes the following three interrelated sub-steps: Step 2.1: Definition of Building Operation and Maintenance Marginal Contribution: Based on Preset Baseline Strategy The joint action value function obtained in step 1 To obtain the intelligent agent In the Building Operation and Maintenance Subsystem Alliance Marginal contribution ; Preferably, based on a preset baseline strategy The joint action value function obtained in step 1 For intelligent agents In the Building Operation and Maintenance Subsystem Alliance Marginal contribution Its calculation directly depends on the joint action value function in step 1. and baseline strategy : ; Among them, baseline strategy This is the default policy; Each agent Each has a corresponding active current operation and maintenance policy; Indicates alliance The specialized operation and maintenance agent in the system executes the current operation and maintenance strategy; Indicates alliance All external agents adopt the building operations and maintenance baseline strategy. ; Indicates alliance Internal intelligent agents Other members execute the current operation and maintenance policy; Indicates that the intelligent agent The action replacement is its building operation and maintenance baseline strategy. ; Represents intelligent agents In the Building Operation and Maintenance Subsystem Alliance S, the marginal contribution of building operation and maintenance to the overall value of the system. The marginal contribution of building operation and maintenance is reflected by the value difference between the two scenarios, which is used by the intelligent agent. Function: First item Indicates the presence of intelligent agents The overall value of the system when the entire building operation and maintenance subsystem S executes the current strategy; Second item This indicates that only intelligent agents exist in the Building Operation and Maintenance Subsystem Alliance S. The overall value of the system when switching to the baseline strategy; The difference between the two is the agent's value. The larger the difference between the additional value increment brought to the system by executing the current strategy, the more significant the difference in the intelligent agent's performance. The more significant the marginal contribution, the better.

[0051] Step 2.2: Building Operation and Maintenance Collaboration Affinity Modeling: Building Operation and Maintenance Comprehensive Rewards Based on Step 1 Obtain the operation and maintenance collaboration affinity matrix among professional operation and maintenance intelligent agents. ; Here, The comprehensive reward for building operation and maintenance includes multi-dimensional evaluations such as energy efficiency, user comfort, safety indicators, and equipment health. Considering the collaborative characteristics among subsystems in a building operation and maintenance system, an operation and maintenance collaboration affinity matrix is ​​defined among professional operation and maintenance agents. Operation and maintenance collaboration affinity matrix The update directly depends on the building operation and maintenance comprehensive reward in step 1. This integrates real-time performance feedback into collaborative memory. ; in, For a moment At that time, intelligent agent With intelligent agents The degree of cooperation and affinity between them; For a moment At that time, intelligent agent With intelligent agents The degree of cooperation and affinity between them; This serves as a weighting factor, used to balance the impact of historical affinity and current collaborative contribution; This is a building operations and maintenance collaboration indicator function used to determine the first... In this interaction, the intelligent agent and Whether collaboration has occurred is indicated by a value of 1 if yes and 0 otherwise. Among them, building operation and maintenance collaboration indication function Based on the following building operation and maintenance scenario definition: Energy-Environment Collaboration: When energy-saving strategies of the energy control agent and comfort control of the environmental condition agent are coordinated, i.e., when collaboration occurs... =1; Security-Facility Collaboration: When personnel detection by the security management agent and area control by the facility management agent are linked, collaboration occurs. =1; Maintenance-Energy Collaboration: This refers to the coordination between equipment maintenance agents and energy control and management agents during preventative maintenance, i.e., when collaboration occurs. =1.

[0052] It is to utilize the previous moment Affinity, combined arrive The actual interaction data of the stage is used to calculate the result through the update formula; Specifically, The acquisition of (the cooperative affinity between agents i and j at time t) is divided into an initial stage and an iterative update stage, which is an initialization → dynamic iteration process: 1) Initial stage ( When =0): Based on domain knowledge preset; At the initial moment of system startup ( =0), These are initial values ​​directly set based on domain experience in building operation and maintenance, for example: This corresponds to the initial cooperative relationship among the five agents, and the numerical value reflects the pre-defined degree of cooperation between different agent pairs; 2) Iterative update phase ( When >=1): This is obtained through the update formula from the previous time step; when When >=1 It utilizes the previous moment. Affinity, combined arrive The actual interaction data of the stage is used to calculate the result through the update formula.

[0053] by For example, =1: ,in: It refers to the initial affinity; , yes =0 to =Interaction data within phase 1 (effective collaboration markers + comprehensive rewards).

[0054] so, The acquisition logic is: initial time ( =0): Preset initial value for domain knowledge; Subsequent moments ( >=1): Based on the affinity at the previous moment. It is obtained by iterative calculation using the affinity update formula, combined with the latest interaction data.

[0055] Preferably, the operation and maintenance collaboration affinity matrix can be dynamically adjusted according to the building's usage patterns (weekdays / holidays, daytime / nighttime). Update frequency: ; Operation and maintenance collaboration affinity matrix Every An update is performed only once per step, ensuring that its update rhythm is synchronized with the building's operating mode.

[0056] Step 2.3: Building Operations-Oriented Monte Carlo Sampling: Based on the marginal contribution of Step 2.1 and the operation and maintenance collaboration affinity matrix among professional operation and maintenance agents in Step 2.2 Based on the preset adjustment items, an operation and maintenance intelligent agent is obtained. Adaptive sampling number ; Preferably, considering the layered characteristics of building operation and maintenance systems, a layered federated sampling strategy is adopted: The first layer of core operation and maintenance system is layer 1, which includes: security management, energy control, and environmental control intelligent agents; The second layer of support and maintenance system is layer 2, which includes: equipment maintenance and facility management intelligent agents; The third layer of auxiliary operation and maintenance system is layer 3, which includes intelligent agents such as cleaning, parking, and communication.

[0057] Sampling probabilities are weighted according to the importance of building operation and maintenance: the sampling probability P of layer 1 layer1 =0.5, sampling probability P of layer 2 layer2 =0.3, sampling probability P of layer 3 layer3 =0.2.

[0058] For operation and maintenance intelligent agents Adaptive sampling number The calculation is as follows, which includes the operation and maintenance collaboration affinity matrix between professional operation and maintenance agents from step 2.2. And preset adjustment items: ; in, For intelligent agents The total number of samples, which is composed of , and This is obtained by adding the three sampling times together; The baseline number of samples for all agents, where, based on P layer1 =0.5, P layer2 =0.3, P layer3 =0.2 and the baseline sampling number are used to assign corresponding sampling numbers to agents at each level; specifically This ensures the basic reliability of Shapley value calculation. The hierarchical sampling probability allocation strategy of Player1=0.5, Player2=0.3, Player3=0.2 guides computing resources to be prioritized for evaluating the contribution of the core system alliance based on the importance and collaborative characteristics of each agent in the building operation and maintenance system. This significantly improves the efficiency and rationality of credit allocation while ensuring fairness. It is a preset constant that represents the base number of samplings when all operation and maintenance agents perform Shapley credit calculations.

[0059] For intelligent agents The variance of operational performance; The variance of the overall system operation and maintenance performance; and Based on the operation and maintenance collaboration affinity matrix in step 2.2 Perform calculations; intelligent agent Operational performance variance Calculation of: Operation and Maintenance Collaboration Affinity Matrix Describes intelligent agents With all other intelligent agents The degree of close collaboration. First, extract the intelligent agents. Cooperation affinity values ​​with all other agents: i.e., the cooperation affinity matrix The Middle All elements of the row N is the total number of agents; Then, calculate the variance of this set of affinity values: the variance reflects the agent's... The greater the variance in the cooperative relationships with different agents, the more volatile the relationship between the agents. The greater the difference in the closeness of cooperation with other intelligent agents, the weaker the stability of its operation and maintenance performance.

[0060] The formula can be expressed as: ,in, It is an intelligent agent The average level of affinity with all agents.

[0061] Variance of overall system operation and maintenance performance The calculation of the overall system variance: The overall system variance is the statistical value of the variance of the individual operational performance of all agents, usually taken as the mean or the variance of the overall distribution. First, calculate the operational performance variance of each agent k. The method is the same as above. Then, calculate the variances of these individuals. The average value or overall variance is used to represent the degree of fluctuation in the overall operational performance of the system. The formula can be expressed as: .

[0062] The additional sampling number coefficient represents the marginal contribution of each agent to building maintenance calculated in section 2.1. The marginal contribution variance of a single agent and the marginal contribution variance of the entire system are obtained; based on the marginal contribution variance of a single agent and the marginal contribution variance of the entire system, the coefficient of the additional sampling number is obtained. .

[0063] Based on the ratio of the marginal contribution variance of an individual agent to the marginal contribution variance of the entire system, determine The specific values ​​are as follows: agents with larger fluctuations in marginal contribution correspond to higher values. This allows for the acquisition of more additional sampling resources. The design aims to make the number of samples "adapt to the stability of the agent's contribution": if the marginal contribution of the agent fluctuates greatly (high variance), it indicates that the uncertainty of its contribution is strong, and more sampling is needed to accurately assess its Shapley credit; if the marginal contribution of the agent fluctuates little (low variance), then there is no need for too many additional samples, thus avoiding waste of resources.

[0064] Specifically, Calculate using the following steps: Step 1): Calculate the marginal contribution of a single agent. ; First, based on the definition of marginal contribution of the Shapley value, calculate the marginal contribution of agent i in the consortium S: ; The building operation and maintenance value of Alliance S (a collaborative group composed of some intelligent agents) is determined by the joint action value function. calculate); The value of the alliance S after the addition of agent i.

[0065] Step 2): Calculate the marginal contribution variance of a single agent. ; For intelligent agents Iterate through all possible alliances S and calculate the variance of their marginal contributions: ; K: The number of alliances S; The k-th alliance; Intelligent agent The mean marginal contribution.

[0066] Step 3: Calculate the overall marginal contribution variance of the system. ; Calculate the global variance based on the marginal contributions of all agents: ; N: Total number of agents; : The total average marginal contribution of all agents.

[0067] Step 4: Calculate the coefficient for the number of additional samples ; Based on the ratio of the variance of a single agent to the total variance of the system, combined with a scaling factor... (Usually 1~5, to match the total amount of sampling resources), resulting in: ; like This indicates that the agent's contribution fluctuates more than the global level. Greater than To obtain more additional samples; like This indicates that the agent's contribution fluctuation is less than the global level. Less than Reduce additional sampling.

[0068] For example, consider 5 agents: Assumption: Energy agent 2 The contribution fluctuates greatly; Security Intelligent Agent 1 The contribution fluctuation is small; Total system variance ; Scaling factor .

[0069] Then: Energy Agent 2 Four additional samples were obtained; Security Intelligent Agent 1 Only one additional sample was obtained; Through this process This achieves the goal of "allowing agents with large contribution fluctuations to obtain more sampling resources," which not only improves the accuracy of Shapley's credit assessment but also avoids resource waste.

[0070] As a seasonal adjustment item, the number of samplings is increased during seasons with significant changes in building load.

[0071] Preferably, to improve the long-term effectiveness of this method in dynamic environments, the key parameters in the aforementioned steps are adaptively adjusted according to changes in the external environment and system operating mode, forming a closed-loop optimization mechanism. The calculated seasonal parameters are then... Used to adjust seasonal sampling increments ,Right now: , that is, use Dynamically adjust seasonal adjustment items .

[0072] Seasonal parameters Adjustment: Adjust the sampling strategy according to the seasonal load changes of the building: ; in, Baseline parameter values; It indicates the day of the year and simulates seasonal cycles using a sine function; By utilizing the periodicity of the sine function (corresponding to the four seasons), let The parameters fluctuate dynamically with the seasons. Seasons with high load (such as summer and winter) correspond to larger parameter values, while seasons with low load (such as spring and autumn) correspond to smaller parameter values.

[0073] During periods of high seasonal load, Increase This increases the number of samples, adapting to the operational needs of high-load scenarios; when the seasonal load is low, sampling is reduced to save resources.

[0074] Step 2.4: Finally, each agent... Shapley credit rating Through the The expected value of the marginal contribution of each sampling is obtained by calculating the Shapley credit score. The formula integrates the intelligent agent from step 2.1. In the Building Operation and Maintenance Subsystem Alliance Marginal contribution and the operation and maintenance intelligent agent in step 2.3 Adaptive sampling number : ; Here, the Building Operation and Maintenance Subsystem Alliance There is An intelligent agent. This will be used as the input for step 3.

[0075] Step 3: Building Operation and Maintenance Agent Policy Update: Based on the Shapley credit scores of each operation and maintenance agent calculated in Step 2. Joint action value function Feedback is sent to each intelligent agent. Individual learning, for the intelligent agent The current strategy is updated accordingly to be better, and the global value of each agent, i.e. the optimized Q value, is calculated. Based on the global value of each agent, i.e. the optimized Q value, efficient and compliant actions are selected for each agent.

[0076] Preferably, the Shapley credit score is fitted based on a value network. intelligent agent The individual value network loss function is: ; in, For intelligent agents The local action value function, based on its own observation and actions Evaluate the value of actions; after value network optimization, Accurately reflects the intelligent agent The overall value; The Shapley's estimate of agent i is used to quantify the agent. Contributions in collaboration; These are the model parameters of the value network for agent i.

[0077] Here, the local action value function of an agent learns its global Shapley credit, enabling individual policy optimization to consider both its own local observations and its contribution requirements in global collaboration, ultimately achieving efficient multi-agent collaboration in building operation and maintenance. By minimizing the squared difference between the "Shapley credit estimate" and the "local action value," the value network learns action values ​​that reflect its own contribution.

[0078] Preferably, the loss function of the policy network, based on the conventional policy gradient in reinforcement learning, incorporates building operation and maintenance constraints into the policy network update:

[0079] Among them, the policy gradient term : This is the policy gradient loss; in conventional reinforcement learning, the policy is updated by maximizing the value of actions. By maximizing "logarithm of policy probability × action value", the agent tends to choose high-value actions and optimize its own policy. Agent i is observing Select action The probability of the strategy; It is an intelligent agent. The local action value function; after value network optimization. Accurately reflects the intelligent agent The overall value; Operation and maintenance constraints middle, : is the constraint weight factor (the priority of balancing strategy optimization and constraint satisfaction).

[0080] : This refers to building operation and maintenance constraint losses, including safety constraints (such as security thresholds), energy consumption limits (such as energy consumption limits), comfort requirements (such as temperature and humidity ranges), etc.

[0081] By integrating "value network learning contribution + policy updates with operational constraints", each agent can optimize its own strategy while considering both the fairness of contribution in collaboration and the actual constraints of building operation and maintenance, such as safety, energy consumption, and comfort. Ultimately, this achieves efficient and compliant collaboration among multiple agents in building operation and maintenance.

[0082] Specifically, taking the "building summer air conditioning operation and maintenance scenario" as a concrete example, and combining the collaborative logic of 5 intelligent agents (security 1, energy 2, equipment 3, environment 4, and user 5), step 3 includes: 1) Prerequisite scenario setup: Summer noon (high building load): User area temperature 30℃ (exceeding the comfort threshold of 26℃), air conditioning equipment operating load 80% (close to overload). Currently, it is necessary to cool down through "environmental regulation + energy control + equipment maintenance" in a coordinated manner, while meeting the operation and maintenance constraints of energy consumption ≤ 110% of rated value, equipment load ≤ 90%, and comfort level ≥ 0.8.

[0083] Step 1: Value Network Loss Function (Basic), aligning local Q-values ​​with global contributions. Value network loss function The core is to teach intelligent agents to recognize their overall value and avoid local misjudgments.

[0084] 1. Key Input: Shapley Credit Value (Global Contribution Quantification): The building operations manager, through... After evaluating the value of the joint action, calculate the global contribution of each agent: Energy Control Agent 2 ( 2=0.35): The core contribution is "dynamically lowering the upper limit of air conditioner power to 85%", which avoids equipment overload without sacrificing cooling effect, and has the highest overall contribution; Environmental regulation agent 4 ( 4=0.3): The contribution is "setting the air conditioner outlet temperature to 18℃ and directing the airflow", which quickly cools the air without wasting energy; Equipment maintenance intelligent agent 3 ( 3=0.2): The contribution is "real-time monitoring of the air conditioner compressor status, with no abnormalities reported", providing a safety basis for the actions of the former two; User service intelligent agent 5 ( 5=0.1): The contribution is "collecting user comfort feedback (0.85), confirming the effectiveness of cooling"; Security Intelligent Agent 1 ( 1=0.05): No direct participation, lowest contribution.

[0085] Local Q value (Agent self-evaluation): Initially, the agent judges value solely based on its own observations, which may lead to "overestimation / underestimation": The initial state of the environmental agent 4 (Q4=0.45): It sees itself "cooling down rapidly" and mistakenly believes that its contribution is extremely high, thus overestimating its value; The initial state of Energy Agent 2 (Q2=0.25): It only knew to "reduce power" and did not realize the global significance of avoiding equipment overload, thus underestimating its value; The initial state of device agent 3 (Q3=0.1): It thinks it is "just monitoring the status", fails to recognize the importance of feedback, and underestimates its value.

[0086] 2. The optimization process of the value network loss function: For each agent, by minimizing Update value network parameters: Environmental intelligent agent 4: ( Q4=0.3) < initial (Q4=0.45), the loss function drives Q4 down from 0.45 to 0.3 (to align with the true global contribution); Energy Agent 2: ( Q2=0.35) > Initial (Q2=0.25), the loss function drives Q2 to increase from 0.25 to 0.35; Device Agent 3: ( Q3=0.2) > Initial (Q3=0.1), the loss function drives Q3 to be increased from 0.1 to 0.2.

[0087] 3. Optimization Results: The local Q values ​​of each agent are precisely matched with the Shapley credit values ​​(Q2=0.35, Q4=0.3, Q3=0.2...). At this point, the Q value is no longer a "self-perceived value" but a "globally recognized real value"—this is the role of the value network as a "foundation": to provide an "accurate value benchmark" for subsequent policy optimization.

[0088] The second step: The policy network loss function is based on the accurate Q value, which is the global value of each agent, to select efficient and compliant actions for each agent.

[0089] Strategy Network Loss Function It uses accurate value benchmarks to guide intelligent agents to choose the right actions without violating the rules.

[0090] 1. Key Input: Optimized Q values ​​(from the value network loss function): Q2=0.35 (energy), Q4=0.3 (environment), Q3=0.2 (equipment); Operation and maintenance constraints Set constraint thresholds: energy consumption ≤ 110% of rated value, equipment load ≤ 90%, comfort level ≥ 0.8, and constraint weights. =0.5 (Constraints and efficiency are equally important); Policy Probability The probability of an agent choosing different actions under the current observation. For example, energy agent 2 can choose "reduce power to 80%", "85%" or "90%".

[0091] Specifically, the optimized Q-value is obtained through the loss function of the value network. The optimization achieves this by aligning the value of local actions with the global Shapley credit value. The specific steps are as follows: Step 1): Obtain the global Shapley credit score ; Shapley credit score quantifies an agent's contribution to global collaboration and requires prior assessment based on the joint action value function. calculate: Supervisor collects global status (such as building temperature, energy consumption, equipment status) and the coordinated actions of all intelligent agents. ; Through the trained joint action value function Calculate the global value of the joint actions; Using the Shapley value algorithm, a global value is allocated to each agent to obtain a Shapley credit value. For example, energy agent 2 4. Environmental intelligent agent .

[0092] Step 2): Initialize the agent's local Q-values. ; Each agent is based on its own local observations (e.g., energy intelligence agents observe current energy consumption and equipment load) and actions (If the power is adjusted to 85%, initialize the local action value function) (Initial values ​​are usually random or set based on simple rules.)

[0093] For example: The initial state of Energy Agent 2 They underestimated their own overall contribution; Initialization of Environmental Agent 4 They overestimated their own overall contribution.

[0094] Step 3): Optimize the Q-value using the value network loss function; Using the loss function of value networks The parameters of the value network are updated using the gradient descent algorithm. Minimize the mean square error between the Shapley credit score and the local Q value: Calculate the loss function with respect to parameters The gradient, i.e., the direction of error change; Adjust along the opposite direction of the gradient For example, through optimizers such as stochastic gradient descent (SGD) and Adam; Repeat the above process until the loss function is small enough, i.e. the error between the local Q value and the Shapley credit value meets the requirements.

[0095] Ultimately, the local Q value will precisely match Shapley's credit score: The Q2 of Energy Agent 2 was adjusted from 0.25 to 0.35, compared with Alignment; The Q4 of the environmental agent 4 was adjusted from 0.45 to 0.3, and... Alignment; The Q3 of device agent 3 was adjusted from 0.1 to 0.2, and... Alignment.

[0096] Q-value optimization is an alignment process of global contribution (Shapley value) → local value (Q-value). Through gradient descent of the value network loss function, the local value evaluation of the agent can accurately reflect its true contribution in global cooperation.

[0097] 2. Optimization process of the loss function of a policy network with a separate agent: (1) Core action selection of energy control agent 2: Initial strategy: It may randomly select "reduce power to 80%", "85%" or "90%" with a probability of 1 / 3 for each; Policy gradient term Function: The driving agent tends to choose actions that "can increase the Q value". Since the optimized Q2=0.35 corresponds to "downgrade to 85%" (it has been verified that this action has the highest global contribution), the loss function will increase the probability of "downgrade to 85%" (from 1 / 3 to 60%) and reduce the probability of other actions. Constraints Function: If Energy Agent 2 wants to select "reduce to 90%" (which seems to reduce temperature more), this action will cause the equipment load to reach 92% (exceeding the ≤90% constraint). This will increase the overall loss function, thereby suppressing the selection of this action (the probability drops from 1 / 3 to 10%). Optimized strategy: Energy agent 2 has a 60% probability of selecting "downgrade to 85%" (high Q value + compliance), a 30% probability of selecting "downgrade to 80%" (compliant but slightly lower Q value), and a 10% probability of selecting "downgrade to 90%" (violation, probability suppressed).

[0098] (2) Environmental regulation agent 4 (action compliance constraints): Initial strategy: Consider setting the air outlet temperature to 16℃ (for faster cooling) or 18℃; The effect of the policy gradient term: Q4=0.3 corresponds to "set to 18℃" (highest global contribution), which drives the probability of this action to increase; Effect of the constraint: If "16℃" is selected, the air conditioner energy consumption will reach 115% (exceeding the constraint of ≤110%). As the probability increases, the loss function rises, thus suppressing the action (the probability drops from 50% to 15%). Optimized strategy: Environmental agent 4 has a 75% probability of selecting "18℃ directional air supply" (high Q value + compliance), a 15% probability of selecting "16℃" (violation, probability suppressed), and a 10% probability of selecting "20℃" (compliant but slow cooling, slightly lower Q value).

[0099] (3) Equipment maintenance intelligent agent 3 (assisted action optimization): Initial strategy: possibly "monitor every 5 minutes" or "monitor every 1 minute"; The effect of the policy gradient term: Q3=0.2 corresponds to "monitoring once every 1 minute" (real-time feedback of status to support the actions of energy / environment intelligent agents), which drives the probability of this action to increase; The purpose of the constraint: "Monitoring every 1 minute" will not violate any constraints (energy consumption is negligible and will not affect the equipment). =0, no inhibition; Optimized strategy: Device agent 3 has a 90% probability of selecting "monitor every 1 minute" (high Q value + compliance).

[0100] Here, based on the accurate Q value, the intelligent agent prioritizes high-value actions, while eliminating illegal actions through constraints. This ultimately forms a collaborative strategy of "energy reduction to 85% + directional air supply to 18°C ​​environment + equipment monitoring for 1 minute". This achieves the efficient goal of "temperature reduction to 26°C (comfort level 0.85)" while meeting the compliance requirements of "108% energy consumption and 85% equipment load", perfectly matching the needs of building operation and maintenance.

[0101] According to another aspect of the present invention, a computer-readable storage medium is also provided, having stored thereon computer-executable instructions, wherein when executed by a processor, the computer-executable instructions cause the processor to perform the method described in any of the above embodiments.

[0102] In summary, the present invention has the following advantages: (1) Improved fairness of building operation and maintenance: Credit allocation based on Shapley value ensures that the contribution of each operation and maintenance subsystem is fairly evaluated, and avoids underestimation of some key but insignificant systems (such as preventive maintenance).

[0103] (2) Cross-system collaboration optimization: By evaluating the value of an alliance of operation and maintenance subsystems of any size, effective cross-system collaboration models can be discovered and strengthened, such as the joint energy-saving strategy of HVAC and smart lighting.

[0104] (3) Improved building operation and maintenance efficiency: The special design for building operation and maintenance scenarios enables each intelligent agent to learn strategies that are more in line with the actual building operation needs, significantly improving the overall operation and maintenance efficiency.

[0105] (4) Enhanced dynamic adaptability: Considering the seasonal and cyclical changes in building usage patterns, the credit allocation mechanism can adapt to changes in collaborative relationships under different operating scenarios.

[0106] For detailed descriptions of the various device embodiments of the present invention, please refer to the corresponding sections of the various method embodiments; they will not be repeated here.

[0107] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

[0108] It should be noted that the present invention can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of the present invention can be executed by a processor to implement the steps or functions described above. Similarly, the software program of the present invention (including associated data structures) can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of the present invention can be implemented in hardware, for example, as circuitry that works with a processor to perform the various steps or functions.

[0109] Furthermore, a portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. The program instructions invoking the methods of the invention may be stored in a fixed or removable recording medium, and / or transmitted via a data stream in a broadcast or other signal-carrying medium, and / or stored in the working memory of a computer device operating according to the program instructions. Here, an embodiment of the invention includes an apparatus comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein, when the computer program instructions are executed by the processor, the apparatus is triggered to operate the methods and / or technical solutions based on the foregoing embodiments of the invention.

[0110] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the apparatus claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.

Claims

1. A method for updating the strategy of a building operation and maintenance intelligent agent, characterized in that, include: Based on trainable parameters, train the function for the current task, and based on the completed training function, obtain the joint action value function of the intelligent agent of the building operation and maintenance system for the current task. Based on the joint action value function of the agents, calculate the Shapley credit value of each agent; Based on each agent's Shapley credit score, efficient and compliant actions are selected for each agent.

2. The building operation and maintenance intelligent agent strategy update method as described in claim 1, characterized in that, Based on trainable parameters, a function for the current task is trained. Based on the completed training function, the joint action value function of the building operation and maintenance system agents for the current task is obtained, including: The function of the current task for: ; in, The comprehensive reward for building operation and maintenance includes a multi-dimensional evaluation of energy efficiency, user comfort, safety indicators, and equipment health. It is a discount factor used to balance the weight of current rewards and future rewards; and These are the current global state and the combined action, respectively. This represents the value function of the joint actions at the current moment; the building operation and maintenance system contains N intelligent agents, each of which... Based on its own local observations Generate action During the training phase, the global state is received. Joint actions with all specialized intelligent agents It outputs a joint action value function at the current moment. ; and These are the global state and the combined action at the next moment, respectively. The value function representing the joint action at the next moment; Represents the value function of joint actions The model parameters can be used to train the parameters. It is by The determined value assessment function is optimized. ,let The valuation results are more in line with the actual benefits of building operation and maintenance; It represents the mathematical expectation.

3. The building operation and maintenance intelligent agent strategy update method as described in claim 2, characterized in that, Based on the joint action value function of the agents, the Shapley credit score of each agent is calculated, including: Based on the preset baseline strategy and joint action value function, the marginal contribution of the agent in the building operation and maintenance subsystem alliance is obtained; Based on the comprehensive reward for building operation and maintenance, the operation and maintenance collaboration affinity matrix among intelligent agents is obtained; Based on the marginal contribution of each agent in the building operation and maintenance subsystem alliance, the operation and maintenance collaboration affinity matrix between agents, and the preset adjustment terms, the adaptive sampling number of the agent is obtained. The Shapley credit value for each agent is obtained based on the marginal contribution of each agent in the building operations and maintenance subsystem alliance and the adaptive sampling number of each agent.

4. The building operation and maintenance intelligent agent strategy update method as described in claim 3, characterized in that, Based on the preset baseline strategy and joint action value function, the marginal contribution of the agent in the building operation and maintenance subsystem alliance is obtained, including: Based on preset baseline strategy and joint action value function To obtain the intelligent agent In the Building Operation and Maintenance Subsystem Alliance Marginal contribution The formula is as follows: ; Among them, baseline strategy This is the default policy; Each agent Each has a corresponding active current operation and maintenance policy; Indicates alliance The specialized operation and maintenance agent in the system executes the current operation and maintenance strategy; Indicates alliance All external agents adopt the building operations and maintenance baseline strategy. ; Indicates alliance Internal intelligent agents Other members execute the current operation and maintenance policy; Indicates that the intelligent agent The action replacement is its building operation and maintenance baseline strategy. ; Represents intelligent agents In the Building Operation and Maintenance Subsystem Alliance S, the marginal contribution of building operation and maintenance to the overall value of the system. The marginal contribution of building operation and maintenance is reflected by the value difference between the two scenarios, which is used by the intelligent agent. Function: First item Indicates the presence of intelligent agents The overall value of the system when the entire building operation and maintenance subsystem S executes the current strategy; Second item This indicates that only intelligent agents exist in the Building Operation and Maintenance Subsystem Alliance S. The overall value of the system when switching to the baseline strategy; and The difference is the agent's value. The additional value increment brought to the system by executing the current strategy.

5. The building operation and maintenance intelligent agent strategy update method as described in claim 3, characterized in that, Based on the comprehensive building operation and maintenance reward, an operation and maintenance collaboration affinity matrix is ​​obtained among professional operation and maintenance intelligent agents, including: Based on comprehensive building operation and maintenance rewards Obtain the operation and maintenance collaboration affinity matrix among professional operation and maintenance intelligent agents. The formula is as follows: ; in, The comprehensive reward for building operation and maintenance includes a multi-dimensional evaluation of energy efficiency, user comfort, safety indicators, and equipment health. For a moment At that time, intelligent agent With intelligent agents The degree of cooperation and affinity between them; For a moment At that time, intelligent agent With intelligent agents The degree of cooperation and affinity between them; It is to utilize the previous moment Affinity, combined arrive The actual interaction data of the stage is used to calculate the result through the update formula; This serves as a weighting factor, used to balance the impact of historical affinity and current collaborative contribution; This is a building operations and maintenance collaboration indicator function used to determine the first... In this interaction, the intelligent agent and Whether collaboration has occurred is indicated by a value of 1 if yes and 0 otherwise. Among them, building operation and maintenance collaboration indication function Based on the following building operation and maintenance scenario definition: When the energy-saving strategy of the energy control agent and the comfort control of the environmental condition agent cooperate, i.e., when collaboration occurs... =1; When the personnel detection of the security management intelligent agent and the area control linkage of the facility management intelligent agent occur, i.e., when collaboration takes place... =1; When the coordination between the equipment maintenance agent and the energy control and management agent in preventive maintenance, i.e., when collaboration occurs, =1.

6. The building operation and maintenance intelligent agent strategy update method as described in claim 5, characterized in that, After obtaining the affinity matrix for operational collaboration among professional operational intelligence agents, it also includes: According to building usage patterns Dynamically adjust the operation and maintenance collaboration affinity matrix Update frequency: 。 7. The building operation and maintenance intelligent agent strategy update method as described in claim 3, characterized in that, Based on the marginal contribution of each agent in the building operations and maintenance subsystem alliance, the operational collaboration affinity matrix between agents, and preset adjustment terms, the adaptive sampling number of the agents is obtained, including: Based on the operational collaboration affinity matrix among intelligent agents and preset adjustment items To obtain the operation and maintenance intelligent agent Adaptive sampling number The formula is as follows: ; in, For intelligent agents The total number of samples, which is composed of , and This is obtained by adding the three sampling times together; The baseline number of samples is given for all agents, where the sampling probability P based on layer 1 is... layer1 =0.5, sampling probability P of layer 2 layer2 =0.3, sampling probability P of layer 3 layer3 =0.2 and the baseline sampling number are used to assign corresponding sampling numbers to agents at each layer; layer 1 includes: security management, energy control and environmental control agents; layer 2 includes: equipment maintenance and facility management agents; layer 3 includes: cleaning, parking and communication agents; For intelligent agents The variance of operational performance; The variance of the overall system operation and maintenance performance; and Based on the operation and maintenance collaboration affinity matrix Perform calculations; The additional sampling coefficient represents the marginal contribution of building operations and maintenance based on each agent. The marginal contribution variance of a single agent and the marginal contribution variance of the entire system are obtained; based on the marginal contribution variance of a single agent and the marginal contribution variance of the entire system, the coefficient of the additional sampling number is obtained. ; As a seasonal adjustment item, the number of samplings is increased during seasons with significant changes in building load.

8. The building operation and maintenance intelligent agent strategy update method as described in claim 7, characterized in that, Based on the marginal contribution of each agent in the building operations and maintenance subsystem alliance, the operational collaboration affinity matrix between agents, and preset adjustment terms, the adaptive sampling number of the agents is obtained, including: The calculated seasonal parameters Used to adjust seasonal adjustment items ,Right now: , ; in, Baseline parameter values; It indicates the day of the year.

9. The building operation and maintenance intelligent agent strategy update method as described in claim 3, characterized in that, Based on each agent's marginal contribution to the building operations and maintenance subsystem consortium and the adaptive sampling number for each agent, the Shapley credit value for each agent is obtained, including: The Shapley credit value for each agent is obtained using the following formula. : ; in, For intelligent agents In the Building Operation and Maintenance Subsystem Alliance Marginal contribution in; Building Operation and Maintenance Subsystem Alliance There is One intelligent agent; For intelligent agents The adaptive number of samplings.

10. The building operation and maintenance intelligent agent strategy update method as described in claim 1, characterized in that, Based on each agent's Shapley credit score, efficient and compliant actions are selected for each agent, including: The global value of each agent is calculated based on the Shapley credit value and joint action value function of each operation and maintenance agent. Based on the global value of each agent, efficient and compliant actions are selected for each agent.

11. The building operation and maintenance intelligent agent strategy update method as described in claim 10, characterized in that, Based on the Shapley credit score and joint action value function of each operational agent, the global value of each agent is calculated, including: Shapley credit score fitted based on value network intelligent agent The individual value network loss function is: ; in, For intelligent agents The local action value function, based on its own observation and actions Evaluate the value of actions; after value network optimization, Accurately reflects the intelligent agent The overall value; For intelligent agents The Shapley estimate is used to quantify the intelligent agent. Contributions in collaboration; These are the model parameters of the value network for agent i.

12. The building operation and maintenance intelligent agent strategy update method as described in claim 11, characterized in that, Based on the global value of each agent, efficient and compliant actions are selected for each agent, including: The loss function of the network based on the following policy Select efficient and compliant actions for each intelligent agent: ; Policy gradient term middle, Intelligent agent In observation Select action The probability of the strategy; It is an intelligent agent. The local action value function; Operation and maintenance constraints middle, : This refers to the loss due to building operation and maintenance constraints.

13. A computer-readable storage medium having stored thereon computer-executable instructions, wherein, When the computer-executable instructions are executed by the processor, the processor causes the processor to perform the method as described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Cooperative multi-agent cooperation method based on explicit credit distribution

    CN118864085A

  • Residential integrated energy system optimization control method based on multi-agent reinforcement learning

    CN119024707A

  • Low-carbon building cluster energy management optimization method and system based on reinforcement learning

    CN119359496A

  • Intelligent business building cluster collaborative operation method based on improved reinforcement learning

    CN119721575A

  • Power-cut plan arrangement method and system based on multi-agent interpretable reinforcement learning

    CN120197915A