Intelligent agent cooperation method and device, equipment and storage medium
By determining the priority factors and semantic similarity of intelligent agents in the industrial park energy management and control system and combining them with collaborative contribution for global optimization, the problem of low efficiency of traditional energy management and control systems is solved, and efficient collaboration of intelligent agents and energy optimization management are achieved.
Patent Information
- Application Number
- CN202511165532.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional industrial park energy management and control systems rely on fixed rules of intelligent entities and are difficult to adapt to complex and changing energy demands, resulting in low utilization efficiency and increased operating costs.
By determining the priority factors of the energy management and control intelligent body, intercepting historical operation instructions and converting them into text data, using pre-trained word vector models for semantic mapping, calculating semantic similarity and assigning local reward indicators, and combining collaborative contribution for global optimization, an optimized management and control strategy is constructed.
It achieves efficient collaboration between intelligent entities, adapts to complex and changing energy demands, improves overall utilization efficiency and system stability, and reduces operating costs.
Smart Images

Figure CN120655069A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of energy management technology, and in particular to an intelligent body collaboration method, device, equipment and storage medium. Background Art
[0002] With the ongoing restructuring of industrial structures and the rapid advancement of smart buildings, industrial parks, as significant energy consumption hubs, are expanding, and electricity demand is rapidly increasing. The energy management system for industrial parks is a complex and extensive system, encompassing multiple energy components such as photovoltaic and biomass power generation systems, which together form a microgrid system.
[0003] The scheduling strategies of traditional industrial park energy management systems generally rely on fixed rules of intelligent agents. However, the energy demands of industrial parks are complex and changeable. This static scheduling model is difficult to ensure the maximization of the overall utilization efficiency of intelligent agents and may even lead to energy waste and increased operating costs. Summary of the Invention
[0004] The main purpose of this application is to provide an intelligent body collaboration method, device, equipment and storage medium, aiming to solve the technical problem that the scheduling strategy of the traditional industrial park energy management system generally relies on the fixed rules of the intelligent body. Since the energy demand of the industrial park is complex and changeable, this static scheduling mode is difficult to ensure the maximization of the overall utilization efficiency of the intelligent body.
[0005] To achieve the above objectives, this application proposes an agent collaboration method, which includes: Determine the priority factors of each energy management agent in the energy management system of the industrial park; Intercepting historical operation instructions output by the energy management and control intelligent agent and converting the historical operation instructions into text data information; Perform semantic mapping on the text data information based on the pre-trained word vector model to obtain the semantic similarity between each energy management and control agent; assigning a local reward indicator to each energy management agent based on the semantic similarity and the priority factor; Based on the local reward index, each energy management and control intelligent agent is globally optimized according to the collaboration contribution to obtain an optimized management and control strategy, so that the energy management and control intelligent agents can collaborate based on the optimized management and control strategy.
[0006] In one embodiment, the step of intercepting the historical operation instructions output by the energy management and control agent and converting the historical operation instructions into text data information includes: Retrieving the operation log of the energy management and control agent from the energy management and control system; Matching historical operation instructions within a preset time period in the operation log using a regular expression; The historical operation instructions are parsed using natural language processing technology to obtain text data information of the energy management and control intelligent body.
[0007] In one embodiment, the step of performing semantic mapping on the text data information according to the pre-trained word vector model to obtain the semantic similarity between each energy management and control agent includes: Performing standardization processing on the text data information to obtain a standardized text; Convert each word in the standardized text into a corresponding word vector using a pre-trained word vector model to obtain a word vector group corresponding to the standardized text; Performing clustering processing on the word vector group to obtain multiple word vector clusters; The word vector cluster is semantically mapped using the cosine similarity algorithm to obtain the semantic similarity between each energy management and control agent.
[0008] In one embodiment, the step of clustering the word vector group to obtain a plurality of word vector clusters includes: Determining the number of types of the energy management and control agents; Randomly select initial cluster centers corresponding to the number of types from the word vector group; The word vector group is clustered using the K-Means clustering algorithm to obtain multiple word vector clusters.
[0009] In one embodiment, the step of allocating a local reward indicator to each energy management agent based on the semantic similarity and the priority factor includes: Defining fuzzy sets corresponding to the semantic similarity and the priority factors in each energy management and control agent; Establishing a relationship between the semantic similarity, the priority factor and the local reward indicator to obtain a fuzzy rule; According to the fuzzy rules, fuzzy reasoning is performed on the fuzzy set by using the Mamdani reasoning method to obtain the membership degree of the fuzzy set; The membership is defuzzified by the maximum membership method to obtain the local reward index corresponding to each energy management agent.
[0010] In one embodiment, the step of performing global optimization on each energy management and control agent according to the collaborative contribution based on the local reward indicator to obtain an optimized management and control strategy includes: According to the collaborative relationship between energy management and control agents, a relationship model between energy management and control agents is constructed; Based on the relationship model, a cost-benefit analysis method is used to quantify the collaborative contribution of each energy management agent to other energy management agents; Based on the collaborative contribution, the energy management and control agents are globally optimized through a multi-agent proximal strategy optimization algorithm to obtain an optimized management and control strategy.
[0011] In one embodiment, the step of performing global optimization on each energy management agent based on the collaborative contribution by using a multi-agent proximal strategy optimization algorithm to obtain an optimized management strategy includes: Modeling the energy management and control environment of the industrial park to determine a state space, wherein the state space includes key energy information and the collaborative contribution of each energy management and control agent; Build a global strategy network and a global value network for each energy management agent; According to the state space, the global policy network and the global value network are trained by a multi-agent proximal policy optimization algorithm, and the training is continued until the network converges; When the global policy network and the global value network converge, an optimized management and control strategy is extracted from the trained global policy network.
[0012] In addition, to achieve the above objectives, the present application also proposes an intelligent agent collaboration device, the device comprising: A priority determination module is used to determine the priority factors of each energy management and control agent in the energy management and control system of the industrial park; An operation interception module, configured to intercept historical operation instructions output by the energy management and control agent and convert the historical operation instructions into text data information; A semantic mapping module is used to perform semantic mapping on the text data information based on a pre-trained word vector model to obtain the semantic similarity between each energy management and control agent; A local reward module, configured to assign a local reward indicator to each energy management agent based on the semantic similarity and the priority factor; A global optimization module is used to perform global optimization on each energy management and control agent according to the collaboration contribution based on the local reward index to obtain an optimized management and control strategy so that each energy management and control agent can collaborate based on the optimized management and control strategy.
[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes an intelligent agent collaboration device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the intelligent agent collaboration method as described above.
[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the intelligent agent collaboration method described above are implemented.
[0015] One or more technical solutions proposed in this application have at least the following technical effects: The intelligent agent collaboration method of this application includes: determining the priority factors of each energy management and control intelligent agent in the energy management and control system of the industrial park; intercepting the historical operation instructions output by the energy management and control intelligent agent, and converting the historical operation instructions into text data information; performing semantic mapping on the text data information according to the pre-trained word vector model to obtain the semantic similarity between each energy management and control intelligent agent; allocating a local reward index to each energy management and control intelligent agent according to the semantic similarity and the priority factor; based on the local reward index, globally optimizing each energy management and control intelligent agent according to the collaboration contribution to obtain an optimized management and control strategy, so that each energy management and control intelligent agent can collaborate based on the optimized management and control strategy.
[0016] Since this application determines the semantic similarity between each energy management and control intelligent agent based on the historical operation instructions output by the energy management and control intelligent agent, allocates local reward indicators to each energy management and control intelligent agent through semantic similarity and priority factors, and finally performs global optimization in combination with the collaborative contribution to obtain the optimized management and control strategy, the collaboration between energy management and control intelligent agents can be optimized in real time, so that the intelligent agents can learn the optimized management and control strategy faster and adapt to the complex and changing energy demand situation, so that these intelligent agents can work together efficiently to achieve optimized management and control of regional energy, ensuring the maximization of the overall utilization efficiency of the intelligent agents. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 A flowchart of the first embodiment of the method for intelligent collaboration provided in this application; Figure 2 A flowchart of the second embodiment of the intelligent agent collaboration method of this application is provided; Figure 3This is a schematic diagram of the module structure of the intelligent collaborative device according to an embodiment of the present application; Figure 4 This is a schematic diagram of the device structure of the hardware operating environment involved in the intelligent agent collaboration method in the embodiment of this application.
[0020] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0021] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0022] In order to better understand the technical solution of this application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0023] It should be noted that the execution entity of this embodiment can be a computing service device with semantic mapping, local reward allocation, and global optimization functions, such as a personal computer or server, or an electronic device capable of implementing the above functions, or an agent collaboration device that executes the agent collaboration method of this application (hereinafter referred to as a collaboration device), etc. This embodiment is not limited to this. The following uses the collaboration device as an example to illustrate this embodiment and the following embodiments.
[0024] Based on this, the first embodiment of the present application is proposed. The embodiment of the present application provides an intelligent agent collaboration method, referring to Figure 1 , Figure 1 A flowchart of the first embodiment of the intelligent agent collaboration method of this application is provided.
[0025] In this embodiment, the agent collaboration method includes steps S10 to S50: Step S10: Determine the priority factors of each energy management and control intelligent entity in the energy management and control system of the industrial park.
[0026] It should be noted that the energy management and control system of an industrial park is an integrated management system that comprehensively monitors, analyzes and controls the production, transmission, distribution and use of various energy sources (such as electricity, natural gas, biomass energy, etc.) within the industrial park.
[0027] The energy management agent is an intelligent entity within the energy management system specifically designed for energy management and control. It can autonomously perceive information related to energy generation, transmission, storage, and usage, and proactively and interactively respond. For example, it can dynamically adjust energy allocation strategies based on energy prices and user demand over time to achieve efficient energy utilization and implement appropriate management and allocation.
[0028] For example, according to the goals of the energy management and control intelligent body, it can be divided into energy production intelligent body (which can stably produce sufficient energy based on the availability, cost and efficiency of various energy resources, such as solar energy, wind energy, biomass energy, etc.), energy distribution intelligent body (which can reasonably distribute the produced energy to various consumption areas or users), energy consumption side management intelligent body (which can monitor the demand patterns of users in industrial parks and adjust users' electricity equipment usage strategies according to electricity prices in different time periods), etc. This embodiment does not impose any restrictions on this.
[0029] It should be understood that the priority factor is an element used to determine the order of importance of the above-mentioned different intelligent entities in the energy management and control operation in the energy management and control system of the industrial park.
[0030] For example, priority factors can be categorized as stability, emergency demand, and cost-effectiveness. For example, in terms of stability, the stability of energy production should be prioritized, as fluctuations or interruptions in energy production can severely impact the entire energy system. Regarding emergency demand, when certain regions or users have urgent energy needs, the energy distribution agent should prioritize meeting those needs.
[0031] In the specific implementation, we first evaluate the importance of each energy management and control intelligent body in each link of energy generation, transmission, storage and use, sort them by importance, and determine the priority factors of each energy management and control intelligent body.
[0032] Step S20: intercepting the historical operation instructions output by the energy management and control intelligent agent, and converting the historical operation instructions into text data information.
[0033] It's important to note that historical operation commands are a record of the operational commands executed by the energy management agent over a period of time. They reflect the actions taken by the energy management agent under different energy demand, supply conditions, and system states. For example, within the energy production agent, the power agent may issue commands to adjust power generation or dispatch power based on peak and trough power consumption within the park, while the natural gas agent may issue commands to adjust delivery volume based on user demand.
[0034] It can be understood that text data is a textual representation of the historical operational instructions output by the energy management agent. By converting historical operational instructions into text, patterns can be discovered and the agent's performance can be evaluated.
[0035] In the specific implementation, the collaborative device can locate the historical operation instructions from the log of the energy management and control intelligent body, and then use the data conversion tool to convert it into text format according to the instruction format to obtain text data information.
[0036] In a feasible implementation, step S20 of this embodiment may include the steps of: retrieving the operation log of the energy management and control intelligent body from the energy management and control system; matching historical operation instructions within a preset time period in the operation log through regular expressions; and parsing the historical operation instructions using natural language processing technology to obtain text data information of the energy management and control intelligent body.
[0037] It should be noted that the operation log is a record of various operations performed by the energy management and control intelligent body during its operation, including information such as the time of the operation, the type of operation, the equipment or parameters involved, etc.
[0038] It's understood that regular expressions are tools for matching, searching, and manipulating text patterns. In energy management, by defining regular expression patterns, you can accurately locate matching operation instructions in operation logs. For example, you can filter historical operation instructions by keyword or time period. The preset time period is a pre-set time range, such as a week or a month, and this embodiment does not impose any restrictions on this.
[0039] It should be understood that natural language processing (NLP) technology enables computers to understand, process, and generate human language. In energy management, using NLP to parse historical operating instructions can convert the unstructured information contained in these instructions into structured information that computers can understand, thereby obtaining the text data corresponding to these historical operating instructions.
[0040] For example, named entity recognition technology can be used to accurately identify entities such as device names and location names in historical operation instructions, and then generate coherent sentences. For example, if the historical operation instruction is "<Operation type: Energy allocation, Source location: Energy Zone A, Target location: Campus Building B, Allocation amount: 1000kW, Timestamp: 2025-06-18 14:00:00>", named entity recognition can be used to determine that "Energy Zone A" and "Campus Building B" are entities, and then generate the text message "At 14:00 on June 18, 2025, 1000 kilowatts of energy were allocated from Energy Zone A to Campus Building B."
[0041] In this embodiment, since the energy management and control system records the operation logs of the energy management and control agent and stores them in a local server or cloud storage, the collaborative device can determine the preset time period of the historical operation instructions to be intercepted according to demand. For example, if the energy demand has changed dramatically in the past week (such as the introduction of new wind energy), the rows with timestamps within the past week can be filtered using regular expressions based on the timestamp in the log file or the date field in the database to obtain the corresponding historical operation instructions. Natural language processing technology can then be used to parse the operation instructions to obtain the corresponding text data information. This helps to conduct in-depth analysis of the operation status of the energy management and control agent, identify potential problems, and optimize the operation strategy of the energy management and control agent.
[0042] Step S30: semantically map the text data information according to the pre-trained word vector model to obtain the semantic similarity between each energy management and control agent.
[0043] It should be noted that a word vector model is a mathematical model that maps words in a text into a low-dimensional vector space, such as word vector models like Word2Vec, GloVe, or BERT. A word vector model assigns a vector representation to each word. This vector representation captures the word's semantic information, placing semantically similar words closer together in the vector space. For example, a pre-trained word vector model can be trained based on a large-scale corpus (such as energy-related news articles and academic literature). When used to process text data converted from historical operational instructions of an energy management agent, it can convert words in the text into vectors with semantic meaning, providing a foundation for subsequent semantic mapping and similarity calculations.
[0044] It's understood that semantic similarity is a measure of the semantic proximity between two text fragments within a text data set. After converting historical operation instructions into text data using a word embedding model, the calculated semantic similarity indicates the degree of semantic similarity between the historical operation instructions of different agents. If the historical operation instructions of two energy management agents are semantically similar, it indicates that they share similar decision-making and operational logic or respond to similar situations, and this can be used as the semantic similarity between the energy management agents.
[0045] In a specific implementation, after converting the text data, the collaborative device can first input the text data into a pre-trained word vector model, which converts the text data into vectors. The model then calculates the distance between the vectors corresponding to different energy management agents. The smaller the distance, the higher the semantic similarity. By combining the vectors, the semantic similarity between each energy management agent can be obtained.
[0046] In a feasible implementation, step S30 of this embodiment may include the steps of: standardizing the text data information to obtain standardized text; converting each word in the standardized text into a corresponding word vector through a pre-trained word vector model to obtain a word vector group corresponding to the standardized text; clustering the word vector group to obtain multiple word vector clusters; and semantically mapping the word vector clusters through a cosine similarity algorithm to obtain the semantic similarity between each energy management and control intelligent entity.
[0047] It should be noted that standardized text is the text data obtained by standardizing the format and removing noise. For example, standardizing uppercase and lowercase letters, removing special symbols, and handling whitespace characters make the text data more standardized in structure and content expression, allowing subsequent word embedding models to process it more effectively, improving accuracy and consistency. For example, energy production data output by energy production agents, transmission capacity and distribution plan data from energy distribution agents, and demand forecast data from energy consumption management agents can all use the same units and data format.
[0048] It's easy to understand that a word vector group is a combination of vector representations of standardized text after it's been transformed using a pre-trained word vector model. A word vector group is a sequential collection of word vectors corresponding to each word in the standardized text. In other words, a word vector group represents a historical operation instruction and contains the representation vectors of the standardized text in the semantic space. Each vector contains the semantic information of the word.
[0049] It should be understood that a vector cluster is a collection of word vector groups formed by clustering them. Each word vector in a word vector group represents the semantic vector representation of a word. The clustering operation groups similar word vectors into a category based on the similarity measure between word vectors. These word vector groups that are grouped into the same category constitute a word vector cluster.
[0050] In another feasible implementation, the step of clustering the word vector group to obtain multiple word vector clusters described in this embodiment includes: determining the number of types of the energy management and control intelligent entities; randomly selecting initial cluster centers corresponding to the number of types from the word vector group; and clustering the word vector group through the K-Means clustering algorithm to obtain multiple word vector clusters.
[0051] It should be noted that the number of types refers to the number of different types of energy management agents in the energy management system. Initial cluster centers are key elements for initiating the clustering process. They serve as the "core" that attracts surrounding similar word vector groups. A number of word vectors corresponding to the number of energy management agent types can be randomly selected from the word vector group as initial cluster centers.
[0052] The K-Means clustering algorithm is a clustering algorithm that iteratively updates cluster centers and redistributes data points (i.e., word vector groups) to the nearest cluster center. Using the K-Means clustering algorithm, word vector groups are divided into K clusters (K is the number of energy management agent types. For example, if energy management agents are divided into energy production agents, energy distribution agents, and energy consumption management agents, then k = 3). The sum of the distances from each word vector group within a cluster to its cluster center is minimized. By continuously iteratively updating the cluster centers, a stable clustering result is ultimately achieved, resulting in multiple word vector clusters.
[0053] In this embodiment, the word vector groups related to the energy management and control agent are effectively classified, which helps to deeply understand the semantic structure of the text related to the energy management and control agent.
[0054] It should be noted that the cosine similarity algorithm is used to measure the similarity between two vectors. The cosine similarity algorithm expresses their similarity by calculating the cosine of the angle between two vectors. The cosine value ranges from -1 to 1, with values closer to 1 indicating greater similarity between the two vectors. In semantic mapping, it is used to determine the semantic similarity between word vector clusters.
[0055] For example, when the Energy Production Agent sends a message about "increasing electricity production," the Energy Distribution Agent needs to understand the semantic connection between this message and its own action of "allocating more electricity." If the calculated semantic similarity exceeds a threshold (e.g., 0.8), the two messages or actions are considered semantically related and can collaborate.
[0056] For example, during the clustering process of the K-Means clustering algorithm, for each energy management agent's standardized text, the word vector group corresponding to it can be calculated with its similarity to (k) cluster centers, and the word vector group can be assigned to the cluster with the closest cluster center. Then, the update phase begins, recalculating the center of each cluster, and repeating the above assignment phase with the new center. This continues until the cluster center no longer changes significantly. After clustering is completed, energy management agents in the same cluster have a high degree of semantic similarity, while agents in different clusters have a relatively low degree of semantic similarity.
[0057] In this embodiment, through the above-mentioned standardization processing, word vector conversion, clustering and cosine similarity calculation, the text semantic information can be deeply mined and the relationship between the optimized intelligent agents can be accurately determined.
[0058] Step S40: allocating a local reward indicator to each energy management and control agent according to the semantic similarity and the priority factor.
[0059] It's important to note that local reward metrics are used to evaluate the performance of each energy management agent. They reflect the performance of the energy management agent within a local scope (relative to its own operations and its relationships with other agents). Local reward metrics focus on the performance of individual agents, motivating each agent to perform its tasks optimally according to its role and functional characteristics within the system.
[0060] By combining semantic similarity (reflecting the degree of correlation between operations between energy management agents) and priority factors (reflecting the importance of energy management agents) to assign local reward indicators, the contribution of local reward indicators can be evaluated more comprehensively.
[0061] In a specific implementation, after obtaining the above-mentioned semantic similarity and priority factors, the collaborative device can set a comprehensive calculation formula, for example, multiplying the semantic similarity score by a weight coefficient (which can be determined based on actual conditions, such as the system's emphasis on semantic matching and priority) plus the priority factor score multiplied by another weight coefficient to obtain a local reward indicator.
[0062] In another feasible implementation, step S40 of this embodiment may include the steps of: defining fuzzy sets corresponding to the semantic similarity and the priority factors in each energy management and control agent; establishing a relationship between the semantic similarity and the priority factors and the local reward index to obtain a fuzzy rule; according to the fuzzy rule, performing fuzzy reasoning on the fuzzy set by the Mamdani reasoning method to obtain the membership of the fuzzy set; defuzzifying the membership by the maximum membership method to obtain the local reward index corresponding to each energy management and control agent.
[0063] It should be noted that fuzzy sets are used to describe the uncertainty of semantic similarity and priority factors within a certain range. For example, for semantic similarity, three fuzzy sets can be defined: "low," "medium," and "high." For priority, three fuzzy sets can be defined: "low priority," "medium priority," and "high priority."
[0064] Fuzzy rules are used to characterize the relationship between semantic similarity and priority factors and local reward indicators. For example, if semantic similarity is "high" and priority is "high priority," the local reward indicator is "high reward." If semantic similarity is "medium" and priority is "medium priority," the local reward indicator is "medium reward." Fuzzy rules can be determined based on the knowledge of domain experts or analysis of energy management systems.
[0065] It can be understood that the Mamdani reasoning method is a method of reasoning in fuzzy logic that can handle fuzziness and uncertainty, transforming fuzzy input information into fuzzy output information, and achieving reasoning from fuzzy input to fuzzy output. Membership is a term used to represent the degree to which an element in a fuzzy set belongs to that fuzzy set.
[0066] For example, using Mamdani reasoning, for each fuzzy rule, the rule's excitation strength is calculated based on the premise (semantic similarity and priority membership). Then, based on the rule's conclusion (a fuzzy set of local reward metrics) and the excitation strength, the membership of the output fuzzy set is determined. For example, if an agent has a semantic similarity of 0.6, based on the defined fuzzy rules, it may have a membership of 0.8 in the semantic similarity fuzzy set and a membership of 0.2 in the "high" semantic similarity fuzzy set.
[0067] It should be understood that the maximum membership method is a defuzzification technique. After obtaining the membership of a fuzzy set, a specific value is ultimately required as the local reward indicator for each energy management agent. The maximum membership method selects the clear value corresponding to the element with the highest membership. For example, if the fuzzy set of local reward indicators obtained through Mamdani reasoning has the highest membership for the element "higher reward," then the specific value corresponding to "higher reward" is used as the final local reward indicator.
[0068] In this implementation, considering that concepts such as semantic similarity and priority factors are often difficult to define with precise numerical values in actual energy management systems, the use of fuzzy sets effectively addresses the uncertainty inherent in these concepts. By establishing fuzzy rules, complex relationships can be represented in an intuitive manner, making them easier to understand and adjust. Finally, defuzzification using the maximum membership method yields a clear local reward metric. This approach not only addresses uncertainty but also ultimately yields a specific numerical value that can be used in practical operations (such as reward allocation), helping to optimize the performance of the energy management agent within the system.
[0069] Step S50: Based on the local reward index, each energy management and control agent is globally optimized according to the collaboration contribution to obtain an optimized management and control strategy, so that each energy management and control agent collaborates based on the optimized management and control strategy.
[0070] It should be noted that collaborative contribution is a metric that measures the degree to which an energy management agent contributes to the achievement of overall energy management goals through collaboration with other agents. For example, if the operations of an energy management agent influence the operations of other energy management agents and promote the stable and efficient operation of the entire campus energy system, resulting in more stable energy supply and more efficient energy utilization, then the collaborative contribution of the energy management agent is high.
[0071] It is understandable that the optimized management and control strategy is a management and control plan obtained in the global optimization process. It clarifies the role, task allocation, resource usage, and mutual collaboration of each energy management and control intelligent agent in the system, weighs the interests and mutual influences between each energy management and control intelligent agent, and can maximize the overall benefits of the system.
[0072] In practice, after evaluating the performance of each energy management agent using a local reward metric, the collaborative contribution of each agent can be quantified. This collaborative contribution and local reward metric are then incorporated into a global optimization model. This optimization calculation takes into account multiple factors, such as energy utilization and cost, to produce an optimized management strategy that enables all agents to collaborate better.
[0073] In the technical solution provided in this embodiment, the importance of each energy management agent in various aspects of energy generation, transmission, storage, and use is first assessed, ranked by importance, and a priority factor for each energy management agent is determined. Historical operation instructions are then located from the energy management agent's logs. Using a data conversion tool, these instructions are converted into text format according to the instruction format, resulting in text data information. The text data information is then input into a pre-trained word vector model, which converts the text data information into vectors. The distance between the corresponding vectors of different energy management agents is calculated. The smaller the distance, the higher the semantic similarity. By combining the vectors, the semantic similarity between each energy management agent can be determined. A comprehensive calculation formula can then be established. For example, the semantic similarity score is multiplied by a weight coefficient (which can be determined based on actual conditions, such as the system's emphasis on semantic matching and priority) and the priority factor score is multiplied by another weight coefficient to obtain a local reward index. Finally, the collaborative contribution of each energy management agent is quantified. The collaborative contribution and local reward index are incorporated into a global optimization model. This optimization calculation takes into account multiple factors, such as energy utilization and cost, to produce an optimized management strategy that enables each energy management agent to collaborate according to this strategy. Since this embodiment determines the semantic similarity between each energy management and control agent based on the historical operation instructions output by the energy management and control agent, allocates local reward indicators to each energy management and control agent through semantic similarity and priority factors, and finally performs global optimization in combination with the collaborative contribution to obtain the optimized management and control strategy, the collaboration between energy management and control agents can be optimized in real time, so that the agents can learn the optimized management and control strategy more quickly and adapt to the complex and changing energy demand, so that these agents can work together efficiently to achieve optimized management and control of regional energy, ensuring the maximization of the overall utilization efficiency of the agents.
[0074] Based on the above embodiment 1 of this application, the second embodiment of this application is proposed. In the second embodiment of this application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be repeated hereafter. Figure 2 , Figure 2 A flow chart of the second embodiment of the intelligent agent collaboration method of this application is provided.
[0075] In this example, step S50 includes steps S51 to S53: Step S51: constructing a relationship model between energy management and control agents based on the collaborative relationship between the energy management and control agents.
[0076] It should be noted that the collaborative relationship refers to the relationship of mutual cooperation, coordination and influence between different energy management and control intelligent entities in the energy management and control system.
[0077] It can be understood that the relational model is a model that abstracts and describes the collaborative relationship between various energy management and control agents. It can be expressed in the form of a mathematical structure or logical framework by analyzing the interaction, information exchange and other relationships between energy management and control agents in energy management tasks.
[0078] For example, there may be a synergistic relationship between the power agent and the heat agent because some cogeneration equipment can generate both electricity and heat.
[0079] Step S52: Based on the relationship model, a cost-benefit analysis method is used to quantify the collaborative contribution of each energy management and control agent to other energy management and control agents.
[0080] It's important to note that cost-benefit analysis is an evaluation method used to measure the inputs (costs) and outputs (benefits) of each energy management and control system in its collaboration with other energy management and control systems. Costs can include energy consumption, computing resource usage, and information transmission costs, while benefits can be reflected in improving energy efficiency, reducing energy waste, and enhancing system stability.
[0081] For example, if a power agent's decision to adjust its power generation plan reduces the heating costs of a thermal agent, the power agent's collaborative contribution to the thermal agent can be quantified based on the magnitude of the cost reduction. A cost-benefit analysis can be used to determine the contribution of each energy management agent based on the impact of its decision on overall system performance.
[0082] Step S53: Based on the collaborative contribution, the energy management and control agents are globally optimized through a multi-agent proximal strategy optimization algorithm to obtain an optimized management and control strategy.
[0083] It's important to note that the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm is used to optimize the behavioral strategies of multiple energy management agents. In energy management scenarios, MAPPO uses data such as the collaborative contributions of each energy management agent as input to continuously adjust the strategies of the energy management agents, exploring optimal strategy combinations locally (at the local scale). This allows the overall strategies of multiple energy management agents to reach a global optimum, thereby improving the performance of the entire energy management system.
[0084] In a feasible implementation, step S53 of this example includes the steps of: modeling the energy management and control environment of the industrial park and determining a state space, wherein the state space includes the key energy information and the collaborative contribution of each energy management and control agent; constructing a global policy network and a global value network for each energy management and control agent; training the global policy network and the global value network through a multi-agent proximal policy optimization algorithm based on the state space, and continuing the training until the network converges; when the global policy network and the global value network converge, extracting the optimized management and control strategy from the trained global policy network.
[0085] It's important to note that the state space describes all possible states of an energy management system. In the context of energy management in an industrial park, the state space includes key energy information and collaborative contributions of each energy management agent. This key energy information can include key information such as energy reserves, energy demand, and energy prices.
[0086] The global policy network is a neural network structure used to generate action strategies for energy management agents. In a multi-agent system, each energy management agent needs to determine its own actions based on the system's state. The global policy network is designed to learn appropriate decision-making strategies from the overall state space. The global policy network receives input from the state space (including key energy information and collaborative contributions of each energy management agent). After internal neural network calculations and weight adjustments, it outputs action strategies tailored to each agent, such as energy allocation ratios and equipment operation instructions, to achieve optimal energy management for the entire industrial park.
[0087] It should be understood that the global value network is a neural network used to evaluate the expected value of an energy management system adopting a specific strategy under specific conditions. The global value network takes the state space as input and outputs a numerical value representing the value, which quantitatively evaluates the achievement of the overall energy management goal (such as cost minimization).
[0088] Specifically, the energy management system's environment must first be modeled to determine the state space. This state space contains key information about each energy management agent, such as energy reserves, energy demand, and energy prices. Collaborative contributions are also incorporated into this state space to form a complete state representation. Next, a global policy network and a global value network are constructed for each energy management agent from a holistic perspective. The global policy network takes the state of the energy management system as input and outputs the probability distribution of different actions taken by the energy management agent. The global value network estimates the long-term value of the energy management agent under a given state.
[0089] Next, the energy management agents interact within the energy management system according to the initial strategy and collect experience samples. These experience samples may include the current state, actions taken, rewards received, and the next state. A multi-agent proximal policy optimization algorithm is then used to update the parameters of the global policy network and global value network based on the collected experience samples. The policy network parameters are updated by maximizing an objective function related to the action probabilities and advantage functions output by the policy network. Simultaneously, the value network parameters are updated by minimizing the mean squared error. Training continues until the global policy network and global value network converge. Finally, when the algorithm converges, optimized management policies are extracted from the trained global policy network. These policies represent the optimal action selection options for each energy management agent under different states, thus achieving global optimization.
[0090] The technical solution provided in this embodiment models the industrial park's energy management environment and determines the state space containing key energy information and collaborative contributions. This allows for a deeper understanding of the operating mechanisms and relationships within the complex energy management system. This enables management strategies to comprehensively consider various factors, more comprehensively reflecting the system state and avoiding situations where a single agent's local optimum results in poor overall performance. Furthermore, the use of a multi-agent proximal strategy optimization algorithm enables global optimization, resulting in an optimized management strategy that improves energy management efficiency.
[0091] It should be noted that the above examples are only used to understand this application and do not constitute a limitation on the intelligent body collaboration method of this application. More simple transformations based on this technical concept are all within the scope of protection of this application.
[0092] This application also provides an intelligent collaborative device, please refer to Figure 3 , Figure 3 This is a schematic diagram of the module structure of the intelligent collaborative device according to an embodiment of the present application; the intelligent collaborative device includes: The priority determination module 301 is used to determine the priority factors of each energy management and control agent in the energy management and control system of the industrial park; An operation interception module 302 is used to intercept historical operation instructions output by the energy management and control agent and convert the historical operation instructions into text data information; A semantic mapping module 303 is used to perform semantic mapping on the text data information according to a pre-trained word vector model to obtain the semantic similarity between each energy management and control agent; A local reward module 304 is configured to allocate a local reward indicator to each energy management agent based on the semantic similarity and the priority factor; The global optimization module 305 is used to perform global optimization on each energy management and control agent according to the collaboration contribution based on the local reward index to obtain an optimized management and control strategy so that each energy management and control agent can collaborate based on the optimized management and control strategy.
[0093] As an embodiment, the operation interception module 302 is also used to retrieve the operation log of the energy management and control intelligent body from the energy management and control system; match historical operation instructions within a preset time period in the operation log through regular expressions; and use natural language processing technology to parse the historical operation instructions to obtain text data information of the energy management and control intelligent body.
[0094] As an embodiment, the semantic mapping module 303 is also used to standardize the text data information to obtain standardized text; convert each word in the standardized text into a corresponding word vector through a pre-trained word vector model to obtain a word vector group corresponding to the standardized text; cluster the word vector group to obtain multiple word vector clusters; and semantically map the word vector clusters through a cosine similarity algorithm to obtain the semantic similarity between each energy management and control intelligent entity.
[0095] As an embodiment, the semantic mapping module 303 is also used to determine the number of types of the energy management and control intelligent agent; randomly select the initial clustering centers corresponding to the number of types from the word vector group; cluster the word vector group through the K-Means clustering algorithm to obtain multiple word vector clusters.
[0096] As an implementation method, the local reward module 304 is also used to define the fuzzy sets corresponding to the semantic similarity and the priority factors in each energy management and control intelligent agent; establish a relationship between the semantic similarity and the priority factors and the local reward indicators to obtain fuzzy rules; according to the fuzzy rules, fuzzy reasoning is performed on the fuzzy sets by the Mamdani reasoning method to obtain the membership of the fuzzy sets; the membership is defuzzified by the maximum membership method to obtain the local reward indicators corresponding to each energy management and control intelligent agent.
[0097] As an implementation method, the global optimization module 305 is also used to construct a relationship model between the energy management and control agents based on the collaborative relationship between the energy management and control agents; based on the relationship model, a cost-benefit analysis method is used to quantify the collaborative contribution of each energy management and control agent to other energy management and control agents; based on the collaborative contribution, the energy management and control agents are globally optimized through a multi-agent proximal strategy optimization algorithm to obtain an optimized management and control strategy.
[0098] As an implementation method, the global optimization module 305 is also used to model the energy management and control environment of the industrial park and determine the state space, which includes the key energy information and the collaborative contribution of each energy management and control agent; construct a global policy network and a global value network for each energy management and control agent; based on the state space, the global policy network and the global value network are trained through a multi-agent proximal policy optimization algorithm, and the training is continued until the network converges; when the global policy network and the global value network converge, the optimized management and control strategy is extracted from the trained global policy network.
[0099] Other embodiments or specific implementation methods of the intelligent body collaboration device of the present application can refer to the above-mentioned method embodiments and will not be repeated here.
[0100] The intelligent agent collaboration device provided by this application adopts the intelligent agent collaboration method in the above embodiment, which can solve the technical problem that the scheduling strategy of the energy management and control system of the traditional industrial park generally relies on the fixed rules of the intelligent agent. Since the energy demand of the industrial park is complex and changeable, this static scheduling mode is difficult to ensure the maximization of the overall utilization efficiency of the intelligent agent. Compared with the existing technology, the beneficial effects of the intelligent agent collaboration device provided by this application are the same as the beneficial effects of the intelligent agent collaboration method provided by the above embodiment, and the other technical features of the intelligent agent collaboration device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0101] The present application provides an intelligent agent collaboration device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the intelligent agent collaboration method in the above-mentioned embodiment one.
[0102] Reference below Figure 4 , Figure 4 The figure is a schematic diagram of the device structure of the hardware operating environment involved in the intelligent agent collaboration method in the embodiment of the present application, which shows a schematic diagram of the structure of the intelligent agent collaboration device suitable for implementing the embodiment of the present application. The intelligent agent collaboration device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4The intelligent agent collaboration device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
[0103] like Figure 4 As shown, the intelligent collaborative device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the intelligent collaborative device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the intelligent collaborative device to communicate with other devices wirelessly or wired to exchange data. Although the figure shows an intelligent collaborative device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have instead.
[0104] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0105] The intelligent agent collaboration device provided by this application adopts the intelligent agent collaboration method in the above embodiment, which can solve the technical problem that the scheduling strategy of the energy management and control system of the traditional industrial park generally relies on the fixed rules of the intelligent agent. Since the energy demand of the industrial park is complex and changeable, this static scheduling mode is difficult to ensure the maximization of the overall utilization efficiency of the intelligent agent. Compared with the existing technology, the beneficial effects of the intelligent agent collaboration device provided by this application are the same as the beneficial effects of the intelligent agent collaboration method provided by the above embodiment, and the other technical features of the intelligent agent collaboration device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0106] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0107] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0108] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the intelligent agent collaboration method in the above-mentioned embodiment.
[0109] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0110] The above-mentioned computer-readable storage medium may be included in the intelligent agent collaboration device; or it may exist independently without being assembled into the intelligent agent collaboration device.
[0111] The computer-readable storage medium carries one or more programs, which, when executed by the agent collaboration device, cause the agent collaboration device to: determine priority factors for each energy management agent in the energy management and control system of the industrial park; Intercepting historical operation instructions output by the energy management and control agent, and converting the historical operation instructions into text data information; performing semantic mapping on the text data information according to a pre-trained word vector model to obtain the semantic similarity between each energy management and control agent; allocating a local reward index to each energy management and control agent according to the semantic similarity and the priority factor; based on the local reward index, globally optimizing each energy management and control agent according to the collaborative contribution to obtain an optimized management and control strategy, so that each energy management and control agent can collaborate based on the optimized management and control strategy.
[0112] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0114] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0115] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned agent collaboration method. This can solve the technical problem that the scheduling strategy of the traditional industrial park energy management and control system generally relies on the fixed rules of the agents. Due to the complex and changeable energy demand of the industrial park, this static scheduling model is difficult to ensure the maximization of the overall utilization efficiency of the agents. Compared with the existing technology, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the agent collaboration method provided in the above-mentioned embodiment, and will not be repeated here.
[0116] The above description is only part of the embodiments of the present application and does not limit the scope of protection of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the scope of protection of the present application.
Claims
1. An agent collaboration method, characterized in that: The method comprises: Determine the priority factors of each energy management agent in the energy management system of the industrial park; Intercepting historical operation instructions output by the energy management and control intelligent agent and converting the historical operation instructions into text data information; Perform semantic mapping on the text data information based on the pre-trained word vector model to obtain the semantic similarity between each energy management and control agent; assigning a local reward indicator to each energy management agent based on the semantic similarity and the priority factor; Based on the local reward index, each energy management and control intelligent agent is globally optimized according to the collaboration contribution to obtain an optimized management and control strategy, so that the energy management and control intelligent agents can collaborate based on the optimized management and control strategy.
2. The method according to claim 1, wherein The step of intercepting the historical operation instructions output by the energy management and control intelligent agent and converting the historical operation instructions into text data information includes: Retrieving the operation log of the energy management and control agent from the energy management and control system; Matching historical operation instructions within a preset time period in the operation log using a regular expression; The historical operation instructions are parsed using natural language processing technology to obtain text data information of the energy management and control intelligent body.
3. The method according to claim 2, wherein The step of performing semantic mapping on the text data information according to the pre-trained word vector model to obtain the semantic similarity between each energy management and control agent includes: Performing standardization processing on the text data information to obtain a standardized text; Convert each word in the standardized text into a corresponding word vector using a pre-trained word vector model to obtain a word vector group corresponding to the standardized text; Performing clustering processing on the word vector group to obtain multiple word vector clusters; The word vector cluster is semantically mapped using the cosine similarity algorithm to obtain the semantic similarity between each energy management and control agent.
4. The method according to claim 3, wherein The step of clustering the word vector group to obtain multiple word vector clusters includes: Determining the number of types of the energy management and control agents; Randomly select initial cluster centers corresponding to the number of types from the word vector group; The word vector group is clustered using the K-Means clustering algorithm to obtain multiple word vector clusters.
5. The method according to claim 4, wherein The step of allocating a local reward indicator to each energy management and control agent according to the semantic similarity and the priority factor includes: Defining fuzzy sets corresponding to the semantic similarity and the priority factors in each energy management and control agent; Establishing a relationship between the semantic similarity, the priority factor and the local reward indicator to obtain a fuzzy rule; According to the fuzzy rules, fuzzy reasoning is performed on the fuzzy set by using the Mamdani reasoning method to obtain the membership degree of the fuzzy set; The membership is defuzzified by the maximum membership method to obtain the local reward index corresponding to each energy management agent.
6. The method according to any one of claims 1 to 5, characterized in that The step of performing global optimization on each energy management and control agent according to the collaborative contribution based on the local reward indicator to obtain an optimized management and control strategy includes: According to the collaborative relationship between energy management and control agents, a relationship model between energy management and control agents is constructed; Based on the relationship model, a cost-benefit analysis method is used to quantify the collaborative contribution of each energy management agent to other energy management agents; Based on the collaborative contribution, the energy management and control agents are globally optimized through a multi-agent proximal strategy optimization algorithm to obtain an optimized management and control strategy.
7. The method according to claim 6, wherein The step of performing global optimization on each energy management agent based on the collaborative contribution by using a multi-agent proximal strategy optimization algorithm to obtain an optimized management strategy includes: Modeling the energy management and control environment of the industrial park to determine a state space, wherein the state space includes key energy information and the collaborative contribution of each energy management and control agent; Build a global strategy network and a global value network for each energy management agent; According to the state space, the global policy network and the global value network are trained by a multi-agent proximal policy optimization algorithm, and the training is continued until the network converges; When the global policy network and the global value network converge, an optimized management and control strategy is extracted from the trained global policy network.
8. An intelligent agent collaboration device, characterized in that: The device comprises: A priority determination module is used to determine the priority factors of each energy management and control agent in the energy management and control system of the industrial park; An operation interception module, configured to intercept historical operation instructions output by the energy management and control agent and convert the historical operation instructions into text data information; A semantic mapping module is used to perform semantic mapping on the text data information based on a pre-trained word vector model to obtain the semantic similarity between each energy management and control agent; A local reward module, configured to assign a local reward indicator to each energy management agent based on the semantic similarity and the priority factor; A global optimization module is used to perform global optimization on each energy management and control agent according to the collaboration contribution based on the local reward index to obtain an optimized management and control strategy so that each energy management and control agent can collaborate based on the optimized management and control strategy.
9. An intelligent collaborative device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the intelligent agent collaboration method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the intelligent agent collaboration method as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Intelligent agent optimal strategy acquisition method and device
CN113128705A
Distributed energy coordination control method based on multi-agent reinforcement learning
CN119323301A
Heterogeneous multi-agent cooperative comprehensive park electrical energy scheduling method and system
CN119476855A
Intelligent collaborative recommendation method and device for multi-modal immersive content
CN120086453A
Power transmission and distribution production task cooperation system and method based on intelligent agent
CN120338452A
Cited By
All-optical computing power network resource dynamic scheduling method based on multi-agent collaboration
CN122248297A