Task collaboration method and device and related equipment
By constructing an intent recognition model in a multi-agent system and performing conflict detection and priority adjustment, the problem of insufficient task collaboration capability is solved, and more efficient task collaboration execution is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-03-27
AI Technical Summary
Existing multi-agent systems have weak task collaboration capabilities, mainly due to the use of master-slave control architecture and static task allocation strategies, which cannot effectively support complex collaboration mechanisms.
By acquiring various types of business data for attribution modeling, multiple intent recognition models are constructed. Conflict detection and task prioritization are used to adjust the agent's task execution parameters, thereby achieving federated collaboration and reducing conflicts between tasks.
It improves the task coordination capability of multi-agent systems, effectively avoids and mitigates conflicts in task execution, and enhances overall coordination efficiency.
Smart Images

Figure CN121743005A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a task collaboration method, apparatus and related equipment. Background Technology
[0002] Currently, Multi-Agent Systems (MAS), as the core architecture for realizing automated decision-making and execution of complex tasks, have been widely applied in tasks such as network scheduling, energy-saving control, and complaint handling. However, in most publicly available technical solutions, task collaboration still adopts a "master-slave" control architecture or a static task allocation strategy. This approach cannot effectively support complex collaboration mechanisms, resulting in weak task collaboration capabilities among multi-agent systems. Summary of the Invention
[0003] This application provides a task collaboration method, apparatus, and related equipment, which can solve the technical problem of weak task collaboration capabilities among multiple agents.
[0004] In a first aspect, embodiments of this application provide a task collaboration method, the method comprising:
[0005] Multiple types of business data are acquired, and multiple intelligent agents are used to perform attribution modeling based on the multiple types of business data and preset task triggering conditions to obtain multiple intent recognition models.
[0006] Conflict detection is performed on the task execution parameters of the multiple intent recognition models to obtain conflict detection results;
[0007] Determine the task priorities of the multiple intent recognition models;
[0008] Based on the conflict detection results and the task priority, the task execution parameters of the multiple intent recognition models are adjusted collaboratively by the multiple intelligent agents to obtain the adjusted multiple intent recognition models.
[0009] Task execution is performed based on the adjusted multiple intent recognition models.
[0010] The optional task execution parameters for each of the multiple intent recognition models include task type, target region, task time window, and expected metric change.
[0011] Optionally, the conflict detection performed on the task execution parameters of the plurality of intent recognition models to obtain conflict detection results includes at least one of the following:
[0012] The target regions of the multiple intent recognition models are respectively subjected to region overlap detection to obtain the first conflict detection result;
[0013] Time window overlap detection is performed on the task time windows of the multiple intent recognition models respectively to obtain the second conflict detection result;
[0014] The direction of index change is detected for the expected index changes of the multiple intent recognition models respectively, and the third conflict detection result is obtained.
[0015] Optionally, determining the task priority of the plurality of intent recognition models includes:
[0016] A first intermediate value is determined based on the preset score of the urgency of the complaint from the multiple intent recognition models and the first preset weight;
[0017] The second intermediate value is determined based on the preset score and second preset weight of the energy-saving benefit of the multiple intent recognition models;
[0018] A third intermediate value is determined based on the network influence preset score and the third preset weight of the multiple intent recognition models;
[0019] The priority scores of the multiple intent recognition models are determined based on the sum of the first intermediate value, the second intermediate value, and the third intermediate value. The priority scores of the multiple intent recognition models are then sorted to obtain the task priority of the multiple intent recognition models.
[0020] Optionally, the method further includes:
[0021] A digital twin simulation platform is constructed based on the state space and action space of the multiple intelligent agents;
[0022] Based on a preset reward function, the result of the task execution, and a preset reinforcement learning algorithm, the parameter combination of the multiple agents is federated on the digital twin simulation platform to obtain a globally optimized federated parameter aggregation result. The parameter combination of the multiple agents is then updated according to the federated parameter aggregation result.
[0023] By utilizing the updated multiple intelligent agents, attribution modeling is performed based on the multiple types of business data and preset task triggering conditions to obtain multiple intent recognition models;
[0024] Based on the conflict detection results and the task priority, the task execution parameters of the multiple intent recognition models are adjusted collaboratively by the updated multiple agents to obtain the adjusted multiple intent recognition models.
[0025] Optionally, the method further includes:
[0026] The adjusted multiple intent recognition models are simulated on the digital twin simulation platform to obtain simulation results, and the risk of task execution is assessed based on the simulation results.
[0027] Optionally, the preset reward function is determined by weighting and summing the service quality score corresponding to the task execution result, the resource consumption corresponding to the task execution result, and the conflict penalty corresponding to the conflict detection result according to preset weights.
[0028] Secondly, embodiments of this application provide a task collaboration device, the device comprising:
[0029] The acquisition module is used to acquire multiple types of business data, and use multiple intelligent agents to perform attribution modeling based on the multiple types of business data and preset task triggering conditions to obtain multiple intent recognition models.
[0030] The first processing module is used to perform conflict detection on the task execution parameters of the multiple intent recognition models respectively, and obtain conflict detection results;
[0031] The second processing module is used to determine the task priority of the plurality of intent recognition models;
[0032] The third processing module is used to adjust the task execution parameters of the multiple intent recognition models in collaboration with the multiple intelligent agents based on the conflict detection results and the task priority, so as to obtain the adjusted multiple intent recognition models.
[0033] The execution module is used to perform task execution based on the adjusted multiple intent recognition models.
[0034] Thirdly, embodiments of this application provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the task coordination method as described in the first aspect.
[0035] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the task coordination method as described in the first aspect.
[0036] Fifthly, embodiments of this application provide a computer program product including computer instructions that, when executed by a processor, implement the steps of the task coordination method as described in the first aspect.
[0037] In this embodiment, multiple agents perform attribution modeling based on the various types of business data and preset task triggering conditions to obtain multiple intent recognition models. The multiple agents coordinate to adjust the task execution parameters of the multiple intent recognition models based on conflict detection results and task priorities. This can effectively reduce or avoid conflicts between different tasks during task execution, thereby improving the task collaboration capability of multiple agents. Attached Figure Description
[0038] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a flowchart of a task collaboration method provided in an embodiment of this application;
[0040] Figure 2 This is a schematic diagram of the structure of a task collaboration device provided in an embodiment of this application;
[0041] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0042] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0043] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "and / or" in this application indicates at least one of the connected objects. For example, the scope of protection of "A and / or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. Additionally, the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0044] See Figure 1 , Figure 1 This is a flowchart of a task collaboration method provided in an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:
[0045] Step 101: Acquire multiple types of business data, and use multiple intelligent agents to perform attribution modeling based on the multiple types of business data and preset task triggering conditions to obtain multiple intent recognition models;
[0046] The various types of business data can be data from multiple sources, such as: Key Performance Indicators (KPIs), alarm logs, user complaints, semantic descriptions, etc. It should be noted that the types of business data can be set as needed by those skilled in the art.
[0047] The multiple intelligent agents can be intelligent agents corresponding to multiple types of business data. For example, one intelligent agent can correspond to one type of business data. It should be noted that the "multiple intelligent agents" here can be not only multiple intelligent agent instances (for example, one intelligent agent can be one intelligent agent instance), but also a collection of multiple types of intelligent agents (for example, one intelligent agent can be one type of intelligent agent, or a collection of intelligent agents composed of one type of intelligent agents).
[0048] For example, the plurality of intelligent agents may include at least one of the following five types of intelligent agents:
[0049] The fault handling intelligent agent can be used to automatically sense the location and type of network faults, and combine historical experience knowledge base to perform root cause localization and closed-loop repair operations.
[0050] The network optimization agent can dynamically configure and schedule wireless resources based on perceived user behavior and load conditions to improve network throughput and quality of service (QoS); it can also be responsible for ensuring network quality (such as perceived negative feedback and abnormal indicators).
[0051] The complaint analysis agent can parse user complaint content using Natural Language Processing (NLP) algorithms, mapping natural language into spatial-temporal-task metadata to achieve rapid location and task backtracking; it can also be responsible for complaint response (such as improving user coverage or optimizing KPIs).
[0052] Energy-optimizing agents can achieve cell energy-saving scheduling and beam power control based on multi-site collaborative strategies, while ensuring communication quality, thereby improving green energy-saving effects; they can also be responsible for energy-saving control (such as shutting down low-load equipment).
[0053] The Fifth Generation Mobile Communication Advanced (5G-A) precision number allocation intelligent agent can combine location information, load balancing and user behavior characteristics to realize intelligent number allocation strategies and interference prediction, and is adapted to the collaborative scheduling of heterogeneous sites.
[0054] The task triggering conditions can be set by those skilled in the art as needed or according to predetermined rules.
[0055] For example, the attribution modeling using multiple agents based on the various types of business data and preset task triggering conditions can employ an intent classifier based on Bidirectional Encoder Representations from Transformers (BERT) and a rule knowledge graph matching strategy to perform attribution modeling on the task triggering conditions, thereby obtaining an intent recognition model.
[0056] The intent recognition model may include multiple task execution parameters related to task execution.
[0057] In this step, since each agent performs attribution modeling for a different type of business data, a distributed modeling and processing mechanism is formed. This allows the features of different types of business data to be fully mined and represented in their respective agents, thereby improving the targeting and accuracy of the overall intent recognition modeling and reducing single points of failure and performance bottlenecks that occur in centralized system processing.
[0058] Step 102: Perform conflict detection on the task execution parameters of the multiple intent recognition models respectively to obtain conflict detection results;
[0059] In this step, conflict detection is performed on the task execution parameters of the multiple intent recognition models to obtain conflict detection results. This can help identify conflicts between different models in terms of task execution parameters (such as target area, task time window and expected index changes, as described later), providing a data basis for adjusting the task parameters of the intent recognition models.
[0060] Step 103: Determine the task priority of the multiple intent recognition models;
[0061] In this step, the task execution of multiple intent recognition models can be prioritized based on factors such as the business importance, real-time requirements, resource consumption, and historical execution results of the tasks corresponding to each intent recognition model, thereby providing a data basis for adjusting the task parameters of the intent recognition models.
[0062] Step 104: Based on the conflict detection results and the task priority, the task execution parameters of the multiple intent recognition models are adjusted collaboratively by the multiple intelligent agents to obtain the adjusted multiple intent recognition models;
[0063] The collaborative adjustment of task execution parameters of multiple intelligent agents by multiple intent recognition models can be achieved by introducing a federated scheduling mechanism, enabling multiple intelligent agents to form a "federated collaboration" logic at the task level. Specifically, multiple intelligent agents can share information, compromise resources, and transfer intents according to a preset communication strategy, thereby realizing cross-intent conflict handling and priority game; based on the conflict detection results and combined with the current task priority, the task execution parameters are dynamically adjusted to achieve adaptive resource allocation and collaborative correction of task execution paths, thereby improving task collaboration capabilities.
[0064] In this step, the task execution parameters of the multiple intent recognition models are adjusted collaboratively by the multiple intelligent agents, so that the task execution parameters of each intent recognition model can be dynamically optimized under global conflict state and priority constraints, thereby effectively eliminating or mitigating resource competition and execution conflicts among multiple intent recognition models, and thus improving task collaboration capabilities.
[0065] Step 105: Perform the task according to the adjusted multiple intent recognition models.
[0066] The task execution based on the adjusted multiple intent recognition models can be performed by the multiple agents independently executing their corresponding tasks; or, under a preset collaborative strategy (such as a federated collaborative strategy), the multiple agents can collaboratively execute the tasks, for example, by decomposing tasks, allocating sub-tasks, and converging results among the agents to jointly complete the execution of complex tasks or cross-intent combination tasks; or, non-agent task execution units (such as task scheduling modules, business processing modules, etc.) can execute the corresponding tasks based on the multiple intent recognition models.
[0067] In this step, since task execution is based on multiple intent recognition models with adjusted task execution parameters, resource contention and mutual interference between different tasks can be avoided during the actual execution phase, thereby improving the overall task coordination capability.
[0068] In this embodiment, multiple agents are used to perform attribution modeling based on the various types of business data and preset task triggering conditions to obtain multiple intent recognition models. The multiple agents coordinate to adjust the task execution parameters of the multiple intent recognition models according to the conflict detection results and task priorities. During task execution, conflicts between different tasks can be effectively reduced or avoided, thereby improving the task coordination capability of multiple agents.
[0069] In some implementations, the task execution parameters of each of the plurality of intent recognition models include task type, target region, task time window, and expected metric change.
[0070] Specifically, the intent recognition model can be represented by the following first formula:
[0071]
[0072] In the first formula, Used to represent intent representations that can represent intent recognition models. It can be used to represent the first in multiple intent recognition models. Task execution parameters for an intent recognition model;
[0073] It can be used to represent task type; for example, it represents the type of task objective performed by the agent, specifically... ="Quality Guaranteed" ="Energy-saving control" =“Complaint Response”;
[0074] It can be used to represent the target area (Location), which corresponds to the geographical area or network unit where the task is performed (such as cell number, base station identifier (ID)).
[0075] It can be used to represent a task time window, which corresponds to the expected start and end times of a task;
[0076] It can be used to represent expected KPI shifts, which are changes in key performance indicators (KPIs) expected to occur due to the mission intent.
[0077] More specifically, in the first formula This can be expressed by the following second formula:
[0078]
[0079] In the second formula, It can be used to represent a target area. It can be used to represent in The first in Each region ; It can be used to represent a geographical area or network unit where a task is performed.
[0080] More specifically, in the first formula This can be expressed using the following third formula:
[0081]
[0082] In the third formula, It can be used to represent task time windows. It can be used to indicate the expected start time of a task. It can be used to indicate the expected completion time of a task.
[0083] More specifically, in the first formula This can be expressed by the following fourth formula:
[0084]
[0085] In the fourth formula, It can represent changes in expected indicators; These can represent the expected changes required to achieve a specific KPI (such as PRB utilization, power consumption, coverage, or user satisfaction). For example, in the fourth formula... Changes in expected indicators ( This can be expressed by the following fifth formula:
[0086]
[0087] In the fifth formula, It can be used to represent the first Expected change in the indicator; It can be used to represent the first Target values for the expected indicators; It can be used to represent the first Current value of the expected indicator.
[0088] For example, a task type of "complaint response," a target area of 102 communities, a task time window of 08:00 to 12:00, and an intent recognition model with expected changes in indicators of a 15% increase in coverage and a 20% increase in satisfaction can be expressed by the first formula:
[0089]
[0090] in, Intent representation used to represent the intent recognition model regarding complaint responses. With the first formula Corresponding, used to indicate task type; With the first formula Corresponding, used to represent the target area; With the first formula Correspondingly, it is used to represent the task time window; With the first formula Correspondingly, it is used to indicate the expected change in the indicator.
[0091] For example, the task type is "energy saving control", the target area is cell 102, the task time window is 08:00 to 20:00, and the expected indicator change is a 30% reduction in power consumption, the intent recognition model can be expressed by the first formula as follows:
[0092]
[0093] in, Intent representation used to represent the intent recognition model regarding energy-saving control. With the first formula Corresponding, used to indicate task type; With the first formula Corresponding, used to represent the target area; With the first formula Correspondingly, it is used to represent the task time window; With the first formula Correspondingly, it is used to indicate the expected change in the indicator.
[0094] In this embodiment, since the task execution parameters of each of the multiple intent recognition models include task type, target region, task time window, and expected indicator changes, on the one hand, it facilitates fine-grained comparison and conflict detection of task execution parameters between different intent recognition models; on the other hand, it also provides a parameter basis for subsequent adjustment of task execution parameters of intent recognition models, thereby better improving task collaboration capabilities.
[0095] In some implementations, the step of performing conflict detection on the task execution parameters of the plurality of intent recognition models to obtain conflict detection results includes at least one of the following:
[0096] The target regions of the multiple intent recognition models are respectively subjected to region overlap detection to obtain the first conflict detection result;
[0097] Time window overlap detection is performed on the task time windows of the multiple intent recognition models respectively to obtain the second conflict detection result;
[0098] The direction of index change is detected for the expected index changes of the multiple intent recognition models respectively, and the third conflict detection result is obtained.
[0099] In this embodiment, since conflict detection is performed on the task execution parameters of multiple intent recognition models from three dimensions—target area, task time window, and expected indicator change—the corresponding first, second, and third conflict detection results are obtained. These results can be used to avoid or mitigate mutual interference or repetitive operations of multiple tasks in the same target area, resource contention in the same time period, and mutual cancellation of execution objectives in the direction of expected indicator change.
[0100] Therefore, potential conflicts can be detected and located in advance in multiple dimensions such as space, time and changes in expected indicators, resulting in more comprehensive and refined conflict detection results. Based on the conflict detection results and the task priority, the task execution parameters of the multiple intent recognition models can be adjusted collaboratively by the multiple intelligent agents, which can further improve the collaborative ability of the multiple intelligent agents.
[0101] For example, region overlap detection is performed on the target regions of two different intent recognition models, and the target regions of the two different intent recognition models can be used as follows: and It means that if This indicates that the two intent recognition models may conflict in task execution within the same spatial region;
[0102] Time window overlap detection is performed on the task time windows of two different intent recognition models. The task time windows of the two different intent recognition models can be used as follows: and It means that if This indicates that the two intent recognition models are executing tasks at the same time, which poses a risk of actual resource conflict.
[0103] The direction of change of the expected index of two different intent recognition models is detected. The changes in the expected index of the two different intent recognition models can be used as... and It means that if This indicates that the two intent recognition models optimize in opposite directions for the same desired change in the metric (for example, one intent recognition model wants to increase the utilization of the Physical Resource Block (PRB), while the other intent recognition model wants to decrease the utilization of the PRB).
[0104] Furthermore, a conflict determination function can be used to detect conflicts in the task execution parameters of multiple intent recognition models. This conflict determination function can be expressed by the following sixth formula:
[0105]
[0106] In the sixth formula, Two intent recognition models and In collision detection, the numerical values "1" and "0" can be used to represent the numerical expression of the collision detection results. , and These respectively indicate that the two intent recognition models may conflict in task execution in the same spatial region, that there is a risk of actual resource conflict when the two intent recognition models execute tasks in the same time period, and that the two intent recognition models optimize the same expected indicator change in opposite directions.
[0107] In the sixth formula, in ( In this case, the collision detection result is represented by the value "1";
[0108] In the sixth formula, in ( In cases other than those described in the sixth formula, the conflict detection result is represented by the value "0".
[0109] Furthermore, the step of collaboratively adjusting the task execution parameters of the multiple intent recognition models based on the conflict detection results and the task priorities using the multiple intelligent agents may include:
[0110] If the conflicts among the multiple intent recognition models are severe (e.g., the conflict detection result is "1" in the sixth formula), the multiple agents can collaboratively adjust the task time windows of the multiple intent recognition models to delay task execution, and / or can collaboratively reduce the changes in the expected indicators of the multiple intent recognition models. ;
[0111] If there are conflicts among the multiple intent recognition models, a collaborative execution strategy can be adopted, for example, for target areas where the task type is energy-saving control. It can exclude target areas whose task type is complaint response. The intersection area will be used to adjust the target area for energy-saving control. To avoid the target area corresponding to the complaint response. .
[0112] In some implementations, determining the task priority of the plurality of intent recognition models includes:
[0113] A first intermediate value is determined based on the preset score of the urgency of the complaint from the multiple intent recognition models and the first preset weight;
[0114] The second intermediate value is determined based on the preset score and second preset weight of the energy-saving benefit of the multiple intent recognition models;
[0115] A third intermediate value is determined based on the network influence preset score and the third preset weight of the multiple intent recognition models;
[0116] The priority scores of the multiple intent recognition models are determined based on the sum of the first intermediate value, the second intermediate value, and the third intermediate value. The priority scores of the multiple intent recognition models are then sorted to obtain the task priority of the multiple intent recognition models.
[0117] The complaint urgency preset score can be a score preset according to historical complaint handling rules and business strategies for different complaint levels (e.g., general, important, urgent); the energy saving benefit preset score can be a benefit score preset according to the energy saving strategy's energy consumption reduction ratio or cost saving effect; the network impact preset score can be a network impact score preset according to the degree of impact on different business indicators (e.g., coverage, throughput, call drop rate, etc.).
[0118] The first preset weight, the second preset weight, and the third preset weight can be weights set by those skilled in the art as needed.
[0119] In this embodiment, priority scores for each intent recognition model are obtained based on a preset score for complaint urgency and a first preset weight, a preset score for energy-saving benefits and a second preset weight, and a preset score for network impact and a third preset weight. These priority scores are then sorted to determine the task priorities of the multiple intent recognition models. This method quantifies and configures the importance of tasks when executing tasks based on intent recognition models, enabling tasks with high complaint urgency, high energy-saving benefits, or high network impact to receive more reasonable priority in the overall scheduling, thereby further improving the collaborative capabilities among multiple tasks.
[0120] In some implementations, a scoring function can be used to determine the priority score of the intent recognition model, and the scoring function can be represented by the following seventh formula:
[0121]
[0122] In the seventh formula, It can be used to represent intent recognition models Priority score; It can be used to represent the first preset weight; It can be used to represent a second preset weight; It can be used to represent the third preset weight.
[0123] In some factual ways, the method further includes:
[0124] A digital twin simulation platform is constructed based on the state space and action space of the multiple intelligent agents;
[0125] Based on a preset reward function, the result of the task execution, and a preset reinforcement learning algorithm, the parameter combination of the multiple agents is federated on the digital twin simulation platform to obtain a globally optimized federated parameter aggregation result. The parameter combination of the multiple agents is then updated according to the federated parameter aggregation result.
[0126] By utilizing the updated multiple intelligent agents, attribution modeling is performed based on the multiple types of business data and preset task triggering conditions to obtain multiple intent recognition models;
[0127] Based on the conflict detection results and the task priority, the task execution parameters of the multiple intent recognition models are adjusted collaboratively by the updated multiple agents to obtain the adjusted multiple intent recognition models.
[0128] The state space of the multiple agents may include the current network state (e.g., load or signal quality), the agent's internal state (e.g., resource utilization or remaining task time), and summaries of other agents' intentions (e.g., summaries of other agents' intentions).
[0129] The action space may include resource request / release, power adjustment, task postponement / termination, or channel switching, etc.
[0130] A digital twin simulation platform can be a simulation system that constructs a virtual network environment and business scenario in a computer environment based on the state space and action space of the multiple intelligent agents. It is used to model and simulate the real network topology, business load distribution, resource constraints, intelligent agent interaction process and task execution, so as to train, verify and optimize multi-agent strategies without affecting the actual network. A digital twin simulation platform can also be called a digital twin simulation model.
[0131] The reward function can be a function that quantifies the contribution of the preset indicators based on the actual task execution results of the intent recognition model.
[0132] The preset reinforcement learning algorithm can be a multi-agent reinforcement learning algorithm under the multi-agent reinforcement learning framework, including but not limited to the multi-agent deep deterministic policy gradient algorithm (MADDPG), multi-agent proximal policy optimization (Multi-agent PPO) algorithm, monotonic value function factorisation (QMIX) algorithm, hierarchical agent-based proximal policy optimization (HAPPO) algorithm, and hierarchical actor-critic algorithm, etc.
[0133] On the digital twin simulation platform, the parameter combinations of the multiple intelligent agents are aggregated using federated parameters to obtain a globally optimized federated parameter aggregation result. The parameter combinations of the multiple intelligent agents are then updated based on the federated parameter aggregation result. This can be achieved by:
[0134] On the digital twin simulation platform, parameter combinations of multiple agents are trained centrally in the digital twin environment, and a federated parameter aggregation mechanism (such as federated averaging algorithm (FedAvg)) is used to optimize the parameter combinations of the distributed multiple agents. The optimized optimal parameter combinations are then distributed or synchronized to the corresponding online agents to obtain the updated multiple agents.
[0135] In this embodiment, based on a preset reward function, task execution results, and a preset reinforcement learning algorithm, a federated collaborative optimization is performed on the parameter combinations of multiple agents on a digital twin simulation platform. This federated collaborative mechanism completes global parameter aggregation and policy updates without sharing original data, ensuring both data privacy and training efficiency, while also guaranteeing that each agent works collaboratively under a consistent global policy. This enables tasks across intents, regions, and businesses to achieve coordinated division of labor and collaborative execution within a unified policy framework, further enhancing the task collaboration capabilities of multiple agents.
[0136] In some implementations, the updated plurality of agents can also be used to perform tasks based on the adjusted plurality of intent recognition models, so as to further improve task collaboration capabilities.
[0137] In some implementations, to further improve the network performance of coordinated scheduling, a digital simulation platform can be built based on real network topology, terminal distribution, base station resource status, etc., and the parameter combination of multiple agents can be optimized by combining the task execution results (such as maximizing coverage, minimizing network energy consumption, minimizing the probability of user complaints, and maximizing the coordination success rate of multiple agents).
[0138] Furthermore, reinforcement learning algorithms or particle swarm optimization (PSO) algorithms can be combined to optimize parameter combinations such as power control granularity, time window allocation, and agent behavior thresholds for multiple agents. The optimized parameter combinations are then distributed to the multiple agents to update them. These updated combinations are then used to generate multiple intent recognition models, and the parameter combinations of these models are collaboratively adjusted. Finally, tasks are performed based on these adjusted intent recognition models, thereby further enhancing the multi-agent task collaboration capability.
[0139] In some embodiments, the method further includes:
[0140] The adjusted multiple intent recognition models are simulated on the digital twin simulation platform to obtain simulation results. The risk of task execution is assessed based on the simulation results, thereby predicting potential risks before task execution and adjusting the parameter combination of multiple agents and / or the task execution parameters of multiple intent recognition models and / or other approval controls to improve the safety and controllability of task execution.
[0141] In some implementations, the preset reward function is determined by weighting and summing the service quality score corresponding to the task execution result, the resource consumption corresponding to the task execution result, and the conflict penalty corresponding to the conflict detection result according to preset weights.
[0142] In this embodiment, by weighting and summing the service quality score corresponding to the task execution result, the resource consumption corresponding to the task execution result, and the conflict penalty corresponding to the conflict detection result according to a preset weight, a preset reward function can be determined. This can guide multiple agents to automatically balance the trade-off between "service quality – resource consumption – conflict risk" during parameter combination update. Under the premise of ensuring overall service quality, resource consumption is reduced and the probability of task conflict is lowered, thereby improving the task collaboration capability of multiple agents.
[0143] Specifically, the preset reward function can be expressed by the following seventh formula:
[0144]
[0145] In the seventh formula, It can be used to represent the first The preset reward function value corresponding to the preset reward function of the result of each task execution; It can be used to represent the first The service quality score corresponds to the result of each task execution. The service quality score can be determined by user satisfaction and task completion rate. It can be used to represent the first Resource consumption corresponding to the result of each task execution; It can be used to represent the first The amount of conflict penalty corresponding to the result of each task execution; , and , respectively , and The preset weighting factors are used to balance the relative importance of service quality, resource consumption, and conflict penalty in the reward function. The specific values can be set as needed by those skilled in the art.
[0146] It should be noted that the task collaboration method described above can be executed by an electronic device, that is, all steps included in the above method are completed by the electronic device, which can be a server, an edge computing node base station control unit, or a terminal with computing and storage capabilities, etc.
[0147] In some implementations, this application also provides a multi-agent federated cooperative architecture, which may include:
[0148] The task planner is used to generate multi-dimensional task plans based on network environment status and system objectives, covering types such as fault recovery, energy saving, network reconstruction, and user complaint response. The task planner can break down the overall task into several sub-tasks and assign them to suitable intelligent agents for execution through an intelligent matching mechanism.
[0149] The agent pool can include multiple types of agents with autonomous decision-making capabilities, each responsible for different types of network tasks. Each agent has a closed-loop capability of perception, analysis, decision-making, and execution, and achieves information sharing through a communication interface. The agent pool includes, but is not limited to, the fault handling agent, network optimization agent, complaint analysis agent, energy consumption optimization agent, and 5G-A precise number allocation agent mentioned above.
[0150] The strategy negotiation module, driven by game theory and reinforcement learning algorithms, is responsible for selecting collaborative strategies among agents. When multiple agents encounter resource conflicts or inconsistent strategies during execution, this module coordinates all parties to reach the optimal collaborative solution. The strategy negotiation module supports two levels of strategy game: intra-task game (such as spectrum resource conflict coordination) and inter-task game (such as the exclusive coordination of complaint priority and energy-saving scheduling resources).
[0151] The conflict detection module is used to monitor the behavior trajectory and resource requests of each agent in real time, and to determine whether there are task conflicts, resource contention or policy contradictions based on the graph model and rule base; if there are task conflicts, resource contention or policy contradictions, the feedback is sent to the policy negotiation module.
[0152] The digital twin simulation engine is used to build a full-element simulation environment. The decision-making results of the agent can be run and verified in advance in the simulation environment, avoiding business risks caused by real network adjustments. Simulation data can also feed back into the agent's strategy optimization, forming a closed-loop learning mechanism.
[0153] This embodiment provides a multi-agent federated collaborative architecture, which can be a method, device, or system. It breaks through the bottleneck of the master-slave control architecture in existing network agent systems, realizes efficient task allocation, autonomous collaboration among agents, and strategy game optimization, and ultimately helps intelligent collaborative management of complex tasks such as network resource allocation, energy efficiency optimization, and fault handling. It is widely applicable to operation and maintenance optimization and resource scheduling in new-generation wireless network environments such as 5G-A and 6th Generation Mobile Communication System (6G).
[0154] Please see Figure 2 , Figure 2 This is a schematic diagram of the structure of a task collaboration device provided in an embodiment of this application, as shown below. Figure 2 As shown, the task coordination device 200 includes:
[0155] The acquisition module 201 is used to acquire multiple types of business data, and use multiple intelligent agents to perform attribution modeling based on the multiple types of business data and preset task triggering conditions to obtain multiple intent recognition models.
[0156] The first processing module 202 is used to perform conflict detection on the task execution parameters of the multiple intent recognition models respectively, and obtain conflict detection results;
[0157] The second processing module 203 is used to determine the task priority of the plurality of intent recognition models;
[0158] The third processing module 204 is used to adjust the task execution parameters of the multiple intent recognition models in collaboration with the multiple intelligent agents according to the conflict detection results and the task priority, so as to obtain the adjusted multiple intent recognition models.
[0159] The execution module 205 is used to perform task execution based on the adjusted multiple intent recognition models.
[0160] Optionally, the task execution parameters of each of the plurality of intent recognition models include task type, target region, task time window, and expected metric change.
[0161] Optionally, the conflict detection performed on the task execution parameters of the plurality of intent recognition models to obtain conflict detection results includes at least one of the following:
[0162] The target regions of the multiple intent recognition models are respectively subjected to region overlap detection to obtain the first conflict detection result;
[0163] Time window overlap detection is performed on the task time windows of the multiple intent recognition models respectively to obtain the second conflict detection result;
[0164] The direction of index change is detected for the expected index changes of the multiple intent recognition models respectively, and the third conflict detection result is obtained.
[0165] Optionally, determining the task priority of the plurality of intent recognition models includes:
[0166] A first intermediate value is determined based on the preset score of the urgency of the complaint from the multiple intent recognition models and the first preset weight;
[0167] The second intermediate value is determined based on the preset score and second preset weight of the energy-saving benefit of the multiple intent recognition models;
[0168] A third intermediate value is determined based on the network influence preset score and the third preset weight of the multiple intent recognition models;
[0169] The priority scores of the multiple intent recognition models are determined based on the sum of the first intermediate value, the second intermediate value, and the third intermediate value. The priority scores of the multiple intent recognition models are then sorted to obtain the task priority of the multiple intent recognition models.
[0170] Optionally, the task coordination device 200 may also include: a fourth processing module 206;
[0171] The fourth processing module 206 is used to construct a digital twin simulation platform based on the state space and action space of the multiple intelligent agents;
[0172] Based on a preset reward function, the result of the task execution, and a preset reinforcement learning algorithm, the parameter combination of the multiple agents is federated on the digital twin simulation platform to obtain a globally optimized federated parameter aggregation result. The parameter combination of the multiple agents is then updated according to the federated parameter aggregation result.
[0173] By utilizing the updated multiple intelligent agents, attribution modeling is performed based on the multiple types of business data and preset task triggering conditions to obtain multiple intent recognition models;
[0174] Based on the conflict detection results and the task priority, the task execution parameters of the multiple intent recognition models are adjusted collaboratively by the updated multiple agents to obtain the adjusted multiple intent recognition models.
[0175] Optionally, the task coordination device 200 may also include: a fifth processing module 207;
[0176] The fifth processing module 207 is used to simulate the adjusted multiple intent recognition models under the digital twin simulation platform, obtain simulation results, and assess the risk of task execution based on the simulation results.
[0177] Optionally, the preset reward function is determined by weighting and summing the service quality score corresponding to the task execution result, the resource consumption corresponding to the task execution result, and the conflict penalty corresponding to the conflict detection result according to preset weights.
[0178] The task coordination device 200 is designed to implement the various processes described above in the various embodiments of the task coordination method. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0179] This application also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described task coordination method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0180] For details, see Figure 3 This application also provides an electronic device, including a bus 301, a transceiver 302, an antenna 303, a bus interface 304, a processor 305, and a memory 306.
[0181] The transceiver 302 is used to acquire multiple types of business data and use multiple intelligent agents to perform attribution modeling based on the multiple types of business data and preset task triggering conditions to obtain multiple intent recognition models.
[0182] The processor 305 is used to perform conflict detection on the task execution parameters of the plurality of intent recognition models respectively, and obtain conflict detection results;
[0183] Determine the task priorities of the multiple intent recognition models;
[0184] Based on the conflict detection results and the task priority, the task execution parameters of the multiple intent recognition models are adjusted collaboratively by the multiple intelligent agents to obtain the adjusted multiple intent recognition models.
[0185] Task execution is performed based on the adjusted multiple intent recognition models.
[0186] Optionally, the task execution parameters of each of the plurality of intent recognition models include task type, target region, task time window, and expected metric change.
[0187] Optionally, the conflict detection performed on the task execution parameters of the plurality of intent recognition models to obtain conflict detection results includes at least one of the following:
[0188] The target regions of the multiple intent recognition models are respectively subjected to region overlap detection to obtain the first conflict detection result;
[0189] Time window overlap detection is performed on the task time windows of the multiple intent recognition models respectively to obtain the second conflict detection result;
[0190] The direction of index change is detected for the expected index changes of the multiple intent recognition models respectively, and the third conflict detection result is obtained.
[0191] Optionally, determining the task priority of the plurality of intent recognition models includes:
[0192] A first intermediate value is determined based on the preset score of the urgency of the complaint from the multiple intent recognition models and the first preset weight;
[0193] The second intermediate value is determined based on the preset score and second preset weight of the energy-saving benefit of the multiple intent recognition models;
[0194] A third intermediate value is determined based on the network influence preset score and the third preset weight of the multiple intent recognition models;
[0195] The priority scores of the multiple intent recognition models are determined based on the sum of the first intermediate value, the second intermediate value, and the third intermediate value. The priority scores of the multiple intent recognition models are then sorted to obtain the task priority of the multiple intent recognition models.
[0196] Optionally, the processor 305 is also used to construct a digital twin simulation platform based on the state space and action space of the plurality of intelligent agents;
[0197] Based on a preset reward function, the result of the task execution, and a preset reinforcement learning algorithm, the parameter combination of the multiple agents is federated on the digital twin simulation platform to obtain a globally optimized federated parameter aggregation result. The parameter combination of the multiple agents is then updated according to the federated parameter aggregation result.
[0198] By utilizing the updated multiple intelligent agents, attribution modeling is performed based on the multiple types of business data and preset task triggering conditions to obtain multiple intent recognition models;
[0199] Based on the conflict detection results and the task priority, the task execution parameters of the multiple intent recognition models are adjusted collaboratively by the updated multiple agents to obtain the adjusted multiple intent recognition models.
[0200] Optionally, the processor 305 is further configured to simulate the adjusted multiple intent recognition models under the digital twin simulation platform, obtain simulation results, and assess the risk of task execution based on the simulation results.
[0201] Optionally, the preset reward function is determined by weighting and summing the service quality score corresponding to the task execution result, the resource consumption corresponding to the task execution result, and the conflict penalty corresponding to the conflict detection result according to preset weights.
[0202] exist Figure 3 In this context, a bus architecture (represented by bus 301) is used. Bus 301 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 305 and memory represented by memory 306. Bus 301 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 304 provides an interface between bus 301 and transceiver 302. Transceiver 302 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 305 is transmitted over a wireless medium via antenna 303, which further receives data and transmits it to processor 305.
[0203] Processor 305 manages bus 301 and general processing, and also provides various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 306 can be used to store data used by processor 305 during operation.
[0204] Optionally, the processor 305 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD).
[0205] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described task coordination method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0206] This application also provides a computer program product, including computer instructions. When these computer instructions are executed by a processor, they implement the various processes of the above-described task coordination method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0207] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0208] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0209] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A task collaboration method, characterized in that, The method includes: Multiple types of business data are acquired, and multiple intelligent agents are used to perform attribution modeling based on the multiple types of business data and preset task triggering conditions to obtain multiple intent recognition models. Conflict detection is performed on the task execution parameters of the multiple intent recognition models to obtain conflict detection results; Determine the task priorities of the multiple intent recognition models; Based on the conflict detection results and the task priority, the task execution parameters of the multiple intent recognition models are adjusted collaboratively by the multiple intelligent agents to obtain the adjusted multiple intent recognition models. Task execution is performed based on the adjusted multiple intent recognition models.
2. The method according to claim 1, characterized in that, The task execution parameters for each of the multiple intent recognition models include task type, target region, task time window, and expected metric change.
3. The method according to claim 2, characterized in that, The conflict detection process, which involves performing conflict detection on the task execution parameters of the multiple intent recognition models to obtain conflict detection results, includes at least one of the following: The target regions of the multiple intent recognition models are respectively subjected to region overlap detection to obtain the first conflict detection result; Time window overlap detection is performed on the task time windows of the multiple intent recognition models respectively to obtain the second conflict detection result; The direction of index change is detected for the expected index changes of the multiple intent recognition models respectively, and the third conflict detection result is obtained.
4. The method according to claim 1, characterized in that, Determining the task priority of the plurality of intent recognition models includes: A first intermediate value is determined based on the preset score of the urgency of the complaint from the multiple intent recognition models and the first preset weight; The second intermediate value is determined based on the preset score and second preset weight of the energy-saving benefit of the multiple intent recognition models; A third intermediate value is determined based on the network influence preset score and the third preset weight of the multiple intent recognition models; The priority scores of the multiple intent recognition models are determined based on the sum of the first intermediate value, the second intermediate value, and the third intermediate value. The priority scores of the multiple intent recognition models are then sorted to obtain the task priority of the multiple intent recognition models.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: A digital twin simulation platform is constructed based on the state space and action space of the multiple intelligent agents; Based on a preset reward function, the result of the task execution, and a preset reinforcement learning algorithm, the parameter combination of the multiple agents is federated on the digital twin simulation platform to obtain a globally optimized federated parameter aggregation result. The parameter combination of the multiple agents is then updated according to the federated parameter aggregation result. By utilizing the updated multiple intelligent agents, attribution modeling is performed based on the multiple types of business data and preset task triggering conditions to obtain multiple intent recognition models; Based on the conflict detection results and the task priority, the task execution parameters of the multiple intent recognition models are adjusted collaboratively by the updated multiple agents to obtain the adjusted multiple intent recognition models.
6. The method according to claim 5, characterized in that, The method further includes: The adjusted multiple intent recognition models are simulated on the digital twin simulation platform to obtain simulation results, and the risk of task execution is assessed based on the simulation results.
7. The method according to claim 5, characterized in that, The preset reward function is determined by weighting and summing the service quality score corresponding to the task execution result, the resource consumption corresponding to the task execution result, and the conflict penalty corresponding to the conflict detection result according to preset weights.
8. A task collaboration device, characterized in that, The device includes: The acquisition module is used to acquire multiple types of business data, and use multiple intelligent agents to perform attribution modeling based on the multiple types of business data and preset task triggering conditions to obtain multiple intent recognition models. The first processing module is used to perform conflict detection on the task execution parameters of the multiple intent recognition models respectively, and obtain conflict detection results; The second processing module is used to determine the task priority of the plurality of intent recognition models; The third processing module is used to adjust the task execution parameters of the multiple intent recognition models in collaboration with the multiple intelligent agents based on the conflict detection results and the task priority, so as to obtain the adjusted multiple intent recognition models. The execution module is used to perform task execution based on the adjusted multiple intent recognition models.
9. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 7.
11. A computer program product, characterized in that, Includes computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 7.