An autonomous scheduling and collaborative decision-making method based on an agent cluster

CN122549480APending Publication Date: 2026-08-11西安圣瞳科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]现有技术模式下,态势感知数据缺乏统一的结构化梳理方式,任务要素与环境状态无法形成规整的感知呈现形式,感知信息易出现杂乱无序的状态

Benefits of technology

对初始态势感知集合进行结构化解析,生成包含任务要素分解图和动态环境状态向量的集群感知模型,态势感知信息可形成标准化的结构化呈现形态,任务要素与环境状态的对应关系可清晰展现,感知数据的规整程度得到优化,智能体集群内部感知信息的传递与识别更贴合集群自主调度的运行状态,感知信息碎片化的状态得到改善,集群感知环节的信息处理逻辑与集群调度的运行逻辑相契合。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122549480A_ABST
    Figure CN122549480A_ABST
Patent Text Reader

Abstract

This invention discloses a method for autonomous scheduling and collaborative decision-making based on intelligent agent clusters, belonging to the field of intelligent agent cluster scheduling technology. The method includes acquiring task announcements and environmental situation maps of the intelligent agent cluster, constructing an initial situation awareness set, and generating a cluster awareness model containing a task element decomposition diagram and a dynamic environmental state vector after structured parsing of the set. Based on this model, an improved consensus initiative algorithm is used to calculate individual preference vectors that integrate task inclination and capability matching. These vectors are aggregated to form a cluster-level preference distribution map. Through multiple rounds of iterative distributed negotiation, task allocation schemes and action coordination plans are generated, ultimately converted into a set of individual action instructions that can be parsed by the intelligent agents to drive cluster operation. This method can standardize the form of situation awareness data, match the capabilities of the intelligent agents with task execution requirements, adapt to dynamic environmental changes, and optimize the connection between intelligent agent cluster scheduling and collaborative actions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent agent cluster scheduling technology, specifically a method based on intelligent agent cluster autonomous scheduling and collaborative decision-making. Background Technology

[0002] Existing technologies related to agent cluster scheduling and collaborative decision-making often directly collect task announcements and environmental situation maps, resulting in fragmented initial perception data without structuring it. Some technologies employ conventional consensus-based proactive algorithms for agent task-related computations, generating task-related parameters from a single dimension without incorporating task inclination and capability matching to form individual preference data. Cluster decision-making often uses centralized allocation or simple one-time negotiation to directly generate task allocation results, without building a cluster-level distributed reference based on preference data before further negotiation.

[0003] Under current technological models, situational awareness data lacks a unified, structured approach, making it difficult to present task elements and environmental states in a regular, orderly manner, resulting in chaotic and disordered information. Conventional algorithms generate individual parameters with limited dimensions, failing to match the actual conditions under which agents execute tasks, and lack a unified reference for preference distribution at the cluster level. The negotiation process lacks a multi-round, iterative, distributed operating mode, making it difficult to adapt the generated task allocation and action coordination schemes to dynamic cluster changes, and hindering the formation of a proper relationship between individual agent commands and overall cluster actions.

[0004] The initial situational awareness set needs to be structurally analyzed to form a clustered awareness model containing task element decomposition diagrams and dynamic environmental state vectors, thus improving the unstructured state of the perceived information. An improved consensus initiative algorithm is needed to calculate and generate individual preference vectors that fuse task inclination and capability matching. These vectors are then aggregated to construct a cluster-level preference distribution map, which initiates a multi-round iterative distributed negotiation process to generate suitable task allocation schemes and action coordination plans. The decision schemes are then converted into a set of individual action instructions that can be parsed by the agents, optimizing the adaptation state of cluster scheduling and decision-making. Summary of the Invention

[0005] This invention aims to solve at least one of the technical problems existing in the prior art; To this end, the present invention proposes a method for autonomous scheduling and collaborative decision-making based on intelligent agent clusters, comprising: Acquire task announcements and environmental situation maps of the intelligent agent cluster to form an initial situational awareness set; The initial situational awareness set is structured and analyzed to generate a clustered awareness model containing a task element decomposition diagram and a dynamic environment state vector. Based on the aforementioned cluster perception model, an improved consensus initiative algorithm is used to calculate and generate an individual preference vector for each agent, which includes task inclination and ability matching degree. Aggregate the individual preference vectors to construct a cluster-level preference distribution map; Based on the cluster-level preference distribution map, a multi-round iterative distributed negotiation process is initiated to generate a cluster decision scheme that includes task allocation schemes and action coordination plans. The cluster decision scheme is converted into a set of individual action instructions that can be parsed by each agent in the agent cluster, and the agent cluster is driven to execute it. The working principle of the improved consensus initiative algorithm includes: During the algorithm initialization phase, dynamic attention weights are injected into each agent. These dynamic attention weights are adaptively adjusted based on the preference cues received by the current agent from neighboring agents. The update method is that the agent counts the number of preference cues and effective interactions from neighboring agents for each atomic task, and updates the weights in real time with the accumulated interaction value. In the individual preference calculation stage, the algorithm not only considers the static matching between the atomic task and the agent's own capabilities and motivations, but also introduces a consensus prediction subprocess. The consensus prediction subprocess requires that when an agent calculates its own preference for a certain atomic task, it must estimate the collective preference tendency of its neighboring agents for the atomic task, and feed back the estimated collective preference tendency in a weighted form into its own ability matching degree and task preference degree calculation. After the adjustment of the dynamic attention weights and the feedback correction of the consensus prediction subprocess, the individual preference vector output by each agent has implicitly predicted the potential consensus of the cluster and actively moved towards it.

[0006] Furthermore, the initial situational awareness set is subjected to structured analysis to generate a clustered awareness model containing a task element decomposition diagram and a dynamic environment state vector, including: Extract the task objective description, task constraints, and task completion indicators from the task announcement, and decompose the task objective description into multiple atomic tasks that can be executed independently or sequentially to form the task element decomposition diagram; Extract agent positions, obstacle outlines, and dynamic event markers from the environmental situation map, and quantify the agent positions, obstacle outlines, and dynamic event markers into time-stamped numerical state variables; The numerical state variables are aligned to a time series and spatially gridded to generate a dynamic environmental state vector that reflects the spatiotemporal changes of environmental elements. The task element decomposition diagram and the dynamic environment state vector are associated and combined to encapsulate a unified cluster perception model.

[0007] Furthermore, based on the aforementioned cluster perception model, and utilizing an improved consensus initiative algorithm, an individual preference vector containing task inclination and capability matching degree is calculated and generated for each agent, including: From the cluster perception model, extract the local task view and local environment view related to the current agent; Guided by the improved consensus initiative algorithm, the atomic tasks in the local task view are compared with the internal capability model of the current agent to calculate the capability matching degree. Meanwhile, guided by the improved consensus initiative algorithm, the conditions in the local environment view that are favorable or unfavorable to completing the atomic task are coupled with the current agent's internal motivation model to calculate the task inclination. The capability matching degree and task preference degree of each atomic task corresponding to the current agent are combined into a numerical pair, and the numerical pairs of all atomic tasks constitute the individual preference vector.

[0008] Further, the individual preference vectors are aggregated to construct a cluster-level preference distribution map, including: Collect the individual preference vectors of all agents in the agent cluster to form a high-dimensional cluster preference matrix; Perform dimensionality reduction visualization calculation on the cluster preference matrix to map the preference values ​​of each agent for each atomic task into a two-dimensional preference space, where one dimension represents the ability matching degree and the other dimension represents the task preference degree. In the preference space, connection edges between agent nodes are constructed based on the physical proximity of agents or the logical association of tasks; The mapped agent nodes and their connecting edges, along with the preference values ​​carried on the nodes, are rendered together to form a visual graph, namely the cluster-level preference distribution map.

[0009] Furthermore, based on the cluster-level preference distribution map, a multi-round iterative distributed negotiation process is initiated to generate a cluster decision-making scheme that includes a task allocation scheme and an action coordination plan, including: The cluster-level preference distribution map is broadcast to every agent in the agent cluster as a common basis for negotiation; Each agent calculates its expected reward for each atomic task based on the cluster-level preference distribution graph and sends a proposal message containing its expected reward to its neighboring agents. Each agent receives proposal messages from neighboring agents, adjusts its expected revenue based on a pre-defined negotiation strategy, and generates a bid or commitment for the atomic task. After a preset number of iterations, when all agents' bids or commitments for the atomic task no longer change, or the changes are below a threshold, the negotiation process terminates, and the final commitments of all agents are aggregated to form a preliminary task allocation map. The initial task allocation mapping is subjected to conflict detection and resolution to ensure that each atomic task is committed to execution by one and only one agent, and that the task committed by each agent does not exceed its capacity, thus forming a determined task allocation scheme. Based on the task allocation scheme and the dynamic environment state vector, an execution time window and spatial path are planned for each allocated atomic task, forming an overall execution sequence and spatial coordination relationship among all atomic tasks, i.e., the action coordination plan.

[0010] Furthermore, conflict detection and resolution are performed on the preliminary task allocation mapping, including: Scan the preliminary task allocation map to identify atomic tasks that multiple agents have committed to execute, and mark them as conflicting tasks; For each conflicting task, obtain the final contribution value of all agents who committed to performing the conflicting task during the negotiation process; Compare the final output value, assign the conflicting task to the agent with the highest final output value, and send task deprivation notices to other agents who have committed to the conflicting task; Upon receiving a task preemption notification, the agent removes the conflicting task from its list of committed tasks and checks whether its current task load is below a minimum threshold. If it is below the minimum threshold, the agent needs to re-enter the local negotiation process for the remaining unassigned atomic tasks. Conflict detection is repeated until there are no conflicting tasks in the initial task allocation map and the task load of all agents is within an acceptable range, thus completing conflict resolution.

[0011] Furthermore, based on the task allocation scheme and the dynamic environment state vector, an execution time window and spatial path are planned for each allocated atomic task, including: Environmental prediction information for a future period is extracted from the dynamic environment state vector, and the environmental prediction information includes obstacle movement trajectory and dynamic event occurrence probability. Based on the task allocation scheme, the executing agent for each atomic task and the logical sequence relationship between the atomic tasks are determined. With the dual objectives of minimizing the overall task completion time and avoiding spatiotemporal conflicts between agents, a start and end time range, i.e., an execution time window, is allocated to each atomic task in the time dimension. Within the execution time window of each atomic task, combined with the environmental prediction information, a collision-free path, i.e. a spatial path, is planned for its executing agent from the starting position to the task target position, and then to the next task point or waiting area. By integrating the execution time windows and spatial paths of all atomic tasks, potential conflicts arising from different agents occupying the same spatial location at the same time are examined and resolved, forming a globally coordinated action coordination plan.

[0012] Furthermore, the integration of execution time windows and spatial paths for all atomic tasks, and the checking and resolution of potential conflicts where different agents occupy the same spatial location at the same time, include: A four-dimensional spatiotemporal grid model is established, which includes a three-dimensional spatial coordinate dimension and a one-dimensional time slice dimension; The spatial path of each agent is mapped to the four-dimensional spatiotemporal grid model according to its execution time window, and the spatial grid occupied by the agent in a specific time slice is marked. Traverse the four-dimensional spatiotemporal grid model to find spatial grids that are occupied by multiple agents in the same time slice and record them as spatiotemporal conflict points. For each spatiotemporal conflict point, the spatial path or execution time window of the relevant intelligent agent is adjusted according to the preset conflict resolution rules, which include the spatial avoidance principle and the time staggering principle. After adjusting all spatiotemporal conflict points, the adjusted information of all agents is remapped into the four-dimensional spatiotemporal grid model for a new round of conflict checks until there are no more spatiotemporal conflict points in the model, thus completing the conflict resolution.

[0013] Further, the cluster decision-making scheme is converted into a set of individual action instructions that can be parsed by each agent in the agent cluster, including: Analyze the task allocation scheme in the cluster decision scheme, and extract all atomic tasks assigned to the current agent and their execution order; The action coordination plan in the cluster decision-making scheme is analyzed, and the detailed parameters of the execution time window and spatial path of each atomic task corresponding to the current agent are extracted. The description of each atomic task, its corresponding execution time window, and spatial path parameters are encoded into a series of action instructions with clear semantics and sequence according to the agent's internal instruction protocol. This series of action instructions is packaged and timestamped according to the task execution order to form the individual action instruction set that belongs to the current intelligent agent and contains a complete action sequence and time arrangement.

[0014] Compared with the prior art, the beneficial effects of the present invention are: The initial situational awareness set is structured and analyzed to generate a clustered awareness model containing task element decomposition diagrams and dynamic environmental state vectors. Situational awareness information can be presented in a standardized and structured form, the correspondence between task elements and environmental states can be clearly shown, the regularity of the perception data is optimized, the transmission and recognition of perception information within the intelligent agent cluster are more in line with the autonomous scheduling operation state of the cluster, the fragmentation of perception information is improved, and the information processing logic of the clustered awareness link is consistent with the operation logic of cluster scheduling.

[0015] Based on the cluster perception model, an improved consensus initiative algorithm is used to calculate and generate individual preference vectors containing task inclination and capability matching. Individual preference vectors are aggregated to construct a cluster-level preference distribution map. Based on the cluster-level preference distribution map, a multi-round iterative distributed negotiation process is initiated to generate a cluster decision scheme containing task allocation schemes and action coordination plans. The cluster decision scheme is converted into a set of individual action instructions that can be parsed by the agents. The task preferences of individual agents can be correlated with their own execution capabilities. The preference distribution at the cluster level can be intuitively presented. The iterative process of distributed negotiation can adapt to the dynamic changes of the cluster. The fit between task allocation and action coordination is optimized. Individual action instructions are matched with the execution logic of agents. The connection state of cluster collaborative actions is more in line with the overall operation requirements of the cluster. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the steps of an intelligent agent cluster-based autonomous scheduling and collaborative decision-making method as described in this invention. Figure 2 A flowchart for generating a cluster-aware model; Figure 3 A flowchart for calculating and generating individual preference vectors. Detailed Implementation

[0017] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] See Figure 1 This invention provides a method for autonomous scheduling and collaborative decision-making based on intelligent agent clusters, the specific method including: The initial situational awareness set is formed by acquiring task announcements from external sources and environmental situation maps perceived through sensor networks. This initial situational awareness set is then structurally analyzed, extracting structured task elements from the task announcements and quantifying dynamic environmental states from the environmental situation maps. These two sets are then combined and encapsulated to form a unified cluster perception model. Based on this cluster perception model, an improved consensus initiative algorithm is used to generate an individual preference vector for each agent in the cluster. This vector contains the agent's ability matching degree and task inclination for each atomic task. The individual preference vectors of all agents are aggregated, and a cluster-level preference distribution map that intuitively reflects the overall preference tendency of the cluster is constructed through dimensionality reduction and visualization techniques. Using this distribution map as a common knowledge base, a multi-round iterative distributed negotiation process is initiated among the agents. Through exchanging proposals and adjusting bids, a cluster decision scheme is finally converged, explicitly including a task allocation scheme and an action coordination plan. This cluster-level decision scheme is parsed and converted into a set of individual action instructions that each agent can independently understand and execute, thereby driving the entire agent cluster to execute tasks in an orderly and collaborative manner.

[0019] In one embodiment of the present invention, this embodiment details how to perform structured analysis on an initial situational awareness set to generate a clustered awareness model. The specific process is as follows: [Refer to...] Figure 2 From the acquired task announcements, three core elements—task objective description, task constraints, and task completion indicators—are extracted using natural language processing or structured parsing techniques. The complex task objective description is then logically broken down into multiple atomic tasks that can be executed independently or sequentially. The temporal, causal, or spatial relationships between these atomic tasks are clearly marked, thus constructing a clearly structured task element decomposition diagram. Simultaneously, key environmental elements are extracted from the environmental situation map, including the real-time position coordinates of each agent, the outlines of static and dynamic obstacles, and the location and type markers of dynamic events. These elements are further quantified into numerical state variables with precise timestamps. To integrate spatiotemporal information, the aforementioned numerical state variables undergo time-series alignment to ensure synchronization across the timeline, and spatial gridding is performed, discretizing the continuous space into unified grid units. Based on this, a dynamic environmental state vector that comprehensively reflects the spatiotemporal evolution of environmental elements is generated. Finally, the task element decomposition diagram is associated and combined with the dynamic environment state vector, and a mapping relationship is established between the task elements and the relevant environmental state regions. This combination is then encapsulated into a cluster perception model with a unified format and structured data for subsequent process calls.

[0020] In practical implementation, the initial situational awareness set is structured and parsed to generate a cluster awareness model. The initial situational awareness set consists of task announcements and an environmental situation map. The task announcements contain task information described in natural language or a structured format, while the environmental situation map contains environmental information collected in real time through a sensor network. Three core elements are extracted from the task announcements: task objective description, task constraints, and task completion indicators. The task objective description defines the final goal the cluster needs to achieve; the task constraints limit resources, time, or rules during task execution; and the task completion indicators provide specific standards for measuring task success. In practical implementation, natural language processing technology or a predefined structured parsing template is used to decompose the task objective description, breaking it down into multiple atomic tasks that can be executed independently or sequentially. The temporal dependencies, causal relationships, or spatial associations between atomic tasks are clearly marked and recorded, forming a clearly structured task element decomposition graph where nodes represent atomic tasks and edges represent relationships between tasks. Key environmental elements are extracted from the environmental situation map. These environmental elements include the real-time position coordinates of all agents in the agent cluster, the outline boundary information of static and dynamic obstacles in the environment, and the position coordinates and event type labels of dynamic events. These environmental elements are further quantified into numerical state variables with precise timestamps. Each numerical state variable corresponds to the digital representation of an environmental element at a specific moment.

[0021] In some embodiments, the numerical state variables undergo time-series alignment processing. This alignment ensures that all numerical state variables from different sensors or different acquisition times are synchronized on a unified time reference axis, making environmental states comparable. The time-series aligned numerical state variables are then spatially gridded, discretizing the continuous geographic or operational space into a series of uniformly sized grid cells. Each numerical state variable is mapped to a corresponding spatial grid cell and appended with a timestamp. In a specific implementation, a dynamic environmental state vector is generated based on the set of time-series aligned and spatially gridded numerical state variables. This dynamic environmental state vector is a multi-dimensional data structure, where each dimension represents the evolution sequence of a specific environmental feature on the spatiotemporal grid. The dynamic environmental state vector comprehensively reflects the changing patterns of environmental elements over time and space. It is understandable that a relationship needs to be established between the task element decomposition diagram and the dynamic environment state vector. This relationship is achieved by mapping each atomic task in the task element decomposition diagram to the relevant environmental region or time segment in the dynamic environment state vector. For example, an atomic task involving moving to a certain location will be associated with a subset of dimensions in the dynamic environment state vector that describe the distribution of obstacles and passage conditions around that location.

[0022] Optionally, the associative combination process is completed through a wrapper function. This wrapper function packages the task element decomposition graph, the dynamic environment state vector, and the mapping table between them into a unified data object. This unified data object is the cluster perception model, which provides a standardized data interface for subsequent processes to access and call. In some embodiments, the cluster perception model is internally organized using a hierarchical or graph structure. The hierarchical structure uses the task element decomposition graph as the logical layer and the dynamic environment state vector as the physical layer, with the two layers linked by pointers or indexes. The graph structure integrates atomic task nodes, environment state nodes, and the edges between them into a unified graph. It can be understood that the process of generating the cluster perception model is fully automated. This automation is executed by a dedicated parsing engine. The parsing engine reads the initial situational awareness set, processes and combines it according to the above steps, and finally generates the cluster perception model without manual intervention.

[0023] In one embodiment of the present invention, this embodiment details how to calculate and generate an individual preference vector for each agent based on a cluster-aware model and using an improved consensus initiative algorithm. For any agent in the cluster, see [reference needed]. Figure 3The algorithm extracts relevant local information from the cluster perception model, including a subset of atomic tasks within its executable or perceptible range forming a local task view, and the environmental state within its sensor detection range forming a local environment view. An improved consensus initiative algorithm plays a role in this process. During algorithm initialization, dynamic attention weights are injected into each agent. These weights are not fixed but adaptively adjusted based on the strength of preference cues received by the agent from neighboring agents for various tasks during information interaction. In each round of information interaction, the agent continuously counts the number of preference cues sent by neighboring agents for each atomic task and the number of effective interactions. The number of preference cues and the number of effective interactions corresponding to the same atomic task are accumulated to obtain the cumulative interaction value for that task. Based on the cumulative interaction value, the dynamic attention weight is updated synchronously. The higher the cumulative interaction value, the greater the dynamic attention weight; the lower the cumulative interaction value, the smaller the dynamic attention weight. The update process is executed in real time with neighbor interactions; a dynamic attention weight update is completed after each round of neighbor information interaction, without the need for a pre-set fixed update cycle. In the individual preference calculation phase, the algorithm comprises two parallel computational processes. First, it compares each atomic task in the local task view with the agent's internal capability model to calculate a capability matching degree, representing the agent's competence in performing the task. Second, it couples favorable or unfavorable conditional factors in the local environment view for completing the atomic task with the agent's internal motivation model to calculate a task inclination degree, representing the agent's willingness to perform the task. This phase also introduces a consensus prediction sub-process. This sub-process requires the agent, when calculating its own preference for a given atomic task, to predict the potential collective preference of its neighboring agents for that task. This prediction result is then fed back in a weighted form to correct the agent's original capability matching degree and task inclination degree. The agent first sends an atomic task preference query command to all neighboring agents within its communication range. Upon receiving the command, each neighboring agent provides feedback on its preliminary capability matching degree and task inclination calculation results for that atomic task. The current agent collects all feedback results, removes abnormal feedback values ​​that deviate significantly from the overall data, and calculates the arithmetic mean of the remaining valid feedback data to obtain the collective preference tendency of the neighboring agents for that atomic task. This collective preference tendency is then fused with the agent's own preliminary preference calculation results to obtain the final corrected preference result used for individual preference vector calculation, completing consensus prediction and preference calibration. After adjustment of dynamic attention weights and feedback correction in the consensus prediction subprocess, each agent ultimately outputs a corrected capability matching degree and task inclination value pair for each atomic task. The numerical pairs of all atomic tasks, arranged in order, constitute the agent's individual preference vector. This vector implicitly contains a prediction of potential consensus within the cluster and a proactive tendency to converge.

[0024] In specific implementations, an individual preference vector is generated for each agent based on a swarm perception model. The swarm perception model includes a structured task element decomposition graph and a dynamic environment state vector. When calculating its own individual preference vector, the agent first extracts relevant local information from the swarm perception model. Specifically, this local information includes a local task view and a local environment view. The local task view consists of a subset of all atomic tasks within the current agent's executable range in the task element decomposition graph. The local environment view consists of a subset of environmental state data corresponding to the current agent's sensor detection range in the dynamic environment state vector. In some embodiments, an improved consensus initiative algorithm is invoked to guide the computation process. During the initialization phase, the improved consensus initiative algorithm injects dynamic attention weights into each participating agent. The dynamic attention weight is a scalar parameter that changes over time or with the interaction state. The value of the dynamic attention weight is adaptively adjusted based on the strength of preference cues received by the current agent from neighboring agents regarding various atomic tasks in the previous round or initial phase. The preference cues are quantified through short messages containing preference tendencies exchanged between neighboring agents. The improved consensus initiative algorithm's individual preference calculation phase comprises two parallel computational sub-processes and one consensus prediction sub-process. The first parallel computational sub-process compares each atomic task in the local task view with the agent's internal capability model. The internal capability model is a pre-defined set of parameters describing the agent's functions and performance. The comparison process calculates a capability matching degree value between 0 and 1. The second parallel computational sub-process couples favorable or unfavorable conditional factors in the local environment view to the agent's internal motivation model. The internal motivation model defines the agent's response preferences to different environmental conditions and task types. The coupling process calculates a task inclination value reflecting the agent's subjective intention.

[0025] Optionally, the preliminary calculation of the ability matching degree and task inclination degree can refer to a basic calculation relationship, which can be expressed as follows:

[0026] Where: characters Represents the median value of agent i's preference for atomic task j; character Represents the internal capability model of agent i; character This represents the task requirement description for atomic task j; characters This represents a function for calculating capability matching; characters Represents the local environment view perceived by agent i; character This represents a function for calculating task preference; characters This represents a balance coefficient between 0 and 1, used to adjust the weights of ability and motivation. The consensus prediction subprocess is the core component of the improved consensus initiative algorithm. It requires agents to estimate the potential collective preference tendencies of all their neighboring agents for a given atomic task before calculating their own final preference. In some embodiments, the estimation of collective preference tendencies is achieved through the exchange of prediction requests and responses between agents and their neighbors. The current agent sends a preference query to its neighbors regarding a specific atomic task, and each neighbor responds with a preference feedback value based on its initial calculation. The current agent collects all responses and calculates their weighted average as the estimated collective preference tendency. In specific implementations, after dynamic attention weight adjustment and feedback correction by the consensus prediction subprocess, each agent generates a corrected ability matching degree and task preference value pair for each atomic task. The correction process involves weightedly adding the estimated collective preference tendency to the initially calculated ability matching degree and task preference value. It can be understood that the dynamic attention weight here adjusts the weight of the collective preference feedback value. The larger the dynamic attention weight, the more the current agent tends to align its own preferences with the perceived cluster consensus. Ultimately, the capability matching degree value and task preference degree value of each atomic task corresponding to the current agent are combined into a numerical pair. All the numerical pairs of atomic tasks are arranged in the order of the atomic tasks' numbers in the task element decomposition diagram, forming an individual preference vector that implicitly contains the cluster consensus prediction.

[0027] In one embodiment of the present invention, individual preference vectors calculated by all agents in an agent cluster are aggregated to construct a cluster-level preference distribution map. The agent cluster contains multiple agents, and the individual preference vector of each agent is an ordered list. Each element in the individual preference vector is a pair of values ​​containing a capability matching degree value and a task inclination degree value. In specific implementations, the collection process is completed through a central node broadcast request or active reporting by agents. All collected individual preference vectors are organized according to agent identifiers and atomic task identifiers, and arranged to form a high-dimensional cluster preference matrix. The row dimension of the cluster preference matrix corresponds to agents, and the column dimension corresponds to atomic tasks. Each element in the matrix is ​​a two-dimensional pair of values. In some embodiments, the mathematical representation of the cluster preference matrix is ​​as follows: Let the agent cluster have... There are [number] intelligent agents, totaling [number] agents. For each atomic task, the cluster preference matrix is... It is A matrix, where elements :

[0028] character: Indicates the first The first agent on the first The capability matching degree of each atomic task, character Indicates the first The first agent on the first The task preference of each atomic task. This can be understood as the dimension of the cluster preference matrix. and The value may be large. In order to intuitively observe the overall preference structure of the cluster and the relationship between agents, it is necessary to perform dimensionality reduction visualization calculation on the cluster preference matrix. Dimensionality reduction visualization calculation projects high-dimensional numerical data onto a two-dimensional plane so that humans can observe or machines can quickly identify patterns.

[0029] In some embodiments, after generating a two-dimensional scatter plot, it is necessary to construct connecting edges based on the physical proximity relationships or task logical associations between agents. Physical proximity relationships are calculated using Euclidean distance based on the real-time position coordinates of the agents; when the distance is less than a set threshold, connecting edges are established between the corresponding agent nodes. Task logical associations are based on the dependencies between atomic tasks in the task element decomposition graph; if two agents are assigned atomic tasks with strong dependencies, connecting edges are established between the corresponding agent nodes. Optionally, the node and edge information can be stored as a graph data structure. The graph data structure includes a node list and an edge list. Each node object in the node list contains its corresponding agent identifier, coordinates in the two-dimensional preference space, and a set of preference value pairs for all atomic tasks. Each edge object in the edge list contains the identifiers of the source and target nodes it connects, as well as the edge type. In practical implementation, the mapped agent nodes, the connecting edges between nodes, and the preference values ​​carried on the nodes are rendered together to form a visual graph. The rendering process uses a visualization library to draw nodes as labeled graphical elements, edges as connecting lines, and the agent's preference values ​​for different atomic tasks are marked around the nodes using heatmaps, labels, or radial lines. The final visual graph is the cluster-level preference distribution map. It can be understood that the cluster-level preference distribution map is the common foundation for subsequent distributed negotiation processes. The cluster-level preference distribution map is broadcast to each agent in the cluster in the form of an image or structured graph data. Each agent can parse the distribution map to understand the preference tendencies of other agents and the overall structure of the cluster. For a detailed explanation of the structure of individual preference vectors, refer to Table 1, which shows a simplified example involving three agents and two atomic tasks.

[0030] Table 1: Aggregated Table of Individual Preference Vectors

[0031] In practice, after the above table data is processed by dimensionality reduction visualization calculation, the two-dimensional numerical pairs in each cell will be mapped to a point on a two-dimensional plane. The points corresponding to different atomic tasks in the same row are arranged around the central node representing the agent, thus forming a visualized preference pattern in the cluster-level preference distribution map.

[0032] In one embodiment of the present invention, a multi-round iterative distributed negotiation process is initiated based on a cluster-level preference distribution map to generate a cluster decision scheme. The starting point of the distributed negotiation process is to broadcast the cluster-level preference distribution map to every agent in the agent cluster. The cluster-level preference distribution map serves as a common knowledge base for all agents to negotiate. In specific implementations, each agent calculates its expected reward for each atomic task based on the received cluster-level preference distribution map and its own internal state. The calculation of the expected reward integrates the agent's ability matching degree to the task, task inclination, and the competitive situation observed from the cluster-level preference distribution map. In some embodiments, the formula for calculating the expected reward is:

[0033] Where: characters Represents intelligent agents For atomic tasks Expected return; character Represents intelligent agents For atomic tasks Capability matching degree; characters Represents intelligent agents For atomic tasks Task preference; character This represents the preference distribution for atomic tasks derived from the cluster-level preference distribution diagram. The estimated value of competition intensity; character This represents the weighting coefficients for each item. Estimated competition intensity. The calculation method is as follows: Based on the cluster-level preference distribution map, the number of agents submitting bids or commitments for atomic task m is counted, and the average expected return of each agent for atomic task m is also counted. The number of agents submitting bids or commitments is used as the base value of popularity, and the average expected return is used as the weighted value of popularity. The base value and the weighted value of popularity are linearly fused to obtain the estimated value of competition popularity for atomic task m. The estimated value of competition popularity is updated synchronously with the update of agents' bids and commitment states, so as to truly reflect the cluster competition degree of atomic task m. After the calculation is completed, the agent sends a proposal message containing its own expected return to the neighboring agents within its communication range. The proposal message usually includes the message sender identifier, the target atomic task identifier, the expected return value, and the possible commitment state. After receiving a proposal message from a neighboring agent, each agent adjusts its expected payoff based on a pre-defined negotiation strategy. This strategy can be either a market auction-based bargaining strategy or a game theory-based strategy update rule. Specifically, the market auction-based bargaining strategy works as follows: Based on its initial expected payoff, the agent receives bids from neighboring agents and compares the difference between its own and its neighbor's expected payoffs. If its payoff is lower, it gradually increases its bid; if its payoff is higher, it maintains its current bid. The bid is adjusted only once per iteration until it stabilizes. The game theory-based strategy update rule works as follows: The agent treats the neighboring agent's proposal as a game opponent's strategy, combines its own capabilities and task preferences, selects the strategy that maximizes its task execution payoff, and simultaneously updates its expected payoff and bid. There is no fixed number of iterations; the adjustment is based solely on the optimal payoff.

[0034] After a preset number of iterations, the negotiation process enters the termination phase. The negotiation process terminates when the system detects that all agents' bids or commitments for the atomic task no longer change, or when the change in bids between two consecutive rounds is less than a set threshold. After termination, the final commitments of all agents are summarized to form a preliminary task allocation map. This preliminary task allocation map is a list recording which agents have committed to executing each atomic task. In practice, conflict detection and resolution are performed on the preliminary task allocation map. Conflict detection identifies atomic tasks with multiple agents committing to execute them by scanning the map and marks these tasks as conflicting tasks. For each conflicting task, the final bid value of all agents committed to executing the conflicting task in the last round of negotiation is obtained. Conflicts are resolved by comparing the final bid values, assigning the conflicting task to the agent with the highest final bid value, and sending task preemption notifications to other agents committed to the conflicting task. Upon receiving a task preemption notification, an agent removes the conflicting task from its own committed task list. After removal, the system checks if the agent's current committed task load is below a preset minimum threshold. If the task load is below the minimum threshold, the agent must re-enter a local negotiation process for the remaining unassigned atomic tasks. This local negotiation process is a simplified version of the main negotiation process and only occurs between agents with low task loads and their neighbors. Conflict detection is repeated until there are no conflicting tasks in the initial task allocation map and all agents' task loads are within acceptable limits. At this point, conflict resolution is complete, and a definitive task allocation scheme is formed.

[0035] It is understandable that after a defined task allocation scheme is formed, a specific action coordination plan needs to be planned in conjunction with the dynamic environment state vector. Environmental prediction information for a future period is extracted from the dynamic environment state vector. This prediction information includes the predicted movement trajectories of obstacles and the probability distribution of dynamic events. Based on the defined task allocation scheme, the executing agent for each atomic task is clarified, as well as the logical sequence between atomic tasks determined based on the task element decomposition diagram. In specific implementation, minimizing the overall task completion time and avoiding spatiotemporal conflicts between agents are the dual optimization objectives. A start and end time range is allocated to each atomic task in the time dimension. The start and end time range of the atomic task is the execution time window. The planning of the execution time window needs to consider the logical order of tasks, the agent's movement speed, and the environmental prediction information. Within the execution time window of each atomic task, a collision-free path is planned for the executing agent, from the agent's starting position to the task target position, and then to the next task point or waiting area, based on the environmental prediction information. This collision-free path is the spatial path, which consists of a series of ordered spatial coordinate points. In some embodiments, the execution time windows and spatial paths of all atomic tasks are integrated. This integration process requires checking and resolving potential conflicts where different agents occupy the same spatial location at the same time. Conflict checking is achieved by establishing a four-dimensional spatiotemporal grid model, which includes three-dimensional spatial coordinates and one-dimensional time slices. Optionally, the spatial path of each agent is mapped to the four-dimensional spatiotemporal grid model according to its execution time window. The spatial grid occupied by the agent in each time slice is marked. The four-dimensional spatiotemporal grid model is traversed to find spatial grids marked and occupied by multiple agents in the same time slice. These spatial grids are recorded as spatiotemporal conflict points. For each spatiotemporal conflict point, the spatial path or execution time window of the relevant agent is adjusted according to preset conflict resolution rules. These rules include spatial avoidance principles and time staggering principles. Spatial avoidance principles refer to planning a detour path for one of the conflicting parties, while time staggering principles refer to delaying or advancing the time when a certain agent passes through the conflict point. After adjusting all spatiotemporal conflict points, the adjusted execution time windows and spatial paths of all agents are remapped into the four-dimensional spatiotemporal grid model for a new round of conflict checks until no spatiotemporal conflict points remain in the model. Finally, conflict resolution is achieved, and a globally coordinated action plan is generated. See Table 2 for a comparison of conflict detection and output value.

[0036] Table 2: Comparison of Value of Conflict Missions

[0037] In one embodiment of the present invention, the system analyzes the task allocation scheme portion of the cluster decision-making scheme, extracts all atomic tasks assigned to the current agent, and clarifies the execution order logic between these tasks. Simultaneously, it analyzes the action coordination plan portion of the cluster decision-making scheme, extracting detailed execution parameters for each assigned atomic task related to the current agent, including the precise execution time window of the task and the key point coordinate sequence of the planned spatial path. Subsequently, the abstract description of each atomic task, its corresponding specific execution time window parameters, and spatial path parameters are translated and encoded into a series of basic action instructions with clear semantics and a strict order, according to a preset instruction encoding protocol within the agent, such as "move to coordinates (X,Y)," "execute operation A at time T," and "move along path P at speed V." Finally, this series of action instructions for different atomic tasks is logically packaged according to the task execution order, and corresponding timestamps are added to key instructions, forming a unique individual action instruction set for the current agent, containing a complete action sequence and precise timing. This instruction set can be directly issued to the agent's controller to drive its execution.

[0038] In practical implementation, the cluster decision-making scheme is converted into a set of individual action instructions that can be parsed by each agent in the agent cluster. The cluster decision-making scheme includes a task allocation scheme and an action coordination plan. The conversion process is performed by the central scheduling module or the agent's own instruction parsing module. The task allocation scheme in the cluster decision-making scheme is parsed to extract all atomic tasks assigned to the current agent and their execution order. The task allocation scheme specifies the responsible agent for each atomic task in the form of a list or mapping. All atomic task entries for the current agent are selected based on the responsible agent, and the logical dependencies between these atomic tasks are obtained from the task element decomposition graph to determine the execution order. The action coordination plan in the cluster decision-making scheme is parsed to extract detailed parameters of the execution time window and spatial path for each atomic task corresponding to the current agent. The action coordination plan plans a specific execution time window and spatial path for each atomic task. By matching the atomic task identifier with the agent identifier, all execution time windows and spatial path data related to the current agent are retrieved. In some embodiments, the detailed parameters of the execution time window include the planned start time stamp and planned end time stamp of the task execution, and the detailed parameters of the spatial path include a series of ordered path point coordinate sequences and the preset movement speed of the agent between adjacent path points.

[0039] In practical implementation, the description of each atomic task, its corresponding detailed execution time window parameters, and detailed spatial path parameters are encoded according to the agent's internal instruction protocol. This internal instruction protocol is a predefined set of low-level commands that the agent controller can directly recognize and execute. The encoding process transforms the high-level task description and parameters into a series of basic action instructions with clear semantics and a strict order. The semantics of these basic action instructions include movement, waiting, performing operations, and communication. The order of these basic action instructions is determined by the execution logic and spatiotemporal constraints of the atomic tasks. The encoding process can be understood as a function mapping, which transforms the task parameter tuple into a sequence of action instructions. The relationship can be expressed as:

[0040] Where: characters Represents intelligent agents To perform atomic tasks The required sequence of action instructions; characters Representing atomic tasks Description information; characters Representing atomic tasks Detailed parameters for the execution time window; character Representing atomic tasks Detailed parameters of the spatial path; characters This represents the encoding function that conforms to the agent's internal instruction protocol. Optional, the encoding function... The specific implementation depends on the hardware platform and control system of the intelligent agent, and the encoding function. It may include subprocesses such as path point interpolation, velocity profile generation, and instruction formatting.

[0041] A series of action instructions generated for different atomic tasks are packaged and timestamped according to the task execution order. The task execution order is determined by the logical order of the atomic tasks when parsing the task allocation scheme. The packaging process concatenates the action instruction sequences corresponding to all atomic tasks into a complete action sequence specific to the current agent. In specific implementation, the timestamp is added to the key instructions in the sequence with absolute or relative time stamps. The absolute time stamp is generated based on the detailed parameters of the execution time window, while the relative time stamp is generated based on the position of the instruction in the sequence and the preset time interval. Optionally, the packaged and timestamped data constitutes a structured instruction package, which is the individual action instruction set. The individual action instruction set contains a complete action sequence and a precise time arrangement. It can be understood that the data format of the individual action instruction set needs to be compatible with the agent's communication interface and instruction parser. The finally generated individual action instruction set is sent to the corresponding agent through the communication link to drive the agent to execute.

[0042] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. An autonomous scheduling and collaborative decision-making method based on a cluster of intelligent agents, characterized in that, The method includes: Acquire task announcements and environmental situation maps of the intelligent agent cluster to form an initial situational awareness set; The initial situational awareness set is structured and analyzed to generate a clustered awareness model containing a task element decomposition diagram and a dynamic environment state vector. Based on the aforementioned cluster perception model, an improved consensus initiative algorithm is used to calculate and generate an individual preference vector for each agent, which includes task inclination and ability matching degree. Aggregate the individual preference vectors to construct a cluster-level preference distribution map; Based on the cluster-level preference distribution map, a multi-round iterative distributed negotiation process is initiated to generate a cluster decision scheme that includes task allocation schemes and action coordination plans. The cluster decision scheme is converted into a set of individual action instructions that can be parsed by each agent in the agent cluster, and the agent cluster is driven to execute it. The working principle of the improved consensus initiative algorithm includes: During the algorithm initialization phase, dynamic attention weights are injected into each agent. These dynamic attention weights are adaptively adjusted based on the preference cues received by the current agent from neighboring agents. The update method is that the agent counts the number of preference cues and effective interactions from neighboring agents for each atomic task, and updates the weights in real time with the accumulated interaction value. In the individual preference calculation stage, the algorithm not only considers the static matching between the atomic task and the agent's own capabilities and motivations, but also introduces a consensus prediction subprocess. The consensus prediction subprocess requires that when an agent calculates its own preference for a certain atomic task, it must estimate the collective preference tendency of its neighboring agents for the atomic task, and feed back the estimated collective preference tendency in a weighted form into its own ability matching degree and task preference degree calculation. After the adjustment of the dynamic attention weights and the feedback correction of the consensus prediction subprocess, the individual preference vector output by each agent has implicitly predicted the potential consensus of the cluster and actively moved towards it.

2. The method according to claim 1, characterized in that, The initial situational awareness set is subjected to structured analysis to generate a clustered awareness model containing a task element decomposition diagram and dynamic environment state vectors, including: Extract the task objective description, task constraints, and task completion indicators from the task announcement, and decompose the task objective description into multiple atomic tasks that can be executed independently or sequentially to form the task element decomposition diagram; Extract agent positions, obstacle outlines, and dynamic event markers from the environmental situation map, and quantify the agent positions, obstacle outlines, and dynamic event markers into time-stamped numerical state variables; The numerical state variables are aligned to a time series and spatially gridded to generate a dynamic environmental state vector that reflects the spatiotemporal changes of environmental elements. The task element decomposition diagram and the dynamic environment state vector are associated and combined to encapsulate a unified cluster perception model.

3. The method of claim 1, wherein, Based on the aforementioned cluster awareness model, and utilizing an improved consensus initiative algorithm, an individual preference vector containing task inclination and capability matching degree is calculated and generated for each agent, including: From the cluster perception model, extract the local task view and local environment view related to the current agent; Guided by the improved consensus initiative algorithm, the atomic tasks in the local task view are compared with the internal capability model of the current agent to calculate the capability matching degree. Meanwhile, guided by the improved consensus initiative algorithm, the conditions in the local environment view that are favorable or unfavorable to completing the atomic task are coupled with the current agent's internal motivation model to calculate the task inclination. The capability matching degree and task preference degree of each atomic task corresponding to the current agent are combined into a numerical pair, and the numerical pairs of all atomic tasks constitute the individual preference vector.

4. The method of claim 1, wherein, Aggregating the individual preference vectors to construct a cluster-level preference distribution map includes: Collect the individual preference vectors of all agents in the agent cluster to form a high-dimensional cluster preference matrix; Perform dimensionality reduction visualization calculation on the cluster preference matrix to map the preference values ​​of each agent for each atomic task into a two-dimensional preference space, where one dimension represents the ability matching degree and the other dimension represents the task preference degree. In the preference space, connection edges between agent nodes are constructed based on the physical proximity of agents or the logical association of tasks; The mapped agent nodes and their connecting edges, along with the preference values ​​carried on the nodes, are rendered together to form a visual graph, namely the cluster-level preference distribution map.

5. The method for autonomous scheduling and collaborative decision-making based on an intelligent agent cluster according to claim 1, characterized in that, Based on the cluster-level preference distribution map, a multi-round iterative distributed negotiation process is initiated to generate a cluster decision-making scheme that includes task allocation schemes and action coordination plans, including: The cluster-level preference distribution map is broadcast to every agent in the agent cluster as a common basis for negotiation; Each agent calculates its expected reward for each atomic task based on the cluster-level preference distribution graph and sends a proposal message containing its expected reward to its neighboring agents. Each agent receives proposal messages from neighboring agents, adjusts its expected revenue based on a pre-defined negotiation strategy, and generates a bid or commitment for the atomic task. After a preset number of iterations, when all agents' bids or commitments for the atomic task no longer change, or the changes are below a threshold, the negotiation process terminates, and the final commitments of all agents are aggregated to form a preliminary task allocation map. The initial task allocation mapping is subjected to conflict detection and resolution to ensure that each atomic task is committed to execution by one and only one agent, and that the task committed by each agent does not exceed its capacity, thus forming a determined task allocation scheme. Based on the task allocation scheme and the dynamic environment state vector, an execution time window and spatial path are planned for each allocated atomic task, forming an overall execution sequence and spatial coordination relationship among all atomic tasks, i.e., the action coordination plan.

6. The method for autonomous scheduling and collaborative decision-making based on an intelligent agent cluster according to claim 5, characterized in that, The preliminary task allocation mapping is subjected to conflict detection and resolution, including: Scan the preliminary task allocation map to identify atomic tasks that have multiple agents committed to execute, and mark them as conflicting tasks; For each conflicting task, obtain the final contribution value of all agents who committed to performing the conflicting task during the negotiation process; Compare the final output value, assign the conflicting task to the agent with the highest final output value, and send task deprivation notices to other agents who have committed to the conflicting task; Upon receiving a task preemption notification, the agent removes the conflicting task from its list of committed tasks and checks whether its current task load is below a minimum threshold. If it is below the minimum threshold, the agent needs to re-enter the local negotiation process for the remaining unassigned atomic tasks. Conflict detection is repeated until there are no conflicting tasks in the initial task allocation map and the task load of all agents is within an acceptable range, thus completing conflict resolution.

7. The method of claim 6, wherein, Based on the task allocation scheme and the dynamic environment state vector, an execution time window and spatial path are planned for each allocated atomic task, including: Environmental prediction information for a future period is extracted from the dynamic environment state vector, and the environmental prediction information includes obstacle movement trajectory and dynamic event occurrence probability. Based on the task allocation scheme, the executing agent for each atomic task and the logical sequence relationship between the atomic tasks are determined. With the dual objectives of minimizing the overall task completion time and avoiding spatiotemporal conflicts between agents, a start and end time range, i.e., an execution time window, is allocated to each atomic task in the time dimension. Within the execution time window of each atomic task, combined with the environmental prediction information, a collision-free path, i.e. a spatial path, is planned for its executing agent from the starting position to the task target position, and then to the next task point or waiting area. By integrating the execution time windows and spatial paths of all atomic tasks, potential conflicts arising from different agents occupying the same spatial location at the same time are examined and resolved, forming a globally coordinated action coordination plan.

8. The method according to claim 7, characterized in that, The process integrates the execution time windows and spatial paths of all atomic tasks, checks and resolves potential conflicts where different agents occupy the same spatial location at the same time, including: A four-dimensional spatiotemporal grid model is established, which includes a three-dimensional spatial coordinate dimension and a one-dimensional time slice dimension; The spatial path of each agent is mapped to the four-dimensional spatiotemporal grid model according to its execution time window, and the spatial grid occupied by the agent in a specific time slice is marked. Traverse the four-dimensional spatiotemporal grid model to find spatial grids that are occupied by multiple agents in the same time slice and record them as spatiotemporal conflict points. For each spatiotemporal conflict point, the spatial path or execution time window of the relevant intelligent agent is adjusted according to the preset conflict resolution rules, which include the spatial avoidance principle and the time staggering principle. After adjusting all spatiotemporal conflict points, the adjusted information of all agents is remapped into the four-dimensional spatiotemporal grid model for a new round of conflict checks until there are no more spatiotemporal conflict points in the model, thus completing the conflict resolution.

9. The method of claim 1, wherein, The cluster decision-making scheme is converted into a set of individual action instructions that can be parsed by each agent in the agent cluster, including: Analyze the task allocation scheme in the cluster decision scheme, and extract all atomic tasks assigned to the current agent and their execution order; The action coordination plan in the cluster decision-making scheme is analyzed, and the detailed parameters of the execution time window and spatial path of each atomic task corresponding to the current agent are extracted. The description of each atomic task, its corresponding execution time window, and spatial path parameters are encoded into a series of action instructions with clear semantics and sequence according to the agent's internal instruction protocol. This series of action instructions is packaged and timestamped according to the task execution order to form the individual action instruction set that belongs to the current intelligent agent and contains a complete action sequence and time arrangement.