Priority scheduling method and related devices based on large-scale multi-agent concurrency
By employing a priority scheduling method for concurrent multi-agent operations in a large model, the problem of user operation continuity and efficiency caused by the serial interaction mode of agents is solved, achieving concurrent execution of tasks and efficient utilization of resources, thereby improving the user operation experience.
Patent Information
- Application Number
- CN202411559573.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-04
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-11-04
AI Technical Summary
In existing technologies, the interaction mode between intelligent agents and users is serial, which limits the continuity and efficiency of user operations. It is impossible to design priorities for different tasks according to the user's needs and urgency, and to automatically identify task priorities. Important tasks cannot be processed immediately.
We adopt a priority scheduling method based on large model and multi-agent concurrency. Through the steward agent, we can personalize and optimize task instructions and use time-sharing technology to allow users to issue tasks multiple times in a row. Dedicated task agents can execute multiple tasks concurrently. We use knowledge graph and hierarchical memory module for task processing and priority scheduling.
It achieves continuity and efficiency improvement in user operations. The dedicated task agent can execute multiple tasks concurrently without waiting for the previous task to complete, thereby improving the system's throughput and resource utilization efficiency.
Smart Images

Figure CN119440842B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a priority scheduling method and related apparatus for multi-agent concurrency based on large models. Background Technology
[0002] With the deepening of the digital transformation of society, the scale of computing power has experienced explosive growth, greatly promoting the development of artificial intelligence technology. In particular, the rapid development of large-scale artificial intelligence models, represented by LLM (Large Language Model), has driven the deep integration of artificial intelligence technology and application systems, empowering application software and setting off a high tide of development and application of large-scale artificial intelligence models.
[0003] The technology of building application systems based on large-scale language models refers to using pre-trained large-scale language models as the core component of artificial intelligence systems, working with intelligent agents to reason about natural language and output reasoning conclusions. However, the current interaction mode between intelligent agents and users is mostly serial, meaning that after a user issues a task, they must wait for the intelligent agent to complete it before issuing a new task. This mode limits the continuity and efficiency of user operations. There is an urgent need for a technology that can prioritize different tasks according to the urgency of user needs, automatically identify task priorities, and ensure that important tasks (such as video conferencing) can be processed immediately. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a priority scheduling method and related apparatus for multi-agent concurrency based on a large model. By using time-sharing technology, it allows users to issue tasks multiple times consecutively, enabling the agent concurrency scheduling system to execute multiple tasks concurrently without waiting for the previous task to complete.
[0005] A priority scheduling method based on multi-agent concurrency for large-scale language models is proposed, applied to a concurrent agent scheduling system. This system includes a steward agent and k dedicated task agents. The large-scale language model is deployed with a pre-built concurrent agent scheduling system. The method includes:
[0006] The butler agent performs personalized optimization on the k first task instructions obtained, resulting in k second task instructions, which correspond one-to-one with the k first task instructions.
[0007] The housekeeper agent sends each of the k second task instructions to the corresponding dedicated task agent;
[0008] Each of the k dedicated task agents enhances the received second task instructions according to the corresponding hierarchical memory module to obtain k third task instructions.
[0009] Each of the k dedicated task agents performs intent recognition and process reasoning decomposition on the corresponding third task instruction to generate the workflow corresponding to the k dedicated task agents respectively.
[0010] The agent concurrent scheduling system executes the workflows corresponding to k dedicated task agents concurrently according to a priority scheduling algorithm.
[0011] Furthermore, priority scheduling algorithms also include:
[0012] The priority scheduling algorithm is characterized by the following formula:
[0013] π agent =α·π pre +β·(L default -L remaining )
[0014] Where, π agent π represents the real-time priority of the workflow corresponding to the intelligent agent executing the dedicated task. pre For predefined scenario priorities of dedicated task agents, L default L represents the default length of the workflow corresponding to the dedicated task agent. remaining This represents the remaining workflow length of the unexecuted dedicated task agent, α represents the predefined priority weight of the dedicated task agent, and β represents the influence weight of the remaining workflow length of the dedicated task agent on the real-time priority.
[0015] Furthermore, the priority weight of the predefined dedicated task agent is set to 0.7, and the weight of the impact of the remaining workflow length of the dedicated task agent on the real-time priority is set to 0.3.
[0016] Furthermore, the intelligent agent performs personalized optimization on the acquired k first task instructions to obtain k second task instructions, including:
[0017] The intelligent agent performs task scenario analysis on each of the k first task instructions it acquires, and determines the optimization information corresponding to each of the k first task instructions.
[0018] The butler agent optimizes each of the k first task instructions based on the optimization information corresponding to each of the k first task instructions, resulting in k optimized task instructions. Each of the k optimized task instructions corresponds one-to-one with the k first task instructions.
[0019] The steward agent calls the knowledge graph to supplement the background context of each of the k optimization task instructions, thereby obtaining k second task instructions.
[0020] Furthermore, the knowledge graph is constructed in the following manner:
[0021] The knowledge graph construction model was obtained by fine-tuning the large language model qwen2 using the LORA method.
[0022] The knowledge graph construction model is based on the GraphRAG method to construct the system text and obtain the knowledge graph.
[0023] Furthermore, the hierarchical memory module includes a core memory area, a primary memory area, and a fuzzy memory area.
[0024] Furthermore, the method also includes:
[0025] Each of the k dedicated task agents uses the ReAct method to decompose the corresponding third task instructions into process reasoning.
[0026] A priority scheduling device based on large-scale multi-agent concurrency includes an optimization unit, a sending unit, an enhancement unit, a generation unit, and an execution unit, wherein:
[0027] The optimization unit is used by the butler intelligence agent to perform personalized optimization on the k first task instructions obtained, so as to obtain k second task instructions, and the k second task instructions correspond one-to-one with the k first task instructions.
[0028] The sending unit is used by the housekeeper agent to send each of the k second task instructions to the corresponding dedicated task agent.
[0029] An enhancement unit is used for each of the k dedicated task agents to enhance the received second task instructions according to the corresponding hierarchical memory module to obtain k third task instructions.
[0030] The generation unit is used by each of the k dedicated task agents to perform intent recognition and process reasoning decomposition on the corresponding third task instruction, and generate the workflow corresponding to each of the k dedicated task agents.
[0031] An execution unit is used by the agent concurrent scheduling system to concurrently execute the workflows corresponding to k dedicated task agents according to a priority scheduling algorithm.
[0032] A terminal includes a processor, an input device, an output device, and a memory, the processor, the input device, the output device, and the memory being interconnected, wherein the memory is used to store a computer program, the computer program including program instructions, and the processor is configured to invoke the program instructions to execute the priority scheduling method for multi-agent concurrency based on a large model as described in any of the preceding claims.
[0033] A computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform a priority scheduling method for multi-agent concurrency based on a large model, as described above.
[0034] The beneficial effects of this invention are as follows: This invention, through time-sharing technology, allows users to issue tasks multiple times consecutively, enabling dedicated task agents to execute multiple tasks concurrently without waiting for the previous task to complete, thus improving the continuity and efficiency of user operations. Attached Figure Description
[0035] To more clearly illustrate the specific embodiments of the present invention, the accompanying drawings used in the description of the specific embodiments or prior art will be briefly introduced below. In all the drawings, the elements or parts are not necessarily drawn to scale.
[0036] Figure 1 A flowchart illustrating a priority scheduling method for multi-agent concurrency based on a large model, provided in an embodiment of the present invention;
[0037] Figure 2 This is a schematic diagram of the structure of a priority scheduling device based on large model multi-agent concurrency provided in an embodiment of the present invention. Detailed Implementation
[0038] The embodiments of the technical solution of the present invention will now be described in detail with reference to the accompanying drawings. These embodiments are merely illustrative of the technical solution of the present invention and are therefore intended to limit the scope of protection of the present invention.
[0039] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application should have the ordinary meaning as understood by those skilled in the art to which this invention pertains.
[0040] In one embodiment, such as Figure 1 As shown, a priority scheduling method based on multi-agent concurrency of large models is provided and applied to large language models. The agent concurrency scheduling system includes a steward agent and k dedicated task agents. The large language model is deployed with a pre-built agent concurrency scheduling system, which allows the user to input the first task command on the front-end interface. The steward agent can cooperate with the downstream dedicated task agents to complete the task issued by the user according to the input first task command.
[0041] Specifically, priority scheduling methods based on large-scale multi-agent concurrency include:
[0042] 1. The butler intelligent agent performs personalized optimization on the k first task instructions obtained, and obtains k second task instructions. The k second task instructions correspond one-to-one with the k first task instructions.
[0043] Preferably, the intelligent butler agent performs personalized optimization on each of the k acquired first task instructions, including:
[0044] (1) The housekeeper intelligent agent performs task scenario analysis on each of the k first task instructions obtained, and determines the optimization information corresponding to each of the k first task instructions.
[0045] (2) The housekeeper intelligent agent optimizes the k first task instructions according to the optimization information corresponding to the k first task instructions respectively, and obtains k optimized task instructions. The k optimized task instructions correspond one-to-one with the k first task instructions.
[0046] (3) The steward agent calls the knowledge graph to supplement the background context of each of the k optimization task instructions, and obtains k second task instructions.
[0047] The administrator agent receives k first-task instructions. Based on user requirements, it analyzes the task scenario of each instruction, first determining optimization information for each instruction. This optimization includes rewriting, error correction, and disambiguation. Error correction involves restoring the user's erroneously entered instructions to their objective form through large-scale model analysis. Disambiguation primarily eliminates referential ambiguity. When a dedicated task agent—an office role agent—is present, and the user needs to create work based on system text, the administrator agent needs to supplement the relevant context of the first-task instructions, i.e., fill in the task-related information in the local system document, analyze the task scenario, and determine the optimization information for each of the k instructions. Then, the administrator agent optimizes each instruction based on this optimization information, resulting in k optimized instructions. Afterward, the administrator agent uses a knowledge graph to supplement the background context of each optimized instruction, resulting in k second-task instructions. By performing personalized optimization on the k first-task instructions, the administrator agent can reduce communication costs and assist the downstream dedicated task agent in completing tasks more effectively.
[0048] Preferably, the knowledge graph is constructed in the following ways:
[0049] (1) The knowledge graph construction model is obtained by fine-tuning the large language model qwen2 using the LORA method;
[0050] (2) The knowledge graph construction model is based on the GraphRAG method to construct the system text and obtain the knowledge graph.
[0051] To reduce costs, this invention selects the fine-tuned large language model qwen2:7b as the base model for constructing the knowledge graph and querying. The qwen2 model is fine-tuned and tested using the LoRA fine-tuning script provided in the official project to obtain the knowledge graph construction model. The qwen2:7b model is loaded, and a new dataset is prepared, dividing it into training and validation sets. Next, fine-tuning parameters are configured, including determining the learning rate (e.g., from 1e-5 to 1e-3), batch size (depending on available GPU memory, e.g., 8 or 16), number of iterations, and evaluation frequency. Then, the knowledge graph construction model is loaded using the official fine-tuning script, and LoRA is applied, defining the optimizer and loss function. In each training epoch, batches in the dataset are iterated, performing forward propagation, calculating the loss, backpropagation, and updating the weights. The performance of the obtained knowledge graph construction model is periodically evaluated on the validation set, and hyperparameters such as the learning rate and batch size are adjusted based on feedback to optimize the knowledge graph construction model. The datasets focus on instruction datasets, including NLP tasks such as Entity Recognition (NER), Relation Extraction (RE), and Event Extraction (EE), as well as various Chinese and English datasets. Compared to the general GPT model, it has advantages in accuracy, completeness, and adherence to specified output formats, while significantly reducing costs.
[0052] The GraphRAG method leverages external structured knowledge graphs to enhance the contextual understanding of the large language model qwen2 and generate more insightful responses. Its goal is to retrieve the most relevant knowledge from databases, thereby improving the quality of answers for downstream tasks.
[0053] 2. The butler agent sends each of the k second task instructions to the corresponding dedicated task agent;
[0054] Each dedicated task agent has its own configuration file containing targeted system prompts. For example, a meeting agent might have a prompt like "You are a meeting assistant, tasked with booking or instantly starting meetings." Setting these system prompts fully leverages the performance potential of the large model in specific tasks. Additionally, a catalog of collected or self-made tools is added to the dedicated task agent, along with specialized tools, enabling it to better complete user-assigned tasks.
[0055] 3. Each of the k dedicated task agents enhances the received second task instructions according to the corresponding hierarchical memory module to obtain k third task instructions;
[0056] Preferably, the hierarchical memory module includes a core memory area, a main memory area, and a fuzzy memory area.
[0057] The hierarchical memory module stores and retrieves memories of varying access frequency and importance in a hierarchical manner. Drawing inspiration from the hierarchical caching design in operating systems, the text memory is managed through hierarchical partitioning, dividing it into a core memory area, a primary memory area, and a fuzzy memory area, with differentiated management strategies applied to each partition.
[0058] Core memory area: This area has the smallest capacity but the fastest retrieval speed. It mainly stores frequently accessed and important memory entries, usually involving user preferences and key information. The core memory area uses a priority replacement strategy; when the area reaches its capacity limit, the lowest priority memory entries will be replaced first.
[0059] Priority of memory entries π memory The calculation formula is as follows:
[0060] π memory =γ·f hit +δ·I importance
[0061] Where f hit Indicates access frequency, i.e., the number of times a memory entry is hit; I importance The importance of information is represented by a large model that rates the importance of memory items; γ represents the weight of access frequency; and δ represents the weight of information importance.
[0062] Primary Memory Area: This area has a large capacity but a relatively slow retrieval speed. It is mainly used to store dialogue records between the user and the dedicated task agent. When the primary memory area is saturated, the system will use an LRU (Least Recently Used) or LFU (Least Frequently Used) strategy for management. Memory entries that have not been used for a long time or have a low usage frequency will be moved to the fuzzy memory area, while memory entries that have been hit more than a threshold can be moved to the core memory area.
[0063] Fuzzy Memory Area: This area stores entries removed from the main memory area, which are then further compressed and summarized before being saved. Because the fuzzy memory area has the largest capacity but a low hit rate and low data frequency, a FIFO (First-In, First-Out) strategy is used for management, effectively reducing time complexity. This strategy achieves automatic forgetting of memories by discarding the oldest entries.
[0064] During interactions between the dedicated task agent and the user, when information outside of short-term memory (i.e., the context window) is involved, the system automatically triggers a memory retrieval mechanism. Following the principle of locality, the system first searches the core memory region, which quickly provides key or frequently used entries. If no match is found, it continues searching the main memory region to obtain historical interaction records between the user and the agent. If still no match is found, it enters the fuzzy memory region. By calculating the cosine similarity between the memory entry and the user's input, the system can assess the relevance of the memory entry to the current task, thus returning the most relevant memory information.
[0065] The hierarchical memory model based on operating system-level caching significantly improves the memory retrieval efficiency and adaptability of task-specific intelligent agents through partitioned storage and hierarchical management strategies. This hierarchical memory management approach not only improves resource utilization efficiency but also enhances the system's flexibility and scalability when handling complex tasks.
[0066] 4. Each of the k dedicated task agents performs intent recognition and process reasoning decomposition on the corresponding third task instruction, generating workflows corresponding to the k dedicated task agents respectively;
[0067] Preferably, each of the k dedicated task agents uses the ReAct method to decompose the corresponding third task instruction into a process reasoning decomposition.
[0068] ReAct uses a chain of thoughts to guide an LLM (Large Language Model) to break down complex problems, performing reasoning and action step by step. It also incorporates an observation phase; after each action, the current state is observed before proceeding with the next reasoning step. Developers guide the LLM step-by-step through the reasoning process and determine the appropriate action based on the results. A dedicated task agent autonomously plans and decomposes the task into a series of steps (workflows). Each step is completed using one or more tools, iterating at different stages, and progressively guiding the final result.
[0069] The steward agent manages numerous downstream dedicated task agents. When each dedicated task agent receives the corresponding second task instruction, it breaks down the task based on existing tool resources, guides the LLM (Large Language Model) to perform reasoning, and then determines which action to take, i.e. which tool to use, based on the reasoning result. Then, based on the ReAct paradigm workflow, it further reasons after observing and reflecting on the results, and automatically generates a complete workflow.
[0070] Preferably, each dedicated task agent will automatically generate a workflow based on the ReAct paradigm, according to the user's target task and its own tool resources. The workflow length is limited to no more than 10, and errors are allowed during workflow execution. If an error occurs, the subsequent workflow will be reorganized, but the number of errors will not exceed three.
[0071] Downstream dedicated task intelligence agents understand tasks, based on the ReAct paradigm, and automatically generate the steps required to complete the task according to existing tool capabilities. They execute the steps, determine whether the task is successful, and provide direct feedback to the user if successful. If unsuccessful, they switch to a template workflow to execute manually predefined steps.
[0072] 5. The agent concurrent scheduling system executes the workflows corresponding to k dedicated task agents concurrently according to the priority scheduling algorithm.
[0073] Preferably, the priority scheduling algorithm further includes:
[0074] The priority scheduling algorithm is characterized by the following formula:
[0075] π agent =α·π pre +β·(L default -L remaining )
[0076] Where, π agent π represents the real-time priority of the workflow corresponding to the intelligent agent executing the dedicated task. pre For predefined scenario priorities of dedicated task agents, L default L represents the default length of the workflow corresponding to the dedicated task agent. remaining This represents the remaining workflow length of the unexecuted dedicated task agent, α represents the predefined priority weight of the dedicated task agent, and β represents the influence weight of the remaining workflow length of the dedicated task agent on the real-time priority.
[0077] Preferably, the priority weight of the predefined dedicated task agent is 0.7, and the weight of the influence of the remaining workflow length of the dedicated task agent on the real-time priority is 0.3.
[0078] In typical multi-agent systems, dedicated task agents usually execute linearly in chronological order. This can lead to a dedicated task agent continuously occupying resources before completing its task, thus hindering the issuance and execution of other tasks, and preventing users from submitting new tasks during the execution process. To address this issue, this invention designs a time-sharing, shared, multi-agent concurrent priority scheduling method, allowing users to continuously issue tasks without waiting. The core of this mechanism lies in subdividing the continuous resource usage cycle into multiple discrete workflow time slices and allocating dedicated resource access periods to different dedicated task agents. When a user submits a natural language instruction (the first task instruction), the steward agent receives and analyzes the instruction, routing it to the corresponding dedicated task agent. The large language model automatically generates a workflow for the dedicated task agent based on the task requirements. The workflow generated by the large model varies depending on the dedicated task agent and the task. This invention treats a single workflow step as the smallest unit of scheduling in the agent concurrent scheduling system and employs First-In-First-Out (FIFO) and priority scheduling algorithms to switch workflows among dedicated task agents.
[0079] First-In-First-Out (FIFO) scheduling algorithm: Based on the classic scheduling method in operating systems, the system allocates resources according to the arrival order of the workflow steps of the dedicated task agent.
[0080] Priority scheduling algorithm: The priority of dedicated task agents is dynamically adjusted based on the weighted average of task urgency and the length of generated unexecuted workflows. Task urgency is based on predefined scenario priorities for dedicated task agents; a smaller priority value indicates a higher priority. The real-time priority π of the workflow corresponding to the dedicated task agent is used for execution. agent The calculation formula is as follows:
[0081] π agent =α·π pre +β·(10-L remaining )
[0082] The default workflow length for the dedicated task agent here is 10, π. pre Assign scenario priorities to predefined dedicated task agents (e.g., 0 for "meeting agent", 1 for "application installation agent", and 2 for "document writing agent").
[0083] Through this scheduling mechanism, the execution of each dedicated task agent changes from linear sequence to interleaved discrete execution, which significantly improves the throughput of the agent concurrent scheduling system, avoids a single dedicated task agent occupying resources for a long time, and thus reduces the average waiting time and turnaround time of dedicated task agents, enabling users to obtain feedback information from dedicated task agents earlier.
[0084] In one embodiment, 1, the user issues a first task instruction: "Write a PPT summarizing my annual work year based on the system document." The administrator agent receives the first task instruction and analyzes whether it needs correction or ambiguity removal (optimization information corresponding to the first task instruction). If the first task instruction is correct and does not require rewriting the system instruction `systemprompt`, it is determined to be an optimized task instruction. Based on this optimized task instruction, the administrator agent analyzes and determines that a knowledge graph is needed. Among numerous system documents, it uses graphrag (graph retrieval enhancement) to find a series of relevant work documents for the user through a global search, generating PPT materials related to the annual summary as the background context for the first task instruction, thus obtaining a second task instruction (first task instruction + background context of the first task instruction). The administrator agent routes the second task instruction to the downstream dedicated task agent – the PPT agent. After receiving the instruction, the PPT agent starts the hierarchical memory module to search for the user's preferences regarding PPTs. The PPT agent then analyzes the theme and combines it with the user's style preferences, automatically generating a workflow using a PPT generation tool, and then executing the workflow to generate the PPT file.
[0085] In one embodiment, 2, the user issues a first task instruction: install a document processing software. The administrator agent receives the first task instruction, analyzes that the user has not explicitly specified the software, and the instruction is ambiguous, requiring disambiguation (optimization information corresponding to the first task instruction) to obtain an optimized task instruction. Then, it calls graphrag (graph retrieval enhancement) to retrieve the user's software usage preferences, finding that the user prefers WPS Office for document processing. The administrator agent modifies the user's first task instruction, adding: "Install a document processing software, WPS Office recommended." This yields a second task instruction. The new second task instruction is then routed to the downstream dedicated task agent—the software installation agent—which generates its own workflow and executes it.
[0086] In one embodiment, 3, the user issues a first task instruction: schedule a meeting for the start of the new semester at 9:00 AM tomorrow (Schedule a Meeting). Upon receiving this instruction, the administrator agent immediately routes it to the dedicated task agent – the meeting agent. The user then immediately issues a second first task instruction: immediately initiate a meeting about China Tourism Day (Immediate Meeting). The administrator agent also immediately routes this to the dedicated task agent – the meeting agent. At this point, two different meeting agents are started on different threads, each completing one of the two different tasks. Each agent generates its own workflow and executes them concurrently. Ultimately, the scheduled meeting and the immediate meeting tasks are completed essentially simultaneously.
[0087] In one embodiment, 4, the user successively issues three first-task instructions: immediately initiate a parent-teacher meeting; install QQ; write a document about the Mid-Autumn Festival. Upon receiving these instructions, the system administrator concurrently launches three dedicated task agents. In daily use, video conferencing typically requires immediate processing, the urgency of software installation depends on the software's purpose and the user's needs, while document writing is a creative and time-intensive task. The three dedicated task agents automatically generate workflows based on their available resources and execute them concurrently. The continuous resource usage cycle is subdivided into multiple discrete workflow time slices, and a priority scheduling algorithm allocates dedicated resource access time slots to different dedicated task agents. A single workflow step is considered the smallest unit of scheduling in the agent concurrent scheduling system, and a priority scheduling algorithm switches workflows between dedicated task agents. The system ultimately achieves the effect of parallel processing of three tasks.
[0088] In one embodiment, a priority scheduling device based on large-scale multi-agent concurrency is also provided, including an optimization unit, a sending unit, an enhancement unit, a generation unit, and an execution unit, wherein:
[0089] The optimization unit is used by the butler intelligence agent to perform personalized optimization on the k first task instructions obtained, so as to obtain k second task instructions, and the k second task instructions correspond one-to-one with the k first task instructions.
[0090] The sending unit is used by the housekeeper agent to send each of the k second task instructions to the corresponding dedicated task agent.
[0091] An enhancement unit is used for each of the k dedicated task agents to enhance the received second task instructions according to the corresponding hierarchical memory module to obtain k third task instructions.
[0092] The generation unit is used by each of the k dedicated task agents to perform intent recognition and process reasoning decomposition on the corresponding third task instruction, and generate the workflow corresponding to each of the k dedicated task agents.
[0093] The execution unit is used by the agent concurrent scheduling system to concurrently execute the workflows corresponding to k dedicated task agents according to the priority scheduling algorithm.
[0094] A terminal includes a processor, an input device, an output device, and a memory, which are interconnected. The memory stores a computer program, which includes program instructions. The processor is configured to invoke the program instructions to execute some or all of the steps of any priority scheduling method based on large-scale multi-agent concurrency as described in the above method embodiments.
[0095] A computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform some or all of the steps of any priority scheduling method based on large model multi-agent concurrency as described in the above method embodiments.
[0096] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0097] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. A priority scheduling method based on large-scale multi-agent concurrency, characterized in that, An agent-based concurrent scheduling system is applied to large-scale language models. This system includes a steward agent and k dedicated task agents. The large-scale language model is deployed with a pre-built agent-based concurrent scheduling system. The method includes: The intelligent butler agent performs personalized optimization on the k first task instructions obtained to obtain k second task instructions, and the k second task instructions correspond one-to-one with the k first task instructions. The butler agent sends each of the k second task instructions to the corresponding dedicated task agent. Each of the k dedicated task agents performs enhancement processing on the received second task instruction according to the corresponding hierarchical memory module to obtain k third task instructions; Each of the k dedicated task agents performs intent recognition and process reasoning decomposition on the corresponding third task instruction to generate workflows corresponding to the k dedicated task agents respectively. The intelligent agent concurrent scheduling system executes the workflows corresponding to k dedicated task intelligent agents concurrently according to a priority scheduling algorithm; The priority scheduling algorithm also includes: The priority scheduling algorithm is characterized by the following formula: in, This indicates the real-time priority of the workflow corresponding to the intelligent agent executing the dedicated task. Scenario priorities for predefined dedicated task agents, This indicates the default length of the workflow corresponding to the dedicated task agent. This indicates the remaining workflow length of the dedicated task agent that has not yet been executed. This represents the priority weight of a predefined task-specific intelligent agent. This represents the weight of the impact of the remaining workflow length of the dedicated task agent on the real-time priority.
2. The priority scheduling method based on large-scale multi-agent concurrency as described in claim 1, characterized in that, The priority weight of the predefined dedicated task agent is 0.7, and the influence weight of the remaining workflow length of the dedicated task agent on the real-time priority is 0.
3.
3. The priority scheduling method based on large-scale multi-agent concurrency as described in any one of claims 1-2, characterized in that, The intelligent butler agent performs personalized optimization on the acquired k first task instructions to obtain k second task instructions, including: The butler intelligent agent performs task scenario analysis on each of the k first task instructions obtained, and determines the optimization information corresponding to each of the k first task instructions. The butler agent optimizes the k first task instructions according to the optimization information corresponding to the k first task instructions respectively, and obtains k optimized task instructions. The k optimized task instructions correspond one-to-one with the k first task instructions. The steward agent invokes the knowledge graph to supplement the background context of each of the k optimization task instructions, thereby obtaining k second task instructions.
4. The priority scheduling method based on large-scale multi-agent concurrency as described in claim 3, characterized in that, The knowledge graph is constructed in the following manner: The knowledge graph construction model was obtained by fine-tuning the large language model qwen2 using the LORA method. The knowledge graph construction model is based on the GraphRAG method to construct the system text and obtain the knowledge graph.
5. The priority scheduling method based on large-scale multi-agent concurrency as described in claim 4, characterized in that, The hierarchical memory module includes a core memory area, a main memory area, and a fuzzy memory area.
6. The priority scheduling method based on large-scale multi-agent concurrency as described in claim 5, characterized in that, Also includes: Each of the k dedicated task agents uses the ReAct method to decompose the corresponding third task instruction into a process reasoning decomposition.
7. A priority scheduling device based on large-scale multi-agent concurrency, characterized in that, It includes an optimization unit, a sending unit, an enhancement unit, a generation unit, and an execution unit, wherein: The optimization unit is used by the butler intelligence agent to perform personalized optimization on the k first task instructions obtained, so as to obtain k second task instructions, and the k second task instructions correspond one-to-one with the k first task instructions. The sending unit is used by the butler agent to send each of the k second task instructions to the corresponding dedicated task agent. The enhancement unit is used to enhance the received second task instruction according to the corresponding hierarchical memory module for each of the k dedicated task agents to obtain k third task instructions. The generation unit is used to perform intent recognition and process reasoning decomposition on the corresponding third task instruction by each of the k dedicated task agents, and generate the workflow corresponding to each of the k dedicated task agents respectively. The execution unit is used by the agent concurrent scheduling system to concurrently execute the workflows corresponding to k dedicated task agents according to the priority scheduling algorithm; The priority scheduling algorithm also includes: The priority scheduling algorithm is characterized by the following formula: in, This indicates the real-time priority of the workflow corresponding to the intelligent agent executing the dedicated task. Scenario priorities for predefined dedicated task agents, This indicates the default length of the workflow corresponding to the dedicated task agent. This indicates the remaining workflow length of the dedicated task agent that has not yet been executed. This represents the priority weight of a predefined task-specific intelligent agent. This represents the weight of the impact of the remaining workflow length of the dedicated task agent on the real-time priority.
8. A terminal, characterized in that, The system includes a processor, an input device, an output device, and a memory, which are interconnected. The memory is used to store a computer program, which includes program instructions. The processor is configured to invoke the program instructions to perform the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-6.